ci: run the IPv6-only e2e job on both sandbox classes [DO NOT MERGE — signal only] - #1123
Draft
Yuan Gao (ygao-g) wants to merge 28 commits into
Draft
ci: run the IPv6-only e2e job on both sandbox classes [DO NOT MERGE — signal only]#1123Yuan Gao (ygao-g) wants to merge 28 commits into
Yuan Gao (ygao-g) wants to merge 28 commits into
Conversation
On a fresh IP_FAMILY=ipv6 cluster nothing resolves from inside a pod and no actor boots: CoreDNS inherits the node's IPv4 resolver, which a v6-only pod cannot reach, and "kind-registry" NXDOMAINs in atelet's own netns. Point the forward at an IPv6 upstream, overridable with IPV6_DNS_UPSTREAM, and give the registry its own server block, so it is asked for nothing but its own name. IPv4 and dual-stack clusters are unchanged, and atenet-egress still crashloops on v6-only for an unrelated Envoy bind bug. Asking once was not enough to prove that: about half of fresh clusters do not answer the first query, and a pod that goes unanswered stays unanswered, so the check re-asks with a new pod and prints what the pod saw when it gives up. It lives in hack/verify-ipv6-dns.sh rather than inline, because the registry block records an address the registry can move off and there was no way to re-check a cluster without rebuilding it.
The HTTP and HTTPS ingress listeners bound 0.0.0.0 only, so on a dual-stack cluster Envoy answered on the router Service's IPv4 ClusterIP and on nothing at all for IPv6. Each primary socket now carries an additional "::" address on the same port. Ipv4Compat stays false on the additional address: clearing IPV6_V6ONLY would collide with the primary already bound to that port. Leaving the primary alone is what keeps an IPv4-only cluster unchanged -- with the caveat that a node lacking AF_INET6 entirely could not bind "::" and the listener would not come up. First of three commits binding atenet's gateways dual-stack. (cherry picked from commit 501991d)
The Envoy admin socket bound 0.0.0.0, and the atenet-router Service carried no ipFamilyPolicy -- which the API server defaults to SingleStack, one IPv4 ClusterIP and nothing else. Between them the router had no IPv6 address to answer on. The socket now binds "::" with ipv4_compat, one socket for both families, and the Service asks for PreferDualStack. bootstrap.v3.Admin takes a single address and has no additional_addresses, so the ingress listeners' shape is not available here; ipv4_compat is what makes the one socket serve both families. It is load-bearing: dataplane.go health-checks the admin listener over http://127.0.0.1:9901/ready, so a bare "::" would report the dataplane component of /statusz unhealthy. Prefer, not Require, keeps the Service valid on a single-stack cluster; spec.ipFamilies is left alone because the primary family is immutable and the API server appends the secondary itself. (cherry picked from commit 2a21292)
The gateway's admin and :443 sockets bound 0.0.0.0, so on an IPv6-primary cluster the kubelet probed the pod on its only address and atenet-egress crashlooped -- Envoy started fine and logged "admin address: 0.0.0.0:15000" -- while an actor's CONNECT had no v6 path in. Both sockets now bind "::" with ipv4_compat, and the Service asks for PreferDualStack so a dual-stack cluster hands out an IPv6 ClusterIP to reach them on. One socket here rather than the ingress listeners' pair: IPv4 peers then arrive as ::ffff: addresses, and nothing on this path reads the peer -- actor identity comes from the client certificate and the access log records the cert SAN. ipv4_compat also has to stay on the admin socket, because the ext-proc sidecar's drainer dials 127.0.0.1:15000 and envoydrain.go reads a refusal there as "Envoy already exited", skipping the drain silently. Last of three. (cherry picked from commit 2549657)
Both gateway admin sockets bind "::" with ipv4_compat, and the flag is what keeps their in-pod callers working: dataplane.go health-checks the router's over IPv4 loopback, and envoydrain.go dials the egress one the same way and reads a refusal as "Envoy already exited", skipping the drain without reporting an error. No Go test, golden file, or verify script read either manifest, so dropping the flag would have failed silently. make verify now rejects an admin socket that binds "::" without it. (cherry picked from commit 4b478a0)
The CONNECT-terminating listeners landed after the first commit of this series, so they kept a bare 0.0.0.0 socket while ingress HTTP and HTTPS gained their "::" pair. Give them the same additional address, so all four of the router's socket listeners answer on both families. Both are port-gated and no e2e suite configures them yet, which is why nothing caught this; the internal main_internal listener has no socket and needs nothing. (cherry picked from commit 54c727d)
TCPOriginalDestination read only the IPv4 SOL_IP/SO_ORIGINAL_DST, so an actor's IPv6 connection redirected into the transparent egress listener had no destination to dial and the proxy failed it. Read IP6T_SO_ORIGINAL_DST too, falling back to it only when the IPv4 lookup returns ENOENT, so unrelated IPv4 failures keep their own error. One step towards dual-stack actor networking; the actor veth and its nftables rules are still IPv4-only. Co-authored-by: Yuan Gao <ypgao@google.com> (cherry picked from commit d8527b5)
Both ateom herders defaulted the actor ingress flags to "0.0.0.0:443" and "0.0.0.0:444", which reads as IPv4-only. It never was: Go treats an unspecified address as a wildcard and binds it dual-stack, so the sockets already served both families. Spell the defaults ":443" and ":444" so the flag says what it does, and note why in a comment. Part of the dual-stack actor networking series; no behavior change. (cherry picked from commit 51b2cbe)
EnableIPv4Forwarding now also writes /proc/sys/net/ipv6/conf/all/forwarding so actor IPv6 traffic (including DNS queries) is routed between the actor veth and pod eth0 instead of being dropped by ip6_forward() on dual-stack / IPv6-only clusters. Factor the sysctl write into writeSysctlIfUnset preserving the original read-only remount/restore behavior, and add unit coverage for its fast paths. Fixes: agent-substrate#945
IPv6 sysctls are absent on kernels with IPv6 disabled (e.g. some containers set net.ipv6.conf.* only when IPv6 is enabled). Treat a missing path as 'nothing to enable' instead of forcing a remount and failing, matching the documented behavior.
The helper has enabled both address families since IPv6 forwarding was added; the name now says so. The single call site in SetupActorNetwork is updated along with the doc comment.
The os.Stat/IsNotExist fallback was never executed by the unit tests: every existing subtest's temp path could be created, so each returned at the os.WriteFile fast path. Point the new subtest at a node under a directory that does not exist — what procfs always does in production — and assert the file stays absent.
The egress Envoy pinned dns_lookup_family to V4_ONLY, so it asked only for A records. On an IPv6-only cluster no upstream name resolves and no actor can reach the internet. AUTO tries AAAA and falls back to A, so IPv4-only clusters behave as before. One step of the IPv6 egress work, and not the one that unblocks it -- actor egress still stops earlier, in atunnel's original-destination lookup. (cherry picked from commit de81578)
Runs the full install plus the demo and networking e2e suites against a single-stack IPv6-only kind cluster, and asserts the cluster really is v6-only so a green run cannot quietly become a second IPv4 run. It stays out of the e2e-test merge gate, so it reports IPv6 status without being able to block a PR, and it runs on every PR for now so the results are visible; the TODO on the trigger records the intended ci/ipv6 label gate. ubuntu-latest has no IPv6 egress, so the job stands up tayga for NAT64 and points CoreDNS at an upstream resolver through the well-known prefix. DNS64 is scoped to a catch-all server block: synthesizing AAAA over the cluster zones destroys the v6-only ClusterIP answers and the control plane never comes up. (cherry picked from commit 748e841)
The probe attached to the pod to collect its markers, and an attach can end before the last write arrives. A CI run lost the registry marker that way, so the check reported a registry it could not reach -- and then refused to re-probe, because only the resolve leg was treated as a settling race. Wait for the pod to terminate and read its log instead, and close the probe with a PROBE_DONE marker so a short read is re-probed rather than read as a failed fetch. A registry that really is down still fails on the first attempt.
Before, the actor zone answered A queries and failed everything else -- AAAA for a valid actor, and any name in the zone that is not an actor. A failure reads as a temporary error rather than an answer, so clients retry it and then give up on the name; Alpine actors could not resolve each other at all, even on an IPv4-only cluster. After, those queries return a correct empty answer, and one that resolvers can cache. A unit test pins the whole rendered zone as a literal, so editing the name pattern or the suffix fails there rather than passing silently.
Splits a Service's cluster IPs into its IPv4 and IPv6 entries, returning "" for a family the Service has no address in. No behavior change on its own -- nothing calls it until the AAAA change later in this series. It is shared rather than package-local because a Service with no ipFamilyPolicy is SingleStack, so one empty family is the steady state on every cluster, not an error, and each caller would otherwise have to decide that for itself. Unit tests cover single- and dual-stack Services and the unallocated and malformed cases.
No behavior change -- buildTemplate() already ran once, from init(). The next commit renders the Corefile on every call instead, where a stamp taken inline would differ each time: reconcile compares the render against the file on disk, so it would rewrite and reload CoreDNS every tick.
Before, an actor name never resolved over IPv6: the zone published the router's primary cluster IP, always as an A record whatever family it was. On a dual-stack cluster the v6 address went unpublished; on an IPv6-only cluster the record was malformed, so every A query for an actor name failed. After, the zone publishes an address record per family the router has an address in, and answers empty for a family it has none in. Unit tests pin the rendered zone for each family combination, verified against the pinned coredns/coredns:1.11.1.
Nothing in the e2e harness could query the actor DNS zone. Suites reach actors by port-forwarding atenet-router and passing the actor name as a Host header, so the zone CoreDNS actually serves went unasserted, and a suite that wanted to check it had no way to distinguish an empty answer from a server failure. Adds a DNS client that port-forwards the atenet DNS Service and reports the rcode class alongside the addresses, plus a helper for the router's ClusterIP in each family. First of two commits; the tests that use these follow. clusterIPsByFamily here is a stopgap that agent-substrate#938 replaces with internal/ipfamily.
The zone answered A queries and failed everything else -- AAAA for a valid actor, and any name in the zone that is not an actor -- and no test caught it, because Go's resolver masks a SERVFAIL that musl treats as fatal. These assert the rcode class rather than the record: a non-A qtype and a name that misses the actor regex must come back NODATA or NXDOMAIN, and an A query must carry the router's ClusterIP. A separate test covers the AAAA record, skipped where the router has no v6 address. Second of two commits. The assertions are red until agent-substrate#874 and agent-substrate#938 land, so this stays a draft until then. Part of agent-substrate#246.
Nothing checked that the router's dataplane listeners bind more than an IPv4 socket, and nothing reached an actor over the router's IPv6 ClusterIP. Every other path a test has into the router -- a port-forward, the pods/proxy and services/proxy subresources -- is mediated by the API server, which picks the family, so no existing test could have caught a listener that lost its IPv6 socket. Reads the bound addresses from Envoy's own admin /listeners, and drives an in-cluster probe pod at the router over each ClusterIP in turn. Red until agent-substrate#911 binds those sockets, so this stays a draft until then. The per-family probe skips on a single-stack cluster. Part of agent-substrate#246.
The actor's NAT and filter rules lived in an ip table, which can only ever carry IPv4. They are now in an inet table, so one table can hold both address families when the actor veth becomes dual-stack. Every IPv4 match opens with an NFPROTO comparison, because a bare payload match in an inet table would read an IPv4 offset out of an IPv6 header. IPv4 behaviour is unchanged, but the move is not a no-op on a dual-stack pod: inet nat chains register the nat hooks for both families, so IPv6 traffic in the worker pod netns is now conntracked, and the forward accept now covers IPv6. NAT in the inet family needs Linux 5.2 or later. Teardown sweeps ip as well as inet. A table name is unique per family, so the ip table an earlier ateom left behind is invisible to an inet-only cleanup: the dump comes back empty, the "already clean" path reports success, and the stale table keeps redirecting alongside the new one. Part of agent-substrate#246
Actor networking was IPv4-only, so an actor on a dual-stack worker pod could not reach an IPv6-only destination at all. SetupActorNetwork now assigns the fd00:169:254::/126 counterparts of the existing point-to-point pair to both ends of the actor veth, installs an IPv6 default route in the interior netns, and adds the matching rules to the inet-family actor table. Whether the actor gets IPv6 is decided once in the worker pod netns and carried into the interior one, which is created fresh and so always reports IPv6 available whatever the cluster's families are. Both halves have to hold: the pod needs a global IPv6 address of its own, and the veth has to accept an IPv6 address -- IPv4-only GKE sets disable_ipv6 and netlink then rejects the assignment with EPERM. Addresses carry IFA_F_NODAD rather than the accept_dad sysctl, which the unprivileged ateom container cannot write. Part of agent-substrate#246
On an IPv6-only cluster the rustfs ClusterIP is a v6 literal, so the endpoint came out as http://fd00:10:96::abcd:9000 — not a valid URL — and micro-VM asset staging failed outright. Bracketing alone is not enough: the botocore vendored into aws-cli 2.17.0 has no IPv6 endpoint validator and rejects a bracketed literal too. Bracket the literal when it contains a colon, and bump this script's aws-cli pin. The rustfs-bucket-init Job addresses rustfs by DNS name and is left alone, so no address family pays for a fresh in-cluster image pull.
The IPv6-only job only ever ran gVisor, so nothing in CI booted a micro-VM guest on a cluster that has IPv6. That is the configuration where the two classes diverge: gVisor's guest adopts the interior netns and gets the actor's IPv6 for free, while a micro-VM's is cross-connected to a tap at L2 and has to be told over the kata-agent channel. The job now enables KVM, asserts the node really got /dev/kvm, deploys the micro-VM counter and egress demos, and replays both suites under E2E_SANDBOX_CLASS=microvm. The vacuous-green guard covers the new logs but deliberately does not gain the dual-stack job's per-test assertions: on a single-stack cluster TestActorIngressPerFamily skips itself by design, so naming it would fail every run. Part of agent-substrate#246.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Signal only, on top of #1084: the same IPv6-only job with the micro-VM sandbox
class added. Everything below the top two commits is #1084's. Until now no job
anywhere booted a micro-VM guest on a cluster that has IPv6, which is exactly
where the two classes diverge — gVisor's guest adopts the interior netns and
picks the actor's IPv6 up for free, while a micro-VM's is cross-connected to a
tap at L2 and has to be told over the kata-agent channel.
Verified locally on an IPv6-only kind cluster: the golden snapshot reaches Ready
and both suites pass on both classes. The vacuous-green guard covers the new logs
but deliberately gains no per-test assertions —
TestActorIngressPerFamilyskipsitself by design on single-stack, so naming it would fail every run.
hack/microvm-assets: stage assets to rustfs over IPv6ci: run the IPv6-only e2e job on both sandbox classes🤖 Generated with Claude Code