Skip to content

CI: kernel-matrix workflow from yeet new, one row per object and program - #2

Merged
julian-goldstein merged 8 commits into
mainfrom
kernel-matrix-ci
Oct 7, 2026
Merged

julian-goldstein merged 8 commits into
mainfrom
kernel-matrix-ci

Conversation

@julian-goldstein

@julian-goldstein julian-goldstein commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Ports the kernel-matrix CI that yeet new scaffolds, adapted to this repo's per-directory BPF objects, and fixes the two verifier problems its first runs found.

  • .github/workflows/kernel-matrix.yml: builds every bin/*.bpf.o on the runner, resolves the newest date-stamped quay.io/lvh-images/kind tag for 6.6, 6.12, 6.18, 7.2 and bpf-next, boots each under cilium's little-vm-helper, and runs the vendored static veristat in the VM. Per-kernel tables and the final grid are keyed by (object, program), since every TLS tap carries its own peer_sendmsg/peer_recvmsg. bpf-next runs with continue-on-error.
  • build/verify-kernel.sh: the in-VM gate. Defaults to all bin/*.bpf.o, reads veristat's verdict column because veristat exits 0 on a rejected program, and reloads rejected objects verbosely so the verifier's reason lands in the job log.
  • build/kernel-matrix.sh: the local lvh + static qemu harness, same multi-object treatment. make veristat-matrix now depends on bpf instead of the bin/probe.bpf.o this repo never links.

What the matrix found, and the fixes

Floor is 6.6, not 6.1. 6.1 refuses the wire tap's TCX attach type at load. Every other object loads there; the socket tap's iov_iter reads are CO-RE guarded, so its 6.4 note is for the build host only.

Wire tap on 6.6 (bpf/wire/wire.bpf.c): rejected with R4 invalid zero-sized read: u64=[0,4094] at bpf_skb_load_bytes. The size was guarded by a != 0 test, but verifiers before 6.9 do not narrow on that branch. The capture loop now submits the empty record first and rebuilds the bound as ((cap - 1) & CAP_MASK) + 1 behind a barrier.

Walk VM from 7.2.7 on (bpf/walk/walk.bpf.c): the lvh 7.2 image is 7.2.8 and, like bpf-next, ran the object to the 1M instruction limit, step alone at 850k, against 14k on 6.12 and on this host's 7.2.6. The 7.2.7 changelog carries a batch of verifier precision fixes for bpf_loop callbacks. A callback loop is verified once only when its entry state converges with an earlier one, and the state on frame 0's stack kept it from converging. The VM's mutable scalars now live in a one-element per-CPU array the verifier does not track; the callbacks reach them by lookup and recompute the ops and output pointers from (field, section, entry). The event pointer alone rides on the stack as the callback context. Event layout and the op set are unchanged.

Final grid

All 34 (object, program) rows load on 6.6, 6.12, 6.18, 7.2 and bpf-next. The walk VM verifies in 1.6k to 4.7k instructions depending on kernel.

Verified

  • CI: run 37549310163, all jobs green.
  • npm test: 105 pass.
  • End to end on this host (7.2.6, through yeetd): yeet run scripts/selftest-walk.js -- --ports 8093 --data against a local HTTP server captured 11 rows with every op kind exercised: scalar reads (cwnd, srtt, mss), scratch-register arithmetic (inflight), a bitfield (ca), two-level pointer chase into a string (iface = "lo"), proto = "TCP". req is null as before on lo, where the payload sits in frags and the app fills it from the wire-tap join. Counters: seen 686, emitted 11, ring_full 0.

…d program

Port the template's kernel-matrix CI: build every bin/*.bpf.o on the
runner, boot 6.1, 6.6, 6.12 and bpf-next under cilium's little-vm-helper,
and run the vendored static veristat in each VM. The gate reads veristat's
verdict column, since veristat exits 0 on a rejected program.

This repo links one object per bpf/<name>/ directory rather than a single
bin/probe.bpf.o, so the gate script and the local harness take the whole
bin/ set, and both summary grids are keyed by (object, program):
peer_sendmsg and peer_recvmsg recur in every TLS tap, so a bare program
name would collapse 34 rows into 26.

The veristat-matrix make target already existed but depended on
$(BPF_OUT), which this repo never builds; it now depends on `bpf`.

Expected on 6.1: the wire tap (TCX, 6.6+) and the socket tap
(iov_iter.__iov, 6.4+) are rejected. That marks the floor.
… loads on 6.1

First run of the matrix: 6.12 clean, wire rejected on 6.1 and 6.6 after
~1600 processed insns, walk rejected on bpf-next (7.3-rc4), and the socket
tap loaded everywhere, its iov_iter reads being CO-RE guarded. The gate
ran veristat at log level 0, so none of the reasons reached CI. It now
reloads just the rejected objects with -v -l1 and a rotated 64 KiB log,
tail-bounded, before failing.

Correct the README and workflow header: the 6.4 floor for the socket tap
is for the build host's vmlinux.h, not for the kernel it loads on.
…rts at 6.6

The first matrix run put the floor at 6.6, not 6.1: 6.1 refuses the wire
tap's TCX attach type at load, with zero instructions verified, and every
other object loads there. The socket tap's iov_iter reads are CO-RE
guarded, so its 6.4 note is for the build host only.

On 6.6 the wire tap was rejected inside emit(): "R4 invalid zero-sized
read: u64=[0,4094]" at bpf_skb_load_bytes. The size was guarded by
`cap && ...`, but a verifier before 6.9 does not narrow a register on the
`!= 0` branch, so it still saw a possible zero. Submit the empty record
first, then rebuild the bound by arithmetic the verifier does track:
((cap - 1) & CAP_MASK) + 1 is [1, CAP_MASK + 1] and an identity on the
range that reaches it; barrier_var keeps clang from folding it away.

Matrix: 6.6, 6.12, 6.18, 7.2 gate; bpf-next runs with continue-on-error.
On 7.3-rc4 the walk VM's bpf_loop callbacks cost 850k verifier
instructions against 14k on 6.12 and hit the 1M limit. That is worth
seeing on every run, but a moving target should not block a merge until
the change ships in a release. The local harness defaults to the same
spread.
…f_loop backports

The lvh 7.2 image is 7.2.8 and rejects walk.bpf.o the way bpf-next does:
"BPF program is too large. Processed 1000001 insn", with the step
callback alone at 850k instructions, against 14k for the whole object on
6.12 and 6.18. This host runs 7.2.6 and loads it. ChangeLog-7.2.7 carries
a batch of verifier precision fixes for bpf_loop callbacks, which is where
the difference lands. Say so in the workflow header and the README; the
7.2 row keeps gating, since that is a stable kernel.
…converge

From 7.2.7 the verifier simulates every one of the VM's nested bpf_loop
iterations in full, 8 x 16 x 8 of them at ~830 instructions each, and
hits the 1M limit. A callback loop converges only when its entry state
matches an earlier one, and a precise scalar in that state blocks it.
Two leaks: `step` stored the probe-read size into st->wrote, and a size
argument is always marked precise; and walk_node's `wrote == 0` test is
predictable on a range the verifier knows excludes zero, which marks the
operand precise back through the step loop's entry.

Keep the count as a separate opaque copy of arg, never the register that
went to the helper, and test it after a round trip through the event,
where a load from ring buffer memory is an unknown the verifier cannot
predict a branch on.
…to converge on

Keeping precise scalars out of the stack-held state was not enough: on
7.2.8 and bpf-next the nested bpf_loop callbacks still failed to
converge and the object stayed at the 1M instruction limit (the three
passing kernels went from 14k to 35k-107k). The verifier compares the
whole state at a callback's entry, frame-0 stack included, and the log
cannot show which slot differed from inside a callback frame.

So take the state out of its sight. The VM's mutable scalars now live in
a one-element per-CPU array, which the verifier does not track, and
each callback reaches them with a lookup; the ops pointer and the
output pointer are recomputed from (field, section, entry) rather than
stored. The one thing the callbacks need kept typed is the event they
write into, since a ring buffer pointer kept in map memory would come
back as a plain number, so that rides on the stack as the callback
context bpf_loop requires anyway. raw_tp/tcp_probe cannot re-enter on
a CPU, so one element per CPU is enough. Event layout and the op set
are unchanged.
Every kernel rejected the per-CPU-state object at insn 670: "R4 unbounded
memory access" on fcount[key]. clang spilled key as a 4-byte slot and the
verifier treats that fill as unknown. Mask where it is used, as before.
@julian-goldstein
julian-goldstein merged commit d68cf53 into main Oct 7, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant