Skip to content

Headstart: start dependents on early metadata, in cargo check and cargo build - #1

Merged
zmaril merged 19 commits into
mainfrom
early-metadata
Sep 30, 2026
Merged

zmaril merged 19 commits into
mainfrom
early-metadata

Conversation

@zmaril

@zmaril zmaril commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Headstart starts dependent crates as soon as a dependency's item interfaces are checked, instead of after all of its function bodies are, for both cargo check and cargo build.

On rustc's default front end, it makes clean builds of 13 real projects up to 54% faster for cargo check and up to 42% for cargo build, and none slower. Failing builds print the same diagnostics as today.

What changes

  • rustc -Zearly-metadata (patch):

    • A new analysis_interfaces query splits analysis into item interfaces and function bodies.
    • Between the two, the driver writes libfoo.early-rmeta and announces it with an early-metadata artifact notification.
    • Crate loading accepts early metadata. Before code generation, each dependency loaded from it is swapped for its full metadata; before linking, rustc waits for the rlibs.
    • While waiting, rustc announces wait-metadata, polls, then announces resume. It never touches the jobserver.
    • The crate hash is upstream's metadata-bytes hash, computed over the early metadata plus a supplement that includes the HIR hash of every body. Full metadata carries the same hash.
    • If the crate fails, through linking, the early file is deleted, and anything waiting on it stops.
    • -Zearly-metadata-verify reports two things:
      • any left-out definition whose body was type-checked during the early write anyway;
      • any query that only full metadata can answer, asked of a crate still loaded from early metadata.
  • cargo CARGO_HEADSTART=1 (patch):

    • Passes -Zearly-metadata to every compile.
    • Starts every rustc unit on its rlib dependencies' early-metadata notification, in check and build. That covers libraries, binaries, tests, proc macros and build scripts.
    • Doesn't count paused compilations as running, so their job slots go to other work.
    • Reports a unit's output only once its dependencies have succeeded.

    With the variable unset, cargo behaves like upstream.

docs/design.md covers how it works, what early metadata leaves out, the risks, and prior art (#64112).

Correctness

  • Sweeps (scripts/sweep.sh):
    • What they build: all 53 rustc-perf compile benchmarks, headstart off and on, with -Zearly-metadata-verify.
    • Six sweeps: cargo check and cargo build, each in debug, release and debug with -Zthreads=8.
    • Result: every sweep is clean. 52 benchmarks pass in both modes with identical diagnostics and no verify reports. stm32f4 fails in both modes (it needs a device feature).
  • Real projects: 312 builds, all succeeded with identical diagnostics. The only difference is lemmy under -Zthreads, whose overflow warnings move between runs even without headstart.
  • scripts/check-errors.sh: an error in a dependency's body and an error in the binary, under check and build, in human and JSON formats. Diagnostics, exit status and the binary's output are identical.
  • scripts/check-incremental.sh check|build: ten edit steps. Each matches headstart-off, and the final state matches a clean build.
  • scripts/check-swap.sh: a library that starts on early metadata and swaps in the full metadata while paused, at opt-levels 0–3, s and z. It fails at 2 and 3 without the first fix below.
  • Test suites, run with the patches applied and headstart off:
    • rustc's UI suite: 22129 passed, 0 failed.
    • cargo's test suite: 4034 passed. The one failure only flags the missing cross-compile target.

Timing (Linux, 16 cores, median of 3 runs)

project cargo check cargo build on top of -Zthreads=8 (check / build)
rust-analyzer 90 → 42 s (54%) 138 → 80 s (42%) 25% / 22%
polars 172 → 87 s (49%) 341 → 241 s (29%) 13% / 9%
wasmtime 151 → 80 s (47%) 245 → 156 s (36%) 14% / 13%
bevy 121 → 85 s (29%) 196 → 137 s (30%) 4% / 7%
zed 280 → 204 s (27%) 425 → 318 s (25%) 6% / 9%
lemmy 449 → 331 s (26%) 532 → 372 s (30%) 13% / 14%
typst 86 → 63 s (26%) 142 → 123 s (13%) 22% / 2%
atuin 125 → 109 s (13%) 163 → 127 s (22%) −1% / 5%
vaultwarden 116 → 107 s (8%) 165 → 150 s (9%) 3% / 4%
zola 117 → 115 s (2%) 126 → 126 s (0%) −2% / 4%

The full tables are in docs/results.md. They also cover helix, lldap, nushell and 21 rustc-perf benchmarks, where headstart saves 5–46% and, on top of -Zthreads=8, 0–32%.

  • Deep workspaces gain most. They chain their own crates one after another.
  • Wide builds that end in one big crate gain least. zola ends with 81 s of a single crate compiling alone.
  • The parallel front end covers some of the same ground. On top of it, headstart adds up to 25%, and three results are within noise of even.

Bugs found by real projects, all fixed

Until this round, the sweeps built only in debug mode, so none of these paths had been exercised:

  1. Optimized dependencies (typst, zed): "missing optimized MIR". The early write optimized MIR, and the MIR inliner cached "no MIR" answers from dependencies still loaded from early metadata.
  2. Statics evaluated on one thread before the early write. Under -Zthreads, zola was 80% slower with headstart. regex-automata's long-standing small loss had the same cause.
  3. cargo check --release: full metadata computed the reachable set. It asked early dependencies for MIR, and the result was never read.

A review of this PR also found that rustc withdrew early metadata only until code generation started. Its guard now lasts through linking.

Open

The next PR in the stack has the full list of what's left, and how to move this work to Claude Code containers. The main items:

  • Scheduling when the machine is saturated.
  • Fewer cores.
  • Peak memory under cargo build.
  • Tests that could live upstream.
  • A rebase onto a newer rustc.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM

zmaril and others added 9 commits September 27, 2026 06:54
rustc -Zearly-metadata writes a crate's metadata once item interfaces are
checked, before function bodies are. cargo with CARGO_HEADSTART=1 pipelines
check builds on it, so dependents start while their dependencies are
still checking bodies. Includes setup and benchmark scripts, a smoke test,
design notes, and a correctness sweep over rustc-perf on macOS and Linux.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Cargo now holds a provisional unit's diagnostics, JSON messages and result
until every dependency finishes cleanly, and drops them if one fails, so a
failing build prints what it prints today (scripts/check-errors.sh checks
this). Check units no longer get the synthetic full-build edges meant for
linking, and early metadata no longer computes the reachable set, which
forced inline and generic bodies before the write. Adds the error-delay
measurement, prior art, and a Linux rerun.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
rustc: nested definitions are encoded when their body is an interface body
(consts, const fns, coroutines, statics, opaque-type-defining functions),
computed from the code, instead of by peeking at the query cache.
-Zearly-metadata-verify checks the set; the sweep reports no misses.
The early write now runs outside analysis's dependency tracking, depends
on the crate's item list so incremental metadata reuse notices new items,
and the flag is tracked.

cargo: a failing dependency's dropped dependents no longer trigger the
'waiting for other jobs' message.

Evidence: incremental edit sequences match headstart-off and a clean build;
clippy (5534 diagnostics) and rustdoc output match; UI suite passes;
peak memory, -Zthreads=8, and an incremental edit loop measured.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
rustc: a new analysis_interfaces query splits analysis; the driver writes
libfoo.early-rmeta between it and the bodies, outside any query, replacing
the session flag and global. Crate loading accepts early metadata and,
before code generation, swaps in full metadata (same crate hash), pausing
for it with the jobserver token handed back. Both files share an HIR-based
SVH. Early metadata leaves out exported symbols.

cargo: -Zearly-metadata on every compile; dependents start on the
early-metadata notification in check and build; failure markers stop
waiting dependents; emit kinds are parsed explicitly.

Verified: check and build sweeps, UI suite (22129 passed), cargo testsuite,
error and incremental scripts in both modes. Debug builds 6-32% faster,
checks 7-47% faster on 16 cores.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
@zmaril zmaril changed the title Headstart: start dependents on early metadata in cargo check Headstart: start dependents on early metadata, in cargo check and cargo build Sep 28, 2026
zmaril and others added 10 commits September 28, 2026 04:58
A paused rustc no longer hands its jobserver token back itself. Cargo stops
counting it as running (reusing its slot), and once the full metadata it
waits for is written, starts no new work until it has resumed with a
freed token, which cargo returns to the jobserver when it finishes. Before
this, paused compilations were starved by work started in their place:
148 s of cargo-0.87.1's build, and ~17 s on its check critical path.

scripts/log-rustc now logs artifact events per rustc run, and
scripts/critical-path.py reports the critical path, linking crates held
back, and pause time split into waiting for metadata vs. for a slot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
The SVH is upstream's metadata-bytes hash, computed over the early metadata
(plus the same supplement: the HIR hash covering all bodies, and tracked
options), and full metadata carries the same one. That's a complete
identity from the start, so the HIR-based hash override is gone.

Writing early metadata deletes the crate's stale full metadata, so a full
file next to early metadata is always current: crate loading prefers it
when it exists, with no modification-time rules. A failed compilation
retracts its early metadata (rustc on errors or panics, cargo on crashes),
and a dependent waiting for full metadata stops when that happens. The
.failed-rmeta markers are gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Under headstart, every rustc-compiled unit starts on its rlib dependencies'
early metadata; proc-macro and dylib dependencies still need a full build.
A crate that links records where each dependency's rlib will be written,
and waits for them before linking (before code generation with LTO), with
the same pause protocol. Writing early metadata deletes the crate's stale
rlib, so an rlib that exists is current. Cargo drops the synthetic
wait-for-every-rlib edges for these units.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
A paused compilation announces each file it waits for, and resumes as soon
as it's written, without touching the jobserver. Cargo stops counting it
as running while it's paused, so its slot goes to other work, and counts
it again when it resumes, starting nothing new until enough jobs finish.

The previous rule (cargo holds new work until a ready job has taken a
token) stalled builds whose build scripts also draw tokens, such as
libgit2's parallel C compile: cargo-0.87.1 went from 96 s to 101 s
(build) and 51 s to 66 s (check) with headstart. Resuming without a
token: 83 s and 41 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
With optimization on, writing early metadata computed deduced parameter
attributes, which optimizes every function body. The MIR inliner then
asked dependencies still loaded from early metadata whether they had MIR,
and the "no" stayed cached after the swap to full metadata. Dependents
failed with "missing optimized MIR" (typst and zed, which build
dependencies with opt-level 2 or 3).

Early metadata now records optimized MIR only for coroutines, and no
deduced parameter attributes. -Zearly-metadata-verify reports queries
that only full metadata answers when asked of an early crate.
scripts/check-swap.sh reproduces the swap at every opt-level; it fails
at 2 and 3 without this change.

bench.sh takes per-project cargo arguments (dir::args) and deletes each
project's target directory when done.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
- Evaluate statics before the early write with par_hir_body_owners, as
  body checking does, and prefetch consts' and const fns' MIR in parallel
  under -Zthreads. zola's minify_html_common (large generated statics)
  took 78 s single-threaded before its early write, against 26 s for all
  of analysis with -Zthreads=8 and no headstart.
- Without code generation, full metadata doesn't compute the reachable set
  under -Zearly-metadata. It's never read there, and at opt-level >= 1 it
  ran the MIR inliner against dependencies still loaded from early
  metadata (2819 -Zearly-metadata-verify reports in cargo check --release).

Sweeps of all 53 rustc-perf benchmarks, headstart off and on, all clean
with -Zearly-metadata-verify: build --release, check --release, and check
and build with -Zthreads=8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Results
- 13 real projects (rust-analyzer, zed, bevy, lemmy, polars, wasmtime,
  typst, helix, nushell, atuin, vaultwarden, zola, lldap), check and
  build, with and without -Zthreads=8. Default front end: 2-54% (check),
  0-42% (build). On top of -Zthreads: up to 25%, three within noise.
- rustc-perf -Zthreads=8 re-timed after the statics fix: nothing slower.
- Correctness: six sweeps (debug, release, -Zthreads; check and build)
  clean with -Zearly-metadata-verify; UI and cargo test suites.

rustc
- Keep the early-metadata guard through linking, so a code generation
  failure withdraws early metadata without relying on the build tool.
- Fix stale comments (jobserver tokens, choosing metadata by age, empty
  metadata for binaries); -Zearly-metadata-verify help covers both checks.

cargo
- Patch renamed to 0001-headstart.patch; fix stale comments.

Scripts
- scripts/sweep.sh and scripts/real-projects.sh (pinned commits, bevy's
  lockfile in projects/) replace one-off drivers.
- bench.sh: a new output directory per run by default.
- critical-path.py: pauses counted per episode, rlib waits included.
- check-swap.sh covers opt-level z; bench-incremental.sh fails if its
  edit doesn't apply.

Docs: one list of what output can differ; old-version numbers labeled;
retraction, emit parsing and incremental checks described accurately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tet2pfGBx9bM6ge28w8rgM
@zmaril
zmaril added this pull request to stack #3 September 30, 2026 18:25
@zmaril
zmaril merged commit b1d8c9d into main Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant