Skip to content

mirth: second batch of checks; all Python ported to mirth-lab (Rust) - #30

Merged
zmaril merged 31 commits into
mirth/oraclesfrom
mirth/oracles2
Oct 10, 2026
Merged

zmaril merged 31 commits into
mirth/oraclesfrom
mirth/oracles2

Conversation

@zmaril

@zmaril zmaril commented Oct 10, 2026

Copy link
Copy Markdown
Contributor

Stacked on #29.

Second batch of checks (docs/checks.md)

  • suggest-diff (check 18): every machine-applicable suggestion applied one at a time must keep the program compiling. Finding 29: lint fixes that break builds.
  • diag-check (check 13): no compiler-internal debug output in diagnostics, spans in bounds. Finding 30.
  • repro-diff (check 15): outputs depend only on inputs (repeat, other directory, threads, decoy libraries).
  • gate-check (check 17): nothing unstable usable from stable. Finding 31: #[rustc_main] on a struct ICEs.
  • xlink (check 5): every target builds core/alloc and links a probe with no undefined symbols. Timeouts are now their own result class.
  • instr-check (check 21): PGO and coverage instrumentation round trips.
  • scale-check (check 10): compile time, memory, frames and future sizes grow about linearly. Finding 32: on long iterator chains the new solver got 10× slower and uses 4× the memory, bisected to nightly-2026-08-04 (#160254 is the only solver PR in that range; not confirmed with a revert build).

The Python is now Rust: crates/mirth-lab

All 36 scripts in rustc/*.py (about 6,200 lines) are now one library plus 32 clap subcommands.

  • The library has shared modules for UI test headers, compiling and running, typed JSON diagnostics, the sweep driver (rayon, --recheck/--known/--pause-on-finding), output normalization, Miri, artifact comparison, Cargo messages, coverage sites and the edit mutations.
  • mirth-rewrite is split into a library so rewrite-diff can run it in-process.

Each script was compared against its Python version before it was deleted:

  • Same findings: the oracle sweeps.
  • Same counts and gap lists: callgraph on build-blk; only the order of tied rows differs.
  • Same files, or the same lines in a different order: the other coverage tools.
  • Same results on the same inputs: the flag tools and fuzz/replay.
  • Exceptions:
    • Tools that make random edits use a Rust random generator, so seeds from Python runs don't replay the same edits.
    • flag-fuzz was only smoke-tested.

The shell callers build and call mirth-lab, and the docs name the subcommands. ui-solver-diff was removed because solver-diff covers it.

Python still in the repository: one-off research scripts under docs/ and ur/.

Commits

  • flag-walk, flag-fuzz, coverage-flags: --rows must be lo:hi (a malformed value ran every row; Python raised)
  • Finding 32: new-solver compile-time regression on long iterator chains, bisected to nightly-2026-08-04 (#160254 the only solver PR in range); facts and the 200-map test
  • xlink: a timed-out build is its own result class
  • rustc/coverage-flags.py removed (ported to mirth-lab coverage-flags)
  • coverage, coverage-compact, coverage-generators, grammar-coverage and ui-coverage scripts removed (ported to mirth-lab); docs name the subcommands
  • wip: fuzz, fuzz-replay, replay in mirth-lab (own mutations and cargo modules, to reconcile with port-ui)
  • Cargo.lock: mirth-lab's libc dependency
  • mirth-lab: coverage, coverage-compact, coverage-generators, grammar-coverage, ui-coverage (validated against the Python scripts); shell callers build and use mirth-lab
  • wip: flag tools
  • mirth-lab ui-fuzz, and the mutations and Cargo-artifact comparison in the library; ui-solver-diff.py removed (solver-diff covers it)
  • instr-check: executable-no-mangle-strip is expected under IR PGO (clang -fprofile-generate fails the same link)
  • mirth-lab callgraph: the reachability analysis in Rust (same counts and gap lists as callgraph.py on build-blk; ties ordered differently); coverage-report.sh uses it
  • The Python oracle scripts removed (ported to mirth-lab and validated); docs name the mirth-lab subcommands; how to run them in checks.md
  • mirth-lab: release-diff, xlink, scale-check (wait4 rusage through libc), abi-diff (typed differences; known and undecided classes labelled by variant); all 14 oracle tools ported
  • mirth-lab: repro-diff, gate-check, instr-check (validated: same results as Python, instr-check with the harness artifacts it surfaced fixed); spawn retried on ETXTBSY
  • mirth-lab: miri-diff, rewrite-diff (mirth-rewrite in process), suggest-diff (identical to the Python sweep: 7,357 suggestions, 159 tests with findings)
  • mirth-lab: the checks in Rust. Core library (uitest, rustc with typed JSON diagnostics and timeouts, driver with the frontier options, artifacts, miri, normalize) and opt-diff, solver-diff, crash-di
  • scale-check.py (check 10: growth exponents of compile time, memory, future size and stack frames over size-parameterized programs)
  • instr-check.py (check 21: PGO and coverage round trips), xlink.py + xlink-probe (check 5: build-std and link for every target)
  • gate-check.py (check 17: unstable attributes and library items from stable code); finding 31
  • repro-diff.py (check 15: repeat, path, threads, decoy); nothing new beyond the known parallel-frontend issues
  • diag-check.py (check 13: debug output in diagnostics, spans out of bounds, errors without a location, duplicates); finding 30
  • suggest-diff.py (check 18: every machine-applicable suggestion applied and compiled); finding 29: lint fixes that break builds

🤖 Generated with Claude Code

https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT

zmaril and others added 30 commits October 10, 2026 06:09
…d and compiled); finding 29: lint fixes that break builds

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…unds, errors without a location, duplicates); finding 30

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…eyond the known parallel-frontend issues

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…table code); finding 31

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…link-probe (check 5: build-std and link for every target)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…uture size and stack frames over size-parameterized programs)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
… JSON diagnostics and timeouts, driver with the frontier options, artifacts, miri, normalize) and opt-diff, solver-diff, crash-diff, diag-check, validated against the Python sweeps (same findings); mirth-rewrite split into a library and a binary

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…t-diff (identical to the Python sweep: 7,357 suggestions, 159 tests with findings)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…ts as Python, instr-check with the harness artifacts it surfaced fixed); spawn retried on ETXTBSY

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…c), abi-diff (typed differences; known and undecided classes labelled by variant); all 14 oracle tools ported

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…; docs name the mirth-lab subcommands; how to run them in checks.md

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…nd gap lists as callgraph.py on build-blk; ties ordered differently); coverage-report.sh uses it

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…ng -fprofile-generate fails the same link)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
… the library; ui-solver-diff.py removed (solver-diff covers it)

ui-fuzz: same options, outputs and finding classes as ui-fuzz.py on 48 tests
(276 vs 283 edits; debuginfo and io-checks P5s; finding 18 reproduced on
pinned-drop-sugar-no-core). Edits come from rand's StdRng seeded by the
test path, so sequences differ from Python's random.Random. Unified diffs
of the edit history from similar.

mutations: the 16 edits of mutations.py; the lookbehinds done by hand.
artifacts: collect and compare from artifacts.py (Collected), for the
fuzz and replay ports. mutations.py and artifacts.py stay until fuzz.py,
replay.py, flag-walk.py and coverage-flags.py are ported.

ui-solver-diff.py: solver-diff on tests/ui/transmutability compiles the
same 78 tests and finds the same 8 verdict differences; it notes the
E0277/E0521 case. Not covered: a different first-error message under the
same error code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…overage, ui-coverage (validated against the Python scripts); shell callers build and use mirth-lab

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
… flag-tools port

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
# Conflicts:
#	crates/mirth-lab/src/main.rs
…modules, to reconcile with port-ui)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
… ui-coverage scripts removed (ported to mirth-lab); docs name the subcommands

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…y and artifacts.py removed

fuzz, fuzz-replay and replay keep the Python options, output files and
formats. The library's cargo module has what they share: Cargo's JSON
messages, the compiler's verify-reuse and untracked-read lines, a
process-group timeout that tells a hung rustc from a looping build script,
and tree copies that keep modification times (Cargo's freshness depends on
them). Comparison is artifacts::collect/compare and the edits are mutations,
both from port-ui. A --seed names an edit sequence only within the Rust fuzzer.

Validated: fuzz-replay printed the same per-step lines and differing .rmeta
files as fuzz-replay.py on a 12-edit finding, also with --upto 3; replay
wrote the same records as replay.py for bitflags commits 290-296; a 2-worker
fuzz of fixtures/sink with rustc-verify12 found the same class (the known
allocation-sharing verify-reuse report) with the same finding files.

flag-walk.py and coverage-flags.py still import mutations.py and
artifacts.py; they are being ported on mirth/port-flags.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…flag-fuzz, audit-options, coverage-flags; the Python versions removed

Validated against the Python on the same inputs: flag-model byte-identical (15 variants);
flag-universe options/singles/pairs/requires.json byte-identical (covering-array sizes differ
by a row or few: another RNG); flag-rows, audit-options and flag-min identical output;
flag-walk same per-row results on 3 rows; coverage-flags same statuses (site counts differ
by the random edit). flag-fuzz calls `mirth-lab fuzz` (from mirth/port-fuzz).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
# Conflicts:
#	crates/mirth-lab/src/lib.rs
#	crates/mirth-lab/src/main.rs
…in rustc/ removed (artifacts.py, mutations.py); references updated

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…s, bisected to nightly-2026-08-04 (#160254 the only solver PR in range); facts and the 200-map test

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
…ed value ran every row; Python raised)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QXiEXbESemwqMLYKaWLDbT
@zmaril
zmaril added this pull request to stack #32 October 10, 2026 08:50
@zmaril
zmaril merged commit 76ab15f into main Oct 10, 2026
0 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant