Repository navigation
Make UK build outcomes honest in staging, telemetry and the Logbook - #1147
Conversation
Automated review pass (Claude Code, high effort) — round 1 at
|
…ng contract v3) A build whose gates refuse its candidate used to close its staging run as `completed`, and a build refused at the preflight gates read as a pass on the dashboard. Every failure carried `BUILD_FAILED`, and the Logbook and the telemetry could disagree on how a run ended. - Staging documents move to schema version 3: a `blocked` run status with a `block` record (gate phase, blocking gate ids, count) and a `failure_class` on failures. The delivery summary keeps contract version 2, which publication and the assemblers pin. Version 2 documents stay readable; the v2 fixtures are frozen beside a new v3 set (adds a blocked run). - `microcosm.build.run_outcome` classifies a build's end once: a gate block anywhere in the cause chain is `blocked`; Ctrl-C, out-of-memory, graph-node and other errors get distinct codes and classes; the Logbook disposition is derived from the same classification. - Both UK close-outs close the run `blocked` (terminal battery, national seam battery) or `completed` from that classification; the preflight refusal now records its gates, a `preflight_gates` stage, Logbook verdicts that resolve in the preflight gate document, and a `blocked` run at phase `preflight`. - The spine build closes gate blocks as `blocked` and classifies other errors. - The hosted emitter gains a `blocked` run event and `error_code` on failures. Readers must accept version 3 first: PolicyEngine/calibration-diagnostics#206. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A supervisor's SIGTERM killed a UK build without running its failure path: the staging run stayed `running` and no Logbook row was written. The spine build also let Ctrl-C escape without closing its run. - `microcosm.build.termination.raise_on_sigterm` turns the first SIGTERM into `BuildTerminatedError`, a `KeyboardInterrupt` subclass, so the existing interrupt arms record it and kernels' `except Exception` cannot swallow it. The handler fires once and then restores the default, so a second SIGTERM still kills the process at once. - The spine, dense and national entry points install it and, after the close-out, exit with status 143, so supervisors still see a termination. - The spine build gains an interrupt arm: the run closes `failed` (`INTERRUPTED` or `TERMINATED`) and the attempt records a `discarded` row. Known limits: settlement after the first signal can outlast a supervisor's grace period; SIGKILL and out-of-memory kills are covered only by the hosted emitter's heartbeat and write no Logbook row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A graph-built dense candidate could not be assembled: its `*.local_gates.json` held the graph's unsigned 26-gate document, with no `release_id` and no `shippable`, so the dense contract, the release preflight and the size evaluation all failed to read it. - `replay_uk_dense_gate_battery` projects the six local outcomes from the graph's terminal document (nothing is re-evaluated; the projection checks the declarations match), replays them through the gate battery under `MARKS_ARTIFACT`, grafts the scoped-report trio and signs the report with the Logbook build id as `release_id`. Without the key or an attempt it is written unsigned (`signing_error`) and is never shippable. - The dense build writes it before any refusal, so a blocked candidate leaves it too; the graph's own enforcement still sets the exit status. - The full document keeps its evidence name (`uk.full.gates.calibrated.gate_report.json`), is what the package binds, and is recorded as `outputs.full_gate_report`. Gate receipts now resolve: local gates in the signed report, the rest in the full document. - A build whose `--target-geographies` select no local targets writes no local report (`outputs.local_gate_report: null`, with the reason). - `--release-candidate` refuses to start without a 32-byte signing key or with a filter that drops the local targets; the preflight requires exactly 32 bytes. - The assembler and preflight tests now feed the build's own report; no test signs a report by hand. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The spine build never recorded its output file's digest, so nothing tied
a downstream build to the exact H5 a spine run wrote.
- The digest is measured once the H5 is fully written (smoke marking
included) and recorded in the sidecar (`output {filename, sha256,
size_bytes}`), a `<h5>.sha256` file in `sha256sum` format, the
`spine_h5_creation` stage event and the Logbook `pipeline` verdict
(`artifact_sha256`).
- `load_bound_spine_checkpoint` refuses an input H5 whose measured digest
differs from the sidecar's `output.sha256`; older sidecars without the
key still bind by content identity. Both full-build roles pass the
measured digest.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A resumed build reused stored graph results from earlier attempts, but nothing linked it to them or said how much it reused. - Each attempt's `request.json` now names its Logbook build id and staging run id. - Both roles' rowwise manifests gain an `execution` block: the graph store, the attempt directory, the nodes reused from the store versus computed (the first graph run to reach a node decides; later runs in the same attempt see this attempt's own results as hits), and the earlier attempt directories on the same store with their ids. It sits outside the run parameters, so the candidate identity is unchanged. - The counts reach the staging run as a `graph_execution` stage event before it closes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d the lineage Review round 1 on #1147 (findings 1, 3 to 7): - `StagingRunBundleWriterV2.block()` runs every check, the content policy included, before it changes any state, so a refused block leaves the run `running`. One `close_run_blocked` (shared by the rowwise and spine lines) then closes the staging run and the hosted emitter `failed` with `GATE_BLOCK_UNRECORDED` / `unrecorded_gate_block`, so no path leaves a run `running` and the two destinations cannot disagree. - `blocking_gate_ids` may be empty when the refusal named no gate; the count stays at least 1. No code writes `unidentified_gate` any more. - A raised refusal carries the gate statuses from the report the battery wrote beside its block, or from a spine refusal's phase report, so its `blocked` event lists them when the report could be read. - `_run_with_telemetry` classifies a non-zero return with `classify_return` (`BUILD_REFUSED` / `refused`), the attempt close-outs' vocabulary. - Interrupts are found anywhere in the cause chain, like terminations. - `execution.earlier_attempts` lists only attempts of the same request, most recent first, at most twenty, beside counts of every earlier attempt on the store and of the matching ones. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…forever Review round 1 on #1147 (finding 2): an old collector answers a `blocked` event with 422, and the delivery queue retried any non-202 status other than 401 and 403 indefinitely, wedging that run's queue for good. A 4xx other than a timeout (408) or a rate limit (429) on a run's events now makes the run local-only (`collector_rejected_events`) with one warning that names the status; the same rule covers a rejected registration. 401 still refreshes the token; 408 and 429 still retry. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
52732b2 to
0e80b03
Compare
|
Round 1 addressed at 0e80b03 (three commits on top of the five areas, so the delta reviews on its own). #1099 merged at bd886fc, so the branch is rebased onto Should
Nits
Verification at 0e80b03: the four shared suites ( |
Automated review pass (Claude Code, high effort) — round 2 at
|
| # | Item | Status |
|---|---|---|
| 1 | Failed block close left the run running |
Fixed. block() validates everything before changing state (staging_v2.py:1136-1150), and close_run_blocked falls back to fail() with GATE_BLOCK_UNRECORDED on both destinations. The spine now shares it. |
| 2 | 422 retried forever | Fixed. Any 4xx except 408/429 makes the run local-only, for events and registration (collector.py:125-137, :204-209, :267-270), with one warning naming the status. |
| 3 | Missing codes / no error_code on a non-zero return |
Fixed here: classify_return gives BUILD_REFUSED / refused. The three new codes are handed to calibration-diagnostics#206. |
| 4 | unidentified_gate placeholder |
Fixed. Empty ids are allowed with a count ≥ 1. No minItems is deliberate and documented. |
| 5 | gate_statuses on a raised block |
Fixed, read from the battery's written report or the spine's in-memory report. |
| 6 | earlier_attempts unbounded |
Fixed: same-request attempts only, newest first, at most 20, plus two counts. |
| 7 | KeyboardInterrupt only checked at the top |
Fixed, it walks the cause chain (run_outcome.py:199-201). |
Each fix has a test that fails without it. I copied the new test files onto d182a1a2 (before the round-1 commits) and they fail there:
- the refused-block tests in
test_staging_v2; - the 422/404 cases in
test_telemetry_emitter; - the lifecycle, replay and unrecorded-block tests in
test_uk_full_build_cli.
At the head, the six affected suites pass apart from the two tests below, and ruff check is clean.
The CI failure
Both engine-free jobs fail on one test, test_telemetry_emitter.py::test_pre_eligibility_spool_is_not_uploaded_after_upgrade (assert spool.has_pending()). Locally, test_current_pre_alembic_spool_is_adopted_without_losing_events fails the same way.
- The cause is the fixtures. They write events with a fixed
created_atof2026-10-02T10:00:00+00:00(test_telemetry_emitter.py:276, and the pre-Alembic test).EventSpool.__init__callsprune(), which deletes events older thanRETENTION_DAYS = 7(spool.py:216). From 2026-10-09 10:00 UTC the fixture events are pruned on open, sohas_pending()is false. mainis affected too. Both tests fail atmain's heada469ebe3.main's last green run (36557ca) started at 09:11 today, before the cutoff, and its later runs are still in progress.- The PR's emitter changes don't touch
spool.py.
Fix (on main, or carried here): date the fixture events relative to now, e.g. (datetime.now(UTC) - timedelta(hours=1)).isoformat(), for both the event and run rows. Optionally add a test that an event older than RETENTION_DAYS is pruned on open, so the retention behaviour is pinned on purpose.
New
- Nit: a 404 on run events now makes the run local-only (
collector.py:204). That's right for a settled route, but during a collector deploy a briefly missing route would drop the run's hosted view for good. Fine to keep given the deploy order, but worth one line indocs/uk-staging-operations.mdnext to the 422 note.
Originally stacked on #1099 (always-on telemetry emitter), which merged at bd886fc; this now targets
main.Readers must accept staging contract version 3 first: PolicyEngine/calibration-diagnostics#206 has to be deployed and its collector qualified (a
blockedevent returns 2xx) before this merges. An old collector answersblockedwith 422; since review round 1 the emitter treats that as settled and makes the run local-only instead of retrying, so a blocked run would be missing from the hosted view rather than wedging delivery. PolicyEngine/calibration-diagnostics#206 still goes first so no blocked run is lost.Why
After #1115, a review of how UK builds show up in the calibration dashboard found eight gaps:
error_code: BUILD_FAILED, andfailure_classwas never written.completed; the dashboard inferred "blocked" from counts.*.local_gates.jsonheld the graph's unsigned 26-gate document, with norelease_idorshippable.runningwith no Logbook row. There was no SIGTERM handler, and the spine build ignored Ctrl-C.What changes (one commit per area)
A. Staging contract v3 and one outcome classification (gaps 1, 2, 6, 7).
schema_version3. A run can now endblocked, withblock {phase, blocking_failure_count, blocking_gate_ids}. Failures gainfailure_class.contract_version2, whichpublish_cli, both assemblers and the dashboard pin.microcosm.build.run_outcomeclassifies a build's end once, and the staging bundle, the hosted emitter and the Logbook disposition all take it from there:blocked;INTERRUPTED, SIGTERMTERMINATED,MemoryErrorOUT_OF_MEMORY, a graph-node failureGRAPH_NODE_FAILED, and anything elseBUILD_FAILED.failed;discarded.preflight_gatesstage;blockedrun at phasepreflight.blockedrun event anderror_codeon failures.failedwithGATE_BLOCK_UNRECORDED, so no run is leftrunning.blocking_gate_idsmay be empty when the refusal named no gate (never a placeholder), and a raised refusal'sblockedevent carries the gate statuses when its report could be read._run_with_telemetry, each attempt's own close-out closes the hosted emitter with this classification first, so the wrapper's generic complete or fail only runs when nothing closed it. An error raised before the attempt takes over (its preflight digest, say) is classified there too. Add always-on build telemetry delivery #1099's test that expected a gate-blocked dense run to readfailednow expectsblocked.B. Termination signals (gap 5).
raise_on_sigterm()turns the first SIGTERM intoBuildTerminatedError. It is aKeyboardInterruptsubclass, so the existing interrupt arms record it and kernels'except Exceptioncan't swallow it.failedwith its own class and writes adiscardedrow.C. Sign the dense local gate report (gap 4).
MARKS_ARTIFACT, grafts the scoped-report trio, and signs the result with the Logbook build id asrelease_id.uk.full.gates.calibrated.gate_report.json. The package binds it, and it is recorded asoutputs.full_gate_report.--release-candidaterefuses to start without a 32-byte signing key, or with--target-geographiesthat select no local targets.outputs.local_gate_reportis null, with the reason recorded. It never fills the missing local gates in as not applicable.D. Spine output checksum (gap 3).
output);<h5>.sha256file;spine_h5_creationstage event;pipelineverdict.output.sha256. Older sidecars still bind by content identity.E. Resume lineage (gap 8).
request.jsonnames its Logbook build id and its staging run id.executionblock, outside the run parameters, so the identity is unchanged. It records:graph_executionstage event.Review round 1
Two commits on top of the five areas (see the round 1 reply for the item-by-item):
block()validates before mutating, and a refused block falls back tofailedwithGATE_BLOCK_UNRECORDEDon both destinations; no placeholder gate ids; gate statuses on raised refusals;classify_returnin_run_with_telemetry; interrupts found in the cause chain;earlier_attemptsbounded to the same request.Known limits
unexpected_process_exitevent cover them, and they still write no Logbook row. Area E lets the next attempt on the same store name the killed one.Testing
classify_failure, including the cause chain, wrapped spine blocks and every class;logbook_disposition;TERMINATED,discarded) on the spine, dense and national lines._check_uk_dense_gate_report;--resume requirereplay reports reuse and names the cold attempt. Request ids and the stage event are tested too.tools/generate_staging_contract_fixtures.py --checkpasses, the v2SHA256SUMSis unchanged, andtools/ci_test_plan.py verifypasses.mainafter Add always-on build telemetry delivery #1099 merged):test_run_outcome,test_staging_v2,test_telemetry_emitter,test_termination,test_uk_full_build_cli,test_uk_frs_spine,test_uk_rowwise_national_roleandtest_uk_calibration_runall pass except two pre-existing spool-migration tests intest_telemetry_emitter(test_pre_eligibility_spool_is_not_uploaded_after_upgrade,test_current_pre_alembic_spool_is_adopted_without_losing_events), which fail the same way on a clean checkout oforigin/mainon this machine whilemain's CI at 36557ca is green, so they are a local-environment failure and touch nothing this PR changes.🤖 Generated with Claude Code