Skip to content

Add chronicle verify-raw to read back R2 artifacts against manifests - #334

Merged
MaxGhenis merged 1 commit into
mainfrom
nzhub/verifyraw
Oct 11, 2026
Merged

MaxGhenis merged 1 commit into
mainfrom
nzhub/verifyraw

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Closes #332.

What this adds

chronicle verify-raw is the read-only counterpart of publish-raw. It scans manifests under --root the same way. For each file entry with an R2 location, it:

  1. recomputes the canonical key with the same resolve_r2_prefix/build_r2_key path publish-raw uses, country segment included;
  2. streams the object back with wrangler r2 object get --remote --pipe;
  3. compares its SHA-256 and size with the manifest, and with the local file if present.

It never writes to manifests or to R2. The JSON report gives one status per entry: verified, key_mismatch, digest_mismatch, size_mismatch, missing_remote, no_r2_location, fetch_error or invalid_manifest. The command exits 0 only when every entry is verified and the scan reported no errors.

That report is a receipt anyone with read access can regenerate. Five independent reviews of the NZ packages (#328, #329, #330) noted that the R2 upload and read-back were only attested in prose.

Live run (hub, with R2 read credentials, at this branch's code)

Root Result
db/data/ird/wage_salary_distribution_2026 1 entry, verified
db/data/stats_nz/household_net_worth_2024 1 entry, verified
db/data/treasury/estimates_of_appropriations_2026_27 (--r2-prefix raw/nz) 2 entries, verified
db/data/ird/taxable_income_distribution_2025 1 entry, verified
db/data/mbie/tenancy_bond_rents_tla_2026 1 entry, verified

Negative checks ran on scratch copies of the IRD manifest:

  • Key without the nz segment (raw/ird/…, where a stray copy of the object actually exists): key_mismatch, exit 1.
  • Canonical key for an object that doesn't exist: missing_remote, exit 1. Wrangler's real "The specified key does not exist" message is recognised.
  • An all-digit hash, which YAML parses as an integer: invalid_manifest.

No manifest changed during any run.

Invariants and how they're checked

  • Read-only. No file under the root changes, including on failure paths, and the fake wrangler receives only get. There are snapshot tests, and a mutant that writes a sentinel file fails them.
  • verified means the key, digest and size all match. A mutant forcing every status to verified fails 10 tests.
  • Determinism. The report is identical across runs for a fixed fake remote, and entries keep publish-raw's scan order. A mutant reversing the order fails.
  • Property test over seeded random payloads. Verification passes exactly when the remote returns the manifest's bytes at the canonical key. It fails with the matching status for a flipped byte, a truncated payload, a missing object or a wrong-prefix key.
  • The incident as regressions. An object present only at the legacy raw/ird/… key reports missing_remote, and a manifest declaring the legacy key reports key_mismatch.

Tests

  • tests/test_chronicle_verify_raw.py: 25 passed. The build lane ran it, and the hub re-ran it.
  • Six in-memory mutants kill all 25 cases, with no setup errors.
  • tests/test_chronicle_artifacts.py still passes (36).
  • Ruff check and format pass on the changed files.
  • publish-raw is unchanged. On current main it already refuses a declared key that differs from the recomputed one before uploading, so the incident in Add chronicle verify-raw to read back R2 artifacts against manifests #332 needed an older installed chronicle.

Chronicle governance

This is repository tooling, not a source package. The closest approved role is ledger-contract-maintainer, but this PR changes no schema, identity, provenance or consumer contract. It adds a read-only CLI command and its tests in chronicle/artifacts.py, chronicle/cli.py and chronicle/harness.py, plus README and changelog lines. No facts or bundle pins change.

Built by a GPT-6.1 Sol lane from the NZ hub's brief. The hub verified the tests and ran the live checks above. An independent review follows before merge.

🤖 Generated with Claude Code

Add chronicle verify-raw with publish-raw's manifest scan and country-aware
resolve_r2_prefix/build_r2_key routing. Stream declared R2 objects through
Wrangler get --remote --pipe, hash binary bytes, count their length, and
compare manifest and optional local measurements. Emit frozen reports and
JSON with closed statuses and separate no_r2_location counts. Exit 1 for any
unverified entry or manifest error, following artifact command conventions;
missing R2 declarations fail without adding a skip flag. Declared buckets
take precedence over the fallback --r2-bucket option.

Read-only invariant: snapshot all root file bytes and recursive directory
entries before and after success and every failure status; assert fake
Wrangler receives only get calls and never put/delete. A write-sentinel
mutant proves the snapshot checks detect writes.

Verified invariant: require declared key == canonical key, remote SHA-256 ==
manifest SHA-256, and remote size == manifest size. Assert all three on the
successful binary fixture; fake-remote corruption, truncation, missing-object
and wrong-prefix tests exercise failures. Compare present local bytes too.

Determinism invariant: identical fake remote yields identical complete JSON
reports across runs; assert sorted manifest paths and manifest file insertion
order explicitly. A reverse-scan mutant fails that ordering assertion.

Payload property: Hypothesis is absent from pyproject.toml and existing tests;
use seeded random bytes without a dependency. Cover empty and multi-chunk
payloads, byte flips, truncation, missing objects and wrong-prefix keys.

Validation: 25 tests passed in tests/test_chronicle_verify_raw.py using the
assigned Python interpreter and required no-cache/basetemp flags. Ruff check
and format check passed on all changed Python files; git diff --check passed.
Six in-memory mutants ran in two one-file pytest invocations: 33 body failures
and 17 passes, with every one of the 25 new cases killed and no setup errors.
Exact failing IDs and E lines are in .hub-scratch/report.txt. Source-file hashes
are unchanged across mutation runs; pytest scratch was removed after each run.

publish-raw is unchanged. This checkout already rejects differing recorded
keys before upload or manifest writes, so the incident's key rewrite is not
a remaining publish-raw behavior in the code read for this change.

Co-Authored-By: GPT-6.1 Sol <noreply@openai.com>
@MaxGhenis
MaxGhenis merged commit 6d3c37a into main Oct 11, 2026
3 checks passed
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merged by the NZ hub (session local_e407a527) at 6b2a859 with --merge. Gates: gh pr checks exit 0; MERGEABLE; not draft; no CHANGES_REQUESTED review. Reviewed head: 6b2a859. Independent GPT-6.1 Sol review at 6b2a859: APPROVE_WITH_NITS, Needs Max: no (25 new tests and 36 existing pass; eight in-memory mutants kill all 25 cases; argv handling, read-only behaviour and size-versus-digest precedence probed). The hub also ran the command live against R2: six NZ objects verified, and legacy-key and missing-object cases refused. The review's diagnostic nit (ANSI codes inside the missing-object phrase, and over-broad NoSuchKey matching; both still report invalid) goes to a follow-up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add chronicle verify-raw to read back R2 artifacts against manifests

1 participant