feat(flow): retain Jev diagnostics and verify concise handoffs - #167
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Independent PR 167 shipping review.Verdict. PASS+NOTES for the authorized 167, 168, 169 landing sequence. This is not release qualification. Reviewed head is 5d47750. Base is 49921a1. Local HEAD and GitHub PR 167 match that head. The worktree was clean before and after verification. GitHub reports CLEAN and all seven attached CI checks passed. The reviewer did not author repository source. The review reused earlier independent telemetry, replay and timing audits, inspected the exact cumulative base diff and latest presentation refinements, and ran the relevant behavior tests on this frozen head. No source authority regression was found. Domain transitions, Session schema, close transaction, Jev transport, production recovery defaults and package version are unchanged. New recovery information is diagnostic. It does not change selection, mutation grants, cancellation, budgets or retries. Delivery summaries derive from the same closed state and preserve the full report. Native reporting graders remain separate from unsupported decision replay. Fresh verification passed 210 tests with 1446 assertions across nine files. The command was: env -u TYPESAFE_API_KEY bun test tests/jev-decision-provider.test.ts tests/recovery-policy.test.ts tests/runtime-close.test.ts tests/assurance-projection.test.ts tests/delivery-scenarios.test.ts tests/delivery-confirmation.test.ts tests/eval-replay-capability.test.ts tests/scenario-steps.test.ts tests/eval-progress-wait.test.tsThe tests exercise real adapter behavior with mocked transport, recovery tool/status and controller boundaries, close/replay delivery, strict negative presentation cases, CLI replay capability classification, six historical catalog hashes, and unchanged count-only timeout/cancellation timing. The log is verification.log in this directory. No reviewer source edits, model calls, credential access, GitHub comments, merges or publication occurred. CI provides current-head repository, platform persistence, packed smoke and patch-qualification evidence. The local focused run is not a fresh paid model experiment. Notes.
|
Why
Flow discarded Jev timing and decision scores, and asked agents to repeat full delivery reports. Retain those measurements and provide concise handoffs. The exploratory pilot then exposed missing disclosures, overly literal grading and unsupported replay gating. This change addresses those defects while retaining the original measurements.
Scope
Tradeoffs
The full report remains in tool output. Shorter generated text does not establish lower elapsed time. Jev reservation estimates differ from billing. Token usage covers only the final validated response.
Reporting graders support a bounded grammar. They do not infer arbitrary natural-language equivalence. Required limitations use coherent current statements. Full-detail grading preserves substantive values, context and multiplicity rather than requiring identical Markdown.
The collector grades the last manager text part. Missing-summary fallback and unknown native exit retain deterministic coverage. Saved pilot answers are development regressions, so their revised grades do not establish new prompt behavior.
Blast Radius
Production changes affect recovery diagnostics and delivery presentation. Selection thresholds, budgets, mutation authority and Session v5 retain their behavior. Full close replay remains compatible.
Evaluation changes affect followup dispatch, reporting fidelity and replay classification. All six historical release catalog hashes remain unchanged. Required case versions and qualification thresholds are unchanged. These cases establish neither release qualification nor Jev decision quality.
Verification
c92093674de66f3d292a11eb9bc701b3a52f825b. The full repository gate passed 2,206 tests with 20 optional live tests skipped. Focused distribution, schema and prompt checks passed all 78 tests.The original implementation passed the authorized GPT-6.1 Sol probe and plan-only-stops smoke. That two-dispatch budget was consumed. A direct close probe reduced the handoff from 1,464 to 852 characters while preserving the full report byte-for-byte.
Original live pilot
The five-case pilot used GPT-6.1 Sol for manager and reviewer on
d44aa4b305086942635d0529a5b73fe059df396c. All nine authorized top-level dispatches were consumed. Dollar cost was not reported. No Jev calls ran.The original report retains 2/4 scored passes and one excluded host failure. Three completed closures had no detected false completion. Three independent reviewer assignments passed. The report, transcripts and cassette expectations remain unchanged.
Development verification
The revised grader accepts the retained full-detail, deferred and idle answers. The audit still fails for two omitted disclosures and the newly clarified exact-Goal requirement. Adding those facts passes. This is development regrading, not a fresh live score.
Independent probes accepted five positive controls and rejected 30 negative controls. Additional regressions protect feature identity case, record multiplicity, context and dash-prefixed outcome association. Native close and reviewed evidence binding remain strict.
All five pilot cassettes report unsupported native provenance instead of creating a false gated comparison. The original cassette bytes remain intact. Existing decision-only replay remains gated.
The host fixture proves that updates inside existing parts are unmeasured by the current counter. The diagnostic now states this limitation. It does not claim to repair the original interruption. Fresh live confirmation and instrumented stall investigation remain before release. No new paid calls ran during these fixes.
Fresh confirmation on c920936
The second five-case pilot ran on
c92093674de66f3d292a11eb9bc701b3a52f825b, with GPT-6.1 Sol as manager and reviewer and recovery off. All nine authorized top-level dispatches were consumed. No Jev calls ran. Provider dollar cost was unavailable.The frozen report retains 2/5 scored passes. All five cases reached their intended closure without host interruption. Four completed closures had no detected false completion. Four independent reviews passed, with no unsubmitted assignment. All three summary answers retained the exact Goal and all three assurance limitations. Full-detail and idle cases passed.
The completed answer genuinely omitted external-action authority. Its fraction-progress finding, both deferred findings and all four audit findings reject accurate finite wording. Original reports, attempts, transcripts and cassette expectations remain unchanged.
Followup correction
The new instructions explicitly retain closure, progress, assurance conclusion and external-action authority alongside the existing required facts. Labelled and unlabelled fields share whole-value scalar parsers. Assurance qualifiers bind to actual native checks; unavailable-platform reasons bind to declared proof; observation command, exit and nonpass qualification stay in one current record. Unsupported tails and contradictions remain failures.
Development regrading accepts four unchanged answers and leaves the completed answer failing solely for authority omission. Appending its missing authority statement passes. This is oracle verification, not a fresh live score. Independent review passed all 51 probes. The 105 focused reporting/prompt tests passed.
On
6765ee5930fe3dcd85eb6cecdd61c89aa50c3058, contribution preflight passed 2,221 repository tests with 20 optional live tests skipped, plus 78 focused surface tests. The packed plugin passed all 22 provider-free OpenCode checks. All 13 historical gated replays reproduced. The prepared next live proposal is limited to the remaining real prompt omission. Its two top-level dispatches need a new authorization; no such calls have started. Merge and full release qualification remain pending. The 9.6.0 single-provider profile is a separate draft and is not installed.Targeted authority confirmation on 6765ee5
The authorized two-dispatch run completed on
6765ee5930fe3dcd85eb6cecdd61c89aa50c3058. Both GPT-6.1 Sol probe/workflow dispatches were consumed, with no Jev call. Provider dollar cost was unavailable. The workflow passed native validation and independent review, completed and archived with no false completion or unsubmitted review.Independent inspection confirmed that the final handoff contains explicit not-granted external-action authority, the exact Goal, all three assurance limitations and truthful closure/progress/check-count facts. The targeted reporting requirement was observed successfully.
Its original automated result remains FAIL with three findings. The verifier rejected a Flow closure label, embedded progress/zero counts and an accurate required-pass sentence. Original report, transcript, case policy, source identity and tarball remain unchanged. Full gateway/revision identity was unavailable; the host observed OpenAI GPT-6.1 Sol for both roles.
The subsequent correction changes only evaluator code and regression tests. It binds finite compound records, check/zero counts and required-pass claims to accepted complete source evidence. Malformed known prefixes and unregistered Node/Bun pass assertions fail closed. Independent review passed 63 probes and the focused tests passed 127 cases. A separate development regrade accepts the unchanged targeted answer. The old frozen FAIL is not overwritten or promoted to qualification.
On verifier head
5d477502598b951c7dc2e10059d4682c7a49a395, contribution preflight passed 2,235 tests with 20 optional live tests skipped, and all 13 historical gated replays reproduced. The rebuilt tarball SHA256 remains5dae6dbebc1b2fe98649c3267c32f849124f21f9932cda5b99fc306823a31db0, and the unpacked-manifest SHA256 remainse32331bbc85ac9c5f3a1683c6894aced0b45a3db5cd17b24d5c14249d3fa8b1a. The packed bytes are identical to the targeted measurement. Original source/report/transcript and artifact hashes are intact. The separate development regrade records the new verifier commit and source hashes; changed outcomes are not silently resealed as release qualification. No further paid retry is needed for this observed authority requirement. Release 9.6.0 still needs its reviewed profile, frozen artifact, full paid qualification matrix and exact-artifact canary. No merge, tag or publication has occurred.