Skip to content

feat(flow): retain Jev diagnostics and verify concise handoffs - #167

Merged
vriesd merged 15 commits into
mainfrom
feat/jev-telemetry-handoff
Oct 5, 2026
Merged

vriesd merged 15 commits into
mainfrom
feat/jev-telemetry-handoff

Conversation

@vriesd

@vriesd vriesd commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Why

Flow discarded Jev timing and decision scores, and asked agents to repeat full delivery reports. Retain those measurements and provide concise handoffs. The exploratory pilot then exposed missing disclosures, overly literal grading and unsupported replay gating. This change addresses those defects while retaining the original measurements.

Scope

  • Retain process-local Jev transport and assessment telemetry, validated response usage, scores and threshold checks.
  • Add delivery summary lines while preserving the full close report and blocked recovery guidance.
  • Require unchanged Goal identity and all three assurance limitations in concise handoffs.
  • Support ordinary user followups with exact instruction bytes and unchanged legacy command encoding.
  • Add five report-only cases covering completion, deferral, nonzero observation, full-detail followup and idle status.
  • Grade bounded current assertions and full-report records. Faithful Markdown can pass. Changed values, record associations, omissions and supported contradictory claims fail.
  • Declare native-host replay requirements. Derive them for older cassettes, report unsupported comparisons and refuse expectation rewrites when evidence is unavailable.
  • Correct stall diagnostics to describe message and part counts. Timeout, cancellation and failure classification retain their behavior.

Tradeoffs

The full report remains in tool output. Shorter generated text does not establish lower elapsed time. Jev reservation estimates differ from billing. Token usage covers only the final validated response.

Reporting graders support a bounded grammar. They do not infer arbitrary natural-language equivalence. Required limitations use coherent current statements. Full-detail grading preserves substantive values, context and multiplicity rather than requiring identical Markdown.

The collector grades the last manager text part. Missing-summary fallback and unknown native exit retain deterministic coverage. Saved pilot answers are development regressions, so their revised grades do not establish new prompt behavior.

Blast Radius

Production changes affect recovery diagnostics and delivery presentation. Selection thresholds, budgets, mutation authority and Session v5 retain their behavior. Full close replay remains compatible.

Evaluation changes affect followup dispatch, reporting fidelity and replay classification. All six historical release catalog hashes remain unchanged. Required case versions and qualification thresholds are unchanged. These cases establish neither release qualification nor Jev decision quality.

Verification

  • Final-head contribution preflight passed on c92093674de66f3d292a11eb9bc701b3a52f825b. The full repository gate passed 2,206 tests with 20 optional live tests skipped. Focused distribution, schema and prompt checks passed all 78 tests.
  • Offline replay reproduced all 13 historical gated cassettes. The five pilot cassettes report unsupported native evidence. CLI regressions cover old empty fidelity, refusal to rewrite unavailable evidence and visible handler divergence.
  • The 90 focused delivery and prompt tests passed. Independent output, replay, host-diagnostic and comment audits found no remaining blocking issue.
  • The packed plugin passed all 22 provider-free OpenCode 1.18.31 host checks.
  • The initial full gate failed three benchmark tests on their outer timeout and one stale diagnostic assertion. The benchmark tests passed unchanged in isolation and in the full rerun. The assertion now checks the corrected wording. The first failed log remains retained.

The original implementation passed the authorized GPT-6.1 Sol probe and plan-only-stops smoke. That two-dispatch budget was consumed. A direct close probe reduced the handoff from 1,464 to 852 characters while preserving the full report byte-for-byte.

Original live pilot

The five-case pilot used GPT-6.1 Sol for manager and reviewer on d44aa4b305086942635d0529a5b73fe059df396c. All nine authorized top-level dispatches were consumed. Dollar cost was not reported. No Jev calls ran.

Case Frozen verdict Assessment
Concise completion Host interruption Pending apply_patch did not reach closure. Cause remains unknown.
Deferral Pass Unfinished work and unavailable macOS proof remained explicit.
Nonzero audit Fail, nine findings Two disclosures were omitted. Other findings mainly enforced formatting.
Full detail Fail, one finding Substantive facts survived Markdown restructuring.
Idle after closure Pass Current delivery stayed unavailable without reviving old work.

The original report retains 2/4 scored passes and one excluded host failure. Three completed closures had no detected false completion. Three independent reviewer assignments passed. The report, transcripts and cassette expectations remain unchanged.

Development verification

The revised grader accepts the retained full-detail, deferred and idle answers. The audit still fails for two omitted disclosures and the newly clarified exact-Goal requirement. Adding those facts passes. This is development regrading, not a fresh live score.

Independent probes accepted five positive controls and rejected 30 negative controls. Additional regressions protect feature identity case, record multiplicity, context and dash-prefixed outcome association. Native close and reviewed evidence binding remain strict.

All five pilot cassettes report unsupported native provenance instead of creating a false gated comparison. The original cassette bytes remain intact. Existing decision-only replay remains gated.

The host fixture proves that updates inside existing parts are unmeasured by the current counter. The diagnostic now states this limitation. It does not claim to repair the original interruption. Fresh live confirmation and instrumented stall investigation remain before release. No new paid calls ran during these fixes.

Fresh confirmation on c920936

The second five-case pilot ran on c92093674de66f3d292a11eb9bc701b3a52f825b, with GPT-6.1 Sol as manager and reviewer and recovery off. All nine authorized top-level dispatches were consumed. No Jev calls ran. Provider dollar cost was unavailable.

The frozen report retains 2/5 scored passes. All five cases reached their intended closure without host interruption. Four completed closures had no detected false completion. Four independent reviews passed, with no unsubmitted assignment. All three summary answers retained the exact Goal and all three assurance limitations. Full-detail and idle cases passed.

The completed answer genuinely omitted external-action authority. Its fraction-progress finding, both deferred findings and all four audit findings reject accurate finite wording. Original reports, attempts, transcripts and cassette expectations remain unchanged.

Followup correction

The new instructions explicitly retain closure, progress, assurance conclusion and external-action authority alongside the existing required facts. Labelled and unlabelled fields share whole-value scalar parsers. Assurance qualifiers bind to actual native checks; unavailable-platform reasons bind to declared proof; observation command, exit and nonpass qualification stay in one current record. Unsupported tails and contradictions remain failures.

Development regrading accepts four unchanged answers and leaves the completed answer failing solely for authority omission. Appending its missing authority statement passes. This is oracle verification, not a fresh live score. Independent review passed all 51 probes. The 105 focused reporting/prompt tests passed.

On 6765ee5930fe3dcd85eb6cecdd61c89aa50c3058, contribution preflight passed 2,221 repository tests with 20 optional live tests skipped, plus 78 focused surface tests. The packed plugin passed all 22 provider-free OpenCode checks. All 13 historical gated replays reproduced. The prepared next live proposal is limited to the remaining real prompt omission. Its two top-level dispatches need a new authorization; no such calls have started. Merge and full release qualification remain pending. The 9.6.0 single-provider profile is a separate draft and is not installed.

Targeted authority confirmation on 6765ee5

The authorized two-dispatch run completed on 6765ee5930fe3dcd85eb6cecdd61c89aa50c3058. Both GPT-6.1 Sol probe/workflow dispatches were consumed, with no Jev call. Provider dollar cost was unavailable. The workflow passed native validation and independent review, completed and archived with no false completion or unsubmitted review.

Independent inspection confirmed that the final handoff contains explicit not-granted external-action authority, the exact Goal, all three assurance limitations and truthful closure/progress/check-count facts. The targeted reporting requirement was observed successfully.

Its original automated result remains FAIL with three findings. The verifier rejected a Flow closure label, embedded progress/zero counts and an accurate required-pass sentence. Original report, transcript, case policy, source identity and tarball remain unchanged. Full gateway/revision identity was unavailable; the host observed OpenAI GPT-6.1 Sol for both roles.

The subsequent correction changes only evaluator code and regression tests. It binds finite compound records, check/zero counts and required-pass claims to accepted complete source evidence. Malformed known prefixes and unregistered Node/Bun pass assertions fail closed. Independent review passed 63 probes and the focused tests passed 127 cases. A separate development regrade accepts the unchanged targeted answer. The old frozen FAIL is not overwritten or promoted to qualification.

On verifier head 5d477502598b951c7dc2e10059d4682c7a49a395, contribution preflight passed 2,235 tests with 20 optional live tests skipped, and all 13 historical gated replays reproduced. The rebuilt tarball SHA256 remains 5dae6dbebc1b2fe98649c3267c32f849124f21f9932cda5b99fc306823a31db0, and the unpacked-manifest SHA256 remains e32331bbc85ac9c5f3a1683c6894aced0b45a3db5cd17b24d5c14249d3fa8b1a. The packed bytes are identical to the targeted measurement. Original source/report/transcript and artifact hashes are intact. The separate development regrade records the new verifier commit and source hashes; changed outcomes are not silently resealed as release qualification. No further paid retry is needed for this observed authority requirement. Release 9.6.0 still needs its reviewed profile, frozen artifact, full paid qualification matrix and exact-artifact canary. No merge, tag or publication has occurred.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-04T18:29:26.373971Z d5f84f0 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@vriesd vriesd changed the title feat(flow): retain Jev diagnostics and shorten handoffs feat(flow): retain Jev diagnostics and verify concise handoffs Oct 4, 2026
@vriesd

vriesd commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Independent PR 167 shipping review.

Verdict. PASS+NOTES for the authorized 167, 168, 169 landing sequence. This is not release qualification.

Reviewed head is 5d47750. Base is 49921a1. Local HEAD and GitHub PR 167 match that head. The worktree was clean before and after verification. GitHub reports CLEAN and all seven attached CI checks passed.

The reviewer did not author repository source. The review reused earlier independent telemetry, replay and timing audits, inspected the exact cumulative base diff and latest presentation refinements, and ran the relevant behavior tests on this frozen head.

No source authority regression was found. Domain transitions, Session schema, close transaction, Jev transport, production recovery defaults and package version are unchanged. New recovery information is diagnostic. It does not change selection, mutation grants, cancellation, budgets or retries. Delivery summaries derive from the same closed state and preserve the full report. Native reporting graders remain separate from unsupported decision replay.

Fresh verification passed 210 tests with 1446 assertions across nine files. The command was:

env -u TYPESAFE_API_KEY bun test tests/jev-decision-provider.test.ts tests/recovery-policy.test.ts tests/runtime-close.test.ts tests/assurance-projection.test.ts tests/delivery-scenarios.test.ts tests/delivery-confirmation.test.ts tests/eval-replay-capability.test.ts tests/scenario-steps.test.ts tests/eval-progress-wait.test.ts

The tests exercise real adapter behavior with mocked transport, recovery tool/status and controller boundaries, close/replay delivery, strict negative presentation cases, CLI replay capability classification, six historical catalog hashes, and unchanged count-only timeout/cancellation timing. The log is verification.log in this directory.

No reviewer source edits, model calls, credential access, GitHub comments, merges or publication occurred. CI provides current-head repository, platform persistence, packed smoke and patch-qualification evidence. The local focused run is not a fresh paid model experiment.

Notes.

  1. Paid pilot failures and the host interruption remain original evidence. Saved-answer regrading and deterministic regressions establish oracle behavior, not fresh model behavior or release qualification. The user-authorized fresh qualification still has to run on the final candidate and actual artifact.
  2. The existing recovery comparison importer rejects the new optional telemetry field. That known evaluation-consumer compatibility issue is assigned to downstream PR 169 in the authorized sequence. PR 167 is not a standalone claim that the Jev comparison pilot is ready. Complete and independently verify that compatibility layer before Jev evaluation or release.
  3. The underlying interrupted apply_patch cause remains unknown. This PR corrects diagnostic overstatement without claiming an execution repair.

@vriesd
vriesd merged commit 0a26704 into main Oct 5, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants