Skip to content

feat(evals): prepare a separate Sol and Jev advisory pilot - #169

Merged
vriesd merged 5 commits into
mainfrom
feat/jev-shadow-pilot
Oct 5, 2026
Merged

vriesd merged 5 commits into
mainfrom
feat/jev-shadow-pilot

Conversation

@vriesd

@vriesd vriesd commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Why

The current Jev adapter returns transport telemetry that the strict comparison importer rejects. Fix that compatibility gap and add a separate pilot for Sol decisions with and without Jev advice. The existing comparison measures Sol choices versus Jev controller choices; it does not measure Sol consuming advice.

Scope

  • Preserve optional bounded telemetry and resolved-model metadata when importing current advice. Keep old receipts valid and retain strict nested validation. Preserve original telemetry when answered advice becomes controller-rejected.
  • Add offline prepare, bind, and check commands for independent GPT-6.1 Sol decision contexts on identical packets. Jev advice is treatment input. Bind complete prompts, source files, cases, packets and advice. Require shared execution indices to match seeded case and arm slots, with explicit incomplete coverage.
  • Keep forbidden returned choices in paired diagnostics. Separate missing, unavailable, invalid, answered and filtered advice. Preserve scoped nullable telemetry and known legacy answered fields.
  • Reject reviewed holdout input during development. Keep labels outside model prompts. Treat provenance and serving-model identity as unverified declarations.

Tradeoffs

The tool prepares and checks evidence; it dispatches no models. The eight bundled controls remain synthetic and unreviewed. Representative captured blockers, independent labels and a reviewed native collection protocol are required for a real pilot.

These read-only decisions cannot establish successful workflow recovery, fewer interruptions, executed-action safety or qualification. Such rates remain unknown. Reservations are not billed cost.

Blast Radius

This PR depends on #168. It changes evaluation tooling only. The 9.6 Sol-only release matrix, production shadow stopping behavior, recovery thresholds and qualification registry remain unchanged. It adds no delegated recovery profile or paid authorization.

Verification

  • The real production adapter initially failed all three new import controls. The narrow schema correction makes them pass, including missing-key and model-mismatch telemetry.
  • The actual adapter with a mocked transport roundtrips through collection and import under explicit simulation. No live attestation or network call is fabricated.
  • Focused verification passed 44 tests with 332 assertions. Independent review of both review fixes at 807f5f3 passed 29 tests with 206 assertions. Typecheck, affected lint and whitespace checks passed.
  • Independent comment audit found no introduced comments, suppressions or dead wrappers.
  • Full clean-commit push preflight passed at 807f5f3. The main suite passed 2,259 tests and skipped 20 optional live checks, with no failures. Package smoke and path-sensitive focused checks passed.
  • Review regressions reproduced telemetry loss and missing execution-order binding before the fixes. The parent run showed 18 passes and six failures. Strict seeded slots now pass while baseline-first, duplicate, invalid and filtered execution indices fail. Missing slots remain unverified declarations.
  • All 13 gated replays reproduced at 807f5f3.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-05T09:57:01.348515Z 15e4770 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 15e4770a9a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread evals/recovery-decisions/advisory-pilot.ts
Comment thread evals/recovery-decisions/compare.ts
@vriesd

vriesd commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Independent shipping verification PASS at 807f5f3 against f2a7ef7. The reviewer did not author the code. Fresh advisory/import verification passed 15 tests and 113 assertions. Source and diff remained clean. Seeded observation slots and rejected-answer telemetry fixes passed. Synthetic controls, inference origins, model identity and execution order remain explicit unverified declarations. No recovery authority or production qualification entry is added. No paid calls occurred. Parent posts this independent verdict before the authorized sequential landing.

@vriesd
vriesd changed the base branch from release/9.6.0-preparation to main October 5, 2026 13:37
@vriesd
vriesd merged commit 8c5ef35 into main Oct 5, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants