feat(evals): prepare a separate Sol and Jev advisory pilot - #169
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 15e4770a9a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Independent shipping verification PASS at 807f5f3 against f2a7ef7. The reviewer did not author the code. Fresh advisory/import verification passed 15 tests and 113 assertions. Source and diff remained clean. Seeded observation slots and rejected-answer telemetry fixes passed. Synthetic controls, inference origins, model identity and execution order remain explicit unverified declarations. No recovery authority or production qualification entry is added. No paid calls occurred. Parent posts this independent verdict before the authorized sequential landing. |
Why
The current Jev adapter returns transport telemetry that the strict comparison importer rejects. Fix that compatibility gap and add a separate pilot for Sol decisions with and without Jev advice. The existing comparison measures Sol choices versus Jev controller choices; it does not measure Sol consuming advice.
Scope
prepare,bind, andcheckcommands for independent GPT-6.1 Sol decision contexts on identical packets. Jev advice is treatment input. Bind complete prompts, source files, cases, packets and advice. Require shared execution indices to match seeded case and arm slots, with explicit incomplete coverage.Tradeoffs
The tool prepares and checks evidence; it dispatches no models. The eight bundled controls remain synthetic and unreviewed. Representative captured blockers, independent labels and a reviewed native collection protocol are required for a real pilot.
These read-only decisions cannot establish successful workflow recovery, fewer interruptions, executed-action safety or qualification. Such rates remain unknown. Reservations are not billed cost.
Blast Radius
This PR depends on #168. It changes evaluation tooling only. The 9.6 Sol-only release matrix, production shadow stopping behavior, recovery thresholds and qualification registry remain unchanged. It adds no delegated recovery profile or paid authorization.
Verification