Repository navigation
fix(evals): separate command integrity and prioritize delivery gates - #176
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Independent exact-head review: PASS+NOTES for 7e88d8d against bb24b89. A separate inherited GPT-6.1 Sol reviewer verified the unchanged retained native response against parent and head. Script-only integrity stays independent of invocation identity and pass intent. A recorded audit exit 12 with unchanged script remains nonpass beside the verify gate. Stronger combined identity, changed scripts, unavailable workspace, native review/source/output proof, unknown qualifiers and contradictions remain strict. The reviewer ran 305 focused tests and ten independently authored controls. No findings. Model family diversity is not claimed. The five 9.6 delivery cases run first in the declared next catalog. Every prior policy row and all 72 primary attempts, 17 reserves and 114 maximum matrix dispatches remain unchanged. The new catalog/seed/plan/source requires fresh approval; old campaigns remain immutable. Parent full preflight passed 2,535 tests with 20 existing skips and zero failures. Parent offline development diagnostics passed all 335 gradable retained inputs from six campaigns, preserving every official report hash and the historical host exclusion. These diagnostics are not qualification. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7e88d8d192
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Final independent review: PASS+NOTES for a416d29. This supersedes the initial 7e88 verdict after the grounded resource-binding review finding. A separate inherited GPT-6.1 Sol reviewer ran 332 focused tests, ten independent whole-grader resource controls and the prior status/permission/opaque controls. Changed data arguments stay separate from the invoked resource. Actual script changes reject even with quote/escape/./ spelling. Ambiguous forms refuse proof. Raw command identity and source/review/output/native evidence checks remain intact. No findings. Model family diversity is not claimed. Parent full preflight passed 2,562 tests with 20 existing skips and zero failures. Parent diagnostic regrading passed 335 native gradable inputs, preserving all six official report hashes and the one host exclusion. No release qualification or paid authority is inferred from this review. |
Why
Qualification rejected a truthful completed response because “its script remained unchanged” did not fit a qualifier that conflated script integrity with invocation identity. The parser erased an accepted native exit-zero pass. Repeated failures also surfaced only after 57 earlier attempts because delivery cases were appended to the catalog.
Scope
Replace the private unchangedInvocation boolean with one explicit command integrity enum. Recognize bounded script-only claims independently of result qualification. Verify script-only evidence against its registered command and the literal invoked script resource. Edited input arguments are separate from script identity. A closed Node/Bun locator handles literal quote/escape and relative-path forms, and refuses ambiguous wrappers, flags, expansions and compound commands. Raw invocation bytes remain unchanged. Preserve the stronger combined claim’s canonical gate identity requirement. Reject contradictory, conditional and unconsumed qualifiers. Runtime prompts and public package interfaces stay unchanged.
Move the five existing 9.6 delivery rows to the front of the declared catalog. Preserve every prior policy row, all 72 primaries, 17 reserves, model routes and pass thresholds. Scheduler and verifier continue to derive the same canonical order. Existing campaigns and approvals remain bound to their original plans.
Tradeoffs
The grammar stays bounded and rejects ambiguous claims. A script-only claim never establishes invocation identity or turns an observation into a pass. The richer private observation shape replaces the old boolean throughout its callers and tests.
Blast Radius
This changes delivery grading and the next candidate’s declared collection order. It requires fresh source, policy, catalog and plan bindings. Package behavior and release requirements remain unchanged. Original stopped reports are preserved; offline diagnostics cannot qualify them.
Verification
The parser regression had 11 passes and 20 failures before the fix. The order regression had one pass and two failures. After the fixes, the parent focused run passed 332 tests with 1,054 assertions across ten files. The parent offline diagnostic passed all 335 gradable inputs from six campaigns, preserving all report hashes and the one historical host exclusion. Full contribution preflight passed 2,562 tests with 20 existing skips and zero failures. Independent exact-head review on a416d29 is PASS+NOTES with no findings; ten independently authored controls pass. The initial prose-budget failure was corrected by shortening the documentation without raising its budget.