Repository navigation
fix(evals): recognize past-tense authority statements - #175
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Independent exact-head review: PASS+NOTES for ade48d1. A separate GPT-6.1 Sol reviewer verified the unchanged retained native response against parent and head. Parent rejects only authority parsing; head accepts the same input. The response literal in the regression exactly matches the native final text. The reviewer ran 209 focused tests plus 12 independent parser probes and verified that grant, contradiction, conditional permission, historical text, Goal text and registered command arguments remain strict. No findings. Parent full contribution preflight passed 2,494 tests with 20 existing skips and zero failures. Parent offline development diagnostics passed 277 gradable inputs across five prior campaigns; original official reports and the historical host exclusion remain unchanged. These diagnostics do not qualify a release. Model family diversity was unavailable; this is an independent role review on the inherited Sol model. GitHub CI and merge readiness remain separate gates. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ade48d1515
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Final independent exact-head review: PASS+NOTES for 794ba85. This supersedes the earlier ade48d1 verdict after the grounded prefixed-contradiction finding. The separate inherited GPT-6.1 Sol reviewer ran 253 focused tests, all 32 prefixed authority probes and additional native-input/scope controls. The unchanged native input passes; grants, compound and conditional assertions reject. Historical, quoted, Goal, registered-command and canonical disclosure boundaries retain expected behavior. No findings. Model family diversity is not claimed. Parent full preflight passed 2,504 tests with 20 existing skips and zero failures. Offline development diagnostics passed 277 retained gradable inputs, with all original reports unchanged. Qualification and release remain incomplete. |
Why
A completed workflow explicitly reported “External action authority was not granted.” The delivery grader rejected that truthful statement because its recognized verb grammar omitted “was”. This stopped qualification despite the native closure recording no external authority.
Scope
Recognize the past-tense authority statement while retaining strict scalar parsing. Flag unmatched current statements about external-action authority so prefixed grants and unfamiliar conditional assertions cannot bypass the guard. Preserve historical, quoted, Goal and registered-command boundaries.
Add the unchanged native response plus positive and adversarial regressions. Failing tests precede both fixes.
Blast Radius
This changes offline delivery grading. Runtime prompts, package bytes, release thresholds and reserves remain unchanged. Unknown authority assertions are conservatively rejected. The stopped campaign remains immutable and does not become qualifying evidence.
Verification
The original regression showed 19 passes and 5 failures. The prefixed contradiction regression showed 28 passes and 4 failures. The final focused run passed 253 tests with 703 assertions. Full contribution preflight passed 2,504 tests, with 20 existing skips and zero failures.
Independent exact-head review on 794ba85 is PASS+NOTES with no findings. All 32 prefixed authority probes reject unsupported claims. The earlier ade48d1 verdict is superseded by this corrected review.
Parent offline diagnostics passed all 277 gradable retained inputs across five campaigns. One historical host failure stays excluded; all original official report hashes are unchanged. These diagnostics are not fresh qualification.