Skip to content

fix(evals): separate command integrity and prioritize delivery gates - #176

Merged
vriesd merged 9 commits into
mainfrom
fix/delivery-command-integrity
Oct 7, 2026
Merged

vriesd merged 9 commits into
mainfrom
fix/delivery-command-integrity

Conversation

@vriesd

@vriesd vriesd commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Why

Qualification rejected a truthful completed response because “its script remained unchanged” did not fit a qualifier that conflated script integrity with invocation identity. The parser erased an accepted native exit-zero pass. Repeated failures also surfaced only after 57 earlier attempts because delivery cases were appended to the catalog.

Scope

Replace the private unchangedInvocation boolean with one explicit command integrity enum. Recognize bounded script-only claims independently of result qualification. Verify script-only evidence against its registered command and the literal invoked script resource. Edited input arguments are separate from script identity. A closed Node/Bun locator handles literal quote/escape and relative-path forms, and refuses ambiguous wrappers, flags, expansions and compound commands. Raw invocation bytes remain unchanged. Preserve the stronger combined claim’s canonical gate identity requirement. Reject contradictory, conditional and unconsumed qualifiers. Runtime prompts and public package interfaces stay unchanged.

Move the five existing 9.6 delivery rows to the front of the declared catalog. Preserve every prior policy row, all 72 primaries, 17 reserves, model routes and pass thresholds. Scheduler and verifier continue to derive the same canonical order. Existing campaigns and approvals remain bound to their original plans.

Tradeoffs

The grammar stays bounded and rejects ambiguous claims. A script-only claim never establishes invocation identity or turns an observation into a pass. The richer private observation shape replaces the old boolean throughout its callers and tests.

Blast Radius

This changes delivery grading and the next candidate’s declared collection order. It requires fresh source, policy, catalog and plan bindings. Package behavior and release requirements remain unchanged. Original stopped reports are preserved; offline diagnostics cannot qualify them.

Verification

The parser regression had 11 passes and 20 failures before the fix. The order regression had one pass and two failures. After the fixes, the parent focused run passed 332 tests with 1,054 assertions across ten files. The parent offline diagnostic passed all 335 gradable inputs from six campaigns, preserving all report hashes and the one historical host exclusion. Full contribution preflight passed 2,562 tests with 20 existing skips and zero failures. Independent exact-head review on a416d29 is PASS+NOTES with no findings; ten independently authored controls pass. The initial prose-budget failure was corrected by shortening the documentation without raising its budget.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T10:37:06.969070Z 7e88d8d PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@vriesd

vriesd commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Independent exact-head review: PASS+NOTES for 7e88d8d against bb24b89.

A separate inherited GPT-6.1 Sol reviewer verified the unchanged retained native response against parent and head. Script-only integrity stays independent of invocation identity and pass intent. A recorded audit exit 12 with unchanged script remains nonpass beside the verify gate. Stronger combined identity, changed scripts, unavailable workspace, native review/source/output proof, unknown qualifiers and contradictions remain strict. The reviewer ran 305 focused tests and ten independently authored controls. No findings. Model family diversity is not claimed.

The five 9.6 delivery cases run first in the declared next catalog. Every prior policy row and all 72 primary attempts, 17 reserves and 114 maximum matrix dispatches remain unchanged. The new catalog/seed/plan/source requires fresh approval; old campaigns remain immutable.

Parent full preflight passed 2,535 tests with 20 existing skips and zero failures. Parent offline development diagnostics passed all 335 gradable retained inputs from six campaigns, preserving every official report hash and the historical host exclusion. These diagnostics are not qualification.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7e88d8d192

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread evals/delivery-scenario-checks.ts Outdated
@vriesd

vriesd commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Final independent review: PASS+NOTES for a416d29. This supersedes the initial 7e88 verdict after the grounded resource-binding review finding.

A separate inherited GPT-6.1 Sol reviewer ran 332 focused tests, ten independent whole-grader resource controls and the prior status/permission/opaque controls. Changed data arguments stay separate from the invoked resource. Actual script changes reject even with quote/escape/./ spelling. Ambiguous forms refuse proof. Raw command identity and source/review/output/native evidence checks remain intact. No findings. Model family diversity is not claimed.

Parent full preflight passed 2,562 tests with 20 existing skips and zero failures. Parent diagnostic regrading passed 335 native gradable inputs, preserving all six official report hashes and the one host exclusion. No release qualification or paid authority is inferred from this review.

@vriesd
vriesd merged commit dc665c3 into main Oct 7, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants