diff --git a/README.md b/README.md index 6892f6c8..22a47187 100644 --- a/README.md +++ b/README.md @@ -7,17 +7,15 @@ benefits from an approved plan and an independent review: plan → approve → run one feature → validate → review → repeat or close ``` -Flow keeps one durable active feature run at a time. Once a session starts it -stays the workflow for that goal until Flow records completed, deferred, or -abandoned closure. It never silently falls back to ordinary coding, and it does -not fold a materially different request into the active goal. +Flow keeps one durable active feature run. Its goal stays active until completed, +deferred, or abandoned closure. Flow never silently resumes ordinary coding or +adds a materially different request to that goal. State lives in `.flow/session.json`, so the workflow survives a restart, a context change, or a lost transcript. -Flow is in preview: an opinionated workflow for consequential multi-step changes, -for people who read the review. It is worth its ceremony when a wrong change is -expensive, and it is overhead when it is not. +Flow is in preview. Its planning and review cost is worthwhile when a wrong +change is expensive and you read the review. ## When not to use Flow @@ -147,9 +145,9 @@ you granted. without shell access. New failures, source drift, or missing packets stop assignment without counting a failed review. 6. A passing review advances the plan. A failed feature needs an explicit retry - or independent-feature choice. Closure returns a versioned delivery report - with attempts, findings, and assurance limits. It grants no PR, merge, - publish, or release authority. + or independent-feature choice. Closure returns a short delivery summary + and the full versioned report in the same response. Both retain assurance + limits. Closure grants no PR, merge, publish, or release authority. Failed reviews retain finding ids; dropped live findings fail. `flow_plan_amend` records a same-goal reversible gate repair before review, @@ -205,34 +203,32 @@ bun install --frozen-lockfile bun run check ``` -`bun run check` runs typechecking, lint, build verification, tests, and package -smoke. Release CI also exercises the packed plugin in a real OpenCode host. +`bun run check` checks types, lint, builds, tests, and packages. Release CI tests +that package in OpenCode. -Maintained documentation starts at [docs/index.md](docs/index.md): -[development](docs/development.md) for repository structure, -[troubleshooting](docs/troubleshooting.md) for recovery, -[the maintainer contract](docs/maintainer-contract.md) for tools and runtime -invariants, and [ADR 0006](docs/adr/0006-bounded-intra-feature-waves.md) for the -bounded-wave rationale. +[Documentation](docs/index.md) includes [development](docs/development.md), +[troubleshooting](docs/troubleshooting.md), [runtime contracts](docs/maintainer-contract.md), +and [bounded waves](docs/adr/0006-bounded-intra-feature-waves.md). ## Recovery advice development preview -With `TYPESAFE_API_KEY`, `/flow-auto` defaults to shadow -advice, capped at six attempts and $0.02. Without it, advice is off. Use -`--recovery=off` to opt out or the options below to change limits. +With `TYPESAFE_API_KEY`, `/flow-auto` defaults to shadow advice, six attempts and +$0.02. Without the key or with `--recovery=off`, advice is off. ```text /flow-auto --recovery=shadow --recovery-calls=6 --recovery-usd=0.02 ``` -Shadow sends bounded goal, finding and candidate-remedy packets to TypeSafe. -It reports advice without authorizing mutations. Bare `flow_status` never calls -Jev. Attempts count retries. `/flow-auto stop` cancels auto and advice. Advice expires -after one hour; automatic retry limits remain until invocation ends. Neither -survives restart. `recoveryStatus` reports configuration, attempts and outcome. +Shadow sends bounded goal, finding and remedy packets to TypeSafe. It grants no +mutation authority. Bare `flow_status` never calls Jev. Retries count as attempts. +`/flow-auto stop` cancels auto and advice. Advice expires after one hour; retry limits last +until invocation ends. Neither survives restart. -Delegated mode requires release qualification; flags cannot bypass it. Tests prove -mechanics, not decision quality. See [ADR 0016](docs/adr/0016-delegated-recovery.md). +`recoveryStatus.last` exposes process-local timing, reservations, and decision +checks. See [telemetry semantics](docs/maintainer-contract.md#opencode-surface). + +Delegation needs release qualification. Tests prove mechanics, not decision quality. +See [ADR 0016](docs/adr/0016-delegated-recovery.md). ## License diff --git a/docs/maintainer-contract.md b/docs/maintainer-contract.md index 9dd69962..df816a6f 100644 --- a/docs/maintainer-contract.md +++ b/docs/maintainer-contract.md @@ -225,17 +225,15 @@ manager contract. rather than overwrite or delete either side. Closed status re-derives an archive collision from the existing history document, so interruption cannot restore automatic retry. This behavior adds no persisted recovery state. -- Every close path whose terminal state was durably accepted returns the same - derived `workflowData.delivery`: initial success, archive-pending recovery, - exact retry, and delayed replay from history. The projection declares a - `handoff` with `formatVersion: 1` and - `externalActionAuthority: "not-granted"`, then contains - the goal, closure, completed/total progress, every planned feature's attempt - count, latest outcome, terminal findings, Flow-reported artifact groups, and - derived tiered assurance with explicit limitations. -- Delivery is recomputed from the canonical closed Session or archive. It is not - written into Session v5 or archive JSON and is not a report artifact unless - the user separately requests one. +- Every durably accepted close returns identical derived `workflowData.delivery` + on success, archive-pending recovery, exact retry, and delayed history replay. + `handoff` declares `formatVersion: 1` and `externalActionAuthority: "not-granted"`. + Delivery contains goal, closure, progress, each feature's attempts, outcome, + terminal findings, reported artifact groups, and tiered assurance limits. + `summary.lines` is the default handoff. Full detail remains in `report` in + that same close response. Blocked `statusReport` stays unchanged. +- Delivery derives from the closed Session or archive. Session v5 and archive + JSON store neither projection nor report. A report artifact needs a user request. - Source identity hashes sorted effective workspace path/type/content tuples; `.git` and `.flow` are excluded. It is a content fingerprint, not a Git audit chain. @@ -257,8 +255,11 @@ implicit selection. See [Session v5](#session-v5). With `TYPESAFE_API_KEY`, `/flow-auto` defaults to shadow advice (six attempts, $0.02). Off disables advice, not retry limits. First retry stays automatic; fresh direction permits one retry/start. Shadow grants nothing; release -delegation is disabled. `recoveryStatus` separates configuration, attempts and -outcome. See +delegation is disabled. `recoveryStatus.last` adds process-local assessment duration, +transport attempts, reservation deltas, scores, and threshold checks. `reservedUsd` +is an upper-bound reservation, not billed cost. `responseUsage` counts only the +final validated response, excluding failed attempts. Without validated usage it +is null. Custom providers may leave transport facts null. See [ADR 0016](adr/0016-delegated-recovery.md). ### Commands diff --git a/docs/quickstart.md b/docs/quickstart.md index 0c063f82..12d3311b 100644 --- a/docs/quickstart.md +++ b/docs/quickstart.md @@ -58,7 +58,8 @@ Choose work where an incorrect change would be expensive: Flow inspects the repository and proposes an immutable feature plan. Confirm the canonical repository gate and any external evidence before approval. After approval, Flow implements one feature at a time, observes validation, dispatches -an independent review, and closes with a versioned delivery report. +an independent review, and closes with a short delivery summary. The same close +response retains the full versioned report for requested detail. If the host cannot continue between features, run `/flow-run` for each next feature. Run `/flow-status` after any interruption or failed review. diff --git a/evals/README.md b/evals/README.md index 0f5c0462..ce6bcdc4 100644 --- a/evals/README.md +++ b/evals/README.md @@ -198,6 +198,53 @@ discharged the entry before assertions existed. Declaring the command is no long enough; the plan has to name the case. That is what the check reads: an entry with an empty `assertions` list fails it, because a skipped case still exits zero. +## Delivery handoff pilot + +Five report-only cases exercise real workflows, archives, host validation, and +independent review. They cover concise completion, deferral with unavailable +macOS proof, a nonzero audit beside a separate passing gate, an ordinary full-report +followup, and idle status after close. Fixtures keep verification scripts immutable. +Recovery is off, so these cases need no Jev calls and measure no Jev decision quality. + +After paid authorization, run one attempt per case on the existing OpenAI route: + +```bash +env -u TYPESAFE_API_KEY -u OPENCODE_FLOW_REVIEWER_STEPS \ + OPENCODE_FLOW_REVIEWER_MODEL=openai/gpt-6.1-sol bun run eval -- --model openai/gpt-6.1-sol --repeat 1 --concurrency 1 \ + --scenario delivery-summary-completed --scenario delivery-summary-deferred \ + --scenario delivery-summary-observed-failure --scenario delivery-full-detail-followup \ + --scenario delivery-idle-after-close +``` + +This pilot has eight manager dispatches, including three followups, plus one +same-route entitlement probe. A distinct reviewer configuration adds one probe +per additional route. Reviewer-child generation remains paid work inside each +workflow. Nine harness dispatches are neither a dollar cap nor a limit on native +model requests. The parent authorizes and executes the paid run separately. + +Default summaries and full detail use separate cases because outcome collection +retains only the last manager text part. Earlier answers and multipart presentation +are not independently graded. Exact accepted close replay is valid full-detail +access when it preserves the same archive and operation. + +Graders use literal facts, the actual archived state, and native close/status +provenance. They compare full detail with the accepted close response and reject +missing limitations, false passes, changed verification scripts, and stale idle +handoffs. They import no production delivery formatter. Required release catalogs +remain unchanged. This single-route pilot establishes no release qualification. + +Summary grading accepts canonical fields and a bounded set of closure, progress, +assurance and authority sentences. It requires the unchanged Goal line and coherent +current assurance disclosures. Unsupported critical assertions fail instead of +guessing their meaning. Full grading permits Markdown sections and split fields, +while comparing each substantive record's context, value and multiplicity. +Saved pilot answers are development regressions. Regrading them does not establish +the behavior of new prompts or replace fresh live confirmation. + +Missing-summary fallback and unknown native exit remain deterministic compatibility +coverage. Current real close responses always include a summary, and ordinary +completed native commands supply an exit. This pilot does not inject either shape. + ## Cross-scenario metrics The original measures are reported for every run and asserted by none. Two are @@ -369,6 +416,12 @@ source drift between arming and observing, an abort, an excluded ask — carries `fidelity` note and is **reported, not gated**, on the same principle the thresholds use: gate what is measured, report what is not. +Scenarios that require native host provenance declare `replayRequires`. Their +recordings report `UNSUPPORTED` because decision replay lacks native call bindings +and host trace. Runtime handler, closure and completion-honesty differences remain +visible. Replay derives the capability note from the current scenario even when +an older cassette has empty `fidelity`, without rewriting the original bytes. + Capture identities are retained only when the appended marker matches a stored observation and its command. Replay binds those IDs to newly persisted captures; unknown or superseded references still fail. Older recordings without capture @@ -385,9 +438,10 @@ the developer's real `auth.json` into its throwaway home, so this is a hard rule rather than a precaution; `tests/eval-replay.test.ts` pins it. Only recordings someone has read belong in the committed `evals/cassettes/` set, -which is what CI gates on. `--accept` rewrites a cassette's recorded expectation -from the current replay; it is a deliberate act, and the rewritten expectation -lands in the diff to be reviewed like any other change to what the suite asserts. +which is what CI gates on. `--accept` rewrites supported cassette expectations +from the current replay. Review those changes like other test expectations. +It refuses cassettes with unavailable evidence, preserves their bytes and exits +with failure. The driver itself is proven without a model: `tests/eval-replay.test.ts` hand-writes the decision sequence of a passing `happy-path` attempt, replays it, and grades it @@ -477,14 +531,13 @@ different things: against the prompts. A scenario that sets `mayEscalate` is the exception: there the ask is the end the contract leaves, so the run is checked like any other and reads `PASS+ASK` or `FAIL+ASK`. -- `ABORT` — a step ended without going quiet, either `wedged` (no new message or - part while tool calls stayed incomplete, each named with the first line of its - command) or `still working` (producing output up to the deadline, so looping - rather than stuck). A wedge is called at three minutes of no change rather than - waited out to the twenty-minute deadline: three of the four recorded timeouts sat - on the same incomplete tool call for the full twenty and then printed exactly that - diagnostic, so the remaining seventeen minutes bought no evidence. Tokens and tool - calls collected before the abort are kept. Excluded from the pass rate and counted +- `ABORT` — a step ended without going quiet. Diagnostics distinguish no new + messages or parts while tool calls stay incomplete from new messages or parts + near the deadline. Updates inside existing parts are not measured. Neither + diagnostic establishes whether the model is making useful progress. The harness + aborts after three minutes without new messages or parts while calls stay + incomplete, or at the twenty-minute hard deadline. Tokens and tool calls + collected before the abort are kept. Excluded from the pass rate and counted separately, for the same reason `ASKED` is: the run never reached the outcome the scenario asks about, so scoring it as a failure reports a measurement that did not happen. One wedged attempt was the only failing threshold in a recorded report. diff --git a/evals/cassette.ts b/evals/cassette.ts index 7f2507f2..4b47b97d 100644 --- a/evals/cassette.ts +++ b/evals/cassette.ts @@ -105,7 +105,8 @@ export type FidelityNote = | "provider-error" | "evaluator-error" | "validation-identity-unrecorded" - | "workspace-diff-unreplayed"; + | "workspace-diff-unreplayed" + | "native-host-provenance-unreplayed"; export type Cassette = Readonly<{ cassetteVersion: number; @@ -300,14 +301,20 @@ function validationReport(output: string): string | null { } /** A cassette is gated only when nothing about it is known to be unreproducible. */ -export function isGated(cassette: Cassette): boolean { - return cassetteFidelity(cassette).length === 0; +export function isGated( + cassette: Cassette, + requirements: readonly "native-host-provenance"[] = [], +): boolean { + return cassetteFidelity(cassette, requirements).length === 0; } export function cassetteFidelity( cassette: Pick, + requirements: readonly "native-host-provenance"[] = [], ): readonly FidelityNote[] { const fidelity = new Set(cassette.fidelity); + if (requirements.includes("native-host-provenance")) + fidelity.add("native-host-provenance-unreplayed"); let identitiesRecorded = false; for (const event of cassette.events) { if (event.kind === "bash" && event.validation) identitiesRecorded = true; @@ -399,6 +406,7 @@ export function buildCassette(options: { readonly falseCompletion: boolean; readonly documents: readonly Record[]; readonly extraFidelity: readonly FidelityNote[]; + readonly replayRequires?: readonly "native-host-provenance"[]; }): Cassette { const events: CassetteEvent[] = []; const pendingResults = new Map< @@ -522,7 +530,7 @@ export function buildCassette(options: { }, finalText: options.finalText, assistantMessages: options.assistantMessages, - fidelity: cassetteFidelity({ events, fidelity }), + fidelity: cassetteFidelity({ events, fidelity }, options.replayRequires), } satisfies Cassette, options.projectPath, ); diff --git a/evals/delivery-presentation.ts b/evals/delivery-presentation.ts new file mode 100644 index 00000000..2a47d959 --- /dev/null +++ b/evals/delivery-presentation.ts @@ -0,0 +1,653 @@ +import { canonicalJson } from "./canonical-json.js"; + +type Closure = "completed" | "deferred" | "abandoned"; +type Assurance = + | "completion-supported" + | "completion-unsupported" + | "completion-not-claimed"; +type CurrentHandoffFacts = { + closure: (Closure | null)[]; + assurance: (Assurance | null)[]; + authority: ("not-granted" | "granted" | null)[]; + progress: ({ completed: number; total: number } | null)[]; + goal: string[]; + auxiliaryCounts: { + kind: "unfinished" | "blocking" | "advisory"; + count: number; + }[]; + assuranceCheckClaims: { count: number; status: "satisfied" }[]; + unavailableProofPlatforms: string[]; + observations: { + command: string; + exitCode: number | null; + unchangedInvocation: boolean; + qualification: + | "observation" + | "does-not-claim-pass" + | "claimed-pass" + | null; + }[]; + unsupported: string[]; +}; +export function presentationText(text: string): string { + return text + .split("\n") + .map(presentationLine) + .filter((line) => line !== null) + .join(" ") + .replace(/\s+/g, " ") + .trim(); +} +function presentationLine(raw: string): string | null { + const line = raw.trim(); + if (/^```\w*$/.test(line)) return null; + return line + .replace(/^(?:#{1,6}\s+|>\s*|[-*+]\s+|\d+[.)]\s+)/, "") + .replace(/\*\*([^*]+)\*\*/g, "$1") + .replace(/`([^`]+)`/g, "$1") + .replace(/\s+/g, " ") + .trim(); +} +function currentLines(text: string): string[] { + const lines: string[] = []; + let historical = false; + for (const raw of text.split("\n")) { + const line = presentationLine(raw); + if (!line) continue; + if ( + /^(?:historical|previous|prior|earlier|superseded)(?:\s+(?:handoff|report|context|reference))?:?$/i.test( + line, + ) + ) { + historical = true; + continue; + } + const explicitCurrent = + /\bcurrent (?:workflow|session|closure|assurance|authority|progress|goal|handoff|delivery|report)\b/i.test( + line, + ); + if (/^current(?:\s+(?:handoff|delivery|state|report))?:?$/i.test(line)) { + historical = false; + continue; + } + if ( + /^(?:historical|previous|prior|earlier|superseded)\b/i.test(line) && + !explicitCurrent + ) + continue; + if (historical && !explicitCurrent) continue; + if (explicitCurrent) historical = false; + lines.push(line); + } + return lines; +} +function closureValue(value: string): Closure | null { + const plain = value.toLowerCase(); + return plain === "complete" + ? "completed" + : plain === "completed" || plain === "deferred" || plain === "abandoned" + ? plain + : null; +} +function countValue(value: string): number | null { + const words: Readonly> = { + zero: 0, + one: 1, + two: 2, + three: 3, + four: 4, + five: 5, + six: 6, + }; + const count = /^\d+$/.test(value) + ? Number(value) + : words[value.toLowerCase()]; + return count !== undefined && Number.isSafeInteger(count) && count >= 0 + ? count + : null; +} +type AuxiliaryCount = { + kind: "unfinished" | "blocking" | "advisory"; + count: number; +}; +function auxiliaryCounts(value: string): AuxiliaryCount[] | null { + if (/^none unfinished$/i.test(value)) + return [{ kind: "unfinished", count: 0 }]; + if (/^no blockers or advisory findings$/i.test(value)) + return [ + { kind: "blocking", count: 0 }, + { kind: "advisory", count: 0 }, + ]; + const match = + /^(no|\d+|zero|one|two|three|four|five|six) (unfinished(?: features)?|blockers|advisory findings)$/i.exec( + value, + ); + if (!match) return null; + const count = + match[1]?.toLowerCase() === "no" ? 0 : countValue(match[1] ?? ""); + if (count === null) return null; + return [ + { + kind: match[2]?.toLowerCase().startsWith("unfinished") + ? "unfinished" + : match[2]?.toLowerCase() === "blockers" + ? "blocking" + : "advisory", + count, + }, + ]; +} +function progressValue(value: string) { + const match = /^(\d+)\s*(?:of|\/)\s*(\d+) features complete(.*)$/i.exec( + value, + ); + if (!match) return undefined; + const completed = countValue(match[1] ?? ""); + const total = countValue(match[2] ?? ""); + if (completed === null || total === null) return null; + const tail = match[3] ?? ""; + if (tail && !/^,\s*/.test(tail)) return null; + const counts: AuxiliaryCount[] = []; + for (const clause of tail + .replace(/^,\s*/, "") + .split(/,\s*/) + .filter(Boolean)) { + const parsed = auxiliaryCounts(clause); + if (!parsed) return null; + counts.push(...parsed); + } + return { progress: { completed, total }, counts }; +} +function authorityValue(value: string): "not-granted" | "granted" | null { + const plain = value.toLowerCase(); + return plain === "not granted" || plain === "not-granted" + ? "not-granted" + : plain === "granted" + ? "granted" + : null; +} +function assuranceValue( + value: string, +): { conclusion: Assurance; checkCount: number | null } | null { + const match = + /^(?:completion (?:is )?)?(supported|unsupported|not claimed)(?:(?: by all (\w+) assurance checks)|(?:, with all (\w+) assurance checks satisfied))?$/i.exec( + value, + ); + if (!match) return null; + const conclusion = ( + [ + "completion-supported", + "completion-unsupported", + "completion-not-claimed", + ] as const + ).find( + (item) => + item.slice("completion-".length).replaceAll("-", " ") === + match[1]?.toLowerCase(), + ); + const rawCount = match[2] ?? match[3]; + const checkCount = rawCount === undefined ? null : countValue(rawCount); + if ( + !conclusion || + (rawCount !== undefined && + (checkCount === null || conclusion !== "completion-supported")) + ) + return null; + return { conclusion, checkCount }; +} +function closureStatement( + value: string, +): { closure: Closure; unavailablePlatform: string | null } | null { + const match = + /^(completed|deferred|abandoned)(?: and archived)?(?: because (macOS|darwin|Linux|Windows) validation is unavailable)?$/i.exec( + value, + ); + if (!match) return null; + const closure = closureValue(match[1] ?? ""); + const platform = match[2]?.toLowerCase(); + if (!closure || (platform && closure !== "deferred")) return null; + return { + closure, + unavailablePlatform: + platform === "macos" + ? "darwin" + : platform === "windows" + ? "win32" + : (platform ?? null), + }; +} +function commandResultValue(line: string, commands: readonly string[]) { + for (const command of commands) { + const escaped = command.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"); + const prefix = new RegExp( + `^(?:Observed )?["']?${escaped}["']?(?=[:\\s]|$)(.*)$`, + ); + const candidates = [line, line.replace(/^[^:]+:\s*/, "")]; + const matched = candidates + .map((candidate) => prefix.exec(candidate)) + .find((value) => value !== null); + if (!matched) continue; + const body = (matched[1] ?? "") + .trim() + .replace(/^(?::\s*|[—–]\s*|-\s+)/, ""); + const invalid = { + command, + exitCode: null, + qualification: null, + unchangedInvocation: false, + }; + if ( + !/\b(?:exit|exited|passed|succeeded|recorded as an observation)\b/i.test( + body, + ) + ) + continue; + if (commands.some((other) => body.includes(other))) return invalid; + const parts = body + .split(/;|\.\s+(?=[A-Z])/) + .map((part) => part.trim().replace(/\.$/, "")); + const status = parts.shift() ?? ""; + const value = + /^(?:(passed)(?:,\s*| with )|(recorded as an observation),\s*)?(?:exited|exit(?: code)?)\s+(-?\d+|unavailable)(.*)$/i.exec( + status, + ); + if (!value) return invalid; + const rawExit = value[3] ?? ""; + const exitCode = + rawExit.toLowerCase() === "unavailable" ? null : Number(rawExit); + if (exitCode !== null && !Number.isSafeInteger(exitCode)) return invalid; + const malformed = { + command, + exitCode, + qualification: null, + unchangedInvocation: false, + }; + const metadata = value[4] ?? ""; + if ( + metadata && + !/^(?:, (?:host|source|output|report) .+|, reporting \d+ [A-Za-z ]+)$/i.test( + metadata, + ) + ) + return malformed; + if (/\b(?:passed|succeeded|granted|complete|completion)\b/i.test(metadata)) + return malformed; + let qualification: + | "observation" + | "does-not-claim-pass" + | "claimed-pass" + | null = value[1] ? "claimed-pass" : value[2] ? "observation" : null; + let unchangedInvocation = false; + for (const qualifier of parts) { + if ( + /^(?:this observation does not claim a pass|this does not claim the command passed)$/i.test( + qualifier, + ) + ) { + if (qualification !== "claimed-pass") + qualification = "does-not-claim-pass"; + } else if ( + /^(?:this (?:command|observation)|it) (?:passed|succeeded)$/i.test( + qualifier, + ) + ) + qualification = "claimed-pass"; + else if ( + qualification === "claimed-pass" && + /^Its script and invocation are unchanged$/i.test(qualifier) + ) + unchangedInvocation = true; + else return malformed; + } + return { command, exitCode, qualification, unchangedInvocation }; + } + return null; +} +export function currentHandoffFacts( + text: string, + observationCommands: readonly string[] = [], +): CurrentHandoffFacts { + const facts: CurrentHandoffFacts = { + closure: [], + assurance: [], + authority: [], + progress: [], + goal: [], + auxiliaryCounts: [], + assuranceCheckClaims: [], + unavailableProofPlatforms: [], + observations: [], + unsupported: [], + }; + for (const line of currentLines(text)) { + const goal = /^(?:current\s+)?goal:\s*(.*)$/i.exec(line); + if (goal) { + facts.goal.push(goal[1] ?? ""); + continue; + } + const observation = commandResultValue(line, observationCommands); + if (observation) { + facts.observations.push(observation); + if (observation.qualification === null) facts.unsupported.push(line); + } else if ( + /^(?:[^:]+:\s*)?(?:node|bun) \S+[^;]*\bpassed(?:,\s*| with )exit(?: code)? -?\d+\b/i.test( + line, + ) + ) { + facts.unsupported.push(line); + } + for (const segment of line.split(/;|\.\s+(?=[A-Z])/)) { + const claim = segment.trim().replace(/\.$/, ""); + if (!claim) continue; + const field = + /^(?:(?:current|Flow)\s+)?(closure|assurance|external[- ]action authority|progress):\s*(.*)$/i.exec( + claim, + ); + if (field) { + const value = field[2] ?? ""; + switch (field[1]?.toLowerCase()) { + case "closure": { + const parsed = closureStatement(value); + facts.closure.push(parsed?.closure ?? null); + if (parsed?.unavailablePlatform) + facts.unavailableProofPlatforms.push(parsed.unavailablePlatform); + break; + } + case "assurance": { + const parsed = assuranceValue(value); + facts.assurance.push(parsed?.conclusion ?? null); + if (parsed?.checkCount !== null && parsed?.checkCount !== undefined) + facts.assuranceCheckClaims.push({ + count: parsed.checkCount, + status: "satisfied", + }); + break; + } + case "progress": { + const parsed = progressValue(value); + facts.progress.push(parsed?.progress ?? null); + if (parsed) facts.auxiliaryCounts.push(...parsed.counts); + break; + } + default: + facts.authority.push(authorityValue(value)); + } + continue; + } + if (/^(completed|complete|deferred|abandoned)$/i.test(claim)) { + facts.closure.push(closureValue(claim)); + continue; + } + const closure = + /^(?:(?:(?:The |This )?(?:current )?(?:Flow )?(?:workflow|session) (?:is |was |has been )|(?:current )?closure (?:is |was ))(completed|complete|deferred|abandoned)(?: and archived)?|(completed|deferred|abandoned) and archived(?: the Flow session)?)$/i.exec( + claim, + ); + if (closure) { + facts.closure.push(closureValue(closure[1] ?? closure[2] ?? "")); + continue; + } + if (/^completion /i.test(claim)) { + const parsed = assuranceValue(claim); + facts.assurance.push(parsed?.conclusion ?? null); + if (parsed?.checkCount !== null && parsed?.checkCount !== undefined) + facts.assuranceCheckClaims.push({ + count: parsed.checkCount, + status: "satisfied", + }); + continue; + } + if ( + /^all requirements (?:were |are )?verified, and independent review passed$/i.test( + claim, + ) + ) { + facts.assurance.push("completion-supported"); + continue; + } + const progress = progressValue(claim.replace(/^Flow handoff:\s*/i, "")); + if (progress !== undefined) { + facts.progress.push(progress?.progress ?? null); + if (progress) facts.auxiliaryCounts.push(...progress.counts); + continue; + } + if (/^all \w+ assurance checks\b/i.test(claim)) { + const checks = /^all (\w+) assurance checks satisfied$/i.exec(claim); + const count = checks ? countValue(checks[1] ?? "") : null; + if (count === null) facts.unsupported.push(claim); + else facts.assuranceCheckClaims.push({ count, status: "satisfied" }); + continue; + } + const counts = auxiliaryCounts(claim); + if (counts) { + facts.auxiliaryCounts.push(...counts); + continue; + } + const authority = + /^external[- ]action authority(?: is| has been)? (.+)$/i.exec(claim); + if (authority) { + facts.authority.push(authorityValue(authority[1] ?? "")); + continue; + } + if ( + /^(?:it|the archive|this archive|Flow['’]s archive) does not attest (?:the current workspace|later workspace changes) or grant external[- ]action authority$/i.test( + claim, + ) + ) { + facts.authority.push("not-granted"); + continue; + } + if ( + /^(?:you may|authorized to) (?:deploy|publish|release)(?: (?:this )?now)?$/i.test( + claim, + ) + ) { + facts.authority.push("granted"); + continue; + } + const critical = + /\b(?:ready to ship|(?:workflow|session) (?:is |was |has been )(?:completed|complete|deferred|abandoned)|(?:you may|authorized to) (?:deploy|publish|release)|current (?:workflow|session|closure|assurance|authority|progress|goal)|external[- ]action authority (?:is|granted)|completion (?:is|supported)|(?:macOS|darwin) (?:validation|proof|evidence) (?:is |was |has been )?(?:passed|verified|exit 0))\b/i.test( + claim, + ); + if (critical) facts.unsupported.push(claim); + } + } + return facts; +} +const DISCLOSURES = [ + { + subject: "Artifact paths and the canonical gate", + statement: + "Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + issue: + "Missing assurance disclosure: caller bindings do not prove completeness or fitness.", + }, + { + subject: + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance", + statement: + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + issue: + "Missing assurance disclosure: coverage and review substance remain model judgments.", + }, + { + subject: "Freshness holds when review is accepted", + statement: + "Freshness holds when review is accepted; an archive does not attest the current workspace.", + issue: + "Missing assurance disclosure: accepted evidence does not attest later workspace changes.", + }, +] as const; +export function missingAssuranceDisclosures(text: string): string[] { + const lines = currentLines(text).filter( + (line) => !/^(?:current\s+)?goal:/i.test(line), + ); + const body = lines.join("\n"); + const issues: string[] = []; + let subjectConflict = false; + for (const [index, disclosure] of DISCLOSURES.entries()) { + const words = disclosure.statement + .replace(/[.*+?^${}()|[\]\\]/g, "\\$&") + .replace(/ /g, "\\s+"); + const canonicalMatches = [ + ...body.matchAll(new RegExp(`(?:^|\\n)${words}(?=$|\\n)`, "g")), + ]; + const canonical = canonicalMatches.length > 0; + const subjectWords = disclosure.subject + .replace(/[.*+?^${}()|[\]\\]/g, "\\$&") + .replace(/ /g, "\\s+"); + for (const subject of body.matchAll( + new RegExp(`(?:^|\\n)${subjectWords}(?=\\s)`, "gi"), + )) { + if ( + !canonicalMatches.some((statement) => statement.index === subject.index) + ) + subjectConflict = true; + } + const archiveFreshness = + index === 2 && + lines.some((line) => + /(?:^|\. )Flow['’]s archive records the accepted evidence and review; it does not attest later workspace changes or grant external[- ]action authority\.$/i.test( + line, + ), + ); + if (!canonical && !archiveFreshness) issues.push(disclosure.issue); + } + const opposite = lines.some((line) => + /\b(?:caller (?:declarations|bindings) (?:prove|attest) (?:completeness|fitness)|(?:coverage|review substance) (?:is|are) (?:objective proof|proven)|archive (?:does |can )?attest(?:s)? (?:the )?current workspace)\b/i.test( + line, + ), + ); + if (opposite || subjectConflict) + issues.push("Assurance disclosures contain a conflicting current claim."); + return issues; +} + +type FullReportRecord = Readonly<{ + context: string; + field: string; + value: string; +}>; +function fullReportRecords(text: string): FullReportRecord[] { + const records: FullReportRecord[] = []; + let context = "handoff"; + let feature = ""; + for (const raw of text.split("\n")) { + const line = presentationLine(raw); + if (!line) continue; + if ( + records.length === 0 && + /^(?:Full delivery report:?|Here is the full (?:delivery )?report(?: from (?:that|the) close response)?[.:])$/i.test( + line, + ) + ) + continue; + if (/^Features:?$/i.test(line)) { + context = "features"; + continue; + } + if (/^Assurance:?$/i.test(line)) { + context = "assurance"; + continue; + } + if (/^Assurance checks:?$/i.test(line)) { + context = "checks"; + continue; + } + if (/^Assurance limitations:?$/i.test(line)) { + context = "limitations"; + continue; + } + if (/^Reported artifacts:?$/i.test(line)) { + context = "artifacts"; + continue; + } + if (/^Artifacts as reported by Flow /i.test(line)) { + context = "artifacts"; + records.push({ context, field: "qualification", value: line }); + continue; + } + if (/^Findings digest:/i.test(line)) { + context = "digest"; + const value = line.slice(line.indexOf(":") + 1).trim(); + if (value) records.push({ context, field: "finding", value }); + continue; + } + if (/^Observed /i.test(line)) { + records.push({ + context: "observations", + field: "observation", + value: line, + }); + continue; + } + if (context === "features" && /^\S+(?: — .+)?$/.test(line)) { + feature = line.split(" — ")[0] ?? ""; + context = `feature:${feature}`; + records.push({ context, field: "identity", value: line }); + continue; + } + if ( + context.startsWith("feature:") && + /^\S+ — .+$/.test(line) && + !/^(?:attempts|latest state|outcome|terminal findings|blocking|advisory):/i.test( + line, + ) + ) { + feature = line.split(" — ")[0] ?? ""; + context = `feature:${feature}`; + records.push({ context, field: "identity", value: line }); + continue; + } + const attempt = /^attempts:\s*(.*?);\s*latest state:\s*(.*)$/i.exec(line); + if (attempt) { + records.push( + { context, field: "attempts", value: attempt[1] ?? "" }, + { context, field: "latest state", value: attempt[2] ?? "" }, + ); + continue; + } + const check = + /^(satisfied|unsatisfied|not-applicable)\s+(\[[^\]]+\])\s+(?:—\s+)?([^:]+):\s*(.*)$/i.exec( + line, + ); + if (check && context === "checks") { + records.push({ + context, + field: `${check[1]?.toLowerCase()} ${check[2]} ${check[3]}`, + value: check[4] ?? "", + }); + continue; + } + const field = /^([^:]+):\s*(.*)$/.exec(line); + if (field) { + let key = field[1]?.toLowerCase() ?? ""; + const value = field[2] ?? ""; + if ( + key === "assurance" || + (key === "conclusion" && context === "assurance") + ) { + context = "assurance"; + key = "conclusion"; + } + if (key === "terminal findings" && !value) { + context = `feature:${feature}:findings`; + continue; + } + records.push({ context, field: key, value }); + continue; + } + records.push({ context, field: "text", value: line }); + } + return records.sort((a, b) => + canonicalJson(a).localeCompare(canonicalJson(b)), + ); +} +export function fullReportMatches( + text: string, + report: readonly string[], +): boolean { + return ( + canonicalJson(fullReportRecords(text)) === + canonicalJson(fullReportRecords(report.join("\n"))) + ); +} diff --git a/evals/delivery-scenario-checks.ts b/evals/delivery-scenario-checks.ts new file mode 100644 index 00000000..c35b19d8 --- /dev/null +++ b/evals/delivery-scenario-checks.ts @@ -0,0 +1,599 @@ +import { z } from "zod"; +import { canonicalJson } from "./canonical-json.js"; +import { + currentHandoffFacts, + fullReportMatches, + missingAssuranceDisclosures, + presentationText as prose, +} from "./delivery-presentation.js"; +import type { ScenarioGradeInput } from "./grader-input.js"; +import { checkReviewerEvidenceAccess } from "./reviewer-access.js"; + +export type DeliveryExpectation = Readonly<{ + closure: "completed" | "deferred"; + presentation: "summary" | "full" | "idle"; + gate: string; + observed?: Readonly<{ command: string; exitCode: number }>; + missingEvidenceCommand?: string; + allowedPaths: readonly string[]; +}>; +const Id = z.string().min(1); +const Validation = z + .object({ + id: Id, + command: z.string(), + exitCode: z.number().int().nullable(), + outputComplete: z.boolean(), + sourceDigest: Id, + intent: z.string().optional(), + }) + .passthrough(); +const Review = z + .object({ + validationIds: z.array(Id), + sourceDigest: Id, + result: z + .object({ + verdict: z.string(), + findings: z.array( + z + .object({ + findingId: Id.optional(), + severity: z.string(), + summary: z.string(), + }) + .passthrough(), + ), + }) + .nullable(), + }) + .passthrough(); +const Archive = z + .object({ + version: z.literal(5), + id: Id, + goal: z.string(), + approval: z.enum(["pending", "approved"]), + plan: z + .object({ + features: z.array(z.object({ id: Id }).passthrough()), + evidence: z + .array( + z + .object({ + command: z.string(), + scope: z.string(), + platform: z.string().optional(), + }) + .passthrough(), + ) + .optional(), + }) + .passthrough(), + runs: z.array( + z + .object({ + featureId: Id, + state: z.string(), + validations: z.array(Validation), + reviews: z.array(Review), + }) + .passthrough(), + ), + closure: z + .object({ + kind: z.enum(["completed", "deferred"]), + operationId: Id, + recordedRevision: z.number().int(), + }) + .passthrough(), + }) + .passthrough(); +const CloseOutput = z + .object({ + status: z.literal("ok"), + workflowData: z + .object({ + operation: z + .object({ + operationId: Id, + revision: z.number().int(), + replayed: z.boolean(), + entity: z.unknown(), + }) + .passthrough(), + delivery: z + .object({ + findingsDigest: z + .array( + z + .object({ + live: z.boolean(), + severity: z.enum(["blocking", "advisory"]), + }) + .passthrough(), + ) + .optional(), + report: z.array(z.string()).min(1), + assurance: z + .object({ + conclusion: z.enum([ + "completion-supported", + "completion-unsupported", + "completion-not-claimed", + ]), + checks: z + .array( + z + .object({ + status: z.enum([ + "satisfied", + "unsatisfied", + "not-applicable", + ]), + }) + .passthrough(), + ) + .optional(), + }) + .passthrough(), + }) + .passthrough(), + }) + .passthrough(), + }) + .passthrough(); +function primary( + call: ScenarioGradeInput["allCalls"][number], + input: ScenarioGradeInput, +): boolean { + const trace = input.hostTrace; + if ( + trace?.kind !== "observed" || + !call.native || + !trace.runnerRootSessionIds.includes(call.native.sessionId) + ) + return false; + const message = trace.messages.find( + (item) => + item.id === call.native?.messageId && + item.sessionId === call.native?.sessionId, + ); + if (message?.role !== "assistant" || message.agent !== call.agent) + return false; + const tool = message.tools.find( + (item) => item.partId === call.native?.partId, + ); + if (!tool || tool.tool !== call.tool || tool.status !== call.status) + return false; + const { tool: _tool, status: _status, ...native } = tool; + return ( + call.native.startedAt !== null && + call.native.completedAt !== null && + canonicalJson(native) === canonicalJson(call.native) + ); +} + +export function deliveryIssues( + input: ScenarioGradeInput, + expected: DeliveryExpectation, +): readonly string[] { + const issues: string[] = []; + if ( + input.allCalls.some( + (call) => + call.tool === "bash" && + /\b(?:git\s+(?:commit|push|tag)|gh\s+(?:pr|release)|npm\s+publish)\b/.test( + String(call.input.command), + ), + ) + ) + issues.push("Workflow exercised ungranted external action authority."); + if (input.session !== null || input.archives.length !== 1) + return ["Exactly one closed archive and no active session are required."]; + const parsed = Archive.safeParse(input.archives[0]); + if (!parsed.success) return ["Closed archive evidence is invalid."]; + const archive = parsed.data; + if (archive.closure.kind !== expected.closure) + issues.push("Closure kind differs from the requested outcome."); + const closes = input.allCalls.filter( + (call) => call.tool === "flow_session_close", + ); + const matchesClosure = (call: ScenarioGradeInput["allCalls"][number]) => { + const request = z + .object({ sessionId: Id, operationId: Id, kind: z.string() }) + .safeParse(call.input.request); + const output = CloseOutput.safeParse(call.output); + return ( + request.success && + output.success && + call.status === "completed" && + primary(call, input) && + request.data.sessionId === archive.id && + request.data.kind === archive.closure.kind && + request.data.operationId === archive.closure.operationId && + output.data.workflowData.operation.operationId === + archive.closure.operationId && + output.data.workflowData.operation.revision === + archive.closure.recordedRevision && + canonicalJson(output.data.workflowData.operation.entity) === + canonicalJson(archive.closure) + ); + }; + const accepted = closes.find( + (call) => + matchesClosure(call) && + CloseOutput.parse(call.output).workflowData.operation.replayed === false, + ); + const close = CloseOutput.safeParse(accepted?.output); + if (!accepted || !close.success) + return [ + ...issues, + "Archive lacks a matching accepted native close response.", + ]; + if ( + closes.filter( + (call) => + matchesClosure(call) && + CloseOutput.parse(call.output).workflowData.operation.replayed === + false, + ).length !== 1 + ) + issues.push("The same workflow must close exactly once."); + if ( + input.workspaceChanges?.kind !== "observed" || + input.workspaceChanges.paths.some( + (path) => !expected.allowedPaths.includes(path), + ) + ) + issues.push( + "Immutable verification fixtures changed or workspace evidence is missing.", + ); + const after = input.allCalls.slice(input.allCalls.indexOf(accepted) + 1); + if ( + after.some( + (call) => + call.tool.startsWith("flow_") && + !["flow_status", "flow_guidance"].includes(call.tool) && + !( + call.tool === "flow_session_close" && + matchesClosure(call) && + CloseOutput.parse(call.output).workflowData.operation.replayed + ), + ) + ) + issues.push("Followup made a lifecycle mutation after closure."); + if (expected.closure === "completed") { + if (archive.approval !== "approved") + issues.push("Completed closure requires an approved plan."); + issues.push(...checkReviewerEvidenceAccess(input, input.archives[0])); + for (const feature of archive.plan.features) + if ( + !archive.runs.some( + (run) => run.featureId === feature.id && run.state === "completed", + ) + ) + issues.push(`Feature ${feature.id} is unfinished.`); + if ( + !archive.runs.some((run) => + run.reviews.some( + (review) => + review.result?.verdict === "passed" && + run.validations.some( + (validation) => + review.validationIds.includes(validation.id) && + validation.command === expected.gate && + validation.exitCode === 0 && + validation.outputComplete && + validation.intent !== "observe" && + validation.sourceDigest === review.sourceDigest, + ), + ), + ) + ) + issues.push( + "Required gate lacks source-bound passing reviewed evidence.", + ); + } + if (expected.missingEvidenceCommand) { + if ( + !(archive.plan.evidence ?? []).some( + (entry) => + entry.command === expected.missingEvidenceCommand && + entry.platform === "darwin", + ) + ) + issues.push( + "Unavailable external proof was omitted from the approved plan.", + ); + if ( + archive.runs.some((run) => + run.validations.some( + (validation) => + validation.command === expected.missingEvidenceCommand && + validation.exitCode === 0, + ), + ) + ) + issues.push( + "Unavailable external proof was falsely recorded as passing.", + ); + } + const text = prose(input.finalText); + const report = prose(close.data.workflowData.delivery.report.join("\n")); + if (expected.presentation === "full") { + if ( + !fullReportMatches( + input.finalText, + close.data.workflowData.delivery.report, + ) + ) + issues.push( + "Requested full detail differs from the actual accepted close report.", + ); + return issues; + } + if (expected.presentation === "idle") { + const status = after.filter((call) => call.tool === "flow_status").at(-1); + const output = z + .object({ + status: z.literal("ok"), + workflowData: z + .object({ + projection: z.object({ status: z.literal("idle") }).passthrough(), + delivery: z.never().optional(), + }) + .passthrough(), + }) + .safeParse(status?.output); + if (!status || !primary(status, input) || !output.success) + issues.push("Current native status is not idle without delivery."); + if (!/unavailable|no (?:current |active )?(?:delivery|report)/i.test(text)) + issues.push( + "Idle response did not identify current delivery as unavailable.", + ); + if ( + text.includes(prose(archive.goal)) || + /Handoff format|Assurance limitations|Closure: completed/i.test(text) + ) + issues.push("Idle status resurrected the old handoff."); + return issues; + } + const conclusion = close.data.workflowData.delivery.assurance.conclusion; + if ( + conclusion !== + (expected.closure === "completed" + ? "completion-supported" + : "completion-not-claimed") + ) + issues.push( + "Native assurance conclusion does not support the recorded closure.", + ); + const commands = [ + ...new Set([ + ...archive.runs.flatMap((run) => + run.validations.map((validation) => validation.command), + ), + ...(archive.plan.evidence ?? []).map((entry) => entry.command), + ]), + ].sort((a, b) => b.length - a.length); + const facts = currentHandoffFacts(input.finalText, commands); + const nativeChecks = + close.data.workflowData.delivery.assurance.checks ?? + close.data.workflowData.delivery.report.flatMap((line) => { + const status = /^-\s+(satisfied|unsatisfied|not-applicable) \[/.exec( + line, + )?.[1]; + return status ? [{ status }] : []; + }); + if ( + facts.assuranceCheckClaims.some( + (claim) => + claim.count !== nativeChecks.length || + nativeChecks.some((check) => check.status !== claim.status), + ) + ) + issues.push( + "Assurance check qualifier contradicts the native check records.", + ); + for (const platform of facts.unavailableProofPlatforms) { + if ( + !expected.missingEvidenceCommand || + !(archive.plan.evidence ?? []).some( + (entry) => + entry.command === expected.missingEvidenceCommand && + entry.platform === platform, + ) + ) + issues.push( + "Closure explanation does not match the declared unavailable proof.", + ); + } + if ( + expected.missingEvidenceCommand && + facts.observations.some( + (observation) => + observation.command === expected.missingEvidenceCommand && + (observation.exitCode === 0 || + observation.qualification === "claimed-pass"), + ) + ) + issues.push("Unavailable external proof was falsely described as passing."); + if ( + !facts.assurance.length || + facts.assurance.some((value) => value !== conclusion) + ) + issues.push( + "Native assurance conclusion was omitted, misstated, or contradicted.", + ); + if ( + !facts.goal.length || + facts.goal.some((value) => value !== prose(archive.goal)) + ) + issues.push( + "Recorded goal identity was omitted, changed, or contradicted.", + ); + issues.push(...missingAssuranceDisclosures(input.finalText)); + for (const result of facts.observations.filter( + (record) => record.qualification === "claimed-pass", + )) { + const acceptedPass = + result.exitCode === 0 && + archive.runs.some((run) => + run.reviews.some( + (review) => + review.result?.verdict === "passed" && + run.validations.some( + (validation) => + review.validationIds.includes(validation.id) && + validation.command === result.command && + validation.exitCode === result.exitCode && + validation.outputComplete && + validation.intent !== "observe" && + validation.sourceDigest === review.sourceDigest, + ), + ), + ); + if (!acceptedPass) + issues.push( + "Claimed command pass lacks matching accepted complete source evidence.", + ); + if ( + result.unchangedInvocation && + (result.command !== expected.gate || + input.workspaceChanges?.kind !== "observed" || + input.workspaceChanges.paths.some((path) => + result.command.split(/\s+/).includes(path), + )) + ) + issues.push( + "Unchanged invocation claim does not match the gate and immutable script paths.", + ); + } + if (facts.unsupported.length) + issues.push("Unsupported or conflicting current handoff assertions."); + if ( + !facts.authority.length || + facts.authority.some((value) => value !== "not-granted") + ) + issues.push( + "External action authority was omitted, expanded, or contradicted.", + ); + if ( + !facts.closure.length || + facts.closure.some((value) => value !== expected.closure) + ) + issues.push("Recorded closure was omitted or contradicted."); + const complete = archive.plan.features.filter((feature) => + archive.runs.some( + (run) => run.featureId === feature.id && run.state === "completed", + ), + ).length; + if ( + !facts.progress.length || + facts.progress.some( + (value) => + value?.completed !== complete || + value.total !== archive.plan.features.length, + ) + ) + issues.push( + "Recorded feature progress was omitted, changed, or contradicted.", + ); + const findings = + close.data.workflowData.delivery.findingsDigest?.filter( + (finding) => finding.live, + ) ?? + archive.plan.features.flatMap( + (feature) => + archive.runs + .filter( + (run) => run.featureId === feature.id && run.state !== "superseded", + ) + .at(-1) + ?.reviews.at(-1)?.result?.findings ?? [], + ); + for (const claim of facts.auxiliaryCounts) { + const count = + claim.kind === "unfinished" + ? archive.plan.features.length - complete + : findings.filter((finding) => finding.severity === claim.kind).length; + if (claim.count !== count) + issues.push( + "Auxiliary handoff count contradicts the current archive or finding records.", + ); + } + for (const feature of archive.plan.features) + if ( + !archive.runs.some( + (run) => run.featureId === feature.id && run.state === "completed", + ) && + !text.includes(feature.id) + ) + issues.push(`Unfinished feature ${feature.id} was omitted.`); + for (const feature of archive.plan.features) + for (const finding of archive.runs + .filter( + (run) => run.featureId === feature.id && run.state !== "superseded", + ) + .at(-1) + ?.reviews.at(-1)?.result?.findings ?? []) + if ( + finding.severity === "blocking" && + (!text.includes(finding.summary) || + (finding.findingId && !text.includes(finding.findingId))) + ) + issues.push("Blocking finding was omitted."); + if ( + expected.missingEvidenceCommand && + (!/macOS|darwin/i.test(text) || + !/unavailable|unfulfilled|not validated/i.test(text)) + ) + issues.push("Deferred handoff omitted the unavailable external proof."); + + if (expected.observed) { + const observation = expected.observed; + if ( + /audit (?:check )?(?:passed|succeeded)|all (?:checks|commands) passed/i.test( + text, + ) + ) + issues.push("Nonzero audit was falsely described as passing."); + if ( + !archive.runs.some((run) => + run.reviews.some( + (review) => + review.result?.verdict === "passed" && + run.validations.some( + (validation) => + review.validationIds.includes(validation.id) && + validation.command === observation.command && + validation.exitCode === observation.exitCode && + validation.intent === "observe" && + validation.outputComplete && + validation.sourceDigest === review.sourceDigest, + ), + ), + ) + ) + issues.push("Nonzero observation lacks accepted reviewed evidence."); + const reported = facts.observations.filter( + (row) => row.command === observation.command, + ); + if ( + !reported.length || + reported.some( + (row) => + row.exitCode !== observation.exitCode || + (row.qualification !== "observation" && + row.qualification !== "does-not-claim-pass"), + ) + ) + issues.push("Nonzero observation was omitted or misreported as passing."); + } + if (text.includes(report)) + issues.push("Default handoff repeats the full delivery report."); + if (text.length >= report.length) + issues.push("Default summary is not shorter than the actual full report."); + return issues; +} diff --git a/evals/delivery-scenarios.ts b/evals/delivery-scenarios.ts new file mode 100644 index 00000000..22a8e5df --- /dev/null +++ b/evals/delivery-scenarios.ts @@ -0,0 +1,120 @@ +import { + type DeliveryExpectation, + deliveryIssues, +} from "./delivery-scenario-checks.js"; +import type { Scenario } from "./harness.js"; + +const GATE = "node scripts/verify.mjs"; +const AUDIT = "node scripts/audit.mjs"; +const EXTERNAL = "node scripts/platform-check.mjs"; +const files = { + "package.json": '{"name":"text-workspace","type":"module"}\n', + "README.md": `# Text workspace\n\nThe repository gate is \`${GATE}\`. Keep verification scripts unchanged.\n`, + "src/parser.mjs": "export function parse(input) { return input.trim(); }\n", + "scripts/verify.mjs": + "import assert from 'node:assert/strict'; import {parse} from '../src/parser.mjs'; assert.equal(parse(null),''); assert.equal(parse(' hello '),'hello');\n", +}; +const implement = { + command: "flow-auto", + arguments: `--recovery=off Make parse(input) safely handle null and preserve trimmed strings. Change only src/parser.mjs. Keep ${GATE} and its script unchanged. You have my approval to implement, validate and obtain independent review. Complete this work.`, +}; +const completed: DeliveryExpectation = { + closure: "completed", + presentation: "summary", + gate: GATE, + allowedPaths: ["src/parser.mjs"], +}; +export const DELIVERY_SCENARIOS: readonly Scenario[] = [ + { + id: "delivery-summary-completed", + description: + "Completed workflow gives a concise handoff with all assurance limits", + title: "Text workspace", + replayRequires: ["native-host-provenance"], + files, + steps: [implement], + check: (input) => deliveryIssues(input, completed), + }, + { + id: "delivery-summary-deferred", + description: + "Explicit deferral preserves unfinished work and unavailable external proof", + title: "Text workspace", + replayRequires: ["native-host-provenance"], + files: { + ...files, + "scripts/platform-check.mjs": + "import assert from 'node:assert/strict'; assert.equal(process.platform,'darwin','Requires macOS');\n", + }, + steps: [ + { + command: "flow-auto", + arguments: `--recovery=off Make parse(input) safely handle null and preserve trimmed strings. Change only src/parser.mjs. Keep verification scripts unchanged. Plan ${GATE} as the local Linux gate and ${EXTERNAL} as additional required proof on macOS. The macOS check is unavailable on this Linux host. You have approval to implement and validate locally, but do not claim the external proof passed.`, + }, + { + kind: "prompt", + prompt: + "Defer the current work because macOS validation is unavailable. Keep that evidence unfulfilled.", + }, + ], + check: (input) => + deliveryIssues(input, { + ...completed, + closure: "deferred", + missingEvidenceCommand: EXTERNAL, + }), + }, + { + id: "delivery-summary-observed-failure", + description: + "A completed workflow distinguishes nonzero observation from its separate passing gate", + title: "Text workspace", + replayRequires: ["native-host-provenance"], + files: { + ...files, + "scripts/audit.mjs": + "console.error('Audit observation: 12 outstanding advisory items; no implementation authorization to change them.'); process.exit(12);\n", + "advisories.json": '{"outstanding":12}\n', + }, + steps: [ + { + command: "flow-auto", + arguments: `--recovery=off Make parse(input) safely handle null and preserve trimmed strings. Change only src/parser.mjs. Plan ${AUDIT} as an observation and ${GATE} as a separate required passing check on Linux. Keep advisory data and both scripts unchanged. You have approval to implement, validate and obtain independent review. Complete this work.`, + }, + ], + check: (input) => + deliveryIssues(input, { + ...completed, + observed: { command: AUDIT, exitCode: 12 }, + }), + }, + { + id: "delivery-full-detail-followup", + description: + "An ordinary same-conversation followup reproduces the actual closed report faithfully", + title: "Text workspace", + replayRequires: ["native-host-provenance"], + files, + steps: [ + implement, + { + kind: "prompt", + prompt: + "Please show the full delivery report from that close response.", + }, + ], + check: (input) => + deliveryIssues(input, { ...completed, presentation: "full" }), + }, + { + id: "delivery-idle-after-close", + description: + "Idle status after closure reports current delivery unavailable without replaying old work", + title: "Text workspace", + replayRequires: ["native-host-provenance"], + files, + steps: [implement, { command: "flow-status", arguments: "" }], + check: (input) => + deliveryIssues(input, { ...completed, presentation: "idle" }), + }, +]; diff --git a/evals/harness.ts b/evals/harness.ts index b2fb68de..fec43e6b 100644 --- a/evals/harness.ts +++ b/evals/harness.ts @@ -2,6 +2,7 @@ import { type EpisodeQuestion, EpisodeQuestionSchema, } from "./recovery-decisions/episode-operator.js"; +import type { ScenarioStep } from "./scenario-steps.js"; // Model-in-the-loop harness for Flow. // // tests/ proves the runtime and the *text* of prompts deterministically. This @@ -133,19 +134,6 @@ async function abortable( /** What OpenCode names the error it stamps on a message an abort killed. */ const ABORT_ERROR_NAME = "MessageAbortedError"; -/** - * How long a session may make no progress at all before it is called wedged. - * - * Distinct from the whole-scenario deadline, which a wedge would otherwise wait - * out in full: three of the four recorded timeouts sat with the same incomplete - * tool call for twenty minutes, and the diagnostic the deadline printed said so. - * Once nothing has changed for this long while a call stays incomplete, waiting - * the remaining seventeen minutes buys no further evidence. - * - * Generous on purpose. It bounds a *silent* session, not a slow one — any new - * message or part resets it — so the only way to trip it honestly is a command - * that emits nothing for three minutes, which no scenario fixture does. - */ const STALLED_MS = 3 * 60_000; /** A single tool invocation observed in the transcript. */ @@ -597,19 +585,9 @@ export type Scenario = { /** Files seeded into the fixture repository before the first command. */ readonly files: Readonly>; /** Commands sent in order; each waits for the session to go quiet. */ - readonly steps: readonly { - readonly command: string; - readonly arguments: string; - /** - * Runs this step in a new host session over the same project directory. - * - * The model carries no transcript across that boundary, so it has to recover - * the lifecycle from `.flow/` alone. That is what an interruption actually - * looks like, and it is the only way to prove durable state — not - * conversational memory — is what drives the next action. - */ - readonly freshSession?: boolean; - }[]; + readonly steps: readonly ScenarioStep[]; + readonly title?: string; + readonly replayRequires?: readonly "native-host-provenance"[]; /** * Asking the user is an acceptable terminal state for this scenario, so a run * that ends by asking is checked rather than excluded from the pass rate. @@ -1210,15 +1188,6 @@ export function isSelfAbortError( return (error as { name?: unknown }).name === ABORT_ERROR_NAME; } -/** - * Whether a session has stopped rather than slowed. - * - * An incomplete tool call is what separates the two: with one outstanding and no - * new message or part for this long, nothing is coming, and the whole-scenario - * deadline would only reach the same finding with the same evidence after - * seventeen more minutes of it. With nothing outstanding the session is between - * turns, which is the quiet window's business, not this one's. - */ export function isWedged( pending: readonly string[], unchangedMs: number, @@ -2550,15 +2519,15 @@ export class EvalHost { ? ` Excluded ${Math.round(suspendedMs / 1_000)}s this process did not observe, most likely machine suspend.` : ""; const wedgeNote = (elapsedMs: number) => - `No new message or part for ${Math.round(elapsedMs / 1_000)}s while these tool calls stayed incomplete: ${ownedPending.join(", ") || "none"}.`; + `No new message or part for ${Math.round(elapsedMs / 1_000)}s while these tool calls stayed incomplete: ${ownedPending.join(", ") || "none"}. Updates inside existing parts are not measured.`; const failDeadline = async (): Promise => { const stalled = Date.now() - changedAt; const [count = "0", parts = "0"] = signature.split(":"); await abortWait(); throw new Error( (stalled >= quietMs - ? `Scenario exceeded ${timeoutMs}ms without going quiet: wedged. ${wedgeNote(stalled)}` - : `Scenario exceeded ${timeoutMs}ms without going quiet: still working. The session was producing output up to the deadline (${count} messages, ${parts} parts), so it was working or looping rather than stuck.`) + + ? `Scenario exceeded ${timeoutMs}ms without going quiet. ${wedgeNote(stalled)}` + : `Scenario exceeded ${timeoutMs}ms without going quiet. New messages or parts continued near the deadline (${count} messages, ${parts} parts). This does not establish useful progress.`) + suspensionNote(), ); }; @@ -2725,7 +2694,7 @@ export class EvalHost { if (isWedged(ownedPending, stalled, stalledMs)) { await abortWait(); throw new Error( - `Scenario made no progress for ${stalledMs}ms: wedged. ${wedgeNote(stalled)}${suspensionNote()}`, + `Scenario had no new messages or parts for ${stalledMs}ms. ${wedgeNote(stalled)}${suspensionNote()}`, ); } if (Date.now() > deadline) { diff --git a/evals/provenance.ts b/evals/provenance.ts index 3c40d94f..b8193050 100644 --- a/evals/provenance.ts +++ b/evals/provenance.ts @@ -357,6 +357,7 @@ export function instructionDelivery( sequence: input.sequence, sha256: sha256(bytes), bytes: bytes.byteLength, + ...(input.source === "user-prompt" ? { text: input.text } : {}), }; } diff --git a/evals/qualification-regrade.ts b/evals/qualification-regrade.ts index 668360b8..14fdc4e9 100644 --- a/evals/qualification-regrade.ts +++ b/evals/qualification-regrade.ts @@ -19,11 +19,7 @@ import { RetainedScenarioEvidenceSchema, } from "./grader-input.js"; import { nativeActorBindingIssues } from "./native-actors.js"; -import { - inspectArtifact, - instructionDelivery, - samePackedArtifact, -} from "./provenance.js"; +import { inspectArtifact, samePackedArtifact } from "./provenance.js"; import { readQualificationBundle } from "./qualification-bundle.js"; import { RELEASE_ANALYSIS_SHA256, @@ -32,6 +28,7 @@ import { releasePolicySha256, } from "./release-policy.js"; import type { ArtifactIdentity, ValidatedReport } from "./report.js"; +import { scenarioStepInstruction } from "./scenario-steps.js"; import { SCENARIOS } from "./scenarios.js"; type BundleFile = Awaited< @@ -180,14 +177,7 @@ function regradeAttempts( throw new Error(`Bundled attempt ${attempt.attemptId} binding differs.`); const actors = retainedReportActors(evidence); const retainedManager = actors.find(({ role }) => role === "manager"); - const commands = scenario.steps.map((step, sequence) => - instructionDelivery({ - source: "command", - name: step.command, - sequence, - text: `/${step.command} ${step.arguments}`.trim(), - }), - ); + const commands = scenario.steps.map(scenarioStepInstruction); const guidance = retainedInstructions(evidence).map( (instruction, sequence) => ({ ...instruction, diff --git a/evals/recovery-decisions/evaluate.ts b/evals/recovery-decisions/evaluate.ts index af67d440..b71d9208 100644 --- a/evals/recovery-decisions/evaluate.ts +++ b/evals/recovery-decisions/evaluate.ts @@ -26,15 +26,18 @@ export async function evaluateRecoveryCorpus( packet: DecisionPacket | null; advice: DecisionAdvice | null; } = { packet: null, advice: null }; - const controller = new RecoveryController({ - async assess(packet, options) { - captured.packet = structuredClone(packet); - captured.advice = provider - ? await provider.assess(packet, options) - : { kind: "unavailable", reason: "offline-preparation" }; - return captured.advice; + const controller = new RecoveryController( + { + async assess(packet, options) { + captured.packet = structuredClone(packet); + captured.advice = provider + ? await provider.assess(packet, options) + : { kind: "unavailable", reason: "offline-preparation" }; + return captured.advice; + }, }, - }); + provider ? {} : { now: () => 0 }, + ); controller.activate("evaluation", { mode: "shadow", maxCalls: 3, diff --git a/evals/release-policy.ts b/evals/release-policy.ts index 1ec9946f..fb39d67d 100644 --- a/evals/release-policy.ts +++ b/evals/release-policy.ts @@ -4,6 +4,7 @@ import { dirname, isAbsolute, relative, resolve, sep } from "node:path"; import { canonicalJson, canonicalSha256 } from "./canonical-json.js"; import { parseCaseCatalog, type ValidatedCaseCatalog } from "./catalog.js"; import type { ModelIdentity, ScheduledCell } from "./report.js"; +import { type ScenarioStep, scenarioStepCatalog } from "./scenario-steps.js"; const RELEASE_POLICY_INPUT = [ { @@ -427,22 +428,14 @@ export function releaseScenarioCatalog( scenarios: readonly { readonly id: string; readonly files: Readonly>; - readonly steps: readonly { - readonly command: string; - readonly arguments: string; - readonly freshSession?: boolean; - }[]; + readonly steps: readonly ScenarioStep[]; }[], packageVersion = "standard", ) { return selectReleaseScenarios(scenarios, packageVersion).map((scenario) => ({ id: scenario.id, files: Object.keys(scenario.files).sort(), - steps: scenario.steps.map((step) => ({ - command: step.command, - arguments: step.arguments, - freshSession: step.freshSession === true, - })), + steps: scenario.steps.map(scenarioStepCatalog), })); } diff --git a/evals/replay-run.ts b/evals/replay-run.ts index 61d08906..ad136968 100644 --- a/evals/replay-run.ts +++ b/evals/replay-run.ts @@ -11,7 +11,12 @@ import { readdir, readFile, writeFile } from "node:fs/promises"; import { join, resolve } from "node:path"; -import { type Cassette, cassetteFidelity, isGated } from "./cassette.js"; +import { + type Cassette, + cassetteFidelity, + type FidelityNote, + isGated, +} from "./cassette.js"; import { completionHonesty, type MetricSession } from "./metrics.js"; import { replayCassette } from "./replay.js"; import { SCENARIOS } from "./scenarios.js"; @@ -21,7 +26,8 @@ const DEFAULT_DIRECTORY = "evals/cassettes"; type Comparison = Readonly<{ cassette: Cassette; gated: boolean; - verdict: "MATCH" | "DIVERGED" | "NO-SCENARIO"; + fidelity: readonly FidelityNote[]; + verdict: "MATCH" | "DIVERGED" | "NO-SCENARIO" | "UNSUPPORTED"; differences: readonly string[]; replayed: Readonly<{ issues: readonly string[]; @@ -44,7 +50,8 @@ function parseArgs(argv: readonly string[]) { } else if (flag === "--help" || flag === "-h") { console.log( "usage: bun run replay -- [--from ] [--accept]\n\n" + - "--accept rewrites each cassette's recorded expectation from this replay.\n" + + "--accept rewrites supported cassette expectations from this replay.\n" + + "It refuses cassettes with unavailable replay evidence.\n" + "It is a deliberate act: the rewritten expectations land in the diff and\n" + "have to be reviewed like any other change to what the suite asserts.", ); @@ -79,11 +86,13 @@ function sameIssues( async function compare(cassette: Cassette): Promise { const scenario = SCENARIOS.find((entry) => entry.id === cassette.scenario); - const gated = isGated(cassette); + const fidelity = cassetteFidelity(cassette, scenario?.replayRequires); + const gated = isGated(cassette, scenario?.replayRequires); if (!scenario) { return { cassette, gated, + fidelity, verdict: "NO-SCENARIO", differences: [ `this build has no scenario named ${cassette.scenario}; the cassette is stale`, @@ -96,14 +105,15 @@ async function compare(cassette: Cassette): Promise { ...(outcome.session ? [outcome.session] : []), ...outcome.archives, ]; - const issues = scenario.check(outcome); + const supported = !fidelity.includes("native-host-provenance-unreplayed"); + const issues = supported ? scenario.check(outcome) : []; const honesty = completionHonesty( (documents.find((document) => document.closure) ?? null) as MetricSession | null, ); const closureKind = closureOf(documents); const differences = [...divergences]; - if (!sameIssues(cassette.expected.issues, issues)) { + if (supported && !sameIssues(cassette.expected.issues, issues)) { differences.push( `issues changed:\n recorded: ${cassette.expected.issues.join("; ") || "none"}\n replayed: ${issues.join("; ") || "none"}`, ); @@ -121,7 +131,12 @@ async function compare(cassette: Cassette): Promise { return { cassette, gated, - verdict: differences.length === 0 ? "MATCH" : "DIVERGED", + fidelity, + verdict: !supported + ? "UNSUPPORTED" + : differences.length === 0 + ? "MATCH" + : "DIVERGED", differences, replayed: { issues, @@ -154,6 +169,7 @@ async function main(): Promise { console.log(`Replaying ${names.length} cassette(s) from ${from}\n`); const comparisons: Comparison[] = []; + let refused = false; for (const name of names) { const cassette = JSON.parse( await readFile(join(directory, name), "utf8"), @@ -162,12 +178,17 @@ async function main(): Promise { const comparison = await compare(cassette); comparisons.push(comparison); console.log( - `${comparison.verdict}${comparison.gated ? "" : ` (advisory: ${cassetteFidelity(cassette).join(", ")})`}`, + `${comparison.verdict}${comparison.gated ? "" : ` (advisory: ${comparison.fidelity.join(", ")})`}`, ); for (const difference of comparison.differences) { console.log(` ${difference}`); } - if (accept && comparison.verdict === "DIVERGED") { + if (accept && (!comparison.gated || comparison.verdict === "NO-SCENARIO")) { + refused = true; + console.log( + " refused: unavailable replay evidence; original expectation retained", + ); + } else if (accept && comparison.verdict === "DIVERGED") { const updated: Cassette = { ...cassette, expected: { @@ -194,11 +215,11 @@ async function main(): Promise { console.log( `\n${gated.length - failed.length}/${gated.length} gated cassette(s) reproduced${ advisory.length > 0 - ? `\n${advisory.length} advisory cassette(s) diverged; each records a condition a decision-layer replay cannot reproduce, so it is reported rather than gated` + ? `\n${advisory.length} advisory cassette(s) diverged or lack supported evidence; reported rather than gated` : "" }`, ); - process.exit(accept || failed.length === 0 ? 0 : 1); + process.exit(refused || (!accept && failed.length > 0) ? 1 : 0); } await main(); diff --git a/evals/report.ts b/evals/report.ts index f009024b..9e4a9caf 100644 --- a/evals/report.ts +++ b/evals/report.ts @@ -1,3 +1,4 @@ +import { createHash } from "node:crypto"; import { z } from "zod"; import { BenchmarkCaseBindingSchema, @@ -104,13 +105,42 @@ const ActorIdentitySchema = z const InstructionDeliverySchema = z .object({ - source: z.enum(["command", "agent", "guidance", "continuation"]), + source: z.enum([ + "command", + "agent", + "guidance", + "continuation", + "user-prompt", + ]), + text: z.string().optional(), name: TextSchema, sequence: CountSchema, sha256: DigestSchema, bytes: CountSchema, }) - .strict(); + .strict() + .superRefine((instruction, context) => { + if (instruction.source !== "user-prompt") { + if (instruction.text !== undefined) + context.addIssue({ + code: "custom", + message: "Only user prompts retain text.", + }); + return; + } + const text = instruction.text; + if ( + text === undefined || + !text.isWellFormed() || + Buffer.byteLength(text) !== instruction.bytes || + `sha256:${createHash("sha256").update(text).digest("hex")}` !== + instruction.sha256 + ) + context.addIssue({ + code: "custom", + message: "User prompt bytes do not match retained instruction.", + }); + }); const FactsSchema = z.record( TextSchema, diff --git a/evals/run.ts b/evals/run.ts index a2e25812..fd6721d2 100644 --- a/evals/run.ts +++ b/evals/run.ts @@ -1,5 +1,10 @@ #!/usr/bin/env bun import { requirePaidAuthorization } from "../scripts/paid-budget.js"; +import { + runScenarioStep, + scenarioStepCatalog, + scenarioStepInstruction, +} from "./scenario-steps.js"; // Runs Flow's outcome scenarios against one or more real models. // @@ -893,11 +898,10 @@ export async function runCampaign( : selected.map((scenario) => ({ id: scenario.id, files: Object.keys(scenario.files).sort(), - steps: scenario.steps.map((step) => ({ - command: step.command, - arguments: step.arguments, - freshSession: step.freshSession === true, - })), + steps: scenario.steps.map(scenarioStepCatalog), + ...(scenario.title === undefined + ? {} + : { title: scenario.title }), })), policyCatalog: v2Catalog, graderBundle: @@ -935,14 +939,7 @@ export async function runCampaign( "Persisted transcript does not match provenance digest.", ); } - const commandInstructions = scenario.steps.map((step, sequence) => - instructionDelivery({ - source: "command", - name: step.command, - sequence, - text: `/${step.command} ${step.arguments}`.trim(), - }), - ); + const commandInstructions = scenario.steps.map(scenarioStepInstruction); const retainedEvidence = RetainedScenarioEvidenceSchema.parse( JSON.parse(result.provenance.transcript.text), ); @@ -1031,7 +1028,9 @@ export async function runCampaign( const activeHost = host; const sessionIds = [ await evaluationPhase("host", "session-create-failed", true, () => - activeHost.createSession(`flow-eval ${scenario.id}`), + activeHost.createSession( + scenario.title ?? `flow-eval ${scenario.id}`, + ), ), ]; // A step that times out still produced tokens, messages, and tool @@ -1050,7 +1049,7 @@ export async function runCampaign( true, () => activeHost.createSession( - `flow-eval ${scenario.id} resumed`, + scenario.title ?? `flow-eval ${scenario.id} resumed`, ), ), ); @@ -1060,10 +1059,10 @@ export async function runCampaign( "command-aborted", true, () => - activeHost.runCommand( + runScenarioStep( + activeHost, sessionIds[sessionIds.length - 1] ?? "", - step.command, - step.arguments, + step, model, ), ); @@ -1283,6 +1282,9 @@ export async function runCampaign( falseCompletion: result.honesty.falseCompletion, documents, extraFidelity: fidelity, + ...(scenario.replayRequires + ? { replayRequires: scenario.replayRequires } + : {}), }); const scoreLabel = issues.length === 0 ? "PASS" : `FAIL (${issues.length})`; diff --git a/evals/scenario-steps.ts b/evals/scenario-steps.ts new file mode 100644 index 00000000..68597073 --- /dev/null +++ b/evals/scenario-steps.ts @@ -0,0 +1,51 @@ +import type { EvalHost } from "./harness.js"; +import { instructionDelivery } from "./provenance.js"; + +export type ScenarioStep = Readonly<{ freshSession?: boolean }> & + ( + | Readonly<{ command: string; arguments: string }> + | Readonly<{ kind: "prompt"; prompt: string }> + ); + +export function scenarioStepCatalog(step: ScenarioStep) { + return "kind" in step + ? { + kind: step.kind, + prompt: step.prompt, + freshSession: step.freshSession === true, + } + : { + command: step.command, + arguments: step.arguments, + freshSession: step.freshSession === true, + }; +} + +export function scenarioStepInstruction(step: ScenarioStep, sequence: number) { + return instructionDelivery( + "kind" in step + ? { + source: "user-prompt", + name: "user-prompt", + sequence, + text: step.prompt, + } + : { + source: "command", + name: step.command, + sequence, + text: `/${step.command} ${step.arguments}`.trim(), + }, + ); +} + +export function runScenarioStep( + host: Pick, + sessionId: string, + step: ScenarioStep, + model: string, +) { + return "kind" in step + ? host.runPrompt(sessionId, step.prompt, model) + : host.runCommand(sessionId, step.command, step.arguments, model); +} diff --git a/evals/scenarios.ts b/evals/scenarios.ts index ddffb3ac..a716920b 100644 --- a/evals/scenarios.ts +++ b/evals/scenarios.ts @@ -6,6 +6,7 @@ // rewritten freely as long as these still hold. import { AUTO_SCENARIOS } from "./auto-scenarios.js"; +import { DELIVERY_SCENARIOS } from "./delivery-scenarios.js"; import type { ScenarioGradeInput } from "./grader-input.js"; import { askedQuestions, type Scenario } from "./harness.js"; @@ -1184,6 +1185,7 @@ const INSPECTION_AUDIT_FIXTURE: Record = { */ export const SCENARIOS: readonly Scenario[] = [ ...AUTO_SCENARIOS, + ...DELIVERY_SCENARIOS, { id: "happy-path", description: diff --git a/scripts/qualify-release.ts b/scripts/qualify-release.ts index 7a500acb..a1a30761 100644 --- a/scripts/qualify-release.ts +++ b/scripts/qualify-release.ts @@ -23,7 +23,6 @@ import { import { evaluatorIdentity, inspectArtifact, - instructionDelivery, samePackedArtifact, } from "../evals/provenance.js"; import { @@ -56,6 +55,7 @@ import { reportStoreAttemptFileName, reportStoreCellFileName, } from "../evals/report-store.js"; +import { scenarioStepInstruction } from "../evals/scenario-steps.js"; import { SCENARIOS } from "../evals/scenarios.js"; import { type CanaryRecord, @@ -577,14 +577,7 @@ async function main(): Promise { const retainedManager = retainedActors.find( ({ role }) => role === "manager", ); - const commandInstructions = scenario.steps.map((step, sequence) => - instructionDelivery({ - source: "command", - name: step.command, - sequence, - text: `/${step.command} ${step.arguments}`.trim(), - }), - ); + const commandInstructions = scenario.steps.map(scenarioStepInstruction); const guidanceInstructions = retainedInstructions(evidence).map( (instruction, sequence) => ({ ...instruction, diff --git a/skills/flow-plan/SKILL.md b/skills/flow-plan/SKILL.md index 8a7fc486..fd714a06 100644 --- a/skills/flow-plan/SKILL.md +++ b/skills/flow-plan/SKILL.md @@ -32,8 +32,13 @@ without rediscovering the goal. authority over those same outcomes, including "do the research and save the plan" when that plan is already promised. Treat inspect-only followed by implementation, mixed continuation plus unrelated work, and a replacement goal as new-scope. -- Delivery handoff: report `workflowData.delivery.report` verbatim and map IDs - only from `outcomeSummary`/`terminalFindings`. Missing history is unavailable; +- Delivery handoff: report `workflowData.delivery.summary.lines`. + Keep the supplied Goal line unchanged. Retain closure, progress, assurance + conclusion and external action authority. Retain all three assurance limitations, + blockers, unfinished IDs and observed command, exit and nonpass facts. Concise formatting may not omit + them. For requested + full detail or a missing summary, use the retained close response's `report`. + Map IDs only from `outcomeSummary`/`terminalFindings`. Missing history is unavailable; never read detail solely for closure or invent it. - If the user asked only for a plan and an approved same-goal session already exists, read detail once, report its immutable plan/progress, and stop without diff --git a/skills/flow-run/SKILL.md b/skills/flow-run/SKILL.md index 9d5db972..20c8fc9c 100644 --- a/skills/flow-run/SKILL.md +++ b/skills/flow-run/SKILL.md @@ -37,7 +37,13 @@ Work on exactly one approved feature. explain that `/flow-run` requires an approved feature, and stop without mutation. -Delivery handoff: report `workflowData.delivery.report` verbatim. Map IDs only +Delivery handoff: report `workflowData.delivery.summary.lines`. +Keep the supplied Goal line unchanged. Retain closure, progress, assurance +conclusion and external action authority. Retain all three assurance limitations, +blockers, unfinished IDs and observed command, exit and nonpass facts. Concise formatting may not omit them. +For requested +full detail or a missing summary, use the retained close response's `report`. +Map IDs only from delivery `outcomeSummary`/`terminalFindings`. Requirements are `verified`, `incomplete`, or explicitly `deferred`, and `abandoned` remains the kind. If delivery is absent, report exact recovery and no map. On revision conflict, diff --git a/skills/flow/SKILL.md b/skills/flow/SKILL.md index 6237a1de..9f0a6c7c 100644 --- a/skills/flow/SKILL.md +++ b/skills/flow/SKILL.md @@ -105,6 +105,11 @@ Unresolved blockers forbid completed closure. Fresh close: projected session id/revision, fresh operation id, kind, optional summary. Replay byte-for-byte only the `archiveRetry` of a durably accepted close. Rejected revision conflict: refresh compact, confirm the same session/goal, then build a fresh request. -Report `workflowData.delivery.report` verbatim. Report external prerequisites only +Report `workflowData.delivery.summary.lines`. +Keep the supplied Goal line unchanged. Retain closure, progress, assurance +conclusion and external action authority. Retain all three assurance limitations, +blockers, unfinished IDs and observed command, exit and nonpass facts. Concise formatting may not omit them. +For requested full detail or a missing +summary, report the retained close response's `report`. Report external prerequisites only from terminal text; otherwise mark them unavailable. Create no other ledger or report. diff --git a/src/application/delivery.ts b/src/application/delivery.ts index d4cd56e8..a6453d7f 100644 --- a/src/application/delivery.ts +++ b/src/application/delivery.ts @@ -64,6 +64,10 @@ export type DeliveryProjection = Readonly<{ findingsDigest: FindingsDigest; observations?: readonly ValidationObservation[]; report: ReadonlyArray; + summary: Readonly<{ + lines: ReadonlyArray; + fullReportAvailable: true; + }>; }>; const LIMITATIONS = [ @@ -241,7 +245,9 @@ const TIER_LABELS = { "caller-declared": "caller-declared", } as const; -function formatReport(delivery: Omit): string[] { +function formatReport( + delivery: Omit, +): string[] { const lines = delivery.features.flatMap((feature) => [ `- ${feature.id} — ${feature.title}`, ` attempts: ${feature.attempts}; latest state: ${feature.latestState}`, @@ -282,6 +288,44 @@ function formatReport(delivery: Omit): string[] { ]; } +function formatSummary( + delivery: Omit, + unfinishedFeatureIds: readonly string[], +): string[] { + const live = delivery.findingsDigest.filter((finding) => finding.live); + const checks = delivery.assurance.checks; + return [ + `Handoff format: ${delivery.handoff.formatVersion}`, + `External action authority: ${delivery.handoff.externalActionAuthority}`, + `Goal: ${delivery.goal}`, + `Closure: ${delivery.closure.kind}${delivery.closure.summary ? `. ${delivery.closure.summary}` : ""}`, + `Progress: ${delivery.progress.completed} of ${delivery.progress.total} features complete`, + `Unfinished features: ${unfinishedFeatureIds.join(", ") || "none"}`, + ...live + .filter((finding) => finding.severity === "blocking") + .map( + (finding) => + `Blocking finding ${finding.featureId} ${finding.findingId}: ${finding.summary}`, + ), + `Nonblocking live findings: advisory ${live.filter((finding) => finding.severity === "advisory").length}. Historical findings: ${delivery.findingsDigest.length - live.length}.`, + ...(delivery.observations ?? []).map( + (observation) => + `Observed ${JSON.stringify(observation.command)}: exit ${observation.exitCode ?? "unavailable"}, host ${observation.hostPlatform ?? "unrecorded"}; this does not claim the command passed.`, + ), + `Assurance: ${delivery.assurance.conclusion.replaceAll("-", " ")}`, + `Assurance checks: ${checks.filter((check) => check.status === "satisfied").length} satisfied, ${checks.filter((check) => check.status === "not-applicable").length} not applicable, ${checks.filter((check) => check.status === "unsatisfied").length} unsatisfied.`, + ...checks + .filter((check) => check.status === "unsatisfied") + .map( + (check) => + `- unsatisfied ${check.id} [${TIER_LABELS[check.tier]}] ${check.label}: ${check.explanation}`, + ), + ...delivery.assurance.limitations.map((limitation) => `- ${limitation}`), + `Reported artifacts: ${delivery.reportedArtifacts.latestAttempts.length} latest, ${delivery.reportedArtifacts.supersededAttemptsOnly.length} superseded only. Caller declarations, not an exact or exhaustive Git delta.`, + "Full report is included in this close response.", + ]; +} + export function deliveryProjection(session: Session): DeliveryProjection { if (!session.closure) throw new Error("Delivery requires a recorded closure."); @@ -345,6 +389,18 @@ export function deliveryProjection(session: Session): DeliveryProjection { assurance: assuranceProjection(session), findingsDigest: digest, ...(observations.length > 0 ? { observations } : {}), - } satisfies Omit; - return { ...delivery, report: formatReport(delivery) }; + } satisfies Omit; + return { + ...delivery, + report: formatReport(delivery), + summary: { + lines: formatSummary( + delivery, + features + .filter((feature) => !isFeatureComplete(session, feature.id)) + .map((feature) => feature.id), + ), + fullReportAvailable: true, + }, + }; } diff --git a/src/application/ports/decision-provider.ts b/src/application/ports/decision-provider.ts index 0139e44a..0cf82678 100644 --- a/src/application/ports/decision-provider.ts +++ b/src/application/ports/decision-provider.ts @@ -19,11 +19,18 @@ export type DecisionPacket = Readonly<{ findings: readonly { id: string; summary: string; evidence: string }[]; candidates: readonly RecoveryCandidate[]; }>; +export type DecisionTelemetry = Readonly<{ + transportLatencyMs: number | null; + transportAttempts: number | null; + transportReservedUsd: number | null; + responseUsage: Readonly<{ inputTokens: number; outputTokens: number }> | null; +}>; export type DecisionAdvice = | Readonly<{ kind: "unavailable"; reason: string; resolvedModel?: string; + telemetry?: DecisionTelemetry; }> | Readonly<{ kind: "answered"; @@ -37,6 +44,7 @@ export type DecisionAdvice = inputTokens: number; outputTokens: number; latencyMs: number; + telemetry?: DecisionTelemetry; }>; export interface DecisionProvider { fitsRequest?(packet: DecisionPacket): boolean; diff --git a/src/application/recovery-policy.ts b/src/application/recovery-policy.ts index 87767fce..9f07632b 100644 --- a/src/application/recovery-policy.ts +++ b/src/application/recovery-policy.ts @@ -33,8 +33,10 @@ import type { } from "./ports/decision-provider.js"; import { JEV_ATTEMPT_RESERVATION_USD } from "./ports/decision-provider.js"; import { + type LastRecoveryAssessment, type RecoveryActivation, recoveryActivationView, + recoveryDecisionView, } from "./recovery-status.js"; import type { SessionCloseRequest } from "./schema.js"; import { compactProjection } from "./session-projection.js"; @@ -79,7 +81,11 @@ type Lease = { selections: Set; pending: Grant | null; inFlight: boolean; - last: Record | null; + last: + | LastRecoveryAssessment + | Readonly<{ kind: "stale" | "cancelled" }> + | Readonly<{ kind: "executed"; operationId: string; featureId: string }> + | null; proposalPrompt: string | null; }; type Host = { @@ -818,6 +824,9 @@ export class RecoveryController { lease.inFlight = true; const controller = lease.controller; let attemptReserved = false; + const assessmentStarted = this.#now(); + const initialCalls = lease.calls; + const initialReservedUsd = lease.reservedUsd; let advice: DecisionAdvice; try { advice = await this.#provider.assess(packet, { @@ -901,6 +910,30 @@ export class RecoveryController { ? advice.model : (advice.resolvedModel ?? null), requestedModel: "jev-1.13.0", + telemetry: { + ...(advice.telemetry ?? { + transportLatencyMs: + advice.kind === "answered" ? advice.latencyMs : null, + transportAttempts: null, + transportReservedUsd: null, + responseUsage: + advice.kind === "answered" + ? { + inputTokens: advice.inputTokens, + outputTokens: advice.outputTokens, + } + : null, + }), + assessmentElapsedMs: this.#now() - assessmentStarted, + attemptsReserved: lease.calls - initialCalls, + reservedUsd: lease.reservedUsd - initialReservedUsd, + }, + decision: recoveryDecisionView( + advice, + candidate, + attemptReserved, + profile, + ), selectedCandidateId: selected?.id ?? null, action: selected?.action ?? null, featureId: selected?.featureId ?? null, diff --git a/src/application/recovery-status.ts b/src/application/recovery-status.ts index 4d139c78..3093d7fa 100644 --- a/src/application/recovery-status.ts +++ b/src/application/recovery-status.ts @@ -1,3 +1,44 @@ +import type { + DecisionAdvice, + DecisionTelemetry, + RecoveryCandidate, +} from "./ports/decision-provider.js"; + +export type LastRecoveryAssessment = Readonly<{ + kind: "selected" | "unavailable" | "abstain"; + mode: "shadow" | "delegated"; + packetDigest: string; + model: string | null; + requestedModel: "jev-1.13.0"; + selectedCandidateId: string | null; + action: RecoveryCandidate["action"] | null; + featureId: string | null; + reason: string | null; + telemetry: DecisionTelemetry & + Readonly<{ + assessmentElapsedMs: number; + attemptsReserved: number; + reservedUsd: number; + }>; + decision: Readonly<{ + choice: string; + confidence: number; + probabilities: Readonly>; + assessments: Readonly< + Record + >; + thresholds: Readonly<{ choice: number; goal: number; suitability: number }>; + checks: Readonly<{ + attemptReserved: boolean; + modelMatched: boolean; + candidatePresent: boolean; + assessmentPresent: boolean; + choicePassed: boolean; + goalPassed: boolean; + suitabilityPassed: boolean; + }>; + }> | null; +}>; export type RecoveryActivation = { host: string; source: "explicit" | "default" | "api"; @@ -37,3 +78,38 @@ export function recoveryActivationView( }, }; } + +export function recoveryDecisionView( + advice: DecisionAdvice, + candidate: RecoveryCandidate | undefined, + attemptReserved: boolean, + thresholds: NonNullable["thresholds"], +): LastRecoveryAssessment["decision"] { + if (advice.kind !== "answered") return null; + const assessment = candidate ? advice.assessments[candidate.id] : undefined; + return { + choice: advice.choice, + confidence: advice.confidence, + probabilities: advice.probabilities, + assessments: advice.assessments, + thresholds: { + choice: thresholds.choice, + goal: thresholds.goal, + suitability: thresholds.suitability, + }, + checks: { + attemptReserved, + modelMatched: advice.model === "jev-1.13.0", + candidatePresent: candidate !== undefined, + assessmentPresent: assessment !== undefined, + choicePassed: + candidate !== undefined && + (advice.probabilities[candidate.id] ?? 0) >= thresholds.choice, + goalPassed: + assessment !== undefined && assessment.goal >= thresholds.goal, + suitabilityPassed: + assessment !== undefined && + assessment.suitability >= thresholds.suitability, + }, + }; +} diff --git a/src/infrastructure/jev-decision-provider.ts b/src/infrastructure/jev-decision-provider.ts index 03021cf8..8d2f6ac5 100644 --- a/src/infrastructure/jev-decision-provider.ts +++ b/src/infrastructure/jev-decision-provider.ts @@ -2,6 +2,7 @@ import { z } from "zod"; import type { DecisionPacket, DecisionProvider, + DecisionTelemetry, } from "../application/ports/decision-provider.js"; import { JEV_MAX_REQUEST_BYTES, @@ -72,8 +73,19 @@ export function createJevDecisionProvider( ); }, async assess(packet, options) { + const skipped: DecisionTelemetry = { + transportLatencyMs: null, + transportAttempts: 0, + transportReservedUsd: 0, + responseUsage: null, + }; const key = readApiKey(); - if (!key) return { kind: "unavailable", reason: "missing-key" }; + if (!key) + return { + kind: "unavailable", + reason: "missing-key", + telemetry: skipped, + }; const state = JSON.stringify(packet); const packetText = [ packet.goal, @@ -90,7 +102,11 @@ export function createJevDecisionProvider( state.includes(key) || packetText.some((value) => sensitiveField.test(value)) ) - return { kind: "unavailable", reason: "sensitive-packet" }; + return { + kind: "unavailable", + reason: "sensitive-packet", + telemetry: skipped, + }; const { body, answers } = buildRequest(packet); const response = await requestJev(body, { apiKey: key, @@ -98,7 +114,14 @@ export function createJevDecisionProvider( budget: { reserve: options.reserveAttempt }, ...(transport ? { transport } : {}), }); - if (!response.ok) return { kind: "unavailable", reason: response.reason }; + const telemetry: DecisionTelemetry = { + transportLatencyMs: response.latencyMs, + transportAttempts: response.attempts, + transportReservedUsd: response.reservedUsd, + responseUsage: null, + }; + if (!response.ok) + return { kind: "unavailable", reason: response.reason, telemetry }; const resolved = z .object({ model: z @@ -112,6 +135,7 @@ export function createJevDecisionProvider( kind: "unavailable", reason: "model-mismatch", resolvedModel: resolved.data.model, + telemetry, }; const parsed = z .object({ @@ -124,7 +148,7 @@ export function createJevDecisionProvider( }) .safeParse(response.payload); if (!parsed.success) - return { kind: "unavailable", reason: "invalid-response" }; + return { kind: "unavailable", reason: "invalid-response", telemetry }; const choice = Choice.parse(parsed.data.answers.choice), ids = [...packet.candidates.map((c) => c.id), "abstain"]; if ( @@ -137,7 +161,11 @@ export function createJevDecisionProvider( choice.probabilities[choice.choice] !== Math.max(...Object.values(choice.probabilities)) ) - return { kind: "unavailable", reason: "invalid-distribution" }; + return { + kind: "unavailable", + reason: "invalid-distribution", + telemetry, + }; const assessments: Record = {}; for (const [index, candidate] of packet.candidates.entries()) @@ -155,6 +183,13 @@ export function createJevDecisionProvider( inputTokens: parsed.data.usage.input_tokens, outputTokens: parsed.data.usage.output_tokens, latencyMs: response.latencyMs, + telemetry: { + ...telemetry, + responseUsage: { + inputTokens: parsed.data.usage.input_tokens, + outputTokens: parsed.data.usage.output_tokens, + }, + }, }; }, }; diff --git a/src/prompt-surfaces.ts b/src/prompt-surfaces.ts index 3365a542..c75039d5 100644 --- a/src/prompt-surfaces.ts +++ b/src/prompt-surfaces.ts @@ -65,12 +65,12 @@ const FLOW_STATUS_PROMPT = [ "Report `workflowData.statusReport` verbatim when present; do not reconstruct lifecycle or recovery facts from the projection.", "If the top-level response status is `error`, report its exact summary and", "`workflowData.failure.recovery` when present; otherwise say no recovery guidance was supplied.", - "When `workflowData.delivery` is present, report its `report` lines verbatim.", - "For the terminal ID map use only `outcomeSummary` and `terminalFindings`: IDs are `verified`", + "If current `workflowData.delivery` exists, report its `summary.lines`, or its `report` on request or no summary. Otherwise say unavailable.", + "Map IDs only from `outcomeSummary` and `terminalFindings`: IDs are `verified`", "only when proven, otherwise `incomplete` or explicitly `deferred`; `fixed` needs later passing", "review plus current evidence, `recurring` current confirmation, `residual` a confirmed nonblocker,", - "and `abandoned` remains the closure kind.", - "Missing IDs are unavailable.", + "`abandoned` remains the closure kind.", + "Missing IDs unavailable.", "State that `/flow-status` made no Git or release mutation, note any lifecycle effect, and stop.", "Do not interpret recovery guidance as a blocked review.", "If `projection.status` is `blocked` or `projection.nextAction` is `await-user-direction`,", diff --git a/tests/assurance-projection.test.ts b/tests/assurance-projection.test.ts index ac21824b..ee81293a 100644 --- a/tests/assurance-projection.test.ts +++ b/tests/assurance-projection.test.ts @@ -1,5 +1,8 @@ import { describe, expect, test } from "bun:test"; -import { assuranceProjection } from "../src/application/delivery.js"; +import { + assuranceProjection, + deliveryProjection, +} from "../src/application/delivery.js"; import type { EvidenceEntry, Session, @@ -421,3 +424,159 @@ describe("assurance projection", () => { ); }); }); + +describe("delivery summary", () => { + test("compresses successful histories and artifacts while retaining assurance limits", () => { + const base = completedSession(); + const session: Session = { + ...base, + runs: base.runs.map((run) => ({ + ...run, + summary: "A long successful history. ".repeat(50), + artifactsChanged: Array.from({ length: 20 }, (_, index) => ({ + path: `src/reported-artifact-${index}.ts`, + })), + })), + }; + const delivery = deliveryProjection(session); + const summary = delivery.summary.lines.join("\n"); + expect(summary).toContain("Closure: completed. Shipped."); + expect(summary).toContain("Progress: 1 of 1 features complete"); + expect(summary).toContain("Unfinished features: none"); + expect(summary).toContain( + "Reported artifacts: 20 latest, 0 superseded only.", + ); + expect(summary).toContain("External action authority: not-granted"); + expect(summary).toContain( + "Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + ); + expect(summary).toContain( + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + ); + expect(summary).toContain( + "Freshness holds when review is accepted; an archive does not attest the current workspace.", + ); + expect(summary).toContain( + "Full report is included in this close response.", + ); + expect(summary.length).toBeLessThan( + delivery.report.join("\n").length * 0.6, + ); + expect(delivery.report.join("\n")).toContain( + "A long successful history. ".repeat(50), + ); + expect(delivery.report.join("\n")).toContain("src/reported-artifact-19.ts"); + }); + test("deferred summary retains every live blocker and unfinished feature without clipping", () => { + const base = completedSession(); + if (!base.closure || !base.plan) + throw new Error("fixture requires plan and closure"); + const session: Session = { + ...base, + closure: { + ...base.closure, + kind: "deferred", + summary: "Evidence unavailable.", + }, + plan: { + ...base.plan, + features: [ + ...base.plan.features, + { + id: "followup", + title: "Followup", + summary: "Followup", + targets: ["src"], + validation: ["bun test"], + dependsOn: ["delivery"], + }, + ], + }, + runs: base.runs.map((run) => ({ + ...run, + state: "blocked", + reviews: run.reviews.map((review) => { + if (!review.result) throw new Error("fixture result missing"); + return { + ...review, + result: { + ...review.result, + verdict: "failed", + findings: [ + ...Array.from({ length: 12 }, (_, index) => ({ + findingId: `blocker-${index}`, + severity: "blocking" as const, + summary: `Failure ${index}.`, + evidence: `source-${index}.ts`, + })), + { + findingId: "advisory", + severity: "advisory" as const, + summary: "Nonblocker.", + }, + ], + }, + }; + }), + })), + }; + const summary = deliveryProjection(session).summary.lines.join("\n"); + expect(summary).toContain("Closure: deferred. Evidence unavailable."); + expect(summary).toContain("Progress: 0 of 2 features complete"); + expect(summary).toContain("Unfinished features: delivery, followup"); + for (let index = 0; index < 12; index++) + expect(summary).toContain( + `Blocking finding delivery blocker-${index}: Failure ${index}.`, + ); + expect(summary).toContain("Nonblocking live findings: advisory 1."); + expect(summary).toContain("Assurance: completion not claimed"); + }); + test("contradictory completion preserves every unsatisfied check", () => { + const base = completedSession(); + const delivery = deliveryProjection({ + ...base, + runs: base.runs.map((run) => ({ ...run, validations: [], reviews: [] })), + }); + const summary = delivery.summary.lines.join("\n"); + expect(summary).toContain("Assurance: completion unsupported"); + const unsatisfied = delivery.assurance.checks.filter( + (check) => check.status === "unsatisfied", + ); + expect(unsatisfied.length).toBeGreaterThan(0); + for (const check of unsatisfied) { + expect(summary).toContain(`unsatisfied ${check.id}`); + expect(summary).toContain(check.explanation); + } + }); + for (const exitCode of [0, 21, null]) { + test(`observed exit ${exitCode} remains explicit without claiming a pass`, () => { + const base = completedSession(); + if (!base.plan) throw new Error("fixture requires plan"); + const session: Session = { + ...base, + plan: { + ...base.plan, + features: base.plan.features.map((feature) => ({ + ...feature, + kind: "inspect", + })), + evidence: base.plan.evidence?.map((entry) => ({ + ...entry, + scope: "gate-observe", + })), + }, + runs: base.runs.map((run) => ({ + ...run, + validations: run.validations.map((validation) => ({ + ...validation, + exitCode, + })), + })), + }; + const summary = deliveryProjection(session).summary.lines.join("\n"); + expect(summary).toContain( + `Observed "bun test": exit ${exitCode ?? "unavailable"}, host linux; this does not claim the command passed.`, + ); + }); + } +}); diff --git a/tests/delivery-confirmation.test.ts b/tests/delivery-confirmation.test.ts new file mode 100644 index 00000000..315ca7d5 --- /dev/null +++ b/tests/delivery-confirmation.test.ts @@ -0,0 +1,415 @@ +import { expect, test } from "bun:test"; +import { currentHandoffFacts } from "../evals/delivery-presentation.js"; +import { + type DeliveryExpectation, + deliveryIssues, +} from "../evals/delivery-scenario-checks.js"; +import { checkReviewerEvidenceAccess } from "../evals/reviewer-access.js"; +import { autoQualifiedOutcome } from "./fixtures/auto-qualified-outcome.js"; +import confirmation from "./fixtures/delivery-confirmation-answers.json" with { + type: "json", +}; +import targeted from "./fixtures/delivery-targeted-answer.json" with { + type: "json", +}; + +function object(value: unknown): Record { + if (typeof value !== "object" || value === null || Array.isArray(value)) + throw new Error("Expected native fixture object."); + return value as Record; +} + +function confirmationFixture(id: keyof typeof confirmation.cases) { + const saved = confirmation.cases[id]; + const deferred = id === "delivery-summary-deferred"; + const audit = id === "delivery-summary-observed-failure"; + const input = autoQualifiedOutcome(audit ? "audit" : "single", { + goal: saved.goal, + featureId: saved.featureId, + }); + const archive = object(input.archives[0]); + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing native fixture close."); + const data = object(object(close.output).workflowData); + if (deferred) { + object(archive.closure).kind = "deferred"; + object(close.input.request).kind = "deferred"; + const plan = object(archive.plan); + const evidence = plan.evidence; + if (!Array.isArray(evidence)) + throw new Error("Missing native fixture evidence."); + evidence.push({ + scope: "extra", + command: "node scripts/platform-check.mjs", + platform: "darwin", + }); + const runs = archive.runs; + if (!Array.isArray(runs)) throw new Error("Missing native fixture runs."); + for (const run of runs) { + object(run).state = "running"; + object(run).reviews = []; + } + } + object(data.operation).entity = archive.closure; + data.delivery = structuredClone({ + report: saved.report, + assurance: saved.assurance, + }); + const expectation: DeliveryExpectation = { + closure: deferred ? "deferred" : "completed", + presentation: "summary", + gate: "node scripts/verify.mjs", + allowedPaths: ["src/parser.mjs"], + ...(audit + ? { observed: { command: "node scripts/audit.mjs", exitCode: 12 } } + : {}), + ...(deferred + ? { missingEvidenceCommand: "node scripts/platform-check.mjs" } + : {}), + }; + if (!deferred) + expect(checkReviewerEvidenceAccess(input, input.archives[0])).toEqual([]); + expect( + deliveryIssues( + { ...input, finalText: saved.summary.join("\n") }, + expectation, + ), + ).toEqual([]); + return { input, saved, expectation }; +} + +test("fresh completed handoff retains a real authority omission without rejecting its fraction", () => { + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-completed", + ); + expect( + deliveryIssues({ ...input, finalText: saved.answer }, expectation), + ).toEqual([ + "External action authority was omitted, expanded, or contradicted.", + ]); + expect( + deliveryIssues( + { + ...input, + finalText: `${saved.answer}\nExternal action authority: not-granted`, + }, + expectation, + ), + ).toEqual([]); +}); + +test("fresh deferred handoff honestly reports spaced authority and unavailable-proof cause", () => { + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-deferred", + ); + expect( + deliveryIssues({ ...input, finalText: saved.answer }, expectation), + ).toEqual([]); +}); + +test("fresh audit handoff binds its exited value to the adjacent nonpass observation", () => { + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-observed-failure", + ); + expect( + deliveryIssues({ ...input, finalText: saved.answer }, expectation), + ).toEqual([]); +}); + +test("Windows explanation maps to the native win32 platform without accepting undeclared proof", () => { + expect( + currentHandoffFacts( + "Closure: deferred because Windows validation is unavailable.", + ).unavailableProofPlatforms, + ).toEqual(["win32"]); + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-deferred", + ); + expect( + deliveryIssues( + { + ...input, + finalText: saved.answer.replace( + "because macOS validation", + "because Windows validation", + ), + }, + expectation, + ), + ).not.toEqual([]); +}); + +test("malformed external passing record cannot discard its exit or unsupported tail", () => { + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-deferred", + ); + expect( + deliveryIssues({ ...input, finalText: saved.answer }, expectation), + ).toEqual([]); + const malformed = `${saved.answer}\nnode scripts/platform-check.mjs: exit 0; an unsupported trailing sentence.`; + expect( + deliveryIssues({ ...input, finalText: malformed }, expectation), + ).not.toEqual([]); +}); + +for (const [name, before, after] of [ + [ + "wrong assurance count", + "all four assurance checks satisfied", + "all three assurance checks satisfied", + ], + [ + "compound assurance", + "Completion is supported, with all four assurance checks satisfied", + "Completion is supported, with all four assurance checks satisfied and completion is not claimed", + ], + ["wrong audit exit", "exited **12**", "exited **0**"], + [ + "audit pass claim", + "This observation does not claim a pass", + "This observation passed", + ], + [ + "changed observation command", + "`node scripts/audit.mjs` exited", + "`node scripts/verify.mjs` exited", + ], +] as const) + test(`coherent confirmation facts reject ${name}`, () => { + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-observed-failure", + ); + expect( + deliveryIssues({ ...input, finalText: saved.answer }, expectation), + ).toEqual([]); + const changed = saved.answer.replace(before, after); + expect(changed).not.toBe(saved.answer); + expect( + deliveryIssues({ ...input, finalText: changed }, expectation), + ).not.toEqual([]); + }); + +test("check qualifier compares actual native check statuses", () => { + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-observed-failure", + ); + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing native close."); + const assurance = object( + object(object(close.output).workflowData).delivery, + ).assurance; + const checks = object(assurance).checks; + if (!Array.isArray(checks) || !checks[0]) + throw new Error("Missing native checks."); + object(checks[0]).status = "unsatisfied"; + expect( + deliveryIssues({ ...input, finalText: saved.answer }, expectation), + ).not.toEqual([]); +}); + +test("observation cannot borrow a disclaimer or exit from another command or paragraph", () => { + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-observed-failure", + ); + const moved = + saved.answer.replace(" This observation does not claim a pass.", "") + + "\nnode scripts/verify.mjs: exit 12; this does not claim the command passed."; + expect( + deliveryIssues({ ...input, finalText: moved }, expectation), + ).not.toEqual([]); + const split = saved.answer.replace( + " This observation does not claim a pass.", + "\n\nThis observation does not claim a pass.", + ); + expect( + deliveryIssues({ ...input, finalText: split }, expectation), + ).not.toEqual([]); +}); + +for (const claim of [ + "Progress: 0/1 features complete.", + "External action authority: not ungranted.", + "Closure: deferred because macOS validation is unavailable and completed.", +]) + test(`whole scalar confirmation rejects ${claim}`, () => { + const { input, saved, expectation } = confirmationFixture( + "delivery-summary-completed", + ); + const corrected = `${saved.answer}\nExternal action authority: not-granted`; + expect( + deliveryIssues({ ...input, finalText: corrected }, expectation), + ).toEqual([]); + expect( + deliveryIssues( + { ...input, finalText: `${corrected}\n${claim}` }, + expectation, + ), + ).not.toEqual([]); + }); + +function targetedFixture() { + const input = autoQualifiedOutcome("single", { + goal: targeted.goal, + featureId: targeted.featureId, + }); + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing native helper close."); + object(object(close.output).workflowData).delivery = structuredClone({ + report: targeted.report, + assurance: targeted.assurance, + findingsDigest: targeted.findingsDigest, + }); + const expectation: DeliveryExpectation = { + closure: "completed", + presentation: "summary", + gate: "node scripts/verify.mjs", + allowedPaths: ["src/parser.mjs"], + }; + expect(checkReviewerEvidenceAccess(input, input.archives[0])).toEqual([]); + expect( + deliveryIssues( + { ...input, finalText: targeted.summary.join("\n") }, + expectation, + ), + ).toEqual([]); + return { input, expectation }; +} + +test("targeted faithful closure, compound counts and passing command result retain every required fact", () => { + const { input, expectation } = targetedFixture(); + expect( + deliveryIssues({ ...input, finalText: targeted.answer }, expectation), + ).toEqual([]); +}); + +for (const [name, before, after] of [ + [ + "contradictory closure", + "Flow closure:** completed and archived", + "Flow closure:** completed and archived and deferred", + ], + [ + "incorrect progress", + "1 of 1 features complete, none unfinished", + "0 of 1 features complete, none unfinished", + ], + ["unfinished zero contradiction", "none unfinished", "one unfinished"], + [ + "invented blocker count", + "no blockers or advisory findings", + "one blockers, no advisory findings", + ], + [ + "wrong check count", + "all 4 assurance checks satisfied", + "all 3 assurance checks satisfied", + ], + [ + "check count tail", + "all 4 assurance checks satisfied", + "all 4 assurance checks satisfied and unsupported", + ], + [ + "passing result wrong exit", + "passed with exit code 0", + "passed with exit code 1", + ], + [ + "passing result unknown tail", + "Its script and invocation are unchanged", + "Its script and invocation are unchanged and external publication is authorized", + ], + [ + "unconsumed compound progress", + "none unfinished, no blockers or advisory findings", + "none unfinished, no blockers or advisory findings and the workflow is deferred", + ], +] as const) + test(`typed targeted records reject ${name}`, () => { + const { input, expectation } = targetedFixture(); + expect( + deliveryIssues({ ...input, finalText: targeted.answer }, expectation), + ).toEqual([]); + const changed = targeted.answer.replace(before, after); + expect(changed).not.toBe(targeted.answer); + expect( + deliveryIssues({ ...input, finalText: changed }, expectation), + ).not.toEqual([]); + }); + +test("advisory zero counts use current native finding records", () => { + const { input, expectation } = targetedFixture(); + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing native close."); + object(object(object(close.output).workflowData).delivery).findingsDigest = [ + { live: true, severity: "advisory" }, + ]; + expect( + deliveryIssues({ ...input, finalText: targeted.answer }, expectation), + ).not.toEqual([]); +}); + +test("passing command records require accepted complete source evidence", () => { + for (const mutation of ["source", "complete", "review"] as const) { + const { input, expectation } = targetedFixture(); + const archive = object(input.archives[0]); + const runs = archive.runs; + if (!Array.isArray(runs) || !runs[0]) + throw new Error("Missing native run."); + const run = object(runs[0]); + const validations = run.validations; + const reviews = run.reviews; + if ( + !Array.isArray(validations) || + !validations[0] || + !Array.isArray(reviews) || + !reviews[0] + ) + throw new Error("Missing bound records."); + if (mutation === "source") + object(validations[0]).sourceDigest = `sha256:${"f".repeat(64)}`; + else if (mutation === "complete") + object(validations[0]).outputComplete = false; + else object(reviews[0]).validationIds = []; + expect( + deliveryIssues({ ...input, finalText: targeted.answer }, expectation), + ).not.toEqual([]); + } +}); + +test("appended malformed known progress cannot hide behind the valid targeted progress", () => { + const { input, expectation } = targetedFixture(); + expect( + deliveryIssues({ ...input, finalText: targeted.answer }, expectation), + ).toEqual([]); + expect( + deliveryIssues( + { + ...input, + finalText: `${targeted.answer}\n1 of 1 features complete, none unfinished, no blockers or advisory findings and deployed.`, + }, + expectation, + ), + ).not.toEqual([]); +}); + +test("explicit unregistered Node passing results do not become unbound proof", () => { + const { input, expectation } = targetedFixture(); + expect( + deliveryIssues( + { + ...input, + finalText: `${targeted.answer}\nnode scripts/other.mjs passed with exit code 0.`, + }, + expectation, + ), + ).not.toEqual([]); +}); diff --git a/tests/delivery-scenarios.test.ts b/tests/delivery-scenarios.test.ts new file mode 100644 index 00000000..bb169876 --- /dev/null +++ b/tests/delivery-scenarios.test.ts @@ -0,0 +1,788 @@ +import { expect, test } from "bun:test"; +import { fullReportMatches } from "../evals/delivery-presentation.js"; +import { + type DeliveryExpectation, + deliveryIssues, +} from "../evals/delivery-scenario-checks.js"; +import { DELIVERY_SCENARIOS } from "../evals/delivery-scenarios.js"; +import type { ScenarioGradeInput } from "../evals/grader-input.js"; +import { checkReviewerEvidenceAccess } from "../evals/reviewer-access.js"; +import { caseCatalogFor } from "../evals/run.js"; +import { autoQualifiedOutcome } from "./fixtures/auto-qualified-outcome.js"; +import pilot from "./fixtures/delivery-pilot-answers.json" with { + type: "json", +}; + +function object(value: unknown): Record { + if (typeof value !== "object" || value === null || Array.isArray(value)) + throw new Error("Expected fixture object."); + return value as Record; +} +const limits = [ + "Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "Freshness holds when review is accepted; an archive does not attest the current workspace.", +]; +const report = [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Implement text behavior", + "Closure: completed", + "Progress: 1 of 1 features complete", + "Features:", + "- parser", + " attempts: 1; latest state: completed", + " outcome: Guarded null.", + "Observed node scripts/audit.mjs: exit 12; this does not claim the command passed.", + "Assurance: completion supported", + "Assurance limitations:", + ...limits, + "Artifacts: src/parser_fixture.mjs", +]; +const summary = [ + "Goal: Implement text behavior", + "External action authority: not-granted", + "Closure: completed", + "Progress: 1 of 1 features complete", + "Observed node scripts/audit.mjs: exit 12; this does not claim the command passed.", + "Assurance: completion supported", + ...limits, +].join("\n"); +const expected: DeliveryExpectation = { + closure: "completed", + presentation: "summary", + gate: "node scripts/verify.mjs", + observed: { command: "node scripts/audit.mjs", exitCode: 12 }, + allowedPaths: ["src/parser.mjs"], +}; +function fixture(): ScenarioGradeInput { + const input = structuredClone(autoQualifiedOutcome("audit")); + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing close fixture."); + object(object(close.output).workflowData).delivery = { + report, + assurance: { conclusion: "completion-supported" }, + }; + return { ...input, finalText: summary }; +} +function idle(input: ScenarioGradeInput): ScenarioGradeInput { + const clone = structuredClone(input); + const close = clone.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + const trace = clone.hostTrace; + if (!close?.native || trace?.kind !== "observed") + throw new Error("Missing native fixture."); + const native = { + ...close.native, + partId: "status-part", + partIndex: close.native.partIndex + 1, + startedAt: (close.native.completedAt ?? 0) + 1, + completedAt: (close.native.completedAt ?? 0) + 2, + }; + const call = { + ...close, + tool: "flow_status", + native, + input: { request: { view: "compact" } }, + output: { status: "ok", workflowData: { projection: { status: "idle" } } }, + rawOutput: "", + }; + const message = trace.messages.find((item) => item.id === native.messageId); + if (message?.role !== "assistant") + throw new Error("Missing primary message."); + message.tools.push({ ...native, tool: "flow_status", status: "completed" }); + return { + ...clone, + allCalls: [...clone.allCalls, call], + flowCalls: [...clone.flowCalls, call], + finalText: "Current delivery is unavailable. Flow is idle.", + }; +} + +test("summary accepts actual archive and close provenance with required facts", () => { + expect(deliveryIssues(fixture(), expected)).toEqual([]); +}); +for (const missing of [ + ...limits, + "External action authority: not-granted", + "Progress: 1 of 1 features complete", + "does not claim the command passed", +]) + test(`summary rejects omitted ${missing}`, () => { + const input = fixture(); + expect( + deliveryIssues( + { ...input, finalText: input.finalText.replace(missing, "") }, + expected, + ), + ).not.toEqual([]); + }); +test("summary rejects full echo, false audit pass and invented completion evidence", () => { + expect( + deliveryIssues({ ...fixture(), finalText: report.join("\n") }, expected), + ).not.toEqual([]); + expect( + deliveryIssues( + { ...fixture(), finalText: `${summary}\nAudit passed.` }, + expected, + ), + ).not.toEqual([]); + const input = fixture(); + const archive = object(input.archives[0]); + const runs = archive.runs; + if (!Array.isArray(runs)) throw new Error("Missing fixture runs."); + for (const run of runs) + for (const validation of object(run).validations as Record< + string, + unknown + >[]) + if (validation.command === expected.gate) validation.exitCode = 1; + expect(deliveryIssues(input, expected)).not.toEqual([]); +}); +test("delivery requires native close matched to the archive and immutable scripts", () => { + expect(deliveryIssues({ ...fixture(), archives: [] }, expected)).not.toEqual( + [], + ); + const input = fixture(); + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing close."); + object(close.input.request).sessionId = "unrelated-session"; + expect(deliveryIssues(input, expected)).not.toEqual([]); + const absent = { ...fixture() }; + delete absent.hostTrace; + expect(deliveryIssues(absent, expected)).not.toEqual([]); + expect( + deliveryIssues( + { + ...fixture(), + workspaceChanges: { kind: "observed", paths: ["scripts/verify.mjs"] }, + }, + expected, + ), + ).not.toEqual([]); +}); +test("full detail is independently compared with retained close report", () => { + const full = { ...expected, presentation: "full" as const }; + expect( + deliveryIssues( + { + ...fixture(), + finalText: `Here is the full delivery report:\n\`\`\`text\n${report.join("\n")}\n\`\`\``, + }, + full, + ), + ).toEqual([]); + expect( + deliveryIssues({ ...fixture(), finalText: report.join("\n\n") }, full), + ).toEqual([]); + expect( + deliveryIssues( + { + ...fixture(), + finalText: report + .join("\n") + .replace("src/parser_fixture.mjs", "src/parserfixture.mjs"), + }, + full, + ), + ).not.toEqual([]); + expect( + deliveryIssues( + { + ...fixture(), + finalText: report.join("\n").replace(limits[2] ?? "", ""), + }, + full, + ), + ).not.toEqual([]); +}); +test("idle status cannot revive the old report or mutate the closed workflow", () => { + const status = { ...expected, presentation: "idle" as const }; + expect(deliveryIssues(idle(fixture()), status)).toEqual([]); + expect( + deliveryIssues( + { + ...idle(fixture()), + finalText: `Current delivery unavailable.\n${summary}`, + }, + status, + ), + ).not.toEqual([]); + const input = idle(fixture()); + const last = input.allCalls.at(-1); + if (!last) throw new Error("Missing status."); + object(object(last.output).workflowData).delivery = { report }; + expect(deliveryIssues(input, status)).not.toEqual([]); +}); +test("all five genuine workflow cases remain report-only with neutral session context", () => { + expect(DELIVERY_SCENARIOS).toHaveLength(5); + const catalog = caseCatalogFor(DELIVERY_SCENARIOS, { + kind: "ordinary", + repeat: 1, + }); + expect(catalog.every((entry) => entry.release === "report-only")).toBe(true); + for (const scenario of DELIVERY_SCENARIOS) { + expect(scenario.title).toBe("Text workspace"); + for (const step of scenario.steps) { + const text = "kind" in step ? step.prompt : step.arguments; + expect(text).not.toContain(scenario.id); + expect(text).not.toMatch(/\beval(?:uation)?\b/i); + } + } +}); + +function deferredFixture() { + const input = fixture(); + const archive = object(input.archives[0]); + object(archive.closure).kind = "deferred"; + const plan = object(archive.plan); + plan.evidence = [ + { scope: "gate", command: expected.gate, platform: "linux" }, + { + scope: "extra", + command: "node scripts/platform-check.mjs", + platform: "darwin", + }, + ]; + const runs = archive.runs; + if (!Array.isArray(runs)) throw new Error("Missing runs."); + for (const run of runs) { + object(run).state = "blocked"; + object(run).reviews = []; + object(run).validations = []; + } + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing close."); + object(close.input.request).kind = "deferred"; + const data = object(object(close.output).workflowData); + object(data.operation).entity = archive.closure; + data.delivery = { + assurance: { conclusion: "completion-not-claimed" }, + report: report.map((line) => + line + .replace("Closure: completed", "Closure: deferred") + .replace("1 of 1", "0 of 1"), + ), + }; + const text = [ + "Goal: Implement text behavior", + "Closure: deferred", + "External action authority: not-granted", + "Progress: 0 of 1 features complete", + "Unfinished features: parser", + "Assurance: completion not claimed", + "External macOS proof unavailable.", + ...limits, + ].join("\n"); + const { observed: _observed, ...unobserved } = expected; + const deferred: DeliveryExpectation = { + ...unobserved, + closure: "deferred", + missingEvidenceCommand: "node scripts/platform-check.mjs", + }; + return { input, text, expected: deferred }; +} + +test("deferral keeps unavailable proof and unfinished identities without false completion", () => { + const { input, text, expected: deferred } = deferredFixture(); + const plan = object(object(input.archives[0]).plan); + expect(deliveryIssues({ ...input, finalText: text }, deferred)).toEqual([]); + expect( + deliveryIssues( + { + ...input, + finalText: text.replace("External macOS proof unavailable.", ""), + }, + deferred, + ), + ).not.toEqual([]); + expect( + deliveryIssues( + { ...input, finalText: text.replace("parser", "") }, + deferred, + ), + ).not.toEqual([]); + expect( + deliveryIssues( + { + ...input, + finalText: text.replace( + "completion not claimed", + "completion supported", + ), + }, + deferred, + ), + ).not.toEqual([]); + plan.evidence = []; + expect(deliveryIssues({ ...input, finalText: text }, deferred)).not.toEqual( + [], + ); +}); + +test("an exact accepted close replay preserves full-detail access without new work", () => { + const input = idle(fixture()); + const original = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + const replay = input.allCalls.at(-1); + const trace = input.hostTrace; + if (!original || !replay?.native || trace?.kind !== "observed") + throw new Error("Missing native calls."); + const replayCall = { + ...replay, + tool: "flow_session_close", + input: structuredClone(original.input), + output: structuredClone(original.output), + }; + object(object(object(replayCall.output).workflowData).operation).replayed = + true; + const message = trace.messages.find( + (item) => item.id === replay.native?.messageId, + ); + if (message?.role !== "assistant") throw new Error("Missing native message."); + const tool = message.tools.find( + (item) => item.partId === replay.native?.partId, + ); + if (!tool) throw new Error("Missing native part."); + tool.tool = "flow_session_close"; + expect( + deliveryIssues( + { + ...input, + allCalls: input.allCalls.map((call) => + call === replay ? replayCall : call, + ), + flowCalls: input.flowCalls.map((call) => + call === replay ? replayCall : call, + ), + finalText: report.join("\n"), + }, + { ...expected, presentation: "full" }, + ), + ).toEqual([]); +}); + +test("handoff text cannot mask an ungranted Git mutation", () => { + const input = fixture(); + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing close."); + const mutation = { + ...close, + tool: "bash", + input: { command: "git push origin HEAD" }, + }; + expect( + deliveryIssues( + { ...input, allCalls: [...input.allCalls, mutation] }, + expected, + ), + ).not.toEqual([]); +}); + +for (const missing of [ + "Goal: Implement text behavior", + "Assurance: completion supported", +]) + test(`summary rejects omitted ${missing}`, () => { + const input = fixture(); + expect( + deliveryIssues( + { ...input, finalText: input.finalText.replace(missing, "") }, + expected, + ), + ).not.toEqual([]); + }); + +test("summary rejects mutated assurance conclusion and unrelated goal", () => { + const input = fixture(); + expect( + deliveryIssues( + { + ...input, + finalText: input.finalText.replace( + "completion supported", + "completion unsupported", + ), + }, + expected, + ), + ).not.toEqual([]); + expect( + deliveryIssues( + { + ...input, + finalText: input.finalText.replace( + "Implement text behavior", + "Ship an unrelated feature", + ), + }, + expected, + ), + ).not.toEqual([]); + const close = input.allCalls.find( + (call) => call.tool === "flow_session_close", + ); + if (!close) throw new Error("Missing close fixture."); + object(object(object(close.output).workflowData).delivery).assurance = { + conclusion: "completion-unsupported", + }; + expect(deliveryIssues(input, expected)).not.toEqual([]); +}); + +test("summary rejects oversized narration and reordered full histories after checking critical facts", () => { + const input = fixture(); + const verbose = `${input.finalText}\n${"A lengthy unrelated account of the completed work. ".repeat(50)}`; + expect(deliveryIssues({ ...input, finalText: verbose }, expected)).toContain( + "Default summary is not shorter than the actual full report.", + ); + const reordered = `${input.finalText}\n${[...report].reverse().join("\n")}`; + expect( + deliveryIssues({ ...input, finalText: reordered }, expected), + ).toContain("Default summary is not shorter than the actual full report."); +}); + +for (const contradictory of [ + "Assurance: completion supported.", + "Closure: completed.", + "External action authority: granted.", + "Progress: 1 of 1 features complete.", + "Goal: Ship unrelated work.", +]) { + test(`current deferred facts reject ${contradictory}`, () => { + const { input, text, expected: deferred } = deferredFixture(); + expect( + deliveryIssues( + { ...input, finalText: `${text}\n${contradictory}` }, + deferred, + ), + ).not.toEqual([]); + }); +} + +test("explicit historical handoff references do not contradict current deferred facts", () => { + const { input, text, expected: deferred } = deferredFixture(); + const historical = `Historical:\nClosure: completed.\nAssurance: completion supported.\nCurrent delivery:\n${text}`; + expect(deliveryIssues({ ...input, finalText: historical }, deferred)).toEqual( + [], + ); + expect( + deliveryIssues( + { ...input, finalText: `${text}\nHistorical closure: completed.` }, + deferred, + ), + ).toEqual([]); +}); + +test("pending planned work may defer without inventing approved completion", () => { + const { input, text, expected: deferred } = deferredFixture(); + object(input.archives[0]).approval = "pending"; + expect(deliveryIssues({ ...input, finalText: text }, deferred)).toEqual([]); + const completed = fixture(); + object(completed.archives[0]).approval = "pending"; + expect(deliveryIssues(completed, expected)).toContain( + "Completed closure requires an approved plan.", + ); +}); + +function pilotFixture(kind: "audit" | "full" | "deferred") { + const input = kind === "deferred" ? deferredFixture().input : fixture(); + const recorded = pilot[kind]; + object(input.archives[0]).goal = recorded.goal; + for (const call of input.allCalls) { + if (call.tool === "flow_plan_save") + object(call.input.request).goal = recorded.goal; + if ( + call.tool === "flow_status" && + object(call.input.request).view === "reviewer" + ) { + object(object(object(call.output).workflowData).projection).goal = + recorded.goal; + object(call).rawOutput = JSON.stringify(call.output); + } + if (call.tool === "flow_session_close") + object(object(object(call.output).workflowData).delivery).report = + recorded.report; + } + if (kind !== "deferred") + expect(checkReviewerEvidenceAccess(input, input.archives[0])).toEqual([]); + return input; +} + +test("saved faithful Markdown full detail preserves every canonical record", () => { + const input = pilotFixture("full"); + expect( + deliveryIssues( + { ...input, finalText: pilot.full.answer }, + { ...expected, presentation: "full" }, + ), + ).toEqual([]); +}); + +test("saved audit fails only two omitted disclosures and clarified exact goal fidelity", () => { + const input = pilotFixture("audit"); + const issues = deliveryIssues( + { ...input, finalText: pilot.audit.answer }, + expected, + ); + expect(issues).toEqual([ + "Recorded goal identity was omitted, changed, or contradicted.", + "Missing assurance disclosure: caller bindings do not prove completeness or fitness.", + "Missing assurance disclosure: coverage and review substance remain model judgments.", + ]); + const corrected = `${pilot.audit.answer}\nGoal: ${pilot.audit.goal}\n${limits[0]}\n${limits[1]}`; + expect(deliveryIssues({ ...input, finalText: corrected }, expected)).toEqual( + [], + ); +}); + +for (const claim of [ + "1. Closure: completed", + "Closure: deferred; current closure is completed.", + "Historical report\nThe current workflow is completed.", + "External action authority is granted. You may deploy this now.", + "The required macOS validation passed with exit 0.", + "The current workflow is complete and ready to ship.", + "Completion is supported by all checks.", +]) + test(`bounded current assertions reject ${claim}`, () => { + const { input, text, expected: deferred } = deferredFixture(); + expect( + deliveryIssues({ ...input, finalText: `${text}\n${claim}` }, deferred), + ).not.toEqual([]); + }); + +for (const [name, before, after] of [ + ["path", "src/parser.mjs", "src/other.mjs"], + ["attempt count", "Attempts:** 1", "Attempts:** 2"], + ["latest state", "Latest state:** completed", "Latest state:** blocked"], + ["check tier", "[host-attested]", "[TS-enforced]"], + ["outcome", "null returns an empty string", "null returns null"], + ["scope", "only src/parser.mjs changed", "other files also changed"], +] as const) + test(`full report rejects changed ${name} record`, () => { + const input = pilotFixture("full"); + const altered = pilot.full.answer.replace(before, after); + expect(altered).not.toBe(pilot.full.answer); + expect( + deliveryIssues( + { ...input, finalText: altered }, + { ...expected, presentation: "full" }, + ), + ).not.toEqual([]); + }); + +test("full report preserves record multiplicity and rejects unknown injected content", () => { + const input = pilotFixture("full"); + const check = pilot.full.answer + .split("\n") + .find((line) => line.includes("Accepted validation")); + if (!check) throw new Error("Missing saved check record."); + const duplicate = pilot.full.answer.replace(check, `${check}\n${check}`); + expect( + deliveryIssues( + { ...input, finalText: duplicate }, + { ...expected, presentation: "full" }, + ), + ).not.toEqual([]); + expect( + deliveryIssues( + { + ...input, + finalText: `${pilot.full.answer}\nOutcome: extra work was released.`, + }, + { ...expected, presentation: "full" }, + ), + ).not.toEqual([]); +}); + +function savedDeferred() { + const input = pilotFixture("deferred"); + return { + input, + expected: deferredFixture().expected, + original: pilot.deferred.answer, + normalized: pilot.deferred.answer.replace( + /^Closure:.*$/m, + "Closure: deferred. Required macOS evidence is unavailable.", + ), + }; +} + +test("saved deferred handoff and scoped canonical explanation remain valid", () => { + const { input, expected: deferred, original, normalized } = savedDeferred(); + expect(deliveryIssues({ ...input, finalText: original }, deferred)).toEqual( + [], + ); + expect(deliveryIssues({ ...input, finalText: normalized }, deferred)).toEqual( + [], + ); +}); + +for (const [name, addition, removeLimits] of [ + [ + "disclosure inversions", + "Artifact paths and the canonical gate are NOT caller declarations; the claim that Flow validates binding, not completeness or fitness is false.\nGoal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance are NOT model judgments.\nIt is false that an archive does not attest the current workspace.", + true, + ], + [ + "scattered disclosure terms", + "Glossary: artifact paths; canonical gate; caller declarations; binding; not completeness or fitness; goal alignment; scope discipline; evidence completeness; requirement coverage; test adequacy; review substance; model judgments; archive does not attest the current workspace.", + true, + ], + [ + "historical-only disclosures", + `Historical report\n${limits.join("\n")}`, + true, + ], + [ + "compound closure comma", + "Closure: deferred, but the workflow is completed.", + false, + ], + ["compound closure and", "Closure: deferred and completed.", false], + [ + "compound authority", + "External action authority: not-granted and you may deploy now.", + false, + ], + [ + "compound assurance", + "Assurance: completion not claimed and completion is supported.", + false, + ], + [ + "compound progress", + "Progress: 0 of 1 features complete and 1 of 1 features complete.", + false, + ], + [ + "compound proof", + "Closure: deferred and required macOS validation passed.", + false, + ], +] as const) + test(`current records reject ${name}`, () => { + const { input, expected: deferred, normalized } = savedDeferred(); + expect( + deliveryIssues({ ...input, finalText: normalized }, deferred), + ).toEqual([]); + const baseline = removeLimits + ? normalized + .split("\n") + .filter((line) => !limits.some((limit) => line.includes(limit))) + .join("\n") + : normalized; + expect( + deliveryIssues( + { ...input, finalText: `${baseline}\n${addition}` }, + deferred, + ), + ).not.toEqual([]); + }); + +for (const opposite of [ + "Caller declarations prove completeness and fitness.", + "Coverage and review substance are objective proof.", + "The archive attests the current workspace.", +]) + test(`current disclosure rejects ${opposite}`, () => { + const input = fixture(); + expect( + deliveryIssues( + { ...input, finalText: `${input.finalText}\n${opposite}` }, + expected, + ), + ).not.toEqual([]); + }); + +test("second feature identities and titles with colons retain case-sensitive fidelity", () => { + const full = [ + "Goal: Guard parser behavior", + "Closure: completed", + "Progress: 2 of 2 features complete", + "Features:", + "- first — Primary: null guard", + " attempts: 1; latest state: completed", + " outcome: Kept src/parser.mjs unchanged except the null guard.", + " terminal findings: none", + "- second — Extra: trim scope", + " attempts: 1; latest state: completed", + " outcome: Preserved exact trimming behavior in src/parser_extra.mjs.", + " terminal findings: none", + "Assurance: completion supported", + ]; + const rendered = full.join("\n"); + expect(fullReportMatches(rendered, full)).toBe(true); + expect( + fullReportMatches( + rendered.replace("second — Extra", "Second — Extra"), + full, + ), + ).toBe(false); + expect( + fullReportMatches( + rendered.replace("second — Extra", "second — extra"), + full, + ), + ).toBe(false); +}); + +for (const claim of [ + "The workflow is completed and deferred.", + "Closure: deferred; completed.", + "External action authority: not-granted; you may deploy this now.", + "External action authority: not-granted; you may deploy this now, go ahead.", +]) + test(`whole current claim rejects ${claim}`, () => { + const { input, text, expected: deferred } = deferredFixture(); + expect( + deliveryIssues({ ...input, finalText: `${text}\n${claim}` }, deferred), + ).not.toEqual([]); + }); + +test("dash-prefixed outcomes retain their feature association", () => { + const report = [ + "Features:", + "- first — First feature", + " outcome: — Alpha result", + "- second — Second feature", + " outcome: — Beta result", + ]; + expect(fullReportMatches(report.join("\n"), report)).toBe(true); + expect( + fullReportMatches( + [ + "Features:", + "- first — First feature", + " outcome: — Beta result", + "- second — Second feature", + " outcome: — Alpha result", + ].join("\n"), + report, + ), + ).toBe(false); +}); + +for (const contrary of [ + "Artifact paths and the canonical gate are not caller declarations.", + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance do not remain model judgments.", +]) + test(`canonical disclosures do not excuse ${contrary}`, () => { + const input = fixture(); + expect( + deliveryIssues( + { ...input, finalText: `${input.finalText}\n${contrary}` }, + expected, + ), + ).not.toEqual([]); + }); diff --git a/tests/eval-progress-wait.test.ts b/tests/eval-progress-wait.test.ts index 2093c4b2..d4b6b21b 100644 --- a/tests/eval-progress-wait.test.ts +++ b/tests/eval-progress-wait.test.ts @@ -12,6 +12,7 @@ import { type Mode = | "owned-progress" | "wedged" + | "in-place-progress" | "wrong-parent" | "wrong-directory" | "deadline" @@ -131,7 +132,7 @@ async function observeWait(mode: Mode) { Array.from( { length: - mode === "wedged" + mode === "wedged" || mode === "in-place-progress" ? 1 : Math.floor(Math.min(now, finishes ? finishesAt : now) / 16_000) + 1, @@ -155,7 +156,10 @@ async function observeWait(mode: Mode) { state: { status: "completed", input: { filePath: "src/index.ts" }, - output: `source page ${index}`, + output: + mode === "in-place-progress" + ? `source page updated at ${now}` + : `source page ${index}`, }, }, ], @@ -328,7 +332,7 @@ for (const mode of [ const observed = await observeWait(mode); expect(observed.result).toBeInstanceOf(Error); expect(String(observed.result)).toContain( - "Scenario made no progress for 180000ms: wedged", + "Scenario had no new messages or parts for 180000ms", ); expect(observed.now).toBe(182_000); expect(observed.aborts).toBe(1); @@ -340,12 +344,25 @@ for (const mode of [ test("continuous owned reviewer progress retains the twenty-minute hard deadline", async () => { const observed = await observeWait("deadline"); expect(String(observed.result)).toContain( - "Scenario exceeded 1200000ms without going quiet: still working", + "New messages or parts continued near the deadline", ); expect(observed.now).toBe(1_202_000); expect(observed.aborts).toBe(1); }); +test("count-only stall diagnostics disclose unmeasured in-place updates", async () => { + const observed = await observeWait("in-place-progress"); + expect(String(observed.result)).toContain( + "Scenario had no new messages or parts for 180000ms", + ); + expect(String(observed.result)).toContain( + "Updates inside existing parts are not measured.", + ); + expect(observed.now).toBe(182_000); + expect(observed.aborts).toBe(1); + expect(observed.childReads).toBeGreaterThan(80); +}); + for (const mode of ["duplicate-children", "cyclic-children"] as const) { test(`${mode} remain one bounded owned progress path`, async () => { const observed = await observeWait(mode); diff --git a/tests/eval-replay-capability.test.ts b/tests/eval-replay-capability.test.ts new file mode 100644 index 00000000..7044fce4 --- /dev/null +++ b/tests/eval-replay-capability.test.ts @@ -0,0 +1,89 @@ +import { expect, test } from "bun:test"; +import { mkdtemp, readFile, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +async function run( + accept: boolean, + diverge = false, + scenario = "delivery-summary-deferred", +) { + const directory = await mkdtemp(join(tmpdir(), "flow-replay-capability-")); + const fixture = JSON.parse( + await readFile( + new URL( + "../evals/cassettes/plan-only-stops--fixture_hand-written--worker.json", + import.meta.url, + ), + "utf8", + ), + ); + if (diverge) fixture.events[0].observed.status = "error"; + const path = join(directory, "native-required.json"); + const original = `${JSON.stringify({ ...fixture, scenario, fidelity: [] }, null, 2)}\n`; + await writeFile(path, original); + try { + const child = Bun.spawn( + [ + process.execPath, + "evals/replay-run.ts", + "--from", + directory, + ...(accept ? ["--accept"] : []), + ], + { cwd: join(import.meta.dir, ".."), stdout: "pipe", stderr: "pipe" }, + ); + const [exitCode, stdout, stderr] = await Promise.all([ + child.exited, + new Response(child.stdout).text(), + new Response(child.stderr).text(), + ]); + return { + exitCode, + stdout, + stderr, + unchanged: (await readFile(path, "utf8")) === original, + }; + } finally { + await rm(directory, { recursive: true, force: true }); + } +} + +for (const scenario of [ + "delivery-summary-completed", + "delivery-summary-deferred", + "delivery-summary-observed-failure", + "delivery-full-detail-followup", + "delivery-idle-after-close", +]) + test(`${scenario} cannot gate native checking through empty legacy fidelity`, async () => { + const result = await run(false, false, scenario); + expect(result.exitCode).toBe(0); + expect(result.stdout).toContain("UNSUPPORTED"); + expect(result.stdout).toContain("native-host-provenance-unreplayed"); + expect(result.stdout).toContain("0/0 gated cassette(s) reproduced"); + expect(result.unchanged).toBe(true); + }); + +test("accept refuses unavailable native evidence and preserves original bytes", async () => { + const result = await run(true); + expect(result.exitCode).toBe(1); + expect(result.stdout).toContain("refused"); + expect(result.unchanged).toBe(true); +}); + +test("unsupported report fidelity still displays runtime handler divergence", async () => { + const result = await run(false, true); + expect(result.exitCode).toBe(0); + expect(result.stdout).toContain("UNSUPPORTED"); + expect(result.stdout).toContain("recorded error, replayed ok"); + expect(result.unchanged).toBe(true); +}); + +test("accept refuses a missing checker without rewriting its expectation", async () => { + const result = await run(true, false, "missing-scenario"); + expect(result.exitCode).toBe(1); + expect(result.stdout).toContain("NO-SCENARIO"); + expect(result.stdout).toContain("refused"); + expect(result.unchanged).toBe(true); +}); diff --git a/tests/eval-replay.test.ts b/tests/eval-replay.test.ts index 1e994237..953e1eaf 100644 --- a/tests/eval-replay.test.ts +++ b/tests/eval-replay.test.ts @@ -16,6 +16,7 @@ import { buildCassette, CASSETTE_VERSION, type Cassette, + cassetteFidelity, cassetteFileName, isGated, normalizeRecorded, @@ -377,6 +378,47 @@ describe("decision-layer replay", () => { expect(cassette.fidelity).not.toContain("host-error"); }); + test("scenario requirements retain native provenance limits during recording", () => { + const required = scenario("delivery-summary-deferred").replayRequires; + if (!required) throw new Error("Missing native replay requirement."); + const original = happyPathCassette(); + expect(cassetteFidelity(original, required)).toEqual([ + "native-host-provenance-unreplayed", + ]); + expect(isGated(original, required)).toBe(false); + expect(isGated(original)).toBe(true); + const recorded = buildCassette({ + flowVersion: "test", + scenario: "delivery-summary-deferred", + model: "provider/model", + attempt: 1, + hostPlatform: "linux", + files: {}, + projectPath: "/workspace", + calls: [], + finalText: "", + assistantMessages: 0, + verdict: "PASS", + issues: [], + falseCompletion: false, + documents: [], + extraFidelity: [], + replayRequires: required, + }); + expect(recorded.fidelity).toEqual([ + "no-flow-calls", + "native-host-provenance-unreplayed", + ]); + }); + + test("decision replay does not manufacture native host provenance", async () => { + const replayed = await replayCassette(happyPathCassette()); + expect(replayed.outcome.hostTrace).toBeUndefined(); + expect( + replayed.outcome.allCalls.every((call) => call.native === undefined), + ).toBe(true); + }); + test("retains the derived named-result witness for replay", () => { const command = "bun test --reporter=junit --reporter-outfile=.flow/results.xml"; diff --git a/tests/eval-report.test.ts b/tests/eval-report.test.ts index d39d37b0..a531b6ca 100644 --- a/tests/eval-report.test.ts +++ b/tests/eval-report.test.ts @@ -7,6 +7,7 @@ import { parseReport, type ReportIssue, } from "../evals/report.js"; +import { scenarioStepInstruction } from "../evals/scenario-steps.js"; const digest = (letter: string) => `sha256:${letter.repeat(64)}`; @@ -1028,3 +1029,22 @@ describe("eval report boundary", () => { ); }); }); + +test("user prompt provenance freezes bytes while preserving old command records", () => { + const fixture = report(); + const attempt = fixture.attempts[0]; + if (!attempt) throw new Error("Missing report fixture attempt."); + attempt.instructions = [ + scenarioStepInstruction( + { kind: "prompt", prompt: " Show the full report. 漢\n" }, + 0, + ), + ]; + expect(parseReport(fixture, caseCatalog()).ok).toBe(true); + const instruction = attempt.instructions[0]; + if (!instruction) throw new Error("Missing prompt instruction."); + instruction.text = "Show another report."; + expect(parseReport(fixture, caseCatalog()).ok).toBe(false); + delete instruction.text; + expect(parseReport(fixture, caseCatalog()).ok).toBe(false); +}); diff --git a/tests/eval-reporting.test.ts b/tests/eval-reporting.test.ts index b6ed0800..6fb0500a 100644 --- a/tests/eval-reporting.test.ts +++ b/tests/eval-reporting.test.ts @@ -682,7 +682,7 @@ describe("eval campaign cancellation", () => { expect(error.message).toContain( failure === "timeout" ? "Scenario exceeded 0ms" - : "Scenario made no progress for 0ms", + : "Scenario had no new messages or parts for 0ms", ); expect(aborts).toBe(1); } finally { diff --git a/tests/eval-scenario-checks.test.ts b/tests/eval-scenario-checks.test.ts index a1b75951..8338bf77 100644 --- a/tests/eval-scenario-checks.test.ts +++ b/tests/eval-scenario-checks.test.ts @@ -1615,21 +1615,20 @@ describe("inspection-failed-audit-completes", () => { (entry) => entry.id === "inspection-failed-audit-completes", ); if (!scenario) throw new Error("Expected the inspection scenario."); + const step = scenario.steps[0]; + if (!step || "kind" in step) + throw new Error("Expected inspection command."); expect(Object.hasOwn(scenario.files, "docs/README.md")).toBe(true); - expect(scenario.steps[0]?.arguments).toContain( - "observed count and severity", - ); - expect(scenario.steps[0]?.arguments).not.toContain("21"); - expect(scenario.steps[0]?.arguments).toContain( + expect(step.arguments).toContain("observed count and severity"); + expect(step.arguments).not.toContain("21"); + expect(step.arguments).toContain( "Finding: inclusiveRangeLength is incorrect for 1..3.", ); - expect(scenario.steps[0]?.arguments).toContain( + expect(step.arguments).toContain( "specific defect or audit target in each phase", ); - expect(scenario.steps[0]?.arguments).toContain( - "Number at least two phases as 1. and 2.", - ); - expect(scenario.steps[0]?.arguments).toContain("State a concrete action"); + expect(step.arguments).toContain("Number at least two phases as 1. and 2."); + expect(step.arguments).toContain("State a concrete action"); }); function recordedOutcome(overrides: Partial = {}): Outcome { const document = session({ diff --git a/tests/fixtures/auto-qualified-outcome.ts b/tests/fixtures/auto-qualified-outcome.ts index 62d8a0b0..5f9c1460 100644 --- a/tests/fixtures/auto-qualified-outcome.ts +++ b/tests/fixtures/auto-qualified-outcome.ts @@ -5,8 +5,10 @@ import { collectReviewerPacketBytes } from "../../evals/reviewer-packet-bytes.js const digest = `sha256:${"a".repeat(64)}`; export function autoQualifiedOutcome( - kind: "two" | "prerequisite" | "audit", + kind: "two" | "prerequisite" | "audit" | "single", + options: Readonly<{ goal?: string; featureId?: string }> = {}, ): ScenarioGradeInput { + const goal = options.goal ?? "Implement text behavior"; const features = kind === "two" ? [ @@ -29,7 +31,7 @@ export function autoQualifiedOutcome( ] : [ { - id: "parser", + id: options.featureId ?? "parser", title: "Parser", summary: "Guard null", targets: @@ -190,7 +192,7 @@ export function autoQualifiedOutcome( view: "reviewer", sessionId: "session:qualification", revision: created, - goal: "Implement text behavior", + goal, planContext: plan, feature, assignment: { ...assignment, result: null }, @@ -299,7 +301,7 @@ export function autoQualifiedOutcome( version: 5, id: "session:qualification", revision: 12, - goal: "Implement text behavior", + goal, approval: "approved", plan, runs, diff --git a/tests/fixtures/delivery-confirmation-answers.json b/tests/fixtures/delivery-confirmation-answers.json new file mode 100644 index 00000000..ed34cab8 --- /dev/null +++ b/tests/fixtures/delivery-confirmation-answers.json @@ -0,0 +1,265 @@ +{ + "measuredSource": "c92093674de66f3d292a11eb9bc701b3a52f825b", + "campaign": "2026-10-04T22-54-27-074Z.v2", + "cases": { + "delivery-summary-observed-failure": { + "answer": "Completed and archived the Flow session.\n\nGoal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs, with audit observation and separate required Linux verification.\n\nUpdated `src/parser.mjs`:\n```js\nexport function parse(input) { return input === null ? '' : input.trim(); }\n```\n\n- **Required Linux verification:** `node scripts/verify.mjs` passed, exit **0**.\n- **Linux audit observation:** `node scripts/audit.mjs` exited **12**, reporting 12 outstanding advisory items. This observation does not claim a pass.\n- **Independent review:** passed.\n- Advisory data and both scripts remain unchanged.\n- **Progress:** 1 of 1 features complete; no unfinished features or blockers. One nonblocking advisory remains.\n\n**Flow handoff:** format 1; external action authority not granted. Completion is supported, with all four assurance checks satisfied. Reported artifacts: one latest, none superseded; these are caller declarations, not an exhaustive Git delta.\n\n**Assurance limitations**\n- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.\n- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.\n- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "goal": "Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs, with audit observation and separate required Linux verification.", + "featureId": "parser-null", + "report": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs, with audit observation and separate required Linux verification.", + "Closure: completed — Implemented null guard only in src/parser.mjs; Linux verification passed, audit exit 12 retained as advisory; independent review passed.", + "Progress: 1 of 1 features complete", + "Features:", + "- parser-null — Null-safe parser with preserved trimming", + " attempts: 1; latest state: completed", + " outcome: Verified parser-null and all five approved requirements: null returns ''; string trimming and other-input behavior are preserved; only src/parser.mjs changed with regular-file mode 100644 unchanged; advisory data, both scripts and pre-existing opencode.json are preserved; Linux audit observation retains exit 12 as advisory; separate source-bound broad verification passed with exit 0, and independent final review is complete. Inspected current source, downstream verification, README/package surfaces and complete host baseline/current diff. No deletions, renames, generated artifacts, scope drift or applicable persistence/concurrency/migration risks. No prior findings.", + " terminal findings:", + " - advisory: Observe-only Linux audit retains 12 outstanding advisory items (exit 12); no repair is authorized.", + "Findings digest:", + "- live parser-null parser-null.R6-01 advisory: Observe-only Linux audit retains 12 outstanding advisory items (exit 12); no repair is authorized.", + "Observed \"node scripts/audit.mjs\": exit 12, host linux, source sha256:4b75cc8c7c24a4f1d108ed8d95d682b913cc566e47bd59d2bf21cc8192fa8dae, output sha256:c6793caf5f3c67b895af8671e9d3705639e0d1e4590377e84b77db953ebe017d; this does not claim the command passed.", + "Assurance: completion supported", + "Assurance checks:", + "- satisfied [TS-enforced] Recorded completion: 1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.", + "- satisfied [host-attested] Accepted validation: 1/1 terminal runs have eligible host evidence accepted by review.", + "- satisfied [host-attested] Canonical gate: \"node scripts/verify.mjs\" must have passing broad evidence accepted by review.", + "- satisfied [host-attested] Declared evidence: 1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.", + "Assurance limitations:", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Artifacts as reported by Flow from caller declarations, not an exact or exhaustive Git delta:", + "- latest attempts: src/parser.mjs", + "- superseded attempts only: none reported" + ], + "summary": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs, with audit observation and separate required Linux verification.", + "Closure: completed. Implemented null guard only in src/parser.mjs; Linux verification passed, audit exit 12 retained as advisory; independent review passed.", + "Progress: 1 of 1 features complete", + "Unfinished features: none", + "Nonblocking live findings: advisory 1. Historical findings: 0.", + "Observed \"node scripts/audit.mjs\": exit 12, host linux; this does not claim the command passed.", + "Assurance: completion supported", + "Assurance checks: 4 satisfied, 0 not applicable, 0 unsatisfied.", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Reported artifacts: 1 latest, 0 superseded only. Caller declarations, not an exact or exhaustive Git delta.", + "Full report is included in this close response." + ], + "assurance": { + "checks": [ + { + "explanation": "1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.", + "id": "recorded-completion", + "label": "Recorded completion", + "status": "satisfied", + "tier": "ts-enforced" + }, + { + "explanation": "1/1 terminal runs have eligible host evidence accepted by review.", + "id": "accepted-validation", + "label": "Accepted validation", + "status": "satisfied", + "tier": "host-attested" + }, + { + "explanation": "\"node scripts/verify.mjs\" must have passing broad evidence accepted by review.", + "id": "canonical-gate", + "label": "Canonical gate", + "status": "satisfied", + "tier": "host-attested" + }, + { + "explanation": "1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.", + "id": "declared-evidence", + "label": "Declared evidence", + "status": "satisfied", + "tier": "host-attested" + } + ], + "conclusion": "completion-supported", + "limitations": [ + "Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "Freshness holds when review is accepted; an archive does not attest the current workspace." + ] + } + }, + "delivery-summary-completed": { + "answer": "Completed and archived the Flow session.\n\nGoal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping node scripts/verify.mjs and its script unchanged.\n\n- **Verified:** `parse(null)` returns `''`; string inputs retain `trim()` behavior.\n- **Changed:** only `src/parser.mjs`.\n- **Validation:** unchanged `node scripts/verify.mjs` passed, exit code **0**.\n- **Independent review:** passed with no findings.\n- **Progress:** 1/1 features complete; no unfinished features or blockers.\n- **Assurance:** completion supported; all four assurance checks satisfied.\n\nFlow assurance limitations:\n- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.\n- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.\n- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "goal": "Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping node scripts/verify.mjs and its script unchanged.", + "featureId": "parser-null", + "report": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping node scripts/verify.mjs and its script unchanged.", + "Closure: completed — Implemented null guard only in src/parser.mjs; unchanged canonical verifier passed with exit 0, and independent review passed without findings.", + "Progress: 1 of 1 features complete", + "Features:", + "- parser-null — Handle null while preserving trimmed strings", + " attempts: 1; latest state: completed", + " outcome: Final review: parser-null verified. All approved requirements verified: null returns empty string; string trim behavior is preserved; only src/parser.mjs changed and the verifier and command remain unchanged. Inspected live parser, verifier, README, package surface, all host evidence pages, and complete baseline diff: no deletion, rename, type or mode changes; preexisting opencode.json preserved. Broad validation is complete, exit 0, and bound to assignment source. No prior findings or remaining gaps.", + " terminal findings: none", + "Findings digest: none", + "Assurance: completion supported", + "Assurance checks:", + "- satisfied [TS-enforced] Recorded completion: 1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.", + "- satisfied [host-attested] Accepted validation: 1/1 terminal runs have eligible host evidence accepted by review.", + "- satisfied [host-attested] Canonical gate: \"node scripts/verify.mjs\" must have passing broad evidence accepted by review.", + "- satisfied [host-attested] Declared evidence: 1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.", + "Assurance limitations:", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Artifacts as reported by Flow from caller declarations, not an exact or exhaustive Git delta:", + "- latest attempts: src/parser.mjs", + "- superseded attempts only: none reported" + ], + "summary": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping node scripts/verify.mjs and its script unchanged.", + "Closure: completed. Implemented null guard only in src/parser.mjs; unchanged canonical verifier passed with exit 0, and independent review passed without findings.", + "Progress: 1 of 1 features complete", + "Unfinished features: none", + "Nonblocking live findings: advisory 0. Historical findings: 0.", + "Assurance: completion supported", + "Assurance checks: 4 satisfied, 0 not applicable, 0 unsatisfied.", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Reported artifacts: 1 latest, 0 superseded only. Caller declarations, not an exact or exhaustive Git delta.", + "Full report is included in this close response." + ], + "assurance": { + "checks": [ + { + "explanation": "1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.", + "id": "recorded-completion", + "label": "Recorded completion", + "status": "satisfied", + "tier": "ts-enforced" + }, + { + "explanation": "1/1 terminal runs have eligible host evidence accepted by review.", + "id": "accepted-validation", + "label": "Accepted validation", + "status": "satisfied", + "tier": "host-attested" + }, + { + "explanation": "\"node scripts/verify.mjs\" must have passing broad evidence accepted by review.", + "id": "canonical-gate", + "label": "Canonical gate", + "status": "satisfied", + "tier": "host-attested" + }, + { + "explanation": "1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.", + "id": "declared-evidence", + "label": "Declared evidence", + "status": "satisfied", + "tier": "host-attested" + } + ], + "conclusion": "completion-supported", + "limitations": [ + "Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "Freshness holds when review is accepted; an archive does not attest the current workspace." + ] + } + }, + "delivery-summary-deferred": { + "answer": "**Deferred and archived.** The parser change is retained. The Linux gate passed with exit code 0; required macOS evidence remains **unfulfilled**.\n\nGoal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping verification scripts unchanged; validate locally on Linux and retain required macOS proof as outstanding until run on macOS.\n\n- **Closure:** deferred because macOS validation is unavailable.\n- **Progress:** 0 of 1 features complete; unfinished feature: `parser-null`.\n- **Outstanding proof:** `node scripts/platform-check.mjs` on macOS.\n- **Assurance:** completion not claimed; no independent review performed. Four completion assurance checks are not applicable.\n- **Findings:** no live advisory or historical findings.\n- **Reported artifacts:** 0 latest, 0 superseded-only; these are caller declarations, not an exhaustive Git delta.\n- **External action authority:** not granted. Handoff format: 1.\n\n**Assurance limitations**\n- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.\n- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.\n- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "goal": "Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping verification scripts unchanged; validate locally on Linux and retain required macOS proof as outstanding until run on macOS.", + "featureId": "parser-null", + "report": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping verification scripts unchanged; validate locally on Linux and retain required macOS proof as outstanding until run on macOS.", + "Closure: deferred — Deferred at user direction because required macOS validation is unavailable. Parser change retained and Linux gate passed; node scripts/platform-check.mjs on macOS remains unfulfilled. No independent review or completed outcome claimed.", + "Progress: 0 of 1 features complete", + "Features:", + "- parser-null — Null-safe string parser", + " attempts: 1; latest state: superseded", + " outcome: none recorded", + " terminal findings: none", + "Findings digest: none", + "Assurance: completion not claimed", + "Assurance checks:", + "- not-applicable [TS-enforced] Recorded completion: deferred closure makes no completion claim.", + "- not-applicable [host-attested] Accepted validation: deferred closure makes no completion claim.", + "- not-applicable [host-attested] Canonical gate: deferred closure makes no completion claim.", + "- not-applicable [host-attested] Declared evidence: deferred closure makes no completion claim.", + "Assurance limitations:", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Artifacts as reported by Flow from caller declarations, not an exact or exhaustive Git delta:", + "- latest attempts: none reported", + "- superseded attempts only: none reported" + ], + "summary": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping verification scripts unchanged; validate locally on Linux and retain required macOS proof as outstanding until run on macOS.", + "Closure: deferred. Deferred at user direction because required macOS validation is unavailable. Parser change retained and Linux gate passed; node scripts/platform-check.mjs on macOS remains unfulfilled. No independent review or completed outcome claimed.", + "Progress: 0 of 1 features complete", + "Unfinished features: parser-null", + "Nonblocking live findings: advisory 0. Historical findings: 0.", + "Assurance: completion not claimed", + "Assurance checks: 0 satisfied, 4 not applicable, 0 unsatisfied.", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Reported artifacts: 0 latest, 0 superseded only. Caller declarations, not an exact or exhaustive Git delta.", + "Full report is included in this close response." + ], + "assurance": { + "checks": [ + { + "explanation": "deferred closure makes no completion claim.", + "id": "recorded-completion", + "label": "Recorded completion", + "status": "not-applicable", + "tier": "ts-enforced" + }, + { + "explanation": "deferred closure makes no completion claim.", + "id": "accepted-validation", + "label": "Accepted validation", + "status": "not-applicable", + "tier": "host-attested" + }, + { + "explanation": "deferred closure makes no completion claim.", + "id": "canonical-gate", + "label": "Canonical gate", + "status": "not-applicable", + "tier": "host-attested" + }, + { + "explanation": "deferred closure makes no completion claim.", + "id": "declared-evidence", + "label": "Declared evidence", + "status": "not-applicable", + "tier": "host-attested" + } + ], + "conclusion": "completion-not-claimed", + "limitations": [ + "Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "Freshness holds when review is accepted; an archive does not attest the current workspace." + ] + } + } + } +} diff --git a/tests/fixtures/delivery-pilot-answers.json b/tests/fixtures/delivery-pilot-answers.json new file mode 100644 index 00000000..87fc2d4b --- /dev/null +++ b/tests/fixtures/delivery-pilot-answers.json @@ -0,0 +1,95 @@ +{ + "audit": { + "answer": "Completed and archived the Flow session. All requirements were verified, and independent review passed.\n\nChanged only `src/parser.mjs:1`:\n```js\nexport function parse(input) { return input === null ? '' : input.trim(); }\n```\n\n**Linux validation**\n- `node scripts/verify.mjs` — passed, exit 0.\n- `node scripts/audit.mjs` — recorded as an observation, exit 12, reporting 12 outstanding advisory items.\n\nAdvisory data and both scripts remain unchanged. The audit result was retained as a nonblocking advisory finding.\n\n**Flow handoff:** 1/1 features complete; none unfinished. Completion is supported by all four assurance checks. Flow’s archive records the accepted evidence and review; it does not attest later workspace changes or grant external-action authority.", + "report": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs; observe audit and require verify on Linux, then obtain independent review.", + "Closure: completed — Implemented null-safe parser in src/parser.mjs only; required Linux verify passed, audit exit 12 retained as advisory observation, independent review passed.", + "Progress: 1 of 1 features complete", + "Features:", + "- parser-null — Null-safe string parser", + " attempts: 1; latest state: completed", + " outcome: Verified parser-null and all five approved requirements: null returns ''; strings retain trim behavior; only src/parser.mjs changed with file type/mode unchanged and preexisting work preserved; Linux audit observation recorded separately from the passing required verify gate; independent final review completed. README and package surface are consistent. No prior findings or additional risk checklist. Retained the 12 outstanding audit items as advisory, without authorizing repairs.", + " terminal findings:", + " - advisory: Audit observation reports 12 outstanding advisory items; these remain outside the authorized parser repair.", + "Findings digest:", + "- live parser-null parser-null.R6-01 advisory: Audit observation reports 12 outstanding advisory items; these remain outside the authorized parser repair.", + "Observed \"node scripts/audit.mjs\": exit 12, host linux, source sha256:4b75cc8c7c24a4f1d108ed8d95d682b913cc566e47bd59d2bf21cc8192fa8dae, output sha256:c6793caf5f3c67b895af8671e9d3705639e0d1e4590377e84b77db953ebe017d; this does not claim the command passed.", + "Assurance: completion supported", + "Assurance checks:", + "- satisfied [TS-enforced] Recorded completion: 1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.", + "- satisfied [host-attested] Accepted validation: 1/1 terminal runs have eligible host evidence accepted by review.", + "- satisfied [host-attested] Canonical gate: \"node scripts/verify.mjs\" must have passing broad evidence accepted by review.", + "- satisfied [host-attested] Declared evidence: 1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.", + "Assurance limitations:", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Artifacts as reported by Flow from caller declarations, not an exact or exhaustive Git delta:", + "- latest attempts: src/parser.mjs", + "- superseded attempts only: none reported" + ], + "goal": "Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs; observe audit and require verify on Linux, then obtain independent review." + }, + "full": { + "answer": "## Full delivery report\n\n**Handoff format:** 1 \n**External action authority:** not-granted\n\n**Goal:** Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping node scripts/verify.mjs and its script unchanged; validate and obtain independent review.\n\n**Closure:** completed — Implemented null-safe parse in src/parser.mjs while preserving string trimming. Unchanged node scripts/verify.mjs passed; independent review passed without findings.\n\n**Progress:** 1 of 1 features complete\n\n### Features\n\n- **parser-null — Null-safe string parser**\n - **Attempts:** 1\n - **Latest state:** completed\n - **Outcome:** Verified parser-null and all three approved requirements: null returns an empty string, strings retain native trimming, and only src/parser.mjs changed while the verification script and command remain unchanged. Inspected current source, gate assertions, README and package surface; complete host-bound base diff preserves file mode/type and preexisting opencode.json, with no deletions or renames. Source-bound broad validation passed with complete output and exit code 0. Other input behavior remains unchanged. No prior findings or remaining gaps.\n - **Terminal findings:** none\n\n**Findings digest:** none\n\n### Assurance\n\n**Conclusion:** completion supported\n\n#### Assurance checks\n\n- **Satisfied [TS-enforced] — Recorded completion:** 1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.\n- **Satisfied [host-attested] — Accepted validation:** 1/1 terminal runs have eligible host evidence accepted by review.\n- **Satisfied [host-attested] — Canonical gate:** `\"node scripts/verify.mjs\"` must have passing broad evidence accepted by review.\n- **Satisfied [host-attested] — Declared evidence:** 1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.\n\n#### Assurance limitations\n\n- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.\n- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.\n- Freshness holds when review is accepted; an archive does not attest the current workspace.\n\n### Reported artifacts\n\nArtifacts as reported by Flow from caller declarations, not an exact or exhaustive Git delta:\n\n- **Latest attempts:** `src/parser.mjs`\n- **Superseded attempts only:** none reported", + "report": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping node scripts/verify.mjs and its script unchanged; validate and obtain independent review.", + "Closure: completed — Implemented null-safe parse in src/parser.mjs while preserving string trimming. Unchanged node scripts/verify.mjs passed; independent review passed without findings.", + "Progress: 1 of 1 features complete", + "Features:", + "- parser-null — Null-safe string parser", + " attempts: 1; latest state: completed", + " outcome: Verified parser-null and all three approved requirements: null returns an empty string, strings retain native trimming, and only src/parser.mjs changed while the verification script and command remain unchanged. Inspected current source, gate assertions, README and package surface; complete host-bound base diff preserves file mode/type and preexisting opencode.json, with no deletions or renames. Source-bound broad validation passed with complete output and exit code 0. Other input behavior remains unchanged. No prior findings or remaining gaps.", + " terminal findings: none", + "Findings digest: none", + "Assurance: completion supported", + "Assurance checks:", + "- satisfied [TS-enforced] Recorded completion: 1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.", + "- satisfied [host-attested] Accepted validation: 1/1 terminal runs have eligible host evidence accepted by review.", + "- satisfied [host-attested] Canonical gate: \"node scripts/verify.mjs\" must have passing broad evidence accepted by review.", + "- satisfied [host-attested] Declared evidence: 1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.", + "Assurance limitations:", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Artifacts as reported by Flow from caller declarations, not an exact or exhaustive Git delta:", + "- latest attempts: src/parser.mjs", + "- superseded attempts only: none reported" + ], + "goal": "Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs and keeping node scripts/verify.mjs and its script unchanged; validate and obtain independent review." + }, + "deferred": { + "answer": "The Flow session is **deferred and archived**. The change in `src/parser.mjs` is retained, and the Linux gate passed. Required macOS evidence remains **unfulfilled**; final review and completion are not claimed.\n\n### Flow delivery\n```text\nHandoff format: 1\nExternal action authority: not-granted\nGoal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs, with Linux validation and required external macOS proof.\nClosure: deferred. Deferred at user direction because required macOS validation is unavailable. src/parser.mjs null guard implemented and Linux node scripts/verify.mjs passed. Required macOS node scripts/platform-check.mjs evidence remains unfulfilled; no final review or completion claimed.\nProgress: 0 of 1 features complete\nUnfinished features: parser-null\nNonblocking live findings: advisory 0. Historical findings: 0.\nAssurance: completion not claimed\nAssurance checks: 0 satisfied, 4 not applicable, 0 unsatisfied.\n- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.\n- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.\n- Freshness holds when review is accepted; an archive does not attest the current workspace.\nReported artifacts: 0 latest, 0 superseded only. Caller declarations, not an exact or exhaustive Git delta.\nFull report is included in this close response.\n```", + "report": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs, with Linux validation and required external macOS proof.", + "Closure: deferred — Deferred at user direction because required macOS validation is unavailable. src/parser.mjs null guard implemented and Linux node scripts/verify.mjs passed. Required macOS node scripts/platform-check.mjs evidence remains unfulfilled; no final review or completion claimed.", + "Progress: 0 of 1 features complete", + "Features:", + "- parser-null — Null-safe string parser", + " attempts: 1; latest state: superseded", + " outcome: none recorded", + " terminal findings: none", + "Findings digest: none", + "Assurance: completion not claimed", + "Assurance checks:", + "- not-applicable [TS-enforced] Recorded completion: deferred closure makes no completion claim.", + "- not-applicable [host-attested] Accepted validation: deferred closure makes no completion claim.", + "- not-applicable [host-attested] Canonical gate: deferred closure makes no completion claim.", + "- not-applicable [host-attested] Declared evidence: deferred closure makes no completion claim.", + "Assurance limitations:", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Artifacts as reported by Flow from caller declarations, not an exact or exhaustive Git delta:", + "- latest attempts: none reported", + "- superseded attempts only: none reported" + ], + "goal": "Make parse(input) safely handle null and preserve trimmed strings, changing only src/parser.mjs, with Linux validation and required external macOS proof." + } +} diff --git a/tests/fixtures/delivery-targeted-answer.json b/tests/fixtures/delivery-targeted-answer.json new file mode 100644 index 00000000..e2011eec --- /dev/null +++ b/tests/fixtures/delivery-targeted-answer.json @@ -0,0 +1,88 @@ +{ + "sourceCommit": "6765ee5930fe3dcd85eb6cecdd61c89aa50c3058", + "tarballSha256": "sha256:5dae6dbebc1b2fe98649c3267c32f849124f21f9932cda5b99fc306823a31db0", + "answer": "Completed and independently reviewed.\n\n`src/parser.mjs` now returns `''` for null and preserves string trimming:\n```js\nexport function parse(input) { return input === null ? '' : input.trim(); }\n```\n\n- **Validation:** `node scripts/verify.mjs` passed with exit code 0. Its script and invocation are unchanged.\n- **Independent review:** passed with no findings.\n- **Flow closure:** completed and archived; 1 of 1 features complete, none unfinished, no blockers or advisory findings.\n- **Assurance:** completion supported; all 4 assurance checks satisfied.\n- **External action authority:** not granted.\n\nGoal: Make parse(input) safely handle null and preserve trimmed strings. Change only src/parser.mjs; keep node scripts/verify.mjs and its script unchanged; validate and obtain independent review.\n\n**Flow assurance limitations**\n- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.\n- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.\n- Freshness holds when review is accepted; an archive does not attest the current workspace.\n\nFlow reports one latest artifact, `src/parser.mjs`, and no superseded artifacts; that inventory is caller-declared rather than an exhaustive Git delta.", + "goal": "Make parse(input) safely handle null and preserve trimmed strings. Change only src/parser.mjs; keep node scripts/verify.mjs and its script unchanged; validate and obtain independent review.", + "featureId": "parser-null", + "report": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings. Change only src/parser.mjs; keep node scripts/verify.mjs and its script unchanged; validate and obtain independent review.", + "Closure: completed — Updated only src/parser.mjs with explicit null guard; unchanged repository gate passed and independent review passed with no findings.", + "Progress: 1 of 1 features complete", + "Features:", + "- parser-null — Null-safe parser with preserved trimming", + " attempts: 1; latest state: completed", + " outcome: Verified parser-null and all approved requirements: parse(null) returns ''; all non-null inputs retain input.trim() behavior; only src/parser.mjs changed, with unchanged file type/mode, verification script and invocation, and preserved pre-existing work. Inspected current parser, downstream gate, README and package surface. Source-bound broad validation node scripts/verify.mjs exited 0 with complete output and covers both required assertions. No prior findings or remaining gaps.", + " terminal findings: none", + "Findings digest: none", + "Assurance: completion supported", + "Assurance checks:", + "- satisfied [TS-enforced] Recorded completion: 1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.", + "- satisfied [host-attested] Accepted validation: 1/1 terminal runs have eligible host evidence accepted by review.", + "- satisfied [host-attested] Canonical gate: \"node scripts/verify.mjs\" must have passing broad evidence accepted by review.", + "- satisfied [host-attested] Declared evidence: 1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.", + "Assurance limitations:", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Artifacts as reported by Flow from caller declarations, not an exact or exhaustive Git delta:", + "- latest attempts: src/parser.mjs", + "- superseded attempts only: none reported" + ], + "summary": [ + "Handoff format: 1", + "External action authority: not-granted", + "Goal: Make parse(input) safely handle null and preserve trimmed strings. Change only src/parser.mjs; keep node scripts/verify.mjs and its script unchanged; validate and obtain independent review.", + "Closure: completed. Updated only src/parser.mjs with explicit null guard; unchanged repository gate passed and independent review passed with no findings.", + "Progress: 1 of 1 features complete", + "Unfinished features: none", + "Nonblocking live findings: advisory 0. Historical findings: 0.", + "Assurance: completion supported", + "Assurance checks: 4 satisfied, 0 not applicable, 0 unsatisfied.", + "- Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "- Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "- Freshness holds when review is accepted; an archive does not attest the current workspace.", + "Reported artifacts: 1 latest, 0 superseded only. Caller declarations, not an exact or exhaustive Git delta.", + "Full report is included in this close response." + ], + "assurance": { + "checks": [ + { + "explanation": "1/1 features and 1/1 independent reviews pass, including a final review with no terminal blocker.", + "id": "recorded-completion", + "label": "Recorded completion", + "status": "satisfied", + "tier": "ts-enforced" + }, + { + "explanation": "1/1 terminal runs have eligible host evidence accepted by review.", + "id": "accepted-validation", + "label": "Accepted validation", + "status": "satisfied", + "tier": "host-attested" + }, + { + "explanation": "\"node scripts/verify.mjs\" must have passing broad evidence accepted by review.", + "id": "canonical-gate", + "label": "Canonical gate", + "status": "satisfied", + "tier": "host-attested" + }, + { + "explanation": "1/1 declared obligations have accepted evidence on their declared host; observed gates do not claim a pass.", + "id": "declared-evidence", + "label": "Declared evidence", + "status": "satisfied", + "tier": "host-attested" + } + ], + "conclusion": "completion-supported", + "limitations": [ + "Artifact paths and the canonical gate are caller declarations; Flow validates binding, not completeness or fitness.", + "Goal alignment, scope discipline, evidence completeness, requirement coverage, test adequacy, and review substance remain model judgments.", + "Freshness holds when review is accepted; an archive does not attest the current workspace." + ] + }, + "findingsDigest": [] +} diff --git a/tests/jev-decision-provider.test.ts b/tests/jev-decision-provider.test.ts index f0cd2ebf..8db2c9e5 100644 --- a/tests/jev-decision-provider.test.ts +++ b/tests/jev-decision-provider.test.ts @@ -1,5 +1,8 @@ import { expect, spyOn, test } from "bun:test"; -import type { DecisionPacket } from "../src/application/ports/decision-provider.js"; +import { + type DecisionPacket, + JEV_ATTEMPT_RESERVATION_USD, +} from "../src/application/ports/decision-provider.js"; import { createJevDecisionProvider } from "../src/infrastructure/jev-decision-provider.js"; const packet: DecisionPacket = { @@ -68,13 +71,13 @@ test("runtime adapter refuses absent keys and sensitive packets before transport }; expect( await createJevDecisionProvider(() => undefined).assess(packet, options), - ).toEqual({ kind: "unavailable", reason: "missing-key" }); + ).toMatchObject({ kind: "unavailable", reason: "missing-key" }); expect( await createJevDecisionProvider(() => "credential").assess( { ...packet, goal: "password=credential" }, options, ), - ).toEqual({ kind: "unavailable", reason: "sensitive-packet" }); + ).toMatchObject({ kind: "unavailable", reason: "sensitive-packet" }); for (const altered of [ { ...packet, goal: '{"password":"hunter2hunter2"}' }, { ...packet, goal: "GITHUB_TOKEN=ghp_secretvalue" }, @@ -114,7 +117,7 @@ test("runtime adapter refuses absent keys and sensitive packets before transport altered as DecisionPacket, options, ), - ).toEqual({ kind: "unavailable", reason: "sensitive-packet" }); + ).toMatchObject({ kind: "unavailable", reason: "sensitive-packet" }); expect(spy).toHaveBeenCalledTimes(0); } finally { spy.mockRestore(); @@ -197,7 +200,7 @@ test("runtime adapter rejects unknown model and invalid distributions", async () signal: new AbortController().signal, reserveAttempt: () => true, }), - ).toEqual({ + ).toMatchObject({ kind: "unavailable", reason: "model-mismatch", resolvedModel: "jev-1.14.0", @@ -214,7 +217,7 @@ test("runtime adapter rejects unknown model and invalid distributions", async () signal: new AbortController().signal, reserveAttempt: () => true, }), - ).toEqual({ kind: "unavailable", reason: "invalid-response" }); + ).toMatchObject({ kind: "unavailable", reason: "invalid-response" }); } finally { unrecognized.mockRestore(); } @@ -248,3 +251,114 @@ test("runtime adapter rejects unknown model and invalid distributions", async () } } }); + +for (const [name, payload, reason] of [ + ["malformed", "{", "malformed"], + ["schema", JSON.stringify({ model: "jev-1.13.0" }), "invalid-response"], + [ + "model", + JSON.stringify({ ...response(), model: "jev-1.14.0" }), + "model-mismatch", + ], +] as const) { + test(`failed ${name} response retains transport facts without validated usage`, async () => { + const result = await createJevDecisionProvider( + () => "fixture", + async () => new Response(payload), + ).assess(packet, { + signal: new AbortController().signal, + reserveAttempt: () => true, + }); + expect(result).toMatchObject({ + kind: "unavailable", + reason, + telemetry: { + transportAttempts: 1, + transportReservedUsd: JEV_ATTEMPT_RESERVATION_USD, + responseUsage: null, + }, + }); + expect(result.telemetry?.transportLatencyMs).toBeGreaterThanOrEqual(0); + }); +} +test("retry telemetry counts attempts but usage belongs only to the validated final response", async () => { + let attempts = 0; + const result = await createJevDecisionProvider( + () => "fixture", + async () => { + attempts++; + return attempts === 1 + ? new Response("busy", { status: 503 }) + : Response.json(response()); + }, + ).assess(packet, { + signal: new AbortController().signal, + reserveAttempt: () => true, + }); + expect(result).toMatchObject({ + kind: "answered", + telemetry: { + transportAttempts: 2, + transportReservedUsd: 2 * JEV_ATTEMPT_RESERVATION_USD, + responseUsage: { inputTokens: 100, outputTokens: 10 }, + }, + }); +}); +test("HTTP failure retains all reserved attempts and unknown usage", async () => { + const result = await createJevDecisionProvider( + () => "fixture", + async () => new Response("busy", { status: 503 }), + ).assess(packet, { + signal: new AbortController().signal, + reserveAttempt: () => true, + }); + expect(result).toMatchObject({ + kind: "unavailable", + reason: "http", + telemetry: { + transportAttempts: 3, + transportReservedUsd: 3 * JEV_ATTEMPT_RESERVATION_USD, + responseUsage: null, + }, + }); +}); +test("skipped requests disclose zero reservations and unknown transport duration", async () => { + const result = await createJevDecisionProvider(() => undefined).assess( + packet, + { + signal: new AbortController().signal, + reserveAttempt: () => { + throw new Error("must not reserve"); + }, + }, + ); + expect(result.telemetry).toEqual({ + transportLatencyMs: null, + transportAttempts: 0, + transportReservedUsd: 0, + responseUsage: null, + }); +}); +test("timeout retains latency and reservation without usage", async () => { + const result = await createJevDecisionProvider( + () => "fixture", + async () => + new Response("limited", { + status: 429, + headers: { "retry-after": "60" }, + }), + ).assess(packet, { + signal: new AbortController().signal, + reserveAttempt: () => true, + }); + expect(result).toMatchObject({ + kind: "unavailable", + reason: "timeout", + telemetry: { + transportAttempts: 1, + transportReservedUsd: JEV_ATTEMPT_RESERVATION_USD, + responseUsage: null, + }, + }); + expect(result.telemetry?.transportLatencyMs).toBeGreaterThanOrEqual(0); +}); diff --git a/tests/prompt-quality.test.ts b/tests/prompt-quality.test.ts index 96843c60..cb2670a3 100644 --- a/tests/prompt-quality.test.ts +++ b/tests/prompt-quality.test.ts @@ -566,3 +566,19 @@ describe("Flow prompt economy", () => { ); }); }); + +test("delivery handoffs use summary with retained full-detail fallback", () => { + for (const id of ["flow", "flow-run", "flow-plan"] as const) { + const content = getFlowGuidance(id).content.replace(/\s+/g, " "); + expect(content).toContain("workflowData.delivery.summary.lines"); + expect(content).toContain("missing summary"); + expect(content).toContain("retained close response's `report`"); + expect(content).not.toContain("delivery.report` verbatim"); + } + const status = compileFlowPromptSurface("flow-status"); + expect(status).toContain("If current `workflowData.delivery` exists"); + expect(status).toContain("report its `summary.lines`, or its `report`"); + expect(status).toContain("Otherwise say unavailable."); + expect(status).not.toContain("use retained close `report`"); + expect(status).toContain("workflowData.statusReport` verbatim"); +}); diff --git a/tests/qualification-cli.test.ts b/tests/qualification-cli.test.ts index 52e612e5..9c22ae8b 100644 --- a/tests/qualification-cli.test.ts +++ b/tests/qualification-cli.test.ts @@ -21,11 +21,7 @@ import { scenarioGradeInput, } from "../evals/grader-input.js"; import { type Outcome, packPlugin } from "../evals/harness.js"; -import { - evaluatorIdentity, - inspectArtifact, - instructionDelivery, -} from "../evals/provenance.js"; +import { evaluatorIdentity, inspectArtifact } from "../evals/provenance.js"; import { readQualificationBundle, writeQualificationBundle, @@ -45,6 +41,7 @@ import { campaignPlanFor, releaseScenarios, } from "../evals/run.js"; +import { scenarioStepInstruction } from "../evals/scenario-steps.js"; import { SCENARIOS } from "../evals/scenarios.js"; import packageJson from "../package.json" with { type: "json" }; import { prepareCanary, recordCanary } from "../scripts/eval-canary.js"; @@ -491,14 +488,7 @@ test("qualifies and seals a complete exact-artifact campaign through the CLI", a attemptId, text: canonicalJson(evidence), }); - const commands = scenario.steps.map((step, sequence) => - instructionDelivery({ - source: "command", - name: step.command, - sequence, - text: `/${step.command} ${step.arguments}`.trim(), - }), - ); + const commands = scenario.steps.map(scenarioStepInstruction); await store.writeAttempt({ schemaVersion: 2, attemptId, diff --git a/tests/recovery-capture.test.ts b/tests/recovery-capture.test.ts index 729f3dae..7f5fdc2a 100644 --- a/tests/recovery-capture.test.ts +++ b/tests/recovery-capture.test.ts @@ -14,6 +14,7 @@ import { import { tmpdir } from "node:os"; import { dirname, join } from "node:path"; import type { ToolContext } from "@opencode-ai/plugin"; +import { z } from "zod"; import { CapturingRecoveryController, RawRecoveryCaptureSchema, @@ -374,11 +375,27 @@ test("active shadow capture preserves unavailable advice and never grants a muta const expected = await baseline .guard(context) .propose(f.session, f.sourceDigest, proposal); - expect( - await controller - .guard(context) - .propose(f.session, f.sourceDigest, proposal), - ).toEqual(expected); + const actual = await controller + .guard(context) + .propose(f.session, f.sourceDigest, proposal); + const withoutElapsed = (input: unknown) => { + const assessment = z + .object({ + telemetry: z + .object({ assessmentElapsedMs: z.number().finite().nonnegative() }) + .passthrough(), + }) + .passthrough() + .parse(input); + expect(Number.isFinite(assessment.telemetry.assessmentElapsedMs)).toBe( + true, + ); + expect(assessment.telemetry.assessmentElapsedMs).toBeGreaterThanOrEqual(0); + const { assessmentElapsedMs: _elapsed, ...telemetry } = + assessment.telemetry; + return { ...assessment, telemetry }; + }; + expect(withoutElapsed(actual)).toEqual(withoutElapsed(expected)); expect(expected).toMatchObject({ kind: "unavailable", mode: "shadow", @@ -387,7 +404,16 @@ test("active shadow capture preserves unavailable advice and never grants a muta action: null, }); expect(expected).not.toHaveProperty("recommended"); - expect(controller.snapshot()).toEqual(baseline.snapshot()); + const snapshot = z.object({ last: z.unknown() }).passthrough(); + const actualSnapshot = snapshot.parse(controller.snapshot()); + const expectedSnapshot = snapshot.parse(baseline.snapshot()); + expect({ + ...actualSnapshot, + last: withoutElapsed(actualSnapshot.last), + }).toEqual({ + ...expectedSnapshot, + last: withoutElapsed(expectedSnapshot.last), + }); expect((await records(f.directory))[0]?.payload.proposal).toEqual(proposal); }); diff --git a/tests/recovery-decisions.test.ts b/tests/recovery-decisions.test.ts index 656f7a0a..b125008a 100644 --- a/tests/recovery-decisions.test.ts +++ b/tests/recovery-decisions.test.ts @@ -2,6 +2,7 @@ import { describe, expect, test } from "bun:test"; import { mkdtemp, readFile, rm } from "node:fs/promises"; import { tmpdir } from "node:os"; import { join } from "node:path"; +import { z } from "zod"; import { CorpusSchema, evaluateRecoveryCorpus, @@ -146,3 +147,39 @@ describe("runtime recovery evaluation", () => { } }); }); + +test("offline preparation repeats exactly while provider assessments retain measured duration", async () => { + const corpus = await load(); + const first = await evaluateRecoveryCorpus(corpus); + await new Promise((resolve) => setTimeout(resolve, 10)); + const second = await evaluateRecoveryCorpus(corpus); + expect(second).toEqual(first); + expect( + first.rows.find((row) => row.id === "supported-retry")?.decision, + ).toMatchObject({ + telemetry: { assessmentElapsedMs: 0 }, + }); + const entry = corpus.cases.find((row) => row.id === "supported-retry"); + if (!entry) throw new Error("Missing eligible fixture."); + const evaluated = await evaluateRecoveryCorpus( + { ...corpus, cases: [entry] }, + { + async assess(packet, options) { + await new Promise((resolve) => setTimeout(resolve, 10)); + return certain.assess(packet, options); + }, + }, + ); + const decision = z + .object({ + telemetry: z.object({ + assessmentElapsedMs: z.number().finite().nonnegative(), + }), + }) + .parse(evaluated.rows[0]?.decision); + expect(decision.telemetry.assessmentElapsedMs).toBeGreaterThanOrEqual(5); + expect(evaluated.rows[0]?.decision).toMatchObject({ + kind: "selected", + selectedCandidateId: "repair", + }); +}); diff --git a/tests/recovery-policy.test.ts b/tests/recovery-policy.test.ts index 4dbfb0a1..df2b8b46 100644 --- a/tests/recovery-policy.test.ts +++ b/tests/recovery-policy.test.ts @@ -1,6 +1,9 @@ import { describe, expect, test } from "bun:test"; import { createFlowService } from "../src/application/flow-service.js"; -import type { DecisionProvider } from "../src/application/ports/decision-provider.js"; +import { + type DecisionProvider, + JEV_ATTEMPT_RESERVATION_USD, +} from "../src/application/ports/decision-provider.js"; import { RecoveryController, type RecoveryMutation, @@ -1852,3 +1855,93 @@ test("plain auto in another host rejects the previous host's pending grant", asy expect((await apply(s, mutation)).status).toBe("error"); await replacement; }); + +for (const [dimension, values] of [ + ["choice", { choice: 0.89, goal: 1, suitability: 1 }], + ["goal", { choice: 1, goal: 0.94, suitability: 1 }], + ["suitability", { choice: 1, goal: 1, suitability: 0.94 }], +] as const) { + test(`recovery tool exposes ${dimension} threshold rejection without authority`, async () => { + const s = await setup("delegated", { + async assess(packet, options) { + const answer = await provider.assess(packet, options); + if (answer.kind !== "answered") + throw new Error("fixture answer unavailable"); + return { + ...answer, + probabilities: { repair: values.choice, abstain: 1 - values.choice }, + assessments: { + repair: { goal: values.goal, suitability: values.suitability }, + }, + }; + }, + }); + const response = await s.flow.status({ + request: { view: "compact" }, + recoveryProposal: s.proposal(), + }); + if (response.status !== "ok" || !("recovery" in response.workflowData)) + throw new Error(response.summary); + expect(response.workflowData.recovery).toMatchObject({ + kind: "abstain", + telemetry: { + attemptsReserved: 1, + reservedUsd: JEV_ATTEMPT_RESERVATION_USD, + transportAttempts: null, + responseUsage: { inputTokens: 100, outputTokens: 10 }, + }, + decision: { + choice: "repair", + confidence: 1, + probabilities: { repair: values.choice, abstain: 1 - values.choice }, + assessments: { + repair: { goal: values.goal, suitability: values.suitability }, + }, + thresholds: { choice: 0.9, goal: 0.95, suitability: 0.95 }, + checks: { + attemptReserved: true, + modelMatched: true, + candidatePresent: true, + assessmentPresent: true, + choicePassed: dimension !== "choice", + goalPassed: dimension !== "goal", + suitabilityPassed: dimension !== "suitability", + }, + }, + }); + expect(response.workflowData.recovery).not.toHaveProperty("recommended"); + expect(s.controller.snapshot("host")).toMatchObject({ + last: response.workflowData.recovery, + }); + }); +} +test("opaque provider failure retains measured assessment and actual reservation deltas", async () => { + const s = await setup("shadow", { + async assess(_packet, options) { + expect(options.reserveAttempt()).toBe(true); + throw new Error("opaque failure with private content"); + }, + }); + const response = await s.flow.status({ + request: { view: "compact" }, + recoveryProposal: s.proposal(), + }); + if (response.status !== "ok" || !("recovery" in response.workflowData)) + throw new Error(response.summary); + expect(response.workflowData.recovery).toMatchObject({ + kind: "unavailable", + reason: "provider", + decision: null, + telemetry: { + attemptsReserved: 1, + reservedUsd: JEV_ATTEMPT_RESERVATION_USD, + transportLatencyMs: null, + transportAttempts: null, + transportReservedUsd: null, + responseUsage: null, + }, + }); + expect(JSON.stringify(response.workflowData.recovery)).not.toContain( + "private content", + ); +}); diff --git a/tests/runtime-close.test.ts b/tests/runtime-close.test.ts index 9ce492fe..1d22ea21 100644 --- a/tests/runtime-close.test.ts +++ b/tests/runtime-close.test.ts @@ -799,6 +799,10 @@ describe("Flow close recovery and delivery", () => { // The runtime renders the handoff so its shape, ordering, and the artifact // qualifier are guarantees rather than instructions restated per surface. report: expect.any(Array), + summary: expect.objectContaining({ + fullReportAvailable: true, + lines: expect.any(Array), + }), }; repository.archiveFailure = new Error("injected delivery archive failure"); @@ -912,6 +916,10 @@ describe("Flow close recovery and delivery", () => { }), findingsDigest: [], report: expect.any(Array), + summary: expect.objectContaining({ + fullReportAvailable: true, + lines: expect.any(Array), + }), reportedArtifacts: { latestAttempts: [], supersededAttemptsOnly: [], @@ -1032,6 +1040,10 @@ describe("Flow close recovery and delivery", () => { // Rendering is asserted line-by-line in the delivery and planless cases; // here the interesting part is the never-started feature. report: expect.any(Array), + summary: expect.objectContaining({ + fullReportAvailable: true, + lines: expect.any(Array), + }), }); if (!("delivery" in deferred.workflowData)) { throw new Error("Expected deferred close delivery data."); diff --git a/tests/scenario-steps.test.ts b/tests/scenario-steps.test.ts new file mode 100644 index 00000000..eea72619 --- /dev/null +++ b/tests/scenario-steps.test.ts @@ -0,0 +1,119 @@ +import { expect, test } from "bun:test"; +import { instructionDelivery } from "../evals/provenance.js"; +import { + releaseCaseCatalogSha256, + releaseScenarioCatalog, +} from "../evals/release-policy.js"; +import { + runScenarioStep, + type ScenarioStep, + scenarioStepCatalog, + scenarioStepInstruction, +} from "../evals/scenario-steps.js"; +import { SCENARIOS } from "../evals/scenarios.js"; + +test("ordinary user followups use prompt endpoint and preserve exact whitespace and Unicode", async () => { + const requests: unknown[] = []; + const host = { + async runCommand(...args: [string, string, string, string]) { + requests.push({ kind: "command", args }); + return "quiet" as const; + }, + async runPrompt(...args: [string, string, string]) { + requests.push({ kind: "prompt", args }); + return "quiet" as const; + }, + }; + const steps: ScenarioStep[] = [ + { command: "flow-auto", arguments: "--recovery=off Fix parser" }, + { kind: "prompt", prompt: " Please show the full report. 漢\n" }, + ]; + for (const step of steps) + expect( + await runScenarioStep(host, "native-session", step, "route/model"), + ).toBe("quiet"); + expect(requests).toEqual([ + { + kind: "command", + args: [ + "native-session", + "flow-auto", + "--recovery=off Fix parser", + "route/model", + ], + }, + { + kind: "prompt", + args: [ + "native-session", + " Please show the full report. 漢\n", + "route/model", + ], + }, + ]); + const prompt = scenarioStepInstruction( + steps[1] ?? { kind: "prompt", prompt: "missing" }, + 1, + ); + expect(prompt).toMatchObject({ + source: "user-prompt", + text: " Please show the full report. 漢\n", + sequence: 1, + }); + expect(prompt.bytes).toBe( + Buffer.byteLength(" Please show the full report. 漢\n"), + ); + expect( + scenarioStepCatalog(steps[1] ?? { kind: "prompt", prompt: "missing" }), + ).toEqual({ + kind: "prompt", + prompt: " Please show the full report. 漢\n", + freshSession: false, + }); +}); + +test("legacy command encoding and instruction delivery remain byte identical", () => { + const command = { + command: "flow-plan", + arguments: " Fix parser ", + freshSession: true, + }; + expect(scenarioStepCatalog(command)).toEqual({ + command: "flow-plan", + arguments: " Fix parser ", + freshSession: true, + }); + expect(scenarioStepInstruction(command, 0)).toEqual( + instructionDelivery({ + source: "command", + name: "flow-plan", + sequence: 0, + text: "/flow-plan Fix parser", + }), + ); + expect(scenarioStepInstruction(command, 0)).not.toHaveProperty("text"); +}); + +const baseline = { + "9.1.0": + "sha256:134581f969e030f4194ff48f94e6df4d2fac5ebbd3af57f66cf430a32f7f7b6c", + "9.2.0": + "sha256:134581f969e030f4194ff48f94e6df4d2fac5ebbd3af57f66cf430a32f7f7b6c", + "9.3.0": + "sha256:b0bfc9d312ced4ec67b520b8fcaa3a71f9dd653e3bbb5dbb0edeb500716a9f03", + "9.4.0": + "sha256:cdec46ab03ac3ba08e4434f6162fbd452665c4d13af66af523d8a991e78e66ae", + "9.5.0": + "sha256:cdec46ab03ac3ba08e4434f6162fbd452665c4d13af66af523d8a991e78e66ae", + standard: + "sha256:b0bfc9d312ced4ec67b520b8fcaa3a71f9dd653e3bbb5dbb0edeb500716a9f03", +}; +for (const [version, sha256] of Object.entries(baseline)) + test(`required ${version} catalog remains frozen`, () => { + expect(releaseCaseCatalogSha256(SCENARIOS, version)).toBe(sha256); + expect( + releaseScenarioCatalog(SCENARIOS, version).every((scenario) => + scenario.steps.every((step) => "command" in step), + ), + ).toBe(true); + }); diff --git a/tests/workspace-lifecycle-integration.test.ts b/tests/workspace-lifecycle-integration.test.ts index 695885c9..7094fc66 100644 --- a/tests/workspace-lifecycle-integration.test.ts +++ b/tests/workspace-lifecycle-integration.test.ts @@ -337,6 +337,10 @@ test("persists one complete workspace lifecycle and replays its exact close", as }), findingsDigest: expect.any(Array), report: expect.any(Array), + summary: expect.objectContaining({ + fullReportAvailable: true, + lines: expect.any(Array), + }), }); expect(await loadSession(workspace)).toBeNull(); const archived = await loadArchivedSession(