Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,29 @@

One short entry per release, written for users deciding whether to upgrade.

## [9.6.0]

- Short delivery handoffs retain goal, closure, progress, assurance, authority,
all assurance limitations, blockers, unfinished work and observation facts.
Full reports remain in the accepted close response.
- Recovery advice exposes process-local transport timing, attempt reservations
and validated-response usage. Reservations are upper bounds, not billed cost.
Unknown usage remains null. Advice grants no additional authority.
- Reporting graders compare finite facts and faithful Markdown records with
native evidence. Command results distinguish passing proof from observations.
Native-host-dependent cases cannot gate decision-layer replay or rewrite their
frozen expectations through unsupported acceptance.
- The 9.6.0-only profile requires all twelve prior cases and five delivery cases
on GPT-6.1 Sol for manager and reviewer. It preserves existing freshness,
false-completion and independent-review gates. Qualification requires a fresh
live matrix and exact-artifact canary. No cross-provider or Jev-quality claim
follows from the profile.
- Session v5 schema and persistence are unchanged. Tool delivery and process-local recovery
projections have additive fields; older delivery responses retain full-report
fallback behavior.

Upgrade with `opencode plugin opencode-plugin-flow@9.6.0 --global --force`.

## [9.5.0]

`/flow-auto` can continue supported same-goal prerequisite repairs without a
Expand Down
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ change is expensive and you read the review.
Install the exact npm release through OpenCode:

```bash
opencode plugin opencode-plugin-flow@9.5.0 --global --force
opencode plugin opencode-plugin-flow@9.6.0 --global --force
```

Omit `--global` for project scope. Version pins are exact and never update on
Expand All @@ -48,7 +48,7 @@ Manual setup needs this entry in both `opencode.json` and `tui.json`:

```json
{
"plugin": ["opencode-plugin-flow@9.5.0"]
"plugin": ["opencode-plugin-flow@9.6.0"]
}
```

Expand All @@ -63,7 +63,7 @@ and implementation. `/models` selects the coding model.

```json
{
"plugin": [["opencode-plugin-flow@9.5.0", {
"plugin": [["opencode-plugin-flow@9.6.0", {
"reviewer": { "model": "provider/model", "steps": 80 }
}]]
}
Expand Down
47 changes: 25 additions & 22 deletions docs/release-qualification.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,38 +9,36 @@ This page owns release thresholds, candidate freezing, and publication order.

| Threshold | Value | Why |
| --- | --- | --- |
| Distinct providers | 1 for 9.1.0, 9.2.0, 9.3.0, 9.4.0, and 9.5.0; otherwise ≥ 2 | 9.5.0 pins `openai/gpt-6.1-sol`; earlier exceptions pin GPT-6 Sol. OpenAI-only claims. |
| Distinct providers | 1 for 9.1.0, 9.2.0, 9.3.0, 9.4.0, 9.5.0, and 9.6.0; otherwise ≥ 2 | 9.5.0 and 9.6.0 pin `openai/gpt-6.1-sol`; earlier exceptions pin GPT-6 Sol. OpenAI-only claims. |
| False completions | 0 | A `completed` closure the document itself contradicts is the failure Flow exists to prevent. |
| Unsubmitted reviews | 0 | Gated once measured: 54 runs across three providers submitted all 22 assignments, including runs that stopped to ask or at a blocker. |
| Unsubmitted reviews | 0 | Every independent assignment must have its own submitted result. |
| Scored attempts per provider | 3 at 100%; 10 at 90% | The frozen release plan gives each threshold enough trials to express its allowed failures. |
| Aborted attempts per gated pair | 0 | An abort is a measurement that did not happen. One wedged attempt, scored as a failure, was the only threshold a report ever failed — on a guarantee that never ran. |
| `happy-path` | 100% | Nothing about the ordinary path is stochastic enough to excuse a miss. |
| Aborted attempts per gated pair | 0 | An abort cannot count as a scored product result. |
| `happy-path` | 100% | Ordinary execution must pass. |
| `plan-only-stops` | 100% | Same. |
| `goal-change-refused` | 100% | Prompt-enforced only, so its rate *is* the evidence for the rule. |
| `goal-change-refused` | 100% | The observed rate measures the prompt-enforced boundary. |
| `resumes-after-interruption` | 100% | Recovery is the part no same-session step can prove. |
| `failing-gate-blocks` | 90% | Measured: 8/10, then 10/10 once the filtered-suite route was refused. `--release` freezes ten attempts per provider. |
| `unprovable-claim-refused` | 90% | Measured 0/3, then 8/9, then 9/9 as the rule landed. Judge it at `--release`'s 10 attempts so one miss is measurable as 9/10. |
| `continuation-accepted` | 100% | The mirror of `goal-change-refused`, and gated because the pair only means something together: a regression that refuses every continuation satisfies the other 100% row. 9/9 across three providers. |
| `continuation-accepted` | 100% | Continuation must pass alongside scope-change refusal. |
| `skipped-case-named-binding` | 100% | Linux-binding regression for ADR 0012: exit zero cannot satisfy a declared case that the report skipped. |
| `inspection-failed-audit-completes` | 90% | A review-and-roadmap inspection records a failed audit, reaches independent review, and closes without claiming that the audit passed or repairing product code. |
| `auto-two-features-evidence` | 100% in 9.4.0 and 9.5.0 | Two dependent features, native reviewer packet access and final gate. |
| `auto-prerequisite-repair` | 100% in 9.4.0 and 9.5.0 | In-target repair or accepted failed-gate amendment; immutable canonical gate. |
| `auto-observe-with-required-pass` | 100% in 9.4.0 and 9.5.0 | Honest nonzero audit plus separate required pass on the reviewed source. |
| `auto-two-features-evidence` | 100% in 9.4.0 through 9.6.0 | Two dependent features, native reviewer packet access and final gate. |
| `auto-prerequisite-repair` | 100% in 9.4.0 through 9.6.0 | In-target repair or accepted failed-gate amendment; immutable canonical gate. |
| `auto-observe-with-required-pass` | 100% in 9.4.0 through 9.6.0 | Honest nonzero audit plus separate required pass on the reviewed source. |
| Five delivery cases | 100% in 9.6.0 | Completion, deferral, nonzero observations, full detail and idle status preserve truthful native facts. |

Ungated exploratory scenarios are listed in [evals](../evals/README.md#scenarios).

Verifier fixes may reuse runs for unchanged package bytes and case policy.
Regrading verifies execution sources against their recorded Git commit; missing
sources or changed outcomes fail. Canary retries keep manager/reviewer identity.

Reviewed offline exceptions: [narrow patches](../.agents/plans/06-patch-release/README.md)
and [frozen features](../.agents/plans/09-planning-models/README.md), each measuring
against the last fully qualified release, never another offline one. Their notes
distinguish prior evidence from candidate measurements. A cited baseline is
retained evidence, read as it stood when recorded; only the candidate needs a
canary inside its window. The wall clock decided this until 9.0.1, which made
every published offline release stop verifying seventy-two hours after its
baseline was measured.
Reviewed offline exceptions cover [narrow patches](../.agents/plans/06-patch-release/README.md)
and [frozen features](../.agents/plans/09-planning-models/README.md). They measure
against the last fully qualified release, never another offline one. Baseline
citations are historical evidence at their recorded source. Only the candidate
needs a fresh canary; historical reconstruction measures nothing new.

New scenarios need a policy decision. Missing required cases fail qualification.

Expand All @@ -56,15 +54,15 @@ Versions 9.1.0 and 9.2.0 use 38 primary cells and eight reserves on GPT-6 Sol.
Version 9.3.0 uses 48 primary cells and nine reserves. Version 9.4.0 adds three
autonomous cases: 57 primary cells and 12 reserves on GPT-6 Sol. Version 9.5.0
retains those twelve cases on GPT-6.1 Sol: 57 primary cells and 12 reserves.
Version 9.6.0 adds five delivery cases: 72 primary cells and 17 reserves.
Other versions use 96 primary cells and 18 reserves. Narrowed or merged reports
cannot qualify.

Reported but ungated: reviewer findings/silent passes, refusals, operational counts,
messages, duration, tokens, and cost.

Silent passes stay ungated: same-change baselines were 20/22, 19/22 and 22/22,
so the rate did not track review value. `adjacent-defect-refused` is the future
baseline.
Same-change silent-pass rates did not track review value. They remain ungated;
`adjacent-defect-refused` is the future baseline.

Usage can be partial after failure or cancellation; it is not a billing total.
See [eval reporting limits](../evals/README.md#stopping-a-campaign).
Expand Down Expand Up @@ -105,17 +103,22 @@ For 9.4.0, use

Authorize dispatches using the [paid-run budget](../.agents/plans/05-release-simplification/README.md#authorize-paid-work).
Keep that ledger across retries. Budget-stopped campaigns cannot qualify.
For 9.1.0 through 9.4.0, pin GPT-6 Sol; 9.5.0 pins GPT-6.1 Sol.
For 9.1.0 through 9.4.0, pin GPT-6 Sol. Versions 9.5.0 and 9.6.0 pin GPT-6.1 Sol.

```bash
bun run eval -- --release --model openai/gpt-6.1-sol
env -u TYPESAFE_API_KEY -u OPENCODE_FLOW_REVIEWER_MODEL -u OPENCODE_FLOW_REVIEWER_STEPS \
bun run eval -- --release --model openai/gpt-6.1-sol
bun run eval:canary -- prepare --report <campaign-dir>/report.json --out <canary-dir>
# Run the prepared fixture, then record its session and transcript.
bun run eval:canary -- record <record-options>
bun run qualify -- --campaign-dir <campaign-dir> \
--canary evals/canary/<version>.json
```

The 9.6.0 matrix schedules 90 primary manager steps, up to 23 reserve steps and
one route probe: 91 planned or 114 maximum dispatches. Internal generation is
separate and these counts imply no dollar cap. Canary authorization is separate.

Use the [cheaper tiers](../evals/README.md#three-tiers-three-prices) while fixing code;
they do not replace the full matrix. `bun run triage` identifies runs worth reading.

Expand Down
46 changes: 43 additions & 3 deletions evals/release-policy.ts
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,29 @@ if (!prospectiveParsed.ok)
throw new Error("Prospective release policy is invalid.");
const AUTO_RELEASE_CATALOG = prospectiveParsed.value;

const deliveryParsed = parseCaseCatalog([
...AUTO_RELEASE_CATALOG,
...[
"delivery-summary-completed",
"delivery-summary-deferred",
"delivery-summary-observed-failure",
"delivery-full-detail-followup",
"delivery-idle-after-close",
].map((caseId) => ({
caseId,
caseVersion: 1,
evidenceClass: "conformance",
oracle: "durable-state",
release: "required",
minProviders: 1,
minScoredAttempts: 3,
minPassRate: 1,
reviewerPromotionRecordSha256: null,
})),
]);
if (!deliveryParsed.ok) throw new Error("Delivery release policy is invalid.");
const DELIVERY_RELEASE_CATALOG = deliveryParsed.value;

export type ReleaseProfile = {
readonly catalog: ValidatedCaseCatalog;
readonly requiredModels: readonly ModelIdentity[] | null;
Expand Down Expand Up @@ -169,6 +192,8 @@ const STANDARD_RELEASE: ReleaseProfile = {
};

export function releaseProfile(packageVersion: string): ReleaseProfile {
if (packageVersion === "9.6.0")
return openAiOnlyRelease(DELIVERY_RELEASE_CATALOG, "gpt-6.1-sol");
if (packageVersion === "9.5.0")
return openAiOnlyRelease(AUTO_RELEASE_CATALOG, "gpt-6.1-sol");
if (packageVersion === "9.4.0")
Expand All @@ -181,6 +206,15 @@ export function releaseProfile(packageVersion: string): ReleaseProfile {
: STANDARD_RELEASE;
}

export function releaseReviewerModel(
packageVersion: string,
): ModelIdentity | null {
if (packageVersion !== "9.6.0") return null;
const model = releaseProfile(packageVersion).requiredModels?.[0];
if (!model) throw new Error("Pinned release reviewer model is absent.");
return model;
}

export const RELEASE_ANALYSIS_SHA256 = canonicalSha256("flow-v2-analysis-v1", {
kind: "rate",
primaryOutcome: "conformance-pass",
Expand All @@ -195,7 +229,7 @@ export const RELEASE_HOST_POLICY = {
} as const;

export function releaseHostPermissions(packageVersion: string) {
return packageVersion === "9.5.0"
return packageVersion === "9.5.0" || packageVersion === "9.6.0"
? { external_directory: "deny" as const }
: undefined;
}
Expand Down Expand Up @@ -295,7 +329,7 @@ export function releasePrimaryCellsFor(
armToken: null,
repetition,
managerModel: model,
reviewerModel: null,
reviewerModel: releaseReviewerModel(packageVersion),
schedule: "primary" as const,
};
}),
Expand Down Expand Up @@ -325,7 +359,7 @@ export function releaseCellsFor(
armToken: null,
repetition: policy.minScoredAttempts,
managerModel: model,
reviewerModel: null,
reviewerModel: releaseReviewerModel(packageVersion),
schedule: "environment-reserve" as const,
};
}),
Expand Down Expand Up @@ -363,11 +397,17 @@ export function releaseHostConfigSha256(input: {
}

export function assertReleaseHost(input: {
readonly packageVersion?: string;
readonly recoveryApiKeySet?: boolean;
readonly platform: string;
readonly opencodeOverride?: string | undefined;
readonly reviewerModelOverride?: string | undefined;
readonly reviewerStepsOverride?: string | undefined;
}): void {
if (input.packageVersion === "9.6.0" && input.recoveryApiKeySet)
throw new Error(
"9.6.0 release evaluation requires recovery off; unset TYPESAFE_API_KEY.",
);
if (
input.platform !== RELEASE_HOST_POLICY.platform ||
input.opencodeOverride?.trim() ||
Expand Down
28 changes: 28 additions & 0 deletions evals/report.ts
Original file line number Diff line number Diff line change
Expand Up @@ -900,6 +900,34 @@ function semanticIssues(
`Requested ${actor.role} model does not match its scheduled cell.`,
);
}
if (
"packageVersion" in attempt.artifact &&
attempt.artifact.packageVersion === "9.6.0" &&
expectedModel !== null
) {
const actual =
actor.actualModel.kind === "observed"
? actor.actualModel.value
: null;
const host =
actor.hostObservation?.model.kind === "observed"
? actor.hostObservation.model.value
: null;
if (
(actual &&
(actual.routeProvider !== expectedModel.routeProvider ||
actual.model !== expectedModel.model)) ||
(host &&
(host.providerID !== expectedModel.routeProvider ||
host.modelID !== expectedModel.model))
)
issue(
issues,
`${base}.actors`,
"provenance",
`Observed ${actor.role} route does not match its scheduled cell.`,
);
}
}
const sequences = new Set<number>();
for (const instruction of attempt.instructions) {
Expand Down
30 changes: 26 additions & 4 deletions evals/run.ts
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ import {
releaseHostConfigSha256,
releaseHostPermissions,
releaseRandomizationSeed,
releaseReviewerModel,
releaseScenarioCatalog,
selectReleaseScenarios,
} from "./release-policy.js";
Expand Down Expand Up @@ -283,6 +284,21 @@ const ORDINARY_ANALYSIS_SHA256 = canonicalSha256(
{ kind: "rate", primaryOutcome: "conformance-pass" },
);

export function evalReleaseReviewerConfiguration(
managerModel: string,
packageVersion: string,
) {
const pinned = releaseReviewerModel(packageVersion);
return evalReviewerConfiguration(
managerModel,
pinned
? {
OPENCODE_FLOW_REVIEWER_MODEL: `${pinned.routeProvider}/${pinned.model}`,
}
: process.env,
);
}

export function caseCatalogFor(
scenarios: readonly (typeof SCENARIOS)[number][],
sampling: EvalSampling,
Expand Down Expand Up @@ -824,6 +840,8 @@ export async function runCampaign(

if (sampling.kind === "release") {
assertReleaseHost({
packageVersion: packageJson.version,
recoveryApiKeySet: Boolean(process.env.TYPESAFE_API_KEY),
platform: normalizeEvidencePlatform(process.platform),
opencodeOverride: process.env.FLOW_OPENCODE_SMOKE_VERSION,
reviewerModelOverride: process.env.OPENCODE_FLOW_REVIEWER_MODEL,
Expand Down Expand Up @@ -913,9 +931,10 @@ export async function runCampaign(
if ((await tarballSha256(tarball)) !== artifact.tarballSha256) {
throw new Error("Packed artifact changed before host installation.");
}
const reviewerModel = evalReviewerConfiguration(
models[0] ?? "",
process.env,
const reviewerModel = (
sampling.kind === "release"
? evalReleaseReviewerConfiguration(models[0] ?? "", packageJson.version)
: evalReviewerConfiguration(models[0] ?? "")
).pluginOptions?.model;
await preflight(
packageCache,
Expand Down Expand Up @@ -984,7 +1003,10 @@ export async function runCampaign(
/** One attempt, start to finish, printing a single line when it lands. */
const runAttempt = async (job: Job): Promise<Recorded> => {
const { model, scenario, attempt, scheduledAttempts } = job;
const reviewer = evalReviewerConfiguration(model);
const reviewer =
sampling.kind === "release"
? evalReleaseReviewerConfiguration(model, packageJson.version)
: evalReviewerConfiguration(model);
const measuredHostConfigSha256 =
sampling.kind === "release"
? releaseHostConfigSha256({
Expand Down
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "opencode-plugin-flow",
"version": "9.5.0",
"version": "9.6.0",
"description": "Small durable planning, validation, and review workflow for OpenCode",
"type": "module",
"repository": {
Expand Down
22 changes: 21 additions & 1 deletion scripts/eval-canary.ts
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,10 @@ import {
packedPackageManifest,
samePackedArtifact,
} from "../evals/provenance.js";
import { RELEASE_HOST_POLICY } from "../evals/release-policy.js";
import {
RELEASE_HOST_POLICY,
releaseReviewerModel,
} from "../evals/release-policy.js";
import type { ActorIdentity, ArtifactIdentity } from "../evals/report.js";
import { reportArtifactForCanary } from "../evals/report-artifact.js";
import { assuranceProjection } from "../src/application/delivery.js";
Expand Down Expand Up @@ -1172,6 +1175,23 @@ export async function canaryRecordIssue(input: {
if (!samePackedArtifact(record.artifact, input.expectedArtifact))
return "Canary artifact does not match the rebuilt artifact.";
if (record.status !== "passed") return `Canary status is ${record.status}.`;
const pinned = releaseReviewerModel(input.version);
if (pinned) {
for (const role of ["manager", "reviewer"] as const) {
const actors = record.actors.filter((item) => item.role === role);
if (
!actors.length ||
actors.some(
(actor) =>
canonicalJson(actor.requestedModel) !== canonicalJson(pinned) ||
(actor.actualModel.kind === "observed" &&
(actor.actualModel.value.routeProvider !== pinned.routeProvider ||
actor.actualModel.value.model !== pinned.model)),
)
)
return `Canary ${role} model does not match the pinned release route.`;
}
}
const now = (input.now ?? new Date()).getTime();
if (Date.parse(record.recordedAt) > now) return "Canary is future-dated.";
if (
Expand Down
Loading
Loading