Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
52 changes: 24 additions & 28 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,17 +7,15 @@ benefits from an approved plan and an independent review:
plan → approve → run one feature → validate → review → repeat or close
```

Flow keeps one durable active feature run at a time. Once a session starts it
stays the workflow for that goal until Flow records completed, deferred, or
abandoned closure. It never silently falls back to ordinary coding, and it does
not fold a materially different request into the active goal.
Flow keeps one durable active feature run. Its goal stays active until completed,
deferred, or abandoned closure. Flow never silently resumes ordinary coding or
adds a materially different request to that goal.

State lives in `.flow/session.json`, so the workflow survives a restart, a
context change, or a lost transcript.

Flow is in preview: an opinionated workflow for consequential multi-step changes,
for people who read the review. It is worth its ceremony when a wrong change is
expensive, and it is overhead when it is not.
Flow is in preview. Its planning and review cost is worthwhile when a wrong
change is expensive and you read the review.

## When not to use Flow

Expand Down Expand Up @@ -147,9 +145,9 @@ you granted.
without shell access. New failures, source drift, or missing packets stop
assignment without counting a failed review.
6. A passing review advances the plan. A failed feature needs an explicit retry
or independent-feature choice. Closure returns a versioned delivery report
with attempts, findings, and assurance limits. It grants no PR, merge,
publish, or release authority.
or independent-feature choice. Closure returns a short delivery summary
and the full versioned report in the same response. Both retain assurance
limits. Closure grants no PR, merge, publish, or release authority.

Failed reviews retain finding ids; dropped live findings fail.
`flow_plan_amend` records a same-goal reversible gate repair before review,
Expand Down Expand Up @@ -205,34 +203,32 @@ bun install --frozen-lockfile
bun run check
```

`bun run check` runs typechecking, lint, build verification, tests, and package
smoke. Release CI also exercises the packed plugin in a real OpenCode host.
`bun run check` checks types, lint, builds, tests, and packages. Release CI tests
that package in OpenCode.

Maintained documentation starts at [docs/index.md](docs/index.md):
[development](docs/development.md) for repository structure,
[troubleshooting](docs/troubleshooting.md) for recovery,
[the maintainer contract](docs/maintainer-contract.md) for tools and runtime
invariants, and [ADR 0006](docs/adr/0006-bounded-intra-feature-waves.md) for the
bounded-wave rationale.
[Documentation](docs/index.md) includes [development](docs/development.md),
[troubleshooting](docs/troubleshooting.md), [runtime contracts](docs/maintainer-contract.md),
and [bounded waves](docs/adr/0006-bounded-intra-feature-waves.md).

## Recovery advice development preview

With `TYPESAFE_API_KEY`, `/flow-auto` defaults to shadow
advice, capped at six attempts and $0.02. Without it, advice is off. Use
`--recovery=off` to opt out or the options below to change limits.
With `TYPESAFE_API_KEY`, `/flow-auto` defaults to shadow advice, six attempts and
$0.02. Without the key or with `--recovery=off`, advice is off.

```text
/flow-auto --recovery=shadow --recovery-calls=6 --recovery-usd=0.02 <goal>
```

Shadow sends bounded goal, finding and candidate-remedy packets to TypeSafe.
It reports advice without authorizing mutations. Bare `flow_status` never calls
Jev. Attempts count retries. `/flow-auto stop` cancels auto and advice. Advice expires
after one hour; automatic retry limits remain until invocation ends. Neither
survives restart. `recoveryStatus` reports configuration, attempts and outcome.
Shadow sends bounded goal, finding and remedy packets to TypeSafe. It grants no
mutation authority. Bare `flow_status` never calls Jev. Retries count as attempts.
`/flow-auto stop` cancels auto and advice. Advice expires after one hour; retry limits last
until invocation ends. Neither survives restart.

Delegated mode requires release qualification; flags cannot bypass it. Tests prove
mechanics, not decision quality. See [ADR 0016](docs/adr/0016-delegated-recovery.md).
`recoveryStatus.last` exposes process-local timing, reservations, and decision
checks. See [telemetry semantics](docs/maintainer-contract.md#opencode-surface).

Delegation needs release qualification. Tests prove mechanics, not decision quality.
See [ADR 0016](docs/adr/0016-delegated-recovery.md).

## License

Expand Down
27 changes: 14 additions & 13 deletions docs/maintainer-contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -225,17 +225,15 @@ manager contract.
rather than overwrite or delete either side. Closed status re-derives an
archive collision from the existing history document, so interruption cannot
restore automatic retry. This behavior adds no persisted recovery state.
- Every close path whose terminal state was durably accepted returns the same
derived `workflowData.delivery`: initial success, archive-pending recovery,
exact retry, and delayed replay from history. The projection declares a
`handoff` with `formatVersion: 1` and
`externalActionAuthority: "not-granted"`, then contains
the goal, closure, completed/total progress, every planned feature's attempt
count, latest outcome, terminal findings, Flow-reported artifact groups, and
derived tiered assurance with explicit limitations.
- Delivery is recomputed from the canonical closed Session or archive. It is not
written into Session v5 or archive JSON and is not a report artifact unless
the user separately requests one.
- Every durably accepted close returns identical derived `workflowData.delivery`
on success, archive-pending recovery, exact retry, and delayed history replay.
`handoff` declares `formatVersion: 1` and `externalActionAuthority: "not-granted"`.
Delivery contains goal, closure, progress, each feature's attempts, outcome,
terminal findings, reported artifact groups, and tiered assurance limits.
`summary.lines` is the default handoff. Full detail remains in `report` in
that same close response. Blocked `statusReport` stays unchanged.
- Delivery derives from the closed Session or archive. Session v5 and archive
JSON store neither projection nor report. A report artifact needs a user request.
- Source identity hashes sorted effective workspace path/type/content tuples;
`.git` and `.flow` are excluded. It is a content fingerprint, not a Git audit
chain.
Expand All @@ -257,8 +255,11 @@ implicit selection. See [Session v5](#session-v5).
With `TYPESAFE_API_KEY`, `/flow-auto` defaults to shadow advice (six attempts,
$0.02). Off disables advice, not retry limits. First retry stays automatic;
fresh direction permits one retry/start. Shadow grants nothing; release
delegation is disabled. `recoveryStatus` separates configuration, attempts and
outcome. See
delegation is disabled. `recoveryStatus.last` adds process-local assessment duration,
transport attempts, reservation deltas, scores, and threshold checks. `reservedUsd`
is an upper-bound reservation, not billed cost. `responseUsage` counts only the
final validated response, excluding failed attempts. Without validated usage it
is null. Custom providers may leave transport facts null. See
[ADR 0016](adr/0016-delegated-recovery.md).

### Commands
Expand Down
3 changes: 2 additions & 1 deletion docs/quickstart.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,8 @@ Choose work where an incorrect change would be expensive:
Flow inspects the repository and proposes an immutable feature plan. Confirm the
canonical repository gate and any external evidence before approval. After
approval, Flow implements one feature at a time, observes validation, dispatches
an independent review, and closes with a versioned delivery report.
an independent review, and closes with a short delivery summary. The same close
response retains the full versioned report for requested detail.

If the host cannot continue between features, run `/flow-run` for each next
feature. Run `/flow-status` after any interruption or failed review.
75 changes: 64 additions & 11 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -198,6 +198,53 @@ discharged the entry before assertions existed. Declaring the command is no long
enough; the plan has to name the case. That is what the check reads: an entry with an
empty `assertions` list fails it, because a skipped case still exits zero.

## Delivery handoff pilot

Five report-only cases exercise real workflows, archives, host validation, and
independent review. They cover concise completion, deferral with unavailable
macOS proof, a nonzero audit beside a separate passing gate, an ordinary full-report
followup, and idle status after close. Fixtures keep verification scripts immutable.
Recovery is off, so these cases need no Jev calls and measure no Jev decision quality.

After paid authorization, run one attempt per case on the existing OpenAI route:

```bash
env -u TYPESAFE_API_KEY -u OPENCODE_FLOW_REVIEWER_STEPS \
OPENCODE_FLOW_REVIEWER_MODEL=openai/gpt-6.1-sol bun run eval -- --model openai/gpt-6.1-sol --repeat 1 --concurrency 1 \
--scenario delivery-summary-completed --scenario delivery-summary-deferred \
--scenario delivery-summary-observed-failure --scenario delivery-full-detail-followup \
--scenario delivery-idle-after-close
```

This pilot has eight manager dispatches, including three followups, plus one
same-route entitlement probe. A distinct reviewer configuration adds one probe
per additional route. Reviewer-child generation remains paid work inside each
workflow. Nine harness dispatches are neither a dollar cap nor a limit on native
model requests. The parent authorizes and executes the paid run separately.

Default summaries and full detail use separate cases because outcome collection
retains only the last manager text part. Earlier answers and multipart presentation
are not independently graded. Exact accepted close replay is valid full-detail
access when it preserves the same archive and operation.

Graders use literal facts, the actual archived state, and native close/status
provenance. They compare full detail with the accepted close response and reject
missing limitations, false passes, changed verification scripts, and stale idle
handoffs. They import no production delivery formatter. Required release catalogs
remain unchanged. This single-route pilot establishes no release qualification.

Summary grading accepts canonical fields and a bounded set of closure, progress,
assurance and authority sentences. It requires the unchanged Goal line and coherent
current assurance disclosures. Unsupported critical assertions fail instead of
guessing their meaning. Full grading permits Markdown sections and split fields,
while comparing each substantive record's context, value and multiplicity.
Saved pilot answers are development regressions. Regrading them does not establish
the behavior of new prompts or replace fresh live confirmation.

Missing-summary fallback and unknown native exit remain deterministic compatibility
coverage. Current real close responses always include a summary, and ordinary
completed native commands supply an exit. This pilot does not inject either shape.

## Cross-scenario metrics

The original measures are reported for every run and asserted by none. Two are
Expand Down Expand Up @@ -369,6 +416,12 @@ source drift between arming and observing, an abort, an excluded ask — carries
`fidelity` note and is **reported, not gated**, on the same principle the
thresholds use: gate what is measured, report what is not.

Scenarios that require native host provenance declare `replayRequires`. Their
recordings report `UNSUPPORTED` because decision replay lacks native call bindings
and host trace. Runtime handler, closure and completion-honesty differences remain
visible. Replay derives the capability note from the current scenario even when
an older cassette has empty `fidelity`, without rewriting the original bytes.

Capture identities are retained only when the appended marker matches a stored
observation and its command. Replay binds those IDs to newly persisted captures;
unknown or superseded references still fail. Older recordings without capture
Expand All @@ -385,9 +438,10 @@ the developer's real `auth.json` into its throwaway home, so this is a hard rule
rather than a precaution; `tests/eval-replay.test.ts` pins it.

Only recordings someone has read belong in the committed `evals/cassettes/` set,
which is what CI gates on. `--accept` rewrites a cassette's recorded expectation
from the current replay; it is a deliberate act, and the rewritten expectation
lands in the diff to be reviewed like any other change to what the suite asserts.
which is what CI gates on. `--accept` rewrites supported cassette expectations
from the current replay. Review those changes like other test expectations.
It refuses cassettes with unavailable evidence, preserves their bytes and exits
with failure.

The driver itself is proven without a model: `tests/eval-replay.test.ts` hand-writes
the decision sequence of a passing `happy-path` attempt, replays it, and grades it
Expand Down Expand Up @@ -477,14 +531,13 @@ different things:
against the prompts. A scenario that sets `mayEscalate` is the exception: there
the ask is the end the contract leaves, so the run is checked like any other and
reads `PASS+ASK` or `FAIL+ASK`.
- `ABORT` — a step ended without going quiet, either `wedged` (no new message or
part while tool calls stayed incomplete, each named with the first line of its
command) or `still working` (producing output up to the deadline, so looping
rather than stuck). A wedge is called at three minutes of no change rather than
waited out to the twenty-minute deadline: three of the four recorded timeouts sat
on the same incomplete tool call for the full twenty and then printed exactly that
diagnostic, so the remaining seventeen minutes bought no evidence. Tokens and tool
calls collected before the abort are kept. Excluded from the pass rate and counted
- `ABORT` — a step ended without going quiet. Diagnostics distinguish no new
messages or parts while tool calls stay incomplete from new messages or parts
near the deadline. Updates inside existing parts are not measured. Neither
diagnostic establishes whether the model is making useful progress. The harness
aborts after three minutes without new messages or parts while calls stay
incomplete, or at the twenty-minute hard deadline. Tokens and tool calls
collected before the abort are kept. Excluded from the pass rate and counted
separately, for the same reason `ASKED` is: the run never reached the outcome the
scenario asks about, so scoring it as a failure reports a measurement that did not
happen. One wedged attempt was the only failing threshold in a recorded report.
Expand Down
16 changes: 12 additions & 4 deletions evals/cassette.ts
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,8 @@ export type FidelityNote =
| "provider-error"
| "evaluator-error"
| "validation-identity-unrecorded"
| "workspace-diff-unreplayed";
| "workspace-diff-unreplayed"
| "native-host-provenance-unreplayed";

export type Cassette = Readonly<{
cassetteVersion: number;
Expand Down Expand Up @@ -300,14 +301,20 @@ function validationReport(output: string): string | null {
}

/** A cassette is gated only when nothing about it is known to be unreproducible. */
export function isGated(cassette: Cassette): boolean {
return cassetteFidelity(cassette).length === 0;
export function isGated(
cassette: Cassette,
requirements: readonly "native-host-provenance"[] = [],
): boolean {
return cassetteFidelity(cassette, requirements).length === 0;
}

export function cassetteFidelity(
cassette: Pick<Cassette, "events" | "fidelity">,
requirements: readonly "native-host-provenance"[] = [],
): readonly FidelityNote[] {
const fidelity = new Set(cassette.fidelity);
if (requirements.includes("native-host-provenance"))
fidelity.add("native-host-provenance-unreplayed");
let identitiesRecorded = false;
for (const event of cassette.events) {
if (event.kind === "bash" && event.validation) identitiesRecorded = true;
Expand Down Expand Up @@ -399,6 +406,7 @@ export function buildCassette(options: {
readonly falseCompletion: boolean;
readonly documents: readonly Record<string, unknown>[];
readonly extraFidelity: readonly FidelityNote[];
readonly replayRequires?: readonly "native-host-provenance"[];
}): Cassette {
const events: CassetteEvent[] = [];
const pendingResults = new Map<
Expand Down Expand Up @@ -522,7 +530,7 @@ export function buildCassette(options: {
},
finalText: options.finalText,
assistantMessages: options.assistantMessages,
fidelity: cassetteFidelity({ events, fidelity }),
fidelity: cassetteFidelity({ events, fidelity }, options.replayRequires),
} satisfies Cassette,
options.projectPath,
);
Expand Down
Loading
Loading