Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
47 commits
Select commit Hold shift + click to select a range
7585219
[Spike task-Iptx] Findings: Kimi Code CLI feasible with caveats as bu…
mohidmakhdoomi Jul 18, 2026
af1bc53
[Spike task-Iptx] Addendum: task-delivery readiness barrier + correct…
mohidmakhdoomi Jul 18, 2026
1e22513
[Spike task-Iptx] Thread: post-review addendum logged
mohidmakhdoomi Jul 18, 2026
1e8d3c6
chore(porch): 1201 init pir
mohidmakhdoomi Jul 18, 2026
d49c292
[PIR #1201] Plan draft
mohidmakhdoomi Jul 18, 2026
73774d9
chore(porch): 1201 plan-approval gate-requested
mohidmakhdoomi Jul 18, 2026
fe7114e
chore(porch): 1201 plan-approval gate-approved
mohidmakhdoomi Jul 18, 2026
b26c565
chore(porch): 1201 implement phase-transition
mohidmakhdoomi Jul 18, 2026
2cf424c
[PIR #1201] Kimi harness: detection, seed-session launch script, buil…
mohidmakhdoomi Jul 18, 2026
8e86c41
[PIR #1201] Tower: sentinel-gated BEGIN delivery + per-harness Enter …
mohidmakhdoomi Jul 18, 2026
3d40785
[PIR #1201] doctor: kimi presence, truthful auth heuristic, store smo…
mohidmakhdoomi Jul 18, 2026
f075443
[PIR #1201] Docs: kimi builder harness (arch.md + config examples, sk…
mohidmakhdoomi Jul 18, 2026
b27e2d3
[PIR #1201] Pacing resolution is fully best-effort; widen cron sessio…
mohidmakhdoomi Jul 18, 2026
ea6607c
[PIR #1201] Pin Kimi Enter delay with live bisect evidence
mohidmakhdoomi Jul 18, 2026
6b39ca5
[PIR #1201] Live demo driver + results (all 5 checklist steps pass)
mohidmakhdoomi Jul 18, 2026
671e8dd
chore(porch): 1201 dev-approval gate-requested
mohidmakhdoomi Jul 18, 2026
215cda2
chore(porch): 1201 dev-approval gate-approved
mohidmakhdoomi Jul 19, 2026
fee9015
chore(porch): 1201 review phase-transition
mohidmakhdoomi Jul 19, 2026
e8fe2ef
[PIR #1201] Review + retrospective (lessons routed to cold tier)
mohidmakhdoomi Jul 19, 2026
b80dbb4
chore(porch): 1201 record PR #1203
mohidmakhdoomi Jul 19, 2026
ad19c96
chore(porch): 1201 review build-complete
mohidmakhdoomi Jul 19, 2026
732f04b
[PIR #1201] Fix seed-kick confirmation false-positive (codex consulta…
mohidmakhdoomi Jul 19, 2026
b0ad027
chore(porch): 1201 pr gate-requested
mohidmakhdoomi Jul 19, 2026
3c0b6fc
[PIR #1201] Thread: review phase + CMAP disposition logged
mohidmakhdoomi Jul 19, 2026
53bfea0
chore(porch): 1201 pr gate-approved
mohidmakhdoomi Jul 19, 2026
2f889cf
chore(porch): 1201 protocol complete
mohidmakhdoomi Jul 19, 2026
08d4311
[PIR #1201] Thread: pr gate approved, porch wrapped; PR open for main…
mohidmakhdoomi Jul 19, 2026
47d12ba
[PIR #1201] Record CMAP iter-1 disposition (codex finding accepted+fi…
mohidmakhdoomi Jul 19, 2026
2abd362
[PIR #1201] Fix: persist the Kimi pacing marker on the bare launch shape
mohidmakhdoomi Jul 23, 2026
642b172
[PIR #1201] Test: pin touch-before-loop ordering in bare-shape marker…
mohidmakhdoomi Jul 23, 2026
1de55e1
[PIR #1201] Test: complete existence guards on ordering assertions (C…
mohidmakhdoomi Jul 23, 2026
e3bfa56
[PIR #1201] Thread: CMAP loop converged (iter 3 — 3x APPROVE, zero fi…
mohidmakhdoomi Jul 23, 2026
fecd6c9
Merge origin/main into builder/pir-1201
mohidmakhdoomi Jul 23, 2026
99bb044
Merge remote-tracking branch 'origin/main' into builder/pir-1201
mohidmakhdoomi Jul 25, 2026
bfc8d62
[PIR #1201] Adopt shared LAUNCH_LOOP_TAIL (#1244) in Kimi provider-ow…
mohidmakhdoomi Jul 25, 2026
277967a
[PIR #1201] Thread: CMAP + live kimi verification of loop-tail adoption
mohidmakhdoomi Jul 25, 2026
1699119
Merge remote-tracking branch 'origin/main' into builder/pir-1201
mohidmakhdoomi Jul 27, 2026
ae0d034
Merge origin/main into builder/pir-1201; redesign the Kimi builder on…
mohidmakhdoomi Aug 9, 2026
1424213
[Spec 1201] fix: close a false-CLEAN on kimi's multi-row composer
mohidmakhdoomi Aug 9, 2026
2c34cc7
[Spec 1201] fix: make the resume probe answer what kimi actually cont…
mohidmakhdoomi Aug 9, 2026
c8002da
[Spec 1201] docs: record the CMAP round, and fix the demo's role oracle
mohidmakhdoomi Aug 9, 2026
4a7e2af
[Spec 1201] docs: builder thread — post-pivot CMAP round
mohidmakhdoomi Aug 9, 2026
b8df898
[Spec 1201] fix: hold a kimi draft whose every cell is exempt chrome
mohidmakhdoomi Aug 9, 2026
82b888b
[Spec 1201] docs: make the launch loop's task messaging tell the truth
mohidmakhdoomi Aug 9, 2026
1d09234
[Spec 1201] docs: record the architect-review round and its CMAP disp…
mohidmakhdoomi Aug 9, 2026
6520552
[Spec 1201] fix: make a clean exit stick across a pre-mint crash
mohidmakhdoomi Aug 9, 2026
781ca59
[Spec 1201] docs: record the sticky-fresh contract and finding 4's di…
mohidmakhdoomi Aug 9, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 35 additions & 0 deletions codev-skeleton/resources/commands/agent-farm.md
Original file line number Diff line number Diff line change
Expand Up @@ -886,6 +886,41 @@ afx workspace start --architect-cmd "claude --model opus"
afx spawn 42 --protocol spir --builder-cmd "claude --model haiku"
```

### Builder harnesses

The builder CLI's role/prompt mechanics are handled by a harness, auto-detected
from the command basename (`claude`, `codex`, `opencode`, `kimi`) or pinned
explicitly via `shell.builderHarness`. Example — Kimi Code CLI as the builder
(builder-only; requires kimi >= 0.33.0):

```json
{
"shell": {
"builder": "kimi"
}
}
```

Kimi takes no positional prompt, so a Kimi builder gets its role and its task
through two different channels: the role via `--agent-file` (an agent-definition
file written into the worktree, composed around kimi's `${base_prompt}` token so
it extends rather than replaces kimi's own system prompt), and the task via the
`afx send` mailbox, delivered onto a verified-empty composer by the render gate.
A crashed builder resumes with `kimi -c`, but only once a store probe confirms a
conversation exists for that worktree — `kimi -c` with nothing to continue
silently starts a fresh, roleless session, so the probe fails closed to a
role-carrying fresh launch instead.

Two notes specific to Kimi builders. Spawning pre-records workspace trust for the
new worktree, because kimi 0.33.0+ opens on a "Trust this folder?" dialog that an
unattended builder cannot answer (trust gates only whether project-level MCP
servers load; it does not gate tool execution). And Kimi builders do NOT yet get
the worktree write-guard Claude builders have — kimi does have a blocking
`PreToolUse` hook seam, so parity is achievable follow-up work rather than a
permanent limitation.

Architect use of kimi and opencode is unsupported (use claude or codex there).

### Mailbox retention and escalation

`afx send`'s mailbox (Spec 1313) has two Tower-global knobs under a `mailbox` key:
Expand Down
182 changes: 182 additions & 0 deletions codev/plans/1201-support-kimi-code-cli-as-a-bui.md

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
# CMAP dispositions — architect integration review follow-up (2026-08-09)

Round: the architect's three non-blocking findings on PR #1203 at head `4a7e2afe`, plus the
3-way review of the resulting delta. Prior round's dispositions are in
`1201-cmap-postpivot-dispositions.md`; this file covers only this delta.

Verdicts on the delta: **gemini APPROVE · codex REQUEST_CHANGES · claude APPROVE-with-changes**.
Every finding from both non-approving reviews was accepted. Nothing was rejected.

Scope note: a mid-round architect message fenced the PR's three open maintainer decisions (trust
pre-write, 0.33.0 version floor, write-guard parity as follow-up). None were touched.

---

## Architect finding 1 — residual false-CLEAN for an all-exempt draft

**Measure-first, per instruction. The premise holds**, so the rule was implemented rather than
documented as a residual.

Measured on real kimi 0.34.0 (`codev/spikes/pir-1201-kimi-box-growth.mjs`), interior rows =
`endRow - startRow`, the rows the classifier actually scans:

| state | interior rows |
|---|---|
| idle | 1 |
| single-line draft | 1 |
| `/` command menu | 1 |
| `@` file picker | 1 |
| post-reply steady state | 1 |
| newline + bare `>` | **2** |
| newline only | **2** |
| long soft-wrapped single line | **2** (carries text → already busy; verdict unchanged) |

The steady-state row is the load-bearing one: growth on a composer that has already carried a
turn would hold every later message forever — a liveness failure, which is worse than the
fail-safe direction the gate normally errs toward.

The review then surfaced a class the spike had not enumerated (claude Q2), so it was measured
too (`pir-1201-kimi-working-states.mjs`): mid-generation at 5s and 13s, shift+tab mode chrome,
and a draft typed while the agent is working are **all one interior row**. So the rule does not
convert "deliver while busy" into "hold until idle". (`!` bash mode classifies
`no-composer-marker` and holds — pre-existing, fail-safe, and correct: there is unsent input on
that row.)

### Deviation from the suggested implementation

The architect's sketch short-circuited on geometry **before** the cell scan. Implemented that
way it changed an existing fixture's verdict detail — `kimi-multiline-bare` went from
`user-text` to `multi-row-draft`, because that draft is also multi-row — which would have
retired what the older guardrail test was actually testing and demoted the cell scan from
ground truth to dead weight on every multi-row screen. Moved **after** the scan: `userCells > 0`
still wins and still reports `user-text`; `multi-row-draft` is reserved for the case the count
is blind to. Every pre-existing fixture verdict is unchanged. claude independently confirmed
this ordering is not just preferable but *enforced* by the existing assertion at
`render-gate.test.ts:194`.

---

## codex #1 / claude Q5+F1 — arming coupled to `regionStartPatterns` — **ACCEPTED**

Both reviewers independently flagged that arming the rule off `regionStartPatterns` overloads a
field that means "the composer has an upper boundary" with an unrelated claim ("box height
tracks draft lines").

claude supplied evidence that makes this concrete rather than stylistic, **which I verified
myself** with a geometry probe over every shipped fixture: `codex-idle.clean.txt` — a real,
captured, genuinely **empty** codex composer — already spans **two interior rows**
(`marker=18 start=18 end=20`). The rule's geometric predicate is *already true* on a screen that
must stay clean; only the arming gate stands between that capture and codex mail being held
forever. The day anyone declared a region start for codex (a header bound, a boxed redesign),
delivery would die silently.

Decoupled into an explicit profile field, `growsWithDraft?: true`, set only on `KIMI_PROFILE`.
The rule now requires **both**: `growsWithDraft` (the measured promise) and `hasRegionStart`
(what makes the arithmetic mean "interior rows" at all). codex proposed
`maxCleanInteriorRows?: number` instead; chose the boolean because it encodes the *measured
premise* rather than a tunable number, and a wrong threshold under it is caught by the app's own
idle fixture, which must classify clean. Pinned by three tests, all now built on codex's real
capture rather than a constructed screen: inert when neither field is set, inert with either one
alone, and armed only with both.

## codex #2 / claude F4 — fast-fail hint not universally accurate — **ACCEPTED**

My reworded echo asserted unconditionally that "an undelivered task is still queued on the
mailbox". False in a reachable third case: `codev_task_queued` is set only on a **successful**
`afx send`, so if afx is off PATH or Tower is down the flag is still 0, nothing is queued, and
the fresh relaunch really does retry it. The hint now branches on `[ "$codev_task_queued" = 1 ]`
and states the truth in both cases. Behavior still unchanged; this was a message-accuracy fix on
top of a message-accuracy fix.

## codex #3 / claude F5 — tradeoff comment overstates delivery — **ACCEPTED**

"true whenever the operator saw a composer to /quit from" is too strong: seeing a composer is
necessary, not sufficient — the gate also has to have polled it empty at least once. Softened,
and the quit-before-delivery race is now named alongside the trust-dialog case.

## claude F2 — `isClassifierStuck` silently omitted the new detail — **ACCEPTED**

`mailbox-delivery.ts` enumerated stuck details as a closed `||` chain, so widening
`GateVerdict['detail']` did not force a decision. claude checked the resulting behavior and it
was *right* (excluding `multi-row-draft` is correct — it is a human on a draft, and the
premise-failure reading carries no recent output, which `surfaceLiveness` requires to alarm),
but it read as an oversight. Replaced with a `Record<GateVerdict['detail'], boolean>` map, so
the next new detail is a **compile error** rather than a silent `false`, and documented why
`multi-row-draft` sits on the excluded side.

## claude F3 + codex test note — the differential used constructed screens — **ACCEPTED**

codex noted the armed/unarmed halves were "not literally the same bytes despite the comment";
claude asked for the real codex-idle geometry to be cited. Both are answered by the same change:
the test now classifies the **actual `codex-idle.clean.txt` capture** under four profile
variants — shipped, armed, bounded-only, grows-only — so the differential runs on identical real
bytes and the hazard is demonstrated rather than described.

## claude Q1 nit — non-null assertion — **ACCEPTED**

`hasRegionStart` is now a type predicate (`patterns is RegExp[]`), so the `startPatterns!`
assertion is gone and the narrowing is checked rather than conventional. (Applied before
claude's review landed; it had read the pre-edit file.)

---

## Verification

- `pnpm build` clean; `tsc --noEmit` clean.
- Full suite **4906 passed / 48 skipped / 0 failed** (+6 on the pre-round 4900: five new tests
and one new fixture).
- Targeted suites — render-gate, harness, harness-integration, spawn-worktree, mailbox-pacing,
kimi-session-discovery — green.
- Two new live measurements against real kimi 0.34.0, both committed as reproducible spikes.
- No live demo re-run: the rule can only alter verdicts for a kimi composer past one interior
row, and delivery targets the idle composer, measured at one row in every state including
mid-generation.
Original file line number Diff line number Diff line change
@@ -0,0 +1,139 @@
# CMAP dispositions — finding 4, kimi sticky-fresh crash-resume (2026-08-09)

Round: the architect's finding 4 on PR #1203 — after a clean exit, a crash in the pre-mint boot
window makes `kimi -c` resume the conversation the human just ended (#1267's own motivating
defect class). Earlier rounds: `1201-cmap-postpivot-dispositions.md`,
`1201-cmap-architect-review-dispositions.md`.

Verdicts on the delta: **gemini APPROVE · codex REQUEST_CHANGES · claude REQUEST_CHANGES**.
Both REQUEST_CHANGES were right, and they converged on the same blocking defect. Every finding
was accepted; none rejected.

Scope fences respected: no PR comment, and the three parked maintainer decisions (trust
pre-write, 0.33.0 floor, write-guard parity) untouched.

---

## Measurement first — the premise holds

The fix (and the pre-existing resume design) assumes `kimi -c` continues the NEWEST session when
a cwd holds several. The existing continue-probe only covered the ZERO-session case, so this was
measured net-new on real kimi 0.34.0
(`codev/spikes/pir-1201-kimi-continue-newest-probe.mjs`), with two independent oracles because
the model's own answer is not proof:

- **content oracle** — sessions seeded with distinct codewords ALPHA (older) / BRAVO (newer);
`kimi -c` answered **BRAVO**.
- **identity oracle** — snapshot `updatedAt` for every session before and after; the `-c` turn
touched **only** `session_f06c…` (the newest), and **created no new session**. Exit 0, no
prompt.

So identity comparison is well-defined, and the fallback ("document the residual instead") did
not apply.

---

## The fix

The inlined store probe now PRINTS the newest resumable session id instead of exiting 0/1; the
clean-exit branch records that id as superseded; the crash branch takes `-c` only once the
newest id differs. One probe, one mirror — the boolean uses derive from the same output.

---

## codex #2 / claude F1 — BLOCKING: the guard read stdout and discarded exit status — **ACCEPTED**

The delta moved the decision from `$?` onto stdout, so anything else writing to stdout is read
as "a session exists". claude **measured** it: with an empty store and
`NODE_OPTIONS=--require <module that prints>`, the probe printed a banner and exited 1, and the
script read RESUME. That lands on `kimi -c` with nothing to continue — which does not fail, it
starts a session that never saw `--agent-file`: a silently **roleless** builder, the #929 class
the entire guard exists to prevent. **A failure mode the delta introduced** — the pre-delta code
could not produce it. Vectors: `NODE_OPTIONS`, a `node` shim on PATH, corporate instrumentation
preloads.

Fixed by consuming both signals, with the declaration split from the assignment so `local` does
not mask the substitution's status:

```bash
local codev_newest
codev_newest=$(codev_newest_session) || return 1
[ -n "$codev_newest" ] && [ "$codev_newest" != "$codev_superseded_id" ]
```

Pinned by a new test that reproduces the exact vector.

## codex #1 / claude F2 — a transient probe failure at clean exit re-opens the gap — **ACCEPTED**

The architect's sketch said "empty on any error — fail-closed", and my comment repeated it. Both
reviewers showed it is not: if the probe fails transiently (EMFILE, ENOMEM, fork failure, a
throwing preload) the branch records `''`, and the next crash sees the just-ended session as
"different from empty" → resumes it. The very bug the finding is about.

claude's suggested mitigation (keep the previous value) only helps on *iterated* exits; the
first clean exit still records nothing. So the branch now distinguishes **failure** from **empty
store** by status and sets `codev_resume_blocked`, which refuses resume until the next clean
exit re-establishes a baseline. Accepted cost, documented in-code: a later crash restarts fresh
instead of continuing, losing conversation continuity — never the role (fresh always carries it)
and never the task (the mailbox still holds it). It self-heals at the next clean exit.

## claude F3a / codex #4 — `j.cwd ?? j.workDir` is not the mirror — **ACCEPTED**

Discovery's `readStateJson` tests `typeof === 'string'` **per field**; the probe's `??`
short-circuits on any non-null `cwd`, so `{cwd: 12345, workDir: <match>}` was found by discovery
and missed by the probe. Fail-closed in direction, but it disproves the field-for-field claim
the docstring makes — and identity, not just existence, now rides on that claim. Probe changed
to per-field `typeof`; docstring corrected to stop naming `cwd ?? workDir` as the mirror; a
fixture added.

## claude F3b — trailing-slash normalization diverged, in the UNSAFE direction — **ACCEPTED**

The probe's `n()` stripped a trailing slash *before* `realpathSync`; `sameDir` does not. For a
path that does not exist, `/ghost/` canonicalized to `/ghost` in the probe and stayed `/ghost/`
in discovery — so the probe could name a session discovery rejects. claude called it unreachable
(the probe's argument is `$PWD`, which exists) and said record it. Removed instead: the strip
bought nothing, because `realpathSync` already normalizes a trailing slash away for any
directory that exists — which is exactly what the existing trailing-slash fixture covers, and it
still passes. Exact mirror beats documented exception. Fixture added for the ghost case.

## claude F4 — the composition was never executed, only the pieces — **ACCEPTED**

`decideBranch` injects `codev_superseded_id` from the test, so the only evidence the generated
clean-exit branch assigns it was a string match. A refactor wrapping that assignment in a
subshell — an ordinary bash footgun — would pass every test while the contract was dead. Added a
test that drives the **real `while` loop** with stubbed launches and a fed `read -r`, asserting
the branch sequence is `resume, fresh, fresh` (entry resumes; clean exit goes fresh and retires
the id; the pre-mint crash stays fresh).

## claude F5 — "unreadable store" tested an ABSENT store — **ACCEPTED**

The test never wrote a session, so `rmSync` removed nothing and it duplicated the
store-does-not-exist case. Rewritten: write a session that WOULD authorize `-c`, then replace
`sessions/` with a regular file for a deterministic ENOTDIR (root-proof, unlike `chmod 000`).

## claude nits — **ACCEPTED**

Restored the stronger `not.toContain('codev_launch_resume')`; the `afterClean` slice now bounds
on the branch's own two-space-indented `fi` (the earlier `\n\s*fi\n` stopped at the new nested
conditional — the same class of bug as the `"fine"` match it replaced).

## codex edge notes — traced, documented, not engineered against

- **Store GC drops the newest session** while an older abandoned one survives → `-c` reaches the
older one. Requires a retention policy that evicts newest-first. Recorded in-code.
- **`afx spawn --resume` / terminal re-create** resets the in-memory superseded id. Documented
as the intended boundary — and claude noted this is **contract parity**, not a kimi shortfall:
claude's minted id is equally per-process.
- **Two builders in one cwd** — not a real topology (one worktree per builder).

---

## Verification

- `pnpm build` clean; `tsc --noEmit` clean; generated script passes `bash -n`.
- Full suite **4915 passed / 48 skipped / 0 failed** (+9 on the round's starting 4906).
- Targeted suites (harness, harness-integration, spawn-worktree, kimi-session-discovery,
mailbox-pacing, render-gate) green.
- Non-vacuity is demonstrated rather than asserted: `decideBranchLegacy()` runs the pre-fix
existence-only predicate against the same store and the same generated probe, and the
regression test asserts it returns RESUME where the shipped guard returns FRESH.
Loading
Loading