docs(book): explain long-running architecture and validate reader exercises - #5354
huangruiteng merged 38 commits into
Conversation
Register three new chapters in both locale configs and the publication smoke, and renumber the displayed chapter sequence so the book reads as one path. New chapters: - 02b long-horizon-requirements: the four demands long-running work makes (state outlives the context, interruption stops at an identifiable position, one accountable writer, bounded observable consumption) - 04b budget-and-admission: admission, backoff, and monitor-driven waking - 07b when-loopx-is-not-the-answer: the four explicit refusals and the cost table that follows from them The nav now numbers chapters 1-20 continuously in both locales; the topic chapter and appendix stay unnumbered, as before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
Restructure all 25 chapters (zh + en) so the book answers one question instead of touring components: once work runs long enough, what must a system satisfy, and what does LoopX choose to do about it? Each chapter now follows the same shape: a concrete bad ending, why the intuitive fix does not hold, the design, what that design gives up, the named failures with runnable evidence, and checkable invariants. The "cost and boundary" section is the one the previous edition largely lacked, and it is where the answer to "why is it designed this way" actually lives. Substance is preserved, not compressed: every chapter kept its original sections, tables, commands, field names, and protocol anchors. Where a claim had gone stale against the code it was corrected rather than carried forward - the Todo completion compare-and-swap anchor now points at completion_transaction.py, whose field-snapshot comparison replaced the retired event-writeback path in loopx-project#5054. Bilingual parity is enforced per section: heading counts, code blocks, and tables match one for one across all 25 chapters. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
songoow
left a comment
There was a problem hiding this comment.
Request changes conclusion (author-owned PR; GitHub blocks formal self-review)
Reviewed head: c1923db27efc667d0eb021640719978c3dd637d2; base: 0538bf1631a7a9469aa89b0f993ddada3b61244c.
结论:REQUEST_CHANGES。问题引入、架构解释和使用边界的方向正确;当前稿件仍有会误导恢复操作及 Host 选择的事实错误。下面列出 22 个经源码、契约或实际渲染核实的问题:3 个 P1、18 个 P2、1 个 P3。P1 阻断本次批准,其他项用于这次稿件的集中修订;没有把篇幅或风格偏好当作阻断依据。
动机
这次改写要让读者理解 LoopX 如何处理长程任务,并据此判断怎样接入、何时等待、怎样恢复及哪里需要人参与。三篇新增章节把持久状态、中断恢复、执行归属和预算观察串起来,确实比纯功能清单更容易进入。审查依据包括 PR 的目标声明、当前协议、整体 roadmap 的 S6/S7/S10/S12 和实际调用路径;没有将 RFC 已接受等同于其中所有能力已交付。
改动思路
我覆盖了完整 53 文件 diff:25 章中英正文、两份导航,以及 publication smoke 的三个章节条目。正文沿问题场景进入机制,再补代价、不变式与实践路由;现有协议与代码继续拥有行为,书负责解释。源码没有改动,新的 smoke 条目复用原检查器,也没有引入运行时或另一套状态机。
最需要纠正的是解释强度:有些段落从一个正确的实现片段推出了更强的保证,或为了写出“代价”而增加系统没有的限制。真实路径是 current decision → Host → validation → durable writeback → settlement/readback;恢复还依赖身份与未决 effect 的读回。合法前缀校验只是其中一个条件。对接口和调用者的边界,应按具体 contract owner 判断,不能按语言或术语作全局推论。
具体改动
关键内容讲解
- 基础与架构章负责问题、状态、工作图、Turn、恢复及预算;保留了核心状态机图,但需修正当前存储和恢复承诺。
- 接入与 Host 章负责实际启动及持续执行边界;Extension 三章负责 placement、scaffold 和生命周期,需要让连续照做的路径一致。
- 贡献、工程边界及附录负责找 owner、选择验证和提交证据;仍有迁移状态、证据语义、编号及渲染链接问题。
- 导航新增 02b/04b/07b 并保留旧 slug,方向合理;smoke 只增加三个清单项,守住公开路由存在性,不能证明文案语义正确。
实现与规范核对(Standards / implementation)
F1 [P1] 恢复前缀不等于无条件续跑 — docs/book/chapters/03-one-turn.md。
正文说“只要读最后完成阶段,从下一个继续”;实际会按 effect 身份读回 prepared 结果,unknown 时 fail closed。矩阵测试的合法前缀和 mock 回执不能证明任意崩溃点的自动恢复。 依据:loopx/control_plane/turn_driver/settlement.py。修改:区分已提交、确认未提交、结果未知三个恢复分支,说明恢复判定、身份与 provider readback 的前提;同步英文第 61 行及章末不变式。
F2 [P1] 扣额失败时保留写回,不能教读者丢弃结果 — docs/book/chapters/03-one-turn.md。
新增文字称“系统选择丢弃它”,与实现正相反:state_written=true、quota_spent=false,恢复后 Host/写回各仍只执行一次。第 68 行也把 quota 记账写成实际收费保证;docs/quota-allocation.md:228 说明当前为预算 slot,并非模型账单。 依据:tests/test_loopx_turn_executor.py。修改:改为保留已完成写回,按原身份恢复尚未完成的结算;区分 LoopX quota slot 与模型/外部服务真实费用,避免重跑已经生效的操作。
F3 [P1] 区分 TUI 内自主继续与进程退出后的定时唤醒 — docs/book/chapters/07-codex-cli.md。
“你不发消息,就没有下一轮”以及“放弃无人值守继续推进”把无 App heartbeat 扩大成无自主 continuation。生成 Goal body 明确要求结算后重查 quota 并继续允许的工作;rules.py:103 才规定进入 blocked 后需显式 resume。06 章末也重复了这个误导。 依据:loopx/control_plane/heartbeat/task_body.py。修改:开场保留“关闭进程后不会凭空运行”的边界,分别解释 running、blocked、process exited;不要要求用户逐轮触发。
F4 [P2] 恢复被删掉的 Todo event 退役限定 — docs/book/chapters/state-substrate.md。
此次删除了原有退役提示,却保留 event-ledger/replay 内容,导致旧实验重新被读成当前实现;legacy_event_source.py:36 会拒绝非空的旧源。 依据:docs/reference/protocols/event-sourced-state-contract-v0.md。修改:明确 legacy Markdown 与选定 File/SQLite authority 的当前职责,将已退役的 Todo events/replay 移为历史解释或删去。
F5 [P2] 不能把不受支持的手改描述成没有效果 — docs/book/chapters/state-substrate.md。
正文称改成 [x] 不会改变 status。实际 legacy read 仍从 Markdown 解码;合成输入实测同一行由 open 变 done。规范上不应手改,并不等于技术上不会改变源状态。 依据:loopx/control_plane/todos/todo_block_codec.py。修改:按 authority mode 区分 source 和 projection,并写清手改会绕过回执/验收,不能声称它无效。
F6 [P2] 唯一执行实例的保证需要限定生效的 fencing 模式 — docs/book/chapters/work-graph-and-authority.md。
新不变式把唯一实例写成无条件保证;默认 legacy 保留 soft-claim/hard-lease 分裂,todo_lifecycle_decision.ts:476 存在 terminal_fence_not_required 路径。 依据:loopx/control_plane/todos/handoff_mode.py。修改:说明目标性质与实际强制边界的区别,列明 hard_lease/实例围栏生效条件和 legacy 限制。
F7 [P2] 缺少 PR 显式仓库前缀不一定拒绝 — docs/book/chapters/work-graph-and-authority.md。
新失败案例称缺前缀直接拒绝,实际可从 task_repository 绑定仓库;同章 183 行原有解释也与此冲突。 依据:loopx/control_plane/todos/resume_condition.ts。修改:改成显式前缀与 task_repository 均无法提供仓库时才拒绝。
F8 [P2] Workspace 人工动作不能统一归 quota 授权 — docs/book/chapters/workspace-v1.md。
“是否合法仍由 quota decision 判断”错置了 owner lifecycle action 的授权者;当前调用带 actor_kind=owner 和 reviewed fingerprint,quota 决定自动工作准入。 依据:loopx/chat_goal_lifecycle_actions.py。修改:分别写明 owner action、Todo lifecycle、自动 Turn 的授权边界;投影只负责发起请求。
F9 [P2] Capability 不要求先有第二个 Provider — docs/book/chapters/08-extension-placement.md。
新增“没有第二个 Provider 就不建 Capability”给 placement 加了不存在的硬门槛;正式规则要求稳定 caller contract 和真实产品调用。 依据:docs/reference/extensions.md。修改:以调用者合同、内置发行责任、独立生命周期判断放置,不按 provider 数量决定。
F10 [P2] 顺序阅读会重复安装同一 Extension — docs/book/chapters/09-extension-scaffold.md。
第 09 章已执行 install --execute,第 10 章再次 preview/install 会直接报 already installed。 依据:loopx/extensions/runtime.py。修改:第 09 章只生成和修改 scaffold,将一次 package 安装/激活放在第 10 章;前置统一示例 state-file。
F11 [P2] 开场 upgrade 命令参数错误 — docs/book/chapters/10-extension-lifecycle.md。
解析器不接受该 positional id,并要求 --manifest 或 --bundled;纯 parser 复现退出 2,尚未运行 doctor。 依据:loopx/cli_commands/extension.py。修改:改成带真实 manifest 的命令;中英同步,避免把不可执行命令作为机制案例。
F12 [P2] rollback_available=false 表示没有回滚版本 — docs/book/chapters/10-extension-lifecycle.md。
新增指导把 false 解释为环境坏了,应修复环境;该布尔值仅取决于历史 revision 是否存在,rollback_extension 对应逻辑会在 probe 前拒绝无 target。 依据:loopx/extensions/runtime.py。修改:区分无回滚历史、存在 target 但 doctor 失败两种情况,给各自的下一步。
F13 [P2] 保留 activation revision 不保证旧程序仍可运行 — docs/book/en/chapters/10-extension-lifecycle.md。
英文新增保证把元数据原子性扩大成可执行环境原子性;旧 entrypoint identity 可能已变化、readiness 被拒。中文已有同类错误,本次将它传播到英文。 依据:loopx/extensions/runtime.py。修改:只承诺旧 activation 记录保留;恢复先前 package 版本并 doctor/readback 后才能声称旧程序可用。
F14 [P2] 中文附录六个新页内链接没有目标 — docs/book/chapters/appendix-reference.md。
实际 HTML heading id 为 _3 等,新增中文 fragment 不存在;strict build 仅 INFO,不会失败。 依据:mkdocs.yaml。修改:使用显式稳定 heading id,同步路由链接,检查实际渲染 anchor。
目标与写法核对(Spec)
以下问题对应“解释为什么这样设计、怎么使用和边界在哪”及中英语义镜像要求。没有发现足以要求推倒重写的整体定位偏差,也没有把未证实的内容丢失作为 finding。
F15 [P2] 余额仍然参与准入 — docs/book/chapters/04b-budget-and-admission.md。
“余额数字不在其中”与 spent_slots >= allowed_slots → throttled 直接矛盾;should_run_prepare.py:518 用 eligible 判定正常交付。 依据:loopx/quota.py。修改:改成余额是必要输入之一,单看余额不足以授权工作;保留 Gate、能力和 frontier 的优先级解释。
F16 [P2] Backoff 不会消除发现外部变化的延迟 — docs/book/chapters/04b-budget-and-admission.md。
正文承诺外部变化后立即 reset、响应速度不受影响;实际仍需下一次观察或独立 trigger 才能形成新的 identity。02b 的“monitor 替代轮询”也把受治理轮询误写成事件推送。 依据:loopx/control_plane/scheduler/state_transition_rules.ts。修改:写清节省观察成本与发现延迟的取舍;区分按 cadence poll 和已接入的事件唤醒。
F17 [P2] 边界章过度排除了已有偏好记忆能力 — docs/book/chapters/07b-when-loopx-is-not-the-answer.md。
新章称不做偏好记忆,但现有 read/remember 合同支持该场景;roadmap S6 也明确 scoped preference、decision context 和 recall。 依据:loopx/capabilities/semantic_preference/catalog_entry.py。修改:保留“不替代任意知识库”,同时说明 scoped memory 的能力边界、opt-in 和来源要求;不要反向许诺自动保存任意聊天。
F18 [P2] 规则 owner 不能一律按实现语言判定 — docs/book/chapters/source-change-control-plane-rule.md。
正文一律宣称 TypeScript 为 canonical;RFC 明确 task-class、actionability、dependency readiness、agent eligibility 等读规则仍有 Python semantic owner。03:269 同类。 依据:docs/architecture/rfcs/typescript-control-plane-migration-v0.md。修改:按具体已迁移的 bounded context/合同确认 owner,区分架构目标与已交付边界。
F19 [P2] 中英把应记账的规则写反了 — docs/book/chapters/04b-budget-and-admission.md。
中文写“该花的不花”,英文要求应 spend 的必须 spend。结构和表格数相同并不能发现此语义反转。 依据:docs/book/en/chapters/04b-budget-and-admission.md。修改:明确“该记账的必须记,不该记的不能记”,结合所需证据与调用方式逐句核对双语规则。
F20 [P2] 插章后阅读路线仍指向旧章号 — docs/book/chapters/00-reading-guide.md。
导读仍指贡献 11–14、Extension 14–16,实际为 13–16、17–19;英文 Course 的相关表格也保留旧偏移。导读还称基础篇六章,而导航已八章。 依据:docs/book/mkdocs.zh.yaml。修改:两版统一按稳定 slug 链接标识目标,再更新显示编号和前后章承诺。
F21 [P2] 规则修改的开场反例因果相反 — docs/book/chapters/source-change-control-plane-rule.md。
按这个修改应展示越权放行,故事却说无关 Gate 冻结 frontier,后文又解释是放行错误。标成“真实贡献链条”却没有可核验来源。 依据:docs/book/chapters/source-change-control-plane-rule.md。修改:统一反例的因果方向;若无来源,标注为教学情境并删除无依据的精确时间。
F22 [P3] 通过测试的证据不能一概归零 — docs/book/chapters/source-validation-to-pr.md。
“两者产生的证据都为零”把 passing 与 skipped 等同;正确边界是 passing 只证明其实际覆盖的性质,不能证明跳过的 backend。 依据:docs/book/en/chapters/source-validation-to-pr.md。修改:按测试覆盖范围表述证据,不贬低已验证的性质,也不将它扩大到未验证环境。
对主干的风险
运行时代码没有变化,主要风险是用户依照错误保证操作:未知副作用被盲目重试、已写回结果被误当丢弃、活跃 native Goal 被误当需要逐轮人工触发。default-off、authority、typed-state、domain-neutrality、default-behavior disclosure 和 guidance/obligation 均已核对:本 PR 没有修改相应运行合同;文案中的 owner、默认 fencing 和新增 provider-count 限制,已作为具体问题列出,而没有把写作建议冒充机器约束。
已执行的验证:
git diff --check:通过。- 根站、中文书、英文书各自
mkdocs build --strict:全部通过;随后拼装为与 Pages 相同的目录关系。 dev-book-publication-smoke.py --site-dir ...、dev-book-welcome-wagon-smoke.py --site-dir ...、docs-governance-smoke.py:通过。首次独立 book 输出缺少父站课程路由,拼装父站后通过;未据此误报代码缺陷。- settlement fault matrix + settlement parity:22 passed;executor 的 spend 拒绝/权限错误/写回恢复用例:3 passed;TypeScript settlement:19 passed。以上为合成 fixtures/受控 provider 的仓库测试,没有调用真实用户 Goal 或付费模型,不能据此声称所有生产故障组合已被穷尽。
- 纯 parser 复现错误 upgrade 命令 exit 2;正确 manifest 参数可解析。legacy checkbox 合成读回确认 open→done。实际 HTML 确认六个中文 fragment 无目标。
- 同 head 的 Frontstage Pages CI 已成功,包括三个站点构建、渲染检查、Mermaid 浏览器验证和 Welcome Wagon。本地未单独运行浏览器测试。
最后 CI 观察:{'SUCCESS': 22, 'SKIPPED': 3, 'IN_PROGRESS': 3}。仍未完成的检查:test-shard (1), test-shard (2), test-shard (4);不将 pending 当作 passed。
审查工具兼容性:本机 review 适配器要求 policy revision 7,当前源码 CLI 返回 revision 12。本次保留完整 revision 12 packet、按当前 required evidence 执行并校验 REQUEST_CHANGES;该适配器版本不匹配保留为审批证据缺口,没有降级规则或据此出 APPROVE。这是审核环境问题,不要求书稿修改运行时。
我的整体评价
保留现有问题主线,先做一轮事实修订,再做结构收敛,收益高于继续扩写。长程理解和用户操作仍有上述 regression,不能用全绿构建抵消;修复后无需另起一本书或增加教学 runtime。
建议采用更合适的写法:
- 每章围绕一个主问题展开:触发场景 → 可选方案 → LoopX 当前选择 → 保证条件与代价 → 用户怎样判断下一步。使用同一贯穿案例让各章重新汇合,减少反复重复“四个要求”和模板式开场。
- 将事实、设计理由和建议分开:当前实现用契约/读回佐证;RFC 解释动机但标注交付状态;经验判断用条件句。不要把“状态必须可辨识”“没有第二个 provider”“性能与可靠性必有一方变差”统称为已实现不变式。
- 代价必须能落到实际机制:03 的固定结算阶段不禁止 Host 在其内部做多步工作,04b 的 backoff 确实可能延迟发现变化。无需为每章凑出对称的代价数量。
- 证据放在章末的小表:主张、适用模式/版本、契约或测试、未覆盖边界。源码继续作为论据,正文保持架构叙事。行号/路径存在、标题数量相同只能证明结构;强断言和中英义务句需要逐句核对。
- 集中整理现有结构:贡献章节可保留为深入路线,主阅读路径先完成架构与使用判断;使用稳定 slug/显式 anchor 替代散落裸章号。这是当前文档边界内的整理,不需要额外框架。
已做 bounded future-facing pass:未发现应加入本 PR 的生产代码重构;有价值的伴随整理是消除重复论断、把过期模型移出当前路径、修正导航引用,保留公开 URL。旧稿内容与新的正确实现冲突时,应重写或退役,不能以“全部保留”为优先目标。
Standards/implementation:14 项,其中 3 个 P1,最严重的是错误恢复指导。Spec:8 项,主要是边界论证、中英义务及阅读路线不一致。本次仅审查并发布结论,没有修改 PR 文件或合并。
English verdict: REQUEST_CHANGES - c1923db - The problem-led structure is useful, but recovery, spend-failure preservation and native Goal continuation are materially misdescribed; additional state, budget, extension and navigation claims need correction. Local three-site strict builds, rendered smokes and 44 focused tests passed; Pages CI including Mermaid browser checks passed. These checks do not establish semantic correctness. Preserve pending CI and the adapter policy-revision mismatch as evidence gaps.
…acts Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
本次修订已处理前次 review 的问题,更新到
复审还修正了:monitor 可采用 两次独立检查后,已发现的修订项均完成读回。此次跟进相对旧 head 为 +1,020/−2,171 行,按文档与验证分为两个带 DCO 的提交。 验证通过:三个严格站点构建、渲染 publication/Welcome Wagon、docs governance、公开边界检查、50 个章节页内 fragment;32 个 Python 与29个 TypeScript 用例;隔离 Extension 的2个文档示例测试及完整 CLI 生命周期(含禁用拒绝和失败升级)。新增 fragment 检查在旧渲染产物上复现失败,在当前产物上通过。 这是一条作者修订与验证回执,不是合并批准;新 head 的 CI 与最终评审仍以对应检查和记录为准。未运行本地浏览器 smoke,未合并。 English summary: corrected the implementation review findings and follow-up discrepancies while preserving the problem-led architecture narrative. Three-site strict builds, rendered checks, 61 focused contract tests and two extension example tests passed; the disposable real CLI lifecycle also passed. This revision response does not approve or merge the PR. |
Signed-off-by: song <liusongstep@gmail.com>
…t tests Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
当前 head: 已在现有 PR 补充这一轮改进,代码、书稿与验证分成三个带 DCO 的提交:
独立 provider 按现有 placement 规则放入 验证:14 项示例测试、schema 检查、三个严格站点构建、渲染 publication/Welcome Wagon、docs governance、Ruff、公开边界检查通过。超长整数回归在原示例上失败,在修复版上通过。Provider wheel 实际安装,隔离 CLI lifecycle(含禁用拒绝、升级失败、恢复 package 与 rollback)通过。 浏览器在 9 页验证了 27 张图,使用从 CDN 下载的实际 Mermaid 脚本本地缓存;直接访问外网资源在本环境失败,未将其记成通过。额外尝试的根 LoopX wheel 构建因缺少生成的 Chat bundle 未通过,Provider 自身 wheel 已通过,二者分别记录。 Standards / Spec 复审提出的目录问题已修正并读回,当前审查范围内无未解决项。这是作者补充与验证回执,不是合并批准;新 head 的 CI 和最终评审仍按各自记录判断。 English summary: the supplement connects architectural tradeoffs through one running task and ships the complete optional teaching provider. Provider tests, schema checks, isolated real lifecycle, strict builds and cached-CDN browser rendering passed. Direct browser CDN loading and root-product wheel qualification remain explicitly unverified/failed as described above; no merge performed. |
The state-machine chapter's GoalControlSnapshot block is a conceptual view, but two of its rows also appeared in the "actual LoopX expression" table as if they were implemented field names. `completed_requirements` and `pending_requirements` have no definition in the codebase; the alignment output carries `frontier_basis`. Label both rows as conceptual placeholders and name the real field. The reading guide also only pointed at Codex App and Codex CLI while the runtime lists 21 supported agent types and 9 dedicated goal-mode adapter packages. State the teaching tradeoff so the two worked examples do not read as the product's capability boundary. Signed-off-by: song <liusongstep@gmail.com>
|
补充两条未通过项的根因,以及一组已修复的文稿问题。当前 head: 1. 根 LoopX wheel 失败与本改动无关
根包发现是 2. 本次补上的文稿修复(
|
Connect the existing Course entrypoint to four predict/run/explain exercises using current authority, lease, settlement and monitor regression tests. Add diagnostic routing and a synthetic non-coding transfer exercise in both languages without introducing another runtime, schema or chapter route. Guard the documented selectors and bilingual command parity with a small standard-library test. Reference checks do not claim behavioral validation. Signed-off-by: song <liusongstep@gmail.com>
Integrate the four-question standard into the book's reading journey, onboarding-to-return narrative, and counterexample exercises in both locales. Distinguish current release instructions from historical migration context, actual evidence from preview, and source state from displays. Remove the duplicate optional-provider installation detour from basic onboarding. Extend the reference checks and add an isolated Git regression for ignore/index semantics. No production runtime, package, dependency, chapter route, or navigation change. Signed-off-by: song <liusongstep@gmail.com>
Resolve Workspace causality and authority contradictions and separate observation, cadence, waking and replanning. Explain acceptance evidence, original-Turn closeout, execution proof, handoff and integration at their main teaching homes. Move field diagnosis to the appendix and connect the existing reader exercises without adding a second operating runbook. Keep later-main source comparisons distinct from the branch and release. Extend bilingual reference and link checks without changing production runtime, packages, dependencies, navigation or workflows. Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
已把 songoow/loopx#3 完整整合到本 PR,三个原始 DCO 提交均保留。fork PR 已标记合并;后续评审统一在这里进行。 当前 head: 验证与范围:
未运行完整产品套件、真实 Host 从接入到交付的完整旅程或根产品发布资格检查。这是整合与验证回执,不是合并批准;新 head 的 CI 和维护者评审仍需分别确认。未合并到 main。 English summary: absorbed all three signed commits from the fork supplement into this PR and retained one upstream review. Reader, provider, focused behavioral and TypeScript checks passed; one optional PostgreSQL case was skipped. Strict builds, rendered checks and the 9-page/27-diagram browser smoke passed. The final diff contains no production runtime changes. Full product/release qualification and a live Host delivery journey were not performed; no main merge or approval is implied. |
…' into codex/pr-repair-5354-final Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
修复本 head 的 coverage 汇总失败,更新到 CI 中行为分片均通过,但 aggregate 捕获了 development-probe 测试复制到临时仓库的 修复仅在该隔离测试的 CLI helper 中移除继承的 subprocess coverage instrumentation,保留 复现与验证:修复前可读到两个临时源码路径;修复后 development-probe 的 17 项测试通过,coverage 数据没有临时 fixture 源码,XML 汇总通过。CI workflow 合同和书稿阅读检查共 307 项(330 subtests)通过,Ruff/diff 检查通过。没有把旧 CI 结果算作新 head 的通过;新 CI 已启动,未合并。 English summary: isolated copied-fixture CLI subprocesses from inherited coverage startup without lowering checkout coverage gates. The temporary-source regression reproduces before the fix; focused tests and coverage XML aggregation pass afterward. This is a correction receipt, not merge approval. 补充验证修订:已同步后续 main,并纳入两个共享 browser fixture 修复。时钟比较使用暂停的同一时刻,保留严格相等断言;确认场景采用有效的合成 expiry 和完整投递回执,保留无授权时不提供执行控件的断言,并沿当前可见卡片入口读取。 |
…final Signed-off-by: song <liusongstep@gmail.com>
The App chapter and its invariant taught that frequent `should_run=false` means the cadence should slow. A settled turn's replay returns `false` while the typed owner preserves the current schedule and defers work to a fresh Turn identity, so the count of false results is not evidence about the frontier. - App chapter and invariant now separate this turn's receipt from the live frontier in both languages. - Reader checkpoint four names the settled-replay counterexample and runs the native `test_settled_replay_preserves_schedule_without_new_host_effects` selector in both locales. - Chapter 02's retained identity list no longer claims a lone registered Agent defaults to fresh registration; it matches the table and chapter 07: `thread_binding_selection_required` selects an existing lane, and only an explicit `--new-peer` registers a fresh identity. Signed-off-by: song <liusongstep@gmail.com>
…dex/pr-repair-5354-final Re-apply main's unification of the private goals directory to .loopx/goals/ onto this branch's rewritten chapters, and keep both sides where main and the branch each added prose to the same passage. Signed-off-by: song <liusongstep@gmail.com>
…dex/pr-repair-5354-final Signed-off-by: song <liusongstep@gmail.com>
|
两条意见已修复,当前 head F1 [P2,阻断] 不再从频繁
|
huangruiteng
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES — f18be122abd007cc1d2f7319901c171452371e71:上一轮的 scheduler 退避建议已经修正,但身份选择的双语说明仍遗漏首次接入分支。
动机
这不是单纯润色几段文字:整个 PR 重写 Developer Book 的架构、长程要求、接入和扩展练习,增加独立教学包及出版检查。有效结果应是读者能按当前合同完成第一次接入和下一轮继续,而不是仅让链接、章节标记和测试数量正确。此次按完整 65 文件范围重新分类,并重点复核上一轮发现、可执行练习及其真实控制面 owner;不把作者的完成声明视为验收。
改动思路
书籍解释应从当前类型化合同出发,不能自己产生第二套身份或调度规则。06-codex-app 现在正确区分“同一 Turn 已结算”与“下一次触发周期”:频繁得到 false 不授权调慢固定 cadence,先检查 Todo、claim 和状态。这修复了上轮 scheduler 意见。身份页则把上一轮过宽的 fresh 注册规则改得过窄:表格保留了“零个已注册 lane 或显式 new-peer”,紧接着的列表却只保留后一项,两段不能同时作为当前合同。
具体改动
[P2] 为未注册任何 lane 的首次接入保留 fresh 注册分支
中文第 99–102 行及英文第 107–110 行写成:不传 agent-id/new-peer 且 thread 未绑定就选择已有 lane,只有显式 new-peer 才默认 fresh。触发反例是已连接 Goal 的 registered_agents=[]。此时根本没有可供选择的 lane,真实 start-goal --guided 返回 fresh_agent_registration_required / register_fresh_agent,无需 new-peer。按照正文操作的新读者会寻找不存在的身份,或误以为正常首次注册违反合同。
最小修复:第 3 项增加“存在已注册 lane”的前提;第 4 项和相邻表格一致,覆盖“没有已注册 lane 或显式 new-peer”。同步中英文,并保留已注册 lane 存在时不得自动接管的边界。用零 lane、一个 lane、显式 new-peer 三个隔离 CLI 场景核验,不要仅检查关键词存在。
关键代码讲解
这段是本分支复用的既有 bootstrap_command_pack.py owner,并非 PR 新增的 Python 决策源。关键条件是:
has_registered_agents = bool(registered_agents)
# fresh 的默认条件包括零个已注册 lane,不只是显式 new-peer。
# 其余条件仍由原 identity-selection owner 判断。我没有修改真实 Goal/registry,而是分别构造三个隔离项目,从当前 checkout 的 python -m loopx.cli ... start-goal --guided 读取合同:零 lane → fresh、一个 lane → select、显式 new-peer → fresh。返回值与既有测试、页面表格一致,与这次修改的列表不一致。因此这是文档投影错误,不需要改产品身份权限来迎合文字。
教学包本身归属 packages/loopx-text-stats:87 行标准库 provider 只统计输入文本;通用 Extension lifecycle 负责安装和启停,不注册虚假的内置 Capability。decode_request → response → main 拒绝重复键、未知字段、无效 UTF-8 和超限输入,错误不回显原始文本;doctor 不读请求。permissions=[] 被明确标为声明而非 OS sandbox。这种专门计算留在 Python 合理,不要求无关 TS 重写。CI 新增教学包/reader 检查,PR 路径仍不部署生产 Pages。
对主干的风险
主要风险是公开解释让读者错误选择身份或干预调度,不是本 PR 改了核心运行规则。当前差异为 +6586/-4890,其中大量是双语替换、重组和删除旧解释;净行数不证明价值。没有新增 loopx/** 决策 owner,也没有自动加载的 Agent 指令、默认 provider 或权限变化。实际入口是 MkDocs 导航和读者主动运行教学包;Frontend/Lark 不需配套配置修改。
发布前 head 又合入 main,我停止了旧稿发布。当前 base 为 3a2a2919…,PR 仍是 65 文件 +6586/-4890;book、教学包、reader/publication/browser 检查和 bootstrap identity owner 相对 b1a8ae70… 均未改变,但没有仅凭这些 hashes 继承结论,而是在新 checkout 重新执行验证。本 head 本地通过:71 项聚焦测试及 347 项 subtests(包含 reader/provider/语义 fixture/settled replay)、publication source smoke,以及主站/中文书/英文书三次 strict build。另做上述三个真实 CLI 反例。最初构建输出目录放在 docs 内,后来组合站点缺少主站兄弟路由,属于我的验证布局错误,纠正布局后重新验证;不把这种调用错误报成 PR 回归。未查询或等待 GitHub CI。完整可视化浏览器/Mermaid 和 managed Extension 启停本轮未重新执行,旧 head 的证据不能直接认证当前整包;也不据此撤销仍需核验的历史意见。
我的整体评价
方向符合 Developer Book 的可操作教学目标,scheduler 修复有效,教学包边界也比孤立 scaffold 更完整。但首次身份分支是普通读者会实际遇到的错误,因此当前不能批准。未来向重构检查的具体边界是书籍 prose 与身份合同之间的重复知识:优先修正现有章节和三场景验证,不另造文档状态机或追加平行教程。修正这个有界问题后,再以新 head 复核整包及剩余历史意见。未合并、未部署。
English verdict: REQUEST_CHANGES — HEAD f18be12. Preserve the zero-registered-lanes fresh-registration branch in both identity explanations. The scheduler correction is verified; a passing publication check does not establish semantic correctness or full current-head visual/lifecycle qualification.
|
This pull request has merge conflicts with Choose the remote for the base repository, not an out-of-date fork. git fetch upstream
git rebase upstream/main
# Resolve each conflict, git add the resolved files, then git rebase --continue.
git push --force-with-lease origin HEADFor a same-repository clone whose Keep the DCO |
Both sides fixed the same coverage leak in the semantic development probe: the CLI under test runs in a disposable repository, so inherited pytest-cov subprocess instrumentation must not be forwarded. Keep main's broader COV_CORE_/COVERAGE_ prefix filter, which matches the convention already used by scripts/ci/test_impact_plan.py and the CI workflow tests, and drop the branch's narrower duplicate. The conflicting test file is now identical to main. Signed-off-by: song <song@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
动机
[P2] 首次接入的零-lane 例外仍被中英文编号列表遗漏。评审 exact head:c89b54dcc6b90e4c28cbc594575affc07c1d606c。docs/book/chapters/02-session-goal-loopx.md:99–102 写成任何未绑定线程都选择已有 lane、只有显式 --new-peer 才注册 fresh identity;英文同段 107–110 行相同。紧邻的 identity 表格却正确写了“没有已注册 lane,或显式 new-peer”。首次接入者会在空集合中尝试选不存在的身份,或误以为必须额外 new-peer,这与实际 guided CLI 不符。
这就是上一份 REQUEST_CHANGES 保留的问题,最新主线 merge 未修改这两个段落;不能因为 head OID 改变就撤销评审。书稿希望使读者理解长期控制面并执行练习,准确的初始分支比增加术语或测试数量更直接影响可操作性。
改动思路
保留本 PR 的问题导向叙事、双语阅读练习和独立教学 provider,不需要改变生产身份策略。最小修复:第 3 项加上“已有 registered lanes”前提;第 4 项明确“没有已注册 lane,或显式 --new-peer”。中英同时修订,继续保留 existing bound thread 与显式 takeover 的权限区别。
Provider 放在独立 packages/loopx-text-stats 归属合理:它是可选本地文本统计教学包,不是新增默认 Capability/catalog;permissions=[] 不意味着操作系统 sandbox 或新的身份权限。现有 TS/Host identity 合同继续是决策源,不应为了让文案成立改 CLI 或引入平行 Python 分类。
具体改动
关键代码讲解
- 当前
host_loop_activation.py::_identity_state的 662–698 行:缺失绑定仅在非 fresh-default 时选择已有 lane,fresh-default 分支返回注册请求;两者 activation 都仍被 gate 阻止,不隐式注册或接管。 - 当前
test_start_goal_compact_projection的 fresh/no-registered/new-peer 用例刻画这个分支。另用独立合成 registry/runtime、真实 source CLI 验证了三例,不复用实现输出生成 expected:零 lane → fresh_agent_registration_required/register_fresh_agent;一个已有 lane + 未绑定 → thread_binding_selection_required/select_agent_identity;显式 new-peer → fresh/register。三例均保留原 registry/state 字节。 decode_request/response(教学 Provider 37/62 行):限制 bytes/codepoints、拒绝重复/额外字段及错 schema、返回有界错误而不输出原输入;analyze_text是 codepoints/whitespace chunks 统计,不是模型语义判断。test_settled_replay_preserves_schedule_without_new_host_effects:当前中英 06 章正确区分 settled Turn 与 Goal frontier,should_run=false不再推导降低下一次周期;该旧 cadence 问题本次仍保持修复,不重开它。
当前 head 阅读 checkpoint、Provider 和 settled-replay 三组检查:40 passed, 347 subtests passed in 4.44s。另组接入检查为 38 passed、3 failed、53 deselected、347 subtests passed:三项 ambient-bound CLI 测试在本机默认 runtime 路由冲突处失败。分别在 immutable base 与 head 用相同三个 selector 重跑,失败身份和完整错误详情相同,生产路径 diff 为零,归为已有无关失败;不能把整组写成通过,也不将其伪装成该文案缺陷的新增根因。显式隔离 source CLI 的三例均通过。
对精确 PR base 的 loopx/**、apps/** diff 为零,diff whitespace 通过;最终变更是书稿、独立 Provider、维护测试和 frontstage workflow。未查询远端 CI、未重新做全部书稿事实审计、三站 strict build、browser Mermaid 或 installed provider lifecycle;旧作者/旧 head 记录不是本次通过证明。
对主干的风险
这个文案错误不会改变机器的权限 gate,但会把正常第一次接入引向不存在的 lane,增加无意义的人机往返;失败信息和表格再正确也不能抵消编号步骤的误导。所有章节大规模改写还需保持 current-checkout/released-version 区别、命令可执行性和双语一致性,不能因为 focused tests 绿就声称完整出版质量。
有界后续重构审视:教学 provider 的有界 decoder 已局部、独立并复用已有 extension 生命周期;本问题只需消除重复段落中的语义矛盾,不需要新增 bootstrap wrapper、迁移通用 runtime 或更多权限。最有价值的回归是分别钉住三个真实选择分支,并验证两种语言都保留零-lane 例外;不能只搜索列表是否包含 fresh/new-peer 单词。
我的整体评价
REQUEST_CHANGES:只保留上述可复现 P2 文字阻塞;此前 settled replay/cadence 修复有效。请同批修正中英列表并以三分支读回验证。后续 APPROVE 必须逐项核验有效历史 blocking reviews,再按当前 capability 做撤销 closeout;未解决的评审现在不撤销。本次没有更改 PR 内容、未合并,也不把源码检查冒充已安装产品验收。
English verdict: REQUEST_CHANGES - Both identity lists still omit the zero-registered-lane exception. The real isolated guided CLI requests fresh registration in that case, not selection from an empty lane list. Correct both languages without changing production policy.
Signed-off-by: song <liusongstep@gmail.com>
…lists The numbered Goal-start identity steps in both languages said an unbound thread always selects an existing lane and that only an explicit --new-peer defaults to fresh registration. The guided CLI also defaults to fresh registration when the Goal has no registered lane, as the adjacent identity table already states. Scope step 3 to Goals with registered lanes and make step 4 cover both the zero-lane and explicit new-peer branches. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
已修复 review 5390576627 的 P2 并解决与 main 的冲突,当前 head [P2] 零已注册 lane 的 fresh 注册分支中英
第 5 项(只有显式要求才接管某个精确 冲突已合入最新 验证(本 head,本地)
未运行:browser/Mermaid smoke、installed provider lifecycle,以及完整产品测试套件。新 head 的 CI 结果另行确认。 |
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewer: model_agent · GPT-5 · OpenAI
动机
[P2] monitor 练习不能按书中的完整命令执行。评审 exact head:a72fca548f309083c9610437779166ca86e7dae7。中文 docs/book/chapters/12-control-plane-course.md:177–178 和英文 docs/book/en/chapters/12-control-plane-course.md:249–250 都在第三个 selector 后漏了续行符,下一行不是第四个 pytest 参数,而是独立 shell 命令。原样执行只得到 3 passed,随后退出 1、报告 selector 不是可执行文件,关键 settled replay/cadence 练习根本没有运行。
原零 registered-lanes 阻塞已解决,不保留旧问题来凑拒绝。当前两种语言的第 3/4 项都明确已有 lane 的 selection 与首次接入/显式 new-peer 的 fresh registration;本轮独立真实 CLI 三分支全部符合预期,合成 registry/state 字节未变。但“准确且可执行的 reader exercise”仍是本 PR 的公开交付目标,不能以引用存在或两种语言一致替代实际复制运行。
完整 PR 为 64 个路径、+6581/-4890,base 83faf456fa7d5a91f30a2be3e51efa482302be48。它包括长期架构/状态/Host/预算/恢复的双语改写、book 导航、独立教学 provider、reader/出版/browser 检查和 workflow;对该 base 的 loopx/**、apps/** 没有生产差异。它不是一次生产控制面政策修复,也不能声称所有长期运行或真实 onboarding 已验收。
改动思路
问题导向书稿和一个完整可执行 provider 比堆叠 runbook 或未实现示意更有价值。Identity、canonical source、lease、settlement 和 monitor 应继续解释既有 owner,不建立文档版平行决策源。current-checkout 与 released behavior 的区别、未知 effect 的停止条件、original identity 的恢复必须保留。
最小修复是中英文都在第三个 selector 末尾补 \\,或明确拆成两条独立 uv run --extra test pytest 命令。回归应验证 fenced shell 命令边界/实跑,不只是 shlex 汇总出四个存在的 selector。无需修改生产 scheduler、pytest、Host identity,亦无需为这一个 bug 新建 capability/CLI。有界质量改进应落在现有 reader checkpoint guard。
具体改动
关键代码讲解
- 两个
02-session-goal-loopx列表与实际 guided identity owner 对齐。真实 source CLI 的独立合成输入:零 lane → fresh_agent_registration_required/register_fresh_agent;已有 lane 且 thread 未绑定 → thread_binding_selection_required/select_agent_identity;显式 new-peer → fresh/register。三例都不隐式接管,也不写回项目;这是真正修复,不是关键词检查。 checkpoint_sections / pytest_nodes / require_test_node(reader test)验证 checkpoint 顺序、selectors/AST 符号及双语一致性,但pytest_nodes把整个 block 用 shlex 展平。缺失续行符后第四个路径仍能被收集,所以现有绿灯无法证明每个 selector 属于同一个 shell 命令;这就是 tests pass 时仍可失败的具体反例。packages/loopx-text-stats/src/loopx_text_stats/cli.py::decode_request / response拥有 bounded 本地文本协议:bytes/codepoints 上限、重复键、错 schema/类型、固定错误码而非原输入回显。analyze_text统计 Unicode code points / whitespace chunks,不是假装模型语义。它是独立可选 package,复用既有 extension lifecycle,非默认 catalog 或新的 Goal authority。- 双语 lifecycle 区分 package installation 与 activation revision,preview/execute、doctor、disabled refusal、enable 再 probe、upgrade/rollback 与环境恢复;
permissions=[]不等于 OS sandbox。书中的 lease/settlement 练习使用合成 File/SQLite fixture,未将单条 receipt 或历史成功称为当前执行权。 - frontstage workflow 接入 provider/reader tests 与严格双语 build;出版检查验证 source/rendered anchors,browser smoke 核对 Mermaid 数量及脚本请求。它们各有范围,不能互相替代 shell 实跑或 live 用户验收。
本轮验证:reader/provider、fresh identity、in-flight replay、settled construction 组合 62 passed、347 subtests passed;独立 identity CLI 3/3。原样中英文 monitor fenced block 逐一经 bash -e 执行,两者均 3 passed 后退出 1;仅在内存里补续行符的相同命令,两者均 12 passed、exit 0。fixture 使用独立 registry/runtime,没有触碰真实 Goal 或改动候选源码。
主站 strict build 通过;中英文各自 strict build 通过。首次把父 docs 与子 book 目录并发构建造成 reviewer 自己的清理目录竞争,中文出现 FileNotFoundError;按 workflow 的顺序重跑中文成功,不能把这个 reviewer harness 错误归为 PR 缺陷。source publication smoke 和 diff whitespace 通过。未查询远端 CI,也未把旧 head 的 browser/lifecycle 记录升级成当前通过证明。
对主干的风险
当前阻塞属于文档交付,不修改机器排程或权限,但会让读者以为已验证 settled replay,而实际漏跑它。该练习恰好解释为什么 settled Turn 不授权降低下次 cadence;遗漏真实测试比简单排版更影响正确诊断。中英文内容相同、AST 引用存在、strict build 绿,都不能解除这个错误。
书稿较大,新的现场读取/预算/未知结果说明仍需保持 source-version 与承诺范围,不能把人数、测试数或 receipts 当真实长期用户结果。Optional provider 的来源与输入 schema 是局部 package contract,不应进入通用 quota/todo vocabulary。未完整重跑 current-head provider installed lifecycle、浏览器 whole-viewport 或全书事实审计,因此这些更广 APPROVE 要求仍是 unverified,不伪造通过。
Future-facing pass:现有 reader guard 应保留“引用存在”和“命令可执行性”的区别,聚合选择器能捕获重命名漂移,却不能证明 shell continuation。最有价值的有界重构是让现有 guard 验证命令分组,或针对真实 fenced command 执行合成测试;不要为文档引入新的执行权限层、扩大 TS 迁移或堆新的重复 smoke。
我的整体评价
REQUEST_CHANGES。上一轮零-lane 条件有效修复,当前只提出上述新、可复现 P2;请同步修订中英文命令边界,并增加能在漏续行符时失败的回归。完整书稿/providers/publication 的范围已重新检查,但本次不是全部出版/安装验收。保留历史讨论和仍有效阻塞,不因 OID 更新自动撤销;未改 PR、未合并、未调整 heartbeat 周期。
English verdict: REQUEST_CHANGES — head a72fca5. Both monitor exercises omit a shell continuation: only three tests run before exit 1. Adding one backslash runs all 12 parameterized cases successfully. The former zero-lane identity issue is fixed; 62 tests and 347 subtests passed.
The monitor checkpoint in both languages dropped the line continuation after its third selector, so a reader copying the fence ran three monitor tests and then the settled-replay selector as a standalone command, which exits 127 and never runs the cadence counterexample the section relies on. Restore the continuation, and make the reader checkpoint guard split each fence into logical shell commands. A selector that is not an argument of a pytest command is now rejected instead of being credited to the previous command, which is how the broken fence passed the existing check. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
…and scheduler owners Fresh registration in chapter 05 dropped --require-new and told readers a changed=false result could still be accepted. Without the flag an existing id registers idempotently, so that guidance allows a silent takeover. Use the same --require-new command the guided packet emits and continue only on the packet's continuation contract (ok, changed, written, global_sync.ok, registration_readback.verified); a collision or changed=false means the id already exists. Chapter 06 attributed cadence slowing to an owner judgment and argued for manual convergence from wait frequency. The scheduler_hint and stateful backoff propose cadence changes; hand-editing the RRULE shows up as drift_detected. The Gate reminder bound comes from human_gate backoff, with the notice cooldown only after a failed host cadence update, and the step numbers now match the steps the paragraph compares. Chapter 07 said a missing scheduler context neither wakes nor tells the reader; it returns repair_scheduler_execution_context and stops unchanged polling. The identity invariant now matches the code: an unbound thread gets a selection gate, a bound thread keeps its lane, activation without task text may auto-select a sole lane, so pass --agent-id explicitly. Chapter 04 required a baseline revision for an unchanged vision; the CLI takes only --vision-unchanged-reason and binds the existing vision. Chapter 05 also drops merge leftovers: orphaned success bullets, a dangling "Read:", and a duplicate "2." heading that captioned a read-only Git inspection as an install step. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
…aults in chapter 02 The chapter 02 opening story said each wake-up during a pending user Gate spent the quota window again. A user_gate_quiet_wait carries must_attempt_work=false and delivery_allowed=false, and wait or user-action results cannot spend quota, so the story now names the real cost: Host wake-ups, time and model calls. The identity list said an explicit --new-peer always defaults to fresh registration; a thread with a stored binding keeps its bound lane first, which the isolated guided CLI confirms. The same condition now reaches the two places that presented fresh registration as the normal new-session path. "Inference is explicitly banned" overstated the contract: a sole Goal is auto-selected and a bound thread supplies the agent, so the cost names what must be explicit (several Goals, agent identity, unknown host surface). The Host matrix now credits the adapter documentation alongside the Runtime Connector Catalog, since two rows are not in the catalog. Chapter 00 counts the five areas its table lists, and the Chinese edition states the same edition-authority note as the English one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
…tually prove Rollback probes the currently installed entrypoint against the previous revision's recorded manifest and location; it does not reinstall the old package. In an isolated walk, rolling back after a package upgrade reported version 0.1.0 and doctor_verified=true while the installed 0.2.0 code kept running. Chapter 10 now says a passing probe does not prove the matching package is installed and tells readers to install it and read behavior back with run. The troubleshooting row no longer ties rollback availability to the installed package; it depends only on a recorded previous revision. The chapter cited a renamed test; it now names test_failed_enable_remains_disabled_and_rejects_current_artifact and says a failed enable revokes the failed artifact's identity rather than clearing every proof. Chapter 08 attributed the duplicate-id check to the registry smoke, which covers catalog origin and extension composition; the duplicate case is test_duplicate_capability_fails_closed. It also credited the placement guide with endorsing a future finance-value-discovery Capability, while the guide keeps Finance as a standalone extension and the value connector protocol rules out inventing that capability. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
…and fix dangling references source-protocol-map listed event_sourced_state_contract_v0 among the entry points that answer where a fact lives and who may write it, and its symptom list pointed at "status and the event ledger disagree". That protocol page is retired and there is no current Todo event replay, so the bullet now says so and the symptom row is gone. source-validation-to-pr referred to a row of a table that no longer has it and named a `rejects-unclaimed-...` test that does not exist in the repository; both references are restated concretely. Its contribution baseline now uses the documented `uv sync --extra test` / `uv run --extra test` commands with the explicit-environment alternative, matching the testing guide and the current checkout. The rule-change chapter's "nine-stage decision pipeline" had no owner; it now says "ordered decision pipeline". The state-machine timeline claimed nothing binds a Vision checkpoint to an acceptance revision: on the File/SQLite authority path the checkpoint read context does carry the goal and acceptance basis and rejects a changed one as stale, so the timeline now states that and scopes the gap to modes that do not carry acceptance. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
…ng a stale smoke Cleaning up the chapter 05 merge leftovers also removed the observable acceptance list that belonged to the "connection accepted" row, leaving the chapter labelled "reviewable checklist" without the checklist. Restore it as prose under the table: doctor, the two state files, status, exact goal_id reuse, fresh agent_id, and the Git boundary. Chapter 08 cited examples/capability-extension-registry-smoke.py as executable evidence, but that script fails on current main: its hardcoded built-in capability list is missing performance-diagnosis, so it asserts on a stale catalog. No CI job runs it. Point at the registry test module and its duplicate-id case instead, and say plainly that the stale script is not a reliable evidence entrypoint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <liusongstep@gmail.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewer: model_agent · GPT-5 · OpenAI(self-reported;不是身份认证或独立模型溯源)
动机
没有发现本轮新增的阻塞问题,APPROVE 当前 171e8e53836adc4e1876aa4625efd38996a6f4ac。本次重新评估整个 base-to-head PR,而不是把前几轮的逐项修复当作自动批准:旧书把控制面字段、接入命令和实现细节混在一起,读者很难判断“当前事实、为什么允许/拒绝、证据及合法下一步”。当前版本用一个任务串起状态、执行、恢复、协作和观察,并把概念图与可执行练习分开。它交付的是可发布、可练习的双语教学切片,不是证明整个长程路线图、已安装 App 或远端服务已经完成。
Extension 部分按变更前已接受的 docs/reference/extensions.md,固定 revision 4fc30f185e6fbe688b1f0a16b0f7bdafe2666ba9 审查;三项 material criteria 使用原文标题,不另造规格:Runtime Lifecycle——默认 preview、明确执行和真实生命周期读回,已实现并验证;Placement Decision For Agents——standalone provider 复用现有 runtime/lifecycle,不新建虚假的 built-in capability,已实现;Scope Boundaries——激活不安装包、不创造 credentials、不扩大 permission,已实现。教学结构本身没有独立的强制自动运行规格,因此不将新书中出现的示例字段升级成机器义务。
改动思路
主线先解释一次有边界的工作,再回查 source/receipt、claim/lease、Gate scope 和合法恢复责任。需要外部条件时,观察与真正执行分开;一次 settled replay 不替下一轮决定 cadence,接入 Goal 也不等于接管 Agent。源码进阶对照明确标注固定 commit,不把主线功能倒推成发布版 v1.2.3 的保证。当前 book 还明确退役的 Todo event 路径与仍有效的 run history/rollout 不是同一个源。
Extension 教学不是另写控制面:packages/loopx-text-stats 是独立安装的零权限、runtime-only 示例,复用现有 loopx extension 命令。小型 Python provider 只做 Unicode 文本统计、输入限制及结构化错误;共享注册、准入和激活状态仍归现有 owner,没有平行 Python 决策源。相比只放几段无法跑完的伪代码,这个例子让读者实测 preview、拒绝、disable/enable 和 rollback;相比扩展 built-in catalog 或增加 UI 开关,又保持了教学边界和卸载成本。
具体改动
整个 diff 为 65 个文件,+6709/-4966;主要是双语书稿重写、三章补充和导航重排,另有 provider、读者检查及发布 guard。相对上一评审版本,本 head 的零-lane 新执行者分支、settled cadence、monitor 命令续行等修复一并重新检查。评审过程中从 8643a7c0… 更新到本 head,新增四处书稿变更已逐项读完;测试、构建、真实 provider 旅程及最终页面在新 head 重跑,没有把旧 receipt 改一个 hash 冒充新执行。
| 符号/入口 | 行为及验证判断 |
|---|---|
analyze_text,packages/loopx-text-stats/src/loopx_text_stats/cli.py:28 |
code point / whitespace chunk / splitlines 的统计定义明确,不声称语言学分词;固定输入期望 51 字符、44 非空白、8 words、2 lines,真实 managed run 读回相符。 |
decode_request,同文件 :37;main,:71 |
64KiB 输入和 32000 code points 有界,duplicate/unknown fields、schema、类型和空白输入被拒绝;错误不回显原 payload,doctor 不执行统计。 |
pytest_nodes,tests/test_dev_book_reader_checkpoints.py:49 |
按 shell 命令边界识别 test selector,而不是只搜路径;缺失续行不再被错误计作已执行测试。 |
require_test_node,同文件 :69 |
AST 核对真实顶层测试名,不 import 或执行任意被引用文件;缺失、越界和错误 selector 有负例。 |
assert_local_fragments_resolve,examples/dev-book-publication-smoke.py:185 |
从最终 HTML 核对 fragment 与 id,补上构建成功但章节链接失效的发布缺口,保留既有 bilingual/nav/asset guard。 |
本轮亲自完成的决定性证据:
- 新 head 的读者、provider、settled replay、接入投影及 catalog matrix:139 passed,347 subtests passed。它们是各自 fixture 的验证,不是运行中的用户 Goal 已接管。
- 逐字执行中英文 monitor fenced bash:各 12 passed / exit 0;独立删除第三 selector 的续行符后,各只执行 3 passed / exit 1,第四项成为不存在的 shell command。原缺陷由真实 shell 反例检出,不靠 guard 自证。
- 真实 source CLI + 已安装示例 provider,隔离 activation state:14 步 preview/no-write、execute、disable 后拒绝、enable 恢复、非法输入、manifest activation upgrade、rollback、再次运行和独立 list 读回均通过。主干与 PR 采用同一输入、同一 provider 源码和同一 harness,对这 14 步完整响应比较一致;只映射环境绑定的 executable identity,并在每一侧验证实际 identity 与 doctor proof 一致,未删除错误、remediation、receipt 或状态差异。输入 SHA256
6f776e1ffcdaafaebb8389823d2ee5698861634cd5045a5b0f2b703aa61db85a;归一化观察 SHA2562843e0fa80ce5a9522075f19b22d679b8f32c67429f2e88fbf122345b1416e72。这里的 upgrade/rollback 是 activation snapshot,不是假装发布/替换了 provider package。 - 中英文书及主文档 strict build 通过,assembled publication smoke 通过;真实 Chromium 查看最终本地 assembled artifact 的 9 个页面 / 27 个已渲染且有可见尺寸的 Mermaid 图,中英文内容、anchor 点击与 reload 可用。390px 英文 lifecycle 页无页面横向溢出;已查看桌面整屏、实际 diagram 和移动整屏。这是当前源码构建,不是线上站点或已安装 Workspace readback。
接入章现在补回可观察的验收清单:doctor、精确 Goal/source、frontier、fresh Agent/明确 takeover、Git ignore;placement 章改用当前 catalog 测试,不再把过期的硬编码 smoke 宣称为绿证据。Git onboarding fixture 还实际区分 ignore match 与仍在 index 的文件,避免把“被忽略”误说成“已从历史/索引删除”。
对主干的风险
检查整个 diff,loopx/**、apps/**、自动装载 skill 与 core catalog 没有变更;默认 runtime、quota、scheduler 和已有 caller schema 不被这份教学例子改变。包存在、可发现、doctor ready 都不等于激活;preview 不创建 activation state,disabled run 拒绝,恢复由明确 enable/doctor owner 完成。文档是用户主动打开的发布入口,不是自动注入所有 Goal 的 instruction;zero-permission generic run 也不授予 peer、生产写入或跨 Goal 权限。
书稿大幅改写会有翻译/事实漂移成本,因此判断不是“总字数够少”或“链接都存在”就成立:危险的恢复、身份、停等与 package/activation 边界已经按当前 source、真实命令和负例核对,guard 用于防止之后失联。最强未验证项是读者理解效果、任意 OS/Host 集成及真实长时服务故障;此 PR 不承诺这些资格化,也没有 PostgreSQL、跨设备提升、live 模型或外部写入。浏览器 resource 观察与 rendered/语义检查不能替代穷尽 runtime-error 监听。没有查询或等待远端 CI。
本轮几次错误的测试路径、相对 MkDocs 输出目录及隐藏 TOC locator 属于评审 harness 错误,修正后重跑成功,未作为 PR 回归或删除红记录。未来面向改造检查已应用于现有 owner:shell grouping guard 和 bounded decoder 是局部有价值的收敛;没有必要为纯教学 provider 建新 TS orchestrator 或通用状态框架。若将来成为真实共享能力,再按现有 capability/provider contract 增加产品入口,不借此 PR 虚构已交付 UI。
我的整体评价
APPROVE。本 head 的书稿、可执行示例和发布入口形成了完整而可回滚的教学切片;此前具体缺陷已修复并由独立真实路径/负例验证,没有新阻塞 finding。批准只覆盖上述 65 文件和当前 head,不自动接受父路线图或授权合并。批准读回后另执行 capability 的有效旧阻塞 closeout;仅当每一项旧 finding 已独立解决且权限成立才撤销,保留历史讨论、未证实的 dissent 和单独的 merge hold。
English verdict: APPROVE - 171e8e5 delivers a coherent bilingual teaching slice with a bounded standalone provider. Native tests, verbatim shell mutation, real lifecycle/base-head comparison and the built publication journey pass; no new blocker found. This is not qualification of installed hosts, remote services or the entire roadmap, and it grants no merge authority.
Problem and result
The Developer Book should explain why LoopX uses its architecture for long-running work, how readers operate it, and where its guarantees stop. This PR combines the bilingual architecture rewrite, optional teaching provider, and the reader exercises from songoow/loopx#3 into one upstream review.
The supplement's three original DCO-signed commits are preserved. Latest main is incorporated; existing chapter routes and navigation remain intact. The first-reading route follows a complete Turn, then state and authority, recovery, and observation. One synthetic JSON-output task connects artifact revisions, evidence, decisions, handoff, settlement, and result return.
Changes
packages/loopx-text-stats/with manifest, packaging, schemas, examples, bilingual READMEs, and 14 tests. It is explicitly installed, has no default Capability registration, and is not bundled into the LoopX wheel. Long-integer JSON errors return the documented response.BARE_SHA256_PATTERNimport that blocked the documented exercises. After syncing that fix, the final diff against main contains no production runtime change.Current-source wait recovery and independent-worktree examples remain distinct from the released
v1.2.3contract. Maintenance checks prove reference integrity, not bilingual semantic equivalence or complete product qualification.Validation
The runtime-startup failure was reproduced before removing the duplicate import; the previously blocked behavioral tests passed afterward. The identical repair has now landed in main as #5391. Syncing it preserves the validated file tree and removes runtime code from this PR's final diff. Full product suites and root release/wheel qualification were not rerun. New-head CI remains separate from these local results.
The earlier implementation review drove the previous corrections. This consolidated revision also underwent separate Standards and Spec checks; confirmed integration findings were corrected and read back. Because this PR includes an independently installable package and CI changes, it remains for maintainer review and merge.
中文说明
把原架构书修订、完整教学 Provider 和 fork PR #3 的三个补充提交统一到本 PR,不再维护两个待合并的书籍 PR。
正文沿“问题 → 设计选择 → 成立条件与代价 → 使用与恢复”展开。补充稿的十组中英文章节、四组可运行阅读练习和诊断分流全部纳入;重复操作说明集中到附录,Course 保留实现深度。当前源码、发布版承诺、历史事实与新执行权限分别说明。
同步最新 main 后发现一行重复导入阻止 Effect runtime 启动。对应修复已通过上游 #5391 合并,现已同步;本 PR 相对 main 的最终差异不含生产 runtime 改动。完整 checkout 的 17 项阅读检查、14 项 Provider 测试、20 项实际行为用例,以及 20 项 TypeScript 用例通过;1 项可选 PostgreSQL 集成测试未运行。三个站点构建、渲染检查和 9 页 27 张图的浏览器检查通过。
本轮结果不代替完整产品或发布资格验证;新 head 的 CI 和维护者评审仍需分别确认。未合并到 main。