From f6c9b240262105b38917d22f23404695efae77f0 Mon Sep 17 00:00:00 2001 From: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Date: Wed, 30 Sep 2026 14:31:24 +0800 Subject: [PATCH] docs(blog): map long-horizon work by goal and process clarity Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> --- .../blog/application-scenarios/index.html | 154 +++++++++++------- .../public/blog/assets/task-quadrants-en.svg | 55 +++++++ .../public/blog/assets/task-quadrants-zh.svg | 53 ++++++ .../blog/zh/application-scenarios/index.html | 154 +++++++++++------- 4 files changed, 294 insertions(+), 122 deletions(-) create mode 100644 apps/presentation/site/public/blog/assets/task-quadrants-en.svg create mode 100644 apps/presentation/site/public/blog/assets/task-quadrants-zh.svg diff --git a/apps/presentation/site/public/blog/application-scenarios/index.html b/apps/presentation/site/public/blog/application-scenarios/index.html index 2806bdf4c8..1ea94a55c2 100644 --- a/apps/presentation/site/public/blog/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/application-scenarios/index.html @@ -25,10 +25,8 @@ button:focus-visible{outline:2px solid var(--blue);outline-offset:3px}button[aria-pressed=true]{background:var(--ink);color:#fff} .site-header{padding:0 24px}.article-heading{padding-top:64px}.article-heading .eyebrow{margin-top:0} .chapter-line{display:flex;flex-wrap:wrap;gap:16px;align-items:center;color:var(--body);font-size:12px;margin-top:28px} - .scenario-grid{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:16px;margin:32px 0} - .scenario{padding:24px;border:1px solid var(--border);border-radius:12px;background:var(--surface)} - .scenario .number{font:500 12px var(--font-mono);color:var(--blue)}.scenario h3{margin:16px 0 12px;font-size:20px}.scenario p{font-size:14px;margin:0}.scenario small{display:block;border-top:1px solid var(--border);padding-top:16px;margin-top:20px;font-size:12px} .takeaway{font-size:22px!important;line-height:1.65;color:var(--ink);letter-spacing:-.02em} + .quadrant-figure .architecture-viewport a{min-width:640px}.quadrant-figure figcaption{font-size:13px} .loop-row{display:grid;grid-template-columns:repeat(4,minmax(0,1fr));gap:10px;margin:24px 0}.loop-row>div{padding:16px;border:1px solid var(--border);background:var(--surface);border-radius:6px;font-size:14px}.loop-row b{display:block;color:var(--ink)}.loop-row span{display:block;font-size:12px;margin-top:8px} .comparison{display:grid;grid-template-columns:1fr 1fr;gap:16px;margin:24px 0}.comparison>div{border:1px solid var(--border);border-radius:12px;padding:24px;background:var(--surface)}.comparison strong{display:block;font-size:32px;letter-spacing:-.04em;margin:12px 0}.comparison p{font-size:14px;margin:0} .study-signal{margin:32px 0;padding:28px 0 4px;border-top:1px solid var(--border);scroll-margin-top:24px}.study-signal h3{margin:0 0 12px}.study-signal .study-setting{font-size:12px;margin-bottom:16px}.study-signal p{margin-bottom:16px}.study-signal .signal-boundary{font-size:14px}.study-signal .study-link{font-size:14px;margin-bottom:0}.presenting .study-signal .study-setting{font-size:14px}.presenting .study-signal .signal-boundary,.presenting .study-signal .study-link{font-size:17px} @@ -36,10 +34,10 @@ .small-note{font-size:12px!important;color:var(--body)}.layers{margin:24px 0;border-top:1px solid var(--border)}.layer{display:grid;grid-template-columns:150px 1fr;gap:24px;border-bottom:1px solid var(--border);padding:18px 0;font-size:14px}.layer b{color:var(--ink)} .evolution{counter-reset:step;list-style:none!important;padding:0!important}.evolution li{counter-increment:step;display:grid;grid-template-columns:40px 1fr;gap:16px;padding:20px 0;border-bottom:1px solid var(--border);margin:0}.evolution li::before{content:"0" counter(step);font:500 14px var(--font-mono);color:var(--blue);padding-top:4px}.evolution b{display:block}.evolution span{display:block;font-size:14px;margin-top:6px} .source-list{font-size:13px}.site-footer{padding:32px 24px;border-top:1px solid var(--border);font-size:12px}.site-footer a{margin-left:auto} - .presenting .article-layout{max-width:1200px;grid-template-columns:180px minmax(0,960px);gap:40px}.presenting .prose{font-size:21px}.presenting .prose h2{font-size:36px}.presenting .takeaway{font-size:28px!important}.presenting .prose section{padding-bottom:40px;scroll-margin-top:24px}.presenting table{font-size:17px}.presenting .scenario p,.presenting .loop-row>div,.presenting .layer,.presenting .state-demo dd{font-size:17px}.presenting .toc{font-size:13px}.presenting .toc details{top:20px} + .presenting .article-layout{max-width:1200px;grid-template-columns:180px minmax(0,960px);gap:40px}.presenting .prose{font-size:21px}.presenting .prose h2{font-size:36px}.presenting .takeaway{font-size:28px!important}.presenting .prose section{padding-bottom:40px;scroll-margin-top:24px}.presenting table{font-size:17px}.presenting .loop-row>div,.presenting .layer,.presenting .state-demo dd{font-size:17px}.presenting .toc{font-size:13px}.presenting .toc details{top:20px} @media(max-width:1100px){.article-heading,.article-layout{margin-left:24px;margin-right:24px}.presenting .article-layout{grid-template-columns:160px minmax(0,1fr)}} - @media(max-width:760px){.scenario-grid,.comparison{grid-template-columns:1fr}.loop-row{grid-template-columns:1fr 1fr}.article-layout,.presenting .article-layout{display:block}.toc{margin-bottom:40px}.toc details{position:static}.site-header{gap:12px;flex-wrap:wrap;padding:12px 20px}.site-header nav{gap:16px}.header-actions{gap:8px}.article-heading{padding-top:40px}.state-demo dl,.layer{grid-template-columns:1fr;gap:8px}.state-demo dd{margin-bottom:16px}.scenario small{margin-top:12px;padding-top:12px}.presenting .prose{font-size:18px}} - @media print{.site-header,.toc,.stage-buttons,.chapter-line,.site-footer{display:none!important}.article-layout,.presenting .article-layout{display:block;max-width:none;padding:24px 0}.article-heading{padding:0 0 24px}.prose h2{break-after:avoid}.diagram,.comparison,.state-demo,.scenario{break-inside:avoid}body{background:#fff}} + @media(max-width:760px){.comparison{grid-template-columns:1fr}.loop-row{grid-template-columns:1fr 1fr}.article-layout,.presenting .article-layout{display:block}.toc{margin-bottom:40px}.toc details{position:static}.site-header{gap:12px;flex-wrap:wrap;padding:12px 20px}.site-header nav{gap:16px}.header-actions{gap:8px}.article-heading{padding-top:40px}.state-demo dl,.layer{grid-template-columns:1fr;gap:8px}.state-demo dd{margin-bottom:16px}.presenting .prose{font-size:18px}} + @media print{.site-header,.toc,.stage-buttons,.chapter-line,.site-footer{display:none!important}.article-layout,.presenting .article-layout{display:block;max-width:none;padding:24px 0}.article-heading{padding:0 0 24px}.prose h2{break-after:avoid}.diagram,.comparison,.state-demo{break-inside:avoid}body{background:#fff}} @@ -58,18 +56,24 @@

Where does LoopX fit?
Complex tasks, open-ended exploration, and continuo
-

Start with the work, then choose the abstraction

-
-
01 / COMPLETION

Complex tasks with clear acceptance

Refactors, migrations, and protocol implementations. The destination is fairly clear even when the execution path is not.

Preserve: acceptance gaps, versions, repair evidence
Measure: completion rate and cost
-
02 / DISCOVERY

Open-ended exploration

Research, algorithm experiments, and system optimization. New evidence continuously changes the next route.

Preserve: hypotheses, counterevidence, candidate directions
Measure: useful discoveries and validation
-
03 / DELIVERY

Digital workers for continuous delivery

Take an issue, fix it, follow review, handle new feedback, and remain accountable for the result.

Preserve: responsibility, external state, waiting conditions
Measure: delivery quality and human effort
-
-

These categories can nest. An agent maintaining a repository over time may first explore a performance problem, then complete a fix with clear acceptance, and finally follow the pull request through review. The first two categories mainly describe how work is solved; the third adds an ongoing responsibility and a stream of newly arriving work.

-

What should persist is the work—and the reasoning behind its decisions.

+

Ask two questions: what counts as done, and how do we get there?

+

For work handed to an agent, I first separate two uncertainties: clarity of the goal and clarity of the process. Knowing what you want and knowing how to achieve it are different things.

+

A clear goal names the outcome, constraints, and acceptance criteria. A clear process has a reusable method. A workflow with many steps can be well understood; “make the system better” may leave both the outcome and the route unresolved.

+
+
+ Goal clarity increases to the right; process clarity increases upward. Top left: clarify the goal. Top right: routine execution. Bottom left: 02 open exploration. Bottom right: 01 complex tasks. 03 continuous-delivery digital workers operate across these quadrants. +
+
Clarity changes with evidence and human feedback. 01 and 02 describe how work is solved; 03 describes how responsibility continues. Scroll horizontally on narrow screens, or open the image at full size.
+
+

Top right: clear goal, known process. Convert data to a specified format or run a validated release workflow. Reliable execution matters here. A script or workflow can be enough for short work; waiting across days, recovering from interruption, or owning downstream results creates an additional continuity need.

+

Top left: known process, unclear goal. A team may know how to produce a weekly report without having agreed whose decisions it should inform. Clarify the audience, value criteria, and tradeoffs first. Speeding up a familiar process does not automatically make its output useful.

+

Bottom right maps to 01; bottom left maps to 02. The former knows the destination and must find a route. The latter also uses exploration to discover what is worth pursuing. An exploratory round can still define its question, evidence requirements, and stopping conditions; an unclear goal is not a mandate to run indefinitely.

+

03 is a continuing responsibility across quadrants. An agent maintaining a repository may clarify performance needs, investigate a bottleneck, complete a fix, and follow its pull request. It takes on both complex tasks and exploration, while remaining responsible when new feedback arrives.

+

Uncertainty determines what must persist.
Continuing responsibility determines for how long.

@@ -79,11 +83,85 @@

01 / Finish a complex task

Define acceptanceSpecification, baseline, budget
Implement and validateBind artifacts to a specific version
Find the gapsWhat remains beyond the tests?
Continue or stopClose gaps, or close out explicitly

LoopX keeps the goal and its acceptance gaps continuous across turns, then carries new evidence into the next decision. The model implements and judges; domain validators check the result; the control plane records which results were accepted and which work remains.

+

Preserve more than completed steps: the current acceptance version, unresolved gaps, attempted repairs, and the artifact version each check evaluated. If the task list is exhausted but acceptance is unmet, replan. If a person changes compatibility requirements, a pass on the old version cannot establish acceptance under the new criteria.

The best-fitting tasks usually span repeated validation, waiting, or handoffs. Existing agents may already be enough for small work that can be accepted in one session; an added control plane must justify its overhead.

+
+

02 / Find a path through open-ended exploration

+

“Find a better algorithm” does not come with a fully specified task chain. Real progress may validate a hypothesis—or eliminate a promising-looking route. If only successful conclusions survive, the next run can easily repeat ideas that were already disproved.

+
+
Frame the questionConstraints and hypotheses to test
Try several routesIsolate experiments and bound cost
Record positive and negative evidenceSupport, refute, or raise a new question
Update the directionContinue, combine, retire, or pause
+
+

Explore provides an optional evidence graph and bounded branch planning: nodes represent questions and findings, while relationships express supports, refutes, and leads_to. Planning suggestions still pass through normal execution boundaries; the graph neither launches workers nor grants spending authority.

+

Auto Research organizes this pattern as research work: select a topic, propose hypotheses, run experiments, evaluate independently, and produce a report. Existing protocols and commands provide a foundation, while the public showcase still contains blueprints and items awaiting validation. Real effectiveness must be demonstrated on specific tasks.

+

Exploration can be accepted by the uncertainty it removes: compare candidate routes, produce a reproducible experiment, rule out a hypothesis, or identify the evidence still missing. A well-supported negative conclusion can complete a bounded research round. A finite research deliverable can close out while a longer-term direction continues in another round.

+ +

Place RSI here, but keep the claim precise

+

When the object of improvement is the agent’s own tools, strategies, or harness, the work becomes self-improvement research. Changing its own code only creates a candidate; sustained improvement still requires independent evaluation, cross-task generalization, regression checks, and rollback.

+

Applying past experience to a later decision, automatically generating experiments, and modifying the harness are different levels. We can discuss an experimental path toward recursive self-improvement (RSI) here; we cannot equate “automatic iteration” with open-ended self-improvement already achieved.

+
+ +
+

03 / Own continuous delivery

+

For a pull-request or issue-fix digital worker, the job begins by deciding whether the issue is worth fixing. After delivery, CI, review feedback, branch changes, and the merge result still remain. New tasks keep arriving, and external facts keep changing.

+

Responsibility does not end when the patch is generated.

+

This synthetic scenario starts after a fix has been submitted while CI and review are still running. Choose an external state to see how the next step changes.

+ +
+
External fact
GitHub reports that checks are pending for the current revision.
+
Domain state
Record the pull request, revision, and check state; an unchanged observation does not create progress.
+
Capability proposal
Preserve the monitor and recovery condition, then wait for the result.
+
Control-plane check
Admit work by authority, budget, and eligibility; work outside the wait scope can still proceed.
+
Meaning for people
Stay quiet when nothing material changed, so polling does not become nagging.
+
+ +

This interactive diagram is a design explanation. It is not connected to a real repository and performs no GitHub operations.

+

Here, a “digital worker” means continuing responsibility, boundaries, memory, and a feedback loop. Its value should be judged by accepted fixes, reopenings and regressions, handling latency, delivery cost, and how often a person must repeatedly supervise it. Pull-request count and uptime are only process signals.

+

A continuing responsibility also needs rules for when to wake, when to stay quiet, and when to involve a person. New review feedback can become work input; unchanged pending CI does not need repeated interruptions. Missing authority, goal tradeoffs, or exhausted budgets need explicit stopping points. One repair can finish while the maintenance role continues; that role must also be pausable, revocable, and transferable.

+
+ +
+

Work moves between quadrants. State must follow.

+

Consider maintaining a service’s performance. “Make it faster” first requires agreement on workload, metrics, and compatibility constraints. Once the outcome is clear, the bottleneck still needs investigation. Finding a repair route turns the work into implementation and acceptance; submission brings review, feedback, and regression follow-up.

+
+
Clarify expectationsWorkload, metrics, constraints
Explore bottlenecksBaselines, hypotheses, counterevidence
Deliver the fixImplement and verify a specific version
Keep following upReview, regressions, new feedback
+
+

This is not a one-way progression. Feedback can change the goal; a counterexample can invalidate the route. A stable method discovered through exploration can become a routine workflow, while new conditions can send a workflow back into exploration. An agent must recognize whether the remaining issue is disagreement about the goal, uncertainty about the route, or simply waiting.

+

A task’s name does not determine its position; the information currently available does. Research with an agreed question, deliverable, and scoring rule can be managed as a task with clear acceptance. A refactor with a mature method can rely more on an established workflow. Both axes are continuous; the quadrants mark common work patterns.

+

Useful handoff state includes the current goal and constraints, why this route was chosen, rejected approaches, artifact and acceptance versions, and conditions for continuing or stopping. A chat summary or “step three completed” alone cannot carry these changes.

+

The same work may change models, sessions, or executors. Revisions to its goal, evidence of failure, and outstanding responsibilities should remain continuous. This is the shared place of a long-horizon control plane in all three scenarios.

+
+ +
+

The richer the domain, the clearer the division of responsibility must be

+
+
Kernel
Authority and common lifecycle
Which goals and work items are valid, who may act, which authority and budgets apply, and how work is handed off, paused, and committed.
+
Domain State
Domain continuity
Whether this issue is actionable, which revision the pull request names, where CI and review stand, and which conclusions remain valid.
+
Capability
Outcome contract and judgment
Translate domain facts into a verifiable next step: repair, wait, ask a person, or close out—and define acceptance for the scenario.
+
Provider / Runtime
External I/O and execution
Call repositories, tests, and tools; perform authorized actions; and read back the result. GitHub still owns the external facts about code, CI, and review.
+
+

For checks failing, a domain capability can propose a successor task to repair CI. The fact that a repair is needed does not itself grant repository write or merge authority. In an experiment, check state becomes metrics and held-out conclusions; the common claim, authority, and recovery rules should remain consistent.

+

That is the value of a capability: it captures recurring judgments and outcome contracts for a scenario without forcing every business stage into the Kernel.

+
+ +
+

When does a long-horizon control plane help?

+

The quadrants describe uncertainty. The need for continuity across turns determines whether additional state management helps. Exploration that can be accepted within an hour may not need a long-horizon system. A fully understood release process may need one because it spans days, handoffs, and external feedback.

+ +

These are concrete reasons to externalize state. After adding LoopX, check whether verified output increases, regressions and duplicated work decrease, and people spend less attention managing it. Count the additional execution and maintenance costs too.

+

My goal remains to have agents do more work while people spend less attention managing them. Separating exploration, execution, and continuing responsibility helps give each kind of work the right management—and reserve human judgment for decisions worth making.

+
+

Benchmarks: start with LHTB, then three supporting signals

+

Most benchmarks provide acceptance criteria, testing problem solving with a clear goal and an unknown route. Even a task that requires algorithm exploration has a specified deliverable and scoring rule. Choosing an open-ended research direction or maintaining an ongoing responsibility is a different level of work.

Start with recovery and regression across long-horizon domains, then use three software-engineering studies to examine continuation, delivery, and validation. The four studies differ in tasks, model settings, and measurement rules. Each needs to be read on its own terms; the current evidence does not establish a universal LoopX gain.

@@ -136,52 +214,6 @@

DeepSWE × V4 Flash max: let counterexamples change the implementation

Withdrawn SSH Goal and Codex CLI scores and conclusions from SWE-Marathon and DeepSWE × Sol are not used in these comparisons.

-
-

02 / Find a path through open-ended exploration

-

“Find a better algorithm” does not come with a fully specified task chain. Real progress may validate a hypothesis—or eliminate a promising-looking route. If only successful conclusions survive, the next run can easily repeat ideas that were already disproved.

-
-
Frame the questionConstraints and hypotheses to test
Try several routesIsolate experiments and bound cost
Record positive and negative evidenceSupport, refute, or raise a new question
Update the directionContinue, combine, retire, or pause
-
-

Explore provides an optional evidence graph and bounded branch planning: nodes represent questions and findings, while relationships express supports, refutes, and leads_to. Planning suggestions still pass through normal execution boundaries; the graph neither launches workers nor grants spending authority.

-

Auto Research organizes this pattern as research work: select a topic, propose hypotheses, run experiments, evaluate independently, and produce a report. Existing protocols and commands provide a foundation, while the public showcase still contains blueprints and items awaiting validation. Real effectiveness must be demonstrated on specific tasks.

- -

Place RSI here, but keep the claim precise

-

When the object of improvement is the agent’s own tools, strategies, or harness, the work becomes self-improvement research. Changing its own code only creates a candidate; sustained improvement still requires independent evaluation, cross-task generalization, regression checks, and rollback.

-

Applying past experience to a later decision, automatically generating experiments, and modifying the harness are different levels. We can discuss an experimental path toward recursive self-improvement (RSI) here; we cannot equate “automatic iteration” with open-ended self-improvement already achieved.

-
- -
-

03 / Own continuous delivery

-

For a pull-request or issue-fix digital worker, the job begins by deciding whether the issue is worth fixing. After delivery, CI, review feedback, branch changes, and the merge result still remain. New tasks keep arriving, and external facts keep changing.

-

Responsibility does not end when the patch is generated.

-

This synthetic scenario starts after a fix has been submitted while CI and review are still running. Choose an external state to see how the next step changes.

- -
-
External fact
GitHub reports that checks are pending for the current revision.
-
Domain state
Record the pull request, revision, and check state; an unchanged observation does not create progress.
-
Capability proposal
Preserve the monitor and recovery condition, then wait for the result.
-
Control-plane check
Admit work by authority, budget, and eligibility; work outside the wait scope can still proceed.
-
Meaning for people
Stay quiet when nothing material changed, so polling does not become nagging.
-
- -

This interactive diagram is a design explanation. It is not connected to a real repository and performs no GitHub operations.

-

Here, a “digital worker” means continuing responsibility, boundaries, memory, and a feedback loop. Its value should be judged by accepted fixes, reopenings and regressions, handling latency, delivery cost, and how often a person must repeatedly supervise it. Pull-request count and uptime are only process signals.

-
- -
-

The richer the domain, the clearer the division of responsibility must be

-
-
Kernel
Authority and common lifecycle
Which goals and work items are valid, who may act, which authority and budgets apply, and how work is handed off, paused, and committed.
-
Domain State
Domain continuity
Whether this issue is actionable, which revision the pull request names, where CI and review stand, and which conclusions remain valid.
-
Capability
Outcome contract and judgment
Translate domain facts into a verifiable next step: repair, wait, ask a person, or close out—and define acceptance for the scenario.
-
Provider / Runtime
External I/O and execution
Call repositories, tests, and tools; perform authorized actions; and read back the result. GitHub still owns the external facts about code, CI, and review.
-
-

For checks failing, a domain capability can propose a successor task to repair CI. The fact that a repair is needed does not itself grant repository write or merge authority. In an experiment, check state becomes metrics and held-out conclusions; the common claim, authority, and recovery rules should remain consistent.

-

That is the value of a capability: it captures recurring judgments and outcome contracts for a scenario without forcing every business stage into the Kernel.

-
-

Evolution: add one verifiable capability at a time

    diff --git a/apps/presentation/site/public/blog/assets/task-quadrants-en.svg b/apps/presentation/site/public/blog/assets/task-quadrants-en.svg new file mode 100644 index 0000000000..08bb882715 --- /dev/null +++ b/apps/presentation/site/public/blog/assets/task-quadrants-en.svg @@ -0,0 +1,55 @@ + + Goal and process clarity: three kinds of long-running work + Goal clarity increases to the right and process clarity upward. Top left: clarify the goal. Top right: routine execution. Bottom left: open exploration. Bottom right: complex tasks. Continuous-delivery digital workers have an ongoing responsibility across quadrants, rather than occupying a separate quadrant. + + + + + + + Two uncertainties. Three kinds of long-running work. + + + + + + + + + Process clarity + Clear + Unclear + Goal clarity + Unclear + Clear + + ALIGNMENT + Clarify the goal + Known workflow; + unclear success criteria + Which decisions should it inform? + ROUTINE + Routine execution + Format conversion; + a validated workflow + Focus: reliable execution + 02 / DISCOVERY + Open exploration + Find directions and opportunities + Keep hypotheses + counterevidence + 01 / COMPLETION + Complex tasks + Repair · refactor · migrate + Keep acceptance gaps and versions + + + 03 / Continuous-delivery digital workers + Across quadrants: clarify → explore → deliver → follow up + LoopX · 01 and 02: solving work. 03: continuing responsibility. + diff --git a/apps/presentation/site/public/blog/assets/task-quadrants-zh.svg b/apps/presentation/site/public/blog/assets/task-quadrants-zh.svg new file mode 100644 index 0000000000..dd48c2987e --- /dev/null +++ b/apps/presentation/site/public/blog/assets/task-quadrants-zh.svg @@ -0,0 +1,53 @@ + + 目标明确度与过程明确度:三类长程工作的关系 + 目标明确度向右增加,过程明确度向上增加。左上先澄清目标,右上常规执行,左下是开放探索,右下是复杂任务。持续交付的数字员工是一层跨越象限的职责,不属于单独一个象限。 + + + + + + + 两种不确定性,三类长程工作 + + + + + + + + + 过程明确度 + 明确 + 模糊 + 目标明确度 + 模糊 + 明确 + + ALIGNMENT + 先澄清目标 + 流程熟悉,结果标准未定 + 如:周报应帮助谁做什么决策? + ROUTINE + 常规执行 + 格式转换 · 已验证的流程 + 重点:可靠执行 + 02 / DISCOVERY + 开放探索 + 选方向 · 提问题 · 探索机会 + 保留:假设、反证与候选路线 + 01 / COMPLETION + 复杂任务 + 修复 · 重构 · 迁移 + 保留:验收缺口与版本证据 + + + 03 / 持续交付的数字员工 + 跨越象限:澄清 → 探索 → 实现 → 跟进 + LoopX · 01、02 是求解方式;03 是持续职责。 + diff --git a/apps/presentation/site/public/blog/zh/application-scenarios/index.html b/apps/presentation/site/public/blog/zh/application-scenarios/index.html index 8a0465a026..d041f692c8 100644 --- a/apps/presentation/site/public/blog/zh/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/zh/application-scenarios/index.html @@ -22,10 +22,8 @@ button:focus-visible{outline:2px solid var(--blue);outline-offset:3px}button[aria-pressed=true]{background:var(--ink);color:#fff} .site-header{padding:0 24px}.article-heading{padding-top:64px}.article-heading .eyebrow{margin-top:0} .chapter-line{display:flex;flex-wrap:wrap;gap:16px;align-items:center;color:var(--body);font-size:12px;margin-top:28px} - .scenario-grid{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:16px;margin:32px 0} - .scenario{padding:24px;border:1px solid var(--border);border-radius:12px;background:var(--surface)} - .scenario .number{font:500 12px var(--font-mono);color:var(--blue)}.scenario h3{margin:16px 0 12px;font-size:20px}.scenario p{font-size:14px;margin:0}.scenario small{display:block;border-top:1px solid var(--border);padding-top:16px;margin-top:20px;font-size:12px} .takeaway{font-size:22px!important;line-height:1.65;color:var(--ink);letter-spacing:-.02em} + .quadrant-figure .architecture-viewport a{min-width:640px}.quadrant-figure figcaption{font-size:13px} .loop-row{display:grid;grid-template-columns:repeat(4,minmax(0,1fr));gap:10px;margin:24px 0}.loop-row>div{padding:16px;border:1px solid var(--border);background:var(--surface);border-radius:6px;font-size:14px}.loop-row b{display:block;color:var(--ink)}.loop-row span{display:block;font-size:12px;margin-top:8px} .comparison{display:grid;grid-template-columns:1fr 1fr;gap:16px;margin:24px 0}.comparison>div{border:1px solid var(--border);border-radius:12px;padding:24px;background:var(--surface)}.comparison strong{display:block;font-size:32px;letter-spacing:-.04em;margin:12px 0}.comparison p{font-size:14px;margin:0} .study-signal{margin:32px 0;padding:28px 0 4px;border-top:1px solid var(--border);scroll-margin-top:24px}.study-signal h3{margin:0 0 12px}.study-signal .study-setting{font-size:12px;margin-bottom:16px}.study-signal p{margin-bottom:16px}.study-signal .signal-boundary{font-size:14px}.study-signal .study-link{font-size:14px;margin-bottom:0}.presenting .study-signal .study-setting{font-size:14px}.presenting .study-signal .signal-boundary,.presenting .study-signal .study-link{font-size:17px} @@ -33,10 +31,10 @@ .small-note{font-size:12px!important;color:var(--body)}.layers{margin:24px 0;border-top:1px solid var(--border)}.layer{display:grid;grid-template-columns:150px 1fr;gap:24px;border-bottom:1px solid var(--border);padding:18px 0;font-size:14px}.layer b{color:var(--ink)} .evolution{counter-reset:step;list-style:none!important;padding:0!important}.evolution li{counter-increment:step;display:grid;grid-template-columns:40px 1fr;gap:16px;padding:20px 0;border-bottom:1px solid var(--border);margin:0}.evolution li::before{content:"0" counter(step);font:500 14px var(--font-mono);color:var(--blue);padding-top:4px}.evolution b{display:block}.evolution span{display:block;font-size:14px;margin-top:6px} .source-list{font-size:13px}.site-footer{padding:32px 24px;border-top:1px solid var(--border);font-size:12px}.site-footer a{margin-left:auto} - .presenting .article-layout{max-width:1200px;grid-template-columns:180px minmax(0,960px);gap:40px}.presenting .prose{font-size:21px}.presenting .prose h2{font-size:36px}.presenting .takeaway{font-size:28px!important}.presenting .prose section{padding-bottom:40px;scroll-margin-top:24px}.presenting table{font-size:17px}.presenting .scenario p,.presenting .loop-row>div,.presenting .layer,.presenting .state-demo dd{font-size:17px}.presenting .toc{font-size:13px}.presenting .toc details{top:20px} + .presenting .article-layout{max-width:1200px;grid-template-columns:180px minmax(0,960px);gap:40px}.presenting .prose{font-size:21px}.presenting .prose h2{font-size:36px}.presenting .takeaway{font-size:28px!important}.presenting .prose section{padding-bottom:40px;scroll-margin-top:24px}.presenting table{font-size:17px}.presenting .loop-row>div,.presenting .layer,.presenting .state-demo dd{font-size:17px}.presenting .toc{font-size:13px}.presenting .toc details{top:20px} @media(max-width:1100px){.article-heading,.article-layout{margin-left:24px;margin-right:24px}.presenting .article-layout{grid-template-columns:160px minmax(0,1fr)}} - @media(max-width:760px){.scenario-grid,.comparison{grid-template-columns:1fr}.loop-row{grid-template-columns:1fr 1fr}.article-layout,.presenting .article-layout{display:block}.toc{margin-bottom:40px}.toc details{position:static}.site-header{gap:12px;flex-wrap:wrap;padding:12px 20px}.site-header nav{gap:16px}.header-actions{gap:8px}.article-heading{padding-top:40px}.state-demo dl,.layer{grid-template-columns:1fr;gap:8px}.state-demo dd{margin-bottom:16px}.scenario small{margin-top:12px;padding-top:12px}.presenting .prose{font-size:18px}} - @media print{.site-header,.toc,.stage-buttons,.chapter-line,.site-footer{display:none!important}.article-layout,.presenting .article-layout{display:block;max-width:none;padding:24px 0}.article-heading{padding:0 0 24px}.prose h2{break-after:avoid}.diagram,.comparison,.state-demo,.scenario{break-inside:avoid}body{background:#fff}} + @media(max-width:760px){.comparison{grid-template-columns:1fr}.loop-row{grid-template-columns:1fr 1fr}.article-layout,.presenting .article-layout{display:block}.toc{margin-bottom:40px}.toc details{position:static}.site-header{gap:12px;flex-wrap:wrap;padding:12px 20px}.site-header nav{gap:16px}.header-actions{gap:8px}.article-heading{padding-top:40px}.state-demo dl,.layer{grid-template-columns:1fr;gap:8px}.state-demo dd{margin-bottom:16px}.presenting .prose{font-size:18px}} + @media print{.site-header,.toc,.stage-buttons,.chapter-line,.site-footer{display:none!important}.article-layout,.presenting .article-layout{display:block;max-width:none;padding:24px 0}.article-heading{padding:0 0 24px}.prose h2{break-after:avoid}.diagram,.comparison,.state-demo{break-inside:avoid}body{background:#fff}} @@ -55,18 +53,24 @@

    LoopX 用在哪里?
    复杂任务、开放探索与持续交付

    -

    先看工作,再看抽象

    -
    -
    01 / COMPLETION

    验收明确的复杂任务

    重构、迁移、协议实现。终点比较明确,执行路线仍有不确定性。

    保留:验收缺口、版本、修复证据
    关注:完成率与成本
    -
    02 / DISCOVERY

    开放探索类任务

    调研、算法实验、系统优化。下一条路线由新证据不断改变。

    保留:假设、反证、候选方向
    关注:有效发现与验证
    -
    03 / DELIVERY

    持续交付的数字员工

    接收 issue、修复、跟进评审,处理新反馈,并对结果持续负责。

    保留:职责、外部状态、等待条件
    关注:交付质量与人的投入
    -
    -

    这三类可以嵌套。一个长期维护仓库的 Agent,可能先探索性能问题,再完成一项验收明确的修复,最后持续跟进 PR。前两类主要描述工作如何求解,第三类还增加了持续的职责与到达的新任务。

    -

    真正延续下来的,应当是工作及其判断依据。

    +

    先问两件事:什么算完成,怎样做成?

    +

    面对一项交给 Agent 的工作,我会先拆开两种不确定性:目标是否明确,过程是否明确。知道自己要什么,和知道怎样得到它,是两回事。

    +

    目标明确,是能说清结果、约束与验收标准;过程明确,是已有可复用的做法。步骤很多的流程,也可以很明确;一句“把系统优化好”,却可能同时缺少结果标准和求解路线。

    +
    +
    + 横轴是目标明确度,向右增加;纵轴是过程明确度,向上增加。左上先澄清目标,右上常规执行,左下是 02 开放探索,右下是 01 复杂任务。03 持续交付的数字员工跨越这些象限。 +
    +
    明确度会随证据和人的反馈变化。01、02 描述工作如何求解,03 描述责任如何延续。窄屏可横向查看,点击图可打开原图。
    +
    +

    右上:目标、过程都明确。按既定格式转换数据、运行已验证的发布流程,重点在可靠执行。短小的工作用脚本或 Workflow 就很好;如果还要跨天等待、处理中断、对后续结果负责,才出现额外的长程管理需求。

    +

    左上:过程明确,目标仍模糊。团队知道怎样写周报,却未必说清它应该帮助谁做什么决策。这时应先澄清读者、价值标准和取舍。把熟悉的流程加速,并不会自动让结果更有用。

    +

    右下对应 01,左下对应 02。前者知道终点,需要找路;后者连下一步值得追求什么,也要通过探索逐渐弄清。探索仍然可以约定本轮要回答的问题、证据要求和停止条件,目标模糊不等于无限续跑。

    +

    03 是横跨象限的持续职责。一个长期维护仓库的 Agent,可能先澄清性能诉求,再探索瓶颈、完成修复、跟进 PR。它既承担复杂任务,也承担探索,还要在新反馈到来时继续负责。

    +

    工作的不确定性,决定该保留什么;
    责任的持续性,决定要保留多久。

    @@ -76,11 +80,85 @@

    01 / 把复杂任务做完

    定义验收规范、基线、预算
    实现并验证产物绑定具体版本
    找出缺口测试之外还缺什么
    继续或停止补缺口,或明确收口

    LoopX 的作用,是让目标和验收缺口跨轮次保持连续,让新证据进入下一步决策。模型负责实现与判断,领域验证器负责检验结果;控制面记录哪些结果被接受、哪些工作仍未完成。

    +

    这里要保留的不只是完成了哪些步骤,还包括当前验收版本、未解决的缺口、尝试过的修复,以及每条验证绑定的产物版本。任务列表耗尽而验收未满足,下一步仍是重新规划;人修改了兼容性要求,旧版本的通过记录也不能代替新验收。

    适合它的任务,通常会经历多轮验证、等待或交接。短小、一次会话就能验收的工作,现有 Agent 可能已经足够,新增控制面需要证明自身开销值得。

+
+

02 / 在开放探索中找路

+

“找到一个更好的算法方案”没有预先写好的完整任务链。真正的进展可能是验证一个假设,也可能是排除一条看似有希望的路线。只保留成功结论,下一轮就容易重试已被否定的想法。

+
+
提出问题约束与待验证假设
尝试多条路线隔离实验,控制预算
记录正反证据支持、反驳、引出新问题
更新探索方向继续、合并、淘汰或暂停
+
+

Explore 已有可选的证据图和有界分支规划:节点表示问题与发现,关系表达 supports、refutes、leads_to。规划建议需要经过正常执行边界,图本身不启动 Worker,也不授予花费权限。

+

Auto Research 把这一思路组织成研究工作:选题、提出假设、执行实验、独立评价、形成报告。现有协议与命令提供了基础,公开 showcase 文档仍包含蓝图和待验证项;实际效果需要在具体任务中证明。

+

探索的验收可以落在“减少了哪些不确定性”:比较候选路线、给出可复现实验、排除一个假设,或者明确还缺什么证据。一个有依据的否定结论,也可以完成一轮有边界的研究。有限的研究交付能收口,长期的方向探索则可以由下一轮继续承担。

+ +

RSI 放在这里,但收紧含义

+

如果改进对象变成 Agent 自己的工具、策略或 Harness,就进入自改进研究。能修改自身代码,只是能够提出候选版本;持续变强还需要独立评价、跨任务泛化、回归检查和回滚。

+

把历史经验用于下一次决策、自动生成实验、修改 Harness,是不同层次。这里可以讨论通往递归自改进(RSI)的实验路径,不能把“自动迭代”直接说成已经实现开放式自我提升。

+
+ +
+

03 / 持续承担交付责任

+

一个 PR / issue fix 数字员工,工作的起点是“这件事是否值得修”,交付后还要面对 CI、评审意见、分支变化和合并结果。任务不断到来,外部事实也不断变化。

+

生成补丁之后,责任还没有结束。

+

下面用一个构造场景说明:修复已提交,CI 和评审仍在进行。点击外部状态,看下一步如何改变。

+ +
+
外部事实
GitHub 返回当前修订的 checks pending。
+
领域状态
记录 PR、修订与检查状态;相同观察无需制造新进展。
+
能力提出下一步
保留监控和恢复条件,等待结果。
+
控制面检查
按授权、预算和执行资格准入;等待范围之外的工作仍可推进。
+
对人的意义
没有重要变化时保持安静,避免把轮询变成催促。
+
+ +

交互图为设计说明,不连接真实仓库,也不执行任何 GitHub 操作。

+

“数字员工”在这里意味着有持续职责、边界、记忆与反馈闭环。是否值得使用,要看被接受的修复、重开与回归、处理时延、每次交付成本,以及需要人反复盯守多少次。PR 数量和在线时长只是过程信号。

+

持续职责还需要明确何时唤醒、何时保持安静、何时必须交给人:新的评审意见可以成为工作输入,未变化的 CI 等待无需反复打扰;权限不足、目标取舍或超出预算,则需要明确停点。一次修复可以完成,维护职责仍可继续;持续职责也应当能被暂停、撤回或交接。

+
+ +
+

工作会跨越象限,状态必须跟得上

+

以维护服务性能为例。最初只有“让系统更快”,人和 Agent 先确定负载、指标与兼容边界;结果标准明确以后,还要探索瓶颈在哪里。找到修复路线,工作再转为实现与验收;提交之后,又进入评审、反馈与回归跟进。

+
+
澄清期待固定负载、指标与边界
探索瓶颈保留基线、假设与反证
落实修复实现并验证具体版本
持续跟进评审、回归与新反馈
+
+

这种移动不是单向升级:一次新反馈可能改变目标,一条反例可能推翻路线。探索出来的稳定方法可以沉淀为常规流程,常规流程遇到新条件又需要重新探索。Agent 需要识别当前剩下的是目标分歧、路线不确定,还是纯粹等待。

+

任务名并不决定位置,当前掌握的信息才决定位置。同样叫“调研”,如果已经约定问题、交付与评分,就可以按验收明确的任务管理;同样叫“重构”,如果方法已成熟,也可以更多依靠既定流程。两条轴描述程度,象限只标出常见的工作形态。

+

因此,交接时有价值的状态包括:当前目标与约束、为什么选择这条路线、已经被否定的方案、产物与验收版本,以及下一次继续或停止的条件。只保存聊天摘要或“已完成第几步”,很难承接这些变化。

+

同一项工作可以更换模型、会话或执行者,但目标的修订、失败证据与尚未履行的责任应当连续。这也是长程控制面在三个场景中的共同位置。

+
+ +
+

领域越丰富,分工越要清楚

+
+
Kernel
控制权与通用生命周期
哪个目标与工作项有效,谁可执行,哪些授权和预算适用,怎样交接、等待与提交状态。
+
Domain State
领域连续性
这个 issue 是否可修,PR 对应哪个修订,CI 和评审到了哪一步,哪些结论仍然有效。
+
Capability
结果契约与判断
把领域事实翻译成可验证的下一步:修复、等待、交给人判断或收口,并定义这个场景怎样验收。
+
Provider / Runtime
外部读写与实际执行
调用代码仓库、测试和工具,执行已授权动作并回读结果。GitHub 仍拥有代码、CI 和评审的外部事实。
+
+

以 checks failing 为例:领域能力可以提出一个修 CI 的后继任务;它不能因“需要修”就自行得到写仓库或合并权限。换成实验任务,CI 状态变成指标和留出集结论,通用的认领、授权、恢复规则仍应保持一致。

+

这也是 Capability 的价值:把一个场景中反复出现的判断与结果契约沉淀下来。它并不要求把所有业务阶段都写进 Kernel。

+
+ +
+

什么时候值得加一层长程控制面?

+

象限说明工作的不确定性,跨轮次的连续性需求决定是否需要额外的状态管理。一个小时就能验收的探索,未必需要长程系统;路线完全明确的发布,却可能因跨天等待、多人接手和外部反馈而需要它。

+ +

这些地方,状态外置化才有具体的价值。引入 LoopX 后,应检查有效产出是否增加、回退和重复工作是否减少,以及人是否少花了管理注意力;同时计算新增的执行与维护成本。

+

我的目标始终是让 Agent 多干活、人少操心。把探索、执行与持续职责拆开,是为了让不同工作得到合适的管理方式,让人的判断用在值得介入的地方。

+
+

Benchmark:从 LHTB 看长程工作,再看三个补充信号

+

多数 benchmark 已经给定验收,主要检验“目标明确、路线未定”的求解能力。即使一道题要求探索算法,最终要提交什么、如何评分仍然已知;这和开放地选择研究方向、持续承担维护职责,是不同层面的工作。

先看跨领域长程任务中的恢复与回退,再用三项软件工程研究补充续跑、交付和验证行为。四项研究的任务、模型设置与统计口径不同,应分别理解;现有结果还不足以概括 LoopX 的普遍增益。

@@ -133,52 +211,6 @@

DeepSWE × V4 Flash max:让反例改变实现

SWE-Marathon 与 DeepSWE × Sol 已撤回的 SSH Goal、Codex CLI 成绩和结论均不用于这里的比较。

-
-

02 / 在开放探索中找路

-

“找到一个更好的算法方案”没有预先写好的完整任务链。真正的进展可能是验证一个假设,也可能是排除一条看似有希望的路线。只保留成功结论,下一轮就容易重试已被否定的想法。

-
-
提出问题约束与待验证假设
尝试多条路线隔离实验,控制预算
记录正反证据支持、反驳、引出新问题
更新探索方向继续、合并、淘汰或暂停
-
-

Explore 已有可选的证据图和有界分支规划:节点表示问题与发现,关系表达 supports、refutes、leads_to。规划建议需要经过正常执行边界,图本身不启动 Worker,也不授予花费权限。

-

Auto Research 把这一思路组织成研究工作:选题、提出假设、执行实验、独立评价、形成报告。现有协议与命令提供了基础,公开 showcase 文档仍包含蓝图和待验证项;实际效果需要在具体任务中证明。

- -

RSI 放在这里,但收紧含义

-

如果改进对象变成 Agent 自己的工具、策略或 Harness,就进入自改进研究。能修改自身代码,只是能够提出候选版本;持续变强还需要独立评价、跨任务泛化、回归检查和回滚。

-

把历史经验用于下一次决策、自动生成实验、修改 Harness,是不同层次。这里可以讨论通往递归自改进(RSI)的实验路径,不能把“自动迭代”直接说成已经实现开放式自我提升。

-
- -
-

03 / 持续承担交付责任

-

一个 PR / issue fix 数字员工,工作的起点是“这件事是否值得修”,交付后还要面对 CI、评审意见、分支变化和合并结果。任务不断到来,外部事实也不断变化。

-

生成补丁之后,责任还没有结束。

-

下面用一个构造场景说明:修复已提交,CI 和评审仍在进行。点击外部状态,看下一步如何改变。

- -
-
外部事实
GitHub 返回当前修订的 checks pending。
-
领域状态
记录 PR、修订与检查状态;相同观察无需制造新进展。
-
能力提出下一步
保留监控和恢复条件,等待结果。
-
控制面检查
按授权、预算和执行资格准入;等待范围之外的工作仍可推进。
-
对人的意义
没有重要变化时保持安静,避免把轮询变成催促。
-
- -

交互图为设计说明,不连接真实仓库,也不执行任何 GitHub 操作。

-

“数字员工”在这里意味着有持续职责、边界、记忆与反馈闭环。是否值得使用,要看被接受的修复、重开与回归、处理时延、每次交付成本,以及需要人反复盯守多少次。PR 数量和在线时长只是过程信号。

-
- -
-

领域越丰富,分工越要清楚

-
-
Kernel
控制权与通用生命周期
哪个目标与工作项有效,谁可执行,哪些授权和预算适用,怎样交接、等待与提交状态。
-
Domain State
领域连续性
这个 issue 是否可修,PR 对应哪个修订,CI 和评审到了哪一步,哪些结论仍然有效。
-
Capability
结果契约与判断
把领域事实翻译成可验证的下一步:修复、等待、交给人判断或收口,并定义这个场景怎样验收。
-
Provider / Runtime
外部读写与实际执行
调用代码仓库、测试和工具,执行已授权动作并回读结果。GitHub 仍拥有代码、CI 和评审的外部事实。
-
-

以 checks failing 为例:领域能力可以提出一个修 CI 的后继任务;它不能因“需要修”就自行得到写仓库或合并权限。换成实验任务,CI 状态变成指标和留出集结论,通用的认领、授权、恢复规则仍应保持一致。

-

这也是 Capability 的价值:把一个场景中反复出现的判断与结果契约沉淀下来。它并不要求把所有业务阶段都写进 Kernel。

-
-

演进:每一步增加一种可验证的能力