Research, algorithm experiments, and system optimization. New evidence continuously changes the next route.
Preserve: hypotheses, counterevidence, candidate directions Measure: useful discoveries and validation
-
03 / DELIVERY
Digital workers for continuous delivery
Take an issue, fix it, follow review, handle new feedback, and remain accountable for the result.
Preserve: responsibility, external state, waiting conditions Measure: delivery quality and human effort
-
-
These categories can nest. An agent maintaining a repository over time may first explore a performance problem, then complete a fix with clear acceptance, and finally follow the pull request through review. The first two categories mainly describe how work is solved; the third adds an ongoing responsibility and a stream of newly arriving work.
-
What should persist is the work—and the reasoning behind its decisions.
+
Ask two questions: what counts as done, and how do we get there?
+
For work handed to an agent, I first separate two uncertainties: clarity of the goal and clarity of the process. Knowing what you want and knowing how to achieve it are different things.
+
A clear goal names the outcome, constraints, and acceptance criteria. A clear process has a reusable method. A workflow with many steps can be well understood; “make the system better” may leave both the outcome and the route unresolved.
+
+
+
+
+ Clarity changes with evidence and human feedback. 01 and 02 describe how work is solved; 03 describes how responsibility continues. Scroll horizontally on narrow screens, or open the image at full size.
+
+
Top right: clear goal, known process. Convert data to a specified format or run a validated release workflow. Reliable execution matters here. A script or workflow can be enough for short work; waiting across days, recovering from interruption, or owning downstream results creates an additional continuity need.
+
Top left: known process, unclear goal. A team may know how to produce a weekly report without having agreed whose decisions it should inform. Clarify the audience, value criteria, and tradeoffs first. Speeding up a familiar process does not automatically make its output useful.
+
Bottom right maps to 01; bottom left maps to 02. The former knows the destination and must find a route. The latter also uses exploration to discover what is worth pursuing. An exploratory round can still define its question, evidence requirements, and stopping conditions; an unclear goal is not a mandate to run indefinitely.
+
03 is a continuing responsibility across quadrants. An agent maintaining a repository may clarify performance needs, investigate a bottleneck, complete a fix, and follow its pull request. It takes on both complex tasks and exploration, while remaining responsible when new feedback arrives.
+
Uncertainty determines what must persist. Continuing responsibility determines for how long.
@@ -79,11 +83,85 @@
01 / Finish a complex task
Define acceptanceSpecification, baseline, budget
Implement and validateBind artifacts to a specific version
Find the gapsWhat remains beyond the tests?
Continue or stopClose gaps, or close out explicitly
LoopX keeps the goal and its acceptance gaps continuous across turns, then carries new evidence into the next decision. The model implements and judges; domain validators check the result; the control plane records which results were accepted and which work remains.
+
Preserve more than completed steps: the current acceptance version, unresolved gaps, attempted repairs, and the artifact version each check evaluated. If the task list is exhausted but acceptance is unmet, replan. If a person changes compatibility requirements, a pass on the old version cannot establish acceptance under the new criteria.
The best-fitting tasks usually span repeated validation, waiting, or handoffs. Existing agents may already be enough for small work that can be accepted in one session; an added control plane must justify its overhead.
+
+
02 / Find a path through open-ended exploration
+
“Find a better algorithm” does not come with a fully specified task chain. Real progress may validate a hypothesis—or eliminate a promising-looking route. If only successful conclusions survive, the next run can easily repeat ideas that were already disproved.
+
+
Frame the questionConstraints and hypotheses to test
Try several routesIsolate experiments and bound cost
Record positive and negative evidenceSupport, refute, or raise a new question
Update the directionContinue, combine, retire, or pause
+
+
Explore provides an optional evidence graph and bounded branch planning: nodes represent questions and findings, while relationships express supports, refutes, and leads_to. Planning suggestions still pass through normal execution boundaries; the graph neither launches workers nor grants spending authority.
+
Auto Research organizes this pattern as research work: select a topic, propose hypotheses, run experiments, evaluate independently, and produce a report. Existing protocols and commands provide a foundation, while the public showcase still contains blueprints and items awaiting validation. Real effectiveness must be demonstrated on specific tasks.
+
Exploration can be accepted by the uncertainty it removes: compare candidate routes, produce a reproducible experiment, rule out a hypothesis, or identify the evidence still missing. A well-supported negative conclusion can complete a bounded research round. A finite research deliverable can close out while a longer-term direction continues in another round.
+
+
Place RSI here, but keep the claim precise
+
When the object of improvement is the agent’s own tools, strategies, or harness, the work becomes self-improvement research. Changing its own code only creates a candidate; sustained improvement still requires independent evaluation, cross-task generalization, regression checks, and rollback.
+
Applying past experience to a later decision, automatically generating experiments, and modifying the harness are different levels. We can discuss an experimental path toward recursive self-improvement (RSI) here; we cannot equate “automatic iteration” with open-ended self-improvement already achieved.
+
+
+
+
03 / Own continuous delivery
+
For a pull-request or issue-fix digital worker, the job begins by deciding whether the issue is worth fixing. After delivery, CI, review feedback, branch changes, and the merge result still remain. New tasks keep arriving, and external facts keep changing.
+
Responsibility does not end when the patch is generated.
+
This synthetic scenario starts after a fix has been submitted while CI and review are still running. Choose an external state to see how the next step changes.
+
+
+
+
+
External fact
GitHub reports that checks are pending for the current revision.
+
Domain state
Record the pull request, revision, and check state; an unchanged observation does not create progress.
+
Capability proposal
Preserve the monitor and recovery condition, then wait for the result.
+
Control-plane check
Admit work by authority, budget, and eligibility; work outside the wait scope can still proceed.
+
Meaning for people
Stay quiet when nothing material changed, so polling does not become nagging.
+
+
+
This interactive diagram is a design explanation. It is not connected to a real repository and performs no GitHub operations.
+
Here, a “digital worker” means continuing responsibility, boundaries, memory, and a feedback loop. Its value should be judged by accepted fixes, reopenings and regressions, handling latency, delivery cost, and how often a person must repeatedly supervise it. Pull-request count and uptime are only process signals.
+
A continuing responsibility also needs rules for when to wake, when to stay quiet, and when to involve a person. New review feedback can become work input; unchanged pending CI does not need repeated interruptions. Missing authority, goal tradeoffs, or exhausted budgets need explicit stopping points. One repair can finish while the maintenance role continues; that role must also be pausable, revocable, and transferable.
+
+
+
+
Work moves between quadrants. State must follow.
+
Consider maintaining a service’s performance. “Make it faster” first requires agreement on workload, metrics, and compatibility constraints. Once the outcome is clear, the bottleneck still needs investigation. Finding a repair route turns the work into implementation and acceptance; submission brings review, feedback, and regression follow-up.
Deliver the fixImplement and verify a specific version
Keep following upReview, regressions, new feedback
+
+
This is not a one-way progression. Feedback can change the goal; a counterexample can invalidate the route. A stable method discovered through exploration can become a routine workflow, while new conditions can send a workflow back into exploration. An agent must recognize whether the remaining issue is disagreement about the goal, uncertainty about the route, or simply waiting.
+
A task’s name does not determine its position; the information currently available does. Research with an agreed question, deliverable, and scoring rule can be managed as a task with clear acceptance. A refactor with a mature method can rely more on an established workflow. Both axes are continuous; the quadrants mark common work patterns.
+
Useful handoff state includes the current goal and constraints, why this route was chosen, rejected approaches, artifact and acceptance versions, and conditions for continuing or stopping. A chat summary or “step three completed” alone cannot carry these changes.
+
The same work may change models, sessions, or executors. Revisions to its goal, evidence of failure, and outstanding responsibilities should remain continuous. This is the shared place of a long-horizon control plane in all three scenarios.
+
+
+
+
The richer the domain, the clearer the division of responsibility must be
+
+
Kernel Authority and common lifecycleWhich goals and work items are valid, who may act, which authority and budgets apply, and how work is handed off, paused, and committed.
+
Domain State Domain continuityWhether this issue is actionable, which revision the pull request names, where CI and review stand, and which conclusions remain valid.
+
Capability Outcome contract and judgmentTranslate domain facts into a verifiable next step: repair, wait, ask a person, or close out—and define acceptance for the scenario.
+
Provider / Runtime External I/O and executionCall repositories, tests, and tools; perform authorized actions; and read back the result. GitHub still owns the external facts about code, CI, and review.
+
+
For checks failing, a domain capability can propose a successor task to repair CI. The fact that a repair is needed does not itself grant repository write or merge authority. In an experiment, check state becomes metrics and held-out conclusions; the common claim, authority, and recovery rules should remain consistent.
+
That is the value of a capability: it captures recurring judgments and outcome contracts for a scenario without forcing every business stage into the Kernel.
+
+
+
+
When does a long-horizon control plane help?
+
The quadrants describe uncertainty. The need for continuity across turns determines whether additional state management helps. Exploration that can be accepted within an hour may not need a long-horizon system. A fully understood release process may need one because it spans days, handoffs, and external feedback.
+
+
Complex tasks: changing sessions means explaining acceptance again, recovering repair progress, or rediscovering known gaps.
+
Open exploration: several routes run in parallel, negative results disappear, and the next round needs evidence-based allocation rather than recollection.
+
Continuing responsibility: tasks, reviews, or external outcomes keep arriving, leaving a person to act as dispatcher and follow-up manager.
+
+
These are concrete reasons to externalize state. After adding LoopX, check whether verified output increases, regressions and duplicated work decrease, and people spend less attention managing it. Count the additional execution and maintenance costs too.
+
My goal remains to have agents do more work while people spend less attention managing them. Separating exploration, execution, and continuing responsibility helps give each kind of work the right management—and reserve human judgment for decisions worth making.
+
+
Benchmarks: start with LHTB, then three supporting signals
+
Most benchmarks provide acceptance criteria, testing problem solving with a clear goal and an unknown route. Even a task that requires algorithm exploration has a specified deliverable and scoring rule. Choosing an open-ended research direction or maintaining an ongoing responsibility is a different level of work.
Start with recovery and regression across long-horizon domains, then use three software-engineering studies to examine continuation, delivery, and validation. The four studies differ in tasks, model settings, and measurement rules. Each needs to be read on its own terms; the current evidence does not establish a universal LoopX gain.
@@ -136,52 +214,6 @@
DeepSWE × V4 Flash max: let counterexamples change the implementation
Withdrawn SSH Goal and Codex CLI scores and conclusions from SWE-Marathon and DeepSWE × Sol are not used in these comparisons.
-
-
02 / Find a path through open-ended exploration
-
“Find a better algorithm” does not come with a fully specified task chain. Real progress may validate a hypothesis—or eliminate a promising-looking route. If only successful conclusions survive, the next run can easily repeat ideas that were already disproved.
-
-
Frame the questionConstraints and hypotheses to test
Try several routesIsolate experiments and bound cost
Record positive and negative evidenceSupport, refute, or raise a new question
Update the directionContinue, combine, retire, or pause
-
-
Explore provides an optional evidence graph and bounded branch planning: nodes represent questions and findings, while relationships express supports, refutes, and leads_to. Planning suggestions still pass through normal execution boundaries; the graph neither launches workers nor grants spending authority.
-
Auto Research organizes this pattern as research work: select a topic, propose hypotheses, run experiments, evaluate independently, and produce a report. Existing protocols and commands provide a foundation, while the public showcase still contains blueprints and items awaiting validation. Real effectiveness must be demonstrated on specific tasks.
-
-
Place RSI here, but keep the claim precise
-
When the object of improvement is the agent’s own tools, strategies, or harness, the work becomes self-improvement research. Changing its own code only creates a candidate; sustained improvement still requires independent evaluation, cross-task generalization, regression checks, and rollback.
-
Applying past experience to a later decision, automatically generating experiments, and modifying the harness are different levels. We can discuss an experimental path toward recursive self-improvement (RSI) here; we cannot equate “automatic iteration” with open-ended self-improvement already achieved.
-
-
-
-
03 / Own continuous delivery
-
For a pull-request or issue-fix digital worker, the job begins by deciding whether the issue is worth fixing. After delivery, CI, review feedback, branch changes, and the merge result still remain. New tasks keep arriving, and external facts keep changing.
-
Responsibility does not end when the patch is generated.
-
This synthetic scenario starts after a fix has been submitted while CI and review are still running. Choose an external state to see how the next step changes.
-
-
-
-
-
External fact
GitHub reports that checks are pending for the current revision.
-
Domain state
Record the pull request, revision, and check state; an unchanged observation does not create progress.
-
Capability proposal
Preserve the monitor and recovery condition, then wait for the result.
-
Control-plane check
Admit work by authority, budget, and eligibility; work outside the wait scope can still proceed.
-
Meaning for people
Stay quiet when nothing material changed, so polling does not become nagging.
-
-
-
This interactive diagram is a design explanation. It is not connected to a real repository and performs no GitHub operations.
-
Here, a “digital worker” means continuing responsibility, boundaries, memory, and a feedback loop. Its value should be judged by accepted fixes, reopenings and regressions, handling latency, delivery cost, and how often a person must repeatedly supervise it. Pull-request count and uptime are only process signals.
-
-
-
-
The richer the domain, the clearer the division of responsibility must be
-
-
Kernel Authority and common lifecycleWhich goals and work items are valid, who may act, which authority and budgets apply, and how work is handed off, paused, and committed.
-
Domain State Domain continuityWhether this issue is actionable, which revision the pull request names, where CI and review stand, and which conclusions remain valid.
-
Capability Outcome contract and judgmentTranslate domain facts into a verifiable next step: repair, wait, ask a person, or close out—and define acceptance for the scenario.
-
Provider / Runtime External I/O and executionCall repositories, tests, and tools; perform authorized actions; and read back the result. GitHub still owns the external facts about code, CI, and review.
-
-
For checks failing, a domain capability can propose a successor task to repair CI. The fact that a repair is needed does not itself grant repository write or merge authority. In an experiment, check state becomes metrics and held-out conclusions; the common claim, authority, and recovery rules should remain consistent.
-
That is the value of a capability: it captures recurring judgments and outcome contracts for a scenario without forcing every business stage into the Kernel.
-
-
Evolution: add one verifiable capability at a time