Early agentic coding practices (such as early iterations of Superpowers and multi-package agentic-coding suites) pioneered structured AI collaboration. Those designs were highly effective for earlier generations of models that required extensive hand-holding. However, as reasoning and planning capabilities have advanced substantially, past architectural assumptions face new trade-offs:
- Context Overhead from Over-Granular Skill Splitting
Splitting the engineering workflow across a dozen independent skills (brainstorming, task decomposition, TDD, code review, branch management, etc.) helped ensure single-task focus early on. With modern multi-task models, however, numerous competing skill descriptions consume precious prompt space and increase routing ambiguity. - Step-by-Step Checklists Over-Constraining Reasoning
Early rules relied heavily on exhaustive procedural checklists and rigid turn counts to prevent models from skipping steps. Modern models already possess strong self-reflection and analytical abilities; overly mechanical procedures constrain their ability to make pragmatic engineering trade-offs. - Protective Guardrails Inducing Excessive Hesitation
Rigid confirmation checkpoints originally designed to prevent unauthorized edits can cause models to pause prematurely after an initial draft, or mistake task documentation for human approval gates requiring sign-off at every turn. - Continuing Need for Deterministic Verification
Regardless of model capabilities, without objective external tooling, models can still declare tasks complete without running real tests or checking exit codes.
This workflow consolidates fragmented practices into a single, highly cohesive entry point structured around four core principles:
- Progressive Disclosure
The rootSKILL.mdserves as a minimal navigation router. In-depth reference guides (architecture trade-offs, permission boundaries, re-review, Git delivery) are loaded strictly on demand when concrete decisions arise, preserving everyday context. - Invariants over Step Micromanagement
Rather than prescribing rigid conversational turns or procedures, the workflow enforces hard technical and permission boundaries: complete four-document audit semantics, no self-granted permissions, no unapproved force-pushes, and no discarding user uncommitted changes. - Autonomous Momentum After Authorization
The four documents serve as post-hoc traceable audit trails, not bureaucratic roadblocks. Once the user authorizes implementation, the agent works autonomously through coding, real testing, and bug fixing until green delivery, without nagging at every minor step. - Deterministic Read-Only Tooling
Requires zero heavy external frameworks. Built entirely with the Python standard library and local Git: a structural validator that prevents premature completion, and a Git scope tool that deterministically resolves true merge-bases while rejecting shallow histories and dangerous filter executions.
Every multi-phase development, refactoring, or continuation task persists four records within a single directory: tasks/YYYYMMDD-HHMMSS-<name>/:
issue.md → implementation_plan.md → task.md → walkthrough.md
| Document | Phase | Core Constraint |
|---|---|---|
issue.md |
Intake | Locks in original goals, non-negotiable technical constraints, non-goals, and acceptance criteria. |
implementation_plan.md |
Design | Captures non-derivable interfaces, architectural trade-offs, and verification strategies; no hollow fluff. |
task.md |
Execution | Broken down into verifiable delivery units; progress and evidence concentrate here, zero fake checkmarks. |
walkthrough.md |
Delivery | Records actual test commands, exit codes, residual risks, and necessary recovery procedures. |
Requires only Python 3.10+ standard library and local Git. Zero external dependencies.
Verifies document structure integrity and prevents declaring completion with pending work:
# Verify design phase (issue.md + implementation_plan.md)
python scripts/check_task.py tasks/<task-dir> --stage design
# Verify implementation phase (first three documents)
python scripts/check_task.py tasks/<task-dir> --stage implement
# Full delivery check (all four documents complete, no unchecked tasks)
python scripts/check_task.py tasks/<task-dir> --stage completeRead-only, deterministic calculation of exact Git review scope. Accurately finds branch merge bases, rejecting shallow clones and dangerous clean/process filters:
# Branch review mode (calculates merge-base between target branch and HEAD)
python scripts/review_scope.py --repo . --base main --mode branch
# Commit range mode (verifies task-base to task-head)
python scripts/review_scope.py --repo . --base <base-sha> --head HEAD --mode range
# Include local worktree status (separately lists staged / unstaged / untracked)
python scripts/review_scope.py --repo . --base main --mode branch --include-worktree- Copy or symlink this directory into your Agent's skills folder (e.g.,
~/.codex/skills/engineering-workflow). - Invoke explicitly in conversation:
Entry definition: SKILL.md.
$engineering-workflow
The design principles, artifact patterns, and engineering mechanisms in this workflow were informed and adapted from:
- Google Antigravity by Google DeepMind — The tripartite artifact architecture (
implementation_plan.md,task.md,walkthrough.md) for transparent, auditable agent planning and execution. - obra/superpowers by Jesse Vincent (MIT License) — Rigorous test-driven development (TDD), Git merge-base review scopes, and structured branch delivery workflows.
- clawic/skills by Ivan G. Davila (MIT License) — Modular skill layout conventions and packaging best practices.
Distributed under the Apache-2.0 License. See NOTICE for upstream attributions.