Reusable workflows and cross-session memory for coding agents.
Give your agent a shared development method, searchable project knowledge, and tools to evaluate its behavior. The kit combines 37 skills, plain-text instructions, and a Python standard-library memory engine. It is not an agent runtime, a sandbox, or a guarantee of better answers.
Get started · Daily use · Evidence and limitations · Contributing
- Carry decisions across sessions. Save findings in your private memory store and search them from a fresh agent session instead of reconstructing project history from chat.
- Give coding work a repeatable method. Skills cover planning, reproducing bugs, implementation, review, and verification. YAGNI means avoiding functionality and abstractions the task does not need.
- Evaluate behavior rather than trust promises. Run policy scenarios, skill-routing checks, and small coding tasks with deterministic verifiers. These answer different questions; none alone proves overall coding quality.
Developed and tested with Claude Code and Oh My Pi (OMP). Other agents
that read instruction files and SKILL.md skills can be configured manually;
this project does not claim to have tested their behavior. Instructions are
in English; the kit asks the agent to answer in your language.
Installation has two separate steps: create the memory store, then connect your agent. The first step does not install agent instructions.
Requirements: Git and Python 3.12. CI runs on Windows and Ubuntu with
Python 3.12; other Python versions are untested. Runtime scripts use the
standard library. Only contributors running the test suite need pytest.
Run these commands in PowerShell or Bash:
git clone https://github.com/oleg494/coding-kit.git coding-kit
cd coding-kit
python scripts/install.py
The installer creates your user-level ~/.memory/ directory, including
fixtures and indexes, and links its engine to this clone. Keep the clone in
place. Re-running the installer is supported. Look for search smoke: OK.
It does not configure your agent's rules or skills.
Use a different memory location
Set MEMORY_ROOT before installation and in sessions that use the kit.
PowerShell:
$env:MEMORY_ROOT = "$HOME/coding-memory"
python scripts/install.pyBash:
export MEMORY_ROOT="$HOME/coding-memory"
python scripts/install.pyExamples below call the engine from the clone, so they do not depend on
shell expansion of ~ in script arguments.
Merge a pointer to this clone's OPS.md into your agent's rules file. Use an absolute path so the contract remains reachable from other projects. Load it once; AGENTS.md is only a repository router. Preserve existing instructions rather than replacing them wholesale.
Copy or link only the 37 skills declared in profile.yml into the agent's skill directory, preserving unrelated skills. Skill bodies load on demand; no full method chain or memory warmup is required at startup. These user-level changes can affect every project opened with that agent.
| Agent | Rules file | Skills |
|---|---|---|
| Claude Code | ~/.claude/CLAUDE.md |
~/.claude/skills/ |
| Oh My Pi (OMP) | ~/.omp/agent/AGENTS.md |
~/.agents/skills/ with user-agent skill discovery enabled |
| Antigravity | ~/AGENTS.md |
~/.agents/skills/ |
| ZCode | ~/.zcode/AGENTS.md |
~/.zcode/skills/ |
| Hermes | Delimited block in SOUL.md |
Generated owned-skill projection via Hermes adapter |
The OMP paths above match the kit's deployment targets. See the Antigravity and ZCode guides for their environment-specific setup.
From the clone, enter the tools directory, save a note, and retrieve it:
cd memory/db-tools
python findings.py add "first-note" --text "hello" --project coding-kit
python findings.py search "first-note" --project coding-kit
cd ../..
Then start a fresh agent session and ask:
Search my coding-kit memory for "first-note" using the memory tools. Show the command you ran and the stored finding.
The agent should invoke findings.py search or search_all.py against your
memory store and retrieve the note, not merely repeat this example. If it
does not, check that the rules and skill paths are loaded by your agent.
python scripts/doctor.py checks repository and memory consistency.
A green result is not proof that an agent loaded or followed the kit.
Run the memory tools from the clone, or use their absolute paths elsewhere:
python memory/db-tools/search_all.py "deployment decision"
python memory/db-tools/search_all.py "deployment decision" --project coding-kit
python memory/db-tools/search_all.py "deployment decision" --importance high
python memory/db-tools/findings.py projects
Ask the agent to save a decision when it is worth carrying into another session. Memory writes need authorization; a read-only review should not silently modify your knowledge base.
On "continue work on X", the agent first recovers the project's mission, original user grants and next action from memory and available history. Current corrections override old scope; expired or revoked grants do not revive. Missing records require a focused question, not an invented mission. Verification closes once the current objective is evidenced: bounded tasks end with a report; active autonomous missions move to the next useful objective. These are instruction contracts, not runtime enforcement or measured speedups.
Skill cross-references are pointers, not a cascading load order. Load a helper for its actual domain need; host-mandated loads still take precedence. Reuse verification tied to unchanged source, dependencies and environment; rerun when a relevant change or failure invalidates that evidence.
profile.yml declares the kit-owned skills. Third-party skills beside them
(including Firecrawl) are not kit release assets and remain untouched by kit
validation and synchronization. This boundary does not stop your agent from
independently discovering or using those skills.
The kit owns its contract, declared skills, memory scripts and generated routers. Approval mode, sandboxing, model/provider routing, compaction and third-party tools such as Graphify or Firecrawl belong to the host/user setup; kit releases neither change those settings nor claim to repair host rules.
Use targeted source search and LSP for ordinary code navigation. For an
existing indexed project, python memory/db-tools/repomap.py project --db <project.db> --tokens 1500 gives a bounded map without installing a graph
stack. Confirm indexed relationships against current source; an index may be
stale. This is not a replacement for Graphify's multimodal graph-building
features, which ordinary coding tasks do not require.
Memory availability is an explicit check: python memory/scripts/memory-warmup.py; add --full only for diagnostic feeds and
integrity inspection. No automatic global findings feed or memory writes.
Your knowledge lives in ~/.memory/ (or MEMORY_ROOT), outside the kit
repository. The clone contains methodology and tooling, not your personal
project history. Do not commit your memory store or assume it is a sandbox
for untrusted data. See the security policy.
Back up and verify recovery
The store is the one asset this repository cannot rebuild (db/*.db are
gitignored by design). Back it up, then prove the snapshot is usable:
python scripts/tools/backup_memory.py
python scripts/tools/backup_memory.py --list
python scripts/tools/backup_memory.py --restore-drill <backup directory>
Backups land in <memory root>/backups/<timestamp>/. A snapshot whose
database was skipped is tagged DEGRADED and refused as a restore point.
--restore-drill restores into a temporary root, runs PRAGMA integrity_check on every restored database, verifies the findings store and
searches a token taken from the restored rows; it prints JSON and exits 0
only when that verification passes. --restore <dir> performs a live restore
(pre-restore snapshot first; asks unless --yes).
Organize findings by project and importance
Projects are discovered from the memory root's db/*.db files and optional
projects.json. Project slugs use [a-z0-9][a-z0-9_-]{0,63}.
Use portable for reusable cross-project knowledge and unknown for
unclassified notes.
Importance levels:
high: critical boundaries, invariants, and durable release contracts.normal: actionable findings, runbooks, and feature setups.low: temporary checkpoints and scratch notes.unreviewed: findings not yet qualitatively reviewed.
Use findings.py edit --help to update a finding using its returned ID;
IDs are not guaranteed to start at 1. For batch classification,
findings.py classify mapping.json --dry-run validates a mapping before
mutation. Classification is transactional and preserves user-curated records
unless --force is supplied. See findings.py.
The kit is not a demonstrated universal coding-quality improvement. It adds instructions and can add work. Whether that helps depends on the model, task, and integration.
A historical external DeepSWE A/B run
compared the same agent with and without the kit using deepseek-v4-pro,
pier + mini-swe-agent in Docker, and a 10-task seed-0 subset:
| Observation | Reported result |
|---|---|
| Solved tasks in the reported nine-task comparison | 6/9 in both arms |
| Steps across five mutually solved tasks | +21% with the kit |
| Prompt tokens across those five tasks | +41%: 99.5M vs 70.4M |
| Task-level differences | One kit win and one kit loss |
This was a small historical sample using a 36-skill manifest, not a benchmark of the current release. Raw artifacts are retained outside this repository; the repo does not ship a reproduction script for these numbers. The results are descriptive, not a causal explanation or a general reliability claim. Prompt-token counts are not a measured monetary bill.
The 4.5.1 release notes separately document policy-calibration observations and their limits. They are stated-next-action evidence, not an end-to-end coding-quality win rate.
python -m pip install pytest
python scripts/doctor.py
python -m pytest tests -q
python scripts/tools/check_file_sizes.py --ci
CI runs the kit gates on Windows and Ubuntu. Structural validation can run without a model or paid API calls:
python eval/runner.py --inline-skills
python eval/task_runner.py --dry-run
python eval/trigger_eval.py --queries eval/trigger_queries.json
These commands validate evaluation inputs; they do not measure a live
model's behavior. Trap dry-runs record DRY_RUN rows with passed: 0;
exit zero means the inputs are valid, not that the model passed.
Report input validation, stated-next-action text probes, executed tool tasks and controlled comparisons separately. Record the resolved model and reasoning level when observed; otherwise label them unknown. A self-assessed text probe is not a tool-loop success rate, speedup, or Sol-versus-Astra comparison.
Trap and trigger results carry comparison_id: a digest of executed prompts,
expectations, selected cases, scorer code and execution settings. trend.py
compares only matching IDs within a model. Unknown IDs stay visible without
deltas and cannot seed baselines; existing run files are never rewritten.
On --update-baselines, new trap/trigger baselines use
{model: {comparison_id: rate}}; old numeric baselines are not reused.
Matching IDs do not control mutable images, ambient skills or provider aliases,
and do not prove that an instruction change caused an improvement.
Evaluation tools and what they measure
| Tool | Purpose and boundary |
|---|---|
| Trap-suite | Adversarial policy scenarios. The judge defaults to the executor; use a distinct judge to reduce self-judging bias. Answers over 8000 characters fail evaluation without a model verdict, rather than judging an incomplete prefix. Policy adherence is not task superiority. |
| Task smoke | Six coding tasks, including two impossible canaries, with deterministic verify.py oracles. A smoke check, not a statistical benchmark. |
| Trigger evals | Skill activation routing: use --queries auto for current co-located cases; an 80-query central corpus provides fallback coverage. |
| Results store and trend | Structured results, explicit live/dry-run modes, failure categories, and comparisons of recorded runs. |
| Telemetry | Measures duration. Optional usage totals are user-reported, not independently measured token cost. |
| Ablation | Compares prompts with and without an inlined skill. Ambient skills remain uncontrolled; results are descriptive, not causal. |
| Rigor A/B | Policy experiments with isolation probes and an acceptance gate that can reject a candidate. |
Live evaluations require an executor and may incur provider charges.
Prompt evaluations require docker:<image> <argv...>; bare host commands
are refused. The image must contain the executor. Only explicitly declared
read-only mounts are visible alongside a disposable writable task directory;
host credentials are not inherited. Network is off by default. @net enables
unrestricted bridge networking, not endpoint allowlisting: live-model CK-03
acceptance remains open until credential and network restrictions are verified.
Requests such as “do useful work” or “keep going without asking” activate autonomous-work: evidence-backed work selection, verified progress, and immediate stop/revocation. Ordinary bounded requests do not become autonomous missions. Outward, destructive, spending, and memory-writing actions still need authorization.
An optional foreground Python supervisor supports continuation across process boundaries:
python scripts/tools/autonomous.py --help
It accepts an executor and independent verifier, records resumable state,
and reports completion only after verification succeeds. A STOP file
prevents further spawning and interrupts a live child. Read the
supervisor contract before
configuring commands; a workspace is not a security sandbox.
Capture a task brief with scripts/tools/handoff.py. The brief is JSON with
goal, acceptance, constraints, pending, and observations. Each
observation has a claim and a nonempty paths list of workspace-relative
regular files that support it. List only files you intend to fingerprint.
python scripts/tools/handoff.py capture --workspace /path/to/repo --brief brief.json --output handoff.json
python scripts/tools/handoff.py resume --workspace /path/to/repo --handoff handoff.json
Capture stores hashes, not file contents, and refuses to overwrite an existing
handoff. Resume is read-only: it returns the complete task context and marks
observations stale when their supporting files changed, disappeared, or became
unsafe to read. Exit codes: 0 unchanged, 1 drift, 2 invalid input. Use
--json on resume for structured output. Unchanged bytes are not proof that a
claim is true, that tests passed, or that an action is authorized.
Give the fresh agent the resume output and the workspace, not the previous chat.
It must inspect stale evidence and preserve current user edits before continuing.
For automatic worker launches, add --handoff /path/to/handoff.json to the
supervisor invocation. The supervisor refreshes this context before each worker;
the configured independent verifier remains the completion gate. Neither tool
executes commands embedded in a handoff, snapshots the whole repository, or
replaces the host's permission and sandbox controls.
| Component | Entry point |
|---|---|
| Agent instructions and routing | AGENTS.md |
| Operating contract | OPS.md |
| Context-size modes | SKILL_RUNTIME.md |
| Paths and skill manifest | profile.yml |
| Workflow and domain skills | skills/ |
| Memory engine | memory/db-tools/ |
| Evaluation tools | eval/ |
| Changes and contribution guide | Changelog · Contributing |
Windows-first development; CI also runs on Ubuntu. The memory engine link is a junction on Windows and a symlink elsewhere.
MIT. Phase-workflow skills are derived from
obra/superpowers by Jesse Vincent,
reworked and extended for coding-kit; see the
superpowers license.
The ponytail skill is adapted from
DietrichGebert/ponytail;
see its license.