Skip to content

About

Portable agent-brain kit: superpowers methodology, YAGNI minimalism, cross-chat SQLite FTS5 memory, adversarial trap-suite evals. Hermes-compatible skills for OMP/Claude Code/Gemini CLI/Hermes/Antigravity/ZCode.

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Repository files navigation

coding-kit

Reusable workflows and cross-session memory for coding agents.

Give your agent a shared development method, searchable project knowledge, and tools to evaluate its behavior. The kit combines 37 skills, plain-text instructions, and a Python standard-library memory engine. It is not an agent runtime, a sandbox, or a guarantee of better answers.

Kit gates Release License: MIT

Get started · Daily use · Evidence and limitations · Contributing

What you can use it for

  • Carry decisions across sessions. Save findings in your private memory store and search them from a fresh agent session instead of reconstructing project history from chat.
  • Give coding work a repeatable method. Skills cover planning, reproducing bugs, implementation, review, and verification. YAGNI means avoiding functionality and abstractions the task does not need.
  • Evaluate behavior rather than trust promises. Run policy scenarios, skill-routing checks, and small coding tasks with deterministic verifiers. These answer different questions; none alone proves overall coding quality.

Developed and tested with Claude Code and Oh My Pi (OMP). Other agents that read instruction files and SKILL.md skills can be configured manually; this project does not claim to have tested their behavior. Instructions are in English; the kit asks the agent to answer in your language.

Get started

Installation has two separate steps: create the memory store, then connect your agent. The first step does not install agent instructions.

1. Create your memory store

Requirements: Git and Python 3.12. CI runs on Windows and Ubuntu with Python 3.12; other Python versions are untested. Runtime scripts use the standard library. Only contributors running the test suite need pytest.

Run these commands in PowerShell or Bash:

git clone https://github.com/oleg494/coding-kit.git coding-kit
cd coding-kit
python scripts/install.py

The installer creates your user-level ~/.memory/ directory, including fixtures and indexes, and links its engine to this clone. Keep the clone in place. Re-running the installer is supported. Look for search smoke: OK. It does not configure your agent's rules or skills.

Use a different memory location

Set MEMORY_ROOT before installation and in sessions that use the kit.

PowerShell:

$env:MEMORY_ROOT = "$HOME/coding-memory"
python scripts/install.py

Bash:

export MEMORY_ROOT="$HOME/coding-memory"
python scripts/install.py

Examples below call the engine from the clone, so they do not depend on shell expansion of ~ in script arguments.

2. Connect your agent

Merge a pointer to this clone's OPS.md into your agent's rules file. Use an absolute path so the contract remains reachable from other projects. Load it once; AGENTS.md is only a repository router. Preserve existing instructions rather than replacing them wholesale.

Copy or link only the 37 skills declared in profile.yml into the agent's skill directory, preserving unrelated skills. Skill bodies load on demand; no full method chain or memory warmup is required at startup. These user-level changes can affect every project opened with that agent.

Agent Rules file Skills
Claude Code ~/.claude/CLAUDE.md ~/.claude/skills/
Oh My Pi (OMP) ~/.omp/agent/AGENTS.md ~/.agents/skills/ with user-agent skill discovery enabled
Antigravity ~/AGENTS.md ~/.agents/skills/
ZCode ~/.zcode/AGENTS.md ~/.zcode/skills/
Hermes Delimited block in SOUL.md Generated owned-skill projection via Hermes adapter

The OMP paths above match the kit's deployment targets. See the Antigravity and ZCode guides for their environment-specific setup.

3. Verify the connection

From the clone, enter the tools directory, save a note, and retrieve it:

cd memory/db-tools
python findings.py add "first-note" --text "hello" --project coding-kit
python findings.py search "first-note" --project coding-kit
cd ../..

Then start a fresh agent session and ask:

Search my coding-kit memory for "first-note" using the memory tools. Show the command you ran and the stored finding.

The agent should invoke findings.py search or search_all.py against your memory store and retrieve the note, not merely repeat this example. If it does not, check that the rules and skill paths are loaded by your agent.

python scripts/doctor.py checks repository and memory consistency. A green result is not proof that an agent loaded or followed the kit.

Daily use

Run the memory tools from the clone, or use their absolute paths elsewhere:

python memory/db-tools/search_all.py "deployment decision"
python memory/db-tools/search_all.py "deployment decision" --project coding-kit
python memory/db-tools/search_all.py "deployment decision" --importance high
python memory/db-tools/findings.py projects

Ask the agent to save a decision when it is worth carrying into another session. Memory writes need authorization; a read-only review should not silently modify your knowledge base.

On "continue work on X", the agent first recovers the project's mission, original user grants and next action from memory and available history. Current corrections override old scope; expired or revoked grants do not revive. Missing records require a focused question, not an invented mission. Verification closes once the current objective is evidenced: bounded tasks end with a report; active autonomous missions move to the next useful objective. These are instruction contracts, not runtime enforcement or measured speedups.

Skill cross-references are pointers, not a cascading load order. Load a helper for its actual domain need; host-mandated loads still take precedence. Reuse verification tied to unchanged source, dependencies and environment; rerun when a relevant change or failure invalidates that evidence.

profile.yml declares the kit-owned skills. Third-party skills beside them (including Firecrawl) are not kit release assets and remain untouched by kit validation and synchronization. This boundary does not stop your agent from independently discovering or using those skills.

Ownership and lightweight navigation

The kit owns its contract, declared skills, memory scripts and generated routers. Approval mode, sandboxing, model/provider routing, compaction and third-party tools such as Graphify or Firecrawl belong to the host/user setup; kit releases neither change those settings nor claim to repair host rules.

Use targeted source search and LSP for ordinary code navigation. For an existing indexed project, python memory/db-tools/repomap.py project --db <project.db> --tokens 1500 gives a bounded map without installing a graph stack. Confirm indexed relationships against current source; an index may be stale. This is not a replacement for Graphify's multimodal graph-building features, which ordinary coding tasks do not require.

Memory availability is an explicit check: python memory/scripts/memory-warmup.py; add --full only for diagnostic feeds and integrity inspection. No automatic global findings feed or memory writes.

Your knowledge lives in ~/.memory/ (or MEMORY_ROOT), outside the kit repository. The clone contains methodology and tooling, not your personal project history. Do not commit your memory store or assume it is a sandbox for untrusted data. See the security policy.

Back up and verify recovery

The store is the one asset this repository cannot rebuild (db/*.db are gitignored by design). Back it up, then prove the snapshot is usable:

python scripts/tools/backup_memory.py
python scripts/tools/backup_memory.py --list
python scripts/tools/backup_memory.py --restore-drill <backup directory>

Backups land in <memory root>/backups/<timestamp>/. A snapshot whose database was skipped is tagged DEGRADED and refused as a restore point. --restore-drill restores into a temporary root, runs PRAGMA integrity_check on every restored database, verifies the findings store and searches a token taken from the restored rows; it prints JSON and exits 0 only when that verification passes. --restore <dir> performs a live restore (pre-restore snapshot first; asks unless --yes).

Organize findings by project and importance

Projects are discovered from the memory root's db/*.db files and optional projects.json. Project slugs use [a-z0-9][a-z0-9_-]{0,63}. Use portable for reusable cross-project knowledge and unknown for unclassified notes.

Importance levels:

  • high: critical boundaries, invariants, and durable release contracts.
  • normal: actionable findings, runbooks, and feature setups.
  • low: temporary checkpoints and scratch notes.
  • unreviewed: findings not yet qualitatively reviewed.

Use findings.py edit --help to update a finding using its returned ID; IDs are not guaranteed to start at 1. For batch classification, findings.py classify mapping.json --dry-run validates a mapping before mutation. Classification is transactional and preserves user-curated records unless --force is supplied. See findings.py.

Evidence and limitations

The kit is not a demonstrated universal coding-quality improvement. It adds instructions and can add work. Whether that helps depends on the model, task, and integration.

A historical external DeepSWE A/B run compared the same agent with and without the kit using deepseek-v4-pro, pier + mini-swe-agent in Docker, and a 10-task seed-0 subset:

Observation Reported result
Solved tasks in the reported nine-task comparison 6/9 in both arms
Steps across five mutually solved tasks +21% with the kit
Prompt tokens across those five tasks +41%: 99.5M vs 70.4M
Task-level differences One kit win and one kit loss

This was a small historical sample using a 36-skill manifest, not a benchmark of the current release. Raw artifacts are retained outside this repository; the repo does not ship a reproduction script for these numbers. The results are descriptive, not a causal explanation or a general reliability claim. Prompt-token counts are not a measured monetary bill.

The 4.5.1 release notes separately document policy-calibration observations and their limits. They are stated-next-action evidence, not an end-to-end coding-quality win rate.

Run the kit's checks

python -m pip install pytest
python scripts/doctor.py
python -m pytest tests -q
python scripts/tools/check_file_sizes.py --ci

CI runs the kit gates on Windows and Ubuntu. Structural validation can run without a model or paid API calls:

python eval/runner.py --inline-skills
python eval/task_runner.py --dry-run
python eval/trigger_eval.py --queries eval/trigger_queries.json

These commands validate evaluation inputs; they do not measure a live model's behavior. Trap dry-runs record DRY_RUN rows with passed: 0; exit zero means the inputs are valid, not that the model passed.

Report input validation, stated-next-action text probes, executed tool tasks and controlled comparisons separately. Record the resolved model and reasoning level when observed; otherwise label them unknown. A self-assessed text probe is not a tool-loop success rate, speedup, or Sol-versus-Astra comparison.

Trap and trigger results carry comparison_id: a digest of executed prompts, expectations, selected cases, scorer code and execution settings. trend.py compares only matching IDs within a model. Unknown IDs stay visible without deltas and cannot seed baselines; existing run files are never rewritten. On --update-baselines, new trap/trigger baselines use {model: {comparison_id: rate}}; old numeric baselines are not reused. Matching IDs do not control mutable images, ambient skills or provider aliases, and do not prove that an instruction change caused an improvement.

Evaluation tools and what they measure
Tool Purpose and boundary
Trap-suite Adversarial policy scenarios. The judge defaults to the executor; use a distinct judge to reduce self-judging bias. Answers over 8000 characters fail evaluation without a model verdict, rather than judging an incomplete prefix. Policy adherence is not task superiority.
Task smoke Six coding tasks, including two impossible canaries, with deterministic verify.py oracles. A smoke check, not a statistical benchmark.
Trigger evals Skill activation routing: use --queries auto for current co-located cases; an 80-query central corpus provides fallback coverage.
Results store and trend Structured results, explicit live/dry-run modes, failure categories, and comparisons of recorded runs.
Telemetry Measures duration. Optional usage totals are user-reported, not independently measured token cost.
Ablation Compares prompts with and without an inlined skill. Ambient skills remain uncontrolled; results are descriptive, not causal.
Rigor A/B Policy experiments with isolation probes and an acceptance gate that can reject a candidate.

Live evaluations require an executor and may incur provider charges. Prompt evaluations require docker:<image> <argv...>; bare host commands are refused. The image must contain the executor. Only explicitly declared read-only mounts are visible alongside a disposable writable task directory; host credentials are not inherited. Network is off by default. @net enables unrestricted bridge networking, not endpoint allowlisting: live-model CK-03 acceptance remains open until credential and network restrictions are verified.

Autonomous work (opt-in)

Requests such as “do useful work” or “keep going without asking” activate autonomous-work: evidence-backed work selection, verified progress, and immediate stop/revocation. Ordinary bounded requests do not become autonomous missions. Outward, destructive, spending, and memory-writing actions still need authorization.

An optional foreground Python supervisor supports continuation across process boundaries:

python scripts/tools/autonomous.py --help

It accepts an executor and independent verifier, records resumable state, and reports completion only after verification succeeds. A STOP file prevents further spawning and interrupts a live child. Read the supervisor contract before configuring commands; a workspace is not a security sandbox.

Carry unfinished work into a fresh session

Capture a task brief with scripts/tools/handoff.py. The brief is JSON with goal, acceptance, constraints, pending, and observations. Each observation has a claim and a nonempty paths list of workspace-relative regular files that support it. List only files you intend to fingerprint.

python scripts/tools/handoff.py capture --workspace /path/to/repo --brief brief.json --output handoff.json
python scripts/tools/handoff.py resume --workspace /path/to/repo --handoff handoff.json

Capture stores hashes, not file contents, and refuses to overwrite an existing handoff. Resume is read-only: it returns the complete task context and marks observations stale when their supporting files changed, disappeared, or became unsafe to read. Exit codes: 0 unchanged, 1 drift, 2 invalid input. Use --json on resume for structured output. Unchanged bytes are not proof that a claim is true, that tests passed, or that an action is authorized.

Give the fresh agent the resume output and the workspace, not the previous chat. It must inspect stale evidence and preserve current user edits before continuing. For automatic worker launches, add --handoff /path/to/handoff.json to the supervisor invocation. The supervisor refreshes this context before each worker; the configured independent verifier remains the completion gate. Neither tool executes commands embedded in a handoff, snapshots the whole repository, or replaces the host's permission and sandbox controls.

Inside the repository

Component Entry point
Agent instructions and routing AGENTS.md
Operating contract OPS.md
Context-size modes SKILL_RUNTIME.md
Paths and skill manifest profile.yml
Workflow and domain skills skills/
Memory engine memory/db-tools/
Evaluation tools eval/
Changes and contribution guide Changelog · Contributing

Windows-first development; CI also runs on Ubuntu. The memory engine link is a junction on Windows and a symlink elsewhere.

Credits and license

MIT. Phase-workflow skills are derived from obra/superpowers by Jesse Vincent, reworked and extended for coding-kit; see the superpowers license. The ponytail skill is adapted from DietrichGebert/ponytail; see its license.

About

Portable agent-brain kit: superpowers methodology, YAGNI minimalism, cross-chat SQLite FTS5 memory, adversarial trap-suite evals. Hermes-compatible skills for OMP/Claude Code/Gemini CLI/Hermes/Antigravity/ZCode.

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages