Offline foley pipeline — record a long take, get named axis-annotated sound atoms, compose them into game-ready variants.
Status: early, working. Validated on real recordings. Engine-agnostic (Node + ffmpeg only). Companion to AnchorSFX — atom prepares the material, anchor plays it back in-engine.
AI SFX generators (Suno, ElevenLabs) give you a finished sound. That's great until you need control, and then three problems show up at once:
- Every generated clip carries its own room. Sixty of them in one game means sixty different acoustic spaces, and the mix never settles.
- You can't adjust it. Want the shield break a little drier, the tail a little shorter? You can't — you can only reroll and hope.
- Quality is a dice roll. Some takes are great, some are mush, and you don't get to decide which.
Recording your own material fixes all three — but only if the material is atoms, not finished sounds. That's a different discipline, and it needs different tooling.
record one long take → split-takes → mix-layers → game
(talk, then perform) named atoms variants
You record the way a foley artist actually works: pick something up, say what it is and how you're about to hit it, then hit it. Hit M when you change intensity. Keep going.
The splitter figures out the rest. Markers in the wav's cue chunk define the tiers; everything gets named from that.
Separating your narration from your takes works two ways:
- Transients (impacts, taps, clacks) — free. Speech and percussive onsets differ by an order of magnitude in attack steepness (30–80ms vs <5ms), no ASR needed.
- Sustained sources (fire, wind, water, scraping) — needs
--speech. A flame is slow-attack, long, multi-peak — acoustically indistinguishable from talking. So invert it: locate the speech with word-level ASR timestamps, and everything else is material.
python tools/transcribe.py take.wav -o speech.json
atom-split --input take.wav --speech speech.json --min-dur 2000 --dry-runWord-level timestamps are required — segment-level ones fold "talk … 30s of fire … talk" into one span.
No shot list. No recording script. You don't decide what a sound is for until after you've heard it.
Needs Node ≥18 and ffmpeg on PATH.
git clone https://github.com/AgentGameLab/atom-sfx && cd atom-sfx && npm link1. Look at what you recorded — nothing is written yet:
atom-split --input take.wav --dry-run时长 86.5s · 噪声地板 -61.6dBFS · 触发阈值 -40.0dBFS
切到 37 段:slate 10 · take 27
marker 9 个:28.0s 33.5s 39.1s 50.7s ...
合并出 3 段口报
[ 1] 9 个 take (口报 @ 1.28s)
1. 档1 28.57s attack 7ms dur 143ms peak -6.8dBFS
2. 档1 30.46s attack 6ms dur 123ms peak -11.0dBFS
...
2. Describe what each group was — transcribe your own narration (any ASR), write it down:
{
"source": "desk_padded",
"technique": "fingernail",
"tierAxis": "f",
"groups": [
{ "axis": { "r": "center" } },
{ "axis": { "r": "mid" } },
{ "axis": { "r": "edge" } }
]
}3. Cut:
atom-split --input take.wav --map desk.json --out atoms/desk_padded__fingernail__r_center__f1__01.wav
desk_padded__fingernail__r_center__f1__02.wav
desk_padded__fingernail__r_center__f2__01.wav
...
4. Compose — normalize and render a single source, or layer two sources into one sound:
atom-mix --mode direct --src atoms/ --family desk_padded --def-id piece-move
atom-mix --mode layered --src atoms/ --contact piece_tap --body board_body --body-delay 4Output is numbered {def-id}.mp3, {def-id}-2.mp3, … so it drops straight into an
engine that picks randomly from a variant pool.
1. Atoms must be ingredients, not finished sounds. A library "metal impact" is already designed — it has its own tail, its own room, its own mix. Stack two of those and the ear hears two sounds, not one. Record dry, single-event, short, same signal chain throughout.
2. Split at the level where behaviour differs. If a component appears under one game state and not another, it must be its own atom. Debris only exists when the shield actually breaks; armour rattle is the dominant secondary when it doesn't. Mechanical, not taste-based. The converse also holds — components that always appear together can stay fused, and splitting them just costs you recording time.
3. Record along an axis, not as scattered points. A library gives you fifty unrelated "metal impacts". One mic on one object varying only force gives you a coordinate. That's what makes energy → sound a real mapping instead of a pitch-shift guess, and it's the one thing buying can't give you.
- Normalization is per-family, one gain for the whole set — never per-file. Per-file normalization flattens the force axis: soft and hard end up equally loud, and the axis is gone. The scripts scale the whole family by one coefficient so internal relationships survive.
- Short transients use peak/RMS, not LUFS.
ebur128integrated loudness needs a ≥400ms analysis block; impacts land at 300–400ms, right on the boundary, and the reading isn't trustworthy. - Cutting must re-encode, never
-c copy. WAV-c copyseeks at packet granularity, so-ssrounds to a packet boundary and silently eats the attack.pcm_f32le → pcm_f32leis lossless and sample-accurate. - NG detection is off by default. The "not a take and short" heuristic misfires on soft hits (their attack exceeds the transient threshold) and would silently discard a good take. Enable with
--ngif you want it.
atoms/ ships 374 atoms across 32 sources — wood, stone (dry and damped), two metals, ceramic, glass, leather, cardboard, cork, paper, a flamethrower, plus a few generated ones for what a kitchen can't produce (fire, wind, water, a creature growl). Every one is named with its axis coordinates, so atom-mix can use them straight out of a clone. See library.md for the full index.
Raw session recordings (~200MB) and rendered output stay out of the repo — the atoms are the distilled part, and the output is reproducible from atoms + parameters.
- Library index — every source, its axes, and where it came from (auto-generated)
- Composition notes — the hard-won part: why "heavier" needs a new low layer rather than ducking the others, why shatter is sparse-then-dense, why generated material can fill a gap but never form an axis, why per-file normalization kills your force axis (Chinese)
- 录制说明 — how to record so the tooling can read it (Chinese)
- Schema — atom metadata, axes, naming (Chinese)
MIT — see LICENSE.