The operator forge — write kernels once, ship them to both training and inference.
An open-source operator/kernel framework built on one thesis:
- Write JIT — kernels are authored in a Python DSL (Triton / CuTeDSL): fast to iterate, low barrier to entry, friendly to AI assistance.
- Ship AOT — compiled artifacts are written to disk ahead of time and distributed with the package: production does no runtime compilation, so cold start is predictable.
- One kernel, two uses — the same compute kernel serves both training and inference. What genuinely cannot be shared (autograd orchestration, CUDA Graph, paged KV) stays in each side's adapter layer.
- The numeric contract is a first-class citizen —
DETERMINISTIC/HIGH_PRECISION/FASTare declared by the operator and chosen by the caller, instead ofbitwise=Truebeing threaded through every call site.
Status: skeleton + four backends (python / triton / cutedsl / cpp escape hatch) + AOT precompilation — L0–L3 interfaces, three real DSL operators, 109 tests, 9 architecture invariants; green on both CPU-only and GPU configurations. CuTeDSL cross-process artifact reuse is implemented (official
export_to_c+cute.runtime.load_module); the compile half is verified here, but the launch half needs a driver no older than cuda-python's CUDA version (see GPU environment notes). The precompiled-wheel package (opforge-prebuilt) is explicitly deferred — seedocs/design.md§6 for the reasoning and the alternatives.
Training-side and inference-side operator libraries are two separate worlds today: the training side (Megatron / DeepSpeed / …) and the inference side (vLLM / SGLang / TensorRT-LLM) each write their own. But most of the math is identical — norm, activation, quantization, GEMM, the MoE backbone. The differences are not in "how to compute", but in:
| Difference | Training | Inference |
|---|---|---|
| Needs backward | Yes | No |
| Numeric requirement | Bitwise-reproducible / high-precision accumulation | Latency first, can relax |
| Shape distribution | Fixed batch, long sequences | Dynamic batch, decode |
| Execution form | autograd + optimizer | CUDA Graph + paged KV |
| Delivery | Source + fast iteration | Precompiled artifact + cold start |
opforge's claim is to handle these differences by layering: push the shared parts
down, keep the unshared parts apart. See docs/design.md.
L0 language layer backend protocol (triton / cutedsl / cpp escape hatch / python reference)
L1 kernel layer operator declaration (@kernel) and the call pipeline ┐
L2 contract layer numeric contracts, variant resolution, preconditions │ shared
L2 cache layer two-level fingerprints, manifest, sealed compilation │ by both
L2 launch layer per-signature caching of the launch closure ┘
L3 adapter layer training-side to_autograd / inference-side to_custom_op ← separate
L0–L2 have zero module-level dependencies — contract resolution, fingerprints, caching and launching can all be run and tested without any framework or GPU.
- Python ≥ 3.10 (3.12 recommended; CI covers 3.10 / 3.11 / 3.12).
- That is all for the core:
numpyis the only runtime dependency and is installed for you. torchis only needed by the L3 adapter layer;triton/nvidia-cutlass-dslare only needed by the GPU backends.
The project is not published on PyPI yet, so install from a checkout:
git clone <repo-url> opforge && cd opforge
# Core only: contract layer + cache + launch + the pure-Python reference backend.
pip install -e .
# If you also want the L3 adapter layer (autograd / custom-op integration):
pip install -e ".[training,inference]"
# If you want the GPU backends (compile Triton / CuTeDSL kernels):
pip install -e ".[torch,triton,cutedsl]"
# Or everything:
pip install -e ".[all]"An operator is a decorated function. You must supply a contract, and you should supply a
reference (pure-PyTorch ground truth) — the reference is the CI accuracy baseline, the
training-side fallback path, and (with backward="auto") the source of the backward pass.
from opforge import Contract, RequiresDtypeIn, kernel
def rms_norm_fast(x, weight, eps=1e-6, *, zero_centered=False):
"""FAST tier: accumulate in the input dtype — faster, less accurate."""
...
@kernel(
name="rms_norm",
contract=Contract.HIGH_PRECISION, # required: the default tier
variants={Contract.FAST: rms_norm_fast}, # optional: the other tiers
preconditions=(RequiresDtypeIn("float32", "float16", "bfloat16"),),
backward="auto",
)
def rms_norm(x, weight, eps=1e-6, *, zero_centered=False):
"""y = x / sqrt(mean(x**2) + eps) * weight"""
...
def rms_norm_ref(x, weight, eps=1e-6, *, zero_centered=False):
"""Pure-PyTorch ground truth.
Found automatically by the `<operator name>_ref` naming convention, so you never have
to pass it explicitly.
"""
...To make it run on a GPU you additionally declare which backend builds the implementation
(declare_triton / declare_cutedsl / declare_cpp) — see
CONTRIBUTING.md §"Adding an operator" and the three built-in examples
kernels/norm/rms_norm.py (Triton), kernels/elementwise/scale.py (CuTeDSL) and
kernels/reduction/warp_reduce.py (the C++ escape hatch).
from opforge import Contract
from opforge.kernels.norm import rms_norm
out = rms_norm(x, w, contract=Contract.HIGH_PRECISION)The call pipeline is: resolve the contract → check preconditions → cache hit or build → launch. Two failure modes are deliberately distinct:
- the contract cannot be satisfied → raises
ContractUnsatisfied(hard failure; carrying on would silently change the numeric semantics); - a precondition is unmet (no CUDA Graph capture, ranks not in lockstep, wrong dtype …) →
returns
Unavailable(normal fallback, not an exception, so the caller can route elsewhere).
Contracts are ordered, and fallback is only allowed toward a stricter tier:
| Requested | Implementations {DET, FAST} |
Implementations {FAST} only |
|---|---|---|
DETERMINISTIC |
uses DET |
raises ContractUnsatisfied |
HIGH_PRECISION |
uses DET (stricter, acceptable) |
raises ContractUnsatisfied |
FAST |
uses FAST |
uses FAST |
Every fallback is counted, so it can be surfaced as a metric rather than staying invisible:
from opforge import fallback_counts
print(fallback_counts()) # {("rms_norm", "DETERMINISTIC", "HIGH_PRECISION"): 3, ...}# Tests (bundled stdlib runner; if pytest is installed, plain `pytest` works too)
PYTHONPATH=src python3 tools/run_tests.py
# Architecture invariants I1–I9
PYTHONPATH=src python3 tools/check_invariants.pyBoth are also what CI runs. The test runner needs no pytest on purpose: a zero-dependency core layer is one of the design goals, so the core must be testable in a bare environment.
The second half of "write JIT, ship AOT". Operators compile the shapes they will need at build time and write them to disk; runtime only loads.
# Build time (on the same architecture you deploy to) — mount both dirs:
TRITON_CACHE_DIR=/opt/opforge-triton \
OPFORGE_CACHE_DIR=/opt/opforge-kernels \
python3 tools/build_aot.py
# Runtime: mount the same two dirs — installed means warm start.
export TRITON_CACHE_DIR=/opt/opforge-triton
export OPFORGE_CACHE_DIR=/opt/opforge-kernelsThe two directories have different jobs: Triton artifacts can only live in Triton's own
cache (it has no export API), whereas CuTeDSL artifacts are managed by opforge itself
(the object file from export_to_c plus a manifest, at
<OPFORGE_CACHE_DIR>/v1/<target_fp>/<ns>/<source_fp>/<entry_id>.artifact).
Running the build twice therefore shows directly that the second run recompiled nothing:
3 kernels, 3 compiled, 0 loaded, ok=True # cold: everything compiles
rms_norm triton HIGH_PRECISION ok
scale cutedsl HIGH_PRECISION ok
warp_reduce_sum cpp DETERMINISTIC ok
3 kernels, 2 compiled, 1 loaded, ok=True # run again: CuTeDSL hits the disk
rms_norm triton HIGH_PRECISION ok
scale cutedsl HIGH_PRECISION loaded
warp_reduce_sum cpp DETERMINISTIC ok
An operator must declare which shapes to compile, otherwise the build step has nothing
to work from: Triton uses declare_triton(..., prewarm=...), CuTeDSL uses
declare_cutedsl(..., examples=...) (for CuTeDSL the example inputs are a required
argument of cute.compile, not an optional optimization).
At runtime, combine it with KernelCache.seal() to assert "nothing compiles for the rest
of the process lifetime" — after sealing, any build that misses raises CompilationSealed,
so "the second run passed" is itself proof that nothing was recompiled.
python3 -m build --wheel # or: python3 -m pip wheel . -w dist --no-deps
# artifact: dist/opforge-0.0.1.dev0-py3-none-any.whlOn a machine with a GPU, source tools/gpu_env.sh before running the Triton path. It does
two things, both container-environment quirks rather than design choices of this
project:
LD_PRELOAD=/usr/local/nvidia/lib64/libcuda.so.1— in the container/usr/lib/x86_64-linux-gnu/libcuda.sopoints at a zero-byte stub and the real driver lives in/usr/local/nvidia/lib64/. torch loads it with ctypes' defaultRTLD_LOCAL, so the symbols never enter the global scope, and Triton's driver extension then fails withundefined symbol: cuModuleGetFunction— which also breakstorch.compile's inductor backend.PYTHONNOUSERSITE=1— skips a broken torch under~/.localthat would otherwise shadow the working one.
CuTeDSL has one more constraint: compilation only needs the toolchain
(CUTE_DSL_ARCH=sm_90a must be set explicitly), but execution requires the driver
to be no older than cuda-python's CUDA version. When it is older, launching raises
cudaErrorInsufficientDriver — so on such a host CuTeDSL's AOT compile + cross-process
load half still works while the launch half cannot run; the corresponding test skips
itself.
Note that verifying cross-process reuse does not depend on being able to launch: the
producing process writes the artifact with export_to_c, and the consuming process loads
it into a sealed cache (KernelCache.seal()) — getting the artifact back proves there
was no fallback to cute.compile, which would have raised CompilationSealed.
docs/
design.md why: trade-offs and motivations
architecture.md how: layered interfaces and invariants
src/opforge/
contract/ numeric contracts, variant resolution, preconditions
kernel/ @kernel decorator, KernelSpec, call pipeline
cache/ two-level fingerprints, manifest, sealed compilation
launch/ per-signature caching of the launch closure
lang/ backend protocol (python / triton / cutedsl / cpp), escape-hatch registration
kernels/ operator implementations (norm/, elementwise/, reduction/)
training/ L3 training-side adapter
inference/ L3 inference-side adapter
aot.py AOT precompile entry point
tests/ accuracy, contract, cache, backend, AOT, adapter tests
tools/ test runner, invariant checker, AOT build, GPU env script
.github/workflows/ CI (CPU) and the GPU job
| Document | Contents |
|---|---|
docs/design.md |
Why: layered architecture, the numeric contract model, escape-hatch policy, compile cache and fingerprints, delivery forms, CI quality gates, roadmap |
docs/architecture.md |
How: the L0–L4 interfaces — the full @kernel signature, the Contract enum and variant derivation, the Precondition protocol, the cache/launch layers, architecture invariants I1–I9 |
CONTRIBUTING.md |
Layering conventions, how to add an operator / a backend, and the two easy-to-trip pitfalls |
Apache-2.0 (see LICENSE).
算子锻造坊 —— 一次写内核,同时发到训练与推理两侧。
一个开源算子/内核框架,主张只有一条:
- JIT 写 —— 内核用 Python DSL(Triton / CuTeDSL)编写:迭代快、门槛低、易被 AI 辅助。
- AOT 发 —— 编译产物提前落盘并随包分发:生产环境不做运行期编译,冷启动可预测。
- 一套内核,两用 —— 同一份计算内核同时服务训练与推理;真正不共用的部分 (autograd 编排、CUDA Graph、paged KV)留在各自的适配层。
- 数值契约是一等公民 ——
DETERMINISTIC/HIGH_PRECISION/FAST由算子声明、 调用方按场景选择,而不是把bitwise=True一路传遍所有调用点。
状态:骨架 + 四个后端(python / triton / cutedsl / cpp 逃生口)+ AOT 预编译 —— L0–L3 接口、三个真 DSL 算子、109 项测试、9 条架构不变量;CPU-only 与 GPU 两种配置都跑绿。 CuTeDSL 的跨进程产物复用已实现(官方
export_to_c+cute.runtime.load_module): 编译那半已验证,launch 那半要求驱动不旧于 cuda-python 的 CUDA 版本才能跑(见 GPU 环境)。 预编译包(opforge-prebuilt)明确推迟 —— 理由与替代方案见docs/design.md§6。
训练侧与推理侧的算子库今天基本是两套:训练侧(Megatron / DeepSpeed 等)和推理侧 (vLLM / SGLang / TensorRT-LLM)各写各的。但大部分算子的数学是一样的 —— norm、 激活、量化、GEMM、MoE 主干。差别不在"怎么算",而在:
| 差异 | 训练 | 推理 |
|---|---|---|
| 需要 backward | 是 | 否 |
| 数值要求 | 逐位可复现 / 高精度累加 | 延迟优先,可放开 |
| shape 分布 | 固定 batch、长序列 | 动态 batch、decode |
| 执行形态 | autograd + 优化器 | CUDA Graph + paged KV |
| 交付 | 源码 + 快速迭代 | 预编译产物 + 冷启动 |
opforge 的主张是把这些差异分层处理:共用的下沉、不共用的分开。
详见 docs/design.md。
L0 语言层 后端协议(triton / cutedsl / cpp 逃生口 / python 参考后端)
L1 内核层 算子声明(@kernel)与调用流水线 ┐
L2 契约层 数值契约、变体解析、前置条件 │ 两侧共用
L2 缓存层 两级指纹、manifest、封编译 │
L2 启动层 按签名缓存启动闭包 ┘
L3 适配层 训练侧 to_autograd / 推理侧 to_custom_op ← 两侧分开
L0–L2 的模块级依赖为零 —— 契约解析、指纹、缓存、启动都能脱离任何框架与 GPU 运行与测试。
- Python ≥ 3.10(推荐 3.12;CI 覆盖 3.10 / 3.11 / 3.12)。
- 核心层只需要这些:
numpy是唯一的运行期依赖,会随安装自动装上。 torch只有 L3 适配层需要;triton/nvidia-cutlass-dsl只有 GPU 后端需要。
本项目尚未发布到 PyPI,请从源码检出安装:
git clone <repo-url> opforge && cd opforge
# 只装核心:契约层 + 缓存 + 启动 + 纯 Python 参考后端
pip install -e .
# 需要 L3 适配层(autograd / custom op 集成)时:
pip install -e ".[training,inference]"
# 需要 GPU 后端(编译 Triton / CuTeDSL 内核)时:
pip install -e ".[torch,triton,cutedsl]"
# 或者全装:
pip install -e ".[all]"一个算子就是一个被装饰的函数。contract 是必填的;同时应该给出 reference
(纯 PyTorch 真值)—— 它既是 CI 的精度基准,也是训练侧的回退路径,配合
backward="auto" 还是反向实现的来源。
from opforge import Contract, RequiresDtypeIn, kernel
def rms_norm_fast(x, weight, eps=1e-6, *, zero_centered=False):
"""FAST 档:在输入 dtype 上累加 —— 更快、更不精确。"""
...
@kernel(
name="rms_norm",
contract=Contract.HIGH_PRECISION, # 必填:默认档位
variants={Contract.FAST: rms_norm_fast}, # 可选:其它档位
preconditions=(RequiresDtypeIn("float32", "float16", "bfloat16"),),
backward="auto",
)
def rms_norm(x, weight, eps=1e-6, *, zero_centered=False):
"""y = x / sqrt(mean(x**2) + eps) * weight"""
...
def rms_norm_ref(x, weight, eps=1e-6, *, zero_centered=False):
"""纯 PyTorch 真值。
按 `<算子名>_ref` 的命名约定自动找到,不需要显式传入。
"""
...要让它跑在 GPU 上,还要声明由哪个后端构建这个实现
(declare_triton / declare_cutedsl / declare_cpp)—— 见
CONTRIBUTING.md 的「加一个算子」一节,以及三个内置范例:
kernels/norm/rms_norm.py(Triton)、kernels/elementwise/scale.py(CuTeDSL)、
kernels/reduction/warp_reduce.py(C++ 逃生口)。
from opforge import Contract
from opforge.kernels.norm import rms_norm
out = rms_norm(x, w, contract=Contract.HIGH_PRECISION)调用流水线是:解析契约 → 检查前置条件 → 缓存命中或构建 → 启动。 两种失败刻意区分开:
- 契约无法满足 → 抛
ContractUnsatisfied(硬失败:继续执行等于悄悄改变数值语义); - 前置条件不满足(在 CUDA Graph 捕获期、多 rank 未步调一致、dtype 不符……)→
返回
Unavailable(正常回退,不是异常,好让调用方换路走)。
契约是有序的,回退只允许朝更严的方向:
| 请求 | 有 {DET, FAST} |
只有 {FAST} |
|---|---|---|
DETERMINISTIC |
用 DET |
抛 ContractUnsatisfied |
HIGH_PRECISION |
用 DET(更严,可接受) |
抛 ContractUnsatisfied |
FAST |
用 FAST |
用 FAST |
每次回退都会被计数,因此可以当成指标暴露出来,而不是让它无声无息:
from opforge import fallback_counts
print(fallback_counts()) # {("rms_norm", "DETERMINISTIC", "HIGH_PRECISION"): 3, ...}# 测试(自带 stdlib 运行器;装了 pytest 也可以直接 pytest)
PYTHONPATH=src python3 tools/run_tests.py
# 架构不变量 I1–I9
PYTHONPATH=src python3 tools/check_invariants.py这两条也是 CI 跑的。测试运行器刻意不依赖 pytest:核心层零依赖是本项目的设计目标之一, 所以核心必须能在裸环境里被测。
"JIT 写、AOT 发"的后半句。算子在 build 期把要用的形状编好并落盘,运行期只加载:
# 构建期(在与部署相同的架构上)—— 两个目录都值得挂:
TRITON_CACHE_DIR=/opt/opforge-triton \
OPFORGE_CACHE_DIR=/opt/opforge-kernels \
python3 tools/build_aot.py
# 运行期把同一个目录挂进去 —— 装机即热启动
export TRITON_CACHE_DIR=/opt/opforge-triton
export OPFORGE_CACHE_DIR=/opt/opforge-kernels两个目录分工不同:Triton 的产物只能住在它自己的缓存里(它没有 export API),
CuTeDSL 的产物由 opforge 自己管(export_to_c 出的 object + manifest,路径
<OPFORGE_CACHE_DIR>/v1/<target_fp>/<ns>/<source_fp>/<entry_id>.artifact)。
于是预编译跑两遍时,报告能直接体现"第二遍没有再编译":
3 kernels, 3 compiled, 0 loaded, ok=True # 冷启动:全编
rms_norm triton HIGH_PRECISION ok
scale cutedsl HIGH_PRECISION ok
warp_reduce_sum cpp DETERMINISTIC ok
3 kernels, 2 compiled, 1 loaded, ok=True # 再跑一遍:CuTeDSL 命中磁盘
rms_norm triton HIGH_PRECISION ok
scale cutedsl HIGH_PRECISION loaded
warp_reduce_sum cpp DETERMINISTIC ok
算子必须声明要编哪些形状,否则构建期无从下手:Triton 用
declare_triton(..., prewarm=...),CuTeDSL 用 declare_cutedsl(..., examples=...)
(后者的样例输入是 cute.compile 的必需参数,不是可选优化)。
运行期想断言"全程不再编译",配合 KernelCache.seal() 即可 —— 封编译之后任何未命中的
构建都会抛 CompilationSealed,所以"第二遍能跑通"本身就证明了没重编译。
python3 -m build --wheel # 或 python3 -m pip wheel . -w dist --no-deps
# 产物:dist/opforge-0.0.1.dev0-py3-none-any.whl在有 GPU 的机器上跑 Triton 路径前,先 source tools/gpu_env.sh。它做两件事,都是
容器环境的坑,不是本项目的设计:
LD_PRELOAD=/usr/local/nvidia/lib64/libcuda.so.1—— 容器里/usr/lib/x86_64-linux-gnu/libcuda.so指向一个 0 字节桩,真驱动在/usr/local/nvidia/lib64/。torch 用 ctypes 默认的RTLD_LOCAL加载它,符号不进 全局作用域,于是 Triton 的 driver 扩展加载时报undefined symbol: cuModuleGetFunction,连torch.compile的 inductor 后端也一起用不了。PYTHONNOUSERSITE=1—— 避开~/.local下一份损坏的 torch,它会遮蔽正常的版本。
CuTeDSL 另有一条限制:编译只要求工具链(CUTE_DSL_ARCH=sm_90a 需显式指定),
但执行要求驱动版本不旧于 cuda-python 的 CUDA 版本;驱动更旧时 launch 会报
cudaErrorInsufficientDriver —— 所以在那样的机器上,CuTeDSL 的
AOT 编译 + 跨进程加载那半照常可用,只有 launch 跑不了,对应测试自动 skip。
跨进程复用的验证方式与"能不能 launch"无关:构建进程用 export_to_c 落盘,消费进程用
已封编译(KernelCache.seal())的缓存去加载 —— 能拿回产物就证明没有回退到
cute.compile(否则会抛 CompilationSealed)。
docs/
design.md 为什么:取舍与动机
architecture.md 怎么做:分层接口与不变量
src/opforge/
contract/ 数值契约、变体解析、前置条件
kernel/ @kernel 装饰器、KernelSpec、调用流水线
cache/ 两级指纹、manifest、封编译
launch/ 按签名缓存启动闭包
lang/ 后端协议(python / triton / cutedsl / cpp)、逃生口登记
kernels/ 算子实现(norm/、elementwise/、reduction/)
training/ L3 训练侧适配
inference/ L3 推理侧适配
aot.py AOT 预编译入口
tests/ 精度、契约、缓存、后端、AOT、适配层
tools/ 测试运行器、不变量检查、AOT 构建、GPU 环境脚本
.github/workflows/ CI(CPU)与 GPU 作业
| 文档 | 内容 |
|---|---|
docs/design.md |
为什么:分层架构、数值契约模型、逃生口策略、编译缓存与指纹、交付形态、CI 质量门、路线图 |
docs/architecture.md |
怎么做:L0–L4 各层接口 —— @kernel 完整签名、Contract 枚举与变体派生、Precondition 协议、缓存/启动层、架构不变量 I1–I9 |
CONTRIBUTING.md |
分层约定、如何加算子 / 加后端、两个容易踩的坑 |
Apache-2.0(见 LICENSE)。