Lithon is a zero-dependency, dual-tier execution engine for a statically verified, Python-flavored language built for bare-metal performance.
It looks like Python. It reads like Python. But under the hood, Lithon makes a strict promise: every variable has a proven, fixed type before execution begins. No dynamic dispatch, no boxing integers on the heap, and no interpreter overhead on critical paths.
Rather than linking heavy compiler frameworks like LLVM or Cranelift, Lithon’s custom C++ engine compiles typed AST blocks directly into raw, guard-free x86-64 machine instructions and executes them in memory via native OS page allocation (mmap / VirtualAlloc).
The Lithon Promise: If a type flow cannot be mathematically proven safe and static, Lithon will refuse to compile it—ensuring slow execution paths never exist in your binary.
Lithon has no third-party dependencies — no LLVM, no runtime library, nothing
to pip install to make the engine work. It does need a toolchain and a
language, though, so the "zero dependencies" claim is about the engine's
inputs, not its build:
- CMake ≥ 3.20 and a C++20 compiler (tested with GCC 12 and Clang 17)
- Python ≥ 3.10 — only for the compile step, which turns
.pyinto IR. Once IR exists, the native program never touches CPython: no objects, no refcounting, no GC.
# 1. Clone
git clone git@github.com:Project-Lithon/lithon.git
cd lithon
# 2. Build the C++ engine. There is no Makefile; CMake drives everything.
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
# 3. Compile a .py program to IR, then run it on the native tier.
# The frontend is the only step that involves Python.
python3 src/frontend/frontend.py tests/programs/float.py > /tmp/float.ir
./build/tier_runner /tmp/float.ir --strict[tier1] native
0.3333333333333333
0.30000000000000004
1.0
-0.0
1.0000000000000002
...
The [tier1] native line goes to stderr; the program's own output goes to
stdout. --strict means "native only, refuse rather than fall back", so a
successful exit status is proof the code was really emitted and executed —
that is the flag to use when testing, because --auto can hide a total
fallback to the interpreter behind correct output.
pip install -e . adds a lithon wrapper that does the two-step dance for you.
It is optional — the engine works without it — and it installs into whichever
interpreter you point pip at, so check lithon --version agrees with the
python3 you expect.
pip install -e .
lithon tests/programs/float.py # auto: native when provably safe
lithon tests/programs/float.py --strict # native only, rc=3 if refused
lithon tests/programs/float.py --ir # print the IR and stop
lithon tests/programs/float.py -v # show which tier ran, and whyinit.py finds the frontend and build/tier_runner relative to the repo, and
honours LITHON_HOME and LITHON_RUNNER if you need to point it elsewhere.
Two layouts, one of them dead. The tree carries a package layout (
__init__.py,__main__.py, where_ROOTis the repo's parent) alongside the flat layout actually installed (init.py,main.py, where_ROOTis the repo root).pyproject.tomlwires the flat one, so the package pair is dead code andimport lithonraisesModuleNotFoundError. Thelithoncommand works; thelithonmodule does not yet. Worth collapsing to a reallithon/package.
Lithon operates across four precise architectural layers, using each language exclusively where it excels:
| Layer | Language | Primary Responsibility |
|---|---|---|
| Frontend | Python | Expressive syntax, user logic, AST generation target |
| Engine | C++ (C++20) | AST parser, static flow verifier, type checker, memory manager |
| Bridge | C | SysV & Win64 ABI calling convention alignment |
| Backend | x64 Assembly | Bare-metal machine code generation (hand-encoded opcodes) |
graph TD
A[Python Source Code] --> B[Lithon Static Flow Verifier]
B -->|Verified Static Types| C[Tier-1: Hand-Rolled x64 JIT Emitter]
B -->|Dynamic Operations / Unhandled I/O| D[Tier-0: C++ Fallback Interpreter]
C --> E[mmap PROT_EXEC Executable Memory Buffer]
E --> F[Direct CPU Execution]
D --> G[C++ Native Execution]
| Feature | CPython | Cython / mypyc | PyPy | Lithon |
|---|---|---|---|---|
| Primary Target | Dynamic Bytecode | C Extension Source | Tracing JIT | Bare-Metal x64 Machine Code |
| Typing System | Dynamic | Optional / Annotated | Dynamic Tracing | Mandatory Static |
| External Dependencies | CPython Runtime | C/C++ Compiler | Heavy JIT Runtime | Zero (Self-Contained Engine) |
| Startup Overhead | High | Medium | Very High (Warmup) | Instant (< 1ms) |
| Memory Allocation | Boxed Heap Objects | Partial Unboxing | Traced Heap | Raw Stack Registers / Unboxed |
| Memory Page Control | None | OS Standard | Runtime Managed | Direct mmap / VirtualAlloc |
Lithon is purpose-built for low-latency tasks where Python traditionally relies on external C/C++ wrappers:
- 🔢 Numeric & Algorithmic Loops: High-throughput mathematical sequences, signal processing, and matrix transformations.
- ⚡ Quantitative Systems & Trading Logic: Microsecond-level strategy execution without garbage collector pauses or JIT warmup latency.
- 🎮 Performance-Sensitive Engine Tooling: High-frequency game logic, spatial partitioning, and real-time data pipelines.
- 🛡️ Predictable Micro-Utilities: Deterministic, fixed-schema processing where memory and execution bounds must be guaranteed prior to runtime.
Phase I (Current) Phase II (Next) Phase III (Future)
┌──────────────────────┐ ┌──────────────────────┐ ┌──────────────────────┐
│ Dual-Tier JIT Engine │───►│ AOT Binary Synthesis │───►│ Zero-Copy C-FFI │
│ In-memory mmap x64 │ │ Standalone <10KB ELF │ │ Direct Syscall Ops │
└──────────────────────┘ └──────────────────────┘ └──────────────────────┘
-
Phase I: Dual-Tier JIT Architecture (V1)
- Hand-rolled x86-64 machine code emitter in pure C++20.
- Direct execution via
mmap(PROT_READ | PROT_EXEC) andVirtualAlloc. - Static type flow analysis with deterministic Tier-0 interpreter fallback.
- Full SysV and Windows x64 ABI compliance for C-level call compatibility.
-
Phase II: Ahead-Of-Time (AOT) Binary Compiler (V2)
- Direct ELF64 (Linux) and PE32+ (Windows) header synthesis.
- Compile Python scripts into standalone native executables (
.bin/.exe) under 10 KB. - Complete decoupling from the Lithon interpreter engine for standalone deployment.
-
Phase III: Systems & Hardware Integration (V3)
- Direct
syscall(0x0F 0x05) instruction emission from Python syntax. - Zero-copy C pointer exposure and raw memory array mutation.
- AVX2 / SIMD vectorization for parallel list processing.
- Direct
Current status, Phase I (Dual-Tier JIT). "Shipped" means enforced by a test that fails when it regresses — not merely present.
| Area | Status | Evidence |
|---|---|---|
| Hand-rolled x86-64 encoder | Shipped | encoder_test, branch_test, stack_test, tools/check_encoder_vs_as.py |
| SysV x64 ABI | Shipped | src/jit/jit_abi.h; check_stack_alignment.py proves callee-saved + 16-byte stack alignment at every call/ret |
| Windows x64 ABI | Implemented, not yet verified | #if defined(_WIN32) in jit_abi.h; the audit only exercises the host ABI, so the Win64 path has no test evidence yet |
| Tier-0 interpreter fallback | Shipped | tier_runner --auto falls back on unprovable output (print_guard.h) |
| Static type flow verifier | Shipped | tools/typecheck.py, run_typed_regression.py (12/12) |
| Liveness + register allocation | Shipped | liveness_test, regalloc_test |
| IR text format | Shipped | src/ir/text_parser.cpp — no Python dependency in the engine |
| Function calls, recursion, TCO | Shipped | compile_module_call_test, fib_test; self-tail-calls become loops, O(1) stack |
Every optimization is independently switchable, so its effect can be measured rather than assumed.
| Change | Lives in | Isolated measurement vs. HEAD |
|---|---|---|
strength_reduce_multiplies — invariant × induction-variable → repeated add |
optimize.h |
nested 1.062× |
| Callee-saved borrowing — temps live across a call borrow a callee-saved register instead of the stack | register_alloc.h |
fib 1.063× |
Shared virtual-temp liveness — one VirtualTemps set, excluded before live ranges are computed |
liveness.h |
correctness, not speed |
fold_constants, eliminate_dead_code, convert_self_tail_calls |
optimize.h |
bundled, not isolated |
Measured on an Intel i3-3110M (Sandy Bridge), 12 interleaved rounds, CPU-time clock, one pinned core. See Testing to reproduce them yourself.
Phihas no emitter. The IR can express it; the encoder cannot yet. It is the only missing opcode, and is needed forif-as-expression lowering once both arms must merge without a stack round-trip.- Net: 20 of the 21 IR opcodes are emitted. The one missing one is
Phi.
This is the one place where Lithon deliberately does not follow Python, so
it is worth stating plainly rather than leaving to a comment in ir.h.
print(-7 % 3) # Lithon: -1 CPython: 2
print(7 % -3) # Lithon: 1 CPython: -2% truncates toward zero and takes the sign of the dividend, which is what
C, Rust, Java and every other compiled language do. CPython floors instead, so
its remainder has the sign of the divisor. Both engines here implement the
truncating rule, which is what makes the tier diff a meaningful check rather
than two engines agreeing on a shared mistake.
The reason is the one C gives: a remainder never leaves the domain of its
operands, so Mod is typed like Mul (int iff both operands are int) rather
than like Div, which must widen to float because a quotient generally is not
an integer. Typing it as Div would make 7 % 3 a float and lose the point.
Consequences worth knowing, all covered by tests:
- A zero divisor traps on both engines, like
Div, with its own message (modulo by zero) so the two are distinguishable in a diff. INT64_MIN % -1is0. It is the one input that makes hardwareidivraise#DE, since the quotient would be 2⁶³, so the divisor is tested up front and the whole division collapses to a zero.- Float
Modhas no SSE2 instruction —fmodis a libm call, and calling one per modulo would be far more expensive than the interpreter this is meant to be replacing. It is computed asa - n*bforn = trunc(a/b), the definition C uses. The interesting case isn == 0, which happens exactly when|a| < |b|, and it is the only case where the multiply can go wrong: IEEE makes0 * infa NaN, butn*bis 0 for everybwhennis 0, so C'sfmod(1.0, inf)is1.0. The guard that skips the multiply has to distinguish a real zero quotient from a NaN one using the sameZF AND !PFshape as the zero-divisor check, becausecomisdsets ZF for an unordered compare too. - A zero remainder is normalized to
+0.0to match CPython, which does not preserve the dividend's sign for a zero remainder. C'sfmod(-4.0, 2.0)is-0.0; Lithon prints0.0. This is a conscious divergence, chosen so the float and integer paths agree with each other.
Integer Mod is also strength-reduced where it is exact: a constant divisor
that is a power of two becomes a mask plus a sign fixup, since -7 & 3 is 1
and not the -3 that -7 % 4 has to return.
float is implemented end-to-end: ConstFloat, load/store, Add/Sub/Mul/
Div/Mod, the three comparisons, and native print(). Both tiers now run
float.ir, mixed_numeric.ir and comparison.ir natively, with
run_tier_diff.py reporting 34/34 native and zero interpreter fallbacks.
The hard part was not the arithmetic — SSE2 is straightforward once the encoding is right — it was making the two engines agree byte for byte, since that is the property everything else is measured against:
- Formatting is one function, called by both.
host_format_double()infloat_runtime.his what emitted code calls and what the interpreter calls. CPython's rule is the shortest string that round-trips, so neither"%f"(which prints3.500000) nor"%.17g"(which prints0.10000000000000001) is acceptable.%.*gis also wrong in a way that is easy to miss: it chooses exponent notation based on the precision it needed, whereas CPython's threshold is absolute — decimal exponent below −4 or above 16. That is why924966630.0must print in full, not as9.2499663e+08. Divby zero traps on both engines, matching Python'sZeroDivisionErrorrather than IEEEinf/nan. The check cannot be a single branch:comisdsets ZF, PF and CF together when the operands are unordered, so ZF alone cannot separate "equal" from "NaN". The emitted code computesZF AND !PFin a GP register instead. A NaN divisor must not trap — Python propagates — and-0.0must, since it compares equal to0.0.- NaN compares false against everything, including itself. The same
unordered-flag problem applies to
Lt/Gt/Eq, and is excluded withsetcc+AND setnprather than a parity branch:0F 9Ais a byte-for-byte collision betweenjp rel32andsetp r/m8, so a parityJccis not encodable here. - A float live across a call spills. Every XMM in the temp pool is
caller-saved on both ABIs, and
host_format_doubleis an ordinary C function that clobbers all of them, so leaving a float in one across acall printsilently corrupts it. - The guard refuses a variable stored both an
intand afloat. Its kind joins toUnknown, and codegen only asksis_float_value— so it would lower the arithmetic as integer operations over a double's bit pattern. Printing the resultingboolhides this, because a comparison is alwaysbooland so always passes the print check.
These are pinned by float_format_test (CPython repr transcribed by hand),
the SSE2 byte-exact assertions in encoder_test, and a dedicated
tools/fuzz_diff.py --floats mode — the general fuzzer annotates every variable
int[64] and so never reached any of it.
- Arguments are capped at 2 per function and per call.
- Branchy loop bodies are not unrolled. The unroller is implemented,
correct, and fuzzed — but it measured 1.10× slower on an if/else loop
(1.05× with a heavier body), because a diamond's if/else test is irreducible
and unrolling only inflates the loop ~2×. It is therefore opt-in behind
--unroll-diamonds, not deleted. - No AOT backend. Everything runs in-process; there is no
.bin/.exeemit. - x86-64 only. No ARM64 backend.
Phi— the last unemitted opcode, needed forif-as-expression lowering once both arms must merge without a stack round-trip.- Lift the 2-argument cap — most remaining test programs are blocked on it.
- ARM64 backend — the genuinely arch-agnostic layers are
ir/,liveness.h, andoptimize.h(they name no registers at all).register_alloc.hnames registers only viaabi::kPromotionPool. The x86-specific surface isx86_encoder.hplus the emit calls incompile_function.h.float_runtime.his in that last group only in the sense that its formatter is shared — the arithmetic and itsZF AND !PFzero test are not. - AOT emit — Phase II below.
Note: Phase I is a work in progress. The engine is fast and well-tested on the subset it supports, and it refuses what it cannot prove — that refusal is the feature, not a workaround.
Work outward and stop when you are satisfied. Each layer is roughly an order of magnitude slower than the one above it.
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)If you are on a machine with a small /tmp (a 100 MB tmpfs is enough to fail
this), point the compiler's scratch space somewhere roomier:
TMPDIR=/path/to/scratch cmake --build build -j$(nproc)ctest --test-dir build --output-on-failure # 17/17The tests are not all the same kind, and it is worth knowing which is which:
liveness_test,regalloc_test— pure analysis tests. They build IR and call the pass directly, and never emit a byte. A green run here means the data structures are right, not that any machine code ran.compile_function_*,compile_module_*,optimize_lsr_test— compile a module to machine code, mmap it, and call it through a function pointer, checking the returned values. These are the tests that would catch a bad encoding.encoder_test,stack_test,branch_test,print_guard_*— the x86/ABI layer underneath, tested in isolation.float_format_test— pure computation, no JIT, no interpreter. Pinshost_format_doubleagainst CPython'sreprwith expectations transcribed by hand rather than generated from the code under test, since generating them would only prove the formatter agrees with itself.compile_module_float_test— needs real machine code, so it is POSIX-only and usesfork: a float live across a call, a NaN divisor that must propagate, and0.0/-0.0divisors that must trap. The last two run in a child process because the trap handler callsexit(1).
bash tools/verify_all.sh # ABI, stack alignment, callee-saved audit
python3 tools/run_regression.py # 12/12 untyped programs
python3 tools/run_typed_regression.py # 12/12 typed programs
python3 tools/run_tier_diff.py # 34/34run_tier_diff.py is the highest-value of the four. It runs every program
through both tiers and requires byte-identical stdout, and it reports which
tier actually ran — so a green run cannot hide "everything silently fell back
to the interpreter".
python3 tools/fuzz_diff.py --count 300 # general programs
python3 tools/fuzz_diff.py --count 300 --lsr # strength-reduction shapes
python3 tools/fuzz_diff.py --count 300 --diamond # diamond-unroll shapes
python3 tools/fuzz_diff.py --count 300 --floats # int/float mixes, div, calls
python3 tools/fuzz_diff.py --mod --count 300 # modulo, any signs
python3 tools/fuzz_diff.py --mod-negatives --count 300 # modulo, negative operandsThe JIT's output is compared against the interpreter, and the interpreter's
against CPython — reported separately, because Lithon deliberately diverges
from CPython for loop variables (v == n after a loop, not n - 1), so a
CPython disagreement is not by itself a bug in the JIT. The --lsr and
--diamond modes exist because the general
generator almost never reaches those two passes; without them they would be
essentially untested. --floats exists because the general generator annotates
every variable int[64] and so emits no const_f64 at all — that mode found a
real miscompile (an Unknown-kind operand lowered as integer arithmetic) that
four general modes had never approached.
--mod-negatives is the mode to reach for when changing modulo, with one caveat:
because Lithon's % truncates and CPython's floors, most of its CPython
disagreements are expected language gaps rather than bugs, and the mode reports
them separately. The number that must stay at zero is the interpreter-vs-JIT
mismatch count. Neither modulo mode generates infinities or NaNs, so the
0 * inf class of bug needs the adversarial run_tier_diff.py case instead.
Mismatches are written to
fuzz_failures/, minimised,
and printed.
Often the most convincing check, because it needs no timing and no baseline:
./build/lithon_jit nested_loop.ir --dump-code /tmp/n.bin
objdump -D -b binary -mi386:x86-64 -M intel /tmp/n.bin
# Did strength reduction fire? nested_loop has one multiply per inner iteration.
objdump -D -b binary -mi386:x86-64 -M intel /tmp/n.bin | grep -c imul # 0 = fired
# Same program with the pass disabled, to prove the above was the pass's doing.
./build/lithon_jit nested_loop.ir --no-lsr --dump-code /tmp/n2.bin
objdump -D -b binary -mi386:x86-64 -M intel /tmp/n2.bin | grep -c imul # 4Every optimization is toggleable, so a claim can be checked rather than taken on trust. This is how the numbers in the roadmap above were established:
./build/lithon_jit prog.ir --no-opt # constant folding + DCE
./build/lithon_jit prog.ir --no-promote # register promotion
./build/lithon_jit prog.ir --no-rotate # loop rotation
./build/lithon_jit prog.ir --unroll=1 # all loop unrolling off
./build/lithon_jit prog.ir --no-lsr # strength reduction
./build/lithon_jit prog.ir --unroll-diamonds # opt in to diamond unrolling
./build/lithon_jit # full flag listtier_runner takes --no-lsr and --unroll-diamonds too, which is what the
fuzzer uses; the other four are currently lithon_jit only.
Use the existing harness — it ships the four official workloads plus three stress cases, reports a noise column, and warns that results are unreliable above ~15% noise.
python3 tools/native_bench.py --runs 30 --pin 2To compare two states, write a baseline first and compare against it:
python3 tools/native_bench.py --runs 30 --json /tmp/before.json
# ...change something, rebuild...
python3 tools/native_bench.py --runs 30 --compare /tmp/before.jsonTwo traps. Always pass
--pin <cpu>. Unpinned, one run showed 26% noise onbranchyand the numbers were unusable; pinned, the same comparison was stable to ~0.01 ms. And do not point--compareatbenchmarks/results/*.json— those were written bytools/bench.py, which has a different schema and different workload names, so thevs beforecolumn comes out silently empty rather than erroring. Both sides must come fromnative_bench.py.
Read
min, notmedian. Noise from other processes, frequency scaling and cache state only ever add time, so the minimum is the least contaminated estimate. Differences under ~5% are not meaningful on a shared machine.
The test suite covers the subset of the language the engine supports today.
Floats, Div and Mod now work, so they have moved out of the gaps list; what
remains unimplemented is Phi and support for more than two arguments, and
those are listed above precisely so a green run
is not mistaken for a complete one. If you add support for one of them, the
honest next step is to move it out of that list.
Two more limits worth stating plainly, because a passing run can obscure both:
- Opcode coverage is not operand coverage. 20 of 21 opcodes are emitted and
each is exercised through
tier_runner --strict, so a pass proves the opcode was genuinely executed natively rather than fallen back. It does not prove every operand shape is right — that is whatencoder_test's byte-exact assertions and the fuzz modes are for.gtandnot, for instance, have a single native use each in the checked-in.ircorpus.Modis a standing example of why the two are different: the general fuzzer only ever emits finite constants, so it cannot generate1.0 % inf, which was returning NaN natively while the interpreter was correct. That case is pinned by an adversarialrun_tier_diff.pyentry that requires the native tier, so a future guard change cannot make it silently fall back and hide the bug. - The interpreter is an oracle, not a specification. Where Lithon and
CPython disagree,
run_tier_diff.pyreports it separately as a language gap rather than a JIT bug — loop variables are one known case, deliberate. That is correct for this project, but it does mean the suite cannot catch a bug where both engines share the same wrong idea.
| Benchmark | Lithon JIT | Reference | Speedup | Result |
|---|---|---|---|---|
| fib(30) | 6.9917 ms | 1459.0371 ms | 208.7× | 832040 |
Status: PASS
Last updated by Lithon Reporter Mamba.
Lithon is released under the MIT License.

