This repo is an inference engine for spoken language identification, built on
top of GGML. It is a sibling of
transcribe.cpp: same
vendored ggml, same build and packaging mechanics, same numerical-validation
methodology, its own small C API (langid_*).
The project is intentionally conservative. It is a library surface, a model runtime, and a set of package artifacts that users embed in larger processes. Changes should be easy to review, portable across supported platforms, and maintainable after merge. Contributors who submit substantial features, model families, packaging changes, or backend work may be asked to help maintain those areas over time.
notes/PLAN.md (implementation spec) and notes/DECISIONS.md (decision log) are maintainer-local, not in the repository; they record the decisions
behind it and how to reverse them. Read both before proposing a structural
change.
AI tools may be used when a human contributor is driving the design, reviewing the output, and prepared to debug and maintain the result.
If AI meaningfully assisted with code or documentation:
- Disclose that usage in the PR.
- Manually review the generated or assisted changes before submission.
Do not use AI to write PR descriptions, issue reports, commit messages, project discussions, or replies to reviewers. Do not submit automated commits or pull requests. Obviously AI written PR descriptions will almost certainly be rejected.
The style below is adapted from the current llama.cpp and whisper.cpp contribution guidelines and made local so contributors do not need to chase external documents before writing code.
General rules:
- Avoid adding third-party dependencies, extra files, extra headers, or new build-time requirements unless they are necessary and justified.
- Always consider compatibility with supported operating systems, compilers, CPU architectures, accelerators, and package formats.
- Prefer plain, C-like C++ for runtime code. Use basic
forloops and simple helper functions over clever STL, template, or lambda-heavy constructs. - Keep abstractions narrow. Add one only when it removes real duplication, centralizes ownership/lifetime, or matches an established local pattern.
- New model families and backend work bring up CPU support first. The numerical-parity regime is F32 weights, CPU backend, one thread.
- Do not add new ggml operators,
ggml_typevalues, or backend behavior without CPU validation, benchmark or accuracy evidence where relevant, and an upstream plan for ggml/llama.cpp when appropriate. - Do not mix functional changes with broad reformatting or unrelated cleanup.
- Keep comments concise. Explain non-obvious invariants, ABI contracts, tensor layout, numerical choices, and portability traps. Do not preserve local task history or design backstory in source comments.
- Use ASCII in source files and comments unless a file or data format requires otherwise.
Formatting rules:
-
Use 4 spaces for indentation.
-
Put braces on the same line.
-
Use
void * ptrandint & apointer/reference spacing. -
Use vertical alignment when it improves readability and batch editing.
-
Clean up trailing whitespace.
-
Format with the pinned formatter rather than a system clang-format:
scripts/ci/clang-format.sh # format our tree in place scripts/ci/clang-format.sh --check # verify only, no changes
It wraps clang-format
22.1.5(fetched viauvx) against the root.clang-format, copied byte-identical from transcribe.cpp (itself a profile based on llama.cpp's). The version is pinned in the script so local output matches CI byte-for-byte; bump it there and reformat the tree in the same commit. -
Formatting scope is our C/C++ only. The vendored tree (
ggml/) and verbatim upstream copies (examples/common/dr_wav.h) are never reformatted. -
Do not reformat unrelated code as part of a behavior change. CI enforces this through the
clang-formatworkflow (.github/workflows/clang-format.yml), which gates our C/C++ against the pinned formatter.
Naming rules:
- Use
snake_casefor functions, variables, and type names. - Prefer names with the longest common prefix when grouping related values:
number_small,number_bigrather thansmall_number,big_number. - Public symbols are prefixed with
langid_orLANGID_. - Enum values are uppercase and prefixed by the enum name:
enum langid_backend_request {
LANGID_BACKEND_AUTO = 0,
LANGID_BACKEND_CPU = 1,
};- In public C/C++ headers, use sized integer types such as
int32_twhere ABI size matters.size_tis appropriate for allocation sizes and byte offsets. - In C++ code, omit optional
structandenumkeywords when they are not needed:
// OK
langid_model * model;
langid_backend_request backend;
// Not OK in C++ implementation code
struct langid_model * model;
enum langid_backend_request backend;- C/C++ filenames are lowercase. Prefer dashes for ordinary source files.
Family directories and family public headers may use the canonical family key,
including underscores, so API names and file paths stay aligned. Headers use
.h; source files use.cor.cpp. - Python library/module filenames are lowercase with underscores. Command-line scripts may use dashes when following the existing converter/tool convention.
API and tensor rules:
- Keep
include/langid.hfree of ggml includes. Callers should not need to include<ggml.h>. - Public structs crossing the ABI use the documented
struct_sizeconvention. - Tensors store data in row-major order. Dimension 0 is columns, dimension 1 is rows, and dimension 2 is matrices.
ggml_mul_mat(ctx, A, B)follows ggml's convention, not ordinary source-code reading order: treat existing ggml/llama.cpp graph patterns as the reference when adding graph code.
When this repo intentionally differs from llama.cpp/whisper.cpp, keep the difference local and documented near the API or subsystem that needs it. The main accepted internal exception is that the multi-family runtime may use C++ ownership helpers or narrow virtual bases where they centralize lifetime and avoid duplicating model/context teardown logic.
Required before merge:
| Gate | Command / owner | Requirement |
|---|---|---|
| Formatting | scripts/ci/clang-format.sh --check (CI: clang-format workflow) |
All our C/C++ matches the pinned clang-format |
| Default tests | ctest --test-dir build (CI: native-ci cpp-tests, Linux + macOS) |
All enabled default tests pass |
| Sanitizers | -DLANGID_SANITIZE=ON + ctest (CI: native-ci cpp-tests-sanitized) |
Clean under ASan + UBSan |
| Install + link manifest | cmake --install build --prefix <p> then python3 scripts/ci/link_smoke.py --prefix <p> (CI: native-ci, both static and shared) |
An external C consumer links and runs from lib/langid-link.json alone |
| Numerical validation | uv run scripts/validate.py all --family <f> |
Every contract tensor inside tests/tolerances/<f>.json and identical predictions on all manifest cases |
| Dataset gate | uv run scripts/eval/compare.py <ref>.jsonl <cpp>.jsonl |
C++ F32 agrees with the reference on >= 99.9% of decisions, differences only on near-ties (last measured: 12000/12000) |
The last two need a converted GGUF and the pinned reference environment, so
they run locally, not in CI. Run them for any change that can move numerics:
the graph, the front end, the converter, the loader, or the ggml pin. See
docs/validation.md for what each proves and how to re-run it.
Real-model smokes (ctest --test-dir build -R <family> with
-DLANGID_BUILD_REAL_MODEL_TESTS=ON and LANGID_<FAMILY>_GGUF set) are also
local-only, for the same reason; they exit 77 (skip) without a model.
- Process question: open an issue tagged
process. - License or semantic op mapping: escalate to the maintainer; do not guess.