Skip to content

Latest commit

 

History

History
166 lines (133 loc) · 7.91 KB

File metadata and controls

166 lines (133 loc) · 7.91 KB

Contributing to langid.cpp

This repo is an inference engine for spoken language identification, built on top of GGML. It is a sibling of transcribe.cpp: same vendored ggml, same build and packaging mechanics, same numerical-validation methodology, its own small C API (langid_*).

The project is intentionally conservative. It is a library surface, a model runtime, and a set of package artifacts that users embed in larger processes. Changes should be easy to review, portable across supported platforms, and maintainable after merge. Contributors who submit substantial features, model families, packaging changes, or backend work may be asked to help maintain those areas over time.

notes/PLAN.md (implementation spec) and notes/DECISIONS.md (decision log) are maintainer-local, not in the repository; they record the decisions behind it and how to reverse them. Read both before proposing a structural change.

AI-assisted contributions

AI tools may be used when a human contributor is driving the design, reviewing the output, and prepared to debug and maintain the result.

If AI meaningfully assisted with code or documentation:

  1. Disclose that usage in the PR.
  2. Manually review the generated or assisted changes before submission.

Do not use AI to write PR descriptions, issue reports, commit messages, project discussions, or replies to reviewers. Do not submit automated commits or pull requests. Obviously AI written PR descriptions will almost certainly be rejected.

Coding style

The style below is adapted from the current llama.cpp and whisper.cpp contribution guidelines and made local so contributors do not need to chase external documents before writing code.

General rules:

  • Avoid adding third-party dependencies, extra files, extra headers, or new build-time requirements unless they are necessary and justified.
  • Always consider compatibility with supported operating systems, compilers, CPU architectures, accelerators, and package formats.
  • Prefer plain, C-like C++ for runtime code. Use basic for loops and simple helper functions over clever STL, template, or lambda-heavy constructs.
  • Keep abstractions narrow. Add one only when it removes real duplication, centralizes ownership/lifetime, or matches an established local pattern.
  • New model families and backend work bring up CPU support first. The numerical-parity regime is F32 weights, CPU backend, one thread.
  • Do not add new ggml operators, ggml_type values, or backend behavior without CPU validation, benchmark or accuracy evidence where relevant, and an upstream plan for ggml/llama.cpp when appropriate.
  • Do not mix functional changes with broad reformatting or unrelated cleanup.
  • Keep comments concise. Explain non-obvious invariants, ABI contracts, tensor layout, numerical choices, and portability traps. Do not preserve local task history or design backstory in source comments.
  • Use ASCII in source files and comments unless a file or data format requires otherwise.

Formatting rules:

  • Use 4 spaces for indentation.

  • Put braces on the same line.

  • Use void * ptr and int & a pointer/reference spacing.

  • Use vertical alignment when it improves readability and batch editing.

  • Clean up trailing whitespace.

  • Format with the pinned formatter rather than a system clang-format:

    scripts/ci/clang-format.sh            # format our tree in place
    scripts/ci/clang-format.sh --check    # verify only, no changes

    It wraps clang-format 22.1.5 (fetched via uvx) against the root .clang-format, copied byte-identical from transcribe.cpp (itself a profile based on llama.cpp's). The version is pinned in the script so local output matches CI byte-for-byte; bump it there and reformat the tree in the same commit.

  • Formatting scope is our C/C++ only. The vendored tree (ggml/) and verbatim upstream copies (examples/common/dr_wav.h) are never reformatted.

  • Do not reformat unrelated code as part of a behavior change. CI enforces this through the clang-format workflow (.github/workflows/clang-format.yml), which gates our C/C++ against the pinned formatter.

Naming rules:

  • Use snake_case for functions, variables, and type names.
  • Prefer names with the longest common prefix when grouping related values: number_small, number_big rather than small_number, big_number.
  • Public symbols are prefixed with langid_ or LANGID_.
  • Enum values are uppercase and prefixed by the enum name:
enum langid_backend_request {
    LANGID_BACKEND_AUTO = 0,
    LANGID_BACKEND_CPU  = 1,
};
  • In public C/C++ headers, use sized integer types such as int32_t where ABI size matters. size_t is appropriate for allocation sizes and byte offsets.
  • In C++ code, omit optional struct and enum keywords when they are not needed:
// OK
langid_model * model;
langid_backend_request backend;

// Not OK in C++ implementation code
struct langid_model * model;
enum langid_backend_request backend;
  • C/C++ filenames are lowercase. Prefer dashes for ordinary source files. Family directories and family public headers may use the canonical family key, including underscores, so API names and file paths stay aligned. Headers use .h; source files use .c or .cpp.
  • Python library/module filenames are lowercase with underscores. Command-line scripts may use dashes when following the existing converter/tool convention.

API and tensor rules:

  • Keep include/langid.h free of ggml includes. Callers should not need to include <ggml.h>.
  • Public structs crossing the ABI use the documented struct_size convention.
  • Tensors store data in row-major order. Dimension 0 is columns, dimension 1 is rows, and dimension 2 is matrices.
  • ggml_mul_mat(ctx, A, B) follows ggml's convention, not ordinary source-code reading order: treat existing ggml/llama.cpp graph patterns as the reference when adding graph code.

When this repo intentionally differs from llama.cpp/whisper.cpp, keep the difference local and documented near the API or subsystem that needs it. The main accepted internal exception is that the multi-family runtime may use C++ ownership helpers or narrow virtual bases where they centralize lifetime and avoid duplicating model/context teardown logic.

Review gates

Required before merge:

Gate Command / owner Requirement
Formatting scripts/ci/clang-format.sh --check (CI: clang-format workflow) All our C/C++ matches the pinned clang-format
Default tests ctest --test-dir build (CI: native-ci cpp-tests, Linux + macOS) All enabled default tests pass
Sanitizers -DLANGID_SANITIZE=ON + ctest (CI: native-ci cpp-tests-sanitized) Clean under ASan + UBSan
Install + link manifest cmake --install build --prefix <p> then python3 scripts/ci/link_smoke.py --prefix <p> (CI: native-ci, both static and shared) An external C consumer links and runs from lib/langid-link.json alone
Numerical validation uv run scripts/validate.py all --family <f> Every contract tensor inside tests/tolerances/<f>.json and identical predictions on all manifest cases
Dataset gate uv run scripts/eval/compare.py <ref>.jsonl <cpp>.jsonl C++ F32 agrees with the reference on >= 99.9% of decisions, differences only on near-ties (last measured: 12000/12000)

The last two need a converted GGUF and the pinned reference environment, so they run locally, not in CI. Run them for any change that can move numerics: the graph, the front end, the converter, the loader, or the ggml pin. See docs/validation.md for what each proves and how to re-run it.

Real-model smokes (ctest --test-dir build -R <family> with -DLANGID_BUILD_REAL_MODEL_TESTS=ON and LANGID_<FAMILY>_GGUF set) are also local-only, for the same reason; they exit 77 (skip) without a model.

Getting help

  • Process question: open an issue tagged process.
  • License or semantic op mapping: escalate to the maintainer; do not guess.