LexNeedle is a deterministic, Unicode-aware dictionary matcher for Python 3.12+, with zero runtime dependencies. Use it when exact terms, trustworthy original-text offsets, and predictable overlap handling matter.
uv add lexneedleFor the unreleased main branch:
uv add "git+https://github.com/crootman/LexNeedle"from lexneedle import Matcher
matcher = Matcher()
matcher.add("Panadol", value="paracetamol", metadata={"type": "brand"})
for match in matcher.find("Panadol 500mg tablets"):
print(match.text, match.value, match.start, match.end)This prints Panadol paracetamol 0 7. The reported offsets always index the
original Python string:
text = "Panadol 500mg tablets"
[match] = matcher.find(text)
assert text[match.start : match.end] == match.text- Case-insensitive matching uses locale-independent Unicode
casefold(); optional NFC, NFD, NFKC, and NFKD normalization preserves source offsets. - Results are deterministic for the same configuration, terms, input, and Python Unicode data. They do not depend on term insertion order.
- Explicit boundary and overlap policies make substring and collision behavior visible rather than accidental.
- JSON save/load is validated and versioned. New POSIX files are created with
mode
0o600, subject to the process umask.
The documentation sources cover getting started, redacting known identifiers, tagging documents, Unicode and offsets, boundaries and overlap strategies, replacement, JSON persistence, FlashText migration, API reference, and performance limits.
Run the library quality gate with uv sync --all-groups, then:
uv run ruff format --check .
uv run ruff check .
uv run ty check
uv run pytest
uv run sphinx-build -W --keep-going -b html docs docs/_build/html
uv buildSee the contribution guide for development guidance and the release checklist for the maintainer workflow.