Skip to content

Add offline Jina v5 nano retrieval embeddings - #13

Merged
RightL merged 1 commit into
mainfrom
feat/jina-nano-embeddings
Sep 30, 2026
Merged

RightL merged 1 commit into
mainfrom
feat/jina-nano-embeddings

Conversation

@RightL

@RightL RightL commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Summary

Adds --embedding-adapter jina-v5-nano for the merged jinaai/jina-embeddings-v5-text-nano-retrieval checkpoint. The Nemotron default and Jina v3.5 reranker remain available and unchanged.

  • Load the local Safetensors snapshot and its official EuroBert code with local_files_only=True
  • Use Query: / Document: prefixes, right-padding-aware last-token pooling, and FP32 L2-normalized output
  • Use BF16 on CUDA or FP32 on CPU; report 768 dimensions and reject inputs beyond the 8,192-token context without truncating
  • Include adapter settings and snapshot contents in cache identity, so changing embedding models rebuilds derived vectors
  • Document model files, service launch, existing client configuration, health check, cache behavior, and CC BY-NC 4.0 license

This is a separate feature branch based on main after CI fix #12 (fa34141). It is for local evaluation and is not ready to merge on the strength of mocked tests alone.

Verification

GitHub Actions passed on Ubuntu and Windows for afeb9709a4449b19a2bd028096f7399b2db56206. Both ran the 9 new adapter tests with no skips. Full suite: 1,552 tests on each OS, 16 skips on Ubuntu and 44 on Windows; no failures/errors. All 7 native Windows lifecycle tests passed.

  • python -m unittest tests.test_embedding_models tests.test_embedding_retrieval: 51 tests passed, including 9 new adapter tests
  • python -m tests: 1,552 tests, 16 skips, 0 failures/errors on Linux
  • python -m compileall -q rightmemory tests: passed
  • git diff --check: passed

Adapter tests use dependency-free numeric doubles and mocked checkpoint loading. They cover offline loader arguments, CPU/CUDA precision choice, mixed-length pooling, prefixes, context boundaries, cache identity, CLI registration, service capability reporting, and the preserved Nemotron default. No model weights were downloaded. Actual Torch/Transformers checkpoint loading, GPU behavior, retrieval quality, latency, and VRAM usage remain to be tested on the model host.

Local evaluation

From a clean checkout:

git fetch origin
git switch --track origin/feat/jina-nano-embeddings
python -m pip install -e ".[embedding-service]"

Use your own downloaded snapshot of jinaai/jina-embeddings-v5-text-nano-retrieval, including model.safetensors, config.json, tokenizer.json, tokenizer_config.json, configuration_eurobert.py, and modeling_eurobert.py from the same revision. Point the service at the snapshot directory, not a GGUF/ONNX file or the base multi-task model. The existing downloaded Jina reranker v3.5 snapshot is still required.

rightmemory embedding-service --embedding-adapter jina-v5-nano --embedding-model "C:/models/jina-embeddings-v5-text-nano-retrieval" --reranker-model "C:/models/jina-reranker-v3.5" --device cuda:0 --port 8766

In a second terminal:

Invoke-RestMethod http://127.0.0.1:8766/info

Check that embedding_model starts with jina-v5-nano: and dimensions is 768. Keep or add this section in the Memory root configuration, preserving the existing agent/model fallback:

[retrieve]
backend = "embedding"
max_output_chars = 100000

[retrieve.embedding]
url = "http://127.0.0.1:8766"
candidate_count = 40
result_count = 10
timeout_seconds = 60

Then run your normal rightmemory retrieve --session <id> "<complete need>" command for the configured Memory root. First retrieval after switching models rebuilds the derived embedding cache. The service loads both models and does not download either one.

Support the merged nano retrieval checkpoint with last-token pooling and model-specific prefixes while preserving the Nemotron default. Document local setup and add offline mocked adapter, limit, identity, and CLI tests.
@RightL
RightL merged commit 0cd5e86 into main Sep 30, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant