Add offline Jina v5 nano retrieval embeddings - #13
Merged
Merged
Conversation
Support the merged nano retrieval checkpoint with last-token pooling and model-specific prefixes while preserving the Nemotron default. Document local setup and add offline mocked adapter, limit, identity, and CLI tests.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
--embedding-adapter jina-v5-nanofor the mergedjinaai/jina-embeddings-v5-text-nano-retrievalcheckpoint. The Nemotron default and Jina v3.5 reranker remain available and unchanged.local_files_only=TrueQuery:/Document:prefixes, right-padding-aware last-token pooling, and FP32 L2-normalized outputThis is a separate feature branch based on main after CI fix #12 (fa34141). It is for local evaluation and is not ready to merge on the strength of mocked tests alone.
Verification
GitHub Actions passed on Ubuntu and Windows for
afeb9709a4449b19a2bd028096f7399b2db56206. Both ran the 9 new adapter tests with no skips. Full suite: 1,552 tests on each OS, 16 skips on Ubuntu and 44 on Windows; no failures/errors. All 7 native Windows lifecycle tests passed.python -m unittest tests.test_embedding_models tests.test_embedding_retrieval: 51 tests passed, including 9 new adapter testspython -m tests: 1,552 tests, 16 skips, 0 failures/errors on Linuxpython -m compileall -q rightmemory tests: passedgit diff --check: passedAdapter tests use dependency-free numeric doubles and mocked checkpoint loading. They cover offline loader arguments, CPU/CUDA precision choice, mixed-length pooling, prefixes, context boundaries, cache identity, CLI registration, service capability reporting, and the preserved Nemotron default. No model weights were downloaded. Actual Torch/Transformers checkpoint loading, GPU behavior, retrieval quality, latency, and VRAM usage remain to be tested on the model host.
Local evaluation
From a clean checkout:
Use your own downloaded snapshot of
jinaai/jina-embeddings-v5-text-nano-retrieval, includingmodel.safetensors,config.json,tokenizer.json,tokenizer_config.json,configuration_eurobert.py, andmodeling_eurobert.pyfrom the same revision. Point the service at the snapshot directory, not a GGUF/ONNX file or the base multi-task model. The existing downloaded Jina reranker v3.5 snapshot is still required.In a second terminal:
Check that
embedding_modelstarts withjina-v5-nano:anddimensionsis768. Keep or add this section in the Memory root configuration, preserving the existing agent/model fallback:Then run your normal
rightmemory retrieve --session <id> "<complete need>"command for the configured Memory root. First retrieval after switching models rebuilds the derived embedding cache. The service loads both models and does not download either one.