diff --git a/README.md b/README.md index a06974f..cab6472 100644 --- a/README.md +++ b/README.md @@ -167,7 +167,7 @@ For a short recording script, see [docs/DEMO.md](docs/DEMO.md). - Multi-device memory continuity across laptops, desktops, agent clients, and project-specific roots. - Automatic orchestration plus independent explicit skills for map maintenance, Memory maintenance, guidance review, and Manager support in Codex App. - Two executor modes behind the same `rightmemory` CLI: standalone runtime or delegated Codex SDK/Claude Code CLI role execution. -- Source-authored retrieval output, selected by an agent or an optional Nemotron + Jina retrieval service. +- Source-authored retrieval output, selected by an agent or an optional embedding + reranking retrieval service. - One updater for Memory and Agent Corrections, transcript-review candidate extraction, and immutable candidate records for input-to-edit provenance. ## Install Options And Updates @@ -815,7 +815,7 @@ Configure `[sync-reconciler.model]` or `[sync-reconciler.agent_cli]` only if syn ### Embedding Retrieval -Embedding retrieval is an optional alternative to the agent-based retriever in either installation mode. One complete query goes to **Nemotron 3 Embed 1B**, which selects forty source entries; **Jina reranker v3.5** ranks those candidates and RightMemory returns up to ten. There is no relevance-score cutoff or additional LLM selection call. Related entries can be returned even when Memory does not contain the requested answer; the calling agent decides which entries apply. +Embedding retrieval is an optional alternative to the agent-based retriever in either installation mode. One complete query goes to the selected embedding model, **Nemotron 3 Embed 1B** (the default) or **Jina Embeddings v5 Text Nano Retrieval**, which selects forty source entries; **Jina reranker v3.5** ranks those candidates and RightMemory returns up to ten. There is no relevance-score cutoff or additional LLM selection call. Related entries can be returned even when Memory does not contain the requested answer; the calling agent decides which entries apply. Add these retrieval settings, keeping the existing `[retrieve.agent_cli]` or `[retrieve.model]` table as the fallback: @@ -852,11 +852,21 @@ rightmemory embedding-service \ --device cuda:0 --port 8766 ``` -Supply downloaded snapshots, for example from ModelScope. The service loads local files offline and keeps both models resident. The supported model interfaces are documented by [NVIDIA](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16) and [Jina](https://huggingface.co/jinaai/jina-reranker-v3.5); Jina's snapshot includes custom model code. The service serializes GPU requests and defaults to loopback. For direct access on a trusted private network, start it with `--host ` and configure clients with `url = "http://:8766"`. Encoding sends source passages to that configured host; reranking sends the query and candidate passages. +For the smaller Jina nano embedding model, select its adapter and the merged retrieval snapshot; keep the same reranker and client configuration: -Embedding and reranking adapters are selected independently. The defaults above support the tested Nemotron 1B / Jina v3.5 pair. To support another model family, implement the small `EmbeddingAdapter` or `RerankerAdapter` interface in `rightmemory/embedding_models.py` and register it in the corresponding adapter dictionary. An embedding adapter returns dense vectors for cosine search; a reranker returns every candidate index in ranked order. The adapters own model loading, input formatting, and token limits. The service reports embedding dimensions, batch size, and candidate capacity, which the client uses instead of model-specific constants. A compatible external service can also implement `/info`, `/embed`, and `/rerank`. +```bash +rightmemory embedding-service --embedding-adapter jina-v5-nano --embedding-model /path/to/jina-embeddings-v5-text-nano-retrieval --reranker-model /path/to/jina-reranker-v3.5/snapshot --device cuda:0 --port 8766 +``` + +Use [`jinaai/jina-embeddings-v5-text-nano-retrieval`](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano-retrieval), with the Safetensors weights, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `configuration_eurobert.py`, and `modeling_eurobert.py` from the same snapshot. This adapter loads the merged retrieval weights directly; the base multi-task nano checkpoint, adapter-only weights, GGUF, and ONNX files are not substitutes. It uses `Query: ` / `Document: ` prefixes, right padding, last non-padding token pooling, and L2 normalization. The output is 768 dimensions with an 8,192-token input limit including prefixes and tokenizer special tokens; longer inputs fail without truncation. CUDA inference uses BF16, while `--device cpu` uses FP32. No additional PEFT, sentence-transformers, or flash-attention dependency is needed for this adapter. The model is licensed [CC BY-NC 4.0](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano-retrieval#license); check Jina's terms for commercial use. + +After starting the service, check `http://127.0.0.1:8766/info` (PowerShell: `Invoke-RestMethod http://127.0.0.1:8766/info`). With nano selected, it reports an `embedding_model` identity starting with `jina-v5-nano:` and `dimensions = 768`. Then use the normal retrieve command above. Switching from Nemotron rebuilds the derived embedding cache on the next retrieval; the first query may take longer. This does not change the Memory source files. Real-checkpoint quality, latency, and GPU memory use should be evaluated on the model host. + +Supply downloaded snapshots, for example from ModelScope. The service does not download models; it loads local files offline and keeps both models resident. The supported model interfaces are documented by [NVIDIA](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16) and [Jina](https://huggingface.co/jinaai/jina-reranker-v3.5). Both Jina snapshots include custom model code, which the service executes locally; use trusted official snapshots. The service serializes GPU requests and defaults to loopback. For direct access on a trusted private network, start it with `--host ` and configure clients with `url = "http://:8766"`. Encoding sends source passages to that configured host; reranking sends the query and candidate passages. + +Embedding and reranking adapters are selected independently. The default pair remains Nemotron 1B / Jina v3.5; `jina-v5-nano` changes only the embedding model. To support another model family, implement the small `EmbeddingAdapter` or `RerankerAdapter` interface in `rightmemory/embedding_models.py` and register it in the corresponding adapter dictionary. An embedding adapter returns dense vectors for cosine search; a reranker returns every candidate index in ranked order. The adapters own model loading, input formatting, and token limits. The service reports embedding dimensions, batch size, and candidate capacity, which the client uses instead of model-specific constants. A compatible external service can also implement `/info`, `/embed`, and `/rerank`. -Embedding cache identities cover snapshot contents, adapter settings, and an adapter revision. The built-in adapter records its input prefixes, pooling, normalization, precision, and library versions; bump its revision for behavior changes not represented in those settings. Changing embedding identity rebuilds the vector cache; replacing only the reranker reuses it. Adding another model family requires an adapter rather than just pointing an existing adapter at unrelated weights. +Embedding cache identities cover snapshot contents, adapter settings, and an adapter revision. Each built-in embedding adapter records its input prefixes, pooling, normalization, precision, and library versions; bump its revision for behavior changes not represented in those settings. Changing embedding identity rebuilds the vector cache; replacing only the reranker reuses it. Adding another model family requires an adapter rather than just pointing an existing adapter at unrelated weights. An API key is optional when binding to loopback or an explicit private IP address (RFC 1918 IPv4 or IPv6 unique-local). Without a service key, requests are accepted without credentials. To enable authentication, set `RIGHTMEMORY_EMBEDDING_API_KEY` on the service and the matching `[retrieve.embedding].api_key` on clients. Public, wildcard, and other hostname bindings require a key. The service is started separately and is not launched by the normal installer or a retrieve call. diff --git a/rightmemory/embedding_models.py b/rightmemory/embedding_models.py index 3344665..f775afb 100644 --- a/rightmemory/embedding_models.py +++ b/rightmemory/embedding_models.py @@ -99,6 +99,49 @@ def encode(self, texts: list[str], *, kind: Literal["query", "passage"]) -> list return vectors.cpu().tolist() +class JinaV5NanoEmbedding: + """The merged jina-embeddings-v5-text-nano-retrieval checkpoint.""" + + max_batch_size = 8 + + def __init__(self, path: Path, *, device: str): + torch, transformers = _dependencies() + self.torch, self.device = torch, device + path = path.expanduser().resolve(strict=True) + dtype = torch.bfloat16 if device.startswith("cuda") else torch.float32 + self.tokenizer = transformers.AutoTokenizer.from_pretrained( + str(path), local_files_only=True, trust_remote_code=True, padding_side="right", + ) + self.model = transformers.AutoModel.from_pretrained( + str(path), local_files_only=True, use_safetensors=True, trust_remote_code=True, + dtype=dtype, attn_implementation="sdpa", + ).to(device).eval() + self.dimensions = self.model.config.hidden_size + self.max_tokens = min(8192, self.model.config.max_position_embeddings) + self.prefixes = {"query": "Query: ", "passage": "Document: "} + self.identity = adapter_identity(path, "jina-v5-nano", revision=1, settings={ + "checkpoint": "jinaai/jina-embeddings-v5-text-nano-retrieval", + "prefixes": self.prefixes, "pooling": "last-token", "normalization": "l2", + "padding_side": "right", "dtype": str(dtype), "attention": "sdpa", + "dimensions": self.dimensions, "max_tokens": self.max_tokens, + "torch": torch.__version__, "transformers": transformers.__version__, + }) + + def encode(self, texts: list[str], *, kind: Literal["query", "passage"]) -> list[list[float]]: + batch = self.tokenizer([self.prefixes[kind] + text for text in texts], + padding=True, truncation=False, return_tensors="pt") + if batch["input_ids"].shape[1] > self.max_tokens: + raise ValueError("embedding input exceeds the model context; shorten the query or source passage") + batch = batch.to(self.device) + with self.torch.inference_mode(): + hidden = self.model(**batch).last_hidden_state + # With right padding, each row ends at its own last non-padding token. + last_tokens = batch["attention_mask"].sum(dim=1) - 1 + pooled = hidden[self.torch.arange(hidden.shape[0], device=hidden.device), last_tokens] + vectors = self.torch.nn.functional.normalize(pooled.float(), p=2, dim=1) + return vectors.cpu().tolist() + + class Jina35Reranker: max_candidates = 125 max_query_tokens = 1024 @@ -143,5 +186,5 @@ def rerank(self, query: str, documents: list[str]) -> list[int]: return [int(result["index"]) for result in results] -EMBEDDING_ADAPTERS = {"nemotron3": Nemotron3Embedding} +EMBEDDING_ADAPTERS = {"nemotron3": Nemotron3Embedding, "jina-v5-nano": JinaV5NanoEmbedding} RERANKER_ADAPTERS = {"jina-v3.5": Jina35Reranker} diff --git a/tests/test_embedding_models.py b/tests/test_embedding_models.py new file mode 100644 index 0000000..09a63cc --- /dev/null +++ b/tests/test_embedding_models.py @@ -0,0 +1,191 @@ +from __future__ import annotations + +from contextlib import nullcontext +import math +from pathlib import Path +import tempfile +from types import SimpleNamespace +import unittest +from unittest.mock import Mock, patch + +from fastapi.testclient import TestClient + +from rightmemory import embedding_models, embedding_service +from tests.test_embedding_retrieval import FakeRerankerAdapter + + +class Tensor: + """Small numeric test double; the ordinary test suite does not require torch.""" + + def __init__(self, values): + self.values = values + self.shape = (len(values), len(values[0])) if isinstance(values[0], list) else (len(values),) + self.device = "cpu" + + def sum(self, *, dim): + assert dim == 1 + return Tensor([sum(row) for row in self.values]) + + def __sub__(self, value): + return Tensor([item - value for item in self.values]) + + def __getitem__(self, indices): + rows, columns = indices + return Tensor([self.values[row][column] for row, column in zip(rows, columns.values, strict=True)]) + + def float(self): + return self + + def cpu(self): + return self + + def tolist(self): + return self.values + + +def normalize(tensor, *, p, dim): + assert (p, dim) == (2, 1) + return Tensor([[value / max(math.hypot(*row), 1e-12) for value in row] for row in tensor.values]) + + +class Batch(dict): + def to(self, device): + self.device = device + return self + + +class JinaNanoEmbeddingTests(unittest.TestCase): + def setUp(self): + directory = tempfile.TemporaryDirectory() + self.addCleanup(directory.cleanup) + self.path = Path(directory.name) + (self.path / "model.safetensors").write_bytes(b"mock weights") + self.batch = Batch(input_ids=Tensor([[1, 2, 0], [1, 2, 3]]), + attention_mask=Tensor([[1, 1, 0], [1, 1, 1]])) + self.tokenizer = Mock(return_value=self.batch) + self.model = Mock() + self.model.config = SimpleNamespace(hidden_size=768, max_position_embeddings=8192) + self.model.to.return_value = self.model + self.model.eval.return_value = self.model + self.model.return_value = SimpleNamespace(last_hidden_state=Tensor([ + [[9, 9], [3, 4], [100, 100]], [[9, 9], [100, 100], [0, 5]], + ])) + self.torch = SimpleNamespace( + __version__="test", bfloat16="bfloat16", float32="float32", + inference_mode=Mock(side_effect=nullcontext), arange=lambda count, **kwargs: range(count), + nn=SimpleNamespace(functional=SimpleNamespace(normalize=Mock(side_effect=normalize))), + ) + self.transformers = SimpleNamespace( + __version__="test", AutoTokenizer=Mock(), AutoModel=Mock(), + ) + self.transformers.AutoTokenizer.from_pretrained.return_value = self.tokenizer + self.transformers.AutoModel.from_pretrained.return_value = self.model + dependencies = patch.object(embedding_models, "_dependencies", return_value=(self.torch, self.transformers)) + dependencies.start() + self.addCleanup(dependencies.stop) + + def test_loads_local_custom_checkpoint_with_device_precision(self): + for device, dtype in (("cpu", "float32"), ("cuda:0", "bfloat16")): + with self.subTest(device=device): + adapter = embedding_models.JinaV5NanoEmbedding(self.path, device=device) + self.transformers.AutoTokenizer.from_pretrained.assert_called_with( + str(self.path.resolve()), local_files_only=True, trust_remote_code=True, padding_side="right", + ) + self.transformers.AutoModel.from_pretrained.assert_called_with( + str(self.path.resolve()), local_files_only=True, use_safetensors=True, trust_remote_code=True, + dtype=dtype, attn_implementation="sdpa", + ) + self.model.to.assert_called_with(device) + self.model.eval.assert_called_with() + self.assertEqual((adapter.dimensions, adapter.max_tokens, adapter.max_batch_size), (768, 8192, 8)) + + def test_missing_path_fails_before_loading(self): + with self.assertRaises(FileNotFoundError): + embedding_models.JinaV5NanoEmbedding(self.path / "missing", device="cpu") + self.transformers.AutoModel.from_pretrained.assert_not_called() + self.transformers.AutoTokenizer.from_pretrained.assert_not_called() + + def test_query_and_passage_prefixes_and_masked_last_token_pooling(self): + adapter = embedding_models.JinaV5NanoEmbedding(self.path, device="cpu") + for kind, prefix in (("query", "Query: "), ("passage", "Document: ")): + with self.subTest(kind=kind): + vectors = adapter.encode(["short", "longer text"], kind=kind) + self.tokenizer.assert_called_with( + [prefix + "short", prefix + "longer text"], + padding=True, truncation=False, return_tensors="pt", + ) + self.assertEqual(vectors, [[0.6, 0.8], [0.0, 1.0]]) + self.assertEqual(self.batch.device, "cpu") + self.model.assert_called_with(**self.batch) + self.assertEqual(self.torch.inference_mode.call_count, 2) + + def test_context_overflow_fails_before_device_transfer_and_inference(self): + adapter = embedding_models.JinaV5NanoEmbedding(self.path, device="cpu") + self.batch["input_ids"].shape = (2, 8193) + with self.assertRaises(ValueError): + adapter.encode(["too long"], kind="query") + self.assertFalse(hasattr(self.batch, "device")) + self.model.assert_not_called() + + def test_exact_context_boundary_is_allowed(self): + adapter = embedding_models.JinaV5NanoEmbedding(self.path, device="cpu") + self.batch["input_ids"].shape = (2, 8192) + self.assertEqual(adapter.encode(["at limit"], kind="passage"), [[0.6, 0.8], [0.0, 1.0]]) + + def test_context_respects_smaller_checkpoint_limit(self): + self.model.config.max_position_embeddings = 4096 + adapter = embedding_models.JinaV5NanoEmbedding(self.path, device="cpu") + self.assertEqual(adapter.max_tokens, 4096) + self.model.config.max_position_embeddings = 32768 + adapter = embedding_models.JinaV5NanoEmbedding(self.path, device="cpu") + self.assertEqual(adapter.max_tokens, 8192) + + def test_identity_tracks_precision_weights_and_adapter_settings(self): + cpu = embedding_models.JinaV5NanoEmbedding(self.path, device="cpu") + self.assertTrue(cpu.identity.startswith("jina-v5-nano:")) + self.assertEqual(cpu.identity, embedding_models.JinaV5NanoEmbedding(self.path, device="cpu").identity) + self.assertNotEqual(cpu.identity, embedding_models.JinaV5NanoEmbedding(self.path, device="cuda:0").identity) + with patch.object(embedding_models, "adapter_identity", return_value="identity") as identity: + embedding_models.JinaV5NanoEmbedding(self.path, device="cpu") + settings = identity.call_args.kwargs["settings"] + self.assertEqual(settings["pooling"], "last-token") + self.assertEqual(settings["normalization"], "l2") + self.assertEqual(settings["padding_side"], "right") + (self.path / "model.safetensors").write_bytes(b"changed weights") + self.assertNotEqual(cpu.identity, embedding_models.JinaV5NanoEmbedding(self.path, device="cpu").identity) + + def test_cli_selects_nano_and_service_reports_its_identity_and_dimensions(self): + with patch.dict(embedding_service.RERANKER_ADAPTERS, {"jina-v3.5": FakeRerankerAdapter}), \ + patch.dict("os.environ", {"RIGHTMEMORY_EMBEDDING_API_KEY": ""}), \ + patch("uvicorn.run") as run: + self.assertEqual(embedding_service.main([ + "--embedding-adapter", "jina-v5-nano", "--embedding-model", str(self.path), + "--reranker-model", "rerank", "--device", "cpu", + ]), 0) + with TestClient(run.call_args.args[0]) as client: + info = client.get("/info").json() + self.assertTrue(info["embedding_model"].startswith("jina-v5-nano:")) + self.assertEqual((info["dimensions"], info["max_batch_size"]), (768, 8)) + response = client.post("/embed", json={"model": info["embedding_model"], + "kind": "passage", "texts": ["short", "longer text"]}) + self.assertEqual(response.status_code, 200) + self.assertEqual(response.json()["vectors"], [[0.6, 0.8], [0.0, 1.0]]) + self.batch["input_ids"].shape = (2, 8193) + self.assertEqual(client.post("/embed", json={"model": info["embedding_model"], + "kind": "query", "texts": ["too long"]}).status_code, 422) + + def test_nemotron_remains_the_default(self): + self.assertIs(embedding_models.EMBEDDING_ADAPTERS["nemotron3"], embedding_models.Nemotron3Embedding) + from tests.test_embedding_retrieval import FakeEmbeddingAdapter + embedding = Mock(side_effect=FakeEmbeddingAdapter) + with patch.dict(embedding_service.EMBEDDING_ADAPTERS, {"nemotron3": embedding}), \ + patch.dict(embedding_service.RERANKER_ADAPTERS, {"jina-v3.5": FakeRerankerAdapter}), \ + patch.dict("os.environ", {"RIGHTMEMORY_EMBEDDING_API_KEY": ""}), patch("uvicorn.run"): + self.assertEqual(embedding_service.main([ + "--embedding-model", "embed", "--reranker-model", "rerank", "--device", "cpu", + ]), 0) + embedding.assert_called_once_with(Path("embed"), device="cpu") + + +if __name__ == "__main__": + unittest.main()