Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 15 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -167,7 +167,7 @@ For a short recording script, see [docs/DEMO.md](docs/DEMO.md).
- Multi-device memory continuity across laptops, desktops, agent clients, and project-specific roots.
- Automatic orchestration plus independent explicit skills for map maintenance, Memory maintenance, guidance review, and Manager support in Codex App.
- Two executor modes behind the same `rightmemory` CLI: standalone runtime or delegated Codex SDK/Claude Code CLI role execution.
- Source-authored retrieval output, selected by an agent or an optional Nemotron + Jina retrieval service.
- Source-authored retrieval output, selected by an agent or an optional embedding + reranking retrieval service.
- One updater for Memory and Agent Corrections, transcript-review candidate extraction, and immutable candidate records for input-to-edit provenance.

## Install Options And Updates
Expand Down Expand Up @@ -815,7 +815,7 @@ Configure `[sync-reconciler.model]` or `[sync-reconciler.agent_cli]` only if syn

### Embedding Retrieval

Embedding retrieval is an optional alternative to the agent-based retriever in either installation mode. One complete query goes to **Nemotron 3 Embed 1B**, which selects forty source entries; **Jina reranker v3.5** ranks those candidates and RightMemory returns up to ten. There is no relevance-score cutoff or additional LLM selection call. Related entries can be returned even when Memory does not contain the requested answer; the calling agent decides which entries apply.
Embedding retrieval is an optional alternative to the agent-based retriever in either installation mode. One complete query goes to the selected embedding model, **Nemotron 3 Embed 1B** (the default) or **Jina Embeddings v5 Text Nano Retrieval**, which selects forty source entries; **Jina reranker v3.5** ranks those candidates and RightMemory returns up to ten. There is no relevance-score cutoff or additional LLM selection call. Related entries can be returned even when Memory does not contain the requested answer; the calling agent decides which entries apply.

Add these retrieval settings, keeping the existing `[retrieve.agent_cli]` or `[retrieve.model]` table as the fallback:

Expand Down Expand Up @@ -852,11 +852,21 @@ rightmemory embedding-service \
--device cuda:0 --port 8766
```

Supply downloaded snapshots, for example from ModelScope. The service loads local files offline and keeps both models resident. The supported model interfaces are documented by [NVIDIA](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16) and [Jina](https://huggingface.co/jinaai/jina-reranker-v3.5); Jina's snapshot includes custom model code. The service serializes GPU requests and defaults to loopback. For direct access on a trusted private network, start it with `--host <server-private-IP>` and configure clients with `url = "http://<server-private-IP>:8766"`. Encoding sends source passages to that configured host; reranking sends the query and candidate passages.
For the smaller Jina nano embedding model, select its adapter and the merged retrieval snapshot; keep the same reranker and client configuration:

Embedding and reranking adapters are selected independently. The defaults above support the tested Nemotron 1B / Jina v3.5 pair. To support another model family, implement the small `EmbeddingAdapter` or `RerankerAdapter` interface in `rightmemory/embedding_models.py` and register it in the corresponding adapter dictionary. An embedding adapter returns dense vectors for cosine search; a reranker returns every candidate index in ranked order. The adapters own model loading, input formatting, and token limits. The service reports embedding dimensions, batch size, and candidate capacity, which the client uses instead of model-specific constants. A compatible external service can also implement `/info`, `/embed`, and `/rerank`.
```bash
rightmemory embedding-service --embedding-adapter jina-v5-nano --embedding-model /path/to/jina-embeddings-v5-text-nano-retrieval --reranker-model /path/to/jina-reranker-v3.5/snapshot --device cuda:0 --port 8766
```

Use [`jinaai/jina-embeddings-v5-text-nano-retrieval`](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano-retrieval), with the Safetensors weights, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `configuration_eurobert.py`, and `modeling_eurobert.py` from the same snapshot. This adapter loads the merged retrieval weights directly; the base multi-task nano checkpoint, adapter-only weights, GGUF, and ONNX files are not substitutes. It uses `Query: ` / `Document: ` prefixes, right padding, last non-padding token pooling, and L2 normalization. The output is 768 dimensions with an 8,192-token input limit including prefixes and tokenizer special tokens; longer inputs fail without truncation. CUDA inference uses BF16, while `--device cpu` uses FP32. No additional PEFT, sentence-transformers, or flash-attention dependency is needed for this adapter. The model is licensed [CC BY-NC 4.0](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano-retrieval#license); check Jina's terms for commercial use.

After starting the service, check `http://127.0.0.1:8766/info` (PowerShell: `Invoke-RestMethod http://127.0.0.1:8766/info`). With nano selected, it reports an `embedding_model` identity starting with `jina-v5-nano:` and `dimensions = 768`. Then use the normal retrieve command above. Switching from Nemotron rebuilds the derived embedding cache on the next retrieval; the first query may take longer. This does not change the Memory source files. Real-checkpoint quality, latency, and GPU memory use should be evaluated on the model host.

Supply downloaded snapshots, for example from ModelScope. The service does not download models; it loads local files offline and keeps both models resident. The supported model interfaces are documented by [NVIDIA](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16) and [Jina](https://huggingface.co/jinaai/jina-reranker-v3.5). Both Jina snapshots include custom model code, which the service executes locally; use trusted official snapshots. The service serializes GPU requests and defaults to loopback. For direct access on a trusted private network, start it with `--host <server-private-IP>` and configure clients with `url = "http://<server-private-IP>:8766"`. Encoding sends source passages to that configured host; reranking sends the query and candidate passages.

Embedding and reranking adapters are selected independently. The default pair remains Nemotron 1B / Jina v3.5; `jina-v5-nano` changes only the embedding model. To support another model family, implement the small `EmbeddingAdapter` or `RerankerAdapter` interface in `rightmemory/embedding_models.py` and register it in the corresponding adapter dictionary. An embedding adapter returns dense vectors for cosine search; a reranker returns every candidate index in ranked order. The adapters own model loading, input formatting, and token limits. The service reports embedding dimensions, batch size, and candidate capacity, which the client uses instead of model-specific constants. A compatible external service can also implement `/info`, `/embed`, and `/rerank`.

Embedding cache identities cover snapshot contents, adapter settings, and an adapter revision. The built-in adapter records its input prefixes, pooling, normalization, precision, and library versions; bump its revision for behavior changes not represented in those settings. Changing embedding identity rebuilds the vector cache; replacing only the reranker reuses it. Adding another model family requires an adapter rather than just pointing an existing adapter at unrelated weights.
Embedding cache identities cover snapshot contents, adapter settings, and an adapter revision. Each built-in embedding adapter records its input prefixes, pooling, normalization, precision, and library versions; bump its revision for behavior changes not represented in those settings. Changing embedding identity rebuilds the vector cache; replacing only the reranker reuses it. Adding another model family requires an adapter rather than just pointing an existing adapter at unrelated weights.

An API key is optional when binding to loopback or an explicit private IP address (RFC 1918 IPv4 or IPv6 unique-local). Without a service key, requests are accepted without credentials. To enable authentication, set `RIGHTMEMORY_EMBEDDING_API_KEY` on the service and the matching `[retrieve.embedding].api_key` on clients. Public, wildcard, and other hostname bindings require a key. The service is started separately and is not launched by the normal installer or a retrieve call.

Expand Down
45 changes: 44 additions & 1 deletion rightmemory/embedding_models.py
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,49 @@ def encode(self, texts: list[str], *, kind: Literal["query", "passage"]) -> list
return vectors.cpu().tolist()


class JinaV5NanoEmbedding:
"""The merged jina-embeddings-v5-text-nano-retrieval checkpoint."""

max_batch_size = 8

def __init__(self, path: Path, *, device: str):
torch, transformers = _dependencies()
self.torch, self.device = torch, device
path = path.expanduser().resolve(strict=True)
dtype = torch.bfloat16 if device.startswith("cuda") else torch.float32
self.tokenizer = transformers.AutoTokenizer.from_pretrained(
str(path), local_files_only=True, trust_remote_code=True, padding_side="right",
)
self.model = transformers.AutoModel.from_pretrained(
str(path), local_files_only=True, use_safetensors=True, trust_remote_code=True,
dtype=dtype, attn_implementation="sdpa",
).to(device).eval()
self.dimensions = self.model.config.hidden_size
self.max_tokens = min(8192, self.model.config.max_position_embeddings)
self.prefixes = {"query": "Query: ", "passage": "Document: "}
self.identity = adapter_identity(path, "jina-v5-nano", revision=1, settings={
"checkpoint": "jinaai/jina-embeddings-v5-text-nano-retrieval",
"prefixes": self.prefixes, "pooling": "last-token", "normalization": "l2",
"padding_side": "right", "dtype": str(dtype), "attention": "sdpa",
"dimensions": self.dimensions, "max_tokens": self.max_tokens,
"torch": torch.__version__, "transformers": transformers.__version__,
})

def encode(self, texts: list[str], *, kind: Literal["query", "passage"]) -> list[list[float]]:
batch = self.tokenizer([self.prefixes[kind] + text for text in texts],
padding=True, truncation=False, return_tensors="pt")
if batch["input_ids"].shape[1] > self.max_tokens:
raise ValueError("embedding input exceeds the model context; shorten the query or source passage")
batch = batch.to(self.device)
with self.torch.inference_mode():
hidden = self.model(**batch).last_hidden_state
# With right padding, each row ends at its own last non-padding token.
last_tokens = batch["attention_mask"].sum(dim=1) - 1
pooled = hidden[self.torch.arange(hidden.shape[0], device=hidden.device), last_tokens]
vectors = self.torch.nn.functional.normalize(pooled.float(), p=2, dim=1)
return vectors.cpu().tolist()


class Jina35Reranker:
max_candidates = 125
max_query_tokens = 1024
Expand Down Expand Up @@ -143,5 +186,5 @@ def rerank(self, query: str, documents: list[str]) -> list[int]:
return [int(result["index"]) for result in results]


EMBEDDING_ADAPTERS = {"nemotron3": Nemotron3Embedding}
EMBEDDING_ADAPTERS = {"nemotron3": Nemotron3Embedding, "jina-v5-nano": JinaV5NanoEmbedding}
RERANKER_ADAPTERS = {"jina-v3.5": Jina35Reranker}
Loading
Loading