diff --git a/.env.template b/.env.template new file mode 100644 index 0000000..b482030 --- /dev/null +++ b/.env.template @@ -0,0 +1,17 @@ +# Copy to .env and fill in. .env is ignored by git. Run the scripts with: +# uv run --env-file .env scripts/measure_cloud.py +# uv run --env-file .env scripts/measure_local.py + +# measure_cloud.py (M1, M2, M3) +HOTDATA_API_KEY= +# Optional. Without it, the framework picks the active workspace of the key. +HOTDATA_WORKSPACE= +HOTDATA_API_URL=https://api.hotdata.dev +# The throwaway database. The script refuses a name that already exists and deletes +# the database on exit. Use a new name for each run. +HOTMEMORY_MEASURE_DB=hotmemory-measure-1 +HOTMEMORY_EMBEDDING_PROVIDER=sys_emb_openai + +# measure_local.py (M4, M5, M6). It ignores the HOTDATA_* values above. +HOTMEMORY_LOCAL_URL=http://localhost:3000 +HOTMEMORY_M5_SIZES=1000,10000,100000 diff --git a/.github/CODEOWNERS b/.github/CODEOWNERS new file mode 100644 index 0000000..43693f2 --- /dev/null +++ b/.github/CODEOWNERS @@ -0,0 +1 @@ +* @hotdata-dev/engineers diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..0a1cf3b --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,39 @@ +# Agent instructions + +This file is for coding agents that work in this repository. It gives the commands and the +rules, and points to the files that hold everything else. + +## Commands + +```sh +make verify # the full check; run it before each commit +make local-up # start the local RuntimeDB stack (Postgres, RustFS, engine) +make local-down # stop it and delete its data +cp .env.template .env # then fill in .env; it is ignored by git +uv run --env-file .env scripts/measure_cloud.py # M1 to M3 +uv run --env-file .env scripts/measure_local.py # M4 to M6; needs make local-up +uv run --env-file .env scripts/measure_local.py --cloud # M4 to M6 in the cloud +``` + +## Where things are + +- [docs/internal/plan.md](docs/internal/plan.md): the current phase, its tasks, and its stop + conditions. +- [docs/internal/roadmap.md](docs/internal/roadmap.md): every phase and its status. +- [docs/contracts.md](docs/contracts.md): the record and the operations. +- [docs/guarantees.md](docs/guarantees.md): what a consumer can rely on, and the proof. +- [docs/internal/brief.md](docs/internal/brief.md): the design and its reasons. +- [CONTRIBUTING.md](CONTRIBUTING.md): the checks inside `make verify` and the test rules. + +## Rules + +- Work the current GitHub issue in task order, on one branch from `main`. +- Do not push, open a pull request, or post a comment until the repository owner agrees. +- Never name a private repository, a customer, or a deployment detail in a committed file. + Before each commit, search the diff for the names that you know. +- Never write a measured number that you did not observe. Mark a claim that you read from + source, and did not observe, as not observed. +- The measurement scripts read credentials from the environment. Do not put a key, a + workspace id, or a database id in a file. +- If a change alters a behavior, update `docs/contracts.md` and `docs/guarantees.md` in the + same commit. diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..43c994c --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1 @@ +@AGENTS.md diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md new file mode 100644 index 0000000..97d66c0 --- /dev/null +++ b/CONTRIBUTING.md @@ -0,0 +1,74 @@ +# Contributing + +## Prerequisites + +- [uv](https://docs.astral.sh/uv/) installs the development tools and runs the scripts. +- GNU Make runs the targets below. +- Docker Desktop runs the local RuntimeDB stack. You need it only for `make local-up`. + +## The one command + +Run this command before each commit: + +```sh +make verify +``` + +`make verify` is the full check. CI runs the same target. If it passes on your machine, it +passes in CI. It has no tiers, because the full check takes less than five seconds. + +Today, `make verify` runs these checks, in this order: + +1. `ruff check` over the Python files. +2. `ruff format --check` over the Python files and the Python code blocks in Markdown. +3. The link check. It reads each relative link in the Markdown files at the root and + under `docs/`. A link to a file that does not exist makes it fail. + +When the library exists, strict `mypy` and the offline test suite join `make verify` +before the link check. + +## The other targets + +| Target | What it does | +|---|---| +| `make verify` | Runs every check above. | +| `make local-up` | Starts the local RuntimeDB stack: Postgres, RustFS, and the engine. [docs/local.md](docs/local.md) tells you how to point the library at it. | +| `make local-down` | Stops the stack and deletes its data. | +| `make local-pull` | Downloads newer images for the stack. | + +## Rules of the harness + +The library is deterministic. The offline suite needs no clock, no network, and no model. +These rules keep it that way. + +- One conformance suite runs against every driver. Each answered guarantee in + [docs/guarantees.md](docs/guarantees.md) is one test. A driver that fails a conformance + test is not a driver. +- The in-memory driver is the reference for the Hotdata driver. A test builds the same + records in both drivers and compares the results of `search`. +- Four surfaces are frozen: the names in `__all__`, the method set of the `Store` + protocol, the fields and field types of the record for each schema version, and the + filter keys of `list` and `search`. A test compares each surface against a literal set. + A change to a frozen surface is a public contract change, and it needs a changelog + entry. +- A frozen surface is compared against a literal set, never against the thing that it + protects. A test that iterates over the protected thing turns a deletion into one test + fewer and not into a failure. +- Each row in the guarantees ledger names the test that proves it. A test reads the ledger. + A named test that does not exist makes it fail. +- Tests marked `hotdata` run the Hotdata driver against a real database. They need + `HOTMEMORY_TEST_DB` to name a throwaway database. Without it, they skip. They + run once for each pull request. They are the only tests that use the network. +- No test calls a model. `capture` takes a callable, and the tests pass a fake extractor + that returns fixed facts. +- There is no coverage gate, no mutation-testing gate, and no report generator. The output + of `make verify` is the report. + +## Documents + +Documents are checked like code. If you change a behavior, update +[docs/contracts.md](docs/contracts.md) and [docs/guarantees.md](docs/guarantees.md) in the +same pull request. + +No file in this repository names a private repository, a customer, or a deployment detail. +This rule includes `docs/internal/`. diff --git a/Makefile b/Makefile new file mode 100644 index 0000000..2694cf8 --- /dev/null +++ b/Makefile @@ -0,0 +1,35 @@ +export RUNTIMEDB_IMAGE ?= ghcr.io/hotdata-dev/runtimedb:latest +export LOCAL_PORT ?= 3000 + +STORAGE_URL := http://127.0.0.1:9000 + +.PHONY: verify local-up local-down local-pull + +verify: + uv run --group dev ruff check . + uv run --group dev ruff format --check . + uv run --no-project python scripts/check_links.py + +local-up: + docker compose up -d --wait catalog storage + @for i in $$(seq 1 30); do \ + curl -s -o /dev/null $(STORAGE_URL)/ && break; \ + sleep 1; \ + done + curl -fsS -o /dev/null -X PUT --aws-sigv4 "aws:amz:us-east-1:s3" \ + --user hotmemory:hotmemory-local $(STORAGE_URL)/runtimedb + RUNTIMEDB_SECRET_KEY="$$(openssl rand -base64 32)" docker compose up -d runtimedb + @for i in $$(seq 1 60); do \ + curl -fs -o /dev/null http://localhost:$(LOCAL_PORT)/health && \ + echo "RuntimeDB is up at http://localhost:$(LOCAL_PORT)" && exit 0; \ + sleep 1; \ + done; \ + echo "RuntimeDB did not answer /health in 60 seconds"; \ + docker compose logs --tail 50 runtimedb; \ + exit 1 + +local-down: + docker compose down -v + +local-pull: + docker compose pull diff --git a/README.md b/README.md index 4504cd7..385adb5 100644 --- a/README.md +++ b/README.md @@ -2,22 +2,37 @@ Agent memory as tables on [Hotdata](https://hotdata.dev). -A memory record is a row in a managed table. The row carries the text, an embedding, a -scope, tags, source references, and the time span the fact was true. Because the table is -an ordinary Hotdata table, an agent can join its memory to its own data in one SQL query. -No other memory store offers that join. +A memory record is a row in a managed table. The row carries the text, a scope, tags, +source references, and the time span in which the fact was true. The table is an ordinary +Hotdata table, so an agent can join its memory to its own data in one SQL query. None of +the memory systems that we surveyed stores memory as typed columns in the same engine as +the data of the consumer. -Status: design. The contracts are drafted in [docs/internal/brief.md](docs/internal/brief.md). No code -exists yet. +Status: design. No code exists yet. The contracts and the guarantees are written, and +phase 0 measured each guarantee that was open. The library has two layers: - A storage contract. Put, get, list, search, and delete records in a namespace. Records are immutable. A new put under the same key creates a new revision. - A memory contract. Remember facts, recall them inside a context budget, supersede a - subject, and forget by id or by horizon. Extraction from raw text is optional and takes + fact, and forget by id or by horizon. Extraction from raw text is optional, and it takes a model callable that the caller supplies. -The first consumer is an incident investigator built on Hotdata. The library is not -specific to it. A plain Python agent, a LangGraph agent, or any process with a Hotdata API -key can use it. +The library runs in the process of the consumer and calls the Hotdata API with the API key +of the consumer. There is no hotmemory server. The first consumer is an incident +investigator built on Hotdata. The library is not specific to it. A plain Python agent, a +LangGraph agent, or any process with a Hotdata API key can use it. + +## Documents + +- [docs/contracts.md](docs/contracts.md): the record, the storage operations, the memory + operations, and the platform facts behind them. +- [docs/guarantees.md](docs/guarantees.md): each behavior that a consumer can rely on, its + state, and its proof. +- [docs/local.md](docs/local.md): how to run the library against a local RuntimeDB + container. +- [CONTRIBUTING.md](CONTRIBUTING.md): the one command that checks a change, and the rules + of the test suite. +- [docs/internal/](docs/internal/brief.md): the design brief, the survey, the roadmap, and + the plan for the current phase. diff --git a/compose.yaml b/compose.yaml new file mode 100644 index 0000000..5ef96ee --- /dev/null +++ b/compose.yaml @@ -0,0 +1,57 @@ +# Local RuntimeDB for development and measurements. Started by `make local-up`. +# +# runtimedb shares the network namespace of storage, so 127.0.0.1:9000 is RustFS both +# inside runtimedb and on the host. Presigned upload URLs carry that address, so a +# client on the host can upload. +# +# No named volume. `make local-down` removes every container and its anonymous volumes, so +# the data is gone. + +name: hotmemory + +services: + catalog: + image: ${POSTGRES_IMAGE:-postgres:17} + environment: + POSTGRES_USER: runtimedb + POSTGRES_PASSWORD: runtimedb + POSTGRES_DB: runtimedb + healthcheck: + test: ["CMD", "pg_isready", "-U", "runtimedb", "-d", "runtimedb"] + interval: 1s + timeout: 3s + retries: 30 + + storage: + image: ${RUSTFS_IMAGE:-rustfs/rustfs:latest} + environment: + RUSTFS_ACCESS_KEY: hotmemory + RUSTFS_SECRET_KEY: hotmemory-local + ports: + - "127.0.0.1:9000:9000" + - "127.0.0.1:${LOCAL_PORT:-3000}:3000" + + runtimedb: + image: ${RUNTIMEDB_IMAGE:-ghcr.io/hotdata-dev/runtimedb:latest} + network_mode: service:storage + depends_on: + catalog: + condition: service_healthy + storage: + condition: service_started + environment: + RUNTIMEDB_SECRET_KEY: ${RUNTIMEDB_SECRET_KEY:-} + RUNTIMEDB_AUTH__ALLOW_UNAUTHENTICATED: "true" + RUNTIMEDB_ENGINE__SQL_WRITES: "true" + RUNTIMEDB_CATALOG__TYPE: postgres + RUNTIMEDB_CATALOG__HOST: catalog + RUNTIMEDB_CATALOG__PORT: "5432" + RUNTIMEDB_CATALOG__DATABASE: runtimedb + RUNTIMEDB_CATALOG__USER: runtimedb + RUNTIMEDB_CATALOG__PASSWORD: runtimedb + RUNTIMEDB_STORAGE__TYPE: s3 + RUNTIMEDB_STORAGE__BUCKET: runtimedb + RUNTIMEDB_STORAGE__REGION: us-east-1 + RUNTIMEDB_STORAGE__ENDPOINT: http://127.0.0.1:9000 + RUNTIMEDB_STORAGE__ACCESS_KEY: hotmemory + RUNTIMEDB_STORAGE__SECRET_KEY: hotmemory-local diff --git a/docs/contracts.md b/docs/contracts.md new file mode 100644 index 0000000..e5a8cd5 --- /dev/null +++ b/docs/contracts.md @@ -0,0 +1,236 @@ +# Contracts + +hotmemory has two contracts. The storage contract is the stable surface that every driver +implements. The memory contract is the surface that an agent calls, and it is built on one +store. This file states both. The behavior that each contract guarantees, and the proof for +each guarantee, are in [guarantees.md](guarantees.md). + +Status: design. No code exists yet. This file describes schema version 1 as the library +will ship it. + +## The platform under the store + +The Hotdata driver writes to managed tables. A managed table is a table that the Hotdata +API owns and loads from files. The facts below shape the contracts. A fact marked +[measured] was observed against a real Hotdata workspace between 2026-08-18 and +2026-09-18. + +- A write is a file load. The caller uploads parquet and calls a keyed load in replace, + append, upsert, update, or delete mode. [measured] +- A write costs about 2.1 seconds per call, and the cost does not change with row count. + One row costs 1.75 seconds. Ten thousand rows cost 2.5 seconds. Call count is the only + cost that a caller controls. [measured] +- There is no conditional write. No load mode compares the existing value, and no load + mode returns the row that it changed. If two writers load one key, the last writer + wins. [measured] +- A load locks the whole table. A second load at the same time is refused with + `409 RESOURCE_LOCKED`. Distinct keys give no concurrency. [measured] +- A read after a write was never stale in 30 loads. A read costs 0.3 to 0.6 seconds. The + first call to a cold engine costs about 8 to 10 seconds. [measured] +- A table layout is permanent. There is no ALTER. An upsert load with extra columns + succeeds and drops the extra columns without an error. Only a replace load adds + columns. [measured] +- After you delete a managed table, its name stays reserved forever. [measured] +- The `expires_at` value on a database is a record only. Nothing deletes the database + when the time passes. [measured] +- A managed database is a hard isolation boundary. One managed database cannot attach + another. [measured] +- A managed table can have three kinds of index: BM25 over a text column, vector, and + sorted. A plain vector index covers a vector column that the caller fills. A + provider-backed vector index takes a text column and embeds it on the server. In plain + mode, a query whose metric differs from the index metric does a full scan and gives no + error. SQL cannot see indexes. Only the control plane (the management API) + can tell whether an index exists. [measured] +- A provider-backed vector index cannot share its table with any other index. The engine + refuses the second index. A plain vector index and a BM25 index can share a table. + [measured 2026-08-18] +- RuntimeDB, the Hotdata query engine, runs on a laptop with a Postgres container and an + S3-compatible storage container. The bare engine image alone refuses managed tables. + [measured 2026-10-05] [local.md](local.md) tells you how to start the stack. + +The contracts take these decisions from the facts: + +- Records are immutable. A revision is a new row, and the new row marks the old row as + superseded. This gives history, supersession, and safe concurrency on a store whose + write is a keyed row replace. +- Each table has one writer. The driver sends one load at a time and retries on 409. +- A write is synchronous and slow, or buffered and flushed. The API has both. + [guarantees.md](guarantees.md) gives the visibility rule for each one. +- Expiry is a column and a sweeper. The platform deletes nothing. +- A database is the tenant boundary. The platform enforces it through the workspace + allow-list of the API token. A namespace is a column filter inside the database, and + the library enforces it. The library never claims to enforce more than that. +- The table name carries the schema version. A column change makes a new version and + never edits a table. +- Each row holds one fact. A document is not a memory. It is an episode that facts point + back to. The section on tables tells you why. + +## The storage contract + +The storage contract is a Python protocol named `Store`. Every driver implements it, and +the conformance suite runs against every driver. + +### The record + +A record is the unit that the store holds. Schema version 1 fixes these fields. + +| Field | Type | Meaning | +|---|---|---| +| `namespace` | tuple of strings | Where the record lives. The store keeps it as one path string joined with `/`. A match is on whole labels and never on a string prefix. No label contains `.`. | +| `key` | string | The stable identifier of the caller inside the namespace. | +| `revision` | integer | 1 for the first put under a key. Each later put adds 1. | +| `kind` | string | One of `fact`, `profile`, `procedure`, `episode`. | +| `subject` | string | What the record is about, for example an alert key, a person, or a service. Empty if unknown. | +| `content` | string | The text that a model reads. | +| `cues` | list of strings | Questions or phrases that this record answers. Optional. The driver embeds them apart from `content`. | +| `payload` | JSON object | Structured data that the consumer defines. The store never reads it. | +| `tags` | list of strings | Free labels. You can filter on them. | +| `sources` | list of strings | References to the origin of the record: a thread id, a document path, a run id, an episode key. The length of the list is the corroboration count. | +| `actor` | string | Who wrote this revision: a user id, an agent name, or an extractor name. | +| `created_at` | timestamp | When the store wrote this revision. System clock. | +| `observed_at` | timestamp or null | The time of the source. A post-mortem that you load a year later keeps the incident date here. | +| `valid_from` | timestamp or null | When the fact became true. The default is `observed_at`. Null means unknown. | +| `valid_until` | timestamp or null | When the fact stopped being true. World clock. Null means that it is still true. | +| `expired_at` | timestamp or null | When the store learned that the fact stopped being true. System clock. Null on a current record. | +| `superseded_by` | string or null | The id of the revision that replaced this one. Null on the current revision. | +| `forget_after` | timestamp or null | When a sweeper can delete the record. Null means keep. | +| `forget_reason` | string | The reason for `forget_after`. Empty if `forget_after` is null. | +| `id` | string | `namespace/key@revision`. Derived. The load key. | + +The public record has no embedding field. If the caller supplies an embedder, the Hotdata +driver adds embedding columns. If the caller uses a provider-backed index, the driver adds +none. The driver configuration selects one of the two. + +A record has two clocks. `valid_until` records the change in the world. `expired_at` +records the moment that the store found out. To ask what memory held on a given day, read `created_at` and +`expired_at`. To ask what was true in the world on that day, read `valid_from` and +`valid_until`. + +### The operations + +| Operation | Arguments | Behavior | +|---|---|---| +| `put` | namespace, key, record fields | Writes a new revision. If the key exists, the new row gets the next revision, and the previous current row gets `superseded_by`. Returns the id. If the normalized content is equal to the content of the current revision, it writes nothing and returns the current id. | +| `get` | namespace, key, optional revision | Returns the current revision, or the named revision. Returns None if the record does not exist. | +| `history` | namespace, key | Returns every revision, oldest first. | +| `list` | namespace prefix, optional filter, optional since, limit | Returns current revisions under the prefix, newest first. It uses no model and no embedding. | +| `search` | query text or none, namespace prefixes, optional filter, k | Returns up to k current revisions in order of relevance, closest first, each with a distance. It takes the same filter as `list`. With no query text, it is `list`. | +| `delete` | namespace, key | Removes every revision of the key. This is a hard delete. | +| `list_namespaces` | optional prefix | Returns the distinct namespaces under the prefix. | +| `writer` | none | A context manager. It buffers every `put` inside it. The buffer flushes on exit, at a row count, or at an interval, and returns the ids that it flushed. | + +The filter accepts equality on `kind`, `subject`, `tags`, and `actor`. It accepts a range on +`valid_from`, `valid_until`, `created_at`, and `expired_at`. Any other filter raises an +error. + +`list` and `search` never return a deleted revision, a superseded revision, or a record +past `forget_after`. `history` returns superseded revisions. No operation returns a deleted +revision. + +Normalized content is the content in lowercase with each run of whitespace changed to one +space. Deduplication compares the normalized form exactly, and does nothing more. A +paraphrase is a new record. The extractor merges paraphrases, with help from `candidates` +in the memory contract. + +### The drivers + +Version 1 ships two drivers. + +- `MemoryStore` runs in the process, in memory. It is a real driver and not a mock. It + computes relevance with the same cosine distance that the engine uses, and it refuses a + filter that it does not model. The offline test suite runs against it, and it is the + reference for the other driver. +- `HotdataStore` uses one managed database, two tables per schema version, keyed loads, a + serialized writer, and the retrieval query below. With the local RuntimeDB stack, it is also + the development driver. + +### Tables and retrieval + +Each schema version has two tables, because facts and raw material have different shapes +and different write patterns. + +The `memory` table holds facts, profiles, and procedures. Its rows are small, revisioned, +and searched often, and `profile` renders them. The `episode` table holds raw material: a +thread, a document, or a post-mortem, cut into chunks of a fixed size. It is append-only +and never revisioned. If a fact is not enough, a consumer searches it. The `sources` of a +fact name the episode keys that it came from. Thus a consumer can go from a fact to its +evidence in one join. Tabular data is never copied into either table. It stays in the +tables of the consumer, and a fact points at it. + +Each row holds one fact and not one document, for this reason. If a document is one +record, each small revision is a near-duplicate of the whole document. The cosine distance +between two versions is close to zero, and that is the duplicate problem that a memory +layer exists to prevent. Small facts revise, supersede, and render into a profile without +this problem. Markdown is a good format for the `content` of an episode chunk. + +Retrieval is one SQL query over the `memory` table, in three stages. + +1. Exact filters come first: namespace labels, validity at the as-of time, `kind`, + `subject`, `tags`, and `forget_after`. These are plain predicates, and they make the + candidate set smaller before any ranking. +2. The query ranks the remaining rows by BM25 over `content` and by vector distance over + `content` and over `cues`. It fuses the rankings by reciprocal rank fusion (it adds + `1 / (60 + rank)` from each ranking). Fusion is ordinary SQL with common table + expressions and needs no engine function. +3. A sorted index on `created_at` serves recency order and the sweeper. + +Stage 2 needs a BM25 index and vector indexes on one table. A provider-backed vector index +cannot share a table with another index, so stage 2 needs plain vector indexes and an +embedder that the caller supplies. With a provider-backed index and no BM25 index, stage 2 +ranks by meaning only. Measurement M5 in [guarantees.md](guarantees.md) confirmed both +rules on 2026-10-05. + +`bm25_search` ranks the whole table, and the filters of stage 1 apply after it. The BM25 +fetch depth must therefore be wide enough that enough rows survive the filters. Also, +`bm25_search` refuses to run without a BM25 index, so the driver always builds one. + +A cue is the question that a record answers. The extractor or the caller writes it at +capture time. A query that resembles the question matches the cue, even when the +query and the content share no words. Cues are optional, and retrieval on content alone +always works. + +Reads are fast because the table is small. A memory table holds thousands of rows, and a +filtered scan of that is fast without an index. Measurement M5 found that the vector +index saves engine time from about ten thousand rows. In the cloud, one request costs about +400 ms at every size. That cost hides the saving up to at least one hundred thousand rows. [measured 2026-10-05] + +## The memory contract + +The memory contract is what an agent calls. It is a class named `Memory`, built over one +`Store`. It can change more freely than the storage contract while the library is young. + +| Operation | Arguments | Behavior | +|---|---|---| +| `remember` | facts, scope, actor | Writes facts that are already structured. Each fact is a record with `kind` set. The key comes from the subject and a hash of the normalized content, so a retried call writes nothing new. | +| `recall` | query, scopes, optional as_of, budget in characters | Searches the allowed scopes, keeps the records that are valid at `as_of`, and returns the top records inside the budget. It returns them as a list and as one rendered block. The block labels each record with its sources and its validity span, and with nothing else. | +| `candidates` | fact, scopes, k | Returns the k nearest current records with their distances. It makes no decision. A consolidator that the caller writes reads this before it calls `remember` or `supersede`. | +| `supersede` | key of the record to close, new fact, optional valid_from | Closes the named record and writes the new fact as the next revision under its key. The `valid_until` of the old record becomes the `valid_from` of the new record, and the `expired_at` of the old record becomes now. If the `valid_from` of the old record is later than that of the new record, the call refuses. The library decides nothing by itself. | +| `forget` | ids, or a horizon | Deletes the named records, or every record whose `forget_after` is before the horizon. | +| `profile` | subject, scopes, budget in characters | Returns the current records for the subject, grouped by `kind`, as one rendered block inside the budget. The block ends with the namespaces and record counts that `recall` can reach, so an agent knows what it can search for. This is the block that an agent always loads. | +| `capture` | text, scope, actor, observed_at, extractor | Calls the extractor of the caller with the text, `observed_at`, and the current records that `recall` returns for the scope. Then it calls `remember` on the result. The extractor is a plain callable. The library ships no model and names no model. | + +A record is valid at time T when `valid_from` is null or at most T, and `valid_until` is +null or after T. A null `valid_from` means the start of time. + +These three rules apply to every operation: + +- A scope is a namespace prefix. The caller supplies the allowed scopes on each call. The + library filters on them, and does nothing else to enforce access. +- Memory is a record, not a rule. The block from `recall` and from `profile` says what was + recorded and when. It never says what to do. If a consumer wants instructions, it + writes them in its own prompt. +- A write is visible to the next `recall`, and never inside the current turn. A consumer + that captures and recalls in one step reads what was there before the capture. + +## Not in version 1 + +- An adapter for the LangGraph `BaseStore` interface. The operation names above match it, + so the adapter will be thin. +- An MCP server over the memory contract. +- A memory service in front of Hotdata. The library runs in the process of the consumer + and calls the Hotdata API. A service becomes necessary only for a client in another + language, for extraction that must run centrally, or for scope enforcement above the API + key. +- Extraction on the server. `capture` takes a callable and does nothing more. +- A third driver. The local RuntimeDB stack covers offline development. +- Access control beyond scope filtering. diff --git a/docs/guarantees.md b/docs/guarantees.md new file mode 100644 index 0000000..c8d5341 --- /dev/null +++ b/docs/guarantees.md @@ -0,0 +1,223 @@ +# Guarantees + +This file lists each behavior that a consumer can rely on, with its state and its proof. +Read it before you rely on a behavior. The contracts themselves are in +[contracts.md](contracts.md). + +## States + +Each guarantee has one of two states. + +- Answered: a measurement or the design already supports the answer. +- To measure: a measurement must give the answer before the library relies on it. + +A guarantee marked [measured] was observed against a real Hotdata workspace. In phase 1, +each answered guarantee gets one conformance test, and the Test column names it. If a named test +does not exist, a test in the suite fails. Until the suite exists, +the Test column is empty. + +## The ledger + +| Question | Answer | State | Test | +|---|---|---|---| +| Does a second `put` under the same key replace the record? | No. It writes revision n+1 and marks revision n as superseded. `get` returns n+1. | answered | | +| Can a retried `remember` create duplicates? | No. The key comes from the subject and a hash of the normalized content. A put whose normalized content is equal to the current revision writes nothing. | answered | | +| Is a paraphrase a duplicate? | No. Deduplication is exact on normalized content. `candidates` exists so that a caller can decide. | answered | | +| After a synchronous `put` returns, does `list` see the record? | Yes. A read after a write was never stale in 30 trials. | answered [measured] | | +| After a synchronous `put` returns, does `search` see the record? | Yes. Without an index, the retrieval query scans the table. With a provider-backed vector index or a BM25 index, the first search after the load returned the new row. | answered [measured], M1 | | +| After a buffered `put` returns, is the record visible? | No. It is visible after the writer flushes. `writer` returns the ids that it flushed. | answered | | +| Two processes write to the same table. What happens? | The engine refuses the second load with 409. The driver retries with backoff and stops after a bound. Inside one process, the writer sends one load at a time. | answered [measured], M3 | | +| Two writers put the same key. What happens? | The last writer wins at the row level. Revisions are new rows, so both revisions exist and the later one is current. | answered | | +| Does `delete` remove retained revisions and their embeddings? | Yes. An embedding is a column of its row. After a keyed delete, neither a provider-backed vector index nor a BM25 index returned the deleted row. | answered [measured], M2 | | +| Which filters work in `list` and `search`? | Equality on `kind`, `subject`, `tags`, and `actor`. A range on `valid_from`, `valid_until`, `created_at`, and `expired_at`. A prefix on namespace labels. Any other filter raises an error. | answered | | +| Does a higher score mean more relevant? | The store returns a distance, and a lower distance is closer. The memory contract returns records in order, with no score. An adapter that needs a score converts the distance. | answered | | +| What does `recall(as_of=T)` return? | The records that are valid at T by the as-of rule in [contracts.md](contracts.md). If `history` is consulted, this includes records superseded after T. If not, it excludes them. | answered | | +| Who enforces scope? | The library filters on the allowed scopes of the caller. The platform enforces the database boundary through the API token. A caller that holds the token can go around the library. | answered | | +| Can a consumer tell sources, extractions, and hypotheses apart? | Yes, through `kind`, `sources`, and `actor`. An extracted fact carries the name of the extractor in `actor`. | answered | | +| Is a record deleted after its `forget_after` time passes? | No. It stops appearing in `list` and `search`. The sweeper deletes it on its next run. | answered [measured] | | +| Does a write inside a turn reach a `recall` in the same turn? | No, by contract. A consumer reads what was there before its own capture. | answered | | + +## Measurements + +Phase 0 runs six measurements. M1, M2, and M3 run against a throwaway cloud database that +the run creates and deletes. M4, M5, and M6 were planned for a local RuntimeDB container. +The bare container refuses managed tables, so they also run against throwaway cloud +databases, with `--cloud`. The scripts are `scripts/measure_cloud.py` and +`scripts/measure_local.py`. + +Each result below has a number, a date, and the engine that produced it. All six ran on +2026-10-05. + +### M1. Does an index serve rows loaded after its build? + +Build a provider-backed vector index on a table, load ten more rows, and search for one of +the new rows. Record whether the index serves the new row, and after how long. + +Result, 2026-10-05, `api.hotdata.dev`, embedding provider `sys_emb_openai`: the index +served the new row on the first search after the load returned. Each index was built over +50 rows, and then 10 rows were loaded in one upsert. + +| Index | Build | Load of 10 rows | New row served | +|---|---|---|---| +| Provider-backed vector | 5.4 s | 3.7 s | at the first search, which returned 1.0 s after the load | +| BM25 | 3.0 s | 2.4 s | at the first search, which returned 0.4 s after the load | + +The measurement cannot tell whether the index or a scan of the new rows served the row. +For the contract, the result is the same: a synchronous `put` is visible to `search`. + +### M2. Does an index forget a deleted row? + +Delete a row that an index covers, and search for its content. Record whether the index +still returns the row. + +Result, 2026-10-05, `api.hotdata.dev`: no. After a keyed delete load, a scan counted 0 +rows for the deleted id. The provider-backed vector index and the BM25 index did not return +the row, at once or after 30 seconds. This is one trial on each index, on a table of 59 +rows. + +### M3. Is the retry bound of the writer enough? + +Run two processes that load the same table at the same time, each with a serialized +writer. Record how many 409 responses occur, and whether the retry bound of the driver is +enough. + +Result, 2026-10-05, `api.hotdata.dev`: the bound is enough. Two processes each made 10 +single-row upsert loads into one table at the same time. Each process retried on 409 up to +8 attempts, with a backoff that starts at 0.25 seconds and doubles to at most 4 seconds. + +| Process | Loads that succeeded | Loads that gave up | 409 responses | Most attempts for one load | Total time | +|---|---|---|---|---|---| +| 1 | 10 | 0 | 3 | 2 | 26.1 s | +| 2 | 10 | 0 | 5 | 3 | 29.8 s | + +Under this contention, a load costs about 2.6 to 3.0 seconds, against about 2.1 seconds +alone. + +### M4. What does a SQL INSERT cost against a file load? + +With `RUNTIMEDB_ENGINE__SQL_WRITES=true`, time one hundred single-row INSERT statements and +one hundred single-row upload-and-load calls into the same table shape. Record the cost per +call of each. + +Cloud result, 2026-10-05, `api.hotdata.dev`: the engine refused the first INSERT with +`Bad Request`. Production does not turn on SQL writes, so this is the expected result. One +hundred single-row upload-and-append loads took a median of 2,136 ms, a p95 of 2,614 ms, and +a mean of 2,208 ms. This agrees with the 2.1 seconds per call measured earlier. + +Local result, 2026-10-05, the local stack in [local.md](local.md) with SQL writes turned +on: one hundred single-row INSERT statements took a median of 17 ms, a p95 of 22 ms, and a +mean of 18 ms. One hundred single-row upload-and-append loads took a median of 24 ms, a p95 +of 29 ms, and a mean of 24 ms. Both tables held 101 rows after the calls. + +Locally, a load costs about 24 ms, against 2,136 ms in the cloud. So almost all of the +cloud cost is outside the engine. An INSERT is about 30 percent cheaper than a load +locally, but production does not accept it. + +### M5. At what size does an index beat a scan? + +Load one thousand, ten thousand, and one hundred thousand rows in the shape of the `memory` +table. At each size, time the three-stage retrieval query with and without the BM25 and +vector indexes. Record the size at which an index first beats the scan. Also record +whether the engine accepts a BM25 index beside a plain vector index, and beside a +provider-backed vector index. + +Cloud result, 2026-10-05, `api.hotdata.dev`: no index beat the scan by a clear margin +at any size up to one hundred thousand rows. Each query costs about 400 ms at every size, +with or without indexes. That cost is the floor of one query request. The embeddings have +64 dimensions. Each value is the median of 5 runs after one warm-up, with k 10 and a +fusion depth of 100. + +| Query | 1,000 rows, without / with | 10,000 rows, without / with | 100,000 rows, without / with | +|---|---|---|---| +| Filter scan, newest first | 379 / 423 ms | 414 / 394 ms | 695 / 407 ms | +| Filtered vector rank | 384 / 397 ms | 402 / 639 ms | 405 / 411 ms | +| Unfiltered vector rank | 392 / 393 ms | 374 / 488 ms | 380 / 404 ms | +| Text match scan (`LIKE`) | 368 / 401 ms | 386 / 616 ms | 375 / 451 ms | +| `bm25_search` | refused / 380 ms | refused / 716 ms | refused / 431 ms | +| Fused three-stage query | refused / 422 ms | refused / 601 ms | refused / 462 ms | + +| Index build | 1,000 rows | 10,000 rows | 100,000 rows | +|---|---|---|---| +| BM25 on `content` | 0.7 s | 0.7 s | 0.7 s | +| Plain vector on `embedding`, cosine | 0.7 s | 3.0 s | 65.8 s | +| Sorted on `created_at` | 0.7 s | 0.6 s | 3.1 s | + +The loads took 4.0 s, 7.5 s, and 62.9 s. Each load time includes the upload of the parquet +file. At one hundred thousand rows the file holds about 25 MB of embeddings, so the transfer +dominates. + +These results lead to four conclusions. + +- The only clear gain is the sorted index on `created_at`. It cut the newest-first filter + scan at one hundred thousand rows from 695 ms to 407 ms. +- The vector index gave no gain at any size. A scan of one hundred thousand vectors of 64 + dimensions is already inside the 400 ms floor. +- `bm25_search` refuses to run without a BM25 index, so the fused query cannot run without + one. The fused query cost 422 to 462 ms with indexes, close to the floor. +- The 10,000-row column is slower with indexes in three rows. This is one run, and the + measurement did not repeat it to separate noise from a real cost. + +The engine accepted a BM25 index, a plain vector index, and a sorted index on one table. It +refused a provider-backed vector index beside them, with this error: + +```text +Embedding-backed vector indexes cannot coexist with other indexes on the same table. +``` + +Local result, 2026-10-05, the local stack in [local.md](local.md): a local request costs +about 6 to 10 ms, so the cost of the engine shows. The vector index starts to pay between +one thousand and ten thousand rows, and it pays clearly at one hundred thousand rows. The +settings are the same as in the cloud run. + +| Query | 1,000 rows, without / with | 10,000 rows, without / with | 100,000 rows, without / with | +|---|---|---|---| +| Filter scan, newest first | 8 / 9 ms | 9 / 12 ms | 13 / 15 ms | +| Filtered vector rank | 8 / 9 ms | 21 / 13 ms | 48 / 15 ms | +| Unfiltered vector rank | 7 / 6 ms | 14 / 9 ms | 47 / 6 ms | +| Text match scan (`LIKE`) | 7 / 9 ms | 9 / 12 ms | 20 / 19 ms | +| `bm25_search` | refused / 6 ms | refused / 8 ms | refused / 8 ms | +| Fused three-stage query | refused / 17 ms | refused / 21 ms | refused / 21 ms | + +| Index build | 1,000 rows | 10,000 rows | 100,000 rows | +|---|---|---|---| +| BM25 on `content` | 0.2 s | 0.2 s | 0.2 s | +| Plain vector on `embedding`, cosine | 0.2 s | 1.4 s | 21.9 s | +| Sorted on `created_at` | 0.2 s | 0.2 s | 0.4 s | + +The script polls an index build every 0.2 seconds, so 0.2 s means that the build was done +at the first poll. The loads took 0.2 s or less. + +At one hundred thousand rows, the fused query with indexes took 21 ms, against 48 ms for +the filtered vector scan alone. Thus the three-stage query is faster than a scan at that +size. Locally, the sorted index gave no gain. The stack has no embedding provider, so the +local run did not test a provider-backed index beside the others. + +Across both runs, an index saves engine time from about ten thousand rows. In the cloud, +the cost of one request hides that saving up to at least one hundred thousand rows. + +### M6. Does the driver work against a bare container? + +Against a local container with no API key and no control plane, run each framework call +that the driver needs: create a managed database, declare two tables, load with keyed +upsert and delete, build a BM25 index and a provider-backed vector index, and query. Record +which calls work. If all of them work, the integration tests can run against the container +in CI, in place of a throwaway cloud database. + +Bare container result, 2026-10-05, with the `latest` image pulled on that day (digest +`sha256:302371bb1923`): the first call fails. The container refuses to create a managed +database, with `managed catalogs require ducklake.metadata_pg_url to be configured`. +[local.md](local.md) explains the cause. + +Cloud result, 2026-10-05, `api.hotdata.dev`: every call works. The calls create a +database with two keyed tables and load it in replace, upsert, and delete mode. They build +a BM25 index and run `bm25_search`. They build a provider-backed vector index +(`sys_emb_openai`) and run `vector_search`. +The row count after the loads was 3, as expected. This is the reference for the local +result. + +Local result, 2026-10-05, the local stack in [local.md](local.md), with Postgres and +RustFS beside the engine: every call works except the provider-backed vector index. It +fails with `Embedding provider 'sys_emb_openai' not found`, because the stack configures no +embedding provider. The row count after the loads was 3. So the integration tests can run +against the local stack in CI, with plain vector indexes in place of provider-backed +ones. diff --git a/docs/internal/brief.md b/docs/internal/brief.md index 4015795..ee33686 100644 --- a/docs/internal/brief.md +++ b/docs/internal/brief.md @@ -1,6 +1,7 @@ # Brief: hotmemory, agent memory as tables on Hotdata -Status: draft, revised 2026-10-05, phases moved to `roadmap.md` and `plan.md`. Nothing in this brief is built. It fixes the two +Status: draft, revised 2026-10-05, phases moved to `roadmap.md` and `plan.md`, and the phase 0 +measurements applied (results in `docs/guarantees.md`). Nothing in this brief is built. It fixes the two contracts, the guarantees behind them, and the harness that proves them, before the first line of code. The revision applies section 4 of `survey.md`, in this folder. @@ -46,12 +47,15 @@ for MCP are separate work and are listed under out of scope. The library runs inside the consumer's process. The consumer imports it, and it calls the Hotdata API over HTTPS with the consumer's API key. There is no hotmemory server. The hosted API that production needs already exists, and it is Hotdata's own API: files, loads, -indexes, and query. The same library runs against a RuntimeDB container on a laptop with -three environment variables: `HOTDATA_API_URL` pointing at the container, -`HOTDATA_WORKSPACE` set to any id so the framework skips the workspace listing that a bare -RuntimeDB does not serve, and the container started with -`RUNTIMEDB_AUTH__ALLOW_UNAUTHENTICATED=true`. Whether every managed-table call the driver -makes works in that mode is measurement M6. A separate memory service in front of Hotdata is what section 8 rules +indexes, and query. The same library runs against a local RuntimeDB on a laptop, but +not against the bare engine image. [measured, M6] A managed table needs a Postgres catalog, +because DuckLake metadata lives only in Postgres, and storage that can presign, because the +framework uploads parquet through presigned URLs. The local stack in `docs/local.md` is +three containers: Postgres, RustFS, and the engine sharing the RustFS network namespace so +the presigned address resolves on both sides. The library then needs `HOTDATA_API_URL` +pointing at the engine, `HOTDATA_WORKSPACE` set to any id, and any non-empty +`HOTDATA_API_KEY`. Every driver call works on that stack except a provider-backed vector +index, because no embedding provider is configured locally. A separate memory service in front of Hotdata is what section 8 rules out, and the reasons it can become necessary are listed there. ### 2.2 Platform facts that shape the design @@ -92,9 +96,17 @@ These come from measurements against Hotdata between 2026-08-27 and 2026-09-18. provider, and the query passes text. A metric mismatch in plain mode silently reverts to a full scan. Indexes are invisible to SQL, so only the control plane can say whether one exists. (`hotdata_framework/client.py:584-600`, `hotdata_langchain/vectorstore.py:541`) -- RuntimeDB runs standalone from a container image with a SQLite catalog and filesystem - storage, with no cloud dependency. A local container is a real Hotdata for development - and for measurements that production cannot run. +- A provider-backed vector index cannot share its table with any other index. The engine + refuses the second one. A BM25 index, a plain vector index, and a sorted index share a + table. [measured 2026-08-18, confirmed by M5 on 2026-10-05] +- `bm25_search` refuses to run on a column without a BM25 index. There is no scan + fallback. [measured, M5] +- RuntimeDB runs on a laptop with a Postgres container and an S3-compatible storage + container, with no cloud dependency. The bare image, with its SQLite catalog and + filesystem storage, refuses managed tables. [measured, M6] The local stack is a real + Hotdata for development and for measurements that production cannot run. A load there + costs about 24 ms, against about 2.1 seconds in the cloud, so the cloud write cost is + outside the engine. [measured, M4] ### 2.3 What the design takes from these facts @@ -183,8 +195,9 @@ Two drivers ship in version 1. driver. - `HotdataStore`. One managed database, two tables per schema version, keyed loads, a serialized writer, and the ranking query in section 3.4. The integration leg runs against - it. Against a local RuntimeDB container it is also the development driver, so no third - driver is needed for working offline. + it. Against the local RuntimeDB stack it is also the development driver, so no third + driver is needed for working offline. The integration leg can run against that stack in + CI, with plain vector indexes in place of provider-backed ones. [measured, M6] ### 3.4 Tables and retrieval @@ -205,15 +218,20 @@ Small facts revise cleanly, supersede cleanly, and render into a profile. The su systems that extract agree on the unit: 15 to 80 words in mem0, one edge per fact in Graphiti. Markdown remains a fine format for the `content` column of an episode chunk. -Retrieval runs in one SQL query over the `memory` table, in three stages that the engine -can push down. +Retrieval runs in one SQL query over the `memory` table, in three stages. 1. Exact filters first: namespace labels, validity at the as-of time, `kind`, `subject`, `tags`, and `forget_after`. These are plain predicates and cut the candidate set before any ranking. 2. Three rankings over the survivors, each served by its own index: BM25 over `content`, vector distance over `content`, and vector distance over `cues`. Fused by reciprocal rank - fusion, which the engine already does for text plus semantic search. + fusion, written as plain SQL with common table expressions. The engine has no fusion + primitive and needs none. Because a provider-backed index excludes every other index, + the two vector rankings need plain vector indexes over embedding columns, and so an + embedder supplied by the caller. Without an embedder, the driver can offer BM25 alone or + a provider-backed vector ranking alone, not both. `bm25_search` ranks the whole table + and the filters of stage 1 apply after it, so the BM25 fetch depth must be wide enough + to survive the filters. 3. A sorted index on `created_at` for recency ordering and for the sweeper. `cues` is the part that is not in any surveyed system and is cheap here. A cue is the @@ -223,9 +241,15 @@ content. Cues are optional, and the content-only path always works, so a consume writes no cues loses nothing it had. Fast reads come from the table being small, not from cleverness. A consumer's memory table -holds thousands of rows, and a filtered scan of that is fast before any index exists. The -indexes matter above about a hundred thousand rows, and phase 0 measures the point at -which they start to pay. +holds thousands of rows, and a filtered scan of that is fast before any index exists. M5 +measured the crossover. In the engine, the vector index starts to pay between one thousand +and ten thousand rows, and at one hundred thousand rows it cuts the filtered vector rank +from 48 ms to 15 ms. The fused query with indexes takes 21 ms there. In the cloud, one +request costs about 400 ms at every size, and that floor hides the saving up to at least +one hundred thousand rows. Only the sorted index on `created_at` showed a cloud gain, 695 ms +to 407 ms at one hundred thousand rows. The driver therefore builds the BM25 index whatever +the size, because `bm25_search` needs it, and treats the vector and sorted indexes as an +optimisation that matters from about ten thousand rows. [measured, M5] ## 4. The memory contract @@ -269,11 +293,11 @@ gets a conformance test in phase 1, and `docs/guarantees.md` names the test. | Can a retried `remember` create duplicates? | No. The key is derived from subject and normalized content hash, and a put whose normalized content equals the current revision's writes nothing. | answered | | Is a paraphrase a duplicate? | No. Deduplication is exact on normalized content. `candidates` exists so a caller can decide. | answered | | When a synchronous `put` returns, is the record visible to `list`? | Yes. Read after write was never stale in 30 trials. | answered [measured] | -| When a synchronous `put` returns, is the record visible to `search`? | Yes without an index, because the ranking query scans the table. With an index, unknown. | to measure, M1 | +| When a synchronous `put` returns, is the record visible to `search`? | Yes. Without an index the ranking query scans the table. With a provider-backed vector index or a BM25 index, the first search after the load returned the new row. | answered [measured], M1 | | When a buffered `put` returns, is the record visible? | No. It is visible after the writer flushes, and `writer` returns the ids it flushed. | answered | -| What happens when two processes write the same table? | The second load is refused with 409. The driver retries with backoff and gives up after a bound. Within one process the writer serializes. | answered [measured], retry bound to measure, M3 | +| What happens when two processes write the same table? | The second load is refused with 409. The driver retries with backoff and gives up after a bound. Within one process the writer serializes. | answered [measured], M3: with two concurrent writers no load needed more than 3 of 8 attempts | | What happens when two writers put the same key? | Last writer wins at the row level. Because revisions are new rows, both revisions exist and the later one is current. | answered | -| Does `delete` remove retained revisions and derived embeddings? | Yes for rows, because the embedding is a column of the row. Index entries, unknown. | to measure, M2 | +| Does `delete` remove retained revisions and derived embeddings? | Yes. The embedding is a column of the row, and after a keyed delete neither a provider-backed vector index nor a BM25 index returned the row. | answered [measured], M2 | | Which filters work in `list` and `search`? | Equality on `kind`, `subject`, `tags`, and `actor`. Range on `valid_from`, `valid_until`, `created_at`, and `expired_at`. Prefix on namespace labels. Anything else raises. | answered | | Does a higher score mean more relevant? | The store returns distance, and lower is closer. The memory contract returns records in order and no score. An adapter that needs a score converts. | answered | | What does `recall(as_of=T)` return? | Records valid at T by the as-of rule in section 4, including records superseded after T when `history` is consulted, and excluding them otherwise. | answered | @@ -283,7 +307,9 @@ gets a conformance test in phase 1, and `docs/guarantees.md` names the test. | Does a write inside a turn reach a `recall` in the same turn? | No, by contract. A consumer reads what was there before its own capture. | answered | Phase 0 measurements. M1 to M3 run against a throwaway database that the run creates and -deletes. M4, M5, and M6 run against a local RuntimeDB container. +deletes. M4, M5, and M6 ran against throwaway cloud databases and against the local stack, +because the bare container refuses managed tables. All six were run on 2026-10-05, and +`docs/guarantees.md` holds the numbers. - M1. Build a provider-backed vector index on a table, load ten more rows, and search for one of them. Record whether the new row is served, and after how long. @@ -392,7 +418,7 @@ the API. client in another language, for extraction that must run centrally, or for scope enforcement above the API key. None of the three exists yet. - Server-side extraction. `capture` takes a callable and that is all. -- A third driver. A local RuntimeDB container covers offline development with the Hotdata +- A third driver. The local RuntimeDB stack covers offline development with the Hotdata driver itself. - Access control beyond scope filtering. - A Rust implementation. The Rust client has no framework layer to build on. diff --git a/docs/internal/plan.md b/docs/internal/plan.md index 7056bf8..7b62aa7 100644 --- a/docs/internal/plan.md +++ b/docs/internal/plan.md @@ -8,29 +8,39 @@ phase closes. The phases themselves are in `roadmap.md`. Section numbers below r When this phase closes, a reader can learn the contracts and the guarantees from two public files, a contributor can run one command that checks the documents, an engineer can run the -library's target against a RuntimeDB container on a laptop, and every guarantee marked to -measure in the brief has a number and a date. +library's target against a RuntimeDB container on a laptop, agents working here have one +instruction file, and every guarantee marked to measure in the brief has a number and a +date. ## Tasks -Each task is one issue once the repository has a remote. Each is one commit or a few. +GitHub issue #1 holds these tasks as a checklist. They are worked in order on one branch, +and one pull request closes the issue. 1. `docs/contracts.md`: sections 3 and 4 of the brief in public form, with the measured platform facts kept and no private name. -2. `docs/guarantees.md`: every row of section 5 with its state, and the measurement - placeholders M1 to M6. +2. `docs/guarantees.md`: every row of section 5 with its state, and placeholders M1 to M6. 3. `README.md` revised to point at both files, and `CONTRIBUTING.md` with the one command - and the rules from section 6. -4. `Makefile` with `verify` running the link check over `README.md` and `docs/`, which is - the only documentation check that exists before code does. -5. `docs/local.md` and a `make local-up` target: the container command from the RuntimeDB - README, the three environment variables, and how to point the library at it. -6. `scripts/measure_cloud.py`: creates a throwaway database, runs M1, M2, and M3, prints - the numbers, deletes the database. -7. `scripts/measure_local.py`: starts the container with the SQL writes flag on, runs M4, - M5, and M6, prints the numbers. -8. The numbers and dates written into `docs/guarantees.md`, and any guarantee that the - numbers contradict rewritten in the brief. + and the rules from section 6. The make targets are named `verify` and `local-up`. +4. `Makefile` with `verify` running a link check over `README.md` and `docs/`, and a + minimal `pyproject.toml` holding only the development tools so `verify` has something + to run. +5. `docs/local.md` and `make local-up`: the container command from the RuntimeDB README, + the three environment variables (`HOTDATA_API_URL`, `HOTDATA_WORKSPACE`, and the + container's `RUNTIMEDB_AUTH__ALLOW_UNAUTHENTICATED=true`), and how to point the library + at it. +6. `scripts/measure_cloud.py`: creates a throwaway database named by an environment + variable, runs M1, M2, and M3, prints the numbers, deletes the database. Inline script + metadata (PEP 723) so `uv run scripts/measure_cloud.py` works on its own. +7. `scripts/measure_local.py`: against a running container whose URL comes from the + environment, runs M4, M5, and M6, prints the numbers. Same inline metadata. The script + does not start the container. +8. `AGENTS.md` with the commands, the rules, and pointers to `plan.md`, `roadmap.md`, + `contracts.md`, and `guarantees.md`, restating nothing they say. `CLAUDE.md` is one + `@AGENTS.md` line plus any Claude-only note. +9. The numbers and dates written into `docs/guarantees.md`, and any guarantee the numbers + contradict rewritten in the brief. The scripts are run by the repository owner: the + cloud script needs an API key and a workspace, the local script needs Docker Desktop. ## Acceptance criteria @@ -43,6 +53,9 @@ Each task is one issue once the repository has a remote. Each is one commit or a control plane. Proven by the M6 script's first call. - AC4. No file under `docs/` or at the root names a private repository, a customer, or a deployment detail. Proven by a grep for the known names, recorded in the pull request. +- AC5. `AGENTS.md` and `CLAUDE.md` exist, and `AGENTS.md` contains no sentence that + `plan.md`, `roadmap.md`, `contracts.md`, or `guarantees.md` already contains. Proven by + review. ## Verification diff --git a/docs/local.md b/docs/local.md new file mode 100644 index 0000000..2457e91 --- /dev/null +++ b/docs/local.md @@ -0,0 +1,82 @@ +# Run against a local RuntimeDB + +RuntimeDB is the Hotdata query engine. The library can use a local RuntimeDB in place of a +cloud workspace, with no API key and no cloud service. This page tells you how to start it +and how to point the library at it. + +## Start the stack + +You need Docker Desktop. Run this command from the repository root: + +```sh +make local-up +``` + +The target starts three containers from [compose.yaml](../compose.yaml): + +| Service | Image | Purpose | +|---|---|---| +| `catalog` | `postgres:17` | Holds the RuntimeDB catalog and the DuckLake metadata of each managed table. | +| `storage` | `rustfs/rustfs:latest` | S3-compatible object storage for the table data and for uploads. | +| `runtimedb` | `ghcr.io/hotdata-dev/runtimedb:latest` | The engine, on port 3000. | + +The target creates the `runtimedb` bucket in storage, starts the engine, and waits until +`GET /health` answers. To stop the stack and delete all of its data, run `make local-down`. +To download newer images, run `make local-pull`. + +To use a different port or engine image, set the make variables: + +```sh +make local-up LOCAL_PORT=3100 RUNTIMEDB_IMAGE=runtimedb:local +``` + +Port 9000 on the host must be free, because storage always uses it. + +Warning: do not expose this stack to a network. The engine accepts every request with no +token, and the storage credentials are fixed and public. + +## Why three containers + +A managed table needs two services that the bare engine image does not have. + +- A Postgres catalog. RuntimeDB keeps the DuckLake metadata of a managed table in Postgres. + DuckLake metadata records which parquet files make up each version of the table. With + a Postgres `[catalog]`, the engine derives the DuckLake address from it. The bare image + uses a SQLite catalog, so it refuses to create a managed database with + `managed catalogs require ducklake.metadata_pg_url to be configured`. +- Object storage that can presign. The framework uploads parquet through presigned URLs. + A presigned URL is a storage address, signed by the engine, to which the client sends + the file directly. Filesystem storage cannot presign, and a load from the bare image + fails with `PRESIGN_UNSUPPORTED`. + +A presigned URL carries the host name of the storage endpoint. The engine and the client +on the host must reach the same name. The `runtimedb` service shares the network namespace +of `storage`, so `127.0.0.1:9000` is the storage service both inside the engine and on the +host. + +## Point the library at it + +Set these three variables in the shell that runs the library: + +```sh +export HOTDATA_API_URL=http://localhost:3000 +export HOTDATA_WORKSPACE=local +export HOTDATA_API_KEY=local +``` + +- `HOTDATA_API_URL` sends each API call to the local engine. +- `HOTDATA_WORKSPACE` can be any value. If it is set, the framework does not ask for the + workspace list, which the local engine does not serve. +- `HOTDATA_API_KEY` can be any value that is not empty. The engine does not read it, but + `HotdataClient.from_env()` in `hotdata-framework` refuses an empty key. This statement comes from the framework source. + +`scripts/measure_local.py` needs none of these. It reads `HOTMEMORY_LOCAL_URL`, which +defaults to `http://localhost:3000`. + +## What works locally + +Measurement M6 in [guarantees.md](guarantees.md) records each call. On 2026-10-05, every +call that the driver needs worked against this stack, except one. A provider-backed vector +index fails with `Embedding provider 'sys_emb_openai' not found`, because the stack +configures no embedding provider. A plain vector index, over embeddings that the caller +supplies, works. diff --git a/pyproject.toml b/pyproject.toml new file mode 100644 index 0000000..b900a75 --- /dev/null +++ b/pyproject.toml @@ -0,0 +1,20 @@ +[project] +name = "hotmemory" +version = "0.0.0" +description = "Agent memory as tables on Hotdata" +requires-python = ">=3.11" +license = { text = "MIT" } +dependencies = [] + +[dependency-groups] +dev = ["ruff"] + +[tool.uv] +package = false + +[tool.ruff] +line-length = 100 +target-version = "py311" + +[tool.ruff.lint] +select = ["E", "F", "I", "UP", "B", "SIM"] diff --git a/scripts/check_links.py b/scripts/check_links.py new file mode 100644 index 0000000..0f355c3 --- /dev/null +++ b/scripts/check_links.py @@ -0,0 +1,65 @@ +"""Fail when a relative Markdown link points at a file that does not exist. + +Checks every Markdown file at the repository root and under docs/. A link +with a scheme (https:, mailto:) or a bare anchor is skipped. Links inside fenced +code blocks and inline code spans are ignored. Prints one line per broken link +and exits 1 when any is found. +""" + +from __future__ import annotations + +import re +import sys +from pathlib import Path + +ROOT = Path(__file__).resolve().parent.parent + +FENCE = re.compile(r"^\s*(```|~~~)") +CODE_SPAN = re.compile(r"`[^`]*`") +LINK = re.compile(r"!?\[[^\]]*\]\(\s*]+)>?(?:\s+\"[^\"]*\")?\s*\)") +SCHEME = re.compile(r"^[a-zA-Z][a-zA-Z0-9+.-]*:") + + +def markdown_files() -> list[Path]: + files = sorted(ROOT.glob("*.md")) + files.extend(sorted((ROOT / "docs").rglob("*.md"))) + return files + + +def links(path: Path) -> list[tuple[int, str]]: + found: list[tuple[int, str]] = [] + in_fence = False + for number, line in enumerate(path.read_text(encoding="utf-8").splitlines(), start=1): + if FENCE.match(line): + in_fence = not in_fence + continue + if in_fence: + continue + for match in LINK.finditer(CODE_SPAN.sub("", line)): + found.append((number, match.group(1))) + return found + + +def broken(path: Path) -> list[str]: + problems: list[str] = [] + for number, target in links(path): + if SCHEME.match(target) or target.startswith("#"): + continue + file_part = target.split("#", 1)[0] + if not (path.parent / file_part).exists(): + problems.append(f"{path.relative_to(ROOT)}:{number}: broken link {target}") + return problems + + +def main() -> int: + problems = [problem for path in markdown_files() for problem in broken(path)] + for problem in problems: + print(problem) + if problems: + print(f"{len(problems)} broken link(s)") + return 1 + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/measure_cloud.py b/scripts/measure_cloud.py new file mode 100644 index 0000000..8a71235 --- /dev/null +++ b/scripts/measure_cloud.py @@ -0,0 +1,373 @@ +# /// script +# requires-python = ">=3.11" +# dependencies = [ +# "hotdata-framework>=0.14.1", +# "pyarrow>=14.0", +# ] +# /// +"""Run measurements M1, M2, and M3 against a throwaway Hotdata cloud database. + +Reads the connection from the environment: + +- HOTDATA_API_KEY: required. +- HOTDATA_WORKSPACE: optional. Without it, the framework picks the active workspace. +- HOTDATA_API_URL: optional. Defaults to https://api.hotdata.dev. +- HOTMEMORY_MEASURE_DB: required. The name of the throwaway database. The script + refuses to run when a database with this name already exists, creates it, and + deletes it on exit. +- HOTMEMORY_EMBEDDING_PROVIDER: optional. The provider for the provider-backed + vector index. Defaults to sys_emb_openai. + +M1 and M2 run once against a provider-backed vector index and once against a BM25 +index, on separate tables, because a provider-backed index cannot share a table with +another index. M3 runs two processes that load one table at the same time. + +Prints the results as Markdown, ready to paste into docs/guarantees.md. +""" + +from __future__ import annotations + +import multiprocessing as mp +import os +import sys +import tempfile +import time +from dataclasses import dataclass, field +from datetime import UTC, datetime, timedelta +from pathlib import Path +from typing import Any + +import pyarrow as pa +import pyarrow.parquet as pq +from hotdata_framework import HotdataClient, ManagedDatabase +from hotdata_framework.env import default_api_key, default_host, pick_workspace + +SCHEMA = "public" +VECTOR_TABLE = "m1_vector" +BM25_TABLE = "m1_bm25" +LOCK_TABLE = "m3_lock" +BASE_ROWS = 50 +NEW_ROWS = 10 +SERVE_TIMEOUT_S = 120.0 +SERVE_POLL_S = 2.0 +FORGET_WAIT_S = 30.0 + +M3_LOADS_PER_PROCESS = 10 +M3_MAX_ATTEMPTS = 8 +M3_BACKOFF_BASE_S = 0.25 +M3_BACKOFF_CAP_S = 4.0 + +SERVICES = ["billing", "checkout", "search", "ingest", "auth", "ledger", "notify", "export"] +SYMPTOMS = [ + "disk filled on the primary volume", + "connection pool ran out of slots", + "certificate expired at midnight", + "a deploy shipped a bad feature flag", + "the upstream DNS resolver timed out", + "memory grew until the pod was killed", +] +NEW_FACTS = [ + "The aurora gateway rejects tokens signed with the retired zebra key", + "Saturn queue consumers stall when the walrus partition rebalances", + "The marigold cron job double-bills invoices on leap days", + "Penguin cache entries outlive their tenant after a quokka migration", + "The obsidian exporter drops rows whose payload exceeds the tangerine limit", + "Heron workers deadlock when the lantern lock is taken twice", + "The cobalt webhook retries forever after a pelican timeout", + "Falcon shards lose writes when the meadow replica lags", + "The juniper scheduler skips jobs during the otter maintenance window", + "Ivory sessions survive logout when the badger flag is on", +] + + +def env(name: str, default: str | None = None) -> str: + value = os.environ.get(name, default) + if not value: + sys.exit(f"{name} must be set") + return value + + +def now() -> str: + return datetime.now(UTC).strftime("%Y-%m-%d %H:%M:%S UTC") + + +def base_contents() -> list[str]: + return [ + f"{SERVICES[i % len(SERVICES)]} paged because the {SYMPTOMS[i % len(SYMPTOMS)]} " + f"(incident {i})" + for i in range(BASE_ROWS) + ] + + +def write_parquet(directory: Path, name: str, rows: dict[str, list[Any]]) -> str: + path = directory / f"{name}.parquet" + pq.write_table(pa.table(rows), path) + return str(path) + + +def connect() -> HotdataClient: + api_key = env("HOTDATA_API_KEY", default_api_key()) + host = default_host() + return HotdataClient(api_key, pick_workspace(api_key, host), host=host) + + +def is_locked(error: BaseException) -> bool: + cause = error.__cause__ + return getattr(cause, "status", None) == 409 or "RESOURCE_LOCKED" in str(error) + + +def table_ref(table: str) -> str: + return f"default.{SCHEMA}.{table}" + + +def search_ids( + client: HotdataClient, db: ManagedDatabase, kind: str, table: str, text: str +) -> list[str]: + function = "vector_search" if kind == "vector" else "bm25_search" + order = "_distance ASC" if kind == "vector" else "score DESC" + quoted = text.replace("'", "''") + sql = ( + f"SELECT id FROM {function}('{table_ref(table)}', 'content', '{quoted}', 5) " + f"ORDER BY {order}" + ) + result = client.execute_sql(sql, database=db) + return [str(row[0]) for row in result.rows] + + +@dataclass +class IndexResult: + kind: str + build_s: float = 0.0 + load_s: float = 0.0 + served: bool = False + served_after_s: float | None = None + deleted_row_count: int | None = None + deleted_returned_at_once: bool | None = None + deleted_returned_after_wait: bool | None = None + errors: list[str] = field(default_factory=list) + + +def measure_index( + client: HotdataClient, db: ManagedDatabase, kind: str, table: str, work: Path, provider: str +) -> IndexResult: + result = IndexResult(kind=kind) + contents = base_contents() + ids = [f"base-{i}" for i in range(BASE_ROWS)] + client.load_managed_table( + db, + table, + file=write_parquet(work, f"{table}_base", {"id": ids, "content": contents}), + mode="replace", + ) + + started = time.perf_counter() + try: + if kind == "vector": + client.create_index( + db, + table, + columns=["content"], + index_type="vector", + embedding_provider_id=provider, + ) + else: + client.create_index(db, table, columns=["content"], index_type="bm25") + except (RuntimeError, TimeoutError) as e: + result.errors.append(f"index build: {e}") + return result + result.build_s = time.perf_counter() - started + + new_ids = [f"new-{i}" for i in range(NEW_ROWS)] + started = time.perf_counter() + client.load_managed_table( + db, + table, + file=write_parquet(work, f"{table}_new", {"id": new_ids, "content": NEW_FACTS}), + mode="upsert", + key=["id"], + ) + result.load_s = time.perf_counter() - started + + # M1: is a row loaded after the build served by the index? + target_id, target_text = new_ids[3], NEW_FACTS[3] + loaded_at = time.perf_counter() + while time.perf_counter() - loaded_at < SERVE_TIMEOUT_S: + try: + if target_id in search_ids(client, db, kind, table, target_text): + result.served = True + result.served_after_s = time.perf_counter() - loaded_at + break + except RuntimeError as e: + result.errors.append(f"M1 search: {e}") + break + time.sleep(SERVE_POLL_S) + + # M2: does the index still return a row after a keyed delete? + victim_id, victim_text = new_ids[5], NEW_FACTS[5] + try: + client.load_managed_table( + db, + table, + file=write_parquet(work, f"{table}_delete", {"id": [victim_id]}), + mode="delete", + key=["id"], + ) + except RuntimeError as e: + result.errors.append(f"M2 delete load: {e}") + return result + count = client.execute_sql( + f"SELECT count(*) FROM {table_ref(table)} WHERE id = '{victim_id}'", database=db + ) + result.deleted_row_count = int(count.rows[0][0]) + try: + result.deleted_returned_at_once = victim_id in search_ids( + client, db, kind, table, victim_text + ) + time.sleep(FORGET_WAIT_S) + result.deleted_returned_after_wait = victim_id in search_ids( + client, db, kind, table, victim_text + ) + except RuntimeError as e: + result.errors.append(f"M2 search: {e}") + return result + + +@dataclass +class WriterResult: + process: int + loads_ok: int = 0 + loads_given_up: int = 0 + locked_responses: int = 0 + max_attempts_used: int = 0 + elapsed_s: float = 0.0 + errors: list[str] = field(default_factory=list) + + +def serialized_load( + client: HotdataClient, db: ManagedDatabase, path: str, stats: WriterResult +) -> None: + for attempt in range(1, M3_MAX_ATTEMPTS + 1): + try: + client.load_managed_table(db, LOCK_TABLE, file=path, mode="upsert", key=["id"]) + stats.loads_ok += 1 + stats.max_attempts_used = max(stats.max_attempts_used, attempt) + return + except RuntimeError as e: + if not is_locked(e): + stats.errors.append(str(e)) + return + stats.locked_responses += 1 + time.sleep(min(M3_BACKOFF_CAP_S, M3_BACKOFF_BASE_S * 2 ** (attempt - 1))) + stats.loads_given_up += 1 + stats.max_attempts_used = M3_MAX_ATTEMPTS + + +def writer_process(process: int, db_id: str, barrier: Any, results: Any) -> None: + client = connect() + db = client.resolve_managed_database(db_id) + stats = WriterResult(process=process) + with tempfile.TemporaryDirectory() as tmp: + paths = [ + write_parquet( + Path(tmp), + f"p{process}_{n}", + {"id": [f"p{process}-{n}"], "content": [f"writer {process} load {n}"]}, + ) + for n in range(M3_LOADS_PER_PROCESS) + ] + barrier.wait(timeout=60) + started = time.perf_counter() + for path in paths: + serialized_load(client, db, path, stats) + stats.elapsed_s = time.perf_counter() - started + results.put(stats) + + +def measure_lock(client: HotdataClient, db: ManagedDatabase, work: Path) -> list[WriterResult]: + client.load_managed_table( + db, + LOCK_TABLE, + file=write_parquet(work, "lock_seed", {"id": ["seed"], "content": ["seed"]}), + mode="replace", + ) + context = mp.get_context("spawn") + barrier = context.Barrier(2) + results = context.Queue() + processes = [ + context.Process(target=writer_process, args=(n, db.id, barrier, results), daemon=True) + for n in (1, 2) + ] + for process in processes: + process.start() + collected = [results.get(timeout=600) for _ in processes] + for process in processes: + process.join() + return sorted(collected, key=lambda r: r.process) + + +def report_index(label: str, r: IndexResult) -> None: + print( + f"- {label} index ({r.kind}): build {r.build_s:.1f} s, " + f"load of {NEW_ROWS} rows {r.load_s:.1f} s." + ) + if r.served: + print(f" - M1: the new row was served {r.served_after_s:.1f} s after the load returned.") + else: + print(f" - M1: the new row was not served within {SERVE_TIMEOUT_S:.0f} s.") + print( + f" - M2: after the delete, a scan counts {r.deleted_row_count} row(s) for the id. " + f"The index returned it at once: {r.deleted_returned_at_once}. " + f"After {FORGET_WAIT_S:.0f} s: {r.deleted_returned_after_wait}." + ) + for error in r.errors: + print(f" - error: {error}") + + +def main() -> int: + name = env("HOTMEMORY_MEASURE_DB") + provider = os.environ.get("HOTMEMORY_EMBEDDING_PROVIDER", "sys_emb_openai") + client = connect() + if any(d.description == name for d in client.list_managed_databases()): + sys.exit(f"a database named {name!r} already exists; pick a new HOTMEMORY_MEASURE_DB") + + started_at = now() + expires = (datetime.now(UTC) + timedelta(hours=1)).isoformat() + tables = [VECTOR_TABLE, BM25_TABLE, LOCK_TABLE] + db = client.create_managed_database( + name, tables=tables, keys={t: ["id"] for t in tables}, expires_at=expires + ) + print(f"created database {name} ({db.id}) on {client.host}", file=sys.stderr) + try: + with tempfile.TemporaryDirectory() as tmp: + work = Path(tmp) + vector = measure_index(client, db, "vector", VECTOR_TABLE, work, provider) + bm25 = measure_index(client, db, "bm25", BM25_TABLE, work, provider) + writers = measure_lock(client, db, work) + finally: + client.delete_managed_database(db) + print(f"deleted database {name} ({db.id})", file=sys.stderr) + + print(f"## Cloud measurements, {started_at}") + print() + print(f"Host: {client.host}. Embedding provider: {provider}.") + print() + report_index("Provider-backed vector", vector) + report_index("BM25", bm25) + print( + f"- M3: two processes, {M3_LOADS_PER_PROCESS} single-row upsert loads each, " + f"retry bound {M3_MAX_ATTEMPTS} attempts with backoff {M3_BACKOFF_BASE_S} s doubling " + f"to {M3_BACKOFF_CAP_S} s." + ) + for w in writers: + print( + f" - process {w.process}: {w.loads_ok} loaded, {w.loads_given_up} gave up, " + f"{w.locked_responses} responses were 409, most attempts for one load " + f"{w.max_attempts_used}, {w.elapsed_s:.1f} s in total." + ) + for error in w.errors: + print(f" - error: {error}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/measure_local.py b/scripts/measure_local.py new file mode 100644 index 0000000..6ab174e --- /dev/null +++ b/scripts/measure_local.py @@ -0,0 +1,552 @@ +# /// script +# requires-python = ">=3.11" +# dependencies = [ +# "hotdata-framework>=0.14.1", +# "pyarrow>=14.0", +# ] +# /// +"""Run measurements M6, M4, and M5 against a local RuntimeDB container or the cloud. + +By default the script targets a running local container. It does not start the +container. Start it with `make local-up`, which sets RUNTIMEDB_ENGINE__SQL_WRITES=true +for M4. With --cloud, it targets the Hotdata cloud API instead. + +Reads the connection from the environment: + +- HOTMEMORY_LOCAL_URL: optional, local only. The container URL. Defaults to + http://localhost:3000. Locally the script ignores HOTDATA_API_URL, HOTDATA_API_KEY, + and HOTDATA_WORKSPACE, and sends the placeholder key and workspace "local", so a + .env that holds cloud credentials never sends them to the container. +- HOTDATA_API_KEY: required with --cloud. +- HOTDATA_WORKSPACE: optional with --cloud. Without it, the framework picks the + active workspace. +- HOTDATA_API_URL: optional with --cloud. Defaults to https://api.hotdata.dev. +- HOTMEMORY_MEASURE_DB: required with --cloud. The prefix for the throwaway database + names; each measurement appends -m4, -m5, or -m6. The script refuses to run when a + database with one of those names already exists. +- HOTMEMORY_EMBEDDING_PROVIDER: optional. The provider that M6 and M5 try for a + provider-backed vector index. Defaults to sys_emb_openai. +- HOTMEMORY_M5_SIZES: optional. Comma-separated row counts for M5. Defaults to + 1000,10000,100000. + +M6 runs first. Locally, its first two calls are the proof that the container answers +a query with no API key and no control plane. Each measurement creates its own +database and deletes it on exit. + +Prints the results as Markdown, ready to paste into docs/guarantees.md. +""" + +from __future__ import annotations + +import os +import random +import statistics +import sys +import tempfile +import time +from collections.abc import Callable +from datetime import UTC, datetime, timedelta +from pathlib import Path +from typing import Any + +import pyarrow as pa +import pyarrow.parquet as pq +from hotdata_framework import HotdataClient, ManagedDatabase +from hotdata_framework.env import default_api_key, default_host, pick_workspace + +SCHEMA = "public" +M4_CALLS = 100 +M5_DIMENSIONS = 64 +M5_RUNS = 5 +M5_DEPTH = 100 +M5_K = 10 +RRF_K = 60 + +WORDS = [ + "alert", + "billing", + "checkout", + "disk", + "pool", + "certificate", + "deploy", + "flag", + "resolver", + "memory", + "pod", + "queue", + "shard", + "replica", + "latency", + "timeout", + "retry", + "webhook", + "cache", + "tenant", + "session", + "token", + "gateway", + "scheduler", + "exporter", + "partition", + "lock", + "index", + "ingest", + "ledger", + "auth", + "notify", + "search", + "region", + "rollback", + "migration", + "quota", + "throttle", + "backlog", + "cron", + "invoice", + "payload", + "schema", +] +NAMESPACES = ["team/payments", "team/platform", "team/search", "team/identity"] +KINDS = ["fact", "fact", "fact", "profile", "procedure"] + + +def env(name: str, default: str | None = None) -> str: + value = os.environ.get(name, default) + if not value: + sys.exit(f"{name} must be set") + return value + + +def now() -> str: + return datetime.now(UTC).strftime("%Y-%m-%d %H:%M:%S UTC") + + +def connect() -> HotdataClient: + return HotdataClient("local", "local", host=env("HOTMEMORY_LOCAL_URL", "http://localhost:3000")) + + +def connect_cloud() -> HotdataClient: + api_key = env("HOTDATA_API_KEY", default_api_key()) + host = default_host() + return HotdataClient(api_key, pick_workspace(api_key, host), host=host) + + +def write_parquet(directory: Path, name: str, table: pa.Table) -> str: + path = directory / f"{name}.parquet" + pq.write_table(table, path) + return str(path) + + +def table_ref(table: str) -> str: + return f"default.{SCHEMA}.{table}" + + +def vector_literal(values: list[float]) -> str: + return "ARRAY[" + ", ".join(repr(float(v)) for v in values) + "]" + + +def timed(fn: Callable[[], Any]) -> float: + started = time.perf_counter() + fn() + return time.perf_counter() - started + + +def short(error: BaseException) -> str: + return " ".join(str(error).split())[:300] + + +# M6: every framework call the driver needs, against the bare container. + + +def measure_bare( + client: HotdataClient, work: Path, provider: str, prefix: str +) -> list[tuple[str, str]]: + steps: list[tuple[str, str]] = [] + db: ManagedDatabase | None = None + + def step(label: str, fn: Callable[[], Any]) -> Any: + try: + value = fn() + except Exception as e: # noqa: BLE001 + steps.append((label, f"fails: {short(e)}")) + return None + steps.append((label, "works")) + return value + + db = step( + "create a managed database with two keyed tables", + lambda: client.create_managed_database( + f"{prefix}-m6", tables=["memory", "episode"], keys={"memory": ["id"], "episode": ["id"]} + ), + ) + if db is None: + return steps + try: + step("query SELECT 1", lambda: client.execute_sql("SELECT 1", database=db)) + rows = pa.table({"id": ["a", "b", "c"], "content": ["disk full", "pool empty", "cert"]}) + step( + "load in replace mode", + lambda: client.load_managed_table( + db, "memory", file=write_parquet(work, "m6_replace", rows), mode="replace" + ), + ) + upsert = pa.table({"id": ["c", "d"], "content": ["certificate expired", "dns timeout"]}) + step( + "load in upsert mode with key id", + lambda: client.load_managed_table( + db, + "memory", + file=write_parquet(work, "m6_upsert", upsert), + mode="upsert", + key=["id"], + ), + ) + step( + "load in delete mode with key id", + lambda: client.load_managed_table( + db, + "memory", + file=write_parquet(work, "m6_delete", pa.table({"id": ["a"]})), + mode="delete", + key=["id"], + ), + ) + count = step( + "query the loaded table", + lambda: client.execute_sql(f"SELECT count(*) FROM {table_ref('memory')}", database=db), + ) + if count is not None: + steps.append(("row count after the loads (expected 3)", str(count.rows[0][0]))) + step( + "build a BM25 index", + lambda: client.create_index(db, "memory", columns=["content"], index_type="bm25"), + ) + step( + "query bm25_search", + lambda: client.execute_sql( + f"SELECT id, score FROM bm25_search('{table_ref('memory')}', 'content', " + f"'certificate', 5) ORDER BY score DESC", + database=db, + ), + ) + step( + "load the episode table", + lambda: client.load_managed_table( + db, "episode", file=write_parquet(work, "m6_episode", rows), mode="replace" + ), + ) + step( + f"build a provider-backed vector index ({provider})", + lambda: client.create_index( + db, + "episode", + columns=["content"], + index_type="vector", + embedding_provider_id=provider, + timeout_s=120, + ), + ) + step( + "query vector_search", + lambda: client.execute_sql( + f"SELECT id, _distance FROM vector_search('{table_ref('episode')}', 'content', " + f"'certificate', 5) ORDER BY _distance ASC", + database=db, + ), + ) + finally: + step("delete the managed database", lambda: client.delete_managed_database(db)) + return steps + + +# M4: SQL INSERT against upload-and-load, one row per call. + + +def measure_insert(client: HotdataClient, work: Path, prefix: str) -> dict[str, Any]: + out: dict[str, Any] = {} + db = client.create_managed_database(f"{prefix}-m4", tables=["by_load", "by_sql"]) + try: + seed = pa.table({"id": ["seed"], "content": ["seed"]}) + for table in ("by_load", "by_sql"): + client.load_managed_table( + db, table, file=write_parquet(work, f"m4_{table}", seed), mode="replace" + ) + + sql_times: list[float] = [] + for n in range(M4_CALLS): + sql = f"INSERT INTO {table_ref('by_sql')} VALUES ('sql-{n}', 'row {n}')" + try: + sql_times.append(timed(lambda sql=sql: client.execute_sql(sql, database=db))) + except RuntimeError as e: + out["sql_error"] = short(e) + break + out["sql"] = sql_times + + paths = [ + write_parquet( + work, f"m4_load_{n}", pa.table({"id": [f"load-{n}"], "content": [f"row {n}"]}) + ) + for n in range(M4_CALLS) + ] + load_times: list[float] = [] + for path in paths: + load_times.append( + timed( + lambda path=path: client.load_managed_table( + db, "by_load", file=path, mode="append" + ) + ) + ) + out["load"] = load_times + + for table in ("by_load", "by_sql"): + result = client.execute_sql(f"SELECT count(*) FROM {table_ref(table)}", database=db) + out[f"{table}_rows"] = int(result.rows[0][0]) + finally: + client.delete_managed_database(db) + return out + + +# M5: the three-stage retrieval query at three sizes, with and without indexes. + + +def memory_rows(size: int, rng: random.Random) -> pa.Table: + base = datetime(2026, 1, 1, tzinfo=UTC) + created = [base + timedelta(minutes=i) for i in range(size)] + vectors = [[rng.uniform(-1.0, 1.0) for _ in range(M5_DIMENSIONS)] for _ in range(size)] + return pa.table( + { + "id": [f"r{i}" for i in range(size)], + "namespace": [NAMESPACES[i % len(NAMESPACES)] for i in range(size)], + "key": [f"k{i}" for i in range(size)], + "revision": pa.array([1] * size, pa.int64()), + "kind": [KINDS[i % len(KINDS)] for i in range(size)], + "subject": [f"service-{i % 50}" for i in range(size)], + "content": [" ".join(rng.choices(WORDS, k=12)) for _ in range(size)], + "tags": [[WORDS[i % len(WORDS)]] for i in range(size)], + "actor": ["loader"] * size, + "created_at": pa.array(created, pa.timestamp("us", tz="UTC")), + "valid_from": pa.array(created, pa.timestamp("us", tz="UTC")), + "valid_until": pa.array([None] * size, pa.timestamp("us", tz="UTC")), + "superseded_by": pa.array([None] * size, pa.string()), + "forget_after": pa.array([None] * size, pa.timestamp("us", tz="UTC")), + "embedding": pa.array(vectors, pa.list_(pa.float32())), + } + ) + + +FILTERS = ( + "(namespace = 'team/payments' OR namespace LIKE 'team/payments/%') " + "AND kind = 'fact' AND superseded_by IS NULL " + "AND (valid_from IS NULL OR valid_from <= TIMESTAMP '2027-01-01 00:00:00') " + "AND (valid_until IS NULL OR valid_until > TIMESTAMP '2027-01-01 00:00:00') " + "AND (forget_after IS NULL OR forget_after > TIMESTAMP '2027-01-01 00:00:00')" +) + + +def queries(table: str, word: str, vector: list[float]) -> dict[str, str]: + ref = table_ref(table) + lit = vector_literal(vector) + fused = ( + f"WITH near AS (SELECT id, cosine_distance(embedding, {lit}) AS d FROM {ref} " + f"ORDER BY d ASC LIMIT {M5_DEPTH}), " + f"near_ranked AS (SELECT id, ROW_NUMBER() OVER (ORDER BY d ASC) AS r FROM near), " + f"text_ranked AS (SELECT id, ROW_NUMBER() OVER (ORDER BY score DESC) AS r " + f"FROM bm25_search('{ref}', 'content', '{word}', {M5_DEPTH})), " + f"fused AS (SELECT COALESCE(n.id, t.id) AS id, " + f"COALESCE(1.0 / ({RRF_K} + n.r), 0) + COALESCE(1.0 / ({RRF_K} + t.r), 0) AS s " + f"FROM near_ranked n FULL OUTER JOIN text_ranked t ON n.id = t.id) " + f"SELECT b.id, f.s FROM fused f JOIN {ref} b ON b.id = f.id " + f"WHERE {FILTERS} ORDER BY f.s DESC, b.id ASC LIMIT {M5_K}" + ) + return { + "filter scan": ( + f"SELECT id FROM {ref} WHERE {FILTERS} ORDER BY created_at DESC LIMIT {M5_K}" + ), + "filtered vector rank": ( + f"SELECT id, cosine_distance(embedding, {lit}) AS d FROM {ref} " + f"WHERE {FILTERS} ORDER BY d ASC LIMIT {M5_K}" + ), + "unfiltered vector rank": ( + f"SELECT id, cosine_distance(embedding, {lit}) AS d FROM {ref} " + f"ORDER BY d ASC LIMIT {M5_K}" + ), + "text match scan": ( + f"SELECT id FROM {ref} WHERE {FILTERS} AND content LIKE '%{word}%' " + f"ORDER BY created_at DESC LIMIT {M5_K}" + ), + "bm25_search": ( + f"SELECT id, score FROM bm25_search('{ref}', 'content', '{word}', {M5_K}) " + f"ORDER BY score DESC" + ), + "fused three-stage": fused, + } + + +def median_time(client: HotdataClient, db: ManagedDatabase, sql: str) -> str: + try: + client.execute_sql(sql, database=db) + runs = [timed(lambda: client.execute_sql(sql, database=db)) for _ in range(M5_RUNS)] + except RuntimeError as e: + return f"fails: {short(e)}" + return f"{statistics.median(runs) * 1000:.0f} ms" + + +def measure_scale( + client: HotdataClient, work: Path, sizes: list[int], provider: str, prefix: str +) -> dict: + rng = random.Random(7) + tables = [f"m5_{size}" for size in sizes] + out: dict[str, Any] = {"sizes": {}, "coexistence": []} + db = client.create_managed_database(f"{prefix}-m5", tables=tables) + try: + for size, table in zip(sizes, tables, strict=True): + entry: dict[str, Any] = {} + data = memory_rows(size, rng) + path = write_parquet(work, table, data) + entry["load_s"] = timed( + lambda path=path, table=table: client.load_managed_table( + db, table, file=path, mode="replace" + ) + ) + vector = [rng.uniform(-1.0, 1.0) for _ in range(M5_DIMENSIONS)] + sqls = queries(table, "certificate", vector) + entry["before"] = {name: median_time(client, db, sql) for name, sql in sqls.items()} + + builds: dict[str, str] = {} + for label, kwargs in ( + ("bm25 on content", {"columns": ["content"], "index_type": "bm25"}), + ( + "plain vector on embedding (cosine)", + {"columns": ["embedding"], "index_type": "vector", "metric": "cosine"}, + ), + ("sorted on created_at", {"columns": ["created_at"], "index_type": "sorted"}), + ): + try: + seconds = timed( + lambda kwargs=kwargs, table=table: client.create_index( + db, table, timeout_s=1800, poll_interval_s=0.2, **kwargs + ) + ) + builds[label] = f"{seconds:.1f} s" + except Exception as e: # noqa: BLE001 + builds[label] = f"fails: {short(e)}" + entry["builds"] = builds + entry["after"] = {name: median_time(client, db, sql) for name, sql in sqls.items()} + out["sizes"][size] = entry + + first = tables[0] + try: + client.create_index( + db, + first, + columns=["subject"], + index_type="vector", + embedding_provider_id=provider, + timeout_s=120, + ) + out["coexistence"].append( + ("provider-backed vector beside BM25 and plain vector", "accepted") + ) + except Exception as e: # noqa: BLE001 + out["coexistence"].append( + ("provider-backed vector beside BM25 and plain vector", f"refused: {short(e)}") + ) + finally: + client.delete_managed_database(db) + return out + + +def summary(times: list[float]) -> str: + if not times: + return "no calls completed" + ordered = sorted(times) + p95 = ordered[min(len(ordered) - 1, int(len(ordered) * 0.95))] + return ( + f"{len(times)} calls, median {statistics.median(times) * 1000:.0f} ms, " + f"p95 {p95 * 1000:.0f} ms, mean {statistics.mean(times) * 1000:.0f} ms" + ) + + +def main() -> int: + provider = os.environ.get("HOTMEMORY_EMBEDDING_PROVIDER", "sys_emb_openai") + sizes = [int(s) for s in os.environ.get("HOTMEMORY_M5_SIZES", "1000,10000,100000").split(",")] + cloud = "--cloud" in sys.argv[1:] + if cloud: + prefix = env("HOTMEMORY_MEASURE_DB") + client = connect_cloud() + names = {f"{prefix}-{m}" for m in ("m4", "m5", "m6")} + taken = sorted(names & {d.description for d in client.list_managed_databases()}) + if taken: + sys.exit( + f"databases already exist: {', '.join(taken)}; pick a new HOTMEMORY_MEASURE_DB" + ) + else: + prefix = "hotmemory" + client = connect() + started_at = now() + with tempfile.TemporaryDirectory() as tmp: + work = Path(tmp) + bare = measure_bare(client, work, provider, prefix) + print("m6 done", file=sys.stderr) + try: + insert = measure_insert(client, work, prefix) + except RuntimeError as e: + insert = {"setup_error": short(e)} + print("m4 done", file=sys.stderr) + try: + scale = measure_scale(client, work, sizes, provider, prefix) + except RuntimeError as e: + scale = {"sizes": {}, "coexistence": [], "setup_error": short(e)} + print("m5 done", file=sys.stderr) + + print(f"## {'Cloud' if cloud else 'Local'} measurements of M4 to M6, {started_at}") + print() + print(f"Host: {client.host}.") + print() + print("### M6") + print() + for label, outcome in bare: + print(f"- {label}: {outcome}") + print() + print("### M4") + print() + if "setup_error" in insert: + print(f"- setup fails: {insert['setup_error']}") + print(f"- SQL INSERT, one row per statement: {summary(insert.get('sql', []))}.") + if "sql_error" in insert: + print(f" - error: {insert['sql_error']}") + print(f"- Upload and append load, one row per call: {summary(insert.get('load', []))}.") + print( + f"- Rows after the calls: by_sql {insert.get('by_sql_rows')}, " + f"by_load {insert.get('by_load_rows')} (each includes one seed row)." + ) + print() + print("### M5") + print() + if "setup_error" in scale: + print(f"- setup fails: {scale['setup_error']}") + print( + f"Median of {M5_RUNS} runs after one warm-up, {M5_DIMENSIONS}-dimension embeddings, " + f"k {M5_K}, fusion depth {M5_DEPTH}." + ) + for size, entry in scale["sizes"].items(): + print() + print(f"#### {size} rows (load {entry['load_s']:.1f} s)") + print() + for label, outcome in entry["builds"].items(): + print(f"- index build, {label}: {outcome}") + print() + print("| Query | Without indexes | With indexes |") + print("|---|---|---|") + for name in entry["before"]: + print(f"| {name} | {entry['before'][name]} | {entry['after'][name]} |") + print() + for label, outcome in scale["coexistence"]: + print(f"- {label}: {outcome}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/uv.lock b/uv.lock new file mode 100644 index 0000000..2f1c99b --- /dev/null +++ b/uv.lock @@ -0,0 +1,43 @@ +version = 1 +revision = 3 +requires-python = ">=3.11" + +[[package]] +name = "hotmemory" +version = "0.0.0" +source = { virtual = "." } + +[package.dev-dependencies] +dev = [ + { name = "ruff" }, +] + +[package.metadata] + +[package.metadata.requires-dev] +dev = [{ name = "ruff" }] + +[[package]] +name = "ruff" +version = "0.16.10" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/c4/49/23802c45f093eb14bde54b141d2b2f058edfa63a06db7beed047308cc08f/ruff-0.16.10.tar.gz", hash = "sha256:eff4728c4eaae93f0955cd264d24b2ab348e74bf59986ccf282ba6dc16b3b017", size = 4958724, upload-time = "2026-10-01T18:03:21.697Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/2f/21/ebce22e1d90cdb2cd691026b9c6e9bec6499f0b396089481755b5d49efff/ruff-0.16.10-py3-none-linux_armv6l.whl", hash = "sha256:488b0fe3f3574210e5cf80d9f59b9e3ab17a127a8155de3f307b392589cfb511", size = 10095557, upload-time = "2026-10-01T18:02:36.072Z" }, + { url = "https://files.pythonhosted.org/packages/cb/98/a54de85876a8b2612bfa0d84c7b9abfb39c6a3354aee7800b09c1649b9e3/ruff-0.16.10-py3-none-macosx_10_12_x86_64.whl", hash = "sha256:e748ff95c934c4e978783b8e687bc174e7bd84e8ad24e3243e1ecfcda5e0282d", size = 10340577, upload-time = "2026-10-01T18:02:39.258Z" }, + { url = "https://files.pythonhosted.org/packages/9f/16/1a5a4a2657effe4806110f2b907313802f1367fcdbb29e8122766e407fab/ruff-0.16.10-py3-none-macosx_11_0_arm64.whl", hash = "sha256:3031a4a2e8e7b8a46f70be45f198c35a11ece509a94b80334d8d397a33c67550", size = 9774282, upload-time = "2026-10-01T18:02:41.79Z" }, + { url = "https://files.pythonhosted.org/packages/6e/fb/470085af734da396e68cd80fb0a3e7c109a459716ae59588c0f6fab8a17d/ruff-0.16.10-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:494401c86df4c4c25f69b9419605d944467ee98c42fb6ad405ef4fa40b8fb67d", size = 9920895, upload-time = "2026-10-01T18:02:44.458Z" }, + { url = "https://files.pythonhosted.org/packages/57/de/f10cffe4f88a37ea6615ec460f01bf76f0bc1477737a35d9ee61a630a78b/ruff-0.16.10-py3-none-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:d203abc0ff2b773ee33d00ab8df0bb67046fbc7c07b119332f08b9b345cf8221", size = 9902707, upload-time = "2026-10-01T18:02:47.041Z" }, + { url = "https://files.pythonhosted.org/packages/ff/44/3fdcedf83ae60ef837dd239170e480dce606a9142cfa2db911afdf855fd9/ruff-0.16.10-py3-none-manylinux_2_17_i686.manylinux2014_i686.whl", hash = "sha256:bd83d1235a5258d318477bc5b576303974cbdef5df0c01a1bff14efcc23bd12a", size = 10619287, upload-time = "2026-10-01T18:02:49.485Z" }, + { url = "https://files.pythonhosted.org/packages/c1/62/02e76a5574002153618eafbb70e159468addc72a5d2c488e65cbb4639d2d/ruff-0.16.10-py3-none-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:bc2610fb269fa56dd8a68669ae470fa6272902668c0fc2ebc3aa112b2633d5b8", size = 11339942, upload-time = "2026-10-01T18:02:52.008Z" }, + { url = "https://files.pythonhosted.org/packages/6b/c4/cde27d47ad8d4126c587e608c67c46e7feac11763f606d261a1bc8a489e6/ruff-0.16.10-py3-none-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:3e70175e29cc94c26ea296c80e470180b744b7419026898e58f520c6ab32578e", size = 10934316, upload-time = "2026-10-01T18:02:54.811Z" }, + { url = "https://files.pythonhosted.org/packages/e4/03/17234145f302a645a123e8c3bb2411ecf4669fbc350de3b0d803230a1729/ruff-0.16.10-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:f33f43a864a8483eebd160e713336c8bab02c934feaff0a33cf5ccb41546d09a", size = 10387968, upload-time = "2026-10-01T18:02:57.497Z" }, + { url = "https://files.pythonhosted.org/packages/71/29/2493af60240ee7644b38a4b821f5f9c3b5a4fa3178770fe1d0c685217395/ruff-0.16.10-py3-none-manylinux_2_31_riscv64.whl", hash = "sha256:1dfc6f0088149fb6a362c1c446bcbb3fd2157b3852fe2fa68409276eab9ad9b3", size = 10537906, upload-time = "2026-10-01T18:03:00.006Z" }, + { url = "https://files.pythonhosted.org/packages/2e/9a/f56b28f3b143bb9e518c34fb89ab191b626bca8b70dbf8af0d2e2d473572/ruff-0.16.10-py3-none-musllinux_1_2_aarch64.whl", hash = "sha256:6553498afc35f580f036030795810b9e6bcea31604b0fd9e8d352795473042e3", size = 10012938, upload-time = "2026-10-01T18:03:03.006Z" }, + { url = "https://files.pythonhosted.org/packages/c2/c4/fca37362848ea4d4d80712df13632e645e7c7cfb1bdedf140699ac7a090b/ruff-0.16.10-py3-none-musllinux_1_2_armv7l.whl", hash = "sha256:a3b8471dea115d37f123882be852bed13403746d3a76c11de4a19ec5f9ff5a03", size = 9897945, upload-time = "2026-10-01T18:03:05.818Z" }, + { url = "https://files.pythonhosted.org/packages/7d/c7/e0bc57664d6e0c61fd22f620af260fe9662d7e1ddb320ea7f160e76ef165/ruff-0.16.10-py3-none-musllinux_1_2_i686.whl", hash = "sha256:92e59a70bcbd9d3a5483656da906ec28edfdacfce00afd99edb8b4e9d15644be", size = 10333134, upload-time = "2026-10-01T18:03:08.25Z" }, + { url = "https://files.pythonhosted.org/packages/21/aa/5c9f3b68737c0e4a33a1db5dfab44f7b786d96ca91a7233112ddab1e6dc2/ruff-0.16.10-py3-none-musllinux_1_2_x86_64.whl", hash = "sha256:7ae7375f803b5520dc9f546bed7e3a0acb70b91812e9bb4b19927de22f25b77d", size = 10741702, upload-time = "2026-10-01T18:03:10.775Z" }, + { url = "https://files.pythonhosted.org/packages/78/fa/0f9c2020dc316c53d983be011720b8157cc052b97e3f4dfa7db6d0f880a6/ruff-0.16.10-py3-none-win32.whl", hash = "sha256:2a12e01cb9156c10c466f63b46eaae5ecea28dfbd21b5836353ae498e7d1349a", size = 10139176, upload-time = "2026-10-01T18:03:13.267Z" }, + { url = "https://files.pythonhosted.org/packages/99/29/cfb0df9448d4d4ad48c2de029ada9ebd71baa6da983a6c77ee6c6cd0fe82/ruff-0.16.10-py3-none-win_amd64.whl", hash = "sha256:97f2015c92aa97105b0eab19eb5d224884399281cfc5da86a92db4ab5e7fb2ca", size = 10584734, upload-time = "2026-10-01T18:03:16.006Z" }, + { url = "https://files.pythonhosted.org/packages/fc/05/c16957eb287c3fc062e032619a25868d93d408a844b725bcf514f7a378ff/ruff-0.16.10-py3-none-win_arm64.whl", hash = "sha256:25a65fe998c4e6861ec079ada5826a2fc605e6cbccbe9dcd7fac1f54e791621b", size = 10366440, upload-time = "2026-10-01T18:03:19.04Z" }, +]