You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: apps/sim/lib/workspace-files/search/README.md
+2Lines changed: 2 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -16,6 +16,8 @@ The chunk GIN index uses `fastupdate = off`. Each bounded insert updates the mai
16
16
17
17
Indexing transactions have separate limits from search: ten seconds per statement, five seconds waiting for a lock, and thirty seconds total on PostgreSQL 17. The outer limit leaves time for ordinary statement cancellation and rollback instead of terminating the connection at the same ten-second deadline. PostgreSQL 16 uses the compatible idle-transaction guard. A canceled batch remains unpublished; the task retry starts a fresh fenced build, and cleanup retires the previous attempt. One row's direct GIN insert is not interruptible, so when storage is saturated a single ordinary chunk can run well past the statement deadline and the cancellation lands only after it; smaller batches cannot prevent that. A statement, lock, or transaction timeout is therefore treated as missing database capacity rather than a bad file: the task retries it after about 2, 4, 8, 16, and 30 minutes (with jitter), six attempts in all, so the retries outlast a slow window instead of landing inside it. Other failures keep three attempts with the runner's short default delays. A run waiting to retry still holds one of its workspace's two outstanding dispatch slots. Only a revision that exhausts its attempts is marked failed; this does not automatically retry revisions already marked failed.
18
18
19
+
A dispatch claim commits before Trigger.dev accepts its run. A dispatcher that stops in between, for example one killed at its 60-second limit while PostgreSQL is still committing, leaves a claim with no run, and that claim holds one of its workspace's two slots. Each claim therefore carries a two-minute handoff deadline in PostgreSQL time, twice the dispatcher's maximum duration. The deadline is cleared once a run is known to exist: the dispatcher clears it after Trigger.dev accepts the batch, and the worker clears it when the build begins. That write skips rows another transaction holds rather than waiting on them, so it cannot deadlock with a bulk file change. The next dispatch releases a claim whose deadline has passed, logs it, and counts it in its result. The file is claimed again later under a new token, which fences out any run the old claim did get, so nothing re-sent has to be deduplicated. A claim with no deadline, including one made before the column existed, keeps the six-hour stale-dispatch recovery, which covers runs lost after they were handed off. In-process dispatch, used when Trigger.dev is not configured, clears the deadline as soon as it schedules the work in its own process, so a restart there still leaves the scheduled claims to the six-hour window.
20
+
19
21
The indexing task uses an isolated `medium-2x` Trigger worker (4 GB RAM). Document parsers can materialize expanded content before chunking, so source and extracted-text byte limits do not bound parser memory. Parser complexity guards and the worker's memory budget remain separate protections.
20
22
21
23
File edits, context changes, and deletion invalidate metadata and expire builds. Chunks have no cascading foreign key to files or workspaces. Cleanup locks at most 100 expired builds with `SKIP LOCKED`, deletes at most 1,000 chunks per transaction, retires empty builds in the same batch, and stops after 10 batches or five seconds. Dispatch pauses while at least 10,000 expired chunks await cleanup, so sustained revisions cannot keep admitting new builds faster than retirement can drain them. Existing ready files remain searchable. Stale workers cannot revive a reclaimed build.
0 commit comments