Skip to content

Commit 1e7d6fd

Browse files
committed
fix(search): never grow a Search retirement page back to a size that timed out, and pause after a timeout
A timed-out page halved the row limit, but one fast page doubled it straight back, so the run alternated between the size that timed out and half of it, rolling back a full statement-timeout page each time. The limit now grows only up to half of the smallest size that timed out, and a timed-out page is followed by the same pause as any other page.
1 parent ee1d868 commit 1e7d6fd

3 files changed

Lines changed: 31 additions & 9 deletions

File tree

‎packages/db/script-migrations/0027_retire_search_embeddings.integration.ts‎

Lines changed: 14 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -332,12 +332,15 @@ describe('retiring dormant Search embeddings', () => {
332332
CASE WHEN i % 3 = 0 THEN 'ordinary-doc' ELSE 'search-doc' END
333333
FROM generate_series(1003, 5002) i`
334334
await sql`CREATE TABLE committed_statement (rows integer NOT NULL)`
335+
/** Sequences are not transactional, so this counts the timed-out statements that rolled back. */
336+
await sql`CREATE SEQUENCE timed_out_statement`
335337
/** Stands in for write cost: a statement touching more than `bound` rows times out and rolls back. */
336338
await sql.unsafe(`CREATE FUNCTION bound_statement_rows() RETURNS trigger LANGUAGE plpgsql AS $$
337339
DECLARE touched integer;
338340
BEGIN
339341
SELECT count(*) INTO touched FROM changed_rows;
340342
IF touched > ${bound} THEN
343+
PERFORM nextval('timed_out_statement');
341344
RAISE EXCEPTION 'canceling statement due to statement timeout' USING ERRCODE = 'query_canceled';
342345
END IF;
343346
INSERT INTO committed_statement VALUES (touched);
@@ -353,6 +356,13 @@ describe('retiring dormant Search embeddings', () => {
353356
sum(rows)::int AS total FROM committed_statement`
354357
expect(largest).toBeLessThanOrEqual(bound)
355358
expect(largest).toBeGreaterThan(0)
359+
/**
360+
* The limit halves 2,000 → 1,000 → 500 → 250 on the first document page and never grows back
361+
* to a size that timed out, in either phase: exactly three rolled-back statements. A page
362+
* that read past its row limit in either phase would time out again.
363+
*/
364+
const [{ timeouts }] = await sql`SELECT last_value::int AS timeouts FROM timed_out_statement`
365+
expect(timeouts).toBe(3)
356366
/**
357367
* Documents: i % 3 <> 0 gives 2,667 Search rows, of which i % 7 = 0 leaves 381 retired, so
358368
* 2,286 are updated, plus `search-doc`. Chunks: 501 Search rows from the fixture plus the
@@ -387,6 +397,7 @@ describe('retiring dormant Search embeddings', () => {
387397
await sql`DROP TRIGGER IF EXISTS bound_embedding_delete ON embedding`
388398
await sql`DROP FUNCTION bound_statement_rows()`
389399
await sql`DROP TABLE committed_statement`
400+
await sql`DROP SEQUENCE timed_out_statement`
390401
}
391402
}, 60_000)
392403

@@ -423,9 +434,9 @@ describe('retiring dormant Search embeddings', () => {
423434
expect(await pass()).toBe(true)
424435
const reads = (await documentReads()) - before
425436
/**
426-
* About 120 pages retire the 3,001 documents 25 at a time, each after one rolled-back attempt
427-
* at 50 rows. A window of four IDs per row reads about 300 IDs and 75 update lookups per page,
428-
* roughly 15 reads per document, plus the chunk phase and completion rechecks. A fixed
437+
* About 120 pages retire the 3,001 documents 25 at a time once the limit has halved down to
438+
* its floor. A window of four IDs per row reads about 100 IDs and 25 update lookups per page,
439+
* roughly 5 reads per document, plus the chunk phase and completion rechecks. A fixed
429440
* 25,000-ID window re-reads the rest of the table on every attempt, over 100 per document.
430441
*/
431442
expect(reads).toBeLessThan(30 * docs)

‎packages/db/script-migrations/0027_retire_search_embeddings.ts‎

Lines changed: 14 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -23,9 +23,15 @@ const SCAN_ROWS_PER_MUTATION = 4
2323
const ROW_LIMIT = { initial: 2_000, min: 25, max: 8_000 } as const
2424
/** A page slower than this halves the row limit. */
2525
const SLOW_PAGE_MS = 30_000
26-
/** A page faster than this doubles the row limit, widening its scan window with it. */
26+
/**
27+
* A page faster than this doubles the row limit, widening its scan window with it, but never back
28+
* to a size that timed out.
29+
*/
2730
const FAST_PAGE_MS = SLOW_PAGE_MS / 4
28-
/** Each page is followed by a pause as long as the page, up to this, to leave the primary headroom. */
31+
/**
32+
* Each page, committed or timed out, is followed by a pause as long as the page, up to this, to
33+
* leave the primary headroom.
34+
*/
2935
const MAX_PAGE_PAUSE_MS = 5_000
3036
const LOCK_RETRY_BUDGET_MS = 60_000
3137

@@ -135,6 +141,8 @@ export const retireSearchEmbeddingsMigration: ScriptMigration = {
135141
let batches = 0
136142
let mutated = 0
137143
let rowLimit: number = ROW_LIMIT.initial
144+
/** The largest limit the run may still try: half of the smallest limit that timed out. */
145+
let ceiling: number = ROW_LIMIT.max
138146
for (;;) {
139147
/** Timed around the whole call, so the synchronous-replication wait at commit counts. */
140148
const pageStartedAt = performance.now()
@@ -145,8 +153,10 @@ export const retireSearchEmbeddingsMigration: ScriptMigration = {
145153
if (!(error instanceof PageMutationTimeout)) throw error
146154
/** The timed-out page rolled back with its cursor, so it is retried with fewer rows. */
147155
if (rowLimit <= ROW_LIMIT.min) throw error.timeout
148-
rowLimit = halve(rowLimit)
156+
ceiling = halve(rowLimit)
157+
rowLimit = ceiling
149158
logger.warn('Search retirement page timed out; retrying with fewer rows', { rowLimit })
159+
await sleep(Math.min(performance.now() - pageStartedAt, MAX_PAGE_PAUSE_MS))
150160
continue
151161
}
152162
if (page.done) break
@@ -167,7 +177,7 @@ export const retireSearchEmbeddingsMigration: ScriptMigration = {
167177
rowLimit,
168178
})
169179
} else if (pageMs < FAST_PAGE_MS) {
170-
rowLimit = Math.min(ROW_LIMIT.max, rowLimit * 2)
180+
rowLimit = Math.min(ceiling, rowLimit * 2)
171181
}
172182
if (batches % 10 === 0) {
173183
logger.info('Search embedding retirement progress', {

‎packages/db/script-migrations/search-embedding-retirement.md‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -48,8 +48,9 @@ also widens the scan window, so sparse stretches are not crawled in small window
4848
Materialized SQL pages keep the IDs inside PostgreSQL; the migration process receives only a cursor
4949
and a validation result. Each page uses a two-minute statement timeout and a one-second lock
5050
timeout. If a page's mutating statement exceeds the statement timeout, the page rolls back with its
51-
cursor and is retried with half the row limit; one that still times out at 25 rows fails the
52-
migration. Any other statement timeout fails the migration at once, because a smaller page cannot
51+
cursor and is retried with half the row limit after the usual pause. From then on, fast pages grow
52+
the limit only up to that halved size, so a size that timed out is never tried again. A page that
53+
still times out at 25 rows fails the migration. Any other statement timeout fails the migration at once, because a smaller page cannot
5354
speed it up. The completion rechecks, which walk every captured KB once, run with a 30-minute
5455
timeout. Brief lock timeouts retry the rolled-back page with bounded backoff for up to one minute.
5556
Other errors, or exhausted lock retries, fail the migration without a completion receipt.

0 commit comments

Comments
 (0)