You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
fix(search): bound Search retirement pages so the migration finishes under its statement timeout (#8460)
* fix(search): bound retirement pages by mutated rows so they finish under the statement timeout
A retirement page read 25,000 IDs and updated or deleted every target row among them in one
statement. Retiring a document is a non-HOT update that writes every index on `document`, and a
deleted chunk cascades into its projections, so on a KB that dominates the table a page's write
cost, not its scan, outran the two-minute statement timeout and failed the deploy migration.
Each page now mutates at most a row limit of its target rows. A page that reaches the limit
advances the cursor only to its last mutated row, and already-retired documents never spend the
limit. The limit starts at 2,000, halves after a slow page or a statement timeout (the timed-out
page rolls back with its cursor and is retried), and doubles after a fast full page. The completion
rechecks, which walk every captured KB once, run with a 30-minute timeout. The retirement stays
idempotent and resumes from its saved cursor.
* fix(search): shrink only timed-out page mutations and pace Search retirement pages
Only a statement timeout from a page's mutating statement now halves the row limit and retries
the rolled-back page. Any other timeout, such as a completion recheck, fails the run at once
instead of repeating the same statement at every smaller limit.
Pages are timed around the whole call, commit included, so the synchronous-replication wait
counts toward the slow-page threshold. Each page is followed by a pause as long as the page, up
to five seconds, and the row limit is capped at 8,000. Phase changes no longer adjust the limit.
The progress log now carries the phase, cursor and rows mutated, and the migration logs slow-page
halvings, phase changes and the start of the completion recheck.
* test(search): prove retired documents never spend the retirement row limit
A run of already-retired Search documents longer than the row limit must be crossed in one page. The test counts documents-phase statements and fails if the page filter on unretired rows is removed.
* fix(search): scale the retirement scan window with the row limit and time out resumed rechecks like completion
A capped page re-reads its scan from its last mutated row, so a fixed 25,000-ID window re-read
most of the same IDs on every page once the row limit shrank. Each page now reads at most four IDs
per row of its limit, capped at 25,000, and any fast page doubles the limit so sparse stretches
widen the window again.
A retry that finds retirement already complete revalidates every captured KB, as completion does,
so it now runs under the same 30-minute timeout instead of the two-minute page timeout.
* fix(search): never grow a Search retirement page back to a size that timed out, and pause after a timeout
A timed-out page halved the row limit, but one fast page doubled it straight back, so the run
alternated between the size that timed out and half of it, rolling back a full statement-timeout
page each time. The limit now grows only up to half of the smallest size that timed out, and a
timed-out page is followed by the same pause as any other page.
0 commit comments