perf(core): seed the published-step set as soon as handleSuspension returns - #4128
perf(core): seed the published-step set as soon as handleSuspension returns#4128pranaygp wants to merge 2 commits into
Conversation
🦋 Changeset detectedLatest commit: 196f91f The changes in this PR will be included in the next version bump. This PR includes changesets to release 16 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results✅ All tests passed 🛠 Infra Events (absorbed by the harness)Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.
E2E Test SummarySummary
Details by Category✅ ▲ Vercel Production
✅ 💻 Local Development
✅ 📦 Local Production
✅ 🐘 Local Postgres
✅ 🪟 Windows
✅ 🌐 Cross-language Conformance
✅ vercel-http-transport
✅ vercel-multi-region
✅ vercel-ws-transport
|
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 156984ms → this run 156925ms (Δ -59ms, 0%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): 📜 Previous results (1)637e5d1Fri, 11 Sep 2026 21:13:31 GMT · run logs
Streams
ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
There was a problem hiding this comment.
🟡 Changes recommended
Preserve the step name or use a composite identity so a legitimate dispatch is not skipped when correlation IDs are reused.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
This PR moves published-step tracking earlier in replay handling to prevent duplicate step dispatches.
Changes:
- Seeds published-step tracking immediately after suspension handling.
- Adds hook-conflict replay regression coverage.
- Adds a core patch changeset.
File summaries
| File | Summary |
|---|---|
packages/core/src/runtime.ts |
Seeds published steps before early exits; critical review issue concerns correlation-ID-only tracking. |
packages/core/src/runtime.test.ts |
Adds hook-conflict deduplication coverage and harness options. |
.changeset/dispatch-skip-republished-seed.md |
Documents the core patch release. |
Review details
- Files reviewed: 3/3 changed files
- Comments generated: 1
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| for (const correlationId of suspensionResult.queuedStepCorrelationIds) { | ||
| publishedStepCorrelationIds.add(correlationId); |
There was a problem hiding this comment.
Good catch, fixed in f6bc113.
The suspension handler now reports its publishes as dispatch keys (queuedStepDispatchKeys, each stepDispatchIdempotencyKey(correlationId, stepName)), and the runtime seeds, checks and adds the invocation's published set under that same composite key, so it agrees with the idempotency key on what a step's identity is. The resilient path and the batched fold both have the step name at hand, so no new field was needed.
New test (still dispatches a step a later pass bound to a correlation id published under another name) rebinds the queued step's correlation id on the second pass and asserts it is dispatched under its new identity with the new key; it fails against the id-only keying. One note from writing it: a workflow cannot stage that rebinding by itself within one delivery (events reach the VM one macrotask at a time, so the pass that published ordinal n never observed anything a later pass could branch on before n), so the test taps the engine's output to rename the pending item.
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 41 total
Full trace: |
The set that stops a replay pass from re-publishing step messages this invocation already sent was keyed by correlation id alone. The dispatch path treats (correlationId, stepName) as the step's identity (stepDispatchIdempotencyKey) because a corrected replay can rebind a correlation id to a different step; keyed by id alone, a later pass would skip that step's only dispatch. The suspension handler now reports the steps it published as dispatch keys (queuedStepDispatchKeys) and the runtime seeds, checks and adds the same composite key. A runtime test taps the engine's output to rebind the queued step's correlation id on the second pass and asserts it is dispatched under its new identity. Addresses #4128 (comment) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…eturns Follow-up to #4099. The invocation-scoped set of step correlation ids this delivery already published was seeded from the suspension handler's `queuedStepCorrelationIds` only at the dispatch pass, which runs after the hook-conflict, attribute-event and serialization-failure branches `continue` the replay loop in-process. The handler's resilient publishes on such a pass were therefore forgotten: on the next pass the step already existed, the handler no longer reported it, and the dispatch pass published its message a second time. Move the seeding to immediately after `handleSuspension` returns, before any early exit, and drop the dispatch-pass-only placement. Adds a test that drives a hook_conflict pass with resilient dispatch on and asserts the queued step's message is sent exactly once across the invocation (it sends twice without the fix). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The set that stops a replay pass from re-publishing step messages this invocation already sent was keyed by correlation id alone. The dispatch path treats (correlationId, stepName) as the step's identity (stepDispatchIdempotencyKey) because a corrected replay can rebind a correlation id to a different step; keyed by id alone, a later pass would skip that step's only dispatch. The suspension handler now reports the steps it published as dispatch keys (queuedStepDispatchKeys) and the runtime seeds, checks and adds the same composite key. A runtime test taps the engine's output to rebind the queued step's correlation id on the second pass and asserts it is dispatched under its new identity. Addresses #4128 (comment) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
f6bc113 to
196f91f
Compare
About these numbersSizes are gzip; parentheses show the change against
|
Follow-up to #4099, addressing the Copilot review comment #4099 (comment).
Problem
#4099 added an invocation-scoped set of step correlation ids this delivery already published, so the post-inline replay pass stops re-publishing them. That set was seeded from the suspension handler's
queuedStepCorrelationIdsonly at the dispatch pass. But three branchescontinuethe replay loop in-process before the dispatch pass runs: the hook-conflict continuation, the attribute-event detour, and the serialization-failure replay.handleSuspensioncan publish resilient step messages (and the batched fold's eager publishes) on exactly such a pass. On the next pass the step already exists, so the handler no longer reports it, and the dispatch pass re-publishes it. The scenario Copilot described is real: the new test reproduces 2 sends for one step onmain.Change
packages/core/src/runtime.ts: seedpublishedStepCorrelationIdsfromsuspensionResult.queuedStepCorrelationIdsimmediately after thehandleSuspensiontry/catch, before any early exit; remove the dispatch-pass-only seeding. The skip logic itself is unchanged, and a fresh delivery still re-enqueues unconditionally.packages/core/src/runtime.test.ts: the ack-ordering harness gainsworkflowSourceandconflictHookCreateoptions (the mock World answershook_createdwith a recordedhook_conflict). New test drives a fire-and-forget hook plus the two-step fan-out with resilient dispatch on, asserts the first pass hit the conflict, went through more than one log load, and the queued step's message was sent exactly once (carryingstepInput). Fails with 2 sends without the runtime change.Tests
src/runtime.test.ts@workflow/core(FORCE_COLOR=0 pnpm test)tsc --noEmitReview follow-up
The published set is keyed by step identity (
stepDispatchIdempotencyKey(correlationId, stepName)), not by correlation id alone, so it agrees with the dispatch idempotency key: a step that a corrected replay binds to a correlation id an earlier pass published under a different name is still dispatched. The suspension handler reports its publishes asqueuedStepDispatchKeysaccordingly. Covered by the new rebinding test (which taps the engine's output to stage the rebinding, since a workflow cannot observe log state between two synchronous calls within one delivery). See #4128 (comment).🤖 Generated with Claude Code