…ache
The stop-after suite failed 12 of 175 CI runs (Oct 1-6), always in its own
`next dev` app, which restored the Turbopack dev cache (.next/dev) the
SCIM/version-compare app left in the same job. That app ran with other
NEXT_PUBLIC_* values, and `next dev` SIGKILLs its server 100ms after SIGTERM.
Restoring the cache:
- panicked Turbopack at startup (inner_of_upper_lost_follower), 6 runs;
- panicked mid-compile of the execute route ("socket connection was closed
unexpectedly"), 4 runs;
- wedged that compile until the 300s request timeout ("The operation timed
out."), 2 runs.
No run reached the executor: the cleanup check found no open execution log.
- Each app step removes .next/dev before starting, so no app restores another
app's cache.
- A failing step prints the server log tail, not only a startup failure.
- The suite compiles the execute route with a refused request under its own
300s budget and named check; every later request and CLI run is bounded at
60s. A timed-out or dropped request names the route and elapsed time, lands
in the report with a null status, and a CLI timeout is reported as one.
Summary
The "End-to-end over real HTTP" job's stop-after step (
test-workflow-stop-after-e2e.ts) was flaky. This was test infrastructure, not a product bug: no failed run ever reached the executor.Root cause. All the apps in the job share
apps/sim/.next. The stop-after app restored the Turbopack dev cache (.next/dev) that the job's previous app, the SCIM/version-compare one, left behind. That cache was built with differentNEXT_PUBLIC_*values by a server thatnext devSIGKILLs 100ms after SIGTERM (NEXT_EXIT_TIMEOUT_MSdefaults to 100). The step also started while the previous app could still be writing that directory:killpluswaitonnext devreturns before its server process tree is gone. An earlier revision of this PR hit that directly. Itsrm -rf .next/devfailed withDirectory not empty1s after the previous step ended.I classified every failed run of this step across 175 CI runs from Oct 1 to Oct 6. 12 were this flake, and all 12 came from the restored cache:
Local workflow app exited during startupinner_of_upper_lost_follower ... aggregation_update.rsThe socket connection was closed unexpectedly/api/v2/workflows/[workflowId]/executeThe operation timed out.after 5 min○ Compiling /api/v2/workflows/[workflowId]/execute ...that never finishes, with JS heap flatThe other 2 failed runs had real branch errors. In the timed-out runs the cleanup check found no open execution log, so the first request never reached the handler. That rules out a stop-after run that never terminates. Healthy runs of this step took 35–99s (p50 47s).
Changes
test-build.yml), for all three app steps:setsid)..github/scripts/stop-session.shsends SIGTERM to the whole session and returns only once none of it is left running. It escalates to SIGKILL after 10s and fails the step if anything survives that. It logs any process that outlivednext dev.rm -rf .next/dev).status: null, so a status is recorded only for a fully read response.next devcompiling the claim, lease and complete routes. One CI attempt failed that way. A new check now compiles those three routes with refused bodies before the timed checks.Evidence
taskset), unchanged code: 8 of 8 iterations passed. On a quiet box the flake doesn't reproduce, and no process outlivednext devthere.lint,typecheck,audits,migrations,docs,blocksandtestall pass. Intest, 3 unrelated files failed under load and passed when rerun alone.