Skip to content

fix(ci): stop each HTTP end-to-end app fully and start the next from an empty dev cache - #8686

Merged
waleedlatif1 merged 5 commits into
stagingfrom
fix/stop-after-e2e-cold-cache
Oct 6, 2026
Merged

waleedlatif1 merged 5 commits into
stagingfrom
fix/stop-after-e2e-cold-cache

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The "End-to-end over real HTTP" job's stop-after step (test-workflow-stop-after-e2e.ts) was flaky. This was test infrastructure, not a product bug: no failed run ever reached the executor.

Root cause. All the apps in the job share apps/sim/.next. The stop-after app restored the Turbopack dev cache (.next/dev) that the job's previous app, the SCIM/version-compare one, left behind. That cache was built with different NEXT_PUBLIC_* values by a server that next dev SIGKILLs 100ms after SIGTERM (NEXT_EXIT_TIMEOUT_MS defaults to 100). The step also started while the previous app could still be writing that directory: kill plus wait on next dev returns before its server process tree is gone. An earlier revision of this PR hit that directly. Its rm -rf .next/dev failed with Directory not empty 1s after the previous step ended.

I classified every failed run of this step across 175 CI runs from Oct 1 to Oct 6. 12 were this flake, and all 12 came from the restored cache:

Symptom Runs What the server log shows
Local workflow app exited during startup 6 Turbopack panic inner_of_upper_lost_follower ... aggregation_update.rs
The socket connection was closed unexpectedly 4 the same panic while compiling /api/v2/workflows/[workflowId]/execute
The operation timed out. after 5 min 2 ○ Compiling /api/v2/workflows/[workflowId]/execute ... that never finishes, with JS heap flat

The other 2 failed runs had real branch errors. In the timed-out runs the cleanup check found no open execution log, so the first request never reached the handler. That rules out a stop-after run that never terminates. Healthy runs of this step took 35–99s (p50 47s).

Changes

  • CI (test-build.yml), for all three app steps:
    • Each app starts in its own session (setsid).
    • The new .github/scripts/stop-session.sh sends SIGTERM to the whole session and returns only once none of it is left running. It escalates to SIGKILL after 10s and fails the step if anything survives that. It logs any process that outlived next dev.
    • Each app then starts from an empty dev cache (rm -rf .next/dev).
    • A failing step prints the server log tail, not only on a startup failure.
  • Stop-after suite:
    • A named first check compiles the execute route with a refused request under its own 300s budget.
    • Every later request and CLI run is bounded at 60s.
    • A timed-out or dropped request names the route and the elapsed time. It lands in the report with status: null, so a status is recorded only for a fully read response.
    • A CLI timeout is reported as a timeout, and an unexpected CLI failure reports its stderr.
  • Desktop inbox suite (feat(desktop): bind turns to a desktop and enforce its deadlines from the row #8650): a cold cache exposed that its 10s "all four calls" deadline also covered next dev compiling the claim, lease and complete routes. One CI attempt failed that way. A new check now compiles those three routes with refused bodies before the timed checks.

Evidence

  • Local: a CI-faithful harness on the Mac (5 iterations) and a targeted SIGTERM-then-restore loop (5 cycles) both passed. The panic didn't reproduce on darwin-arm64.
  • Linux EC2 (8 cores via taskset), unchanged code: 8 of 8 iterations passed. On a quiet box the flake doesn't reproduce, and no process outlived next dev there.
  • Linux EC2, this PR: 8 of 8 two-step iterations and 6 of 6 three-step iterations passed. A cold cache moves the stop-after step from ~58s to ~105s.
  • CI reruns of the http-e2e job on this PR: see the comment thread.
  • Gate (EC2): lint, typecheck, audits, migrations, docs, blocks and test all pass. In test, 3 unrelated files failed under load and passed when rerun alone.

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@vercel

vercel Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Oct 6, 2026 8:20pm UTC

Request Review

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Turn on auto-fix | Re-trigger cubic

@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High risk] Changes CI process for end-to-end test isolation and timing.

The PR appears safe to merge; no outstanding findings or new actionable defects remain.

Summary

The PR gives each HTTP end-to-end app a separate process session, waits for cleanup before the next app starts, and clears the shared Turbopack dev cache. It also adds route warm-up checks and improves stop-after suite timeouts and diagnostics.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Clear .next/dev] --> B[Start next dev in its own session]
  B --> C[Run HTTP end-to-end checks]
  C --> D[Stop and wait for the session]
  D --> E[Start the next app]
Loading

Reviews (3) · Last reviewed commit: "fix(e2e): compile the desktop executor r..."

Comment thread apps/sim/scripts/test-workflow-stop-after-e2e.ts Outdated
@waleedlatif1
waleedlatif1 force-pushed the fix/stop-after-e2e-cold-cache branch from 358b53b to 2c9bebe Compare October 6, 2026 18:26
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Turn on auto-fix | Re-trigger cubic

Comment thread .github/workflows/test-build.yml Outdated
…ache

The stop-after suite failed 12 of 175 CI runs (Oct 1-6), always in its own
`next dev` app, which restored the Turbopack dev cache (.next/dev) the
SCIM/version-compare app left in the same job. That app ran with other
NEXT_PUBLIC_* values, and `next dev` SIGKILLs its server 100ms after SIGTERM.
Restoring the cache:
- panicked Turbopack at startup (inner_of_upper_lost_follower), 6 runs;
- panicked mid-compile of the execute route ("socket connection was closed
  unexpectedly"), 4 runs;
- wedged that compile until the 300s request timeout ("The operation timed
  out."), 2 runs.
No run reached the executor: the cleanup check found no open execution log.

- Each app step removes .next/dev before starting, so no app restores another
  app's cache.
- A failing step prints the server log tail, not only a startup failure.
- The suite compiles the execute route with a refused request under its own
  300s budget and named check; every later request and CLI run is bounded at
  60s. A timed-out or dropped request names the route and elapsed time, lands
  in the report with a null status, and a CLI timeout is reported as one.
@waleedlatif1 waleedlatif1 changed the title fix(ci): start each HTTP end-to-end app from an empty Turbopack dev cache fix(ci): stop each HTTP end-to-end app fully and start the next from an empty dev cache Oct 6, 2026
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 4 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Turn on auto-fix | Re-trigger cubic

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

CI evidence for head bd1e89d: I reran the "End-to-end over real HTTP" job 10 times, and all 10 attempts passed. Each attempt ran the SCIM, stop-after and desktop inbox steps, every app started from an empty .next/dev, and stop-session.sh never reported a process outliving next dev. Healthy runs before the fix took 35–99s for the stop-after step. With a cold cache it takes about 75–115s.

@waleedlatif1
waleedlatif1 merged commit 11f8190 into staging Oct 6, 2026
277 of 278 checks passed

This branch was previously deployed

1 inactive deployment
Preview — bd1e89dc Deployed Oct 6, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant