Skip to content

BeginFrame probe times out with 6+ workers on one GPU, and the render silently falls back to screenshot capture #4584

Description

@crmne

Summary

With --browser-gpu on a hardware GPU, once 6 or more workers start together, every worker's BeginFrame probe times out. Each worker relaunches in screenshot mode, and the render takes about twice as long but still exits 0. The only signs are the summary line (screenshot capture · hardware gpu) and one [BrowserManager] warning per worker. --workers auto picks 6 on this machine, so a default GPU render always takes the slow path.

Environment

hyperframes 0.8.79 (same on main at e00baef), chrome-headless-shell 152.0.7977.30, ANGLE on EGL, Linux 7.2.5, NVIDIA RTX 3090 (driver 610.57.04), Node 25.9.0. hyperframes doctor passes.

Reproduction

A 10 s 1080p30 composition with one GSAP tween (no WebGL, no media):

index.html
<!doctype html>
<html lang="en">
  <head>
    <meta charset="UTF-8" />
    <script src="https://cdn.jsdelivr.net/npm/gsap@3.14.2/dist/gsap.min.js"></script>
    <style>
      html, body { margin: 0; background: #111; }
      #root { position: relative; width: 1920px; height: 1080px; overflow: hidden; }
      .box { position: absolute; top: 490px; left: 100px; width: 100px; height: 100px; background: #3af; }
    </style>
  </head>
  <body>
    <div id="root" data-composition-id="main" data-start="0" data-width="1920" data-height="1080" data-duration="10" data-fps="30">
      <div class="box" id="box"></div>
    </div>
    <script>
      window.__timelines = window.__timelines || {};
      const tl = gsap.timeline({ paused: true });
      tl.to("#box", { x: 1620, rotation: 360, duration: 10, ease: "none" }, 0);
      window.__timelines["main"] = tl;
    </script>
  </body>
</html>
npx hyperframes@0.8.79 render --browser-gpu --gpu --workers 6 --output out.mp4; echo "exit $?"
# [BrowserManager] HeadlessExperimental.beginFrame probe failed after 2001ms: beginFrame probe timeout during warm-up beginFrame; falling back to screenshot mode.  (x6)
# screenshot capture · hardware gpu · ... capture 7.5s
# exit 0
Workers Capture path Wall time (median) Probe timeouts
1 to 5 beginframe (every run) 5.4 to 6.5 s 0
6, 7, 8 screenshot (every run) 9.6 to 11.0 s all workers
auto (6) screenshot 10.1 s 5 or 6 of 6
60 s clip, 5 / 6 workers beginframe / screenshot 10.3 s / 22.2 s 0 / 6 of 6

5 workers is close to the edge: one run at 5 also timed out, and with a CPU-heavy job running alongside, 4 workers sometimes did too.

Cause

The first HeadlessExperimental.beginFrame does complete, just slowly while other browsers start on the same GPU. Timing it outside Hyperframes, with the exact worker flags and a 30 s deadline, it takes about 0.4 s for one browser and about 0.33 s more for each browser started alongside: 1.9 to 2.2 s at 6, and 2.7 to 3.0 s at 8. On SwiftShader the same step takes 21 to 159 ms.

  • Parallel capture starts all workers at once (parallelCoordinator.ts:1120), each probed with a single 2000 ms deadline (BEGINFRAME_PROBE_TIMEOUT_MS, browserManager.ts:256).
  • A probe failure only logs a warning and relaunches the browser in screenshot mode (browserManager.ts:729), so it never reaches the exit code.
  • defaultSafeMaxWorkers() never goes below 6 (parallelCoordinator.ts:165), whatever the GPU mode.
  • formatScreenshotFallbackHint only prints for a software GPU, so this case gets no hint.

Related, but not the same: #410 (auto chose 6 on a WebGL-heavy composition), #955 (SwiftShader probe contention at 6 workers per pod), #2810 (a Chrome build without BeginFrame falling back silently on Cloud Run), and #3482 (a fallback-ratio guard for drawElement).

Proposal

  1. Make the fallback an error on request: --require-beginframe (env PRODUCER_REQUIRE_BEGINFRAME) fails the render with the reason instead of continuing in screenshot capture. PR to follow.
  2. Fix the cause: retry a timed-out probe once. Locally, 6 workers, 8 workers and auto then ran BeginFrame every time (9 of 9 renders), and the 60 s clip at 6 workers took 9.7 s instead of 22 s. A probe that genuinely fails still falls back at once. Happy to send this as a PR too. Probing the workers one at a time wasn't enough on its own (7 of 8 passed).
  3. Print the fallback hint for hardware GPUs too.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions