Skip to content

fix(runtime): a blocking WASM pipe read waits for its writer instead of failing at maxBlockingReadMs - #1966

Open
abcxff wants to merge 1 commit into
mainfrom
stack/fix-runtime-a-blocking-wasm-pipe-read-waits-for-its-writer-instead-of-failing-at-maxblockingreadms-zxsvyrlk
Open

fix(runtime): a blocking WASM pipe read waits for its writer instead of failing at maxBlockingReadMs#1966
abcxff wants to merge 1 commit into
mainfrom
stack/fix-runtime-a-blocking-wasm-pipe-read-waits-for-its-writer-instead-of-failing-at-maxblockingreadms-zxsvyrlk

Conversation

@abcxff

@abcxff abcxff commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator

After the #1959 fix WASM process.fd_read is non-blocking and the runner poll+retries. It still hard-failed a blocking fd with EAGAIN once maxBlockingReadMs elapsed, so a slow-but-alive writer (e.g. a stage that idles longer than the configured cap) surfaced a guest-visible 'Resource temporarily unavailable'/'Would block' and lost output — the same EAGAIN-on-a-blocking-fd contract violation, just relocated to the runner loop.

The kernel returns an empty read (EOF) as soon as writers == 0, so every EAGAIN in the loop means the writer is still alive. A blocking read must therefore keep waiting for data or EOF rather than failing at the cap. maxBlockingReadMs no longer fails a blocking read (reads are non-blocking so the reactor is never parked, and vm.exec's own timeoutMs still bounds a genuinely stuck pipeline); it now emits a one-shot 'still waiting' warning and only O_NONBLOCK fds return EAGAIN. Bumps NODE_IMPORT_CACHE_ASSET_VERSION so the runner re-materializes.

…of failing at maxBlockingReadMs

After the #1959 fix WASM process.fd_read is non-blocking and the runner poll+retries. It still hard-failed a *blocking* fd with EAGAIN once maxBlockingReadMs elapsed, so a slow-but-alive writer (e.g. a stage that idles longer than the configured cap) surfaced a guest-visible 'Resource temporarily unavailable'/'Would block' and lost output — the same EAGAIN-on-a-blocking-fd contract violation, just relocated to the runner loop.

The kernel returns an empty read (EOF) as soon as writers == 0, so every EAGAIN in the loop means the writer is still alive. A blocking read must therefore keep waiting for data or EOF rather than failing at the cap. maxBlockingReadMs no longer fails a blocking read (reads are non-blocking so the reactor is never parked, and vm.exec's own timeoutMs still bounds a genuinely stuck pipeline); it now emits a one-shot 'still waiting' warning and only O_NONBLOCK fds return EAGAIN. Bumps NODE_IMPORT_CACHE_ASSET_VERSION so the runner re-materializes.
@abcxff

abcxff commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator Author

Stack for rivet-dev/agentos

Get stack: forklift get 1966
Push local edits: forklift submit
Merge when ready: forklift merge 1966

change zxsvyrlk

@railway-app

railway-app Bot commented Sep 9, 2026

Copy link
Copy Markdown

🚅 Environment agentos-pr-1966 in rivet-frontend has no services deployed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant