fix(runtime): a blocking WASM pipe read waits for its writer instead of failing at maxBlockingReadMs - #1966
Open
abcxff wants to merge 1 commit into
Conversation
…of failing at maxBlockingReadMs After the #1959 fix WASM process.fd_read is non-blocking and the runner poll+retries. It still hard-failed a *blocking* fd with EAGAIN once maxBlockingReadMs elapsed, so a slow-but-alive writer (e.g. a stage that idles longer than the configured cap) surfaced a guest-visible 'Resource temporarily unavailable'/'Would block' and lost output — the same EAGAIN-on-a-blocking-fd contract violation, just relocated to the runner loop. The kernel returns an empty read (EOF) as soon as writers == 0, so every EAGAIN in the loop means the writer is still alive. A blocking read must therefore keep waiting for data or EOF rather than failing at the cap. maxBlockingReadMs no longer fails a blocking read (reads are non-blocking so the reactor is never parked, and vm.exec's own timeoutMs still bounds a genuinely stuck pipeline); it now emits a one-shot 'still waiting' warning and only O_NONBLOCK fds return EAGAIN. Bumps NODE_IMPORT_CACHE_ASSET_VERSION so the runner re-materializes.
Collaborator
Author
|
Stack for rivet-dev/agentos
Get stack: change zxsvyrlk |
|
🚅 Environment agentos-pr-1966 in rivet-frontend has no services deployed. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
After the #1959 fix WASM process.fd_read is non-blocking and the runner poll+retries. It still hard-failed a blocking fd with EAGAIN once maxBlockingReadMs elapsed, so a slow-but-alive writer (e.g. a stage that idles longer than the configured cap) surfaced a guest-visible 'Resource temporarily unavailable'/'Would block' and lost output — the same EAGAIN-on-a-blocking-fd contract violation, just relocated to the runner loop.
The kernel returns an empty read (EOF) as soon as writers == 0, so every EAGAIN in the loop means the writer is still alive. A blocking read must therefore keep waiting for data or EOF rather than failing at the cap. maxBlockingReadMs no longer fails a blocking read (reads are non-blocking so the reactor is never parked, and vm.exec's own timeoutMs still bounds a genuinely stuck pipeline); it now emits a one-shot 'still waiting' warning and only O_NONBLOCK fds return EAGAIN. Bumps NODE_IMPORT_CACHE_ASSET_VERSION so the runner re-materializes.