Skip to content

sandbox exec reads piped stdin to EOF before starting the command; hangs forever when stdin is an open pipe #3993

Description

@fede-kamel

User Story

As an operator driving OpenShell from scripts, CI runners, or agent harnesses, I want openshell sandbox exec to start the remote command without waiting for stdin to reach EOF, so that non-interactive execs behave like docker exec and kubectl exec do when no input is given.

Problem Statement

When stdin is not a terminal, the CLI reads stdin to EOF (up to the 4 MiB cap) before it sends ExecSandbox. In crates/openshell-cli/src/run.rs (sandbox_exec_grpc), stdin_prefix is filled by a spawn_blocking read_to_end and only then is the request built. If the parent keeps the pipe open, the CLI blocks forever in read(2), the gateway never receives an exec RPC, and there is no error. --no-tty does not change this. Even when the pipe does close, the command only starts after EOF: with a producer that closes 5 s later, the remote command starts 5 s later.

Impact / Why This Matters

  • Any harness whose stdin is an open pipe (CI runners, supervisors, agent tooling) hangs indefinitely on every exec. The gateway log shows GetSandbox and nothing else, so the symptom looks like a relay or supervisor problem; it cost about an hour to diagnose here.
  • Current workaround: </dev/null on every non-interactive exec, or pipe real input. It is undocumented, easy to forget, and forgetting it produces a silent hang rather than an error.
  • Commands are delayed behind slow stdin producers even when those eventually close.

Acceptance Criteria

  • sleep 300 | openshell sandbox exec -n sb -- true returns exit 0 promptly: the command runs while stdin is still open, and remote stdin is closed when the pipe ends or the command exits.
  • Or: stdin forwarding becomes opt-in (-i / --no-stdin, following docker exec and kubectl exec), and the docs state that piped stdin is read before the command starts.
  • The small-pipe unary path for older gateways keeps working.

Reproduction Steps

  1. openshell sandbox create --name sb --detach -- sleep infinity (any image).
  2. sleep 300 | timeout 12 openshell sandbox exec -n sb -- true → killed by timeout, exit 124. Same result with --no-tty.
  3. openshell sandbox exec -n sb -- true </dev/null → exit 0 in 0.1 s.
  4. echo hi | openshell sandbox exec -n sb -- cat → prints hi (EOF is reached).
  5. (sleep 5; echo late) | openshell sandbox exec -n sb -- sh -c 'date +%s; cat' → the printed epoch is 5 s after the pipeline started, so the command starts only after stdin EOF.
  6. Gateway log: ExecSandbox RPCs appear for steps 3, 4, and 5 only; none for step 2.

Environment

  • OpenShell 0.1.2 CLI and gateway (Homebrew install), macOS 26.7 arm64.
  • Docker compute driver through Rancher Desktop (Docker 29.5, Alpine VM).
  • Sandbox image ghcr.io/astral-sh/uv:python3.12-bookworm-slim; any image reproduces.

Logs

Stack sample of the hung CLI: one non-tokio thread parked in read (libsystem_kernel), tokio workers idle, lsof shows an established connection to the gateway. Gateway log during the hang: GetSandbox only, no ExecSandbox or RelayStream. Supervisor log: no ssh relay open for the hung call.

1. never-closing pipe as stdin            -> rc=124 (timeout)
2. same with --no-tty                     -> rc=124 (timeout)
3. stdin closed                           -> rc=0   0.1s
4. small piped input with EOF             -> rc=0   0.1s  (prints hi)
5. EOF after 5 s; command prints its start time -> start_epoch = pipeline start + 5 s
6. gateway ExecSandbox RPCs in the window -> 3 (cases 3, 4, 5)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions