Skip to content

feat(sandbox): support stop requests conditional on sandbox execution identity #4009

Description

@danehans

User Story

As an operator building an event-driven controller, I want to stop only the sandbox execution that produced an observation, so delayed events cannot stop a newer execution.

Problem Statement

The public StopSandboxRequest accepts workspace_scope, name, and request_id. It cannot express an expected sandbox incarnation or execution. The request ID provides durable at-most-once admission; it is not a target precondition.

Consider this source-derived scenario (not yet reproduced in a live end-to-end test):

  1. Execution A produces an observation that is queued for processing.
  2. The sandbox is stopped and started, producing execution B under the same sandbox identity and name.
  3. The controller processes A's observation and requests a stop by workspace and name. It cannot require the Gateway to reject the request if B is now current.

Deleting and recreating a sandbox with the same name presents a similar stale-target problem.

Impact / Why This Matters

Controllers must either disable automatic stop or risk stopping a replacement execution. Reading current state before stopping leaves a race between the read and the mutation. Direct Kubernetes Pod UID checks/deletes are backend-specific and bypass the Gateway lifecycle contract.

This blocks portable automation that must bind an action to the execution observed. The finding is a missing API capability; this investigation has not established an authorization bypass or sandbox-isolation vulnerability.

Proposed Design

Retain the public workspace/name addressing convention and support a conditional stop workflow:

  • A controller receives trustworthy, Gateway/supervisor-attested execution identity with the relevant extension observation. Define its lifecycle semantics, including stop/start and runtime replacement.
  • The controller supplies an expected-target precondition when requesting stop. An opaque token is one possible public representation; internal identifiers need not become public entity references.
  • The Gateway checks that precondition as part of the serialized lifecycle operation. A stale or unverifiable target produces a distinguishable result without stopping the current replacement.
  • Responses and retries distinguish an action on the intended execution from a stale target or replay of an earlier result.

The exact field names, token representation, and internal implementation are open for maintainers to choose.

Acceptance Criteria

  • Supported extension observations expose authenticated metadata sufficient to bind a stop to the observed sandbox incarnation and execution.
  • A matching conditional request stops the intended execution through the normal authorized Gateway lifecycle.
  • A condition from before stop/start or same-name delete/recreate cannot stop the replacement, including when replacement races with the request.
  • Missing, stale, or unverifiable conditions in the conditional workflow never silently fall back to an unconditional stop.
  • Retry/replay behavior preserves the original target and does not imply that a replacement was stopped.
  • The contract documents identity lifetime, result semantics, and compatibility with existing unconditional stop, and has regression coverage plus at least one real-driver end-to-end test.

Alternatives Considered

  • Client-side read/check before stop: cannot close the read-to-mutation race.
  • Use only the persistent sandbox ID: distinguishes same-name recreation, but not stop/start of the same sandbox.
  • Use backend-native deletion: sacrifices portability and Gateway lifecycle semantics.
  • Use a durable stop hold: addresses subsequent activation, but does not establish which execution a delayed stop should target.

Agent Investigation

Reviewed current main at 912a077bd641272016fb8b2fd58209f6c7c6f194:

Related: #3556 concerns driver-facing sandbox references; #3153 concerns durable activation holds. Neither defines this conditional execution-targeting workflow.

Checklist

  • I've reviewed existing issues and the published docs.
  • This is a design proposal, not a "please build this" request.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions