From aaca7e5ca1682f903880b5e52754f83fa4413705 Mon Sep 17 00:00:00 2001 From: openhands Date: Fri, 25 Sep 2026 12:42:28 +0000 Subject: [PATCH 1/5] docs: restructure custom sandbox images into landing page with sub-pages - Replace single enterprise/custom-sandbox-image.mdx with a five-page structure under enterprise/custom-sandbox-images/ - Landing page: warm runtime concept, prerequisites, choose-your-path cards - building-custom-images: Dockerfile, versioning, what to bake, private registries - multiple-images-warm-pools: full step-by-step guide (VM + Helm tabbed), inline scripts for both install types, config format, troubleshooting, API ref - single-image-admin-console: deprecated notice at top, Admin Console steps - using-custom-images: per-user selection, per-conversation API, warm claim behavior - docs.json: replace flat nav entry with group; add redirects from old URL Co-authored-by: openhands --- docs.json | 22 +- enterprise/custom-sandbox-image.mdx | 482 ------------- .../building-custom-images.mdx | 110 +++ enterprise/custom-sandbox-images/index.mdx | 84 +++ .../multiple-images-warm-pools.mdx | 642 ++++++++++++++++++ .../single-image-admin-console.mdx | 44 ++ .../using-custom-images.mdx | 64 ++ 7 files changed, 964 insertions(+), 484 deletions(-) delete mode 100644 enterprise/custom-sandbox-image.mdx create mode 100644 enterprise/custom-sandbox-images/building-custom-images.mdx create mode 100644 enterprise/custom-sandbox-images/index.mdx create mode 100644 enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx create mode 100644 enterprise/custom-sandbox-images/single-image-admin-console.mdx create mode 100644 enterprise/custom-sandbox-images/using-custom-images.mdx diff --git a/docs.json b/docs.json index e65defdbe..49f127457 100644 --- a/docs.json +++ b/docs.json @@ -538,7 +538,17 @@ "enterprise/enterprise-vs-oss", "enterprise/sizing-guide", "enterprise/quick-start", - "enterprise/custom-sandbox-image", + { + "group": "Custom Sandbox Images", + "icon": "box", + "pages": [ + "enterprise/custom-sandbox-images/index", + "enterprise/custom-sandbox-images/building-custom-images", + "enterprise/custom-sandbox-images/multiple-images-warm-pools", + "enterprise/custom-sandbox-images/single-image-admin-console", + "enterprise/custom-sandbox-images/using-custom-images" + ] + }, "enterprise/docker-in-sandbox", "enterprise/external-postgres", "enterprise/troubleshooting" @@ -881,6 +891,14 @@ { "source": "/openhands/usage/automations/examples", "destination": "/openhands/usage/automations/overview" + }, + { + "source": "/enterprise/custom-sandbox-image", + "destination": "/enterprise/custom-sandbox-images" + }, + { + "source": "/enterprise/custom-sandbox-image#run-multiple-custom-images-with-warm-runtime-pools", + "destination": "/enterprise/custom-sandbox-images/multiple-images-warm-pools" } ] -} \ No newline at end of file +} diff --git a/enterprise/custom-sandbox-image.mdx b/enterprise/custom-sandbox-image.mdx deleted file mode 100644 index 271f45a56..000000000 --- a/enterprise/custom-sandbox-image.mdx +++ /dev/null @@ -1,482 +0,0 @@ ---- -title: Custom Sandbox Images -description: Preload repos, dependencies, and tooling into custom sandbox images, and run multiple images side by side with warm runtime pools. -icon: box ---- - -Custom sandbox images let you prebake the repository, dependencies, compiled output, and test harness -your agents need. Instead of spending minutes provisioning a workspace on every run, your agents start -on the actual task immediately. - -This page covers two levels of customization: - -1. **[A single custom image](#configure-a-single-custom-image-admin-console)** that replaces the default - sandbox image for the whole installation. Configured in the Replicated Admin Console; no cluster access needed. -2. **[Multiple custom images](#run-multiple-custom-images-with-warm-runtime-pools)** running side by side, - each with its own warm pool, selectable per user. Configured through the Runtime API; requires `kubectl` access. - -## Why Use a Custom Image - -Custom images eliminate cold-start setup work (clone, install, transpile, and bootstrap) so agents -spend their time on the actual task. They also reduce setup variance and lower sandbox memory requirements -by keeping only what the agent needs. - -With **multiple** custom images, different teams get different environments: a PHP image with Composer and -MySQL client for the web team, a JDK and Maven image for the Java services team, a data science image with -pinned Python packages for the analytics team. Each image is kept ready in its own warm pool so conversations -start in seconds regardless of which environment they use. - -## Build Your Own Custom Image - -The [OpenHands agent-server sandbox guide](https://docs.openhands.dev/sdk/guides/agent-server/docker-sandbox) -provides full documentation on building custom sandbox images. The approach is the same for the Enterprise -Replicated VM deployment. - -### Basic Pattern - -1. Start from the OpenHands agent-server base image. -2. Keep the normal OpenHands entrypoint intact: extend the image, do not replace the entrypoint. -3. Add your repo, docs, tools, and verification wrappers. -4. Pre-run the expensive setup you do not want to repeat at task time. -5. Publish the image to a registry reachable from your OpenHands cluster. - - - Do not override the entrypoint or replace the runtime contract of the base image. The installer - expects standard OpenHands agent-server behavior. Only extend, do not replace. - - -### Base Image - -```dockerfile -FROM ghcr.io/openhands/agent-server:1.46.0-python -``` - -This example matches OpenHands Enterprise 0.64.0. Pin a specific version tag to ensure reproducible -builds, and replace it with the tag expected by your installed release. Check -[ghcr.io/openhands/agent-server](https://github.com/OpenHands/OpenHands/pkgs/container/agent-server) -for available tags. - -### Version Compatibility - -Each OpenHands Enterprise release expects a specific agent-server version. The base image tag you -build from must match the release you run: the `openhands-sdk` inside the sandbox and the one inside -the OpenHands application must agree on major and minor version. - -To find the expected tag, enable **Use a Custom Sandbox Image** in the Admin Console. The -**Sandbox Image Tag** field defaults to the tag the current release expects. - -When a conversation starts on a custom image, OpenHands checks the sandbox's agent-server version. -If it does not match the release, the conversation fails with an error naming the expected and -actual versions. Rebuild your image from the expected tag and update the **Sandbox Image Tag** -field to fix it. - - - Rebuild your custom image before each upgrade. The agent-server base image changes with every - OHE release, and an image built for an older release will be rejected by the version check. - - -### Example: Build and Push - -```bash -docker buildx build \ - --platform linux/amd64 \ - -f your-project/Dockerfile \ - -t ghcr.io//openhands-custom-image: \ - --push \ - . -``` - -Use `--platform linux/amd64` because the Enterprise Replicated VM runs on `x86-64`. - -### What to Bake In - -Good candidates for prebaking: - -- Pinned repository checkouts -- Package manager caches and installed dependencies (`node_modules`, Python virtualenvs, etc.) -- Compiled or transpiled output -- Native system packages (`xvfb`, `libkrb5-dev`, `pkg-config`, etc.) -- Browser or Electron artifacts -- Stable helper scripts such as `prepare-*` and `*-verify` wrappers - -### What to Keep Out - - - Do not bake the following into your image: - - - Secrets, API keys, or personal credentials - - Machine-specific paths or environment assumptions - - Uncommitted source changes or task-specific fixes - - Rapidly changing dependencies (use a lightweight `prepare-*` helper instead) - - -If the repository or dependencies change frequently, include a `prepare-*` script in the image -so the agent can refresh only the parts that need updating without a full rebuild. - -## Configure a Single Custom Image (Admin Console) - -Once your image is built and pushed to a registry, point the Replicated Admin Console at it. - -1. Open the **Admin Console** at `https://admin.:30000`. -2. Navigate to **Config** and find the **Sandbox Configuration** section. -3. Set the following fields: - -| Field | Value | -|---|---| -| **Use a Custom Sandbox Image** | Enabled | -| **Sandbox Image Repository** | Your image repository (e.g. `ghcr.io/your-org/openhands-custom-image`) | -| **Sandbox Image Tag** | Your image tag (e.g. `v1.2.0`) | -| **Registry Server** | If your registry requires authentication | -| **Registry Username** | If your registry requires authentication | -| **Registry Password or Credentials** | If your registry requires authentication | - -4. Click **Save config** and then **Deploy** to apply the change. - -This single image becomes both the default image for new conversations and the image kept ready in the -installer-managed warm pool. - - - This setting applies to the **sandbox / agent-server image** only (the image that runs inside each - agent's isolated workspace). It does not replace the other OpenHands service images. - - -## Run Multiple Custom Images with Warm Runtime Pools - -To offer several sandbox images at once, configure **warm runtime pools** through the Runtime API. -Each configuration names one image and keeps a pool of pre-started sandbox pods ready for it. The -OpenHands application automatically exposes every configuration as a selectable sandbox, so users can -pick their environment without any redeployment. - -**Requirements:** - -- OpenHands Enterprise **0.64.0 or later**. -- `kubectl` access to the cluster. On a Replicated VM install, get a shell with - `sudo /var/lib/embedded-cluster/bin/openhands shell`; on a Helm install, use your normal kubeconfig. -- Custom images built and pushed as described above (all on the agent-server version your release expects). - -### How It Works - -- The installer-managed configuration remains the base configuration. Configurations saved through the - Runtime API are overlaid by name: a new name adds a pool, while an existing name overrides that - installer-managed entry. -- You manage database configurations with the admin REST endpoints - (`PUT` / `DELETE /api/admin/warm-runtime-configs/{name}`). Deleting an override reveals the - installer-managed entry again. -- A reconciler job runs **every minute** and creates or removes warm sandbox pods so each - configuration has `count` unclaimed pods ready. -- The OpenHands application polls the configuration list (cached for 60 seconds) and exposes each - configuration as a **sandbox spec**. Users choose their default in **Settings → Application → Default Sandbox**. -- When a conversation starts, the Runtime API hands it a matching warm pod in a few seconds. If no - warm pod is available, the sandbox cold-starts from the image instead (20+ seconds), and the - reconciler replenishes the pool. - -Changes take effect within about a minute, with no application restarts and no redeployments. - - - The installer-managed `v1_current` pool remains active when you add API-managed configurations. Do - not save a `v1_current` configuration unless you intentionally want to override the installer default. - - -### Step 1: Confirm the Admin Password - -The Runtime API's admin endpoints authenticate with an admin password. Replicated generates a durable -password, stores it in the `admin-password` secret, and injects it into the runtime-api pod. The helper -script in Step 2 uses that pod environment. If the value is empty, the script reports an error before a -save or delete. - -To set or rotate the password: - -1. Open the `Admin Console` and select `Config`. -2. In `Sandbox Configuration`, set `Runtime API Admin Password`. -3. Select `Save config`, then deploy the new configuration. - -The value persists across later deploys. Changing it automatically restarts runtime-api so the new -password takes effect. For a Helm installation, populate the chart's `admin-password` Secret before -using the admin endpoints and restart runtime-api after changing it. - -### Step 2: Save the Helper Script - -The Runtime API is not exposed outside the cluster by default. Download the maintained -[`warm-runtime-configs.sh`](https://github.com/OpenHands/runtime-api/blob/main/scripts/warm-runtime-configs.sh) -helper, which runs each API call inside the runtime-api pod with `kubectl exec`: - -```bash -curl -fsSLo warm-runtime-configs.sh \ - https://raw.githubusercontent.com/OpenHands/runtime-api/main/scripts/warm-runtime-configs.sh -chmod +x warm-runtime-configs.sh -./warm-runtime-configs.sh list -``` - - - The helper uses `DEFAULT_API_KEY` and `ADMIN_PASSWORD` from the runtime-api pod without printing or - copying either value. Listing authenticates with the regular API key. Saving and deleting use the - admin password via a challenge-response login that returns a 24-hour JWT. - - -List responses identify each configuration's `source`. When `v1_current` is not overridden, it appears -with `"source": "file"`. - -### Step 3: Start From the Installer's Default Configuration - -Do not write configurations from scratch. The environment in a warm runtime configuration is what its -sandbox pods actually boot with; the default configuration contains install-specific values (webhook -callback URL, CA bundles, workspace paths) that sandboxes need to function. Export the default from the -installer-managed ConfigMap and use it as your template: - -```bash -kubectl -n openhands get configmap warm-runtimes-config \ - -o jsonpath='{.data.warm-runtimes\.json}' \ - | jq '.configs[] | select(.name == "v1_current") | del(.name)' > default-config.json -``` - -(If the ConfigMap has a different name in your install, find it with -`kubectl -n openhands get configmap | grep warm-runtimes`.) - -The installer-managed `v1_current` entry remains live and follows Admin Console changes. Derive each -custom image configuration from the exported template, changing only the image and pool size: - -```bash -jq '.image = "ghcr.io/your-org/openhands-php:8.4-v1" | .count = 1' \ - default-config.json > php-web.json -./warm-runtime-configs.sh save php-web php-web.json - -./warm-runtime-configs.sh list -``` - -### Configuration Format - -| Field | Type | Required | Description | -|-------|------|----------|-------------| -| `image` | string | Yes | Full image reference (e.g. `ghcr.io/your-org/openhands-php:8.4-v1`) | -| `working_dir` | string | Yes | Working directory inside the sandbox (copy from the default) | -| `command` | array | Yes | Agent-server start command (copy from the default) | -| `environment` | object | Yes | Environment variables the sandbox boots with (copy from the default) | -| `count` | integer | No | Warm pods to keep ready for this image. When omitted, uses the installer-wide **Warm Runtime Count** (default: `1` on Replicated installs) | -| `run_as_user` | integer | No | Copy from the installer default (currently `10001`) so warm pods match application start requests | -| `run_as_group` | integer | No | Copy from the installer default (currently `10001`) so warm pods match application start requests | -| `fs_group` | integer | No | Copy from the installer default (currently `10001`) so warm pods match application start requests | - -The configuration name comes from the URL path (the `save ` argument), not the body. Saving -creates or replaces a database entry. If an installer-managed entry has the same name, the database -entry overrides it. List responses also include a read-only `source` field: `file` for installer-managed -entries and `db` for API-managed entries and overrides. Do not add `source` to a saved configuration. - -The application uses the image reference as the sandbox spec ID. Give every selectable configuration a -distinct image reference; configurations that share an image reference cannot be selected independently, -even if their commands or environments differ. - - - Set `count` explicitly. Every warm pod reserves the full sandbox resource envelope (including 10Gi of - ephemeral storage by default) whether or not it is in use, so the sum of all pool sizes must fit your - node capacity. Pools that exceed capacity show up as `Pending` pods. Start with `count: 1` per image - and grow the pools that see real traffic. - - -### Step 4: Verify the Warm Pools - -The reconciler runs every minute. Watch it create the pods: - -```bash -# Warm (unclaimed) sandboxes: runtime deployments with no session_id label yet -kubectl -n openhands get deploy -l 'runtime_id,!session_id' \ - -o custom-columns='NAME:.metadata.name,READY:.status.readyReplicas,IMAGE:.spec.template.spec.containers[0].image' -``` - -You should see one `runtime-` deployment per warm pod, with your configured images. To see -the reconciler's own view (per-pool counts, pull failures, culling decisions), read the latest -reconciler job log: - -```bash -JOB=$(kubectl -n openhands get jobs --sort-by=.metadata.creationTimestamp -o name \ - | grep warm-runtimes | tail -1) -kubectl -n openhands logs "$JOB" -``` - -If a pod is stuck pulling your image, `kubectl -n openhands describe pod ` shows the pull -error. For private registries, either fill in the **Registry Server / Username / Password** fields in -the Admin Console's **Sandbox Configuration** section (they render an image pull secret that runtime -pods use), or add your own secret name to the runtime-api `RUNTIME_IMAGE_PULL_SECRETS` setting. - -### Step 5: Pick an Image and Start a Conversation - -Within a minute of saving configurations (the application caches the list for 60 seconds): - -- **Per user**: each user opens **Settings → Application** and picks an image in the **Default Sandbox** - dropdown (entries are the image references). Before a user picks an image, the application uses the - configuration named `v1_current`, or the first configuration if no `v1_current` exists. All of the - user's new conversations use their selected image. -- **Per conversation (API)**: start a sandbox for a specific image, then attach a conversation to it: - - ```bash - # 1. Start a sandbox from a specific spec (the spec id is the image reference) - curl -X POST "https://app./api/v1/sandboxes?sandbox_spec_id=ghcr.io/your-org/openhands-php:8.4-v1" \ - -H "Authorization: Bearer $API_KEY" - # 2. Create the conversation on that sandbox, using "id" from the response - curl -X POST "https://app./api/v1/app-conversations" \ - -H "Authorization: Bearer $API_KEY" \ - -H "Content-Type: application/json" \ - -d '{"sandbox_id": ""}' - ``` - -To confirm a conversation claimed a warm pod rather than cold-starting, note that its sandbox was -ready in a few seconds, or check the cluster: the claimed runtime deployment now carries a -`session_id` label, and the reconciler creates a fresh warm pod to replace it within a minute. - -### How Warm Pods Are Claimed - -A conversation claims a warm pod only when the pod **exactly matches** the requested image, command, -working directory, environment (ignoring a fixed set of session-specific variables), and -`run_as_user` / `run_as_group` / `fs_group`. Because the application requests exactly what the -selected configuration declares, conversations started through OpenHands match automatically. - -Cold starts still happen when: - -- All warm pods for the image are already claimed (`count` too low for the traffic). -- The configuration changed in the last minute, so the old pods no longer match and replacements are - still starting. -- Warm pods cannot become ready (image pull failures, insufficient node resources). - -Cold-started conversations run the same image and work normally; they just take 20+ seconds to begin. - -### Updating and Deleting Configurations - -Update by saving the same name again. To roll out a new image version: - -```bash -jq '.image = "ghcr.io/your-org/openhands-php:8.4-v2"' php-web.json > php-web-v2.json -./warm-runtime-configs.sh save php-web php-web-v2.json -``` - -Within a minute the reconciler stops the old pods and starts pods on the new image. Delete a -database-only configuration to remove its pool: - -```bash -./warm-runtime-configs.sh delete php-web -``` - -If the deleted name overrides an installer-managed entry, the underlying entry becomes effective again -instead of disappearing. Confirm the result with `./warm-runtime-configs.sh list`; its `source` changes -from `db` to `file`. - -Keep superseded image tags available in your registry while conversations that used them can still -resume: a paused conversation resumes on its **original** image. Delete old tags only after the -conversations that used them are gone (by default, stopped sandboxes are cleaned up after 10 days). - -### After Upgrading OpenHands Enterprise - - - API-managed configurations are **frozen snapshots**; upgrades do not touch them. The installer-managed - `v1_current` entry updates automatically unless a database entry with that name overrides it. Each - release expects a specific agent-server version and may add or change sandbox environment variables. - After every OpenHands Enterprise upgrade: - - 1. Rebuild your custom images on the release's new agent-server base version. - 2. Re-export the default template (Step 3) from the refreshed ConfigMap. - 3. Re-derive and save each API-managed custom configuration from the new template. - 4. If you intentionally override `v1_current`, refresh or delete that override so the new - installer-managed entry can take effect. - - Skipping this leaves configurations pointing at the previous agent-server version, and new - conversations fail with a version mismatch error until the configurations are updated. - - -### Return an Entry to Installer Management - -Delete a same-named database override to restore the installer-managed entry on the next reconciler -cycle. For example, if `v1_current` was intentionally overridden: - -```bash -./warm-runtime-configs.sh delete v1_current -./warm-runtime-configs.sh list # v1_current now reports "source": "file" -``` - -Other API-managed configurations continue running. Delete them individually when you no longer want -their pools or images in the application's selector. - -### Troubleshooting - -| Symptom | Cause and fix | -|---|---| -| `HTTP 403: Admin functionality is disabled` | The runtime-api deployment has no admin password, or the configured value is empty. On Replicated installs, set `Runtime API Admin Password` per Step 1 and deploy. | -| `HTTP 401` on login | Wrong password, or the challenge expired. Challenges are single-use and expire after 5 minutes; the script fetches a fresh one per call. | -| `HTTP 401: ...provide a valid API key...` on list | The list endpoint authenticates with `X-API-Key`, not the admin JWT. Use the helper script. | -| Installer-managed default is missing from the list | The release does not include overlay support, or the overlay is not enabled. Upgrade OpenHands Enterprise and confirm that the list reports `source` before saving configurations. | -| Saved a config but the dropdown does not show it | The application caches the list for 60 seconds; wait a minute and reload. Also confirm with `./warm-runtime-configs.sh list`. | -| No warm pods appear | Read the latest reconciler job log (Step 4). Look for image pull errors or scheduling failures. | -| Warm pods `Pending` | Insufficient node resources. Every warm pod reserves the full sandbox resource envelope; lower the pool `count`s or add capacity. | -| Conversations cold-start despite warm pods | Pool exhausted or configuration recently changed; see [How Warm Pods Are Claimed](#how-warm-pods-are-claimed). | -| Sandbox fails at start with an agent-server version error | The custom image's base version does not match the release. Rebuild on the expected agent-server version (see [Base Image](#base-image)). | -| Conversations on a custom image start but never show agent output | The configuration's `environment` is missing install-specific values (webhook callback URL, CA bundles). Rebuild the configuration from the default template (Step 3). | - -### API Reference - -The endpoints below are served by the runtime-api service (in-cluster: `http://:5000`). - -**Admin authentication** (for save and delete): - -1. `GET /api/admin/challenge` returns `{challenge, salt, iterations}`. Challenges are single-use and - expire after 5 minutes. -2. Compute `PBKDF2-HMAC-SHA256(password, salt + challenge, iterations, dklen=32)` and hex-encode it. -3. `POST /api/admin/login` with `{"challenge": ..., "hash": ...}` returns `{"token": ...}`, a JWT - valid for 24 hours. -4. Send `Authorization: Bearer ` on admin requests. - -**List configurations** (regular API key, not admin): - -```http -GET /api/warm-runtime-configs -X-API-Key: {api-key} -``` - -Returns `200` with the effective configuration set: - -```json -{ - "configs": [ - {"name": "v1_current", "image": "...", "source": "file"}, - {"name": "php-web", "image": "...", "source": "db"} - ] -} -``` - -Installer-managed entries have `source: "file"`. API-managed entries have `source: "db"`; a database -entry with the same name replaces the file entry in this effective list. - -**Create or update a configuration** (admin): - -```http -PUT /api/admin/warm-runtime-configs/{name} -Authorization: Bearer {admin-jwt} -Content-Type: application/json - -{"image": "...", "working_dir": "...", "command": [...], "environment": {...}, "count": 1} -``` - -Returns `200` with the saved configuration. Creates or overwrites; the name in the URL is the identity. - -**Delete a configuration** (admin): - -```http -DELETE /api/admin/warm-runtime-configs/{name} -Authorization: Bearer {admin-jwt} -``` - -Returns `200` with a confirmation message, or `404` if no database configuration has that name. When -the deleted name also exists in the installer-managed file, that file entry becomes effective again. - -## Reference - - - - Full SDK documentation on building custom sandbox images - - - Dockerfile, benchmark scripts, and analysis tooling for the VS Code custom image example - - - How conversations, sandboxes, and their lifecycle fit together - - - Capacity planning, including headroom for warm pools - - diff --git a/enterprise/custom-sandbox-images/building-custom-images.mdx b/enterprise/custom-sandbox-images/building-custom-images.mdx new file mode 100644 index 000000000..e58aa198e --- /dev/null +++ b/enterprise/custom-sandbox-images/building-custom-images.mdx @@ -0,0 +1,110 @@ +--- +title: Building a Custom Sandbox Image +description: How to build, version, and push a custom sandbox image for use with OpenHands Enterprise. +icon: docker +--- + +All custom sandbox image approaches — single-image Admin Console and multi-image +warm runtime pools — start here. Build your image once and then point whichever +configuration approach you use at it. + +## Basic Pattern + +1. Start from the OpenHands agent-server base image. +2. Keep the normal OpenHands entrypoint intact: extend the image, do not replace it. +3. Add your repo, docs, tools, and verification wrappers. +4. Pre-run the expensive setup you do not want to repeat at task time. +5. Push the image to a registry reachable from your OpenHands cluster. + + + Do not override the entrypoint or replace the runtime contract of the base + image. OpenHands expects standard agent-server behavior. Only extend, do not + replace. + + +## Base Image + +```dockerfile +FROM ghcr.io/openhands/agent-server:1.46.0-python +``` + +Pin a specific version tag to ensure reproducible builds, and replace it with +the tag expected by your installed release. See [Version Compatibility](#version-compatibility) +below to find the right tag. + +## Version Compatibility + +Each OpenHands Enterprise release expects a specific agent-server version. The +base image tag you build from must match the release you run: the +`openhands-sdk` inside the sandbox and the one inside the OpenHands application +must agree on major and minor version. + +To find the expected tag for your release, enable **Use a Custom Sandbox +Image** in the Admin Console. The **Sandbox Image Tag** field defaults to the +tag the current release expects. Check +[ghcr.io/openhands/agent-server](https://github.com/OpenHands/OpenHands/pkgs/container/agent-server) +for available tags. + +When a conversation starts on a custom image, OpenHands checks the sandbox's +agent-server version. If it does not match, the conversation fails with an +error naming the expected and actual versions. Rebuild your image from the +expected tag to fix it. + + + Rebuild your custom image before each OHE upgrade. The agent-server base + image changes with every release, and an image built for an older release + will be rejected by the version check. + + +## Build and Push + +```bash +docker buildx build \ + --platform linux/amd64 \ + -f your-project/Dockerfile \ + -t ghcr.io//openhands-custom-image: \ + --push \ + . +``` + +Use `--platform linux/amd64` because the Enterprise Replicated VM runs on x86-64. + +## What to Bake In + +Good candidates for prebaking: + +- Pinned repository checkouts +- Package manager caches and installed dependencies (`node_modules`, Python virtualenvs, etc.) +- Compiled or transpiled output +- Native system packages (`xvfb`, `libkrb5-dev`, `pkg-config`, etc.) +- Browser or Electron artifacts +- Stable helper scripts such as `prepare-*` and `*-verify` wrappers + +## What to Keep Out + + + Do not bake the following into your image: + + - Secrets, API keys, or personal credentials + - Machine-specific paths or environment assumptions + - Uncommitted source changes or task-specific fixes + - Rapidly changing dependencies (use a lightweight `prepare-*` script instead) + + +If the repository or dependencies change frequently, include a `prepare-*` +script in the image so the agent can refresh only the parts that need updating +without a full rebuild. + +## Private Registries + +If your image lives in a private registry, provide pull credentials so the +cluster can fetch it at pod start time. + +**Replicated VM installs:** set **Registry Server**, **Registry Username**, and +**Registry Password or Credentials** in **Config → Sandbox Configuration** in +the Admin Console and deploy. The installer renders an image pull secret that +runtime pods automatically use. + +**Helm installs:** create a pull secret in the `openhands` namespace and add +its name to the runtime-api `RUNTIME_IMAGE_PULL_SECRETS` environment variable +(comma-separated list of secret names). diff --git a/enterprise/custom-sandbox-images/index.mdx b/enterprise/custom-sandbox-images/index.mdx new file mode 100644 index 000000000..cd334ff52 --- /dev/null +++ b/enterprise/custom-sandbox-images/index.mdx @@ -0,0 +1,84 @@ +--- +title: Custom Sandbox Images +description: Prebake your repository, dependencies, and tooling into custom sandbox images so agents start on the actual task instead of spending time on setup. +icon: box +--- + +Custom sandbox images let you prebake the repository, dependencies, compiled +output, and test harness your agents need. Instead of spending minutes +provisioning a workspace on every run, your agents start on the actual task +immediately. + +## How Warm Runtime Pools Work + +The runtime-api keeps a set of pre-started sandbox pods ready before any +conversation is requested. When a user starts a conversation, it claims a +waiting pod in seconds instead of cold-starting one from scratch (which takes +20 seconds or more). Each configuration names one image and a pool size; a +reconciler runs every minute to maintain that count. Multiple configurations +run side by side, each with its own pool, each independently selectable by +users. + +## Prerequisites + +Before configuring any custom image, the following must be in place: + +**An image registry reachable from your OpenHands cluster.** The cluster must +be able to pull your custom image at pod start time. Public registries (GitHub +Container Registry, Docker Hub) work without extra configuration. Private +registries require credentials — either set via **Config → Sandbox +Configuration → Registry Server / Username / Password** in the Admin Console, +or via the `RUNTIME_IMAGE_PULL_SECRETS` setting on Helm installs. + +**A custom image built from the correct agent-server base.** See +[Building a Custom Image](/enterprise/custom-sandbox-images/building-custom-images). +The image must be pushed to your registry before you configure it. + +**OpenHands Enterprise 0.64.0 or later** for the warm runtime pool approach. + +**`kubectl` access to the cluster** for the warm runtime pool approach. On +Replicated VM installs: + +```bash +sudo /var/lib/embedded-cluster/bin/openhands shell +``` + +On Helm installs, use your normal kubeconfig. + +## Configuration Approaches + + + + One warm pool per image, selectable per user. Changes take effect within a + minute with no restarts. **Recommended.** + + + **Deprecated.** Configures one image for the whole installation via the + Replicated Admin Console. Superseded by the warm runtime pool approach. + + + +## Reference + + + + Dockerfile pattern, version pinning, and what to bake in + + + How users select an image and how to target one via the API + + + How conversations, sandboxes, and their lifecycle fit together + + + Capacity planning, including headroom for warm pools + + diff --git a/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx new file mode 100644 index 000000000..74476129d --- /dev/null +++ b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx @@ -0,0 +1,642 @@ +--- +title: Multiple Images with Warm Runtime Pools +description: Configure multiple custom sandbox images through the Runtime API, each kept ready in its own warm pool and independently selectable by users. +icon: layer-group +--- + +To offer several sandbox images at once, configure **warm runtime pools** +through the Runtime API. Each configuration names one image and keeps a pool +of pre-started sandbox pods ready for it. The OpenHands application +automatically exposes every configuration as a selectable sandbox so users can +pick their environment — no redeployment required. + +**Requirements:** + +- OpenHands Enterprise **0.64.0 or later** +- `kubectl` access to the cluster (see [Prerequisites](/enterprise/custom-sandbox-images#prerequisites)) +- Custom images built and pushed as described in [Building a Custom Image](/enterprise/custom-sandbox-images/building-custom-images) + +## How It Works + + + + Warm runtime configurations run in **overlay mode** on Replicated installs + (`WARM_RUNTIME_CONFIG_OVERLAY=1` is set by default). API-managed + configurations layer on top of the installer's ConfigMap entries rather than + replacing them. + + - Adding a new name creates an additional pool alongside the + installer-managed `v1_current` pool. + - Saving a name that already exists in the ConfigMap overrides that entry; + deleting the override reverts to the ConfigMap value. + - The installer-managed `v1_current` pool keeps running throughout — you + never lose the default pool by adding custom configurations. + + Changes take effect within about a minute with no application restarts and + no redeployments. + + + + **The first configuration you save takes over warm pool management.** + While the Runtime API database holds any configurations, the + installer-managed ConfigMap is ignored entirely. Always re-declare the + default image as a configuration (Step 3 does this). To hand control + back to the installer, delete **all** configurations. + + To enable overlay mode instead — so API configs layer on top of the + installer pool rather than replacing it, matching the Replicated + behavior — add `WARM_RUNTIME_CONFIG_OVERLAY: "1"` to the runtime-api + environment and restart the deployment. + + + Changes take effect within about a minute with no application restarts and + no redeployments. + + + +--- + +## Step 1: Confirm the Admin Password + +The Runtime API admin endpoints require an admin password. + + + + The password is **auto-generated at install** (`{{repl RandomString 32}}`) + and stored in the `admin-password` Kubernetes secret. The helper script in + Step 2 reads it from the pod environment automatically — no action + required for a standard installation. + + To set a memorable password or rotate the generated one: + + 1. Open the **Admin Console** at `https://admin.:30000`. + 2. Navigate to **Config → Sandbox Configuration → Runtime API Admin Password**. + 3. Enter your new password and click **Save config**, then **Deploy**. + + The Admin Console updates the secret and rolls out the runtime-api + automatically. The password persists across all future Admin Console + deploys. + + + Do not use `kubectl patch` to set the password. The Admin Console manages + the `admin-password` secret and overwrites it on every deploy, so a + patched value is lost the next time you save any config change. Always + use the Admin Console field. + + + + The password was set when you created the `admin-password` secret during + installation: + + ```bash + kubectl -n openhands create secret generic admin-password \ + --from-literal=admin-password= + ``` + + The helper script reads it automatically — no action required. + + To rotate the password: + + ```bash + # Store the new value somewhere secure before running this + kubectl -n openhands delete secret admin-password + kubectl -n openhands create secret generic admin-password \ + --from-literal=admin-password=$(openssl rand -base64 24) + kubectl -n openhands rollout restart deployment \ + -l app.kubernetes.io/name=runtime-api + kubectl -n openhands rollout status deployment \ + -l app.kubernetes.io/name=runtime-api + ``` + + + +--- + +## Step 2: Save the Helper Script + + + + The runtime-api is exposed externally at + `https://runtime-api.`. All API calls go directly to + that URL — no cluster shell or `kubectl exec` needed for day-to-day + operations. + + Save the script below as `warm-runtime-configs.sh`. On first run it reads + credentials from Kubernetes secrets; export them afterward and you can run + the script entirely without cluster access. + + ```bash + #!/usr/bin/env bash + # warm-runtime-configs.sh — manage warm runtime configurations (VM install). + # + # Usage: + # ./warm-runtime-configs.sh list + # ./warm-runtime-configs.sh save + # ./warm-runtime-configs.sh delete + # + # Credentials are read from Kubernetes secrets on first run. To run without + # cluster access afterward, export these before calling the script: + # export RUNTIME_API_URL=https://runtime-api. + # export API_KEY= + # export ADMIN_PASSWORD= + set -euo pipefail + + NAMESPACE="${NAMESPACE:-openhands}" + COMMAND="${1:?usage: $0 list|save |delete }" + CONFIG_NAME="${2:-}" + CONFIG_FILE="${3:-}" + + if [ -z "${RUNTIME_API_URL:-}" ]; then + RUNTIME_API_URL="https://$(kubectl get ingress -n "$NAMESPACE" \ + -l app.kubernetes.io/name=runtime-api \ + -o jsonpath='{.items[0].spec.rules[0].host}')" + fi + if [ -z "${API_KEY:-}" ]; then + API_KEY=$(kubectl get secret default-api-key -n "$NAMESPACE" \ + -o jsonpath='{.data.default-api-key}' | base64 -d) + fi + if [ "$COMMAND" != "list" ] && [ -z "${ADMIN_PASSWORD:-}" ]; then + ADMIN_PASSWORD=$(kubectl get secret admin-password -n "$NAMESPACE" \ + -o jsonpath='{.data.admin-password}' | base64 -d) + fi + if [ "$COMMAND" != "list" ] && [ -z "${ADMIN_PASSWORD:-}" ]; then + echo "Error: admin password is not set. See Step 1." >&2 + exit 1 + fi + + PYSCRIPT=' + import binascii, hashlib, json, os, sys, urllib.error, urllib.parse, urllib.request + + API_URL = os.environ["RUNTIME_API_URL"].rstrip("/") + + def req(path, method="GET", data=None, headers=None): + h = {"Content-Type": "application/json", **(headers or {})} + r = urllib.request.Request(f"{API_URL}{path}", method=method, headers=h) + if data is not None: + r.data = json.dumps(data).encode() + try: + with urllib.request.urlopen(r) as resp: + return json.loads(resp.read().decode()) + except urllib.error.HTTPError as e: + sys.exit(f"HTTP {e.code}: {e.read().decode()}") + except urllib.error.URLError as e: + sys.exit(f"Could not connect to {API_URL}: {e.reason}") + + def admin_token(): + chal = req("/api/admin/challenge") + dk = hashlib.pbkdf2_hmac( + "sha256", + os.environ["ADMIN_PASSWORD"].encode(), + (chal["salt"] + chal["challenge"]).encode(), + chal["iterations"], + dklen=32, + ) + resp = req("/api/admin/login", "POST", + {"challenge": chal["challenge"], "hash": binascii.hexlify(dk).decode()}) + return resp["token"] + + action = os.environ["ACTION"] + name = urllib.parse.quote(os.environ.get("CONFIG_NAME", ""), safe="") + + if action == "list": + configs = req("/api/warm-runtime-configs", + headers={"X-API-Key": os.environ["API_KEY"]})["configs"] + print(json.dumps(configs, indent=2)) + elif action == "save": + body = json.load(sys.stdin) + token = admin_token() + saved = req(f"/api/admin/warm-runtime-configs/{name}", "PUT", body, + headers={"Authorization": f"Bearer {token}"}) + print("Saved", saved["name"], "image:", saved["image"], "count:", saved.get("count")) + elif action == "delete": + token = admin_token() + resp = req(f"/api/admin/warm-runtime-configs/{name}", "DELETE", + headers={"Authorization": f"Bearer {token}"}) + print(resp["message"]) + ' + + case "$COMMAND" in + list) + ACTION=list API_KEY="$API_KEY" RUNTIME_API_URL="$RUNTIME_API_URL" \ + python3 -c "$PYSCRIPT" + ;; + save) + if [ -z "$CONFIG_NAME" ] || [ ! -f "$CONFIG_FILE" ]; then + echo "usage: $0 save " >&2; exit 1 + fi + ACTION=save CONFIG_NAME="$CONFIG_NAME" ADMIN_PASSWORD="$ADMIN_PASSWORD" \ + RUNTIME_API_URL="$RUNTIME_API_URL" python3 -c "$PYSCRIPT" < "$CONFIG_FILE" + ;; + delete) + if [ -z "$CONFIG_NAME" ]; then + echo "usage: $0 delete " >&2; exit 1 + fi + ACTION=delete CONFIG_NAME="$CONFIG_NAME" ADMIN_PASSWORD="$ADMIN_PASSWORD" \ + RUNTIME_API_URL="$RUNTIME_API_URL" python3 -c "$PYSCRIPT" + ;; + esac + ``` + + ```bash + chmod +x warm-runtime-configs.sh + ./warm-runtime-configs.sh list + ``` + + After the first run, export the discovered values so subsequent runs need + no cluster access at all: + + ```bash + NAMESPACE=openhands + echo "export RUNTIME_API_URL=https://$(kubectl get ingress -n "$NAMESPACE" \ + -l app.kubernetes.io/name=runtime-api \ + -o jsonpath='{.items[0].spec.rules[0].host}')" + echo "export API_KEY=$(kubectl get secret default-api-key -n "$NAMESPACE" \ + -o jsonpath='{.data.default-api-key}' | base64 -d)" + echo "export ADMIN_PASSWORD=$(kubectl get secret admin-password -n "$NAMESPACE" \ + -o jsonpath='{.data.admin-password}' | base64 -d)" + ``` + + + The runtime-api is not exposed outside the cluster by default on Helm + installs. The script tunnels each API call into the runtime-api pod via + `kubectl exec`. You need `kubectl` access for every operation. + + Save the script below as `warm-runtime-configs.sh`: + + ```bash + #!/usr/bin/env bash + # warm-runtime-configs.sh — manage warm runtime configurations (Helm install). + # + # Usage: + # ./warm-runtime-configs.sh list + # ./warm-runtime-configs.sh save + # ./warm-runtime-configs.sh delete + set -euo pipefail + + NAMESPACE="${NAMESPACE:-openhands}" + COMMAND="${1:?usage: $0 list|save |delete }" + CONFIG_NAME="${2:-}" + CONFIG_FILE="${3:-}" + + POD=$(kubectl get pods -n "$NAMESPACE" -l app.kubernetes.io/name=runtime-api \ + -o jsonpath='{.items[0].metadata.name}') + if [ -z "$POD" ]; then + echo "Error: no runtime-api pod found in namespace $NAMESPACE." >&2; exit 1 + fi + + PYSCRIPT=' + import binascii, hashlib, json, os, sys, urllib.error, urllib.parse, urllib.request + + API_URL = "http://localhost:5000" + + def required_env(name): + value = os.environ.get(name) + if not value: + sys.exit(f"Error: {name} is not set in the runtime-api pod.") + return value + + def req(path, method="GET", data=None, headers=None): + h = {"Content-Type": "application/json", **(headers or {})} + r = urllib.request.Request(f"{API_URL}{path}", method=method, headers=h) + if data is not None: + r.data = json.dumps(data).encode() + try: + with urllib.request.urlopen(r) as resp: + return json.loads(resp.read().decode()) + except urllib.error.HTTPError as e: + sys.exit(f"HTTP {e.code}: {e.read().decode()}") + except urllib.error.URLError as e: + sys.exit(f"Could not connect: {e.reason}") + + def admin_token(): + chal = req("/api/admin/challenge") + dk = hashlib.pbkdf2_hmac( + "sha256", + required_env("ADMIN_PASSWORD").encode(), + (chal["salt"] + chal["challenge"]).encode(), + chal["iterations"], + dklen=32, + ) + resp = req("/api/admin/login", "POST", + {"challenge": chal["challenge"], "hash": binascii.hexlify(dk).decode()}) + return resp["token"] + + action = os.environ["ACTION"] + name = urllib.parse.quote(os.environ.get("CONFIG_NAME", ""), safe="") + + if action == "list": + configs = req("/api/warm-runtime-configs", + headers={"X-API-Key": required_env("DEFAULT_API_KEY")})["configs"] + print(json.dumps(configs, indent=2)) + elif action == "save": + body = json.load(sys.stdin) + token = admin_token() + saved = req(f"/api/admin/warm-runtime-configs/{name}", "PUT", body, + headers={"Authorization": f"Bearer {token}"}) + print("Saved", saved["name"], "image:", saved["image"], "count:", saved.get("count")) + elif action == "delete": + token = admin_token() + resp = req(f"/api/admin/warm-runtime-configs/{name}", "DELETE", + headers={"Authorization": f"Bearer {token}"}) + print(resp["message"]) + ' + + case "$COMMAND" in + list) + kubectl exec -n "$NAMESPACE" "$POD" -- \ + env ACTION=list RUNTIME_API_URL="http://localhost:5000" python3 -c "$PYSCRIPT" + ;; + save) + if [ -z "$CONFIG_NAME" ] || [ ! -f "$CONFIG_FILE" ]; then + echo "usage: $0 save " >&2; exit 1 + fi + kubectl exec -i -n "$NAMESPACE" "$POD" -- \ + env ACTION=save CONFIG_NAME="$CONFIG_NAME" \ + RUNTIME_API_URL="http://localhost:5000" python3 -c "$PYSCRIPT" < "$CONFIG_FILE" + ;; + delete) + if [ -z "$CONFIG_NAME" ]; then + echo "usage: $0 delete " >&2; exit 1 + fi + kubectl exec -n "$NAMESPACE" "$POD" -- \ + env ACTION=delete CONFIG_NAME="$CONFIG_NAME" \ + RUNTIME_API_URL="http://localhost:5000" python3 -c "$PYSCRIPT" + ;; + esac + ``` + + ```bash + chmod +x warm-runtime-configs.sh + ./warm-runtime-configs.sh list + ``` + + + + + Listing authenticates with the regular API key (`X-API-Key`). Saving and + deleting use the admin password via a PBKDF2 challenge-response login that + returns a 24-hour JWT. The script handles both flows automatically. + + +--- + +## Step 3: Export the Installer's Default Configuration + +Do not write configurations from scratch. The environment block in a warm +runtime configuration is what its sandbox pods actually boot with; the default +configuration contains install-specific values (webhook callback URL, CA +bundles, workspace paths) that sandboxes need to function. Export the default +from the installer-managed ConfigMap and use it as your template: + +```bash +kubectl -n openhands get configmap warm-runtimes-config \ + -o jsonpath='{.data.warm-runtimes\.json}' \ + | jq '.configs[] | select(.name == "v1_current") | del(.name)' > default-config.json +``` + +If the ConfigMap has a different name in your install, find it with: + +```bash +kubectl -n openhands get configmap | grep warm-runtimes +``` + + + + The installer's `v1_current` pool keeps running while you add API-managed + entries. You do not need to save `v1_current` explicitly — the ConfigMap + entry stays live. Skip to saving your first custom configuration below. + + + Because the first API-managed configuration takes over warm pool + management, re-declare the default pool explicitly before adding custom + images. This ensures the default pool survives the takeover: + + ```bash + jq '.count = 1' default-config.json > v1_current.json + ./warm-runtime-configs.sh save v1_current v1_current.json + ``` + + + +Then derive each custom image configuration from the same template, changing +only the image and the pool size: + +```bash +jq '.image = "ghcr.io/your-org/openhands-php:8.4-v1" | .count = 1' \ + default-config.json > php-web.json +./warm-runtime-configs.sh save php-web php-web.json + +./warm-runtime-configs.sh list +``` + +--- + +## Configuration Format + +| Field | Type | Required | Description | +|---|---|---|---| +| `image` | string | Yes | Full image reference (e.g. `ghcr.io/your-org/openhands-php:8.4-v1`) | +| `working_dir` | string | Yes | Working directory inside the sandbox — copy from the default | +| `command` | array | Yes | Agent-server start command — copy from the default | +| `environment` | object | Yes | Environment variables the sandbox boots with — copy from the default | +| `count` | integer | No | Warm pods to keep ready. Falls back to the installer-wide **Warm Runtime Count** setting — on Replicated installs this defaults to **1** (adjustable in **Config → Sandbox Configuration**); the code-level fallback when nothing is configured is **3** | +| `run_as_user` | integer | No | Copy from the installer default so warm pods match application start requests | +| `run_as_group` | integer | No | Copy from the installer default so warm pods match application start requests | +| `fs_group` | integer | No | Copy from the installer default so warm pods match application start requests | + +The configuration name comes from the URL path (the `save ` argument), +not the body. A `source` field appears in list responses (`file` for +installer-managed entries, `db` for API-managed entries) but must not be +included in saved configurations. + +The application uses the image reference as the sandbox spec ID. Give every +selectable configuration a distinct image reference; configurations that share +an image reference cannot be selected independently. + + + Set `count` explicitly. Every warm pod reserves the full sandbox resource + envelope (25 Gi of ephemeral storage by default) whether or not it is in + use, so the sum of all pool sizes must fit your node capacity. Pools that + exceed capacity show up as `Pending` pods. Start with `count: 1` per image + and grow the pools that see real traffic. + + +--- + +## Step 4: Verify the Warm Pools + +The reconciler runs every minute. Watch it create the pods: + +```bash +# Warm (unclaimed) sandboxes: runtime deployments with no session_id label yet +kubectl -n openhands get deploy -l 'runtime_id,!session_id' \ + -o custom-columns='NAME:.metadata.name,READY:.status.readyReplicas,IMAGE:.spec.template.spec.containers[0].image' +``` + +You should see one `runtime-` deployment per warm pod with your +configured images. To see the reconciler's own view — per-pool counts, pull +failures, culling decisions — read the latest reconciler job log: + +```bash +JOB=$(kubectl -n openhands get jobs --sort-by=.metadata.creationTimestamp \ + -o name | grep warm-runtimes | tail -1) +kubectl -n openhands logs "$JOB" +``` + +If a pod is stuck pulling your image, `kubectl -n openhands describe pod ` +shows the pull error. See [Building a Custom Image — Private Registries](/enterprise/custom-sandbox-images/building-custom-images#private-registries) +for pull secret configuration. + +--- + +## Updating and Deleting Configurations + +Update by saving the same name again: + +```bash +jq '.image = "ghcr.io/your-org/openhands-php:8.4-v2"' php-web.json > php-web-v2.json +./warm-runtime-configs.sh save php-web php-web-v2.json +``` + +Within a minute the reconciler stops the old pods and starts pods on the new +image. Delete a configuration to remove its pool: + +```bash +./warm-runtime-configs.sh delete php-web +``` + +If the deleted name overrides an installer-managed entry, the underlying +installer entry becomes effective again. Confirm with +`./warm-runtime-configs.sh list` — its `source` changes from `db` to `file`. + +Keep superseded image tags available in your registry while conversations that +used them can still resume: a paused conversation resumes on its **original** +image. Delete old tags only after the conversations that used them are gone +(stopped sandboxes are cleaned up after 10 days by default). + +--- + +## After Upgrading OpenHands Enterprise + + + API-managed configurations are **frozen snapshots** — upgrades do not touch + them. The installer-managed `v1_current` entry updates automatically unless + a database entry with that name overrides it. Each release expects a + specific agent-server version and may add or change sandbox environment + variables. After every OHE upgrade: + + 1. Rebuild your custom images on the release's new agent-server base version. + 2. Re-export the default template (Step 3) from the refreshed ConfigMap. + 3. Re-derive and save each API-managed custom configuration from the new template. + 4. If you intentionally override `v1_current`, refresh or delete that override + so the new installer-managed entry can take effect. + + Skipping this leaves configurations pointing at the previous agent-server + version, and new conversations fail with a version mismatch error until + the configurations are updated. + + +--- + +## Returning an Entry to Installer Management + +Delete a same-named database override to restore the installer-managed entry +on the next reconciler cycle: + +```bash +./warm-runtime-configs.sh delete v1_current +./warm-runtime-configs.sh list # v1_current now reports "source": "file" +``` + +Other API-managed configurations continue running. Delete them individually +when you no longer want their pools or images in the application's selector. + +--- + +## Troubleshooting + +| Symptom | Cause and fix | +|---|---| +| `HTTP 403: Admin functionality is disabled` | The runtime-api deployment has no admin password configured. On Replicated installs, set **Runtime API Admin Password** per Step 1 and deploy. | +| `HTTP 401` on login | Wrong password, or the challenge expired (challenges are single-use and expire after 5 minutes; the script fetches a fresh one per call). To verify the current password value: `kubectl get secret admin-password -n openhands -o jsonpath='{.data.admin-password}' \| base64 -d`. On Replicated installs, the password only changes if you update it in the Admin Console and deploy. | +| `HTTP 401: ...provide a valid API key...` on list | The list endpoint authenticates with `X-API-Key`, not the admin JWT. Use the helper script. | +| Saved a config but the dropdown does not show it | The app server caches the config list for 60 seconds; the UI may cache it for up to 5 minutes. Wait, then navigate away from and back to the Settings page to prompt a fresh fetch. Confirm the config was saved with `./warm-runtime-configs.sh list`. | +| No warm pods appear | Read the latest reconciler job log (Step 4). Look for image pull errors or scheduling failures. | +| Warm pods `Pending` | Insufficient node resources. Every warm pod reserves the full sandbox resource envelope; lower the pool `count`s or add capacity. | +| Conversations cold-start despite warm pods | Pool exhausted or configuration recently changed. See [How Warm Pods Are Claimed](/enterprise/custom-sandbox-images/using-custom-images#how-warm-pods-are-claimed). | +| Sandbox fails with an agent-server version error | The custom image's base version does not match the release. Rebuild on the expected agent-server version. See [Version Compatibility](/enterprise/custom-sandbox-images/building-custom-images#version-compatibility). | +| Conversations start but never show agent output | The configuration's `environment` is missing install-specific values. Rebuild the configuration from the default template (Step 3). | + +--- + +## API Reference + +The endpoints below are served by the runtime-api service. + +**Admin authentication** (required for save and delete): + +1. `GET /api/admin/challenge` returns `{challenge, salt, iterations}`. Challenges + are single-use and expire after 5 minutes. +2. Compute `PBKDF2-HMAC-SHA256(password, salt + challenge, iterations, dklen=32)` + and hex-encode the result. +3. `POST /api/admin/login` with `{"challenge": ..., "hash": ...}` returns + `{"token": ...}`, a JWT valid for 24 hours. +4. Send `Authorization: Bearer ` on admin requests. + +**List configurations** (regular API key, not admin): + +```http +GET /api/warm-runtime-configs +X-API-Key: {api-key} +``` + +Returns `200` with the effective configuration set: + +```json +{ + "configs": [ + { + "name": "v1_current", + "image": "ghcr.io/openhands/agent-server:1.46.0-python", + "source": "file", + "count": 1 + }, + { + "name": "php-web", + "image": "ghcr.io/your-org/openhands-php:8.4-v1", + "source": "db", + "count": 1 + } + ] +} +``` + +`source: "file"` — installer-managed entry. `source: "db"` — API-managed +entry. On Replicated installs (overlay mode), the list is the full effective +set: ConfigMap entries merged with same-named API entries overriding them. On +Helm installs without overlay mode, only API-saved entries appear while the +database contains any rows. + +**Create or update a configuration** (admin): + +```http +PUT /api/admin/warm-runtime-configs/{name} +Authorization: Bearer {admin-jwt} +Content-Type: application/json + +{"image": "...", "working_dir": "...", "command": [...], "environment": {...}, "count": 1} +``` + +Returns `200` with the saved configuration. Creates or overwrites; the name +in the URL is the identity. + +**Delete a configuration** (admin): + +```http +DELETE /api/admin/warm-runtime-configs/{name} +Authorization: Bearer {admin-jwt} +``` + +Returns `200` with a confirmation message, or `404` if no database +configuration has that name. When the deleted name also exists in the +installer-managed ConfigMap, that ConfigMap entry becomes effective again. diff --git a/enterprise/custom-sandbox-images/single-image-admin-console.mdx b/enterprise/custom-sandbox-images/single-image-admin-console.mdx new file mode 100644 index 000000000..65ea9e2f5 --- /dev/null +++ b/enterprise/custom-sandbox-images/single-image-admin-console.mdx @@ -0,0 +1,44 @@ +--- +title: Single Image via Admin Console +description: Configure a single custom sandbox image for the whole installation through the Replicated Admin Console. +icon: triangle-exclamation +--- + + + **This approach is deprecated.** It configures one image for the entire + installation and does not support per-user image selection or multiple + simultaneous environments. + + New installations should use + [Multiple Images with Warm Runtime Pools](/enterprise/custom-sandbox-images/multiple-images-warm-pools) + instead. This page is retained for installations that have not yet migrated. + Both approaches coexist — you can adopt warm runtime pools without removing + this setting. + + +Once your image is built and pushed to a registry, point the Replicated Admin +Console at it. + +1. Open the **Admin Console** at `https://admin.:30000`. +2. Navigate to **Config** and find the **Sandbox Configuration** section. +3. Set the following fields: + +| Field | Value | +|---|---| +| **Use a Custom Sandbox Image** | Enabled | +| **Sandbox Image Repository** | Your image repository (e.g. `ghcr.io/your-org/openhands-custom-image`) | +| **Sandbox Image Tag** | Your image tag (e.g. `v1.2.0`) | +| **Registry Server** | If your registry requires authentication | +| **Registry Username** | If your registry requires authentication | +| **Registry Password or Credentials** | If your registry requires authentication | + +4. Click **Save config** and then **Deploy** to apply the change. + +This single image becomes both the default for new conversations and the image +kept ready in the installer-managed warm pool. + + + This setting applies to the **sandbox / agent-server image** only — the + image that runs inside each agent's isolated workspace. It does not replace + the other OpenHands service images. + diff --git a/enterprise/custom-sandbox-images/using-custom-images.mdx b/enterprise/custom-sandbox-images/using-custom-images.mdx new file mode 100644 index 000000000..ebf756224 --- /dev/null +++ b/enterprise/custom-sandbox-images/using-custom-images.mdx @@ -0,0 +1,64 @@ +--- +title: Using Custom Images +description: How users select a custom sandbox image and how to target a specific image per conversation via the API. +icon: play +--- + +Once warm runtime pool configurations are saved, the application makes them +available to users and the API within about a minute (the application caches +the configuration list for 60 seconds). + +## Per-User Selection + +Each user opens **Settings → Application** and picks an image in the +**Default Sandbox** dropdown. Entries are the image references from your +saved configurations. + +Leaving the setting on **System default** uses the configuration named +`v1_current`, or the first configuration in the list if no `v1_current` +exists. All of the user's new conversations use their selected image. + +## Per-Conversation via the API + +To target a specific image for a single conversation regardless of the user's +default: + +```bash +# 1. Start a sandbox from a specific image (the spec ID is the image reference) +curl -X POST \ + "https://app./api/v1/sandboxes?sandbox_spec_id=ghcr.io/your-org/openhands-php:8.4-v1" \ + -H "Authorization: Bearer $API_KEY" + +# 2. Create the conversation on that sandbox, using "id" from the response above +curl -X POST \ + "https://app./api/v1/app-conversations" \ + -H "Authorization: Bearer $API_KEY" \ + -H "Content-Type: application/json" \ + -d '{"sandbox_id": ""}' +``` + +## How Warm Pods Are Claimed + +A conversation claims a warm pod only when the pod **exactly matches** the +requested image, command, working directory, environment (ignoring a fixed set +of session-specific variables), and `run_as_user` / `run_as_group` / +`fs_group`. Because the application requests exactly what the selected +configuration declares, conversations started through the OpenHands UI match +automatically. + +Cold starts still happen when: + +- All warm pods for the selected image are already claimed (`count` too low + for current traffic). +- The configuration changed in the last minute, so old pods no longer match + and replacements are still starting. +- Warm pods cannot reach `Ready` (image pull failures, insufficient node + resources). + +Cold-started conversations run the same image and work normally — they just +take 20 seconds or more to begin rather than a few seconds. + +To confirm a conversation claimed a warm pod, note that its sandbox was ready +in a few seconds. To verify from the cluster: the claimed runtime deployment +acquires a `session_id` label, and the reconciler creates a fresh warm pod to +replace it within a minute. From baa8d0c79b12ccfd7af12069c39609825589312b Mon Sep 17 00:00:00 2001 From: openhands Date: Fri, 25 Sep 2026 12:59:40 +0000 Subject: [PATCH 2/5] fix: clarify sandbox pools heading and kubectl prereq by install type - Rename 'How Warm Runtime Pools Work' to 'How Sandbox Pools Work' and introduce 'warm runtime pools' as the internal term, grounding the concept in user-facing language first - Split kubectl prerequisite by install type: VM needs kubectl once to export credentials, Helm needs it for every operation Co-authored-by: openhands --- enterprise/custom-sandbox-images/index.mdx | 29 +++++++++++----------- 1 file changed, 14 insertions(+), 15 deletions(-) diff --git a/enterprise/custom-sandbox-images/index.mdx b/enterprise/custom-sandbox-images/index.mdx index cd334ff52..4f55111da 100644 --- a/enterprise/custom-sandbox-images/index.mdx +++ b/enterprise/custom-sandbox-images/index.mdx @@ -9,15 +9,15 @@ output, and test harness your agents need. Instead of spending minutes provisioning a workspace on every run, your agents start on the actual task immediately. -## How Warm Runtime Pools Work +## How Sandbox Pools Work -The runtime-api keeps a set of pre-started sandbox pods ready before any -conversation is requested. When a user starts a conversation, it claims a -waiting pod in seconds instead of cold-starting one from scratch (which takes -20 seconds or more). Each configuration names one image and a pool size; a -reconciler runs every minute to maintain that count. Multiple configurations -run side by side, each with its own pool, each independently selectable by -users. +Each custom sandbox image can be kept ready in its own **pool** of +pre-started sandboxes (called warm runtime pools internally). When a user +starts a conversation, it claims a waiting sandbox from the pool in seconds +instead of cold-starting one from scratch (which takes 20 seconds or more). +Each configuration names one image and a pool size; a reconciler runs every +minute to maintain that count. Multiple pools run side by side, each +independently selectable by users. ## Prerequisites @@ -36,14 +36,13 @@ The image must be pushed to your registry before you configure it. **OpenHands Enterprise 0.64.0 or later** for the warm runtime pool approach. -**`kubectl` access to the cluster** for the warm runtime pool approach. On -Replicated VM installs: +**`kubectl` access** for initial setup, with different requirements by install type: -```bash -sudo /var/lib/embedded-cluster/bin/openhands shell -``` - -On Helm installs, use your normal kubeconfig. +- **Replicated VM installs:** kubectl is needed once to read the initial + credentials. After that the management script calls the runtime-api HTTPS + endpoint directly and can run from any machine without cluster access. +- **Helm installs:** kubectl is required for every management operation. + Use your normal kubeconfig. ## Configuration Approaches From 9c5dc9ee87d40692751bb51f768be8254b5c4b1b Mon Sep 17 00:00:00 2001 From: openhands Date: Fri, 25 Sep 2026 13:01:56 +0000 Subject: [PATCH 3/5] fix: rename warm pools page to 'Configuring Custom Sandbox Images' Co-authored-by: openhands --- enterprise/custom-sandbox-images/index.mdx | 2 +- .../custom-sandbox-images/multiple-images-warm-pools.mdx | 4 ++-- .../custom-sandbox-images/single-image-admin-console.mdx | 2 +- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/enterprise/custom-sandbox-images/index.mdx b/enterprise/custom-sandbox-images/index.mdx index 4f55111da..ceb66346a 100644 --- a/enterprise/custom-sandbox-images/index.mdx +++ b/enterprise/custom-sandbox-images/index.mdx @@ -48,7 +48,7 @@ The image must be pushed to your registry before you configure it. diff --git a/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx index 74476129d..108df5d09 100644 --- a/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx +++ b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx @@ -1,6 +1,6 @@ --- -title: Multiple Images with Warm Runtime Pools -description: Configure multiple custom sandbox images through the Runtime API, each kept ready in its own warm pool and independently selectable by users. +title: Configuring Custom Sandbox Images +description: Configure custom sandbox images through the Runtime API, each kept ready in its own warm pool and independently selectable by users. icon: layer-group --- diff --git a/enterprise/custom-sandbox-images/single-image-admin-console.mdx b/enterprise/custom-sandbox-images/single-image-admin-console.mdx index 65ea9e2f5..0d97b00ba 100644 --- a/enterprise/custom-sandbox-images/single-image-admin-console.mdx +++ b/enterprise/custom-sandbox-images/single-image-admin-console.mdx @@ -10,7 +10,7 @@ icon: triangle-exclamation simultaneous environments. New installations should use - [Multiple Images with Warm Runtime Pools](/enterprise/custom-sandbox-images/multiple-images-warm-pools) + [Configuring Custom Sandbox Images](/enterprise/custom-sandbox-images/multiple-images-warm-pools) instead. This page is retained for installations that have not yet migrated. Both approaches coexist — you can adopt warm runtime pools without removing this setting. From d1a31ea19d43c9c08009b5667f47bcc070d3fe5a Mon Sep 17 00:00:00 2001 From: openhands Date: Fri, 25 Sep 2026 13:10:31 +0000 Subject: [PATCH 4/5] fix: simplify How It Works to UX overview; move Helm detail to Step 3 - Replace tabbed technical overview (overlay mode, ConfigMap semantics, env var names) with a plain three-step UX summary: set admin password, register images via API, users pick their environment - Remove 'Replicated' and 'Kubernetes' from tab labels throughout - Helm takeover behavior remains where it is actionable: Step 3 Co-authored-by: openhands --- .../multiple-images-warm-pools.mdx | 68 ++++++------------- 1 file changed, 21 insertions(+), 47 deletions(-) diff --git a/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx index 108df5d09..500d91413 100644 --- a/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx +++ b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx @@ -4,55 +4,29 @@ description: Configure custom sandbox images through the Runtime API, each kept icon: layer-group --- -To offer several sandbox images at once, configure **warm runtime pools** -through the Runtime API. Each configuration names one image and keeps a pool -of pre-started sandbox pods ready for it. The OpenHands application -automatically exposes every configuration as a selectable sandbox so users can -pick their environment — no redeployment required. - **Requirements:** - OpenHands Enterprise **0.64.0 or later** -- `kubectl` access to the cluster (see [Prerequisites](/enterprise/custom-sandbox-images#prerequisites)) - Custom images built and pushed as described in [Building a Custom Image](/enterprise/custom-sandbox-images/building-custom-images) ## How It Works - - - Warm runtime configurations run in **overlay mode** on Replicated installs - (`WARM_RUNTIME_CONFIG_OVERLAY=1` is set by default). API-managed - configurations layer on top of the installer's ConfigMap entries rather than - replacing them. - - - Adding a new name creates an additional pool alongside the - installer-managed `v1_current` pool. - - Saving a name that already exists in the ConfigMap overrides that entry; - deleting the override reverts to the ConfigMap value. - - The installer-managed `v1_current` pool keeps running throughout — you - never lose the default pool by adding custom configurations. - - Changes take effect within about a minute with no application restarts and - no redeployments. - - - - **The first configuration you save takes over warm pool management.** - While the Runtime API database holds any configurations, the - installer-managed ConfigMap is ignored entirely. Always re-declare the - default image as a configuration (Step 3 does this). To hand control - back to the installer, delete **all** configurations. - - To enable overlay mode instead — so API configs layer on top of the - installer pool rather than replacing it, matching the Replicated - behavior — add `WARM_RUNTIME_CONFIG_OVERLAY: "1"` to the runtime-api - environment and restart the deployment. - +Custom sandbox images are registered through the **Runtime API** — a +management interface built into OpenHands Enterprise. The process has three +steps: - Changes take effect within about a minute with no application restarts and - no redeployments. - - +1. **Set an admin password** in the install admin UI. This secures the Runtime + API so only authorized administrators can register or remove images. + +2. **Register images via the API.** Use the helper script below to give each + image a name and tell OpenHands where to pull it from. OpenHands pulls the + image from the registry you specify and keeps a pool of ready sandboxes for + it. No restarts or redeployments are needed — new images become available + within about a minute. + +3. **Users choose their environment.** Each registered image appears in the + user's **Settings → Application → Default Sandbox** dropdown. Users pick + their default and all their new conversations start in that environment. --- @@ -61,7 +35,7 @@ pick their environment — no redeployment required. The Runtime API admin endpoints require an admin password. - + The password is **auto-generated at install** (`{{repl RandomString 32}}`) and stored in the `admin-password` Kubernetes secret. The helper script in Step 2 reads it from the pod environment automatically — no action @@ -84,7 +58,7 @@ The Runtime API admin endpoints require an admin password. use the Admin Console field. - + The password was set when you created the `admin-password` secret during installation: @@ -115,7 +89,7 @@ The Runtime API admin endpoints require an admin password. ## Step 2: Save the Helper Script - + The runtime-api is exposed externally at `https://runtime-api.`. All API calls go directly to that URL — no cluster shell or `kubectl exec` needed for day-to-day @@ -256,7 +230,7 @@ The Runtime API admin endpoints require an admin password. -o jsonpath='{.data.admin-password}' | base64 -d)" ``` - + The runtime-api is not exposed outside the cluster by default on Helm installs. The script tunnels each API call into the runtime-api pod via `kubectl exec`. You need `kubectl` access for every operation. @@ -401,12 +375,12 @@ kubectl -n openhands get configmap | grep warm-runtimes ``` - + The installer's `v1_current` pool keeps running while you add API-managed entries. You do not need to save `v1_current` explicitly — the ConfigMap entry stays live. Skip to saving your first custom configuration below. - + Because the first API-managed configuration takes over warm pool management, re-declare the default pool explicitly before adding custom images. This ensures the default pool survives the takeover: From 6ea8f572fff30545d411a0530fa3fd5315ef8817 Mon Sep 17 00:00:00 2001 From: openhands Date: Fri, 25 Sep 2026 13:28:28 +0000 Subject: [PATCH 5/5] fix: remove kubectl from Steps 3 and 4; use API throughout MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Step 3: fetch default template via script list command (API) instead of kubectl ConfigMap read — works for VM installs with no cluster access - Step 4: replace kubectl pod inspection with script list + Settings UI verification; kubectl diagnostic commands moved to troubleshooting table - Helm takeover warning stays in Step 3 where it is actionable Co-authored-by: openhands --- .../multiple-images-warm-pools.mdx | 90 ++++++++----------- 1 file changed, 39 insertions(+), 51 deletions(-) diff --git a/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx index 500d91413..a5de750fc 100644 --- a/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx +++ b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx @@ -354,55 +354,50 @@ The Runtime API admin endpoints require an admin password. --- -## Step 3: Export the Installer's Default Configuration +## Step 3: Save Your First Configuration -Do not write configurations from scratch. The environment block in a warm -runtime configuration is what its sandbox pods actually boot with; the default -configuration contains install-specific values (webhook callback URL, CA -bundles, workspace paths) that sandboxes need to function. Export the default -from the installer-managed ConfigMap and use it as your template: +Do not write configurations from scratch. The default configuration contains +install-specific values (callback URLs, CA bundles, workspace paths) that +sandboxes need to function. Fetch it from the API and use it as your template: ```bash -kubectl -n openhands get configmap warm-runtimes-config \ - -o jsonpath='{.data.warm-runtimes\.json}' \ - | jq '.configs[] | select(.name == "v1_current") | del(.name)' > default-config.json -``` - -If the ConfigMap has a different name in your install, find it with: - -```bash -kubectl -n openhands get configmap | grep warm-runtimes +./warm-runtime-configs.sh list \ + | jq '.configs[] | select(.name == "v1_current") | del(.name, .source)' \ + > default-config.json ``` - The installer's `v1_current` pool keeps running while you add API-managed - entries. You do not need to save `v1_current` explicitly — the ConfigMap - entry stays live. Skip to saving your first custom configuration below. + The default `v1_current` pool keeps running while you add configurations. + Derive your custom configuration from the template, changing only the image + and pool size, then save it: + + ```bash + jq '.image = "ghcr.io/your-org/openhands-php:8.4-v1" | .count = 1' \ + default-config.json > php-web.json + ./warm-runtime-configs.sh save php-web php-web.json + ``` - Because the first API-managed configuration takes over warm pool - management, re-declare the default pool explicitly before adding custom - images. This ensures the default pool survives the takeover: + + The first configuration you save takes over warm pool management — the + installer's default pool is ignored while any API configurations exist. + Save `v1_current` explicitly first so the default pool keeps running. + ```bash + # Re-declare the default pool before adding custom images jq '.count = 1' default-config.json > v1_current.json ./warm-runtime-configs.sh save v1_current v1_current.json + + # Now add your custom image + jq '.image = "ghcr.io/your-org/openhands-php:8.4-v1" | .count = 1' \ + default-config.json > php-web.json + ./warm-runtime-configs.sh save php-web php-web.json ``` -Then derive each custom image configuration from the same template, changing -only the image and the pool size: - -```bash -jq '.image = "ghcr.io/your-org/openhands-php:8.4-v1" | .count = 1' \ - default-config.json > php-web.json -./warm-runtime-configs.sh save php-web php-web.json - -./warm-runtime-configs.sh list -``` - --- ## Configuration Format @@ -437,29 +432,22 @@ an image reference cannot be selected independently. --- -## Step 4: Verify the Warm Pools +## Step 4: Verify -The reconciler runs every minute. Watch it create the pods: +Confirm your configurations were saved: ```bash -# Warm (unclaimed) sandboxes: runtime deployments with no session_id label yet -kubectl -n openhands get deploy -l 'runtime_id,!session_id' \ - -o custom-columns='NAME:.metadata.name,READY:.status.readyReplicas,IMAGE:.spec.template.spec.containers[0].image' +./warm-runtime-configs.sh list ``` -You should see one `runtime-` deployment per warm pod with your -configured images. To see the reconciler's own view — per-pool counts, pull -failures, culling decisions — read the latest reconciler job log: - -```bash -JOB=$(kubectl -n openhands get jobs --sort-by=.metadata.creationTimestamp \ - -o name | grep warm-runtimes | tail -1) -kubectl -n openhands logs "$JOB" -``` +The response shows each saved configuration with its name, image, pool size, +and source. Within about a minute the pool is ready. Open +**Settings → Application → Default Sandbox** — your image name appears in the +dropdown. Select it and start a conversation to confirm it loads in a few +seconds rather than 20 or more. -If a pod is stuck pulling your image, `kubectl -n openhands describe pod ` -shows the pull error. See [Building a Custom Image — Private Registries](/enterprise/custom-sandbox-images/building-custom-images#private-registries) -for pull secret configuration. +If the image does not appear or conversations cold-start, see +[Troubleshooting](#troubleshooting) below. --- @@ -535,8 +523,8 @@ when you no longer want their pools or images in the application's selector. | `HTTP 401` on login | Wrong password, or the challenge expired (challenges are single-use and expire after 5 minutes; the script fetches a fresh one per call). To verify the current password value: `kubectl get secret admin-password -n openhands -o jsonpath='{.data.admin-password}' \| base64 -d`. On Replicated installs, the password only changes if you update it in the Admin Console and deploy. | | `HTTP 401: ...provide a valid API key...` on list | The list endpoint authenticates with `X-API-Key`, not the admin JWT. Use the helper script. | | Saved a config but the dropdown does not show it | The app server caches the config list for 60 seconds; the UI may cache it for up to 5 minutes. Wait, then navigate away from and back to the Settings page to prompt a fresh fetch. Confirm the config was saved with `./warm-runtime-configs.sh list`. | -| No warm pods appear | Read the latest reconciler job log (Step 4). Look for image pull errors or scheduling failures. | -| Warm pods `Pending` | Insufficient node resources. Every warm pod reserves the full sandbox resource envelope; lower the pool `count`s or add capacity. | +| No warm pods appear | Check the reconciler log: `JOB=$(kubectl -n openhands get jobs --sort-by=.metadata.creationTimestamp -o name \| grep warm-runtimes \| tail -1) && kubectl -n openhands logs "$JOB"`. Look for image pull errors or scheduling failures. | +| Warm pods `Pending` | Insufficient node resources. Check with `kubectl -n openhands get deploy -l 'runtime_id,!session_id'`. Every warm pod reserves the full sandbox resource envelope; lower the pool `count`s or add capacity. | | Conversations cold-start despite warm pods | Pool exhausted or configuration recently changed. See [How Warm Pods Are Claimed](/enterprise/custom-sandbox-images/using-custom-images#how-warm-pods-are-claimed). | | Sandbox fails with an agent-server version error | The custom image's base version does not match the release. Rebuild on the expected agent-server version. See [Version Compatibility](/enterprise/custom-sandbox-images/building-custom-images#version-compatibility). | | Conversations start but never show agent output | The configuration's `environment` is missing install-specific values. Rebuild the configuration from the default template (Step 3). |