diff --git a/docs.json b/docs.json index 3cd5eaa36..32fc6ed62 100644 --- a/docs.json +++ b/docs.json @@ -538,7 +538,17 @@ "enterprise/enterprise-vs-oss", "enterprise/sizing-guide", "enterprise/quick-start", - "enterprise/custom-sandbox-image", + { + "group": "Custom Sandbox Images", + "icon": "box", + "pages": [ + "enterprise/custom-sandbox-images/index", + "enterprise/custom-sandbox-images/building-custom-images", + "enterprise/custom-sandbox-images/multiple-images-warm-pools", + "enterprise/custom-sandbox-images/single-image-admin-console", + "enterprise/custom-sandbox-images/using-custom-images" + ] + }, "enterprise/docker-in-sandbox", "enterprise/external-postgres", "enterprise/troubleshooting" @@ -884,6 +894,14 @@ { "source": "/openhands/usage/automations/examples", "destination": "/openhands/usage/automations/overview" + }, + { + "source": "/enterprise/custom-sandbox-image", + "destination": "/enterprise/custom-sandbox-images" + }, + { + "source": "/enterprise/custom-sandbox-image#run-multiple-custom-images-with-warm-runtime-pools", + "destination": "/enterprise/custom-sandbox-images/multiple-images-warm-pools" } ] -} \ No newline at end of file +} diff --git a/enterprise/custom-sandbox-image.mdx b/enterprise/custom-sandbox-image.mdx deleted file mode 100644 index 271f45a56..000000000 --- a/enterprise/custom-sandbox-image.mdx +++ /dev/null @@ -1,482 +0,0 @@ ---- -title: Custom Sandbox Images -description: Preload repos, dependencies, and tooling into custom sandbox images, and run multiple images side by side with warm runtime pools. -icon: box ---- - -Custom sandbox images let you prebake the repository, dependencies, compiled output, and test harness -your agents need. Instead of spending minutes provisioning a workspace on every run, your agents start -on the actual task immediately. - -This page covers two levels of customization: - -1. **[A single custom image](#configure-a-single-custom-image-admin-console)** that replaces the default - sandbox image for the whole installation. Configured in the Replicated Admin Console; no cluster access needed. -2. **[Multiple custom images](#run-multiple-custom-images-with-warm-runtime-pools)** running side by side, - each with its own warm pool, selectable per user. Configured through the Runtime API; requires `kubectl` access. - -## Why Use a Custom Image - -Custom images eliminate cold-start setup work (clone, install, transpile, and bootstrap) so agents -spend their time on the actual task. They also reduce setup variance and lower sandbox memory requirements -by keeping only what the agent needs. - -With **multiple** custom images, different teams get different environments: a PHP image with Composer and -MySQL client for the web team, a JDK and Maven image for the Java services team, a data science image with -pinned Python packages for the analytics team. Each image is kept ready in its own warm pool so conversations -start in seconds regardless of which environment they use. - -## Build Your Own Custom Image - -The [OpenHands agent-server sandbox guide](https://docs.openhands.dev/sdk/guides/agent-server/docker-sandbox) -provides full documentation on building custom sandbox images. The approach is the same for the Enterprise -Replicated VM deployment. - -### Basic Pattern - -1. Start from the OpenHands agent-server base image. -2. Keep the normal OpenHands entrypoint intact: extend the image, do not replace the entrypoint. -3. Add your repo, docs, tools, and verification wrappers. -4. Pre-run the expensive setup you do not want to repeat at task time. -5. Publish the image to a registry reachable from your OpenHands cluster. - - - Do not override the entrypoint or replace the runtime contract of the base image. The installer - expects standard OpenHands agent-server behavior. Only extend, do not replace. - - -### Base Image - -```dockerfile -FROM ghcr.io/openhands/agent-server:1.46.0-python -``` - -This example matches OpenHands Enterprise 0.64.0. Pin a specific version tag to ensure reproducible -builds, and replace it with the tag expected by your installed release. Check -[ghcr.io/openhands/agent-server](https://github.com/OpenHands/OpenHands/pkgs/container/agent-server) -for available tags. - -### Version Compatibility - -Each OpenHands Enterprise release expects a specific agent-server version. The base image tag you -build from must match the release you run: the `openhands-sdk` inside the sandbox and the one inside -the OpenHands application must agree on major and minor version. - -To find the expected tag, enable **Use a Custom Sandbox Image** in the Admin Console. The -**Sandbox Image Tag** field defaults to the tag the current release expects. - -When a conversation starts on a custom image, OpenHands checks the sandbox's agent-server version. -If it does not match the release, the conversation fails with an error naming the expected and -actual versions. Rebuild your image from the expected tag and update the **Sandbox Image Tag** -field to fix it. - - - Rebuild your custom image before each upgrade. The agent-server base image changes with every - OHE release, and an image built for an older release will be rejected by the version check. - - -### Example: Build and Push - -```bash -docker buildx build \ - --platform linux/amd64 \ - -f your-project/Dockerfile \ - -t ghcr.io//openhands-custom-image: \ - --push \ - . -``` - -Use `--platform linux/amd64` because the Enterprise Replicated VM runs on `x86-64`. - -### What to Bake In - -Good candidates for prebaking: - -- Pinned repository checkouts -- Package manager caches and installed dependencies (`node_modules`, Python virtualenvs, etc.) -- Compiled or transpiled output -- Native system packages (`xvfb`, `libkrb5-dev`, `pkg-config`, etc.) -- Browser or Electron artifacts -- Stable helper scripts such as `prepare-*` and `*-verify` wrappers - -### What to Keep Out - - - Do not bake the following into your image: - - - Secrets, API keys, or personal credentials - - Machine-specific paths or environment assumptions - - Uncommitted source changes or task-specific fixes - - Rapidly changing dependencies (use a lightweight `prepare-*` helper instead) - - -If the repository or dependencies change frequently, include a `prepare-*` script in the image -so the agent can refresh only the parts that need updating without a full rebuild. - -## Configure a Single Custom Image (Admin Console) - -Once your image is built and pushed to a registry, point the Replicated Admin Console at it. - -1. Open the **Admin Console** at `https://admin.:30000`. -2. Navigate to **Config** and find the **Sandbox Configuration** section. -3. Set the following fields: - -| Field | Value | -|---|---| -| **Use a Custom Sandbox Image** | Enabled | -| **Sandbox Image Repository** | Your image repository (e.g. `ghcr.io/your-org/openhands-custom-image`) | -| **Sandbox Image Tag** | Your image tag (e.g. `v1.2.0`) | -| **Registry Server** | If your registry requires authentication | -| **Registry Username** | If your registry requires authentication | -| **Registry Password or Credentials** | If your registry requires authentication | - -4. Click **Save config** and then **Deploy** to apply the change. - -This single image becomes both the default image for new conversations and the image kept ready in the -installer-managed warm pool. - - - This setting applies to the **sandbox / agent-server image** only (the image that runs inside each - agent's isolated workspace). It does not replace the other OpenHands service images. - - -## Run Multiple Custom Images with Warm Runtime Pools - -To offer several sandbox images at once, configure **warm runtime pools** through the Runtime API. -Each configuration names one image and keeps a pool of pre-started sandbox pods ready for it. The -OpenHands application automatically exposes every configuration as a selectable sandbox, so users can -pick their environment without any redeployment. - -**Requirements:** - -- OpenHands Enterprise **0.64.0 or later**. -- `kubectl` access to the cluster. On a Replicated VM install, get a shell with - `sudo /var/lib/embedded-cluster/bin/openhands shell`; on a Helm install, use your normal kubeconfig. -- Custom images built and pushed as described above (all on the agent-server version your release expects). - -### How It Works - -- The installer-managed configuration remains the base configuration. Configurations saved through the - Runtime API are overlaid by name: a new name adds a pool, while an existing name overrides that - installer-managed entry. -- You manage database configurations with the admin REST endpoints - (`PUT` / `DELETE /api/admin/warm-runtime-configs/{name}`). Deleting an override reveals the - installer-managed entry again. -- A reconciler job runs **every minute** and creates or removes warm sandbox pods so each - configuration has `count` unclaimed pods ready. -- The OpenHands application polls the configuration list (cached for 60 seconds) and exposes each - configuration as a **sandbox spec**. Users choose their default in **Settings → Application → Default Sandbox**. -- When a conversation starts, the Runtime API hands it a matching warm pod in a few seconds. If no - warm pod is available, the sandbox cold-starts from the image instead (20+ seconds), and the - reconciler replenishes the pool. - -Changes take effect within about a minute, with no application restarts and no redeployments. - - - The installer-managed `v1_current` pool remains active when you add API-managed configurations. Do - not save a `v1_current` configuration unless you intentionally want to override the installer default. - - -### Step 1: Confirm the Admin Password - -The Runtime API's admin endpoints authenticate with an admin password. Replicated generates a durable -password, stores it in the `admin-password` secret, and injects it into the runtime-api pod. The helper -script in Step 2 uses that pod environment. If the value is empty, the script reports an error before a -save or delete. - -To set or rotate the password: - -1. Open the `Admin Console` and select `Config`. -2. In `Sandbox Configuration`, set `Runtime API Admin Password`. -3. Select `Save config`, then deploy the new configuration. - -The value persists across later deploys. Changing it automatically restarts runtime-api so the new -password takes effect. For a Helm installation, populate the chart's `admin-password` Secret before -using the admin endpoints and restart runtime-api after changing it. - -### Step 2: Save the Helper Script - -The Runtime API is not exposed outside the cluster by default. Download the maintained -[`warm-runtime-configs.sh`](https://github.com/OpenHands/runtime-api/blob/main/scripts/warm-runtime-configs.sh) -helper, which runs each API call inside the runtime-api pod with `kubectl exec`: - -```bash -curl -fsSLo warm-runtime-configs.sh \ - https://raw.githubusercontent.com/OpenHands/runtime-api/main/scripts/warm-runtime-configs.sh -chmod +x warm-runtime-configs.sh -./warm-runtime-configs.sh list -``` - - - The helper uses `DEFAULT_API_KEY` and `ADMIN_PASSWORD` from the runtime-api pod without printing or - copying either value. Listing authenticates with the regular API key. Saving and deleting use the - admin password via a challenge-response login that returns a 24-hour JWT. - - -List responses identify each configuration's `source`. When `v1_current` is not overridden, it appears -with `"source": "file"`. - -### Step 3: Start From the Installer's Default Configuration - -Do not write configurations from scratch. The environment in a warm runtime configuration is what its -sandbox pods actually boot with; the default configuration contains install-specific values (webhook -callback URL, CA bundles, workspace paths) that sandboxes need to function. Export the default from the -installer-managed ConfigMap and use it as your template: - -```bash -kubectl -n openhands get configmap warm-runtimes-config \ - -o jsonpath='{.data.warm-runtimes\.json}' \ - | jq '.configs[] | select(.name == "v1_current") | del(.name)' > default-config.json -``` - -(If the ConfigMap has a different name in your install, find it with -`kubectl -n openhands get configmap | grep warm-runtimes`.) - -The installer-managed `v1_current` entry remains live and follows Admin Console changes. Derive each -custom image configuration from the exported template, changing only the image and pool size: - -```bash -jq '.image = "ghcr.io/your-org/openhands-php:8.4-v1" | .count = 1' \ - default-config.json > php-web.json -./warm-runtime-configs.sh save php-web php-web.json - -./warm-runtime-configs.sh list -``` - -### Configuration Format - -| Field | Type | Required | Description | -|-------|------|----------|-------------| -| `image` | string | Yes | Full image reference (e.g. `ghcr.io/your-org/openhands-php:8.4-v1`) | -| `working_dir` | string | Yes | Working directory inside the sandbox (copy from the default) | -| `command` | array | Yes | Agent-server start command (copy from the default) | -| `environment` | object | Yes | Environment variables the sandbox boots with (copy from the default) | -| `count` | integer | No | Warm pods to keep ready for this image. When omitted, uses the installer-wide **Warm Runtime Count** (default: `1` on Replicated installs) | -| `run_as_user` | integer | No | Copy from the installer default (currently `10001`) so warm pods match application start requests | -| `run_as_group` | integer | No | Copy from the installer default (currently `10001`) so warm pods match application start requests | -| `fs_group` | integer | No | Copy from the installer default (currently `10001`) so warm pods match application start requests | - -The configuration name comes from the URL path (the `save ` argument), not the body. Saving -creates or replaces a database entry. If an installer-managed entry has the same name, the database -entry overrides it. List responses also include a read-only `source` field: `file` for installer-managed -entries and `db` for API-managed entries and overrides. Do not add `source` to a saved configuration. - -The application uses the image reference as the sandbox spec ID. Give every selectable configuration a -distinct image reference; configurations that share an image reference cannot be selected independently, -even if their commands or environments differ. - - - Set `count` explicitly. Every warm pod reserves the full sandbox resource envelope (including 10Gi of - ephemeral storage by default) whether or not it is in use, so the sum of all pool sizes must fit your - node capacity. Pools that exceed capacity show up as `Pending` pods. Start with `count: 1` per image - and grow the pools that see real traffic. - - -### Step 4: Verify the Warm Pools - -The reconciler runs every minute. Watch it create the pods: - -```bash -# Warm (unclaimed) sandboxes: runtime deployments with no session_id label yet -kubectl -n openhands get deploy -l 'runtime_id,!session_id' \ - -o custom-columns='NAME:.metadata.name,READY:.status.readyReplicas,IMAGE:.spec.template.spec.containers[0].image' -``` - -You should see one `runtime-` deployment per warm pod, with your configured images. To see -the reconciler's own view (per-pool counts, pull failures, culling decisions), read the latest -reconciler job log: - -```bash -JOB=$(kubectl -n openhands get jobs --sort-by=.metadata.creationTimestamp -o name \ - | grep warm-runtimes | tail -1) -kubectl -n openhands logs "$JOB" -``` - -If a pod is stuck pulling your image, `kubectl -n openhands describe pod ` shows the pull -error. For private registries, either fill in the **Registry Server / Username / Password** fields in -the Admin Console's **Sandbox Configuration** section (they render an image pull secret that runtime -pods use), or add your own secret name to the runtime-api `RUNTIME_IMAGE_PULL_SECRETS` setting. - -### Step 5: Pick an Image and Start a Conversation - -Within a minute of saving configurations (the application caches the list for 60 seconds): - -- **Per user**: each user opens **Settings → Application** and picks an image in the **Default Sandbox** - dropdown (entries are the image references). Before a user picks an image, the application uses the - configuration named `v1_current`, or the first configuration if no `v1_current` exists. All of the - user's new conversations use their selected image. -- **Per conversation (API)**: start a sandbox for a specific image, then attach a conversation to it: - - ```bash - # 1. Start a sandbox from a specific spec (the spec id is the image reference) - curl -X POST "https://app./api/v1/sandboxes?sandbox_spec_id=ghcr.io/your-org/openhands-php:8.4-v1" \ - -H "Authorization: Bearer $API_KEY" - # 2. Create the conversation on that sandbox, using "id" from the response - curl -X POST "https://app./api/v1/app-conversations" \ - -H "Authorization: Bearer $API_KEY" \ - -H "Content-Type: application/json" \ - -d '{"sandbox_id": ""}' - ``` - -To confirm a conversation claimed a warm pod rather than cold-starting, note that its sandbox was -ready in a few seconds, or check the cluster: the claimed runtime deployment now carries a -`session_id` label, and the reconciler creates a fresh warm pod to replace it within a minute. - -### How Warm Pods Are Claimed - -A conversation claims a warm pod only when the pod **exactly matches** the requested image, command, -working directory, environment (ignoring a fixed set of session-specific variables), and -`run_as_user` / `run_as_group` / `fs_group`. Because the application requests exactly what the -selected configuration declares, conversations started through OpenHands match automatically. - -Cold starts still happen when: - -- All warm pods for the image are already claimed (`count` too low for the traffic). -- The configuration changed in the last minute, so the old pods no longer match and replacements are - still starting. -- Warm pods cannot become ready (image pull failures, insufficient node resources). - -Cold-started conversations run the same image and work normally; they just take 20+ seconds to begin. - -### Updating and Deleting Configurations - -Update by saving the same name again. To roll out a new image version: - -```bash -jq '.image = "ghcr.io/your-org/openhands-php:8.4-v2"' php-web.json > php-web-v2.json -./warm-runtime-configs.sh save php-web php-web-v2.json -``` - -Within a minute the reconciler stops the old pods and starts pods on the new image. Delete a -database-only configuration to remove its pool: - -```bash -./warm-runtime-configs.sh delete php-web -``` - -If the deleted name overrides an installer-managed entry, the underlying entry becomes effective again -instead of disappearing. Confirm the result with `./warm-runtime-configs.sh list`; its `source` changes -from `db` to `file`. - -Keep superseded image tags available in your registry while conversations that used them can still -resume: a paused conversation resumes on its **original** image. Delete old tags only after the -conversations that used them are gone (by default, stopped sandboxes are cleaned up after 10 days). - -### After Upgrading OpenHands Enterprise - - - API-managed configurations are **frozen snapshots**; upgrades do not touch them. The installer-managed - `v1_current` entry updates automatically unless a database entry with that name overrides it. Each - release expects a specific agent-server version and may add or change sandbox environment variables. - After every OpenHands Enterprise upgrade: - - 1. Rebuild your custom images on the release's new agent-server base version. - 2. Re-export the default template (Step 3) from the refreshed ConfigMap. - 3. Re-derive and save each API-managed custom configuration from the new template. - 4. If you intentionally override `v1_current`, refresh or delete that override so the new - installer-managed entry can take effect. - - Skipping this leaves configurations pointing at the previous agent-server version, and new - conversations fail with a version mismatch error until the configurations are updated. - - -### Return an Entry to Installer Management - -Delete a same-named database override to restore the installer-managed entry on the next reconciler -cycle. For example, if `v1_current` was intentionally overridden: - -```bash -./warm-runtime-configs.sh delete v1_current -./warm-runtime-configs.sh list # v1_current now reports "source": "file" -``` - -Other API-managed configurations continue running. Delete them individually when you no longer want -their pools or images in the application's selector. - -### Troubleshooting - -| Symptom | Cause and fix | -|---|---| -| `HTTP 403: Admin functionality is disabled` | The runtime-api deployment has no admin password, or the configured value is empty. On Replicated installs, set `Runtime API Admin Password` per Step 1 and deploy. | -| `HTTP 401` on login | Wrong password, or the challenge expired. Challenges are single-use and expire after 5 minutes; the script fetches a fresh one per call. | -| `HTTP 401: ...provide a valid API key...` on list | The list endpoint authenticates with `X-API-Key`, not the admin JWT. Use the helper script. | -| Installer-managed default is missing from the list | The release does not include overlay support, or the overlay is not enabled. Upgrade OpenHands Enterprise and confirm that the list reports `source` before saving configurations. | -| Saved a config but the dropdown does not show it | The application caches the list for 60 seconds; wait a minute and reload. Also confirm with `./warm-runtime-configs.sh list`. | -| No warm pods appear | Read the latest reconciler job log (Step 4). Look for image pull errors or scheduling failures. | -| Warm pods `Pending` | Insufficient node resources. Every warm pod reserves the full sandbox resource envelope; lower the pool `count`s or add capacity. | -| Conversations cold-start despite warm pods | Pool exhausted or configuration recently changed; see [How Warm Pods Are Claimed](#how-warm-pods-are-claimed). | -| Sandbox fails at start with an agent-server version error | The custom image's base version does not match the release. Rebuild on the expected agent-server version (see [Base Image](#base-image)). | -| Conversations on a custom image start but never show agent output | The configuration's `environment` is missing install-specific values (webhook callback URL, CA bundles). Rebuild the configuration from the default template (Step 3). | - -### API Reference - -The endpoints below are served by the runtime-api service (in-cluster: `http://:5000`). - -**Admin authentication** (for save and delete): - -1. `GET /api/admin/challenge` returns `{challenge, salt, iterations}`. Challenges are single-use and - expire after 5 minutes. -2. Compute `PBKDF2-HMAC-SHA256(password, salt + challenge, iterations, dklen=32)` and hex-encode it. -3. `POST /api/admin/login` with `{"challenge": ..., "hash": ...}` returns `{"token": ...}`, a JWT - valid for 24 hours. -4. Send `Authorization: Bearer ` on admin requests. - -**List configurations** (regular API key, not admin): - -```http -GET /api/warm-runtime-configs -X-API-Key: {api-key} -``` - -Returns `200` with the effective configuration set: - -```json -{ - "configs": [ - {"name": "v1_current", "image": "...", "source": "file"}, - {"name": "php-web", "image": "...", "source": "db"} - ] -} -``` - -Installer-managed entries have `source: "file"`. API-managed entries have `source: "db"`; a database -entry with the same name replaces the file entry in this effective list. - -**Create or update a configuration** (admin): - -```http -PUT /api/admin/warm-runtime-configs/{name} -Authorization: Bearer {admin-jwt} -Content-Type: application/json - -{"image": "...", "working_dir": "...", "command": [...], "environment": {...}, "count": 1} -``` - -Returns `200` with the saved configuration. Creates or overwrites; the name in the URL is the identity. - -**Delete a configuration** (admin): - -```http -DELETE /api/admin/warm-runtime-configs/{name} -Authorization: Bearer {admin-jwt} -``` - -Returns `200` with a confirmation message, or `404` if no database configuration has that name. When -the deleted name also exists in the installer-managed file, that file entry becomes effective again. - -## Reference - - - - Full SDK documentation on building custom sandbox images - - - Dockerfile, benchmark scripts, and analysis tooling for the VS Code custom image example - - - How conversations, sandboxes, and their lifecycle fit together - - - Capacity planning, including headroom for warm pools - - diff --git a/enterprise/custom-sandbox-images/building-custom-images.mdx b/enterprise/custom-sandbox-images/building-custom-images.mdx new file mode 100644 index 000000000..98d407506 --- /dev/null +++ b/enterprise/custom-sandbox-images/building-custom-images.mdx @@ -0,0 +1,192 @@ +--- +title: Building a Custom Sandbox Image +description: How to build, version, and push a custom sandbox image for use with OpenHands Enterprise. +icon: docker +--- + +All custom sandbox image approaches — single-image Admin Console and multi-image +warm runtime pools — start here. Build your image once and then point whichever +configuration approach you use at it. + +## Basic Pattern + +1. Start from the OpenHands agent-server base image. +2. Keep the normal OpenHands entrypoint intact: extend the image, do not replace it. +3. Add your repo, docs, tools, and verification wrappers. +4. Pre-run the expensive setup you do not want to repeat at task time. +5. Push the image to a registry reachable from your OpenHands cluster. + + + Do not override the entrypoint or replace the runtime contract of the base + image. OpenHands expects standard agent-server behavior. Only extend, do not + replace. + + +## Base Image + +```dockerfile +FROM ghcr.io/openhands/agent-server:1.46.0-python +``` + +Pin a specific version tag to ensure reproducible builds, and replace it with +the tag expected by your installed release. See [Version Compatibility](#version-compatibility) +below to find the right tag. + +### What is already inside the base image? + +| | | +|---|---| +| **Runtime user** | `openhands` (UID 10001), home `/home/openhands` | +| **Working directory** | `/workspace/project` | +| **Languages** | Python 3.13, Node.js 24 | +| **Preinstalled tooling** | `git`, `curl`, VS Code server, headless browser, Docker CLI | +| **ACP providers** | `claude-code`, `codex`, `gemini-cli` | +| **Entrypoint** | `tini -- /usr/local/bin/openhands-agent-server` (do not override) | + +The full recipe - every preinstalled package, capability flag, and build stage - lives in +[`openhands-agent-server/openhands/agent_server/docker/Dockerfile`](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-agent-server/openhands/agent_server/docker/Dockerfile) +in the `OpenHands/software-agent-sdk` repository. Read it before duplicating +something that is already present. + +## Version Compatibility + +Agent-server versions are largely forward and backward compatible: newer +agent-servers work with older OpenHands releases and vice versa for the +features they have in common. There is no strict version match at conversation +start. + +The one exception is a per-feature minimum-version check. A handful of newer +APIs (Hooks, MCP test, MCP OAuth, and any subsequent feature that ships with a +declared floor) fail with an `AGENT_SERVER_VERSION_TOO_OLD` error, naming the +feature and its required version, if the sandbox's agent-server predates that +feature. Building your image from too old a base image only affects those +specific features - everything else keeps working. + +Each OpenHands Enterprise release still ships with a **recommended** default +tag. To find it, enable **Use a Custom Sandbox Image** in the Admin Console; +the **Sandbox Image Tag** field defaults to that recommended tag. See +[ghcr.io/openhands/agent-server](https://github.com/OpenHands/OpenHands/pkgs/container/agent-server) +for the full tag list. + + + Best practice is to stay reasonably current with the recommended tag so + you keep access to newer features without having to think about which ones + have a version floor. A refresh at each OHE upgrade is a good cadence, but + it is not required for existing functionality to keep working. + + +## Build and Push + +```bash +docker buildx build \ + --platform linux/amd64 \ + -f your-project/Dockerfile \ + -t ghcr.io//openhands-custom-image: \ + --push \ + . +``` + +Use `--platform linux/amd64` because the Enterprise Replicated VM runs on x86-64. + +## What to Bake In + +Good candidates for prebaking: + +- Pinned repository checkouts +- Package manager caches and installed dependencies (`node_modules`, Python virtualenvs, etc.) +- Compiled or transpiled output +- Native system packages (`xvfb`, `libkrb5-dev`, `pkg-config`, etc.) +- Browser or Electron artifacts +- Stable helper scripts such as `prepare-*` and `*-verify` wrappers + +## What to Keep Out + + + Do not bake the following into your image: + + - Secrets, API keys, or personal credentials + - Machine-specific paths or environment assumptions + - Uncommitted source changes or task-specific fixes + - Rapidly changing dependencies (use a lightweight `prepare-*` script instead) + + +If the repository or dependencies change frequently, include a `prepare-*` +script in the image so the agent can refresh only the parts that need updating +without a full rebuild. + +## Complete Example + +A realistic customization that bakes a pinned repository checkout, installs +its dependencies, and ships a `prepare-repo` refresh script. The `ENTRYPOINT` +from the base image is inherited unchanged. + +```dockerfile +# syntax=docker/dockerfile:1.7 +FROM ghcr.io/openhands/agent-server:1.46.0-python + +ARG REPO_URL=https://github.com/your-org/your-service.git +ARG REPO_REF=v2.4.1 + +# System packages require root; drop back to the openhands user before the +# entrypoint runs so the sandbox does not execute as root at task time. +USER root +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + libkrb5-dev \ + libpq-dev \ + pkg-config \ + && rm -rf /var/lib/apt/lists/* + +# Refresh script for the fast-changing bits. The agent runs this at task start +# instead of paying for a full image rebuild every time the branch moves. +COPY --chown=openhands:openhands prepare-repo /usr/local/bin/prepare-repo +RUN chmod +x /usr/local/bin/prepare-repo + +USER openhands +WORKDIR /workspace/project + +# Clone once at a pinned ref so the image is reproducible. `prepare-repo` +# fast-forwards this checkout at task time if the caller passes a newer ref. +RUN git clone --depth 50 "${REPO_URL}" . \ + && git checkout "${REPO_REF}" \ + && git config --global --add safe.directory /workspace/project + +# Prebake dependencies so the first task does not pay install time. Pin the +# lockfile so the image and the runtime resolve the same versions. +RUN --mount=type=cache,target=/home/openhands/.cache/uv,uid=10001,gid=10001 \ + uv sync --frozen + +# Do NOT set ENTRYPOINT or CMD. The base image's +# `tini -- /usr/local/bin/openhands-agent-server` is required for the sandbox +# to register with the runtime-api. +``` + +An accompanying `prepare-repo` script (fetches and fast-forwards without +losing the prebaked dependency cache): + +```bash +#!/usr/bin/env bash +set -euo pipefail +cd /workspace/project +git fetch --depth 50 origin "${1:-$(git rev-parse --abbrev-ref HEAD)}" +git reset --hard FETCH_HEAD +uv sync --frozen +``` + +Build and push it with the command from [Build and Push](#build-and-push) +above, then point your Admin Console or warm runtime configuration at the +resulting tag. + +## Private Registries + +If your image lives in a private registry, provide pull credentials so the +cluster can fetch it at pod start time. + +**Replicated VM installs:** set **Registry Server**, **Registry Username**, and +**Registry Password or Credentials** in **Config → Sandbox Configuration** in +the Admin Console and deploy. The installer renders an image pull secret that +runtime pods automatically use. + +**Helm installs:** create a pull secret in the `openhands` namespace and add +its name to the runtime-api `RUNTIME_IMAGE_PULL_SECRETS` environment variable +(comma-separated list of secret names). diff --git a/enterprise/custom-sandbox-images/index.mdx b/enterprise/custom-sandbox-images/index.mdx new file mode 100644 index 000000000..ceb66346a --- /dev/null +++ b/enterprise/custom-sandbox-images/index.mdx @@ -0,0 +1,83 @@ +--- +title: Custom Sandbox Images +description: Prebake your repository, dependencies, and tooling into custom sandbox images so agents start on the actual task instead of spending time on setup. +icon: box +--- + +Custom sandbox images let you prebake the repository, dependencies, compiled +output, and test harness your agents need. Instead of spending minutes +provisioning a workspace on every run, your agents start on the actual task +immediately. + +## How Sandbox Pools Work + +Each custom sandbox image can be kept ready in its own **pool** of +pre-started sandboxes (called warm runtime pools internally). When a user +starts a conversation, it claims a waiting sandbox from the pool in seconds +instead of cold-starting one from scratch (which takes 20 seconds or more). +Each configuration names one image and a pool size; a reconciler runs every +minute to maintain that count. Multiple pools run side by side, each +independently selectable by users. + +## Prerequisites + +Before configuring any custom image, the following must be in place: + +**An image registry reachable from your OpenHands cluster.** The cluster must +be able to pull your custom image at pod start time. Public registries (GitHub +Container Registry, Docker Hub) work without extra configuration. Private +registries require credentials — either set via **Config → Sandbox +Configuration → Registry Server / Username / Password** in the Admin Console, +or via the `RUNTIME_IMAGE_PULL_SECRETS` setting on Helm installs. + +**A custom image built from the correct agent-server base.** See +[Building a Custom Image](/enterprise/custom-sandbox-images/building-custom-images). +The image must be pushed to your registry before you configure it. + +**OpenHands Enterprise 0.64.0 or later** for the warm runtime pool approach. + +**`kubectl` access** for initial setup, with different requirements by install type: + +- **Replicated VM installs:** kubectl is needed once to read the initial + credentials. After that the management script calls the runtime-api HTTPS + endpoint directly and can run from any machine without cluster access. +- **Helm installs:** kubectl is required for every management operation. + Use your normal kubeconfig. + +## Configuration Approaches + + + + One warm pool per image, selectable per user. Changes take effect within a + minute with no restarts. **Recommended.** + + + **Deprecated.** Configures one image for the whole installation via the + Replicated Admin Console. Superseded by the warm runtime pool approach. + + + +## Reference + + + + Dockerfile pattern, version pinning, and what to bake in + + + How users select an image and how to target one via the API + + + How conversations, sandboxes, and their lifecycle fit together + + + Capacity planning, including headroom for warm pools + + diff --git a/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx new file mode 100644 index 000000000..9af5dd64e --- /dev/null +++ b/enterprise/custom-sandbox-images/multiple-images-warm-pools.mdx @@ -0,0 +1,505 @@ +--- +title: Configuring Custom Sandbox Images +description: Configure custom sandbox images through the Runtime API, each kept ready in its own warm pool and independently selectable by users. +icon: layer-group +--- + +**Requirements:** + +- OpenHands Enterprise **0.64.0 or later** +- Custom images built and pushed as described in [Building a Custom Image](/enterprise/custom-sandbox-images/building-custom-images) + +## How It Works + +Custom sandbox images are registered through the **Runtime API** — a +management interface built into OpenHands Enterprise. The process has three +steps: + +1. **Set an admin password** in the install admin UI. This secures the Runtime + API so only authorized administrators can register or remove images. + +2. **Register images via the API.** Use the helper script below to give each + image a name and tell OpenHands where to pull it from. OpenHands pulls the + image from the registry you specify and keeps a pool of ready sandboxes for + it. No restarts or redeployments are needed — new images become available + within about a minute. + +3. **Users choose their environment.** Each registered image appears in the + user's **Settings → Application → Default Sandbox** dropdown. Users pick + their default and all their new conversations start in that environment. + +--- + +## Step 1: Confirm the Admin Password + +The Runtime API admin endpoints require an admin password. + + + + The password is **auto-generated at install** (`{{repl RandomString 32}}`) + and stored in the `admin-password` Kubernetes secret. The helper script in + Step 2 reads it from the pod environment automatically — no action + required for a standard installation. + + To set a memorable password or rotate the generated one: + + 1. Open the **Admin Console** at `https://admin.:30000`. + 2. Navigate to **Config → Sandbox Configuration → Runtime API Admin Password**. + 3. Enter your new password and click **Save config**, then **Deploy**. + + The Admin Console updates the secret and rolls out the runtime-api + automatically. The password persists across all future Admin Console + deploys. + + + Do not use `kubectl patch` to set the password. The Admin Console manages + the `admin-password` secret and overwrites it on every deploy, so a + patched value is lost the next time you save any config change. Always + use the Admin Console field. + + + + The password was set when you created the `admin-password` secret during + installation: + + ```bash + kubectl -n openhands create secret generic admin-password \ + --from-literal=admin-password= + ``` + + The helper script reads it automatically — no action required. + + To rotate the password: + + ```bash + # Store the new value somewhere secure before running this + kubectl -n openhands delete secret admin-password + kubectl -n openhands create secret generic admin-password \ + --from-literal=admin-password=$(openssl rand -base64 24) + kubectl -n openhands rollout restart deployment \ + -l app.kubernetes.io/name=runtime-api + kubectl -n openhands rollout status deployment \ + -l app.kubernetes.io/name=runtime-api + ``` + + + +--- + +## Enable Overlay Mode (Helm Installs Only) + +By default on Helm installs, saving any configuration via the API takes over +warm pool management and the installer-managed `v1_current` default is +ignored. Enable overlay mode so API-saved configurations sit alongside +`v1_current` rather than replacing it. VM installs have overlay mode enabled +by default and can skip this section. + +Add to your `site-values.yaml`: + +```yaml +runtime-api: + env: + WARM_RUNTIME_CONFIG_OVERLAY: "1" +``` + +Apply and roll out: + +```bash +helm upgrade openhands openhands/openhands -f site-values.yaml -n openhands +kubectl -n openhands rollout restart deployment \ + -l app.kubernetes.io/name=runtime-api +kubectl -n openhands rollout status deployment \ + -l app.kubernetes.io/name=runtime-api +``` + +You can confirm overlay mode is active after Step 2 by running +`python3 scripts/warm_runtime_configs.py list` — `v1_current` should appear +with `"source": "file"`. + +--- + +## Step 2: Install the Management CLI + +We publish a small Python CLI — the **`runtime-api-configs`** plugin — that +wraps the runtime-api admin endpoints. It uses only the Python standard +library and runs from any host that can reach the runtime-api (or from +inside the pod on Helm installs). Install it once with OpenHands +[extensions](https://github.com/OpenHands/extensions): + +```bash +git clone --depth 1 https://github.com/OpenHands/extensions +cd extensions/plugins/runtime-api-configs +``` + + + Prefer to run everything as raw HTTP calls? The [API Reference](#api-reference) + section at the bottom of this page documents the endpoints so you can drive + them directly from `curl` or any HTTP client. All following steps show the + CLI form because it is shorter and handles the PBKDF2 handshake for you. + + + + + The runtime-api is exposed externally at + `https://runtime-api.`. Export the two env vars the + CLI needs — the URL, and the admin password you confirmed in Step 1: + + ```bash + export RUNTIME_API_URL=https://runtime-api. + export ADMIN_PASSWORD= + python3 scripts/warm_runtime_configs.py list + ``` + + That is the full setup. No `kubectl`, no SSH, no cluster access. The + CLI uses the admin password directly for `save` and `delete` (via the + PBKDF2 handshake), and for `list` and `template` it logs in as admin + and fetches the read-only API key over HTTPS from + `/api/admin/api-keys`. + + Prefer to pull the credentials straight from Kubernetes secrets in one + shot? That is the `bootstrap` subcommand — kept as an optional + convenience for cluster operators and CI, and documented in + [Advanced: bootstrap from Kubernetes](#advanced-bootstrap-from-kubernetes) + at the bottom of this page. + + + The runtime-api is not exposed outside the cluster on Helm installs. + Port-forward it, then export the admin password from Step 1: + + ```bash + kubectl -n openhands port-forward svc/runtime-api 5000:5000 & + export RUNTIME_API_URL=http://localhost:5000 + export ADMIN_PASSWORD=$(kubectl -n openhands get secret admin-password \ + -o jsonpath='{.data.admin-password}' | base64 -d) + python3 scripts/warm_runtime_configs.py list + ``` + + As on VM installs, `list` and `template` fetch the read-only API key + over HTTPS via the admin login — `ADMIN_PASSWORD` is the only credential + you need to export. + + + + + For `save` and `delete`, the CLI runs a PBKDF2 challenge-response login + with `ADMIN_PASSWORD` to obtain a 24-hour JWT and calls the admin routes + as `Authorization: Bearer `. For `list` and `template`, it uses that + same admin login to fetch the `default` read-only API key from + `/api/admin/api-keys` and sends it as `X-API-Key` on + `/api/warm-runtime-configs`. Export `API_KEY` explicitly if you would + rather skip the extra login round trip on reads. See the plugin's + [`SKILL.md`](https://github.com/OpenHands/extensions/blob/main/plugins/runtime-api-configs/SKILL.md) + for the full subcommand reference. + + +--- + +## Step 3: Save Your First Configuration + +Do not write configurations from scratch. The default configuration contains +install-specific values (callback URLs, CA bundles, workspace paths) that +sandboxes need to function. The CLI's `template` subcommand fetches an +existing configuration, strips the identity fields, and lets you override the +image and pool size in one step. + + + + The default `v1_current` pool keeps running while you add configurations. + Derive your custom configuration from the template and save it: + + ```bash + python3 scripts/warm_runtime_configs.py template v1_current \ + --image ghcr.io/your-org/openhands-php:8.4-v1 --count 1 \ + --output php-web.json + python3 scripts/warm_runtime_configs.py save php-web --file php-web.json + ``` + + Piping directly into `save` works too: + + ```bash + python3 scripts/warm_runtime_configs.py template v1_current \ + --image ghcr.io/your-org/openhands-php:8.4-v1 --count 1 \ + | python3 scripts/warm_runtime_configs.py save php-web --file - + ``` + + + With overlay mode enabled (see [Enable Overlay Mode](#enable-overlay-mode-helm-installs-only) + above), the default `v1_current` pool keeps running while you add + configurations. Derive your custom configuration from the template and + save it: + + ```bash + python3 scripts/warm_runtime_configs.py template v1_current \ + --image ghcr.io/your-org/openhands-php:8.4-v1 --count 1 \ + --output php-web.json + python3 scripts/warm_runtime_configs.py save php-web --file php-web.json + ``` + + + +--- + +## Configuration Format + +| Field | Type | Required | Description | +|---|---|---|---| +| `image` | string | Yes | Full image reference (e.g. `ghcr.io/your-org/openhands-php:8.4-v1`) | +| `working_dir` | string | Yes | Working directory inside the sandbox — copy from the default | +| `command` | array | Yes | Agent-server start command — copy from the default | +| `environment` | object | Yes | Environment variables the sandbox boots with — copy from the default | +| `count` | integer | No | Warm pods to keep ready. Falls back to the installer-wide **Warm Runtime Count** setting — on Replicated installs this defaults to **1** (adjustable in **Config → Sandbox Configuration**); the code-level fallback when nothing is configured is **3** | +| `run_as_user` | integer | No | Copy from the installer default so warm pods match application start requests | +| `run_as_group` | integer | No | Copy from the installer default so warm pods match application start requests | +| `fs_group` | integer | No | Copy from the installer default so warm pods match application start requests | +| `fuse_s3_mount` | boolean | No | Use the fusey S3 workspace instead of a PVC. Copy from the installer default when unsure — most installs do not set it. | + +The configuration name comes from the URL path (the `save ` argument), +not the body. A `source` field appears in list responses (`file` for +installer-managed entries, `db` for API-managed entries) but must not be +included in saved configurations. + +The application uses the image reference as the sandbox spec ID. Give every +selectable configuration a distinct image reference; configurations that share +an image reference cannot be selected independently. + + + Set `count` explicitly. Every warm pod reserves the full sandbox resource + envelope (25 Gi of ephemeral storage by default) whether or not it is in + use, so the sum of all pool sizes must fit your node capacity. Pools that + exceed capacity show up as `Pending` pods. Start with `count: 1` per image + and grow the pools that see real traffic. + + +--- + +## Step 4: Verify + +Confirm your configurations were saved: + +```bash +python3 scripts/warm_runtime_configs.py list +``` + +The response shows each saved configuration with its name, image, pool size, +and source. Within about a minute the pool is ready. Open +**Settings → Application → Default Sandbox** — your image name appears in the +dropdown. Select it and start a conversation to confirm it loads in a few +seconds rather than 20 or more. + +If the image does not appear or conversations cold-start, see +[Troubleshooting](#troubleshooting) below. + +--- + +## Updating and Deleting Configurations + +Update by re-deriving from the current default and saving under the same +name: + +```bash +python3 scripts/warm_runtime_configs.py template v1_current \ + --image ghcr.io/your-org/openhands-php:8.4-v2 --count 1 \ + | python3 scripts/warm_runtime_configs.py save php-web --file - +``` + +Within a minute the reconciler stops the old pods and starts pods on the new +image. Delete a configuration to remove its pool: + +```bash +python3 scripts/warm_runtime_configs.py delete php-web +``` + +If the deleted name overrides an installer-managed entry, the underlying +installer entry becomes effective again. Confirm with +`python3 scripts/warm_runtime_configs.py list` — its `source` changes from `db` to `file`. + +Keep superseded image tags available in your registry while conversations that +used them can still resume: a paused conversation resumes on its **original** +image. Delete old tags only after the conversations that used them are gone +(stopped sandboxes are cleaned up after 10 days by default). + +--- + +## After Upgrading OpenHands Enterprise + + + API-managed configurations are **frozen snapshots** — upgrades do not touch + them. The installer-managed `v1_current` entry updates automatically unless + a database entry with that name overrides it. Each release expects a + specific agent-server version and may add or change sandbox environment + variables. After every OHE upgrade: + + 1. Rebuild your custom images on the release's new agent-server base version. + 2. Re-export the default template (Step 3) from the refreshed ConfigMap. + 3. Re-derive and save each API-managed custom configuration from the new template. + 4. If you intentionally override `v1_current`, refresh or delete that override + so the new installer-managed entry can take effect. + + Skipping this leaves configurations pinned to the previous agent-server + version. Existing features keep working - agent-server is largely forward + and backward compatible - but any newer feature with a declared minimum + version (Hooks, MCP test, MCP OAuth, and future additions) fails on those + sandboxes with an `AGENT_SERVER_VERSION_TOO_OLD` error until the + configurations are refreshed. See + [Version Compatibility](/enterprise/custom-sandbox-images/building-custom-images#version-compatibility) + for details. + + +--- + +## Returning an Entry to Installer Management + +Delete a same-named database override to restore the installer-managed entry +on the next reconciler cycle: + +```bash +python3 scripts/warm_runtime_configs.py delete v1_current +python3 scripts/warm_runtime_configs.py list # v1_current now reports "source": "file" +``` + +Other API-managed configurations continue running. Delete them individually +when you no longer want their pools or images in the application's selector. + +--- + +## Troubleshooting + +| Symptom | Cause and fix | +|---|---| +| `HTTP 403: Admin functionality is disabled` | The runtime-api deployment has no admin password configured. On Replicated installs, set **Runtime API Admin Password** per Step 1 and deploy. | +| `HTTP 401` on login | Wrong password, or the challenge expired (challenges are single-use and expire after 5 minutes; the script fetches a fresh one per call). To verify the current password value: `kubectl get secret admin-password -n openhands -o jsonpath='{.data.admin-password}' \| base64 -d`. On Replicated installs, the password only changes if you update it in the Admin Console and deploy. | +| `HTTP 401: ...provide a valid API key...` on list | The list endpoint authenticates with `X-API-Key`, not the admin JWT. Use the helper script. | +| Saved a config but the dropdown does not show it | The app server caches the config list for 60 seconds; the UI may cache it for up to 5 minutes. Wait, then navigate away from and back to the Settings page to prompt a fresh fetch. Confirm the config was saved with `python3 scripts/warm_runtime_configs.py list`. | +| No warm pods appear | Check the reconciler log: `JOB=$(kubectl -n openhands get jobs --sort-by=.metadata.creationTimestamp -o name \| grep warm-runtimes \| tail -1) && kubectl -n openhands logs "$JOB"`. Look for image pull errors or scheduling failures. | +| Warm pods `Pending` | Insufficient node resources. Check with `kubectl -n openhands get deploy -l 'runtime_id,!session_id'`. Every warm pod reserves the full sandbox resource envelope; lower the pool `count`s or add capacity. | +| Conversations cold-start despite warm pods | Pool exhausted or configuration recently changed. See [How Warm Pods Are Claimed](/enterprise/custom-sandbox-images/using-custom-images#how-warm-pods-are-claimed). | +| Sandbox fails with an agent-server version error | The custom image's base version does not match the release. Rebuild on the expected agent-server version. See [Version Compatibility](/enterprise/custom-sandbox-images/building-custom-images#version-compatibility). | +| Conversations start but never show agent output | The configuration's `environment` is missing install-specific values. Rebuild the configuration from the default template (Step 3). | + +--- + +## API Reference + +The endpoints below are served by the runtime-api service. + +**Admin authentication** (required for save, delete, and — if you skip the +`X-API-Key` header on reads — for `GET /api/admin/api-keys`): + +1. `GET /api/admin/challenge` returns `{challenge, salt, iterations}`. Challenges + are single-use and expire after 5 minutes. `salt` is returned as an ASCII + hex string. +2. Compute `PBKDF2-HMAC-SHA256(password_utf8, (salt_hex + challenge).utf8, iterations, dklen=32)` + and hex-encode the result. The `salt` value returned above goes into PBKDF2 + as **its ASCII hex string**, not decoded to raw bytes first — concatenate + `salt` and `challenge` as strings, then UTF-8 encode. +3. `POST /api/admin/login` with `{"challenge": ..., "hash": ...}` returns + `{"token": ...}`, a JWT valid for 24 hours. +4. Send `Authorization: Bearer ` on admin requests. + +**Fetch the read-only API key over HTTPS** (admin — lets you drive +`/api/warm-runtime-configs` without any cluster access): + +```http +GET /api/admin/api-keys +Authorization: Bearer {admin-jwt} +``` + +Returns `200` with `[{"id": ..., "name": "default", "key_value": "...", ...}, ...]`. +Use the `key_value` of the `name: "default"` entry as your `X-API-Key`. + +**List configurations** (regular API key, not admin): + +```http +GET /api/warm-runtime-configs +X-API-Key: {api-key} +``` + +Returns `200` with the effective configuration set: + +```json +{ + "configs": [ + { + "name": "v1_current", + "image": "ghcr.io/openhands/agent-server:1.46.0-python", + "source": "file", + "count": 1 + }, + { + "name": "php-web", + "image": "ghcr.io/your-org/openhands-php:8.4-v1", + "source": "db", + "count": 1 + } + ] +} +``` + +`source: "file"` — installer-managed entry. `source: "db"` — API-managed +entry. The list is the full effective set: ConfigMap entries merged with +same-named API entries overriding them. + +**Create or update a configuration** (admin): + +```http +PUT /api/admin/warm-runtime-configs/{name} +Authorization: Bearer {admin-jwt} +Content-Type: application/json + +{"image": "...", "working_dir": "...", "command": [...], "environment": {...}, "count": 1} +``` + +Returns `200` with the saved configuration. Creates or overwrites; the name +in the URL is the identity. + +**Delete a configuration** (admin): + +```http +DELETE /api/admin/warm-runtime-configs/{name} +Authorization: Bearer {admin-jwt} +``` + +Returns `200` with a confirmation message, or `404` if no database +configuration has that name. When the deleted name also exists in the +installer-managed ConfigMap, that ConfigMap entry becomes effective again. + +--- + +## Advanced: bootstrap from Kubernetes + +The CLI's `bootstrap` subcommand pulls `RUNTIME_API_URL`, `API_KEY`, and +`ADMIN_PASSWORD` from Kubernetes secrets in one step. It is optional — the +HTTPS-only flow in Step 2 is preferred for interactive administration. Use +`bootstrap` when you have `kubectl` access anyway and want a single +one-liner for a CI job or an operator runbook. + + + + The Replicated embedded k0s cluster stores its kubeconfig at + `/var/lib/k0s/pki/admin.conf`, which is root-owned. Run under `sudo -E` + so the CLI's `kubectl` calls can read it: + + ```bash + eval "$(sudo -E python3 scripts/warm_runtime_configs.py bootstrap \ + --namespace openhands)" + python3 scripts/warm_runtime_configs.py list + ``` + + `bootstrap` prints three `export` lines. After the `eval`, subsequent + commands run from any host with network access to the ingress — no + further cluster access needed. + + + Run against your own `kubectl` context (no `sudo` needed on typical + Helm-managed clusters). Port-forward first, then bootstrap with + `--skip-url` so your `port-forward` target is not overwritten: + + ```bash + kubectl -n openhands port-forward svc/runtime-api 5000:5000 & + export RUNTIME_API_URL=http://localhost:5000 + eval "$(python3 scripts/warm_runtime_configs.py bootstrap \ + --namespace openhands --skip-url)" + python3 scripts/warm_runtime_configs.py list + ``` + + diff --git a/enterprise/custom-sandbox-images/single-image-admin-console.mdx b/enterprise/custom-sandbox-images/single-image-admin-console.mdx new file mode 100644 index 000000000..0d97b00ba --- /dev/null +++ b/enterprise/custom-sandbox-images/single-image-admin-console.mdx @@ -0,0 +1,44 @@ +--- +title: Single Image via Admin Console +description: Configure a single custom sandbox image for the whole installation through the Replicated Admin Console. +icon: triangle-exclamation +--- + + + **This approach is deprecated.** It configures one image for the entire + installation and does not support per-user image selection or multiple + simultaneous environments. + + New installations should use + [Configuring Custom Sandbox Images](/enterprise/custom-sandbox-images/multiple-images-warm-pools) + instead. This page is retained for installations that have not yet migrated. + Both approaches coexist — you can adopt warm runtime pools without removing + this setting. + + +Once your image is built and pushed to a registry, point the Replicated Admin +Console at it. + +1. Open the **Admin Console** at `https://admin.:30000`. +2. Navigate to **Config** and find the **Sandbox Configuration** section. +3. Set the following fields: + +| Field | Value | +|---|---| +| **Use a Custom Sandbox Image** | Enabled | +| **Sandbox Image Repository** | Your image repository (e.g. `ghcr.io/your-org/openhands-custom-image`) | +| **Sandbox Image Tag** | Your image tag (e.g. `v1.2.0`) | +| **Registry Server** | If your registry requires authentication | +| **Registry Username** | If your registry requires authentication | +| **Registry Password or Credentials** | If your registry requires authentication | + +4. Click **Save config** and then **Deploy** to apply the change. + +This single image becomes both the default for new conversations and the image +kept ready in the installer-managed warm pool. + + + This setting applies to the **sandbox / agent-server image** only — the + image that runs inside each agent's isolated workspace. It does not replace + the other OpenHands service images. + diff --git a/enterprise/custom-sandbox-images/using-custom-images.mdx b/enterprise/custom-sandbox-images/using-custom-images.mdx new file mode 100644 index 000000000..ebf756224 --- /dev/null +++ b/enterprise/custom-sandbox-images/using-custom-images.mdx @@ -0,0 +1,64 @@ +--- +title: Using Custom Images +description: How users select a custom sandbox image and how to target a specific image per conversation via the API. +icon: play +--- + +Once warm runtime pool configurations are saved, the application makes them +available to users and the API within about a minute (the application caches +the configuration list for 60 seconds). + +## Per-User Selection + +Each user opens **Settings → Application** and picks an image in the +**Default Sandbox** dropdown. Entries are the image references from your +saved configurations. + +Leaving the setting on **System default** uses the configuration named +`v1_current`, or the first configuration in the list if no `v1_current` +exists. All of the user's new conversations use their selected image. + +## Per-Conversation via the API + +To target a specific image for a single conversation regardless of the user's +default: + +```bash +# 1. Start a sandbox from a specific image (the spec ID is the image reference) +curl -X POST \ + "https://app./api/v1/sandboxes?sandbox_spec_id=ghcr.io/your-org/openhands-php:8.4-v1" \ + -H "Authorization: Bearer $API_KEY" + +# 2. Create the conversation on that sandbox, using "id" from the response above +curl -X POST \ + "https://app./api/v1/app-conversations" \ + -H "Authorization: Bearer $API_KEY" \ + -H "Content-Type: application/json" \ + -d '{"sandbox_id": ""}' +``` + +## How Warm Pods Are Claimed + +A conversation claims a warm pod only when the pod **exactly matches** the +requested image, command, working directory, environment (ignoring a fixed set +of session-specific variables), and `run_as_user` / `run_as_group` / +`fs_group`. Because the application requests exactly what the selected +configuration declares, conversations started through the OpenHands UI match +automatically. + +Cold starts still happen when: + +- All warm pods for the selected image are already claimed (`count` too low + for current traffic). +- The configuration changed in the last minute, so old pods no longer match + and replacements are still starting. +- Warm pods cannot reach `Ready` (image pull failures, insufficient node + resources). + +Cold-started conversations run the same image and work normally — they just +take 20 seconds or more to begin rather than a few seconds. + +To confirm a conversation claimed a warm pod, note that its sandbox was ready +in a few seconds. To verify from the cluster: the claimed runtime deployment +acquires a `session_id` label, and the reconciler creates a fresh warm pod to +replace it within a minute.