Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/notes/lint-rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -667,7 +667,7 @@ shipped. An id is a permanent documentation anchor: a suppression comment carryi
an older branch, review or commit message must never start meaning something new. CE023 is
retired the same way, after the rule was deleted with the package it guarded.

Claim 068 next, and note 065 IS TAKEN without being in `ALL_RULES`: doc-surface and
Claim 069 next, and note 065 IS TAKEN without being in `ALL_RULES`: doc-surface and
whole-tree rules are `@pytest.mark.lint` classes in `tests/test_custom_lint.py` rather than
`BaseRule`s, so `runner.py`'s uniqueness assert cannot see them. Enumerating them in a
comment is how that note fell behind CE044, so grep instead:
Expand Down
8 changes: 0 additions & 8 deletions .claude/notes/permissions.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,10 +29,6 @@ read-write throwaway copy; when it was bind-mounted `:ro` the chmod returned ERO
no tmpfs mask and has not been one since: see
[isolation.md](isolation.md) § Why the framework mounts are writable copies.

What the window does NOT hide is the task DEFINITION. `task.yaml` is also staged at
`/work/input` for the in-container orchestrator and that mount is untouched, so hiding the
criteria from the agent remains a separate, unsolved problem.

Criteria address reference files with the `$REFERENCE_DIR` token (same resolver as
`$TASK_DIR`) and the `REFERENCE_DIR` env var for `run_command`; `reference_comparison` names
one file via `reference_file`.
Expand Down Expand Up @@ -64,10 +60,6 @@ criteria and judges run. The task directory is shielded by the same window: unde
it is mounted as a throwaway COPY at a fixed container path, which is what makes it
chmod-able without touching the user's checked-out `tasks/` tree.

It shields grading MATERIAL that happens to live in the task directory (a `reference/`
subdirectory, fixtures), not the task DEFINITION: `task.yaml` is separately staged at
`/work/input`, which the agent can still read.

### Why a stack and not a refcount

Windows nest with *different* modes, so what an exit has to restore is the mode of the
Expand Down
39 changes: 39 additions & 0 deletions docs/DOCKER_ISOLATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -339,6 +339,45 @@ If the task declares a `reference:` block, a throwaway **copy** of its directory

See [Reference Solutions](TASK_DEFINITION_GUIDE.md#reference-solutions).

### Two more passive-read anti-cheat blocks

These are `driver: docker` only. `driver: tempdir` shares the host uid and has no
filesystem isolation, so neither applies there (nor can — there is nothing to
mask). Both are defense-in-depth passive-read blocks, consistent with the
reference window's posture above; neither contains an adversarial agent.

- **The staged grading inputs are deleted after load.** The host stages the
post-override `TaskDefinition` (with `success_criteria`) at `/work/input/task.yaml`
**and** a `context.json` whose `source_yaml` is the raw task text — criteria
verbatim, both at the top level and inside every `config_lineage` entry — for the
in-container orchestrator to load once at startup. The agent runs in the same
container, so leaving *either* readable would hand it the grading answer key
(deleting only `task.yaml` leaves the identical criteria one file over in
`context.json`). The in-container entry point deletes **both** immediately after
they are consumed — `context.json` is parsed into memory in the command body and
`task.yaml` by `load_task`, both before the delete (gated on
`CODER_EVAL_IN_CONTAINER`). They are read exactly once — grading reads criteria
from the in-memory task, never from disk. The `/work/input` mount is therefore
read-write (a `:ro` mount rejects `rm` with EROFS). `prior.json` is kept: it is
read later on the regrade path, and a regrade runs no agent so it is not a leak.

- **Auto-mounted plugin trees are default-deny masked.** An `agent.plugins[].path`
(or a `TemplateDirSource.path` that is itself a plugin root) is auto-mounted at
its host path `:ro` so the plugin loads. Eval material colocated under that tree
as siblings of the skills dir — sibling task YAMLs, reference solutions, test
fixtures — would otherwise be readable. So the runner keeps the whole root
mounted but layers an empty `--tmpfs` over every child dir OUTSIDE the keep-set
(`.claude-plugin` + the manifest-declared skill dirs). Everything that is not
the plugin surface is masked by default, so an unknown or new eval layout can
never leak; `tests/`, `node_modules/`, and reference solutions are masked for
free. A root agent cannot `umount` a tmpfs (`CAP_SYS_ADMIN` is not in Docker's
default set), so this mask is *stronger* than the mode-000 reference window. Two
residuals the mask cannot cover — an eval def or reference COLOCATED inside a
skill dir (masking it would hide the skill), and a `task_id:` YAML **file** loose
at the plugin root (a tmpfs masks a directory, not a single file) — are caught by
lint rule CE068 (keep eval material out of skill dirs and off the plugin root;
put it under a sibling `tests/`).

Inside the container, the entrypoint invokes `coder-eval _run-task-internal` (hidden subcommand), which loads the staged YAML + context, runs the standard in-process Orchestrator (driver auto-coerced back to `tempdir`), and writes `task.json` to the output mount. Host reads it and feeds the existing aggregation pipeline.

A `result_kind` discriminator on `CriterionResult` ensures `ClassificationCriterionResult` subclasses survive the JSON round-trip — without it, host-side aggregation would silently lose `observed_label`/`expected_label`.
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -305,6 +305,7 @@ external = [
"CE064",
"CE065",
"CE066",
"CE068",
] # custom architectural lint rules (tests/lint/)

[tool.ruff.lint.pylint]
Expand Down
4 changes: 2 additions & 2 deletions src/coder_eval/agents/_skills.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@
_SKILL_FILE = "SKILL.md"


def _manifest_skill_dirs(root: Path) -> list[Path]:
def manifest_skill_dirs(root: Path) -> list[Path]:
"""Skill directories a Claude-plugin root declares, in manifest order.

Reads the ``skills`` field of ``<root>/.claude-plugin/plugin.json`` (a string
Expand Down Expand Up @@ -81,7 +81,7 @@ def _plugin_skill_dirs(
hint,
)
continue
candidates = [directory for directory in _manifest_skill_dirs(root) if directory.is_dir()]
candidates = [directory for directory in manifest_skill_dirs(root) if directory.is_dir()]
# A path that is ALREADY a bare skills directory has no `skills/` subdir,
# so use it as-is. Deliberately NOT a fallback for a root that HAS one.
# Rationale: .claude/notes/agents.md § Skills, per harness
Expand Down
13 changes: 13 additions & 0 deletions src/coder_eval/cli/run_task_internal_command.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@
import asyncio
import contextlib
import logging
import os
from pathlib import Path
from typing import Any

Expand Down Expand Up @@ -45,6 +46,17 @@
logger = logging.getLogger(__name__)


def _scrub_staged_inputs(task_yaml: Path, context_json: Path) -> None:
"""Delete the staged ``task.yaml`` and ``context.json`` once loaded, inside a container only.

Both carry the task's ``success_criteria``, and the agent runs in this same container.
``prior.json`` is left for the regrade path, which reads it later and runs no agent.
"""
if os.environ.get(IN_CONTAINER_ENV) == "1":
task_yaml.unlink(missing_ok=True)
context_json.unlink(missing_ok=True)


def heartbeat_is_alive(current: str, last_counter: str, current_mtime: float, last_mtime: float) -> bool:
"""True when the heartbeat shows a fresh signal of life.

Expand Down Expand Up @@ -176,6 +188,7 @@ def run_task_internal_command(
# `task_file` is then pointed under the task_dir mount, so the `TASK_DIR` the
# Orchestrator exposes to `run_command` criteria resolves there, not /work/input.
task, _ = load_task(task_yaml)
_scrub_staged_inputs(task_yaml, context_json)
# The path below is never re-read; it only seeds Orchestrator's TASK_DIR.
runtime_task_file = task_dir / "task.yaml" if task_dir.is_dir() else task_yaml

Expand Down
138 changes: 91 additions & 47 deletions src/coder_eval/isolation/docker_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -140,6 +140,13 @@ def _rewrite_loopback_for_container(url: str) -> str | None:
# Rationale: .claude/notes/isolation.md § The stdout line limit
STDOUT_LINE_LIMIT_BYTES = 64 * 1024 * 1024 # 64 MiB

_MASK_WARNING = "Masking non-skill path %s under plugin root %s (anti-cheat: only skills stay readable)."
_MASK_STANDDOWN_WARNING = (
"Anti-cheat mask stood down for plugin root %s: its whole tree is the declared skill surface "
'(e.g. manifest `skills: "."`), so nothing is masked. Any eval material colocated here is READABLE '
"to the agent — move it outside the plugin root."
)


async def _heartbeat_loop(heartbeat_path: Path) -> None:
"""Write a monotonic counter to ``heartbeat_path`` every interval until cancelled.
Expand Down Expand Up @@ -656,8 +663,9 @@ async def run(self) -> EvaluationResult:
await asyncio.to_thread(self._prepare_task_dir_mount, staging)
# AFTER staging, BEFORE the container starts: the DAC caps are dropped, so
# every framework-owned mount must be reachable through its `other` bits.
# Writable so the entry point can delete the staged task.yaml/context.json.
# Rationale: .claude/notes/isolation.md § grant_container_access
await asyncio.to_thread(grant_container_access, input_dir, writable=False)
await asyncio.to_thread(grant_container_access, input_dir, writable=True)
await asyncio.to_thread(grant_container_access, output_dir, writable=True)
if self.grade_workspace is not None:
# The one mount whose files the harness did NOT create, so the owner
Expand Down Expand Up @@ -1218,6 +1226,84 @@ def _reference_mount_args(self) -> list[str]:
# every turn, which a `:ro` mount rejects with EROFS.
return ["-v", f"{self._reference_mount_src}:{CONTAINER_REFERENCE_DIR}"]

def _resolve_mount_path(self, raw_path: str) -> Path:
"""Resolve an auto-mount source to an absolute host path.

A relative path resolves against the task-file dir, as CE068 resolves it, not the CWD.
"""
expanded = Path(os.path.expandvars(os.path.expanduser(raw_path)))
if not expanded.is_absolute() and self.rt.task_file is not None:
expanded = self.rt.task_file.parent / expanded
return expanded.resolve()

def _append_auto_mounts(self, argv: list[str]) -> None:
"""Bind-mount the plugin, template and system-prompt paths a task references, ``:ro`` at their host path.

A plugin root also gets every non-skill child dir masked with an empty tmpfs. The
reference is deliberately NOT here: it has its own mount (``_reference_mount_args``).
"""
mounted: set[Path] = set()
# Warned, not refused: `plugin.path` / `reference.directory` /
# `template_sources` are user-controlled strings, and legitimate uses exist.
# Rationale: .claude/notes/isolation.md § Extra mounts and reserved destinations
sensitive_sources = self._sensitive_source_paths()

# Lazy: eval_material imports agents._skills, whose package imports this module.
from coder_eval.isolation.eval_material import mask_dirs

# Masked dir -> its plugin root. Emitted only once every bind is known, so a
# nested plugin root's bind can win over its parent's mask of the same path.
mask_targets: dict[Path, Path] = {}

def _auto_mount(raw_path: str | None, *, dir_only: bool = True) -> None:
if not raw_path:
return
resolved = self._resolve_mount_path(raw_path)
# File paths get mounted as the parent dir so a single -v covers
# the file; container-side reads still resolve at the same path.
target = resolved if (dir_only or resolved.is_dir()) else resolved.parent
if target in mounted or not target.is_dir():
return
for sensitive in sensitive_sources:
if target == sensitive or sensitive in target.parents:
logger.warning(
"Auto-mounting sensitive host path %s into container; fix task YAML if unintended.",
target,
)
break
mounted.add(target)
argv.extend(["-v", f"{target}:{target}:ro"])
masks = mask_dirs(target)
if not masks and (target / ".claude-plugin" / "plugin.json").is_file():
logger.warning(_MASK_STANDDOWN_WARNING, target)
for masked_dir in masks:
mask_targets.setdefault(masked_dir, target)

plugins = (self.rt.task.agent.plugins if self.rt.task.agent else None) or []
for plugin in plugins:
_auto_mount(plugin.get("path") if isinstance(plugin, dict) else None)

from coder_eval.models import TemplateDirSource

sandbox_cfg = self.rt.task.sandbox
for source in (sandbox_cfg.template_sources or []) if sandbox_cfg else []:
if isinstance(source, TemplateDirSource):
_auto_mount(source.path)

# Defensive: normally inlined into system_prompt by load_task / experiment
# resolution, but a variant could inject an absolute path that survives.
agent_cfg = self.rt.task.agent
if agent_cfg and agent_cfg.system_prompt_file:
_auto_mount(agent_cfg.system_prompt_file, dir_only=False)

# A deeper --tmpfs wins over the enclosing :ro bind regardless of argv order;
# a mask that is also a bind would be a duplicate mount point, so the bind wins.
for masked_dir, root in sorted(mask_targets.items()):
if masked_dir in mounted:
continue
argv.extend(["--tmpfs", str(masked_dir)])
logger.warning(_MASK_WARNING, masked_dir, root)

def _build_argv(
self, input_dir: Path, output_dir: Path, *, container_name: str, image: str | None = None
) -> list[str]:
Expand Down Expand Up @@ -1300,7 +1386,9 @@ def _build_argv(
# Rationale: .claude/notes/isolation.md § Environment forwarding
argv += ["--env", "TELEMETRY_ENABLED=false"]

argv += ["-v", f"{input_dir.resolve()}:{CONTAINER_INPUT_DIR}:ro"]
# Read-WRITE: the entry point deletes the staged task.yaml and context.json
# after load, and `rm` fails with EROFS on a `:ro` bind mount.
argv += ["-v", f"{input_dir.resolve()}:{CONTAINER_INPUT_DIR}"]
# The host run_dir at the container's standard output location, so the
# in-container Orchestrator writes straight to the host filesystem.
argv += ["-v", f"{output_dir}:{CONTAINER_OUTPUT_DIR}"]
Expand Down Expand Up @@ -1330,51 +1418,7 @@ def _build_argv(
host_claude_dir = Path.home() / ".claude"
argv += ["-v", f"{self._claude_mount_src}:{host_claude_dir}"]

# Host paths the task references (plugin dirs, resolved template dirs), at
# the SAME path inside the container. The reference is deliberately NOT
# here -- it has its own mount and is masked out of the task_dir mount.
# ``mounted`` dedupes overlapping entries.
mounted: set[Path] = set()
# Warned, not refused: `plugin.path` / `reference.directory` /
# `template_sources` are user-controlled strings, and legitimate uses exist.
# Rationale: .claude/notes/isolation.md § Extra mounts and reserved destinations
sensitive_sources = self._sensitive_source_paths()

def _auto_mount(raw_path: str | None, *, dir_only: bool = True) -> None:
if not raw_path:
return
resolved = Path(os.path.expandvars(os.path.expanduser(raw_path))).resolve()
# File paths get mounted as the parent dir so a single -v covers
# the file; container-side reads still resolve at the same path.
target = resolved if (dir_only or resolved.is_dir()) else resolved.parent
if target in mounted or not target.is_dir():
return
for sensitive in sensitive_sources:
if target == sensitive or sensitive in target.parents:
logger.warning(
"Auto-mounting sensitive host path %s into container; fix task YAML if unintended.",
target,
)
break
mounted.add(target)
argv.extend(["-v", f"{target}:{target}:ro"])

plugins = (self.rt.task.agent.plugins if self.rt.task.agent else None) or []
for plugin in plugins:
_auto_mount(plugin.get("path") if isinstance(plugin, dict) else None)

from coder_eval.models import TemplateDirSource

sandbox_cfg = self.rt.task.sandbox
for source in (sandbox_cfg.template_sources or []) if sandbox_cfg else []:
if isinstance(source, TemplateDirSource):
_auto_mount(source.path)

# Defensive: normally inlined into system_prompt by load_task / experiment
# resolution, but a variant could inject an absolute path that survives.
agent_cfg = self.rt.task.agent
if agent_cfg and agent_cfg.system_prompt_file:
_auto_mount(agent_cfg.system_prompt_file, dir_only=False)
self._append_auto_mounts(argv)

# HAZARD: task.reference.directory is deliberately NOT auto-mounted at its
# host path. That would bind the REAL tree in beside the shielded copy, so
Expand Down
58 changes: 58 additions & 0 deletions src/coder_eval/isolation/eval_material.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
"""Allowlist (default-deny) mask for auto-mounted Claude-plugin trees.

``mask_dirs`` names the child dirs of a mounted plugin root to tmpfs-mask so that only the
plugin surface (``.claude-plugin`` + the manifest-declared skill dirs) stays readable.
See docs/DOCKER_ISOLATION.md for the two residuals the mask cannot cover.
"""

from __future__ import annotations

from pathlib import Path

from coder_eval.agents._skills import manifest_skill_dirs


_PLUGIN_MANIFEST_RELPATH = (".claude-plugin", "plugin.json")


def mask_dirs(root: Path) -> list[Path]:
"""Directories under a mounted plugin ``root`` to tmpfs-mask, or ``[]`` for a non-plugin root.

A nested skills path keeps its ancestors unmasked but masks their other children;
symlinked children are skipped.
"""
root = root.resolve()
if not (root / Path(*_PLUGIN_MANIFEST_RELPATH)).is_file():
return []

keep = {(root / ".claude-plugin").resolve(), *manifest_skill_dirs(root)}

# A kept path or an ancestor of one: never masked, descended into instead.
protected: set[Path] = set()
for kept in keep:
protected.add(kept)
for ancestor in kept.parents:
if ancestor == root:
break
if root in ancestor.parents:
protected.add(ancestor)

# Never descend into a kept path. Under manifest `skills: "."` the root itself
# is kept, so nothing is masked (the caller warns).
descend = {p for p in ({root} | protected) if p not in keep}

masked: set[Path] = set()
for directory in descend:
if not directory.is_dir() or directory.is_symlink():
continue
for child in directory.iterdir():
if child.is_symlink():
continue
if not child.is_dir():
continue
resolved = child.resolve()
if resolved in protected:
continue
masked.add(resolved)

return sorted(masked)
4 changes: 2 additions & 2 deletions src/coder_eval/orchestrator.py
Original file line number Diff line number Diff line change
Expand Up @@ -1957,8 +1957,8 @@ async def _communicate_attempt() -> TurnRecord:
# ANTI-CHEAT WINDOW. The reference and the task dir sit at mode 000 for every
# communicate attempt, retries included (this wrapper is outside
# execute_with_retry), and are restored on every exit path. The SANDBOX owns
# whether a chmod window means anything for its driver. It does NOT hide the
# task DEFINITION: task.yaml is also staged at /work/input, an unsolved gap.
# whether a chmod window means anything for its driver. The task DEFINITION
# staged at /work/input is not windowed; the entry point deletes it after load.
# Rationale: .claude/notes/permissions.md § Reference solutions and the anti-cheat window
assert self.sandbox is not None
async with self.sandbox.set_permissions([self._reference_dir, self.sandbox.task_dir]):
Expand Down
Loading
Loading