fix(docker): close two driver:docker channels that leaked grading material to the agent - #179
Conversation
…(anti-cheat) Fix A: make /work/input a writable mount (drop :ro; grant input dir writable) and delete the staged task.yaml immediately after load_task in the in-container entry point, gated on IN_CONTAINER_ENV. The staged file is the post-override TaskDefinition with success_criteria, and the agent runs in the same container, so leaving it readable hands over the grading answer key. It is read exactly once; grading reads criteria from memory. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fix B: default-deny mask over an auto-mounted Claude-plugin root. The whole root stays :ro-mounted so the plugin loads, but every child dir outside the keep-set (.claude-plugin + manifest-declared skill dirs) is masked with an empty tmpfs. Colocated eval material — sibling task YAMLs, reference solutions, tests/, node_modules/ — is masked by default so an unknown layout can never leak. New pure helper eval_material.mask_dirs; shared skill-dir resolver manifest_skill_dirs (renamed public in agents/_skills.py) is the SSOT for "what is a skill dir". Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fix C: a static lint rule (pytest class, CE055 template) that flags the one residual the Fix B allowlist cannot close — a task_id: YAML or a resolved reference.directory colocated INSIDE a skill dir reached through an in-repo agent.plugins[].path / sandbox.template_sources[].path. Masking a skill dir would hide the skill, so such material stays readable and leaks the grading answer key under driver: docker. Reuses the shared manifest_skill_dirs resolver (SSOT, import-identity asserted) with positive+negative sensors. Documents both docker anti-cheat fixes in docs/DOCKER_ISOLATION.md and the CE065 entry in CLAUDE.md. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Code review found Fix A closed only /work/input/task.yaml but left the identical success_criteria readable in /work/input/context.json via source_yaml (top-level AND inside every config_lineage entry / ConfigLineageEntry.source_yaml) — in a dir this change makes world-readable+writable. The agent could cat context.json to recover its grading answer key. Delete both staged files after they are read into memory (_scrub_staged_task_yaml -> _scrub_staged_inputs); correct the docstring and docs/DOCKER_ISOLATION.md (the 'context.json carries no criteria' claim was false). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
L1 — CE065 now flags ANY task_id: YAML or reference dir left readable under a plugin root (not just inside a skill dir): it reuses the runtime mask_dirs (SSOT), walks the readable surface, and catches a loose task_id: YAML *file* at the plugin root too (a tmpfs masks a dir, not a single file). Renamed the class to TestCE065EvalMaterialReadableUnderPluginRoot; added a loose-root-file sensor. L2 — docker_runner resolves a relative agent.plugins[].path against the task-file dir (new _resolve_mount_path), not CWD, matching reference/template/CE065 resolution so the static rule and the runtime inspect the same tree. L3 — eval_material.mask_dirs no longer descends a kept path, so a degenerate manifest skills: '.' masks nothing (fail-safe) instead of hiding the whole skill surface; added a test. Updated CLAUDE.md CE065 entry and docs. make verify green (the one failure is a pre-existing flaky wall-clock datetime test in test_agent_telemetry, passes on rerun, unrelated to these files). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-model review (Opus + Fable + Sonnet) found no Critical/High and no leak bypass; addressed the 3 Mediums and the doc drift: M2 — nested plugin roots (plugin B under plugin A) produced a duplicate Docker mount destination: A's mask emitted --tmpfs <A>/B while B's own mount bound <A>/B:ro. Defer mask emission until all binds are known and drop any mask whose path is also a bind (the bind wins so B loads and masks its own children). Extracted the auto-mount block into _append_auto_mounts to stay under the per-method lint limits. Added a nested-plugin test. M1 — added direct tests for _resolve_mount_path (relative->task-file dir, absolute unchanged, task_file None->cwd, env-var expansion); the L2 fix had no coverage. M3 — warn when a plugin root's mask stands down (manifest skills: '.') so the anti-cheat mask never voids silently. Added a test. Doc drift (L-a/L-b/L-c): corrected the now-false 'task.yaml readable at /work/input' / 'hiding criteria is unsolved' claims in orchestrator.py, fs_permissions.py and CLAUDE.md (Fix A deletes task.yaml + context.json after load); widened DOCKER_ISOLATION.md's residual note to include the loose-root-file case; marked the harness-candidates entry RESOLVED (nothing deferred). make verify green (5758 passed). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-leaks # Conflicts: # CLAUDE.md
|
Claude finished @CarlesUIPath's task in 1m 21s —— View job Code Review in Progress
|
bai-uipath
left a comment
There was a problem hiding this comment.
Fix A is good, Fix B breaks real plugins. Suggest landing Fix A on its own and reworking the rest.
-
The mask hides the plugin, not just eval material. The keep-set is only
.claude-pluginplus the skill dirs, sohooks/,commands/andagents/(directories Claude Code loads by default) get masked too. On the skills nightly's staged.plugin-root(UiPath/skills#3546) that drops the plugin's hooks and commands and hidestests/scripts/stage-preview-sdk-workspace.sh, which the nightly pre_run calls, so the eval grades a different plugin than the one that ships. Fix before merge: keep the full plugin surface (default component dirs plus every manifest-declared path), or replace the mask with a preflight that errors when a mounted plugin root containstask_id:YAMLs. -
CE065 checks nothing real. It only walks this repo's
tasks/, where no task resolves to a plugin root, and never sees the suites where this layout actually lives. Fix: cut it, or fold it into the runtime preflight above so it covers every suite. -
Needs a rebase, and the comments won't pass main's new prose gate. Several new docstrings and comment blocks exceed the caps, and the same rationale is repeated across about six files. Fix: one section in
.claude/notes/isolation.mdplus aRationale:pointer. -
Fix A removes one copy of the criteria, not all of them. The same YAML (and its siblings) is still in
/work/task_dir, behind a chmod window that gemini-3.8-flash bypassed in the BPMN sweep. Not a blocker, but the comments saying the agent "cannot read its criteria" should be scoped to/work/input.
Brings the branch up to 0.12.7 (main 101bb5c). - Conflicts took main's slimmer comment style; the anti-cheat claims were updated in place (task.yaml is no longer "an unsolved gap"). - .claude/architecture-notes.md was deleted on main, so the edit to it is dropped; the two now-false permissions.md sentences are deleted. - The eval-material lint rule is renumbered CE065 -> CE068, since main took CE065 for the pricing mirror. It is added to ruff's external list, and "claim 069 next" is updated. - The branch's long comments and docstrings were trimmed to pass main's prose budget (8-line comment runs). - Dropped the branch's "resolved in-branch" harness-candidates entry. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bai-uipath
left a comment
There was a problem hiding this comment.
Verdict: good to merge. It blocks the in-container leaks and leaves the web hole open. It breaks one skills task, which needs a small fix before the coder-eval pin bump.
Gemini reruns
Gemini 3.8 Flash on Antigravity, the paper's settings. Two SkillsBench tasks, run in docker.
| build | runs | cheated | how |
|---|---|---|---|
| 0.10.2 (paper) | 4 | 3 | read tests/ answers in the container |
| main | 20 | 3 | downloaded ground_truth.json from GitHub with Antigravity's web tools |
| PR #179 | 19 | 0 | none |
The gap is a small sample (p = 0.21). All 3 main cheats went through the web, which #179 doesn't touch, so given enough runs they will show up on #179 too.
What it blocks
- Closed: the staged
task.yaml/context.jsonin/work/input(now deleted after load), and eval files sitting next to the skills in a plugin root (now masked). A probe task confirmed that 4 of the 8 routes open on main are closed. - Still open: Antigravity's web tools and the network, which is the route Gemini actually uses. The root
chmodreopen of/work/task_diris also still open, though no run tried it.
Skills-suite check
Run against skills main with the nightly config: docker, whole repo as the plugin.
- Unchanged: skills load identically (28 UiPath skills, 6 agents) and template fixtures are still visible. 45 of 46 real tasks set up exactly as on main, and the activation rate is 42 vs 43 of 91.
- Lost: the plugin's SessionStart hooks (telemetry only) and
/uipath:install-permissions. - Breaks one task:
skill-troubleshoot-smoke-manifest-commands. Itspre_runreads$SKILLS_REPO_PATH/tests/..., which is now masked.- Fix: a task-level
extra_mountsfortests/tasks/uipath-troubleshoot, restating the~/.uipathmount. - Timing: land it before bumping the pin, currently 0.12.7.
- Fix: a task-level
Status
- Merge: main merged into the PR (
8f6a364a), CI 17/17 green, waiting on review. - Next: close the web hole by passing
CapabilitiesConfig(disabled_tools=[SEARCH_WEB, READ_URL_CONTENT])inantigravity_agent.py. The SDK supports it, and it's about 30 lines.

Summary
Under
driver: docker, an eval agent could read its own grading material as follows:This PR closes both channels, adds a lint guard so the skills layout can't recur, and is verified end-to-end with a live adversarial agent. That is, a test which asks the agent to cheat by finding reference files.
The existing reference anti-cheat window #106 (mode-000 on the staged reference copy) did not cover either channel: it guards the reference copies, while the criteria sat unshielded at
/work/inputand the plugin mount exposed the originals.The fixes
/work/inputis now writable, and the in-container entry point deletes bothtask.yamlandcontext.jsonimmediately afterload_task(_scrub_staged_inputs), before any agent turn on all paths (normal + regrade). Deleting should be a safe op since graders do not read at task.yaml after agent finishes.isolation/eval_material.py::mask_dirs: the plugin root stays mounted:ro, but every child dir outside the keep-set (.claude-plugin+ the manifest-declared skill dirs) is--tmpfs-masked. Everything that isn't the skill surface is hidden by default, so an unknown/new eval layout can never leak;tests/, references,node_modules/are masked for free. This is stronger than a chmod window. Graders do not look into the mounted tasks, since these resources are present in task dirs from the previous anti-cheat PR feat(reference)!: directory-only references + anti-cheat permission window #106.task_id:YAML orreference.directoryleft readable under an in-repo plugin root (colocated inside a skill dir, or a loose file at the root that a tmpfs can't mask). Reusesmask_dirs+manifest_skill_dirsas the single source of truth, pinned by an import-identity test.Verification
make verify: 5758 passed, 4 skipped (green; 2 pre-existing env-driven failures only appear when a local.envleaksTASK_DIR/model — unrelated files, unchanged from main).find …/tests…andfind … reference→ no output;cat /work/input/task.yamlandcat /work/input/context.json→ No such file or directory;verdict.txt: DENIED; anti-cheat criterion PASS, task SUCCESS.mask_dirs,_resolve_mount_path, nested-plugin dedup, theskills:"."stand-down, and_scrub_staged_inputs.Known residuals (documented defense-in-depth)
chmodback; full containment needs a non-root uid) is unchanged — out of scope.skills:"."stands the mask down (fail-safe, now logged + CE065-flagged).Files
src/:isolation/eval_material.py(new),isolation/docker_runner.py,cli/run_task_internal_command.py,agents/_skills.py(manifest_skill_dirsmade public/SSOT),orchestrator.py,fs_permissions.py.tests/:test_eval_material.py(new),test_run_task_internal.py(new),test_docker_runner_mounts.py,test_custom_lint.py(CE065).Docs:
docs/DOCKER_ISOLATION.md,.claude/architecture-notes.md,CLAUDE.md.🤖 Generated with Claude Code