|
| 1 | +# Agent harness evaluations |
| 2 | + |
| 3 | +Measurement for the agent harness — the code that turns a model's tool calls |
| 4 | +into executed tools, feeds the results back, and keeps the turn alive when a |
| 5 | +tool fails. Unit and integration tests prove the harness handles the cases we |
| 6 | +already know about; evals measure whether it still behaves across a suite of |
| 7 | +scenarios when the harness changes. |
| 8 | + |
| 9 | +## What runs |
| 10 | + |
| 11 | +The first suite lives in [`agent-tool-use/`](./agent-tool-use) and drives the |
| 12 | +real OpenAI-compatible streaming tool loop |
| 13 | +(`apps/sim/providers/openai-compat/streaming-tool-loop.ts`) — the loop that |
| 14 | +serves OpenAI, DeepSeek, Groq, Cerebras, and the other OpenAI-compatible |
| 15 | +providers. The model is **scripted**: each scenario supplies the assistant turns |
| 16 | +(tool calls or a final answer) and the result of each tool call. That keeps the |
| 17 | +suite deterministic and runnable in CI with no provider key, while the thing |
| 18 | +being measured — tool dispatch, result feedback, error recovery — is real |
| 19 | +production code. |
| 20 | + |
| 21 | +The suites cover four behaviors: |
| 22 | + |
| 23 | +| Category | What it measures | |
| 24 | +| --- | --- | |
| 25 | +| `tool-selection` | The loop dispatches the tool the model asked for, including from a set of distractors. | |
| 26 | +| `planning` | Multi-turn, dependent and parallel tool calls execute in the right order and all results reach the next turn. | |
| 27 | +| `retrieval` | Values returned by a tool survive into the final answer instead of being dropped or invented. | |
| 28 | +| `recovery` | Tool errors, unknown tool names, and malformed argument JSON are fed back to the model rather than thrown out of the loop. | |
| 29 | + |
| 30 | +## Run it |
| 31 | + |
| 32 | +From `apps/sim`: |
| 33 | + |
| 34 | +```sh |
| 35 | +bun run test:evals |
| 36 | +``` |
| 37 | + |
| 38 | +The command writes a JSON report and a Markdown summary to |
| 39 | +`test-results/evals/agent-tool-use.{json,md}` (gitignored) and fails the process |
| 40 | +if any scenario fails. To point the report somewhere else, run Vitest directly: |
| 41 | + |
| 42 | +```sh |
| 43 | +EVAL_REPORT_PATH=/tmp/agent-tool-use.json bunx vitest run evals/agent-tool-use |
| 44 | +``` |
| 45 | + |
| 46 | +The suite is also picked up by the normal `bun run test` run, so a regression |
| 47 | +fails CI even without the dedicated command. |
| 48 | + |
| 49 | +## Add a case |
| 50 | + |
| 51 | +1. Open [`agent-tool-use/scenarios.ts`](./agent-tool-use/scenarios.ts) and add |
| 52 | + an entry to `AGENT_TOOL_USE_SCENARIOS`. |
| 53 | +2. Declare the `tools` the model may call and the `script` it produces. A |
| 54 | + `tools` turn lists the calls the model emits; an `answer` turn ends the run. |
| 55 | + Attach each call's stub `result` (or leave it to default to a successful |
| 56 | + empty output). |
| 57 | +3. Add the assertions you care about under `expect`: the ordered |
| 58 | + `toolCallSequence`, `requiredTools`/`forbiddenTools`, `finalContent`, |
| 59 | + `maxIterations`, and tool call counts. Every assertion becomes a named check |
| 60 | + in the report. |
| 61 | +4. Run `bun run test:evals`. |
| 62 | + |
| 63 | +A scenario is data, not code — there is no harness change needed for a new case. |
| 64 | + |
| 65 | +### Simulating a failure |
| 66 | + |
| 67 | +- **Tool error:** give the call `result: { success: false, error: '...' }`. |
| 68 | +- **Unknown tool:** call a `name` that is not in `tools`; the loop returns a |
| 69 | + tool-not-found error to the model. |
| 70 | +- **Malformed arguments:** set `argumentsJson` to an invalid or non-object JSON |
| 71 | + string. The loop must not execute the call and must return the parse error to |
| 72 | + the model. |
| 73 | + |
| 74 | +## Report shape |
| 75 | + |
| 76 | +`report.json` is machine-readable for dashboards and trend tracking; `report.md` |
| 77 | +is the same data as a table. Each result carries the scenario id, pass/fail, |
| 78 | +every named check with a failure detail, the final content, the executed tool |
| 79 | +invocations, and metrics: iterations, tool call counts (success/error), latency, |
| 80 | +model/tool time, first-response time, and token usage. |
| 81 | + |
| 82 | +## Scope and next steps |
| 83 | + |
| 84 | +This suite evaluates the tool loop directly. The next layer is a scenario that |
| 85 | +runs the same scripted model through the full `DAGExecutor` so agent block |
| 86 | +wiring, variable resolution, and the executor's retry/fallback policy are |
| 87 | +measured alongside the loop. The `AgentToolUseResult` shape is deliberately |
| 88 | +independent of the harness entry point so both can share scoring and reporting. |
0 commit comments