Skip to content

Commit 86c79d7

Browse files
committed
feat(evals): add agent tool-use evaluation harness
Add a deterministic eval layer for the agent harness. Scenarios script the OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery without a provider key. - apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report - `bun run test:evals` from apps/sim runs the suite and writes the report - picked up by the normal vitest run so a regression fails CI - README documents the contract and how to add a case
1 parent ae75043 commit 86c79d7

7 files changed

Lines changed: 1029 additions & 0 deletions

File tree

‎apps/sim/evals/README.md‎

Lines changed: 88 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,88 @@
1+
# Agent harness evaluations
2+
3+
Measurement for the agent harness — the code that turns a model's tool calls
4+
into executed tools, feeds the results back, and keeps the turn alive when a
5+
tool fails. Unit and integration tests prove the harness handles the cases we
6+
already know about; evals measure whether it still behaves across a suite of
7+
scenarios when the harness changes.
8+
9+
## What runs
10+
11+
The first suite lives in [`agent-tool-use/`](./agent-tool-use) and drives the
12+
real OpenAI-compatible streaming tool loop
13+
(`apps/sim/providers/openai-compat/streaming-tool-loop.ts`) — the loop that
14+
serves OpenAI, DeepSeek, Groq, Cerebras, and the other OpenAI-compatible
15+
providers. The model is **scripted**: each scenario supplies the assistant turns
16+
(tool calls or a final answer) and the result of each tool call. That keeps the
17+
suite deterministic and runnable in CI with no provider key, while the thing
18+
being measured — tool dispatch, result feedback, error recovery — is real
19+
production code.
20+
21+
The suites cover four behaviors:
22+
23+
| Category | What it measures |
24+
| --- | --- |
25+
| `tool-selection` | The loop dispatches the tool the model asked for, including from a set of distractors. |
26+
| `planning` | Multi-turn, dependent and parallel tool calls execute in the right order and all results reach the next turn. |
27+
| `retrieval` | Values returned by a tool survive into the final answer instead of being dropped or invented. |
28+
| `recovery` | Tool errors, unknown tool names, and malformed argument JSON are fed back to the model rather than thrown out of the loop. |
29+
30+
## Run it
31+
32+
From `apps/sim`:
33+
34+
```sh
35+
bun run test:evals
36+
```
37+
38+
The command writes a JSON report and a Markdown summary to
39+
`test-results/evals/agent-tool-use.{json,md}` (gitignored) and fails the process
40+
if any scenario fails. To point the report somewhere else, run Vitest directly:
41+
42+
```sh
43+
EVAL_REPORT_PATH=/tmp/agent-tool-use.json bunx vitest run evals/agent-tool-use
44+
```
45+
46+
The suite is also picked up by the normal `bun run test` run, so a regression
47+
fails CI even without the dedicated command.
48+
49+
## Add a case
50+
51+
1. Open [`agent-tool-use/scenarios.ts`](./agent-tool-use/scenarios.ts) and add
52+
an entry to `AGENT_TOOL_USE_SCENARIOS`.
53+
2. Declare the `tools` the model may call and the `script` it produces. A
54+
`tools` turn lists the calls the model emits; an `answer` turn ends the run.
55+
Attach each call's stub `result` (or leave it to default to a successful
56+
empty output).
57+
3. Add the assertions you care about under `expect`: the ordered
58+
`toolCallSequence`, `requiredTools`/`forbiddenTools`, `finalContent`,
59+
`maxIterations`, and tool call counts. Every assertion becomes a named check
60+
in the report.
61+
4. Run `bun run test:evals`.
62+
63+
A scenario is data, not code — there is no harness change needed for a new case.
64+
65+
### Simulating a failure
66+
67+
- **Tool error:** give the call `result: { success: false, error: '...' }`.
68+
- **Unknown tool:** call a `name` that is not in `tools`; the loop returns a
69+
tool-not-found error to the model.
70+
- **Malformed arguments:** set `argumentsJson` to an invalid or non-object JSON
71+
string. The loop must not execute the call and must return the parse error to
72+
the model.
73+
74+
## Report shape
75+
76+
`report.json` is machine-readable for dashboards and trend tracking; `report.md`
77+
is the same data as a table. Each result carries the scenario id, pass/fail,
78+
every named check with a failure detail, the final content, the executed tool
79+
invocations, and metrics: iterations, tool call counts (success/error), latency,
80+
model/tool time, first-response time, and token usage.
81+
82+
## Scope and next steps
83+
84+
This suite evaluates the tool loop directly. The next layer is a scenario that
85+
runs the same scripted model through the full `DAGExecutor` so agent block
86+
wiring, variable resolution, and the executor's retry/fallback policy are
87+
measured alongside the loop. The `AgentToolUseResult` shape is deliberately
88+
independent of the harness entry point so both can share scoring and reporting.
Lines changed: 34 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,34 @@
1+
import { providersMock } from '@sim/testing/mocks/providers.mock'
2+
import { providersConversationHistoryMock } from '@sim/testing/mocks/providers-conversation-history.mock'
3+
import { providersUtilsMock } from '@sim/testing/mocks/providers-utils.mock'
4+
import { toolsMock } from '@sim/testing/mocks/tools.mock'
5+
import { afterAll, describe, expect, it, vi } from 'vitest'
6+
import { runScenario } from '@/evals/agent-tool-use/harness'
7+
import { writeEvalReport } from '@/evals/agent-tool-use/report'
8+
import { AGENT_TOOL_USE_SCENARIOS } from '@/evals/agent-tool-use/scenarios'
9+
import type { AgentToolUseResult } from '@/evals/agent-tool-use/types'
10+
11+
vi.mock('@/providers/conversation-history', () => providersConversationHistoryMock)
12+
vi.mock('@/tools', () => toolsMock)
13+
vi.mock('@/providers/utils', () => providersUtilsMock)
14+
vi.mock('@/providers', () => providersMock)
15+
16+
const results: AgentToolUseResult[] = []
17+
18+
afterAll(() => {
19+
const reportPath = process.env.EVAL_REPORT_PATH
20+
if (reportPath) writeEvalReport(results, reportPath)
21+
})
22+
23+
describe('agent tool-use eval suite', () => {
24+
it.each(AGENT_TOOL_USE_SCENARIOS)('$id: $name', async (scenario) => {
25+
const result = await runScenario(scenario)
26+
results.push(result)
27+
28+
const failed = result.checks.filter((entry) => !entry.passed)
29+
expect(
30+
failed,
31+
failed.map((entry) => `${entry.name}: ${entry.detail}`).join('; ') || undefined
32+
).toEqual([])
33+
})
34+
})

0 commit comments

Comments
 (0)