feat(evals): add first agent tool-use evaluation suite - #8409
sudoKrishna wants to merge 3 commits into
Conversation
Add a deterministic eval layer for the agent harness. Scenarios script the OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery without a provider key. - apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report - `bun run test:evals` from apps/sim runs the suite and writes the report - picked up by the normal vitest run so a regression fails CI - README documents the contract and how to add a case
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
|
| return async () => { | ||
| const turn = scenario.script[turnIndex] |
There was a problem hiding this comment.
Tool feedback goes unchecked The scripted model emits its next hardcoded turn without reading the messages it receives. If the loop stops forwarding a tool result or error, the retrieval, dependent-planning, and recovery cases can still produce their expected answers and pass. Check the tool messages received on later turns so these cases can catch that regression.
Knowledge Base Used: Agent execution and sandbox tasks
| const actualSequence = toolCalls.map((call) => call.name) | ||
| const checks: EvalCheck[] = [] |
There was a problem hiding this comment.
Tool arguments are not checked The harness records the arguments sent to each tool, but scoring checks only tool names, counts, iterations, and final content. A call to the right tool with missing or incorrect arguments can therefore pass, limiting the suite’s ability to catch tool-use regressions. Add argument expectations to the relevant scenarios.
|
|
||
| const logger = createLogger('AgentToolUseEval') | ||
|
|
||
| interface CapturedToolCall { |
There was a problem hiding this comment.
Relative imports violate app requirement This file imports
./types, and report.ts and scenarios.ts do the same. The Sim app’s import directive requires absolute imports and prohibits relative imports. Use the @/evals/agent-tool-use/types alias in all three files; this repository requirement must be satisfied before merging.
Context Used: Import patterns for the Sim application (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Replay the same scenarios against a real model. The model is the only thing that changes: runScenario now takes an optional completion transport and a live mode that relaxes exact assertions (ordered subsequence, minimum successes) and skips scripted-only recovery cases. - live.ts: OpenAI-compatible transport + DeepSeek factory - agent-tool-use.live.test.ts: K trials per scenario, gated on EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI - live report with pass rates, avg iterations, latency, failed checks - test:evals:live script and README knobs
…ve mode The first live DeepSeek run exposed brittle assertions, not harness bugs: the model chained the tools correctly but the checks were case-sensitive and required an internal order id. Match the retrieved value case-insensitively and let live runs accept the grounded status rather than the internal id.
|
@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel. A member of the Team first needs to authorize it. |
Summary
Adds a deterministic evaluation layer for agent tool use. Scenarios script the real OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery — no provider key required.
Closes #8408
What changed
apps/sim/evals/agent-tool-use/— scenario contract, 8 scenarios, harness, scoring, JSON + Markdown reportapps/sim/evals/README.md— how the suite works and how to add a caseapps/sim/package.json—test:evalscommand that runs the suite and writes the reportScenarios
tool-selectionplanningretrievalrecoveryHow it works
The model is scripted and the tools are stubbed, but the code under test is the real loop (
apps/sim/providers/openai-compat/streaming-tool-loop.ts) with its real dispatch, result feedback, and error handling. That keeps the suite deterministic and CI-friendly while still exercising production behavior.How to run
cd apps/sim bun run test:evalsWrites
test-results/evals/agent-tool-use.{json,md}and fails on any scenario failure. The suite is also collected by the normalbun run test, so a regression fails CI without the dedicated command.Test plan
bun run test:evals→ 8/8 pass, report writtenbun run check:test-patternspassesharness.tstype-checks against the real loop signaturesbun run type-check— run in CI (local run OOMs in this environment)Follow-up
Run the same scripted model through the full
DAGExecutorso agent-block wiring, variable resolution, and the executor retry/fallback policy are measured alongside the loop. The result/scoring shape is entry-point independent so both can share the report.