Skip to content

feat(evals): add first agent tool-use evaluation suite - #8409

Open
sudoKrishna wants to merge 3 commits into
simstudioai:mainfrom
sudoKrishna:feat/agent-tool-use-evals
Open

sudoKrishna wants to merge 3 commits into
simstudioai:mainfrom
sudoKrishna:feat/agent-tool-use-evals

Conversation

@sudoKrishna

Copy link
Copy Markdown

Summary

Adds a deterministic evaluation layer for agent tool use. Scenarios script the real OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery — no provider key required.

Closes #8408

What changed

  • apps/sim/evals/agent-tool-use/ — scenario contract, 8 scenarios, harness, scoring, JSON + Markdown report
  • apps/sim/evals/README.md — how the suite works and how to add a case
  • apps/sim/package.json — test:evals command that runs the suite and writes the report

Scenarios

Category Cases
tool-selection single relevant tool; correct tool among distractors
planning multi-step dependent chain; parallel independent calls
retrieval retrieved value survives into the final answer
recovery tool error then retry; unknown tool; malformed argument JSON

How it works

The model is scripted and the tools are stubbed, but the code under test is the real loop (apps/sim/providers/openai-compat/streaming-tool-loop.ts) with its real dispatch, result feedback, and error handling. That keeps the suite deterministic and CI-friendly while still exercising production behavior.

How to run

cd apps/sim
bun run test:evals

Writes test-results/evals/agent-tool-use.{json,md} and fails on any scenario failure. The suite is also collected by the normal bun run test, so a regression fails CI without the dedicated command.

Test plan

  • bun run test:evals → 8/8 pass, report written
  • Negative check: intentionally broke an expectation → suite failed with the failing check, then reverted
  • bun run check:test-patterns passes
  • harness.ts type-checks against the real loop signatures
  • Full bun run type-check — run in CI (local run OOMs in this environment)

Follow-up

Run the same scripted model through the full DAGExecutor so agent-block wiring, variable resolution, and the executor retry/fallback policy are measured alongside the loop. The result/scoring shape is entry-point independent so both can share the report.

Screenshot From 2026-09-29 15-20-31

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
@sudoKrishna
sudoKrishna requested a review from a team as a code owner September 29, 2026 09:51
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 29, 2026 9:51am UTC

Request Review

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds evaluation suite for the agent tool-use loop.

The evaluation suite has no identified runtime blocker, but its import paths must satisfy the repository requirement before merging.

Findings

  1. P2 Tool feedback goes unchecked ▶
  2. P2 Tool arguments are not checked ▶
  3. P2 Relative imports violate app requirement ▶

Summary

The PR adds eight deterministic scenarios that drive the production OpenAI-compatible streaming tool loop, score outcomes, and write JSON and Markdown reports.

  • The suite exercises dispatch and error accounting without a provider key.
  • Its scripted model does not inspect tool feedback, and scoring does not check dispatched arguments, limiting the regressions it can detect.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Scenario script] --> B[Scripted model turns]
  B --> C[Production streaming tool loop]
  C --> D[Stub tool results]
  D --> C
  C --> E[Scoring]
  E --> F[JSON and Markdown reports]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): add agent tool-use evaluati..."

Comment on lines +110 to +111
return async () => {
const turn = scenario.script[turnIndex]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Tool feedback goes unchecked The scripted model emits its next hardcoded turn without reading the messages it receives. If the loop stops forwarding a tool result or error, the retrieval, dependent-planning, and recovery cases can still produce their expected answers and pass. Check the tool messages received on later turns so these cases can catch that regression.

Knowledge Base Used: Agent execution and sandbox tasks

Comment on lines +195 to +196
const actualSequence = toolCalls.map((call) => call.name)
const checks: EvalCheck[] = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Tool arguments are not checked The harness records the arguments sent to each tool, but scoring checks only tool names, counts, iterations, and final content. A call to the right tool with missing or incorrect arguments can therefore pass, limiting the suite’s ability to catch tool-use regressions. Add argument expectations to the relevant scenarios.


const logger = createLogger('AgentToolUseEval')

interface CapturedToolCall {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Relative imports violate app requirement This file imports ./types, and report.ts and scenarios.ts do the same. The Sim app’s import directive requires absolute imports and prohibits relative imports. Use the @/evals/agent-tool-use/types alias in all three files; this repository requirement must be satisfied before merging.

Context Used: Import patterns for the Sim application (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

This branch was previously deployed

1 inactive (outdated) deployment
Preview — 86c79d78 Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evals): add first agent tool-use evaluation suite

1 participant