Skip to content

feat(search-events): Require dataset so Seer runs for natural language queries - #1371

Closed
sentry-junior[bot] wants to merge 1 commit into
mainfrom
feat/search-events-require-dataset
Closed

sentry-junior[bot] wants to merge 1 commit into
mainfrom
feat/search-events-require-dataset

Conversation

@sentry-junior

@sentry-junior sentry-junior Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Seer's search agent (/search-agent/start/) translates into a single fixed strategy and can't pick one on its own. Because of that, search_events only calls Seer when the caller passes a supported dataset. dataset was optional, so client models often left it out. The handler then fell back to errors and skipped Seer, even for plain natural language queries like slowest api calls in the last 24 hours.

The calling model already reads the question, so this PR makes it choose the dataset:

  • dataset is now required in the search_events input schema. The embedded agent can still correct it when one is configured.
  • The tool description has a short routing hint: log messages go to logs; API/HTTP calls, DB queries and latency go to spans. It's merged with the existing logs hint so the description stays under the 2048-character limit.
  • The params.dataset ?? "errors" fallback and the hasExplicitDataset branch are removed, since dataset is always set now.
  • Tests that relied on the implicit errors default now pass dataset: "errors", so their behavior is unchanged. A new schema test covers the missing-dataset case.
  • Docs and examples that omitted dataset are updated, and tool/skill definitions are regenerated.

Breaking-ish: calls that omit dataset now fail schema validation instead of defaulting to errors.

Validation

  • tsc passes for every package except packages/cli.
  • pnpm run lint passes.
  • mcp-core tests: 1737 passed. mcp-cloudflare tests: 430 passed.
  • packages/cli typecheck and tests fail in this sandbox with ERR_INVALID_RETURN_PROPERTY_VALUE from the Node loader hook. The failure also happens on a clean main, so it isn't caused by this change.
  • Not run end to end against a live org or a Seer session.

via shaun.kaasten.

--

View Junior Session [Sentry]

…e queries

Seer's search agent translates into a single strategy, so search_events
only calls it when a supported dataset is passed explicitly. Because
`dataset` was optional, client models often omitted it, the handler fell
back to errors, and Seer was skipped even for plain natural language
queries like "slowest api calls in the last 24 hours".

Make `dataset` required so the calling model picks it up front, and add
a short description hint that routes API/HTTP calls, DB queries, and
latency questions to spans. The embedded agent can still correct the
dataset when configured.

Co-Authored-By: Shaun Kaasten <shaun.kaasten@sentry.io>

@sentry-junior sentry-junior Bot left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

search_events evals vs main

Compared PR eval run 37019830331 with the latest main eval run 36883320840. The baseline commit is an ancestor of the PR base, with no intervening changes to eval-related paths. Overall eval average stayed 0.73.

The search_events scenario scores were mostly unchanged, with these shifts among repeated scored entries:

  • Improved: “p95 span duration grouped by span.op”, “all errors from today”, and “top 10 most frequent error types” each had one entry move from 0.5 to 1.0.
  • Improved: “spans where custom.db.pool_size > 10” had one entry move from 0.0 to 0.5.
  • Regressed: typo case spon.duration:>100 had one entry move from 0.5 to 0.0.

These are repeated eval entries; the changes largely offset and don’t indicate a clear overall gain or regression. This suite doesn’t exercise the Seer path (no experimental mode or search-agent endpoint mock), so it can’t verify the behavior this PR targets.

@skaasten
skaasten marked this pull request as ready for review October 2, 2026 14:44
@github-actions github-actions Bot added the risk: medium PR risk score: medium label Oct 2, 2026
@skaasten

skaasten commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Closing in favour of #1374

@skaasten skaasten closed this Oct 2, 2026
skaasten added a commit that referenced this pull request Oct 2, 2026
Supersedes #1371. Instead of making `dataset` required on
`search_events`, this splits search into one tool per dataset, so the
agent picks the dataset by picking the tool and has fewer ways to get it
wrong.

## What changed

**New direct tools**, all backed by the same handler:

- `search_errors`: `errors`, Seer `Errors`
- `search_logs`: `logs`, Seer `Logs`
- `search_traces`: `spans`, Seer `Traces`
- `search_metrics`: `metrics`, Seer `Metrics`
- `search_profiles`: `profiles`, no Seer
- `search_replays`: `replays`, no Seer

- The handler moved from `tools/catalog/search-events.ts` to
`tools/support/search-events/search.ts` (`runSearchEvents`). Each tool
sets its own dataset and passes `lockDataset: true`. The embedded agent
prompt says the dataset is fixed, and the handler ignores any dataset
switch the agent suggests. This way `search_logs` can't return spans.
- Each tool has a shorter description that covers only its own dataset.
The event tools have no `dataset` param, and environment filters go in
`query`. `search_replays` has no `fields` and keeps its separate
`environment` param.
- `search_traces` documents same-trace cross-event queries ("checkout
requests that also have an error log"). Seer only applies cross-event
filters with the `Traces` strategy.
- **Compatibility:** `search_events` is still in the catalog as a
deprecated alias (`execute_sentry_tool` can still call it). It is off
the direct surface and left out of skill definitions.
- Direct surface (`surfaces.ts`) grows from 9 to 14 tools. That's still
under the target of 20 and the hard limit of 25.
- Updated next-step hints that pointed at `search_events`: trace
details, profile formatter, Seer unsupported-issue message,
search_issues formatters, get_sentry_resource.
- Updated the plugin agent prompts, evals (`search-events.eval.ts` now
expects the specific tool), the Cloudflare stdio setup copy, and the
docs (spec, testing, architecture, READMEs). Regenerated the tool/skill
definitions.

## Tradeoffs

- **Token cost:** the direct-surface definitions go from 5,786 to 8,906
tokens (+3,120), per `pnpm run measure-tokens`. The six search tools
total 4,238 tokens, compared with 1,130 for `search_events`, because
each one repeats the shared params. Possible follow-ups: keep
`search_profiles`/`search_replays` catalog-only, or trim the shared
param descriptions.
- **Misrouting still happens, just earlier.** Picking a tool by name is
usually more reliable than picking from an enum, but "slow requests with
an error log" can still go to `search_logs` instead of `search_traces`.
The cross-event hint on `search_traces` is meant to catch that.
- **Out of scope:** the Seer gate still skips structured-looking queries
and explicit fields/sort. That's a separate change.

## Testing

- `mcp-core`: 1749 tests pass. That includes new tests for each tool,
covering the locked dataset and the schema with no `dataset` param.
`search_traces` also has a test that sends a natural-language query to
Seer's `Traces` strategy.
- `mcp-cloudflare`: 430 tests pass.
- Typecheck passes for core, cloudflare, server, evals and test-client.
`packages/cli` typecheck fails in my sandbox with a Node loader error
(`ERR_INVALID_RETURN_PROPERTY_VALUE`). The same failure happens on clean
`main`.
- Lint passes.
- Not yet run: CI evals, and a live check against an org that has Seer
enabled.

<!-- junior-request-attribution:start -->
via **shaun.kaasten**.
<!-- junior-request-attribution:end -->

<!-- junior-session-footer:start -->
<!-- junior-conversation-id:slack%3AD0BAS2BU2TC%3A1790949170.109589 -->

--

[View Junior
Session](https://junior-prod.sentry.dev/conversations/slack%3AD0BAS2BU2TC%3A1790949170.109589)
[[Sentry]](https://sentry.sentry.io/explore/conversations/slack%3AD0BAS2BU2TC%3A1790949170.109589/?project=4510944073809921)

<!-- junior-session-footer:end -->

---------

Co-authored-by: sentry-junior[bot] <264270552+sentry-junior[bot]@users.noreply.github.com>
Co-authored-by: Shaun Kaasten <shaun.kaasten@sentry.io>

This branch was successfully deployed

1 active deployment
Actions — 22109bf3 Deployed Oct 2, 2026 by sentry-junior[bot] via eval #1167
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

risk: medium PR risk score: medium

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant