feat(search-events): Require dataset so Seer runs for natural language queries - #1371
Closed
sentry-junior[bot] wants to merge 1 commit into
Closed
sentry-junior[bot] wants to merge 1 commit into
sentry-junior[bot] wants to merge 1 commit into
Conversation
…e queries Seer's search agent translates into a single strategy, so search_events only calls it when a supported dataset is passed explicitly. Because `dataset` was optional, client models often omitted it, the handler fell back to errors, and Seer was skipped even for plain natural language queries like "slowest api calls in the last 24 hours". Make `dataset` required so the calling model picks it up front, and add a short description hint that routes API/HTTP calls, DB queries, and latency questions to spans. The embedded agent can still correct the dataset when configured. Co-Authored-By: Shaun Kaasten <shaun.kaasten@sentry.io>
Contributor
Author
There was a problem hiding this comment.
search_events evals vs main
Compared PR eval run 37019830331 with the latest main eval run 36883320840. The baseline commit is an ancestor of the PR base, with no intervening changes to eval-related paths. Overall eval average stayed 0.73.
The search_events scenario scores were mostly unchanged, with these shifts among repeated scored entries:
- Improved: “p95 span duration grouped by
span.op”, “all errors from today”, and “top 10 most frequent error types” each had one entry move from 0.5 to 1.0. - Improved: “spans where
custom.db.pool_size> 10” had one entry move from 0.0 to 0.5. - Regressed: typo case
spon.duration:>100had one entry move from 0.5 to 0.0.
These are repeated eval entries; the changes largely offset and don’t indicate a clear overall gain or regression. This suite doesn’t exercise the Seer path (no experimental mode or search-agent endpoint mock), so it can’t verify the behavior this PR targets.
Contributor
|
Closing in favour of #1374 |
skaasten
added a commit
that referenced
this pull request
Oct 2, 2026
Supersedes #1371. Instead of making `dataset` required on `search_events`, this splits search into one tool per dataset, so the agent picks the dataset by picking the tool and has fewer ways to get it wrong. ## What changed **New direct tools**, all backed by the same handler: - `search_errors`: `errors`, Seer `Errors` - `search_logs`: `logs`, Seer `Logs` - `search_traces`: `spans`, Seer `Traces` - `search_metrics`: `metrics`, Seer `Metrics` - `search_profiles`: `profiles`, no Seer - `search_replays`: `replays`, no Seer - The handler moved from `tools/catalog/search-events.ts` to `tools/support/search-events/search.ts` (`runSearchEvents`). Each tool sets its own dataset and passes `lockDataset: true`. The embedded agent prompt says the dataset is fixed, and the handler ignores any dataset switch the agent suggests. This way `search_logs` can't return spans. - Each tool has a shorter description that covers only its own dataset. The event tools have no `dataset` param, and environment filters go in `query`. `search_replays` has no `fields` and keeps its separate `environment` param. - `search_traces` documents same-trace cross-event queries ("checkout requests that also have an error log"). Seer only applies cross-event filters with the `Traces` strategy. - **Compatibility:** `search_events` is still in the catalog as a deprecated alias (`execute_sentry_tool` can still call it). It is off the direct surface and left out of skill definitions. - Direct surface (`surfaces.ts`) grows from 9 to 14 tools. That's still under the target of 20 and the hard limit of 25. - Updated next-step hints that pointed at `search_events`: trace details, profile formatter, Seer unsupported-issue message, search_issues formatters, get_sentry_resource. - Updated the plugin agent prompts, evals (`search-events.eval.ts` now expects the specific tool), the Cloudflare stdio setup copy, and the docs (spec, testing, architecture, READMEs). Regenerated the tool/skill definitions. ## Tradeoffs - **Token cost:** the direct-surface definitions go from 5,786 to 8,906 tokens (+3,120), per `pnpm run measure-tokens`. The six search tools total 4,238 tokens, compared with 1,130 for `search_events`, because each one repeats the shared params. Possible follow-ups: keep `search_profiles`/`search_replays` catalog-only, or trim the shared param descriptions. - **Misrouting still happens, just earlier.** Picking a tool by name is usually more reliable than picking from an enum, but "slow requests with an error log" can still go to `search_logs` instead of `search_traces`. The cross-event hint on `search_traces` is meant to catch that. - **Out of scope:** the Seer gate still skips structured-looking queries and explicit fields/sort. That's a separate change. ## Testing - `mcp-core`: 1749 tests pass. That includes new tests for each tool, covering the locked dataset and the schema with no `dataset` param. `search_traces` also has a test that sends a natural-language query to Seer's `Traces` strategy. - `mcp-cloudflare`: 430 tests pass. - Typecheck passes for core, cloudflare, server, evals and test-client. `packages/cli` typecheck fails in my sandbox with a Node loader error (`ERR_INVALID_RETURN_PROPERTY_VALUE`). The same failure happens on clean `main`. - Lint passes. - Not yet run: CI evals, and a live check against an org that has Seer enabled. <!-- junior-request-attribution:start --> via **shaun.kaasten**. <!-- junior-request-attribution:end --> <!-- junior-session-footer:start --> <!-- junior-conversation-id:slack%3AD0BAS2BU2TC%3A1790949170.109589 --> -- [View Junior Session](https://junior-prod.sentry.dev/conversations/slack%3AD0BAS2BU2TC%3A1790949170.109589) [[Sentry]](https://sentry.sentry.io/explore/conversations/slack%3AD0BAS2BU2TC%3A1790949170.109589/?project=4510944073809921) <!-- junior-session-footer:end --> --------- Co-authored-by: sentry-junior[bot] <264270552+sentry-junior[bot]@users.noreply.github.com> Co-authored-by: Shaun Kaasten <shaun.kaasten@sentry.io>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Seer's search agent (
/search-agent/start/) translates into a single fixedstrategyand can't pick one on its own. Because of that,search_eventsonly calls Seer when the caller passes a supporteddataset.datasetwas optional, so client models often left it out. The handler then fell back toerrorsand skipped Seer, even for plain natural language queries likeslowest api calls in the last 24 hours.The calling model already reads the question, so this PR makes it choose the dataset:
datasetis now required in thesearch_eventsinput schema. The embedded agent can still correct it when one is configured.params.dataset ?? "errors"fallback and thehasExplicitDatasetbranch are removed, sincedatasetis always set now.errorsdefault now passdataset: "errors", so their behavior is unchanged. A new schema test covers the missing-dataset case.datasetare updated, and tool/skill definitions are regenerated.Breaking-ish: calls that omit
datasetnow fail schema validation instead of defaulting toerrors.Validation
tscpasses for every package exceptpackages/cli.pnpm run lintpasses.mcp-coretests: 1737 passed.mcp-cloudflaretests: 430 passed.packages/clitypecheck and tests fail in this sandbox withERR_INVALID_RETURN_PROPERTY_VALUEfrom the Node loader hook. The failure also happens on a cleanmain, so it isn't caused by this change.via shaun.kaasten.
--
View Junior Session [Sentry]