## Summary Add reportable metrics for model-backed behavioral runs so teams can compare baseline and candidate tool interfaces. ## Acceptance criteria - [ ] Report tool-selection accuracy by probe and aggregate - [ ] Report argument-validity rate and validation failures - [ ] Report risk and confirmation expectation compliance - [ ] Results are available in JSON and Markdown - [ ] Metrics distinguish missing data from failed evaluations - [ ] Tests use deterministic fixture results
Summary
Add reportable metrics for model-backed behavioral runs so teams can compare baseline and candidate tool interfaces.
Acceptance criteria