Separate Brunch-authored Voice speech from on-screen responses - #9622
Draft
kostandinang wants to merge 2 commits into
Draft
Separate Brunch-authored Voice speech from on-screen responses#9622kostandinang wants to merge 2 commits into
kostandinang wants to merge 2 commits into
Conversation
Co-authored-by: Kostandin Angjellari <ka@hash.ai>
Co-authored-by: Kostandin Angjellari <ka@hash.ai>
kostandinang
added this pull request to stack #9623
September 9, 2026 11:27
|
The latest updates on your projects. Learn more about Vercel for GitHub.
1 Skipped Deployment
|
|
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🌟 What is the purpose of this PR?
Test whether Brunch can author a useful brief spoken answer or takeaway alongside a complete on-screen response. Voice playback waits for the whole correlated reply, including browser-tool continuations, and Realtime delivers the selected Brunch text verbatim.
This is a draft follow-up stacked on #9585. Local deterministic and browser checks establish routing, persistence, completion gating, and replay—not model adherence, audible usefulness, fidelity, tolerable delay, or deployed acceptance. Typed effective instructions and core/SDCPN prompts remain unchanged.
🔗 Related links
libs/@hashintel/brunch-agent/docs/evidence/implementations/separate-voice-speech/verification.md🚫 Blocked by
🔍 What does this change?
Brunch authors speech separately from full visible prose using the existing structured-data and conversation-history paths. Speech stays inspectable beside its response. The application releases only the final correlated speech after the whole reply completes; missing or unusable speech gets a fixed reading notice, never a fabricated summary or automatic report reading. Full-response and exact-question replay retain their existing meanings.
🏗️ Agent notes
The stack contains a separate mission-authority commit followed by the implementation commit. The parent history is unchanged. The new implementation commit has no Amp session-ID trailer.
The existing data writer and generic tool-result card carry the authored speech; no new conversation store or Realtime domain authority is introduced. A later substantive tool invalidates an earlier speech draft unless Brunch replaces it. The panel's derived busy state covers browser-tool continuations. Successful routing is not proof that a model's content is useful or true.
The inherited CORS mission is preserved as historical and unadjudicated, not accepted or discarded. There is exactly one live mission on this branch. The current contract follows.
Separate Brunch-authored Voice speech and display
Status
Live; scope and implementation authorized by the owner. This is the sole mission on local
voice/separate-brunch-speech, stacked on #9585at 0902dddb.
Authority was committed separately before product changes. The owner has authorized committing,
pushing, and opening this follow-up as a draft stacked PR. Issue creation, paid provider activity,
manual deployment changes, and mission acceptance remain unauthorized.
The accepted scope and timing decision are in the
owner conversation.
The prior planning conversation contains the pasted meeting transcript; the Notion proposal and
Slack discussion remain unreviewed. The inherited CORS contract is preserved without adjudicating
its acceptance in its historical record.
No provisional future draft is consumed.
Imperative
Determine whether Brunch can give a useful brief spoken answer or takeaway alongside complete
on-screen content while retaining domain authority. Judge usefulness, fidelity, and delay
separately. This tests separate Brunch-authored outputs under whole-correlated-reply completion
gating, not the best possible latency of a relay.
Visible advance: a Voice clarification gives a useful answer rather than only a reading notice;
a long analysis gives a substantive takeaway while preserving the full report on screen.
Demo: open the prepared crew-reservation fixture, ask the two comparison inputs below, inspect
the spoken content and report, request full reading, repeat a marked question, interrupt, Stop,
and reopen. The local panel is the initial proof boundary; no deployed claim follows from it.
Throughline
Existing Voice admission and delivery-scoped context → Brunch tools and browser continuations →
Brunch-authored spoken and displayed outputs → completion of the whole correlated reply →
application-selected verbatim Realtime playback.
modelling goal; clarify first when ambiguity would materially change the answer.
speech is not proof that it was heard. Read full response selects the complete displayed text.
replay; a spoken question preserves that wording.
An earlier completed message/submission does not suffice. The gate does not wait for the next
user answer. Failed/aborted replies do not release pending automatic speech.
report reading. A delivery notice is not successful substantive delivery.
Ownership and permitted changes
App-owned ChatAgent Voice instructions own output separation. Core SYSTEM.md, its question
semantics, and SDCPN prompts/skills remain unchanged, including full recoverable-workpiece and
prepared-fixture obligations. Realtime remains a delivery-only renderer with no domain tools,
independent questions, conclusions, or summaries. Typed effective instructions remain unchanged.
Existing structured data writers/transport are a candidate, not a preselected schema. First pin
live completion, persisted history, response identity, continuation folding, and replay. Use the
existing conversation route/store. Stop if a new store or broader runtime redesign is required.
Expected owners: ChatAgent and its tests; core's shared data contract if required without changing
universal prompts; AI SDK streaming/history/correlation; website Voice selection, bridge,
controller, session/policy, and minimal response-associated inspection UI. Update relevant user
documentation if exposed behavior requires it. No unrelated prompt or infrastructure cleanup.
Proof
apps/brunch-agent/test/voice-context.test.tsand transport admissions distinguishtyped → Voice → browser continuation → typed. Typed instructions and tool availability remain
unchanged. Unknown preferences do not enable Voice behavior.
ui-stream.test.ts,transcript.test.ts, andchat-transport.test.ts, plus websitecanonical-speech.test.ts,realtime-brunch-bridge.test.ts,voice-turn-controller.test.ts,openai-realtime-session.test.ts, and browser-tool integration tests. Exercise speech databefore report completion, earlier completions followed by continuations, identical text on
distinct replies, failed/aborted continuations, missing/invalid speech, full-report selection,
exact question replay, interruption versus durable Stop, and reload without autoplay,
duplicate admission, or duplicate content. Test actual outputs, not only absence of crashes.
fixture and tool evidence. A simple clarification must be useful without gratuitous follow-up;
a consequential gap must ask a relevant marked question; a long takeaway must not contradict
the report or omit qualifications that change its meaning. A browser-tool continuation must
report success, rejection, and no-op truthfully. Deterministic checks cannot accept this leaf.
“Give me a detailed analysis of this model, including assumptions, possible bottlenecks,
missing constraints, and what still needs validation. Do not change the model.” Use
crew-reservation-v1and record model/configuration differences from FE-1630: Optimize and measure the Brunch Voice relay #9585. Inspect asynchronized audible browser recording plus representative UI/accessibility states.
Paid execution is blocked until an explicit bounded owner authorization; no campaign.
speech request, and first substantive audible answer separately. Synchronized audio/human
inspection is the first-audible oracle; provider buffer events and notices are not answers.
Report no answer when none is heard. No invented word or latency acceptance threshold.
test:unit,lint:tsc,lint:eslint,build, changed-fileOxfmt and
git diff --check. Render and inspect the affected UI. Report blocked/failed checkshonestly. Local and mocked checks do not establish deployed end-to-end behavior.
The #9585 evidence records 192 → 151 clarification words and only a notice spoken afterward;
the long report stayed complete and opt-in reading worked in the recorded run. These were
individual synthetic-input real-provider diagnostics, not a statistical campaign or human
acceptance. They do not prove core caused verbosity or exhaust relay prompt alternatives.
Constraints
interruption versus durable Stop, and reload without autoplay/duplication.
store, workpiece/provenance redesign, unrelated infrastructure work, or paid campaign.
not an upstream-supported API; do not broaden that exception silently.
Fog-line
The existing structured data writer is now selected for the local implementation: runtime
restart, transport, real-panel continuation, and replay checks establish the tested routing
contract. Local verification
records the evidence and its limits; no real-provider or deployed acceptance follows.
Model adherence, useful brevity, speech/report consistency, actual audio fidelity, and tolerable
delay remain experimental. Completion gating avoids speculative delivery, not semantic errors.
Historical preview configuration and backend-deployment verification remain unresolved; any
remote claim requires a new real deployed witness. Human acceptance and paid ceilings are
owner-held. No separate Linear issue is linked, and no Linear integration is available in this
orb. Draft publication uses the repository's descriptive-title contribution workflow; Linear
writes still require explicit approval.
Stop or reorient
Stop if typed behavior inherits Voice, speech loses response identity, cancellation allows later
autoplay, replay duplicates content, workpieces are incomplete, or claims exceed tool evidence.
Reorient if outputs repeatedly contradict, substantive speech is absent, or the small routing
change requires broader mechanisms. Do not weaken the oracle or repair content in Realtime.
Verdicts update this relay variant only. Content success with unacceptable delay leaves earlier
delivery unresolved for a follow-up; it does not select another architecture. Prepare evidence
and stop for owner acceptance rather than declaring naturalness or mission closure.
Deferred
MISSION.next.md retains the existing future spine and inherited limitations.
Its CORS transition pointer preserves deployment, authentication, and rate-limit owners.
Earlier delivery re-enters only if measured delay is unacceptable despite content success;
its safety and benefit need a separate scope and audible oracle. Alternative architecture
selection remains owner-held, not an automatic consequence of any experimental failure.
Pre-Merge Checklist 🚀
🚢 Has this modified a publishable library?
This PR:
📜 Does this require a change to the docs?
The changes in this PR:
🕸️ Does this require a change to the Turbo Graph?
The changes in this PR:
NODE_OPTIONS=--no-experimental-webstorageunder this orb's Node 26.5.1 to avoid the Node/jsdom localStorage conflict. No application or test changes were made to hide that failure.hash-backend-utilsbuild reports missing graph-workspace dependencies; its emitted OpenTelemetry entry allowed the Brunch build and full tests to pass. No clean whole-monorepo build is claimed.🐾 Next steps
After bounded paid-run authorization, repeat #9585's two recorded inputs on
crew-reservation-v1, then inspect a consequential modelling gap, browser-tool continuation, mixed typed/Voice use, interruption, Stop, and reload. Record configuration differences. Judge usefulness, fidelity, and delay independently with synchronized audible evidence.If content succeeds but delay is unacceptable, record earlier delivery as an unresolved follow-up. Failure of this variant does not select delegation, Brunch-as-client-tool, or a broad core-prompt redesign. Keep this PR in draft pending the outstanding gates.
🛡 What tests cover this?
git diff --checkpassed. Petrinaut library build passed.❓ How to test this?
Check out this child branch, not the parent alone, and build its workspace dependencies.
Run:
In an authorized real-provider Voice trial, ask “What does reserving a dispatch crew mean here?” and “Give me a detailed analysis of this model, including assumptions, possible bottlenecks, missing constraints, and what still needs validation. Do not change the model.” Verify useful speech, complete visible content, preserved consequential qualifications, and one delivery after the entire correlated reply completes.
Check Read full response selects the full visible report, Repeat question selects only the exact marked question, interruption does not durably stop Brunch, and Stop aborts active Flue work while preserving the settled-step limitation above. Check reload does not autoplay or duplicate turns. Measure first substantive audible answer separately from transcription, acknowledgements, and notices; no owner-approved word or latency threshold is assumed.
📹 Demo
The actual website/Petrinaut panel was rendered in Chromium with a seeded local fixture. The inspected screenshot shows complete authored speech, including its qualification, separately from the full report. Reload/reopen DOM checks found both texts and exactly one user message.