fix(simulation): send the simulated user's utterance, not its tagged transcript - #204
Merged
Merged
Conversation
…transcript The user simulator is a ClaudeCodeAgent. Its `agent_output` is a tagged transcript that repeats the reply under `[ASSISTANT]` and `[RESULT - SUCCESS]`. `UserSimulator.next_user_message` sent that raw string to the coding agent. So every simulated-user turn reached the agent twice, wrapped in harness tags. conversation.log never showed it, because the log path already collapses the transcript with `_extract_utterance`. Every simulated task in adhoc-2026-09-28_16-14-30 (14/14) and every earlier pulled run since 08-26 carries the markup in `iterations[].user_input`. - Move `_extract_utterance` and its tag regex into `coder_eval.simulation.utterance`. The orchestrator imports it under its old name, so the log path is unchanged. - `next_user_message` sends the collapsed utterance. It still detects the stop token on the raw text, and `raw_text` keeps the transcript. - Two tests feed the real tagged shape. TextStubAgent returns plain text, which is why the existing tests never caught this. 🤖 Generated with Claude Code Co-Authored-By: [Claude](mailto:noreply@anthropic.com) Claude-Session: https://claude.ai/code/session_014Tw2Pqkugyik4uVJHo2Ref
tmatup
requested review from
CarlesUIPath,
akshaylive,
bai-uipath and
uipreliga
as code owners
September 28, 2026 20:32
|
Claude finished @tmatup's task in 1m 25s —— View job PR Review in Progress
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
Short version: in dialog mode, the coding agent receives each simulated-user message twice, wrapped in harness tags. This PR sends the plain message.
What the agent actually got as turn 1 (from
task.jsoniterations[0].user_input, adhoc-2026-09-28_16-14-30,skill-flow-customer-escalation-simulated):ClaudeCodeAgent. Itsagent_outputis a tagged transcript (_format_messages), not the model's reply.UserSimulator.next_user_messagesent it through unchanged.conversation.logcollapses the transcript with_extract_utterance, so the log looked clean. The raw form only shows intask.jsonuser_input.Change
_extract_utteranceand its tag regex move tocoder_eval/simulation/utterance.pyasextract_utterance. The orchestrator imports it under the old name, so the log path is unchanged.next_user_messagesends the collapsed utterance. The stop token is still detected on the raw text, andSimulatorResult.raw_textstill holds the transcript.Test plan
TestTaggedTranscriptOutputcases use the real tagged shape: an opener, and a stop turn. Both fail without the fix; the existing stubs return plain text, which hid the bug.make verify: 6103 passed, 3 skipped.Note for reviewers
This changes what agents see in every simulated task, so simulated scores may shift slightly. That is the intended fix, not a regression.
🤖 Generated with Claude Code
https://claude.ai/code/session_014Tw2Pqkugyik4uVJHo2Ref