fix(codex): record provider error notifications on the turn - #203
Merged
Merged
Conversation
Codex retries a stalled or failed provider stream itself and reports each attempt only through its app-server `error` notification (`willRetry`). `_CodexTurnState.dispatch` dropped that method, so a provider stall left no trace in task.json or task.log: adhoc-2026-09-28_16-14-30 v2 skill-flow-slack-channel-description lost 1,125 s of its 1,200 s budget to two stalls on a ~12.5K-token request before the first tool call, and the record shows only two long "thinking" messages. - dispatch routes `error` to a new `on_error`, which logs a WARNING and appends a `ProviderError` (at, message, kind, http_status, will_retry, details) to the turn state. - `AgentEndEvent` / `TurnRecord` carry `provider_errors`; EventCollector passes it through, so crashed partial turns keep it too. - REPORT_SCHEMA documents the field; golden snapshots gain the empty list. Closes part 1 of #151 (part 2, per-request TTFT timing, stays open). 🤖 Generated with Claude Code Co-Authored-By: [Claude](mailto:noreply@anthropic.com) Claude-Session: https://claude.ai/code/session_014Tw2Pqkugyik4uVJHo2Ref
tmatup
requested review from
CarlesUIPath,
akshaylive,
bai-uipath and
uipreliga
as code owners
September 28, 2026 19:52
|
Claude finished @tmatup's task in 1m 35s —— View job Code Review in Progress
|
bai-uipath
approved these changes
Sep 28, 2026
bai-uipath
left a comment
Collaborator
There was a problem hiding this comment.
Approve: makes sense, small and safe. It only adds visibility: how the run behaves and how it's scored stay the same.
| Before | After | |
|---|---|---|
| Codex retries a stalled stream | Same | Same |
| Task outcome | TIMEOUT, scored as a miss | TIMEOUT, scored as a miss |
task.log |
Nothing | One WARNING per provider error |
task.json |
Long "thinking" messages, no errors | One provider_errors row per error (kind, HTTP status, will_retry) |
Worth fixing, not blockers:
- The WARNING line leaves out
details. Codex can put the underlying cause there instead of inmessage. Fix: include it, sotask.logshows the cause whichever field it lands in. - An unparsed payload writes a misleading row. If the SDK can't validate the notification, it arrives as
UnknownNotificationand still reaches the error handler, which records an empty message andwill_retry=False. Fix: read the rawparams, or skip the row.
Follow-up: a provider-caused timeout still counts against the model, and a will_retry=False error still ends the turn as COMPLETED. Reclassifying those as infra errors is the payoff of this data.
Minor: the test class docstring describes the old behavior, and the clean-turn test is already covered by the golden snapshots.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
Short version: codex already tells us when the model provider stalls or fails and it is retrying. We threw that message away. This PR keeps it in
task.jsonandtask.log.errornotification ({error: TurnError, willRetry}) for every provider stream error. That includes the ones it retries internally (for example, the default 300 s stream idle timeout)._CodexTurnState.dispatchrouted only 5 methods, soerrorfell through. A provider stall read as "the model was slow".dispatchrouteserrortoon_error. It logs a WARNING and records aProviderErrorrow:at,message,kind(thecodexErrorInfocategory),http_status,will_retry,details.AgentEndEvent, then land inTurnRecord.provider_errors, so crashed partial turns keep them too.Example of a new
task.logline:Why now
Run
adhoc-2026-09-28_16-14-30, arm v2, testskill-flow-slack-channel-description(gpt-5.6-luna, Azure custom provider) ended in TIMEOUT. It spent 553 s + 572 s on two model calls for a ~12.5K-token request, before the agent ran its first command. The other ~36 tasks on the same endpoint in the same 20 minutes had no model call over 60 s. The run record holds only two long "thinking" messages, so we cannot say whether codex hit its idle timeout, got a 5xx, or waited on headers. This PR makes that visible next time.Part 1 of #151. Part 2 (per-request time-to-first-token) stays open.
Not in this PR: changing the codex provider retry settings (
stream_idle_timeout_ms300 s ×stream_max_retries5 is larger than the 1200 s task budget). That is a behaviour change for every eval, so it is proposed separately.Test plan
TestProviderErrorNotifications(3 cases), built from realopenai_codexErrorNotificationpayloads: retried stream error with kind + HTTP status, a bare category, missing info, and a clean turn (empty list).GOLDEN_REGEN=1). The only diff is"provider_errors": []in all 34.make verify: 6104 passed, 3 skipped.🤖 Generated with Claude Code
https://claude.ai/code/session_014Tw2Pqkugyik4uVJHo2Ref