From 35bbc872b2b08fa3b2ae137bbd42d2147b6a6d64 Mon Sep 17 00:00:00 2001 From: jmangel Date: Fri, 3 Jul 2026 07:56:15 -0400 Subject: [PATCH] Fix Bedrock Converse input_tokens double-subtracting cache tokens (#832) Fixes #828. ### Problem `Converse::Chat#input_tokens` subtracts `cacheReadInputTokens` and `cacheWriteInputTokens` from `inputTokens`. But AWS Bedrock's `inputTokens` is **already** the non-cached count. The cache buckets are reported separately, not folded in. Per the [AWS prompt-caching docs](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html): > When prompt caching is enabled, the `inputTokens` field represents **only the non-cached input tokens**... `total input tokens = inputTokens + cacheReadInputTokens + cacheWriteInputTokens` So the subtraction removes tokens that were never there, understating input on every cache hit and flooring it to `0` whenever the cached prefix exceeds the fresh input (the common multi-turn case). Any cost or usage tracking on `message.input_tokens` under-reports. ### Why pass-through is correct The `Anthropic` provider already handles the same Claude models correctly: `anthropic/chat.rb` passes `input_tokens` through raw and reads the cache buckets separately. Bedrock serves those same models with the same semantics, so this just brings Converse into line. The same fix landed in LiteLLM ([#15292](https://github.com/BerriAI/litellm/pull/15292)) and was reported in LangSmith ([#1858](https://github.com/langchain-ai/langsmith-sdk/issues/1858)). ### Change - `input_tokens` returns `usage['inputTokens']` unchanged. `cached_tokens` and `cache_creation_tokens` are still parsed separately, so nothing is lost. Streaming delegates to this method, so both paths are covered. - Updated the converse chat spec. The previous test asserted the subtraction on a payload that cannot occur under AWS semantics (`inputTokens` already excludes cache), so its fixture is corrected to a realistic one, and a regression is added for the floor-to-zero case. Verified: `bundle exec rspec spec/ruby_llm/protocols/converse/` plus `spec/ruby_llm/providers/bedrock_spec.rb` green (39 examples), rubocop clean. Co-authored-by: Carmine Paolino (cherry picked from commit dc79ce62949d5648a838d4bdffda7499e55186d7) Co-Authored-By: Claude Opus 5.5 Co-authored-by: Sam Boland --- lib/ruby_llm/protocols/converse/chat.rb | 9 +++--- spec/ruby_llm/protocols/converse/chat_spec.rb | 30 +++++++++++++++++-- 2 files changed, 33 insertions(+), 6 deletions(-) diff --git a/lib/ruby_llm/protocols/converse/chat.rb b/lib/ruby_llm/protocols/converse/chat.rb index 83addb6c1..08629c4b5 100644 --- a/lib/ruby_llm/protocols/converse/chat.rb +++ b/lib/ruby_llm/protocols/converse/chat.rb @@ -102,10 +102,11 @@ def parse_completion_body(data, raw:) end def input_tokens(usage) - input_tokens = usage['inputTokens'] - return unless input_tokens - - [input_tokens.to_i - usage['cacheReadInputTokens'].to_i - usage['cacheWriteInputTokens'].to_i, 0].max + # AWS Bedrock reports inputTokens as already non-cached; cacheReadInputTokens and + # cacheWriteInputTokens are separate buckets, not folded into inputTokens. Subtracting + # them (as inclusive providers require) understates input and floors to zero on cache + # hits. See https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html + usage['inputTokens'] end def reasoning_tokens(usage) diff --git a/spec/ruby_llm/protocols/converse/chat_spec.rb b/spec/ruby_llm/protocols/converse/chat_spec.rb index c90368308..df2d77d60 100644 --- a/spec/ruby_llm/protocols/converse/chat_spec.rb +++ b/spec/ruby_llm/protocols/converse/chat_spec.rb @@ -4,7 +4,9 @@ RSpec.describe RubyLLM::Protocols::Converse::Chat do describe '.parse_completion_response' do - it 'normalizes cache read and write tokens out of input tokens' do + it 'exposes AWS inputTokens as-is (already non-cached) and keeps cache buckets separate' do + # Per AWS, inputTokens already excludes cache; a real payload sends the non-cached count + # directly, with cache read/write reported separately. response_body = { 'modelId' => 'anthropic.claude-sonnet-4-5-20250929-v1:0', 'output' => { @@ -13,7 +15,7 @@ } }, 'usage' => { - 'inputTokens' => 100, + 'inputTokens' => 50, 'outputTokens' => 5, 'cacheReadInputTokens' => 40, 'cacheWriteInputTokens' => 10 @@ -29,6 +31,30 @@ expect(message.cache_creation_tokens).to eq(10) end + it 'does not subtract cache buckets or floor to zero when the cached prefix exceeds fresh input' do + response_body = { + 'modelId' => 'anthropic.claude-sonnet-4-5-20250929-v1:0', + 'output' => { + 'message' => { + 'content' => [{ 'text' => 'Hi!' }] + } + }, + 'usage' => { + 'inputTokens' => 3, + 'outputTokens' => 5, + 'cacheReadInputTokens' => 7714, + 'cacheWriteInputTokens' => 327 + } + } + + response = instance_double(Faraday::Response, body: response_body) + message = described_class.parse_completion_response(response) + + expect(message.input_tokens).to eq(3) + expect(message.cached_tokens).to eq(7714) + expect(message.cache_creation_tokens).to eq(327) + end + it 'preserves raw stopReason as finish_reason' do response_body = { 'modelId' => 'amazon.nova-lite-v1:0',