fix(anthropic): omit thinking blocks from Responses API message history replay

Fixes #26916

When `anthropic_messages()` was called with a prior assistant `thinking` block
in the conversation history, the Responses-API adapter was serializing the
thinking content as an `output_text` block. This leaks internal model
reasoning back to the model as if it were ordinary visible output.

Thinking blocks represent private internal reasoning and have no equivalent
history item in the OpenAI Responses API, so they are now silently dropped
during history replay. This matches the behavior of the Anthropic API itself
(which strips `thinking` blocks on replay when extended thinking is enabled).
This commit is contained in:
PRABHU KIRAN VANDRANKI 2026-06-08 13:20:10 -04:00
parent a72414a061
commit 189c530b90

View file

@ -166,11 +166,9 @@ class LiteLLMAnthropicToResponsesAPIAdapter:
}
)
elif btype == "thinking":
thinking_text = block.get("thinking", "")
if thinking_text:
asst_parts.append(
{"type": "output_text", "text": thinking_text}
)
# Thinking blocks are internal model reasoning and must not be
# forwarded as output_text in replayed history (fixes #26916).
pass
if asst_parts:
input_items.append(
{