mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-03 02:22:24 +00:00
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking
Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.
Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.
(cherry picked from commit
|
||
|---|---|---|
| .. | ||
| __init__.py | ||
| test_function_call_output_normalization.py | ||
| test_handler.py | ||
| test_image_generation_output.py | ||
| test_litellm_completion_responses.py | ||
| test_reasoning_input_item_preservation.py | ||
| test_session_handler.py | ||
| test_session_handler_with_cold_storage.py | ||
| test_streaming_iterator_transformation.py | ||
| test_tool_output_order_preserved_for_gemini.py | ||