mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-07 08:26:10 +00:00
For multi-turn conversations, convert thinking_blocks on assistant messages into content blocks prepended before the rest of the content, so reasoning context is passed back to the hosted_vllm API. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| chat | ||
| embedding | ||
| test_hosted_vllm_rerank_transformation.py | ||