mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-02 02:11:58 +00:00
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm chat templates consume it, so popping it made reasoning models lose earlier reasoning across tool loops. thinking_blocks is still removed for vLLM compatibility. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(hosted_vllm): forward replayed reasoning_content only when it is a string * test(integration): cover hosted_vllm reasoning_content replay across endpoints * test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell --------- Co-authored-by: mateo <mateo@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: yassin <yassin@berri.ai> |
||
|---|---|---|
| .. | ||
| chat_bridge | ||
| providers | ||
| responses_bridge | ||