litellm/tests/integration/messages_endpoint
devin-ai-integration[bot] 657bb777fa
fix(hosted_vllm): keep reasoning_content on replayed assistant messages (#43599)
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages

vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm
chat templates consume it, so popping it made reasoning models lose
earlier reasoning across tool loops. thinking_blocks is still removed
for vLLM compatibility.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(hosted_vllm): forward replayed reasoning_content only when it is a string

* test(integration): cover hosted_vllm reasoning_content replay across endpoints

* test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-09-30 13:12:28 -07:00
..
chat_bridge fix(hosted_vllm): keep reasoning_content on replayed assistant messages (#43599) 2026-09-30 13:12:28 -07:00
providers fix(router): stream /v1/messages lifecycle frames live when no fallback can take over (#43600) 2026-09-29 09:22:05 -07:00
responses_bridge test(integration): group /v1/messages contracts under tests/integration/messages_endpoint (#43352) 2026-09-26 16:00:32 -07:00