litellm/tests/test_litellm/llms/hosted_vllm
Chesars e382351d4a feat(hosted_vllm): support thinking parameter for /v1/messages endpoint
Adds support for Anthropic-style 'thinking' parameter in hosted_vllm,
converting it to OpenAI-style 'reasoning_effort' since vLLM is
OpenAI-compatible.

This enables users to use Claude Code CLI with hosted vLLM models
like GLM-4.6/4.7 through the /v1/messages endpoint.

Mapping (same as Anthropic adapter):
- budget_tokens >= 10000 -> "high"
- budget_tokens >= 5000  -> "medium"
- budget_tokens >= 2000  -> "low"
- budget_tokens < 2000   -> "minimal"

Fixes #19761
2026-01-26 13:49:55 -03:00
..
chat feat(hosted_vllm): support thinking parameter for /v1/messages endpoint 2026-01-26 13:49:55 -03:00
test_hosted_vllm_rerank_transformation.py fix: fix vllm test 2025-09-27 10:01:48 -07:00