mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-09 03:18:44 +00:00
* fix(ollama): read the JSON thinking field on non-streaming completions Ollama's /api/generate returns reasoning in a top-level `thinking` field, but the completion transport only looked for inline <think> tags. reasoning_content was therefore always null, and a model that spent its whole turn reasoning returned an empty assistant message with tokens billed. Port the precedence the ollama_chat transport already uses: the field wins and inline tags stay the fallback. Applied to both non-streaming paths, including the JSON-mode text fallback. The two fields are read through a small validated model rather than off the untyped JSON, so absent and explicitly null `response` stay distinct exactly as before. * fix(ollama): keep the thinking field on JSON-mode completions The first pass read `thinking` for plain replies and for JSON-mode text that failed to parse, but the three JSON-mode branches that succeed still dropped it: an empty `response`, a valid JSON object, and a function-call shaped one. A model that spent its whole turn reasoning under `format: json` therefore still came back blank with the tokens billed. Carry the field on all three, type the new test helper's parameters, and cover the null and malformed `response` fallbacks. --------- Co-authored-by: Pawan-Shahane <shahanepawan511@gmail.com> |
||
|---|---|---|
| .. | ||
| test_ollama_chat_transformation.py | ||
| test_ollama_completion_transformation.py | ||
| test_ollama_embedding.py | ||
| test_ollama_model_info.py | ||