fix(ollama): extract thinking field from /api/generate response

Ollama's /api/generate endpoint returns a top-level 'thinking' field
for reasoning models (Qwen3, DeepSeek-R1) but the non-streaming
transform_response() only parsed <think> XML tags from the response
text. This caused reasoning_content to always be null.

Check for the 'thinking' field first, matching the pattern already used
in the /api/chat path and the streaming iterator.

Fixes #27956
This commit is contained in:
sharziki 2026-05-14 21:06:32 -04:00
parent e58a561caa
commit 0a7c50c238

View file

@ -370,7 +370,12 @@ class OllamaConfig(BaseConfig):
response_text = response_json.get("response", "")
content = None
reasoning_content = None
if response_text is not None and isinstance(response_text, str):
# Check for top-level "thinking" field (Ollama returns this for
# reasoning models like Qwen3/DeepSeek-R1 via /api/generate)
if "thinking" in response_json:
reasoning_content = response_json["thinking"]
content = response_text
elif response_text is not None and isinstance(response_text, str):
reasoning_content, content = _parse_content_for_reasoning(response_text)
else:
content = response_text # type: ignore