fix: prompt cache misses after a model calls tools one after another (#31593)

When a model called one tool, got its result and then called a second tool with no text in between, the next request merged both calls into an earlier assistant message that the provider had already seen. That changed the conversation's beginning, so the provider's prompt cache stopped matching from there for the rest of the chat. Each tool call and its result are now sent as their own messages, so the start of the conversation stays identical from one request to the next.

Fixes #31588
This commit is contained in:
Classic298 2026-10-01 05:34:59 +02:00 • committed by GitHub
parent 1c233f1be6
commit e88e1f8119
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -479,7 +479,7 @@ def convert_output_to_messages(
for item in output:
item_type = item.get('type', '')
if item_type not in {'function_call', 'function_call_output'}:
if item_type != 'function_call_output':
flush_tool_outputs()
flush_tool_images()