When litellm converts Responses API function_call input items to Chat
Completion messages, each function_call item becomes a separate assistant
message with a single tool_calls entry. For providers like Gemini that
issue parallel tool calls in one model turn, this produces multiple
assistant messages instead of one with multiple tool_calls.
The Gemini message converter later walks the history looking for the
'last assistant message with tool calls' to match a tool result. With
split messages, it only finds the last call's assistant message and
fails to match earlier calls' tool results, producing:
Missing corresponding tool call for tool response message.
Received - message={'role': 'tool', 'content': '...', 'tool_call_id': 'call_xxx'}
Fix: merge consecutive function_call-derived assistant messages into a
single assistant message with multiple tool_calls entries, re-indexing
as needed. This preserves the one-turn-one-message invariant that
downstream converters expect.
Affects: Gemini 3 models with parallel tool calls via the Responses API
(litellm.aresponses / /v1/responses endpoint).
- Fix Atlassian config: use http transport and correct URL (/v1/mcp not /v1/sse)
- Fix URL pattern: use standard /mcp/<server_name> not legacy /<server_name>/mcp
- Add parameter breakdown table explaining each claude mcp add argument
- Add warning that server name in proxy config must match URL path
- Add ngrok step for OAuth callback accessibility
- Add ~/.claude.json config option alongside claude mcp add
- Fix auth header guidance: use x-litellm-api-key for OAuth servers
The test-complete aggregate job adds no value as GitHub Actions
already provides visibility into matrix job results.
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
poetry.lock was out of sync with pyproject.toml, causing CI failures
across all PRs with "Run `poetry lock` to fix the lock file" error.
Regenerated the lock file to resolve the sync issue.
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
* feat(bedrock): add DeepSeek V3.2 pricing and region support
* feat(bedrock): add minimax.minimax-m2.1 pricing and region support
* feat(bedrock): add moonshotai.kimi-k2.5 pricing and region support
* feat(bedrock): add qwen.qwen3-coder-next
pricing and region support
* add some sanity unit tests for the bedrock beta models added
* --amend
* resolve greptileai comments and suggestions
* docs: add reference to example_openai_endpoint repo for self-hosting fake OpenAI proxy (#21006)
- Updated benchmarks.md with a section on setting up fake OpenAI endpoints
- Updated load_test.md to mention the self-hosted option
- Updated load_test_advanced.md with a tip box about the example repo
Reference: https://github.com/BerriAI/example_openai_endpoint
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* MCP fixes
* fix(oldteams.tsx): show policies when creating
* fix(proxy/_types.py): ensure mcp rest endpoints can be called by virtual key
ensures UI works with virtual key testing mcp endpoints
* refactor: migrate get object permissions table logic to happen in user api key auth - allows functions to trust user api key object they receive has what they need
* fix(rest_endpoints.py): filter for allowed tools based on what key has access to
* fix(mcp_server_manager.py): ensure only allowed MCP's are returned to the user, via rest endpoints
* Guardrails - add toxic/abusive content filter guardrails
* fix(streaming): preserve usage data from post-finish_reason chunks in OpenAI-compatible streaming
Fixes#16112
OpenRouter and other OpenAI-compatible providers send a usage chunk after
the finish_reason='stop' chunk when stream_options.include_usage is True.
The OpenAIChatCompletionStreamingHandler.chunk_parser() was not passing
the usage field to ModelResponseStream, causing real token counts from the
provider to be lost and falling back to inaccurate estimates.
* fix: resolve merge conflict in test file
- Fix typo in test method name (extra space)
- Move test_prompt_cache_key_in_optional_params to its own class
---------
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* fix(vertex_ai): forward extra_body to completion transformation handler
The responses() function accepted extra_body as a named parameter but
did not pass it to response_api_handler when responses_api_provider_config
was None (completion transformation path), silently dropping it.
Also adds deep-merge support for extra_body in Vertex AI Gemini
transformation, so dict values like generationConfig are merged rather
than replaced.
* refactor(vertex_ai): extract _merge_extra_body to fix PLR0915 lint
Move the extra_body merge loop into a helper function to keep
_transform_request_body under the 50-statement limit.