Commit graph

34478 commits

Author SHA1 Message Date
michelligabriele
4630793fb0 fix(websearch_interception): preserve thinking blocks in agentic loop follow-up messages
When extended thinking is enabled, the websearch interception agentic loop
builds a follow-up assistant message with only tool_use blocks. Anthropic's
API requires assistant messages to start with thinking/redacted_thinking
blocks when thinking is enabled, causing a 400 Bad Request.

Extract thinking blocks from the model's initial response, thread them
through the agentic loop, and prepend them to the follow-up assistant
message — matching the pattern used by anthropic_messages_pt in factory.py.

Fixes the error: "Expected 'thinking' or 'redacted_thinking', but found
'tool_use'"
2026-02-19 21:51:00 +01:00
jtsaw
093a67f774
Update litellm/llms/anthropic/chat/transformation.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-19 12:48:37 -08:00
Ishaan Jaff
18f8a2cee3
docs: add latency overhead troubleshooting guide (#21603)
* add latency overhead troubleshooting doc

* add latency_overhead to troubleshooting sidebar

* docs: add x-litellm-overhead-duration-ms to latency troubleshooting guide
2026-02-19 12:42:33 -08:00
Chesars
3fe331ed7d fix(gemini): preserve $ref in JSON Schema for Gemini 2.0+ to avoid nesting depth errors 2026-02-19 17:37:53 -03:00
Ishaan Jaff
2c8fcf854a
docs: add latency overhead troubleshooting guide (#21600)
* add latency overhead troubleshooting doc

* add latency_overhead to troubleshooting sidebar
2026-02-19 12:34:23 -08:00
jtsaw
063563608d fix lint 2026-02-19 12:32:56 -08:00
Julio Quinteros Pro
ab381a61f5 fix(lint): remove redundant explicit router import in policy_endpoints __init__
The `router` name is already re-exported via `from .endpoints import *`,
making the explicit `from .endpoints import (router,)` on the following
lines redundant and triggering ruff F401 (imported but unused).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-19 17:31:58 -03:00
Ishaan Jaff
cd95e54c10
Add OpenAPI-to-MCP support via API and UI (#21575)
* add spec_path column to LiteLLM_MCPServerTable schema

* add spec_path to MCP request types and table model

* wire spec_path through build_mcp_server_from_table

* add openapi transport type constant

* add OpenAPI Spec as first-class transport option in create form

* add OpenAPI transport support to edit form with auto-detection

* support spec_path in connection status component

* support spec_path in tool configuration component

* support OpenAPI transport in test connection hook

* register OpenAPI tools on server add/update/reload

* preview OpenAPI tools in test/tools/list endpoint
2026-02-19 12:23:24 -08:00
Ishaan Jaff
e9a07347dc
fix: reduce proxy overhead for large base64 payloads (#21594)
* fix aviation safety topic filter: remove overly broad exceptions, add cockpit access block words

* fix airline brand protection filter: add identifier words, competitor/ops block words, tighten exceptions

* add constants for large payload handling and detailed timing

* add base64 truncation for logging payloads

* use shallow copy for messages, track copy and callback timing

* add callback duration and detailed timing to response metadata

* add callback duration header, size-gate debug logging, detailed timing headers

* add tests for callback timing, base64 truncation, and detailed timing

* fix code quality: extract helpers, fix regex, clean up imports

* rewrite _truncate_base64_in_value iteratively to satisfy recursive detector
2026-02-19 12:10:52 -08:00
Ishaan Jaff
b209b11522
feat: AI policy template suggestions (#21589)
* fix aviation safety topic filter: remove overly broad exceptions, add cockpit access block words

* fix airline brand protection filter: add identifier words, competitor/ops block words, tighten exceptions

* add example_sentences to all policy templates + topic-filtering and prompt-injection templates

* add policy_endpoints package with AI policy suggester

* update test patch targets for policy_endpoints package move

* add unit tests for AI policy suggester

* add suggestPolicyTemplates networking function

* add AI suggestion modal component

* add Use AI button and template loading callback to PolicyTemplates

* wire up AI suggestion modal in policies page

* fix policy_templates_backup.json path after package move

* add estimated_latency field to all policy templates

* use llm_router and accept model parameter in ai_policy_suggester

* add model param to suggest templates endpoint

* pass model param in suggestPolicyTemplates

* polish ai suggestion modal: model selector, auto-growing textareas, latency badges

* add template queue for processing multiple AI-suggested templates

* show template progress badge in guardrail selection modal
2026-02-19 12:00:26 -08:00
jtsaw
b28ec1c75e support reasoning + effort on sonnet 4.6 2026-02-19 11:54:55 -08:00
Chesars
c5cec60fd0 fix(moonshot): preserve image_url blocks in multimodal messages
Moonshot's _transform_messages unconditionally flattened content arrays
to plain text, dropping image_url blocks. Vision models like kimi-k2.5
accept the standard OpenAI content array format.

Now checks for image_url blocks before flattening — if any message
contains images the content array is preserved intact.

Fixes #20862
2026-02-19 16:44:38 -03:00
michelligabriele
16cfdccc7b
fix(bedrock): add Accept header for AgentCore MCP server requests (#21551)
AgentCore MCP server endpoints require the Accept header to contain
both application/json and text/event-stream per the MCP specification
(Streamable HTTP transport). Without this header, requests are rejected
with a 406 Not Acceptable error (JSON-RPC code -32011).

Sets the Accept header at the top of sign_request() so both JWT/Bearer
and SigV4 authentication paths include it.
2026-02-19 11:35:14 -08:00
Cesar Garcia
3b7236d618
Update litellm/llms/chatgpt/chat/streaming_utils.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-19 16:24:19 -03:00
milan-berri
506d1dff43
fix: handle deprovisioning operations without path field (#21571)
* fix(scim): handle deprovisioning operations without path field

When SCIM providers send deprovisioning requests without a path field
(e.g., {"op": "replace", "value": {"active": false}}), the code was
storing the value under an empty string key in metadata.

This fix:
- Detects operations with no path where value is a dict
- Extracts and handles known fields like 'active' correctly
- Sets metadata["scim_active"] = false instead of metadata[""] = {"active": false}

Fixes: SCIM deprovisioning creating empty string keys in user metadata

* fix(scim): handle all known fields in operations without path

Extended the fix to handle all SCIM fields (not just active) when
operations have no path field:
- active -> scim_active
- displayName -> user_alias
- externalId -> user_id
- name.givenName/familyName -> scim_metadata

Added comprehensive test for multiple fields without path.

Addresses Greptile review feedback on RFC 7644 compliance.

* trigger PR update
2026-02-19 11:09:04 -08:00
Chesars
27413790e6 fix(openrouter): use provider-reported usage in streaming without stream_options
When providers like OpenRouter send a usage chunk after the finish_reason
chunk, _hidden_params["usage"] was already calculated (with zeros) before
the usage data arrived. The StopIteration handler now recalculates usage
from stream_chunk_builder and updates the shared _hidden_params dict so
the user's copy reflects the real provider-reported token counts.

Fixes #20760
2026-02-19 16:07:30 -03:00
Sameer Kankute
ddc1371cea
Merge pull request #21588 from BerriAI/geminii_blog
Fix release
2026-02-20 00:31:01 +05:30
Sameer Kankute
4d392cacb8 Fix release 2026-02-20 00:27:12 +05:30
yuneng-jiang
aa4a95c889
Merge pull request #21587 from BerriAI/litellm_fix_key_delete_feb19
[Infra] Add project_id to DeletedVerificationTable
2026-02-19 10:51:28 -08:00
yuneng-jiang
6c0cc4fb4f adding build 2026-02-19 10:50:47 -08:00
Chesars
3aea9c81c9 fix(openrouter): prevent double-stripping of native model names in get_llm_provider
Move the fix to the OpenRouter level: define native OpenRouter models
(openrouter/auto, openrouter/free, openrouter/bodybuilder) and check
them in get_llm_provider() before the provider_list stripping logic.
This prevents the second strip across all bridges without modifying
each adapter/handler individually.

Fixes #16353
2026-02-19 15:50:39 -03:00
yuneng-jiang
2eded1dda6 bump: version 0.4.43 → 0.4.44 2026-02-19 10:50:23 -08:00
yuneng-jiang
a5fba84ae3
Merge pull request #21586 from BerriAI/infra_feb19
[Infra] Fixing Merge Artifacts
2026-02-19 10:47:39 -08:00
yuneng-jiang
2d928f8c54 fixing merge 2026-02-19 10:46:58 -08:00
Chesars
155bed57e8 fix(vertex_ai): pass through native Gemini imageConfig params for image generation
aspectRatio and imageSize were silently dropped because they weren't
listed in get_supported_openai_params(), so the validation layer filtered
them out before they could reach transform_image_generation_request().

Fixes #21070
2026-02-19 15:39:27 -03:00
yuneng-jiang
bac1b6b2e0
Merge pull request #21545 from BerriAI/litellm_key_last_active_tracking
[Feature] Key Last Active Tracking
2026-02-19 10:29:23 -08:00
yuneng-jiang
caa931b167 adding build 2026-02-19 10:29:01 -08:00
yuneng-jiang
9eaaa9740b bump: version 0.4.42 → 0.4.43 2026-02-19 10:28:35 -08:00
yuneng-jiang
c911cfbabf Merge remote-tracking branch 'origin' into litellm_key_last_active_tracking 2026-02-19 10:27:48 -08:00
Chesars
a185182086 fix(openai): validate logprobs/top_p against reasoning_effort for gpt-5.1/5.2
logprobs, top_p, top_logprobs are only accepted by OpenAI when
reasoning_effort="none". Add validation matching the existing
temperature logic: raise UnsupportedParamsError or drop when
reasoning_effort is set to other values.
2026-02-19 15:20:14 -03:00
Chesars
26d27803eb fix(models): disable function calling for PublicAI Apertus models
The Apertus 8B and 70B models do not support standard OpenAI-style
tool calling. Per Swiss AI's docs, tool use integration into inference
engines is still in development. Set supports_function_calling and
supports_tool_choice to false.

Fixes #21124
2026-02-19 15:17:27 -03:00
Chesars
c6240ff621 fix(azure_ai): resolve api_base from env var in get_complete_url for Document Intelligence OCR
validate_environment() resolved AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT
from the env var but only returned headers. get_complete_url() still
received None and raised ValueError. Now get_complete_url() also
resolves the env var as a fallback.

Fixes #21034
2026-02-19 15:14:09 -03:00
yuneng-jiang
2df7d1770a
Merge pull request #21580 from BerriAI/proxy_extras_feb19
[Infra] bump proxy extras
2026-02-19 10:12:33 -08:00
yuneng-jiang
c3317ea2b3 build files 2026-02-19 10:11:47 -08:00
yuneng-jiang
23547df8a1 bump: version 0.4.41 → 0.4.42 2026-02-19 10:11:19 -08:00
yuneng-jiang
0465067055
Merge pull request #21579 from BerriAI/proxy_extras_feb19
bump: version 0.4.40 → 0.4.41
2026-02-19 10:07:00 -08:00
yuneng-jiang
35797d1f75 bump: version 0.4.40 → 0.4.41 2026-02-19 10:05:38 -08:00
Chesars
5596545fb9 fix(gemini): correct streaming finish_reason for tool calls
Gemini returns finishReason="STOP" even when tool calls are present,
and sends tool_calls and finishReason in separate streaming chunks.
The ModelResponseIterator now tracks tool_calls across chunks and
correctly maps finish_reason to "tool_calls" per the OpenAI spec.

Fixes #21041
2026-02-19 15:01:41 -03:00
Chesars
5c4c085353 fix(openai): correct supported_openai_params for GPT-5 model family
Remove logit_bias, modalities, prediction, audio, web_search_options
from supported params for all GPT-5 reasoning models (OpenAI rejects
them). Add logprobs, top_p, top_logprobs for gpt-5.1/5.2 which support
them when reasoning_effort="none".

Related to #21572
2026-02-19 14:56:49 -03:00
Chesars
f3f731a678 fix(openai): restrict supported params for gpt-5-search models
gpt-5-search-api models were routed through OpenAIGPT5Config which
listed params like n, temperature, tools, reasoning_effort as supported,
but OpenAI rejects all of these for search models.

Fixes #21572
2026-02-19 14:37:43 -03:00
Sameer Kankute
2e34870df9
Merge pull request #21568 from BerriAI/litellm_gemini-3.1-pro-preview_day_0
[Feat]Add gemini 3.1 pro preview day 0 support
2026-02-19 22:25:28 +05:30
jquinter
f20b5fc62a
Merge pull request #21566 from BerriAI/fix/pass-through-mypy-type-errors
fix(types): fix mypy errors in pass-through endpoint query param types
2026-02-19 13:51:39 -03:00
Sameer Kankute
c123dc5c24 Fix vercel build 2026-02-19 22:19:34 +05:30
Sameer Kankute
9f66a4c122 Fix test_reasoning_effort_dict_format_gemini_3 2026-02-19 22:14:20 +05:30
Sameer Kankute
884c763fb1 Fix date in docs 2026-02-19 22:14:20 +05:30
Sameer Kankute
a951d6c681 Update docs/my-website/blog/gemin_3.1/index.md
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-19 22:14:20 +05:30
Sameer Kankute
e27725a8b5 Update docs/my-website/blog/gemin_3.1/index.md
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-19 22:14:20 +05:30
Sameer Kankute
468be6f5a8 Fix date in docs 2026-02-19 22:14:20 +05:30
Sameer Kankute
2133a97e97 Add gemini-3.1-pro-preview pricing data 2026-02-19 22:14:19 +05:30
Sameer Kankute
8305bbee21 Add mapping for medium thinking level for gemini-3.1-pro-preview 2026-02-19 22:14:19 +05:30