Commit graph

52143 commits

Author SHA1 Message Date
berriai-litellm-provider-info-sync[bot]
c5a388a94c
chore(prices): sync Google Gemini prices: 1 model [enrichment failed: Google Gemini, sync failed: AWS Bedrock, 6 held]
gemini/gemini-3.1-pro-preview-customtools: input_cost_per_token_flex, output_cost_per_token_flex, cache_read_input_token_cost_flex
2026-09-16 01:46:17 +00:00
Yuneng Jiang
aa15e9f23f
test(e2e): reject expired cache entries before physical eviction 2026-09-15 18:45:43 -07:00
yassin
15f2e25e8a refactor(proxy): replace configurable model access denied message with a fixed clean client message
Drop the model_access_denied_message setting, its {model} template, the DB
override entry and the Admin UI field. Model access denials now always return
the fixed client message while the allowlist diagnostic is logged at the final
HTTP, realtime and MCP boundaries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:44:36 +00:00
kerry
91c7f75864 docs(fireworks-ai): describe the cache-read fallback without asserting provider billing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:44:30 +00:00
Yassin Kortam
67cb0bc089
Merge pull request #41313 from BerriAI/litellm_model_activity_response_time
feat(ui): show average response time per model in usage model activity
2026-09-15 18:43:48 -07:00
ryan
ab99be9dad test(team): cover team_member_budget propagation and per-member isolation end to end
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:43:21 +00:00
kerry
091a38cce6 test(anthropic): import Final for the annotated locals
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:43:08 +00:00
ryan
9afac68995 test(router): cover _register_router_selector and _replace_routing_groups directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:43:04 +00:00
kerry
c4c96180e3 fix(gemini): strip version suffix from modelVersion and keep it on blocked streams
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:42:36 +00:00
kerry
8352045f32 style(fireworks-ai): apply repository conventions to the cost component change
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:39:24 +00:00
Joshua Valluru
ece1de73b4 Merge remote-tracking branch 'origin/main' into litellm_fix_mcp_jwt_oauth_persistence 2026-09-15 18:36:02 -07:00
Devin AI
8d625d0400 fix(xai): keep 'instructions' on the xAI Responses API so system messages survive web_search bridging
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:35:13 +00:00
kerry
10aef224e7 style(anthropic): apply repository conventions to the missing-usage change
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:34:28 +00:00
yucheng
3c972cb31f fix(proxy): apply source overrides to IPv4-mapped IPv6 sign-in sources
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:33:54 +00:00
kerry
91c8d1cdc1 chore: merge origin/main into litellm-providers/price-sync
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:30:24 +00:00
kerry
d1aa37b0a8 merge main into litellm_fix_fireworks_cost_components
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:28:29 +00:00
Joshua Valluru
8a3add3c6a fix(mcp): reject scheme-only Basic credentials 2026-09-15 18:27:25 -07:00
Yuneng Jiang
45d5e6b833
fix(e2e): start cache CI service and count bypass calls 2026-09-15 18:26:44 -07:00
yassin
d90e7b3aec fix(team): resolve model aliases in team admin model_max_budget authority check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:26:40 +00:00
ryan
9f990c4f86 fix(router): validate routing_groups at save time and keep invalid DB groups from blocking SSO load
Overlapping routing_groups persisted from the Admin UI raised inside
Router._init_routing_groups during the DB config reconcile, which skipped
loading SSO, guardrails and the other DB-backed settings while leaving the
proxy healthy. /config/update now returns 400 for overlapping models,
duplicate names, the reserved default name and unknown strategies before
writing, the Router builds every group selector before replacing its state
so a rejected update keeps the previous groups routing, and the proxy applies
routing_groups separately from the other router settings so an already
persisted invalid value is logged and skipped instead of aborting the
reconcile. The Admin UI modal blocks picking a model another group owns.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:25:55 +00:00
kerry-berri
171888b716
Merge pull request #41298 from BerriAI/litellm_drop_remaining_vendor_fact_pins
test: drop remaining tests that pin cost-map vendor facts
2026-09-15 18:25:14 -07:00
kerry
a6d2332f7a test(responses): drive the usage-estimate failure path without patching litellm
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:24:35 +00:00
kerry
5a9fe56aff fix(responses): count custom-tool and MCP argument deltas in the streamed usage fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:24:05 +00:00
yucheng-berri
a9a654cb60
Merge pull request #38241 from BerriAI/litellm_agent365_mcp_guardrail
feat(guardrails): add Microsoft Agent 365 MCP tool-call guardrail
2026-09-15 18:21:46 -07:00
Yuneng Jiang
f214555068
fix(e2e): expect models filters to persist after reload 2026-09-15 18:21:43 -07:00
ryan
5b5bbac769 fix(team): link new members to the shared team member budget so /team/update applies to them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:16:21 +00:00
kerry
44a6d19889 test(gemini): drop review narration from the wrapper test docstring
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:14:37 +00:00
Yuneng Jiang
4587dbeae9
feat(e2e): cache exact provider responses for 24 hours 2026-09-15 18:14:18 -07:00
kerry
1484fd7600 fix(responses): keep the usage estimate best-effort when token counting raises
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:13:26 +00:00
yuneng-jiang
113a416619
Merge pull request #41319 from BerriAI/litellm_fix_ui_onboarding
fix(e2e): onboard dashboard users through invitations
2026-09-15 18:06:52 -07:00
yassin
d8ef940232 chore(ui): regenerate schema.d.ts for organization_id on archived key records
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:06:36 +00:00
Joshua Valluru
13553473aa fix(mcp): reject missing upstream authentication credentials 2026-09-15 18:06:00 -07:00
yucheng
f66ffc387e fix(guardrails): only honor the judge call-origin stamp on logging_only in llm_as_a_judge
On pre_call, during_call and post_call the hook data is the client request body, so a
client-supplied litellm_params.metadata.internal_call_origin must not skip enforcement.
Type the during_call helper's request payload as dict[str, object]

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
kerry
801f6a92b4 test(gemini): assert the served modelVersion reaches the assembled stream through CustomStreamWrapper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
dd4829398a fix(guardrails): log the configured mode when logging_only is mixed with an enforcing llm_as_a_judge mode
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
10506ab904 fix(guardrails): judge the whole latest user turn, run every during_call guardrail, log combined modes as configured
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
1b8b17cc8a fix(guardrails): judge only the latest request turn in pre_call and during_call llm_as_a_judge
The request-side prompt told the judge to focus on the most recent user turn but the text under review was every extracted request message joined together, so a multi-turn request with an off-topic earlier turn and an on-topic latest turn scored 50 and was blocked. Request-side judging now evaluates the last extracted request text (after the configured message scoping) and passes the full role-labelled conversation only as context. Response-side judging still evaluates all extracted response text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
e9357d9a6f fix(guardrails): stop logging_only llm_as_a_judge from judging its own judge calls
Judge sub-calls now carry the internal_call_origin metadata stamp and the
guardrail skips any logged call bearing it, so a logging_only judge no longer
recurses into an unbounded chain of judge requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
a658f20005 test(guardrails): grant premium_user for the tagged Mode case in llm_as_a_judge mode-shape test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
4d1330ea0b test(guardrails): drop callback-manager patch from llm_as_a_judge mode-shape test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
28bd00a004 fix(guardrails): label llm_as_a_judge logging_only verdicts with their mode
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
b93f22265c fix(guardrails): give the request-side judge role-labelled conversation context
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
6fccf00d93 fix(guardrails): keep list and tagged mode shapes for llm_as_a_judge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:52 +00:00
yucheng
0b34800501 fix(guardrails): resolve llm_as_a_judge request-side log mode from event_hook lists, drop unused alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:52 +00:00
yucheng
77a6327675 feat(guardrails): support pre_call and during_call modes for llm_as_a_judge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:52 +00:00
mateo-berri
5f51937107 Merge remote-tracking branch 'origin/main' into litellm_azure_ai_responses_native
Resolve the azure_ai common_utils conflict onto main's api_key_header_for_base
helper, gate native Responses routing on the resolved api_base host, derive the
/openai/v1/responses URL from the host and project prefix, keep websocket mode
on the managed emulation, and cover the routing end to end with respx
2026-09-15 18:04:36 -07:00
yassin
3e8566d878 fix(keys): keep organization_id on archived key records and /key/info
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:03:40 +00:00
kerry
de9aa48cd6 fix(responses): count multimodal input and tool-call output in the streamed usage fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:02:29 +00:00
Tin Chi Lo
5aa2adbacd fix(proxy): preserve Anthropic pricing modifiers in router savings 2026-09-15 18:02:05 -07:00
mateo-berri
3c15f64fd4 fix(bedrock): make prompt caching work on the Nova InvokeModel route
Nova InvokeModel rejects the standalone cachePoint blocks the shared Converse transform emits, so each one is folded into the block it caches and tool_config injection points are dropped before the transform runs, since this route has no tool caching to credit. Usage reads Bedrock's Count-suffixed cache keys and adds cached tokens into prompt_tokens, streaming routes every wrapped InvokeModel event through the Converse chunk parser and tolerates the missing totalTokens, and the Nova 1 cost-map entries gain cache_read_input_token_cost at a quarter of the input rate
2026-09-15 18:01:21 -07:00