Commit graph

51406 commits

Author SHA1 Message Date
yassin
260ff5f491 feat(team): team-level model_max_budget with key-level overrides
A team can now carry a per-model budget map that every key on the team
inherits. A key's own model_max_budget entry for the same model takes
precedence, so it is gated on and billed to the key alone.

Backend: NewTeamRequest/UpdateTeamRequest accept model_max_budget (validated
like the key-level field, enterprise gated); the value is hydrated onto
UserAPIKeyAuth via the token view, TeamGrants and the carried budget state;
_check_team_model_budget enforces it in the centralized common checks; the
limiter meters spend under team_model_spend:<team>:<model>:<duration> and
skips the team counter when the key overrides; /team/update lets only a
proxy admin raise, re-window or drop a cap; /team/info exposes usage.
The Anthropic context-management compaction summary subrequest runs the
same team gate. Both fallback token-view SQL definitions project the column.

UI: team create and edit forms reuse the key-level ModelMaxBudgetEditor,
premium gated, sending {} to clear and omitting unchanged fields.

A key entry overrides the team cap only when it spend-gates the model
(non-negative max_budget); a row that only carries tpm/rpm limits or a
negative cap leaves the team cap in force.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:40:58 +00:00
Devin AI
232233f654 refactor(fireworks_ai): extract reasoning_effort mapping into helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:39:11 +00:00
Devin AI
a7a61db78d fix(fireworks_ai): flatten dict-form reasoning_effort to its effort string
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:37:55 +00:00
tin-berri
878716f806
Merge pull request #41326 from BerriAI/litellm_forecast_classifier_entitlement
feat(router): limit unlicensed Capability and Fuse v2 routers to one each
2026-09-15 17:36:22 -07:00
kerry
d5a36eb2ca fix(gemini): propagate provider modelVersion onto model responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:36:07 +00:00
yassin
9541b0734b refactor(keys): drop status helper docstrings and test /key/list status through the endpoint
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:35:45 +00:00
kerry
ba6bcd747f chore: merge origin/main into test cleanup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:33:25 +00:00
kerry
f5c1c82f81 fix(responses): estimate usage from text when streamed completed event omits usage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:32:24 +00:00
Mateo Wang
aa710dc6a9
Merge pull request #35388 from BerriAI/litellm_session_total_duration
fix(spend): sum multi-round session duration in logs UI
2026-09-15 17:31:41 -07:00
Joshua Valluru
61e3b5ddae fix(mcp): persist OAuth credentials for rowless JWT admins 2026-09-15 17:31:07 -07:00
kerry-berri
daa98cd6ff
Merge pull request #41320 from BerriAI/litellm_gemini_fallback_generalizations
feat(model_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization
2026-09-15 17:28:58 -07:00
Yassin Kortam
7eeba69016
Merge pull request #41316 from BerriAI/litellm_nvidia_nim_infer_passthrough
feat(proxy): add /nvidia_nim passthrough route for NIM object detection and OCR /v1/infer
2026-09-15 17:28:25 -07:00
kerry
4657d43fb9 fix(anthropic): tolerate message_delta chunks without usage in streams
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:27:09 +00:00
mateo
e62ab28561 chore(codeowners): add ryan and kerry as owners of the cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:26:34 +00:00
Yassin Kortam
5071d96d16
Merge pull request #41255 from BerriAI/litellm_org_member_spend_tracking
fix(proxy): track per-member organization spend
2026-09-15 17:25:40 -07:00
Yassin Kortam
e4a7d2aa0b
Merge pull request #41302 from BerriAI/litellm_fix_key_model_rpm_override_precedence
fix(proxy): key model rpm/tpm override takes precedence over team model limit
2026-09-15 17:24:38 -07:00
kerry-berri
95c0c39b05
Merge pull request #41318 from BerriAI/litellm_ws_responses_service_tier_pricing
fix(cost): price native Responses WebSocket turns at their returned service_tier
2026-09-15 17:23:35 -07:00
yassin
241b177f05 fix(proxy): log the internal model access denial reason for MCP sampling denials
MCP sampling catches the denial itself and returns ErrorData, so the central ProxyException handler never sees it. Log the sanitized internal reason there and share the CR/LF stripping through ModelAccessDeniedProxyException.sanitized_internal_message

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:22:48 +00:00
Yassin Kortam
3413efcd7f
Merge pull request #41306 from BerriAI/litellm_lit2435_litellm_log_error_silences_info
fix(proxy): honor LITELLM_LOG for uvicorn and proxy extras loggers
2026-09-15 17:22:39 -07:00
Yassin Kortam
636313e9fc
Merge pull request #41322 from BerriAI/litellm_gcp_video_usage
fix(vertex_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough
2026-09-15 17:21:15 -07:00
Joshua Valluru
72a4824acb fix(mcp): preserve JWT agent validation after main merge 2026-09-15 17:20:06 -07:00
ryan-crabbe-berri
f34a6eda92
Merge pull request #41294 from BerriAI/litellm_usage_export_gating
fix(ui): block usage export and flag the range when a spend page fails
2026-09-15 17:19:38 -07:00
yucheng
7c2c59da90 test(proxy): explain the internal patches in the login throttle tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:13:51 +00:00
Joshua Valluru
53318796fd fix(mcp): separate JWT identity lookup from request authorization 2026-09-15 17:13:10 -07:00
yassin
b7af51cc4a Merge remote-tracking branch 'origin/main' into litellm_model_access_denied_message 2026-09-16 00:12:53 +00:00
yassin
f60a603519 fix(proxy): log configured model access denials at the final response boundary
Post-auth denials from can_key_call_resolved_model (per-request alias
rewrite, MCP sampling, realtime) never reach the auth exception handler,
so the internal allowlist reason was dropped when
model_access_denied_message was set. Log it once from the ProxyException
response handler and the realtime rejection path instead, and convert
JWT ModelAccessDeniedHTTPException into the specialized ProxyException so
the same boundary covers it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:12:49 +00:00
yassin
f25272bfa0 fix(gateway): expose /nvidia_nim/ on the gateway data-plane allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:05:17 +00:00
yassin
65f9d9bbbd test(proxy): use the module-level HTTPException import in the /nvidia_nim route tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:57:16 +00:00
yassin
251aeea97d fix(ui): show the S3 label when editing the s3_v2 callback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:56:31 +00:00
yucheng
aa7f1e16b8 feat(proxy): throttle failed Admin UI sign-ins per source and source/username
Replace the username-global lockout with counters keyed by source address and by
source/username pair. Each has a fixed counting window (60s) and a separate
block TTL (300s). Blocks are soft: a correct password still signs in, wrong
passwords from a blocked key take one of 5 held slots per worker and are held
30s before a 429. Once a pair is blocked its failures stop counting against the
source. The source scope runs only when trusted_proxy_ranges is set, IPv6 is
grouped by /64, and per-source limits accept IP and CIDR overrides with
longest-prefix matching. Redis is authoritative through one Lua script per
failure, with bounded per-worker fallback when Redis raises.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:53:14 +00:00
kerry
c9dd4b44f8 revert(model_info): keep Gemini baseline majors single-digit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:51:56 +00:00
Yassin Kortam
24153b5f29
Merge pull request #41308 from BerriAI/litellm_resolve_model_group_alias_before_auth 2026-09-15 16:51:28 -07:00
Joshua Valluru
88d0371a46 fix(mcp): reuse the standard JWT auth builder for OAuth ownership 2026-09-15 16:50:55 -07:00
yucheng
5c0757e990 Merge remote-tracking branch 'origin/main' into litellm_lit5285_login_rate_limit_v2 2026-09-15 23:50:43 +00:00
Yassin Kortam
474563a4ba
Merge pull request #41303 from BerriAI/litellm_passthrough_auth_false_db_overlay 2026-09-15 16:50:41 -07:00
kerry
945adfb603 refactor(model_info): support multi-digit Gemini majors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:49:18 +00:00
yucheng-berri
c1a3784e6c
Merge pull request #41205 from BerriAI/litellm_lit_5856_error_log_call_id
fix(proxy): include litellm_call_id in LLM API exception logs
2026-09-15 16:47:57 -07:00
yassin
ad8de0e192 fix(proxy): roll up only closed days into LiteLLM_DailyGlobalSpend and split the key-free read at the marker
The write path no longer dual-writes the global table. The cron rolls up closed UTC days
only, so a pod still flushing the current day can never leave the global table short. The
key-free arm reads days through the marker from the global table and later days from
LiteLLM_DailyUserSpend in one UNION ALL, and the marker comes from the config cache
rather than a per-request database lookup.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:42:19 +00:00
yassin
6e7de5fd20 test(s3): type the prompts-only test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:40:35 +00:00
yassin
7267c6bed7 Merge remote-tracking branch 'origin/main' into litellm_model_access_denied_message 2026-09-15 23:36:30 +00:00
kerry
5b54bf2328 refactor(model_info): restore provider-neutral Gemini chat baseline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:36:22 +00:00
yassin
b99f0c812f fix(proxy): log configured model access denial only at the auth error boundary
Move the WARNING that carries the internal denial detail out of the message formatter and into the auth exception handler. The denial exceptions now carry internal_message so access-group probes and fallback paths that catch and recover from the denial no longer log a false denial for an allowed request

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:36:12 +00:00
yassin
74a2c2eab1 refactor(vertex_ai): tighten Interactions route match and type the usage boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:34:31 +00:00
yassin
771b2509d1 fix(proxy): reject mixed NIM model groups and strip the deployment model before the group in /nvidia_nim URLs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:34:22 +00:00
kerry
99bf8e9b2f test(e2e): add cost calculation CI proxy config
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:32:39 +00:00
berriai-litellm-provider-info-sync[bot]
083ecb3c0b
chore(prices): sync prices for 3 providers: 26 models, 26 deprecated [enrichment failed: Google Gemini, sync failed: AWS Bedrock, 34 held]
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: deprecation_date
fireworks_ai/deepseek-v4-pro: deprecation_date
fireworks_ai/accounts/fireworks/models/minimax-m2p7: deprecation_date
fireworks_ai/minimax-m2p7: deprecation_date
together_ai/deepseek-ai/deepseek-coder-33b-instruct: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Llama-70B: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B: deprecation_date
together_ai/google/gemma-2-27b-it: deprecation_date
vertex_ai/imagegeneration@006: deprecation_date
vertex_ai/imagen-3.0-capability-001: deprecation_date
vertex_ai/imagen-3.0-fast-generate-001: deprecation_date
vertex_ai/imagen-3.0-generate-001: deprecation_date
vertex_ai/imagen-3.0-generate-002: deprecation_date
vertex_ai/imagen-4.0-fast-generate-001: deprecation_date
vertex_ai/imagen-4.0-generate-001: deprecation_date
vertex_ai/imagen-4.0-ultra-generate-001: deprecation_date
together_ai/meta-llama/Llama-3-8b-chat-hf: deprecation_date
together_ai/meta-llama/Meta-Llama-3-70B-Instruct-Turbo: deprecation_date
together_ai/meta-llama/Meta-Llama-3-8B-Instruct: deprecation_date
together_ai/NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO: deprecation_date
together_ai/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF: deprecation_date
together_ai/Qwen/Qwen2-72B-Instruct: deprecation_date
together_ai/Qwen/Qwen2-VL-72B-Instruct: deprecation_date
together_ai/Qwen/Qwen2.5-Coder-32B-Instruct: deprecation_date
together_ai/Qwen/Qwen2.5-VL-72B-Instruct: deprecation_date
2026-09-15 23:31:18 +00:00
yucheng
7c24c8bc1d Merge remote-tracking branch 'origin/main' into litellm_lit5285_login_rate_limit_v2 2026-09-15 23:31:07 +00:00
kerry
269afbe382 test(e2e): apply review nits to cost calculation suite
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:28:44 +00:00
yassin
d61908e241 refactor(proxy): type request_data on the alias rewrite helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:28:01 +00:00
mateo-berri
790c9ab77d test(ui): name the session column fixtures so the inline-object lint budget holds
Some checks failed
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
2026-09-15 16:25:01 -07:00