Commit graph

50986 commits

Author SHA1 Message Date
yassin
f25272bfa0 fix(gateway): expose /nvidia_nim/ on the gateway data-plane allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:05:17 +00:00
yassin
65f9d9bbbd test(proxy): use the module-level HTTPException import in the /nvidia_nim route tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:57:16 +00:00
yassin
251aeea97d fix(ui): show the S3 label when editing the s3_v2 callback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:56:31 +00:00
yucheng
aa7f1e16b8 feat(proxy): throttle failed Admin UI sign-ins per source and source/username
Replace the username-global lockout with counters keyed by source address and by
source/username pair. Each has a fixed counting window (60s) and a separate
block TTL (300s). Blocks are soft: a correct password still signs in, wrong
passwords from a blocked key take one of 5 held slots per worker and are held
30s before a 429. Once a pair is blocked its failures stop counting against the
source. The source scope runs only when trusted_proxy_ranges is set, IPv6 is
grouped by /64, and per-source limits accept IP and CIDR overrides with
longest-prefix matching. Redis is authoritative through one Lua script per
failure, with bounded per-worker fallback when Redis raises.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:53:14 +00:00
kerry
c9dd4b44f8 revert(model_info): keep Gemini baseline majors single-digit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:51:56 +00:00
Yassin Kortam
24153b5f29
Merge pull request #41308 from BerriAI/litellm_resolve_model_group_alias_before_auth 2026-09-15 16:51:28 -07:00
Joshua Valluru
88d0371a46 fix(mcp): reuse the standard JWT auth builder for OAuth ownership 2026-09-15 16:50:55 -07:00
yucheng
5c0757e990 Merge remote-tracking branch 'origin/main' into litellm_lit5285_login_rate_limit_v2 2026-09-15 23:50:43 +00:00
Yassin Kortam
474563a4ba
Merge pull request #41303 from BerriAI/litellm_passthrough_auth_false_db_overlay 2026-09-15 16:50:41 -07:00
kerry
945adfb603 refactor(model_info): support multi-digit Gemini majors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:49:18 +00:00
yucheng-berri
c1a3784e6c
Merge pull request #41205 from BerriAI/litellm_lit_5856_error_log_call_id
fix(proxy): include litellm_call_id in LLM API exception logs
2026-09-15 16:47:57 -07:00
yassin
ad8de0e192 fix(proxy): roll up only closed days into LiteLLM_DailyGlobalSpend and split the key-free read at the marker
The write path no longer dual-writes the global table. The cron rolls up closed UTC days
only, so a pod still flushing the current day can never leave the global table short. The
key-free arm reads days through the marker from the global table and later days from
LiteLLM_DailyUserSpend in one UNION ALL, and the marker comes from the config cache
rather than a per-request database lookup.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:42:19 +00:00
yassin
6e7de5fd20 test(s3): type the prompts-only test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:40:35 +00:00
yassin
7267c6bed7 Merge remote-tracking branch 'origin/main' into litellm_model_access_denied_message 2026-09-15 23:36:30 +00:00
kerry
5b54bf2328 refactor(model_info): restore provider-neutral Gemini chat baseline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:36:22 +00:00
yassin
b99f0c812f fix(proxy): log configured model access denial only at the auth error boundary
Move the WARNING that carries the internal denial detail out of the message formatter and into the auth exception handler. The denial exceptions now carry internal_message so access-group probes and fallback paths that catch and recover from the denial no longer log a false denial for an allowed request

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:36:12 +00:00
yassin
74a2c2eab1 refactor(vertex_ai): tighten Interactions route match and type the usage boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:34:31 +00:00
yassin
771b2509d1 fix(proxy): reject mixed NIM model groups and strip the deployment model before the group in /nvidia_nim URLs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:34:22 +00:00
berriai-litellm-provider-info-sync[bot]
083ecb3c0b
chore(prices): sync prices for 3 providers: 26 models, 26 deprecated [enrichment failed: Google Gemini, sync failed: AWS Bedrock, 34 held]
fireworks_ai/accounts/fireworks/models/deepseek-v4-pro: deprecation_date
fireworks_ai/deepseek-v4-pro: deprecation_date
fireworks_ai/accounts/fireworks/models/minimax-m2p7: deprecation_date
fireworks_ai/minimax-m2p7: deprecation_date
together_ai/deepseek-ai/deepseek-coder-33b-instruct: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Llama-70B: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B: deprecation_date
together_ai/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B: deprecation_date
together_ai/google/gemma-2-27b-it: deprecation_date
vertex_ai/imagegeneration@006: deprecation_date
vertex_ai/imagen-3.0-capability-001: deprecation_date
vertex_ai/imagen-3.0-fast-generate-001: deprecation_date
vertex_ai/imagen-3.0-generate-001: deprecation_date
vertex_ai/imagen-3.0-generate-002: deprecation_date
vertex_ai/imagen-4.0-fast-generate-001: deprecation_date
vertex_ai/imagen-4.0-generate-001: deprecation_date
vertex_ai/imagen-4.0-ultra-generate-001: deprecation_date
together_ai/meta-llama/Llama-3-8b-chat-hf: deprecation_date
together_ai/meta-llama/Meta-Llama-3-70B-Instruct-Turbo: deprecation_date
together_ai/meta-llama/Meta-Llama-3-8B-Instruct: deprecation_date
together_ai/NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO: deprecation_date
together_ai/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF: deprecation_date
together_ai/Qwen/Qwen2-72B-Instruct: deprecation_date
together_ai/Qwen/Qwen2-VL-72B-Instruct: deprecation_date
together_ai/Qwen/Qwen2.5-Coder-32B-Instruct: deprecation_date
together_ai/Qwen/Qwen2.5-VL-72B-Instruct: deprecation_date
2026-09-15 23:31:18 +00:00
yucheng
7c24c8bc1d Merge remote-tracking branch 'origin/main' into litellm_lit5285_login_rate_limit_v2 2026-09-15 23:31:07 +00:00
yassin
d61908e241 refactor(proxy): type request_data on the alias rewrite helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:28:01 +00:00
mateo-berri
790c9ab77d test(ui): name the session column fixtures so the inline-object lint budget holds
Some checks failed
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
2026-09-15 16:25:01 -07:00
ryan-crabbe-berri
69c3212220 fix(ui): gate the usage export on range coverage, not on a fetch being in flight
A loading flag only flips once the fetch effect runs, so the render right after a
date or filter change still reported the previous range as loaded and let an export
read its rows. Stamp the completed range on the hook and compare it during render
instead, the way the tiles already do.

Also stop the failure banner claiming a page loaded when the first request is what
failed, which left it reading 1/1.
2026-09-15 16:21:37 -07:00
mateo-berri
3441f73118 Merge remote-tracking branch 'origin/main' into litellm_session_total_duration
# Conflicts:
#	litellm/proxy/spend_tracking/spend_management_endpoints.py
#	tests/test_litellm/proxy/spend_tracking/test_spend_management_endpoints.py
#	ui/litellm-dashboard/src/components/view_logs/RequestLogsTableColumns.tsx
#	ui/litellm-dashboard/src/components/view_logs/columns.tsx
2026-09-15 16:17:18 -07:00
Tin Chi Lo
39cf1f302d feat(router): apply entitlement limits to forecast classifiers 2026-09-15 16:17:00 -07:00
kerry
726430c0e7 refactor(model_info): scope Gemini baseline to first-party chat ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:16:32 +00:00
yassin
0c65dcff22 fix(proxy): strip line breaks from alias resolution debug log
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:15:34 +00:00
yassin
f6f782ff67 feat(s3): add s3_log_prompts_only option to log prompts without responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:11:44 +00:00
tin-berri
8cd00d2d6e
Merge pull request #41282 from BerriAI/litellm_fast_mode_toggle_0915
Some checks failed
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
ai-gateway image / ai-gateway release image (push) Has been cancelled
feat(auto-router): add per-model Fast mode toggle
2026-09-15 16:10:06 -07:00
Mateo Wang
a056370e1e
Merge pull request #41179 from IToSSc/feat/TASK-2BK38Y-model-prices-docs
feat: add aihubmix provider pricing entries
2026-09-15 16:09:05 -07:00
yuneng-jiang
fa9177d5f2
Merge pull request #41321 from BerriAI/litellm_/release-version-bump-940c92
chore: bump litellm-enterprise 0.1.67 -> 0.1.68, litellm-proxy-extras 0.4.97 -> 0.4.98, litellm 1.102.0 -> 1.103.0
2026-09-15 16:06:39 -07:00
kerry
e64371062c test: remove test files left empty by the cleanup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:04:40 +00:00
yassin
c8a2d8c349 feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard
Adds a daily spend table without api_key or user_id, written atomically alongside
LiteLLM_DailyUserSpend from the batched writer, reconciled from history by a
scheduled job that advances a marker in LiteLLM_Config, and read by the key-free
arm of the aggregated usage query once the marker covers the requested range.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:00:36 +00:00
yassin
c007fb9928 fix(proxy): skip alias rewrite only for dispatched pass-through handlers
Match the pass-through skip to what FastAPI actually dispatched (the user-defined
endpoint marker or a provider handler's {endpoint:path} param) instead of the
mapped route prefixes, which also cover native routes such as /openai/v1/responses
and /cursor/chat/completions. Wrap the added test lines to the 120-column limit.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:00:08 +00:00
Mateo Wang
6bc821492b
Merge pull request #40476 from BerriAI/litellm_codex_model_catalog_sync
feat(cli): sync Codex /model picker from proxy /v1/models in lite codex
2026-09-15 15:59:04 -07:00
yassin
557c237b81 test(proxy-extras): move the LITELLM_LOG logger test into the CI-run proxy-extras shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:57:23 +00:00
Yuneng Jiang
56685ddeae
Merge remote-tracking branch 'origin/main' into litellm_/release-version-bump-940c92 2026-09-15 15:56:29 -07:00
Yassin Kortam
0bb95e298c
Merge pull request #41309 from BerriAI/litellm_playground_custom_request_headers 2026-09-15 15:56:22 -07:00
Yuneng Jiang
818fd7bbb1
Merge remote-tracking branch 'origin/main' into litellm_/release-version-bump-940c92 2026-09-15 15:56:18 -07:00
Yassin Kortam
9d75cdd502
Merge pull request #41304 from BerriAI/litellm_lit1698_openai_system_messages_first 2026-09-15 15:56:13 -07:00
kerry
082f02bd73 refactor(model_info): drop Gemini vertex routing rule, keep baseline to certain capabilities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:55:29 +00:00
yassin
b50a22b370 fix(proxy): restrict /nvidia_nim route to NIM-backed model groups and inject router in tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:55:14 +00:00
Yassin Kortam
9bb83fdaea
Merge pull request #41281 from BerriAI/litellm_lit_7417_jwt_key_mapping_issuer_scope
fix(jwt-auth): scope JWT key mappings by issuer to prevent cross-issuer collisions
2026-09-15 15:51:50 -07:00
yassin
fd90eeb3c6 test(proxy): cover disjoint-method db/yaml pass-through entries on a shared path
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:50:29 +00:00
shivam
a654034f6a test(ci): allow completion_cost recursion for mixed-tier WS pricing split
The recursion is bounded to depth 1: split parts each carry a single
service_tier, so the recursive call's partition has one key and the split
helper returns None.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:50:28 +00:00
kerry
c94a4f8696 test: restore xai reported-cost passthrough tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:49:57 +00:00
kerry
024a887521 feat(model_info): add Gemini fallback generalization rules (vertex routing + 2.5+ family baseline)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:49:35 +00:00
Yuneng Jiang
ad74c3ede7
bump: litellm-enterprise 0.1.67 -> 0.1.68, litellm-proxy-extras 0.4.97 -> 0.4.98, litellm 1.102.0 -> 1.103.0 2026-09-15 15:48:54 -07:00
yassin
91dde59bc3 refactor(proxy): validate request metadata before reconstructing key rate-limit view
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 22:48:47 +00:00
Yassin Kortam
eec954bf9a
Merge pull request #41296 from BerriAI/litellm_models_table_url_state
feat(ui): persist Models table search, filters, sort and page in the URL
2026-09-15 15:46:27 -07:00