Joshua Valluru
743684bdbe
fix(mcp): preserve request-selected guardrails during tool execution
2026-09-17 09:57:08 -07:00
yucheng
b227a8c4c9
refactor(proxy): move login throttle sentinels into constants
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 16:42:04 +00:00
Yuneng Jiang
44bc3d1436
test(management): avoid pinning tenant error disclosure
2026-09-17 09:41:49 -07:00
mateo
4f3b90b588
fix(proxy): expose TypeSafe passthrough on gateway
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 16:32:04 +00:00
Joshua Valluru
5a105657c1
test(mcp): isolate health assertions to owned servers
2026-09-17 09:25:05 -07:00
mateo
7fca7fae37
fix(proxy): satisfy TypeSafe CI gates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 16:20:47 +00:00
Joshua Valluru
664b1f16bb
style(tests): wrap MCP health regression setup
2026-09-17 09:19:05 -07:00
Joshua Valluru
e21db01d67
fix(mcp): scope health discovery for route-restricted keys
2026-09-17 09:17:55 -07:00
yucheng
9a365d2021
feat(proxy): hard-block throttled Admin UI sign-ins with no credential bypass
...
A blocked source, or source and username pair, is now refused with 429 before the database lookup and password check, in place of the soft block that held wrong guesses for 30 seconds and let a correct password through. The env admin credentials and the master key typed into the login form are refused like any other credential while blocked; recovery is the master key as an API bearer token, which never goes through the sign-in path
trusted_proxy_ranges: [] now means clients connect directly, so the peer address is the source and the per-source limit stays on. Only an unset or malformed value leaves the topology unknown, warns at startup and turns the per-source limit off
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 16:10:23 +00:00
mateo
9470aa47f9
fix(proxy): satisfy TypeSafe CI gates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 16:01:27 +00:00
Yujong Lee
f1ea94fee7
make test pass
2026-09-17 08:55:59 -07:00
mateo
78eb92ca55
refactor(proxy): simplify TypeSafe passthrough pricing lookup and route
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 15:55:57 +00:00
mateo
2dc9697381
feat(proxy): add TypeSafe Jev passthrough spend tracking
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 15:53:19 +00:00
Yujong Lee
370cdaabf9
encode failing tests
2026-09-17 08:04:52 -07:00
Yujong Lee
b063ffe883
providers folder is gone
2026-09-17 07:52:14 -07:00
Yujong Lee
27ccf7326b
mistral alignment
2026-09-17 07:33:06 -07:00
Yujong Lee
ab1f966a17
test coverage
2026-09-17 07:10:49 -07:00
Yujong Lee
0e5f41bc93
refactor(rust): drop unused OpaqueParams body-composition helpers
2026-09-17 06:54:18 -07:00
Yujong Lee
3ad91fc27e
fix(ocr): run hooks on completed Azure poll
2026-09-17 06:53:19 -07:00
Yujong Lee
0f636c5db1
refactor(core): use string for DeepSeek model
2026-09-17 06:51:52 -07:00
Devin AI
f79c3ebfee
fix(models): align Bedrock Mantle Grok 4.3 GovCloud context window with model card
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:48:59 +00:00
Yujong Lee
cde34d2b39
fix(rust): decode Anthropic citation deltas
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:43:28 +00:00
Devin AI
b1b6747869
fix(models): Azure retirement dates and Bedrock Mantle Grok 4.3 context window
...
Azure schedule: https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule
AWS card: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-3.html
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:31:57 +00:00
yucheng
48b25e448d
fix(grafana): sum redis circuit breaker state across workers instead of taking the max
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:44:26 +00:00
Devin AI
2876ac03ce
fix: keep fastapi import within proxy
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:41:58 +00:00
Devin AI
25949a87ac
test: cover deferred import branches
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:26:35 +00:00
yucheng
880897826f
fix(grafana): aggregate provider remaining budget with min like the other remaining budget panels
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:21:19 +00:00
yucheng
f89fb20709
test(prometheus): restore the full registry and build admission metrics fresh in the dashboard fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:21:19 +00:00
yucheng
8ae2ebdfcf
test(prometheus): reset the admission control metric owner in the dashboard consistency fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:00:22 +00:00
Devin AI
293a96332c
perf: defer fastapi and tiktoken BPE imports out of import litellm
...
This defers FastAPI, Starlette, and the cl100k BPE table until the paths that use them run
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:59:48 +00:00
yucheng
4aa8d06eda
test(prometheus): restore unrelated collectors after the dashboard consistency fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:45:46 +00:00
yucheng
9fa85c5da5
fix(proxy): name the blocking guardrail in x-litellm-applied-guardrails
...
When a guardrail hook raises, the common ProxyLogging dispatch (sequential and parallel pre_call, pipeline block, during_call and post_call metrics wrapper, streaming iterator wrapper) now records that guardrail in applied_guardrails before re-raising, and pre_call_hook folds request-declared guardrails in on its raising path. Buffered streams rebuild their response headers after the first chunk so a post_call block reached while buffering carries the blocker too
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:39:01 +00:00
yucheng
aa1fedbfdd
fix(grafana): reset lazy Prometheus collectors in the dashboard test fixture and state the overhead panel unit in seconds
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:20:50 +00:00
yucheng
184add7cee
fix(grafana): hide the batch cost last-run panel until the job has run once
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:13:31 +00:00
yucheng
8164189237
test(ui): cover a negative policy attachment priority typed keystroke by keystroke
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:56:43 +00:00
yucheng
5451c38dcc
feat(grafana): add all-metrics dashboard and fix stale dashboard_v2 gauges
...
Fixes the litellm_remaining_requests and litellm_remaining_tokens queries in
dashboard_v2 (renamed to *_metric in v1.80.15) and adds dashboard_all_metrics
with a panel for every litellm_* family the proxy can emit, including the
prometheus_system service metrics, admission control, Redis circuit breaker and
spend log cleanup metrics. dashboard_1 charted a metric that is never emitted
and is superseded, so it is removed. A test fails when a dashboard references a
metric the proxy does not emit or when an emitted family has no panel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:50:36 +00:00
kerry-berri
4b368bf066
Merge pull request #41576 from BerriAI/litellm_openrouter_union_alpha
...
feat(openrouter): add stealth/union-alpha to the model cost map
2026-09-17 00:37:39 -07:00
yucheng-berri
e40b90bbfa
Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context
...
fix(guardrails): give post-call scans the scoped request conversation and tools
2026-09-17 00:31:46 -07:00
kerry
3d0fd127d5
feat(openrouter): add stealth/union-alpha to the model cost map
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:20:16 +00:00
yucheng
b237c185db
test(guardrails): type the recording guardrail logging_obj as the logging object
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:12:27 +00:00
yucheng-berri
d8d5437f55
Merge pull request #41558 from BerriAI/litellm_lit_6568_streaming_redaction
...
fix(guardrails): stream Prompt Security post_call redactions in incremental_diff mode
2026-09-17 00:10:45 -07:00
yucheng
b536289233
refactor(guardrails): type the request scan context helpers as read-only mappings
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:54:17 +00:00
yucheng
669a66499c
feat(policy_engine): bound priority to int32 and expose it in the Admin UI
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:47:53 +00:00
Yuneng Jiang
a40b6b3e44
test(management): close project lifecycle coverage gaps
2026-09-16 23:43:25 -07:00
yucheng
c05095373d
fix(proxy): keep mapped-notation trusted proxy ranges matching mapped peers
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:41:23 +00:00
berriai-litellm-provider-info-sync[bot]
5a82ed9bca
chore(prices): sync prices for 2 providers: 27 models
...
fireworks_ai/accounts/fireworks/models/minimax-m3: supports_vision
fireworks_ai/minimax-m3: supports_vision
wandb/deepseek-ai/DeepSeek-V4-Flash: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Flash-0731: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro-0813: max_input_tokens
wandb/google/gemma-4-31B-it: max_input_tokens
wandb/ibm-granite/granite-4.1-8b: max_input_tokens
wandb/ibm-granite/granite-4.2-8b: max_input_tokens
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-70B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-8B-Instruct: max_input_tokens
wandb/MiniMaxAI/MiniMax-M3: max_input_tokens
wandb/moonshotai/Kimi-K2.6: max_input_tokens
wandb/moonshotai/Kimi-K2.7-Code: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: max_input_tokens
wandb/openai/gpt-oss-120b: max_input_tokens
wandb/openai/gpt-oss-20b: max_input_tokens
wandb/OpenPipe/Qwen3-14B-Instruct: max_input_tokens
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: max_input_tokens
wandb/Qwen/Qwen3.5-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.6-27B: max_input_tokens
wandb/Qwen/Qwen3.6-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.8-27B: max_input_tokens
wandb/zai-org/GLM-5.2: max_input_tokens
wandb/zai-org/GLM-5.3-Flash: max_input_tokens
2026-09-17 06:31:05 +00:00
yuneng-jiang
5c486126ac
Merge pull request #41568 from BerriAI/litellm_/ci-test-failures-investigation-99785b
...
test(e2e): drop the auto-router select "opens below" spec
2026-09-16 23:23:36 -07:00
Devin AI
1b1f6ada46
fix(policy_engine): make priority migration idempotent
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:12:42 +00:00
yucheng
0986f404f8
fix(proxy): match IPv4-mapped IPv6 peers against IPv4 trusted proxy ranges
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:08:12 +00:00
Yuneng Jiang
4210f586c2
test(management): cover project authorization lifecycle
2026-09-16 23:08:00 -07:00