Yujong Lee
27ccf7326b
mistral alignment
2026-09-17 07:33:06 -07:00
Yujong Lee
ab1f966a17
test coverage
2026-09-17 07:10:49 -07:00
Yujong Lee
0e5f41bc93
refactor(rust): drop unused OpaqueParams body-composition helpers
2026-09-17 06:54:18 -07:00
Yujong Lee
3ad91fc27e
fix(ocr): run hooks on completed Azure poll
2026-09-17 06:53:19 -07:00
Yujong Lee
0f636c5db1
refactor(core): use string for DeepSeek model
2026-09-17 06:51:52 -07:00
Devin AI
f79c3ebfee
fix(models): align Bedrock Mantle Grok 4.3 GovCloud context window with model card
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:48:59 +00:00
Yujong Lee
cde34d2b39
fix(rust): decode Anthropic citation deltas
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:43:28 +00:00
Devin AI
b1b6747869
fix(models): Azure retirement dates and Bedrock Mantle Grok 4.3 context window
...
Azure schedule: https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule
AWS card: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-3.html
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:31:57 +00:00
Devin AI
48bde68781
fix(key_generate): use user's budget for UI session personal keys
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:47:14 +00:00
yucheng
48b25e448d
fix(grafana): sum redis circuit breaker state across workers instead of taking the max
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:44:26 +00:00
Devin AI
2876ac03ce
fix: keep fastapi import within proxy
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:41:58 +00:00
Devin AI
25949a87ac
test: cover deferred import branches
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:26:35 +00:00
yucheng
880897826f
fix(grafana): aggregate provider remaining budget with min like the other remaining budget panels
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:21:19 +00:00
yucheng
f89fb20709
test(prometheus): restore the full registry and build admission metrics fresh in the dashboard fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:21:19 +00:00
yucheng
8ae2ebdfcf
test(prometheus): reset the admission control metric owner in the dashboard consistency fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:00:22 +00:00
Devin AI
293a96332c
perf: defer fastapi and tiktoken BPE imports out of import litellm
...
This defers FastAPI, Starlette, and the cl100k BPE table until the paths that use them run
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:59:48 +00:00
yucheng
4aa8d06eda
test(prometheus): restore unrelated collectors after the dashboard consistency fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:45:46 +00:00
yucheng
9fa85c5da5
fix(proxy): name the blocking guardrail in x-litellm-applied-guardrails
...
When a guardrail hook raises, the common ProxyLogging dispatch (sequential and parallel pre_call, pipeline block, during_call and post_call metrics wrapper, streaming iterator wrapper) now records that guardrail in applied_guardrails before re-raising, and pre_call_hook folds request-declared guardrails in on its raising path. Buffered streams rebuild their response headers after the first chunk so a post_call block reached while buffering carries the blocker too
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:39:01 +00:00
yucheng
aa1fedbfdd
fix(grafana): reset lazy Prometheus collectors in the dashboard test fixture and state the overhead panel unit in seconds
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:20:50 +00:00
yucheng
184add7cee
fix(grafana): hide the batch cost last-run panel until the job has run once
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:13:31 +00:00
yucheng
8164189237
test(ui): cover a negative policy attachment priority typed keystroke by keystroke
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:56:43 +00:00
yucheng
5451c38dcc
feat(grafana): add all-metrics dashboard and fix stale dashboard_v2 gauges
...
Fixes the litellm_remaining_requests and litellm_remaining_tokens queries in
dashboard_v2 (renamed to *_metric in v1.80.15) and adds dashboard_all_metrics
with a panel for every litellm_* family the proxy can emit, including the
prometheus_system service metrics, admission control, Redis circuit breaker and
spend log cleanup metrics. dashboard_1 charted a metric that is never emitted
and is superseded, so it is removed. A test fails when a dashboard references a
metric the proxy does not emit or when an emitted family has no panel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:50:36 +00:00
kerry-berri
4b368bf066
Merge pull request #41576 from BerriAI/litellm_openrouter_union_alpha
...
feat(openrouter): add stealth/union-alpha to the model cost map
2026-09-17 00:37:39 -07:00
yucheng-berri
e40b90bbfa
Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context
...
fix(guardrails): give post-call scans the scoped request conversation and tools
2026-09-17 00:31:46 -07:00
kerry
3d0fd127d5
feat(openrouter): add stealth/union-alpha to the model cost map
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:20:16 +00:00
David Steele
bd222bd8d9
fix(azure): avoid mutable request mapping
...
DEVX-829
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-17 08:18:36 +01:00
yucheng
b237c185db
test(guardrails): type the recording guardrail logging_obj as the logging object
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:12:27 +00:00
yucheng-berri
d8d5437f55
Merge pull request #41558 from BerriAI/litellm_lit_6568_streaming_redaction
...
fix(guardrails): stream Prompt Security post_call redactions in incremental_diff mode
2026-09-17 00:10:45 -07:00
yucheng
b536289233
refactor(guardrails): type the request scan context helpers as read-only mappings
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:54:17 +00:00
yucheng
669a66499c
feat(policy_engine): bound priority to int32 and expose it in the Admin UI
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:47:53 +00:00
Yuneng Jiang
a40b6b3e44
test(management): close project lifecycle coverage gaps
2026-09-16 23:43:25 -07:00
yucheng
c05095373d
fix(proxy): keep mapped-notation trusted proxy ranges matching mapped peers
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:41:23 +00:00
berriai-litellm-provider-info-sync[bot]
5a82ed9bca
chore(prices): sync prices for 2 providers: 27 models
...
fireworks_ai/accounts/fireworks/models/minimax-m3: supports_vision
fireworks_ai/minimax-m3: supports_vision
wandb/deepseek-ai/DeepSeek-V4-Flash: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Flash-0731: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro-0813: max_input_tokens
wandb/google/gemma-4-31B-it: max_input_tokens
wandb/ibm-granite/granite-4.1-8b: max_input_tokens
wandb/ibm-granite/granite-4.2-8b: max_input_tokens
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-70B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-8B-Instruct: max_input_tokens
wandb/MiniMaxAI/MiniMax-M3: max_input_tokens
wandb/moonshotai/Kimi-K2.6: max_input_tokens
wandb/moonshotai/Kimi-K2.7-Code: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: max_input_tokens
wandb/openai/gpt-oss-120b: max_input_tokens
wandb/openai/gpt-oss-20b: max_input_tokens
wandb/OpenPipe/Qwen3-14B-Instruct: max_input_tokens
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: max_input_tokens
wandb/Qwen/Qwen3.5-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.6-27B: max_input_tokens
wandb/Qwen/Qwen3.6-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.8-27B: max_input_tokens
wandb/zai-org/GLM-5.2: max_input_tokens
wandb/zai-org/GLM-5.3-Flash: max_input_tokens
2026-09-17 06:31:05 +00:00
David Steele
48712f733a
style(azure): format request parameter filter
...
DEVX-829
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-17 07:30:06 +01:00
yuneng-jiang
5c486126ac
Merge pull request #41568 from BerriAI/litellm_/ci-test-failures-investigation-99785b
...
test(e2e): drop the auto-router select "opens below" spec
2026-09-16 23:23:36 -07:00
Devin AI
1b1f6ada46
fix(policy_engine): make priority migration idempotent
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:12:42 +00:00
yucheng
0986f404f8
fix(proxy): match IPv4-mapped IPv6 peers against IPv4 trusted proxy ranges
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:08:12 +00:00
Yuneng Jiang
4210f586c2
test(management): cover project authorization lifecycle
2026-09-16 23:08:00 -07:00
David Steele
e254377049
fix(azure): drop tool_choice without tools
...
DEVX-829
2026-09-17 07:07:57 +01:00
Devin AI
85444b56d9
fix(guardrails): hand the input scan context to the logging_only response scan
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:04:38 +00:00
yuneng-jiang
f58389c0e8
Merge pull request #41563 from BerriAI/litellm_budget-null-clear-tests
...
test(budgets): cover management null handling
2026-09-16 23:02:04 -07:00
berriai-litellm-provider-info-sync[bot]
b4212b949b
chore(prices): sync prices for 5 providers: 34 models, 1 new, 19 deprecated [1 with gaps]
...
fireworks_ai/accounts/fireworks/routers/glm-5p3-fast:
azure_ai/FW-Kimi-K3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/deepseek-ai/DeepSeek-R1-0528: deprecation_date
wandb/deepseek-ai/DeepSeek-V3-0324: deprecation_date
wandb/deepseek-ai/DeepSeek-V4-Flash: deprecation_date
wandb/deepseek-ai/DeepSeek-V4-Pro: deprecation_date
together_ai/deepseek-ai/DeepSeek-V4.1-Flash:
azure/eu/gpt-5.5-2026-04-24:
gemini-3.8-live: supports_response_schema
gemini-3.8-live-extended-thinking: supports_response_schema
azure/gpt-5.5-2026-04-24:
azure/gpt-5.6-luna-2026-07-09:
azure/gpt-5.6-sol-2026-07-09:
azure/gpt-5.6-terra-2026-07-09:
azure/gpt-6-astra-2026-09-03:
wandb/ibm-granite/granite-4.1-8b: deprecation_date
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: deprecation_date
wandb/meta-llama/Llama-3.1-70B-Instruct: deprecation_date
wandb/meta-llama/Llama-4-Scout-17B-16E-Instruct: deprecation_date
wandb/microsoft/Phi-4-mini-instruct: deprecation_date
wandb/MiniMaxAI/MiniMax-M2.5: deprecation_date
wandb/moonshotai/Kimi-K2-Instruct: deprecation_date
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/OpenPipe/Qwen3-14B-Instruct: deprecation_date
wandb/Qwen/Qwen3-235B-A22B-Instruct-2507: deprecation_date
wandb/Qwen/Qwen3-235B-A22B-Thinking-2507: deprecation_date
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: deprecation_date
wandb/Qwen/Qwen3-Coder-480B-A35B-Instruct: deprecation_date
wandb/Qwen/Qwen3.5-35B-A3B: deprecation_date
wandb/Qwen/Qwen3.6-27B: deprecation_date
azure/us/gpt-5.5-2026-04-24:
wandb/zai-org/GLM-4.5: deprecation_date
wandb/zai-org/GLM-5.3-Flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, supports_function_calling, supports_tool_choice, supports_response_schema, supports_prompt_caching, supports_reasoning
2026-09-17 06:01:14 +00:00
Devin AI
f5dea4de76
refactor(policy_engine): shorten attachment priority field descriptions
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:00:25 +00:00
Devin AI
21ffbdc7ea
feat(policy_engine): explicit priority for policy attachment execution order
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:59:01 +00:00
yucheng
3406913ca0
test(spend-logs): expect azure_spillover in spend log metadata golden
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:58:19 +00:00
Devin AI
8ce2887888
fix(responses): close the message content part as output_text on reasoning turns
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:56:28 +00:00
yucheng
7b855bd53f
feat(spend-logs): record Azure spillover source deployment in spend log metadata
...
SpendLogsMetadata gains a typed azure_spillover key so a request Azure
served off pay-as-you-go capacity is visible in spend tracking, stamped
from the provider response headers or the processed llm_provider- headers
on the standard logging payload. The header parsing moves into a shared
azure_spillover() helper that is_spilled_over_ptu_request() now wraps.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:48:51 +00:00
yucheng
e3c8f74a4f
fix(proxy): default prompt injection heuristics executor to a single worker
...
SequenceMatcher holds the GIL, so extra heuristic threads add contention with the event loop without adding throughput. One worker drains scans in arrival order and keeps the loop responsive; PROMPT_INJECTION_HEURISTICS_MAX_THREADS remains an env override
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:45:33 +00:00
Yuneng Jiang
25445e8b5c
test(e2e): drop the auto-router select "opens below" spec
...
The spec pinned Base UI's collision behaviour, not our code: it only passes
while the template popup happens to fit under the trigger at 1280x900, and
#41315 's taller Add Auto Router form broke that premise for the second time
in three weeks. #41527 tried to scroll the trigger into the upper half, but
the dialog content is shorter than its max height, so nothing scrolls and CI
still fails 3/3 with the trigger at y=487
The guarantee #38554 introduced is that the popup never covers the trigger,
and the sibling spec keeps asserting that at a viewport with no room below
2026-09-16 22:43:51 -07:00
Devin AI
5861240946
test(responses): type new streaming bridge test parameters
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:43:14 +00:00