Commit graph

53363 commits

Author SHA1 Message Date
Devin AI
2876ac03ce fix: keep fastapi import within proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:41:58 +00:00
Devin AI
25949a87ac test: cover deferred import branches
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:26:35 +00:00
yucheng
880897826f fix(grafana): aggregate provider remaining budget with min like the other remaining budget panels
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:21:19 +00:00
yucheng
f89fb20709 test(prometheus): restore the full registry and build admission metrics fresh in the dashboard fixture
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:21:19 +00:00
yucheng
8ae2ebdfcf test(prometheus): reset the admission control metric owner in the dashboard consistency fixture
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:00:22 +00:00
Devin AI
293a96332c perf: defer fastapi and tiktoken BPE imports out of import litellm
This defers FastAPI, Starlette, and the cl100k BPE table until the paths that use them run

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:59:48 +00:00
yucheng
4aa8d06eda test(prometheus): restore unrelated collectors after the dashboard consistency fixture
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:45:46 +00:00
yucheng
9fa85c5da5 fix(proxy): name the blocking guardrail in x-litellm-applied-guardrails
When a guardrail hook raises, the common ProxyLogging dispatch (sequential and parallel pre_call, pipeline block, during_call and post_call metrics wrapper, streaming iterator wrapper) now records that guardrail in applied_guardrails before re-raising, and pre_call_hook folds request-declared guardrails in on its raising path. Buffered streams rebuild their response headers after the first chunk so a post_call block reached while buffering carries the blocker too

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:39:01 +00:00
yucheng
aa1fedbfdd fix(grafana): reset lazy Prometheus collectors in the dashboard test fixture and state the overhead panel unit in seconds
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:20:50 +00:00
yucheng
184add7cee fix(grafana): hide the batch cost last-run panel until the job has run once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:13:31 +00:00
yucheng
8164189237 test(ui): cover a negative policy attachment priority typed keystroke by keystroke
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:56:43 +00:00
yucheng
5451c38dcc feat(grafana): add all-metrics dashboard and fix stale dashboard_v2 gauges
Fixes the litellm_remaining_requests and litellm_remaining_tokens queries in
dashboard_v2 (renamed to *_metric in v1.80.15) and adds dashboard_all_metrics
with a panel for every litellm_* family the proxy can emit, including the
prometheus_system service metrics, admission control, Redis circuit breaker and
spend log cleanup metrics. dashboard_1 charted a metric that is never emitted
and is superseded, so it is removed. A test fails when a dashboard references a
metric the proxy does not emit or when an emitted family has no panel

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:50:36 +00:00
kerry-berri
4b368bf066
Merge pull request #41576 from BerriAI/litellm_openrouter_union_alpha
feat(openrouter): add stealth/union-alpha to the model cost map
2026-09-17 00:37:39 -07:00
yucheng-berri
e40b90bbfa
Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context
fix(guardrails): give post-call scans the scoped request conversation and tools
2026-09-17 00:31:46 -07:00
kerry
3d0fd127d5 feat(openrouter): add stealth/union-alpha to the model cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:20:16 +00:00
David Steele
bd222bd8d9
fix(azure): avoid mutable request mapping
DEVX-829

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-17 08:18:36 +01:00
yucheng
b237c185db test(guardrails): type the recording guardrail logging_obj as the logging object
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:12:27 +00:00
yucheng-berri
d8d5437f55
Merge pull request #41558 from BerriAI/litellm_lit_6568_streaming_redaction
fix(guardrails): stream Prompt Security post_call redactions in incremental_diff mode
2026-09-17 00:10:45 -07:00
yucheng
b536289233 refactor(guardrails): type the request scan context helpers as read-only mappings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:54:17 +00:00
yucheng
669a66499c feat(policy_engine): bound priority to int32 and expose it in the Admin UI
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:47:53 +00:00
Yuneng Jiang
a40b6b3e44
test(management): close project lifecycle coverage gaps 2026-09-16 23:43:25 -07:00
yucheng
c05095373d fix(proxy): keep mapped-notation trusted proxy ranges matching mapped peers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:41:23 +00:00
berriai-litellm-provider-info-sync[bot]
5a82ed9bca
chore(prices): sync prices for 2 providers: 27 models
fireworks_ai/accounts/fireworks/models/minimax-m3: supports_vision
fireworks_ai/minimax-m3: supports_vision
wandb/deepseek-ai/DeepSeek-V4-Flash: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Flash-0731: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro-0813: max_input_tokens
wandb/google/gemma-4-31B-it: max_input_tokens
wandb/ibm-granite/granite-4.1-8b: max_input_tokens
wandb/ibm-granite/granite-4.2-8b: max_input_tokens
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-70B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-8B-Instruct: max_input_tokens
wandb/MiniMaxAI/MiniMax-M3: max_input_tokens
wandb/moonshotai/Kimi-K2.6: max_input_tokens
wandb/moonshotai/Kimi-K2.7-Code: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: max_input_tokens
wandb/openai/gpt-oss-120b: max_input_tokens
wandb/openai/gpt-oss-20b: max_input_tokens
wandb/OpenPipe/Qwen3-14B-Instruct: max_input_tokens
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: max_input_tokens
wandb/Qwen/Qwen3.5-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.6-27B: max_input_tokens
wandb/Qwen/Qwen3.6-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.8-27B: max_input_tokens
wandb/zai-org/GLM-5.2: max_input_tokens
wandb/zai-org/GLM-5.3-Flash: max_input_tokens
2026-09-17 06:31:05 +00:00
David Steele
48712f733a
style(azure): format request parameter filter
DEVX-829

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-17 07:30:06 +01:00
yuneng-jiang
5c486126ac
Merge pull request #41568 from BerriAI/litellm_/ci-test-failures-investigation-99785b
test(e2e): drop the auto-router select "opens below" spec
2026-09-16 23:23:36 -07:00
Devin AI
1b1f6ada46 fix(policy_engine): make priority migration idempotent
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:12:42 +00:00
yucheng
0986f404f8 fix(proxy): match IPv4-mapped IPv6 peers against IPv4 trusted proxy ranges
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:08:12 +00:00
Yuneng Jiang
4210f586c2
test(management): cover project authorization lifecycle 2026-09-16 23:08:00 -07:00
David Steele
e254377049
fix(azure): drop tool_choice without tools
DEVX-829
2026-09-17 07:07:57 +01:00
Devin AI
85444b56d9 fix(guardrails): hand the input scan context to the logging_only response scan
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:04:38 +00:00
yuneng-jiang
f58389c0e8
Merge pull request #41563 from BerriAI/litellm_budget-null-clear-tests
test(budgets): cover management null handling
2026-09-16 23:02:04 -07:00
berriai-litellm-provider-info-sync[bot]
b4212b949b
chore(prices): sync prices for 5 providers: 34 models, 1 new, 19 deprecated [1 with gaps]
fireworks_ai/accounts/fireworks/routers/glm-5p3-fast: 
azure_ai/FW-Kimi-K3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/deepseek-ai/DeepSeek-R1-0528: deprecation_date
wandb/deepseek-ai/DeepSeek-V3-0324: deprecation_date
wandb/deepseek-ai/DeepSeek-V4-Flash: deprecation_date
wandb/deepseek-ai/DeepSeek-V4-Pro: deprecation_date
together_ai/deepseek-ai/DeepSeek-V4.1-Flash: 
azure/eu/gpt-5.5-2026-04-24: 
gemini-3.8-live: supports_response_schema
gemini-3.8-live-extended-thinking: supports_response_schema
azure/gpt-5.5-2026-04-24: 
azure/gpt-5.6-luna-2026-07-09: 
azure/gpt-5.6-sol-2026-07-09: 
azure/gpt-5.6-terra-2026-07-09: 
azure/gpt-6-astra-2026-09-03: 
wandb/ibm-granite/granite-4.1-8b: deprecation_date
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: deprecation_date
wandb/meta-llama/Llama-3.1-70B-Instruct: deprecation_date
wandb/meta-llama/Llama-4-Scout-17B-16E-Instruct: deprecation_date
wandb/microsoft/Phi-4-mini-instruct: deprecation_date
wandb/MiniMaxAI/MiniMax-M2.5: deprecation_date
wandb/moonshotai/Kimi-K2-Instruct: deprecation_date
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/OpenPipe/Qwen3-14B-Instruct: deprecation_date
wandb/Qwen/Qwen3-235B-A22B-Instruct-2507: deprecation_date
wandb/Qwen/Qwen3-235B-A22B-Thinking-2507: deprecation_date
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: deprecation_date
wandb/Qwen/Qwen3-Coder-480B-A35B-Instruct: deprecation_date
wandb/Qwen/Qwen3.5-35B-A3B: deprecation_date
wandb/Qwen/Qwen3.6-27B: deprecation_date
azure/us/gpt-5.5-2026-04-24: 
wandb/zai-org/GLM-4.5: deprecation_date
wandb/zai-org/GLM-5.3-Flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, supports_function_calling, supports_tool_choice, supports_response_schema, supports_prompt_caching, supports_reasoning
2026-09-17 06:01:14 +00:00
Devin AI
f5dea4de76 refactor(policy_engine): shorten attachment priority field descriptions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:00:25 +00:00
Devin AI
21ffbdc7ea feat(policy_engine): explicit priority for policy attachment execution order
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:59:01 +00:00
yucheng
3406913ca0 test(spend-logs): expect azure_spillover in spend log metadata golden
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:58:19 +00:00
Devin AI
8ce2887888 fix(responses): close the message content part as output_text on reasoning turns
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:56:28 +00:00
yucheng
7b855bd53f feat(spend-logs): record Azure spillover source deployment in spend log metadata
SpendLogsMetadata gains a typed azure_spillover key so a request Azure
served off pay-as-you-go capacity is visible in spend tracking, stamped
from the provider response headers or the processed llm_provider- headers
on the standard logging payload. The header parsing moves into a shared
azure_spillover() helper that is_spilled_over_ptu_request() now wraps.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:48:51 +00:00
yucheng
e3c8f74a4f fix(proxy): default prompt injection heuristics executor to a single worker
SequenceMatcher holds the GIL, so extra heuristic threads add contention with the event loop without adding throughput. One worker drains scans in arrival order and keeps the loop responsive; PROMPT_INJECTION_HEURISTICS_MAX_THREADS remains an env override

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:45:33 +00:00
Yuneng Jiang
25445e8b5c
test(e2e): drop the auto-router select "opens below" spec
The spec pinned Base UI's collision behaviour, not our code: it only passes
while the template popup happens to fit under the trigger at 1280x900, and
#41315's taller Add Auto Router form broke that premise for the second time
in three weeks. #41527 tried to scroll the trigger into the upper half, but
the dialog content is shorter than its max height, so nothing scrolls and CI
still fails 3/3 with the trigger at y=487

The guarantee #38554 introduced is that the popup never covers the trigger,
and the sibling spec keeps asserting that at a viewport with no room below
2026-09-16 22:43:51 -07:00
Devin AI
5861240946 test(responses): type new streaming bridge test parameters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:43:14 +00:00
yucheng
2d925e5dde fix(guardrails): scope the logging_only reply scan with the request's own translation
The chat-shaped output handler now takes the input translation as its
request scoping, so the logged request is scoped exactly once and with
the pre-call semantics of the surface it arrived on. This drops the
unscoped chat_shaped_request_conversation detour from af312dc8, which
made the Anthropic response scan remove in-sequence system turns under
skip_system while the request scan kept them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:40:44 +00:00
yucheng
d6f6f64c0f fix(proxy): derive prompt injection heuristics thread count from CPU count with env override
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:30:41 +00:00
Devin AI
af312dc8d7 fix(guardrails): scope the logging_only response scan once
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:30:06 +00:00
yucheng
a57483d1c8 fix(cost): price Azure PTU spillover requests at standard token rates
Azure PTU deployments carry zeroed per-token pricing because the reservation
is billed flat by the hour. When Azure spills a request onto pay-as-you-go
capacity it returns x-ms-is-spilled-over: true, and that traffic was still
priced at zero. The response cost calculator now detects the spillover header
on the result's hidden params or the logged provider response headers and
skips the zeroed custom pricing only for genuine PTU deployments while the
feature flag is on. Azure sync streaming now also records response headers on
the logging object, matching the async paths.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:24:05 +00:00
Devin AI
e9625ad069 fix(responses): default the message output index when no reasoning item exists
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:22:41 +00:00
Devin AI
42c4c81633 fix(responses): allocate the message output index from the shared item allocator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:15:07 +00:00
yucheng
060abd263e fix(guardrails): keep usage chunk and defer tool_calls finish_reason behind held text in incremental_diff
A stream_options.include_usage usage chunk (empty delta plus usage) was folded into the final
transform round and rebuilt without its usage, so token counts and cost vanished from clients.
Metadata-only chunks are now replayed after the final text flush.

A terminal tool-call chunk arriving while earlier text was still held back carried
finish_reason=tool_calls ahead of that text. The finish_reason is now deferred to the final
text chunk whenever the choice has held text.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:11:04 +00:00
yucheng
87263cefca Merge remote-tracking branch 'origin/main' into litellm_post_call_guardrail_context
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	tests/test_litellm/llms/openai/responses/test_openai_responses_guardrail_handler.py
2026-09-17 05:08:02 +00:00
Yuneng Jiang
a0869fe835
test(budgets): avoid mutable fixture state 2026-09-16 22:06:31 -07:00
Devin AI
d4e54a0f34 fix(responses): keep sync text deltas and give the message item its own output index
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:05:54 +00:00