mateo
78eb92ca55
refactor(proxy): simplify TypeSafe passthrough pricing lookup and route
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 15:55:57 +00:00
mateo
2dc9697381
feat(proxy): add TypeSafe Jev passthrough spend tracking
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 15:53:19 +00:00
Yujong Lee
370cdaabf9
encode failing tests
2026-09-17 08:04:52 -07:00
Yujong Lee
b063ffe883
providers folder is gone
2026-09-17 07:52:14 -07:00
Yujong Lee
27ccf7326b
mistral alignment
2026-09-17 07:33:06 -07:00
Yujong Lee
ab1f966a17
test coverage
2026-09-17 07:10:49 -07:00
Yujong Lee
0e5f41bc93
refactor(rust): drop unused OpaqueParams body-composition helpers
2026-09-17 06:54:18 -07:00
Yujong Lee
3ad91fc27e
fix(ocr): run hooks on completed Azure poll
2026-09-17 06:53:19 -07:00
Yujong Lee
0f636c5db1
refactor(core): use string for DeepSeek model
2026-09-17 06:51:52 -07:00
Devin AI
f79c3ebfee
fix(models): align Bedrock Mantle Grok 4.3 GovCloud context window with model card
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:48:59 +00:00
Yujong Lee
cde34d2b39
fix(rust): decode Anthropic citation deltas
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:43:28 +00:00
Devin AI
b1b6747869
fix(models): Azure retirement dates and Bedrock Mantle Grok 4.3 context window
...
Azure schedule: https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-schedule
AWS card: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-3.html
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 13:31:57 +00:00
yucheng
48b25e448d
fix(grafana): sum redis circuit breaker state across workers instead of taking the max
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:44:26 +00:00
Devin AI
2876ac03ce
fix: keep fastapi import within proxy
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:41:58 +00:00
Devin AI
25949a87ac
test: cover deferred import branches
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:26:35 +00:00
yucheng
880897826f
fix(grafana): aggregate provider remaining budget with min like the other remaining budget panels
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:21:19 +00:00
yucheng
f89fb20709
test(prometheus): restore the full registry and build admission metrics fresh in the dashboard fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:21:19 +00:00
yucheng
8ae2ebdfcf
test(prometheus): reset the admission control metric owner in the dashboard consistency fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 09:00:22 +00:00
Devin AI
293a96332c
perf: defer fastapi and tiktoken BPE imports out of import litellm
...
This defers FastAPI, Starlette, and the cl100k BPE table until the paths that use them run
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:59:48 +00:00
yucheng
4aa8d06eda
test(prometheus): restore unrelated collectors after the dashboard consistency fixture
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:45:46 +00:00
yucheng
aa1fedbfdd
fix(grafana): reset lazy Prometheus collectors in the dashboard test fixture and state the overhead panel unit in seconds
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:20:50 +00:00
yucheng
184add7cee
fix(grafana): hide the batch cost last-run panel until the job has run once
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 08:13:31 +00:00
yucheng
8164189237
test(ui): cover a negative policy attachment priority typed keystroke by keystroke
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:56:43 +00:00
yucheng
5451c38dcc
feat(grafana): add all-metrics dashboard and fix stale dashboard_v2 gauges
...
Fixes the litellm_remaining_requests and litellm_remaining_tokens queries in
dashboard_v2 (renamed to *_metric in v1.80.15) and adds dashboard_all_metrics
with a panel for every litellm_* family the proxy can emit, including the
prometheus_system service metrics, admission control, Redis circuit breaker and
spend log cleanup metrics. dashboard_1 charted a metric that is never emitted
and is superseded, so it is removed. A test fails when a dashboard references a
metric the proxy does not emit or when an emitted family has no panel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:50:36 +00:00
kerry-berri
4b368bf066
Merge pull request #41576 from BerriAI/litellm_openrouter_union_alpha
...
feat(openrouter): add stealth/union-alpha to the model cost map
2026-09-17 00:37:39 -07:00
yucheng-berri
e40b90bbfa
Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context
...
fix(guardrails): give post-call scans the scoped request conversation and tools
2026-09-17 00:31:46 -07:00
kerry
3d0fd127d5
feat(openrouter): add stealth/union-alpha to the model cost map
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:20:16 +00:00
yucheng
b237c185db
test(guardrails): type the recording guardrail logging_obj as the logging object
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 07:12:27 +00:00
yucheng-berri
d8d5437f55
Merge pull request #41558 from BerriAI/litellm_lit_6568_streaming_redaction
...
fix(guardrails): stream Prompt Security post_call redactions in incremental_diff mode
2026-09-17 00:10:45 -07:00
yucheng
b536289233
refactor(guardrails): type the request scan context helpers as read-only mappings
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:54:17 +00:00
yucheng
669a66499c
feat(policy_engine): bound priority to int32 and expose it in the Admin UI
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:47:53 +00:00
Yuneng Jiang
a40b6b3e44
test(management): close project lifecycle coverage gaps
2026-09-16 23:43:25 -07:00
berriai-litellm-provider-info-sync[bot]
5a82ed9bca
chore(prices): sync prices for 2 providers: 27 models
...
fireworks_ai/accounts/fireworks/models/minimax-m3: supports_vision
fireworks_ai/minimax-m3: supports_vision
wandb/deepseek-ai/DeepSeek-V4-Flash: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Flash-0731: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro: max_input_tokens
wandb/deepseek-ai/DeepSeek-V4-Pro-0813: max_input_tokens
wandb/google/gemma-4-31B-it: max_input_tokens
wandb/ibm-granite/granite-4.1-8b: max_input_tokens
wandb/ibm-granite/granite-4.2-8b: max_input_tokens
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-70B-Instruct: max_input_tokens
wandb/meta-llama/Llama-3.1-8B-Instruct: max_input_tokens
wandb/MiniMaxAI/MiniMax-M3: max_input_tokens
wandb/moonshotai/Kimi-K2.6: max_input_tokens
wandb/moonshotai/Kimi-K2.7-Code: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: max_input_tokens
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: max_input_tokens
wandb/openai/gpt-oss-120b: max_input_tokens
wandb/openai/gpt-oss-20b: max_input_tokens
wandb/OpenPipe/Qwen3-14B-Instruct: max_input_tokens
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: max_input_tokens
wandb/Qwen/Qwen3.5-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.6-27B: max_input_tokens
wandb/Qwen/Qwen3.6-35B-A3B: max_input_tokens
wandb/Qwen/Qwen3.8-27B: max_input_tokens
wandb/zai-org/GLM-5.2: max_input_tokens
wandb/zai-org/GLM-5.3-Flash: max_input_tokens
2026-09-17 06:31:05 +00:00
yuneng-jiang
5c486126ac
Merge pull request #41568 from BerriAI/litellm_/ci-test-failures-investigation-99785b
...
test(e2e): drop the auto-router select "opens below" spec
2026-09-16 23:23:36 -07:00
Devin AI
1b1f6ada46
fix(policy_engine): make priority migration idempotent
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:12:42 +00:00
Yuneng Jiang
4210f586c2
test(management): cover project authorization lifecycle
2026-09-16 23:08:00 -07:00
Devin AI
85444b56d9
fix(guardrails): hand the input scan context to the logging_only response scan
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:04:38 +00:00
yuneng-jiang
f58389c0e8
Merge pull request #41563 from BerriAI/litellm_budget-null-clear-tests
...
test(budgets): cover management null handling
2026-09-16 23:02:04 -07:00
berriai-litellm-provider-info-sync[bot]
b4212b949b
chore(prices): sync prices for 5 providers: 34 models, 1 new, 19 deprecated [1 with gaps]
...
fireworks_ai/accounts/fireworks/routers/glm-5p3-fast:
azure_ai/FW-Kimi-K3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/deepseek-ai/DeepSeek-R1-0528: deprecation_date
wandb/deepseek-ai/DeepSeek-V3-0324: deprecation_date
wandb/deepseek-ai/DeepSeek-V4-Flash: deprecation_date
wandb/deepseek-ai/DeepSeek-V4-Pro: deprecation_date
together_ai/deepseek-ai/DeepSeek-V4.1-Flash:
azure/eu/gpt-5.5-2026-04-24:
gemini-3.8-live: supports_response_schema
gemini-3.8-live-extended-thinking: supports_response_schema
azure/gpt-5.5-2026-04-24:
azure/gpt-5.6-luna-2026-07-09:
azure/gpt-5.6-sol-2026-07-09:
azure/gpt-5.6-terra-2026-07-09:
azure/gpt-6-astra-2026-09-03:
wandb/ibm-granite/granite-4.1-8b: deprecation_date
wandb/JetBrains/Mellum2-12B-A2.5B-Instruct: deprecation_date
wandb/meta-llama/Llama-3.1-70B-Instruct: deprecation_date
wandb/meta-llama/Llama-4-Scout-17B-16E-Instruct: deprecation_date
wandb/microsoft/Phi-4-mini-instruct: deprecation_date
wandb/MiniMaxAI/MiniMax-M2.5: deprecation_date
wandb/moonshotai/Kimi-K2-Instruct: deprecation_date
wandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
wandb/OpenPipe/Qwen3-14B-Instruct: deprecation_date
wandb/Qwen/Qwen3-235B-A22B-Instruct-2507: deprecation_date
wandb/Qwen/Qwen3-235B-A22B-Thinking-2507: deprecation_date
wandb/Qwen/Qwen3-30B-A3B-Instruct-2507: deprecation_date
wandb/Qwen/Qwen3-Coder-480B-A35B-Instruct: deprecation_date
wandb/Qwen/Qwen3.5-35B-A3B: deprecation_date
wandb/Qwen/Qwen3.6-27B: deprecation_date
azure/us/gpt-5.5-2026-04-24:
wandb/zai-org/GLM-4.5: deprecation_date
wandb/zai-org/GLM-5.3-Flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, supports_function_calling, supports_tool_choice, supports_response_schema, supports_prompt_caching, supports_reasoning
2026-09-17 06:01:14 +00:00
Devin AI
f5dea4de76
refactor(policy_engine): shorten attachment priority field descriptions
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 06:00:25 +00:00
Devin AI
21ffbdc7ea
feat(policy_engine): explicit priority for policy attachment execution order
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:59:01 +00:00
yucheng
3406913ca0
test(spend-logs): expect azure_spillover in spend log metadata golden
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:58:19 +00:00
yucheng
7b855bd53f
feat(spend-logs): record Azure spillover source deployment in spend log metadata
...
SpendLogsMetadata gains a typed azure_spillover key so a request Azure
served off pay-as-you-go capacity is visible in spend tracking, stamped
from the provider response headers or the processed llm_provider- headers
on the standard logging payload. The header parsing moves into a shared
azure_spillover() helper that is_spilled_over_ptu_request() now wraps.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:48:51 +00:00
yucheng
e3c8f74a4f
fix(proxy): default prompt injection heuristics executor to a single worker
...
SequenceMatcher holds the GIL, so extra heuristic threads add contention with the event loop without adding throughput. One worker drains scans in arrival order and keeps the loop responsive; PROMPT_INJECTION_HEURISTICS_MAX_THREADS remains an env override
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:45:33 +00:00
Yuneng Jiang
25445e8b5c
test(e2e): drop the auto-router select "opens below" spec
...
The spec pinned Base UI's collision behaviour, not our code: it only passes
while the template popup happens to fit under the trigger at 1280x900, and
#41315 's taller Add Auto Router form broke that premise for the second time
in three weeks. #41527 tried to scroll the trigger into the upper half, but
the dialog content is shorter than its max height, so nothing scrolls and CI
still fails 3/3 with the trigger at y=487
The guarantee #38554 introduced is that the popup never covers the trigger,
and the sibling spec keeps asserting that at a viewport with no room below
2026-09-16 22:43:51 -07:00
yucheng
2d925e5dde
fix(guardrails): scope the logging_only reply scan with the request's own translation
...
The chat-shaped output handler now takes the input translation as its
request scoping, so the logged request is scoped exactly once and with
the pre-call semantics of the surface it arrived on. This drops the
unscoped chat_shaped_request_conversation detour from af312dc8 , which
made the Anthropic response scan remove in-sequence system turns under
skip_system while the request scan kept them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:40:44 +00:00
yucheng
d6f6f64c0f
fix(proxy): derive prompt injection heuristics thread count from CPU count with env override
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:30:41 +00:00
Devin AI
af312dc8d7
fix(guardrails): scope the logging_only response scan once
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:30:06 +00:00
yucheng
a57483d1c8
fix(cost): price Azure PTU spillover requests at standard token rates
...
Azure PTU deployments carry zeroed per-token pricing because the reservation
is billed flat by the hour. When Azure spills a request onto pay-as-you-go
capacity it returns x-ms-is-spilled-over: true, and that traffic was still
priced at zero. The response cost calculator now detects the spillover header
on the result's hidden params or the logged provider response headers and
skips the zeroed custom pricing only for genuine PTU deployments while the
feature flag is on. Azure sync streaming now also records response headers on
the logging object, matching the async paths.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:24:05 +00:00
yucheng
060abd263e
fix(guardrails): keep usage chunk and defer tool_calls finish_reason behind held text in incremental_diff
...
A stream_options.include_usage usage chunk (empty delta plus usage) was folded into the final
transform round and rebuilt without its usage, so token counts and cost vanished from clients.
Metadata-only chunks are now replayed after the final text flush.
A terminal tool-call chunk arriving while earlier text was still held back carried
finish_reason=tool_calls ahead of that text. The finish_reason is now deferred to the final
text chunk whenever the choice has held text.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 05:11:04 +00:00