Commit graph

49802 commits

Author SHA1 Message Date
kerry
801f6a92b4 test(gemini): assert the served modelVersion reaches the assembled stream through CustomStreamWrapper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
dd4829398a fix(guardrails): log the configured mode when logging_only is mixed with an enforcing llm_as_a_judge mode
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
10506ab904 fix(guardrails): judge the whole latest user turn, run every during_call guardrail, log combined modes as configured
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
1b8b17cc8a fix(guardrails): judge only the latest request turn in pre_call and during_call llm_as_a_judge
The request-side prompt told the judge to focus on the most recent user turn but the text under review was every extracted request message joined together, so a multi-turn request with an off-topic earlier turn and an on-topic latest turn scored 50 and was blocked. Request-side judging now evaluates the last extracted request text (after the configured message scoping) and passes the full role-labelled conversation only as context. Response-side judging still evaluates all extracted response text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
e9357d9a6f fix(guardrails): stop logging_only llm_as_a_judge from judging its own judge calls
Judge sub-calls now carry the internal_call_origin metadata stamp and the
guardrail skips any logged call bearing it, so a logging_only judge no longer
recurses into an unbounded chain of judge requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
a658f20005 test(guardrails): grant premium_user for the tagged Mode case in llm_as_a_judge mode-shape test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
4d1330ea0b test(guardrails): drop callback-manager patch from llm_as_a_judge mode-shape test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
28bd00a004 fix(guardrails): label llm_as_a_judge logging_only verdicts with their mode
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
b93f22265c fix(guardrails): give the request-side judge role-labelled conversation context
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:53 +00:00
yucheng
6fccf00d93 fix(guardrails): keep list and tagged mode shapes for llm_as_a_judge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:52 +00:00
yucheng
0b34800501 fix(guardrails): resolve llm_as_a_judge request-side log mode from event_hook lists, drop unused alias
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:52 +00:00
yucheng
77a6327675 feat(guardrails): support pre_call and during_call modes for llm_as_a_judge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:05:52 +00:00
kerry
de9aa48cd6 fix(responses): count multimodal input and tool-call output in the streamed usage fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:02:29 +00:00
Tin Chi Lo
5aa2adbacd fix(proxy): preserve Anthropic pricing modifiers in router savings 2026-09-15 18:02:05 -07:00
kerry
405a127838 fix(fireworks-ai): drop banned typing.cast to a suppressed import for the copied pricing entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:57:05 +00:00
mateo-berri
3a8679d3d8 fix(proxy): never forward the LiteLLM virtual key to Anthropic on the /anthropic passthrough
The /anthropic/{endpoint} route forwarded every incoming header upstream, so
the header carrying the caller's LiteLLM virtual key (Authorization, x-api-key,
x-litellm-api-key, or the operator-configured key header) reached Anthropic and
was rejected there as an invalid credential, with or without a proxy-side
Anthropic key layered on top.

Share the Vertex credential-less header filter: drop the proxy-only credential
headers by name, drop the value that authenticated the caller (virtual key,
master key, or JWT) from Authorization / x-api-key, keep a caller's own
Anthropic credential, layer the proxy's Anthropic credential on top, and fail
with a clean 401 when neither the proxy nor the caller supplied one.

Resolves LIT-3550
2026-09-15 17:54:58 -07:00
mateo-berri
b4f9e319fc fix(proxy): surface the upstream status code when a RAG query fails 2026-09-15 17:53:57 -07:00
yucheng
9bd3f7b885 Merge remote-tracking branch 'origin/main' into litellm_agent365_mcp_guardrail 2026-09-16 00:52:27 +00:00
yucheng-berri
0b3e56448f
Merge pull request #41126 from BerriAI/litellm_custom_code_guardrail_identity
fix(guardrails): resolve caller identity from metadata buckets in custom code guardrail
2026-09-15 17:47:17 -07:00
yassin
14d239fc84 test(proxy): pin JWT scope denial message shape through auth exception conversion
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:46:59 +00:00
Yuneng Jiang
388eef4fbb
fix(e2e): onboard dashboard fixtures through invitations 2026-09-15 17:44:18 -07:00
kerry
5eab1feb20 fix(fireworks-ai): bill cache-write, reasoning, and audio tokens via the shared cost calculator
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:42:13 +00:00
Devin AI
232233f654 refactor(fireworks_ai): extract reasoning_effort mapping into helper
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:39:11 +00:00
Devin AI
a7a61db78d fix(fireworks_ai): flatten dict-form reasoning_effort to its effort string
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:37:55 +00:00
tin-berri
878716f806
Merge pull request #41326 from BerriAI/litellm_forecast_classifier_entitlement
feat(router): limit unlicensed Capability and Fuse v2 routers to one each
2026-09-15 17:36:22 -07:00
kerry
d5a36eb2ca fix(gemini): propagate provider modelVersion onto model responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:36:07 +00:00
kerry
ba6bcd747f chore: merge origin/main into test cleanup
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:33:25 +00:00
kerry
f5c1c82f81 fix(responses): estimate usage from text when streamed completed event omits usage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:32:24 +00:00
Mateo Wang
aa710dc6a9
Merge pull request #35388 from BerriAI/litellm_session_total_duration
fix(spend): sum multi-round session duration in logs UI
2026-09-15 17:31:41 -07:00
Joshua Valluru
61e3b5ddae fix(mcp): persist OAuth credentials for rowless JWT admins 2026-09-15 17:31:07 -07:00
kerry-berri
daa98cd6ff
Merge pull request #41320 from BerriAI/litellm_gemini_fallback_generalizations
feat(model_info): add provider-neutral Gemini 2.5+ chat baseline fallback generalization
2026-09-15 17:28:58 -07:00
Yassin Kortam
7eeba69016
Merge pull request #41316 from BerriAI/litellm_nvidia_nim_infer_passthrough
feat(proxy): add /nvidia_nim passthrough route for NIM object detection and OCR /v1/infer
2026-09-15 17:28:25 -07:00
kerry
4657d43fb9 fix(anthropic): tolerate message_delta chunks without usage in streams
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:27:09 +00:00
mateo
e62ab28561 chore(codeowners): add ryan and kerry as owners of the cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:26:34 +00:00
Yassin Kortam
5071d96d16
Merge pull request #41255 from BerriAI/litellm_org_member_spend_tracking
fix(proxy): track per-member organization spend
2026-09-15 17:25:40 -07:00
Yassin Kortam
e4a7d2aa0b
Merge pull request #41302 from BerriAI/litellm_fix_key_model_rpm_override_precedence
fix(proxy): key model rpm/tpm override takes precedence over team model limit
2026-09-15 17:24:38 -07:00
kerry-berri
95c0c39b05
Merge pull request #41318 from BerriAI/litellm_ws_responses_service_tier_pricing
fix(cost): price native Responses WebSocket turns at their returned service_tier
2026-09-15 17:23:35 -07:00
yassin
241b177f05 fix(proxy): log the internal model access denial reason for MCP sampling denials
MCP sampling catches the denial itself and returns ErrorData, so the central ProxyException handler never sees it. Log the sanitized internal reason there and share the CR/LF stripping through ModelAccessDeniedProxyException.sanitized_internal_message

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:22:48 +00:00
Yassin Kortam
3413efcd7f
Merge pull request #41306 from BerriAI/litellm_lit2435_litellm_log_error_silences_info
fix(proxy): honor LITELLM_LOG for uvicorn and proxy extras loggers
2026-09-15 17:22:39 -07:00
Yassin Kortam
636313e9fc
Merge pull request #41322 from BerriAI/litellm_gcp_video_usage
fix(vertex_ai): bill Gemini Omni Interactions usage and Veo sampleCount on passthrough
2026-09-15 17:21:15 -07:00
Joshua Valluru
72a4824acb fix(mcp): preserve JWT agent validation after main merge 2026-09-15 17:20:06 -07:00
ryan-crabbe-berri
f34a6eda92
Merge pull request #41294 from BerriAI/litellm_usage_export_gating
fix(ui): block usage export and flag the range when a spend page fails
2026-09-15 17:19:38 -07:00
Joshua Valluru
53318796fd fix(mcp): separate JWT identity lookup from request authorization 2026-09-15 17:13:10 -07:00
yassin
b7af51cc4a Merge remote-tracking branch 'origin/main' into litellm_model_access_denied_message 2026-09-16 00:12:53 +00:00
yassin
f60a603519 fix(proxy): log configured model access denials at the final response boundary
Post-auth denials from can_key_call_resolved_model (per-request alias
rewrite, MCP sampling, realtime) never reach the auth exception handler,
so the internal allowlist reason was dropped when
model_access_denied_message was set. Log it once from the ProxyException
response handler and the realtime rejection path instead, and convert
JWT ModelAccessDeniedHTTPException into the specialized ProxyException so
the same boundary covers it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:12:49 +00:00
yassin
f25272bfa0 fix(gateway): expose /nvidia_nim/ on the gateway data-plane allowlist
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 00:05:17 +00:00
yassin
65f9d9bbbd test(proxy): use the module-level HTTPException import in the /nvidia_nim route tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:57:16 +00:00
kerry
c9dd4b44f8 revert(model_info): keep Gemini baseline majors single-digit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 23:51:56 +00:00
Yassin Kortam
24153b5f29
Merge pull request #41308 from BerriAI/litellm_resolve_model_group_alias_before_auth 2026-09-15 16:51:28 -07:00
Joshua Valluru
88d0371a46 fix(mcp): reuse the standard JWT auth builder for OAuth ownership 2026-09-15 16:50:55 -07:00