Yassin Kortam
dbecd11d99
Merge pull request #41468 from BerriAI/litellm_key_alias_update_secret_sync
...
fix(proxy): rename AWS Secrets Manager secret when key alias changes
2026-09-16 13:20:53 -07:00
mateo-berri
6015437d67
test(bedrock): assert the Nova cache-read rate as a discount instead of pinning the vendor ratio
2026-09-16 13:19:58 -07:00
yassin
e7d537442d
fix(prometheus): read the cached customer row for request-time budget gauges
...
The request path used get_end_user_object, which falls back to a database
lookup on a cache miss. Read the LiteLLM_EndUserTable row auth already cached
instead, with the default budget already attached, and leave misses to the
periodic refresh
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 20:13:52 +00:00
Mateo Wang
365c6de875
Merge pull request #41340 from BerriAI/litellm_anthropic_passthrough_strip_virtual_key
...
fix(proxy): never forward the LiteLLM virtual key to Anthropic on the /anthropic passthrough
2026-09-16 13:12:42 -07:00
Mateo Wang
737929f338
Merge pull request #41335 from BerriAI/litellm_fireworks_dict_reasoning_effort
...
fix(fireworks_ai): flatten dict-form reasoning_effort to its effort string
2026-09-16 13:12:37 -07:00
yassin
1979533901
test(router): force the callback to observe the usage stamp before the pre-header increment fails
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 20:10:22 +00:00
yujonglee
560065df16
Merge pull request #41480 from BerriAI/litellm_rust_ci_nextest_rust_cache
...
ci(rust): split rust jobs, use nextest and Swatinem/rust-cache
2026-09-16 13:08:20 -07:00
Mateo Wang
2b33201a09
Merge pull request #41475 from BerriAI/litellm_bedrock_kb_user_context
...
fix(bedrock): forward userContext in Knowledge Base Retrieve requests
2026-09-16 13:07:03 -07:00
Devin AI
a229f99ac8
test(models): drive Grok prompt caching coverage through litellm APIs
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:59:07 +00:00
yassin
9ebd55e53e
test(router): cover success callback recovering the count when the pre-header increment fails
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:57:55 +00:00
mateo-berri
884087f01c
test(bedrock): type the vector store search test helper
2026-09-16 12:55:29 -07:00
Yujong Lee
48df3d5a48
ci(rust): install nextest via pinned taiki-e/install-action
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:54:46 +00:00
mateo-berri
043c954aa9
fix(proxy): keep a caller's own Anthropic key when the proxy has no master key
...
Without a master key the auth layer echoes whatever key the caller presented as the authenticated key, so the passthrough's strip-by-value matched the caller's own Anthropic key and dropped it: a bring-your-own-key request that returned 200 on main answered 401 telling the caller to send the key they had just sent. Only the auth module's own no-auth dev-mode definition, shared through is_no_auth_dev_mode, decides that nothing was authenticated, and only when no custom auth is installed; JWTs, OAuth2 tokens, and custom-auth credentials are still stripped there. The sk- prefix heuristic goes with it.
The Vertex credential-less test now sets a master key, since a virtual key can only authenticate under one: the auth layer returns before any key lookup when the master key is unset.
2026-09-16 12:52:37 -07:00
mateo-berri
a2724e7f15
fix(otel): drop None attributes before they reach the OTLP encoder
...
The metric attribute filter now removes every attribute whose value is
None, and the content and inference-details events pass their attributes
through drop_none before emitting, so a call with no provider label or no
model name never hands the OTLP exporter a NoneType attribute. This closes
the gen_ai.request.model report on #36759 the same way the gen_ai.system
one was closed, and the regression tests cover both keys.
2026-09-16 12:52:14 -07:00
Devin AI
ba03f60f71
fix(models): add batch and flex tier prices on Azure dated and Gemini latest alias rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:50:28 +00:00
Yujong Lee
3f15dcd96b
ci(rust): fold fmt into the clippy job
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:47:48 +00:00
yassin
7ff8945da4
fix(router): sync both usage keys from Redis even when one increment is zero
...
A stream counted before its usage is known increments TPM by zero, so the
worker that served it never refreshed its local TPM value from Redis and
the first byte headers reported the token count another worker had already
consumed. Both pipeline operations now always run, matching the pre-change
callback, so the returned values refresh both worker local keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:45:48 +00:00
Yujong Lee
e850232f02
test(rust): assert merge cost scales linearly instead of a wall-clock bound
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:43:42 +00:00
yassin
0d605b7b45
Merge remote-tracking branch 'origin/main' into litellm_streaming_buffer_release_on_scan
2026-09-16 19:42:36 +00:00
mateo-berri
0dff64ce1a
fix(bedrock): move a Nova invoke cache point behind an image or tool result to the last text block
2026-09-16 12:42:15 -07:00
kerry-berri
b6143b3711
Merge pull request #41457 from BerriAI/litellm-providers/price-sync
...
chore(prices): sync Google Gemini prices: 22 models
2026-09-16 12:37:16 -07:00
yassin
6faeb19c0c
Merge remote-tracking branch 'origin/main' into litellm_streaming_buffer_release_on_scan
2026-09-16 19:33:16 +00:00
yassin
1e169bb912
refactor(guardrails): mark release_on_scan flag as Final
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:33:07 +00:00
mateo-berri
c052d7816e
fix(bedrock): rename the cache point inliner so the recursion detector stops flagging it
2026-09-16 12:30:33 -07:00
mateo-berri
aa6d053a37
Merge remote-tracking branch 'origin/main' into litellm_otel_gen_ai_system_none
2026-09-16 12:30:20 -07:00
yassin
cb26454034
test(router): cover success callback racing the pre-header count and drop redundant docstrings
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:28:50 +00:00
Mateo Wang
a1b1f1ef9d
Merge pull request #34455 from BerriAI/litellm_lit4767_empty_choices_streaming_guard
...
fix(responses): guard empty-choices chunks in the Responses API streaming bridge
2026-09-16 12:28:46 -07:00
yassin
8de51dfaab
test(otel): describe which request metadata the pre-call seed reads
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:25:40 +00:00
yassin
938b782b8a
Merge remote-tracking branch 'origin/main' into litellm_customer_budget_prometheus_metrics
2026-09-16 19:25:23 +00:00
yassin
d5837ab97a
test(guardrails): pin tool-call-only scan keys as non-empty
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:24:14 +00:00
yassin
c22916377f
test(prometheus): cover customer budget series expiry by end_user ttl
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:22:30 +00:00
Yujong Lee
e8b5632c20
ci(rust): run the token counter timing test alone under nextest
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:22:02 +00:00
kerry
e46106e20b
test: drop gemini-3.1-flash-lite-image capability pins
...
The per-route capability test hardcoded vendor facts, including function calling support on the gemini route, which the live model card says is not supported. Keep the backup-matches-main invariant and the routing tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:21:50 +00:00
Devin AI
93eefc5922
fix(models): dedupe Fireworks deprecation keys and sync Azure dated snapshot service tiers
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:20:20 +00:00
Devin AI
8caf2cb61f
fix(models): flag prompt caching on Vertex and Azure AI Grok rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:17:45 +00:00
Mateo Wang
734038a83c
Merge pull request #41333 from BerriAI/litellm_cost_json_codeowners
...
chore(codeowners): add ryan and kerry as owners of the cost map
2026-09-16 12:17:13 -07:00
Devin AI
7734e3e186
fix(models): sonnet 4.5 1M input, daybreak alias, Mistral GLM 5.3, Azure dated snapshots, Together/OpenRouter sync
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:17:13 +00:00
Yujong Lee
b70ddc2fd8
ci(rust): split rust jobs, use nextest and Swatinem/rust-cache
...
Split the Rust workflow into fmt, clippy, nextest and wheel jobs so they run in parallel, replace manual actions/cache with Swatinem/rust-cache, and install a pinned checksum-verified cargo-nextest. Make two python-bridge tests self-contained so they pass when nextest runs each test in its own process.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:14:32 +00:00
yassin
30f02aa6da
fix(otel): read only the caller's requester_metadata snapshot in the v2 pre-call hook
...
The pre-call hook passed the proxy's whole per-request metadata dict into the
request identity, so proxy-owned siblings such as requester_ip_address were
promoted alongside the caller's keys. Only the requester_metadata mapping is
read now, keyed under its wrapper, which keeps the default allowlist behaviour
unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:11:25 +00:00
yassin
73a0c3bb8a
Merge remote-tracking branch 'origin/main' into litellm_streaming_buffer_release_on_scan
2026-09-16 19:09:19 +00:00
yassin
2821ba9667
fix(prometheus): refresh default-budget customers and honor independent customer gauges
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:08:42 +00:00
berriai-litellm-provider-info-sync[bot]
b893e6b926
chore(prices): sync Google Gemini prices: 22 models
...
gemini/gemini-2.5-flash: max_tokens, max_output_tokens, supports_audio_input
gemini/gemini-2.5-flash-image: max_input_tokens, supports_web_search, supports_audio_input, supports_response_schema, supports_function_calling
gemini/gemini-2.5-flash-lite: max_tokens, max_output_tokens, supports_audio_input
gemini-2.5-flash-native-audio-preview-12-2025: supports_vision, max_input_tokens, supports_web_search, supports_response_schema, supports_function_calling
gemini/gemini-2.5-flash-native-audio-preview-12-2025: supports_vision, max_input_tokens, supports_web_search, supports_response_schema, supports_function_calling
gemini-2.5-flash-preview-tts: max_tokens, max_input_tokens, max_output_tokens, supports_web_search, supports_audio_input, supports_response_schema, supports_function_calling
gemini/gemini-2.5-flash-preview-tts: max_tokens, max_input_tokens, max_output_tokens, supports_web_search, supports_audio_input, supports_response_schema, supports_function_calling
gemini/gemini-2.5-pro: max_tokens, max_output_tokens
gemini/gemini-2.5-pro-preview-tts: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_web_search, supports_audio_input, supports_response_schema, supports_function_calling
gemini/gemini-3-flash-preview: max_tokens, max_output_tokens, supports_audio_input
gemini/gemini-3-pro-image: supports_response_schema
gemini/gemini-3.1-flash-image: max_input_tokens, supports_response_schema
gemini/gemini-3.1-flash-lite-image: supports_web_search, supports_function_calling
gemini-3.1-flash-live-preview: supports_response_schema
gemini/gemini-3.1-flash-live-preview: supports_response_schema
gemini/gemini-3.1-flash-tts-preview: supports_web_search, supports_response_schema, supports_function_calling
gemini/gemini-3.5-flash: max_tokens, max_output_tokens
gemini/gemini-3.5-live-translate-preview: supports_web_search, supports_response_schema, supports_function_calling
gemini/gemini-3.5-transcribe: supports_function_calling
gemini/gemini-3.5-transcribe-live: supports_function_calling
gemini/gemini-embedding-2: supports_vision, supports_audio_input
gemini/gemini-omni-1.1-flash: max_input_tokens
2026-09-16 19:06:19 +00:00
Devin AI
03931180d7
merge main into litellm_registry_audit_2026_09_14
2026-09-16 19:03:36 +00:00
yassin
e105cde56a
fix(router): sync worker-local usage cache to shared post-increment counters
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:02:15 +00:00
yassin
5e2d9e1d5c
refactor(otel): build promoted baggage without local dict mutation
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:59:21 +00:00
yucheng-berri
b04d530ecf
Merge pull request #41386 from BerriAI/litellm_otel_trace_correlation
...
fix(proxy): default litellm_trace_id to the OTel server span trace id
2026-09-16 11:58:45 -07:00
yassin
ad25fe1886
fix(guardrails): hold unscannable Responses windows and key terminal envelopes by output items
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:56:56 +00:00
mateo-berri
878fe17735
chore(proxy): restore the lazy OpenAPI snapshot main renders under Python 3.12
2026-09-16 11:56:18 -07:00
yassin
a2108f02fb
test(router): cover _increment_deployment_usage delta and unlimited deployment behavior
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:51:40 +00:00
mateo-berri
033aa8ba6d
fix(bedrock): forward userContext in Knowledge Base Retrieve requests
...
The Bedrock vector store search only lifted retrievalConfiguration out of extra_body, so the caller's userContext (the Retrieve API's ACL identity) never reached Bedrock and ACL-enabled data sources answered with zero results. The transform now forwards userContext, taken from extra_body first and then from the top-level params where the OpenAI SDK's extra_body merge lands, as the caller sent it.
2026-09-16 11:50:36 -07:00