yassin
8de51dfaab
test(otel): describe which request metadata the pre-call seed reads
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:25:40 +00:00
yassin
30f02aa6da
fix(otel): read only the caller's requester_metadata snapshot in the v2 pre-call hook
...
The pre-call hook passed the proxy's whole per-request metadata dict into the
request identity, so proxy-owned siblings such as requester_ip_address were
promoted alongside the caller's keys. Only the requester_metadata mapping is
read now, keyed under its wrapper, which keeps the default allowlist behaviour
unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 19:11:25 +00:00
yassin
5e2d9e1d5c
refactor(otel): build promoted baggage without local dict mutation
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:59:21 +00:00
yassin
8a059cd4b4
fix(otel): promote nested metadata keys under the caller's dotted path
...
Strip only the proxy's requester_metadata. wrapper from an allowlisted key so
requester_metadata.trace_id lands as litellm.metadata.trace_id while other
dotted keys keep their full path and cannot collide on a shared leaf name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:39:55 +00:00
yassin
8cab3a7846
refactor(otel): walk nested metadata iteratively instead of recursively
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:20:10 +00:00
yassin
17059564a8
feat(otel): promote nested request metadata keys to litellm.metadata.* span attributes
...
baggage_metadata_keys entries such as requester_metadata.trace_id now resolve the caller's nested metadata.trace_id and stamp it on the LLM-call span as litellm.metadata.trace_id, in both the OTEL v2 logger and the legacy OpenTelemetry callback. Nested metadata mappings are flattened to dotted paths, only allowlisted leaves are promoted, and the requester_metadata blob itself is never promoted
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 18:08:14 +00:00
kerry-berri
4e996400e2
Merge pull request #41154 from BerriAI/litellm-providers/price-sync
...
chore(prices): sync Azure, Azure AI, Gemini, OpenAI, Bedrock, Together AI, Fireworks and Vertex prices: 278 models, 59 new, 30 deprecated
2026-09-16 10:37:21 -07:00
Yassin Kortam
8aebd4ff63
Merge pull request #40571 from BerriAI/litellm_presidio_mcp_mode_no_post_call_scan
...
fix(guardrails): don't add post_call output scan for MCP-only Presidio modes
2026-09-16 10:09:12 -07:00
yujonglee
a7cfdd4cd5
Merge pull request #41432 from BerriAI/litellm_remove_rust_gateway_router_realtime
...
refactor(rust): remove gateway, config, router, realtime, and Rust trace-parity instrumentation
2026-09-16 10:06:31 -07:00
Yassin Kortam
8491d01668
Merge pull request #41268 from BerriAI/litellm_outbound_http2_opt_in
...
feat(http): opt-in outbound HTTP/2 for httpx clients
2026-09-16 10:05:37 -07:00
kerry
b9ab36279b
fix(prices): add tpm and rpm to gemini 3.8 live rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:49:17 +00:00
kerry
eed760a37d
test: drop azure gpt-5.6 rate tests pinned to a dated price page
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:49:15 +00:00
yassin
8a43fed20c
fix(guardrails): use explicit returns in _configured_event_hooks
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:48:15 +00:00
Yassin Kortam
f0474bb70e
Merge pull request #41310 from BerriAI/litellm_model_access_denied_message
...
fix(proxy): hide model allowlist from client-facing model access denied errors
2026-09-16 09:38:29 -07:00
kerry
cab6732928
test: drop price-pinning tests that break on catalog updates
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:32:37 +00:00
yassin
7b3582aa66
fix(guardrails): treat tag-based Mode as MCP-only when all hooks are MCP hooks
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:31:44 +00:00
yassin
e48dde7b9b
fix(tests): drop leftover merge markers in test_handle_jwt
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:28:13 +00:00
yassin
5fee1c8710
Merge remote-tracking branch 'origin/main' into litellm_model_access_denied_message
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
# Conflicts:
# tests/test_litellm/proxy/auth/test_handle_jwt.py
2026-09-16 16:27:18 +00:00
berriai-litellm-provider-info-sync[bot]
809685603a
chore(prices): sync Azure prices: 247 models
...
azure_ai/Codestral-2501:
azure_ai/cohere-command-a:
azure_ai/deepseek-r1:
azure_ai/deepseek-v3:
azure_ai/deepseek-v3-0324:
azure_ai/deepseek-v3.1:
azure_ai/deepseek-v3.2:
azure_ai/deepseek-v3.2-speciale:
azure_ai/deepseek-v4-flash:
azure_ai/DeepSeek-V4-Flash-0731:
azure_ai/deepseek-v4-pro:
azure_ai/embed-v-4-0:
azure_ai/FW-DeepSeek-V3.2:
azure_ai/FW-DeepSeek-V4-Pro:
azure_ai/FW-GLM-5:
azure_ai/FW-GLM-5.1:
azure_ai/FW-GLM-5.2:
azure_ai/FW-GLM-5.2-Fast:
azure_ai/FW-Inkling:
azure_ai/FW-Kimi-K2.5:
azure_ai/FW-Kimi-K2.6:
azure_ai/FW-Kimi-K2.7-Code:
azure_ai/FW-Kimi-K3:
azure_ai/FW-MiniMax-M2.5:
azure_ai/FW-MiniMax-M3:
azure_ai/FW-Nemotron-3-Ultra-NVFP4:
azure_ai/FW-Nemotron-Lightning-3.5-30B-A3B:
azure_ai/gpt-oss-120b:
azure_ai/grok-3:
azure_ai/global/grok-3:
azure_ai/grok-3-mini:
azure_ai/global/grok-3-mini:
azure_ai/grok-4:
azure_ai/grok-4-1-fast-non-reasoning:
azure_ai/grok-4-1-fast-reasoning:
azure_ai/grok-4-20-non-reasoning:
azure_ai/grok-4-20-reasoning:
azure_ai/grok-4-fast-non-reasoning:
azure_ai/grok-4-fast-reasoning:
azure_ai/grok-4.3:
azure_ai/grok-4.6:
azure_ai/grok-code-fast-1:
azure_ai/kimi-k2.5:
azure_ai/kimi-k2.6:
azure_ai/kimi-k2.7-code:
azure_ai/Llama-3.3-70B-Instruct:
azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8:
azure_ai/MAI-DS-R1:
azure_ai/MAI-Image-2.5:
azure_ai/MAI-Image-2.5-Flash:
azure_ai/MAI-Image-2e:
azure_ai/MAI-Thinking-1:
azure_ai/mistral-large-3:
azure_ai/Phi-3-medium-128k-instruct:
azure_ai/Phi-3-medium-4k-instruct:
azure_ai/Phi-3-mini-128k-instruct:
azure_ai/Phi-3-mini-4k-instruct:
azure_ai/Phi-3-small-128k-instruct:
azure_ai/Phi-3-small-8k-instruct:
azure_ai/Phi-3.5-mini-instruct:
2026-09-16 16:26:23 +00:00
Yassin Kortam
e941edd08a
Merge pull request #41407 from BerriAI/litellm_content_filter_stream_bounded_scan
...
perf(content_filter): scan a bounded window per streamed chunk
2026-09-16 09:25:23 -07:00
kerry
d398c9199c
Merge remote-tracking branch 'origin/main' into HEAD
2026-09-16 16:18:47 +00:00
jesus
47117d880c
fix(guardrails): don't add post_call output scan for MCP-only Presidio modes
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:15:49 +00:00
Yujong Lee
6330d80efa
refactor(rust): glob workspace members
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:15:46 +00:00
Yassin Kortam
29a712b186
Merge pull request #41368 from BerriAI/litellm_shadow_streaming_multi_target
...
feat(router): stream shadow traffic and fan out silent_model to multiple targets
2026-09-16 09:08:24 -07:00
Yujong Lee
96baeb8b04
refactor(rust): remove gateway, config, router, realtime, and trace-parity infrastructure
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:00:07 +00:00
joshua-berri
9cd787386e
Merge pull request #41314 from BerriAI/litellm_fix_mcp_jwt_oauth_persistence
...
fix(mcp): authorize JWT OAuth credential persistence
2026-09-16 06:47:40 -07:00
berriai-litellm-provider-info-sync[bot]
0443605c40
chore(prices): sync Together AI prices: 4 models, 4 deprecated
...
together_ai/google/gemma-4-31B-it: deprecation_date
together_ai/intfloat/multilingual-e5-large-instruct: deprecation_date
together_ai/openai/gpt-oss-20b: deprecation_date
together_ai/thinkingmachines/Inkling-Small: deprecation_date
2026-09-16 13:16:12 +00:00
yassin
a8fff5b091
fix(content_filter): refuse a trim that splits a conditional word across the cut
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 12:45:58 +00:00
yassin
4fbe631146
test(content_filter): annotate streaming test locals as Final and type the logging metadata
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 12:17:47 +00:00
yassin
60642e875b
perf(content_filter): back off refused streamed buffer cuts by one context length
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 11:51:06 +00:00
yassin
7e429dee87
fix(content_filter): keep exception phrases and open conditional sentences in the streamed buffer
...
Trimming the streamed buffer to the retained tail could drop a category
exception phrase that suppresses a later keyword, or the identifier word
of an unfinished sentence that a conditional category pairs with a later
block word. Refuse the cut while either would leave the buffer so the
bounded scan masks and blocks exactly like a scan of the full text
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 11:18:10 +00:00
yassin
52f06906fe
fix(content_filter): widen the streamed scan tail to the longest configured keyword
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 10:58:30 +00:00
yassin
62ecb11ab9
perf(content_filter): scan a bounded window per streamed chunk
...
The streaming post-call hook rescanned the whole accumulated choice buffer on every chunk, so scan cost grew quadratically with output length. Keep a bounded per-choice buffer instead: once it exceeds twice the scan context, drop the head when masking the head and tail separately yields the same output as masking the whole buffer, so no pattern, phrase or exception straddles the cut. Detections from the dropped head are kept and merged, deduplicated, into the final log row
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 09:59:53 +00:00
Joshua Valluru
ee676d59f2
fix(mcp): preserve browser OAuth for unrelated bearer tokens
2026-09-15 23:08:14 -07:00
Joshua Valluru
e035682ed1
refactor(auth): separate JWT identity and OAuth authorization
2026-09-15 22:15:57 -07:00
yassin
252c69b532
fix(router): snapshot shadow kwargs per target so concurrent shadows never share metadata
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 04:31:43 +00:00
tin-berri
a8979fe054
Merge pull request #41371 from BerriAI/litellm_forecast_simplify_advanced_ui
...
fix(ui): simplify Capability and Fuse advanced routing options
2026-09-15 21:19:50 -07:00
kerry-berri
5960881640
Merge pull request #41337 from BerriAI/litellm_fix_responses_stream_absent_usage_recount
...
fix(responses): recount tokens when a streamed response completes without usage
2026-09-15 21:19:42 -07:00
yuneng-jiang
bbffddd517
Merge pull request #41366 from BerriAI/litellm_/back-002-litellm-e2e-replay-8de5b1
...
fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live
2026-09-15 21:18:14 -07:00
yuneng-jiang
e3c1f78e28
Merge branch 'main' into litellm_/back-002-litellm-e2e-replay-8de5b1
2026-09-15 21:08:29 -07:00
yassin
0f7ed4433b
fix(router): snapshot shadow kwargs before fan-out so shadows never see primary mutations
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 04:04:01 +00:00
Tin Chi Lo
9d640b86ed
fix(ui): simplify Capability and Fuse routing options
2026-09-15 20:58:38 -07:00
kerry
ff878e7df0
refactor(responses): copy the terminal event instead of mutating stubbed chunks
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:54:04 +00:00
yassin
9cabde90dd
test(router): cover _run_silent_experiment directly for router coverage gate
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:48:32 +00:00
kerry
7cc07d437a
fix(responses): build the billed terminal response immutably and guard the cache dump
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:47:03 +00:00
berriai-litellm-provider-info-sync[bot]
641dcd5f8f
chore(prices): sync Google Gemini prices: 1 model
...
gemini/gemini-robotics-er-2-preview: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
2026-09-16 03:46:34 +00:00
yuneng-jiang
174c1ac4ed
Merge pull request #41348 from BerriAI/litellm_fix_models_reload_test
...
fix(e2e): expect models filters to persist after reload
2026-09-15 20:31:47 -07:00
yassin
06fcc1f733
feat(router): stream shadow traffic and fan out silent_model to multiple targets
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:28:08 +00:00
yuneng-jiang
9d54c2e159
Merge pull request #41359 from BerriAI/litellm_fix_budget_reset_decrement_tests
...
test(proxy): assert budget resets decrement the cleared spend
2026-09-15 20:27:35 -07:00
yuneng-jiang
ef06790e39
Merge pull request #41358 from BerriAI/litellm_fix_stale_router_tests
...
test(router): ignore deployment-selection logs in the fallback log assertion
2026-09-15 20:27:21 -07:00