Commit graph

49857 commits

Author SHA1 Message Date
Yassin Kortam
29a712b186
Merge pull request #41368 from BerriAI/litellm_shadow_streaming_multi_target
feat(router): stream shadow traffic and fan out silent_model to multiple targets
2026-09-16 09:08:24 -07:00
yassin
04aad317c7 Merge remote-tracking branch 'origin/main' into pr34829 2026-09-16 16:02:08 +00:00
Yujong Lee
96baeb8b04 refactor(rust): remove gateway, config, router, realtime, and trace-parity infrastructure
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 16:00:07 +00:00
joshua-berri
9cd787386e
Merge pull request #41314 from BerriAI/litellm_fix_mcp_jwt_oauth_persistence
fix(mcp): authorize JWT OAuth credential persistence
2026-09-16 06:47:40 -07:00
Devin AI
ba820783ac fix(tests): import Bedrock usage types
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:25:14 +00:00
Devin AI
642525f571 fix(registry): add azure gpt-image-2.5 entries and together/azure deprecation dates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 13:19:53 +00:00
berriai-litellm-provider-info-sync[bot]
0443605c40
chore(prices): sync Together AI prices: 4 models, 4 deprecated
together_ai/google/gemma-4-31B-it: deprecation_date
together_ai/intfloat/multilingual-e5-large-instruct: deprecation_date
together_ai/openai/gpt-oss-20b: deprecation_date
together_ai/thinkingmachines/Inkling-Small: deprecation_date
2026-09-16 13:16:12 +00:00
Devin AI
4fd69b25e1 Merge remote-tracking branch 'origin/main' into litellm_registry_audit_2026_09_14 2026-09-16 13:05:28 +00:00
yassin
a8fff5b091 fix(content_filter): refuse a trim that splits a conditional word across the cut
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 12:45:58 +00:00
yassin
4fbe631146 test(content_filter): annotate streaming test locals as Final and type the logging metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 12:17:47 +00:00
yassin
60642e875b perf(content_filter): back off refused streamed buffer cuts by one context length
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 11:51:06 +00:00
yassin
7e429dee87 fix(content_filter): keep exception phrases and open conditional sentences in the streamed buffer
Trimming the streamed buffer to the retained tail could drop a category
exception phrase that suppresses a later keyword, or the identifier word
of an unfinished sentence that a conditional category pairs with a later
block word. Refuse the cut while either would leave the buffer so the
bounded scan masks and blocks exactly like a scan of the full text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 11:18:10 +00:00
yassin
52f06906fe fix(content_filter): widen the streamed scan tail to the longest configured keyword
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 10:58:30 +00:00
yassin
62ecb11ab9 perf(content_filter): scan a bounded window per streamed chunk
The streaming post-call hook rescanned the whole accumulated choice buffer on every chunk, so scan cost grew quadratically with output length. Keep a bounded per-choice buffer instead: once it exceeds twice the scan context, drop the head when masking the head and tail separately yields the same output as masking the whole buffer, so no pattern, phrase or exception straddles the cut. Detections from the dropped head are kept and merged, deduplicated, into the final log row

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 09:59:53 +00:00
yucheng
3628025aae fix(proxy): keep caller metadata.trace_id ahead of the OTel fallback on litellm_metadata routes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 08:30:54 +00:00
yucheng
898fbd37a7 test(proxy): mark locals Final in the OTel trace id fallback tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 08:02:04 +00:00
yucheng
4cb4493fa7 fix(proxy): let the OTel trace id fallback fill a null litellm_trace_id
A body that serializes litellm_trace_id as null or an empty string carries no identity, so it must not
block the server span fallback. Also mark the nested metadata write as an out-param store

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:48:58 +00:00
yucheng
49417d4fa2 fix(proxy): ignore non-span parent_otel_span when deriving litellm_trace_id
UserAPIKeyAuth.parent_otel_span is Any at runtime (opentelemetry is an optional extra), so the OTel
trace-id fallback must only format an int trace id, otherwise an object that merely quacks like a span
turns the whole request into a 500

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:41:48 +00:00
yucheng
f1fd1c8996 fix(proxy): default litellm_trace_id to the OTel server span trace id
When the otel callback is enabled and the client sends no trace or session identity, the request now inherits the W3C trace id of the proxy's server span as litellm_trace_id and metadata.trace_id. The missing_session_id policy and SpendLogs then persist that value as session_id, so a trace in the OTel backend and its row in the Logs UI carry the same id. Explicit x-litellm-trace-id, traceparent, body metadata.trace_id and litellm_trace_id keep priority.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 07:22:03 +00:00
Joshua Valluru
ee676d59f2 fix(mcp): preserve browser OAuth for unrelated bearer tokens 2026-09-15 23:08:14 -07:00
Joshua Valluru
e035682ed1 refactor(auth): separate JWT identity and OAuth authorization 2026-09-15 22:15:57 -07:00
yassin
252c69b532 fix(router): snapshot shadow kwargs per target so concurrent shadows never share metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 04:31:43 +00:00
tin-berri
a8979fe054
Merge pull request #41371 from BerriAI/litellm_forecast_simplify_advanced_ui
fix(ui): simplify Capability and Fuse advanced routing options
2026-09-15 21:19:50 -07:00
kerry-berri
5960881640
Merge pull request #41337 from BerriAI/litellm_fix_responses_stream_absent_usage_recount
fix(responses): recount tokens when a streamed response completes without usage
2026-09-15 21:19:42 -07:00
yuneng-jiang
bbffddd517
Merge pull request #41366 from BerriAI/litellm_/back-002-litellm-e2e-replay-8de5b1
fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live
2026-09-15 21:18:14 -07:00
yuneng-jiang
e3c1f78e28
Merge branch 'main' into litellm_/back-002-litellm-e2e-replay-8de5b1 2026-09-15 21:08:29 -07:00
yassin
0f7ed4433b fix(router): snapshot shadow kwargs before fan-out so shadows never see primary mutations
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 04:04:01 +00:00
Tin Chi Lo
9d640b86ed fix(ui): simplify Capability and Fuse routing options 2026-09-15 20:58:38 -07:00
kerry
ff878e7df0 refactor(responses): copy the terminal event instead of mutating stubbed chunks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:54:04 +00:00
yassin
9cabde90dd test(router): cover _run_silent_experiment directly for router coverage gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:48:32 +00:00
kerry
7cc07d437a fix(responses): build the billed terminal response immutably and guard the cache dump
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:47:03 +00:00
berriai-litellm-provider-info-sync[bot]
641dcd5f8f
chore(prices): sync Google Gemini prices: 1 model
gemini/gemini-robotics-er-2-preview: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, input_cost_per_token_batches, output_cost_per_token_batches
2026-09-16 03:46:34 +00:00
yuneng-jiang
174c1ac4ed
Merge pull request #41348 from BerriAI/litellm_fix_models_reload_test
fix(e2e): expect models filters to persist after reload
2026-09-15 20:31:47 -07:00
yassin
06fcc1f733 feat(router): stream shadow traffic and fan out silent_model to multiple targets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:28:08 +00:00
yuneng-jiang
9d54c2e159
Merge pull request #41359 from BerriAI/litellm_fix_budget_reset_decrement_tests
test(proxy): assert budget resets decrement the cleared spend
2026-09-15 20:27:35 -07:00
yuneng-jiang
ef06790e39
Merge pull request #41358 from BerriAI/litellm_fix_stale_router_tests
test(router): ignore deployment-selection logs in the fallback log assertion
2026-09-15 20:27:21 -07:00
Yuneng Jiang
a6fb21c3f8
fix(e2e): record cookie-setting provider responses and keep prompt-caching tests live
The first cache-enabled litellm-e2e build (211) showed three gaps in the shared provider cache:

Every OpenAI response carries Cloudflare bot-management Set-Cookie headers, and the capture rejected any response with Set-Cookie, so no OpenAI response was ever recorded (179 of 372 misses rejected). The edge already withholds Set-Cookie from the proxy, so drop it before validating and storing instead of rejecting.

The provider prompt-caching tests need fresh provider state: a replayed priming response reports cache creation rather than a cache read, and the TPM test then trips the key limit. Mark both modules provider_live.

TestApiBaseSeam::test_live_mode_returns_none ran inside the cache-enabled runner and saw the shared edge; isolate it from E2E_PROVIDER_CACHE.
2026-09-15 20:26:20 -07:00
kerry-berri
4b84fa9230
Merge pull request #41336 from BerriAI/litellm_fix_anthropic_stream_absent_usage
fix(anthropic): tolerate message_delta events without usage when streaming
2026-09-15 20:23:01 -07:00
tin-berri
3b14631f06
Merge pull request #41315 from BerriAI/litellm_capability_fuse_classifier_ui
feat(ui): configure capability and Fuse v2 classifiers
2026-09-15 20:13:57 -07:00
kerry
7121e64db4 fix(responses): type the dict terminal response so the estimated usage is billed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 03:13:53 +00:00
Yassin Kortam
d108cdc431
Merge pull request #38254 from BerriAI/litellm_fix_xai_responses_instructions 2026-09-15 19:55:52 -07:00
Yuneng Jiang
a8fba14e10
test(proxy): assert budget resets decrement the cleared spend
The reset job moved from zeroing spend to an atomic decrement of the amount
it cleared, so every batched write now carries {"decrement": <cleared>}
instead of 0. Four tests still pinned 0 and had been failing since, which
also meant they no longer checked the amount at all. Assert the decrement
equals each row's own pre-reset spend, so a wrong amount fails the test.
2026-09-15 19:48:25 -07:00
yuneng-jiang
65d607d161
Merge pull request #41346 from BerriAI/litellm_e2e_provider_cache
feat(e2e): reuse exact provider responses for 24 hours
2026-09-15 19:45:56 -07:00
Yuneng Jiang
6fd988cc63
test(router): ignore deployment-selection logs in the fallback log assertion
simple_shuffle logs the selected deployment at INFO whenever a weight set
applies, so the fallback group's selection line lands between the fallback
notice and the success notice and pushed the notice out of the tail-3 window.
Filter it the same way the neighbouring get_available_deployment noise is
already filtered.
2026-09-15 19:45:33 -07:00
yuneng-jiang
c64b821a7e
Merge pull request #41353 from BerriAI/litellm_/litellm-exception-oct-15-e61cb0
ci(image-scan): ignore zlib CVE-2026-85091 until Wolfi ships the fix
2026-09-15 19:32:53 -07:00
Yuneng Jiang
9170183087
fix(e2e): exclude unknown routes from upstream counters 2026-09-15 19:32:50 -07:00
Tin Chi Lo
39b7812810 feat(ui): configure capability and Fuse v2 classifiers 2026-09-15 19:30:50 -07:00
Yuneng Jiang
3dda798c8a
ci(e2e): consolidate cache contracts in filtered CircleCI job 2026-09-15 19:28:50 -07:00
yucheng-berri
c14ac1dbe9
Merge pull request #41329 from BerriAI/litellm_singulr_v2_contract 2026-09-15 19:27:30 -07:00
Yuneng Jiang
064d5d61da
ci(image-scan): ignore zlib CVE-2026-85091 until Wolfi ships the fix
Wolfi's security database names zlib 1.3.3-r0 as the fix for CVE-2026-85091,
but the newest zlib published to the Wolfi apk repo is 1.3.2-r7. Every
wolfi-base digest, including the current latest, still reports the CVE, so no
base image bump or apk upgrade can clear it and image-scan fails on every PR
touching a Dockerfile or the lockfile, and on the nightly schedule.

Ignore that CVE and its GHSA alias for the zlib apk package only, so a fixable
High in anything else still fails the job.
2026-09-15 19:23:21 -07:00