Commit graph

53121 commits

Author SHA1 Message Date
Yujong Lee
4a8826e03b refactor(cache-disk): tidy python adapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:23:56 +00:00
Yujong Lee
2b257bd9a3 feat(cache-redis): expose the pooled connection handling for reuse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:23:28 +00:00
Yujong Lee
6237dd51cb fix(cache-redis-semantic): isolate on an incompatible index after a lost create race
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:59 +00:00
Yujong Lee
e729820301 revert(rust): drop wheel-size release profile tweaks, gate moves in #42300
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:27 +00:00
Yujong Lee
59675bc0c9 Revert "ci(rust): raise native wheel size gate to 35 MB"
This reverts commit c1382086d6.
2026-09-21 21:22:26 +00:00
Yujong Lee
3ce436af5e refactor(cache-disk): isolate python compatibility behind a value adapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:22 +00:00
Yujong Lee
769917a7e8 revert(rust-wheel): restore release optimization
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:15 +00:00
Yujong Lee
974d9f97ff revert(rust): restore thin LTO in the release profile
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:09 +00:00
Yujong Lee
4063658168 build(rust): drop native extension size limit bump, deferred to #42300
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:02 +00:00
Yujong Lee
4509eb9914 fix(cache): align native semantic cache scope keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:01 +00:00
Yujong Lee
d2f8e8c835 fix(cache-redis-semantic): harden index initialization
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:21:57 +00:00
Mateo Wang
4b9aa222a6
Merge pull request #42317 from BerriAI/litellm_redis_cluster_topology
feat(rust-cache): serve RedisClusterCache natively as a Redis topology
2026-09-21 14:21:54 -07:00
mateo-berri
a2ae80ec9b fix(router): wrap every Responses and Messages fallback hop for mid-stream failover
The /v1/responses and /v1/messages streaming wrappers only ever wrapped the
primary's stream, so a hop reached through the regular fallback chain had no
mid-stream handler: its failure re-raised, or the outer wrapper retried the
same entry with a fresh attempted set and never reached the rest of the list.
Every attempt of the chain now runs through a per-endpoint attempt function
that wraps its own stream, mirroring chat completions, and the per-request
fallback and retry overrides ride a frozen carrier so each hop's re-entry
still sees them after the retry layer pops them.
2026-09-21 14:19:52 -07:00
Yujong Lee
7c8f7d3f2c fix(rust): reduce native wheel size
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:18:16 +00:00
Joshua Valluru
64078689d3 refactor(mcp): clear server list lint warnings 2026-09-21 14:18:01 -07:00
mateo-berri
2f8bee053d test(alerting): inject the webhook client and extend the mapped test files 2026-09-21 14:16:37 -07:00
Yassin Kortam
f6c69af427
Merge pull request #41101 from hMED22/litellm_add_edenai_provider
feat(edenai): add Eden AI provider across chat, Responses, Messages, embeddings, audio, images and video
2026-09-21 16:16:28 -05:00
Yujong Lee
0e281d458f fix(rust): treat ConditionNotMet as an existing blob on sync Azure Blob writes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:16:19 +00:00
kerry
5ad36d2b7e Merge remote-tracking branch 'origin/litellm_cost_shard_proxy_behaviour' into litellm_cost_shard_batches_realtime 2026-09-21 21:15:56 +00:00
kerry
2760ea2e6c Merge remote-tracking branch 'origin/litellm_cost_shard_provider_wires' into litellm_cost_shard_proxy_behaviour 2026-09-21 21:15:49 +00:00
Yujong Lee
793efe3eb4 fix(rust): restore hot-path opt levels and validate s3 binding destination
Keep sigv4 signing, eventstream decoding, smithy runtime api and types at opt-level 3 since they serve Bedrock request and streaming hot paths, and make the S3 facade binding reject region and endpoint mismatches between the projected configuration and the native handle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:15:42 +00:00
kerry
54a1e85d38 Merge remote-tracking branch 'origin/main' into litellm_cost_shard_provider_wires 2026-09-21 21:15:39 +00:00
Mateo Wang
fc82f6e8fa
Merge pull request #42288 from BerriAI/litellm_safeguards_bedrock_vertex_messages
fix(anthropic): forward Claude Code safeguards and dangerous-tool-use beta to Bedrock Invoke and Vertex on /v1/messages
2026-09-21 14:15:36 -07:00
kerry-berri
b5e47936e5
Merge pull request #42035 from BerriAI/litellm_cost_shard_pricing_dimensions
test(integration): pricing dimension and provider reported cost cases
2026-09-21 14:15:21 -07:00
kerry-berri
c20d2c2803
Merge pull request #42320 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 1 model
2026-09-21 14:13:30 -07:00
Yuneng Jiang
7590012a03
feat(docker): add a quickstart compose file served from the product repo
The docs quickstart pipes a compose file hosted on the docs site straight
into `docker compose -f -`. That puts the content users execute in the docs
repo rather than here, and nothing lands on disk for them to read first.

Move the two-service stack (gateway + Postgres) into docker/ so it ships and
is reviewed alongside the code it starts, and pin the image to main-stable
instead of latest. The docs change to download-then-run follows separately.

Verified: `docker compose up -d` brings the stack healthy, /health/liveliness
returns 200, /v1/models returns 200 with the placeholder key and 401 without,
and /ui/ serves.
2026-09-21 14:12:55 -07:00
Yujong Lee
b5548082c6 build(rust): raise native extension size limit to 30 MB for the Azure Blob SDK
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:12:07 +00:00
Yujong Lee
96224f1c0f fix(cache-redis): fan cluster PING out to every node and skip pool return pings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:11:22 +00:00
mateo-berri
1b568319d0 fix(types): blank non-string datadog tool text fields, look the spend table up by name 2026-09-21 14:10:09 -07:00
Yujong Lee
a9ff1de42b build(rust): switch release LTO to fat for wheel size headroom
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:09:36 +00:00
Yassin Kortam
da1ccaec67
Merge pull request #40322 from BerriAI/litellm_lit7351_reservation_lease_renewal
fix(proxy): renew budget reservation counter TTL while the request is in flight
2026-09-21 16:07:14 -05:00
Yujong Lee
3f11d9787d build(ci): raise native extension size limit to 30 MB for cluster support
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:06:12 +00:00
yuneng-jiang
5e0512b611
Merge pull request #42291 from BerriAI/litellm_lit7597_detach_credential
fix(proxy): detach stored credential when model editor selects None
2026-09-21 14:05:28 -07:00
Yujong Lee
c1382086d6 ci(rust): raise native wheel size gate to 35 MB
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:05:21 +00:00
Yujong Lee
0d09e9d892 fix(rust-wheel): reduce native extension size
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:04:57 +00:00
kerry
6ef1003f7d Merge remote-tracking branch 'origin/litellm_cost_shard_proxy_behaviour' into litellm_cost_shard_batches_realtime 2026-09-21 21:04:53 +00:00
kerry
3204bf0df9 Merge remote-tracking branch 'origin/litellm_cost_shard_provider_wires' into litellm_cost_shard_proxy_behaviour 2026-09-21 21:04:45 +00:00
kerry
5034b96fb6 Merge remote-tracking branch 'origin/litellm_cost_shard_pricing_dimensions' into litellm_cost_shard_provider_wires 2026-09-21 21:04:37 +00:00
kerry
b0070892da Merge remote-tracking branch 'origin/main' into litellm_cost_shard_pricing_dimensions 2026-09-21 21:04:26 +00:00
kerry-berri
b5f02f72fe
Merge pull request #42028 from BerriAI/litellm_cost_shard_passthrough
test(integration): passthrough route cost cases
2026-09-21 14:04:03 -07:00
berriai-litellm-provider-info-sync[bot]
a8cd414377
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 21:01:19 +00:00
mateo-berri
a64febb3e7 fix(streaming): keep an explicit provider prompt_tokens=0 or completion_tokens=0 in streamed usage
The stream chunk builder started its per-chunk accumulators at 0 and adopted only nonzero counts, then fell back to litellm's tokenizer whenever the accumulated value was falsy, so a provider that reported an explicit 0 for prompt or completion tokens was billed the estimate instead. The accumulators now start at None, a usage chunk that reports a count marks it reported (a later chunk's 0 never replaces a reported nonzero), and the estimate only runs when no chunk reported the count. The Anthropic message_start cursor reset now yields None so the estimate still covers a cancelled stream, and Ollama chat streaming only attaches usage on the done chunk when both counts are present instead of inventing 0/0 on every chunk
2026-09-21 14:01:08 -07:00
Yujong Lee
5acfcfa4fe feat(rust): native Azure Blob response cache backend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:00:48 +00:00
Yujong Lee
0672fcbafe fix(rust): shrink release wheel under the 25 MB native limit
Optimize the aws-sdk dependency tree and cache-s3 for size in the release profile and switch to fat LTO so the native extension stays under the wheel verification gate (27.87 -> 21.88 MB)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:59:13 +00:00
kerry-berri
89de508d6e
Merge pull request #42306 from BerriAI/litellm_fal_ai_surface_video_errors
fix(fal_ai): surface fal errors in video status and content instead of completed and generic 500
2026-09-21 13:59:04 -07:00
Yujong Lee
d3f2ddba05 fix(python-bridge): allow instance shadowing only for validated config attributes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:58:02 +00:00
mateo-berri
239bff3315 test(responses): type the MCP lifecycle test helpers 2026-09-21 13:57:50 -07:00
kerry-berri
f5f53a4cf4
Merge pull request #40429 from BerriAI/litellm_upgrade_banner_changelog_stats
feat(ui): add upgrade banner with latest release changelog stats
2026-09-21 13:56:05 -07:00
kerry
aeff48238a Merge remote-tracking branch 'origin/litellm_cost_shard_proxy_behaviour' into litellm_cost_shard_batches_realtime 2026-09-21 20:56:02 +00:00
kerry
efd7f3dadd Merge remote-tracking branch 'origin/litellm_cost_shard_provider_wires' into litellm_cost_shard_proxy_behaviour 2026-09-21 20:55:55 +00:00