Commit graph

52992 commits

Author SHA1 Message Date
Yujong Lee
ef2ab74c7a refactor(rust): back the vault secret manager with vaultrs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:33:38 +00:00
Yujong Lee
c7c0afb1f0 ci(rust): raise the native wheel size gate to 40 MB
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:33:06 +00:00
yucheng-berri
506cecfb0b
Merge pull request #42267 from BerriAI/litellm_otel_v2_langfuse_ocr_output
fix(otel v2): map OCR page markdown onto the generation output
2026-09-21 14:33:02 -07:00
Yujong Lee
b6b0e58ba3 refactor(cache-valkey-semantic): reuse cache-redis connection layer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:32:42 +00:00
Yujong Lee
2a5112d146 feat(cache-redis): expose the pooled connection handling for reuse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:32:40 +00:00
Yujong Lee
9c48e137dc fix(rust): preserve HTTP host and TLS error context
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:32:33 +00:00
Yujong Lee
cc6ca5ee04 Merge origin/main into litellm_valkey_semantic_native_cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:32:33 +00:00
Yujong Lee
27808b51a0 chore(rust): drop interop planning note
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:32:33 +00:00
Yujong Lee
5aeb367d2a fix(rust): preserve Python settings coercion at the native boundary 2026-09-21 21:32:33 +00:00
Yujong Lee
52216df1ff fix(rust): preserve structured conversion semantics 2026-09-21 21:32:33 +00:00
Yujong Lee
a41061be43 docs(rust): plan Python interop foundation 2026-09-21 21:32:33 +00:00
Joshua Valluru
4c1a6309c2 test(mcp): type server fixtures and role queries 2026-09-21 14:31:28 -07:00
berriai-litellm-provider-info-sync[bot]
8c64f82bb3
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 21:31:16 +00:00
yucheng-berri
79d6e236f9
Merge pull request #39805 from BerriAI/litellm_mcp_admin_api_preserve_oauth_scopes
fix(mcp): keep oauth scopes in admin api credential redaction
2026-09-21 14:29:13 -07:00
Yuneng Jiang
d1b160310f
fix(docker): require a generated master key in the quickstart stack
The committed sk-1234 placeholder is in PUBLICLY_KNOWN_MASTER_KEYS, so once an
image ships fe480533e8 the proxy refuses to boot and the quickstart stops
working. It also meant the documented stack came up on port 4000 with a
credential anyone could guess.

Both keys now come from .env and compose refuses to render without them. The
salt key is generated alongside so it stays stable across restarts, which
keeps stored credentials readable.

Verified: no .env -> compose fails closed naming the missing variable; with a
generated .env the stack is healthy, /v1/models returns 200 for the generated
key, 401 for sk-1234 and 401 unauthenticated, /ui/ serves, and the key still
works after a restart.
2026-09-21 14:27:29 -07:00
Yujong Lee
a5571333fb refactor(cache-valkey-semantic): reuse cache-redis connection layer
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:25:22 +00:00
kerry
1e54b4f286 Merge remote-tracking branch 'origin/litellm_cost_shard_proxy_behaviour' into litellm_cost_shard_batches_realtime 2026-09-21 21:25:05 +00:00
kerry
924ca78d06 Merge remote-tracking branch 'origin/main' into litellm_cost_shard_proxy_behaviour 2026-09-21 21:24:54 +00:00
Mateo Wang
124e5d6d53
Merge pull request #34358 from BerriAI/claude/e2e-tests-custom-endpoints-qxoi1o
test(e2e): replace custom endpoints_client with provider SDK clients
2026-09-21 14:24:38 -07:00
kerry-berri
34d5f9d41b
Merge pull request #42052 from BerriAI/litellm_cost_shard_provider_wires
test(integration): provider wire cost cases
2026-09-21 14:24:35 -07:00
Yuneng Jiang
4602376977
fix(proxy): report a stored alerting value as db even when it is null
A stored null or empty list for a nested alerting field is still the value
the proxy serves when the config file leaves alerting_args alone, so the
source is db. Keying off the value rather than its presence reported those
fields as default and hid a stored setting that is genuinely in effect.

Presence in the stored row now decides, with the config file still checked
first so a config-owned key keeps reporting config. Test helpers are typed
and the router test injects a stub rather than patching a class attribute.
2026-09-21 14:24:32 -07:00
Yujong Lee
4a8826e03b refactor(cache-disk): tidy python adapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:23:56 +00:00
Yujong Lee
2b257bd9a3 feat(cache-redis): expose the pooled connection handling for reuse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:23:28 +00:00
Yujong Lee
6237dd51cb fix(cache-redis-semantic): isolate on an incompatible index after a lost create race
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:59 +00:00
Yujong Lee
e729820301 revert(rust): drop wheel-size release profile tweaks, gate moves in #42300
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:27 +00:00
Yujong Lee
59675bc0c9 Revert "ci(rust): raise native wheel size gate to 35 MB"
This reverts commit c1382086d6.
2026-09-21 21:22:26 +00:00
Yujong Lee
3ce436af5e refactor(cache-disk): isolate python compatibility behind a value adapter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:22 +00:00
Yujong Lee
769917a7e8 revert(rust-wheel): restore release optimization
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:15 +00:00
Yujong Lee
974d9f97ff revert(rust): restore thin LTO in the release profile
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:09 +00:00
Yujong Lee
4063658168 build(rust): drop native extension size limit bump, deferred to #42300
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:02 +00:00
Yujong Lee
4509eb9914 fix(cache): align native semantic cache scope keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:22:01 +00:00
Yujong Lee
d2f8e8c835 fix(cache-redis-semantic): harden index initialization
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:21:57 +00:00
Mateo Wang
4b9aa222a6
Merge pull request #42317 from BerriAI/litellm_redis_cluster_topology
feat(rust-cache): serve RedisClusterCache natively as a Redis topology
2026-09-21 14:21:54 -07:00
mateo-berri
a2ae80ec9b fix(router): wrap every Responses and Messages fallback hop for mid-stream failover
The /v1/responses and /v1/messages streaming wrappers only ever wrapped the
primary's stream, so a hop reached through the regular fallback chain had no
mid-stream handler: its failure re-raised, or the outer wrapper retried the
same entry with a fresh attempted set and never reached the rest of the list.
Every attempt of the chain now runs through a per-endpoint attempt function
that wraps its own stream, mirroring chat completions, and the per-request
fallback and retry overrides ride a frozen carrier so each hop's re-entry
still sees them after the retry layer pops them.
2026-09-21 14:19:52 -07:00
Yujong Lee
7c8f7d3f2c fix(rust): reduce native wheel size
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:18:16 +00:00
Joshua Valluru
64078689d3 refactor(mcp): clear server list lint warnings 2026-09-21 14:18:01 -07:00
mateo-berri
2f8bee053d test(alerting): inject the webhook client and extend the mapped test files 2026-09-21 14:16:37 -07:00
Yassin Kortam
f6c69af427
Merge pull request #41101 from hMED22/litellm_add_edenai_provider
feat(edenai): add Eden AI provider across chat, Responses, Messages, embeddings, audio, images and video
2026-09-21 16:16:28 -05:00
Yujong Lee
0e281d458f fix(rust): treat ConditionNotMet as an existing blob on sync Azure Blob writes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:16:19 +00:00
kerry
5ad36d2b7e Merge remote-tracking branch 'origin/litellm_cost_shard_proxy_behaviour' into litellm_cost_shard_batches_realtime 2026-09-21 21:15:56 +00:00
kerry
2760ea2e6c Merge remote-tracking branch 'origin/litellm_cost_shard_provider_wires' into litellm_cost_shard_proxy_behaviour 2026-09-21 21:15:49 +00:00
Yujong Lee
793efe3eb4 fix(rust): restore hot-path opt levels and validate s3 binding destination
Keep sigv4 signing, eventstream decoding, smithy runtime api and types at opt-level 3 since they serve Bedrock request and streaming hot paths, and make the S3 facade binding reject region and endpoint mismatches between the projected configuration and the native handle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:15:42 +00:00
kerry
54a1e85d38 Merge remote-tracking branch 'origin/main' into litellm_cost_shard_provider_wires 2026-09-21 21:15:39 +00:00
Mateo Wang
fc82f6e8fa
Merge pull request #42288 from BerriAI/litellm_safeguards_bedrock_vertex_messages
fix(anthropic): forward Claude Code safeguards and dangerous-tool-use beta to Bedrock Invoke and Vertex on /v1/messages
2026-09-21 14:15:36 -07:00
kerry-berri
b5e47936e5
Merge pull request #42035 from BerriAI/litellm_cost_shard_pricing_dimensions
test(integration): pricing dimension and provider reported cost cases
2026-09-21 14:15:21 -07:00
kerry-berri
c20d2c2803
Merge pull request #42320 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 1 model
2026-09-21 14:13:30 -07:00
Yuneng Jiang
7590012a03
feat(docker): add a quickstart compose file served from the product repo
The docs quickstart pipes a compose file hosted on the docs site straight
into `docker compose -f -`. That puts the content users execute in the docs
repo rather than here, and nothing lands on disk for them to read first.

Move the two-service stack (gateway + Postgres) into docker/ so it ships and
is reviewed alongside the code it starts, and pin the image to main-stable
instead of latest. The docs change to download-then-run follows separately.

Verified: `docker compose up -d` brings the stack healthy, /health/liveliness
returns 200, /v1/models returns 200 with the placeholder key and 401 without,
and /ui/ serves.
2026-09-21 14:12:55 -07:00
Yujong Lee
b5548082c6 build(rust): raise native extension size limit to 30 MB for the Azure Blob SDK
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:12:07 +00:00
Yujong Lee
96224f1c0f fix(cache-redis): fan cluster PING out to every node and skip pool return pings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 21:11:22 +00:00
mateo-berri
1b568319d0 fix(types): blank non-string datadog tool text fields, look the spend table up by name 2026-09-21 14:10:09 -07:00