Commit graph

52992 commits

Author SHA1 Message Date
Yujong Lee
df89665919 test: narrow Azure parity exception assertion
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:37:57 +00:00
Yujong Lee
94b8bf7fd6 feat(rust): add native disk cache backend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:37:26 +00:00
Yujong Lee
7d597b2dd4 fix(rust): coalesce CyberArk authentication
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:36:59 +00:00
kerry
0afb9bb587 chore: merge main into litellm_cost_shard_audio_images
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:36:48 +00:00
Yujong Lee
bfb4a8a2b3 refactor(rust): build Azure Key Vault auth inputs directly and keep credential tests offline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:36:10 +00:00
Yujong Lee
788e24655a feat(rust): add native S3 cache backend
Mirror the Redis vertical slice for S3Cache: a litellm-cache-s3 crate built on aws-sdk-s3 with path-style custom endpoints, python-identical put_object metadata (cache-control, expires, content headers), expires-aware get_object, and no-op flush/unsupported test_connection. Wire it through python-bridge config projection, facade guards, test handle, and binding dispatch, plus an in-process S3 stub and parity tests.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:35:08 +00:00
Yujong Lee
ac1c4a399f refactor(cache-redis-semantic): drop the unused default index name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:34:57 +00:00
Yujong Lee
c311073a17 feat(cache-redis-semantic): expose backend accessors for bridge binding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:34:51 +00:00
Yujong Lee
a69ebb7ac4 feat(cache): add the unsupported operation error
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:34:39 +00:00
Yujong Lee
f6db876a3d test(cache): disambiguate Valkey semantic test module
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:34:21 +00:00
kerry-berri
dcb195e805
Merge pull request #42020 from BerriAI/litellm_cost_shard_embeddings_rerank
test(integration): embeddings, rerank, completions and moderations cost cases
2026-09-21 13:33:42 -07:00
Yujong Lee
4db3481244 feat(python-bridge): serve ValkeySemanticCache natively
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:33:32 +00:00
kerry-berri
00c41e8cfb
Merge pull request #41999 from BerriAI/litellm_cost_shard_harness_extensions
test(integration): endpoint, breakdown component and failure support in the cost harness
2026-09-21 13:32:48 -07:00
berriai-litellm-provider-info-sync[bot]
07aadfd192
chore(prices): sync OpenRouter prices: 4 models, 3 new
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/xiaomi/mimo-v2.6-flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/xiaomi/mimo-v2.6-pro: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/xiaomi/mimo-v2.6-pro-ultraspeed: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_pdf_input, supports_reasoning, supports_web_search, supports_audio_input, supports_tool_choice, supports_prompt_caching, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 20:31:35 +00:00
Yujong Lee
ae69a8c79a feat(rust): add Azure Key Vault secret manager backend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:30:48 +00:00
mateo-berri
6b8e988ff0 test(e2e): settle for the replica propagation window and trim the disconnect cell's prose 2026-09-21 13:30:25 -07:00
Yujong Lee
6ab121a3e7 feat(cache-redis-semantic): add native Redis Semantic cache backend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:29:49 +00:00
Yujong Lee
beba2576be fix(rust): align vault namespace handling with python and drop unused settings hook
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:29:06 +00:00
Yujong Lee
ca31149040 feat(rust): add native GCS object-store cache backend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:28:27 +00:00
kerry
eaa6936f13 fix(fal_ai): carry fal response into content errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:27:08 +00:00
Yujong Lee
3ba4a60d5e feat(rust): add HashiCorp Vault secret manager crate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:26:28 +00:00
yucheng
0a88658227 chore: retrigger ci after docs merge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:25:56 +00:00
Yuneng Jiang
ddf6565970
Merge remote-tracking branch 'origin/main' into litellm_config_read_source 2026-09-21 13:25:45 -07:00
Yuneng Jiang
be2f0d081b
fix(proxy): report sources only on the read endpoints main does not cover
/config/field/info and /config/list already report per-key source on main,
so this drops the branch's versions of those and keeps /alerting/settings,
/get/ui_settings and /router/settings.

Read endpoints no longer write the freshly read database row back into the
shared settings store; the reload path already keeps it current, and a GET
that mutates global state leaks across callers.

Regenerates the lazy OpenAPI snapshot on Python 3.12, matching CI, and the
dashboard API types for the two new response fields.
2026-09-21 13:25:39 -07:00
kerry-berri
89e13ee959 chore: merge litellm_cost_shard_proxy_behaviour into litellm_cost_shard_batches_realtime 2026-09-21 20:25:25 +00:00
kerry-berri
1f71e4c8f5 chore: merge litellm_cost_shard_provider_wires into litellm_cost_shard_proxy_behaviour 2026-09-21 20:25:21 +00:00
kerry-berri
5d3dfe9b70 chore: merge litellm_cost_shard_pricing_dimensions into litellm_cost_shard_provider_wires 2026-09-21 20:25:17 +00:00
kerry-berri
5b51be83b8 chore: merge litellm_cost_shard_passthrough into litellm_cost_shard_pricing_dimensions 2026-09-21 20:25:13 +00:00
kerry-berri
f63a85e9cc chore: merge litellm_cost_shard_audio_images into litellm_cost_shard_passthrough 2026-09-21 20:25:10 +00:00
Yujong Lee
fe6804ea74 refactor(rust): reject negative CyberArk refresh intervals
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:25:07 +00:00
mateo-berri
61fcfd986d fix(anthropic): map the dangerous-tool-use beta for Bedrock Mantle so safeguards never reach it without the beta 2026-09-21 13:24:52 -07:00
kerry-berri
134c7111a5 chore: merge litellm_cost_shard_harness_extensions into litellm_cost_shard_audio_images 2026-09-21 20:24:37 +00:00
Yujong Lee
1f86bb8e46 feat(cache-valkey-semantic): add native Valkey semantic cache backend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:24:34 +00:00
kerry-berri
7b1283ba18 chore: merge litellm_cost_shard_harness_extensions into litellm_cost_shard_embeddings_rerank 2026-09-21 20:24:32 +00:00
Yujong Lee
8d9ab9eeaa feat(cache-redis): expose the pooled connection handling for reuse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:24:23 +00:00
kerry
1ae2c0938a chore: merge main into litellm_cost_shard_harness_extensions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:23:58 +00:00
Yujong Lee
4e2d4b5ff9 feat(rust): add CyberArk Conjur secret manager backend
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:23:54 +00:00
Yujong Lee
1320eeeb41 refactor(cache-response): generalize ResponseCache over the backend context
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:23:29 +00:00
Yujong Lee
8dc960c928 feat(cache): add SemanticCacheContext
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:22:13 +00:00
Yujong Lee
df97b274fc refactor(cache-response): generalize ResponseCache over the backend context
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:21:19 +00:00
Yujong Lee
d2f457f144 feat(cache): add semantic cache context and unsupported operation error
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:20:31 +00:00
mateo-berri
bdbb4cf527 test(e2e): send no-cache on rerank bodies like the other request models 2026-09-21 13:20:06 -07:00
Mateo Wang
662e5b6e32
Merge pull request #42284 from BerriAI/litellm_qianwen_ai_platform_rename
fix: rename the mainland China brand to Qianwen AI Platform
2026-09-21 13:18:20 -07:00
kerry
ffa1cceb11 refactor(fal_ai): share result request derivation between status fetchers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:17:40 +00:00
kerry
a909a7908e fix(fal_ai): surface fal errors in video status and content
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 20:17:29 +00:00
mateo-berri
56d1f042ef fix(types): read upstream headers through a typed helper 2026-09-21 13:16:58 -07:00
Mateo Wang
13c604c124
Merge pull request #42296 from BerriAI/litellm_ci_smoke_test_master_key
fix(ci): let the install smoke test boot its key-less proxy config
2026-09-21 13:16:44 -07:00
kerry-berri
5216844c40
Merge pull request #42286 from BerriAI/litellm_fal_ai_minimax_h3
feat(fal_ai): add MiniMax H3 text-to-video and reference-to-video
2026-09-21 13:16:08 -07:00
Yuneng Jiang
dc85812971
test(migrations): close the gaps the upgrade assertions left open
Three holes in the new suite, all of which let a test pass without proving
what its name claims:

- A migration recorded twice, once per replica, each with
  applied_steps_count = 1, slipped past both the step-count check and
  migration_names(), which collapses the history into a set. Reject
  duplicate migration_name rows outright.
- auth_traffic only asserted the failures it had seen by the time
  keep_serving hit its target. A request failing after that, or on the
  other replica while the test waited on one stream, was recorded and
  never read. Assert the recorded failures once the thread has joined.
- The rolling test warmed the baseline replica's virtual-key cache before
  the upgrade, and that cache holds for 60 seconds by default
  (UserAPIKeyCacheTTLEnum.in_memory_cache_ttl). The candidate migrates
  well inside that window, so the post-upgrade requests could be served
  from cache without ever repeating the whole-row token lookup that the
  stale prepared statement breaks. Drive the baseline replica with a key
  minted after the schema moved, which it has never seen and must resolve
  from the database.

Re-ran against v1.101.0 -> v1.102.0: 6 passed.
2026-09-21 13:15:46 -07:00
mateo-berri
b0651d52ec Merge remote-tracking branch 'origin/main' into litellm_safeguards_bedrock_vertex_messages 2026-09-21 13:15:33 -07:00