Commit graph

52076 commits

Author SHA1 Message Date
Yuneng Jiang
45d22dc5e1
test(migrations): cover the release-to-release upgrade path
The migration e2e harness only ever used one image: it seeded the database
with the candidate build and then applied synthetic migrations on top. That
proves the migration machinery (locking, crash recovery, legacy baselining,
pooling) but never executes the real schema of release N against the real
migrations of release N+1, which is the path operators actually run.

Adds a baseline image alongside the candidate, so a test can seed with a
published release and upgrade with the build under test.

Suites:

- test_upgrade.py: the candidate applies the pending release migrations,
  keys minted by the baseline release survive, and concurrent replicas
  upgrade a baseline database exactly once.
- test_rolling_upgrade.py: a baseline replica keeps serving virtual-key
  auth while the candidate migrates underneath it, and both releases serve
  and resolve each other's keys during the overlap. This is the reported
  failure: a new column on LiteLLM_VerificationToken invalidates prepared
  plans on pods still running the old release, which the proxy reads
  whole-row, and auth starts failing until those pods leave service.
- test_shaped_database.py: the upgrade completes and preserves rows on a
  populated spend log, rather than on the empty database every other
  migration test starts from.

Every upgrade assertion is gated on the candidate having actually applied
migrations the baseline had not, so a stale pin fails loudly instead of
passing on an empty delta.

CI adds two jobs to the migration_startup workflow. The baseline defaults
to a committed release pin and is overridable per pipeline, matching how
migration_candidate_image already works; only the upgrade jobs pull it.

Verified against a real v1.101.0 -> v1.102.0 upgrade: 6 passed, with the
baseline seeding 165 migrations and the candidate applying the 6 that
landed between the two releases.
2026-09-21 12:50:58 -07:00
tin-berri
457b01e96d
Merge pull request #42055 from BerriAI/litellm_prompt_caching_request_table
feat(ui): show prompt caching requests and net savings
2026-09-21 12:45:11 -07:00
Mateo Wang
e8d97d381a
Merge pull request #42080 from BerriAI/litellm_prompt_caching_affinity_lookback
fix(router): keep prompt caching affinity when the breakpoint moves
2026-09-21 12:44:34 -07:00
kerry-berri
c3663aaac8
Merge pull request #42289 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 1 model
2026-09-21 12:41:02 -07:00
kerry-berri
2692a275fa
Merge pull request #42290 from BerriAI/litellm-providers/price-sync-aws-bedrock
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 18 held]
2026-09-21 12:40:58 -07:00
joshua-berri
211ff96943
Merge pull request #39189 from BerriAI/litellm_mcp_list_pagination_lit5594
fix(mcp): paginate prompt and resource discovery
2026-09-21 19:35:17 +00:00
Mateo Wang
0e7cf5113e
Merge pull request #42069 from BerriAI/litellm_redacted_thinking_prompt_caching_pin
fix(token_counter): count replayed redacted_thinking blocks so prompt_caching keeps pinning
2026-09-21 12:35:14 -07:00
berriai-litellm-provider-info-sync[bot]
f851e6ddb7
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 18 held]
us.moonshotai.kimi-k3: 
zai.glm-4.7: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
zai.glm-4.7-flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
2026-09-21 19:31:34 +00:00
berriai-litellm-provider-info-sync[bot]
cbf5bb5e71
chore(prices): sync OpenRouter prices: 1 model
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 19:31:29 +00:00
Mateo Wang
d8267d507d
Merge pull request #41956 from BerriAI/litellm_explicit_cache_injection_points_survive_client_marks
fix: apply configured cache_control_injection_points beside client cache_control marks
2026-09-21 12:28:02 -07:00
Mateo Wang
e7e4df9098
Merge pull request #42041 from BerriAI/litellm_azure_ai_gpt5_tools_responses_bridge
fix(azure_ai): bridge gpt-5.4+ function tools with reasoning to the Foundry Responses API
2026-09-21 12:21:02 -07:00
kerry-berri
f92ff60ebd
Merge pull request #42280 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 3 models
2026-09-21 12:11:43 -07:00
kerry-berri
b619cc22bd
Merge pull request #42279 from BerriAI/litellm-providers/price-sync-aws-bedrock
chore(prices): sync AWS Bedrock prices: 4 models [enrichment failed: AWS Bedrock, 26 held]
2026-09-21 12:11:39 -07:00
Mateo Wang
9a90adad32
Merge pull request #42275 from BerriAI/litellm_claude_platform_messages_beta_passthrough
fix(bedrock): forward anthropic-beta headers verbatim on the Claude platform messages path
2026-09-21 12:01:39 -07:00
berriai-litellm-provider-info-sync[bot]
a6842da112
chore(prices): sync OpenRouter prices: 3 models
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 19:01:36 +00:00
berriai-litellm-provider-info-sync[bot]
ae06a6478f
chore(prices): sync AWS Bedrock prices: 4 models [enrichment failed: AWS Bedrock, 26 held]
global.moonshotai.kimi-k3: 
qwen.qwen3-coder-next: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
qwen.qwen3-next-80b-a3b: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
qwen.qwen3-vl-235b-a22b: max_tokens, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
2026-09-21 19:01:36 +00:00
Yassin Kortam
5673f67727
Merge pull request #42022 from BerriAI/litellm_redis_durable_spend_log_buffer
fix(proxy): park requeued spend logs in Redis so they survive a pod restart during a DB outage
2026-09-21 14:01:34 -05:00
mateo-berri
5db2a97829 chore: keep main's lazy OpenAPI snapshot
The snapshot check runs on Python 3.12, which keeps the indentation of a route docstring that Python 3.13+ strips at compile time, so regenerating it locally on 3.14 produces a file CI rejects.
2026-09-21 12:01:26 -07:00
mateo-berri
5ea4fe620f test(router): give each prompt caching check test a fresh callback registry 2026-09-21 11:55:58 -07:00
moe-berri
a83773cfa5
Merge pull request #41886 from BerriAI/litellm_jev_autorouter_launch_1789767495
feat(auto-router): add JEV classifier alongside LLM classifier
2026-09-21 11:55:44 -07:00
mateo-berri
f549c87091 Merge branch 'main' into litellm_prompt_caching_affinity_lookback 2026-09-21 11:55:40 -07:00
Mateo Wang
7cb884cd31
Merge pull request #42120 from BerriAI/litellm_cadence_58065d4_google_genai_proxy_master_key
test(google): boot the unified Google proxy fixture with a real master key
2026-09-21 11:54:03 -07:00
kerry-berri
8ee8613b07
Merge pull request #42271 from BerriAI/litellm_bedrock_kimi_k3_us_cris
feat(bedrock): add us.moonshotai.kimi-k3 pricing and fill the global Kimi K3 entry
2026-09-21 11:51:54 -07:00
kerry-berri
12379aa1e3
Merge pull request #42095 from BerriAI/litellm_fal_gpt_image_25_flux_dev_edits
feat(fal_ai): add gpt-image-2.5 flare/sunburst, flux/dev and image edits
2026-09-21 11:45:45 -07:00
kerry-berri
24f0f373fc
Merge pull request #42270 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 6 models
2026-09-21 11:41:51 -07:00
kerry-berri
9b65ec206a
Merge pull request #42269 from BerriAI/litellm-providers/price-sync-aws-bedrock
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 32 held]
2026-09-21 11:41:45 -07:00
mateo-berri
e51ccbc759 fix(bedrock): forward anthropic-beta headers verbatim on the Claude platform messages path 2026-09-21 11:39:26 -07:00
mateo-berri
f567fe230e fix: reserve cap slots for direct marks on /v1/messages when extra_body unmarks them 2026-09-21 11:39:25 -07:00
Mateo Wang
4b2e96a5f5
Merge pull request #42143 from BerriAI/litellm_e2e_changed_keep_pytest_log
ci(e2e): fix the stage-mirror batch reds and keep a redacted pytest log
2026-09-21 11:39:25 -07:00
Joshua Valluru
1499d84f5a fix(mcp): paginate optional discovery lists 2026-09-21 11:37:55 -07:00
kerry
29837b422e feat(bedrock): add us.moonshotai.kimi-k3 pricing and fill the global Kimi K3 entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 18:32:55 +00:00
berriai-litellm-provider-info-sync[bot]
9f041f3ea3
chore(prices): sync OpenRouter prices: 6 models
openrouter/~deepseek/deepseek-pro-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/~deepseek/deepseek-v4-flash-latest: output_cost_per_token
openrouter/deepseek/deepseek-v4-flash-0731: output_cost_per_token
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/nvidia/nemotron-3-nano-30b-a3b: supports_prompt_caching, input_cost_per_token, output_cost_per_token
2026-09-21 18:31:34 +00:00
berriai-litellm-provider-info-sync[bot]
950f28fa63
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 32 held]
google.gemma-3-27b-it: max_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
mistral.ministral-3-3b-instruct: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
nvidia.nemotron-nano-12b-v2: max_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
2026-09-21 18:31:32 +00:00
Moe Khalil
83ec5d6101 fix(auto-router): skip JEV for encrypted delegated tasks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 18:29:21 +00:00
mateo-berri
79d1af9d3c chore: merge origin/main to pick up the pre-call check test move 2026-09-21 11:28:04 -07:00
mateo-berri
c1ba76154e fix(litellm): pass a flat Responses-style function tool through the chat bridge unchanged 2026-09-21 11:14:59 -07:00
kerry-berri
36b8be7d81
Merge pull request #42264 from BerriAI/litellm_add_grok_4_7
feat(xai): add grok-4.7 to the cost map
2026-09-21 11:10:37 -07:00
kerry-berri
f1362193f3
Merge pull request #42265 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 4 models
2026-09-21 11:10:05 -07:00
mateo-berri
c8e42c2ac3 ci(e2e): give string_leaves a single trailing return 2026-09-21 11:02:28 -07:00
berriai-litellm-provider-info-sync[bot]
3cc0948dc0
chore(prices): sync OpenRouter prices: 4 models
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/~z-ai/glm-flash-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/z-ai/glm-5.3-flash: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 18:01:32 +00:00
kerry
795239de20 fix: add xai/grok-4.7 to model cost map backup file
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:59:09 +00:00
kerry
38fc7d6dca feat(xai): add grok-4.7 to the cost map
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:57:32 +00:00
ryan-crabbe-berri
cc1a3157d3
Merge pull request #42121 from BerriAI/litellm_utils_model_info_lookup
feat(proxy): add GET /utils/model_info to look up cost map info for unregistered models
2026-09-21 10:42:28 -07:00
kerry
052d93d6dd test(integration): assert full fal image payloads and use existing catalog rows
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:42:02 +00:00
kerry-berri
1fbfb460e1
Merge pull request #42261 from BerriAI/litellm-providers/price-sync-openrouter
chore(prices): sync OpenRouter prices: 7 models
2026-09-21 10:40:56 -07:00
kerry-berri
01bdda72ba
Merge pull request #42254 from BerriAI/litellm-providers/price-sync-aws-bedrock
chore(prices): sync AWS Bedrock prices: 13 models, 1 new [1 with gaps, enrichment failed: AWS Bedrock, 38 held]
2026-09-21 10:40:33 -07:00
kerry
7510697355 test(integration): fal image generation and edit wire contracts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:39:53 +00:00
ryan-crabbe-berri
5573265013
Merge pull request #41906 from BerriAI/litellm_team_member_budget_source_reset
feat(team): show whether a member follows the team default budget and allow resetting to it
2026-09-21 10:34:08 -07:00
berriai-litellm-provider-info-sync[bot]
00b298a36d
chore(prices): sync AWS Bedrock prices: 13 models, 1 new [1 with gaps, enrichment failed: AWS Bedrock, 38 held]
deepseek.v3.2: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
global.moonshotai.kimi-k3: supports_vision, max_input_tokens, supports_audio_input, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, supports_tool_choice, supports_prompt_caching
google.gemma-3-12b-it: max_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
google.gemma-3-4b-it: max_tokens, max_output_tokens, supports_audio_input, supports_function_calling
mistral.devstral-2-123b: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
mistral.magistral-small-2509: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
mistral.ministral-3-14b-instruct: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
mistral.ministral-3-8b-instruct: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
mistral.mistral-large-3-675b-instruct: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
moonshotai.kimi-k2.5: max_tokens, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
nvidia.nemotron-nano-3-30b: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
nvidia.nemotron-nano-9b-v2: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
nvidia.nemotron-super-3-120b: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
2026-09-21 17:31:49 +00:00
berriai-litellm-provider-info-sync[bot]
bed94a48d9
chore(prices): sync OpenRouter prices: 7 models
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/~moonshotai/kimi-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/ibm-granite/granite-4.2-8b: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/moonshotai/kimi-k3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/qwen/qwen3.8-27b: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 17:31:40 +00:00