Yuneng Jiang
45d22dc5e1
test(migrations): cover the release-to-release upgrade path
...
The migration e2e harness only ever used one image: it seeded the database
with the candidate build and then applied synthetic migrations on top. That
proves the migration machinery (locking, crash recovery, legacy baselining,
pooling) but never executes the real schema of release N against the real
migrations of release N+1, which is the path operators actually run.
Adds a baseline image alongside the candidate, so a test can seed with a
published release and upgrade with the build under test.
Suites:
- test_upgrade.py: the candidate applies the pending release migrations,
keys minted by the baseline release survive, and concurrent replicas
upgrade a baseline database exactly once.
- test_rolling_upgrade.py: a baseline replica keeps serving virtual-key
auth while the candidate migrates underneath it, and both releases serve
and resolve each other's keys during the overlap. This is the reported
failure: a new column on LiteLLM_VerificationToken invalidates prepared
plans on pods still running the old release, which the proxy reads
whole-row, and auth starts failing until those pods leave service.
- test_shaped_database.py: the upgrade completes and preserves rows on a
populated spend log, rather than on the empty database every other
migration test starts from.
Every upgrade assertion is gated on the candidate having actually applied
migrations the baseline had not, so a stale pin fails loudly instead of
passing on an empty delta.
CI adds two jobs to the migration_startup workflow. The baseline defaults
to a committed release pin and is overridable per pipeline, matching how
migration_candidate_image already works; only the upgrade jobs pull it.
Verified against a real v1.101.0 -> v1.102.0 upgrade: 6 passed, with the
baseline seeding 165 migrations and the candidate applying the 6 that
landed between the two releases.
2026-09-21 12:50:58 -07:00
tin-berri
457b01e96d
Merge pull request #42055 from BerriAI/litellm_prompt_caching_request_table
...
feat(ui): show prompt caching requests and net savings
2026-09-21 12:45:11 -07:00
Mateo Wang
e8d97d381a
Merge pull request #42080 from BerriAI/litellm_prompt_caching_affinity_lookback
...
fix(router): keep prompt caching affinity when the breakpoint moves
2026-09-21 12:44:34 -07:00
kerry-berri
c3663aaac8
Merge pull request #42289 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 1 model
2026-09-21 12:41:02 -07:00
kerry-berri
2692a275fa
Merge pull request #42290 from BerriAI/litellm-providers/price-sync-aws-bedrock
...
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 18 held]
2026-09-21 12:40:58 -07:00
joshua-berri
211ff96943
Merge pull request #39189 from BerriAI/litellm_mcp_list_pagination_lit5594
...
fix(mcp): paginate prompt and resource discovery
2026-09-21 19:35:17 +00:00
Mateo Wang
0e7cf5113e
Merge pull request #42069 from BerriAI/litellm_redacted_thinking_prompt_caching_pin
...
fix(token_counter): count replayed redacted_thinking blocks so prompt_caching keeps pinning
2026-09-21 12:35:14 -07:00
berriai-litellm-provider-info-sync[bot]
f851e6ddb7
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 18 held]
...
us.moonshotai.kimi-k3:
zai.glm-4.7: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
zai.glm-4.7-flash: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
2026-09-21 19:31:34 +00:00
berriai-litellm-provider-info-sync[bot]
cbf5bb5e71
chore(prices): sync OpenRouter prices: 1 model
...
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 19:31:29 +00:00
Mateo Wang
d8267d507d
Merge pull request #41956 from BerriAI/litellm_explicit_cache_injection_points_survive_client_marks
...
fix: apply configured cache_control_injection_points beside client cache_control marks
2026-09-21 12:28:02 -07:00
Mateo Wang
e7e4df9098
Merge pull request #42041 from BerriAI/litellm_azure_ai_gpt5_tools_responses_bridge
...
fix(azure_ai): bridge gpt-5.4+ function tools with reasoning to the Foundry Responses API
2026-09-21 12:21:02 -07:00
kerry-berri
f92ff60ebd
Merge pull request #42280 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 3 models
2026-09-21 12:11:43 -07:00
kerry-berri
b619cc22bd
Merge pull request #42279 from BerriAI/litellm-providers/price-sync-aws-bedrock
...
chore(prices): sync AWS Bedrock prices: 4 models [enrichment failed: AWS Bedrock, 26 held]
2026-09-21 12:11:39 -07:00
Mateo Wang
9a90adad32
Merge pull request #42275 from BerriAI/litellm_claude_platform_messages_beta_passthrough
...
fix(bedrock): forward anthropic-beta headers verbatim on the Claude platform messages path
2026-09-21 12:01:39 -07:00
berriai-litellm-provider-info-sync[bot]
a6842da112
chore(prices): sync OpenRouter prices: 3 models
...
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, off_peak_pricing, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 19:01:36 +00:00
berriai-litellm-provider-info-sync[bot]
ae06a6478f
chore(prices): sync AWS Bedrock prices: 4 models [enrichment failed: AWS Bedrock, 26 held]
...
global.moonshotai.kimi-k3:
qwen.qwen3-coder-next: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
qwen.qwen3-next-80b-a3b: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
qwen.qwen3-vl-235b-a22b: max_tokens, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
2026-09-21 19:01:36 +00:00
Yassin Kortam
5673f67727
Merge pull request #42022 from BerriAI/litellm_redis_durable_spend_log_buffer
...
fix(proxy): park requeued spend logs in Redis so they survive a pod restart during a DB outage
2026-09-21 14:01:34 -05:00
mateo-berri
5db2a97829
chore: keep main's lazy OpenAPI snapshot
...
The snapshot check runs on Python 3.12, which keeps the indentation of a route docstring that Python 3.13+ strips at compile time, so regenerating it locally on 3.14 produces a file CI rejects.
2026-09-21 12:01:26 -07:00
mateo-berri
5ea4fe620f
test(router): give each prompt caching check test a fresh callback registry
2026-09-21 11:55:58 -07:00
moe-berri
a83773cfa5
Merge pull request #41886 from BerriAI/litellm_jev_autorouter_launch_1789767495
...
feat(auto-router): add JEV classifier alongside LLM classifier
2026-09-21 11:55:44 -07:00
mateo-berri
f549c87091
Merge branch 'main' into litellm_prompt_caching_affinity_lookback
2026-09-21 11:55:40 -07:00
Mateo Wang
7cb884cd31
Merge pull request #42120 from BerriAI/litellm_cadence_58065d4_google_genai_proxy_master_key
...
test(google): boot the unified Google proxy fixture with a real master key
2026-09-21 11:54:03 -07:00
kerry-berri
8ee8613b07
Merge pull request #42271 from BerriAI/litellm_bedrock_kimi_k3_us_cris
...
feat(bedrock): add us.moonshotai.kimi-k3 pricing and fill the global Kimi K3 entry
2026-09-21 11:51:54 -07:00
kerry-berri
12379aa1e3
Merge pull request #42095 from BerriAI/litellm_fal_gpt_image_25_flux_dev_edits
...
feat(fal_ai): add gpt-image-2.5 flare/sunburst, flux/dev and image edits
2026-09-21 11:45:45 -07:00
kerry-berri
24f0f373fc
Merge pull request #42270 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 6 models
2026-09-21 11:41:51 -07:00
kerry-berri
9b65ec206a
Merge pull request #42269 from BerriAI/litellm-providers/price-sync-aws-bedrock
...
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 32 held]
2026-09-21 11:41:45 -07:00
mateo-berri
e51ccbc759
fix(bedrock): forward anthropic-beta headers verbatim on the Claude platform messages path
2026-09-21 11:39:26 -07:00
mateo-berri
f567fe230e
fix: reserve cap slots for direct marks on /v1/messages when extra_body unmarks them
2026-09-21 11:39:25 -07:00
Mateo Wang
4b2e96a5f5
Merge pull request #42143 from BerriAI/litellm_e2e_changed_keep_pytest_log
...
ci(e2e): fix the stage-mirror batch reds and keep a redacted pytest log
2026-09-21 11:39:25 -07:00
Joshua Valluru
1499d84f5a
fix(mcp): paginate optional discovery lists
2026-09-21 11:37:55 -07:00
kerry
29837b422e
feat(bedrock): add us.moonshotai.kimi-k3 pricing and fill the global Kimi K3 entry
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 18:32:55 +00:00
berriai-litellm-provider-info-sync[bot]
9f041f3ea3
chore(prices): sync OpenRouter prices: 6 models
...
openrouter/~deepseek/deepseek-pro-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/~deepseek/deepseek-v4-flash-latest: output_cost_per_token
openrouter/deepseek/deepseek-v4-flash-0731: output_cost_per_token
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/nvidia/nemotron-3-nano-30b-a3b: supports_prompt_caching, input_cost_per_token, output_cost_per_token
2026-09-21 18:31:34 +00:00
berriai-litellm-provider-info-sync[bot]
950f28fa63
chore(prices): sync AWS Bedrock prices: 3 models [enrichment failed: AWS Bedrock, 32 held]
...
google.gemma-3-27b-it: max_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
mistral.ministral-3-3b-instruct: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
nvidia.nemotron-nano-12b-v2: max_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
2026-09-21 18:31:32 +00:00
Moe Khalil
83ec5d6101
fix(auto-router): skip JEV for encrypted delegated tasks
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 18:29:21 +00:00
mateo-berri
79d1af9d3c
chore: merge origin/main to pick up the pre-call check test move
2026-09-21 11:28:04 -07:00
mateo-berri
c1ba76154e
fix(litellm): pass a flat Responses-style function tool through the chat bridge unchanged
2026-09-21 11:14:59 -07:00
kerry-berri
36b8be7d81
Merge pull request #42264 from BerriAI/litellm_add_grok_4_7
...
feat(xai): add grok-4.7 to the cost map
2026-09-21 11:10:37 -07:00
kerry-berri
f1362193f3
Merge pull request #42265 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 4 models
2026-09-21 11:10:05 -07:00
mateo-berri
c8e42c2ac3
ci(e2e): give string_leaves a single trailing return
2026-09-21 11:02:28 -07:00
berriai-litellm-provider-info-sync[bot]
3cc0948dc0
chore(prices): sync OpenRouter prices: 4 models
...
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/~z-ai/glm-flash-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/z-ai/glm-5.3-flash: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 18:01:32 +00:00
kerry
795239de20
fix: add xai/grok-4.7 to model cost map backup file
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:59:09 +00:00
kerry
38fc7d6dca
feat(xai): add grok-4.7 to the cost map
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:57:32 +00:00
ryan-crabbe-berri
cc1a3157d3
Merge pull request #42121 from BerriAI/litellm_utils_model_info_lookup
...
feat(proxy): add GET /utils/model_info to look up cost map info for unregistered models
2026-09-21 10:42:28 -07:00
kerry
052d93d6dd
test(integration): assert full fal image payloads and use existing catalog rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:42:02 +00:00
kerry-berri
1fbfb460e1
Merge pull request #42261 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 7 models
2026-09-21 10:40:56 -07:00
kerry-berri
01bdda72ba
Merge pull request #42254 from BerriAI/litellm-providers/price-sync-aws-bedrock
...
chore(prices): sync AWS Bedrock prices: 13 models, 1 new [1 with gaps, enrichment failed: AWS Bedrock, 38 held]
2026-09-21 10:40:33 -07:00
kerry
7510697355
test(integration): fal image generation and edit wire contracts
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:39:53 +00:00
ryan-crabbe-berri
5573265013
Merge pull request #41906 from BerriAI/litellm_team_member_budget_source_reset
...
feat(team): show whether a member follows the team default budget and allow resetting to it
2026-09-21 10:34:08 -07:00
berriai-litellm-provider-info-sync[bot]
00b298a36d
chore(prices): sync AWS Bedrock prices: 13 models, 1 new [1 with gaps, enrichment failed: AWS Bedrock, 38 held]
...
deepseek.v3.2: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
global.moonshotai.kimi-k3: supports_vision, max_input_tokens, supports_audio_input, supports_response_schema, supports_function_calling, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, cache_creation_input_token_cost, supports_tool_choice, supports_prompt_caching
google.gemma-3-12b-it: max_tokens, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
google.gemma-3-4b-it: max_tokens, max_output_tokens, supports_audio_input, supports_function_calling
mistral.devstral-2-123b: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
mistral.magistral-small-2509: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
mistral.ministral-3-14b-instruct: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
mistral.ministral-3-8b-instruct: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
mistral.mistral-large-3-675b-instruct: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
moonshotai.kimi-k2.5: max_tokens, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
nvidia.nemotron-nano-3-30b: max_tokens, supports_vision, max_input_tokens, max_output_tokens, supports_audio_input, supports_response_schema
nvidia.nemotron-nano-9b-v2: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema, supports_function_calling
nvidia.nemotron-super-3-120b: max_tokens, supports_vision, max_output_tokens, supports_audio_input, supports_response_schema
2026-09-21 17:31:49 +00:00
berriai-litellm-provider-info-sync[bot]
bed94a48d9
chore(prices): sync OpenRouter prices: 7 models
...
openrouter/~deepseek/deepseek-pro-latest: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/~moonshotai/kimi-latest: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/deepseek/deepseek-v4-pro-0813: max_tokens, max_output_tokens, input_cost_per_token, output_cost_per_token, cache_read_input_token_cost, off_peak_pricing
openrouter/ibm-granite/granite-4.2-8b: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/moonshotai/kimi-k3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/qwen/qwen3.8-27b: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-21 17:31:40 +00:00