mateo-berri
2c3fc4cbff
test: drop narrating docstrings and wrap long lines in the cache hook tests
2026-09-19 18:34:23 -07:00
mateo-berri
875f015e24
fix(token_counter): count replayed redacted_thinking blocks so prompt_caching keeps pinning
...
A conversation that replays a redacted_thinking block (Anthropic redacted reasoning, or the
/v1/messages bridge's stand-in for a reasoning item that carries no summary) made
_count_content_list raise, is_prompt_caching_valid_prompt swallowed that to False, and the
prompt_caching pre-call check neither recorded nor pinned the serving deployment, so the
conversation bounced across the group and paid a cache write on every deployment. The block
now counts like a thinking block with no text: zero tokens for the encrypted payload.
2026-09-19 18:33:56 -07:00
mateo-berri
fba179f2c0
fix(google_genai): forward response schema and tool parameters through the generateContent adapter
2026-09-19 18:31:09 -07:00
kerry
2e16cd76c1
test(integration): tighten batch and realtime cost assertions
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 01:29:46 +00:00
joshua-berri
daecea3eb8
Merge pull request #42050 from BerriAI/litellm_mcp_scoped_regressions_4506_rework
...
test(mcp): restore scoped execution and credential isolation regressions
2026-09-20 01:29:44 +00:00
Tin Chi Lo
94b2fd827b
feat(ui): show prompt caching requests and net savings
2026-09-19 18:27:38 -07:00
kerry
807541291d
test(integration): batch and realtime cost cases
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 01:25:28 +00:00
Mateo Wang
93e39d5042
Merge pull request #42062 from BerriAI/litellm_pr38499_batch_retrieve_model_group
...
fix(router): stamp model_group when retrieving a batch, so batch tokens are attributable (internal copy of #38499 )
2026-09-19 18:23:06 -07:00
mateo-berri
b0971ee0ba
fix: count extra_body tools and cache_control in place of the direct ones
2026-09-19 18:22:41 -07:00
Yassin Kortam
56b3422cc2
Merge pull request #42046 from BerriAI/litellm_cost_poll_404_no_cooldown
2026-09-19 18:20:27 -07:00
Yassin Kortam
755f5c535d
Merge pull request #41994 from BerriAI/litellm_redis_spend_requeue_safety
2026-09-19 18:19:57 -07:00
yuneng-jiang
fba00f5084
Merge pull request #42054 from BerriAI/litellm_/release-ui-build-95a688
...
chore: rebuild Admin UI bundle from main (build kXnLzJ6ylsRPmgSkCkCKM)
2026-09-19 18:19:33 -07:00
mateo-berri
e833bdccde
fix(azure_ai): bridge Foundry function-tool requests only where the chat surface rejects them
...
Foundry's OpenAI v1 chat surface rejects function tools with an explicit
reasoning_effort from gpt-5.6 on and with reasoning left on from gpt-6 on,
while gpt-5.4, gpt-5.5 and unset-effort gpt-5.6 serve them. Key the
azure_ai bridge on those measured boundaries instead of the azure
provider's gpt-5.4+ rule so working chat traffic keeps its n, logprobs,
seed and chatcmpl ids.
2026-09-19 18:19:02 -07:00
Joshua Valluru
fdc4ad54c5
docs(mcp): remove duplicate regression coverage inventory
2026-09-19 18:18:37 -07:00
kerry-berri
faa2a71fba
Merge pull request #42063 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 2 models
2026-09-19 18:13:28 -07:00
ryan-crabbe-berri
99f99cfb46
fix(proxy): stop the boot when the requested master key migration fails, unless allow_requests_on_db_unavailable tolerates the outage
2026-09-19 18:10:52 -07:00
Joshua Valluru
a82f0a0bd2
fix(auth): reject deactivated JWT users and invalidate cached status
2026-09-19 18:10:49 -07:00
mateo-berri
368a839640
fix(bedrock_mantle): send anthropic betas in the header Mantle reads on /v1/messages
2026-09-19 18:08:41 -07:00
Moe Khalil
24b7a38b5f
fix(auto-router): validate saved JEV probe payloads without credentials
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 01:06:41 +00:00
berriai-litellm-provider-info-sync[bot]
b652aaad4a
chore(prices): sync OpenRouter prices: 2 models
...
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.3: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-20 01:00:52 +00:00
kerry
96ca550377
test(integration): register deployments for bedrock passthrough cases
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 01:00:25 +00:00
mateo-berri
2a35dc5217
fix(guardrails): keep the undeliverable rewrite reason through copies and name the responses mismatch
2026-09-19 18:00:03 -07:00
Moe Khalil
8898d11f6e
test(auto-router): keep editor probe on unsaved configuration
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:59:48 +00:00
Yuneng Jiang
98a3f45d21
chore: update Next.js build artifacts (2026-09-20 00:59 UTC, node v24.19.0)
2026-09-19 17:59:07 -07:00
mateo-berri
a114af2a26
fix(proxy): resolve the view setup gate through the search_path and set the row count before the views
2026-09-19 17:54:20 -07:00
Moe Khalil
97c54e278e
fix(auto-router): resolve saved JEV probes on the server
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:54:17 +00:00
ryan-crabbe-berri
a6c51ba3de
fix(proxy): never treat plaintext that base64-decodes to nothing as a ciphertext during the master key migration
...
A string such as "*" or "..." has no base64 characters, so it decoded to no bytes and read as an empty plaintext under any key. The migration would have counted it and overwritten it with a ciphertext of the empty string. Also read from the writer database instead of a read replica, report a database error during the migration instead of crashing the boot, skip columns the connected schema lacks across every schema on the search path, cap the JSON walk depth for the recursion detector, and move the boot wiring into one tested function.
2026-09-19 17:53:17 -07:00
kerry
6e77d23f4d
test(integration): settle rollup and fallback row polling
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:53:09 +00:00
kerry
dc9889a481
test(integration): restore the concurrent rollup and disconnect fix from cc93a37
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:51:08 +00:00
kerry
7d9a95462e
test(integration): reconcile concurrent cost tracking changes
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:50:24 +00:00
yucheng
2e23c2d653
fix(user_update): evict cached user on max_budget change so the personal key ceiling refreshes on every worker
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:50:18 +00:00
kerry
8e3bb5daab
test(integration): wait for all rollup writes and pin fallback and disconnect rows
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:49:51 +00:00
kerry
cc93a37322
test(integration): wait for every rollup write and bound the disconnect recount
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:48:42 +00:00
kerry-berri
d1773d96e9
Merge pull request #42058 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 5 models
2026-09-19 17:41:47 -07:00
mateo-berri
867a4df347
fix(proxy): claim a finished batch's per-model budget charge atomically
...
A finished batch reports its whole cost on every poll. The charge-once
marker is now taken with one atomic increment on the shared cache, so two
workers polling the same batch at once cannot both charge it, and the
marker's TTL is refreshed on every poll so a batch polled within every
budget window is never charged again after the marker's first expiry.
2026-09-19 17:41:19 -07:00
mateo
144cf9a9ba
test(logging): add autorouter estimate keys to the GCS pub/sub spend-log golden
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:38:42 +00:00
yucheng-berri
82ddab2405
Merge pull request #41840 from BerriAI/litellm_team_audit_lifecycle
...
fix(team): emit audit events for member_delete and role changes and carry the final roster on team create
2026-09-19 17:37:16 -07:00
kerry
2edea0be08
test(integration): assert recount pins and unique fixture request ids
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:36:18 +00:00
mateo-berri
114fc16554
merge: origin/main into litellm_lit_7346_multi_choice_stream_guardrails
2026-09-19 17:34:43 -07:00
ryan-crabbe-berri
38d776bd2b
feat(proxy): re-encrypt stored secrets at boot from LITELLM_MIGRATE_FROM_MASTER_KEY so an unsafe key can be replaced while the proxy refuses to start
...
Rotating through POST /key/regenerate needs a running proxy, which a refused boot does not have. The refusal now counts the stored values that decrypt under the unsafe key. When there are none it only asks for a new key. When there are some it also asks for LITELLM_MIGRATE_FROM_MASTER_KEY, and the next boot with a safe key re-encrypts them and logs that the variable can be deleted. Leaving the variable set afterwards is a no-op with one notice.
2026-09-19 17:32:58 -07:00
yuneng-jiang
1f28813c4d
Merge pull request #42053 from BerriAI/litellm_nova_sonic_v2_e2e_model
...
test(e2e): point the Nova Sonic realtime test at nova-2-sonic
2026-09-19 17:31:07 -07:00
berriai-litellm-provider-info-sync[bot]
260990629a
chore(prices): sync OpenRouter prices: 5 models
...
openrouter/deepseek/deepseek-v4-flash: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/qwen/qwen3.5-35b-a3b: max_tokens, max_output_tokens, supports_prompt_caching, input_cost_per_token, output_cost_per_token
openrouter/qwen/qwen3.5-9b: max_tokens, max_output_tokens
openrouter/qwen/qwen3.8-27b: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
openrouter/z-ai/glm-5.2: input_cost_per_token, output_cost_per_token, cache_read_input_token_cost
2026-09-20 00:30:47 +00:00
mateo-berri
b220178e9d
Merge remote-tracking branch 'origin/main' into litellm_explicit_cache_injection_points_survive_client_marks
2026-09-19 17:30:21 -07:00
Moe Khalil
401baf32c3
fix(auto-router): preserve JEV transport across dashboard edits
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:26:11 +00:00
yucheng-berri
2ec5c2c7cd
Merge pull request #41740 from BerriAI/litellm_otel_v2_langfuse_llm_spans_only
...
feat(otel v2): opt-in llm_only span scope for Langfuse destinations and the operator Langfuse exporter
2026-09-19 17:24:50 -07:00
Yuneng Jiang
9f84382a24
test(e2e): cite the source and date for the pinned Nova Sonic model id
...
AGENTS.md allows a vendor-owned literal only when its source and date are cited
next to it. A live realtime test cannot avoid naming a model, so record how the
id was checked, and record that a retired id fails as a hang rather than an
error so the next reader does not start by suspecting litellm.
2026-09-19 17:20:00 -07:00
kerry
a841750d46
test(integration): proxy behaviour cost cases
...
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:19:56 +00:00
mateo-berri
1b8f704035
fix(proxy): await the cancelled view setup task quietly and assert it starts at boot
...
Use contextlib.suppress for the cancelled task in stop_view_setup_task, make the legacy prisma setup test inject a plain mock for the synchronous start_view_setup_task and assert it is called, and drop the docstrings the branch added to tests
2026-09-19 17:17:18 -07:00
kerry-berri
75f4c11444
Merge pull request #42006 from BerriAI/litellm-providers/price-sync-openrouter
...
chore(prices): sync OpenRouter prices: 2 models
2026-09-19 17:11:50 -07:00
yucheng
de70cf842a
fix(team): run the role update and budget upsert in one transaction under the team lock
...
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 00:04:32 +00:00