Commit graph

50471 commits

Author SHA1 Message Date
ryan-crabbe-berri
4647cd1215 fix(realtime): keep a Muse turn active for turnless partials after speechEnd
Muse partials carry no turnId and belong to the most recent speechStart,
and the docs say the model may keep post processing a turn after speechEnd
until speechComplete. Releasing the active turn on speechEnd made any
partial arriving in that window raise and get dropped in ENDPOINTING mode.
The turn now stays active until its speechComplete or final transcript.
2026-09-12 15:13:44 -07:00
mateo-berri
c9a3c5a414 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_unified_key_policy_hook 2026-09-12 15:08:57 -07:00
ryan-crabbe-berri
ba6a0c9fc6
Merge pull request #40652 from BerriAI/litellm_key_activity_search
feat(ui): search Key Activity by key alias, key hash, user id, or email
2026-09-12 15:05:53 -07:00
mateo-berri
3c2342bfd3 refactor(cost): return a new prompt token details wrapper when combining usage 2026-09-12 15:05:41 -07:00
ryan
a3ebeae28b chore(ui): regenerate schema.d.ts for /key/list substring_matching descriptions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 22:05:20 +00:00
kerry
df272d7e2f fix(registry): limit reasoning fallback to single-digit gpt majors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 22:01:58 +00:00
kerry
ca35119168 style(tests): wrap oversized model tuples
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:56:11 +00:00
kerry-berri
0f03fc2985
Merge pull request #40901 from BerriAI/litellm_remove_fireworks_price_snapshot_tests
test(fireworks): stop pinning prices in the cost-map tests
2026-09-12 14:53:58 -07:00
ryan
20787ba186 fix(proxy): allow key_alias substring matching on /key/list for non-admins
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:52:49 +00:00
yuneng-jiang
626d7aaf6a
Merge pull request #40905 from BerriAI/litellm_/release-version-bump-5f3482
chore: bump litellm-enterprise 0.1.66 -> 0.1.67, litellm-proxy-extras 0.4.96 -> 0.4.97
2026-09-12 14:49:46 -07:00
ryan-crabbe-berri
9a3f752724
Merge pull request #40656 from BerriAI/litellm_ui_table_search_pending_state
fix(ui): show loading state instead of stale rows while a table search is pending
2026-09-12 14:49:24 -07:00
Yuneng Jiang
f3ecdccb59
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/release-version-bump-5f3482 2026-09-12 14:40:07 -07:00
Yuneng Jiang
147eb23aab
bump: litellm-enterprise 0.1.66 -> 0.1.67, litellm-proxy-extras 0.4.96 -> 0.4.97 2026-09-12 14:40:01 -07:00
devin-ai-integration[bot]
5b36de4646
fix(guardrails): log mask when a guardrail adds request keys (#40882)
* fix(guardrails): log mask when a guardrail adds request keys

_inputs_were_modified only compared keys present in the pre-hook baseline, so a
guardrail that injected a new key such as tools was logged as allow. Compare over
the union of both key sets, and narrow the pre_call return value to the same
prompt-bearing keys the baseline holds so passthrough stays allow.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): snapshot apply_guardrail inputs before the hook mutates them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 14:39:40 -07:00
kerry
543ed2f6da fix(registry): scope codex/deep-research/chat-latest markers to gpt bases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:23:24 +00:00
yuneng-jiang
883b722fd2
Merge pull request #40895 from BerriAI/litellm_budget_clear_persistence
fix(ui): persist cleared budgets and reset intervals
2026-09-12 14:22:24 -07:00
kerry
fbcc602122 test(fireworks): stop pinning prices in the cost-map tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:19:35 +00:00
devin-ai-integration[bot]
311d9bba37
perf(policy_engine): dedup attachments in one pass after sorting (#40883)
get_attached_policies_with_reasons rescanned the sorted matches with next() once
per distinct policy, which is quadratic and misses the one second budget past a
few thousand global attachments. Build a policy to broadest attachment map in one
pass instead, keeping the specificity sort and result order.

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 14:19:20 -07:00
kerry
db6b851884 feat(registry): add openai reasoning-family fallback generalization
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:14:24 +00:00
kerry
09d63b259c test(fireworks): drop hardcoded price snapshot tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 21:11:02 +00:00
devin-ai-integration[bot]
261807114a
fix(proxy): accept both deferred stream logging arg shapes on native routes (#40869)
* fix(proxy): accept both deferred stream logging arg shapes on native routes

_arm_deferred_stream_dispatch armed a one-argument closure on every
anthropic_messages/aresponses stream that was not a CustomStreamWrapper or a
LiteLLMCompletionStreamingIterator. The bridged /v1/messages path returns a
plain SSE generator that shares its inner CustomStreamWrapper logging_obj, so
it stores (assembled_response, cache_hit) and _fire_deferred_stream_logging
raised TypeError, dropping spend logs and callbacks and ending the stream with
an error. The closure now dispatches on the stored args shape: a single
coroutine is enqueued, a two-tuple runs success handlers, anything else is
logged and dropped

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): assert dropped deferred payload via caplog instead of patching the logger

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 14:10:58 -07:00
Tin Chi Lo
cba843cc16 feat(proxy): predict prompt-cache costs across deployments 2026-09-12 14:02:45 -07:00
Yuneng Jiang
c8bb54993e
test: enforce isolated actors and stop OIDC process groups 2026-09-12 13:49:49 -07:00
Yuneng Jiang
b24480941b
fix(ui): retain cleared user budget values after save 2026-09-12 13:43:55 -07:00
Yuneng Jiang
ad966d8340
fix(projects): persist explicit budget cap clears 2026-09-12 13:43:55 -07:00
Yuneng Jiang
f5f81e973a
fix(budgets): preserve explicit reset interval clears 2026-09-12 13:43:55 -07:00
Yuneng Jiang
88de192dcf
test: bind management E2E callers and isolate JWT actors 2026-09-12 13:29:04 -07:00
mateo-berri
c7b607c46e fix(databricks): keep the Claude fallback when gating the anthropic thinking payload
Some checks failed
ai-gateway image / ai-gateway release image (push) Has been cancelled
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Gate the reasoning_effort translation on the cost-map flag or the model name containing
claude, so unmapped Claude serving endpoints keep translating. Flag the newer Claude
entries that were missing it. Expose supports_anthropic_thinking_payload as a public
helper next to the other supports_* wrappers instead of importing the private factory.
Drop the adaptive-only guard, since the adaptive flags only ever match Claude ids, and
add regression tests for an unmapped Claude endpoint and an adaptive Claude model
2026-09-12 13:13:38 -07:00
mateo-berri
10a0da7a32 Merge litellm_internal_staging into devin/1784568628-databricks-gemini-reasoning-effort 2026-09-12 13:08:22 -07:00
mateo-berri
68e6089950 Merge branch 'litellm_internal_staging' into litellm_codex_model_catalog_sync 2026-09-12 13:07:23 -07:00
mateo-berri
036a380fa0 chore: merge litellm_internal_staging into litellm_e2e_reliability_module_cells
Some checks failed
ai-gateway image / ai-gateway release image (push) Waiting to run
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Modules / fmt, validate, test (gcp) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
2026-09-12 13:05:02 -07:00
mateo-berri
92a544b079 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_unified_key_policy_hook 2026-09-12 13:03:40 -07:00
ryan-crabbe-berri
a2e383a1a5 fix(realtime): ignore a late speechStart for a finished Muse turn
A duplicate speechStart for a turn that already stopped used to make that
closed turn active again, so the next turnless PUSH_TO_TALK transcript was
routed to the finished item and dropped.
2026-09-12 12:53:03 -07:00
mateo-berri
9fd1ef01fe Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix_realtime_cached_audio_cost
# Conflicts:
#	litellm/responses/litellm_completion_transformation/transformation.py
2026-09-12 12:46:37 -07:00
mateo
414442cd06 fix(registry): mark computer-use-preview as supporting pdf input
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:46:13 +00:00
yucheng
fdb7da7c61 fix(mcp): bound per-caller listed-tool catalogs to 256 identities per server
Some checks are pending
ai-gateway image / ai-gateway release image (push) Waiting to run
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:41:20 +00:00
ryan-crabbe-berri
6b78438c99 fix(realtime): close every Muse turn on its own terminal signal
Turns no longer wait behind each other in a FIFO queue, so an empty
server_vad turn (speechStart then speechEnd with no transcript) cannot
stall every later turn, and a PUSH_TO_TALK speechComplete now closes its
turn without waiting for a speechEnd that never arrives. Each turn keeps
its own idempotent emit state, so late or duplicate speechEnd,
speechComplete and transcript frames are no-ops, and finished turns are
remembered in a bounded map instead of a separate tombstone deque.

The session.created ack and the sanitized error frame are now typed as
members of OpenAIRealtimeEvents, which removes the typing.cast calls
that the strict ruff budget flagged.
2026-09-12 12:39:59 -07:00
yucheng
95b8e2aaa8 fix(mcp): scope listed-tool metadata per caller on per-user MCP servers
Servers whose upstream catalog depends on the caller (oauth2 per-user, token exchange, id_jag, per-user env vars, delegated auth) now keep one listed-tool mapping per (user_id, api_key) hash under the server id, so one caller's tools/list cannot supply another caller's description or inputSchema to pre-call guardrails. Shared servers and OpenAPI-backed servers keep a single server-wide entry

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:32:03 +00:00
yucheng
d620d1ffea test(mcp): register OpenAPI listing fixtures directly instead of patching the global tool registry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:22:02 +00:00
ryan-crabbe-berri
559247fa84
Merge pull request #40880 from BerriAI/litellm_tf_key_cascade_delete_recovery
fix(key): recover from a cascade-deleted key instead of failing the apply
2026-09-12 12:20:53 -07:00
mateo
9aa06cb4a3 fix(registry): correct computer-use-preview provider/schema flag and OpenRouter deepseek-v3.2 / claude-opus-4.6 metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:20:05 +00:00
yucheng
4296a5d5d7 fix(mcp): list OpenAPI tools by exact prefix boundary so overlapping server prefixes stay separate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:16:55 +00:00
yuneng-jiang
5f73f837ec
Merge pull request #40774 from BerriAI/litellm_cache_response_semantics
test(e2e): verify cached answers and upstream request count
2026-09-12 12:11:57 -07:00
yucheng
eba05f4333 fix(mcp): record OpenAPI tool definitions for pre-call guardrail metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 19:09:17 +00:00
devin-ai-integration[bot]
49f94cf832
fix(ui): clarify blank TPM/RPM hint on budget modals (#40697)
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 12:08:08 -07:00
mateo
085db0b266 Merge litellm_internal_staging into rolling registry PR 2026-09-12 19:04:09 +00:00
yucheng
c23ae322c1 chore: merge litellm_internal_staging into litellm_agent365_mcp_guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 18:59:45 +00:00
ryan-crabbe-berri
a109553909 fix(key): recover from a cascade-deleted key instead of failing the apply
Deleting a team cascade-deletes its keys, so `terraform apply -replace` on a
team left the key's `/key/update` 404ing and aborted the apply with the key
resource stuck. The update now confirms the key is really gone and recreates it
under the new team; a `team_id` change between two live teams stays an in-place
update, and an unrelated failure still errors out.

Rebased onto current staging, which added a typed `apiError` and `isNotFound`,
so the recovery matches on the status code plus a re-read rather than on the
error string. The metadata pre-read, which fails before `/key/update` is ever
reached when the key is gone, routes through the same recovery.

Original work by @matthowardcohere in #39747.

Claude-Session: https://claude.ai/code/session_01XT1qsbjLwnhiN5sQ2hNUxr
2026-09-12 11:58:44 -07:00
yuneng-jiang
27f8d2a9ba
Merge pull request #40826 from BerriAI/litellm_local_form_clear_adapters
fix(ui): preserve clear and default semantics in local forms
2026-09-12 11:57:29 -07:00
yuneng-jiang
a73454b8fc
fix(ui): restore MCP catalog provider logos (#40781)
Resolves LIT-7390
2026-09-12 11:56:59 -07:00