Commit graph

48758 commits

Author SHA1 Message Date
yuneng
da246e9282 chore(deps): bump anyio to 4.14.2 and soupsieve to 2.9.0
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 23:03:52 +00:00
Yassin Kortam
9a46f1f53e Merge pull request #41462 from BerriAI/litellm_otel_promote_nested_request_metadata_keys
feat(otel): promote nested request metadata keys to litellm.metadata.* span attributes

(cherry picked from commit 79fc5153d3)
2026-09-22 23:02:58 +00:00
devin-ai-integration[bot]
96e7739dcb fix(realtime): surface an upstream handshake refusal as an error event and policy close (#42388)
* fix(realtime): surface an upstream handshake refusal as an error event and policy close

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(realtime): tidy the handshake refusal e2e

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(realtime): keep upstream exception text out of the Azure client error

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(realtime): map handshake refusal close codes with a lookup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit 2bab39e374)
2026-09-22 23:02:53 +00:00
Mateo Wang
89000e402f
Merge pull request #42538 from BerriAI/litellm_cherrypick_safeguards_1_102_x
fix(anthropic): backport #42152 and #42288 to stable/1.102.x for v1.102.1
2026-09-22 14:29:04 -07:00
kerry
0fee6fcc76 fix(test): run the all-beta-headers bedrock cases on Claude Fable 5.1
Backport of #42048 to stable/1.102.x.
Cherry-picked from 7966f50c34 (main). The safeguards backport maps the dangerous-tool-use-2026-09-03 beta for Bedrock, which Claude Opus 4.5 on Bedrock Invoke rejects as an invalid beta flag, so the all-beta-headers Bedrock cases run on Claude Fable 5.1 as they do on main.
2026-09-22 12:54:44 -07:00
mateo-berri
712428eded chore(types): keep the backported safeguards annotations within the line's budgets
The picked TypedDict fields use read-only Sequence[Mapping[str, object]] annotations and the picked Vertex test carries a test-quality-ok marker, so stable/1.102.x's LIT001, LIT012 and TQ008 budgets hold. Static typing only, no runtime change.
2026-09-22 12:18:27 -07:00
mateo-berri
4f4ed1121a bump: version 1.102.1 2026-09-22 10:46:02 -07:00
mateo-berri
60c019166a test: add the local_beta_headers_config fixture the safeguards tests use
Hand-ported to stable/1.102.x from 47b2479c94 on main (fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag), the one prerequisite the #42288 handler tests need; the rest of that commit stays on main.
2026-09-22 10:45:59 -07:00
mateo-berri
b6f77c32b5 fix(anthropic): forward Claude Code safeguards and dangerous-tool-use beta to Bedrock Invoke and Vertex on /v1/messages
Backport of #42288 to stable/1.102.x.
Cherry-picked from merge commit fc82f6e8fa (litellm_safeguards_bedrock_vertex_messages).
The line has no bedrock_mantle beta-header mapping and no Mantle /v1/messages route, so the Mantle mapping, its test file, and the bedrock_mantle test parameter are left out.
2026-09-22 10:45:57 -07:00
yassin
e714ab4b20 fix(anthropic): forward safeguards and anthropic-beta unchanged on native /v1/messages
Backport of #42152 to stable/1.102.x.
Cherry-picked from merge commit e912ebe999 (litellm_claude_code_safeguards_passthrough).
2026-09-22 10:38:43 -07:00
yuneng-jiang
95293834e8
Merge pull request #42059 from BerriAI/litellm_backport_42053_rc_1_102_0
test(e2e): backport the Nova Sonic nova-2-sonic model fix to rc/1.102.0
2026-09-19 17:36:17 -07:00
Yuneng Jiang
64db6e1d3b
test(e2e): cite the source and date for the pinned Nova Sonic model id
AGENTS.md allows a vendor-owned literal only when its source and date are cited
next to it. A live realtime test cannot avoid naming a model, so record how the
id was checked, and record that a retired id fails as a hang rather than an
error so the next reader does not start by suspecting litellm.

(cherry picked from commit 9f84382a24)
2026-09-19 17:32:32 -07:00
Yuneng Jiang
c48532e21b
test(e2e): point the Nova Sonic realtime test at nova-2-sonic
AWS retired amazon.nova-sonic-v1:0. GetFoundationModel now answers
ResourceNotFoundException "This model version has reached the end of its life",
and opening a bidirectional stream against it fails with ValidationException
"The provided model identifier is invalid". amazon.nova-2-sonic-v1:0 is the
active replacement.

The test passed on builds 219 (2026-09-16) and 246 (2026-09-17) and has failed
every run since, three attempts per build, with no litellm change to the
realtime path in between. The symptom was a clean websocket close: Bedrock ends
the stream rather than erroring, the forwarder treats a None receive as a normal
stream end and closes the client socket, so the client sees ConnectionClosedOK
and the test fails waiting for response.done.

litellm already carries both models in the cost map, with
"deprecation_date": "2026-09-14" on the old one, and the promptStart
transformation already sends the audioOutputConfiguration that nova-2-sonic
requires; only the test constant was left behind.

Verified against live Bedrock with the promptStart shape the transformation
builds: amazon.nova-sonic-v1:0 raises "The provided model identifier is
invalid", amazon.nova-2-sonic-v1:0 opens a session and returns a usageEvent.

The mocked handler and provider-cache tests keep the old id: it is only a label
there, no call reaches AWS.

(cherry picked from commit ca18755b64)
2026-09-19 17:32:31 -07:00
Mateo Wang
eebb3cb17d
Merge pull request #41935 from BerriAI/litellm_backport_41347_rc_1_102_0
Some checks failed
ai-gateway image / ai-gateway release image (push) Has been cancelled
fix(team): apply team_member_budget updates to members still on the team default (backport of #41347 to rc/1.102.0)
2026-09-19 00:01:54 -07:00
mateo-berri
868052b1ac fix(team): apply team_member_budget updates to members still on the team default
Backport of #41347 to rc/1.102.0.
Cherry-picked from merge commit 4bb1ae115b (main), originally by app/devin-ai-integration.

Conflicts: main's #41349 (membership rows written through upsert) is not on this line, so add_new_member keeps its create call and still writes no membership row when no budget resolves. The picked tests are adapted to that and to this line's _check_team_member_budget, which loads the membership itself.
2026-09-18 23:18:38 -07:00
Mateo Wang
1a993c3d28
Merge pull request #41854 from BerriAI/litellm_backport_stable_batch_rc_1_102_0
Some checks failed
ai-gateway image / ai-gateway release image (push) Waiting to run
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
fix: backport eight backport-stable fixes to rc/1.102.0 (#40596, #41046, #41086, #41171, #41178, #41283, #41495, #41689)
2026-09-18 12:16:09 -07:00
mateo-berri
3a7a204b9b fix(responses): keep the addressed response id off bridged provider requests
Backport of #41689 to rc/1.102.0.
Cherry-picked from merge commit 07b5051c0d (main), originally by app/devin-ai-integration.

Both conflicts were in test files: the rc line lacks the Codex additional_tools tests that sit next to the new test in test_handler.py on main, and its test_utils.py imports all_litellm_params on a separate line, so only this PR's own additions (the imports, the recording handler, and the two new tests) are taken.
2026-09-18 11:15:56 -07:00
mateo-berri
707a81aeae fix(utils): run post-call deployment hook on converted chat streams
Backport of #41495 to rc/1.102.0.
Cherry-picked from merge commit 0add8c0083 (main), originally by app/devin-ai-integration.

The only conflict was the test import header: the rc line never gained the logging_executor import that main carries next to this change, so only the MockResponseIterator import comes along.
2026-09-18 11:11:07 -07:00
mateo-berri
8f8b47d6f0 fix(proxy): retry rate-limit fallbacks from a pristine request snapshot
Backport of #40596 to rc/1.102.0.
Cherry-picked from merge commit c5325b1492 (main), originally by app/devin-ai-integration.
2026-09-18 11:10:33 -07:00
mateo-berri
e6f29fbe9f fix(logging): track spend for streams a deployment hook converted to non-streaming
Backport of #41171 to rc/1.102.0.
Cherry-picked from merge commit c3222ec110 (main), originally by app/devin-ai-integration.
2026-09-18 11:10:31 -07:00
mateo-berri
24e583c7f6 fix(proxy): keep access-group raw SQL writes on the writer while writer_unavailable is stale
Backport of #41283 to rc/1.102.0.
Cherry-picked from merge commit 0e5be275b0 (main), originally by app/devin-ai-integration.
2026-09-18 11:10:30 -07:00
mateo-berri
7c1186e880 fix(router): bind per-request routing_strategy override selectors to the request's callbacks
Backport of #41178 to rc/1.102.0.
Cherry-picked from merge commit 1ca4579375 (main), originally by app/devin-ai-integration.
2026-09-18 11:10:28 -07:00
mateo-berri
034998be71 fix(cli): drop enum.StrEnum so the CLI imports on Python 3.10
Backport of #41046 to rc/1.102.0.
Cherry-picked from merge commit 1ba97665b2 (main), originally by app/devin-ai-integration.
2026-09-18 11:10:27 -07:00
mateo-berri
31d78fa084 fix(proxy): keep org admins' own team memberships in other orgs visible on team list
Backport of #41086 to rc/1.102.0.
Cherry-picked from merge commit 4123b4bc2b (main), originally by app/devin-ai-integration.

rc/1.102.0 has no litellm/proxy/auth/resolvers/grants.py, so its one-line import change is dropped; auth_checks.py imports UserNotFoundError from the new types module, so every importer on this line still resolves.
2026-09-18 11:10:08 -07:00
Mateo Wang
4fd5bc6b22
Merge pull request #41701 from BerriAI/litellm_cherrypick_rc_1_102_0
fix(license): backport the wildcard license auto_router grant to rc/1.102.0 (#41684)
2026-09-17 17:09:09 -07:00
Mateo Wang
e1ef8b336e
Merge pull request #41697 from BerriAI/litellm_backport_39512_rc_1_102_0
fix(images): backport the image[] and mask[] form key drop to rc/1.102.0 (#39512)
2026-09-17 16:58:46 -07:00
mateo-berri
e71b78f01c fix(license): let a wildcard allowed_features license grant the auto_router feature
Backport of #41684 to rc/1.102.0.
Cherry-picked from c2fbb11dca (litellm_wildcard_license_auto_router).
2026-09-17 16:42:25 -07:00
mateo-berri
24d44530d4 test(images): type the monkeypatch fixture on the backported edit tests 2026-09-17 16:34:29 -07:00
mateo-berri
0e03b62ab8 fix(images): backport the image[] and mask[] form key drop to rc/1.102.0 (#39512)
(cherry picked from commit 289c52bcd6)
2026-09-17 16:21:06 -07:00
yuneng-jiang
3f965660a1
Merge pull request #41342 from BerriAI/litellm_backport_param_leaks_rc_1_102_0
fix(responses): backport request-param leak fixes to rc/1.102.0 (#41018, #41141, #41144)
2026-09-15 18:06:20 -07:00
Yassin Kortam
262978eebc
Merge pull request #41144 from BerriAI/litellm_responses_bridge_filters_unknown_params
fix(responses): filter bridged kwargs like the native Responses path

(cherry picked from commit 41b5d47c71)
2026-09-15 17:40:38 -07:00
Yassin Kortam
8a23af9bc6
Merge pull request #41141 from BerriAI/litellm_lit7694_forwarded_headers_body_leak
fix(openai): keep extra_headers out of the chat request body on the httpx handler path

(cherry picked from commit de55e22899)
2026-09-15 17:40:38 -07:00
Yassin Kortam
4afb271ba3
Merge pull request #41018 from BerriAI/litellm_fix_model_alias_map_leak
fix(utils): keep litellm params out of provider request bodies

(cherry picked from commit 1f46e58494)
2026-09-15 17:40:26 -07:00
yuneng-jiang
16f6d92453
Merge pull request #40933 from BerriAI/litellm_internal_staging
chore(ci): promote internal staging to main
2026-09-12 18:51:00 -07:00
yuneng-jiang
0f16cff575
Merge pull request #40931 from BerriAI/litellm_/release-version-bump-6cf851
chore: rebuild Admin UI bundle from staging
2026-09-12 18:42:18 -07:00
Yuneng Jiang
d15edb8cc9
chore: update Next.js build artifacts (2026-09-13 01:20 UTC, node v24.19.0) 2026-09-12 18:20:05 -07:00
ryan-crabbe-berri
a969319fb5
Merge pull request #36222 from BerriAI/litellm_lit_5292_model_info_pricing_filter
fix(model_management): stop persisting cost map pricing as a deployment override
2026-09-12 18:16:06 -07:00
Mateo Wang
9d984371fd
Merge pull request #40909 from BerriAI/litellm_databricks_reasoning_effort_thinking
fix(databricks): translate reasoning_effort to thinking for Gemini 2.5
2026-09-12 17:46:44 -07:00
ryan-crabbe-berri
76cb0fec1c fix(model_management): stop persisting cost map pricing as a deployment override
/model/info fills a deployment's missing pricing in from the model cost map so the
Admin UI has a rate to display. Clients echo that whole model_info blob back on save,
and update_db_model merged it into the row, so editing an unrelated setting turned
that day's catalog price into a real per-deployment override. After that the
deployment ignored the cost map and Reload Price Data could no longer move it,
because the reload replays each deployment's stored pricing over the fresh catalog.

Drop the derived pricing from incoming model_info on the two write paths. The
drop-set is read off the same objects the read path uses, CustomPricingLiteLLMParams
plus the tiered *_above_N_tokens pattern that get_model_info passes through and no
model declares, so it cannot drift as new rates are added. output_vector_size is
exempt: it lives on the pricing model but is an embedding dimension, not a rate.

A deployment's own pricing still rides litellm_params, which is untouched, as is the
explicit-null clear, which reads the incoming model rather than the filtered dict.
The filter sits in the endpoint bodies rather than _add_model_to_db, which master-key
rotation reuses to re-serialize every stored deployment.
2026-09-12 17:29:35 -07:00
ryan-crabbe-berri
d4a72e7372
Merge pull request #40907 from BerriAI/litellm_key_alias_substring_non_admin
fix(proxy): allow key_alias substring matching on /key/list for non-admins
2026-09-12 16:48:17 -07:00
kerry-berri
d565031b9c
Merge pull request #40902 from BerriAI/litellm_openai_reasoning_fallback_rule
feat(registry): add openai reasoning-family fallback generalization
2026-09-12 16:27:16 -07:00
ryan
8ccde3879d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_key_alias_substring_non_admin 2026-09-12 23:06:23 +00:00
devin-ai-integration[bot]
b1360efc2f
fix(proxy): attribute gate-rejected requests to their endpoint in cache analytics (#40824)
* fix(proxy): attribute gate-rejected requests to their endpoint in cache analytics

Requests rejected before dispatch (bad key, blocked key, budget, rate limit, malformed body) were spend-logged with an empty call_type because the synthesized logging object never reached the failure lifter. The caching dashboard rolled all of them, plus failed calls on info routes such as /model/info, into one Unknown group.

Resolve call_type from the matched route first, falling back to body shape, and keep the synthesized logging object on request_data so the lifter sees it. Log bare auth exceptions with the 401 ProxyException the client gets so error_code is never empty. Exclude info routes from the cache analytics groups and error breakdown. The dashboard explains the Unknown group when older rows still produce one.

Resolves LIT-5884

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep the raw auth exception for failure callbacks

Record the client-facing status in the spend log through a separate client_exception argument so custom failure callbacks still receive the exception auth raised.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep the route for multi-operation endpoints and exclude info routes from cache filter options

Routes such as /v1/files map to several operations (create, list) and the
method is not available in the failure hook, so a rejected request there is
filed under its route instead of the first mapped call type. The key alias and
model filter-option queries now apply the same info-route exclusion as the
groups and error breakdown, so every offered filter value returns data. The
info-route exclusion and Unknown grouping are now covered against a real
Postgres in tests/proxy_behavior/spend/test_cache_activity.py

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): drop client_exception, the spend log row never used it

The DB spend row for a gate rejection is written by _ProxyDBLogger from the
original exception, so the status-bearing copy only reached the in-memory
logging payload. Live runs at the tip still recorded bare auth exceptions as
Unknown/Exception, the same as the base branch. Removing the plumbing keeps
this PR to endpoint attribution and the info-route exclusion

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 16:03:38 -07:00
ryan-crabbe-berri
c134fb7a38
Merge pull request #39395 from seyeong-han/litellm_meta_muse_voice_realtime
Some checks failed
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
feat(realtime): add Meta Muse Voice transcription
2026-09-12 15:58:24 -07:00
Mateo Wang
566f026c1c
Merge pull request #40767 from BerriAI/litellm_ocr_custom_pricing
fix(cost): honour deployment custom pricing for OCR calls
2026-09-12 15:55:44 -07:00
mateo-berri
a7ebc12673 ci: drop the removed rust bridge test path from the ocr job 2026-09-12 15:53:03 -07:00
mateo-berri
d7158ba795 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_ocr_custom_pricing 2026-09-12 15:42:59 -07:00
Mateo Wang
426e675c49
Merge pull request #40764 from BerriAI/litellm_redis_pool_timeout_counts_as_timeout
fix(redis): count pool wait timeouts as breaker timeouts
2026-09-12 15:42:57 -07:00
ryan-crabbe-berri
0e435e4148 fix(realtime): run transcription guardrails on transcription-only sessions
The provider_config path skipped run_realtime_guardrails for transcription
sessions to avoid sending response.create, which also dropped every
realtime_input_transcription guardrail: no violation error reached the
client and on_violation / end_session_after_n_fails never fired. Run the
guardrail for every completed transcript and only suppress response.create
when the session has no assistant turn.
2026-09-12 15:32:45 -07:00
ryan
02dbff485b test(proxy): mark prisma_client patch in /key/list alias test with test-quality reason
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-12 22:32:14 +00:00