Commit graph

46149 commits

Author SHA1 Message Date
mateo-berri
dd2e1cf7a8 fix: read a single-quoted DO body for routine calls, not only dollar-quoted ones 2026-08-24 11:08:59 -07:00
mateo-berri
03a676995a feat(search): add Grounding with Bing Search (bing_grounding) as a search provider 2026-08-24 11:08:24 -07:00
mateo-berri
3337a0a01f fix: match OpenAI SDK wire format on image/video routes (#36493)
POST /v1/videos without an input_reference file now goes out as
multipart/form-data the way the OpenAI SDK always sends it, instead of a
JSON body that OpenAI-compatible backends (SGLang Diffusion, vLLM-Omni)
reject; gemini, vertex, and runwayml keep their JSON bodies

/v1/images/edits on the openai/azure/openai-compatible path now forwards
unknown provider params (e.g. seed) and honors extra_body, matching
/v1/images/generations, and aimage_edit forwards
extra_headers/extra_query/extra_body instead of dropping them

Generic pass-through no longer downgrades a file-less multipart form to
application/x-www-form-urlencoded
2026-08-24 11:08:21 -07:00
mateo-berri
d0dd24ed6d fix(health): apply model_info.health_check_params to health check probes 2026-08-24 11:01:34 -07:00
mateo-berri
b9e5eec28c test(e2e): pin require_managed_files enforcement behind a marker-gated stack phase 2026-08-24 10:59:13 -07:00
Felipe Rodrigues Gare Carnielli
6a0e7fe10f fix(tencent): route thinking through extra_body in chat completions
Tencent chat completions route through the OpenAI SDK's
chat.completions.create(), which raises TypeError on unknown kwargs -
so a top-level 'thinking' optional param crashed every reasoning
request with a 500 before any HTTP call was made.

Nest the resolved thinking object in extra_body instead: the SDK merges
extra_body into the top-level JSON payload, so TokenHub still receives
the documented thinking field (type/budget_tokens) in the request body.

Also align the param mapping with TokenHub's documented behavior:
- reasoning_effort="none" now maps to thinking={"type": "disabled"}
  instead of being dropped (deepseek-v4-* default to thinking enabled,
  so dropping it never actually disabled thinking)
- MiniMax models only accept thinking.type "adaptive"/"disabled",
  so "enabled" is coerced to "adaptive" instead of returning a 400

Refs: https://www.tencentcloud.com/document/product/1300/82345
2026-08-24 14:58:13 -03:00
Mateo Wang
ddf4c8e58b
Merge pull request #37953 from BerriAI/litellm_fix_24985_thinking_roundtrip
fix(anthropic): round-trip thinking blocks to OpenAI backends on /v1/messages
2026-08-24 10:57:23 -07:00
mateo-berri
fe567bd846 fix(a2a): speak the 0.3 dialect to servers with mis-cased protocol bindings 2026-08-24 10:51:14 -07:00
mateo-berri
91b2a9c360 fix(proxy): keep video reference normalization within lint budgets 2026-08-24 10:40:44 -07:00
mateo-berri
e0511e9384 Merge branch 'litellm_internal_staging' into litellm_bedrock_converse_no_trailing_empty_chunk
Resolves the test-file conflict by keeping both sides, extends the
finish-reason gate to trace-bearing metadata events so guardrail trace
chunks keep their pre-regression delta shape, parametrizes the
regression test over tool-call, mixed, and reasoning streams, and
repairs the one ant-design icon usage the lucide-react migration left
behind in skill_detail.tsx (semantic conflict on the base branch)
2026-08-24 10:37:32 -07:00
mateo-berri
ed8480a821 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_lit5458_rerank_sigv4_bearer_fix
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Waiting to run
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Waiting to run
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
LiteLLM Rust / rustfmt, clippy, test (push) Has been cancelled
2026-08-24 10:37:13 -07:00
Mateo Wang
3122600e21
Merge pull request #37975 from BerriAI/litellm_databricks_cache_token_pricing
fix(databricks): bill cached tokens at cache rates and add missing Claude pricing
2026-08-24 10:36:19 -07:00
Mateo Wang
28b433a007
Merge pull request #37966 from BerriAI/litellm_1787426863_strategy_router_health_check
fix(proxy): skip health checks for strategy routers
2026-08-24 10:26:09 -07:00
mateo-berri
20a82cad8a fix: read a wrapped group before VALUES can end the search, and blank comments in restored bodies 2026-08-24 10:23:17 -07:00
ryan-crabbe-berri
4913b2a3ca
Merge pull request #33514 from ozolam/litellm_fix_skills_marketplace_commands_v2
fix(UI): correct skill install command and marketplace setup UX
2026-08-24 10:17:25 -07:00
yuneng-jiang
3fb1009f81
fix(ui): make playground chat bubbles theme-aware (#37978)
The playground message bubble painted its fill, border and avatar circle from
inline hex values, so in dark mode both bubbles stayed near-white while the text
inherited the dark foreground: the message body was unreadable. The MCP-events
placeholder bubble in ChatUI carried the same three fills.

They move onto the tokens the rest of the sweep already uses, so the assistant
surface is bg-card over border-border and the user surface is the info tint at
the same weight the other selected-state surfaces take. Light mode keeps the
same colour family it had.

The regression test asserts the token classes and that no inline style survives
on either surface, which is the exact shape the bug took.
2026-08-24 10:13:24 -07:00
yuneng-jiang
7113685a76
fix(ui): repoint the key detail URL to the rotated hash after regenerating (#37968)
Regenerating a key from the key info page left the ?key= query param on the
old hash, so dismissing the dialog or reloading landed on a key that no longer
exists and the page rendered "Key not found".

Two defects had to line up. POST /key/{key}/regenerate returns the rotated
hash in token_id and leaves token null, but RegenerateKeyModal read
response.token || response.key_id, neither of which the endpoint populates, so
it always reported the old hash back to its parent. And KeyInfoView's
onKeyDataUpdate prop had no caller anywhere in the tree: VirtualKeysTable owns
the ?key= param and mounts the view but never passed it, so even a correct
hash went nowhere.

VirtualKeysTable now handles the update by pointing ?key= at the rotated hash
and refetching. KeyInfoView holds that callback until the regenerate dialog is
dismissed rather than firing it on the API response, because swapping the
selected key mid-dialog unmounts the view and tears down the one-time
plaintext key before the user can copy it.
2026-08-24 10:12:55 -07:00
yuneng-jiang
5b1c142c6e
fix(ui): render team and org tpm/rpm limits of 0 as 0 instead of Unlimited (#37916)
* fix(ui): render team and org tpm/rpm limits of 0 as 0 instead of Unlimited

A tpm_limit or rpm_limit of 0 is a hard block on the backend (every request 429s) and only null means unlimited, but the team and organization views rendered both as "Unlimited" (and a team-member limit of 0 as "No Limit") because every display site used a falsy || fallback. The team member edit dialog also seeded its form with `tpm_limit || null`, so opening Edit Member on a member stored with 0 and clicking Save sent null to /team/member_update and silently turned the hard block into unlimited

Every limit display site in TeamInfo, organization_view, the organizations list cell and the team members table now uses a nullish check, and both member form seeding paths keep 0 for max_budget_in_team, tpm_limit and rpm_limit. Regression tests cover each site and the existing memberFormValues test that asserted 0 -> null is flipped to assert 0 survives

Resolves LIT-5760

* test(ui): assert a stored 0 member limit survives an untouched save

The EditMembership integration test named the old 0 -> null collapse as the expected payload, so the related-tests CI job went red once the form kept 0. It now asserts 0 survives and only the empty budget_duration collapses to null. The TeamMemberTab fixture is built with a map instead of mutating the nested membership
2026-08-24 10:12:52 -07:00
yuneng-jiang
6db5a5d660
fix(ui): restore the public model name tooltip layout in the add model flow (#37986)
The tooltip popup is an inline-flex row, so the four sibling blocks passed as a fragment laid out side by side in four columns. Wrap them in a single flex-col container instead.

The inline code samples also used bg-muted, which is defined against the page surface, not the inverted tooltip surface, so they rendered as near-white chips carrying near-white text. Tint them from the popup's own token instead.
2026-08-24 10:12:46 -07:00
yuneng-jiang
5f56be3294
fix(ui): theme the created-key box so it follows dark mode (#37985)
The virtual key shown after creating a key sits in a div with a
hardcoded #f8f8f8 inline background, so in dark mode the box keeps
the light background while the key text inherits the light foreground
color, leaving the key nearly unreadable. Swap the inline styles for
the bg-muted and text-foreground tokens, which resolve per theme.
2026-08-24 10:12:16 -07:00
yuneng-jiang
a72203eae4
fix(terraform): add soft_budget, tags, and soft_budget_alerting_emails to litellm_team (#37918)
* fix(terraform): add soft_budget, tags, and soft_budget_alerting_emails to litellm_team

The team resource rejected soft_budget and tags at plan time and had no way
to express the list-valued metadata.soft_budget_alerting_emails the proxy
reads for soft-budget alerts, even though /team/new and /team/update accept
all three. Add the attributes, forward them in buildTeamData (alert emails
merged under metadata, where the proxy stores them), and send the full
metadata map whenever either half changes because /team/update replaces
metadata wholesale.

Read was decoding /team/info as if the team fields were top-level, but the
proxy nests them under team_info, so every attribute silently fell back to
prior state. Decode the envelope and split the proxy's metadata back into
tags / soft_budget_alerting_emails / string metadata, dropping the
server-managed team_member_budget_id.

Verified with OpenTofu plan/apply against a live proxy: the attributes are
accepted, land on the proxy, refresh into state, re-plan clean, propagate
on update, and clear when removed from HCL.

* fix(terraform): clear litellm_team.soft_budget in state when the proxy returns null

Read only wrote soft_budget when the proxy returned a value, so a soft
budget cleared outside Terraform stayed in state and never surfaced as
drift. Set it from the response unconditionally so a null clears it.
2026-08-24 10:12:10 -07:00
ryan-crabbe-berri
7d5a2c1a0d Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_ruff_dead_test_code
# Conflicts:
#	ruff-tests.toml
2026-08-24 09:46:56 -07:00
ryan-crabbe-berri
ee935cec23 refactor(proxy): trim the multi_items comment to the non-obvious clause 2026-08-24 09:46:35 -07:00
Thijmen Stavenuiter
41ab7e57e8 ci(test-unit): raise job timeout to 60m for the three 55m shards
caching-local, proxy-extras and enterprise-package gave pytest 20m but
capped the job at 55m; with a 35m setup ceiling plus 5m of runner
overhead the job deadline could preempt pytest itself.
2026-08-24 09:09:30 +02:00
Devin AI
cffde8d218 chore(lint): ratchet TQ001 budget for the assertions added to the aiohttp transport tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-24 06:57:46 +00:00
Thijmen Stavenuiter
f4406b4060 fix(proxy): clear LIT002 in /v2/key/info batch-cap guard
The two `or []` fallbacks only feed len(), so tuples do the job without
a mutable literal. The detail dict mirrors the sibling HTTPException
payload above and is never mutated, so it carries a mutable-ok reason.
2026-08-24 08:52:18 +02:00
Devin AI
1d22faf408 test(litellm_utils_tests): give aiohttp transport tests real assertions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-24 06:51:17 +00:00
KnyazSh
ef6cb77f91 fix(lint): Partly fix basedpyright lint issues 2026-08-23 19:35:51 +00:00
Thijmen Stavenuiter
ab9a3f4c6f Merge remote-tracking branch 'origin/litellm_internal_staging' into key-budget-window-usage
# Conflicts:
#	tests/test_litellm/proxy/management_endpoints/test_key_management_endpoints.py
2026-08-23 21:28:52 +02:00
KnyazSh
0fc3bccef8 fix(lint): Partly fix TQ lint issues 2026-08-23 14:35:12 +00:00
KnyazSh
950267e3cd fix(lint): Partly fix LIT issues 2026-08-23 14:19:50 +00:00
KnyazSh
069423ce00 fix(lint): Partly fix ANN401, S110, TID251, TRY300,UP028 2026-08-23 13:40:52 +00:00
KnyazSh
91df9845ca fix(lint): Remove definition Union in litellm_logging.py 2026-08-23 12:46:17 +00:00
KnyazSh
88747ef70c fix(test): add tests for gigachat 2026-08-23 12:03:01 +00:00
KnyazSh
29a251bee8 fix(lint): Remove definition Union in litellm_logging.py 2026-08-23 11:46:16 +00:00
KnyazSh
b99a038ea8 fix(tests): remove stale gigachat authenticator tests from old path, add utils tests 2026-08-23 11:15:06 +00:00
KnyazSh
eae7695c99 fix(tests): remove dead code in test_allm_passthrough_route_429_streaming_raises 2026-08-23 11:10:22 +00:00
KnyazSh
6b5c1d0afb fix(lint): Fix Ruff lint issues 2026-08-23 11:04:13 +00:00
Devin AI
8deade4f34 test(ptu): drop the assertion on the flag removed upstream
Some checks failed
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-23 10:02:41 +00:00
KnyazSh
d8477bb665 Merge branch 'litellm_internal_staging' into feature/improve-gigachat-provider 2026-08-23 09:40:28 +00:00
Devin AI
134b6252e0 refactor(ci): simplify PyPI license retries
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-23 09:23:35 +00:00
Devin AI
e4a72c587d fix(ci): retry transient PyPI license lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-23 09:23:35 +00:00
Devin AI
0b938e37f4 test: invalidate memoized model-cost lookups between unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-23 09:23:35 +00:00
Devin AI
3cd49b9e0c Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_deflake_20260821 2026-08-23 09:23:30 +00:00
Devin AI
cb2f5c6641 style(typing): format reasoning extraction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-23 08:07:22 +00:00
Devin AI
f44e7ad9fb chore(typing): drop fresh tech debt suppressions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-23 07:54:07 +00:00
Devin AI
d6feb35a04 chore(typing): tighten annotations added in the last day and ratchet budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-23 07:49:54 +00:00
Devin AI
e45c084c1c chore(typing): replace Any and bare containers added in the last day
Type the annotations that landed in the last 24 hours and ratchet the lint budgets down accordingly. No behavior change.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-23 07:49:44 +00:00
Devin AI
0d1722f604 merge: bring litellm_internal_staging into the rolling techdebt branch 2026-08-23 07:44:52 +00:00
yuneng-jiang
f005afa146
test(exception-mapping): pin the status and error-shape table every provider maps to (#37807)
Some checks are pending
Postgres Tests / proxy-behavior (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
`exception_type` decides the class and status a caller sees for every provider
failure, across 190 raise sites, and the tests for it were written one incident
at a time. Nothing said what a plain 401 from any given provider should be, so
mutating a raise site went unnoticed: swapping the class at each of the 190 in
turn, the mapped test file caught 15.

Adds two tables asserted end to end through `exception_type`: 25 providers by
the 9 upstream statuses, and the three error shapes the router branches on
(a full context window, a content policy block, a timeout). The same 190
mutants now fail 97 of them.

The tables record today's behavior, uneven where it is uneven. cloudflare,
ollama and vllm map no status at all, so every failure reaches the caller as a
500. A full context window is recognised by 15 of the 25, and a content policy
block by 11, which bounds where `context_window_fallbacks` and the content
policy retry policy can fire.
2026-08-22 22:59:02 -07:00