Commit graph

44707 commits

Author SHA1 Message Date
mateo
bfc52b94db fix(ocr): return 400 for an unknown x-req-format header value
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-17 18:11:57 +00:00
mateo
57d739b433 feat(ocr): add req_format=native to return Azure Document Intelligence's own analyzeResult payload
Callers can opt into the provider's raw operation response on /v1/ocr with the x-req-format: native header (or req_format in the body) while page-based cost tracking keeps reading usage_info off the normalized response.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-17 18:03:53 +00:00
ryan-crabbe-berri
9d40cd4df7
feat(ui): link key info header to its user, creator, team, and organization (#37187)
* feat(ui): link key info header to its user, creator, team, and organization

The key info page showed the owning user and creator as plain text and never surfaced the team or organization at all, so walking from a key to its parent entities meant copying ids into other pages. The header now renders User and Created By as links to the user detail page, and gains a far-right column with Team and Organization links (alias when known, id otherwise). Client-side navigation logic shared by BadgeLink and IdentityCell moves into a reusable EntityLink so all entity links behave the same

* test(ui): mock next/navigation in VirtualKeysTable test

KeyInfoView now renders EntityLink, which calls useRouter, so the table test that opens the key detail needs the app router mocked
2026-08-17 10:17:28 -07:00
tin-berri
3c3ada9af0
feat(ui): add Lite mixed-provider auto-router preset (#37068)
* feat(ui): add Lite mixed-provider auto-router preset

* feat(ui): disable classifier context window in Lite preset
2026-08-17 10:07:01 -07:00
tin-berri
5ecc6af541
fix(UI): add default model pin to complexity router UI (#36615)
* Add default model pin to complexity router UI

A complexity router's default model was only ever derived from the tiers, so
operators had no way to point the fallback at a model that is not first in the
Simple or Medium tier. Adds a Default Model select that records an explicit pin.

The pin is stored in complexity_router_config.default_model, which the backend
already reads, and mirrored onto complexity_router_default_model on save. Both
paths resolve through one helper that mirrors init_complexity_router_deployment:
a pin wins, otherwise MEDIUM or SIMPLE. Recording the pin in the config keeps it
distinguishable from a derived value, so a pin that happens to match the tiers
survives a round trip instead of being read back as tier tracking.

The edit modal only requires one non-empty tier, so a router with models in
COMPLEX alone can reach save with nothing the backend would pick. That now
blocks with an inline message rather than saving a router that raises at init.

* fix(UI): probe the pinned default model in the auto router connection test

The connection test built its targets from the tiers and the embedding model
only, so a Default Model pin outside every tier was never reached and a green
result could hide an unreachable default. model_info_view had already hand
rolled the dedupe and append locally, so the rule moved into
buildAutoRouterTestTargets and both call sites now share it.

* fix(ui): mirror backend precedence when resolving a complexity router default

The edit modal only recognized a pin stored in complexity_router_config.default_model,
so a router whose default lived solely in litellm_params.complexity_router_default_model
lost it on the next save. That field cannot be trusted outright either: before this PR
every save wrote a tier-derived value into it, so treating any value as a pin would
freeze legacy routers away from their tiers. Hydration now takes the config marker as
authoritative and falls back to litellm_params only when it diverges from what the tiers
alone derive, which is only reachable through an external API or config write.

Test Connection had the mirror-image bug: it fell back to complexity_router_config.default_model,
a UI-only marker init_complexity_router_deployment never reads, so it could probe a model
the router would never call. It now follows router.py exactly: litellm_params, else pure
tier-derivation.

Also reword a tooltip that hardcoded the Default Model select's position on the page, and
document the dual write and the create-vs-edit validation asymmetry.
2026-08-17 10:04:23 -07:00
Mateo Wang
2bc87ec3cc
Merge pull request #34067 from MUSE-CODE-SPACE/fix/batch-logging-null-output-file
fix(batches): don't crash logging when a completed batch has no output file
2026-08-17 10:01:02 -07:00
Mateo Wang
41c3133d0e
Merge pull request #34087 from ArjunPakhan/fix/bedrock-cancel-batch
fix(batches): support AWS Bedrock batch cancellation via `StopModelInvocationJob`
2026-08-17 10:00:57 -07:00
Mateo Wang
a1644eaf84
Merge pull request #36392 from BerriAI/devin_ai_fix_bedrock_batch_file_bytes_36388
fix(bedrock): report uploaded size in the FileObject returned by managed batch uploads
2026-08-17 10:00:53 -07:00
Mateo Wang
cfb2eba7f9
Merge pull request #36151 from LHMQ878/fix/36088-openai-ws-passthrough
fix(proxy): register WebSocket passthrough for OpenAI prefixes
2026-08-17 10:00:49 -07:00
Mateo Wang
1dbed6eb60
Merge pull request #37073 from BerriAI/litellm_decrease_anys_fable_round2 2026-08-17 09:40:40 -07:00
Marty Sullivan
5fe7793a14 refactor(bedrock): own the Converse batch usage shape in the provider layer
Shape detection and block normalization sat in the generic batch layer, which
let batch and live parsing of the same wire format drift apart. Both now live on
AmazonConverseConfig as is_converse_usage_shape and usage_from_batch_output, so
batch_utils asks the provider adapter rather than knowing Bedrock's field names.

Adds direct coverage for the shape predicate, the completion of an incomplete
block, cache-count inflation, and the streaming usage event that shares the
public transform. Drops the narrative banner from the batch tests.
2026-08-16 22:32:56 -04:00
mateo-berri
5965648547 fix(proxy): close websocket cleanly when OpenAI credentials are missing 2026-08-16 14:40:37 -07:00
mateo-berri
4ba9d6b136 fix(proxy): expose url join helper at module level for websocket route 2026-08-16 14:29:21 -07:00
mateo-berri
81aefe4b3c Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_pr36151_ws_passthrough
# Conflicts:
#	litellm/proxy/pass_through_endpoints/pass_through_endpoints.py
#	tests/test_litellm/proxy/pass_through_endpoints/test_pass_through_endpoints.py
2026-08-16 14:21:57 -07:00
mateo-berri
862f33bbaa fix(proxy): negotiate client subprotocol on OpenAI websocket passthrough 2026-08-16 14:11:12 -07:00
mateo-berri
a258b2b130 fix(proxy): harden OpenAI websocket passthrough
- decode upstream first frame as utf-8 instead of ascii
- reject model-restricted keys at connect to match HTTP model enforcement
- log the actual request path for /openai_passthrough traffic
2026-08-16 13:51:50 -07:00
mateo-berri
5764c0daed Merge fork updates, keep typed kwargs and status-based cancel idempotency
# Conflicts:
#	litellm/llms/bedrock/batches/handler.py
2026-08-16 13:44:35 -07:00
mateo-berri
1b13957776 fix(bedrock): treat ConflictException on stop as idempotent cancel 2026-08-16 13:44:27 -07:00
mubashir1osmani
782746553c fix(perplexity): accept float usage.cost in cost_per_token, not just dict
ResponseAPIUsage.parse_cost already flattens Perplexity's
usage.cost.total_cost dict down to a float before it reaches the
perplexity cost calculator, so the isinstance(cost_info, dict) check
was always False on that path. Every Responses-mode Perplexity model
was silently falling back to manual token-rate calculation and
recording $0 spend whenever static per-token rates were missing.
2026-08-16 14:57:11 -04:00
mubashir1osmani
539a61be08 feat(perplexity): add Agent API third-party models (DeepSeek V4 Flash, GLM 5.2, Kimi K3, Kimi K2.7 Code) 2026-08-16 14:45:22 -04:00
mubashir1osmani
ffa37d05b7 feat(mistral): add zai-glm-5-2 model pricing and metadata 2026-08-16 14:28:50 -04:00
Arjun Pakhan
4ff4e12557 fix(bedrock): add missing kwargs type annotations and refine conflict handling 2026-08-16 17:00:21 +00:00
Arjun Pakhan
88ac6ae63b style(bedrock): format handler.py with ruff to fix CI linting 2026-08-16 16:37:30 +00:00
Arjun Pakhan
ba11eebb2c fix(bedrock): handle ConflictException during idempotent batch cancel 2026-08-16 16:06:06 +00:00
zhanghanduo
a40b9983c5 fix(anthropic): reopen a content block when a resumed item id returns
_get_or_start_block trusted the item_id -> block index map without checking
whether that block was still open, so a provider that reuses one item id for
a whole run and interleaves channels got a delta addressed to a stopped
block. Replaying reasoning, text, reasoning, text produced
content_block_delta index=0 after content_block_stop index=0, which is not a
valid Anthropic stream.

Treat the mapping as valid only while it points at the open block, and
rebind the item to a fresh block otherwise. Items registered through
response.output_item.added still reuse their block, since that block is the
open one while its deltas arrive.

Apodex Deep Research is what surfaced this: it labels every reasoning delta
of a run rs_<response_id> and every answer delta msg_<response_id>, so any
interleaving hits the stale mapping.
2026-08-16 21:25:36 +08:00
zhanghanduo
428380fb1e fix(apodex): preserve reasoning block lifecycle 2026-08-16 20:47:02 +08:00
zhanghanduo
ec3e293ac2 feat(apodex): stream Deep Research reasoning as reasoning deltas
A Deep Research run streams two agents. The worker emits its chain of
thought on the `reasoning` channel and a draft answer on a channel-less
delta; the reporter emits its own reasoning plus the single `output_text`
delta that matches the final response.completed snapshot.

Only `output_text` was mapped, so 176 of 181 deltas in a sample run
surfaced as GenericEvent and the reasoning was effectively lost. Map the
`reasoning` channel to response.reasoning_summary_text.delta, which
LiteLLM already translates into an Anthropic thinking_delta on the
/v1/messages route Deep Research takes, and give it its own item id.

The channel-less deltas stay unclaimed on purpose: splicing the worker's
draft into the answer would corrupt the text. The remaining
response.swarm.* lifecycle events keep passing through, since
transform_streaming_response has no way to drop a chunk and run_finished
carries the final content.
2026-08-16 20:26:47 +08:00
zhanghanduo
42dcfa12a2 fix(apodex): isolate chat routing 2026-08-16 19:47:24 +08:00
Shivi Jain
a6dc447470 fix(proxy): stop Responses batch rows from bypassing project OTPM
Embeddings rows were identified by body shape (has `input`, no
`messages`/`prompt`), which also matches a `/v1/responses` batch row
and reserved zero output tokens for it -- letting a project caller run
large Responses generations against a quota-limited model without
consuming OTPM. Classify embeddings by the row's own `url` instead,
and read `max_output_tokens` as a Responses output cap alongside
`max_tokens`/`max_completion_tokens`.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 16:26:30 +05:30
Shivi Jain
8a44c14928 Merge upstream/litellm_internal_staging and fix batch quota review comments
Resolves conflicts from the upstream merge and addresses the Veria-AI
review comment on this PR: batch rows could bypass a project's
per-model ITPM/OTPM quota when the batch's file-bound/routing model
had no quota configured. Charges each row's own model against its own
project quota instead of only the routing model's, and fixes rate
limit error messages to attribute the correct model via a new
descriptor_value field on RateLimitStatus/AtomicCounterMeta. Also
re-syncs the ruff-strict, type-discipline, and basedpyright budgets
against the correct (non-stale) merge base.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-16 15:47:46 +05:30
zhanghanduo
190a3e01ec fix(apodex): harden provider routing metadata 2026-08-16 18:14:53 +08:00
zhanghanduo
64dcb95268 Revert "drop the undocumented response.swarm.llm_delta mapping"
A live stream against apodex-1-1-deep-research shows the event is real and
load-bearing: 181 of the 193 events are response.swarm.llm_delta and
response.output_text.delta never appears, so without the mapping the answer
text only arrives in the final response.completed snapshot.

Restores the transform with the provenance recorded in a docstring, and
covers it with the payload shape captured off the wire, including the
reasoning channel that carries 176 of those deltas and must not be mistaken
for the answer.
2026-08-16 18:14:53 +08:00
zhanghanduo
c3f8e5aa67 fix(apodex): correct model metadata and the cancel-response rebuild
Cross-checked the provider against platform.apodex.ai/docs and a live
GET /v1/models call.

- apodex-1.1 and apodex-1.1-mini advertised 256K max output; /v1/models
  reports 65536, and max_tokens is the legacy alias of max_output_tokens
- apodex-1-1-deep-discover is Responses-API-only; /v1/chat/completions
  answers 400 unsupported_api for the Discover tiers
- core models do not support response_format, so state it explicitly
- transform_cancel_response_api_response carried Content-Encoding over to
  a response whose body it had already replaced, so httpx tried to
  decompress plain JSON on read. A non-JSON body (the gateway answers a
  timed-out cancel with an HTML 504) also escaped as a pydantic
  ValidationError instead of the provider error
- drop the undocumented response.swarm.llm_delta mapping
- only the Deep Research tiers default stream to true; the core models
  follow OpenAI and default it to false
2026-08-16 18:14:53 +08:00
zhanghanduo
4ace5c8db3 fix(apodex): route deep research through responses 2026-08-16 18:14:53 +08:00
zhanghanduo
eb42c9ea06 refactor(apodex): limit registration to 1.1 models 2026-08-16 18:14:53 +08:00
zhanghanduo
3b4ff9127b refactor(apodex): move to a Python provider with model-aware transformations
The JSON provider path applies one contract to a whole provider, which is wrong
for Apodex: its two model families take different parameters. Replaces the
providers.json entry with litellm/llms/apodex/, reverting the shared
openai_like machinery to its original state.

/v1/responses is now keyed off the model. Core models are a stateless subset,
so store is pinned false and previous_response_id / background are rejected
rather than passed upstream to fail with a 400. The deep research tiers keep
all three, so background survives a client disconnect.

/v1/messages resolves per model too. Apodex serves the protocol natively for
the core models only, so the deep research tiers get no native config and fall
back to translation instead of hitting a path that does not serve them.

Chat completions pin stream to false for both families, drop tool params on the
deep research tiers, and rename max_completion_tokens to max_tokens. The
responses config also stops inheriting OpenAI's OPENAI_API_KEY fallback, which
would otherwise forward an unrelated OpenAI key to Apodex.

Tests live under tests/test_litellm/llms/apodex/ and touch no existing test file.
2026-08-16 18:14:53 +08:00
zhanghanduo
20ef10e920 feat(providers): add Apodex as an OpenAI-compatible provider
Registers apodex via providers.json with /v1/chat/completions, /v1/responses
and native /v1/messages, plus price map entries for the two core models and
the six deep research tiers.

Apodex defaults `stream` to true on both /v1/chat/completions and /v1/responses,
so a non-streaming litellm call would get SSE back and fail to parse it. Adds a
`send_explicit_stream_false` special-handling flag that pins the field on the
wire, and rewrites the JSON provider param mapping to build its result instead
of mutating the caller's dict.
2026-08-16 18:14:53 +08:00
Marty Sullivan
7dbf2d57c5 fix(bedrock): read batch usage by payload shape, not by provider name
Every bedrock batch output line went through the Anthropic usage parser, which
reads snake_case input_tokens/output_tokens. Converse-family models (Nova and
friends) report camelCase inputTokens/outputTokens, so their usage came back
0/0/0 and the batch billed $0 despite real token consumption.

Usage is now selected by the shape of the payload: a Converse-shaped block goes
through the same transform the live Converse path uses, so a batch and an
equivalent non-batch call agree on tokens, including cache reads and writes.
Anthropic-shaped bedrock output is unchanged.

A shape neither parser understands (an InvokeModel-native payload from Titan,
Cohere, or Llama, which name their counts differently again) still reads zero,
but now warns with the keys it saw instead of silently billing $0.

Exposes the Converse usage transform as public, since batch parsing is a second
legitimate caller; that also removes the private-member access invoke_handler
was already making.
2026-08-16 04:18:22 -04:00
mateo-berri
255dad6716 chore(typing): drop 1.3k basedpyright errors across 30 Any hotspot files
Replace Any-typed seams with real types in files carrying the highest
remaining reportAny/reportExplicitAny density after #34745: the proxy
server and its utils, the router, the streaming handler and chunk builder,
litellm_logging, the redis cache, the MCP db/tool-registry/spend-writer
layer, the anthropic pass-through adapters and guardrail translation, the
lasso and presidio guardrail hooks, the azure_ai agents handler, the
management endpoints (keys, users, ui_sso, model access groups, config
override, MCP, projects), the responses MCP handlers, response polling
background streaming, and the containers and vector stores mains

No casts, no type: ignore, no noqa, no new suppressions, and no Any
annotations that were not already at base. Whole-tree basedpyright:
reportAny 14,610 -> 14,009, reportExplicitAny 5,100 -> 4,780, total
144,743 -> 143,471, with no rule increasing repo-wide or in any file.
Budgets ratcheted: basedpyright -1,272 across 48 rules, ruff-strict -85,
type-discipline -37
2026-08-16 03:56:02 +00:00
yuneng-jiang
973329e986
Merge pull request #37069 from BerriAI/litellm_/frosty-goldwasser-93a2a1
Some checks failed
Code Quality Checks / code-quality (push) Has been cancelled
UI Unit Tests / ui-unit-tests (push) Has been cancelled
CI Coverage / assert-ci-coverage (push) Has been cancelled
Unit Tests: Core Utilities / core-utils (push) Has been cancelled
Publish basedpyright base counts / publish (push) Has been cancelled
GitHub Actions Security Analysis / zizmor (push) Has been cancelled
Unit Tests: Documentation Validation / documentation (push) Has been cancelled
Unit Tests: Enterprise, Google GenAI & Routing / enterprise-routing (push) Has been cancelled
Unit Tests: Integrations (Callbacks & Logging) / integrations (push) Has been cancelled
Unit Tests: LLM Provider Transformations / Vertex AI (push) Has been cancelled
Unit Tests: LLM Provider Transformations / All Other Providers (push) Has been cancelled
Unit Tests: MCP, Secrets, Containers & Misc / misc (push) Has been cancelled
Unit Tests: Proxy Auth & Key Management / proxy-auth (push) Has been cancelled
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Has been cancelled
Unit Tests: Proxy API Endpoints / proxy-endpoints (push) Has been cancelled
Unit Tests: Proxy API Endpoints / proxy-server (push) Has been cancelled
Unit Tests: Proxy Infrastructure / proxy-infra (push) Has been cancelled
Unit Tests: Responses, Caching & Types / responses-caching-types (push) Has been cancelled
Unit Tests: Proxy DB Operations / auth-checks (push) Has been cancelled
Unit Tests: Proxy DB Operations / budgets (push) Has been cancelled
Unit Tests: Proxy DB Operations / custom-logging (push) Has been cancelled
Unit Tests: Proxy DB Operations / db-and-spend (push) Has been cancelled
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Has been cancelled
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Has been cancelled
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Has been cancelled
Unit Tests: Proxy DB Operations / key-generation (push) Has been cancelled
Unit Tests: Proxy DB Operations / logging-misc (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-runtime (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-server-core (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-utils (push) Has been cancelled
test(e2e/ui): assert the log drawer chevrons by their lucide classes
2026-08-15 17:59:12 -07:00
Yuneng Jiang
de2b220c36
test(e2e/ui): assert the log drawer chevrons by their lucide classes
The log details drawer moved off Ant Design in 03d2b16bc, so its section
header renders lucide ChevronUp/ChevronDown rather than antd's UpOutlined
and DownOutlined. The collapse test still waited on .anticon-up and
.anticon-down, which no longer exist anywhere under view_logs, so it
failed on every run and burned all three attempts identically.

Point the three assertions at .lucide-chevron-up and .lucide-chevron-down,
matching how the dashboard's other suites address lucide icons.
2026-08-15 17:50:40 -07:00
yuneng-jiang
ae8afec7c1
Merge pull request #37066 from BerriAI/litellm_/release-ui-build-1e15d9
chore: rebuild Admin UI bundle from litellm_internal_staging
2026-08-15 17:38:54 -07:00
mateo-berri
904ff9efa7 Merge branch 'litellm_internal_staging' into devin_ai_fix_bedrock_batch_file_bytes_36388
Some checks failed
Terraform Modules / fmt, validate, test (aws) (push) Has been cancelled
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Resolves the transform_create_file_response conflict by keeping the
_uploaded_object_size handoff over the response Content-Length read,
and adds the rebind-ok justification LIT011 now requires for the
upload-size litellm_params handoff after the base budget ratcheted.
2026-08-15 17:31:52 -07:00
mateo-berri
4114f907ea fix(bedrock): reraise cancel validation errors for non-terminal jobs, allow bedrock in acancel_batch typing 2026-08-15 17:25:16 -07:00
Yuneng Jiang
b3077c9dc0
chore: update Next.js build artifacts (2026-08-16 00:22 UTC, node v24.19.0) 2026-08-15 17:22:48 -07:00
yuneng-jiang
992a8123ac
Merge pull request #37010 from BerriAI/litellm_shadcn_next_0814
fix(ui): de-duplicate the reset budget option and polish shadcn surfaces
2026-08-15 17:18:25 -07:00
mateo-berri
57e946f279 Merge origin/litellm_internal_staging into fix/bedrock-cancel-batch 2026-08-15 17:18:08 -07:00
yuneng-jiang
91aee78e78
Merge pull request #37065 from BerriAI/litellm_/nice-wilson-9fbed6
test(e2e): assert provider error shape instead of pinned prose
2026-08-15 17:15:30 -07:00
Mateo Wang
dddee7d848
Merge pull request #37063 from BerriAI/litellm_pr_template_proof_format
docs(github): proof-of-fix section shows only the latest run as Before/After with nested cases
2026-08-15 17:13:12 -07:00
yuneng-jiang
1968562733
Merge pull request #37059 from BerriAI/litellm_/circleci-pipeline-triage-9b92e5
test: unstick the suites CircleCI is failing on
2026-08-15 17:12:46 -07:00