Commit graph

53260 commits

Author SHA1 Message Date
mateo-berri
79bda61212 fix(spend-logs): read used_client_oauth_token from the bucket the route stamped
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Waiting to run
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Waiting to run
A guardrail on the unified path adds litellm_metadata to a chat request after
the proxy stamped metadata, so both spend row writers read the new bucket and
stored null. The success row now resolves the flag the same way the callback
payload does, and the failure row picks the bucket from the request route.
2026-09-28 17:16:34 -07:00
mateo-berri
c4f6d82a4f fix(logging): read used_client_oauth_token from the proxy-stamped metadata slot
On routes that carry proxy metadata in litellm_metadata, metadata is the
caller's own body field, and merge_litellm_metadata lets it win. Resolve the
flag from litellm_metadata when the proxy stamped it there so a caller cannot
set it in the standard logging payload
2026-09-28 15:30:07 -07:00
mateo-berri
f538d135f0 Merge remote-tracking branch 'origin/main' into litellm_spend_log_client_oauth_flag 2026-09-28 14:33:20 -07:00
devin-ai-integration[bot]
317430db4e
fix(panw_prisma_airs): honor experimental_use_latest_role_message_only on every request shape (#42447)
* fix(panw_prisma_airs): apply experimental_use_latest_role_message_only to every request shape

Explicit true/false now applies to chat completions, Anthropic /v1/messages and /v1/responses alike; unset keeps latest-only for Anthropic and full history otherwise. Text indices are mapped back to their source message by value instead of by count, so Responses instructions, function_call_output and reasoning items no longer derail the alignment and silently rescan the whole history

Co-authored-by: scthornton <scthornton@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(panw_prisma_airs): type latest-message helpers against AllMessageValues

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): require forward and reverse text attribution to agree

A Responses function_call_output whose text equals the latest user turn could claim that turn's slot in a forward-only walk and demote the latest-only scan to an earlier message. Walk both directions and fall back to the full role-filter scan when they disagree

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): pick the latest human turn from messages, not from aligned texts

An image-only latest user turn no longer promotes an earlier user turn into the
latest-only scan; it scans nothing on the request side, as the Anthropic path did
before. A latest user/developer message whose text never reached texts (a trailing
Responses reasoning item) falls back to the role-filter scan instead of narrowing.
Types the test helpers, drops the narrating docstrings and adds regressions for both
shapes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): log when latest-only selection leaves nothing to scan

An image-only latest user turn with experimental_use_latest_role_message_only=true intentionally yields zero scanner calls. Emit a debug line naming the call_id so operators can tell this apart from the guardrail not firing.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): keep Responses reasoning items out of latest-turn selection

The Responses translation handler gives reasoning input items the default user
role, so a reasoning item with text content after the latest prompt was picked
as the latest human turn and the real prompt went unscanned under
experimental_use_latest_role_message_only. Map reasoning items back to their
texts positions from the raw input and exclude them; fall back to the
role-filter scan when the raw items do not account for every text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: scthornton <scthornton@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:16:18 +00:00
Louis Vauterin
b9ba36c231
fix(provider): accept 2xx status codes in unified_access_group create (#42461)
* fix(provider): accept 2xx status in handleResponse for access group create

handleResponse (resource_team.go) only accepted exactly HTTP 200, but
POST /v1/access_group (and its /v1/unified_access_group alias) legitimately
answers 201 Created. litellm_access_group and litellm_unified_access_group
create both succeeded on the proxy and failed in the provider, leaving the
group out of state and forcing an import to recover on the next apply's
409 for the now-duplicate name.

Same fix and shape as #40723, which widened this exact check in
sendRequest/handleAPIResponse/handleMCPAPIResponse for mcp_server, model,
key and organization_member. handleResponse is the one shared status-check
helper that fix didn't reach -- it's a different function in a different
file (resource_team.go, not client.go/utils.go), so this is fully
independent of that PR and can land before, after, or alongside it with
no conflict.

handleResponse is also used by agent, budget, guardrail, organization,
prompt, search_tool, tag, team, team_block, key_block, team_member(_add)
and user -- all unaffected in practice, since every one of their own
endpoints already answers exactly 200. Widening the check costs them
nothing and only changes behavior for the two resources that were
actually broken.

Verified: go test ./... passes, including a new
TestHandleResponseAcceptsFullSuccessRange table test covering
200/201/202/204 (accepted) and 400/404/409/500 (still rejected), mirroring
#40723's own TestHandleAPIResponseAcceptsFullSuccessRange.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: address Greptile review on PR 42461

- CHANGELOG: narrowed the fix's scope to unified_access_group only.
  litellm_access_group (legacy) calls /access_group/new, a completely
  different, unrelated endpoint (model_access_group_management_endpoints.py)
  that already returns 200 -- it was never affected. I'd wrongly assumed
  both resources shared the same /v1/access_group route; they don't.
- resource_team_test.go: removed the preamble comment above the new test,
  which restated what the table test already shows -- against repository
  guidance (AGENTS.md) that reserves comments for complex logic, tool
  inputs, or TODOs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: retrigger CI

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-28 13:54:37 -07:00
tin-berri
6d4ccf7e97
refactor(mcp): consolidate hub publication predicate (#43394) 2026-09-28 13:26:59 -07:00
berriai-litellm-provider-info-sync[bot]
9fd78ff6f4
fix(cost-map): add web search flag and model page source to anthropic claude-sonnet-5-5 (#43584)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-28 19:36:17 +00:00
Yassin Kortam
fe76c2473d
Revert "feat(proxy): server-side Team Usage export beyond the top-N key cap (#42996)" (#43376)
This reverts commit 77eccaca78.
2026-09-28 12:33:18 -07:00
berriai-litellm-provider-info-sync[bot]
875685bf26
chore(cost-map): take azure limits for deepseek-v4-flash-0731 and v3.2-speciale (#43597)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-28 12:31:39 -07:00
yuneng-jiang
f20c400374
fix(proxy): unregister logging callbacks removed from the stored config (#43428)
* fix(proxy): unregister logging callbacks removed from the stored config

POST /config/callback/delete saved the config and resynced, but the resync only
ever added callbacks, so a deleted callback kept exporting and kept showing in
/get/config/callbacks as read-only on every worker.

ProxyConfig now tracks which callback list entries each DB config sync
registered and unregisters them once the stored config stops listing them.
Callbacks it did not register (YAML, code) are never touched, and a failed
config load skips the sync instead of treating the config as empty.

* refactor(proxy): keep callback sync comprehensions to one for clause

* fix(proxy): restore code-registered callbacks the DB sync replaced

Registering a custom-logger callback from the DB swaps an existing string
entry for a logger instance. Deleting the DB entry then removed the instance
and left the code-registered callback gone. The sync now records the entries
it displaced and puts them back when it unregisters.
2026-09-28 12:22:01 -07:00
devin-ai-integration[bot]
e4190d86a6
refactor(rust): centralize host execution and compose callbacks (#43515)
* refactor(rust): extract litellm-host-native as the shared Rust host driver

Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): interrupt the machine when the in-process stream consumer fails

Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): separate the machine contract from coroutine execution

* auth update

* refactor(rust): use standard flow control for host requests

* style(rust): keep host driver imports formatted

* chores

* mostly relocation

* refactor(rust): separate interceptors from queued observers

* refactor(rust): centralize legacy callback mappings and lifecycle

* docs: define Python host boundaries and migration plan

* refactor: enforce Python host and bridge boundaries

* refactor(rust): separate operations from callback composition

* refactor(rust): compose SDK policy through call hooks

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:20:27 +00:00
yuneng-jiang
37be82e45e
test(unit): stop test modules from putting their own directory on sys.path (#43421)
test_ssl_verify_unit.py inserted tests/unit at the front of sys.path, so any later import of litellm_proxy_extras resolved to the tests/unit/litellm_proxy_extras test package. Whenever the CircleCI shard split collected that file before test_litellm_proxy_extras_logging.py, collection failed with ModuleNotFoundError. test_gemini_session_leak.py had the same insert for its own directory
2026-09-28 12:20:20 -07:00
devin-ai-integration[bot]
2e5034016f
fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys (#43587)
* feat(anthropic): add Claude Sonnet 5.5

Adds the anthropic cost map entry for claude-sonnet-5-5 mirroring
claude-sonnet-5 pricing and capabilities, with prompt_cache_min_tokens
at 512, thinking_always_on (thinking cannot be disabled on this model),
and supports_forced_tool_use false (tool_choice required/named returns
400 upstream). Omits thinking cache preservation, same as Opus 5.5, and
registers the model in the setup wizard provider list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys

Sets prompt_cache_min_tokens 512, thinking_always_on, and
supports_forced_tool_use false on every anthropic, bedrock, vertex_ai,
and azure_ai Sonnet 5.5 key, dropping the thinking cache preservation
flag cloned from Sonnet 5. Removes unpublished deprecation dates on
azure_ai and vertex_ai, renames the OpenRouter key to the live
anthropic/claude-sonnet-5.5 id and drops its batch variant, and removes
the aihubmix, deepinfra, and databricks keys for vendors that do not
list the model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drop vendor-absence assertions for Sonnet 5.5 keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:12:31 +00:00
devin-ai-integration[bot]
3726ce2cfc
refactor(guardrails): fix agent 365 to the production endpoint and log the opt-in fail_open at error level (#43189)
* feat(guardrails): fail open by default when Agent 365 cannot evaluate and count it in Prometheus

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(tests): ruff format the Prometheus fail-open registry test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(guardrails): add Agent 365 authority host override, fail-open integration test and per-guardrail YAML default

Add `authority_host` to the Agent 365 config (also read from AGENT365_AUTHORITY_HOST, then AZURE_AUTHORITY_HOST) so sovereign clouds and the integration test can point the OBO exchange at a different Entra host.

Add tests/integration/mcp/test_mcp_agent_365_guardrail.py, a real proxy test with Postgres, Redis, a scripted MCP upstream and local Entra and Agent 365 doubles covering the default fail-open, explicit fail-closed and fail-open, Defender Skipped, policy denial, persisted status and Prometheus counter.

Use PrometheusLogger.get_instance for the fail-open metric lookup instead of a hand-rolled callback scan. Clarify the config description: gateway credential failures fail open, caller token failures block.

Extract the dashboard YAML preview into teamGuardrailConfigYaml.ts so the effective per-guardrail default is unit tested and the "default" hint only shows when nothing was set explicitly.

Regenerate the lazy OpenAPI snapshot and schema.d.ts for the new field.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): default a scheme-less Agent 365 authority host to https and treat a null fallback as unset in the YAML preview

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for the Agent 365 fail-open default across entry points, Entra faults, throttling and two workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prove both Agent 365 workers serve and that a killed worker is replaced

Each fresh connection reports its worker pid from /debug/memory/summary and its MCP catalog on the same
connection, so the two-worker readiness wait covers both workers by identity. The kill test now kills a
pid the proxy reported as a worker and waits for a replacement pid, instead of the first psutil child

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): drop the prometheus fail-open counter from the agent 365 guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): keep agent 365 fail closed by default and make fail_open an explicit opt-in

Restores the shared unreachable_fallback default and the sibling guardrail initializers, drops the Admin UI YAML preview that only existed for the per-guardrail default, and reworks the unit and integration tests so the default blocks with HTTP 503 while unreachable_fallback: fail_open lets availability failures through as Unscanned

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): append authority_host after the existing Agent365Guardrail parameters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(guardrails): agent 365 fails open by default and hides the production overrides from the UI form

Agent 365 sits in the runtime path of every MCP tool call, so an Entra or
Agent 365 outage now lets the call through unscanned (logged at error level,
recorded as Unscanned with guardrail_failed_to_respond) instead of blocking it.
unreachable_fallback: fail_closed stays as the opt-in strict mode. Policy
blocks, throttling, 4xx rejections and a rejected caller token still block

The shared unreachable_fallback field becomes nullable so each guardrail owns
its default; every sibling still resolves None to fail_closed and typesafe
keeps failing open

api_base, resource_app_id and agent_id have production defaults and leave the
dashboard form (ui_hidden); they stay available in config.yaml and env. The
authority_host override and its env keys are gone, the OBO exchange always
uses login.microsoftonline.com. The integration suite keeps only the cells
that need no Entra double, the evaluation paths live in unit tests with an
injected handler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate openapi snapshot and schema.d.ts for the nullable unreachable_fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): fix agent 365 to the production endpoint and keep fail_closed as the default

Remove api_base, resource_app_id and agent_id from the Agent 365 config model, their AGENT365_* env fallbacks and the _is_ui_hidden helper: the evaluation URL and the Agent Tools app id are fixed production constants and the agent identity is always the caller's key alias. Revert the fail_open default; unreachable_fallback: fail_open stays an explicit opt-in. Restore the shared unreachable_fallback field, the sibling guardrail initializers and typesafe to main. Move the Entra dependent cells from the subprocess integration suite to unit tests with an injected HTTP handler.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): warn when agent 365 yaml still carries the removed override keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): inject the http handler into the agent 365 initializer instead of assigning it after construction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 12:05:47 -07:00
ryan-crabbe-berri
6c34c6e7ee
test(integration): pin org-admin status codes in the team-admin matrix (#43592)
Run every route in the team-admin matrix again on a team that belongs to an
organization, as an admin of that organization and as an admin of another
organization. 90 new cases pin what org admins get today, so collapsing the
team-admin helpers into one gate can prove parity for org admins too

Scenario gains organization() and org_member() helpers that clean up the
organization, its budget row and the member users
2026-09-28 19:04:05 +00:00
Krrish Dholakia
0f96d09588
feat(model_prices): add claude-sonnet-5-5 model pricing and capabilities (#43586)
Source: https://docs.anthropic.com/en/docs/about-claude/models
2026-09-28 18:19:41 +00:00
devin-ai-integration[bot]
703eb4fa68
security(proxy): keep team callback credentials out of the stored request body (#43217)
* security(proxy): keep team callback credentials out of the stored request body

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: allow the security conventional commit type

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 11:16:58 -07:00
devin-ai-integration[bot]
6f5ad78a1f
fix(cost-map): registry audit 2026-09-28, openai deep-research shutdown dates, azure deepseek v4.1 flash direct price, vertex gemini 3.8 live avatar video price, drop azure_ai/muse-spark-1.3 (#43566)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: KaiyiQuan <KaiyiQuan@users.noreply.github.com>
Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-28 09:46:37 -07:00
devin-ai-integration[bot]
f191e08d67
fix(proxy): log key owner identity on expired key auth failures (#43105)
* fix(proxy): log key owner identity on expired key auth failures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): escape control characters in logged key identity fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): gate key identity in auth failure logs behind log_auth_failure_key_identity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): apply log_auth_failure_key_identity from DB config reloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 08:18:48 -07:00
devin-ai-integration[bot]
c60c714278
test(rust): enforce shared upstream error contract in wheel checks (#43520)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-28 08:15:41 -07:00
devin-ai-integration[bot]
2c9b0e00ac
refactor(types): replace Any with proven types in 8 files (#43551)
* refactor(types): replace Any with proven types in 8 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep email logger untyped where its alert types differ

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 04:26:41 -07:00
devin-ai-integration[bot]
90e4962c81
refactor: clean up fresh tech debt from 2026-09-27 (#43538)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 01:28:40 -07:00
berriai-litellm-provider-info-sync[bot]
b79fc9f1b0
feat(azure): add mistral ocr pricing and azure max output limits (#43530) 2026-09-27 23:07:45 -07:00
devin-ai-integration[bot]
74cad08997
refactor(rust): remove delivery routing abstraction (#43514)
* wip

* refactor: finish removing delivery routing abstraction

* refactor(rust): remove chat completion decline admission

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-28 03:31:22 +00:00
devin-ai-integration[bot]
f184ace25b
feat(rust): add the MCP gateway (#43470)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:37:41 -07:00
devin-ai-integration[bot]
6e0926edde
feat(rust): add gateway UI login and sessions (#43469)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:37:41 -07:00
devin-ai-integration[bot]
ed43556e92
feat(rust): add virtual key storage contracts (#43468)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:18:59 -07:00
devin-ai-integration[bot]
5f637a2b11
feat(rust): separate gateway authentication and authorization (#43467)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 19:18:59 -07:00
devin-ai-integration[bot]
876539e1b3
feat(rust): add structured route lifecycle tracing (#43466)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 19:18:58 -07:00
devin-ai-integration[bot]
8ab124309f
feat(rust): connect Python inference bindings to shared routes (#43465)
* feat(rust): support the HTTP Responses API

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(rust): enable native Python inference opt-in

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust_bridge): cover only python-only routes in the native load guard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep Python inference rollout disabled

* fix(rust): preserve inference defaults and continuation IDs

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 19:05:39 -07:00
devin-ai-integration[bot]
9b08112ed2
feat(rust): add litellm-db and litellm-db-testing workspace scaffolding (#43504)
* feat(rust): add litellm-db and litellm-db-testing workspace scaffolding

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(db-testing): apply the real Prisma migrations in a test and drop the sort mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 18:48:45 -07:00
berriai-litellm-provider-info-sync[bot]
cc30d93984
fix(cost-map): set together_ai gpt-oss-20b and gemma-4-31B-it deprecation_date to 2026-09-15 (#43509)
Some checks are pending
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-27 17:49:49 -07:00
berriai-litellm-provider-info-sync[bot]
2566b33cea
fix(cost-map): sync openrouter prices from the models API (#43506)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-27 17:38:47 -07:00
devin-ai-integration[bot]
0a21f24d50
feat(rust): support the HTTP Responses API (#43464)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 17:38:15 -07:00
berriai-litellm-provider-info-sync[bot]
0d6b8b5ab4
fix(cost-map): add deprecation_date to together_ai Salesforce/Llama-Rank-V1 (#43507)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-27 17:25:56 -07:00
devin-ai-integration[bot]
1ceeefbf84
refactor(rust): use shared execution in gateway inference (#43463)
* feat(rust): add the HTTP host driver

* refactor(rust): use shared execution in gateway inference

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): apply rustfmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:50:05 -07:00
devin-ai-integration[bot]
438bffc26e
build(rust): package the gateway container (#43471)
* build(rust): package the gateway container

* ci: exempt the gateway Dockerfile from the CI coverage gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:49:39 -07:00
devin-ai-integration[bot]
18933c8a21
feat(rust): add the HTTP host driver (#43462)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-09-27 16:18:10 -07:00
devin-ai-integration[bot]
36784e3b79
refactor(rust): share call lifecycle across route-owned inference (#43461)
* feat(rust): expand gateway configuration parsing

* refactor(rust): unify core calls and host lifecycle

* fix(config): accept environment references for model rate limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): unify core calls and host lifecycle

* style(rust): apply rustfmt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): read environment secrets when litellm is not importable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(core): drop the duplicate rstest attribute

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): make shared route dispatch route-owned

* fix(rust): satisfy Clippy in messages regression test

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:05:30 -07:00
devin-ai-integration[bot]
b94f5bdbed
feat(rust): expand gateway configuration parsing (#43460)
* feat(rust): expand gateway configuration parsing

* fix(config): accept environment references for model rate limits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:54:04 -07:00
devin-ai-integration[bot]
268e8bb735
refactor(rust): share anthropic types, request helpers, and streaming contracts across crates (#43426)
* refactor(rust): standardize Azure Messages module path

* docs(rust): define shared types crate boundaries

* refactor(rust): share request helpers and type Anthropic blocks

* docs(rust): format shared type invariants as bullets

* test(rust): parameterize repeated cases with rstest

* refactor(rust): move Responses transform result into llms

* fix(anthropic): validate chat and batch responses

* docs(rust): clarify API format ownership boundaries

* docs: clarify Rust error message construction

* refactor(auth): keep shared Rust errors provider-neutral

* refactor(rust): separate format contracts from provider policy

* fix(rust): type Anthropic chat response text collection

* fix(rust): pass audio secret sources through hosts

* fix(rust): unblock batch lint and OCR error assertions

* test(rust): assert response failures at the adapter boundary

* refactor(rust): declare error messages with typed context

* wip

* fix(rust): adapt Bedrock error details

* style(rust): cargo fmt bedrock audio transcription

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): adapt tests and dead code to typed error details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep converse error contracts and read env secrets without litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 45 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): tolerate missing usage in converse responses on the transcription route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:53:12 -07:00
berriai-litellm-provider-info-sync[bot]
22b36cbcf6
chore(cost-map): update azure_ai/grok-4.6 input price and add azure_ai/MAI-Cyber-1-Flash (#43446)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-27 09:14:33 -07:00
berriai-litellm-provider-info-sync[bot]
ff462f7a77
chore(cost-map): update azure_ai/grok-4.6 input price from Azure pricing page (#43440) 2026-09-27 08:56:50 -07:00
devin-ai-integration[bot]
f4308bc124
refactor(types): replace Any with proven types in 5 files (#43304)
* refactor(types): replace Any with proven types in 6 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(types): keep enterprise email import inside try-except for unsafe-import check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(types): keep email_logging_instance annotation as Any pending a guarded alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert iterator override typing in proxy utils

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 01:28:02 -07:00
devin-ai-integration[bot]
8e6d99d74a
fix(token_counter): count Gemini function_declarations tools (#43417)
* fix(token_counter): count Gemini function_declarations tools

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(token_counter): skip non-dict tools when formatting definitions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 22:28:31 -07:00
Anmol Jaiswal
4274bdda44
fix(vertex_ai): stop importing the vertexai SDK in partner-model completion (#42274)
completion() imported vertexai only to check that the package exists. Partner
models are reached with an authenticated httpx client and never use that SDK,
the same reasoning count_tokens in this file already follows (#28084). The
import loads all of google-cloud-aiplatform on the first request of every
process and made a google-auth-only install fail with a 400
2026-09-26 22:14:53 -07:00
devin-ai-integration[bot]
b831e9b4ac
fix(bedrock): keep the provider status code on unprocessable image errors (#43416)
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: dbalintx <dbalintx@amazon.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 05:10:24 +00:00
devin-ai-integration[bot]
c1f761eba5
test(vertex_ai): move stray Gemma streaming tests to tests/unit so CI coverage passes (#43422)
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 05:02:44 +00:00
devin-ai-integration[bot]
9a0ff249d5
fix(anthropic): forward the per-turn-control beta to Azure AI Foundry (#43415)
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 05:02:35 +00:00
devin-ai-integration[bot]
491d454826
fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking (#43414)
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking

Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.

Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.

(cherry picked from commit bc9b6f8a5c)

* test(vertex_ai): move orphaned gemma streaming tests into the llm-vertex-ai shard

PR #43147 left a copy of the Gemma streaming tests under
tests/test_litellm/llms, a tree no CI shard claims, which broke
assert-ci-coverage and assert-shard-coverage on main. Fold the two
streaming tests into the existing tests/unit/llms/vertex_ai file so the
llm-vertex-ai shard runs them

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chloe Lu <chloe.lxd@gmail.com>
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 04:51:56 +00:00