Commit graph

54084 commits

Author SHA1 Message Date
Mateo Wang
3069a467d0
docs(agents): ban semicolons and colons in human-facing text (#45684)
* docs(agents): ban semicolons and colons in human-facing text

Tighten the writing rule so ";" and ":" are only allowed where the sentence would be nonsensical without them, rewrite the existing semicolons in AGENTS.md to follow it, and note that docs (litellm-docs) keep their trailing "."

* docs(agents): rewrite semicolons in nested AGENTS.md files

* docs(agents): hand-maintain the dashboard Next.js warning and sync e2e CONTRIBUTING lines

* docs(agents): point at the PR template path that exists

* docs(agents): restore the generated Next.js block in the dashboard AGENTS.md
2026-10-09 16:36:32 -07:00
yujonglee
4cb47edd2a
feat(rust): add the DeepSeek Anthropic Messages config (#45614)
* feat(rust): add the DeepSeek Anthropic Messages config

* refactor(rust): drop the tool discriminator without mutation and use rstest in DeepSeek tests
2026-10-09 16:29:46 -07:00
yujonglee
a0b3c46cf0
feat(rust): add the Vertex AI Anthropic Messages config (#45626)
* feat(rust): add the Vertex AI Anthropic Messages config

* fix(rust): normalize Vertex system turns the way Python does

Vertex uses the shared normalize_system_role_messages: a leading run is
hoisted, a later turn stays or is converted in place by the model's
supports_mid_conversation_system, and billing blocks are filtered after the
hoist so one inside a system-role message no longer reaches Vertex.

* fix(rust_bridge): keep non-Claude Vertex Messages calls on Python

The catalog admits every vertex_ai Messages call, but Rust serves only Claude
there and reports anything else as a terminal InvalidProvider instead of a
decline. The dispatcher now bypasses Rust for those models, as
get_provider_anthropic_messages_config does.

* refactor(rust): template the Vertex location and project errors

The location failure is the shared InvalidType detail and the missing project
is MissingSetting with the provider, setting, param and variable name, so no
sentence is spelled at the return site.

* test(rust): use rstest in the Vertex AI Messages tests

* refactor(rust): derive the missing Vertex project error from its spec
2026-10-09 16:29:45 -07:00
yujonglee
ee295aea68
refactor(rust): give Messages configs typed litellm params (#45611)
* refactor(rust): give Messages configs typed litellm params

* feat(rust): port _normalize_system_role_messages for Messages hosts

mid_conversation_system mirrors Python's module function for function: a
leading run of system turns is hoisted into the top-level system, a later
turn stays in place when the model map flags supports_mid_conversation_system
and otherwise becomes a user turn where it was, never between a tool_use and
its tool_result, and billing blocks are stripped from the hoisted system. The
capability joins MessagesModelCapabilities and the bridge projection. Azure AI
uses it in place of fold_system_role_messages, which hoisted every turn.

* fix(rust): keep the Bedrock runtime endpoint behind a blank api_base

* refactor(rust): import LitellmParams from its crate instead of a re-export

* refactor(rust): template the Messages request decode errors

Both decode failures go through ErrorDetail::invalid with the subject and the
serde error as source instead of a preformatted sentence.

* refactor(rust_bridge): read the module globals the param specs name

The Messages host no longer keeps its own list of litellm globals; a spec
whose spellings the call leaves out names the global to read, so the next
family that has one needs no host change.

* refactor(rust): give each route lookup Python's implicit None

ProviderConfigManager's per-route methods list only the providers the route
serves and fall through to None; the Rust lookups now do the same, so a new
LlmProviders variant touches only the route that serves it.
2026-10-09 16:29:45 -07:00
yujonglee
3d2abf9a44
feat(rust): one LitellmParams type for a config and a caller (#45642)
litellm-router-types mirrors litellm/types/router.py: LitellmParams holds the
model, the credentials, one flattened group per credential family from
auth-types (aws, vertex), the deployment settings a config spells, and every
other key in extra, as Python's extra="allow" keeps it. Config parses it
directly, so config::LiteLlmParams and its untyped additional_fields are gone
and a mistyped provider param is a parse error. fields() and specs() derive
from the families' param specs for hosts that project kwargs and fold module
globals. Spelled<T> is the one untagged shape for a typed value or the text a
config spells, organization takes the list Python's router expands, and the
callback shorthand keeps a value-level one-or-many. No credentials wrapper
group: serde's flatten only consumes keys for struct-shaped children.
2026-10-09 16:29:45 -07:00
yujonglee
04a95f60c1
feat(rust): type the Vertex AI connection params and where Python reads them (#45641)
VertexParams in auth-types holds the vertex_* fields of CredentialLiteLLMParams
in both spellings, with one ParamSpec per setting: the wire names, the litellm
module global VertexBase.get_vertex_ai_* consults, and the environment names.
A credential deserializes from text or the JSON object a config spells, an
empty object counts as absent, and both credential fields are redacted in
Debug. auth-gcp is split so lib.rs is the entrypoint (config, auth,
constants), its lookups resolve through the specs, secret_names() derives
from them, and get_vertex_ai_project_from_credentials reads a service
account's project_id for hosts that need it before a token exchange. The OCR
host folds the module globals in by spec at projection, so OcrSettings no
longer carries vertex_project and vertex_location.
2026-10-09 16:29:44 -07:00
yujonglee
0410abea8b
feat(rust): type the AWS connection params and where Python reads them (#45610)
AwsParams in auth-types holds the aws_* fields of CredentialLiteLLMParams
with one ParamSpec per field: its wire name and the environment names it
falls back to. fields() and secret_names() derive from the specs, the AWS
helpers take the struct instead of a map, aws_auth_config and the region
resolution read through the specs, and a missing region is
AwsParams::REGION.missing("AWS"), whose message is rendered from the spec.
Bedrock Converse, Bedrock transcription, Textract OCR and the Bedrock
Messages config build the struct at their one untyped boundary.
2026-10-09 16:29:44 -07:00
Mateo Wang
883e3210ee
feat(decisions): add Databricks ai_decide as a /v1/decisions provider and auto-router decider (#45200)
* feat(decisions): add Databricks ai_decide as a /v1/decisions provider and auto-router decider

* fix(decisions): reject Databricks endpoint names carrying URL separators and keep the provider wiring under llms/

* fix(decisions): reject dot-segment Databricks endpoint names and add Databricks to the dashboard's OSS classifier editor

* fix(ui): name every OSS classifier provider in the classifier radio copy

* feat(decisions): register the Databricks OpenJev decider as an evaluation-mode cost-map entry

* fix(cost-map): declare zero cache rates on the databricks-openjev-qwen35-4b entry

* fix(proxy): reconcile every DB deployment when a registry miss lands on an auto router

A /model/new for an auto router lands on one replica. A request routed to
that auto router on a sibling replica missed it, the model read-through
loaded only the auto router's own row, and the pre-routing hook then picked
a tier the sibling had not loaded, so the caller got a 400 "no healthy
deployments for <tier>" until the next periodic reload. An auto router miss
now runs the full add_deployment reconcile, tiers included. Found by the
audit's two-worker rig on this PR's routing cells; CircleCI's one-worker
proxy never opens the window.

* test(integration): audit cells for Databricks decisions and the Databricks auto-router classifier

Wire cells for /v1/decisions with databricks/<endpoint> deployments, routing
cells for the complexity router's databricks classifier on all three
endpoints with the Postgres rows read back, management cells for the
classifier's spend writes and for the evaluation-mode health check of a
Databricks serving endpoint.
2026-10-09 16:19:23 -07:00
moe-berri
22ad3c5fff
refactor(lens): connect LiteLLM to the independent Lens service (#45529)
* feat(lens): consume the independent Lens service and shared UI

* fix(lens): keep embedded setup stable and vendor UI before builds

* chore(lens): pin embedded UI provenance to published Lens source

* fix(lens): complete embedded UI extraction across CI builds

* docs(lens): record latest main and removal-boundary validation

* fix(lens): wire external chart credentials and ingestion modes

* test(lens): cover adapter failures and fix extraction CI

* docs(lens): refresh bundled chart setup instructions

* test(lens): restore gateway adapter test package marker

* build(lens): pin clean-cut UI and supported storage chart

* test(ui): await guardrail scope menu before checking options

* perf(lens): omit unused trace-team lookup from product requests

* docs(lens): record final-image adapter permission checks

* test(lens): assert delegated investigation authorization

* refactor(lens): remove copied trace runtime and preserve spend logging

* test(lens): verify independent gateway image upgrades and outages

* test(lens): pin the chart qualification database image

* fix(lens): update embedded UI metadata filters

* test: wait for Lens chart pod identity convergence

* fix(e2e): select Lens chart pods by deployment ownership

* fix(e2e): select Helm migration ownership for Lens upgrades

* test(lens): retain migration Jobs through rollout assertions

* docs(helm): require existing database for migration hooks

* test(lens): qualify release boundaries with embedded navigation update

* docs(lens): record browser qualification for embedded tabs

* fix(lens): remove unrelated CI and guardrail test changes

* test(lens): give qualification tenants unique key aliases

* feat(lens): embed the latest canonical Lens interface

* fix(lens): qualify split inference through the gateway

* fix(lens): preserve scoped auth and isolate delegated credentials

* chore(lens): remove unrelated documentation and lint changes

* fix(lens): bound concurrent forwarding buffers through response delivery

* fix(lens): preserve OAuth2 dispatch without inference policies

* chore(lens): adopt current shared onboarding UI

* fix(lens): refresh shared setup UI and review context

* fix(lens): bound uploads before gateway authentication

* docs(lens): remove extraction evidence from the gateway repo

* docs(lens): refresh shared setup prompt for paired releases
2026-10-09 16:17:05 -07:00
devin-ai-integration[bot]
5b8f0a4803
fix(model_prices): registry audit, perplexity agent api models and claude-sonnet-4-5 200k context (#45647)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 16:07:13 -07:00
devin-ai-integration[bot]
3ea8069624
refactor(cost-map): remove openrouter rows, part 5 of 5 (#45669)
* refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 2 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 3 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 4 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 5 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 22:48:55 +00:00
devin-ai-integration[bot]
b1d0e45f8d
feat(policy-engine): stream detect-only post_call pipeline steps live when buffering is off (#45516)
* feat(policy-engine): stream detect-only post_call pipeline steps live when buffering is off

* fix(policy-engine): scan live pipeline streams a provider error cuts short and discard rewrites per step

A live detect-only pipeline now scans the chunks the client received when the provider fails mid-stream, then re-raises the error. Each step's rewrite is discarded right after the step, so later steps scan the text the client actually received.

* fix(proxy): keep a guardrail verdict recorded after a stream failure on the failure spend row

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-09 15:46:51 -07:00
Yassin Kortam
06006932c7
fix(ui): show the Microsoft 365 Copilot logo instead of the Azure one (#45687)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 22:42:31 +00:00
devin-ai-integration[bot]
da59afee44
fix(proxy): sign RDS IAM tokens for the database's own region (#44782)
* fix(proxy): sign RDS IAM tokens for the database's own region

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): add per-connection RDS IAM signing region overrides for writer and reader

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): integration cells for per-connection RDS IAM signing regions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): opt-in real RDS e2e cells for cross-region IAM signing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): address review on RDS IAM region cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): drop recording-front integration tests, mark RDS e2e suite e2e and register its model via /model/new

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): assert the spend read lands on the replica, not just a live connection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): type generate_iam_auth_token params and retry the replica spend-read check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): derive RDS region from custom RDS Proxy endpoint hostnames

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): add Subject, step labels, frozen rows and Final to the RDS IAM e2e suite

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): drop the opt-in RDS IAM e2e suite in favor of unit coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
2026-10-09 15:41:07 -07:00
devin-ai-integration[bot]
fe3ace9a02
refactor(cost-map): remove openrouter rows, part 4 of 5 (#45668)
* refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 2 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 3 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 4 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 22:40:46 +00:00
devin-ai-integration[bot]
ed22fd929c
refactor(cost-map): remove openrouter rows, part 3 of 5 (#45667)
* refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 2 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 3 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 15:33:16 -07:00
tin-berri
8a348dafd2
fix(cost): preserve zero hourly cache write rates in pricing tiers (#45545) 2026-10-09 15:29:37 -07:00
Yassin Kortam
d1945f69da
chore(github_copilot): sync model catalog with Copilot's served models (#45386)
* chore(github_copilot): sync model catalog with Copilot's served models

* test(github_copilot): assert catalog invariants instead of pinning vendor facts

* fix(github_copilot): keep thinking capabilities on served Claude catalog rows

* fix(github_copilot): keep supports_reasoning on catalog rows that shadow the family fallback

* fix(model-catalog): preserve merged Copilot Haiku metadata

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 15:19:08 -07:00
Yassin Kortam
d36b66e19e
feat(github_copilot): per-user GitHub OAuth connections for Copilot credentials (#45241)
* feat(github_copilot): per-user GitHub OAuth connections for Copilot credentials

Related to LIT-9306

* fix(ui): make per-user GitHub Copilot OAuth credential-only in Add Model

* fix(proxy): render LiteLLM_UserProviderCredentials prisma spans

* style(ui): prettier-format add_model tests

* fix(lint): annotate connection endpoints and narrow per-user credential excepts

* fix(lint): satisfy type-discipline gate for per-user copilot code

* fix(types): clear basedpyright and strict-lint gate regressions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(types): keep runtime guards on untyped responses input

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: ruff format per-user Copilot files

* fix(lint): keep type-discipline suppressions on flagged lines

* fix(github_copilot): resolve team and deployment-id per-user routes, share token cache invalidation

* test(github_copilot): scope response-cache mutation via monkeypatch, assert route allowlist

* fix(github_copilot): fall back to the database on Redis errors and skip worker-local token caching

* fix(lint): suppress mutable return annotation after base merge

* style: shorten rebind-ok reason so the dispatch line stays formatted

* fix(github_copilot): harden per-user token cache, overrides, fallbacks, headers and device flow

* fix(lint): suppress MutableMapping param on session auth pinning

* test(github_copilot): narrow pytest.raises to HTTPException for service-key override

* fix(github_copilot): fail closed on revoke errors, strip caller auth headers, widen fallback discovery

* fix(github_copilot): verify revoke tombstone, bound fallback discovery, cover aliased fallbacks

* fix(github_copilot): make fallback discovery a router superset and guard credential override at deployment

* style(router): drop explanatory comment on per-user credential guard

* fix(github_copilot): gate fallback discovery on per-user credentials and bound aggregate targets

* fix(github_copilot): count every fallback mapping toward the discovery limit

* fix(github_copilot): suppress banned typing cast imports with per-line reasons

* test(github_copilot): cover per-user 401 cooldown skip through the router failure callback

* test(github_copilot): cover per-user 401 cooldown skip through router and fallback paths

* fix(github_copilot): cover auto-router discovery, per-user fallback cooldown and connect cache race

* fix(github_copilot): cover every strategy router kind and verify the connect cache write

* fix merge conflict marker remnant in credential endpoint tests

* fix(github_copilot): do not purge user connections on label-only credential updates

* fix(github_copilot): read user provider credentials from the writer engine

* test(github_copilot): move per-user auth type form test to the integration tier

* test(github_copilot): annotate new credential regression test locals as Final

* fix(github_copilot): type per-user OAuth JSON and DB boundaries

* fix(github_copilot): validate per-user OAuth boundaries and restore truncated comments

* style(github_copilot): complete cut-off suppression reasons on new casts

* refactor(github_copilot): annotate new boundary-typing locals as Final

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 15:19:07 -07:00
Yassin Kortam
fbf4af301c
feat(microsoft_365_copilot): add Microsoft 365 Copilot chat provider with OAuth token exchange (#45158)
* feat(microsoft_365_copilot): add Microsoft 365 Copilot chat provider with OAuth token exchange

* fix(ui): create a credential for credential-only auth types in Add Model

* fix(microsoft_365_copilot): accept max_tokens and list the chat model

* fix(microsoft_365_copilot): register provider model set

* fix(ui): hide other auth types' fields when credential-only types are filtered

* fix(health): do not inherit stored auth settings when a connection test brings its own api_key

* docs(health): restore connection-test inheritance docstrings

* fix(lint): suppress justified provider boundary casts

* fix(lint): satisfy M365 type-discipline gate

* fix(types): eliminate type-check gate regressions

* style: shorten suppression reasons so ruff format leaves them on one line

* refactor(types): drop shared-file type widening unrelated to the Copilot provider

* fix(health): pass caller headers as secret_fields to connection-test probes

* fix(types): type connection-test probe and token counter locally

* fix(lint): drop cast import from connection-test header getter

* chore: remove explanatory comments flagged by review

* fix(proxy): gate OAuth client credentials behind proxy admin

* fix(router): skip cooldown on caller-scoped OAuth auth failures

* fix(router): skip fallback cooldown on caller-scoped OAuth auth failures

* fix(router): type caller-scoped cooldown lookups

* fix(ci): include Microsoft 365 Copilot tests in unit shard

* fix(proxy): keep saved endpoint when test_connection retests a deployment by id

* fix(m365): collapse doubled Graph replies

* fix(m365): preserve trailing system messages

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 15:19:07 -07:00
devin-ai-integration[bot]
2248fcefda
refactor(cost-map): remove openrouter rows, part 2 of 5 (#45666)
* refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost-map): remove openrouter rows, part 2 of 5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 15:00:51 -07:00
devin-ai-integration[bot]
9f5de05b9d
fix(bedrock): drop stop sequences for GPT 5.6 and newer instead of forwarding them to a 400 (#45664)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 15:00:24 -07:00
devin-ai-integration[bot]
dcd0630e98
perf(proxy): keep stream end output and Redis cache debug strings off the event loop (#45252)
* perf(proxy): keep end-of-stream output assembly and Redis pipeline logging off the event loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(caching): serialize Redis pipeline values with orjson

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(proxy): resolve model-level guardrails once per streamed request

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): type the streamed guardrail cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): fall back to json.dumps when orjson is not installed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): cover the json.dumps fallback for values orjson rejects

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep new guardrail cache and orjson helper within the type gates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(caching): guard Redis print_verbose value interpolation behind is_debugging_on

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert(caching): drop the orjson fast path from Redis pipeline writes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): drop the debug-on print_verbose test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* revert(proxy): drop the per-request guardrail lookup cache from this PR

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(logging): let print_verbose take %s args so Redis cache values format only when verbose

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): type set_cache key/value and redis_version so print_verbose args are known to pyright

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 14:56:16 -07:00
devin-ai-integration[bot]
231627ecd9
test(mcp): make scope-discovery server aliases digits-only (#45675)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-09 21:53:43 +00:00
devin-ai-integration[bot]
1f1b28fc8e
refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block) (#45665)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 21:51:48 +00:00
Yassin Kortam
aa64b07281
feat(ui): add Errors tab to Usage with failure rate, status and identity breakdowns (#45396)
Moves failure analytics off the cost overview into a dedicated admin-only
Errors tab: error rate with 4xx/5xx/429/401-403 callouts, error rate over
time, failed requests per day stacked by HTTP status code with a legend,
failures by status code, and a Failures by identity table switchable between
virtual keys, teams, users and models with error rate and top status per
row, linking to Logs. Reads GET /gateway/errors/activity. The Overview keeps
its plain ok/failed counts.

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 14:27:57 -07:00
Yassin Kortam
2df5b07308
feat(proxy): track failed requests by HTTP status and caller (#45244)
Counts failed gateway requests by HTTP status at the ASGI edge
(LiteLLM_DailyGatewayFailedRequests) and rolls failures up per virtual key,
team, user and model group with the status the logging callbacks recorded
(LiteLLM_DailyRequestErrors, fed from the spend writer, flushed like the
gateway counters, directly or through Redis with a pod lease, and on
shutdown). GET /gateway/daily/activity gains by_status_code; the new
admin-only GET /gateway/errors/activity returns per-day totals with
per-status counts, failures by status code, and keys, teams, users and
model groups ranked by failures with their top status. Dashboard schema.d.ts
regenerated.

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 14:27:56 -07:00
devin-ai-integration[bot]
43f07d0bd8
fix(bedrock): list the account's invocable models behind bedrock/* when check_provider_endpoint is on (#43981)
* fix(bedrock): list the account's invocable models behind bedrock/* when check_provider_endpoint is on

* fix(bedrock): sign the listing query RFC 3986 style and raise BedrockError on a failed listing

* test(utils): keep test_utils.py at main's formatting, only the Bedrock discovery test is new

* fix(bedrock): chain BedrockModelInfo's constructor so the passthrough config keeps its event parser

* fix(bedrock): list discovered ids without the provider prefix so partial wildcards filter

get_valid_models returns bare vendor ids for every other provider and for the static bedrock catalog, and the proxy's wildcard expansion adds the provider prefix itself. The lister prefixed its ids, so a partial wildcard like bedrock/anthropic.* never matched the proxy's filter and listed uncallable bedrock/anthropic.bedrock/<id> entries

* test(bedrock): integration cells for wildcard discovery through the proxy and the SDK

* fix(proxy): keep a partial wildcard a filter when its deployment repeats the prefix

The wildcard expansion guessed filter-or-alias by whether any provider id carried the prefix. A deployment such as model_name bedrock/anthropic.* over model bedrock/anthropic.* can only ever route names that share the prefix, so when the account has no on-demand anthropic.* id (every Anthropic model behind an inference profile) the guess fell into the alias branch and listed bedrock/anthropic.<every id>, none of them callable. A deployment whose model repeats the suffix now always filters, and lists nothing when nothing matches

* fix(proxy): prefix every discovered id under a custom wildcard prefix that starts a vendor id

* test(bedrock): Final and read-only annotations in the wildcard discovery tests

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-09 14:14:17 -07:00
devin-ai-integration[bot]
8080601a95
feat(bedrock): default Grok on Bedrock to native Chat Completions, mint unique tool call ids and drop stop (#45473)
* fix(bedrock): route Grok 4.7 tools with reasoning to native Chat Completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): mint unique tool call ids on native Chat Completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): send a blank text block for Converse tool results with no supported content

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(bedrock): default Grok on Bedrock to native Chat Completions and leave Converse unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): gate tool call id minting to Grok on native Chat Completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(bedrock): check the Grok tool call id gate at both call sites

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): remint positional tool call ids on native Chat Completions for every model

AWS returned call_0, call_1 for GPT 5.6 as well as Grok on 2026-10-09 (and unique ids for GPT 5.6 earlier the same day), so the id format is not fixed per model. Drop the Grok-only gate and remint any call_<digits> id on the route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): drop stop sequences for Grok instead of forwarding them to a 400

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 20:56:03 +00:00
devin-ai-integration[bot]
097d018dbf
test(decisions): add hosted_vllm /v1/decisions and provider-400 translation cases (#45661)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 20:42:58 +00:00
ryan-crabbe-berri
4c78023db2
feat(teams): let proxy admins grant team admins the right to raise their team budget (#45404)
* feat(teams): let proxy admins grant team admins the right to raise their team budget

Adds a raise_max_budget entry to team_admin_editable_team_fields. With max_budget alone a team admin can still only keep or lower the team budget. Adding raise_max_budget lets them raise it, capped by the organization's max_budget when the team belongs to one, while removing the budget stays with proxy admins. The entry is rejected without max_budget, /team/info reports may_raise_max_budget to the caller, and the Admin UI nests the new checkbox under Max Budget with a tooltip and shows the team admin a hint under the budget field.

* fix(ui): shorten the raise_max_budget tooltip and link it to the docs

* fix(teams): pin the organization in the guarded team budget write

A granted raise is checked against the team's organization at read time, so the write now also requires organization_id to be unchanged. A concurrent move into a budgeted org returns 409 instead of landing an uncapped raise. Also types the new test helpers.

* test(teams): type the team row store methods and the touched race test arguments

* test(teams): annotate the last fixture and parametrize arguments in the raise_max_budget tests
2026-10-09 13:39:08 -07:00
moe-berri
3ebc71be2a
fix(rust): update serde_with for serialization advisory (#45498) 2026-10-09 13:31:41 -07:00
devin-ai-integration[bot]
465f20f274
fix(openrouter): bill responses, decisions and pass-through from usage.cost and always allow reasoning params (#45640)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 13:28:31 -07:00
berriai-litellm-provider-info-sync[bot]
05166e2460
fix(cost-map): add 2026-10-22 deprecation_date to two together_ai rows (#45657)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-09 13:25:58 -07:00
berriai-litellm-provider-info-sync[bot]
aa31c0b962
feat(pricing): add github_copilot/claude-haiku-5.5 (#45656)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-09 13:22:35 -07:00
Mateo Wang
3c56cd31f5
chore(greptile): load nested AGENTS.md files as review context scoped to their directories (#45636)
* chore(greptile): load nested AGENTS.md files as review context scoped to their directories

* refactor(greptile): use a stdlib dataclass so the files.json check runs without uv

* test(greptile): write fixture AGENTS.md files through a helper instead of rebinding a loop variable
2026-10-09 13:22:29 -07:00
berriai-litellm-provider-info-sync[bot]
bbc8d66323
chore(cost-map): add openai gpt-rosalind-discovery from the pricing page (#45653)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-09 12:51:58 -07:00
devin-ai-integration[bot]
07e315ebfb
fix(ui): show team alias in key edit Team dropdown and allow clearing it (#45629)
* fix(ui): show team alias in key edit Team dropdown and allow clearing it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): show alias-less teams by ID once in the key edit Team dropdown

* fix(ui): keep the key organization when its team is cleared

* fix(ui): show a dash for teams without an alias, keeping the ID underneath

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 19:47:03 +00:00
devin-ai-integration[bot]
dafa9cbd4c
ci: run every build_and_test job in every CircleCI pipeline (#43152)
* ci: run the main branch CircleCI job set on rc/* branches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: let litellm_rc_* branches run the main-only CircleCI jobs

The 26 build_and_test jobs filtered to main now share one anchor that also admits branches named litellm_rc_*, so a run-ci pipeline on such a branch at an rc SHA runs the full suite. Every other PR keeps the integration-only subset, and the migration schedule stays on main

* ci: run every build_and_test job in every CircleCI pipeline

Drop the main-only branch filter from the 26 build_and_test jobs, so a run-ci PR pipeline into any base, an rc branch pipeline and main's scheduled pipeline all run the full suite. The migration cron stays on main

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 12:33:14 -07:00
berriai-litellm-provider-info-sync[bot]
de6a0717b7
feat(azure): add azure_ai/Microsoft-Decision-1 (#45639)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-09 12:01:13 -07:00
devin-ai-integration[bot]
fff86d6376
fix(proxy): classify /v1/messages pass-through streams as Anthropic on any host (#45602)
Co-authored-by: Sunny-Soni00 <sunny.s@atriauniversity.edu.in>
2026-10-09 11:57:03 -07:00
devin-ai-integration[bot]
f2f8df2858
feat(ui): show the public JWKS for LiteLLM-signed Anthropic credentials (#45527)
* feat(ui): show the public JWKS for LiteLLM-signed Anthropic credentials

The edit modal of a saved Anthropic credential whose identity source is the
LiteLLM internal issuer now fetches GET /credentials/{name}/jwks and shows the
document with a copy button, so an admin can register it in the Claude Console
without curl. Picking the internal issuer before the credential is saved shows
a hint to reopen it once saved

* fix(ui): wait for a fresh JWKS before offering one cached from an earlier open

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-09 11:53:10 -07:00
yuneng-jiang
cddc7cde97
test: fix stale and flaky tests across CircleCI, GHA and Buildkite (#45530)
* test(integration): poll the partition lock witness off the event loop

The witness poll ran a blocking psycopg connect and query on the test's event loop, so the partition DDL task only progressed during the 20ms sleeps and missed the 3s deadline on loaded runners. Polling through asyncio.to_thread lets the DDL reach the held lock concurrently, and the deadline is 10s since it only bounds how long the DDL takes to start waiting

* test(e2e): retry the reasoning turn until the model emits a reasoning item

Whether gpt-5.4-mini emits a reasoning item next to a forced function call is up to the model, and it skipped it on four scheduled runs while a same-SHA rerun passed. The first turn now retries up to three times before the unchanged reasoning-item assertion

* test(e2e): retry a Bedrock Converse model once when it reports no cache tokens

Opus on Bedrock Converse reported zero cache tokens on half of the scheduled runs while a same-SHA rerun passed. A model that comes back uncached is asked once more inside the 5 minute window, and the non-zero cache token assertion is unchanged

* test(e2e): poll until a deleted stored response stops being retrievable

Azure kept serving a just-deleted streamed response for a moment, so a single retrieve did not raise. Both delete tests now poll retrieve until it returns an error and still require a 4xx

* test(e2e): give the Together prompt cache five primed attempts

Together documents its prefix cache as best-effort, and three primed attempts came back with zero cached tokens once while a same-SHA rerun passed

* test(e2e): allow 30s for the idle RSS reading at session start

A loaded router replica took longer than 10s to answer the session-start memory read, which pytest reruns cannot recover. The RSS budget assertion is unchanged

* test(e2e): scope the MCP submission check to its own card

The register call timed out at 15s on a loaded stack and leaked a submitted server, so the retries matched two "3 passing, 1 failing" cards. Registration gets 60s and the check reads only the submitted server's card

* test(router): wait for the primary success record before counting it

The shadow fan-out test waited only for the two shadow success events and then asserted exactly one primary success, which the logging worker sometimes had not delivered yet. It needed a rerun on 5 of 59 main runs

* test(proxy): give the spend-log cleanup run a 1s budget

A 0.25s budget could expire before the first batch on a loaded xdist worker, leaving rows_deleted at 0. One second still stops the 50-batch, 5-second loop on the deadline, which is what the test proves

* test(autorouter): release the slow token count at 2.05s instead of 2.3s

The count only has to finish after a quarter of the 8s worker budget, and the planner gives it 3s, so releasing at 2.3s left 0.7s of slack that a loaded runner used up. 2.05s is still past the quarter mark with almost 1s of slack

* test(xai): run the reasoning-effort tests on grok-4.6 with a prompt that needs reasoning

grok-4.7 now returns reasoning tokens without any reasoning text, so reasoning_content was never set and both tests failed on every scheduled run. Called directly, grok-4.6 returned reasoning text 6 of 6 times on a step-by-step arithmetic prompt but only some of the time on a bare greeting

* test(integration): give the two worker-kill chaos tests the owned-proxy time budget

Each test spends about 50s on the burst and then up to graceful_stop_seconds() stopping its owned proxy, which overran the 90s default pytest timeout on loaded runners. They now use the same 2 * graceful_stop_seconds() + 120 budget as other owned-proxy tests

* test(integration): compare regional image responses without their created timestamp

Two identical generations that straddle a second boundary differ only in created, which failed the byte-for-byte comparison. Every other field and both spend rows are still compared

* test(integration): count only the S3 logger's flush task

The set of asyncio tasks created while the logger starts also picked up client close finalizers left by earlier tests, so the count ranged from 1 to 8. The test now counts the periodic_flush task it owns and cancels

* test(integration): send guardrail timeout probes eight at a time

All 44 probes went at once to a two-worker proxy, so on a loaded runner some guardrails used their whole 1s timeout before their request left the proxy, and the sink never saw them. A different set of providers failed on each run. Eight in flight still overlaps the waits

* test(integration): count a killed Arize chaos worker as gone once it is a zombie

psutil reports an unreaped zombie as still running, so the 10s death check failed whenever the supervisor was slow to reap the SIGKILLed worker, and the stop then overran the 90s default timeout. The test now accepts a zombie or a missing process and has the owned-proxy time budget

* test(integration): send the Typesafe connection test its real model and endpoint

The test passed the proxy alias as litellm_params.model, and /health/test_connection lays the request over the stored deployment, so the alias replaced the real typesafe/ model and the check failed with "LLM Provider NOT provided" on every run since it landed in #45481. It now sends the provider model, api_base and api_key directly

* test(integration): keep DB-stored models out of the usage-routing Redis read test

The owned proxy inherited store_model_in_db and the job's shared database, so a model another test left behind joined the router and added its cooldown key to the MGET the test compares exactly. Reproduced locally with one /model/new model present (3 failed), and green with model loading from the database turned off

* test(vertex_ai): prove the batch upload streams by laziness instead of peak memory

The two tracemalloc ratio tests flaked on unrelated PRs because peak memory on a shared xdist worker includes other threads' allocations and garbage from earlier tests. They are replaced by deterministic checks of the same property: the upload stream parses and maps a row only when it is pulled, so a body whose tail is not JSON yields its valid rows first, and a Path source reflects a row rewritten on disk after the upload started. Making the parse, the output, or the file read eager fails these tests

* test(integration): keep the alias test_connection call as a known bug

* test: type the counting mapper and the Bedrock rerun helper

* test: annotate the new test locals as Final and build them as tuples

* test(vertex_ai): prove pull-driven transforms with the garbage tail alone
2026-10-09 11:40:14 -07:00
nate-berri
0519882fef
chore: ignore .mypy_cache (#45628)
Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
2026-10-09 18:26:19 +00:00
tin-berri
1dbed8e1e2
fix(liteadmin): detach keys from teams and default to Sonnet 5.5 (#45625) 2026-10-09 11:17:53 -07:00
berriai-litellm-provider-info-sync[bot]
fe955f5e42
fix(bedrock): add the bare Pegasus 1.5 row and take GPT-5.x context windows from the model cards (#45621)
* fix(bedrock): add the bare Pegasus 1.5 row and take GPT-5.x context windows from the model cards

Price-Sync: litellm-providers

* fix(bedrock): whitelist the bare Pegasus 1.5 id for the converse routing check

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: Kerry Lu <klu@berri.ai>
2026-10-09 11:06:54 -07:00
yuneng-jiang
6afdf482de
test: move offline anthropic prompt caching tests to tests/unit (#45616)
The two offline tests in tests/local_testing/test_anthropic_prompt_caching.py failed on main because
#24071 started passing logging_obj to client.post, so assert_called_with on the patched
AsyncHTTPHandler.post no longer matched. The request bodies themselves were still correct

Both now live in tests/unit/llms/anthropic/chat and assert the body and headers that reach the
wire through respx instead of patching our own HTTP handler. The coverage allowlist entry keeps the
five tests that still need provider credentials
2026-10-09 10:53:08 -07:00
devin-ai-integration[bot]
61e5f2dd3c
feat(decisions): add hosted_vllm provider (#45501)
* feat(decisions): add hosted_vllm provider

Adds HostedVLLMDecisionsConfig so hosted_vllm/<served model> works on
POST /v1/systemone and POST /v1/decisions against a vLLM server that
serves the System One body (vllm-project/vllm#59299, after v0.31.0).
vLLM only supports choice questions and rejects others, so the base
decisions config gains a per-provider health_check_questions attribute
and the evaluation-mode health probe sends a choice question for
hosted_vllm while every other provider keeps the noul probe.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move the hosted_vllm placeholder key to constants

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 10:41:51 -07:00
Krrish Dholakia
18e87d3039
fix(ui): title usage overview chart "Daily usage" and label the model table "Top models" (#45604)
* fix(ui): title usage overview chart "Daily usage" and label the model table "Top models"

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): clarify usage chart subtitle as top 8 models plus Other

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-09 09:56:59 -07:00
yuneng-jiang
5e1c5c0bd1
test: move offline proxy auth, hook, spend and pass-through tests to tests/unit (#45568)
* test: move offline proxy auth, guardrail hook, spend, budget reset, GCS payload and pass-through tests into tests/unit

Relocate 53 legacy tests that pass with no network or keys into the tests/unit
files that mirror the code they exercise, drop 2 llm_guard tests already covered
by test_llm_guard_call_type_aliases, and delete the emptied legacy files.

* test: freeze the reset clock and run the real reset methods in the moved budget reset tests

Pin datetime in reset_budget_job and timezone_utils to a fixed instant and await
the service-hook tasks the job spawns instead of sleeping. Drive the per-user
failure with a malformed budget_duration row from the prisma mock so the real
_reset_budget_for_key/user/team run instead of patched replacements.

* test: use fixed timestamps in the moved pass-through and spend log tests

* test: mock the moderation HTTP call and type the moved moderation and spend log tests

Serve the OpenAI moderation response through an httpx MockTransport so the real
async_make_request runs, annotate the moved tests' locals with Final, and pass
get_logging_payload typed kwargs instead of an untyped input_args dict.

* test: drive the moved proxy tests through boundaries and annotate their locals with Final

Serve GCS downloads and the Router moderation call through httpx mocks with fake
Google credentials, let the real pass-through success handler fail on the mocked
upstream response, and fold the Vertex live route endpoint check into the existing
route test. Annotate moved locals with Final, build the budget reset rows without
rebinding, and type the remaining fakes without **kwargs
2026-10-09 09:49:13 -07:00