* feat(ui): build System One playground requests with a form
* test(ui): fill System One form fields with single change events
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix: support OpenAI SDK 3.x while retaining 2.x compatibility
* fix: preserve HTTPX compatibility across OpenAI SDK versions
* refactor: trim OpenAI SDK compatibility patch
* refactor: simplify SDK client defaults
* test: cover SDK compatibility across provider routes
* fix: use shared names for SDK transport helpers
* fix: address SDK compatibility review feedback
* test(openai): send requests through the SDK API factory clients
* fix(exceptions): keep the provider response on PaymentRequiredError under httpx 2
* ci(base-sdk): count httpx2 as base-only when a base dependency brings it in
OpenAI SDK 3 depends on httpx2, so the litellm-core install at highest resolution pulls it in through openai. The base-only check now resolves the declared base dependency closure of litellm and litellm-core and accepts extras-only modules owned by a distribution inside that closure
* test: cover OpenAI SDK 3 client paths end to end and read SDK HTTP headers case-insensitively
* test: read C02 spend rows once, wait out the worker healthcheck on respawn, and resolve module owners on Python 3.10
---------
Co-authored-by: Marcus Wood <marcuswood@openai.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(decisions): add the OpenAI Decisions spec types and the System One translation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): share one DecisionsModel config and require model and usage on responses
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): rename the shared pydantic parent to DecisionsObjectBase
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): dispatch /v1/decisions through provider configs and the shared HTTP handler
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): map OpenRouter connection failures to APIConnectionError and drop explanatory docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): switch /v1/decisions to the OpenAI Decisions schema
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(decisions): build the usage the OpenAI response schema now requires
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lens): ask signal questions through the OpenAI Decisions schema
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): refuse safety_identifier for System One providers unless drop_params
* feat(decisions): refuse safety_identifier on providers that cannot take it unless drop_params drops it
* test(decisions): cover hosted_vllm in the safety_identifier and per-provider wire tests
* fix(decisions): check a System One request's safety_identifier against the provider too
* fix(decisions): drop a non-string safety_identifier under drop_params
A malformed safety_identifier now follows the drop_params convention in both body shapes: it answers 400 without drop_params and is dropped before the provider call with it. A string identifier on OpenAI stays on the wire either way.
* fix(decisions): let /v1/decisions drop a non-string safety_identifier under drop_params
The route checked the whole body before routing, so a non-string
safety_identifier answered 400 even when the deployment or the body set
drop_params. The route now leaves that field to the Decisions call, which
knows every drop_params source
* refactor(decisions): return the unsupported safety_identifier as a value and raise it in _prepare_call
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(gemini): map OpenAI voice names to Gemini prebuilt TTS voices
* test(gemini): type the voice mapping test helpers
* fix(gemini): leave nova out of the voice mapping since Gemini accepts it as given
* test(gemini): add integration cells for the OpenAI voice mapping on Gemini and Vertex TTS
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(databricks): keep #/$defs refs in json_schema response_format for non-Claude models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): type the json_schema response_format override
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(databricks): integration cells for json_schema $defs refs across endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(databricks): split stacked loop in streaming json_schema cell
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(databricks): mark BaseConfig dict contract on json_schema override
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(databricks): cover json_schema refs on every endpoint, edge refs, sad paths and an outage burst
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* ci: delete legacy required-checks shim workflow
* ci: fold semgrep into lint and helm into unit, drop the duplicate UI build
semgrep runs the repo's own rules over the source, so it becomes lint / semgrep
behind lint passed. The helm chart tests become unit / helm behind unit passed.
UI Build Check ran the same docker build --target ui-builder as
smoke / dashboard-build, so it is deleted
* ci: merge the alphabetical llms shards into one llms-providers shard
The two shards were split by letter range, which says nothing about what they
run. Every provider without its own shard now runs in llms-providers, with the
same 137 paths as before
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* feat(ui): announce decision model support in Models + Endpoints
* feat(ui): make decision models findable, callable, and testable from the Admin UI
Add Model provider search now matches decision model names and shows a note on where to call a decision model. The banner links to the docs and opens Add Model. The System One playground defaults to /v1/systemone with a picker of the proxy's decision models, and the model hub shows a /v1/systemone usage example for evaluation models.
* fix(ui): keep the admin signed in when a typed virtual key cannot list decision models
The System One playground looked up a typed virtual key's decision models through
the dashboard's shared request client, so an expired key signed the admin out, and
typing a key by hand sent one lookup per keystroke. The lookup now uses the same
client as Send, against the same proxy, and waits for typing to stop
* fix(ui): hide decision model Playground links from view-only sessions
A proxy admin viewer cannot open the Playground, so the Models + Endpoints
decision models banner and the AI Hub's Try it in the Playground link sent
them to an Access Denied page. Hide both for view-only sessions.
* fix(ui): send the AI Hub decision example's key in the proxy's configured header
A proxy with litellm_key_header_name set reads keys only from that header,
so the example's hardcoded Authorization header was rejected. Use the header
name the dashboard already reads from the session.
* fix(ui): stop the decision model notice from ruling out /chat/completions for dual-mode models
* fix(ui): point only users who can add models at Add Model in the decision models banner
---------
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(ui): resolve the provider dropdown key to the slug the backend declares
The Add Model form submits the `provider` field from /public/providers/fields
verbatim, but provider_map is keyed by hand on the frontend, so the two sides
disagree for 14 of the 124 providers the backend serves. Some differ only in
case ("MINIMAX" vs "MiniMax", "CURSOR" vs "Cursor"), others have no key at all
("MILVUS", "LANGFUSE"). getProviderModels looked the key up, got undefined,
matched no models, and the model field silently degraded from a dropdown into
a free-text input with no candidates
Both call sites now go through resolveLitellmProviderSlug, which prefers an
exact provider_map key and otherwise lowercases the value. That resolves every
provider in the catalog to the litellm_provider slug the backend declares for
it. Matching provider_map case-insensitively instead would look tempting and be
wrong: "SAGEMAKER" and "SageMaker" are two distinct providers with two distinct
slugs, and a case-insensitive lookup collapses both onto sagemaker_chat
The regression test drives the real provider_create_fields.json rather than a
fixture, so a provider added to the catalog with a mismatched key fails here
instead of reaching users as an empty dropdown
Co-authored-by: AaronHowell <237895480@qq.com>
* chore(ui): drop explanatory comments from provider slug fix
---------
Co-authored-by: AaronHowell <237895480@qq.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* fix(bedrock): keep GPT tool-result images beside the tool result
Bedrock gpt-6.1-sol rejects an image nested in toolResult.content. Put it on the same user turn instead.
* fix(bedrock): read tool-result image support from the price map
OpenAI GPT rows on Bedrock Converse reject an image inside toolResult.content, so those images sit beside the tool result.
* fix: sync price-map test schema and import order
CI validates model_prices_and_context_window.json against an inline schema in test_utils, and ruff checks import order in converse_transformation.
* fix(bedrock): satisfy basedpyright in tool-result image helper
Drop Final assignments inside loops and use TypedDict key checks so the lint type-check gate stays within its merge-base ceiling.
* test(bedrock): cover tool-result image placement helper
Exercise Claude no-op paths, text-only GPT tool results, status preservation, and split text/image tool content for codecov.
* refactor(bedrock): reuse the shared tool-result image placeholder text
---------
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
* test(e2e): cover tool_reference tool results on bedrock converse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): keep tool_reference results as text on converse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(bedrock): format tool reference parser
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover bedrock converse tool_result of only tool_reference blocks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): tag converse tool_reference case with e2e metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): avoid unchecked cast in converse tool_reference mapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic): keep tool_reference results as text on responses paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(anthropic): tighten tool_reference mapping on responses paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(completion-extras): expect tool_reference names in responses input
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(bedrock): drop responses converter changes, keep converse tool_reference fix only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
* fix(proxy): keep the mapped status on assistants and threads route errors
* test(proxy): type and annotate the assistants and threads route error tests
* fix(proxy): redact internal details from assistants and threads route error messages
* test(proxy): integration cells for assistants route error mapping
* test(proxy): drop malformed JSON cases that never reach the assistants route handler
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(azure_ai): support Microsoft-Decision-1 on the decisions API
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(azure_ai): strip project path and API suffix together in decisions base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(agents): ban semicolons and colons in human-facing text
Tighten the writing rule so ";" and ":" are only allowed where the sentence would be nonsensical without them, rewrite the existing semicolons in AGENTS.md to follow it, and note that docs (litellm-docs) keep their trailing "."
* docs(agents): rewrite semicolons in nested AGENTS.md files
* docs(agents): hand-maintain the dashboard Next.js warning and sync e2e CONTRIBUTING lines
* docs(agents): point at the PR template path that exists
* docs(agents): restore the generated Next.js block in the dashboard AGENTS.md
* feat(rust): add the DeepSeek Anthropic Messages config
* refactor(rust): drop the tool discriminator without mutation and use rstest in DeepSeek tests
* feat(rust): add the Vertex AI Anthropic Messages config
* fix(rust): normalize Vertex system turns the way Python does
Vertex uses the shared normalize_system_role_messages: a leading run is
hoisted, a later turn stays or is converted in place by the model's
supports_mid_conversation_system, and billing blocks are filtered after the
hoist so one inside a system-role message no longer reaches Vertex.
* fix(rust_bridge): keep non-Claude Vertex Messages calls on Python
The catalog admits every vertex_ai Messages call, but Rust serves only Claude
there and reports anything else as a terminal InvalidProvider instead of a
decline. The dispatcher now bypasses Rust for those models, as
get_provider_anthropic_messages_config does.
* refactor(rust): template the Vertex location and project errors
The location failure is the shared InvalidType detail and the missing project
is MissingSetting with the provider, setting, param and variable name, so no
sentence is spelled at the return site.
* test(rust): use rstest in the Vertex AI Messages tests
* refactor(rust): derive the missing Vertex project error from its spec
* refactor(rust): give Messages configs typed litellm params
* feat(rust): port _normalize_system_role_messages for Messages hosts
mid_conversation_system mirrors Python's module function for function: a
leading run of system turns is hoisted into the top-level system, a later
turn stays in place when the model map flags supports_mid_conversation_system
and otherwise becomes a user turn where it was, never between a tool_use and
its tool_result, and billing blocks are stripped from the hoisted system. The
capability joins MessagesModelCapabilities and the bridge projection. Azure AI
uses it in place of fold_system_role_messages, which hoisted every turn.
* fix(rust): keep the Bedrock runtime endpoint behind a blank api_base
* refactor(rust): import LitellmParams from its crate instead of a re-export
* refactor(rust): template the Messages request decode errors
Both decode failures go through ErrorDetail::invalid with the subject and the
serde error as source instead of a preformatted sentence.
* refactor(rust_bridge): read the module globals the param specs name
The Messages host no longer keeps its own list of litellm globals; a spec
whose spellings the call leaves out names the global to read, so the next
family that has one needs no host change.
* refactor(rust): give each route lookup Python's implicit None
ProviderConfigManager's per-route methods list only the providers the route
serves and fall through to None; the Rust lookups now do the same, so a new
LlmProviders variant touches only the route that serves it.
litellm-router-types mirrors litellm/types/router.py: LitellmParams holds the
model, the credentials, one flattened group per credential family from
auth-types (aws, vertex), the deployment settings a config spells, and every
other key in extra, as Python's extra="allow" keeps it. Config parses it
directly, so config::LiteLlmParams and its untyped additional_fields are gone
and a mistyped provider param is a parse error. fields() and specs() derive
from the families' param specs for hosts that project kwargs and fold module
globals. Spelled<T> is the one untagged shape for a typed value or the text a
config spells, organization takes the list Python's router expands, and the
callback shorthand keeps a value-level one-or-many. No credentials wrapper
group: serde's flatten only consumes keys for struct-shaped children.
VertexParams in auth-types holds the vertex_* fields of CredentialLiteLLMParams
in both spellings, with one ParamSpec per setting: the wire names, the litellm
module global VertexBase.get_vertex_ai_* consults, and the environment names.
A credential deserializes from text or the JSON object a config spells, an
empty object counts as absent, and both credential fields are redacted in
Debug. auth-gcp is split so lib.rs is the entrypoint (config, auth,
constants), its lookups resolve through the specs, secret_names() derives
from them, and get_vertex_ai_project_from_credentials reads a service
account's project_id for hosts that need it before a token exchange. The OCR
host folds the module globals in by spec at projection, so OcrSettings no
longer carries vertex_project and vertex_location.
AwsParams in auth-types holds the aws_* fields of CredentialLiteLLMParams
with one ParamSpec per field: its wire name and the environment names it
falls back to. fields() and secret_names() derive from the specs, the AWS
helpers take the struct instead of a map, aws_auth_config and the region
resolution read through the specs, and a missing region is
AwsParams::REGION.missing("AWS"), whose message is rendered from the spec.
Bedrock Converse, Bedrock transcription, Textract OCR and the Bedrock
Messages config build the struct at their one untyped boundary.
* feat(decisions): add Databricks ai_decide as a /v1/decisions provider and auto-router decider
* fix(decisions): reject Databricks endpoint names carrying URL separators and keep the provider wiring under llms/
* fix(decisions): reject dot-segment Databricks endpoint names and add Databricks to the dashboard's OSS classifier editor
* fix(ui): name every OSS classifier provider in the classifier radio copy
* feat(decisions): register the Databricks OpenJev decider as an evaluation-mode cost-map entry
* fix(cost-map): declare zero cache rates on the databricks-openjev-qwen35-4b entry
* fix(proxy): reconcile every DB deployment when a registry miss lands on an auto router
A /model/new for an auto router lands on one replica. A request routed to
that auto router on a sibling replica missed it, the model read-through
loaded only the auto router's own row, and the pre-routing hook then picked
a tier the sibling had not loaded, so the caller got a 400 "no healthy
deployments for <tier>" until the next periodic reload. An auto router miss
now runs the full add_deployment reconcile, tiers included. Found by the
audit's two-worker rig on this PR's routing cells; CircleCI's one-worker
proxy never opens the window.
* test(integration): audit cells for Databricks decisions and the Databricks auto-router classifier
Wire cells for /v1/decisions with databricks/<endpoint> deployments, routing
cells for the complexity router's databricks classifier on all three
endpoints with the Postgres rows read back, management cells for the
classifier's spend writes and for the evaluation-mode health check of a
Databricks serving endpoint.
* refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 2 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 3 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 4 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 5 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(policy-engine): stream detect-only post_call pipeline steps live when buffering is off
* fix(policy-engine): scan live pipeline streams a provider error cuts short and discard rewrites per step
A live detect-only pipeline now scans the chunks the client received when the provider fails mid-stream, then re-raises the error. Each step's rewrite is discarded right after the step, so later steps scan the text the client actually received.
* fix(proxy): keep a guardrail verdict recorded after a stream failure on the failure spend row
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(proxy): sign RDS IAM tokens for the database's own region
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(proxy): add per-connection RDS IAM signing region overrides for writer and reader
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): integration cells for per-connection RDS IAM signing regions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): opt-in real RDS e2e cells for cross-region IAM signing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): address review on RDS IAM region cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): drop recording-front integration tests, mark RDS e2e suite e2e and register its model via /model/new
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): assert the spend read lands on the replica, not just a live connection
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type generate_iam_auth_token params and retry the replica spend-read check
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): derive RDS region from custom RDS Proxy endpoint hostnames
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): add Subject, step labels, frozen rows and Final to the RDS IAM e2e suite
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): drop the opt-in RDS IAM e2e suite in favor of unit coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
* refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 2 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 3 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 4 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 2 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 3 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(github_copilot): sync model catalog with Copilot's served models
* test(github_copilot): assert catalog invariants instead of pinning vendor facts
* fix(github_copilot): keep thinking capabilities on served Claude catalog rows
* fix(github_copilot): keep supports_reasoning on catalog rows that shadow the family fallback
* fix(model-catalog): preserve merged Copilot Haiku metadata
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(microsoft_365_copilot): add Microsoft 365 Copilot chat provider with OAuth token exchange
* fix(ui): create a credential for credential-only auth types in Add Model
* fix(microsoft_365_copilot): accept max_tokens and list the chat model
* fix(microsoft_365_copilot): register provider model set
* fix(ui): hide other auth types' fields when credential-only types are filtered
* fix(health): do not inherit stored auth settings when a connection test brings its own api_key
* docs(health): restore connection-test inheritance docstrings
* fix(lint): suppress justified provider boundary casts
* fix(lint): satisfy M365 type-discipline gate
* fix(types): eliminate type-check gate regressions
* style: shorten suppression reasons so ruff format leaves them on one line
* refactor(types): drop shared-file type widening unrelated to the Copilot provider
* fix(health): pass caller headers as secret_fields to connection-test probes
* fix(types): type connection-test probe and token counter locally
* fix(lint): drop cast import from connection-test header getter
* chore: remove explanatory comments flagged by review
* fix(proxy): gate OAuth client credentials behind proxy admin
* fix(router): skip cooldown on caller-scoped OAuth auth failures
* fix(router): skip fallback cooldown on caller-scoped OAuth auth failures
* fix(router): type caller-scoped cooldown lookups
* fix(ci): include Microsoft 365 Copilot tests in unit shard
* fix(proxy): keep saved endpoint when test_connection retests a deployment by id
* fix(m365): collapse doubled Graph replies
* fix(m365): preserve trailing system messages
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 1 of 5 (anthropic to google block)
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(cost-map): remove openrouter rows, part 2 of 5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(proxy): keep end-of-stream output assembly and Redis pipeline logging off the event loop
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(caching): serialize Redis pipeline values with orjson
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(proxy): resolve model-level guardrails once per streamed request
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): type the streamed guardrail cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): fall back to json.dumps when orjson is not installed
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): cover the json.dumps fallback for values orjson rejects
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep new guardrail cache and orjson helper within the type gates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(caching): guard Redis print_verbose value interpolation behind is_debugging_on
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(caching): drop the orjson fast path from Redis pipeline writes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(caching): drop the debug-on print_verbose test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(proxy): drop the per-request guardrail lookup cache from this PR
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(logging): let print_verbose take %s args so Redis cache values format only when verbose
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): type set_cache key/value and redis_version so print_verbose args are known to pyright
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Moves failure analytics off the cost overview into a dedicated admin-only
Errors tab: error rate with 4xx/5xx/429/401-403 callouts, error rate over
time, failed requests per day stacked by HTTP status code with a legend,
failures by status code, and a Failures by identity table switchable between
virtual keys, teams, users and models with error rate and top status per
row, linking to Logs. Reads GET /gateway/errors/activity. The Overview keeps
its plain ok/failed counts.
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Counts failed gateway requests by HTTP status at the ASGI edge
(LiteLLM_DailyGatewayFailedRequests) and rolls failures up per virtual key,
team, user and model group with the status the logging callbacks recorded
(LiteLLM_DailyRequestErrors, fed from the spend writer, flushed like the
gateway counters, directly or through Redis with a pod lease, and on
shutdown). GET /gateway/daily/activity gains by_status_code; the new
admin-only GET /gateway/errors/activity returns per-day totals with
per-status counts, failures by status code, and keys, teams, users and
model groups ranked by failures with their top status. Dashboard schema.d.ts
regenerated.
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): list the account's invocable models behind bedrock/* when check_provider_endpoint is on
* fix(bedrock): sign the listing query RFC 3986 style and raise BedrockError on a failed listing
* test(utils): keep test_utils.py at main's formatting, only the Bedrock discovery test is new
* fix(bedrock): chain BedrockModelInfo's constructor so the passthrough config keeps its event parser
* fix(bedrock): list discovered ids without the provider prefix so partial wildcards filter
get_valid_models returns bare vendor ids for every other provider and for the static bedrock catalog, and the proxy's wildcard expansion adds the provider prefix itself. The lister prefixed its ids, so a partial wildcard like bedrock/anthropic.* never matched the proxy's filter and listed uncallable bedrock/anthropic.bedrock/<id> entries
* test(bedrock): integration cells for wildcard discovery through the proxy and the SDK
* fix(proxy): keep a partial wildcard a filter when its deployment repeats the prefix
The wildcard expansion guessed filter-or-alias by whether any provider id carried the prefix. A deployment such as model_name bedrock/anthropic.* over model bedrock/anthropic.* can only ever route names that share the prefix, so when the account has no on-demand anthropic.* id (every Anthropic model behind an inference profile) the guess fell into the alias branch and listed bedrock/anthropic.<every id>, none of them callable. A deployment whose model repeats the suffix now always filters, and lists nothing when nothing matches
* fix(proxy): prefix every discovered id under a custom wildcard prefix that starts a vendor id
* test(bedrock): Final and read-only annotations in the wildcard discovery tests
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(bedrock): route Grok 4.7 tools with reasoning to native Chat Completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): mint unique tool call ids on native Chat Completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): send a blank text block for Converse tool results with no supported content
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(bedrock): default Grok on Bedrock to native Chat Completions and leave Converse unchanged
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): gate tool call id minting to Grok on native Chat Completions
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(bedrock): check the Grok tool call id gate at both call sites
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): remint positional tool call ids on native Chat Completions for every model
AWS returned call_0, call_1 for GPT 5.6 as well as Grok on 2026-10-09 (and unique ids for GPT 5.6 earlier the same day), so the id format is not fixed per model. Drop the Grok-only gate and remint any call_<digits> id on the route
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): drop stop sequences for Grok instead of forwarding them to a 400
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(teams): let proxy admins grant team admins the right to raise their team budget
Adds a raise_max_budget entry to team_admin_editable_team_fields. With max_budget alone a team admin can still only keep or lower the team budget. Adding raise_max_budget lets them raise it, capped by the organization's max_budget when the team belongs to one, while removing the budget stays with proxy admins. The entry is rejected without max_budget, /team/info reports may_raise_max_budget to the caller, and the Admin UI nests the new checkbox under Max Budget with a tooltip and shows the team admin a hint under the budget field.
* fix(ui): shorten the raise_max_budget tooltip and link it to the docs
* fix(teams): pin the organization in the guarded team budget write
A granted raise is checked against the team's organization at read time, so the write now also requires organization_id to be unchanged. A concurrent move into a budgeted org returns 409 instead of landing an uncapped raise. Also types the new test helpers.
* test(teams): type the team row store methods and the touched race test arguments
* test(teams): annotate the last fixture and parametrize arguments in the raise_max_budget tests