Commit graph

53604 commits

Author SHA1 Message Date
moe-berri
8520626e7a
fix(lens): preserve approved worker digests and harden its image (#44467)
* fix(lens): pin worker dependencies and support approved image digests

* fix(lens): include locked dependencies and release identity in build context

* fix(lens): select the dev worker package for SHA-tagged charts
2026-10-03 17:54:17 -07:00
berriai-litellm-provider-info-sync[bot]
a66adb4ff8
chore(cost-map): sync openrouter prices from the models API (#44466)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 17:34:05 -07:00
devin-ai-integration[bot]
9d16412341
fix(caching): count tool_call cache_control marks in the injection census (#43556)
* fix(caching): count tool_call cache_control marks in the injection census

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): remove cache census casts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): only skip injection on message or content marks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): skip injection on messages whose tool calls carry marks

Reverts 9f08d8aef8. A default 5m mark injected on assistant text lands
before the client's 1h tool_use mark, which Anthropic rejects with a 400
because a 1h breakpoint must not follow a 5m one. Keeping the full census
in the skip check leaves the client's tool_call breakpoint as the only
one on that message.

* fix(caching): count every client tool_call cache mark in the breakpoint census

The census gated tool_call marks on type function and dict shape, so a client mark on a call without a type or with a string cache_control slipped past the count and injection overflowed the 4 breakpoint cap. Count any non-None tool_call mark, keep the server tool exclusion, and add integration cells for the capped surfaces, the yaml stand-down, Bedrock and Gemini, and router affinity

* test(integration): hold the upstream so the worker kill lands mid-burst

* refactor(caching): reuse the transform's server tool lookup in the breakpoint census

The census now calls the same helper the Anthropic transform uses to decide
whether a marked tool call becomes a server tool block, so the two cannot
drift apart. The owned-proxy burst test waits up to 90 seconds for the burst
to reach the wire before it kills a worker

* refactor(anthropic): move the server tool rebuild check under llms/anthropic

* test(integration): audit the tool call mark census across chat, messages, responses, and chaos

Adds the /audit cells for the breakpoint census on assistant tool_calls marks: the
Responses stream bridge, the OpenAI and Anthropic SDK clients, in-process Pydantic
messages, response cache twins, request-level points, a provider 401 on a capped request,
malformed provider_specific_fields and tool_call ids, null or empty points, a points
update mid-burst, a proxy restart mid-burst, and a worker SIGKILL that picks the worker
holding the burst's upstream connections

* test(integration): close the SDK clients the cache census cells open

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:33:48 -07:00
tin-berri
6370104c53
feat(enterprise): bundle LiteAdmin Slack with native gateway login (#44444)
* feat(enterprise): bundle LiteAdmin Slack worker with native gateway login

* fix(enterprise): preserve gateway prefixes during native Slack linking

* fix(enterprise): retain native Slack linking on the admin backend

* fix(enterprise): reuse shared native Slack connection services
2026-10-03 17:28:18 -07:00
moe-berri
4b67a2b845
feat(roi): default people and branch lists to matched accounts (#44465)
* feat(roi): show matched people by default in contributor lists

* fix(roi): keep matched filter tabs readable on narrow screens

* fix(roi): retain spend-only users and support older browsers
2026-10-03 17:24:18 -07:00
devin-ai-integration[bot]
f850b2c324
test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment (#44451)
* test(integration): exact four-part translation cases on a shared fake provider and shared YAML deployment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): compare every non-transport provider header and check for late provider requests at session end

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): name TranslationTestCase fields after litellm and provider sides and drop regressions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prefix checked TranslationTestCase fields with expected_ and name the fake reply mock_provider_response

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(integration): name TranslationTestCase fields in the translation README

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): add a claude-opus-5-5 base case and deployment next to claude-sonnet-4-6

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): name translation cases <MODEL>_TEST_CASE and document the naming rule

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(integration): move translation test rules into tests/integration/translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 00:09:58 +00:00
devin-ai-integration[bot]
0ed1c08f02
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources

Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.

Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.

The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.

Fixes #28607
Resolves LIT-6107

Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>

* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries

The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.

LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.

* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle

* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries

* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly

* fix(proxy): decrypt stored litellm_params before the WIF write gate

* fix(proxy): hide WIF secret references from /health output

* fix(proxy): keep the proxy error shape on credential endpoint refusals

* fix(proxy): hide identity token file paths from /health output

* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working

The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader

* fix(auth): share one exchanged token across workers reading the same assertion

Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it

* fix: keep anthropic federation from being shadowed or leaked

An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.

The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.

The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.

* fix: unlink a staged token file a failed write leaves behind

The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.

* refactor: move anthropic jwks derivation behind a provider-owned tagged union

* fix: unlink the staged token file when its write fails at close

A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.

* fix(anthropic): close the staging descriptor before writing the shared token file

* fix(wif): judge federation writes by what they set, not what is stored

The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500

The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL

* fix(proxy): let a deployment write name a federated credential

reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.

is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.

* refactor(proxy): derive health display policy from the federation key sets

The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places

WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one

* fix(anthropic_wif): treat blank identity-source fields as unset

* test(proxy): classify the federation params in the credential slot registry

main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret

* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics

Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.

* fix(credentials): gate PATCH on WIF fields resolved from model_id

The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path

* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header

Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
2026-10-03 17:08:30 -07:00
devin-ai-integration[bot]
dfdd496db8
fix(tests): match the OS bind error in the owned-proxy port-race retry (#44462)
* fix(tests): match the OS bind error in the owned-proxy port-race retry

* test(integration): keep the port-race predicate pure so its unit tests stay in-process

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:07:10 -07:00
devin-ai-integration[bot]
f0eda6d2a6
fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API (#44419)
* fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API

Bedrock Mantle serves Claude ids only on /anthropic/v1/messages, but health
checks probed every chat-mode deployment over /v1/chat/completions, so a
bedrock_mantle Claude deployment showed unhealthy while real /v1/messages
traffic to it succeeded

Add an anthropic_messages health check mode and make it the default for
bedrock_mantle Claude models. An explicit model_info.mode still wins, and
/health/test_connection and the Add Model form accept the new mode

* fix(health): resolve the test connection mode from the deployment when the request omits it

The Admin UI model page sent the mode /model/info had filled in from the cost
map back as the probe mode, so Test Connection on a Bedrock Mantle Claude
deployment still went over chat completions. The page now forwards only the
row's id, and /health/test_connection resolves a missing mode the way /health
does: the stored model_info.mode, then the mode the provider requires, then the
cost map.

* fix(health): resolve an omitted ahealth_check mode the way the proxy does

* fix(health): test connection honors a stored mode only for the stored model and rejects a non-string mode

A request that selects a stored deployment and sends a different litellm_params.model now resolves the probe mode from that model instead of the stored model_info.mode. A litellm_params.mode that is not a string answers 400 instead of 500. The Bedrock Mantle rule that Claude models are probed over the Messages API moves into the provider package.

* fix(health): shape test connection probe params for the model the request probes

A request that selects a stored deployment by id and overrides the model
resolved its probe mode from the overridden model but still injected
max_tokens from the stored mode, so an embedding override of an
anthropic_messages deployment failed with a Mistral 422 extra_forbidden

* fix(health): report an early ahealth_check failure as itself, not as a missing mode

With the mode resolved automatically when the caller omits it, a failure
before that resolution (no model, a non-string model, a provider that does
not resolve) was wrapped as "Missing mode", a hint that pointed at the wrong
fix and dropped raw_request_typed_dict from the result. Every failure now
returns the same shape.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:05:47 -07:00
berriai-litellm-provider-info-sync[bot]
4d30f8c59b
chore(cost-map): add azure_ai/kimi-k2-thinking retirement date from the Azure retired models page (#44455)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 16:44:13 -07:00
devin-ai-integration[bot]
21ecd0af55
fix(traces): reject conflicting spend aliases and unrelated HTTP siblings (#44456)
Co-authored-by: Yujong Lee <yujong@berri.ai>
2026-10-03 23:43:10 +00:00
devin-ai-integration[bot]
6d73fa6b49
fix(vertex_ai): forward system and tools to partner model count_tokens (#43900)
* fix(vertex_ai): forward system and tools to partner model count_tokens

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vertex_ai): avoid mutable token request construction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vertex_ai): return partner count_tokens provider errors as values so the proxy falls back locally

* test(integration): cover Vertex AI partner count_tokens forwarding and local fallbacks

Add wire-level cells for /v1/messages/count_tokens, /utils/token_counter,
/v1/responses/input_tokens and the Gemini countTokens route on a Vertex AI
Claude deployment: the system prompt and tools reach the partner
count-tokens endpoint verbatim, null fields stay out of the body, malformed
tools are rejected before any peer call, peer, token-endpoint and connection
failures fall back to the local tokenizer unless disable_token_counter is
set, generation on the same deployment keeps working, and concurrent bursts
survive a peer outage, a slow peer and a worker SIGKILL. The sdk cells cover
litellm.acount_tokens the same way.

The _support/process.py and _support/client.py harness files are brought to
main's content so the self-booting cells read INTEGRATION_PROXY_READY_SECONDS
instead of a fixed 70 s boot budget.

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 16:36:14 -07:00
devin-ai-integration[bot]
1ef0fe9790
fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check (#43741)
* fix(auth): resolve hidden model_group_alias entries in the zero-cost budget check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(auth): key the zero-cost cache by the resolved model group

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(auth): keep the zero-cost verdict per requested name and include hidden aliases

* test(integration): audit the zero-cost bypass through hidden model_group_alias names

* test(router): cover the extracted routing strategy switch

* test(integration): record a pre-flip burst before the alias flip

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 16:35:11 -07:00
moe-berri
cb17588276
feat(lens): coordinate worker releases and bundled installs (#44428)
* feat(lens): coordinate worker versions and bundled installs

* test(lens): exercise bundled Compose startup and restart in CI

* fix(lens): refund failed model requests without a response

* test(lens): verify trace persistence in the bundled stack

* fix(lens): align Helm images and isolate Compose storage

* fix(lens): reject worker builds without release identity

* fix(lens): encode Compose credentials and normalize worker versions

* fix(lens): refuse worker recommendations for unidentified builds
2026-10-03 16:33:45 -07:00
yuneng-jiang
427158eb5b
feat(ui): show invitation and reset password links in a copyable field (#44454)
* feat(ui): show invitation and reset password links in a copyable field

Put the link in a read-only input with a Copy button beside it, stack the
User ID and link labels above their values, and focus Copy on open so the
field shows the start of the URL. Copy now goes through the shared
copyToClipboard helper, which falls back to a selection copy where the
Clipboard API is unavailable.

* fix(ui): keep focus on the copy control after a fallback clipboard copy

The execCommand fallback focused a temporary textarea and removed it, so
focus fell to the page body and a second Enter on Copy did nothing.
Restore focus to the element that had it. Move the rendered dialog
tests to the integration tier.
2026-10-03 16:23:13 -07:00
berriai-litellm-provider-info-sync[bot]
5f969982fc
fix(azure): set text-embedding max input to 8192 from the models sold directly page (#44449)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 16:15:02 -07:00
moe-berri
1a7023366f
feat(roi): measure shipping velocity, quality, and recorded spend (#44426)
* feat(ui): prototype observed engineering ROI dashboard

* feat(roi): replace effort estimates with measured repository metrics

* fix(roi): finish connection recovery and generated API contracts

* fix(roi): show merged changes before accounts are linked

* fix(roi): preserve selected report tab across refreshes

* fix(roi): recover app authorization and keep detail values readable

* fix(roi): reuse the shared OAuth HTTP client

* feat(roi): combine providers and compare equal reporting periods

* docs: explain ROI metrics for first-time readers

* fix(roi): preserve connections and scheduled reports during setup

* ci(roi): assign database contracts to the active Postgres shard

* fix(roi): preserve issue counts and normalized connections

* fix(ui): compact ROI dashboard header and metrics

* fix(ui): show ROI repository count with expandable list

* fix(ui): wrap ROI controls within narrow panels

* fix(roi): restore sample report preview and simplify setup
2026-10-03 23:07:34 +00:00
yujonglee
c7e60f03de
fix(tracing): preserve spend identity and gateway correlation (#44421)
* fix(tracing): preserve spend identity and gateway correlation

* test(tracing): refresh real SDK spend captures

* fix(tracing): resolve complete gateway attempt costs across SDKs

* docs: add trace cost screenshot for PR 44421

* update fixtures

* wip

* docs: remove trace cost screenshot from PR evidence
2026-10-03 16:00:49 -07:00
devin-ai-integration[bot]
01b4ffe16b
fix(bedrock_mantle): route Claude chat completions to the native Messages endpoint (#43646)
* fix(bedrock_mantle): route Claude chat completions to the native Messages endpoint

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock_mantle): price Claude chat on the Mantle row and route region-prefixed ids

A Mantle request always carries a region, so a Claude id with no bedrock_mantle/<region>/ row fell through the model-info lookup to the bare Bedrock row, which the bedrock provider family also matches, and billed about 10 percent under the Mantle price. The lookup now tries the provider's region-free row before the bare model. The Claude route test asserts the Mantle row, a region-prefixed Claude id is covered end to end, and the provider config map references the Mantle config directly.

* test(bedrock_mantle): cover supported params for Claude and open-weight Mantle ids

* test(bedrock_mantle): audit the Claude chat bridge on the integration rig

Adds the deterministic cells from the /audit of the Mantle Claude chat
bridge: wire-level translation on every chat, responses, and messages
route, SigV4 and bearer auth, region prefixes, api_base suffixes,
unsupported params with and without drop_params, malformed model ids,
upstream errors, the response-cache hit, the health check, pricing from
the Mantle row for Claude and non-Claude ids with a bare Bedrock twin,
and chaos cells for a mixed burst, an upstream outage, slow streams,
and a worker kill on an owned two-worker proxy

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 15:24:41 -07:00
joshua-berri
0ea166c160
fix(mcp): preserve upstream tool schemas and parameter headers (#44425)
* fix(mcp): preserve tool schemas and modern parameter headers

* fix(mcp): allow bounded cold schema worker startup

* fix(mcp): align catalog deadlines and bound schema traversal

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-03 15:00:43 -07:00
devin-ai-integration[bot]
98337c9334
fix(lens): release budget reservations when the analysis model call fails (#44431)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:39:27 -07:00
devin-ai-integration[bot]
41a3781d4e
fix(proxy): treat Postgres connection exhaustion as backpressure, not poison rows (#44266)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:28:52 -07:00
devin-ai-integration[bot]
9a4b1951a2
fix(proxy): treat Postgres connection exhaustion as DB unavailable, not poison spend-log rows (#44270)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:28:52 -07:00
devin-ai-integration[bot]
e9bf2cfd01
refactor(types): replace Any with proven types in 9 files (#44389)
* refactor(types): replace Any with proven types in 13 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep provider error paths for malformed prefetch and poll JSON

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's login body parsing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's Copilot auth and budget alert typing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for Any sweep 20261003_2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): load audit video deployments at proxy start

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's video-edit prefetch handling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): fix audit chaos Responses cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): probe every route after audit worker kill

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 21:03:16 +00:00
devin-ai-integration[bot]
fe683ea139
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models

GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.

* fix(proxy): offer a Codex service tier only when every deployment of the model lists it

* fix(codex-catalog): an invalid service_tiers value offers no tier for the model

* fix(codex-catalog): read service tiers off the deployments the key's team can route to

A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them

The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default

* test(codex-catalog): drop the redundant module docstring and sort the imports

* test(integration): add the Codex catalog audit cells and the multi-worker convergence note

* test(integration): clean up every catalog test model and answer the refresh GET

* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers

Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.

* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut

The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns

The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 20:42:55 +00:00
Misbah Syed
fd14e51ecf
feat(docker): add a Windows quickstart in PowerShell, and a styled terminal for both quickstarts (#44310)
* feat(docker): Windows quickstart in PowerShell, and a styled terminal for both quickstarts

scripts/quickstart.ps1 does what scripts/quickstart.sh does, for Windows:
  powershell -c "irm https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/quickstart.ps1 | iex"
It asks the same two questions, prints the same lines, writes the same .env
(UTF-8 without a byte order mark, LF endings, readable only by the current
Windows account), and runs in Windows PowerShell 5.1 and PowerShell 7.

Both scripts now give a person at a terminal colors, a check mark per step, a
spinner while Docker starts, and a framed summary. Agents, CI, log files, and
NO_COLOR get the same lines as plain text.

* feat(docker): support podman and rancher desktop in the quickstart scripts

Both quickstarts hard-required the docker CLI. They now pick the first
available engine among docker, podman, and nerdctl (Rancher Desktop in
containerd mode; its dockerd mode already provides a docker CLI), route
every invocation through it, and tailor the start hint (podman machine
start) and the printed stop/logs/volume commands to that engine

* fix(docker): test port availability by binding instead of connecting

On a WSL2-backed engine (Podman, Rancher Desktop), the Windows localhost
relay swallows connection refusals on closed ports, so every connect
waits out its 2-second timeout and Test-PortFree reported ports 4000 to
4099 all taken on a machine with none of them in use. Binding the port
answers instantly and accurately

* fix(quickstart): address review findings

- The PowerShell script names the project with the same POSIX cksum as the
  shell script, so the earlier-install check sees the database from either
  script.
- Both scripts stop when git tracks .env in the install folder.
- .env is written to a temp file with owner-only permissions and moved into
  place, so a failed write never leaves a partial file.
- Errors stay plain text when stderr is redirected.
- The spinners remove their temp files on exit and Ctrl+C, and the shell
  spinner no longer runs date on every frame.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(quickstart): offer to copy the admin password to the clipboard

After the summary, the quickstarts ask "What next?": copy the admin
password, open the admin UI, or finish. The password is piped to the
clipboard (pbcopy, wl-copy, xclip, xsel, clip.exe, or Set-Clipboard), so
it never appears on screen or in the process list. Over SSH, or with no
clipboard, the option is left out and the summary points to .env as
before. The PowerShell menu now honours Ctrl+C.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(quickstart): write .env through mktemp, let Ctrl+C cancel the PowerShell menu, drop redundant comments

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(quickstart): remove the in-progress .env temp file when interrupted

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Mubashir Osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:35:32 -07:00
ishaan-berri
dd31692282
feat(lens): always-on investigations with findings and investigations tables (#44418)
* feat(lens): record run steps, trigger and exact run windows on investigations

* feat(lens): scan only traces since the last run and keep a capped step log

* feat(lens): log each analysis model call with its model, tokens and cost

* feat(lens): accept agent and time window on run now and add turn-all-on

* test(lens): cover new-traces-only windows, manual runs and the step cap

* test(lens): cover run now overrides and turning paused investigations on

* chore(ui): regenerate api types for lens run steps and run now options

* feat(lens): group open findings into one row per problem with agent filters

* test(lens): cover the findings table grouping, filters and schedule labels

* feat(lens): add a findings table across all investigations

* feat(lens): show a live step feed with the model behind each call

* feat(lens): offer turning all paused investigations on

* feat(lens): let run now pick an agent and time window

* test(lens): cover run now request building

* feat(lens): show the step feed and run now dialog on an investigation

* feat(lens): open run now choices instead of running immediately

* feat(lens): open findings first and peek a finding without leaving the table

* feat(lens): fold investigation actions into the findings toolbar

* feat(lens): show each investigation's schedule and open findings

* feat(lens): name the agent on a finding

* feat(lens): keep new investigations watching every 15 minutes by default

* feat(lens): show the watch schedule outside advanced options

* feat(lens): send run now options and turn-all-on from the dashboard

* test(lens): give demo runs steps and a trigger

* feat(lens): let the findings table fill the screen

* test(lens): add steps and trigger to progress fixtures

* test(lens): add steps and trigger to status fixtures

* test(lens): cover the default watch schedule in setup

* test(lens): cover run now choices from an investigation

* test(lens): open saved investigations from the manage view

* feat(lens): use one tab bar for traces, findings and investigations

* feat(lens): place page actions on the lens tab row

* feat(lens): drop the nested tabs and edit investigations in place

* feat(lens): show investigations as a table with run now and edit

* feat(lens): name each findings row for screen readers

* test(lens): open saved investigation links on findings

* test(lens): reach findings and investigations from the top tabs

* fix(lens): mark run now jobs manual and keep them from moving the scheduled scan

* fix(lens): keep run now since-last-run windows even with an agent override

* test(lens): cover that manual runs never skip scheduled traces

* test(lens): cover run now windows with agent and lookback overrides

* fix(lens): group findings without Map.groupBy and expose sampled runs

* fix(lens): open older findings and review every merged copy from one row

* fix(lens): hide edit and run now from read-only viewers

* chore(lens): drop restating comments from the findings table

* chore(lens): drop restating comments from the step feed

* chore(lens): drop restating comments from the paused banner

* chore(lens): drop restating comments from header actions

* chore(lens): drop restating comments from run now

* test(lens): cover merged findings and read-only investigation rows

* fix(lens): record a model step even when the response has no usage

* test(lens): cover model steps with and without reported usage

* fix(lens): keep a merged finding open when one of its updates fails

* refactor(lens): accept update results from the investigations view

* refactor(lens): accept update results in investigation actions

* test(lens): cover retrying a merged finding after a failed update
2026-10-03 20:27:02 +00:00
devin-ai-integration[bot]
f445e466b4
refactor(traces): extract snapshot cache (#44424)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 13:03:30 -07:00
devin-ai-integration[bot]
671748067d
fix(model_prices): registry audit 2026-10-03, add cohere embed v5, grok-imagine-video-1.5-lite and openrouter gpt-image rows (#44376)
* fix(model_prices): registry audit 2026-10-03, add cohere embed v5 and absorb openrouter gpt-image rows

Co-authored-by: tinysolver <iam.tinysolver@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): add grok-imagine-video-1.5-lite and gemini 3.5 transcribe limits, drop deprecation-only date

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: tinysolver <iam.tinysolver@gmail.com>
2026-10-03 12:53:07 -07:00
devin-ai-integration[bot]
d306d6d70b
feat(traces): render curated SQL examples from shared files (#44388)
* feat(traces): render curated SQL examples from shared files

* fix(traces): filter trace-summary example after grouping so boundary-spanning traces keep full totals

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(traces): order failed-spans example by timestamp so newest failures survive the limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 12:40:01 -07:00
devin-ai-integration[bot]
1b98748528
fix(bedrock): honor the per-request timeout on Converse and Invoke streaming (internal copy of #38210) (#44134)
* fix(bedrock): propagate timeout to streaming requests

* test(bedrock): prove streaming fails at the request timeout against a slow upstream

* test(bedrock): simulate the slow upstream in process instead of over a local socket

* test(bedrock): audit the Converse and Invoke stream timeout on the proxy

Two integration files drive the per-request timeout on Bedrock streams
through the real proxy against an owned wire peer: the wire file covers
every surface (chat, messages, responses, invoke, pass-through), the
sad, edge and precedence rows, and the chaos file covers bursts, a
dropping upstream, a killed worker and a proxy stopped mid-burst.

The wire peer gains Reply.drop_connection so a cell can close the
socket before any response, and the harness's graceful stop grace is
now INTEGRATION_PROXY_STOP_SECONDS (default unchanged at 30), since a
two-worker supervisor's interpreter finalization takes longer than that
on a loaded box.

* test(bedrock): pin the fallback audit cell to one proxy worker

The fallback cell created both deployments through /model/new on one
worker and sent the chat request to the other, whose registry
read-through loads only the requested model, so the fallback target
was unknown there until the periodic DB poll. The cell now warms the
fallback model and sends the request over one keep-alive client, so
one TCP connection stays with one uvicorn worker, and it expects the
fallback upstream to see both requests.

---------

Co-authored-by: Sainyam Kapoor <hello@sainyam.me>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 19:29:15 +00:00
ishaan-berri
9e9c29f404
feat: add make lens-dev for one-command lens local dev (#44413)
* feat(lens): add scripts/lens_dev.sh for one-command lens local dev

* feat: add make lens-dev target

* chore: gitignore .lens-dev local state

* docs(lens): mention make lens-dev in the developer note

* fix(lens): resolve a relative LENS_DEV_CONFIG against the caller's cwd

* fix(lens): random lens-dev master key, verify reused services, drop inherited REDIS_*

* test(lens): cover lens-dev worker token, master key, env and cleanup

* fix(lens): skip compose postgres when LENS_DEV_DATABASE_URL is set

* test(lens): external LENS_DEV_DATABASE_URL never starts compose postgres
2026-10-03 19:10:04 +00:00
moe-berri
af36e5c693
fix(lens): preserve full trace access and expose investigation failures (#44406)
* fix(lens): preserve full trace access and expose investigation failures

* chore(lens): sync worker registration schema

* fix(lens): support durations without a configured maximum

* fix(lens): expose every page of fetched investigation evidence

* fix(lens): preserve repeated trace content and interrupt cancelled runs

* fix(lens): retry transient heartbeat failures during analysis
2026-10-03 11:52:55 -07:00
moe-berri
50190134c3
fix(lens): batch run reads and reset trace pagination (#44398)
* fix(lens): batch run reads and reset trace pagination

* fix(lens): scope batched list spend to each run
2026-10-03 18:43:23 +00:00
berriai-litellm-provider-info-sync[bot]
9fe6442172
fix(azure): add azure_ai/flux.2-pro input token limit from the models sold directly page (#44405)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 11:40:44 -07:00
devin-ai-integration[bot]
a57cdfddf8
fix(ci): give the pass-through auth regression test a real FastAPI app (#44403)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 18:31:48 +00:00
devin-ai-integration[bot]
a849093d89
test(cost): pin OpenAI reported web search count with mixed actions (#44414)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 11:31:45 -07:00
devin-ai-integration[bot]
6c32384d8c
fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3 (#44292)
* fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3

Bedrock prices Kimi K3 cache reads through implicit caching but Converse
rejects the explicit cachePoint marker ("This model doesn't support the
cachePoint field"), so any cache_control on the request answered 400.
Mark the three K3 rows supports_prompt_cache_breakpoint: false and have
bedrock_model_accepts_cache_points honor that flag before falling back to
supports_prompt_caching, keeping cached-token pricing intact.

* test(bedrock): assert cache points per request section

* fix(bedrock): honor a deployment's cache breakpoint flag for unmapped models

* fix(bedrock): read a converse-routed deployment's cache breakpoint flag

* refactor(bedrock): look up cache breakpoint flags by key

* test(bedrock): add the Kimi K3 cache point wire audit

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 18:11:46 +00:00
tin-berri
4732647de2
feat(ui): make LiteAdmin enterprise-only (#44399) 2026-10-03 11:07:54 -07:00
ishaan-berri
f0abe1bea1
feat(lens): live dot field timeline and full-screen traces view (#44390)
* feat(lens): add dot field layout and live status helpers for the traces timeline

* test(lens): cover dot field layout, agent colors and live status

* feat(lens): draw the traces timeline as a live dot field

* feat(lens): add the sweep animation for the live traces timeline

* feat(lens): put the lens tabs in a compact header and fill the screen with traces
2026-10-03 10:56:48 -07:00
moe-berri
ad8babae33
fix(lens): paginate trace reads within ClickHouse limits (#44384) 2026-10-03 10:38:40 -07:00
devin-ai-integration[bot]
8b1990b4bc
feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers (#44236)
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register typesafe as a provider so Jev deployments load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move provider endpoints under llms and validate proxy bodies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add Cloudflare Clef and Strands Decider backends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register decisions routes for managed agents and gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(decisions): use raw regex for cloudflare missing account match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): avoid cast in Cloudflare response unwrapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): default model, evaluation health probe, short Cloudflare names

The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.

* fix(decisions): let health_check_params override the evaluation probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit the decisions endpoint across providers, limits, health and chaos

Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).

The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.

* fix(decisions): send env API keys to a configured api_base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register the routes through the lazy feature registry

The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.

The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.

* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell

* fix(proxy): let a config pass-through beat a lazily registered route in eager mode

With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:38:38 +00:00
devin-ai-integration[bot]
b024950353
feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop (#44391)
* feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(harness): name per-tool spec FunctionTool

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-03 17:33:01 +00:00
berriai-litellm-provider-info-sync[bot]
231a46e40b
feat(azure): add azure_ai/kimi-k2-thinking from Azure Kimi pricing page (#44382)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 10:04:59 -07:00
devin-ai-integration[bot]
a5e9f275ea
fix(proxy): keep the database error when the log_db_metrics failure hook raises (#44383)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 16:38:15 +00:00
devin-ai-integration[bot]
4c6c84afb7
perf(proxy): stop prompt-cache eligibility from tokenizing the whole conversation (#44221)
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them

Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
797353f13a
fix(otel): name postgres service spans by operation and table (#44240)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
564d236985
fix(otel): nest cache spans under their operation and name service spans by purpose (#44150)
Response cache reads and writes open cache.get llm_response and cache.set llm_response phase spans with their Redis spans nested underneath, on the Python path and on the native Rust path, and deployment selection runs inside a route {model_group} phase so the cooldown, usage and model-id reads the router issues nest under it before chat {model}. The autorouter classifier call nests under that route phase as well and carries its typed internal origin on litellm.request.purpose, so it is told apart from the provider attempt. Service spans are named {service}.{verb} {target} from a low-cardinality key family the producer declares (llm_response, auth_objects, spend_counters, router_cooldowns, claude_code_session_router_binding, rate_limits, pod_lock, budget_reset, ...) instead of the raw method or a per-request pipeline length; a pipeline flush is targeted by the one family its ops share or by mixed with the sorted families on litellm.redis.families, a batch op keeps the family it was declared under whichever pipeline or standalone read settles it, and the ambient family labels Redis spans only, never the DB write-back a task spawned inside that context performs later. The raw method stays on litellm.service.call_type and on the Prometheus and Datadog labels. Caller attribution is carried across asyncio task boundaries on a ContextVar so forwarder-only chains no longer surface, the raw cache key is dropped from Redis span metadata, pipeline op counts land as an integer attribute, every call_type the Redis cache layer emits maps to a verb, and a scan over litellm/ and enterprise/ fails when a Redis producer, batch reservation included, declares no key family.

A V2 logger built for a key or team logging entry while the operator's V2 logger is already registered keeps only the exporters its own preset contributed, whether or not the operator holds credentials for that backend, so every chat span no longer reaches the operator's collector twice. A span the success callback has to open itself, with no pre-call carrier, starts at the provider handoff (api_call_start_time) instead of the logging object's creation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
devin-ai-integration[bot]
53d2ab6b05
fix(proxy): emit postgres service spans only on real DB reads in auth cache helpers (#44148)
log_db_metrics wrapped whole cache-first auth helpers and always emitted a ServiceTypes.DB success event, so in-memory cache hits showed up as postgres <fn> spans and DB service metrics. The decorator now installs a ContextVar witness that _TrackedPrismaEngine marks on every Prisma query and transaction call, and the DB event is emitted only when the witness was marked. Real reads keep their existing call_type names, the failure path and the PROXY batch-write branch are unchanged, and Redis instrumentation is untouched.

A decorated helper that reaches Prisma only through another decorated helper (get_key_object -> get_object_permission, get_team_object_by_alias -> get_object_permission, get_tag_object -> get_tag_objects_batch) used to emit two events for one query. The inner wrapper now marks its witness as reported when it emits a success or DB failure event, and only unreported activity is handed up to the enclosing witness, so the inner event is the one that survives. An outer helper that also queries Prisma directly or through undecorated callees still gets its own event.

Tests: get_user_object and get_org_object cache hits emit no DB event; a get_user_object miss through the generated Prisma client emits exactly one postgres get_user_object event; decorator-level tests cover nested calls emitting only the inner event, outer calls with their own query, inner non-DB failures, bounded lookups and sibling-request isolation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
berriai-litellm-provider-info-sync[bot]
66a422ea50
fix(azure): add MAI-Image max_input_tokens from models sold directly page (#44375)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 08:49:09 -07:00