Commit graph

19 commits

Author SHA1 Message Date
devin-ai-integration[bot]
fa2c8984ba
test: move the unit half of 126 mixed legacy files into tests/unit (#45090)
* test: move the unit half of 126 mixed legacy files into tests/unit

* test: restore litellm globals that moved tests set

* test: finalize migration test cleanup

* test: restore original bodies of moved legacy tests

The move into tests/unit had rewritten 612 test bodies, and some of the rewrites dropped assertions. Each moved test now carries its original body from the legacy file, with only the imports, helpers, fake provider credentials and monkeypatched env it needs to run under tests/unit

test_timeout_streaming goes back to tests/local_testing because it needs the fake OpenAI endpoint server. The image payload fixture moves with its only user, and two tests that leaked global state (a registered model cost entry and queued logging tasks) are now isolated

* test: drop module imports shadowed by restored local imports

* test: assert on LiteLLM output in no-assertion moved tests and isolate leaks

Twenty no-assertion candidates get one assertion on the value LiteLLM returns, with the original lines unchanged. Four tests go back to their legacy files because they only check types or imports, write into the working directory, or cannot assert without a body change

Two moved tests leaked globals into later tests in the same worker, so monkeypatch fixtures now restore the retry-after header parser and the end user cost tracking flags

* test: drain queued logging tasks before the Phoenix span test

The moved Phoenix test counted spans from logging tasks that earlier tests had queued, so the drain fixture moves to tests/unit/conftest.py and both it and the Datadog batch test use it. test_factory_function goes back to its legacy file because its returned wrapper calls the real Assistants API and cannot be asserted on without a body change

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 14:07:43 -07:00
devin-ai-integration[bot]
e2971e0af4
refactor(llms): expose public names for private provider helpers (#45037)
* refactor(litellm): migrate private usage in llms

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): preserve Bedrock batch signature marker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(litellm): migrate llms private usage symbols

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): preserve UUID Watsonx project IDs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(litellm): retarget llms mocks to public names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(litellm): limit llms changes to renames and forwarders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(litellm): keep GCS mock client patching private Vertex auth methods

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(litellm): assert forwarder arguments and type forwarder signatures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 08:49:25 -07:00
devin-ai-integration[bot]
dfe4df8c8a
test: move whole-unit legacy test files into tests/unit and delete dead skips (#44807)
* test: delete unconditionally skipped legacy tests

* test: move whole-unit legacy test files into tests/unit

* test: keep moved legacy tests free of import-time global state

* test: keep the module-level invocation scan pointed at tests/local_testing

* ci: drop the agent_testing CircleCI job emptied by the move

* test: fix moved-test isolation and router coverage

* test: add Tinyfish search package marker

* test: isolate moved tests from logger state leaks

* test: isolate Helicone logging fixture state

* test: isolate Vertex pass-through credentials between moved tests

* test: cancel S3 periodic flush tasks started by moved tests

* ci: restore CircleCI assistant test selection after move

* test: prevent Datadog datetime import shadowing

* ci: drop the litellm_assistants_api_testing CircleCI job emptied by the move

* test: deduplicate imports in rebased unit tests

* test: remove duplicate passthrough router patch import

* test: remove moved legacy source files after rebase

* test: align moved tests with rebased main

* test: carry main's legacy-file edits into moved destinations

* test: make the moved cost map fallback tests assert the fetch and the backup

The four fallback cases only checked the result was non-empty, so they still
passed with integrity validation disabled. They now inject a mock client, assert
one fetch happened, and assert the result is exactly the local backup with the
fallback reason recorded.

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 05:02:28 +00:00
Mateo Wang
d4a791d650
refactor(types): replace Any with proven types in 157 files (#44798)
* refactor(types): replace Any with proven types in 299 files

Clears 871 basedpyright Any errors (reportAny 6,703 to 6,240, reportExplicitAny 1,651 to 1,243) without adding a cast, an ignore or a suppression, and without touching any budget file

Most edits are annotation-only: a parameter, return or local goes from Any to object, Mapping[str, object] or the concrete type the value always held. Fourteen files validate untyped JSON once where it enters, through a module-level pydantic TypeAdapter or model_validate, and then use real types

No HTTP status, error type or response shape changes. Mistral speech and fal.ai Bria image generation now report a pydantic ValidationError instead of an AttributeError when the provider answers 2xx with a body that is not a JSON object

* refactor(types): make the config locals fix effective and trim no-op edits

The Mapping[str, object] annotation on `locals().copy()` removed no error,
because the checker narrows the variable back to the dict[str, Any] the call
returns. 21 provider config constructors now build the same copy with
dict(locals()), which the checker infers as object values under that
annotation, so each file loses one reportAny.

The same annotation is reverted in 31 other config files where it stayed a
no-op, together with the tests that were added only to cover those lines, and
the one MCP server manager line that no CI coverage shard executes is
reverted too. The pull request drops from 345 to 297 changed files.

* refactor(types): accept only int in the proxy state setter

get_proxy_state_variable is annotated to return int, but
set_proxy_state_variable still took Any, so the checker could not hold
callers to the type the getter promises. The setter now takes int, which is
what its only caller already passes.

* refactor(types): index the proxy state key so the getter returns int

* refactor(types): keep public annotations and provider error text unchanged

Restore every public return, public method parameter, public attribute and exported
alias to its annotation on main so code that type-checks against the package keeps
type-checking, and take the Mistral speech and fal.ai Bria changes back out so no
provider error message differs from main

* refactor(types): leave the Vertex RAG chunking read as it is on main

Take the chunking format validation back out of the Vertex RAG ingestion path. It needs the vertexai SDK, a storage bucket and a RAG corpus to execute, so nothing here could run it end to end, and it cleared only two errors

* test(integration): pin the validated provider boundaries on a live proxy

* test(integration): give the held burst a client that outlasts the gate

The fault cell holds a burst at the upstream for up to 60 seconds while it kills a worker, but sent the burst through the shared 15 second client, so a slow box could time the survivors out before the gate opened. The burst now goes through its own client whose timeout is twice the gate, and the gate length is one named constant.

* refactor(types): keep the license reply handling and experimental MCP signatures as they were

The license check validated the whole reply as a mapping, which changed the error text logged for a reply that is not an object. It now validates only the verify value, so every reply is handled and logged exactly as before while the value is still typed.

Three files under the experimental MCP server changed annotations on public functions and methods (three returns and three parameters). They go back to their previous content so no public signature in the diff is narrowed.

* test(integration): answer the proxy's model-list call in the OpenAI stand-ins

Every 300 seconds each proxy worker asks an OpenAI deployment for GET /v1/models. Four new cells own an OpenAI stand-in that accepted only the call under test, so a refresh landing inside a cell failed it. The stand-ins now answer that call through the suite's own helper and the cells count only the provider calls they drive.
2026-10-06 15:31:54 +00:00
Mateo Wang
d80f8c28ca
refactor(types): replace Any with proven types in 137 files (#44478)
* refactor(types): prove runtime types at harness, search, rag and client boundaries

Replace Any with adapter-validated types in the litellm.agent() harness, the
search provider transformations, RAG ingestion and query, the vector store
pre-call hook and registry, the galileo and opik logging integrations and the
proxy client CLI. Each boundary gets unit tests for well-formed and malformed
payloads.

* chore(typing): prove types at more provider boundaries and restore search transformations

Second pass over a2a, embedding, rerank, image, audio and small provider
modules. Search transformations go back to their previous form because
validating their response bodies would change the proxy status for
malformed upstream bodies from 400 to 500.

* refactor(types): prove types at logging, files, rerank, image, audio and management boundaries

Replace Any with validated or annotation-only types in 37 more files: logging
integrations, token counters, provider files/rerank/image generation/audio
transcription transformations, pass-through logging handlers and management
endpoints. No proxy HTTP status or error type changes.

* refactor(types): prove types at repository, spend, files and router boundaries

Replace Any with repository table accessors, validated mappings and
annotation-only types in 33 more files: Prisma repositories, the enterprise
batch and responses cost checkers, budget reservation, files endpoints,
management endpoints, the policy registry, the adaptive and complexity routers
and the secret managers. No proxy HTTP status or error type changes.

* test(types): run the aiohttp transformation test in-process and cover repository row conversion

The aiohttp chat transformation test no longer starts a server. It feeds the
transformation a response whose json() returns the body under test.

The proxy unit shards now exercise stored model rows whose params are JSON
strings and the object permission create and update paths.
2026-10-05 11:01:45 -07:00
devin-ai-integration[bot]
7a7d27c550
fix(guardrails): run the end-of-stream post_call scan when the client disconnects mid-stream (#43839)
* fix(guardrails): run end-of-stream post_call scan when the client disconnects mid-stream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): close the guardrail stream chain in async_data_generator on client disconnect

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): leave the raw upstream response to the shielded finalizer on client disconnect

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep disconnect cleanup going when a streaming callback cleanup raises

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): assert the refund through a recorder instead of the mock

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): inspect tool calls released before a disconnect under incremental_diff and record a failed scan marker

The incremental_diff transform stream now scans tool calls it already released when the client disconnects, and a disconnect scan whose translation raises after the guardrail recorded success also records guardrail_failed_to_respond

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): pin that a text-only disconnect scan is not handed a tool_calls finish

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): cover disconnect scans on every streaming endpoint and client, plus outage, worker-kill and cache-hit cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): prove the cache-hit twin is served from the cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): type the disconnect-close streams so basedpyright stops reporting unknown arguments

* fix(guardrails): give the guardrail metadata cast a reason so the type discipline gate accepts it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): scan released Messages and Responses tool calls on disconnect

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(guardrails): format the disconnect scan unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(guardrails): drop mutable-ok markers that no longer suppress a rule

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): scan released Responses output after a finished item and end only in-flight Chat choices

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): type the disconnect scan test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): type request_data in the disconnect scan helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): type request_data in the disconnect scan test doubles

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): pin that chat streams with no tool call in flight end as released

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): scan only released chunks on disconnect and skip it once a block owns the verdict

The disconnect scan now uses the chunks actually yielded to the client, copies them before scanning,
skips when a mid-stream block or HTTP error already settled the verdict, and the iterator wrapper
only closes hooks that are async generators so plain async iterator hooks keep working

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): pin that a delivered guardrail error or final chunk settles the disconnect verdict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): close any hook iterator that exposes aclose when the stream ends early

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): accept a synchronous aclose on custom streaming hook iterators

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): swallow callback aclose errors at end of stream

A custom callback whose async_post_call_streaming_iterator_hook returns a
non-generator async iterator with a raising aclose() failed the finished
stream: content plus usage reached the client and then the stream surfaced
an error SSE with no [DONE], or aborted a post_call pipeline's buffering
loop into a 500 with an empty body. Wrap the aclose invocation in
_wrap_streaming_iterator_with_enrichment in try/except and log a warning
naming the callback and the cleanup error, matching close_guarded_stream
and _close_guarded_layers. Iteration-time hook exceptions still propagate.

* fix(proxy): log only the error type when a callback aclose raises

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 01:44:43 -07:00
devin-ai-integration[bot]
0ed1c08f02
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources

Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.

Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.

The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.

Fixes #28607
Resolves LIT-6107

Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>

* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries

The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.

LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.

* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle

* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries

* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly

* fix(proxy): decrypt stored litellm_params before the WIF write gate

* fix(proxy): hide WIF secret references from /health output

* fix(proxy): keep the proxy error shape on credential endpoint refusals

* fix(proxy): hide identity token file paths from /health output

* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working

The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader

* fix(auth): share one exchanged token across workers reading the same assertion

Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it

* fix: keep anthropic federation from being shadowed or leaked

An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.

The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.

The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.

* fix: unlink a staged token file a failed write leaves behind

The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.

* refactor: move anthropic jwks derivation behind a provider-owned tagged union

* fix: unlink the staged token file when its write fails at close

A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.

* fix(anthropic): close the staging descriptor before writing the shared token file

* fix(wif): judge federation writes by what they set, not what is stored

The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500

The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL

* fix(proxy): let a deployment write name a federated credential

reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.

is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.

* refactor(proxy): derive health display policy from the federation key sets

The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places

WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one

* fix(anthropic_wif): treat blank identity-source fields as unset

* test(proxy): classify the federation params in the credential slot registry

main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret

* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics

Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.

* fix(credentials): gate PATCH on WIF fields resolved from model_id

The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path

* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header

Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
2026-10-03 17:08:30 -07:00
devin-ai-integration[bot]
fe910889f7
fix(responses): drop bridge-minted reasoning items from OpenAI replays (#44132)
* fix(responses): drop bridge-minted reasoning items from OpenAI replays

* fix(responses): send id-less stored reasoning items without a made-up id

The chat-to-Responses bridge gave a stored reasoning item with no id an rs_<n> id that OpenAI rejects (404 without encrypted content, 400 with it); an id-less item is accepted and verified by OpenAI itself. Decoding encrypted_content now keeps the verifiable thinking blocks of a mixed array instead of rejecting the whole array, and the verifiable-block rule lives in the shared module.

* test(integration): audit minted reasoning item replay across responses, chat bridge and messages

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 19:39:48 -07:00
devin-ai-integration[bot]
a308a8e579
feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (#43134)
* feat(guardrails): honor litellm_params.timeout in every HTTP guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): accept timeout kwarg in presidio and responses-handler post stubs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): bound hiddenlayer startup jwt call by configured timeout, drop akto from timeout coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): narrow hiddenlayer startup auth timeout without cast

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): bound hiddenlayer jwt refresh by configured timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep provider timeout defaults when unset and bound only rubrik moderation calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): cover model_armor and run timeout probes concurrently

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): match sink calls to the exact guardrail name

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 21:18:47 -07:00
devin-ai-integration[bot]
264b09ac8d
fix(responses): scan and mask top-level instructions with guardrails (#43629)
* fix(responses): scan and mask top-level instructions with guardrails

The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.

Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.

Resolves LIT-8931

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): reject empty guardrail rewrites instead of forwarding raw input

An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the texts-replacing guardrail helper explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): cover skip_system_message_in_guardrail on the live proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system

Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): annotate new guardrail tests with return types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the guardrail test doubles explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 11:44:35 -07:00
devin-ai-integration[bot]
e814532033
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): satisfy type-discipline and strict ruff budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): stamp the served service_tier on every Responses bridge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): cover anthropic and responses served-tier billing paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): bill disconnects through the router's anthropic stream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: apply ruff format to the anthropic stream changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(coverage): ignore delegating properties the ast scan cannot see

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: keep the cast-ok reasons on the cast call line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(streaming): parameterize delegated chunks and messages types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): follow the anthropic pass_through rename after merging main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drain the logging worker between response cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): keep the served service_tier on streamed chunks and bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): type the served service_tier chunk without a loose kwargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill the served service_tier over the requested one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): drop explanatory comment from the tier resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-29 12:54:17 -07:00
devin-ai-integration[bot]
b248b1c7dc
fix(openai): exclude fine-tuned and custom gpt-5-chat aliases from gpt-5 reasoning path (#43185)
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
* fix(openai): exclude fine-tuned and custom gpt-5-chat aliases from gpt-5 reasoning path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(openai): keep gpt-5-chat alias regression test diff minimal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(openai): cover temperature pass-through for gpt-5-chat aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(openai): annotate locals and wrap long lines in gpt-5-chat alias test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:09:09 -07:00
yuneng-jiang
5e6dc89ba1
test: move tests/test_litellm/llms into tests/unit/llms (#43191)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: move tests/test_litellm/llms into tests/unit/llms

Rename-only. Moves the provider tests and the fine-tuning fixtures they
load, mirroring the old paths. Follow-up commits merge, split and wire them.

* test: merge, split and prune the moved llms tests

Merges the Databricks chat transformation tests into the existing unit
file, keeps the tests that need real keys or the network in
tests/test_litellm, deletes the audited tests a stronger unit test
already covers, and points imports at tests.unit.llms.

* ci: run the moved llms tests under their legacy flags

The Vertex AI and All Other Providers shards keep their legacy test-path
for the retained files and add the llm-vertex-ai and llm-other-providers
unit selections. CircleCI gets matching unit jobs.

* test: make the tests/unit/llms directories packages

Adds __init__.py to the moved dirs and drops the legacy ones whose
directories no longer hold tests.

* test: drop script runners and path hacks the llms split left dangling

The __main__ runners in the split openai_like files and the Databricks e2e
runner called tests that now live in the other half of the split or were
deleted. The retained legacy halves also no longer need sys.path edits.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

* test: keep the Databricks manual e2e runner and fix the SageMaker Nova run path

The Databricks e2e file is a manual script whose main() calls the tests
that were pruned, so pruning them broke the documented run. It is back to
its main version. The SageMaker Nova docstring now points at the file's
real location in tests/local_testing.

* test: keep the job's UNIT_FLAG out of the shard-script tests

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 12:43:23 -07:00
devin-ai-integration[bot]
1edc4ba580
fix(logging): pass provider response headers to callbacks on every endpoint (#42824)
* fix(logging): pass provider response headers to callbacks on every endpoint

Custom callbacks only received kwargs["response_headers"] for chat
completions. Responses, image generation and edit, speech, and
transcription calls either never recorded the provider's headers or
recorded them in one place and not the other.

Every handler now records the provider's httpx headers on the response's
hidden params as "headers" (raw) and "additional_headers" (processed,
with LiteLLM's own entries winning on a clash), and the logging object
derives model_call_details["response_headers"] from those hidden params
before cost calculation on the non-stream and both streaming success
paths, keeping a handler-set value authoritative. Binary speech responses
expose their hidden params to the standard logging payload, and the sync
OpenAI transcription request always fetches the raw response.

* test(images): point the legacy image and speech fakes at the raw response surface

Image generation now goes through the SDK's raw response so the provider headers can be read, and the speech binary response now carries hidden params. The unit fakes in the image generation, xinference, proxy provider, image edit, Vertex speech, and otel suites still pinned the old call surface and the old "no hidden params" assertion, so they read an uncalled mock or a fake response without headers.

* test(images): drop the rewritten mock comments and the generated edit PNGs

* test(images): move the llm-span test's image fake to the raw response surface

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 13:01:12 -07:00
yuneng
1ca0a662f0 Merge remote-tracking branch 'origin/main' into litellm_migrate_tests_p11 2026-09-20 12:55:05 +00:00
yuneng
5599c59923 test: add package initializers to migrated unit test directories
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 11:52:07 +00:00
yuneng
e4a58ef91a test(unit): make every tests/unit directory a package so pytest collection is unique
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 11:50:59 +00:00
yuneng
729a96e4ea test: migrate openai, openai_like and openrouter legacy tests to tests/unit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 10:40:27 +00:00
yuneng
443c9f8385 test: migrate nvidia, oci, ocr, oobabooga and openai legacy tests to tests/unit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 08:09:49 +00:00