litellm/tests/test_litellm
PhimmStraiker e0af9917a1
feat(guardrails): straiker guardrail speaks the v3 platform API (/api/v3/detect) (#41880)
* feat(guardrails): speak the Straiker v3 platform API (/api/v3/detect)

The Straiker guardrail posted a webhook envelope to /api/v1/detect/webhook.
The v3 platform exposes /api/v3/detect instead, and its integration keys
(sk_agt_…) are rejected by the v1 route with an empty 401, so a tenant on
the v3 platform could not run this guardrail at all. Measured on a
customer gateway on 2026-09-17 after they rotated to a v3 key.

v3 parses the gateway's own traffic server-side, the same contract as
Straiker's unified Kong plugin. So on v3 the guardrail relays: the
request phase posts the provider body LiteLLM received (Anthropic
Messages or OpenAI chat), the response phase posts
{straiker_phase, sse, model, request}, the answer beside the request it
answers, and Straiker derives prompt, answer, agent and archetype. Both
phases also carry the flat prompt / app_response pair: a gateway-mode
integration key scores only the flat pair and an api-mode key only the
relayed body, each ignoring the other, so one payload serves whichever
key the console issued and it is one turn either way (measured on tenant
123, both key modes, 2026-09-18).

- api_version: "v1" | "v3", unset follows the key prefix, so a v3 key
  needs no extra configuration. Explicit override still wins.
- The relayed body is an allowlist of provider fields. The hook sees the
  client body merged with proxy state: `deployment` carries the resolved
  provider credential and `proxy_server_request` the client's own
  Authorization header. Neither travels. Identity survives as the
  metadata subset Straiker's LiteLLM adapter reads.
- Identity never sends a proxy placeholder. `default_user_id` and the
  master-key alias were being forwarded as a user and became the
  session's identity on the platform.
- Headers: x-tool: litellm (ingress), x-straiker-phase, x-straiker-user,
  and x-claude-code-session-id forwarded when the client sent it.
- Verdict: hookSpecificOutput.permissionDecision on the gateway envelope,
  `action` on the flat one; block on block/deny, and on a non-empty
  blocked_by as a backstop. A detect-mode control reads NONE.
- An error status from Straiker is now a webhook failure. LiteLLM's HTTP
  client raises on any non-2xx and the retry loop caught only connection
  errors, so a 401 or 503 from Straiker escaped the guardrail as an
  exception and was relayed raw to the client, bypassing fail_open /
  fail_closed. Retryable statuses retry; the rest are final.
- v1 is unchanged: same envelope, same X-Straiker-Webhook-Format header.

Tests: 15 new, fixtures from the request dict a hook sees on 1.98.0 and
the verdict envelopes the v3 platform returned on 2026-09-18. Each fix
was mutation-checked (handling removed, the test fails). Live: the same
eight-case battery (chat, /v1/messages, streaming, tool call; benign,
injection, PII) passes on a gateway-mode and an api-mode key, blocks at
pre_call with the tenant's block message, and lands under the declared
agent with the end user attributed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(guardrails): name the agent per application on v3 (x-s6r-agent)

One integration key can front several applications. Straiker enumerates them
as separate agents when the turn names one, which is what the unified Kong
plugin sends as x-s6r-agent. Without it every application on a gateway
collapses onto a single agent.

- Forwards a client-supplied x-s6r-agent.
- New `agent_ref` config names one agent for a route when the client sends
  nothing. The client wins, matching Kong's precedence.
- Neither set: no header, and the platform derives the agent from the traffic.

Verified live on tenant 123 against an integration whose connector is
`gateway`: three distinct values minted three observed agents, and a turn
with no hint derived one from the traffic shape. An integration whose
connector is `custom-agent` declares its agent, so every turn attributes to
that one agent and the hint is ignored (agent_ref_source: attested).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(guardrails): which v3 shape is scored depends on the connector, not the key mode

The earlier comment said a gateway-mode key scores only the flat pair. Re-measured
on tenant 123 across all three integration types with one injection prompt:

  custom-agent connector (Add Agent)  raw body ignored   flat prompt scored
  gateway connector                   raw body scored    flat prompt scored
  api mode                            raw body scored    flat prompt ignored

Behaviour unchanged: the payload already carries both shapes, which is why it works
on every type. Comment only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(guardrails): send exactly what the unified Kong plugin sends on v3

The v3 platform parses the gateway's traffic itself and derives agent,
archetype and identity from it. The earlier commits added to the relayed
body (a flat prompt / app_response pair, source, user_name) and to the
headers (x-tool, x-straiker-phase, x-straiker-user). None of that is in
the Kong v0.12 contract, and traffic through this guardrail was not
classifying by shape the way the same traffic through Kong does. Match
Kong byte for byte and leave classification to the platform.

Request phase: the provider body, plus session_id and
original.processed.Meta.user. Response phase: {straiker_phase, sse,
model, request} plus the same two. No flat fields, no phase or user
headers, no x-tool.

Session id follows Kong's precedence: the client's x-claude-code-session-id,
then the session LiteLLM resolved, then an md5 of system prompt + first
message so a conversation that states no session still groups across its
replays.

Routing hints complete the Kong set: x-s6r-agent (client header, else
`agent_ref`), and new `client` (x-s6r-client) and `format_hint`
(x-s6r-format) config, both optional.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(guardrails): sort imports in the v3 session test

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): send a streamed Messages answer back in the Messages shape on v3

On a streamed /v1/messages call the proxy rebuilds the answer as a chat
completion before the post-call hook runs, and that is what the plugin put in
the response envelope's sse field. Straiker's coding-agent reader parses a
Messages answer, so a Claude Code turn relayed this way came back
coding_agent/claude with no session and zero events scored: the model's tool
calls were never screened on the response phase. Captured live on 2026-09-18
against tenant 123, a real Claude Code Bash tool call through the proxy.

The proxy's own Anthropic adapter turns the rebuilt answer back into a Messages
response when the call arrived on the anthropic_messages route, which is what a
transport relay forwards. Chat completions calls keep the chat completion shape
and a buffered Messages answer is relayed untouched.

The regression test's fixture is the chat completion the proxy actually built
for that captured turn. After the fix the same turn scores on the response
phase (session resolved, one event, the Bash tool_use block present).

* style(straiker): ruff format the v3 guardrail and its tests

* refactor(straiker): one attempt per call in the webhook retry loop

The HTTPStatusError branch added for v3 duplicated the non-200 branch and put
_post_webhook over the strict complexity ceiling. One attempt is now its own
method that returns the verdict or a failure marked retryable, and the loop only
decides whether to try again. Behaviour is unchanged: retryable statuses and
transport errors retry, everything else is final.

* fix(straiker): name Claude Code's client and agent on v3 so its session lands under one coding agent

Straiker types a gateway turn as a coding agent from the "You are Claude Code"
preamble, which only the main agent turns carry. Claude Code's title and
topic-detection sidecars have their own system prompts, so they resolved by
shape as autonomous, and because they share the session id with the main turns
the whole session was filed under Autonomous rather than under a coding agent.
Kong does not hit this because its plugin config names the client and agent on
every call.

The User-Agent (claude-cli/...) is on every call including the sidecars, so the
plugin now reads it and sends x-s6r-client: claude plus, when the route names no
agent, x-s6r-agent: "Claude (LiteLLM)". A client-supplied x-s6r-agent or the
agent_ref config still wins. Verified live on tenant 123: a real Claude Code
session now lands as one coding_agent labelled "Claude (LiteLLM)" with its turns
scored, where before it split across Autonomous.

Identity: the key's own user (email then id) now outranks the end user the
request named. LiteLLM resolves Claude Code's hashed metadata.user_id as the end
user when nothing better is set, so a per-user key was being shadowed by a
session token. The key is the authenticated principal, the way a Kong consumer
is, so it wins; the request end user is the fallback.

* refactor(straiker): build the v3 request, envelope and headers as frozen mappings

The v3 builders seeded dicts and grew them, which the type-discipline gate
counts as mutable accumulators. Each is now one expression over a tuple of
pairs, frozen with MappingProxyType, and the JSON encoder unwraps a frozen
mapping through a default. The session seed and the verdict parser no longer
rebind locals. The wire is unchanged: 36 live calls through the proxy on this
commit carry the same fields, shapes, headers and identities as before, with
no mappingproxy text in any body.

* fix(straiker): satisfy basedpyright on the v3 builders

The frozen-mapping refactor left a shadowed headers local, a Mapping handed to
an HTTP client that takes a dict, an unguarded optional response, a turn id
typed object, and a redundant isinstance on already-typed texts. No behaviour
change: 4 live calls (chat, Messages, Bedrock, injection) return 200 with the
expected verdicts on this commit.

* fix(straiker): type the v3 config fields at the initializer and keep the verbose log as JSON

The four v3 routing fields (api_version, agent_ref, client, format_hint)
travelled through the untyped kwargs passthrough, which basedpyright counts
against the budget. They are now validated through a small Pydantic model at
the initializer and passed by name.

The verbose log serialized the frozen payload with default=str, which printed
a Python repr instead of JSON once the builders returned MappingProxyType.
Every serializer now unwraps a frozen mapping first. A test asserts the logged
payload parses as JSON and carries the identity; mutating the log site back to
default=str fails it.

* fix(straiker): address review findings on the v3 relay

Text completions relay their prompt: `prompt`, `suffix`, `echo` and `best_of`
join the provider allowlist, so /v1/completions traffic is screened.

The route's `agent_ref` now outranks the caller's `x-s6r-agent` header. The
header is caller-supplied, and letting it beat a pinned route would let any key
file its traffic under another application's agent and controls. On a route
that names nothing the header still names the application, which is how
several applications enumerate behind one key.

Credentials inside `tools` and `mcp_servers` (an OpenAI `mcp` tool's `headers`,
Anthropic's `authorization_token`) are replaced with `[redacted]` before the
body leaves the proxy, on both phases and in the verbose log. Detection reads
tool names, descriptions and schemas, never these.

A 200 whose body is valid JSON but not an object now reports an invalid
schema and follows the failure policy instead of raising out of the hook.

Comments that restated a constant are gone. Tests cover each change and the
failure paths (unreadable error body, client exceptions, missing response,
unmodellable request, session seeds from Anthropic block shapes); every fix
fails its test when reverted.

* fix(straiker): scrub tool credentials one level deep, without recursion

* fix(straiker): scrub only the fields that carry a credential, never a schema

The credential set is now the three fields that actually hold one on a tools
or mcp_servers entry (headers, authorization, authorization_token), read one
level deep. A function tool whose parameter schema defines a token, headers or
api_key property is relayed exactly as sent; a test pins that, and fails
against the recursive version.

* test(straiker): use example.com identities; drop a comment that restated its branch

* fix(straiker): present a legacy completion as the chat exchange it is

Straiker scores chat on both phases of a gateway turn but has no reader for a
text_completion answer: the request phase of a /v1/completions call was
scored and the response phase was refused with 501, whether or not the call
named an agent. A completion is one user turn and one assistant turn, so both
phases now present that exchange: the prompt becomes the single user message
and the TextCompletionResponse becomes a chat completion. Measured through the
proxy on this commit, both phases return 200 and score, and the derived
session is shared between them.

The derived session seed accepts the tuple the conversion produces; the test
pins the session on both phases and fails against the list-only check. The
unreachable "parsed is None" branch is folded into the failure branch, and a
malformed tools value is shown to relay as sent.

* fix(straiker): screen a completions prompt as the text the model receives

LiteLLM's /v1/completions accepts a string, a list of strings, a list of
token ids or a list of token-id lists, and decodes token ids with the
text-davinci-003 tokenizer before calling the model. The relay now renders
the prompt the same way, one user message per prompt, so a pre-tokenized
prompt is screened as the text it stands for rather than as digit strings.
A prompt in a shape this cannot render (empty, mixed, or with no tokenizer
available) is relayed untouched instead of being replaced with something
else. Tests cover all four accepted shapes and six unrenderable ones.

* fix(straiker): seed the derived session on the preamble and the first user turn

An OpenAI chat body carries its system prompt as messages[0], and the derived
session seeded on the Anthropic `system` field plus messages[0] with no role
check. For that shape the seed was the system prompt twice and the first user
turn never counted, so every unnamed conversation behind one system prompt
collapsed into one Straiker session. The seed now takes the preamble from
wherever the API puts it (`system`, `instructions`, or a leading system or
developer message) and the first message with role `user`, else a Responses
`input` string, else `prompt`. Two conversations sharing a system prompt are
two sessions again; a replayed conversation stays one.

* fix(straiker): seed the derived session on every text block of the first turn

A user turn that opens with an image or a document block and carries its
text later seeded the session on an empty string, so two different
conversations under the same preamble shared one Straiker session. Read
every text block of the turn instead of only the first block. A plain
string or a single text block seeds exactly as before.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(straiker): cover the tokenizer fallback, a textless first turn and Responses instructions

Three branches of the v3 relay had no test: a token-id prompt relayed as
sent when the tokenizer cannot be fetched, a first user turn with no text
seeding the session on the preamble alone, and a Responses API body
seeding on its instructions and first input turn. Each test fails when
its branch is mutated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): seed the derived session on the principal as well as the conversation

Straiker skips turns it has already scored for a session. The derived
session hashed the system prompt and the first user turn alone, so two
users who opened a conversation with the same words shared one session,
and the second user's copy of an attack came back as a replay: unscored
and allowed. Measured live on 2026-09-20: the first user's SSN turn was
blocked (`social_security_number`, scored=2), the second user's identical
turn was allowed (`controls: []`, replayed=2).

The principal now joins the seed. Explicit session ids, the Claude Code
header and LiteLLM's own session are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): derive the session id with sha256 and drop comments that restated constants

The derived session now hashes the principal, and CodeQL flags MD5 over an
identity as a weak hash on sensitive data. SHA-256 truncated to the same
32 hex characters keeps the id shape. Comments that only labelled the
allowlist groups or restated a constant are removed; the two that explain
a non-obvious choice stay.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): keep a blocked conversation blocked when it is replayed

Straiker de-duplicates turns it has already scored per session and
answers a replay `allow`, whatever the first verdict was. A client that
resends a blocked request, or grows the conversation past the blocked
turn, was let through: measured on 2026-09-20, `block` then `allow,
events_replayed=2` for the same session and body, and Claude Code's
automatic retry after the 400 turned a blocked poisoned-file read into
a pass.

The guardrail now remembers, per session, a fingerprint of every
conversation it blocked (a bounded, day-long in-memory cache) and blocks
a request that repeats or extends one without asking again. A different
session with the same words is a new conversation and is scored afresh.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): scope the block memory by session or principal, never by content alone

A request with no derivable session keyed the replay memory on the
conversation fingerprint alone, so one caller's block could answer
another caller's identical request. The memory is now scoped by the
session, else by the principal, and a request with neither is not
remembered at all.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(straiker): remember only a block that names a control, never one that comes from state

The replay memory kept every block, including one the platform returns
because a kill switch is engaged (`action: block` with `blocked_by: []`).
An administrator lifting the kill switch then left the conversation
refused by the remembered copy: measured on 2026-09-21, traffic stayed
blocked after `POST /inventory/agents/{id}/restore` returned `engaged:
false`.

The same words are the same attack tomorrow, so a control-named block is
still worth remembering; state is not ours to cache. The parsed verdict
now carries `blocked_by` so the two can be told apart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Phimmasone Phonpaseuth <PhimmStraiker@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 12:21:09 -07:00
..
a2a_protocol fix(a2a): send message/stream for Bedrock AgentCore streaming requests 2026-09-21 13:39:32 +00:00
batches test: migrate wave 1 phase 1 legacy tests to tests/unit 2026-09-20 09:17:57 +00:00
caching fix(caching): keep embedding cache hits aligned with request inputs (#42571) 2026-09-22 17:14:44 -05:00
chat_completions test: migrate wave 1 phase 1 legacy tests to tests/unit 2026-09-20 09:17:57 +00:00
completion_extras fix(completion_extras): forward non-enum reasoning_effort through the Responses bridge instead of dropping it (#42452) 2026-09-23 11:04:07 -07:00
containers test: delete assertions that pin vendor cost map facts 2026-09-18 00:28:49 +00:00
endpoints test: migrate wave 1 phase 1 legacy tests to tests/unit 2026-09-20 09:17:57 +00:00
enterprise test: migrate wave 1 phase 1 legacy tests to tests/unit 2026-09-20 09:17:57 +00:00
expected_fine_tuning_api
expected_responses_api_request
experimental_mcp_client test: deflake fuzzy picker, breached-password HIBP, and MCP stdio timeout tests (rolling deflake 2026-09-22) (#42125) 2026-09-23 08:44:59 -07:00
fixtures/together_ai_sync feat(models): add daily Together AI model registry sync script and workflow 2026-08-25 13:12:06 -07:00
google_genai fix(google_genai): drop non-object tool parameters instead of forwarding them 2026-09-19 19:19:33 -07:00
images test(images): pin scalar-array edit params survive as repeated multipart fields 2026-08-24 12:49:54 -07:00
integrations feat(otel): emit gen_ai.conversation.id from the caller's session id on v2 LLM spans (#42486) 2026-09-23 00:49:29 -07:00
interactions refactor(interactions): remove expired use_legacy_interactions_schema shim 2026-09-17 20:14:07 +00:00
litellm_core_utils ci: add merge smoke checks workflow with loopback-only harness and 11 curated cases (#42709) 2026-09-23 11:01:08 -07:00
llms ci: add merge smoke checks workflow with loopback-only harness and 11 curated cases (#42709) 2026-09-23 11:01:08 -07:00
messages test: migrate phase 15 legacy tests to tests/unit 2026-09-20 11:55:43 +00:00
ocr test: migrate phase 15 legacy tests to tests/unit 2026-09-20 11:55:43 +00:00
passthrough test: migrate phase 15 legacy tests to tests/unit 2026-09-20 11:55:43 +00:00
proxy feat(guardrails): straiker guardrail speaks the v3 platform API (/api/v3/detect) (#41880) 2026-09-23 12:21:09 -07:00
rag test: migrate phase 15 legacy tests to tests/unit 2026-09-20 11:55:43 +00:00
rerank_api fix: answer get_api_base for github_copilot and chatgpt without running the login flow (#42602) 2026-09-22 17:23:44 -07:00
responses fix(completion_extras): forward non-enum reasoning_effort through the Responses bridge instead of dropping it (#42452) 2026-09-23 11:04:07 -07:00
router_strategy feat(router): add group-scoped priority routing strategy (#42378) 2026-09-22 13:09:38 -07:00
router_utils feat(cost-map): add Azure Foundry pricing for gpt-6-sol and gpt-6-luna (#42747) 2026-09-23 15:17:36 +00:00
rust_bridge feat(secrets): route secret resolution through native Rust backends (#42619) 2026-09-23 08:24:57 -07:00
secret_managers fix(aws_secret_manager_v2): restore secret scheduled for deletion instead of failing CreateSecret (#42454) 2026-09-22 14:12:20 -05:00
types feat: add configurable provider affinity header mapping (#41033) 2026-09-22 13:31:33 -07:00
vector_stores fix(vector_stores): keep config-defined vector stores listed and read-only (#42574) 2026-09-23 04:02:16 +00:00
videos test: remove phase 16 legacy test files from tests/test_litellm 2026-09-20 10:59:22 +00:00
__init__.py
conftest.py fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag 2026-09-16 18:23:14 -07:00
log.txt
readme.md
test_a2a_registry_lookup.py fix(a2a): Entra credentials own the chat route bearer over a stored api_key or authorization header 2026-09-17 16:09:36 -07:00
test_acompletion_session_reuse_e2e.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_add_deployment_no_master_key.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_aembedding_session_reuse_e2e.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_anthropic_beta_headers_filtering.py feat(router): native compact-to-fit across conversation APIs (#42074) 2026-09-21 22:52:29 -07:00
test_anthropic_skills_transformation.py
test_assert_ci_coverage.py test(integration): add MCP gateway coverage wave 1 with a dedicated mcp shard and proxy coverage artifact (#42711) 2026-09-23 07:48:46 -07:00
test_assert_workflow_dir_hygiene.py feat(ci): assert .github/workflows holds only workflows, correctly named (#37616) 2026-08-20 21:36:26 +00:00
test_audio_transcription_rust_bridge.py refactor(rust_bridge): group route modules into packages and split ocr into main and rust 2026-09-16 20:34:51 +00:00
test_auto_update_price_and_context_window_file.py feat(cost_map): derive source_revision from the loaded bytes instead of a _metadata stamp 2026-09-07 17:47:51 -07:00
test_azure_ad_token_credential_resolution.py test(router): cover s3_output_bucket_name surviving the trusted credential snapshot 2026-08-17 14:51:10 -07:00
test_azure_ai_grok_4_3_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_azure_ai_grok_4_6_model_metadata.py test: keep behavior tests that read the cost map for a later fixture rewrite 2026-09-18 04:52:06 +00:00
test_baseten_glm_5_3_model_metadata.py test: drop tests that pin provider-owned cost map values 2026-09-18 03:55:51 +00:00
test_batch_completion_models_all_responses.py
test_bedrock_marengo_embed_3_model_metadata.py test: drop tests that pin provider-owned cost map values 2026-09-18 03:55:51 +00:00
test_budget_ratchet_check.py test(ci): annotate new test locals as Final 2026-09-20 09:07:51 +00:00
test_chat_ui_responses_session.py
test_check_licenses.py fix(ci): retry transient PyPI license lookups 2026-08-23 09:23:35 +00:00
test_check_mcp_operation_boundary.py refactor(mcp): extract explicit operation context and dispatch 2026-09-21 12:24:15 -07:00
test_check_migrations_no_data_rewrites.py fix(migrations-check): read the table name past comments, ignore referential SET DEFAULT 2026-09-16 18:14:36 +00:00
test_check_py310_typing_imports.py fix: keep litellm importable on Python 3.10 and guard 3.11-only typing imports in CI (#39448) 2026-09-02 18:27:19 -07:00
test_check_test_quality.py test(ci): annotate new test locals as Final 2026-09-20 09:07:51 +00:00
test_check_type_discipline.py perf(ci): fan the budget checkers out across cores (#37784) 2026-08-22 22:44:58 -07:00
test_circleci_path_filter.py fix(mcp): preserve session expiry signals and scope dependency CI 2026-09-18 22:52:10 -07:00
test_circleci_rust_toolchain.py fix(ci): pin workflow toolchain dependencies 2026-09-02 12:16:25 -07:00
test_claude_fable_5_config.py test: keep the pinning-test removal free of unrelated reformatting 2026-09-18 04:27:28 +00:00
test_claude_opus_4_6_config.py test: keep the pinning-test removal free of unrelated reformatting 2026-09-18 04:27:28 +00:00
test_claude_opus_4_8_config.py test: keep the pinning-test removal free of unrelated reformatting 2026-09-18 04:27:28 +00:00
test_claude_opus_5_config.py feat(cost-map): add Claude Opus 5.5 for Vertex AI and Azure AI (#42599) 2026-09-22 21:48:07 +00:00
test_claude_sonnet_5_config.py test: keep the pinning-test removal free of unrelated reformatting 2026-09-18 04:27:28 +00:00
test_cloudflare_workers_ai_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_completion_timeout_resolution.py
test_component_entrypoint.py feat(proxy): share database connections across workers with an in-container pgbouncer (#39683) 2026-09-10 22:14:37 +00:00
test_compression.py
test_conftest.py test: trim the PROXY_BASE_URL fixture and regression docstrings 2026-08-19 00:56:37 -07:00
test_conftest_isolation.py test: roll back live router replay membership between tests (#36278) 2026-08-08 10:45:43 -07:00
test_constants.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_container_router.py
test_cost_calculation_log_level.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_cost_calculator.py fix(prices): add baseten/zai-org/GLM-5.3-Fast pricing (#42764) 2026-09-23 11:31:27 -07:00
test_cost_map_guard.py ci: skip cost map file checks on PRs that leave the cost map untouched (#42406) 2026-09-21 21:40:48 -07:00
test_count_tokens_public_api.py chore(cost-map): remove models past their deprecation date (#42435) 2026-09-22 21:19:26 +00:00
test_dashscope_image_generation.py test: keep the pinning-test removal free of unrelated reformatting 2026-09-18 04:27:28 +00:00
test_daybreak_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_deepseek_model_metadata.py test: drop tests that pin provider-owned cost map values 2026-09-18 03:55:51 +00:00
test_default_branch.py chore(ci): drop litellm_internal_staging and litellm_oss_staging references, main is the only trunk (#42745) 2026-09-23 08:14:11 -07:00
test_detect_changes.py perf(ci): gate the lint, MCP and dashboard jobs on the pull request's file list (#37559) 2026-08-19 18:32:21 -07:00
test_dockerfile_apk_repository.py fix(docker): add public Wolfi apk repo to runtime image (#39033) 2026-09-01 15:11:15 -07:00
test_dockerfile_bedrock_realtime_extra.py fix(bedrock): keep realtime SDK error range inside websocket close reason 2026-09-17 01:57:33 +00:00
test_dockerfile_non_root.py
test_drop_params_env_var.py fix(init): keep non-flag LITELLM_DROP_PARAMS values on with a warning 2026-09-07 22:13:23 -07:00
test_e2e_egress_sentinel.py ci(e2e): record the e2e suite weekly and replay it on weekdays with zero egress (#38163) 2026-08-24 23:49:03 -04:00
test_eager_tiktoken_load.py
test_env_key_doc_gate.py fix(ci): make the env-key doc gate see bare get_secret and get_secret_str reads (#35996) 2026-08-05 14:49:55 -07:00
test_exception_exports.py
test_exception_header_preservation.py fix(bedrock): keep x-amzn-RequestId on chat error responses (#40089) 2026-09-07 17:16:47 -07:00
test_exception_mapping_request_attribute.py
test_filter_out_litellm_params.py
test_fireworks_serverless_model_costs.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_gate_slot_lock.py ci: avoid duplicate default branch fetches 2026-09-07 15:28:01 -07:00
test_gemini_3_1_flash_lite_image_pricing.py test: drop gemini-3.1-flash-lite-image capability pins 2026-09-16 19:21:50 +00:00
test_gemini_tts_native_audio_pricing.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_get_blog_posts.py test(lint): ban blind pytest.raises(Exception) with ruff B017 (#37731) 2026-08-20 18:09:42 -07:00
test_git_hooks.py chore(ci): drop litellm_internal_staging and litellm_oss_staging references, main is the only trunk (#42745) 2026-09-23 08:14:11 -07:00
test_gpt_5_4_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_gpt_5_5_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_gpt_image_cost_calculator.py chore(cost-map): remove models past their deprecation date (#42435) 2026-09-22 21:19:26 +00:00
test_gpt_realtime_mode.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_groq_streaming_encoding.py
test_guardrail_exception_status_codes.py
test_lazy_imports.py perf: defer fastapi and tiktoken BPE imports out of import litellm 2026-09-17 08:59:48 +00:00
test_lint_workflow_diff_gates.py ci(lint): gate top-level tests/e2e and litellm files in the diff-scoped lint steps 2026-09-01 22:30:46 -07:00
test_litellm_params_reserved_keys.py
test_logging.py feat(logger): dispatch Python logging through the Rust diagnostics processor (#42616) 2026-09-22 18:44:15 -07:00
test_lowest_latency_zero_tokens.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_main.py fix(edenai): restore unrelated files the squashed PR commit had reverted 2026-09-21 19:42:37 +00:00
test_main_module_header.py
test_mistral_medium_3_5_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_mistral_small_4_0_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_mistral_zai_glm_5_2_model_metadata.py test: keep behavior tests that read the cost map for a later fixture rewrite 2026-09-18 04:52:06 +00:00
test_model_block_unblock.py fix(proxy): serialize model reconciles so concurrent model writes stop evicting each other (#36687) 2026-08-12 13:42:26 -07:00
test_model_cost_aliases.py
test_model_param_helper.py
test_model_prices_schema.py fix(cost): bill DeepSeek V4.1 Flash and V4 Pro at their off-peak rates outside peak hours 2026-09-19 04:31:55 -07:00
test_model_response_normalization.py
test_muse_spark_1_1_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_muse_spark_1_2_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_muse_spark_1_3_model_metadata.py test: delete assertions that pin vendor cost map facts 2026-09-18 00:28:49 +00:00
test_mutation_report.py fix(ci): stop the mutation report publishing a score it never measured (#37825) 2026-08-21 20:15:52 -07:00
test_nested_drop_params.py
test_non_chat_routes_open_llm_spans.py fix(otel): emit LLM Call spans for speech, image, moderation, ocr and transcription (#37752) 2026-08-22 11:11:21 -07:00
test_openai_embedding_encoding_format_default.py test(embeddings): move legacy intercepts to the wire for the omitted-format path 2026-08-29 12:04:44 -07:00
test_openai_service_tier_long_context_pricing.py feat(openai): add GPT-6 Sol and GPT-6 Luna (#42515) 2026-09-22 11:34:01 -07:00
test_pre_commit_lint.py chore(ci): drop litellm_internal_staging and litellm_oss_staging references, main is the only trunk (#42745) 2026-09-23 08:14:11 -07:00
test_prisma_generate_if_needed.py fix(lint): generate the prisma client into the gate-owned venv 2026-08-06 01:54:26 -07:00
test_process_helpers.py test: count a zombie grandchild as gone in the migrate deploy timeout test (#42570) 2026-09-22 14:59:26 -07:00
test_project_alias_tracking.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_project_tags_pydantic.py test(lint): ban blind pytest.raises(Exception) with ruff B017 (#37731) 2026-08-20 18:09:42 -07:00
test_proxy_auth.py test: run the 30 test files stranded in the second mirror (#37595) 2026-08-20 10:59:43 -07:00
test_rag_openai_ingestion.py
test_rate_limit_error_unification.py feat(proxy): add budget_exceeded_status_code setting to restore 429 for budget refusals 2026-09-20 07:00:10 +00:00
test_redact_string_in_error_paths.py test(realtime): drop legacy InvalidStatusCode tests and pin websockets imports (#42624) 2026-09-22 17:42:12 -07:00
test_redis.py test(redis): pin ElastiCache IAM signing and TLS coercion invariants 2026-09-10 10:13:03 -04:00
test_redis_credential_provider.py test(redis): pin ElastiCache IAM signing and TLS coercion invariants 2026-09-10 10:13:03 -04:00
test_register_model_custom_pricing.py fix(cost): bill off-peak rates for deployments that set only off_peak_pricing 2026-09-01 12:14:46 -07:00
test_register_model_zero_cost_persistence.py
test_replicate_model_key_format.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_responses_api_bridge_non_stream.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_responses_id_security.py fix(proxy): authorize every Responses API id, not only the ones the proxy issued (#39548) 2026-09-11 11:47:05 -07:00
test_responses_streaming_container_ownership.py
test_retrieve_batch_bedrock_dispatch.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_router.py chore(cost-map): remove models past their deprecation date (#42435) 2026-09-22 21:19:26 +00:00
test_router_block_helpers.py
test_router_exception_redaction.py fix(router): explain fallback outcome in plain words in the raised error (#42509) 2026-09-22 18:37:42 +00:00
test_router_google_genai.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_router_model_cost_isolation.py fix(router): preserve discovered limits and model info fallbacks 2026-09-17 00:23:50 +00:00
test_router_order_fallback.py fix(router): skip the refusing deployment when retrying a non-transient error 2026-09-05 22:25:13 -07:00
test_router_per_deployment_num_retries.py refactor(router): resolve retry policy by exception MRO and add DefaultRetries 2026-09-04 16:09:01 -07:00
test_router_redis_init.py
test_router_retry_backoff_headers.py test: run the 30 test files stranded in the second mirror (#37595) 2026-08-20 10:59:43 -07:00
test_router_retry_non_retryable_errors.py fix(timing): union provider timing windows and anchor detailed pre-processing at receive time 2026-09-19 00:48:17 +00:00
test_router_retry_policy_update.py fix(proxy): persist only the router settings keys the request set 2026-09-18 14:01:05 -07:00
test_router_silent_experiment.py fix(router): snapshot shadow kwargs per target so concurrent shadows never share metadata 2026-09-16 04:31:43 +00:00
test_router_streaming_fallback_metadata.py
test_router_weighted_failover.py fix(router): skip the refusing deployment when retrying a non-transient error 2026-09-05 22:25:13 -07:00
test_ruff_strict_gate.py fix: address cross-version CI failures 2026-09-02 14:17:19 -07:00
test_sambanova_model_metadata.py test: keep behavior tests that read the cost map for a later fixture rewrite 2026-09-18 04:52:06 +00:00
test_secret_redaction.py feat(logger): dispatch Python logging through the Rust diagnostics processor (#42616) 2026-09-22 18:44:15 -07:00
test_select_ui_test_scope.py chore(ci): drop litellm_internal_staging and litellm_oss_staging references, main is the only trunk (#42745) 2026-09-23 08:14:11 -07:00
test_service_logger.py
test_setup_wizard.py
test_shared_session_integration.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_ssl_verify_unit.py refactor(bedrock): remove the dead BedrockLLM invoke code path 2026-07-29 20:25:36 -07:00
test_stream_chunk_builder_annotations.py
test_stream_chunk_builder_citations.py fix(streaming): join block-list citation deltas without extra nesting 2026-08-28 13:56:07 -07:00
test_stream_chunk_builder_images.py test: run the 30 test files stranded in the second mirror (#37595) 2026-08-20 10:59:43 -07:00
test_streaming_connection_cleanup.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_sync_together_ai_models.py feat(cost_map): derive source_revision from the loaded bytes instead of a _metadata stamp 2026-09-07 17:47:51 -07:00
test_system_message_format_bug.py
test_test_quality_gate.py ci(tests): wire tests/unit into CircleCI and drain legacy unit shards green 2026-09-20 07:05:42 +00:00
test_thinking_enabled.py test: drop restating comment and wrap long call in thinking tests 2026-08-18 19:55:22 -07:00
test_together_ai_model_metadata.py chore(cost-map): remove models past their deprecation date (#42435) 2026-09-22 21:19:26 +00:00
test_type_check_gate.py fix(lint): retire the single-slot base-counts cache 2026-08-06 02:23:52 -07:00
test_type_discipline_gate.py fix(lint): pick the merge-aware base so in-progress merges are not blamed for base drift 2026-08-04 17:58:41 -07:00
test_typesafe_model_metadata.py feat(proxy): add TypeSafe Jev passthrough spend tracking 2026-09-17 15:53:19 +00:00
test_unit_shard_missing_paths.py ci(test-unit): drop dead misc shard paths and skip missing paths with a warning (#42603) 2026-09-23 00:56:32 +00:00
test_unit_shard_per_test_timeout.py test(ci): drop the structure-only assertion on the shard script; the parametrized hang test covers both invocations 2026-09-19 03:20:48 -07:00
test_utils.py feat(gemini): add gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts prices (#42752) 2026-09-23 09:28:09 -07:00
test_utils_module_docstring.py
test_uuid_helper.py
test_vcr_safe_body_matcher.py
test_vertex_ai_xai_grok_prompt_caching_metadata.py test(pricing): assert cache-priced vertex grok rows advertise supports_prompt_caching (#41526) 2026-09-21 20:58:11 -07:00
test_video_generation.py test: delete assertions that pin vendor cost map facts 2026-09-18 00:28:49 +00:00
test_with_dashboard_node.py fix(bootstrap): fail fast when nvm cannot activate the pinned node 2026-08-04 21:18:58 -07:00
test_xai_grok_4_3_model_metadata.py test: keep tests that survive correct cost-map updates 2026-09-15 22:16:08 +00:00
test_xai_responses_auto_routing.py chore(cost-map): remove models past their deprecation date (#42435) 2026-09-22 21:19:26 +00:00

Testing for litellm/

This directory 1:1 maps the the litellm/ directory, and can only contain mocked tests.

The point of this is to:

  1. Increase test coverage of litellm/
  2. Make it easy for contributors to add tests for the litellm/ package and easily run tests without needing LLM API keys.

File name conventions

  • litellm/proxy/test_caching_routes.py maps to litellm/proxy/caching_routes.py
  • test_<filename>.py maps to litellm/<filename>.py