Existing tests pinned exact kwargs on `PrismaManager.setup_database`,
but the opt-in v2 resolver added `use_v2_resolver=False` to every call.
Update the three assertions to reflect the new signature.
Fixes:
- TestHealthAppFactory::test_use_prisma_db_push_flag_behavior
- TestHealthAppFactory::test_startup_fails_when_db_setup_fails
Adds end-to-end CI coverage for `--use_v2_migration_resolver` via a new
job `installing_litellm_on_python_v2_migration_resolver`:
- Clones the pytest smoke path from `installing_litellm_on_python` but
uses a local Postgres sidecar instead of the shared DB to prevent
collisions with the v1 variant.
- Runs only the new `test_litellm_proxy_server_config_no_general_settings_v2_resolver`
which spawns the proxy with `--use_v2_migration_resolver` and smoke-tests
`/health/liveliness` and `/chat/completions`.
Refactors `test_basic_python_version.py`:
- Extracts the proxy spawn + smoke-test body into `_run_proxy_server_smoke_test`
so the v1 and v2 tests share the same code path.
- The existing `test_litellm_proxy_server_config_no_general_settings` is
now a thin wrapper that passes no extra args (v1 default, unchanged).
- Adds `..._v2_resolver` variant that passes `--use_v2_migration_resolver`.
The existing `installing_litellm_on_python` / `installing_litellm_on_python_3_13`
jobs filter out the v2 variant via `-k "not v2_resolver"` so they keep
running only against their shared DB, unchanged behavior.
Default behavior (v1) is unchanged. Users who have seen schema thrashing
during rolling deploys can opt into the v2 resolver with
`--use_v2_migration_resolver`.
Why v2 is safer:
- Runs `prisma migrate deploy` only.
- Recovers from P3005 (baseline) and idempotent P3009/P3018 errors, same
as v1.
- Never calls `_resolve_all_migrations`, which generates a schema diff
between the live DB and the shipped schema.prisma and applies it via
`prisma db execute`. That path bypassed every migration's SQL and was
the root cause of thrashing when two LiteLLM versions contended for
the same DB.
- Logs a non-blocking warning when the DB has migrations applied that
are newer than anything this build ships (ahead-of-HEAD). It does not
refuse to start — many users have unusual ledger state from past
thrashing, and blocking startup would be a breaking change.
Also prints a message on startup when the default (v1) resolver is in
use, pointing operators at the opt-in flag.
Adds unit tests covering the v2 fail-fast paths, the stripping of
Prisma-specific query params from DATABASE_URL (needed for psycopg),
the timestamp helpers, and pins the default: v1 still invokes
`_resolve_all_migrations`, v2 must not.
Adds total_spend column to LiteLLM_TeamMembership that accumulates
continuously and is not zeroed by the budget cycle reset job. This
enables UI surfaces to distinguish current-cycle spend (the existing
spend column, which resets) from lifetime spend per team member.
Also exposes budget_reset_at on LiteLLM_BudgetTable so /team/info
callers can see when a member's budget window next resets. The field
was already stored in the DB but stripped by the response Pydantic
model.
Includes regression tests that:
- Guard the reset job against ever writing total_spend: 0
- Verify the spend writer increments both spend and total_spend in
one UPDATE statement.
- _find_destructive_statements: add DROP INDEX to the docstring (the
regex already detects it; only the docstring lagged).
- create_migration: correct the base_branch default documented in the
docstring from "main" to "litellm_internal_staging".
Generating a migration from a stale branch could silently emit DROP
COLUMN for columns the stale branch did not know about, and the
script would write that SQL to a new migration file with no warning.
Adds two guards to ci_cd/run_migration.py:
- Branch freshness check: fetches origin/<base-branch> and exits 3 if
HEAD is behind. Default base is litellm_internal_staging. New
flags: --base-branch, --skip-freshness-check.
- Destructive guard: refuses (exit 2) if the generated diff contains
DROP COLUMN / DROP TABLE / DROP INDEX, unless --allow-destructive
is passed.
Refusal banners include guidance and an explicit callout instructing
AI agents not to auto-bypass the flags. Also treats Prisma's
"-- This is an empty migration." output as a no-op rather than
writing an empty file.
Updates litellm-proxy-extras/migration_runbook.md with the new
workflow, flag documentation, and agent warnings.
Two fail-safes for the /v1/messages → Bedrock Invoke pass-through so new
Anthropic-only extensions Claude Code starts sending can't reach Bedrock
and trigger a 400 "Extra inputs are not permitted":
1. Top-level body fields are filtered to a typed allowlist. New
`BedrockInvokeAnthropicMessagesRequest` TypedDict (in
`litellm/types/llms/bedrock.py`) captures the Bedrock Invoke Anthropic
Messages body schema; the runtime allowlist is derived from its
`__annotations__` so the type and the filter can't drift. Anchored to
the AWS reference page in docstrings + transform comment. An
exact-set test pins the resolved allowlist so any future edit forces
conscious review.
Drops context_management, output_config, speed, mcp_servers,
container, inference_geo, internal litellm_metadata, and any future
Anthropic addition. output_format stays as an active inline-schema
conversion (not just a strip).
2. The anthropic-beta header list is filtered + transformed against the
bedrock mapping for ALL betas, not just auto-injected ones. The
previous code union'd user-provided betas back in unfiltered, so a
client on a new Anthropic-direct beta (e.g. advisor-tool-…,
context-management-…) could still pin the request to fail. In a proxy
context the client can't know the backend is Bedrock; the provider
mapping is authoritative. User-provided drops are logged at WARNING
so intentional overrides leave a breadcrumb.
Updates one existing test that happened to assert on the old buggy
pass-through (it used output-128k-2025-02-19, which is null in the
bedrock mapping and would 400 at runtime); rewrote it against a
bedrock-supported beta.
Scope: messages/invoke only. The same user-beta bypass exists in
chat/invoke but that's a different code path with different
user-expectation trade-offs — follow-up.
Brings user personal budget and organization budget enforcement
in line with the existing key and team patterns, which already
read spend from the atomic cross-pod Redis counter.
Add a router utils unit test that directly exercises _filter_deployments_by_model_access_groups for access-group-only key permissions.
Made-with: Cursor
Filter router candidate deployments by caller-authorized model access groups when access is granted via group membership, preventing cross-group load balancing for shared public model names.
Made-with: Cursor
VertexAIImagenImageEditConfig.get_complete_url was resolving vertex_project
and vertex_location only from env vars and global settings, ignoring
litellm_params. Users supplying project/location exclusively via YAML
config would get a ValueError or wrong URL even after auth headers were fixed.
Mirrors the pattern already used by VertexAIGeminiImageEditConfig and
image_generation counterpart (safe_get_vertex_ai_project/location).
Also fixes api_key type hint in MockImageEditConfig (str -> Optional[str])
and adds a test covering get_complete_url credential resolution.
Made-with: Cursor
Adds three test cases to prevent regression of the Vertex AI image_edit
credentials bug:
1. test_validate_environment_signature_includes_litellm_params: ensures
all image-edit configs accept litellm_params (contract for the handler)
2. test_vertex_gemini_image_edit_reads_credentials_from_litellm_params:
verifies Gemini config reads from litellm_params first
3. test_vertex_imagen_image_edit_reads_credentials_from_litellm_params:
verifies Imagen config reads from litellm_params first
These tests catch if the fix is accidentally reverted or if new image-edit
configs are added without the litellm_params parameter.
Made-with: Cursor
When aimage_edit or image_edit was called with Vertex AI Gemini/Imagen models
via YAML-style config (vertex_project / vertex_credentials in proxy YAML),
the credentials were dropped during handler-to-config plumbing, causing
fallback to Application Default Credentials and DefaultCredentialsError.
Root cause: image_edit_handler and async_image_edit_handler did not pass
litellm_params to validate_environment, unlike image_generation_handler.
Fixes:
1. Widen BaseImageEditConfig.validate_environment signature to accept
litellm_params and api_base (optional kwargs).
2. Forward dict(litellm_params) and litellm_params.api_base from both
sync and async image_edit handlers to validate_environment.
3. Update VertexAIImagenImageEditConfig.validate_environment to read
vertex_ai_project/vertex_ai_credentials from litellm_params first,
matching Gemini config pattern (secondary latent bug fix).
4. Widen all image-edit config override signatures to match base.
Made-with: Cursor
Addresses review feedback on the snapshot approach:
1. Class-instance mutable state
The snapshot only covers primitives + collections + None. Class
instances (DualCache, LLMClientCache) weren't reset between tests,
so in-place cache mutations could leak. Can't deepcopy these — they
hold thread locks — but they expose flush_cache(). Collect every
module attribute whose value implements flush_cache() at conftest
import, and invoke it per-test alongside the snapshot restore.
2. Silent skips are now warnings
_snapshot_mutable_state and _restore_mutable_state previously
swallowed exceptions, so if a future attr gained a property without
a setter (or other non-round-trippable state), an isolation gap
would have no signal. Emit warnings.warn on each failure path.
3. Docstring
Explicitly documents what IS and IS NOT reset, and tells authors to
use monkeypatch.setattr() for in-place mutations of instances
without flush_cache() (ProxyLogging, JWTHandler, etc.).
The Build UI from source step used:
cp -r out/ ../../litellm/proxy/_experimental/out/
GNU cp (CircleCI's Ubuntu image, coreutils 8.32) interprets this as
copy the source directory as a CHILD of the destination when the
destination already exists — so the command silently created
litellm/proxy/_experimental/out/out/ instead of replacing the served
bundle at litellm/proxy/_experimental/out/*.
The proxy continued serving whatever bundle was checked in, so every
e2e_ui_testing run between this job's introduction (d09d98a70a,
2026-04-08) and the bundle-rebuild commit (de790fd273, 2026-04-18) was
effectively testing a STALE bundle — not the fresh build. That is why
the double-prefix regression (NEXT_PUBLIC_BASE_URL="ui/" combined with
networking.tsx reading the env var) was never caught in CI even though
the source contained the trigger the whole time: the bundle the proxy
served never picked up the source change.
Replace cp -r with rm + mv so the destination is cleanly swapped.
Verified end-to-end on an Ubuntu 22.04 / GNU coreutils 8.32 container:
- Before fix: fresh build has 9 "ui/" literals in chunks; after cp,
_experimental/out/* still has 0 (stale); _experimental/out/out/ is a
nested dir the proxy does not serve.
- After fix: _experimental/out/* has 9 "ui/" literals — the proxy now
serves the freshly built (broken, in this repro) bundle, so
globalSetup fails at login and every spec is blocked. Removing the
bug from .env.production and rebuilding brings the count back to 0
and the suite passes.
No spec changes, no fixtures, no new infrastructure. The existing
Playwright suite already catches this class of regression via the
login flow in globalSetup; it just needs the CI to actually hand it
the freshly built bundle.
The previous snapshot only tracked list/dict/set values. Tests mutate
scalar module attrs too — master_key, premium_user, prisma_client — and
importlib.reload used to reset those implicitly. Under the snapshot
approach they were leaking between tests, so test_active_callbacks
failed in CI with "No api key passed in." once an earlier test left
master_key set to sk-1234.
Expand the snapshot to cover primitives (str/int/float/bool/bytes/tuple)
and None-valued attributes. Complex object instances are still skipped
to avoid deepcopy issues.
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Has been cancelled
Add regression tests that mock make_bedrock_api_request and verify
input_type=request uses source=INPUT with user messages, and
input_type=response uses source=OUTPUT with synthetic ModelResponse.
Made-with: Cursor
BedrockGuardrail.apply_guardrail hardcoded source="INPUT" regardless of the
input_type parameter. On the non-streaming post-call path (unified_guardrail
-> OpenAIChatCompletionsHandler.process_output_response -> apply_guardrail),
the model response text was sent to Bedrock as INPUT, so guardrail policies
configured for Output (e.g. PII/NAME blocking) returned action=NONE and the
response passed through unblocked. The streaming path was unaffected because
it calls make_bedrock_api_request(source="OUTPUT", ...) directly.
Map input_type to the correct Bedrock source ("request" -> INPUT,
"response" -> OUTPUT) and build a synthetic ModelResponse for the OUTPUT
path so _create_bedrock_output_content_request produces the correct payload.
Made-with: Cursor
tests/proxy_unit_tests/conftest.py was calling importlib.reload(litellm) in an
autouse function-scoped fixture, which cost ~17s per test because it re-ran
the full litellm __init__ import chain. With 400+ proxy unit tests, this was
the single biggest driver of CI wall time — 18 of the top 20 slowest durations
in a typical run were just the 17s fixture setup.
Replace the reload with a snapshot-and-restore approach: snapshot the mutable
lists/dicts/sets on litellm and litellm.proxy.proxy_server once at conftest
import, then deep-copy that snapshot back before each test. Callback lists,
caches, router state, etc. still get reset between tests, but the expensive
import chain only runs once per worker.
Local measurement on test_proxy_utils.py: 188 tests in 3.50s (previously took
~15 minutes of CI wall time on a single worker).
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Waiting to run
- _owner_cache now opportunistically sweeps expired entries past
_OWNER_CACHE_SWEEP_THRESHOLD live entries. Previously sessions that never
came back piled up forever.
- flush_session_to_db strips session_id/router_name/model_name from the update
payload. Prisma rejects writes to @@id fields.
- record_turn no longer persists last_user_content / last_assistant_content /
tool_call_history / pending_tool_calls. Those are needed only in-memory for
the next turn's signal detection; writing user prompts and tool payloads to
the DB would store PII for every conversation.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- SessionState now carries clean_credit_awarded + last_processed_turn (matching
the DB schema). Satisfaction only fires once per session AND only after
MIN_TURNS_FOR_CLEAN_CREDIT turns of context — early "thanks" no longer
inflates alpha.
- _detect_failure no longer treats empty content as failure. Many tools
legitimately return empty output (zero-result searches, silent bash);
penalizing those corrupted the bandit posterior. Only is_error fires now.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Waiting to run
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Waiting to run
The "remaining" proxy-db job was consistently timing out at ~98% because
--dist=loadscope pins every test in test_proxy_utils.py (168+ parametrized
tests) to a single xdist worker. 7 workers finished their files in ~15
minutes, then one worker ran alone for another 8+ minutes and hit the
30-minute job cap.
Give test_proxy_utils.py its own matrix entry so its tests spread across
all 8 workers, and add it to the "remaining" ignore list.
* feat(router): add auto_router/quality_router for quality-tier routing (#25987)
* feat(router): add auto_router/quality_router for quality-tier routing
Adds a new auto-router type that routes a request to a model at a target
quality tier. The quality tier is inferred by re-using the existing
ComplexityRouter's classification, then mapped through an admin-configured
complexity_to_quality table. Each candidate model declares its own
quality_tier in model_info.litellm_routing_preferences.
Resolution strategy: exact tier match, else round up to the next higher
tier, else fall back to default_model.
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): add capability-based filtering
Each deployment can declare a `capabilities: List[str]` field in
`model_info.litellm_routing_preferences` (e.g. ["vision",
"function_calling"]). Requests can pass `litellm_capabilities` in
`request_kwargs` to require specific capabilities — the router will only
route to deployments whose declared capabilities are a superset.
Resolution still walks tier (exact → round up), but at each tier filters
by capability before picking. Falls back to default_model only when it
also satisfies the required capabilities; otherwise raises rather than
silently routing to a model that lacks a required capability.
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): expose routing decision in response headers
For transparency, expose the QualityRouter's routing decision in the
proxy response headers:
x-litellm-quality-router-model → picked model_name (e.g. "haiku-vision")
x-litellm-quality-router-tier → resolved quality tier (e.g. "1")
x-litellm-quality-router-complexity → ComplexityTier name (e.g. "SIMPLE")
Mechanism: the pre-routing hook stashes the decision in
request_kwargs["metadata"]["quality_router_decision"]. After the call
returns, Router.set_response_headers lifts the decision into
response._hidden_params["additional_headers"] alongside the existing
x-litellm-model-group / x-litellm-model-id headers. Existing metadata
keys (trace_id, user_id, etc.) are preserved.
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): replace capabilities with keyword override
Drops the capability-based filtering in favor of a keyword-based override
for v0:
- RoutingPreferences.keywords: List[str] (replaces capabilities) — each
deployment can declare substring keywords.
- If any declared keyword (case-insensitive) appears in the user message,
the router short-circuits the complexity-classification flow and routes
to the matching deployment.
- Tiebreaker for overlapping keyword matches: quality_tier DESC, then
cheapest model_info.input_cost_per_token ASC. Unpriced models lose ties
to priced ones.
Decision metadata + headers now expose the override:
x-litellm-quality-router-via → "keyword" | "quality_tier"
x-litellm-quality-router-keyword → matched keyword (only on keyword route)
x-litellm-quality-router-complexity → complexity tier (only on tier route)
Removes:
- request_kwargs["litellm_capabilities"] reading
- _model_capabilities, _model_supports_capabilities,
_first_capable_model_at_tier, capability filter in
_resolve_model_for_quality_tier
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): add explicit `order` to RoutingPreferences
Adds an explicit priority field to RoutingPreferences for resolving
collisions deterministically:
RoutingPreferences.order: Optional[int] # lower wins; unset = +inf
Used as the PRIMARY tiebreaker in two places:
1. Keyword overlap: when multiple deployments declare the same matching
keyword, sort by (order ASC, quality_tier DESC, input_cost_per_token
ASC, model_name ASC). Explicit always beats implicit.
2. Tier resolution: when multiple deployments share a quality tier,
`_resolve_model_for_quality_tier` picks the one with the lowest
order. The tier list is now sorted at index-build time.
This lets admins make routing decisions explicit when the natural
quality-and-price ordering would pick the wrong model.
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* feat(quality_router): reorder tiebreak to (quality, order, price)
Changes the tiebreak ordering so quality_tier always wins first, then
explicit `order` is used to break ties within the same tier, then price
breaks the rest:
1. quality_tier DESC ← best model wins first
2. order ASC ← explicit priority within a tier
3. input_cost_per_token ASC
4. model_name ASC
Previously `order` was the primary key — that meant a tier-2 model with
`order=1` would beat a tier-3 model with no `order`, which is the wrong
default. Now `order` only resolves collisions among same-tier candidates.
Tier resolution (within a single tier) keeps the same key minus quality:
(order ASC, cost ASC, name).
Test renames + flips:
- test_explicit_order_overrides_quality_tier → test_quality_wins_over_explicit_order
- new: test_order_breaks_tie_within_same_quality_tier
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* fix(quality_router): resolve Greptile review feedback
Addresses four P1 findings from PR review plus test coverage:
1. set_model_list missing quality_routers reset
- Hot-reloading the Router would leave stale QualityRouter instances
pointing at the old model_list. `set_model_list` now clears
`self.quality_routers` alongside the other indices.
2. Round-down fallback before default_model
- `_resolve_model_for_quality_tier` now rounds DOWN to the closest
lower tier after round-up fails, before falling back to
`default_model`. Degrades gracefully rather than jumping straight
off-tier.
3. RoutingPreferences validation bypass
- `_build_tier_index` now instantiates `RoutingPreferences(**prefs)`
so invalid shapes (e.g. non-int quality_tier) raise a clear
ValueError instead of silently succeeding.
4. Config-ordering dependency
- `_tier_to_models` is now built lazily on first access. Previously,
eager construction in `__init__` meant a QualityRouter deployment
had to appear AFTER all its referenced models in config.yaml,
because `Router._create_deployment` populates `model_list`
incrementally. Any `available_models` defined after the router
entry would silently be reported as missing.
Also adds 6 new tests covering each fix:
- test_invalid_quality_tier_type_raises_clear_error
- test_router_can_be_instantiated_before_its_targets_exist
- test_set_model_list_clears_quality_routers_registry
- test_rounds_down_when_no_higher_tier_exists
- test_rounds_down_prefers_closest_lower_tier
- test_prefers_round_up_over_round_down
Co-Authored-By: Claude Opus 4 (1M context) <noreply@anthropic.com>
* style: apply black 24.10.0 formatting to pre-existing offenders
Unblocks the LiteLLM Linting check for this PR — these 12 files are already
failing `black --check` on main (the lint workflow only runs on PRs, so main
drifts). No behavior changes; formatting-only.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Update litellm/router.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Opus 4 (1M context) <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Support /v1/responses in complexity router (#26137)
* feat(proxy): add --reload flag for uvicorn hot reload (dev only)
Opt-in CLI flag, off by default, no env var. Only affects the uvicorn
run path; gunicorn/hypercorn paths and prod (which doesn't pass the
flag) are unaffected.
* Feature/add audio support for scaleway (#26110)
* feat(scaleway): add SCALEWAY to LlmProviders enum
* feat(scaleway): add audio transcription config and dispatch wiring
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(scaleway): add behavior tests for audio transcription config
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(scaleway): advertise audio_transcriptions in endpoint-support JSON
* docs(scaleway): document audio transcription support
* fix(scaleway): address PR review — plain-text response_format + missing-key fail-fast
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(scaleway): cover new response paths, drop gettysburg.wav coupling
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* Prompt Compression - add it to the proxy (#25729)
* refactor: new agentic loop event hook
simplifies how to create logic for tool based multi llm calls
* fix: compress - make it work on anthropic input as well
* fix(compress.py): working prompt compression for claude code
ensures claude code messages can run through proxy easily
* docs: add agentic loop hook guide
* docs: add agentic_loop_hook to sidebar
* fix: fix multiple arguments error
* fix: fix tool call loop for compression on streaming /v1/messages
* fix: fix linting errors
* fix: fix ci/cd errors
* feat(litellm_pre_call_utils.py): use claude code session for litellm session id
allows claude code logs to be stitched together, making it easy to know they were all part of the same conversation
* fix: suppress incorrect mypy warning rE: module
* revert: drop PR's changes to litellm/proxy/_experimental/out/
Restores the 34 HTML files under _experimental/out/ to their pre-PR
paths (X/index.html -> X.html). All renames are R100 (content
unchanged); no other files are touched.
* fix: address greptile review comments on PR #25729
- Skip ``kwargs["tools"] = []`` injection when compression is a no-op —
Anthropic Messages rejects empty tool arrays on requests that did not
originally declare tools.
- Move agentic-loop safety guards (fingerprint cycle / max depth) out of
the per-callback try/except so they propagate instead of being swallowed
by the generic exception handler. Extracted _check_agentic_loop_safety.
- Gate generic ``x-<vendor>-session-id`` capture behind the
LITELLM_CAPTURE_VENDOR_SESSION_HEADERS env var (off by default) to
preserve backwards compatibility; explicit x-litellm-* headers are
unaffected.
- Fix monkeypatch target in pre-call-hook test to patch the actual
module-level binding
(litellm.integrations.compression_interception.handler.compress).
- Add regression tests for empty-tools skip and opt-in session capture.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* revert: drop LITELLM_CAPTURE_VENDOR_SESSION_HEADERS flag
Generic x-<vendor>-session-id header capture is a new feature and only
runs *after* the explicit x-litellm-trace-id / x-litellm-session-id
checks, so it does not change behavior for any existing caller that was
already using the LiteLLM headers — no backwards-incompatibility to gate.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor(compress): replace input_type with CallTypes call_type
Drop the bespoke ``CompressionInputType`` literal and use the existing
``litellm.types.utils.CallTypes`` enum instead. ``litellm.compress()``
now takes ``call_type: Union[CallTypes, str]`` (default
``CallTypes.completion``) — no new concept to learn, and the enum is
already the way the rest of the codebase talks about request shapes.
Supported values: ``completion`` / ``acompletion`` (OpenAI chat-completions
shape) and ``anthropic_messages`` (Anthropic structured content blocks).
Updated: compress(), the compression_interception handler, tests, docs,
and the two eval scripts.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* Support /v1/responses in complexity router
Adds cross-format support to the complexity router via the guardrail
translation handler dispatch. Adds get_structured_messages to base
translation plus OpenAI chat, Responses, and Anthropic handlers.
Auto-router helper _extract_text_from_messages handles tool-call and
multimodal messages. Widens async_pre_routing_hook messages type to
Dict[str, Any].
Fixes https://github.com/BerriAI/litellm/issues/25134
* chore: apply black formatting
* fix: fallback to trying each handler when route inference fails
---------
Co-authored-by: Ryan Crabbe <ryan@berri.ai>
Co-authored-by: nhyy244 <106547304+nhyy244@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: cover _is_quality_router_deployment and init_quality_router_deployment
* fix: reset auto_routers on set_model_list to prevent hot-reload ValueError
* style: apply black formatting to websearch_interception and agentic_streaming_iterator
---------
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
Co-authored-by: Claude Opus 4 (1M context) <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Ryan Crabbe <ryan@berri.ai>
Co-authored-by: nhyy244 <106547304+nhyy244@users.noreply.github.com>
Run auth_ui_unit_tests against a per-job cimg/postgres:16.0 sidecar
with DATABASE_URL pointing at localhost:5432, matching the pattern
used by e2e_ui_testing. Seed the schema via 'litellm --skip_server_startup
--use_prisma_db_push' so each run starts on a clean DB with the current
schema.prisma.
Six tests in test_hooks.py were written against an older API and had been
failing in CI. Updated:
- test_resolve_session_key_* (4 tests): _resolve_session_key now requires
at least SIGNAL_GATE_MIN_MESSAGES messages before deriving a hash (it
returns None on shorter convos to match the signal-processing gate).
Switched the tests to use _long_messages() so they hit the hash path.
- test_post_call_success_hook_* (2 tests): the hook was migrated from
async_post_call_success_hook (mutates response._hidden_params) to
async_post_call_response_headers_hook (returns a headers dict) because
the former fires too late for streaming responses. Rewrote the tests
against the new API; added a metadata-not-dict noop case.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>