litellm/tests/e2e/router/test_complexity_router_e2e.py
devin-ai-integration[bot] fac43df9b9
fix(complexity_router): return empty dict from _classifier_call_metadata when metadata is absent (#33452)
* fix(complexity_router): return empty dict from _classifier_call_metadata when metadata is absent

The LLM classifier reads request_kwargs.get("litellm_metadata"), but the proxy stores request metadata under "metadata", so this returned None. _classifier_call_metadata then passed None straight through to the classifier acompletion call, which assumes a dict and blows up with 'NoneType' object has no attribute 'update'; the router swallowed it and silently fell back to heuristic scoring, so the configured LLM classifier never ran. Returning an empty dict keeps the classifier call well-formed.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): cover complexity-router LLM classifier routes over the proxy

Add a live e2e regression for the complexity auto-router: a lexically simple but hard prompt ("Is P equal to NP?") is routed by the LLM classifier to the higher-tier anthropic backend, read back from the spend log's model. Before the metadata fix the classifier silently crashed and the router fell back to heuristic SIMPLE scoring on the openai backend, so this test fails pre-fix and passes post-fix.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-15 17:46:00 -07:00

62 lines
3 KiB
Python

"""Live e2e: the v2 auto-router's LLM complexity classifier actually runs over the
proxy and drives routing, instead of silently crashing and falling back to the
local heuristic scorer.
The regression this guards (complexity_router.py `_classifier_call_metadata`
returning None when the request carries no `litellm_metadata`, which the classifier
sub-call then fed into a `.update`, raising `'NoneType' object has no attribute
'update'`) was invisible from the outside: the router caught the error and answered
from heuristic scoring, so every request still returned 200. The only tell is which
tier, and therefore which backend, served the request.
`complexity-smart-router` (see the inline config in docker-compose.yml) pins SIMPLE
to the openai backend and every higher tier to the anthropic backend. "Is P equal
to NP?" is lexically trivial, so the heuristic scorer lands it in SIMPLE (openai),
but any competent LLM classifier reads it as a hard reasoning question and lands it
above SIMPLE (anthropic). The served deployment is read back from the spend log's
`model`, so anthropic proves the classifier ran and openai proves it silently fell
back - the exact failure before the fix.
"""
import pytest
from complexity_router_client import ComplexityRouterClient
from e2e_http import unwrap
from models import ChatBody, ChatMessage
pytestmark = pytest.mark.e2e
ROUTER_MODEL = "complexity-smart-router"
# Lexically simple (heuristic -> SIMPLE) but a hard reasoning question (LLM -> above SIMPLE).
LEXICALLY_SIMPLE_HARD_PROMPT = "Is P equal to NP?"
# SIMPLE tier backend; served only when the classifier silently falls back to heuristic.
HEURISTIC_TIER_MODEL = "openai/gpt-5.5"
# MEDIUM/COMPLEX/REASONING tier backend; served only when the LLM classifier runs.
LLM_TIER_MODEL = "anthropic/claude-haiku-4-5"
class TestComplexityRouterLlmClassifier:
@pytest.mark.covers("reliability.routing.complexity_llm_classifier.routes_by_llm_tier")
def test_llm_classifier_runs_and_routes_by_semantic_tier(
self, client: ComplexityRouterClient, scoped_key: str
) -> None:
chat = unwrap(
client.gateway.chat(
scoped_key,
ChatBody(
model=ROUTER_MODEL,
messages=[ChatMessage(role="user", content=LEXICALLY_SIMPLE_HARD_PROMPT)],
max_tokens=16,
),
)
)
assert chat.choices, f"router returned no choices: {chat}"
rows = client.gateway.poll_logs_for_key(scoped_key, min_rows=1)
served = [row.model for row in rows]
assert served == [LLM_TIER_MODEL], (
f"expected the request to be served by {LLM_TIER_MODEL!r} (the higher-tier "
f"backend the LLM classifier picks for a hard prompt), but the spend log shows "
f"{served!r}. {HEURISTIC_TIER_MODEL!r} means the LLM classifier silently failed "
f"and the router fell back to heuristic scoring (SIMPLE) - the pre-fix regression"
)