mirror of
https://github.com/BerriAI/litellm.git
synced 2026-08-28 05:25:59 +00:00
The ComplexityRouter's LLM classifier saw only the last user message, so on a multi-turn conversation it classified whatever happened to be last rather than what the human actually asked, and a near-constant classifier input pinned a whole session to one tier. The blindness turned out to be narrower than first diagnosed, and the fix is correspondingly smaller. Tool output was never the problem: on the Messages surface it rides a user turn as tool_result content blocks, which are not text parts, so flattening to `type == "text"` already dropped those turns; on chat completions it arrives on a `tool` role the extractor never read. Both surfaces were already handled before this change. What actually leaked through was the harness `<system-reminder>` block, which arrives as ordinary text, survives flattening, and became the current ask on any turn that carried one. So reminders are stripped rather than used to reject the turn, because a harness injects them alongside the live ask and not as a turn of their own; rejecting the turn would lose the ask, and keeping the block would feed the classifier the near-constant boilerplate that flattens tier selection in the first place. An earlier revision of this change also pattern-matched serialized tool_result payloads. That check only ever fired on a hand-serialized string neither request surface produces, it was where every review finding in this PR lived, and it is deleted here; the tests now pin the real shapes instead of the synthetic one they were built on. The classifier call is split into a system role carrying the rubric plus the caller's own system prompt, which stays byte-stable across a session so a provider can prompt-cache it, and a user role carrying the variable context: a bounded window of prior user turns, a conversation-depth signal, and the current ask. The caller's system prompt rides every turn, so task constraints are never dropped. The depth signal measures content-parts messages too, since counting only string content reported ~0 tokens for exactly the deep Messages-surface conversations that most need an expensive tier, and it is omitted entirely on the prompt-only path rather than asserting a false zero. Prior turns are excluded by matching the current ask rather than by dropping the newest turn positionally, because `aclassify` takes `prompt` and `messages` separately and a caller may classify something other than the newest turn. Truncated turns carry a marker so the classifier can tell a turn was clipped. Only the LLM classifier's input changes. The heuristic scorer, keyword overrides, escalation matching and semantic embedding still read the extracted current ask, which is why that extraction has to yield one clean human-authored string: those are substring and vector matchers, and an escalation keyword sitting inside a reminder blob would otherwise trip a tier jump on its own. Defaults keep single-turn classification equivalent to before. The prior-turn window is on by default so existing LLM-classifier deployments actually get the fix; the config field documents that those turns reach the classifier model, which may be a different provider than the routed completion model, and that the call already carries the current ask and the caller's system prompt in full. Scoped to the ComplexityRouter; the semantic AutoRouter is not touched. |
||
|---|---|---|
| .. | ||
| agent_tests | ||
| audio_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| e2e | ||
| enterprise | ||
| guardrails_tests | ||
| image_gen_tests | ||
| integration | ||
| litellm | ||
| litellm-proxy-extras | ||
| litellm_core_utils | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| old_proxy_tests/tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_migration_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| scim_tests | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| unified_google_tests | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _fake_openai_endpoint_server.py | ||
| _flush_vcr_cache.py | ||
| _live_test_helpers.py | ||
| _openai_record_replay_proxy.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| _ws_vcr.py | ||
| eval_swe_bench.py | ||
| fake_openai_endpoint.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| pyrightconfig.json | ||
| README.MD | ||
| test_anthropic_compaction_usage.py | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_config.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_entrypoint.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_organizations.py | ||
| test_otel_thread_leak.py | ||
| test_passthrough_endpoints.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/test_litellm
This folder can only run mock tests.