mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-15 23:31:29 +00:00
* fix(router): rank streaming latency routing by raw TTFT, not TTFT per token Latency-based routing divided time-to-first-token by completion_tokens before storing it, so a deployment that streamed a long answer looked faster to first token than one that answered briefly. TTFT is now stored as plain seconds (first token time minus request start) in both the sync and async success handlers, which is what the routing decision compares. Non-streaming latency normalization per output token is unchanged. Claude-Session: https://claude.ai/code/session_01Ttd5Q9ZhRPB4ch5guos3rj * fix(router): store streaming TTFT under a seconds-only cache key Workers on the previous release keep writing seconds-per-token samples under "time_to_first_token" in the shared router cache during a rolling deploy, so mixing the new raw-seconds samples into the same list averaged incompatible units. Raw TTFT now lives under "time_to_first_token_seconds" and routing reads only that key. Also fix the regression test's token counts: with 50 tokens on the fast deployment and 500 on the slow one the old per-token formula picks the slow deployment, so the routing assertion now catches the bug. Claude-Session: https://claude.ai/code/session_01Ttd5Q9ZhRPB4ch5guos3rj * test(router): cover the TTFT sliding window from the unit-test shard Move the TTFT list trimming checks from the CircleCI-only suite into the mapped unit test file as one sync/async parametrized test, so the changed lines in lowest_latency.py are exercised by the GitHub unit-test shard that reports patch coverage. Claude-Session: https://claude.ai/code/session_01Ttd5Q9ZhRPB4ch5guos3rj |
||
|---|---|---|
| .. | ||
| agent_tests | ||
| audio_tests | ||
| base_sdk_tests | ||
| basic_proxy_startup_tests | ||
| batches_tests | ||
| benchmarks | ||
| code_coverage_tests | ||
| documentation_tests | ||
| e2e | ||
| enterprise | ||
| guardrails_tests | ||
| image_gen_tests | ||
| integration | ||
| litellm-proxy-extras | ||
| litellm_utils_tests | ||
| llm_responses_api_testing | ||
| llm_translation | ||
| load_tests | ||
| local_testing | ||
| logging_callback_tests | ||
| mcp_tests | ||
| multi_instance_e2e_tests | ||
| ocr_tests | ||
| openai_endpoints_tests | ||
| otel_tests | ||
| pass_through_tests | ||
| pass_through_unit_tests | ||
| proxy_admin_ui_tests | ||
| proxy_behavior | ||
| proxy_e2e_anthropic_messages_tests | ||
| proxy_migration_tests | ||
| proxy_security_tests | ||
| proxy_unit_tests | ||
| router_unit_tests | ||
| rust-python-harness | ||
| search_tests | ||
| spend_tracking_tests | ||
| store_model_in_db_tests | ||
| test_litellm | ||
| test_litellm_rust | ||
| unified_google_tests | ||
| vector_store_tests | ||
| windows_tests | ||
| __init__.py | ||
| _fake_openai_endpoint_server.py | ||
| _flush_vcr_cache.py | ||
| _live_test_helpers.py | ||
| _openai_record_replay_proxy.py | ||
| _vcr_conftest_common.py | ||
| _vcr_redis_persister.py | ||
| _wait_helpers.py | ||
| _ws_vcr.py | ||
| eval_swe_bench.py | ||
| fake_openai_endpoint.py | ||
| gettysburg.wav | ||
| large_text.py | ||
| openai_batch_completions.jsonl | ||
| pyrightconfig.json | ||
| README.MD | ||
| test_anthropic_compaction_usage.py | ||
| test_budget_management.py | ||
| test_callbacks_on_proxy.py | ||
| test_debug_warning.py | ||
| test_default_encoding_non_root.py | ||
| test_end_users.py | ||
| test_fallbacks.py | ||
| test_gpt5_azure_temperature_support.py | ||
| test_health.py | ||
| test_keys.py | ||
| test_litellm_proxy_responses_config.py | ||
| test_logging.conf | ||
| test_models.py | ||
| test_new_vector_store_endpoints.py | ||
| test_openai_endpoints.py | ||
| test_organizations.py | ||
| test_otel_thread_leak.py | ||
| test_presidio_latency.py | ||
| test_proxy_server_non_root.py | ||
| test_ratelimit.py | ||
| test_resource_cleanup.py | ||
| test_rust_python_harness.py | ||
| test_service_logger_otel.py | ||
| test_spend_logs.py | ||
| test_team.py | ||
| test_team_logging.py | ||
| test_team_members.py | ||
| test_users.py | ||
In total litellm runs 1000+ tests
[02/20/2025] Update:
To make it easier to contribute and map what behavior is tested,
we've started mapping the litellm directory in tests/test_litellm
This folder can only run mock tests.