litellm/litellm/router_utils
devin-ai-integration[bot] 4c6c84afb7
perf(proxy): stop prompt-cache eligibility from tokenizing the whole conversation (#44221)
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them

Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
..
pre_call_checks perf(proxy): stop prompt-cache eligibility from tokenizing the whole conversation (#44221) 2026-10-03 09:20:28 -07:00
router_callbacks fix(router): count TPM/RPM usage before building rate-limit headers 2026-09-16 18:45:01 +00:00
access_windows.py feat(router): time-windowed team reservation of deployments via model_info.access_windows (#42398) 2026-09-22 13:22:15 -05:00
add_retry_fallback_headers.py feat(router): add group-scoped priority routing strategy (#42378) 2026-09-22 13:09:38 -07:00
auto_router_model_naming.py feat: add Laya gateway and OSS classifier providers (#43626) 2026-10-02 11:29:36 -07:00
auto_router_tuning_baseline.py feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00
batch_utils.py feat(vertex): native batch JSONL passthrough with cost tracking (#42810) 2026-09-24 12:35:34 -07:00
client_initalization_utils.py feat(router): reject with 429 when a deployment's max_parallel_requests slots are all in use 2026-09-17 18:14:09 +00:00
clientside_credential_handler.py feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011) 2026-08-04 12:54:39 -07:00
common_utils.py fix(router): explain fallback outcome in plain words in the raised error (#42509) 2026-09-22 18:37:42 +00:00
cooldown_cache.py fix(otel): nest cache spans under their operation and name service spans by purpose (#44150) 2026-10-03 09:20:27 -07:00
cooldown_callbacks.py feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011) 2026-08-04 12:54:39 -07:00
cooldown_handlers.py fix(otel): nest cache spans under their operation and name service spans by purpose (#44150) 2026-10-03 09:20:27 -07:00
fallback_event_handlers.py chore(lint): remove the LIT002 mutable-construction rule (#43971) 2026-10-01 12:24:02 -07:00
get_retry_from_policy.py fix(router): add NotFoundErrorRetries so a retry policy can pin 404 retries 2026-09-19 16:25:47 -07:00
handle_error.py fix(router): name the all-deployments-in-cooldown error on 429 responses 2026-09-13 09:34:16 +00:00
health_state_cache.py fix(otel): nest cache spans under their operation and name service spans by purpose (#44150) 2026-10-03 09:20:27 -07:00
pattern_match_deployments.py fix(router): match provider-prefixed fallback keys for bare model groups served by wildcard deployments (#43062) 2026-09-24 18:08:56 -07:00
prompt_caching_cache.py fix(otel): nest cache spans under their operation and name service spans by purpose (#44150) 2026-10-03 09:20:27 -07:00
reasoning_effort_capability.py fix(mistral): never round reasoning_effort none onto the strength ladder 2026-09-18 12:22:34 -07:00
response_headers.py LiteLLM Minor Fixes & Improvements (11/26/2024) (#6913) 2024-11-28 00:01:38 +05:30
routing_groups.py feat(router): add group-scoped priority routing strategy (#42378) 2026-09-22 13:09:38 -07:00
routing_read_batch.py fix(otel): nest cache spans under their operation and name service spans by purpose (#44150) 2026-10-03 09:20:27 -07:00
search_api_router.py fix(types): drop restating docstrings, import search tool types at runtime, close the spend-table match 2026-09-21 13:44:58 -07:00