litellm/tests
Alexsander Hamir 6e70c279f8
[Fix] - Router's Cache: Fix routing for requests with same cacheable prefix but different user messages (#16951)
* fix(router): use cacheable prefix for prompt caching cache keys

Fix issue where requests with same cacheable prefix but different user
messages were routing to different deployments, preventing cached token
reuse. The cache key now correctly includes only the cacheable prefix
(up to and including the last cache_control block) instead of the
entire messages array.

## New Functions

### extract_cacheable_prefix()
Static method that extracts the cacheable prefix from messages for
prompt caching. The cacheable prefix is defined as everything UP TO
AND INCLUDING the LAST content block (across all messages) that has
cache_control with type "ephemeral". This includes ALL blocks
before the last cacheable block (even if they don't have cache_control
themselves).

- Finds the last content block with cache_control across all messages
- Returns all messages and content blocks up to and including that
  last cacheable block
- Excludes everything after the last cacheable block (including user
  messages that come after)
- Returns empty list if no cacheable blocks are found

## Changed Functions

### get_prompt_caching_cache_key()
Modified to use the cacheable prefix instead of the full messages array
when generating cache keys. This ensures that requests with the same
cacheable prefix but different user messages generate the same cache
key, enabling proper routing to the same deployment.

- Now calls extract_cacheable_prefix() to get only cacheable content
- Returns None if no cacheable prefix is found (can't generate key)
- Cache key is now based on cacheable prefix only, not full messages

### async_get_model_id()
Completely refactored to use the cacheable prefix directly instead of
the previous workaround that checked progressively shorter message
slices. The previous implementation was inefficient and unreliable.

- Removed progressive message slicing logic (messages[:-1], messages[:-2], etc.)
- Now uses single direct cache lookup with cacheable prefix-based key
- More efficient (1 lookup instead of up to 4)
- More reliable (uses correct cache key based on cacheable prefix)
- Returns None if no cacheable prefix found

### add_model_id()
Added None check for cache_key to prevent caching when no cacheable
prefix is found. This ensures we don't attempt to cache when there's
no meaningful cache key to use.

- Added guard: returns early if cache_key is None
- Prevents attempting to cache when no cacheable prefix exists

### async_add_model_id()
Added None check for cache_key to prevent caching when no cacheable
prefix is found. Matches the behavior of add_model_id() for consistency.

- Added guard: returns early if cache_key is None
- Prevents attempting to cache when no cacheable prefix exists

### get_model_id()
Added None check for cache_key to handle cases where no cacheable
prefix is found. Ensures consistent behavior across all cache methods.

- Added guard: returns None if cache_key is None
- Prevents calling get_cache() with None key

## Test

### test_router_prompt_caching_same_cacheable_prefix_routes_to_same_deployment()
New end-to-end test that validates the fix. Tests that requests with
the same cacheable prefix (system blocks with cache_control) but
different user messages:
1. Generate the same cache key
2. Successfully perform cache lookup
3. Route to the same deployment

This test reproduces the exact scenario from the user's bug report
where three requests with different user messages should route to the
same deployment but were previously routing to different ones.

Fixes issue where cached tokens couldn't be reused because requests
were routed to different providers due to different cache keys.

* fix(router): use cast() for proper type handling in extract_cacheable_prefix

Replace type annotation with type: ignore comment with proper cast()
from typing module, matching the pattern used throughout the
codebase for creating modified AllMessageValues dictionaries.
2025-11-21 19:13:40 -08:00
..
audio_tests add runwayml/eleven_multilingual_v2 pricing 2025-11-13 16:45:35 -08:00
basic_proxy_startup_tests test_health_and_chat_completion 2025-10-31 19:28:59 -07:00
batches_tests [Feat] Bedrock Batches - Add support for custom KMS encryption keys in Bedrock Batch operations (#16662) 2025-11-14 16:00:43 -08:00
code_coverage_tests [Infra] CI/CD Fixes (#16937) 2025-11-21 13:58:19 -08:00
documentation_tests Litellm docs readme fixes (#16107) 2025-10-30 17:05:32 -07:00
enterprise Add managed files support for responses API (#16733) 2025-11-17 18:41:26 -08:00
guardrails_tests Add Zscaler AI Guard hook (#15691) 2025-11-11 15:34:27 -08:00
image_gen_tests fix img gen 2025-11-21 17:18:48 -08:00
litellm/llms Fix image_config.aspect_ratio not working for gemini-2.5-flash-image (#15999) 2025-11-03 08:48:36 -08:00
litellm-proxy-extras Prisma Migrate - support setting custom migration dir (#10336) 2025-04-26 12:05:06 -07:00
litellm_utils_tests [Feat] Adds IAM role assumption support for AWS Secret Manager (#16887) 2025-11-20 12:38:48 -08:00
llm_responses_api_testing test_openai_streaming_logging 2025-11-08 11:49:36 -08:00
llm_translation feat(bedrock): Add Claude 4.5 to US Gov Cloud (#16957) 2025-11-21 19:06:26 -08:00
load_tests test: test_embedding_performance 2025-05-14 21:31:07 -07:00
local_testing feat(audio_transcriptions/): calculate duration of audio file for cost calculation + feat (image_generations): cost tracking accuracy improved with output_format, quality, size values fixed per openai model 2025-11-08 16:24:31 -08:00
logging_callback_tests test logging tests + mcp server QA checks 2025-11-15 08:58:46 -08:00
mcp_tests [Feat] mcp resources support (#16800) 2025-11-20 14:53:44 -08:00
multi_instance_e2e_tests fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
ocr_tests [Feat] Add Azure AI Doc Intelligence OCR (#16219) 2025-11-03 17:22:19 -08:00
old_proxy_tests/tests claude-sonnet-4-5-20250929 fix 2025-10-31 18:20:52 -07:00
openai_endpoints_tests claude-sonnet-4-5-20250929 fix 2025-10-31 18:20:52 -07:00
otel_tests test_team_budget_metrics 2025-11-01 09:21:17 -07:00
pass_through_tests fix: use vllm passthrough config for hosted vllm provider instead of raising error (#16537) 2025-11-12 09:12:14 -08:00
pass_through_unit_tests fix test claude-sonnet-4-5-20250929 2025-10-31 18:13:29 -07:00
proxy_admin_ui_tests [Infra] CI/CD Fixes (#16937) 2025-11-21 13:58:19 -08:00
proxy_security_tests test fix 2025-10-04 10:57:02 -07:00
proxy_unit_tests [Feat] UI - Prompt Management - Allow testing prompts with Chat UI (#16898) 2025-11-21 08:53:18 -08:00
router_unit_tests [Fix] - Router's Cache: Fix routing for requests with same cacheable prefix but different user messages (#16951) 2025-11-21 19:13:40 -08:00
scim_tests [Feat SSO] Add LiteLLM SCIM Integration for Team and User management (#10072) 2025-04-16 19:21:47 -07:00
search_tests [Bug Fix]: Search APIs - error in firecrawl-search "Invalid request body" (#16943) 2025-11-21 14:56:19 -08:00
spend_tracking_tests fix ocr test 2025-10-31 20:32:03 -07:00
store_model_in_db_tests Feat/persist mcp credentials in db (#16308) 2025-11-07 19:22:49 -08:00
test_litellm Allow partial matches for user id in user table (#16952) 2025-11-21 19:12:16 -08:00
unified_google_tests fix: gooogle GenAI route tests 2025-10-04 10:18:25 -07:00
vector_store_tests (feat) Milvus - search vector store support + (fix) Passthrough Endpoints - support multi-part form data on passthrough (#16035) 2025-11-01 12:00:29 -07:00
windows_tests [Bug Fix] UnicodeDecodeError: 'charmap' on Windows during litellm import (#10542) 2025-05-03 21:31:05 -07:00
__init__.py [Feat] Add github co-pilot as a new LLM API provider (#12325) 2025-07-04 13:12:16 -07:00
gettysburg.wav feat(main.py): support openai transcription endpoints 2024-03-08 10:25:19 -08:00
large_text.py fix(router.py): check for context window error when handling 400 status code errors 2024-03-26 08:08:15 -07:00
openai_batch_completions.jsonl feat(router.py): Support Loadbalancing batch azure api endpoints (#5469) 2024-09-02 21:32:55 -07:00
README.MD [Feat] MCP Gateway Fine-grained Tools Addition (#15153) 2025-10-03 10:16:29 -07:00
test_budget_management.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_callbacks_on_proxy.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_config.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_debug_warning.py fix(utils.py): fix togetherai streaming cost calculation 2024-08-01 15:03:08 -07:00
test_end_users.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_entrypoint.py (fix) clean up root repo - move entrypoint.sh and build_admin_ui to /docker (#6110) 2024-10-08 11:34:43 +05:30
test_fallbacks.py Ollama Chat - parse tool calls on streaming (#11171) 2025-05-27 16:14:49 -07:00
test_gpt5_azure_temperature_support.py Fix: Azure GPT-5 incorrectly routed to O-series config (temperature parameter unsupported) (#16246) 2025-11-07 10:24:27 -08:00
test_health.py (test) /health/readiness 2024-01-29 15:27:25 -08:00
test_keys.py test: temporarily skip test due to change testing model change - need to update test for new model 2025-05-09 09:02:08 -07:00
test_litellm_proxy_responses_config.py Add native Responses API support for litellm_proxy provider (#15347) 2025-10-08 18:31:26 -07:00
test_logging.conf feat(proxy_cli.py): add new 'log_config' cli param (#6352) 2024-10-21 21:25:58 -07:00
test_models.py test_add_model_run_health 2025-09-27 10:59:25 -07:00
test_openai_endpoints.py test fix claude-sonnet-4-5-20250929 2025-10-28 19:05:13 -07:00
test_organizations.py UI - fix adding vertex models with reusable credentials + fix pagination on keys table + fix showing org budgets on table (#10528) 2025-05-03 08:16:53 -07:00
test_passthrough_endpoints.py fix: update authorization header to use 'Bearer' instead of 'bearer' 2025-09-21 10:44:47 +00:00
test_ratelimit.py test_async_rate_limit 2025-11-15 10:34:24 -08:00
test_resource_cleanup.py Fix: Properly close aiohttp client sessions to prevent resource leaks (#12251) 2025-07-09 09:25:17 -07:00
test_spend_logs.py Revert "Allow configuration to on what threshold to try truncating request content in db" 2025-08-28 16:12:42 -06:00
test_team.py build: publish new litellm-proxy-extras file 2025-05-27 17:44:23 -07:00
test_team_logging.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_team_members.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00
test_users.py fix: use fastuuid helper (#14903) 2025-09-25 15:47:01 -07:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.