litellm/tests
devin-ai-integration[bot] a173657dfb
fix(caching): keep embedding cache hits aligned with request inputs (#42571)
* fix(caching): keep embedding cache hits aligned with request inputs

Partial hits now send only the uncached inputs to the provider and merge
fresh vectors back into their original positions. Responses whose item
count differs from the input count (one input scoring many documents)
are no longer written to the per-input cache, since a later hit would
return a single item.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): drop mutable collection builds flagged by the type discipline gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): bypass embedding cache entries written before the per input cardinality check

Embedding cache entries now carry format_version and readers treat entries without it as
misses, so entries that only hold the first row of a multi row response are refetched instead
of served until their TTL expires. The provider call also receives a copy of the request kwargs
with the uncached inputs rather than mutating the caller's mapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): assert a partial embedding cache hit becomes a full hit on repeat

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(caching): await pending embedding cache writes before asserting on cache hits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): validate cached embeddings without mutating responses or request kwargs

Validate cache rows through a frozen pydantic model so import does not depend on
TypeAdapter support for ReadOnly TypedDicts, accept string embeddings, build the
merged partial hit response instead of mutating the cached one, and hand the
provider request mapping to post call hooks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep cache_hit and response_ms on merged partial embedding hits

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 17:14:44 -05:00
..
agent_tests fix(a2a): keep the upstream status on card discovery failures and inject the card client in tests 2026-09-16 23:55:57 +00:00
audio_tests
base_sdk_tests fix(mcp): explain missing public client dependencies 2026-09-19 19:52:13 -07:00
basic_proxy_startup_tests
batches_tests
benchmarks feat(tokenizer): preserve Python defaults with opt-in Rust dispatch (#42174) 2026-09-22 04:41:11 +00:00
code_coverage_tests ci(code-quality): allowlist _render_json in the recursive detector (#42442) 2026-09-22 01:05:58 -07:00
documentation_tests fix(docs-test): restore line-anchored table regex in router settings check 2026-09-18 20:10:20 +00:00
e2e test(e2e): run the memory cell alone on the shared stack (#42518) 2026-09-22 15:13:37 -07:00
enterprise
guardrails_tests Revert "Merge pull request #41220 from BerriAI/litellm_post_call_guardrail_context" 2026-09-19 17:45:15 +00:00
image_gen_tests test: remove fully commented-out test files that collect no tests 2026-09-17 20:05:27 +00:00
integration fix(otel): keep text completion choice fields beside the synthesized message (#42537) 2026-09-22 14:24:32 -07:00
litellm-proxy-extras Merge remote-tracking branch 'origin/main' into litellm_v2_migration_startup 2026-09-18 20:55:05 -07:00
litellm_utils_tests Merge remote-tracking branch 'origin/main' into litellm_invalid_tool_choice_400 2026-09-19 02:35:29 -07:00
llm_responses_api_testing test: fix five tests left stale by #41311, #41337, #39996 and #41310 2026-09-16 17:47:27 -07:00
llm_translation test(vcr): guard leaked cassette patches and make injected-transport embedding tests immune (#42542) 2026-09-22 14:20:18 -07:00
load_tests
local_testing test(vcr): guard leaked cassette patches and make injected-transport embedding tests immune (#42542) 2026-09-22 14:20:18 -07:00
logging_callback_tests fix(langsmith): json.dumps with default=str so non-serializable metadata does not crash batch flush (#42424) 2026-09-22 12:21:11 -07:00
mcp_tests fix(mcp): apply post-call rewrites without stale structured output (#41530) 2026-09-22 13:44:25 -07:00
multi_instance_e2e_tests
ocr_tests test(ocr): replace per-provider OCR test classes with a declarative provider x auth x input matrix 2026-09-18 21:27:03 +00:00
openai_endpoints_tests
otel_tests test: fix stale budget-status and bad-database-url assertions 2026-09-21 15:00:54 -07:00
pass_through_tests fix(mcp): preserve legacy behavior on SDK2 and streamline verification 2026-09-18 22:28:31 -07:00
pass_through_unit_tests feat(proxy): add TinyFish Agent API passthrough with per-step billing (#41099) 2026-09-21 21:21:43 -07:00
proxy_admin_ui_tests
proxy_behavior Merge pull request #42026 from BerriAI/litellm_user_jwt_savings 2026-09-21 16:41:16 -07:00
proxy_e2e_anthropic_messages_tests fix(test): run the all-beta-headers bedrock cases on Claude Fable 5.1 2026-09-19 23:33:29 +00:00
proxy_migration_tests test: count a zombie grandchild as gone in the migrate deploy timeout test (#42570) 2026-09-22 14:59:26 -07:00
proxy_security_tests refactor(proxy): rename the local development override to dangerously_permit_weak_or_unset_master_key so the name says exactly what it permits 2026-09-19 18:53:14 -07:00
proxy_unit_tests feat(tokenizer): preserve Python defaults with opt-in Rust dispatch (#42174) 2026-09-22 04:41:11 +00:00
router_unit_tests Merge pull request #42283 from BerriAI/litellm_mid_stream_fallback_walks_full_list 2026-09-21 15:14:04 -07:00
rust-python-harness test: drop tests that pin provider-owned cost map values 2026-09-18 03:55:51 +00:00
search_tests test: remove fully commented-out test files that collect no tests 2026-09-17 20:05:27 +00:00
spend_tracking_tests
store_model_in_db_tests
test_gateway
test_litellm fix(caching): keep embedding cache hits aligned with request inputs (#42571) 2026-09-22 17:14:44 -05:00
test_litellm_rust refactor(rust): align the cache crates with Python and wire every native backend (#42530) 2026-09-22 13:13:02 -07:00
unified_google_tests test(google_genai): move unified_google_tests to gemini-3.5-flash-lite (#42520) 2026-09-22 13:07:37 -07:00
unit chore(cost-map): remove models past their deprecation date (#42435) 2026-09-22 21:19:26 +00:00
vector_store_tests
windows_tests
__init__.py
_fake_openai_endpoint_server.py
_flush_vcr_cache.py
_live_test_helpers.py refactor(test): validate cost-map entries into a typed model 2026-09-16 14:32:08 -07:00
_openai_record_replay_proxy.py
_process_helpers.py test: count a zombie grandchild as gone in the migrate deploy timeout test (#42570) 2026-09-22 14:59:26 -07:00
_vcr_conftest_common.py test(vcr): guard leaked cassette patches and make injected-transport embedding tests immune (#42542) 2026-09-22 14:20:18 -07:00
_vcr_redis_persister.py
_wait_helpers.py
_ws_vcr.py
AGENTS.md ci(tests): wire tests/unit into CircleCI and drain legacy unit shards green 2026-09-20 07:05:42 +00:00
capturing_transport.py test(vcr): guard leaked cassette patches and make injected-transport embedding tests immune (#42542) 2026-09-22 14:20:18 -07:00
eval_swe_bench.py
fake_openai_endpoint.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
pyrightconfig.json
README.MD
test_anthropic_compaction_usage.py
test_budget_management.py
test_callbacks_on_proxy.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py
test_keys.py
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_new_vector_store_endpoints.py
test_openai_endpoints.py test: fix five tests left stale by #41311, #41337, #39996 and #41310 2026-09-16 17:47:27 -07:00
test_organizations.py
test_otel_thread_leak.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_rust_python_harness.py
test_service_logger_otel.py
test_spend_logs.py
test_team.py
test_team_logging.py
test_team_members.py
test_users.py

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.