litellm/litellm/caching
user a2473ef0c2
chore(caching): remove allow_legacy_unscoped_cache_hits opt-in
The flag was an opt-in escape hatch for the cross-tenant leak the rest
of the patch closes — flipping it on (env var or constructor param)
re-enables exactly the VERIA-54 primitive on either backend. There is
no operational need that the secure path doesn't already meet:

- Qdrant: legacy points without ``litellm_cache_key`` payload are
  excluded by the must-clause filter and treated as misses; new sets
  populate the cache key, so cold-start lasts only as long as the
  natural cache rebuild.
- Redis: existing unscoped index can't carry the new schema; the init
  path falls back to ``{name}_isolated`` (and recreates it on stale
  schema), leaving the legacy index untouched.

Drop the constructor param, env-var fallback, ``_using_legacy_unscoped_index``
flag, the legacy-reuse branch in ``_init_semantic_cache``, and the
matching guards in set/get paths. Update tests to drop the legacy-mode
cases and assert the secure-only behaviour.
2026-05-04 22:16:30 +00:00
..
__init__.py Add GCS bucket caching support (#13122) 2025-08-04 16:09:33 -07:00
_internal_lru_cache.py (litellm SDK perf improvements) - handle cases when unable to lookup model in model cost map (#7750) 2025-01-13 19:58:46 -08:00
azure_blob_cache.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
base_cache.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
caching.py fix(cache): persist and replay streamed Responses API requests (#24580) 2026-05-01 11:55:36 +05:30
caching_handler.py fix(caching): defer streaming cache-hit callbacks for all stream=True 2026-05-01 17:03:32 +05:30
disk_cache.py [Bug Fix] No module named 'diskcache' (#11600) 2025-06-10 14:54:11 -07:00
dual_cache.py Merge pull request #26202 from BerriAI/litellm_token_verification_query_opt 2026-05-01 10:10:07 -07:00
gcs_cache.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
in_memory_cache.py fix: prune expired in-memory cache heap entries (#25664) 2026-04-14 23:37:49 +05:30
llm_caching_handler.py fix: don't close HTTP/SDK clients on LLMClientCache eviction (#22925) 2026-03-05 12:00:38 -08:00
qdrant_semantic_cache.py chore(caching): remove allow_legacy_unscoped_cache_hits opt-in 2026-05-04 22:16:30 +00:00
Readme.md add azure blob cache support (#12587) 2025-07-15 11:47:38 -07:00
redis_cache.py Merge pull request #26202 from BerriAI/litellm_token_verification_query_opt 2026-05-01 10:10:07 -07:00
redis_cluster_cache.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
redis_semantic_cache.py chore(caching): remove allow_legacy_unscoped_cache_hits opt-in 2026-05-04 22:16:30 +00:00
s3_cache.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00

Caching on LiteLLM

LiteLLM supports multiple caching mechanisms. This allows users to choose the most suitable caching solution for their use case.

The following caching mechanisms are supported:

  1. RedisCache
  2. RedisSemanticCache
  3. QdrantSemanticCache
  4. InMemoryCache
  5. DiskCache
  6. S3Cache
  7. AzureBlobCache
  8. DualCache (updates both Redis and an in-memory cache simultaneously)

Folder Structure

litellm/caching/
├── base_cache.py
├── caching.py
├── caching_handler.py
├── disk_cache.py
├── dual_cache.py
├── in_memory_cache.py
├── qdrant_semantic_cache.py
├── redis_cache.py
├── redis_semantic_cache.py
├── s3_cache.py

Documentation