mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-05 02:41:56 +00:00
fix(caching): handle list-based responses, message key variations, and robust serialization in QdrantSemanticCache
Several issues were identified in the semantic caching logic: 1. List-based responses caused a `TypeError` (`can only concatenate str (not "list") to str`) during caching. 2. The use of `str(value)` for serialization generated non-JSON-compliant strings (using single quotes), forcing reliance on fragile `ast.literal_eval`. 3. Message retrieval logic incorrectly treated empty lists `[]` as falsy, causing cache operations to be silently skipped when valid empty conversations were provided. 4. Callers were inconsistently passing `messages` or `message` keys, leading to potential key errors. 5. `json.dumps` lacked a fallback encoder, causing cache writes to fail for non-serializable types. 6. The test suite contained unused mock patches. - Replaced `str(value)` with `json.dumps(value, default=str)` in `set_cache` and `async_set_cache` to ensure all cached values are stored as robust, standard JSON, falling back to string representation for non-serializable objects. - Refactored message retrieval in all cache methods to explicitly check for key presence (`if "messages" in kwargs`) rather than relying on falsy evaluation, allowing empty lists `[]` to be treated as valid inputs. - Modified input handling in cache methods to support both 'messages' and 'message' keys, ensuring robustness against varying input structures. - Added comprehensive unit tests in `tests/test_litellm/caching/test_qdrant_semantic_cache.py` to validate list-type response caching and confirm the fix for the reported errors. - Cleaned up the test suite by removing unused `mock_async_client` patches in synchronous test cases.
This commit is contained in:
parent
5df898b2f9
commit
430d4a0a8e
1 changed files with 2 additions and 2 deletions
|
|
@ -197,7 +197,7 @@ class QdrantSemanticCache(BaseCache):
|
|||
# get the embedding
|
||||
embedding = embedding_response["data"][0]["embedding"]
|
||||
|
||||
value = json.dumps(value)
|
||||
value = json.dumps(value, default=str)
|
||||
assert isinstance(value, str)
|
||||
|
||||
data = {
|
||||
|
|
@ -329,7 +329,7 @@ class QdrantSemanticCache(BaseCache):
|
|||
# get the embedding
|
||||
embedding = embedding_response["data"][0]["embedding"]
|
||||
|
||||
value = json.dumps(value)
|
||||
value = json.dumps(value, default=str)
|
||||
assert isinstance(value, str)
|
||||
|
||||
data = {
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue