litellm/litellm-rust/crates/cache-qdrant-semantic/tests/contract.rs
devin-ai-integration[bot] 4677f1028e
refactor(rust): align the cache crates with Python and wire every native backend (#42530)
* refactor(rust): align the cache crates with Python and activate every backend

The cache port had drifted: lifecycle and Redis-only operations sat on
`BaseCache`, counters were pinned to `f64`, each semantic backend defined its
own embedder and prompt handling, and only the in-memory backend could be
selected natively.

- Split `disconnect` and `test_connection` out of `BaseCache` into optional
  capabilities, implemented only where the Python class defines them, and give
  every Redis-only operation its own capability trait.
- Decouple counters from the stored value type, so one backend can serve both
  responses and counters as Python's `RedisCache` does.
- Share one `Embedder` and prompt contract in `litellm_cache::semantic`, and
  make the Redis and Valkey semantic backends generic over their codec.
- Port the Python operations that were missing: `async_refresh_ttl`,
  `async_rpush_and_trim`, `async_set_cache_pipeline_with_ttls`, the DualCache
  pipeline, sadd, bulk delete and TTL reads, and the semantic-similarity
  write-back.
- Take the HTTP client from the host pool in the GCS, S3 and Azure backends.
- Activate all nine backends through the Rust catalog, whose rules all stay
  `PYTHON_ONLY`, and route the `Cache` facade's storage calls to the native
  runtime when one is selected.
- Give every crate the same layout, move all tests to `tests/` on rstest, and
  add the shared `litellm-cache-testing` contract suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: freeze native cache request kwargs and batch entries for type discipline

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: declare semantic lookup methods in the native stub

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): align the cache crates with Python and activate every backend

The cache port had drifted: lifecycle and Redis-only operations sat on
`BaseCache`, counters were pinned to `f64`, each semantic backend defined its
own embedder and prompt handling, and only the in-memory backend could be
selected natively.

- Split `disconnect` and `test_connection` out of `BaseCache` into optional
  capabilities, implemented only where the Python class defines them, and give
  every Redis-only operation its own capability trait.
- Decouple counters from the stored value type, so one backend can serve both
  responses and counters as Python's `RedisCache` does.
- Share one `Embedder` and prompt contract in `litellm_cache::semantic`, and
  make the Redis and Valkey semantic backends generic over their codec.
- Port the Python operations that were missing: `async_refresh_ttl`,
  `async_rpush_and_trim`, `async_set_cache_pipeline_with_ttls`, the DualCache
  pipeline, sadd, bulk delete and TTL reads, and the semantic-similarity
  write-back.
- Take the HTTP client from the host pool in the GCS, S3 and Azure backends.
- Activate all nine backends through the Rust catalog, whose rules all stay
  `PYTHON_ONLY`, and route the `Cache` facade's storage calls to the native
  runtime when one is selected.
- Give every crate the same layout, move all tests to `tests/` on rstest, and
  add the shared `litellm-cache-testing` contract suite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: freeze native cache request kwargs and batch entries for type discipline

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: declare semantic lookup methods in the native stub

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): opt the native Messages and tokenizer suites into Rust explicitly

#42517 made the Messages, token counter and tokenizer routes Python-only, so
tests/test_litellm_rust silently exercised the Python path or failed outright.
Each suite now prepends a RUST_OPT_IN rule for its route, keeping native
coverage without changing the shipped default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(rust): pop one at a time in the Redis 6 lpop pipeline and drop explanatory comments

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-22 13:13:02 -07:00

91 lines
2.7 KiB
Rust

//! `overwrite_replaces` does not apply: like Python, every write upserts a new `uuid4` point,
//! so a second write with the same prompt adds a tie instead of replacing the first.
mod support;
use std::future::Future;
use litellm_cache::{JsonCodec, SemanticCacheContext, semantic::PreparedEmbedding};
use litellm_cache_qdrant_semantic::{QdrantSemanticCache, QdrantSemanticConfig, Quantization};
use litellm_cache_testing as contract;
use qdrant_client::Qdrant;
use rstest::{fixture, rstest};
use serde_json::{Value, json};
use support::{FakeQdrant, FakeState};
type Cache = QdrantSemanticCache<PreparedEmbedding, JsonCodec<Value>>;
const PREFIX: &str = "contract:";
#[fixture]
fn context() -> SemanticCacheContext {
SemanticCacheContext {
messages: Some(json!([{"role": "user", "content": "contract prompt"}])),
..Default::default()
}
}
/// Runs a contract against a fresh fake Qdrant. The sync cache methods block on the runtime, so
/// the contract is polled on a blocking thread outside the runtime's own executor.
async fn run<F, Fut>(check: F)
where
F: FnOnce(Cache) -> Fut + Send + 'static,
Fut: Future<Output = ()>,
{
let server = FakeQdrant::start(FakeState::default()).await;
let runtime = tokio::runtime::Handle::current();
let cache = QdrantSemanticCache::connect(
Qdrant::from_url(&server.url()).build().unwrap(),
PreparedEmbedding(vec![0.6, 0.8]),
JsonCodec::new(),
QdrantSemanticConfig {
collection_name: "contract".to_owned(),
similarity_threshold: 0.9,
vector_size: 2,
quantization: Quantization::Binary,
},
runtime.clone(),
)
.await
.unwrap();
tokio::task::spawn_blocking(move || {
let _guard = runtime.enter();
futures_executor::block_on(check(cache));
})
.await
.unwrap();
server.stop();
}
#[rstest]
#[tokio::test(flavor = "multi_thread")]
async fn hit_and_miss(context: SemanticCacheContext) {
run(|cache| async move {
contract::hit_and_miss(&cache, context, PREFIX, json!({"answer": 42})).await;
})
.await;
}
#[rstest]
#[tokio::test(flavor = "multi_thread")]
async fn sync_async_equivalence(context: SemanticCacheContext) {
run(|cache| async move {
contract::sync_async_equivalence(&cache, context, PREFIX, json!("first"), json!([2])).await;
})
.await;
}
#[rstest]
#[tokio::test(flavor = "multi_thread")]
async fn pipeline_writes_every_entry(context: SemanticCacheContext) {
run(|cache| async move {
contract::pipeline_writes_every_entry(
&cache,
context,
PREFIX,
vec![json!("a"), json!(2), json!({"c": true})],
)
.await;
})
.await;
}