Non-streaming responses served by the Rust core already carry
x-litellm-rust: true through _hidden_params.additional_headers, which the
SDK exposes and the gateway renders as a response header. Native streams
did not, because the lifecycle Stream and SyncStream objects had nowhere
to hold hidden params and the marker writer skips objects without them.
Give both stream classes the same _hidden_params bag every other litellm
response has, so the existing marker attaches without wrapping the stream
or changing its identity.
Co-authored-by: Yujong Lee <yujong@berri.ai>
* refactor(rust): align the cache crates with Python and activate every backend
The cache port had drifted: lifecycle and Redis-only operations sat on
`BaseCache`, counters were pinned to `f64`, each semantic backend defined its
own embedder and prompt handling, and only the in-memory backend could be
selected natively.
- Split `disconnect` and `test_connection` out of `BaseCache` into optional
capabilities, implemented only where the Python class defines them, and give
every Redis-only operation its own capability trait.
- Decouple counters from the stored value type, so one backend can serve both
responses and counters as Python's `RedisCache` does.
- Share one `Embedder` and prompt contract in `litellm_cache::semantic`, and
make the Redis and Valkey semantic backends generic over their codec.
- Port the Python operations that were missing: `async_refresh_ttl`,
`async_rpush_and_trim`, `async_set_cache_pipeline_with_ttls`, the DualCache
pipeline, sadd, bulk delete and TTL reads, and the semantic-similarity
write-back.
- Take the HTTP client from the host pool in the GCS, S3 and Azure backends.
- Activate all nine backends through the Rust catalog, whose rules all stay
`PYTHON_ONLY`, and route the `Cache` facade's storage calls to the native
runtime when one is selected.
- Give every crate the same layout, move all tests to `tests/` on rstest, and
add the shared `litellm-cache-testing` contract suite.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: freeze native cache request kwargs and batch entries for type discipline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix: declare semantic lookup methods in the native stub
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): align the cache crates with Python and activate every backend
The cache port had drifted: lifecycle and Redis-only operations sat on
`BaseCache`, counters were pinned to `f64`, each semantic backend defined its
own embedder and prompt handling, and only the in-memory backend could be
selected natively.
- Split `disconnect` and `test_connection` out of `BaseCache` into optional
capabilities, implemented only where the Python class defines them, and give
every Redis-only operation its own capability trait.
- Decouple counters from the stored value type, so one backend can serve both
responses and counters as Python's `RedisCache` does.
- Share one `Embedder` and prompt contract in `litellm_cache::semantic`, and
make the Redis and Valkey semantic backends generic over their codec.
- Port the Python operations that were missing: `async_refresh_ttl`,
`async_rpush_and_trim`, `async_set_cache_pipeline_with_ttls`, the DualCache
pipeline, sadd, bulk delete and TTL reads, and the semantic-similarity
write-back.
- Take the HTTP client from the host pool in the GCS, S3 and Azure backends.
- Activate all nine backends through the Rust catalog, whose rules all stay
`PYTHON_ONLY`, and route the `Cache` facade's storage calls to the native
runtime when one is selected.
- Give every crate the same layout, move all tests to `tests/` on rstest, and
add the shared `litellm-cache-testing` contract suite.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: freeze native cache request kwargs and batch entries for type discipline
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix: declare semantic lookup methods in the native stub
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(rust): opt the native Messages and tokenizer suites into Rust explicitly
#42517 made the Messages, token counter and tokenizer routes Python-only, so
tests/test_litellm_rust silently exercised the Python path or failed outright.
Each suite now prepends a RUST_OPT_IN rule for its route, keeping native
coverage without changing the shipped default.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(rust): pop one at a time in the Redis 6 lpop pipeline and drop explanatory comments
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>