litellm/tests/unit/caching/test_evicted_client_closer.py
yuneng-jiang a11a93f44a
test: move tests/test_litellm core utils, routing, responses, caching and rust_bridge into tests/unit (#43199)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: move tests/test_litellm/llms into tests/unit/llms

Rename-only. Moves the provider tests and the fine-tuning fixtures they
load, mirroring the old paths. Follow-up commits merge, split and wire them.

* test: merge, split and prune the moved llms tests

Merges the Databricks chat transformation tests into the existing unit
file, keeps the tests that need real keys or the network in
tests/test_litellm, deletes the audited tests a stronger unit test
already covers, and points imports at tests.unit.llms.

* ci: run the moved llms tests under their legacy flags

The Vertex AI and All Other Providers shards keep their legacy test-path
for the retained files and add the llm-vertex-ai and llm-other-providers
unit selections. CircleCI gets matching unit jobs.

* test: make the tests/unit/llms directories packages

Adds __init__.py to the moved dirs and drops the legacy ones whose
directories no longer hold tests.

* test: drop script runners and path hacks the llms split left dangling

The __main__ runners in the split openai_like files and the Databricks e2e
runner called tests that now live in the other half of the split or were
deleted. The retained legacy halves also no longer need sys.path edits.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

* test: move tests/test_litellm integrations and secret_managers into tests/unit

Rename-only. Mirrors the old paths, including the directory conftests
and the prompt and JSON fixtures. Follow-up commits prune and wire them.

* test: prune and repoint the moved integrations tests

Deletes the 7 audited tests a stronger test in the same tree already
covers, imports the TLS sink helpers from their new conftest path, and
restores os.environ after each integrations test. Some presets write
OTEL_EXPORTER_OTLP_HEADERS straight into os.environ, and without the
legacy tree's test ordering that header leaked into the AgentOps tests.

* ci: run the moved integrations tests under their legacy flag

The integrations GHA shard and a new CircleCI job run the integrations
unit selection. secret_managers joins the misc selection.

* docs: point integrations and secret_managers references at tests/unit

* test: make the moved integrations directories packages

* test: keep the Databricks manual e2e runner and fix the SageMaker Nova run path

The Databricks e2e file is a manual script whose main() calls the tests
that were pruned, so pruning them broke the documented run. It is back to
its main version. The SageMaker Nova docstring now points at the file's
real location in tests/local_testing.

* test: move tests/test_litellm core utils, routing, responses, caching and rust_bridge into tests/unit

Rename-only. Mirrors the old paths, including fixtures, the stubtest config
and the native-route wheel script. Two files that collide with existing unit
files are merged in a follow-up commit.

* test: merge, prune and repoint the moved core, routing, responses, caching and rust_bridge tests

Merges the two files that collided with existing unit files, folding the
legacy extra case into test_is_chat_completion_cached_dict, and deletes the
9 audited tests a stronger test in the same file already covers.

Keeps what needs the network in tests/test_litellm: test_tokenizers pulls a
tokenizer from the Hugging Face hub, and the gpt2 and r50k_base tokenizer
cases download their BPE files. The unit core_utils conftest points
TIKTOKEN_CACHE_DIR at litellm's bundled encodings so the rest never depend on
import order to stay offline, and FakeSecretVault moves to a shared module
so both trees can build it.

* ci: run the moved core, routing, responses, caching and rust_bridge tests under their flags

core_utils gets a core-utils flag and CircleCI job, and its GHA shard keeps
the legacy path for the retained network tests. router_utils and
router_strategy join enterprise-routing, responses joins
responses-caching-types (minus responses/mcp, which mcp-integration owns),
caching joins caching-local and rust_bridge joins misc. The redis-compat,
test-rust, stubtest and merge-smoke paths follow the move.

* docs: point the Rust crate references at tests/unit

* test: make the moved core, routing and rust_bridge directories packages

* test: keep the no-loop DualCache batch_get_cache regression test

It runs the sync path outside any event loop, which the inside-loop test
cannot, so a change that picks the Redis client by loop state would only
show up there.

* test: keep the job's UNIT_FLAG out of the shard-script tests

* fix(url_utils): block 192.0.0.0/24 on every Python patch release

* test: move the new budget limiter tests into tests/unit/router_strategy

* test: move the new sentry scrubbing tests into tests/unit/litellm_core_utils

* test: move the new zerobus tests into tests/unit/integrations

* test: make tests/unit/integrations/zerobus a package

* test: load litellm's own tiktoken cache setup once instead of resetting it per test

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 17:10:13 -07:00

440 lines
13 KiB
Python

"""
Tests for EvictedClientCloser.
An evicted client must stay open long enough for a request that already holds it
to finish, and must then actually be closed, otherwise its connection pool is
retained until a generational collection runs. A client the caller supplied is
never closed, because litellm does not own its lifecycle.
"""
import asyncio
import gc
import weakref
from unittest.mock import AsyncMock
import httpx
import pytest
from redis.asyncio import ConnectionPool, Redis
from litellm.caching.evicted_client_closer import EvictedClientCloser
from litellm.llms.custom_httpx.http_handler import AsyncHTTPHandler
class FakeClock:
"""Hand-advanced monotonic clock, so grace windows need no real waiting."""
def __init__(self) -> None:
self.now = 1000.0
def __call__(self) -> float:
return self.now
def advance(self, seconds: float) -> None:
self.now += seconds
class AsyncClient:
def __init__(self) -> None:
self.closed = False
async def close(self) -> None:
self.closed = True
class SyncClient:
def __init__(self) -> None:
self.closed = False
def close(self) -> None:
self.closed = True
class CountingDeadline(float):
"""A clock reading that tallies every deadline comparison made against it.
Deadline comparisons are the work a reap does, so counting them says whether
that work tracks the entries that are due or the size of the whole queue.
"""
comparisons = 0
def __add__(self, other: float) -> "CountingDeadline":
return CountingDeadline(float(self) + other)
def __le__(self, other: float) -> bool:
CountingDeadline.comparisons += 1
return float(self) <= float(other)
def __gt__(self, other: float) -> bool:
CountingDeadline.comparisons += 1
return float(self) > float(other)
def make_closer(clock: FakeClock, grace_seconds: float = 60.0) -> EvictedClientCloser:
return EvictedClientCloser(grace_seconds=grace_seconds, clock=clock)
@pytest.mark.asyncio
async def test_redis_client_is_closed_only_after_its_subscription_releases_the_connection():
closer = EvictedClientCloser(grace_seconds=0)
pool = ConnectionPool()
client = Redis.from_pool(pool)
closed = asyncio.Event()
connection = AsyncMock()
connection.disconnect.side_effect = closed.set
pool._available_connections.append(connection)
borrowed = pool.get_available_connection()
closer.mark_owned(client)
closer.schedule(client)
closer.reap()
await asyncio.sleep(0.05)
connection.disconnect.assert_not_awaited()
assert closer.pending_count == 1
await pool.release(borrowed)
closer.reap()
await asyncio.wait_for(closed.wait(), timeout=1)
connection.disconnect.assert_awaited_once()
assert closer.pending_count == 0
async def _trickling_upstream(reader: asyncio.StreamReader, writer: asyncio.StreamWriter) -> None:
"""Serves a chunked body slowly, so a request stays on the wire long enough to observe."""
await reader.read(4096)
writer.write(b"HTTP/1.1 200 OK\r\nTransfer-Encoding: chunked\r\n\r\n")
await writer.drain()
for _ in range(6):
writer.write(b"5\r\nhello\r\n")
await writer.drain()
await asyncio.sleep(0.1)
writer.write(b"0\r\n\r\n")
await writer.drain()
@pytest.mark.asyncio
async def test_owned_client_is_closed_once_the_grace_window_elapses():
clock = FakeClock()
closer = make_closer(clock)
client = AsyncClient()
closer.mark_owned(client)
closer.schedule(client)
clock.advance(61.0)
closer.reap()
await asyncio.sleep(0.05)
assert client.closed is True
assert closer.pending_count == 0
@pytest.mark.asyncio
async def test_owned_client_stays_open_inside_the_grace_window():
"""A request handed the client just before eviction is still using it."""
clock = FakeClock()
closer = make_closer(clock)
client = AsyncClient()
closer.mark_owned(client)
closer.schedule(client)
clock.advance(59.0)
closer.reap()
await asyncio.sleep(0.05)
assert client.closed is False
assert closer.pending_count == 1
@pytest.mark.asyncio
async def test_caller_supplied_client_is_never_closed():
clock = FakeClock()
closer = make_closer(clock)
client = AsyncClient()
closer.schedule(client)
clock.advance(3600.0)
closer.reap()
await asyncio.sleep(0.05)
assert client.closed is False
assert closer.pending_count == 0
@pytest.mark.asyncio
async def test_sync_client_is_closed_once_the_grace_window_elapses():
clock = FakeClock()
closer = make_closer(clock)
client = SyncClient()
closer.mark_owned(client)
closer.schedule(client)
clock.advance(61.0)
closer.reap()
assert client.closed is True
@pytest.mark.asyncio
async def test_a_failing_close_does_not_propagate_or_block_the_others():
class ExplodingClient:
async def close(self) -> None:
raise RuntimeError("connection already gone")
clock = FakeClock()
closer = make_closer(clock)
exploding, healthy = ExplodingClient(), AsyncClient()
for client in (exploding, healthy):
closer.mark_owned(client)
closer.schedule(client)
clock.advance(61.0)
closer.reap()
await asyncio.sleep(0.05)
assert healthy.closed is True
@pytest.mark.asyncio
async def test_an_unhashable_cached_value_does_not_break_eviction():
"""The cache holds arbitrary values; an ownership test must never raise on one."""
class Unhashable:
__hash__ = None # pyright: ignore[reportAssignmentType] # unhashable by construction
clock = FakeClock()
closer = make_closer(clock)
closer.mark_owned(Unhashable())
closer.schedule(Unhashable())
assert closer.pending_count == 0
@pytest.mark.asyncio
async def test_values_with_nothing_to_close_are_never_queued():
"""The cache holds plain values too; those have nothing to reclaim."""
class NotAClient:
pass
clock = FakeClock()
closer = make_closer(clock)
value = NotAClient()
closer.mark_owned(value)
closer.schedule(value)
assert closer.pending_count == 0
@pytest.mark.asyncio
async def test_a_queued_client_is_not_kept_alive_by_the_queue():
"""Waiting out a grace window must not retain what the collector would free first."""
clock = FakeClock()
closer = make_closer(clock)
client = AsyncClient()
gone = weakref.ref(client)
closer.mark_owned(client)
closer.schedule(client)
del client
gc.collect()
assert gone() is None, "the pending queue is holding the client alive"
clock.advance(61.0)
closer.reap()
assert closer.pending_count == 0
def test_sync_client_evicted_outside_an_event_loop_is_still_closed():
"""The sync httpx handler is cached and evicted from call sites with no loop."""
clock = FakeClock()
closer = make_closer(clock)
client = SyncClient()
closer.mark_owned(client)
closer.schedule(client)
assert closer.pending_count == 1
clock.advance(61.0)
closer.reap()
assert client.closed is True
assert closer.pending_count == 0
@pytest.mark.asyncio
async def test_an_async_client_waits_for_a_loop_rather_than_being_dropped():
clock = FakeClock()
closer = make_closer(clock)
client = AsyncClient()
closer.mark_owned(client)
def schedule_outside_a_loop() -> None:
closer.schedule(client)
clock.advance(61.0)
closer.reap()
await asyncio.to_thread(schedule_outside_a_loop)
assert client.closed is False, "no loop was running, so it could not have been closed"
assert closer.pending_count == 1
closer.reap()
await asyncio.sleep(0.05)
assert client.closed is True
@pytest.mark.asyncio
async def test_a_client_evicted_on_another_event_loop_is_left_alone():
"""Closing a client bound to a different loop would schedule work on that loop."""
clock = FakeClock()
closer = make_closer(clock)
client = AsyncClient()
closer.mark_owned(client)
def schedule_on_its_own_loop() -> None:
asyncio.run(_schedule())
async def _schedule() -> None:
closer.schedule(client)
await asyncio.to_thread(schedule_on_its_own_loop)
assert closer.pending_count == 1
clock.advance(61.0)
closer.reap()
await asyncio.sleep(0.05)
assert client.closed is False
assert closer.pending_count == 1
@pytest.mark.asyncio
async def test_a_client_serving_a_request_is_not_closed_when_its_grace_window_ends():
"""The grace window on its own cannot promise that a request has finished.
``litellm.request_timeout`` defaults to 6000 seconds and a streaming response
is bounded only by how long the upstream keeps sending, so a client past its
deadline is closed only once its own pool reports nothing in flight.
"""
server = await asyncio.start_server(_trickling_upstream, "127.0.0.1", 0)
port = server.sockets[0].getsockname()[1]
clock = FakeClock()
closer = make_closer(clock)
client = httpx.AsyncClient()
closer.mark_owned(client)
closer.schedule(client)
async def read_the_stream() -> int:
received = 0
async with client.stream("GET", f"http://127.0.0.1:{port}/") as response:
async for chunk in response.aiter_bytes():
received += len(chunk)
return received
streaming = asyncio.create_task(read_the_stream())
await asyncio.sleep(0.25) # the request is on the wire
clock.advance(3600.0) # and its grace window is long gone
closer.reap()
await asyncio.sleep(0.05)
assert client.is_closed is False, "closed a client that was serving a request"
assert await streaming > 0, "the in-flight request did not survive the reap"
clock.advance(3600.0)
closer.reap()
await asyncio.sleep(0.05)
assert client.is_closed is True, "an idle client past its grace window must be closed"
assert closer.pending_count == 0
server.close()
@pytest.mark.asyncio
async def test_the_aiohttp_backed_handler_is_not_closed_mid_request():
"""The default async path is aiohttp-backed, whose pool accounts for its own leases."""
server = await asyncio.start_server(_trickling_upstream, "127.0.0.1", 0)
port = server.sockets[0].getsockname()[1]
clock = FakeClock()
closer = make_closer(clock)
handler = AsyncHTTPHandler()
held_client = handler.client
closer.mark_owned(handler)
closer.schedule(handler)
request = asyncio.create_task(handler.get(f"http://127.0.0.1:{port}/"))
await asyncio.sleep(0.25)
clock.advance(3600.0)
closer.reap()
await asyncio.sleep(0.05)
assert held_client.is_closed is False, "closed a handler that was serving a request"
assert (await request).status_code == 200
clock.advance(3600.0)
closer.reap()
await asyncio.sleep(0.05)
assert held_client.is_closed is True
assert handler.client.is_closed is False, "a held handler must self-heal after its evicted client is closed"
server.close()
def test_the_pending_queue_cannot_grow_past_its_bound():
"""A caller that churns the client cache must not be able to grow this queue."""
clock = FakeClock()
closer = EvictedClientCloser(grace_seconds=60.0, max_pending=8, clock=clock)
clients = tuple(SyncClient() for _ in range(50))
for client in clients:
closer.mark_owned(client)
closer.schedule(client)
assert closer.pending_count == 8, "the queue grew past max_pending"
clock.advance(61.0)
closer.reap()
assert closer.pending_count == 0
assert sum(client.closed for client in clients) == 8, "everything queued should have been closed"
def test_a_reap_looks_at_what_is_due_rather_than_at_the_whole_queue():
"""Sustained churn evicts a client per request, and every read of the cache reaps.
So the cost of a reap has to track the entries that are due, not the length of
the queue; a reap that filters the whole queue makes the pair quadratic. Each
bucket is ordered by deadline, so an up-to-date reap compares one entry per
bucket and stops. Counting the comparisons measures that directly, where a
wall-clock budget would only measure the machine.
"""
evictions = 1_000
clock = FakeClock()
closer = EvictedClientCloser(
grace_seconds=60.0,
max_pending=evictions,
clock=lambda: CountingDeadline(clock.now),
)
clients = tuple(SyncClient() for _ in range(evictions))
for client in clients:
closer.mark_owned(client)
CountingDeadline.comparisons = 0
for client in clients:
closer.schedule(client)
closer.reap() # nothing is due yet, which is the hot path
clock.advance(61.0)
closer.reap()
assert closer.pending_count == 0
assert all(client.closed for client in clients)
assert CountingDeadline.comparisons < 10 * evictions, (
f"{CountingDeadline.comparisons} deadline comparisons for {evictions} evictions; "
"a reap is walking the whole queue"
)