litellm/tests/unit/test_lowest_latency_zero_tokens.py
yuneng-jiang f6882246d4
test: move tests/test_litellm root and small trees into tests/unit (#43186)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 11:30:43 -07:00

131 lines
3.9 KiB
Python

#### What this tests ####
# This tests the router's handling of zero completion tokens in lowest latency routing
import time
import pytest
import litellm
from litellm.caching.caching import DualCache
from litellm.router_strategy.lowest_latency import LowestLatencyLoggingHandler
def test_zero_completion_tokens_no_division_error():
"""
Test that log_success_event handles zero completion tokens without ZeroDivisionError
This tests the fix for issue #12641 where responses with zero completion tokens
(e.g., from Gemini with long contexts) caused ZeroDivisionError
"""
test_cache = DualCache()
lowest_latency_logger = LowestLatencyLoggingHandler(router_cache=test_cache)
deployment_id = "1234"
kwargs = {
"litellm_params": {
"metadata": {
"model_group": "gemini-2.5-flash",
"deployment": "gemini/gemini-2.5-flash",
},
"model_info": {"id": deployment_id},
}
}
# Create a ModelResponse with zero completion tokens (as reported in issue)
response_obj = litellm.ModelResponse(
id="9p13aIGDDNmPmLAP5-23mQQ",
created=1752669685,
model="gemini-2.5-flash",
object="chat.completion",
choices=[
litellm.Choices(
finish_reason="stop",
index=0,
message=litellm.Message(
content=None, role="assistant", tool_calls=None
),
)
],
usage=litellm.Usage(
completion_tokens=0, # This causes the ZeroDivisionError
prompt_tokens=245537,
total_tokens=245537,
),
)
start_time = time.time()
time.sleep(0.1) # Simulate some response time
end_time = time.time()
# This should not raise ZeroDivisionError
try:
lowest_latency_logger.log_success_event(
response_obj=response_obj,
kwargs=kwargs,
start_time=start_time,
end_time=end_time,
)
except ZeroDivisionError:
pytest.fail(
"log_success_event raised ZeroDivisionError with zero completion tokens"
)
# Verify the deployment was logged (even with zero completion tokens)
cached_value = test_cache.get_cache(
key=f"{kwargs['litellm_params']['metadata']['model_group']}_map"
)
assert cached_value is not None
assert deployment_id in cached_value
def test_zero_completion_tokens_with_time_to_first_token():
"""
Test that time_to_first_token calculation also handles zero completion tokens
"""
test_cache = DualCache()
lowest_latency_logger = LowestLatencyLoggingHandler(router_cache=test_cache)
deployment_id = "1234"
kwargs = {
"litellm_params": {
"metadata": {
"model_group": "gemini-2.5-flash",
"deployment": "gemini/gemini-2.5-flash",
},
"model_info": {"id": deployment_id},
"stream": True,
},
"completion_start_time": time.time() + 0.05, # Simulate time to first token
}
# Create a ModelResponse with zero completion tokens
response_obj = litellm.ModelResponse(
usage=litellm.Usage(
completion_tokens=0, prompt_tokens=100000, total_tokens=100000
)
)
start_time = time.time()
time.sleep(0.1)
end_time = time.time()
# This should not raise ZeroDivisionError
try:
lowest_latency_logger.log_success_event(
response_obj=response_obj,
kwargs=kwargs,
start_time=start_time,
end_time=end_time,
)
except ZeroDivisionError:
pytest.fail(
"log_success_event raised ZeroDivisionError with zero completion tokens in streaming"
)
if __name__ == "__main__":
test_zero_completion_tokens_no_division_error()
test_zero_completion_tokens_with_time_to_first_token()
print("All tests passed!")