litellm/tests/unit/test_router_retry_backoff_headers.py
yuneng-jiang f6882246d4
test: move tests/test_litellm root and small trees into tests/unit (#43186)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 11:30:43 -07:00

88 lines
2.7 KiB
Python

"""
Tests for router retry backoff behavior.
"""
from unittest.mock import patch
import httpx
import pytest
import litellm
from litellm import Router
@pytest.mark.asyncio
async def test_retry_backoff_uses_current_exception_headers():
"""
Ensure retry backoff uses the current retry exception, not the initial one.
"""
router = Router(
model_list=[
{
"model_name": "gpt-3.5-turbo",
"litellm_params": {
"model": "gpt-3.5-turbo",
"api_key": "sk-test",
},
}
],
num_retries=2,
)
first_error = litellm.RateLimitError(
message="Rate limited on first attempt",
model="gpt-3.5-turbo",
llm_provider="openai",
)
first_error.litellm_response_headers = httpx.Headers({"retry-after": "1"})
second_error = litellm.RateLimitError(
message="Rate limited on second attempt",
model="gpt-3.5-turbo",
llm_provider="openai",
)
second_error.litellm_response_headers = httpx.Headers({"retry-after": "15"})
third_error = litellm.RateLimitError(
message="Rate limited on third attempt",
model="gpt-3.5-turbo",
llm_provider="openai",
)
third_error.litellm_response_headers = httpx.Headers({"retry-after": "30"})
raised_errors = [first_error, second_error, third_error]
captured_backoff_errors = []
async def mock_make_call(*args, **kwargs):
raise raised_errors.pop(0)
def mock_time_to_sleep_before_retry(*args, **kwargs):
captured_backoff_errors.append(kwargs["e"])
return 0.01
with patch.object(router, "make_call", side_effect=mock_make_call):
with patch.object(
router,
"_async_get_healthy_deployments",
return_value=(
[{"model_info": {"id": "test-id"}}],
[{"model_info": {"id": "test-id"}}],
),
):
with patch.object(
router,
"_time_to_sleep_before_retry",
side_effect=mock_time_to_sleep_before_retry,
):
with pytest.raises(litellm.RateLimitError):
await router.acompletion(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "Hello"}],
)
# Router computes backoff once after the initial failure, then once per failed retry.
# With num_retries=2 and all attempts failing, that's 1 + 2 = 3 invocations.
assert len(captured_backoff_errors) == router.num_retries + 1
assert captured_backoff_errors[0] is first_error
assert captured_backoff_errors[1] is second_error
assert captured_backoff_errors[2] is third_error