litellm/tests/unit/test_batch_completion_models_all_responses.py
yuneng-jiang f6882246d4
test: move tests/test_litellm root and small trees into tests/unit (#43186)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 11:30:43 -07:00

122 lines
3.7 KiB
Python

import concurrent.futures
import litellm
from litellm.batch_completion.main import batch_completion_models_all_responses
def test_batch_completion_models_all_responses_submits_before_waiting(monkeypatch):
"""
Regression test for issue #20704.
Ensures all model calls are submitted to the thread pool before waiting on results.
"""
models = ["model-a", "model-b", "model-c"]
called_models = []
class _AssertingFuture:
def __init__(self, result, executor, expected_submissions):
self._result = result
self._executor = executor
self._expected_submissions = expected_submissions
def result(self):
if self._executor.submit_count != self._expected_submissions:
raise AssertionError(
"Not all model calls were submitted before waiting"
)
return self._result
class _RecordingThreadPoolExecutor:
def __init__(self, max_workers, *args, **kwargs):
self.max_workers = max_workers
self.submit_count = 0
def __enter__(self):
return self
def __exit__(self, exc_type, exc, tb):
return False
def submit(self, fn, *args, **kwargs):
self.submit_count += 1
result = fn(*args, **kwargs)
return _AssertingFuture(
result=result,
executor=self,
expected_submissions=len(models),
)
def _mock_completion(*args, model, **kwargs):
called_models.append(model)
return {"model": model}
monkeypatch.setattr(litellm, "completion", _mock_completion)
monkeypatch.setattr(
concurrent.futures, "ThreadPoolExecutor", _RecordingThreadPoolExecutor
)
responses = batch_completion_models_all_responses(
models=models,
messages=[{"role": "user", "content": "hello"}],
)
assert sorted(called_models) == sorted(models)
assert len(responses) == len(models)
assert sorted(response["model"] for response in responses) == sorted(models)
def test_batch_completion_models_all_responses_continues_on_model_error(monkeypatch):
models = ["model-a", "model-error", "model-b"]
def _mock_completion(*args, model, **kwargs):
if model == "model-error":
raise RuntimeError("simulated model failure")
return {"model": model}
monkeypatch.setattr(litellm, "completion", _mock_completion)
responses = batch_completion_models_all_responses(
models=models,
messages=[{"role": "user", "content": "hello"}],
)
assert len(responses) == 2
assert sorted(response["model"] for response in responses) == ["model-a", "model-b"]
def test_batch_completion_models_all_responses_returns_empty_for_empty_models(
monkeypatch,
):
called = False
def _mock_completion(*args, model, **kwargs):
nonlocal called
called = True
return {"model": model}
monkeypatch.setattr(litellm, "completion", _mock_completion)
responses = batch_completion_models_all_responses(
models=[],
messages=[{"role": "user", "content": "hello"}],
)
assert responses == []
assert called is False
def test_batch_completion_models_all_responses_accepts_single_model_string(monkeypatch):
called_models = []
def _mock_completion(*args, model, **kwargs):
called_models.append(model)
return {"model": model}
monkeypatch.setattr(litellm, "completion", _mock_completion)
responses = batch_completion_models_all_responses(
models="model-a",
messages=[{"role": "user", "content": "hello"}],
)
assert called_models == ["model-a"]
assert responses == [{"model": "model-a"}]