mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-02 02:11:58 +00:00
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden #43063 stamps used_client_oauth_token into spend-log metadata, so test_async_gcs_pub_sub_v1 failed on main with an extra metadata key * test(ui): give the auto-router threshold save wait room for the availability debounce #42625 keeps Save disabled while a 300ms-debounced availability check runs. This test waits for Save right after the change, so the whole debounce lands inside waitFor's 1s default and it times out under CI load. It is the recurring UI Unit Tests failure on main since #42625 landed * test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default #42870 added both the rule that a served default or standard tier bills at base pricing and records no service_tier, and streamed tests expecting the row to record 'default'. They have failed on every scheduled litellm-e2e run since. The tests now map the served tier to the pricing basis the bill must record and check input is billed at that basis's rate; the messages case registers custom rates so the rate check has something to compare against * test(e2e-ui): wait for the call-id search before hovering the logs row The row the spec hovers is already on the unfiltered first page, so it was found before the search request returned. The search response then re-rendered the table under the mouse, and the Base UI tooltip never opened. Reproduced with Playwright against a local proxy: hovering right after the fill never shows the tooltip, hovering after the search response shows the call id every time * test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off The case picked the cheapest Together row flagged supports_response_schema. DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model the reasoning_effort=none case already exercises, and Together lists it with structured output support * test(integration): read the agent 365 guardrail status by its own name in spend logs The MCP shard runs under xdist against one database, and a sibling file creates a default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so that filter's 'success' entry could land first in guardrail_information and the test read it instead of the agent 365 verdict * test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check gc.collect() inside the caplog window can collect a pending task an earlier test left on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this test's records. The check still counts every LiteLLM logger, and unretrieved task exceptions on this loop still go through the asserted exception handler * test(e2e-ui): fill the create-tag fields inside the dialog #42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description') match two elements and Playwright's strict mode fails the create step * test(integration): run integration proxies with the CI license Multi-worker proxies start each uvicorn worker in a fresh process, so every worker reads the license from its environment. Forward LITELLM_LICENSE into the proxy and test runner environments * ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5 Every pull request saved its own uv, maturin, Rust and Prisma caches, about 4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries within minutes. Pull request jobs then missed every cache, downloaded all dependencies from PyPI and hit the install step timeouts. Pull requests now restore only, and main keeps the caches warm for them. test-linting and check-ui-api-types run only on pull requests and keep saving codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity keybase account, so every upload failed signature verification. 5.5.5 reads it from codecovsecops; the key ID matches the one signing the current CLI * test(unit): join the session-minting thread before collecting the handler asyncio.to_thread resumes the test as soon as the worker sets its result, while the pool thread can still hold the work item and through it the handler. gc.collect() then cannot finalize the handler and the session stays open. A pool that shuts down before the test continues drops that reference * test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early owned_proxy_process released its reserved port and the proxy bound it only after full startup, so another xdist worker or an outgoing connection could take it first and the proxy exited with 'address already in use'. The launch now retries on a fresh port when that happens and stops every failed attempt. uvicorn closes idle keep-alive connections after 5 seconds and httpx expired them at the same 5 seconds, so a request sent right at that mark could reuse a socket the server was closing and get 'Connection reset by peer'. Gateway clients now drop idle connections after 2 seconds * ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build The release profile builds with fat LTO and one codegen unit, so the final link of litellm-cache-s3 runs silently for minutes. Successful builds take 711 to 749 seconds, right at the default 10 minute no-output limit, and about 30% of recent runs were killed there * test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections The proxy retries the database about every 30 seconds and each retry opens roughly one connection, so a 5-connection outage took 3 to 4 retries to clear and recovery landed between 60 and 90 seconds, straddling the test's 80 second reset window. A fixed 10 second outage still refuses the immediate reconnect and recovers on the next retry * ci: move the unit-test uv cache split into a composite action check_workflow_startup_safety sums every setup step's timeout, so the save and restore variants each counted 5 minutes although only one runs. One composite step keeps the setup ceiling at 35 minutes * test(unit): point tiktoken at the bundled cache for every unit test The rust_bridge tokenizer tests loaded o200k_base before any test in their xdist worker had imported default_encoding, so tiktoken fell back to the temp cache and tried to download under pytest-socket. Move the session fixture from litellm_core_utils/conftest.py to the root unit conftest. * test(integration): answer model discovery probes in the hosted_vllm wire tests The router's periodic upstream model info refresh sends GET /v1/models to hosted_vllm deployments, so a wire server that is live during a refresh sees an extra request. Answer the probe with an empty model list and leave it out of the provider-call assertions, matching the responses bridge tests.
334 lines
12 KiB
Python
334 lines
12 KiB
Python
import asyncio
|
|
import base64
|
|
import importlib
|
|
import os
|
|
from collections.abc import Coroutine, Iterator
|
|
from dataclasses import dataclass, field
|
|
from pathlib import Path
|
|
from typing import Final
|
|
|
|
import boto3
|
|
import httpx
|
|
import pytest
|
|
from pytest_socket import enable_socket, socket_allow_hosts
|
|
|
|
HOST_ENVIRONMENT_ALLOWLIST: Final = frozenset(
|
|
(
|
|
"PATH",
|
|
"HOME",
|
|
"USER",
|
|
"LOGNAME",
|
|
"TMPDIR",
|
|
"TEMP",
|
|
"TMP",
|
|
"LANG",
|
|
"LC_ALL",
|
|
"LC_CTYPE",
|
|
"TZ",
|
|
"VIRTUAL_ENV",
|
|
"LITELLM_LOCAL_MODEL_COST_MAP",
|
|
"TIKTOKEN_CACHE_DIR",
|
|
)
|
|
)
|
|
HOST_ENVIRONMENT_ALLOWED_PREFIXES: Final = ("PYTEST_", "PYTHON", "COV_CORE_", "COVERAGE_")
|
|
HOST_ONLY_ENVIRONMENT: Final = frozenset(
|
|
name
|
|
for name in os.environ
|
|
if name not in HOST_ENVIRONMENT_ALLOWLIST and not name.startswith(HOST_ENVIRONMENT_ALLOWED_PREFIXES)
|
|
)
|
|
|
|
os.environ["PYTHON_DOTENV_DISABLED"] = "1"
|
|
os.environ["LITELLM_LOCAL_MODEL_COST_MAP"] = "True"
|
|
|
|
import litellm # noqa: E402 # litellm reads LITELLM_LOCAL_MODEL_COST_MAP at import
|
|
import litellm.router as litellm_router_module # noqa: E402 # same import-time dependency
|
|
import litellm.utils as litellm_utils_module # noqa: E402 # same import-time dependency
|
|
from litellm._logging import ALL_LOGGERS # noqa: E402 # same import-time dependency
|
|
from litellm.anthropic_beta_headers_manager import reload_beta_headers_config # noqa: E402 # same import-time dependency
|
|
from litellm.litellm_core_utils.prompt_templates import factory as prompt_factory_module # noqa: E402 # same import-time dependency
|
|
from litellm.litellm_core_utils.prompt_templates import ( # noqa: E402 # same import-time dependency
|
|
image_handling as image_handling_module,
|
|
)
|
|
from litellm.llms.gemini.chat import transformation as gemini_chat_transformation_module # noqa: E402 # same import-time dependency
|
|
from litellm.llms.custom_httpx.async_client_cleanup import ( # noqa: E402 # same import-time dependency
|
|
close_litellm_async_clients,
|
|
)
|
|
from litellm.proxy.db import tool_registry_writer as tool_registry_writer_module # noqa: E402 # same import-time dependency
|
|
|
|
LOOPBACK_HOSTS: Final = ["127.0.0.1", "::1", "localhost"]
|
|
AMBIENT_AZURE_CREDENTIAL_ENV_VARS: Final = (
|
|
"AZURE_AD_TOKEN",
|
|
"AZURE_TENANT_ID",
|
|
"AZURE_CLIENT_ID",
|
|
"AZURE_CLIENT_SECRET",
|
|
"AZURE_USERNAME",
|
|
"AZURE_PASSWORD",
|
|
)
|
|
AMBIENT_AWS_ENV_VARS: Final = (
|
|
"AWS_PROFILE",
|
|
"AWS_DEFAULT_PROFILE",
|
|
"AWS_CONTAINER_CREDENTIALS_FULL_URI",
|
|
"AWS_CONTAINER_CREDENTIALS_RELATIVE_URI",
|
|
"AWS_SESSION_TOKEN",
|
|
"AWS_ROLE_ARN",
|
|
"AWS_WEB_IDENTITY_TOKEN_FILE",
|
|
"AWS_BEARER_TOKEN_BEDROCK",
|
|
"AWS_REGION_NAME",
|
|
"AWS_DEFAULT_REGION",
|
|
)
|
|
MODULES_WITH_AWS_AUTH_HANDLERS: Final = (
|
|
"litellm.main",
|
|
"litellm.files.main",
|
|
"litellm.rerank_api.main",
|
|
"litellm.realtime_api.main",
|
|
)
|
|
CALLBACK_LISTS: Final = (
|
|
"callbacks",
|
|
"success_callback",
|
|
"failure_callback",
|
|
"input_callback",
|
|
"_async_success_callback",
|
|
"_async_failure_callback",
|
|
"_async_input_callback",
|
|
)
|
|
RESET_TO_NONE_GLOBALS: Final = ("model_fallbacks", "cache")
|
|
RESTORED_GLOBALS: Final = (
|
|
"disable_aiohttp_transport",
|
|
"force_ipv4",
|
|
"drop_params",
|
|
"secret_manager_client",
|
|
"_key_management_system",
|
|
"_key_management_settings",
|
|
"api_base",
|
|
"num_retries",
|
|
"modify_params",
|
|
"ssl_verify",
|
|
"credential_list",
|
|
"model_group_settings",
|
|
"default_internal_user_params",
|
|
"default_team_params",
|
|
"prometheus_emit_stream_label",
|
|
"vector_store_registry",
|
|
"model_cost",
|
|
"cost_margin_config",
|
|
"cost_discount_config",
|
|
"disable_hf_tokenizer_download",
|
|
"disable_copilot_system_to_assistant",
|
|
"cohere_models",
|
|
"anthropic_models",
|
|
"token_counter",
|
|
"initialized_langfuse_clients",
|
|
)
|
|
MODULE_LEVEL_CLIENTS: Final = ("module_level_client", "module_level_aclient")
|
|
SESSION_CLIENTS: Final = ("base_llm_aiohttp_handler", "httpx_client", "aclient", "client")
|
|
ONE_PIXEL_PNG: Final = base64.b64decode(
|
|
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mNkYPhfDwAChwGA60e6kgAAAABJRU5ErkJggg=="
|
|
)
|
|
|
|
|
|
def _allow_loopback_only() -> None:
|
|
socket_allow_hosts(LOOPBACK_HOSTS, allow_unix_socket=True)
|
|
|
|
|
|
_allow_loopback_only()
|
|
|
|
|
|
def pytest_collectstart() -> None:
|
|
_allow_loopback_only()
|
|
|
|
|
|
@pytest.hookimpl(trylast=True)
|
|
def pytest_runtest_setup() -> None:
|
|
_allow_loopback_only()
|
|
|
|
|
|
def _run_coroutine_if_needed(result: object) -> None:
|
|
if not asyncio.iscoroutine(result):
|
|
return
|
|
coroutine: Final[Coroutine[object, object, object]] = result
|
|
try:
|
|
asyncio.run(coroutine)
|
|
except RuntimeError:
|
|
try:
|
|
loop: Final = asyncio.get_running_loop()
|
|
except RuntimeError:
|
|
coroutine.close()
|
|
return
|
|
loop.create_task(coroutine)
|
|
|
|
|
|
def _close_handler_if_needed(handler: object) -> None:
|
|
close: Final = getattr(handler, "close", None)
|
|
if not callable(close):
|
|
return
|
|
_run_coroutine_if_needed(close())
|
|
|
|
|
|
def _reset_aws_auth_caches() -> None:
|
|
modules: Final = tuple(importlib.import_module(name) for name in MODULES_WITH_AWS_AUTH_HANDLERS)
|
|
flushes: Final = (
|
|
getattr(getattr(getattr(module, attr_name), "iam_cache", None), "flush_cache", None)
|
|
for module in modules
|
|
for attr_name in dir(module)
|
|
)
|
|
for flush in filter(callable, flushes):
|
|
flush()
|
|
boto3.DEFAULT_SESSION = None
|
|
|
|
|
|
def _flush_client_caches() -> None:
|
|
litellm.in_memory_llm_clients_cache.flush_cache()
|
|
image_handling_module.in_memory_cache.flush_cache()
|
|
_reset_aws_auth_caches()
|
|
|
|
|
|
@pytest.fixture(autouse=True, scope="session")
|
|
def bundled_tiktoken_cache() -> None:
|
|
importlib.import_module("litellm.litellm_core_utils.default_encoding")
|
|
|
|
|
|
@pytest.fixture(scope="session")
|
|
def isolated_aws_config_files(tmp_path_factory: pytest.TempPathFactory) -> tuple[Path, Path]:
|
|
aws_dir: Final = tmp_path_factory.mktemp("aws-config")
|
|
credentials: Final = aws_dir / "credentials"
|
|
config: Final = aws_dir / "config"
|
|
credentials.write_text("", encoding="utf-8")
|
|
config.write_text("", encoding="utf-8")
|
|
return credentials, config
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def isolate_host_environment(isolated_aws_config_files: tuple[Path, Path]) -> Iterator[None]:
|
|
credentials, config = isolated_aws_config_files
|
|
with pytest.MonkeyPatch.context() as environment:
|
|
for name in HOST_ONLY_ENVIRONMENT:
|
|
environment.delenv(name, raising=False)
|
|
environment.setenv("AWS_SHARED_CREDENTIALS_FILE", str(credentials))
|
|
environment.setenv("AWS_CONFIG_FILE", str(config))
|
|
environment.setenv("AWS_EC2_METADATA_DISABLED", "true")
|
|
for name in AMBIENT_AWS_ENV_VARS:
|
|
environment.delenv(name, raising=False)
|
|
environment.delenv("PROXY_BASE_URL", raising=False)
|
|
environment.setenv("LITELLM_CLI_DISABLE_KEYRING", "1")
|
|
yield
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def isolate_litellm_globals() -> Iterator[None]:
|
|
original_callbacks: Final = {name: list(getattr(litellm, name) or []) for name in CALLBACK_LISTS}
|
|
original_reset: Final = {name: getattr(litellm, name) for name in RESET_TO_NONE_GLOBALS}
|
|
original_restored: Final = {name: getattr(litellm, name) for name in RESTORED_GLOBALS if hasattr(litellm, name)}
|
|
original_clients: Final = {name: litellm.__dict__[name] for name in MODULE_LEVEL_CLIENTS if name in litellm.__dict__}
|
|
original_loggers: Final = {
|
|
logger: (logger.level, logger.disabled, logger.propagate, list(logger.handlers), list(logger.filters))
|
|
for logger in ALL_LOGGERS
|
|
}
|
|
original_tool_policy_registry: Final = tool_registry_writer_module._tool_policy_registry
|
|
_flush_client_caches()
|
|
for name in CALLBACK_LISTS:
|
|
setattr(litellm, name, [])
|
|
for name in RESET_TO_NONE_GLOBALS:
|
|
setattr(litellm, name, None)
|
|
for name in MODULE_LEVEL_CLIENTS:
|
|
litellm.__dict__.pop(name, None)
|
|
tool_registry_writer_module._tool_policy_registry = None
|
|
yield
|
|
_flush_client_caches()
|
|
leaked_clients: Final = tuple(litellm.__dict__.pop(name, None) for name in MODULE_LEVEL_CLIENTS)
|
|
for name, client in zip(MODULE_LEVEL_CLIENTS, leaked_clients):
|
|
if client is not original_clients.get(name):
|
|
_close_handler_if_needed(client)
|
|
litellm.__dict__.update(original_clients)
|
|
for name, value in (original_callbacks | original_reset | original_restored).items():
|
|
setattr(litellm, name, value)
|
|
for logger, (level, disabled, propagate, handlers, filters) in original_loggers.items():
|
|
logger.setLevel(level)
|
|
logger.disabled = disabled
|
|
logger.propagate = propagate
|
|
logger.handlers = handlers
|
|
logger.filters = filters
|
|
tool_registry_writer_module._tool_policy_registry = original_tool_policy_registry
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def isolate_router_model_cost_state() -> Iterator[None]:
|
|
original_live_routers: Final = frozenset(litellm_router_module._live_routers)
|
|
original_runtime_registered_model_cost: Final = {
|
|
model_key: dict(model_value)
|
|
for model_key, model_value in litellm_utils_module._runtime_registered_model_cost.items()
|
|
}
|
|
litellm_utils_module._invalidate_model_cost_lowercase_map()
|
|
yield
|
|
for router in tuple(litellm_router_module._live_routers):
|
|
litellm_router_module._live_routers.discard(router)
|
|
for router in original_live_routers:
|
|
litellm_router_module._live_routers.add(router)
|
|
litellm_utils_module._runtime_registered_model_cost.clear()
|
|
litellm_utils_module._runtime_registered_model_cost.update(original_runtime_registered_model_cost)
|
|
litellm_utils_module._invalidate_model_cost_lowercase_map()
|
|
litellm.get_model_info.cache_clear()
|
|
|
|
|
|
@pytest.fixture
|
|
def local_model_cost_map(monkeypatch: pytest.MonkeyPatch) -> Iterator[None]:
|
|
monkeypatch.setenv("LITELLM_LOCAL_MODEL_COST_MAP", "True")
|
|
monkeypatch.setattr(litellm, "model_cost", litellm.get_model_cost_map(url=""))
|
|
litellm.get_model_info.cache_clear()
|
|
yield
|
|
litellm.get_model_info.cache_clear()
|
|
|
|
|
|
@pytest.fixture
|
|
def local_beta_headers_config(monkeypatch: pytest.MonkeyPatch) -> Iterator[None]:
|
|
monkeypatch.setenv("LITELLM_LOCAL_ANTHROPIC_BETA_HEADERS", "True")
|
|
reload_beta_headers_config()
|
|
yield
|
|
monkeypatch.delenv("LITELLM_LOCAL_ANTHROPIC_BETA_HEADERS", raising=False)
|
|
reload_beta_headers_config()
|
|
|
|
|
|
@dataclass(slots=True)
|
|
class AsyncOnlyImageFetch:
|
|
fetched: list[str] = field(default_factory=list) # mutable-ok: tests assert on the URLs fetched, in order
|
|
base64_png: str = base64.b64encode(ONE_PIXEL_PNG).decode()
|
|
data_url: str = "data:image/png;base64," + base64.b64encode(ONE_PIXEL_PNG).decode()
|
|
|
|
|
|
@pytest.fixture
|
|
def async_only_image_fetch(monkeypatch: pytest.MonkeyPatch) -> AsyncOnlyImageFetch:
|
|
fetch: Final = AsyncOnlyImageFetch()
|
|
|
|
def forbid_sync_fetch(client: object, url: str, **kwargs: object) -> httpx.Response:
|
|
raise litellm.ImageFetchError(f"sync image fetch ran on the event loop: {url}")
|
|
|
|
async def serve_png(client: object, url: str, **kwargs: object) -> httpx.Response:
|
|
fetch.fetched.append(url)
|
|
return httpx.Response(
|
|
200, content=ONE_PIXEL_PNG, headers={"content-type": "image/png"}, request=httpx.Request("GET", url)
|
|
)
|
|
|
|
def forbid_sync_convert(url: str, *args: object, **kwargs: object) -> str:
|
|
if url.startswith(("http://", "https://")):
|
|
raise litellm.ImageFetchError(f"sync convert_url_to_base64 ran on the request path: {url}")
|
|
return url
|
|
|
|
monkeypatch.setattr(image_handling_module, "safe_get", forbid_sync_fetch)
|
|
monkeypatch.setattr(image_handling_module, "async_safe_get", serve_png)
|
|
for module in (image_handling_module, prompt_factory_module, gemini_chat_transformation_module):
|
|
monkeypatch.setattr(module, "convert_url_to_base64", forbid_sync_convert)
|
|
return fetch
|
|
|
|
|
|
@pytest.fixture
|
|
def no_ambient_azure_credentials(monkeypatch: pytest.MonkeyPatch) -> None:
|
|
for name in AMBIENT_AZURE_CREDENTIAL_ENV_VARS:
|
|
monkeypatch.delenv(name, raising=False)
|
|
|
|
|
|
def pytest_sessionfinish() -> None:
|
|
for name in MODULE_LEVEL_CLIENTS:
|
|
_close_handler_if_needed(litellm.__dict__.pop(name, None))
|
|
for name in SESSION_CLIENTS:
|
|
_close_handler_if_needed(getattr(litellm, name, None))
|
|
_run_coroutine_if_needed(close_litellm_async_clients())
|
|
enable_socket()
|