litellm/tests/integration/_support/client.py
yuneng-jiang 6ca90b927c
test(ci): repair stale tests and flaky CI infrastructure (#43983)
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden

#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key

* test(ui): give the auto-router threshold save wait room for the availability debounce

#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed

* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default

#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against

* test(e2e-ui): wait for the call-id search before hovering the logs row

The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time

* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off

The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support

* test(integration): read the agent 365 guardrail status by its own name in spend logs

The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict

* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check

gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler

* test(e2e-ui): fill the create-tag fields inside the dialog

#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step

* test(integration): run integration proxies with the CI license

Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments

* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5

Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving

codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI

* test(unit): join the session-minting thread before collecting the handler

asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference

* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early

owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.

uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds

* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build

The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there

* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections

The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry

* ci: move the unit-test uv cache split into a composite action

check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes

* test(unit): point tiktoken at the bundled cache for every unit test

The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.

* test(integration): answer model discovery probes in the hosted_vllm wire tests

The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.
2026-10-01 17:46:43 +00:00

264 lines
11 KiB
Python

from __future__ import annotations
import os
import time
import uuid
from collections.abc import Callable, Iterator, Mapping, Sequence
from contextlib import ExitStack, contextmanager
from dataclasses import dataclass
from hashlib import sha256
from typing import Final, TypeVar
import httpx
from pydantic import JsonValue, TypeAdapter
from tests.integration._support.database import read_rows
JSON_OBJECT: Final = TypeAdapter(dict[str, JsonValue])
GATEWAY_LIMITS: Final = httpx.Limits(keepalive_expiry=2)
T = TypeVar("T")
def object_value(value: JsonValue) -> dict[str, JsonValue]:
return JSON_OBJECT.validate_python(value)
def string_value(value: JsonValue) -> str:
assert isinstance(value, str), f"Expected a string, received {type(value).__name__}"
return value
def delete_key_if_present(candidate: Gateway, key: str) -> None:
digest: Final = sha256(key.encode()).hexdigest()
if read_rows('SELECT token FROM "LiteLLM_VerificationToken" WHERE token=%s', (digest,)):
candidate.post("/key/delete", {"keys": [key]})
assert read_rows('SELECT token FROM "LiteLLM_VerificationToken" WHERE token=%s', (digest,)) == []
def eventually(
read: Callable[[], T],
satisfied: Callable[[T], bool],
seconds: float = 10,
return_last_on_timeout: bool = False,
) -> T:
deadline: Final = time.monotonic() + seconds
while True:
observed: Final = read()
if satisfied(observed):
return observed
if return_last_on_timeout and time.monotonic() >= deadline:
return observed
assert time.monotonic() < deadline, f"State did not converge: {observed!r}"
time.sleep(0.1)
@dataclass(frozen=True, slots=True)
class Gateway:
client: httpx.Client
key: str
upstream_url: str
def request(
self,
method: str,
path: str,
body: Mapping[str, JsonValue] | None = None,
*,
key: str | None = None,
params: Mapping[str, str] | None = None,
headers: Mapping[str, str] | None = None,
) -> httpx.Response:
request_headers: Final = {
"Authorization": f"Bearer {self.key if key is None else key}",
**(headers or {}),
}
return self.client.request(
method,
path,
json=body,
params=params,
headers=request_headers,
)
def request_multipart(
self,
path: str,
fields: Mapping[str, str],
files: Mapping[str, tuple[str, bytes, str]],
*,
key: str | None = None,
) -> httpx.Response:
return self.client.post(
path,
data=fields,
files=files,
headers={"Authorization": f"Bearer {self.key if key is None else key}"},
)
def post(self, path: str, body: Mapping[str, JsonValue], *, key: str | None = None) -> dict[str, JsonValue]:
response: Final = self.request("POST", path, body, key=key)
assert response.status_code == 200, f"POST {path}: {response.status_code} {response.text}"
return JSON_OBJECT.validate_json(response.content)
def get(self, path: str, params: Mapping[str, str] | None = None) -> dict[str, JsonValue]:
response: Final = self.request("GET", path, params=params)
assert response.status_code == 200, f"GET {path}: {response.status_code} {response.text}"
return JSON_OBJECT.validate_json(response.content)
def chat(self, model: str, *, key: str | None = None, text: str = "integration control") -> dict[str, JsonValue]:
return self.post(
"/v1/chat/completions",
{"model": model, "messages": [{"role": "user", "content": text}]},
key=key,
)
@contextmanager
def scenario(self) -> Iterator[Scenario]:
with ExitStack() as cleanups:
yield Scenario(self, cleanups)
@dataclass(frozen=True, slots=True)
class Scenario:
gateway: Gateway
cleanups: ExitStack
def key(self, **fields: JsonValue) -> str:
created: Final = self.gateway.post("/key/generate", fields)
token: Final = string_value(created["key"])
self.cleanups.callback(self.delete_key, token)
return token
def team(self, **fields: JsonValue) -> str:
created: Final = self.gateway.post("/team/new", {"team_alias": f"integration-{uuid.uuid4().hex}", **fields})
identity: Final = string_value(created["team_id"])
self.cleanups.callback(self.delete_team, identity)
return identity
def delete_team(self, identity: str) -> None:
self.gateway.post("/team/delete", {"team_ids": [identity]})
assert read_rows('SELECT team_id FROM "LiteLLM_TeamTable" WHERE team_id = %s', (identity,)) == []
def project(self, team_id: str, **fields: JsonValue) -> str:
created: Final = self.gateway.post(
"/project/new", {"team_id": team_id, "project_alias": f"integration-{uuid.uuid4().hex}", **fields}
)
identity: Final = string_value(created["project_id"])
self.cleanups.callback(self.delete_project, identity)
return identity
def delete_project(self, identity: str) -> None:
response: Final = self.gateway.request("DELETE", "/project/delete", {"project_ids": [identity]})
assert response.status_code == 200, response.text
assert read_rows('SELECT project_id FROM "LiteLLM_ProjectTable" WHERE project_id = %s', (identity,)) == []
def budget(self, **fields: JsonValue) -> str:
created: Final = self.gateway.post("/budget/new", fields)
identity: Final = string_value(created["budget_id"])
self.cleanups.callback(self.delete_budget, identity)
return identity
def delete_budget(self, identity: str) -> None:
self.gateway.post("/budget/delete", {"id": identity})
assert read_rows('SELECT budget_id FROM "LiteLLM_BudgetTable" WHERE budget_id = %s', (identity,)) == []
def organization(self, **fields: JsonValue) -> str:
created: Final = self.gateway.post(
"/organization/new", {"organization_alias": f"integration-{uuid.uuid4().hex}", **fields}
)
identity: Final = string_value(created["organization_id"])
self.cleanups.callback(self.delete_organization, identity, string_value(created["budget_id"]))
return identity
def delete_organization(self, identity: str, budget_id: str) -> None:
response: Final = self.gateway.request("DELETE", "/organization/delete", {"organization_ids": [identity]})
assert response.status_code == 200, response.text
assert (
read_rows('SELECT organization_id FROM "LiteLLM_OrganizationTable" WHERE organization_id = %s', (identity,))
== []
)
self.delete_budget(budget_id)
def org_member(self, organization_id: str, role: str) -> str:
user_id: Final = self.user(user_role="internal_user")
self.gateway.post(
"/organization/member_add",
{"organization_id": organization_id, "member": {"role": role, "user_id": user_id}},
)
return user_id
def user(self, **fields: JsonValue) -> str:
created: Final = self.gateway.post(
"/user/new", {"user_id": f"integration-{uuid.uuid4().hex}", "auto_create_key": False, **fields}
)
identity: Final = string_value(created["user_id"])
self.cleanups.callback(self.delete_user, identity)
return identity
def delete_user(self, identity: str) -> None:
response: Final = self.gateway.request("POST", "/user/delete", {"user_ids": [identity]})
assert response.status_code == 200 and response.json() == 1, response.text
assert read_rows('SELECT user_id FROM "LiteLLM_UserTable" WHERE user_id = %s', (identity,)) == []
def member(self, team_id: str, role: str = "user") -> str:
"""Create an internal user and add them to ``team_id``; deleting the user later removes the membership."""
user_id: Final = self.user(user_role="internal_user")
self.gateway.post("/team/member_add", {"team_id": team_id, "member": {"role": role, "user_id": user_id}})
return user_id
def delete_key(self, token: str) -> None:
self.gateway.post("/key/delete", {"keys": [token]})
hashed: Final = sha256(token.encode()).hexdigest()
assert read_rows('SELECT token FROM "LiteLLM_VerificationToken" WHERE token = %s', (hashed,)) == []
info: Final = object_value(self.gateway.get("/key/info", {"key": hashed})["info"])
assert info["status"] == "deleted", f"Deleted key still served as live: {info['status']}"
def delete_model(self, identity: str) -> None:
self.gateway.post("/model/delete", {"id": identity})
entries: Final = self.gateway.get("/model/info")["data"]
assert isinstance(entries, list)
assert all(object_value(object_value(entry)["model_info"])["id"] != identity for entry in entries)
assert read_rows('SELECT model_id FROM "LiteLLM_ProxyModelTable" WHERE model_id = %s', (identity,)) == []
def model(self, *, model_info: Mapping[str, JsonValue] | None = None, **parameters: JsonValue) -> str:
name: Final = f"integration-{uuid.uuid4().hex}"
created: Final = self.gateway.post(
"/model/new",
{
"model_name": name,
"litellm_params": {
"model": "openai/gpt-4o-mini",
"api_key": "integration-provider-key",
"api_base": f"{self.gateway.upstream_url}/v1",
**parameters,
},
"model_info": dict(model_info) if model_info is not None else {},
},
)
identity: Final = string_value(object_value(created["model_info"])["id"])
self.cleanups.callback(self.delete_model, identity)
return name
@contextmanager
def gateway_from_environment() -> Iterator[Gateway]:
url: Final = os.environ["INTEGRATION_PROXY_URL"]
upstream: Final = os.environ["INTEGRATION_UPSTREAM_URL"]
with httpx.Client(base_url=url, timeout=15, trust_env=False, limits=GATEWAY_LIMITS) as client:
yield Gateway(client, os.environ["INTEGRATION_MASTER_KEY"], upstream)
def _set_team_admin_permissions(gateway: Gateway, fields: Sequence[str]) -> None:
response: Final = gateway.request("PATCH", "/update/ui_settings", {"team_admin_editable_team_fields": list(fields)})
assert response.status_code == 200, response.text
@contextmanager
def team_admin_permissions(gateway: Gateway, fields: Sequence[str]) -> Iterator[None]:
"""Grant team admins ``fields`` proxy-wide for the block, then restore the prior grant."""
original: Final = object_value(gateway.get("/get/ui_settings")["values"]).get("team_admin_editable_team_fields")
_set_team_admin_permissions(gateway, fields)
try:
yield
finally:
_set_team_admin_permissions(gateway, [str(field) for field in original] if isinstance(original, list) else ())