test(e2e): repair two suites broken by intentional behaviour changes

Both of these are e2e assumptions that PRs #31731 and #39532 invalidated, not
product regressions. They have been red in litellm-e2e builds 119-123.

Wildcard readiness probe (6 errors in test_model_access_group_e2e.py)

#31731 made _get_wildcard_models drop a wildcard route from /v1/models
unconditionally; before it, a wildcard with a matching router deployment stayed
in the list and only the no-router / no-deployment fallbacks removed it. The
shared readiness helper polls /v1/models for an exact id match, so registering
openai/gpt-5.4* now times out at model_servable_timeout every run and every
test in the class errors in setup.

return_wildcard_routes=True still re-adds the route, so the poll asks for it.
The flag is a no-op for a concrete model name -- it only ever adds wildcard
entries -- so it is set unconditionally rather than sniffing the name.

Semantic auto-router spend assertion

#39532 bills the routing embedding to the caller's key on purpose, so the key's
spend logs now legitimately carry an openai/text-embedding-3-small row and
_assert_served_only_by rejects it.

Widening the allowlist would have weakened the assertion this test exists for --
that the request reached the target deployment. Instead the embedding row is
split off and asserted separately, which turns the break into coverage for
#39532. The poll gains a predicate so it waits for the embedding row rather than
racing whichever row is written first.
This commit is contained in:
Yuneng Jiang 2026-09-04 13:48:42 -07:00
parent 300d335255
commit 9e0659212a
No known key found for this signature in database
3 changed files with 24 additions and 3 deletions

View file

@ -851,6 +851,15 @@ class ModelListEntry(BaseModel):
id: str
class ModelsListParams(BaseModel):
"""Query for GET /v1/models. A wildcard route such as ``openai/gpt-5.4*`` is
listed only under ``return_wildcard_routes``; without it the route is dropped
and only its expansions remain, so a readiness poll for the pattern itself
never resolves."""
return_wildcard_routes: bool = True
class ModelsListResponse(BaseModel):
"""GET /v1/models on the data plane: the deployments the gateway can actually
serve right now. Used to confirm a freshly created model has propagated from

View file

@ -55,6 +55,7 @@ from models import (
ModelMode,
ModelNewBody,
ModelNewResponse,
ModelsListParams,
ModelsListResponse,
ModelUpdateBody,
OcrBody,
@ -336,7 +337,7 @@ class ProxyClient:
lambda poll_timeout: self.transport.get(
"/v1/models",
headers=headers,
params=NoBody(),
params=ModelsListParams(),
response_type=ModelsListResponse,
timeout=poll_timeout,
),

View file

@ -597,9 +597,20 @@ class TestSemanticAutoRouterResponses:
)
)
assert answer.id, "/v1/responses through the semantic auto-router returned no response id"
rows: Final = proxy.poll_logs_for_key(key, min_rows=1)
rows: Final = proxy.poll_logs_for_key(
key,
min_rows=2,
predicate=lambda logged: any(row.model == EMBEDDING_MODEL for row in logged),
)
embedding_rows: Final = tuple(row for row in rows if row.model == EMBEDDING_MODEL)
assert embedding_rows, (
"the routing embedding was not billed to the caller's key; "
f"spend logs show {tuple(row.model for row in rows)}"
)
_assert_served_only_by(
rows, CHEAP_SERVED | {semantic_auto_router.target}, "semantic auto-router /v1/responses string input"
[row for row in rows if row.model != EMBEDDING_MODEL],
CHEAP_SERVED | {semantic_auto_router.target},
"semantic auto-router /v1/responses string input",
)