mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-06 02:48:13 +00:00
test(e2e): add parameterized inference endpoint matrix over providers and auth modes
One scenario per non-management inference endpoint, run over every catalog provider that serves it and every way the proxy can hold that provider's secret (os.environ reference, literal value, stored credential). Cells carry exact coverage registry ids and the OpenAI and Anthropic cells join the record/replay lane. The provider edge stops recording or owing the proxy's scheduled GET /v1/models discovery, and ProxyClient gains register_models so the matrix pays the propagation budget once per run instead of once per cell Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
This commit is contained in:
parent
c21ab96741
commit
a11037eb63
8 changed files with 685 additions and 16 deletions
|
|
@ -134,10 +134,20 @@ One sharp edge: a replayed response reuses the recorded provider response id, an
|
|||
|
||||
Another sharp edge, same root: record and replay derive every per-test token deterministically (the model name included, so a replay regenerates the exact requests the record run sent), which means an edge-wired deployment left in the database by an interrupted earlier run carries the same model name as the fresh one the current run registers. The proxy then holds two deployments under one model group and load-balances across both, and because the leftover's `api_base` points at the earlier run's edge process, which is gone, the calls that land on it fail with a connection error that reads like a transport bug rather than the stale row it is. Give each record or replay run a fresh database, or let a run finish so its own teardown deletes what it registered, and never reuse one long-lived proxy across back-to-back record/replay sessions. CI hands every job its own empty database and its own proxy, so it never sees this
|
||||
|
||||
Replay answers any provider call that drifted from the recording with an HTTP 599 whose body names the computed and closest recorded keys, so the test fails loudly instead of silently going live, and a bundle older than seven days fails at collection time naming its age; either way the fix is to re-record. Only tests that register edge-wired deployments participate: everything else hits its provider live in every mode, so record exactly the suite you replay. If the proxy runs in a container, set `E2E_PROVIDER_EDGE_ADVERTISE_HOST` (e.g. `host.docker.internal`) so the api_base the proxy stores can reach the edge on the pytest host, and `E2E_PROVIDER_EDGE_BIND_HOST=0.0.0.0` so the edge accepts it. The suites wired to the edge today are `quota_management/spend_tracking/test_provider_edge_spend_e2e.py`, `llm_translation/test_chat_completions_contract_e2e.py`, the OpenAI registrations in `llm_translation/test_embeddings_endpoint_e2e.py`, the Anthropic tests in `llm_translation/test_messages_e2e.py`, streamed and not, and the OpenAI batch deployment behind `batches/`. A streamed response replays as the chunk sequence the provider sent rather than one buffered body. See `AGENTS.md` in this directory for the bundle format, the edge design, and the current limits (Bedrock). The scheduled CI record/replay lane is described above
|
||||
Replay answers any provider call that drifted from the recording with an HTTP 599 whose body names the computed and closest recorded keys, so the test fails loudly instead of silently going live, and a bundle older than seven days fails at collection time naming its age; either way the fix is to re-record. Only tests that register edge-wired deployments participate: everything else hits its provider live in every mode, so record exactly the suite you replay. If the proxy runs in a container, set `E2E_PROVIDER_EDGE_ADVERTISE_HOST` (e.g. `host.docker.internal`) so the api_base the proxy stores can reach the edge on the pytest host, and `E2E_PROVIDER_EDGE_BIND_HOST=0.0.0.0` so the edge accepts it. The suites wired to the edge today are `quota_management/spend_tracking/test_provider_edge_spend_e2e.py`, `llm_translation/test_chat_completions_contract_e2e.py`, the OpenAI registrations in `llm_translation/test_embeddings_endpoint_e2e.py`, the Anthropic tests in `llm_translation/test_messages_e2e.py`, streamed and not, the OpenAI and Anthropic cells of `llm_translation/test_endpoint_matrix_e2e.py`, and the OpenAI batch deployment behind `batches/`. A streamed response replays as the chunk sequence the provider sent rather than one buffered body. See `AGENTS.md` in this directory for the bundle format, the edge design, and the current limits (Bedrock). The scheduled CI record/replay lane is described above
|
||||
|
||||
Tests marked `@pytest.mark.e2e` hard-fail when no proxy answers `/health/liveliness`, so a run that goes red with `No live proxy` at setup means the proxy isn't up; they never skip for a missing proxy, so an absent proxy can't be mistaken for a pass
|
||||
|
||||
### The endpoint matrix
|
||||
|
||||
`llm_translation/test_endpoint_matrix_e2e.py` runs one scenario per inference endpoint (`/chat/completions`, `/v1/completions`, `/v1/messages`, `/v1/responses`, `/embeddings`, `/v1/images/generations`, `/v1/images/edits`, `/v1/audio/speech`, `/v1/audio/transcriptions`, `/v1/moderations`, streamed and not where the endpoint streams) over every provider in `llm_translation/endpoint_matrix.py` that serves it, and over every way the proxy can hold that provider's secret: `os.environ/` references, literal values, and a stored `litellm_credential_name`. Each cell is one pytest parameter carrying its exact registry id, so a translation bug fixed on `/chat/completions` but not on `/v1/messages` shows up as one red cell next to green ones instead of a test gap. Adding a provider is one `Provider` entry in the catalog naming its route, credential environment variables, and backend model per endpoint; it needs no new test code. `E2E_MATRIX_PROVIDERS` and `E2E_MATRIX_AUTH_MODES` narrow a run to a comma-separated subset of catalog routes and auth modes, unknown names fail at collection, and a provider whose secret is missing from the runner fails its inline and stored-credential cells by variable name rather than skipping
|
||||
|
||||
```bash
|
||||
E2E_MATRIX_PROVIDERS=openai,anthropic E2E_MATRIX_AUTH_MODES=env_ref uv run pytest tests/e2e/llm_translation/test_endpoint_matrix_e2e.py
|
||||
```
|
||||
|
||||
Providers with a generic edge mount (OpenAI and Anthropic today) are edge-wired and marked `replayable`, so their cells join the scheduled record/replay lane and run with zero provider egress on weekdays; the rest hit their provider live in every fixture mode. The module registers the deployments its selected cells need as one `register_models` batch in a module fixture, so a run pays the data-plane reload and propagation budget once rather than once per cell, which is what keeps a 60-cell replay in the tens of seconds
|
||||
|
||||
## What a complete test looks like
|
||||
|
||||
A feature test is complete only when it walks the feature end to end, in this order
|
||||
|
|
|
|||
|
|
@ -97,3 +97,16 @@
|
|||
- {id: llm.messages.together_ai.basic.stream.works, module: llm, tier: P1, subject_endpoint: messages, route: together_ai, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_together_ai_e2e.py", rationale: "Together over /v1/messages streaming"}
|
||||
- {id: llm.messages.together_ai.tool_use.nonstream.works, module: llm, tier: P1, subject_endpoint: messages, route: together_ai, capability: tool_use, streaming: nonstream, assertions: [works], source: "llm_translation/test_together_ai_e2e.py", rationale: "Together tool calls over /v1/messages"}
|
||||
- {id: llm.messages.together_ai.multi_turn.nonstream.works, module: llm, tier: P1, subject_endpoint: messages, route: together_ai, capability: multi_turn, streaming: nonstream, assertions: [works], source: "llm_translation/test_together_ai_e2e.py", rationale: "Together tool result round trip over /v1/messages"}
|
||||
- {id: llm.chat_completions.together_ai.basic.nonstream.works, module: llm, tier: P0, subject_endpoint: chat_completions, route: together_ai, capability: basic, streaming: nonstream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Together chat, plain visible answer over /chat/completions; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.chat_completions.together_ai.basic.stream.works, module: llm, tier: P0, subject_endpoint: chat_completions, route: together_ai, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Together chat streaming; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.chat_completions.gemini.basic.stream.works, module: llm, tier: P0, subject_endpoint: chat_completions, route: gemini, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Gemini chat streaming; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.messages.openai.basic.nonstream.works, module: llm, tier: P0, subject_endpoint: messages, route: openai, capability: basic, streaming: nonstream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "OpenAI over /v1/messages translation; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.messages.openai.basic.stream.works, module: llm, tier: P0, subject_endpoint: messages, route: openai, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "OpenAI over /v1/messages streaming; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.messages.gemini.basic.nonstream.works, module: llm, tier: P0, subject_endpoint: messages, route: gemini, capability: basic, streaming: nonstream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Gemini over /v1/messages translation; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.messages.gemini.basic.stream.works, module: llm, tier: P0, subject_endpoint: messages, route: gemini, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Gemini over /v1/messages streaming; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.messages.together_ai.basic.nonstream.works, module: llm, tier: P0, subject_endpoint: messages, route: together_ai, capability: basic, streaming: nonstream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Together over /v1/messages translation; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.responses.anthropic.basic.stream.works, module: llm, tier: P0, subject_endpoint: responses, route: anthropic, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Responses streaming w/ Anthropic translation; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.responses.bedrock_converse.basic.stream.works, module: llm, tier: P0, subject_endpoint: responses, route: bedrock_converse, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Responses streaming w/ Bedrock Converse; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.responses.vertex.basic.stream.works, module: llm, tier: P0, subject_endpoint: responses, route: vertex, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Responses streaming w/ Vertex; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.responses.gemini.basic.nonstream.works, module: llm, tier: P0, subject_endpoint: responses, route: gemini, capability: basic, streaming: nonstream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Responses w/ Gemini translation; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
- {id: llm.responses.gemini.basic.stream.works, module: llm, tier: P0, subject_endpoint: responses, route: gemini, capability: basic, streaming: stream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Responses streaming w/ Gemini; same request and assertion as every other endpoint/provider cell in the matrix"}
|
||||
|
|
|
|||
|
|
@ -92,3 +92,4 @@
|
|||
- {id: llm.audio_transcriptions.nvidia_riva.basic.nonstream.works, module: llm, tier: P2, subject_endpoint: audio_transcriptions, route: openai, capability: basic, streaming: nonstream, assertions: [works], source: "nvidia_riva/audio_transcription/handler.py", rationale: "NVIDIA Riva (smoke)"}
|
||||
- {id: llm.moderations.openai.basic.nonstream.works, module: llm, tier: P1, subject_endpoint: moderations, route: openai, capability: basic, streaming: nonstream, assertions: [works], source: "proxy_server.py", rationale: "OpenAI moderations (only provider)"}
|
||||
- {id: llm.moderations.openai.input_validation.nonstream.works, module: llm, tier: P1, subject_endpoint: moderations, route: openai, capability: input_validation, streaming: nonstream, assertions: [works], source: "vendor strategy §9.8 / LIT-4778", rationale: "Moderations missing input rejected"}
|
||||
- {id: llm.embeddings.gemini.basic.nonstream.works, module: llm, tier: P0, subject_endpoint: embeddings, route: gemini, capability: basic, streaming: nonstream, assertions: [works], source: "llm_translation/test_endpoint_matrix_e2e.py", rationale: "Gemini embeddings; same request and assertion as every other embeddings cell in the matrix"}
|
||||
|
|
|
|||
213
tests/e2e/llm_translation/endpoint_matrix.py
Normal file
213
tests/e2e/llm_translation/endpoint_matrix.py
Normal file
|
|
@ -0,0 +1,213 @@
|
|||
"""Provider catalog behind test_endpoint_matrix_e2e.py.
|
||||
|
||||
One `Provider` row per route the matrix drives. Adding a provider is adding a row
|
||||
here (its credential fields, its backend model per endpoint family, and the edge
|
||||
mount when the provider edge can record and replay it) plus the registry cells
|
||||
those combinations claim. The test module never names a provider.
|
||||
|
||||
`E2E_MATRIX_PROVIDERS` and `E2E_MATRIX_AUTH_MODES` narrow the default of every
|
||||
catalog row and every auth mode; an unknown name fails collection instead of
|
||||
quietly selecting nothing.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
from collections.abc import Mapping
|
||||
from dataclasses import dataclass
|
||||
from types import MappingProxyType
|
||||
from typing import Final, Literal, assert_never
|
||||
|
||||
from coverage_registry.schema import LlmRoute
|
||||
from models import LiteLLMParamsBody
|
||||
|
||||
type MatrixEndpoint = Literal[
|
||||
"chat_completions",
|
||||
"completions",
|
||||
"messages",
|
||||
"responses",
|
||||
"embeddings",
|
||||
"images_generations",
|
||||
"images_edits",
|
||||
"audio_speech",
|
||||
"audio_transcriptions",
|
||||
"moderations",
|
||||
]
|
||||
type Streaming = Literal["stream", "nonstream"]
|
||||
type AuthMode = Literal["env_ref", "inline", "stored_credential"]
|
||||
|
||||
AUTH_MODES: Final[tuple[AuthMode, ...]] = ("env_ref", "inline", "stored_credential")
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class Backend:
|
||||
model: str
|
||||
params: Mapping[str, str] = MappingProxyType({})
|
||||
id_route: str | None = None
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class Provider:
|
||||
route: LlmRoute
|
||||
credential: Mapping[str, str]
|
||||
backends: Mapping[MatrixEndpoint, Backend]
|
||||
edge_mount: str | None = None
|
||||
edge_path: str = ""
|
||||
|
||||
def id_route(self, endpoint: MatrixEndpoint) -> str:
|
||||
return self.backends[endpoint].id_route or self.route
|
||||
|
||||
def replayable(self) -> bool:
|
||||
return self.edge_mount is not None
|
||||
|
||||
|
||||
_ANTHROPIC_HAIKU: Final = "claude-haiku-4-5"
|
||||
_BEDROCK_HAIKU: Final = "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0"
|
||||
_GEMINI_FLASH: Final = "gemini-3.8-flash"
|
||||
_TOGETHER_LLAMA: Final = "together_ai/meta-llama/Llama-3.3-70B-Instruct-Turbo"
|
||||
_VERTEX_GLOBAL: Final[Mapping[str, str]] = MappingProxyType({"vertex_location": "global"})
|
||||
_VERTEX_US_CENTRAL: Final[Mapping[str, str]] = MappingProxyType({"vertex_location": "us-central1"})
|
||||
_BEDROCK_REGION: Final[Mapping[str, str]] = MappingProxyType({"aws_region_name": "us-east-1"})
|
||||
|
||||
PROVIDERS: Final[tuple[Provider, ...]] = (
|
||||
Provider(
|
||||
route="openai",
|
||||
credential=MappingProxyType({"api_key": "OPENAI_API_KEY"}),
|
||||
edge_mount="openai",
|
||||
edge_path="/v1",
|
||||
backends=MappingProxyType(
|
||||
{
|
||||
"chat_completions": Backend("openai/gpt-5.4-mini"),
|
||||
"completions": Backend("openai/gpt-3.5-turbo-instruct"),
|
||||
"messages": Backend("openai/gpt-5.4-mini"),
|
||||
"responses": Backend("openai/gpt-5.4-mini"),
|
||||
"embeddings": Backend("openai/text-embedding-3-small"),
|
||||
"images_generations": Backend("openai/gpt-image-1-mini"),
|
||||
"images_edits": Backend("openai/gpt-image-1-mini"),
|
||||
"audio_speech": Backend("openai/gpt-4o-mini-tts"),
|
||||
"audio_transcriptions": Backend("openai/gpt-4o-mini-transcribe"),
|
||||
"moderations": Backend("openai/omni-moderation-latest"),
|
||||
}
|
||||
),
|
||||
),
|
||||
Provider(
|
||||
route="anthropic",
|
||||
credential=MappingProxyType({"api_key": "ANTHROPIC_API_KEY"}),
|
||||
edge_mount="anthropic",
|
||||
backends=MappingProxyType(
|
||||
{
|
||||
"chat_completions": Backend(f"anthropic/{_ANTHROPIC_HAIKU}"),
|
||||
"messages": Backend(f"anthropic/{_ANTHROPIC_HAIKU}"),
|
||||
"responses": Backend(f"anthropic/{_ANTHROPIC_HAIKU}"),
|
||||
}
|
||||
),
|
||||
),
|
||||
Provider(
|
||||
route="gemini",
|
||||
credential=MappingProxyType({"api_key": "GEMINI_API_KEY"}),
|
||||
backends=MappingProxyType(
|
||||
{
|
||||
"chat_completions": Backend(f"gemini/{_GEMINI_FLASH}"),
|
||||
"messages": Backend(f"gemini/{_GEMINI_FLASH}"),
|
||||
"responses": Backend(f"gemini/{_GEMINI_FLASH}"),
|
||||
"embeddings": Backend("gemini/gemini-embedding-001"),
|
||||
}
|
||||
),
|
||||
),
|
||||
Provider(
|
||||
route="vertex",
|
||||
credential=MappingProxyType(
|
||||
{"vertex_credentials": "VERTEXAI_CREDENTIALS", "vertex_project": "VERTEXAI_PROJECT"}
|
||||
),
|
||||
backends=MappingProxyType(
|
||||
{
|
||||
"chat_completions": Backend(f"vertex_ai/{_GEMINI_FLASH}", _VERTEX_GLOBAL),
|
||||
"messages": Backend(f"vertex_ai/{_GEMINI_FLASH}", _VERTEX_GLOBAL),
|
||||
"responses": Backend(f"vertex_ai/{_GEMINI_FLASH}", _VERTEX_GLOBAL),
|
||||
"embeddings": Backend("vertex_ai/text-embedding-005", _VERTEX_US_CENTRAL),
|
||||
}
|
||||
),
|
||||
),
|
||||
Provider(
|
||||
route="bedrock_converse",
|
||||
credential=MappingProxyType({"api_key": "AWS_BEARER_TOKEN_BEDROCK"}),
|
||||
backends=MappingProxyType(
|
||||
{
|
||||
"chat_completions": Backend(_BEDROCK_HAIKU, _BEDROCK_REGION),
|
||||
"messages": Backend(_BEDROCK_HAIKU, _BEDROCK_REGION),
|
||||
"responses": Backend(_BEDROCK_HAIKU, _BEDROCK_REGION),
|
||||
"embeddings": Backend("bedrock/amazon.titan-embed-text-v2:0", _BEDROCK_REGION, id_route="bedrock"),
|
||||
}
|
||||
),
|
||||
),
|
||||
Provider(
|
||||
route="together_ai",
|
||||
credential=MappingProxyType({"api_key": "TOGETHER_API_KEY"}),
|
||||
backends=MappingProxyType(
|
||||
{
|
||||
"chat_completions": Backend(_TOGETHER_LLAMA),
|
||||
"messages": Backend(_TOGETHER_LLAMA),
|
||||
}
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def selected_providers() -> tuple[Provider, ...]:
|
||||
by_route: Final = {provider.route: provider for provider in PROVIDERS}
|
||||
return tuple(by_route[route] for route in _selection("E2E_MATRIX_PROVIDERS", tuple(by_route)))
|
||||
|
||||
|
||||
def selected_auth_modes() -> tuple[AuthMode, ...]:
|
||||
chosen: Final = _selection("E2E_MATRIX_AUTH_MODES", AUTH_MODES)
|
||||
return tuple(mode for mode in AUTH_MODES if mode in chosen)
|
||||
|
||||
|
||||
def _selection(variable: str, known: tuple[str, ...]) -> tuple[str, ...]:
|
||||
raw: Final = os.environ.get(variable, "").strip()
|
||||
if not raw:
|
||||
return known
|
||||
chosen: Final = tuple(name.strip() for name in raw.split(",") if name.strip())
|
||||
unknown: Final = tuple(name for name in chosen if name not in known)
|
||||
if unknown:
|
||||
raise ValueError(f"{variable} names unknown entries {unknown}; known: {', '.join(known)}")
|
||||
return chosen
|
||||
|
||||
|
||||
def credential_values(provider: Provider) -> Mapping[str, str]:
|
||||
"""The provider's real secrets as read from this process's environment, for the
|
||||
auth modes that hand the proxy literal values instead of `os.environ/` references.
|
||||
Missing variables fail by name rather than registering a deployment that will
|
||||
fail later with a less specific provider error."""
|
||||
missing: Final = tuple(env for env in provider.credential.values() if not os.environ.get(env))
|
||||
if missing:
|
||||
raise RuntimeError(f"{provider.route} inline/stored auth needs {', '.join(missing)} in the test environment")
|
||||
return MappingProxyType({field: os.environ[env] for field, env in provider.credential.items()})
|
||||
|
||||
|
||||
def deployment_params(
|
||||
provider: Provider,
|
||||
endpoint: MatrixEndpoint,
|
||||
auth_mode: AuthMode,
|
||||
*,
|
||||
edge_base: str | None,
|
||||
credential_name: str | None,
|
||||
) -> LiteLLMParamsBody:
|
||||
backend: Final = provider.backends[endpoint]
|
||||
auth: Final = _auth_fields(provider, auth_mode, credential_name)
|
||||
api_base: Final = {} if edge_base is None else {"api_base": f"{edge_base}{provider.edge_path}"}
|
||||
return LiteLLMParamsBody.model_validate({"model": backend.model, **backend.params, **auth, **api_base})
|
||||
|
||||
|
||||
def _auth_fields(provider: Provider, auth_mode: AuthMode, credential_name: str | None) -> Mapping[str, str]:
|
||||
match auth_mode:
|
||||
case "env_ref":
|
||||
return MappingProxyType({field: f"os.environ/{env}" for field, env in provider.credential.items()})
|
||||
case "inline":
|
||||
return credential_values(provider)
|
||||
case "stored_credential":
|
||||
if credential_name is None:
|
||||
raise ValueError("stored_credential auth needs the name of the credential registered for the case")
|
||||
return MappingProxyType({"litellm_credential_name": credential_name})
|
||||
case _:
|
||||
assert_never(auth_mode)
|
||||
376
tests/e2e/llm_translation/test_endpoint_matrix_e2e.py
Normal file
376
tests/e2e/llm_translation/test_endpoint_matrix_e2e.py
Normal file
|
|
@ -0,0 +1,376 @@
|
|||
"""One scenario per non-management inference endpoint, run over every catalog
|
||||
provider that serves it and every way the proxy can hold that provider's secret.
|
||||
|
||||
A cell is (endpoint case, provider, auth mode). The case knows how to call the
|
||||
endpoint and what a meaningful answer looks like; the provider knows its backend
|
||||
model and credential fields; the auth mode decides whether the deployment carries
|
||||
`os.environ/` references, literal values, or a `litellm_credential_name`. The
|
||||
same scenario therefore hits `/v1/messages`, `/v1/responses`, `/chat/completions`
|
||||
and the rest identically, so a translation bug fixed on one endpoint but not
|
||||
another shows up as one red cell next to green ones.
|
||||
|
||||
OpenAI and Anthropic cells are edge-wired and replayable; every other provider
|
||||
runs live in every fixture mode. Streaming cases only assert grammar and content,
|
||||
never provider timing, so replay stays fast.
|
||||
|
||||
The selected cells' deployments are registered as one batch before the first cell
|
||||
runs, so the module pays the data-plane reload budget once instead of once per
|
||||
cell; each cell still calls through its own fresh virtual key.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
from collections.abc import Callable, Iterator, Mapping
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from types import MappingProxyType
|
||||
from typing import Final
|
||||
|
||||
import pytest
|
||||
from _pytest.mark import ParameterSet
|
||||
from e2e_config import provider_edge_base, unique_marker
|
||||
from e2e_http import BinaryStream, StreamingResponse, require_successful_call, unwrap
|
||||
from endpoint_matrix import (
|
||||
AuthMode,
|
||||
LlmRoute,
|
||||
MatrixEndpoint,
|
||||
Provider,
|
||||
Streaming,
|
||||
credential_values,
|
||||
deployment_params,
|
||||
selected_auth_modes,
|
||||
selected_providers,
|
||||
)
|
||||
from endpoints_client import (
|
||||
CompletionsResult,
|
||||
EmbeddingsResult,
|
||||
EndpointsClient,
|
||||
ImagesResult,
|
||||
MessagesResult,
|
||||
ResponsesOutputTextDeltaEvent,
|
||||
ResponsesResult,
|
||||
ResponsesStreamEventType,
|
||||
)
|
||||
from lifecycle import ResourceManager
|
||||
from models import AnthropicMessagesBody, ChatBody, ChatMessage, CredentialCreateBody, ModelInfoBody, ModelNewBody
|
||||
from pydantic import BaseModel
|
||||
|
||||
pytestmark = [pytest.mark.e2e]
|
||||
|
||||
THIS_MODULE: Final = Path(__file__).resolve()
|
||||
MAX_TOKENS: Final = 256
|
||||
WEATHER_WAV: Final = Path(__file__).resolve().parent / "realtime" / "fixtures" / "weather_question_24k.wav"
|
||||
EDIT_PNG: Final = base64.b64decode(
|
||||
"iVBORw0KGgoAAAANSUhEUgAAAEAAAABACAIAAAAlC+aJAAAAS0lEQVR42u3PMQ0AAAwDoPo3"
|
||||
"3UrYvQQckD4XAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEB"
|
||||
"AYHLAMpT0sIcNbcEAAAAAElFTkSuQmCC"
|
||||
)
|
||||
|
||||
|
||||
def _greeting_prompt() -> str:
|
||||
return f"Reply with one short friendly sentence. Request {unique_marker()}"
|
||||
|
||||
|
||||
def _counting_prompt() -> str:
|
||||
return f"Count from 1 to 10, one number per line. Request {unique_marker()}"
|
||||
|
||||
|
||||
class _ChatDelta(BaseModel):
|
||||
content: str | None = None
|
||||
|
||||
|
||||
class _ChatChunkChoice(BaseModel):
|
||||
delta: _ChatDelta = _ChatDelta()
|
||||
|
||||
|
||||
class _ChatChunk(BaseModel):
|
||||
choices: list[_ChatChunkChoice] = []
|
||||
|
||||
|
||||
class _MessagesDelta(BaseModel):
|
||||
text: str = ""
|
||||
|
||||
|
||||
class _MessagesEvent(BaseModel):
|
||||
type: str
|
||||
delta: _MessagesDelta | None = None
|
||||
|
||||
|
||||
def _assert_stream_established(result: StreamingResponse) -> None:
|
||||
assert result.ok and result.is_streaming, f"stream was not established: {result}"
|
||||
assert result.stream_error is None, f"stream carried an error event: {result.stream_error}"
|
||||
assert len(result.stream_events) > 1, f"stream delivered a single event: {result.stream_events}"
|
||||
|
||||
|
||||
def _chat(client: EndpointsClient, key: str, model: str) -> None:
|
||||
body: Final = ChatBody(
|
||||
model=model, messages=[ChatMessage(role="user", content=_greeting_prompt())], max_tokens=MAX_TOKENS
|
||||
)
|
||||
response: Final = unwrap(client.proxy.chat(key, body))
|
||||
assert response.choices, f"/chat/completions returned no choices: {response}"
|
||||
message: Final = response.choices[0].message
|
||||
assert message is not None and (message.content or "").strip(), (
|
||||
f"/chat/completions returned no assistant text: {response}"
|
||||
)
|
||||
|
||||
|
||||
def _chat_stream(client: EndpointsClient, key: str, model: str) -> None:
|
||||
body: Final = ChatBody(
|
||||
model=model,
|
||||
messages=[ChatMessage(role="user", content=_counting_prompt())],
|
||||
max_tokens=MAX_TOKENS,
|
||||
stream=True,
|
||||
)
|
||||
result: Final = client.proxy.chat_stream(key, body)
|
||||
_assert_stream_established(result)
|
||||
text: Final = "".join(
|
||||
choice.delta.content or ""
|
||||
for event in result.stream_events
|
||||
for choice in _ChatChunk.model_validate_json(event).choices
|
||||
)
|
||||
assert text.strip(), f"/chat/completions stream carried no content deltas: {result.stream_events[:3]}"
|
||||
|
||||
|
||||
def _completions(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = client.text_completions(key, model, _greeting_prompt(), max_tokens=MAX_TOKENS)
|
||||
require_successful_call(result)
|
||||
parsed: Final = CompletionsResult.model_validate_json(result.body)
|
||||
assert parsed.choices and (parsed.choices[0].text or "").strip(), (
|
||||
f"/v1/completions returned no completion text: {result.body[:300]}"
|
||||
)
|
||||
|
||||
|
||||
def _messages(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = client.messages(key, model, _greeting_prompt(), max_tokens=MAX_TOKENS)
|
||||
require_successful_call(result)
|
||||
parsed: Final = MessagesResult.model_validate_json(result.body)
|
||||
assert parsed.role == "assistant", f"/v1/messages did not answer as the assistant: {result.body[:300]}"
|
||||
assert parsed.text.strip(), f"/v1/messages returned no text block: {result.body[:300]}"
|
||||
|
||||
|
||||
def _messages_stream(client: EndpointsClient, key: str, model: str) -> None:
|
||||
body: Final = AnthropicMessagesBody(
|
||||
model=model,
|
||||
max_tokens=MAX_TOKENS,
|
||||
stream=True,
|
||||
messages=[ChatMessage(role="user", content=_counting_prompt())],
|
||||
)
|
||||
result: Final = client.proxy.messages_stream(key, body)
|
||||
_assert_stream_established(result)
|
||||
events: Final = tuple(_MessagesEvent.model_validate_json(event) for event in result.stream_events)
|
||||
types: Final = tuple(event.type for event in events)
|
||||
assert types[0] == "message_start", f"/v1/messages stream did not open with message_start: {types[:3]}"
|
||||
assert types[-1] == "message_stop", f"/v1/messages stream did not close with message_stop: {types[-3:]}"
|
||||
text: Final = "".join(
|
||||
event.delta.text for event in events if event.type == "content_block_delta" and event.delta is not None
|
||||
)
|
||||
assert text.strip(), f"/v1/messages stream carried no text deltas: {types}"
|
||||
|
||||
|
||||
def _responses(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = client.responses(key, model, _greeting_prompt())
|
||||
require_successful_call(result)
|
||||
parsed: Final = ResponsesResult.model_validate_json(result.body)
|
||||
assert parsed.status == "completed", f"/v1/responses did not complete: {result.body[:300]}"
|
||||
assert parsed.text.strip(), f"/v1/responses returned no output text: {result.body[:300]}"
|
||||
|
||||
|
||||
def _responses_stream(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = client.responses(key, model, _counting_prompt(), stream=True)
|
||||
_assert_stream_established(result)
|
||||
types: Final = tuple(ResponsesStreamEventType.model_validate_json(event).type for event in result.stream_events)
|
||||
assert types[-1] == "response.completed", f"/v1/responses stream did not end with response.completed: {types[-3:]}"
|
||||
text: Final = "".join(
|
||||
ResponsesOutputTextDeltaEvent.model_validate_json(event).delta
|
||||
for event, event_type in zip(result.stream_events, types, strict=True)
|
||||
if event_type == "response.output_text.delta"
|
||||
)
|
||||
assert text.strip(), f"/v1/responses stream carried no output_text deltas: {types}"
|
||||
|
||||
|
||||
def _embeddings(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = client.embeddings(key, model, f"the quick brown fox {unique_marker()}")
|
||||
require_successful_call(result)
|
||||
vector: Final = EmbeddingsResult.model_validate_json(result.body).first_vector
|
||||
assert len(vector) > 1, f"/embeddings returned no vector: {result.body[:300]}"
|
||||
assert any(component != 0 for component in vector), "/embeddings returned an all-zero vector"
|
||||
|
||||
|
||||
def _audio_speech(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = client.audio_speech(key, model, "The matrix speaks.")
|
||||
require_successful_call(result)
|
||||
assert "audio" in (result.content_type or ""), (
|
||||
f"/v1/audio/speech content-type is not audio: {result.content_type!r}"
|
||||
)
|
||||
assert result.body, "/v1/audio/speech returned an empty body"
|
||||
|
||||
|
||||
def _audio_speech_stream(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final[BinaryStream] = client.audio_speech_stream(key, model, "The matrix speaks, streamed.")
|
||||
assert result.status_code == 200, f"/v1/audio/speech stream failed: {result.status_code} {result.error_body}"
|
||||
assert "audio" in (result.content_type or ""), f"streamed speech content-type is not audio: {result.content_type!r}"
|
||||
assert result.chunk_count > 0 and result.total_bytes > 0, f"streamed speech delivered no audio bytes: {result}"
|
||||
|
||||
|
||||
def _audio_transcriptions(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = unwrap(client.transcribe(key, model, filename=WEATHER_WAV.name, content=WEATHER_WAV.read_bytes()))
|
||||
text: Final = result.text.strip()
|
||||
assert "weather" in text.lower(), f"transcript of a spoken weather question does not mention weather: {text!r}"
|
||||
|
||||
|
||||
def _images_generations(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = client.images(key, model, "A single red circle on a white background")
|
||||
require_successful_call(result)
|
||||
images: Final = ImagesResult.model_validate_json(result.body)
|
||||
assert images.data, f"/v1/images/generations returned no images: {result.body[:300]}"
|
||||
assert images.data[0].b64_json or images.data[0].url, f"generated image has no payload: {images.data[0]}"
|
||||
|
||||
|
||||
def _images_edits(client: EndpointsClient, key: str, model: str) -> None:
|
||||
edited: Final = unwrap(client.image_edit(key, model, "Add a small red circle in the center", EDIT_PNG))
|
||||
assert edited.data, f"/v1/images/edits returned no images: {edited}"
|
||||
assert edited.data[0].b64_json or edited.data[0].url, f"edited image has no payload: {edited.data[0]}"
|
||||
|
||||
|
||||
def _moderations(client: EndpointsClient, key: str, model: str) -> None:
|
||||
result: Final = unwrap(client.moderations(key, model, f"I enjoy long walks on sunny days. {unique_marker()}"))
|
||||
assert len(result.results) == 1, f"/v1/moderations did not return one verdict for one input: {result}"
|
||||
assert result.results[0].categories, f"/v1/moderations verdict carries no categories: {result.results[0]}"
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class EndpointCase:
|
||||
endpoint: MatrixEndpoint
|
||||
streaming: Streaming
|
||||
run: Callable[[EndpointsClient, str, str], None]
|
||||
|
||||
|
||||
ENDPOINT_CASES: Final[tuple[EndpointCase, ...]] = (
|
||||
EndpointCase("chat_completions", "nonstream", _chat),
|
||||
EndpointCase("chat_completions", "stream", _chat_stream),
|
||||
EndpointCase("completions", "nonstream", _completions),
|
||||
EndpointCase("messages", "nonstream", _messages),
|
||||
EndpointCase("messages", "stream", _messages_stream),
|
||||
EndpointCase("responses", "nonstream", _responses),
|
||||
EndpointCase("responses", "stream", _responses_stream),
|
||||
EndpointCase("embeddings", "nonstream", _embeddings),
|
||||
EndpointCase("audio_speech", "nonstream", _audio_speech),
|
||||
EndpointCase("audio_speech", "stream", _audio_speech_stream),
|
||||
EndpointCase("audio_transcriptions", "nonstream", _audio_transcriptions),
|
||||
EndpointCase("images_generations", "nonstream", _images_generations),
|
||||
EndpointCase("images_edits", "nonstream", _images_edits),
|
||||
EndpointCase("moderations", "nonstream", _moderations),
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True, slots=True)
|
||||
class MatrixCell:
|
||||
case: EndpointCase
|
||||
provider: Provider
|
||||
auth_mode: AuthMode
|
||||
|
||||
@property
|
||||
def registry_id(self) -> str:
|
||||
route: Final = self.provider.id_route(self.case.endpoint)
|
||||
return f"llm.{self.case.endpoint}.{route}.basic.{self.case.streaming}.works"
|
||||
|
||||
@property
|
||||
def test_id(self) -> str:
|
||||
return f"{self.case.endpoint}-{self.case.streaming}-{self.provider.route}-{self.auth_mode}"
|
||||
|
||||
@property
|
||||
def deployment(self) -> DeploymentKey:
|
||||
return (self.provider.route, self.case.endpoint, self.auth_mode)
|
||||
|
||||
|
||||
type DeploymentKey = tuple[LlmRoute, MatrixEndpoint, AuthMode]
|
||||
|
||||
|
||||
def _cells() -> tuple[MatrixCell, ...]:
|
||||
return tuple(
|
||||
MatrixCell(case, provider, auth_mode)
|
||||
for provider in selected_providers()
|
||||
for case in ENDPOINT_CASES
|
||||
if case.endpoint in provider.backends
|
||||
for auth_mode in selected_auth_modes()
|
||||
)
|
||||
|
||||
|
||||
def _param(cell: MatrixCell) -> ParameterSet:
|
||||
replay: Final = (pytest.mark.replayable,) if cell.provider.replayable() else ()
|
||||
return pytest.param(cell, id=cell.test_id, marks=(pytest.mark.covers(cell.registry_id), *replay))
|
||||
|
||||
|
||||
def _selected_cells(session: pytest.Session) -> tuple[MatrixCell, ...]:
|
||||
"""The cells pytest will actually run from this module, after -m / -k deselection,
|
||||
so replay registers no deployment for a provider it holds no credential for."""
|
||||
return tuple(
|
||||
cell
|
||||
for item in session.items
|
||||
if isinstance(item, pytest.Function) and item.path == THIS_MODULE
|
||||
for cell in (item.callspec.params.get("cell"),)
|
||||
if isinstance(cell, MatrixCell)
|
||||
)
|
||||
|
||||
|
||||
def _deployment_body(key: DeploymentKey, provider: Provider, credential_name: str | None) -> ModelNewBody:
|
||||
_, endpoint, auth_mode = key
|
||||
edge_base: Final = None if provider.edge_mount is None else provider_edge_base(provider.edge_mount)
|
||||
return ModelNewBody(
|
||||
model_name=f"e2e-matrix-{endpoint}-{unique_marker()}",
|
||||
litellm_params=deployment_params(
|
||||
provider, endpoint, auth_mode, edge_base=edge_base, credential_name=credential_name
|
||||
),
|
||||
model_info=ModelInfoBody(),
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def deployments(
|
||||
request: pytest.FixtureRequest, endpoints_client: EndpointsClient
|
||||
) -> Iterator[Mapping[DeploymentKey, str]]:
|
||||
"""One deployment per (provider, endpoint, auth mode) among the selected cells, written
|
||||
as a single batch, plus one stored credential per provider that any selected cell
|
||||
references by name. Yields deployment key -> model alias; tears everything down."""
|
||||
cells: Final = _selected_cells(request.session)
|
||||
providers: Final[Mapping[LlmRoute, Provider]] = MappingProxyType(
|
||||
{cell.provider.route: cell.provider for cell in cells}
|
||||
)
|
||||
stored: Final[frozenset[LlmRoute]] = frozenset(
|
||||
cell.provider.route for cell in cells if cell.auth_mode == "stored_credential"
|
||||
)
|
||||
credentials: Final[Mapping[LlmRoute, str]] = MappingProxyType(
|
||||
{route: f"e2e-matrix-cred-{unique_marker()}" for route in stored}
|
||||
)
|
||||
for route, name in credentials.items():
|
||||
endpoints_client.proxy.create_credential(
|
||||
CredentialCreateBody(credential_name=name, credential_values=dict(credential_values(providers[route])))
|
||||
)
|
||||
try:
|
||||
keys: Final = tuple(dict.fromkeys(cell.deployment for cell in cells))
|
||||
bodies: Final = tuple(
|
||||
_deployment_body(key, providers[key[0]], credentials[key[0]] if key[2] == "stored_credential" else None)
|
||||
for key in keys
|
||||
)
|
||||
model_ids: Final = endpoints_client.proxy.register_models(bodies)
|
||||
try:
|
||||
yield MappingProxyType(dict(zip(keys, (body.model_name for body in bodies), strict=True)))
|
||||
finally:
|
||||
for model_id in model_ids:
|
||||
endpoints_client.delete_model(model_id)
|
||||
finally:
|
||||
for name in credentials.values():
|
||||
endpoints_client.proxy.delete_credential(name)
|
||||
|
||||
|
||||
class TestEndpointMatrix:
|
||||
@pytest.mark.parametrize("cell", tuple(_param(cell) for cell in _cells()))
|
||||
def test_endpoint_answers_through_every_provider_and_auth_mode(
|
||||
self,
|
||||
cell: MatrixCell,
|
||||
endpoints_client: EndpointsClient,
|
||||
resources: ResourceManager,
|
||||
deployments: Mapping[DeploymentKey, str],
|
||||
) -> None:
|
||||
cell.case.run(endpoints_client, resources.key(), deployments[cell.deployment])
|
||||
|
|
@ -140,6 +140,15 @@ def resolve_mount(path: str, mounts: Mapping[str, str]) -> ResolvedMount | None:
|
|||
|
||||
REPLAY_MISS_STATUS: Final = 599
|
||||
|
||||
# The proxy's `refresh_model_info` job lists `{api_base}/v1/models` on its own wall clock and
|
||||
# swallows failures, so whether it lands inside a test is chance: never record it, never owe it.
|
||||
_SCHEDULED_DISCOVERY: Final = ("get", "v1/models", "")
|
||||
|
||||
|
||||
def is_scheduled_discovery(method: str, upstream_path: str, query: str) -> bool:
|
||||
return (method.lower(), upstream_path.strip("/"), query) == _SCHEDULED_DISCOVERY
|
||||
|
||||
|
||||
_HOP_BY_HOP_HEADERS: Final[frozenset[str]] = frozenset(
|
||||
{
|
||||
"connection",
|
||||
|
|
@ -842,6 +851,12 @@ def handle_edge_request(
|
|||
mount: Final = resolved.mount
|
||||
upstream_base: Final = resolved.upstream_base
|
||||
test_key, upstream_path = split_test_segment(resolved.upstream_path)
|
||||
if isinstance(backend, RecordEdge | ReplayEdge) and is_scheduled_discovery(method, upstream_path, split.query):
|
||||
return _text_reply(
|
||||
REPLAY_MISS_STATUS,
|
||||
f"{method.upper()} /{upstream_path} is the proxy's scheduled model discovery, not test traffic; "
|
||||
"the edge neither records nor replays it",
|
||||
)
|
||||
profile: Final = (
|
||||
backend.recorder.profile
|
||||
if isinstance(backend, RecordEdge)
|
||||
|
|
|
|||
|
|
@ -11,7 +11,7 @@ from __future__ import annotations
|
|||
import os
|
||||
import time
|
||||
import warnings
|
||||
from collections.abc import Callable, Mapping
|
||||
from collections.abc import Callable, Mapping, Sequence
|
||||
from dataclasses import dataclass, field, replace
|
||||
from datetime import datetime
|
||||
from functools import reduce
|
||||
|
|
@ -654,18 +654,7 @@ class ProxyClient:
|
|||
balancer address is configured (every request opens a fresh connection, so
|
||||
the caller's next request re-rolls), so waiting out PROPAGATION_TIMEOUT is
|
||||
what makes the model safe to use anywhere."""
|
||||
model_id = unwrap(
|
||||
self.transport.post(
|
||||
"/model/new",
|
||||
headers=self.management_headers(),
|
||||
json=body.model_copy(update={"litellm_params": route_cache_model(
|
||||
body.litellm_params, provider_edge_base,
|
||||
enabled=os.environ.get("E2E_PROVIDER_CACHE", "0") == "1" and not provider_live,
|
||||
mode=body.model_info.mode,
|
||||
)}),
|
||||
response_type=ModelNewResponse,
|
||||
)
|
||||
).model_id
|
||||
model_id = unwrap(self._write_model(body, provider_live=provider_live)).model_id
|
||||
written_at = time.monotonic()
|
||||
try:
|
||||
self._await_model_servable(body.model_name, listed_for)
|
||||
|
|
@ -675,6 +664,37 @@ class ProxyClient:
|
|||
settle_propagation(written_at)
|
||||
return model_id
|
||||
|
||||
def register_models(self, bodies: Sequence[ModelNewBody]) -> tuple[str, ...]:
|
||||
"""`register_model` for a batch of proxy-wide deployments: every row is written
|
||||
first, then each is awaited on the data plane, and one propagation wait covers
|
||||
them all, so a suite registering many deployments pays the reload budget once.
|
||||
A failure anywhere deletes every deployment the batch already created."""
|
||||
results: Final = tuple(self._write_model(body) for body in bodies)
|
||||
written_at = time.monotonic()
|
||||
model_ids: Final = tuple(result.data.model_id for result in results if isinstance(result, Success))
|
||||
try:
|
||||
for body, result in zip(bodies, results, strict=True):
|
||||
_ = unwrap(result)
|
||||
self._await_model_servable(body.model_name)
|
||||
except BaseException:
|
||||
for model_id in model_ids:
|
||||
self.delete_model(model_id)
|
||||
raise
|
||||
settle_propagation(written_at)
|
||||
return model_ids
|
||||
|
||||
def _write_model(self, body: ModelNewBody, *, provider_live: bool = False) -> Result[ModelNewResponse]:
|
||||
return self.transport.post(
|
||||
"/model/new",
|
||||
headers=self.management_headers(),
|
||||
json=body.model_copy(update={"litellm_params": route_cache_model(
|
||||
body.litellm_params, provider_edge_base,
|
||||
enabled=os.environ.get("E2E_PROVIDER_CACHE", "0") == "1" and not provider_live,
|
||||
mode=body.model_info.mode,
|
||||
)}),
|
||||
response_type=ModelNewResponse,
|
||||
)
|
||||
|
||||
def _await_model_servable(self, model_name: str, listed_for: str | None = None) -> None:
|
||||
"""Block until every replica lists `model_name`, or fail at model_servable_timeout."""
|
||||
headers: Final = self.management_headers(listed_for)
|
||||
|
|
|
|||
|
|
@ -906,14 +906,14 @@ class TestReplayLeftover:
|
|||
with fake_provider() as provider:
|
||||
with running_edge(record_backend(root), {"openai": provider_url(provider)}) as edge:
|
||||
call_edge(edge, "POST", CHAT_PATH, body=chat_body("hi"))
|
||||
call_edge(edge, "GET", "/openai/v1/models")
|
||||
call_edge(edge, "GET", "/openai/v1/files/file-1")
|
||||
source = replay_source(root)
|
||||
with running_edge(ReplayEdge(source=source), REPLAY_MOUNTS) as edge:
|
||||
call_edge(edge, "POST", CHAT_PATH, body=chat_body("hi"))
|
||||
error = source.leftover_error(current_test_key())
|
||||
assert error is not None
|
||||
assert "1 of 2 recorded interactions never consumed" in error
|
||||
assert "e.g. get /openai/v1/models #" in error
|
||||
assert "e.g. get /openai/v1/files/file-1 #" in error
|
||||
assert "re-record with E2E_FIXTURE_MODE=record" in error
|
||||
|
||||
def test_fully_consumed_recording_leaves_nothing(self, tmp_path: Path) -> None:
|
||||
|
|
@ -936,6 +936,27 @@ class TestReplayLeftover:
|
|||
assert replay_leftover_error(mode_raw="", bundle_dir=missing, test_key="k") is None
|
||||
assert replay_leftover_error(mode_raw="record", bundle_dir=missing, test_key="k") is None
|
||||
|
||||
def test_the_proxys_scheduled_model_discovery_is_neither_recorded_nor_owed(self, tmp_path: Path) -> None:
|
||||
"""The proxy lists `{api_base}/v1/models` on a wall-clock schedule, so whether
|
||||
that call lands inside a test's window is chance: recording it would make
|
||||
replay owe an interaction the schedule may never make, and a miss on it
|
||||
would fail a replay run over housekeeping the proxy itself ignores."""
|
||||
root = tmp_path / "bundle"
|
||||
with fake_provider() as provider:
|
||||
with running_edge(record_backend(root), {"openai": provider_url(provider)}) as edge:
|
||||
call_edge(edge, "POST", CHAT_PATH, body=chat_body("hi"))
|
||||
recorded = call_edge(edge, "GET", "/openai/v1/models")
|
||||
assert provider.hits == ["POST /v1/chat/completions"]
|
||||
assert recorded.status_code == REPLAY_MISS_STATUS
|
||||
assert [file.name for file in this_tests_files(root)] == ["0000-post-openai-v1-chat-completions.json"]
|
||||
source = replay_source(root)
|
||||
with running_edge(ReplayEdge(source=source), REPLAY_MOUNTS) as edge:
|
||||
call_edge(edge, "POST", CHAT_PATH, body=chat_body("hi"))
|
||||
replayed = call_edge(edge, "GET", "/openai/v1/models")
|
||||
assert replayed.status_code == REPLAY_MISS_STATUS
|
||||
assert b"scheduled model discovery" in replayed.body
|
||||
assert source.leftover_error(current_test_key()) is None
|
||||
|
||||
|
||||
class TestConcurrentReplay:
|
||||
def test_parallel_identical_calls_serve_each_recording_exactly_once(self, tmp_path: Path) -> None:
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue