feat(e2e): read management routes back from the control plane replicas (#43373)

* feat(e2e): read management routes back from the control plane replicas

The suite's management read-backs (/key/info, /team/info and friends)
polled the same replica list as the data plane. On a componentized stack
whose LITELLM_PROXY_REPLICA_URLS names the gateway pods directly, that
list answers those routes 404, since a gateway pod trims the management
routes at startup. A new LITELLM_CONTROL_PLANE_REPLICA_URLS names the
addresses a management read-back polls instead: an exported list wins,
and when it is unset the old rule stands, the data-plane replicas while
the control plane shares the suite's base URL and the control-plane base
alone once it is split.

build_proxy_client takes the list as control_replica_urls and
read_back_everywhere picks its replicas per path, the way the rest of the
client already does.

* fix(e2e): derive the control replicas of a client built for another proxy

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
This commit is contained in:
devin-ai-integration[bot] 2026-09-26 18:05:55 -07:00 • committed by GitHub
parent eea1d0f269
commit b396b0b724
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
7 changed files with 215 additions and 30 deletions

View file

@ -105,7 +105,7 @@ A couple of logging destinations are configured on the proxy rather than by the
### The pull request check
Every same-repository PR that adds, modifies, or renames a `tests/e2e/**/test_*.py` file runs those changed files three times. A change to the harness itself, meaning a root-level `tests/e2e/*.py` file or `pytest.ini`, `tests/e2e/gateway/`, `.github/e2e-stack/`, or the workflow, also runs the `access_control` suite and both JWT suites as canaries, because those files have no test of their own that exercises the stack. `.github/e2e-stack/select_tests.py` applies both rules. The stack config at `tests/e2e/gateway/stage_mirror_ci_config.yml` must declare every model the selected suites use; a missing one shows up as a failed test id in the public log. The suite's own single rerun for network errors and 5xx responses (see `pytest.ini`) applies on every pass, so a transport blip does not fail the check while a race inside a test still does. The stage-mirror stack has a control-plane backend, two gateways behind nginx, Postgres, Keycloak, Jaeger, and TLS cluster-mode Valkey. Realm-only edits also trigger these canaries. The stack exports every gateway address in `LITELLM_PROXY_REPLICA_URLS`, so model registration waits until each gateway lists the new model rather than whichever one the load balancer answered from. Documentation, deleted-file, and application-only changes do not start the stack or request environment approval. The `ui/`, `claude_code/`, `load/`, and `secret_manager/` directories, `batches/test_managed_files_enforcement_e2e.py`, `llm_translation/realtime/test_realtime_pipecat_audio_e2e.py`, and `guardrails/test_presidio_masking_e2e.py` remain outside this check because they use separate tooling or need a differently configured stack: the pipecat audio suite skips itself at import time unless the NLTK `punkt_tab` data is installed, and the presidio suite fails without the analyzer and anonymizer services this stack does not start. `logging/test_otel_v2_langfuse_generation_output_e2e.py` is marked `otel_v2` and deselects itself unless `E2E_OTEL_V2` is set, because it needs a gateway booted with `LITELLM_OTEL_V2=true` and Langfuse credentials, neither of which this stack provides, so run it with `E2E_OTEL_V2=1` against a local OTel v2 proxy. The Redis chaos test under `load/` needs a proxy it can pause the Redis of on the same host (`gateway/redis_chaos_ci_config.yml`), which `.github/workflows/test-e2e-redis-chaos.yml` boots, and which the Buildkite `e2e-redis-chaos` step in project-releaser runs co-located with Postgres and Valkey in one pod; it is deselected unless `E2E_REDIS_CHAOS` is set. The `secret_manager/` lanes each need a proxy configured against their own secret manager (see Secret manager lanes below)
Every same-repository PR that adds, modifies, or renames a `tests/e2e/**/test_*.py` file runs those changed files three times. A change to the harness itself, meaning a root-level `tests/e2e/*.py` file or `pytest.ini`, `tests/e2e/gateway/`, `.github/e2e-stack/`, or the workflow, also runs the `access_control` suite and both JWT suites as canaries, because those files have no test of their own that exercises the stack. `.github/e2e-stack/select_tests.py` applies both rules. The stack config at `tests/e2e/gateway/stage_mirror_ci_config.yml` must declare every model the selected suites use; a missing one shows up as a failed test id in the public log. The suite's own single rerun for network errors and 5xx responses (see `pytest.ini`) applies on every pass, so a transport blip does not fail the check while a race inside a test still does. The stage-mirror stack has a control-plane backend, two gateways behind nginx, Postgres, Keycloak, Jaeger, and TLS cluster-mode Valkey. Realm-only edits also trigger these canaries. The stack exports every gateway address in `LITELLM_PROXY_REPLICA_URLS`, so model registration waits until each gateway lists the new model rather than whichever one the load balancer answered from. The Buildkite PR stack exports its two gateway pods the same way and, because those pods sit behind one router base that also fronts the backend, names that base in `LITELLM_CONTROL_PLANE_REPLICA_URLS` so management read-backs poll the plane that serves them instead of the gateway pods, which trim management routes at startup and answer them 404. Documentation, deleted-file, and application-only changes do not start the stack or request environment approval. The `ui/`, `claude_code/`, `load/`, and `secret_manager/` directories, `batches/test_managed_files_enforcement_e2e.py`, `llm_translation/realtime/test_realtime_pipecat_audio_e2e.py`, and `guardrails/test_presidio_masking_e2e.py` remain outside this check because they use separate tooling or need a differently configured stack: the pipecat audio suite skips itself at import time unless the NLTK `punkt_tab` data is installed, and the presidio suite fails without the analyzer and anonymizer services this stack does not start. `logging/test_otel_v2_langfuse_generation_output_e2e.py` is marked `otel_v2` and deselects itself unless `E2E_OTEL_V2` is set, because it needs a gateway booted with `LITELLM_OTEL_V2=true` and Langfuse credentials, neither of which this stack provides, so run it with `E2E_OTEL_V2=1` against a local OTel v2 proxy. The Redis chaos test under `load/` needs a proxy it can pause the Redis of on the same host (`gateway/redis_chaos_ci_config.yml`), which `.github/workflows/test-e2e-redis-chaos.yml` boots, and which the Buildkite `e2e-redis-chaos` step in project-releaser runs co-located with Postgres and Valkey in one pod; it is deselected unless `E2E_REDIS_CHAOS` is set. The `secret_manager/` lanes each need a proxy configured against their own secret manager (see Secret manager lanes below)
Every selected file must execute at least one passing test in each pass, and any test failure, collection error, or entirely skipped or deselected file fails the check. A file whose tests are all marked skip therefore cannot pass this check, so unskip at least one of them, or add the file to `UNSUPPORTED` in `select_tests.py` with the reason, before changing one. A failed pass stops the run. The public log prints pytest's one-line summary for each pass, including the rerun count, and names each failed or errored test as `classname::name`, so a retried network error or a failing test is visible without the raw output. The final `e2e-changed-tests` job succeeds only when no supported test files changed or the approved run completed all three passes. Fork PRs with selected tests fail this gate until a maintainer brings the reviewed change onto a same-repository branch

View file

@ -100,5 +100,6 @@ def require_proxy_client(
master_key=cfg.api_key,
control_plane_base_url=cfg.base_url,
replica_urls=(cfg.base_url,),
control_replica_urls=(cfg.base_url,),
)
return ProxyClientConfig(client=client, api_key=cfg.api_key)

View file

@ -595,6 +595,7 @@ def _build_control_plane_client(proxy_config: ProxyConfig):
master_key=proxy_config.api_key,
control_plane_base_url=proxy_config.base_url,
replica_urls=(proxy_config.base_url,),
control_replica_urls=(proxy_config.base_url,),
)

View file

@ -7,6 +7,7 @@ environment so the same tests run against localhost or a deployed proxy.
from __future__ import annotations
import os
from dataclasses import dataclass
import time
import uuid
from pathlib import Path
@ -33,12 +34,68 @@ CONTROL_PLANE_BASE_URL = os.environ.get(
).rstrip("/")
def split_replica_urls(raw: str) -> tuple[str, ...]:
return tuple(dict.fromkeys(url.strip().rstrip("/") for url in raw.split(",") if url.strip()))
def parse_replica_urls(raw: str, fallback: str) -> tuple[str, ...]:
urls: Final = tuple(dict.fromkeys(url.strip().rstrip("/") for url in raw.split(",") if url.strip()))
return urls or (fallback,)
return split_replica_urls(raw) or (fallback,)
def parse_control_plane_replica_urls(
raw: str, *, control_plane_base_url: str, base_url: str, replica_urls: tuple[str, ...]
) -> tuple[str, ...]:
"""The replicas a management read-back polls. LITELLM_CONTROL_PLANE_REPLICA_URLS
names them outright; unset, they follow the two base URLs: every data-plane
replica when the planes share a base (a monolith serves every route from every
replica) and the control-plane base alone when they differ. A stack sets it when
LITELLM_PROXY_REPLICA_URLS names gateway pods behind a shared router base, since
a gateway trims the management routes at startup and answers them 404."""
explicit: Final = split_replica_urls(raw)
if explicit:
return explicit
return replica_urls if control_plane_base_url == base_url else (control_plane_base_url,)
PROXY_REPLICA_URLS: Final = parse_replica_urls(os.environ.get("LITELLM_PROXY_REPLICA_URLS", ""), PROXY_BASE_URL)
CONTROL_PLANE_REPLICA_URLS: Final = parse_control_plane_replica_urls(
os.environ.get("LITELLM_CONTROL_PLANE_REPLICA_URLS", ""),
control_plane_base_url=CONTROL_PLANE_BASE_URL,
base_url=PROXY_BASE_URL,
replica_urls=PROXY_REPLICA_URLS,
)
@dataclass(frozen=True, slots=True)
class StackEndpoints:
base_url: str
control_plane_base_url: str
replica_urls: tuple[str, ...]
control_replica_urls: tuple[str, ...]
def control_replica_urls_for(
self, *, base_url: str, control_plane_base_url: str, replica_urls: tuple[str, ...]
) -> tuple[str, ...]:
"""The control replicas a client built for these endpoints polls when its caller names none:
this stack's own list for this stack's endpoints, since an exported list describes one stack only,
and the base-URL rule for any other proxy."""
if (base_url, control_plane_base_url, replica_urls) == (
self.base_url,
self.control_plane_base_url,
self.replica_urls,
):
return self.control_replica_urls
return parse_control_plane_replica_urls(
"", control_plane_base_url=control_plane_base_url, base_url=base_url, replica_urls=replica_urls
)
ENV_STACK: Final = StackEndpoints(
base_url=PROXY_BASE_URL,
control_plane_base_url=CONTROL_PLANE_BASE_URL,
replica_urls=PROXY_REPLICA_URLS,
control_replica_urls=CONTROL_PLANE_REPLICA_URLS,
)
UI_USERNAME = os.environ.get("E2E_UI_USERNAME", "admin")
UI_PASSWORD = os.environ.get("E2E_UI_PASSWORD", MASTER_KEY)

View file

@ -187,6 +187,7 @@ def owned_gateway(idp: Keycloak, directory: Path, cleanup: ExitStack) -> OAuthGa
base_url=base_url,
control_plane_base_url=base_url,
replica_urls=(base_url,),
control_replica_urls=(base_url,),
master_key=os.environ["LITELLM_MASTER_KEY"],
),
_environment=environment,

View file

@ -20,6 +20,7 @@ from typing import Final, Literal
from e2e_config import (
CONTROL_PLANE_BASE_URL,
ENV_STACK,
MASTER_KEY,
POLL_INTERVAL,
POLL_TIMEOUT,
@ -530,16 +531,16 @@ class ProxyClient:
response_type: type[R],
converged: Callable[[Result[R]], bool],
) -> Mapping[str, Result[R]]:
"""GET `path` under the master key on every replica in PROXY_REPLICA_URLS (the
data-plane URL alone when the stack exports no per-gateway addresses), polling
each to poll_timeout until its read satisfies `converged`. Returns that read per
replica, or fails naming the first replica that never converged and its last
read. Behind a load balancer the single address proves one replica converged,
not all of them; only per-gateway addresses make this a fleet-wide proof."""
"""GET `path` under the master key on every replica that serves it (see
replicas_for), polling each to poll_timeout until its read satisfies
`converged`. Returns that read per replica, or fails naming the first replica
that never converged and its last read. Behind a load balancer the single
address proves one replica converged, not all of them; only per-replica
addresses make this a fleet-wide proof."""
outcomes: Final = await_converged_everywhere(
{
url: self._body_poller(transport, path, params, response_type)
for url, transport in self.replicas.items()
for url, transport in self.replicas_for(path).items()
},
converged=converged,
timeout=self.poll_timeout,
@ -765,13 +766,11 @@ class ProxyClient:
def replicas_for(self, path: str) -> Mapping[str, Transport]:
"""The replicas that serve `path`: every data-plane replica for an LLM route,
and for a management route the control-plane replicas, since the data-plane
replicas trim management routes and answer them 404. A monolith serves both
from every replica, so a management read-back polls all of them; a split
deployment exposes one control-plane address (there is one backend process
behind it on the stack these suites run against), so it polls that. A
control plane fronting several backends would need its own replica list to
prove each one converged, the way PROXY_REPLICA_URLS does for the gateways.
Never empty: a read-back against no replica would assert nothing and pass."""
replicas trim management routes and answer them 404. CONTROL_PLANE_REPLICA_URLS
names those (see e2e_config): every data-plane replica for a monolith, the
control plane's own address for a split deployment, and the stack's own list
when its gateway pods sit behind a shared router base. Never empty: a
read-back against no replica would assert nothing and pass."""
replicas: Final = self.control_replicas if is_control_plane_path(path) else self.replicas
assert replicas, f"no replica is configured to serve {path}, so a read-back there would prove nothing"
return replicas
@ -1132,6 +1131,7 @@ def build_proxy_client(
master_key: str = MASTER_KEY,
control_plane_base_url: str = CONTROL_PLANE_BASE_URL,
replica_urls: tuple[str, ...] = PROXY_REPLICA_URLS,
control_replica_urls: tuple[str, ...] | None = None,
) -> ProxyClient:
"""The ProxyClient every suite's client is built from: a SplitTransport that routes
LLM calls to the data plane (PROXY_BASE_URL) and management/admin calls to the
@ -1139,15 +1139,24 @@ def build_proxy_client(
base URLs are the same for a monolithic proxy, so routing is then a no-op.
``replica_urls`` (PROXY_REPLICA_URLS) names every data-plane replica the model
barrier polls directly; it is the data-plane URL itself unless the stack
exports each gateway's own address. Management read-backs poll those same
replicas when the two planes share a base URL (a monolith, where every replica
serves every route) and the control plane alone when they differ (a split
deployment, where the data-plane replicas do not serve management routes).
exports each gateway's own address. ``control_replica_urls``
(CONTROL_PLANE_REPLICA_URLS) names the replicas a management read-back polls:
those same replicas when the two planes share a base URL (a monolith, where
every replica serves every route), the control plane alone when they differ (a
split deployment, where the data-plane replicas do not serve management
routes), or the list the stack exports when its gateway pods sit behind a
shared router base, since a gateway pod trims management routes and its
address cannot stand in for the control plane.
The endpoints are injectable for callers that resolve the proxy some other
way than ``e2e_config``'s env names (see ``claude_code/_env.py``); they must
pass all four together, since a caller that overrides only the data plane
would leave management calls and the replica poll pointed at the env defaults.
way than ``e2e_config``'s env names (see ``claude_code/_env.py``); they pass
the three URL parameters together, since a caller that overrides only the
data plane would leave management calls and the replica polls pointed at the
env defaults. An omitted ``control_replica_urls`` is derived from those three
(``ENV_STACK.control_replica_urls_for``): the env stack's own endpoints take
its exported list, any other proxy follows the base-URL rule above, so a
client built for a local test server never reads management state back from
the env proxy.
Test-to-proxy traffic always goes over the wire, in every E2E_FIXTURE_MODE:
record and replay scope to the proxy's provider-bound calls via the
@ -1170,8 +1179,18 @@ def build_proxy_client(
for url in replica_urls
}
)
control_replicas: Final = (
replicas if control_plane_base_url == base_url else MappingProxyType({control_plane_base_url: split.control})
control_replica_urls_named: Final = (
control_replica_urls
if control_replica_urls is not None
else ENV_STACK.control_replica_urls_for(
base_url=base_url, control_plane_base_url=control_plane_base_url, replica_urls=replica_urls
)
)
control_replicas: Final = MappingProxyType(
{
url: HttpTransport(base_url=url, master_key=master_key, request_timeout=REQUEST_TIMEOUT)
for url in control_replica_urls_named
}
)
return ProxyClient(
transport=split,

View file

@ -24,7 +24,7 @@ from types import MappingProxyType
from typing import Final, cast
import pytest
from e2e_config import parse_replica_urls
from e2e_config import StackEndpoints, parse_control_plane_replica_urls, parse_replica_urls
from e2e_http import NoBody, Result, Success, without_retries
from idp import Keycloak
from lifecycle import ResourceManager
@ -110,7 +110,7 @@ def caller_boundary(
thread.start()
url: Final = f"http://127.0.0.1:{server.server_port}"
proxy: Final = build_proxy_client(
base_url=url, control_plane_base_url=url, replica_urls=(url,), master_key="bootstrap"
base_url=url, control_plane_base_url=url, replica_urls=(url,), control_replica_urls=(url,), master_key="bootstrap"
)
try:
yield ManagementClient(proxy=proxy, master_key="bootstrap"), received
@ -343,6 +343,50 @@ class TestParseReplicaUrls:
assert parse_replica_urls(raw, "http://lb") == ("http://127.0.0.1:4010", "http://127.0.0.1:4011")
class TestParseControlPlaneReplicaUrls:
def test_an_exported_list_wins_over_the_base_url_rule(self) -> None:
assert parse_control_plane_replica_urls(
" http://router/, http://router ",
control_plane_base_url="http://router",
base_url="http://router",
replica_urls=("http://10.0.0.1:4000", "http://10.0.0.2:4000"),
) == ("http://router",)
def test_unset_with_one_shared_base_follows_the_data_plane_replicas(self) -> None:
assert parse_control_plane_replica_urls(
"", control_plane_base_url="http://lb", base_url="http://lb", replica_urls=("http://pod-1", "http://pod-2")
) == ("http://pod-1", "http://pod-2")
def test_unset_with_a_split_control_plane_polls_its_base_alone(self) -> None:
assert parse_control_plane_replica_urls(
"", control_plane_base_url="http://backend", base_url="http://lb", replica_urls=("http://gateway-1",)
) == ("http://backend",)
class TestStackEndpointsControlReplicas:
STACK: Final = StackEndpoints(
base_url="http://router",
control_plane_base_url="http://router",
replica_urls=("http://10.0.0.1:4000", "http://10.0.0.2:4000"),
control_replica_urls=("http://router",),
)
def test_the_stacks_own_endpoints_take_its_exported_control_list(self) -> None:
assert self.STACK.control_replica_urls_for(
base_url="http://router",
control_plane_base_url="http://router",
replica_urls=("http://10.0.0.1:4000", "http://10.0.0.2:4000"),
) == ("http://router",)
def test_any_other_endpoints_follow_the_base_url_rule(self) -> None:
assert self.STACK.control_replica_urls_for(
base_url="http://router", control_plane_base_url="http://router", replica_urls=("http://10.0.0.1:4000",)
) == ("http://10.0.0.1:4000",)
assert self.STACK.control_replica_urls_for(
base_url="http://lb", control_plane_base_url="http://backend", replica_urls=("http://gateway-1",)
) == ("http://backend",)
def _answers(answers: Iterable[str]) -> ReplicaRead[str]:
it: Final = iter(answers)
return lambda _timeout: next(it)
@ -390,6 +434,7 @@ class TestReplicasFor:
base_url="http://lb",
control_plane_base_url="http://backend",
replica_urls=("http://gateway-1", "http://gateway-2"),
control_replica_urls=("http://backend",),
)
assert set(client.replicas_for("/key/info")) == {"http://backend"}
assert set(client.replicas_for("/project/info")) == {"http://backend"}
@ -400,9 +445,63 @@ class TestReplicasFor:
base_url="http://lb",
control_plane_base_url="http://lb",
replica_urls=("http://pod-1", "http://pod-2"),
control_replica_urls=("http://pod-1", "http://pod-2"),
)
assert set(client.replicas_for("/key/info")) == {"http://pod-1", "http://pod-2"}
def test_gateway_pods_behind_one_router_read_management_routes_back_from_the_router(self) -> None:
"""The Buildkite PR stack names each gateway pod in PROXY_REPLICA_URLS while
both planes share the router base, so a management read-back polls the
router (CONTROL_PLANE_REPLICA_URLS) rather than the pods, which trim
management routes, while a data-plane read-back still polls every pod."""
client: Final = build_proxy_client(
base_url="http://router",
control_plane_base_url="http://router",
replica_urls=("http://10.0.0.1:4000", "http://10.0.0.2:4000"),
control_replica_urls=("http://router",),
)
assert set(client.replicas_for("/key/info")) == {"http://router"}
assert set(client.replicas_for("/v1/models")) == {"http://10.0.0.1:4000", "http://10.0.0.2:4000"}
def test_a_client_built_for_another_proxy_reads_management_routes_back_from_that_proxy(self) -> None:
"""A caller that points the client at its own server (test_provider_cache.py)
names no control list, so the derived one has to follow that server rather
than the env proxy, on a shared base and on split ones alike."""
local: Final = build_proxy_client(
base_url="http://local", control_plane_base_url="http://local", replica_urls=("http://local",)
)
assert set(local.replicas_for("/key/info")) == {"http://local"}
assert set(local.replicas_for("/v1/models")) == {"http://local"}
split: Final = build_proxy_client(
base_url="http://lb", control_plane_base_url="http://backend", replica_urls=("http://gateway-1",)
)
assert set(split.replicas_for("/key/info")) == {"http://backend"}
assert set(split.replicas_for("/v1/models")) == {"http://gateway-1"}
def test_management_read_backs_poll_the_control_replicas_only(self) -> None:
"""A gateway pod answers /key/info 404 even after the write landed on the
control plane, so a read-back that polled the data-plane replicas for it
would never converge there."""
with caller_boundary(status=404) as (pod, pod_headers), caller_boundary() as (router, router_headers):
pod_url: Final = next(iter(pod.proxy.replicas))
router_url: Final = next(iter(router.proxy.replicas))
proxy: Final = build_proxy_client(
base_url=router_url,
control_plane_base_url=router_url,
replica_urls=(pod_url,),
control_replica_urls=(router_url,),
master_key="bootstrap",
)
read: Final = proxy.read_back_everywhere(
"/key/info",
params=NoBody(),
response_type=KeyInfoResponse,
converged=lambda result: isinstance(result, Success),
)
assert set(read) == {router_url}
assert router_headers.get_nowait() == "Bearer bootstrap"
assert router_headers.empty() and pod_headers.empty()
def test_mcp_admin_routes_read_back_from_every_data_plane_replica(self) -> None:
"""/v1/mcp/* is a lazily mounted feature, so a data-plane replica serves it
too and answers from its own in-memory registry. Routing it to the control
@ -412,6 +511,7 @@ class TestReplicasFor:
base_url="http://lb",
control_plane_base_url="http://backend",
replica_urls=("http://gateway-1", "http://gateway-2"),
control_replica_urls=("http://backend",),
)
assert set(client.replicas_for("/v1/mcp/server/abc")) == {"http://gateway-1", "http://gateway-2"}
assert set(client.replicas_for("/v1/mcp/toolset/abc")) == {"http://gateway-1", "http://gateway-2"}
@ -531,6 +631,7 @@ class TestSplitCallerPropagation:
base_url=data_url,
control_plane_base_url=control_url,
replica_urls=(data_url,),
control_replica_urls=(control_url,),
master_key="bootstrap",
).with_caller(Caller(credential="tenant-token", kind="direct_jwt", role="team_member"))
proxy.key_info("owned")
@ -543,8 +644,13 @@ class TestSplitCallerPropagation:
response_type=KeyInfoResponse,
converged=lambda result: isinstance(result, Success),
)
assert control_headers.get_nowait() == "Bearer tenant-token"
assert control_headers.get_nowait() == "Bearer tenant-token"
proxy.read_back_everywhere(
"/v1/models",
params=NoBody(),
response_type=ModelsListResponse,
converged=lambda result: isinstance(result, Success),
)
assert tuple(control_headers.get_nowait() for _ in range(3)) == ("Bearer tenant-token",) * 3
assert data_headers.get_nowait() == "Bearer tenant-token"
assert control_headers.empty() and data_headers.empty()