Credential create/update/delete and reads were gated to the proxy admin for every
credential, which tightened provider-credential management for keys an admin had
delegated /credentials to via allowed_routes; a backward-incompatible change beyond
the OTEL trace-destination feature.
Restrict the gate to trace destinations (credential_type="logging"). Those require the
proxy admin regardless of allowed_routes, so a non-admin cannot create a global
destination and receive other tenants' traces, nor convert a provider credential into
one. Provider credentials keep their existing route-level authorization. Reads follow
suit: the list hides destinations from non-admins and by_name returns 403 for a
destination, while provider credentials read as before.
The Add Logging Callback dropdown was filtering out langfuse (v2 SDK),
otel (Open Telemetry), and the other non-OTEL callbacks along with the
OTEL trace destinations, so those integrations disappeared from the UI.
Restrict the destination filtering to the OTEL trace-destination
backends only (arize, langfuse_otel, weave_otel, generic), so non-OTEL
logging callbacks keep their existing global-callback path.
Also relabel the langfuse_otel destination to "Langfuse OTEL" so it is
distinct from the Langfuse SDK entry, collapse the duplicate
destination-id set into a single LOGGING_BACKEND_IDS, and format the
logging UI files.
Three review findings on the auth/key hot paths, all resolver/DB redundancy rather
than correctness bugs.
_hoist_request_destinations fires early in the auth builder and again as an outer
catch-all, so _resolve_logging_exporters ran twice per authenticated request. It now
early-returns when request.state already holds the result, keeping the early hoist
(for span timing) and the catch-all (a failed first call leaves state unset and still
retries) while resolving once.
OtelDestinationParams gains the resource_attributes key the resolver populates and the
hoist reads, so the resolver's dict type-checks.
regenerate_key_fn loaded the key's team via get_team_object on every team-key
regeneration, but only uses it when the request sets access_group_ids or
object_permission. The fetch moves back inside that guard, so an ordinary regenerate
does no extra DB round-trip; an access-group/object-permission regenerate still fetches.
Verified live: one chat resolves destinations once (was twice); a plain regenerate does
zero team fetches (was one) while an object_permission regenerate still does one; fan-out
isolation unchanged.
Several docstrings and comment blocks in the otel v2 plumbing ran 10-35 lines of
rationale where a sentence or two states what the code does; the how and why belong
in the PR and design notes. Trims them across routing, context, providers, logger,
and the metadata model, keeping the load-bearing invariants (no-op-when-empty, the
MCP anti-spoof rule, post-auth ordering, gen-AI-span skipping) in compressed form.
No code changes: the executable tokens are byte-identical, only docstrings and
comments moved.
* fix(mcp): resolve call_tool by registry without requiring tool map
Multi-worker reloads put MCP servers in the registry from the DB but do
not re-run tools/list on every process. Gating call_tool on
tool_name_to_mcp_server_name_mapping made cold workers 500 with Tool not
found after another worker had already listed the tool. Treat a registry
match on server id/name/alias as enough; upstream rejects unknown tools
* test(e2e): poll MCP register, tools/list, and tools/call across multi-worker lag
Stage multi-worker gateways only load MCP servers and tool maps on the
process that handled the request. Poll until the server is listed, the
tool appears on tools/list, and tools/call is not a cold-worker 500 so
key-access and Datadog MCP e2e stop racing the LB
* Revert "fix(mcp): resolve call_tool by registry without requiring tool map"
This reverts commit 8b56e51e39.
* test(e2e): tighten MCP multi-worker lag classifier
Only retry tools/call on gateway shapes Tool <name> not found and
server_not_found, not any 500 that mentions tool/server not found, so
upstream failures are not retried until the poll deadline
* test(e2e): drop unit file for MCP lag classifier
The live await_call_tool polls already cover multi-worker lag; a separate
string-match unit module is not worth keeping
The two regression-test docstrings recounted the bug history and repro, which
belongs in the PR description rather than the code; they now state only what the
test checks. Also drops a leftover section comment and updates a few docstring
phrases that still said "assignment" to describe the access grant, since access is
the only routing input now.
A logging destination's credential_info.access (global / teams / orgs) now fully
decides which requests it receives; a destination fires for a request exactly when
its access grants the request's team or org. This removes the second, redundant way
to express the same team-to-destination mapping that the admin-only model left
behind: the auto_enable flag and the per-team/org logging_exporters assignment
column both existed for tenant self-service opt-in, and once assignment became
proxy-admin-only they only duplicated what access already says.
Removed: the auto_enable field on CredentialInfo; the team and organization
logging_exporters columns and their assignment gate (validate_logging_exporter_field
/ validate_logging_exporter_assignment); the request-time naming union in
litellm_pre_call_utils; and the dashboard's per-team/org destination picker and the
"Enable for entire scope" toggle. The access-shape validator stays, the credential's
access fields stay, and /team/info and /organization/info still disclose
resolved_logging_exporters computed from access alone.
This also removes the /v2/organization write that two review bots flagged (there is
no longer a logging_exporters field on that endpoint) and the "(via scope)" UI
ambiguity that came from carrying two representations of the same mapping.
Verified live on a 2-org / 4-team matrix against Langfuse, Arize, Weave, a generic
OTLP collector, and a self-hosted Phoenix: per-team and per-org isolation, empty
access as deny-all, injection defense, admin-only credential management, and
complete trace trees read back from each destination's own API.
An auto-router deployment's litellm_params.model (auto_router/...) is the
discriminator the router loads it by, but the model management endpoints
accepted any client-supplied value verbatim; a doubled or stripped prefix
made router init fail on the next load and ignore_invalid_deployments
silently dropped the deployment. Validate writes that supply
litellm_params.model at all three endpoints against the merged params and
reject incoherent values with an actionable 400. Classification is
extracted to router_utils/auto_router_model_naming.py so the Router
predicates and the validation share one source
Every model-write endpoint returned 200 off the DB write alone; a model the
reload dropped (ignore_invalid_deployments, or a wholesale reload failure)
stayed invisible on every channel at once, which is how the registry-leak
defect went undiagnosed for three weeks. ProxyConfig.add_deployment and
clear_cache now return whether the reload pass completed, and each write
endpoint verifies the rows it wrote are live in this pod's router afterwards,
distinguishing a deliberately environment-inactive model via the same
predicate the Router's own gate uses. The access-group writers return the
mutated id set instead of discarding it
A request rejected before any upstream call (rate limit, budget, pre-call
guardrail) fires the failure callback with a built payload while the
auth-hoisted destinations are still resolved, so _close_llm_call's carrier-None
branch fabricated a gen-AI span for a call that never reached a provider. It now
re-reads call.is_no_upstream_call, the same marker log_pre_api_call already
honors, so a rejected request no longer lands a fake "chat <model>" span in the
tenant's destination. Reproduced live against a real destination (a 429 emitted
a chat span in Arize before the fix, none after) and pinned by a regression test
that fails on the pre-fix code
The pre-call resolver re-run in _apply_admin_logging_exporters was not wrapped,
unlike the auth-time hoist, so a non-HTTPException from the org fallback lookup
could abort a real request; it is now best-effort so admin-owned telemetry setup
can never break request handling
Also drops the explanatory inline comments this feature added across the otel v2
modules, keeping only docstrings and lint suppressions per the repository
convention that new code carries no comments
Bitnami retired the versioned tags under docker.io/bitnami and republished
the archived builds under docker.io/bitnamilegacy, so every install and
upgrade of the chart with the bundled database fails to pull
docker.io/bitnami/postgresql:16.2.0-debian-12-r6. Repoint the subchart
images at the bitnamilegacy copies of the exact builds those subchart
versions shipped with, so the on-disk data directory layout is unchanged
for existing installs.
Pin the subchart dependency ranges to the versions already in Chart.lock.
The current bitnami postgresql chart defaults to `tag: latest`, which is
PostgreSQL 18 today, so an open-ended range turns a dependency refresh
into a major-version jump on an existing volume.
Refuse to render when postgresql.image.tag is empty or `latest` while the
bundled database is deployed. Starting a different PostgreSQL major
against an existing data directory leaves the server unable to boot with
no in-place way back, which is how the reported install lost its data.
Resolves LIT-4708
Destinations now bind at tenancy granularity only. The per-key logging_exporters
surface was half-shipped (edit-form picker but no create flow) and unrequested,
so it goes: the column leaves the key tables and the migration, /key/generate,
/key/update, and /key/regenerate stop accepting the field, the resolver's union
reads team and org columns only, and the key pages drop the picker and exporter
badges. A key's traces route by its team and org, which the live check confirms:
a brand-new team key exports to the team's destinations with no assignment.
Subset targeting below a whole scope remains available at team granularity
(a multi-team scope with enable-for-entire-scope off, named on specific teams).
Re-adding key granularity later is a purely additive column and field.
Viewers only need to know which destinations receive their traces; whether the
binding came from an assignment or the destination's scope is admin plumbing, so
the (via scope) suffix goes. The add-destination toggle is relabeled Enable for
entire scope with a tooltip spelling out on (whole scope exports automatically)
versus off (only explicitly assigned keys, teams, or orgs export)
The team and org pages showed different Logging Exporters lists per role: the
(via scope) badges were derived client-side from GET /credentials, which is
proxy-admin only, so non-admin viewers saw only the identity's own assignments.
Add resolved_logging_exporters to the /team/info and /organization/info
responses: the destination names that will receive the identity's traces,
computed server-side with the same selection the request-time resolver uses
(access grants the identity AND auto_enable or named). Names only; endpoints,
headers, and the access map stay proxy-admin information. The UI renders the
badges from this field, deleting the client-side credentials derivation, so
every role sees the identical list.
* fix(proxy): warm rotate Prisma client for IAM refresh
* fix(proxy): drain Prisma operations during IAM rotation
* fix(proxy): bound the drain wait when retiring a replaced prisma engine
A replaced engine waited indefinitely for its drain tracker to empty.
Hung queries self-release via prisma's 30s default HTTP timeout, but a
transaction whose owner is hard-cancelled before commit/rollback leaks
its drain count forever, keeping the retired engine and its DB
connection pool alive indefinitely; at one rotation per 12 minutes such
engines accumulate. Cap the wait at 90 seconds, which exceeds every
legitimate operation bound (30s HTTP timeout, 60s max interactive
transaction timeout in this codebase), then kill the engine anyway.
Work killed at the deadline degrades to the pre-drain behavior and is
retried by the existing reconnect/backoff layers.
---------
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
Scrub aliases on delete only when the deleted deployment's model_name no
longer resolves in the router. A legacy load-balanced team model can have
several deployment rows sharing one internal name; deleting one replica
must not remove aliases that still route to the survivors, in any team
A team's model_aliases can map a public name like gpt-4 to the internal
routing key (model_name_{team_id}_{uuid}) of a team deployment that has
since been deleted, e.g. after replacing per-team duplicates with one
gateway-level model. The pre-call rewrite then sent every request to a
name the router cannot serve, failing with "no healthy deployments for
model_name_..." even though the requested name still resolves at the
gateway level. The rewrite is now skipped when the alias target has no
live deployment in the router
delete_model also skipped the team alias scan for internal-shaped names
on the assumption they can never be alias values, which is exactly the
shape legacy team model aliases have, so deleting a legacy team model
left the stale alias behind. The scan now always runs, and a public
name that still resolves to a live router deployment (e.g. a shared
gateway-level model group) stays in team.models so the delete does not
revoke the team's access to it
The cache-hit and paid rows for the two driver calls flush from different
pods on independent update_spend timers, so waiting only for the cache-hit
row can return a half-arrived result set where the paid-row assertion then
fails on an empty list. Requiring both row kinds in the poll predicate lets
the existing deadline absorb the slower flush without weakening any assertion
The lint job failed on LIT001: the branch added 36 mutable-collection annotations
and the merge picked up staging's ratcheted budget. Re-annotate the new surface
with read-only views (Mapping/Sequence/frozenset/tuple), build the resolver's
union and destination results functionally, and freeze the span-router grouping
at its boundary; the four genuinely mutable LRU caches and in-place merge targets
carry mutable-ok reasons. Ratchet both budgets down by what the branch now fixes.
Also drop the wholesale credential mask for the admin viewer: it existed for the
removed tenant read, and the viewer principle is read parity with the proxy
admin. Both admin-tier readers now go through _get_masked_values exactly as on
staging, and otel_headers joins the masker's sensitive keys so the collector
auth it carries is masked for every reader.
Full-diff audit against the admin-only design surfaced leftovers in three layers.
Backend: drop dead code the refactor stranded (is_admin_gated_credential_info,
is_destination_visible, the CredentialInfo decider-era fields, the unused
LLMCallEvent.dynamic_params carrier, the _is_user_org_admin_for_org_id extraction)
and correct every comment/docstring still describing the removed team-admin
self-service or scoped-read designs.
UI: gate the Logging Exporters form rows behind a shared proxy-admin-only
LoggingExportersFormItem so non-admin forms no longer render an orphaned label;
skip the GET /credentials fetch for roles it would 403 (team/org/key views); make
the callbacks table read-only for the admin viewer (no Add/Edit/Delete actions
that would 401); reword access tooltips to routing-scope semantics; regenerate
schema.d.ts from the corrected endpoint docstrings.
Tests: flip the credential-migration non-admin expectation to the route-gate 401,
port the visibility tests to access_grants, drop tests of the removed helpers,
and update copy assertions and stale rationale to the admin-only contract.
* fix(gateway): route /a2a through the gateway component
A2A message-send runs the completion bridge, an outbound LLM call, but the
ingress only listed /v1/a2a so the serving routes at /a2a/{agent_id} fell to
the backend catch-all. Backend pods hold no provider credentials, so every
invocation died with a missing-provider-key auth error while the same call
succeeds on the gateway fleet. Adds /a2a to the ingress gateway prefixes and
the gateway route allowlist, plus a parity test so an ingress prefix that the
gateway trims can never reappear
* revert(test): drop the allowlist parity tests
---------
Co-authored-by: yuneng-jiang <yuneng@berri.ai>
* Fix cache leakage card layout to keep date picker on right and prevent content overlap
Removes flex-wrap and mt-3 to ensure date picker stays pinned to the right side of the card header regardless of zoom level, preventing it from covering card content below
* Remove overflow-hidden from Card to allow dropdowns and overlays to display fully
Fixes date picker dropdown being clipped when opened in cards like the Cache Leakage Card. By removing overflow-hidden from the Card container, popovers, dropdowns, and other overflow content can now display properly without being clipped by the card boundaries.
* Make cache leakage card descriptions consistent with line clamping
Adds line-clamp-2 to ensure both 'by model' and 'by virtual key' cards maintain consistent height. Removes conditional anthropic-specific text that caused height variations between dimensions.
The /key/update path fetched the key's team into _key_team only to feed the
team-admin/org-admin assignment flags removed in the admin-only refactor; it is
now unused, so drop the wasted DB round-trip. The resolver's org lookup passed
check_db_only=True, forcing a synchronous DB query on every team-key request even
though the team is already resident in the cache from auth; drop it so the lookup
is cache-first. Also correct two comments that still claimed team or org admins
may assign logging_exporters.
Admin-owned OTEL v2 logging destinations and their access scoping are now managed
only by the proxy admin. Trace routing to identity-scoped destinations is unchanged
because it runs server-side in the resolver; this removes only the tenant-facing
read/write surface the earlier revision exposed.
GET/POST/PATCH/DELETE /credentials are proxy-admin only again (a proxy-admin-viewer
may read); the two credential routes leave self_managed_routes, and the non-admin
scoped list, the team-admin PATCH self-service grant, and the access_decision decider
are deleted. Assigning logging_exporters on a key, team, or org is proxy-admin only,
dropping the team-admin and org-admin widening. In the UI the logging destinations
table and the exporter picker render only for a proxy admin, so non-admins no longer
call GET /credentials.
* test(e2e): harden harness and tests against data-plane pod churn
A stage autoscaler scale-down produced a 2s window of ALB 502s that killed six
budget tests on their first management call, and a freshly scaled-up pod that
had not run its 30s DB object sync yet failed two MCP tests and one prometheus
cardinality test. Retry transient gateway errors (502/503/504, connection
errors) once at the shared e2e_http dispatch seam, poll MCP server registration
to the poll deadline instead of asserting a single-shot listing, anchor the MCP
guardrail full-sync wait to the later of the guardrail and server writes, and
turn the prometheus alias poll into a drive-and-scrape convergence loop that
re-sends traffic for missing aliases and unions results across scrapes
* test(e2e): drain request body in retry stub handler so keep-alive reuse cannot misparse leftovers as requests
* revert(e2e): drop the transient-502 retry seam
A raw 502 during a pod scale-down is what a real client sees, so the suite
retrying past it hides an availability gap instead of flagging it. The
gateway-side fix is graceful drain on the deployment; until then the failures
are signal
* test(e2e): cap per-alias driver re-drives in the prometheus cardinality poll
Bounds worst-case provider spend to 4 completions per alias while scrapes keep
polling to the deadline; counters persist on whichever pod served them, so the
cap costs no convergence unless that pod dies
* test(e2e): drop driver re-drives from the prometheus cardinality poll
The per-key cardinality contract is process-local and counters persist on
whichever pod served the driver call, so unioning aliases across free scrape
polls converges without re-sending billable traffic. The residual gap, a pod
dying inside the poll window, is deferred to direct per-pod scraping
* test(e2e): let the ui suite run from a read-only cwd
The playwright suite never executed on stage. It died in globalSetup before a
single test ran, and the reported error was a red herring.
/app/e2e/ui is a read-only filesystem in the packaged e2e image (the image
runner already redirects playwright's own artifacts to TMPDIR for this reason),
but the suite wrote three things relative to cwd: the per-role storageState
files, the failure-screenshot directory, and the html report. Reproduced in the
pod: storageState raises EROFS, mkdir test-results raises ENOENT.
Worse, the catch block that exists to capture a screenshot threw its own ENOENT
while handling a failure, so the real login error was replaced by a filesystem
error. That is why the run looked like a missing directory rather than whatever
actually went wrong.
Route every artifact through ARTIFACT_DIR (E2E_UI_ARTIFACT_DIR, default "." to
keep run_e2e.sh behavior unchanged), make the diagnostic screenshot best-effort
so it can never mask the underlying failure, and point playwright's reporter and
outputDir at the same place so a bare `npx playwright test` works there too.
fixtures/users.ts had its own copy of the five storageState filenames; it now
re-exports the ones from constants so the paths have a single definition.
Verified in the read-only pod: both writes fail before, both succeed after.
85 tests enumerate and tsc --noEmit is clean.
Refs LIT-4821
* fix(e2e): create the ui artifact root before writing into it
storageState() does not create missing parents, and nothing created ARTIFACT_DIR
itself. Pointing E2E_UI_ARTIFACT_DIR at a writable path that did not exist yet
therefore failed with ENOENT on the very first role's snapshot, before any UI
test ran; the same class of failure the artifact-dir change was meant to remove,
just moved one level up.
Reproduced: writing admin.storageState.json into a missing directory raises
ENOENT. My earlier pod verification masked this because the probe called
mkdirSync itself, which the real code path never did.
mkdir the root once at the top of globalSetup, before the login loop. recursive
makes it idempotent, handles nested paths, and keeps the default "." a no-op.
Playwright creates its own outputDir lazily, so globalSetup is the only place
that needs this, and migration.serverRootPath.globalSetup delegates here so it is
covered too.
* test(e2e): skip the mid-conversation cache checks pending LIT-4873
A mid-conversation role="system" reminder invalidates the prompt cache on the
vertex_ai, azure_ai and bedrock_invoke Messages paths. Measured on the reminder
turn, same conversation shape throughout:
direct to api.anthropic.com 7013 read cache preserved
litellm -> anthropic/claude-opus-4-8 7013 read cache preserved
litellm -> vertex_ai/claude-opus-4-8 0 read cache destroyed
and the Vertex control with the same added assistant/user turns but no reminder
reads 7013, so it is the reminder on the non-first-party paths and not the extra
turns. Anthropic keeping the cache rules out provider behavior; litellm's
first-party anthropic path keeping it rules out the shared Messages transform.
That makes these assertions correct and the failure a real billing bug, so the
tests are skipped rather than weakened; the bodies stay intact and must be
restored unchanged with the fix. Registry rows are left in place, so the three
mid_conversation_system.nonstream.cache_hit cells report as uncovered gaps.
Skips are decorators rather than a pytest.skip() inside the shared helper: a
mid-function skip fires only after setup has already registered a real
deployment via /model/new and left the rest of the body unreachable.
Only Vertex was measured end to end. Azure Foundry and Bedrock Invoke are
inferred from matching nightly failures and should be confirmed with the fix.
Refs LIT-4821, LIT-4873
Both specs assert against UI that has since moved, so they fail on selectors
rather than on behavior.
The MCP discovery modal became a shadcn/Base UI dialog when mcp-servers
migrated off antd, so `.ant-modal` no longer matches it; locate it by its
dialog role instead. The create form below it is still an antd Modal and keeps
its existing locator.
The no-team internal user has no keys, and a keyless non-admin is now sent to
/ui/connect on the post-login landing, which has no sidebar. Wait for that
redirect to settle, then navigate to the keys page explicitly; the redirect is
gated on the ?login=success marker that the fresh navigation drops, so the
dashboard sticks and the rest of the test is unchanged.
* test(reasoning-effort-grid): bump cell-count assertion for claude-opus-5
The claude-opus-5 grid entry added in ae81625ee6 raised the Anthropic direct
route to 31 model combos, but test_grid_cell_count still expected 30, so the
suite went red on the tripwire rather than on any behavior change.
* test(openai): swap the retired deep-research model out of the bridge test
OpenAI shut down o3-deep-research and o4-mini-deep-research on 2026-07-23, so
the live call in this test now comes back as a 400 'Model not found'. The test
was never about deep research specifically; the bridge fires on any model whose
cost-map mode is "responses", so it now uses gpt-5.5-pro, the newest
responses-only OpenAI model, and is renamed to say that.
gpt-5.5-pro was confirmed present on the CI account with an authenticated
GET /v1/models before being picked.
The langtrace v1 exporter host/header fix (https://app.langtrace.ai/api/trace +
x-api-key) and the shared endpoint-normalizer /api/trace carve-out are unrelated
to admin-owned OTEL v2 destinations; they change behavior on the legacy langtrace
v1 path. Moved to its own PR (#34865) so this PR only touches OTEL v2. The v2
langtrace preset and its normalizer in otel/plumbing/providers.py are unaffected.