Commit graph

41719 commits

Author SHA1 Message Date
Yucheng Zhu
6829988c70 fix(credentials): scope the proxy-admin gate to trace destinations
Credential create/update/delete and reads were gated to the proxy admin for every
credential, which tightened provider-credential management for keys an admin had
delegated /credentials to via allowed_routes; a backward-incompatible change beyond
the OTEL trace-destination feature.

Restrict the gate to trace destinations (credential_type="logging"). Those require the
proxy admin regardless of allowed_routes, so a non-admin cannot create a global
destination and receive other tenants' traces, nor convert a provider credential into
one. Provider credentials keep their existing route-level authorization. Reads follow
suit: the list hides destinations from non-admins and by_name returns 403 for a
destination, while provider credentials read as before.
2026-07-29 16:34:17 -07:00
Yucheng Zhu
440277c97b fix(ui): keep non-OTEL logging callbacks in the Add Callback list
The Add Logging Callback dropdown was filtering out langfuse (v2 SDK),
otel (Open Telemetry), and the other non-OTEL callbacks along with the
OTEL trace destinations, so those integrations disappeared from the UI.
Restrict the destination filtering to the OTEL trace-destination
backends only (arize, langfuse_otel, weave_otel, generic), so non-OTEL
logging callbacks keep their existing global-callback path.

Also relabel the langfuse_otel destination to "Langfuse OTEL" so it is
distinct from the Langfuse SDK entry, collapse the duplicate
destination-id set into a single LOGGING_BACKEND_IDS, and format the
logging UI files.
2026-07-29 16:19:46 -07:00
Yucheng Zhu
738a37613e Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_otel_v2_admin_owned_destinations
# Conflicts:
#	tests/test_litellm/proxy/test_litellm_pre_call_utils.py
2026-07-29 00:55:59 -07:00
Yucheng Zhu
8cc922aca5 perf(otel/v2): resolve destinations once per request; scope regenerate team fetch
Three review findings on the auth/key hot paths, all resolver/DB redundancy rather
than correctness bugs.

_hoist_request_destinations fires early in the auth builder and again as an outer
catch-all, so _resolve_logging_exporters ran twice per authenticated request. It now
early-returns when request.state already holds the result, keeping the early hoist
(for span timing) and the catch-all (a failed first call leaves state unset and still
retries) while resolving once.

OtelDestinationParams gains the resource_attributes key the resolver populates and the
hoist reads, so the resolver's dict type-checks.

regenerate_key_fn loaded the key's team via get_team_object on every team-key
regeneration, but only uses it when the request sets access_group_ids or
object_permission. The fetch moves back inside that guard, so an ordinary regenerate
does no extra DB round-trip; an access-group/object-permission regenerate still fetches.

Verified live: one chat resolves destinations once (was twice); a plain regenerate does
zero team fetches (was one) while an object_permission regenerate still does one; fan-out
isolation unchanged.
2026-07-29 00:23:14 -07:00
Yucheng Zhu
da4583a110 docs(otel/v2): tighten verbose docstrings and comments
Several docstrings and comment blocks in the otel v2 plumbing ran 10-35 lines of
rationale where a sentence or two states what the code does; the how and why belong
in the PR and design notes. Trims them across routing, context, providers, logger,
and the metadata model, keeping the load-bearing invariants (no-op-when-empty, the
MCP anti-spoof rule, post-auth ordering, gen-AI-span skipping) in compressed form.
No code changes: the executable tokens are byte-identical, only docstrings and
comments moved.
2026-07-28 23:47:06 -07:00
mubashir1osmani
c274cf321c
test(e2e): poll MCP tools across multi-worker lag (#35047)
* fix(mcp): resolve call_tool by registry without requiring tool map

Multi-worker reloads put MCP servers in the registry from the DB but do
not re-run tools/list on every process. Gating call_tool on
tool_name_to_mcp_server_name_mapping made cold workers 500 with Tool not
found after another worker had already listed the tool. Treat a registry
match on server id/name/alias as enough; upstream rejects unknown tools

* test(e2e): poll MCP register, tools/list, and tools/call across multi-worker lag

Stage multi-worker gateways only load MCP servers and tool maps on the
process that handled the request. Poll until the server is listed, the
tool appears on tools/list, and tools/call is not a cold-worker 500 so
key-access and Datadog MCP e2e stop racing the LB

* Revert "fix(mcp): resolve call_tool by registry without requiring tool map"

This reverts commit 8b56e51e39.

* test(e2e): tighten MCP multi-worker lag classifier

Only retry tools/call on gateway shapes Tool <name> not found and
server_not_found, not any 500 that mentions tool/server not found, so
upstream failures are not retried until the poll deadline

* test(e2e): drop unit file for MCP lag classifier

The live await_call_tool polls already cover multi-worker lag; a separate
string-match unit module is not worth keeping
2026-07-28 22:15:21 -07:00
Napuh
2f7574d7c1
fix(anthropic-adapter): open the first content block with the real upstream type so reasoning-first streams start with thinking (#34433)
* fix(anthropic-adapter): open first content block with the real upstream type

* fix(anthropic): defer blank leading stream deltas
2026-07-28 21:51:59 -07:00
tin-berri
7ac5172686
Merge pull request #33290 from mihidumh/fix/latency-routing-timedelta
fix(router_strategy): serialize latency for non-chat responses in lowest-latency routing
2026-07-28 21:50:50 -07:00
Yucheng Zhu
762b9eb1b4 chore(otel/v2): trim narrative test docstrings and stale assignment wording
The two regression-test docstrings recounted the bug history and repro, which
belongs in the PR description rather than the code; they now state only what the
test checks. Also drops a leftover section comment and updates a few docstring
phrases that still said "assignment" to describe the access grant, since access is
the only routing input now.
2026-07-28 21:24:41 -07:00
tin-berri
9b48bf6084
Merge pull request #34151 from BerriAI/litellm_lit4663_autorouter_prefix
fix(proxy): reject model writes that corrupt an auto-router pseudo-model
2026-07-28 21:01:42 -07:00
Yucheng Zhu
4ff358fcb8 refactor(otel/v2): make destination access the sole routing determinant
A logging destination's credential_info.access (global / teams / orgs) now fully
decides which requests it receives; a destination fires for a request exactly when
its access grants the request's team or org. This removes the second, redundant way
to express the same team-to-destination mapping that the admin-only model left
behind: the auto_enable flag and the per-team/org logging_exporters assignment
column both existed for tenant self-service opt-in, and once assignment became
proxy-admin-only they only duplicated what access already says.

Removed: the auto_enable field on CredentialInfo; the team and organization
logging_exporters columns and their assignment gate (validate_logging_exporter_field
/ validate_logging_exporter_assignment); the request-time naming union in
litellm_pre_call_utils; and the dashboard's per-team/org destination picker and the
"Enable for entire scope" toggle. The access-shape validator stays, the credential's
access fields stay, and /team/info and /organization/info still disclose
resolved_logging_exporters computed from access alone.

This also removes the /v2/organization write that two review bots flagged (there is
no longer a logging_exporters field on that endpoint) and the "(via scope)" UI
ambiguity that came from carrying two representations of the same mapping.

Verified live on a 2-org / 4-team matrix against Langfuse, Arize, Weave, a generic
OTLP collector, and a self-hosted Phoenix: per-team and per-org isolation, empty
access as deny-all, injection defense, admin-only credential management, and
complete trace trees read back from each destination's own API.
2026-07-28 20:52:30 -07:00
Tin Chi Lo
1e04aee089 fix(proxy): reject model writes that corrupt an auto-router pseudo-model
An auto-router deployment's litellm_params.model (auto_router/...) is the
discriminator the router loads it by, but the model management endpoints
accepted any client-supplied value verbatim; a doubled or stripped prefix
made router init fail on the next load and ignore_invalid_deployments
silently dropped the deployment. Validate writes that supply
litellm_params.model at all three endpoints against the merged params and
reject incoherent values with an actionable 400. Classification is
extracted to router_utils/auto_router_model_naming.py so the Router
predicates and the validation share one source
2026-07-28 20:25:07 -07:00
tin-berri
1a6642ee2e
Merge pull request #34861 from BerriAI/litellm_lit4872_surface_reload_drop
fix(proxy): report when a model write does not survive the post-write reload
2026-07-28 20:18:02 -07:00
Mateo Wang
2bb297efa0
Merge pull request #34993 from BerriAI/claude/auto-til-blocked-cwalrj
fix(proxy): skip team model aliases that point at deleted deployments
2026-07-28 20:09:20 -07:00
Mateo Wang
c542e74b68
Merge pull request #34222 from BerriAI/litellm_jwt_v1_messages_team_route_1784693761
fix(jwt_auth): allow /v1/messages for JWT teams by default
2026-07-28 19:59:26 -07:00
mateo-berri
b592a37b8d test(proxy): cover stale-alias warning dedup and key-cache eviction 2026-07-28 19:39:30 -07:00
Tin Chi Lo
9f4e3c6009 fix(proxy): report when a model write does not survive the post-write reload
Every model-write endpoint returned 200 off the DB write alone; a model the
reload dropped (ignore_invalid_deployments, or a wholesale reload failure)
stayed invisible on every channel at once, which is how the registry-leak
defect went undiagnosed for three weeks. ProxyConfig.add_deployment and
clear_cache now return whether the reload pass completed, and each write
endpoint verifies the rows it wrote are live in this pod's router afterwards,
distinguishing a deliberately environment-inactive model via the same
predicate the Router's own gate uses. The access-group writers return the
mutated id set instead of discarding it
2026-07-28 18:52:06 -07:00
Mateo Wang
711be72512
Merge pull request #34816 from BerriAI/litellm_model_prices_json_schema
ci: publish a generated JSON schema for model_prices_and_context_window.json
2026-07-28 18:03:59 -07:00
tin-berri
32a4377acd
Merge pull request #34589 from BerriAI/litellm_lit4798_glm_stop_thinking
fix(anthropic-adapter): translate stop_sequences and disabled thinking for non-Claude targets
2026-07-28 17:42:27 -07:00
tin-berri
898c4e93bc
Merge pull request #34672 from BerriAI/litellm_lit4761_vertex_passthrough_stream
fix(vertex): decide rawPredict passthrough streaming from the request body
2026-07-28 17:40:58 -07:00
Yucheng Zhu
a0622a2aab fix(otel/v2): skip phantom LLM span on rejected requests; harden telemetry setup
A request rejected before any upstream call (rate limit, budget, pre-call
guardrail) fires the failure callback with a built payload while the
auth-hoisted destinations are still resolved, so _close_llm_call's carrier-None
branch fabricated a gen-AI span for a call that never reached a provider. It now
re-reads call.is_no_upstream_call, the same marker log_pre_api_call already
honors, so a rejected request no longer lands a fake "chat <model>" span in the
tenant's destination. Reproduced live against a real destination (a 429 emitted
a chat span in Arize before the fix, none after) and pinned by a regression test
that fails on the pre-fix code

The pre-call resolver re-run in _apply_admin_logging_exporters was not wrapped,
unlike the auth-time hoist, so a non-HTTPException from the org fallback lookup
could abort a real request; it is now best-effort so admin-owned telemetry setup
can never break request handling

Also drops the explanatory inline comments this feature added across the otel v2
modules, keeping only docstrings and lint suppressions per the repository
convention that new code carries no comments
2026-07-28 16:40:03 -07:00
Yassin Kortam
cd9c410ae2
fix(helm): pin bundled postgres and redis to the bitnamilegacy images (#34963)
Some checks are pending
CodSpeed Benchmarks / benchmarks (push) Waiting to run
UI Unit Tests / ui-unit-tests (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Bitnami retired the versioned tags under docker.io/bitnami and republished
the archived builds under docker.io/bitnamilegacy, so every install and
upgrade of the chart with the bundled database fails to pull
docker.io/bitnami/postgresql:16.2.0-debian-12-r6. Repoint the subchart
images at the bitnamilegacy copies of the exact builds those subchart
versions shipped with, so the on-disk data directory layout is unchanged
for existing installs.

Pin the subchart dependency ranges to the versions already in Chart.lock.
The current bitnami postgresql chart defaults to `tag: latest`, which is
PostgreSQL 18 today, so an open-ended range turns a dependency refresh
into a major-version jump on an existing volume.

Refuse to render when postgresql.image.tag is empty or `latest` while the
bundled database is deployed. Starting a different PostgreSQL major
against an existing data directory leaves the server unable to boot with
no in-place way back, which is how the reported install lost its data.

Resolves LIT-4708
2026-07-28 16:19:20 -07:00
Yassin Kortam
caede1c5a0
fix(aiohttp): keep keep-alive connector config when a session is rebuilt (#34962) 2026-07-28 16:18:34 -07:00
Yassin Kortam
86ba228d92
feat(prometheus): add service_tier label to latency and spend metrics (#34966) 2026-07-28 16:18:22 -07:00
mateo-berri
00e8691064 ci: enforce format assertions so calendar-impossible deprecation dates fail validation 2026-07-28 16:11:22 -07:00
mateo-berri
3d01c39d00 ci: tighten deprecation_date pattern to reject impossible months and days 2026-07-28 15:24:24 -07:00
Yucheng Zhu
36faa5de5a refactor(otel/v2): drop per-key destination assignment; keys inherit from team and org
Destinations now bind at tenancy granularity only. The per-key logging_exporters
surface was half-shipped (edit-form picker but no create flow) and unrequested,
so it goes: the column leaves the key tables and the migration, /key/generate,
/key/update, and /key/regenerate stop accepting the field, the resolver's union
reads team and org columns only, and the key pages drop the picker and exporter
badges. A key's traces route by its team and org, which the live check confirms:
a brand-new team key exports to the team's destinations with no assignment.

Subset targeting below a whole scope remains available at team granularity
(a multi-team scope with enable-for-entire-scope off, named on specific teams).
Re-adding key granularity later is a purely additive column and field.
2026-07-28 14:41:13 -07:00
Yucheng Zhu
69ba1840d5 chore(ui): drop the via-scope badge suffix and clarify the auto-enable toggle
Viewers only need to know which destinations receive their traces; whether the
binding came from an assignment or the destination's scope is admin plumbing, so
the (via scope) suffix goes. The add-destination toggle is relabeled Enable for
entire scope with a tooltip spelling out on (whole scope exports automatically)
versus off (only explicitly assigned keys, teams, or orgs export)
2026-07-28 13:50:08 -07:00
Yucheng Zhu
670396ee95 feat(otel/v2): disclose resolved trace destinations on team and org info
The team and org pages showed different Logging Exporters lists per role: the
(via scope) badges were derived client-side from GET /credentials, which is
proxy-admin only, so non-admin viewers saw only the identity's own assignments.

Add resolved_logging_exporters to the /team/info and /organization/info
responses: the destination names that will receive the identity's traces,
computed server-side with the same selection the request-time resolver uses
(access grants the identity AND auto_enable or named). Names only; endpoints,
headers, and the access map stay proxy-admin information. The UI renders the
badges from this field, deleting the client-side credentials derivation, so
every role sees the identical list.
2026-07-28 13:39:08 -07:00
mubashir1osmani
7cd009caf7
fix(proxy): avoid DB outage during planned RDS IAM rotation (#34749)
* fix(proxy): warm rotate Prisma client for IAM refresh

* fix(proxy): drain Prisma operations during IAM rotation

* fix(proxy): bound the drain wait when retiring a replaced prisma engine

A replaced engine waited indefinitely for its drain tracker to empty.
Hung queries self-release via prisma's 30s default HTTP timeout, but a
transaction whose owner is hard-cancelled before commit/rollback leaks
its drain count forever, keeping the retired engine and its DB
connection pool alive indefinitely; at one rotation per 12 minutes such
engines accumulate. Cap the wait at 90 seconds, which exceeds every
legitimate operation bound (30s HTTP timeout, 60s max interactive
transaction timeout in this codebase), then kill the engine anyway.
Work killed at the deadline degrades to the pre-drain behavior and is
retried by the existing reconnect/backoff layers.

---------

Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-07-28 13:27:50 -07:00
mateo-berri
d409fec6de
fix(proxy): keep team model aliases while a surviving replica serves the deleted name
Scrub aliases on delete only when the deleted deployment's model_name no
longer resolves in the router. A legacy load-balanced team model can have
several deployment rows sharing one internal name; deleting one replica
must not remove aliases that still route to the survivors, in any team
2026-07-28 20:07:30 +00:00
mateo-berri
5e1d9705db
fix(proxy): skip team model aliases that point at deleted deployments
A team's model_aliases can map a public name like gpt-4 to the internal
routing key (model_name_{team_id}_{uuid}) of a team deployment that has
since been deleted, e.g. after replacing per-team duplicates with one
gateway-level model. The pre-call rewrite then sent every request to a
name the router cannot serve, failing with "no healthy deployments for
model_name_..." even though the requested name still resolves at the
gateway level. The rewrite is now skipped when the alias target has no
live deployment in the router

delete_model also skipped the team alias scan for internal-shaped names
on the assumption they can never be alias values, which is exactly the
shape legacy team model aliases have, so deleting a legacy team model
left the stale alias behind. The scan now always runs, and a public
name that still resolves to a live router deployment (e.g. a shared
gateway-level model group) stays in team.models so the delete does not
revoke the team's access to it
2026-07-28 19:45:00 +00:00
yuneng-jiang
f4a68a75ff
feat(ui): mark Cost Optimization as beta in the left nav (#34984) 2026-07-28 12:04:45 -07:00
Yucheng Zhu
9043a39327 chore(ui): regenerate schema.d.ts for the get_credentials docstring change 2026-07-28 11:54:55 -07:00
ryan-crabbe-berri
51ad1b0a57
test(e2e): skip passthrough headers test until stage can route custom paths to provider creds (#34980) 2026-07-28 11:36:09 -07:00
ryan-crabbe-berri
01ffd1296b
fix(e2e): poll for both spend rows before asserting the cache-hit contract (#34968)
The cache-hit and paid rows for the two driver calls flush from different
pods on independent update_spend timers, so waiting only for the cache-hit
row can return a half-arrived result set where the paid-row assertion then
fails on an empty list. Requiring both row kinds in the poll predicate lets
the existing deadline absorb the slower flush without weakening any assertion
2026-07-28 11:21:36 -07:00
Yucheng Zhu
0b0bcb43ee fix(otel/v2): clear the type-discipline gate and restore admin-viewer read parity
The lint job failed on LIT001: the branch added 36 mutable-collection annotations
and the merge picked up staging's ratcheted budget. Re-annotate the new surface
with read-only views (Mapping/Sequence/frozenset/tuple), build the resolver's
union and destination results functionally, and freeze the span-router grouping
at its boundary; the four genuinely mutable LRU caches and in-place merge targets
carry mutable-ok reasons. Ratchet both budgets down by what the branch now fixes.

Also drop the wholesale credential mask for the admin viewer: it existed for the
removed tenant read, and the viewer principle is read parity with the proxy
admin. Both admin-tier readers now go through _get_masked_values exactly as on
staging, and otel_headers joins the masker's sensitive keys so the collector
auth it carries is masked for every reader.
2026-07-28 11:12:00 -07:00
Yucheng Zhu
53d745a4b6 chore(otel/v2): sweep remaining tenant-surface leftovers after admin-only refactor
Full-diff audit against the admin-only design surfaced leftovers in three layers.

Backend: drop dead code the refactor stranded (is_admin_gated_credential_info,
is_destination_visible, the CredentialInfo decider-era fields, the unused
LLMCallEvent.dynamic_params carrier, the _is_user_org_admin_for_org_id extraction)
and correct every comment/docstring still describing the removed team-admin
self-service or scoped-read designs.

UI: gate the Logging Exporters form rows behind a shared proxy-admin-only
LoggingExportersFormItem so non-admin forms no longer render an orphaned label;
skip the GET /credentials fetch for roles it would 403 (team/org/key views); make
the callbacks table read-only for the admin viewer (no Add/Edit/Delete actions
that would 401); reword access tooltips to routing-scope semantics; regenerate
schema.d.ts from the corrected endpoint docstrings.

Tests: flip the credential-migration non-admin expectation to the route-gate 401,
port the visibility tests to access_grants, drop tests of the removed helpers,
and update copy assertions and stale rationale to the admin-only contract.
2026-07-28 10:35:44 -07:00
ryan-crabbe-berri
b930e2fc2b
fix(gateway): route /a2a through the gateway component (#34958)
* fix(gateway): route /a2a through the gateway component

A2A message-send runs the completion bridge, an outbound LLM call, but the
ingress only listed /v1/a2a so the serving routes at /a2a/{agent_id} fell to
the backend catch-all. Backend pods hold no provider credentials, so every
invocation died with a missing-provider-key auth error while the same call
succeeds on the gateway fleet. Adds /a2a to the ingress gateway prefixes and
the gateway route allowlist, plus a parity test so an ingress prefix that the
gateway trims can never reappear

* revert(test): drop the allowlist parity tests

---------

Co-authored-by: yuneng-jiang <yuneng@berri.ai>
2026-07-28 10:22:49 -07:00
tin-berri
d91fd084f7
Fix cache leakage card layout to keep date picker on right (#34885)
* Fix cache leakage card layout to keep date picker on right and prevent content overlap

Removes flex-wrap and mt-3 to ensure date picker stays pinned to the right side of the card header regardless of zoom level, preventing it from covering card content below

* Remove overflow-hidden from Card to allow dropdowns and overlays to display fully

Fixes date picker dropdown being clipped when opened in cards like the Cache Leakage Card. By removing overflow-hidden from the Card container, popovers, dropdowns, and other overflow content can now display properly without being clipped by the card boundaries.

* Make cache leakage card descriptions consistent with line clamping

Adds line-clamp-2 to ensure both 'by model' and 'by virtual key' cards maintain consistent height. Removes conditional anthropic-specific text that caused height variations between dimensions.
2026-07-28 10:12:00 -07:00
Yucheng Zhu
e5760dc265 fix(otel/v2): drop dead team lookups flagged in review
The /key/update path fetched the key's team into _key_team only to feed the
team-admin/org-admin assignment flags removed in the admin-only refactor; it is
now unused, so drop the wasted DB round-trip. The resolver's org lookup passed
check_db_only=True, forcing a synchronous DB query on every team-key request even
though the team is already resident in the cache from auth; drop it so the lookup
is cache-first. Also correct two comments that still claimed team or org admins
may assign logging_exporters.
2026-07-28 10:06:47 -07:00
Yucheng Zhu
04e4effba3 refactor(otel/v2): scope admin-owned trace destinations to proxy admin
Admin-owned OTEL v2 logging destinations and their access scoping are now managed
only by the proxy admin. Trace routing to identity-scoped destinations is unchanged
because it runs server-side in the resolver; this removes only the tenant-facing
read/write surface the earlier revision exposed.

GET/POST/PATCH/DELETE /credentials are proxy-admin only again (a proxy-admin-viewer
may read); the two credential routes leave self_managed_routes, and the non-admin
scoped list, the team-admin PATCH self-service grant, and the access_decision decider
are deleted. Assigning logging_exporters on a key, team, or org is proxy-admin only,
dropping the team-admin and org-admin widening. In the UI the logging destinations
table and the exporter picker render only for a proxy admin, so non-admins no longer
call GET /credentials.
2026-07-27 19:41:03 -07:00
ryan-crabbe-berri
daf22ec871
test(e2e): make MCP and prometheus e2e tests robust to data-plane sync lag (#34854)
* test(e2e): harden harness and tests against data-plane pod churn

A stage autoscaler scale-down produced a 2s window of ALB 502s that killed six
budget tests on their first management call, and a freshly scaled-up pod that
had not run its 30s DB object sync yet failed two MCP tests and one prometheus
cardinality test. Retry transient gateway errors (502/503/504, connection
errors) once at the shared e2e_http dispatch seam, poll MCP server registration
to the poll deadline instead of asserting a single-shot listing, anchor the MCP
guardrail full-sync wait to the later of the guardrail and server writes, and
turn the prometheus alias poll into a drive-and-scrape convergence loop that
re-sends traffic for missing aliases and unions results across scrapes

* test(e2e): drain request body in retry stub handler so keep-alive reuse cannot misparse leftovers as requests

* revert(e2e): drop the transient-502 retry seam

A raw 502 during a pod scale-down is what a real client sees, so the suite
retrying past it hides an availability gap instead of flagging it. The
gateway-side fix is graceful drain on the deployment; until then the failures
are signal

* test(e2e): cap per-alias driver re-drives in the prometheus cardinality poll

Bounds worst-case provider spend to 4 completions per alias while scrapes keep
polling to the deadline; counters persist on whichever pod served them, so the
cap costs no convergence unless that pod dies

* test(e2e): drop driver re-drives from the prometheus cardinality poll

The per-key cardinality contract is process-local and counters persist on
whichever pod served the driver call, so unioning aliases across free scrape
polls converges without re-sending billable traffic. The residual gap, a pod
dying inside the poll window, is deferred to direct per-pod scraping
2026-07-27 19:22:52 -07:00
mubashir1osmani
328e41b1f9
test(e2e): unblock the ui suite, fix the mcp registration race, park two known product bugs (#34853)
* test(e2e): let the ui suite run from a read-only cwd

The playwright suite never executed on stage. It died in globalSetup before a
single test ran, and the reported error was a red herring.

/app/e2e/ui is a read-only filesystem in the packaged e2e image (the image
runner already redirects playwright's own artifacts to TMPDIR for this reason),
but the suite wrote three things relative to cwd: the per-role storageState
files, the failure-screenshot directory, and the html report. Reproduced in the
pod: storageState raises EROFS, mkdir test-results raises ENOENT.

Worse, the catch block that exists to capture a screenshot threw its own ENOENT
while handling a failure, so the real login error was replaced by a filesystem
error. That is why the run looked like a missing directory rather than whatever
actually went wrong.

Route every artifact through ARTIFACT_DIR (E2E_UI_ARTIFACT_DIR, default "." to
keep run_e2e.sh behavior unchanged), make the diagnostic screenshot best-effort
so it can never mask the underlying failure, and point playwright's reporter and
outputDir at the same place so a bare `npx playwright test` works there too.

fixtures/users.ts had its own copy of the five storageState filenames; it now
re-exports the ones from constants so the paths have a single definition.

Verified in the read-only pod: both writes fail before, both succeed after.
85 tests enumerate and tsc --noEmit is clean.

Refs LIT-4821

* fix(e2e): create the ui artifact root before writing into it

storageState() does not create missing parents, and nothing created ARTIFACT_DIR
itself. Pointing E2E_UI_ARTIFACT_DIR at a writable path that did not exist yet
therefore failed with ENOENT on the very first role's snapshot, before any UI
test ran; the same class of failure the artifact-dir change was meant to remove,
just moved one level up.

Reproduced: writing admin.storageState.json into a missing directory raises
ENOENT. My earlier pod verification masked this because the probe called
mkdirSync itself, which the real code path never did.

mkdir the root once at the top of globalSetup, before the login loop. recursive
makes it idempotent, handles nested paths, and keeps the default "." a no-op.
Playwright creates its own outputDir lazily, so globalSetup is the only place
that needs this, and migration.serverRootPath.globalSetup delegates here so it is
covered too.

* test(e2e): skip the mid-conversation cache checks pending LIT-4873

A mid-conversation role="system" reminder invalidates the prompt cache on the
vertex_ai, azure_ai and bedrock_invoke Messages paths. Measured on the reminder
turn, same conversation shape throughout:

  direct to api.anthropic.com            7013 read  cache preserved
  litellm -> anthropic/claude-opus-4-8   7013 read  cache preserved
  litellm -> vertex_ai/claude-opus-4-8      0 read  cache destroyed

and the Vertex control with the same added assistant/user turns but no reminder
reads 7013, so it is the reminder on the non-first-party paths and not the extra
turns. Anthropic keeping the cache rules out provider behavior; litellm's
first-party anthropic path keeping it rules out the shared Messages transform.

That makes these assertions correct and the failure a real billing bug, so the
tests are skipped rather than weakened; the bodies stay intact and must be
restored unchanged with the fix. Registry rows are left in place, so the three
mid_conversation_system.nonstream.cache_hit cells report as uncovered gaps.

Skips are decorators rather than a pytest.skip() inside the shared helper: a
mid-function skip fires only after setup has already registered a real
deployment via /model/new and left the rest of the body unreachable.

Only Vertex was measured end to end. Azure Foundry and Bedrock Invoke are
inferred from matching nightly failures and should be confirmed with the fix.

Refs LIT-4821, LIT-4873
2026-07-27 19:20:57 -07:00
yuneng-jiang
3c0b1db633
test(e2e): realign Admin UI specs with the MCP dialog and keyless landing (#34870)
Both specs assert against UI that has since moved, so they fail on selectors
rather than on behavior.

The MCP discovery modal became a shadcn/Base UI dialog when mcp-servers
migrated off antd, so `.ant-modal` no longer matches it; locate it by its
dialog role instead. The create form below it is still an antd Modal and keeps
its existing locator.

The no-team internal user has no keys, and a keyless non-admin is now sent to
/ui/connect on the post-login landing, which has no sidebar. Wait for that
redirect to settle, then navigate to the keys page explicitly; the redirect is
gated on the ?login=success marker that the fresh navigation drops, so the
dashboard sticks and the rest of the test is unchanged.
2026-07-27 18:11:18 -07:00
yuneng-jiang
5c95017bc1
test: unstale the reasoning-effort grid count and the responses bridge test (#34868)
* test(reasoning-effort-grid): bump cell-count assertion for claude-opus-5

The claude-opus-5 grid entry added in ae81625ee6 raised the Anthropic direct
route to 31 model combos, but test_grid_cell_count still expected 30, so the
suite went red on the tripwire rather than on any behavior change.

* test(openai): swap the retired deep-research model out of the bridge test

OpenAI shut down o3-deep-research and o4-mini-deep-research on 2026-07-23, so
the live call in this test now comes back as a 400 'Model not found'. The test
was never about deep research specifically; the bridge fires on any model whose
cost-map mode is "responses", so it now uses gpt-5.5-pro, the newest
responses-only OpenAI model, and is renamed to say that.

gpt-5.5-pro was confirmed present on the CI account with an authenticated
GET /v1/models before being picked.
2026-07-27 18:07:04 -07:00
Yucheng Zhu
b0484bae70 chore(otel/v2): drop bundled langtrace v1 endpoint fix
The langtrace v1 exporter host/header fix (https://app.langtrace.ai/api/trace +
x-api-key) and the shared endpoint-normalizer /api/trace carve-out are unrelated
to admin-owned OTEL v2 destinations; they change behavior on the legacy langtrace
v1 path. Moved to its own PR (#34865) so this PR only touches OTEL v2. The v2
langtrace preset and its normalizer in otel/plumbing/providers.py are unaffected.
2026-07-27 17:37:34 -07:00
yuneng-jiang
f2cda740f7
chore: update Next.js build artifacts (2026-07-28 00:06 UTC, node v20.20.2) (#34859) 2026-07-27 17:20:51 -07:00
mateo-berri
c2f0014a63 fix(jwt_auth): grant only /v1/messages routes to JWT teams by default, not all anthropic_routes 2026-07-27 17:15:09 -07:00
Yucheng Zhu
1042b56d2f merge: resolve conflict with litellm_internal_staging (ruff budget ceilings) 2026-07-27 16:51:24 -07:00