Commit graph

775 commits

Author SHA1 Message Date
devin-ai-integration[bot]
797353f13a
fix(otel): name postgres service spans by operation and table (#44240)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
michelligabriele
119179942e
fix(guardrails): run unified guardrails on /v1/images/edits (#44195) 2026-10-02 21:14:35 -07:00
devin-ai-integration[bot]
4b9f9903f3
refactor(ui): compose dashboard pages with shared layouts (#44306)
* refactor(ui): compose logs tabs directly in the route

* feat(ui): share composable dashboard page layouts

* refactor(ui): compose page header and logs toolbar from parts

PageHeader drops its icon/title/subtitle/primaryAction/tabs/utilities props and the
leadingControls render prop in favor of PageHeaderTitle, PageHeaderDescription and
PageHeaderControls that each wrap one element and forward native props.

LogsTableToolbar's 15 props collapse into one LogsTimeRange value plus composable
LogsToolbar, LogsTimeRangePicker and LogsToolbarSwitch parts assembled in the panel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): express DataTable layout classes as cva variants

Replaces the hand-rolled class-pair constants with boolean cva variants,
which also brings DataTable back under the complexity budget.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 01:35:34 +00:00
yuneng-jiang
a292fd409f
test: fix three order-dependent and timing-flaky tests (#44271)
* test(integration): answer the model-info refresh GET in the mixed MCP responses wire peer

The proxy's periodic model-info refresh sends GET /v1/models to the deployment api_base, which tripped the peer's /responses-only assertion when a tick landed mid-test

* test(e2e): wait for the api-keys URL after clicking Virtual Keys in onboarding

/ui already renders the Virtual Keys heading, so the helper returned before navigation finished. The late route change moved focus and closed the account menu popover in hideLiteAdmin

* test(proxy): stop two unit modules leaking app.openapi_schema and a session-wide Router

test_custom_openapi cached a stripped schema on app.openapi_schema and never cleared it, breaking later openapi route tests. test_proxy_reject_logging built a module-level Router that stayed in the live router registry all session and re-added cost-map keys during a reload. Reset the schema via monkeypatch and make the Router a function-scoped fixture
2026-10-02 16:29:57 -07:00
yuneng-jiang
9b5562f89b
test: repair stale and polluting tests red on scheduled main CI (#44229)
* test(proxy): stop the proxy_server app fixture leaking LITELLM_LOG

The session app fixture set LITELLM_LOG=ERROR with os.environ.setdefault and never removed it, so later tests on the same xdist worker inherited it. test_drop_params_env_var spawns a subprocess with os.environ and lost the warning it asserts on. Scope the variable to the import with a MonkeyPatch context

* test(secret-detection): give the hand-built redaction request an ASGI path

Since #43975 _read_request_body checks the route path via request.scope, and a scope without path raised KeyError that was swallowed into an empty body, so chat_completion failed with a missing messages parameter. Real ASGI scopes always carry path

* test(integration): isolate litellm callback lists per sdk test

usage-based-routing-v2 Routers register their selector in litellm.callbacks and nothing removes it, not even Router.reset(). The counter TTL and Redis service metrics tests left their selectors behind, and the next usage routing test ran their pre-call checks against its own rpm=1 deployments, raising "Deployment over defined rpm limit". An autouse fixture now gives each sdk test copies of the callback lists and restores the originals afterwards

* test(integration): keep the owner-lookup fault proxy off the shared read replica

The owned proxy points DATABASE_URL at a scratch database but inherited
DATABASE_URL_READ_REPLICA from the replica job, so auth read the shared
database and rejected the freshly created key with token_not_found_in_db.
Drop the replica variable like the other scratch-database owned proxies

* test(integration): request every seeded key in the team owner breakdown

The aggregated team activity endpoint now caps breakdown.api_keys at the top
100 keys by default (#43398), so the 300 seeded keys came back as 100 rows.
The test guarantees each key is reported with its own owner, so ask for an
api_key_limit that covers all seeded keys

* test(integration): give every owned Redis its own port in the redis-cache container

On CircleCI every owned Redis ran on the fixed port 16379 inside the shared
redis-cache container. When an earlier server still held that port, the new
one failed to bind, readiness pinged the old server, the pidfile read failed
and cleanup then reported "Owned Redis still serves after shutdown"

Reserve an ephemeral port for the docker-exec path the same way the local
binary path already does, and refuse to start when something already serves
the chosen port so the failure names the real cause

* test(e2e): skip the Vertex Mistral partner case the e2e project cannot reach

The e2e Vertex project gets a 404 publisher model not found for vertex_ai/mistral-small-2503, so the case can only fail

* test(e2e): skip the Vertex gpt-oss partner case the e2e project never serves

vertex_ai/openai/gpt-oss-120b-maas has hit a 60s read timeout with no response headers on every run in the e2e Vertex project since the case was ported, and no other Vertex partner chat model passes there to switch to

* test(e2e): check only stored message content for a leaked card number

The Presidio spend-log check ran the card-number pattern over the whole serialized response, so a Luhn-valid usage.cost float (0.0003466000000000001) failed the streaming /v1/messages case although the stored content was <CREDIT_CARD>. The check now reads the content and text strings of the stored response, which is where a raw card would land, and still requires the placeholder there

* test(e2e): assert the proxy decodes token-array embeddings for titan

The port in #44120 carried over a legacy SDK-direct test that expected Bedrock to reject token ids with a 400. Through the proxy, /embeddings decodes token arrays to text for providers that cannot embed tokens, so titan answers 200. The test now sends a token array and its decoded sentence and requires the two vectors to match, which fails if the proxy stops decoding or decodes with the wrong tokenizer

* test(e2e): run the Bedrock extended-thinking round trip on a model that honors enabled thinking

us.anthropic.claude-sonnet-5-5 is adaptive-only, so litellm sends thinking.type=enabled with a 1024 budget as adaptive with low effort, and Bedrock returned no reasoning blocks on 5 of 5 identical Converse calls (boto3 direct agreed). us.anthropic.claude-sonnet-4-6 accepts the legacy shape verbatim and returned reasoning on 5 of 5. The non-thinking Bedrock case stays on sonnet-5-5

* test(proxy): stop unit modules forcing DEBUG logging into the event-loop lag tests

Five tests/unit modules set verbose_proxy_logger to DEBUG at import, so every xdist worker that collected them logged the 2.4MB pass-through response from a worker thread, and secret redaction of that line held the GIL for ~0.8s+ inside the timed window. The lag tests now pin the LiteLLM loggers to WARNING and freeze gc while timing, and the module-level DEBUG overrides are removed

* test(e2e): cite the tokenizer and date behind the titan token-array fixture

* test(e2e): let migration seed replicas finish their request-log indexes before cloning

Since #43948 a serving proxy builds the two LiteLLM_SpendLogs indexes on a background thread after it reports ready. The seed fixtures stopped the replica at readiness, so every cloned legacy database lacked an index no real deployment would be missing, and the v2 baseline diff refused it. Seeds now wait until both indexes exist and are valid in the database's schema

* test(passthrough): give the pass-through MockRequest an httpx URL and ASGI scope

#43626 made get_request_route read request.scope during pass-through kwarg setup; the MockRequest in tests/unit/passthrough had neither a scope nor a URL object, so both stream-param tests raised before reaching the code they check. Mirrors the repair #43626 made to the tests/pass_through_unit_tests fake

* test(integration): ignore foreign allow_all_keys MCP servers in the access matrix tool list

test_toolset_gateway_url_serves_a_team_granted_toolset_to_a_key_without_its_own_grant (#43908) registers an allow_all_keys server on the shared gateway, and allow_all_keys servers are listed to every key by design, so a matrix case running on another xdist worker at the same time saw its tools. The matrix now drops tools of allow_all_keys servers it did not create, read from LiteLLM_MCPServerTable before and after listing, and still compares everything else exactly
2026-10-02 21:32:17 +00:00
devin-ai-integration[bot]
0238ec9721
fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key (#43978)
* fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key

A path-less add/replace op (RFC 7644 3.5.2, what Okta Push Groups sends on a
rename) carries a partial Group resource. Each of its attributes now applies as
if sent with that path, so displayName updates the team alias and externalId
and members get their usual handling, and the pushed attributes merge into the
scim_data snapshot the PUT path already writes. A path-less remove or a
path-less op without an object value is rejected with a 400. Any group PATCH
drops an empty metadata key an earlier push left behind, and the Admin UI
metadata form skips an empty key so an affected team can save its settings.

* fix(scim): let a later path op win over an earlier path-less value in the group snapshot

* fix(scim): type the stored team metadata before the JSON object check

* test(scim): run the real group transformation in the path-less replace test

* test(scim): assert the renamed group comes back from the path-less replace

* test(scim): audit the path-less group PATCH on the live proxy

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 14:20:04 -07:00
devin-ai-integration[bot]
b21e44cbf9
feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key (#42375)
* test(e2e): jwt auto_register map-existing-key repro

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): exclude blocked keys from auto_register_map_existing_key reuse

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(jwt): route existing-key lookup through VerificationTokenRepository

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): stop requiring LITELLM_SALT_KEY for the owned JWT gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): gate the owned JWT gateway tests behind E2E_OWNED_GATEWAY

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): only reuse keys that can call LLM routes in auto_register_map_existing_key

Skip Admin UI session keys and keys whose allowed_routes restrict them to
anything other than llm_api_routes (management, read_only, password-reset
sessions). Mapping a JWT to one of those left the user with 401s or 403s on
every LLM call, since the mapping persists.

* fix(jwt): scope auto_register_map_existing_key reuse to the JWT-resolved team

Only reuse a key whose team_id matches the team auth_builder resolved for
the JWT (no team matches no team), so a personal key can no longer bypass
the resolved team's model and budget limits.

With the flag on, the first JWT request now falls through to the same
virtual-key checks later mapped requests get, instead of returning early,
so a reused key's own limits apply from request one rather than 200 then
403. Flag off keeps the early return unchanged.

* fix(jwt): keep the early return when no master key is set

Without a master key the generic virtual-key path returns a bare
INTERNAL_USER object, so falling through on the first auto-registered
request dropped the key's team, models and budgets. Only fall through when
a master key is configured.

Tests now assert the reused key per team rather than the query shape, and
cover the flag-off early return and the no-master-key case.

* test(jwt): assert on race-loser's returned key, not only mocks (TQ002)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(jwt): close the auto_register_map_existing_key race, shared-claim and expiry holes

A key auto_register just minted is never adopted by a concurrent request, so the race loser's cleanup can no longer delete a key another request mapped and cascade its mapping away (503, user left with no key)

Reuse only happens when the claim value is the JWT-resolved user_id. A shared claim such as azp or client_id falls back to minting, so one user can no longer land on another user's personal key and budget

Only keys that never expire are reused, so an expiring key can no longer pin the claim to a permanent 401

Integration tests on a real proxy and Postgres cover all three. The race test holds the first mapping insert in a Postgres relay, so the interleaving is forced rather than timed. The where-clause shape unit tests are replaced by these, since only a real database proves the filter

* test(e2e): create the reused key in the team the JWT resolves to

The flag only reuses a key in the JWT-resolved team, and this identity's groups claim resolves to its team, so a teamless key was never eligible and the test could not pass

* test(integration): match the held statement across TCP reads

The relay looked for the trigger inside one read, so an insert split across two reads was never held and the race test would fail waiting for it. It now matches one exact trigger over a window that keeps the end of the previous read

* fix(jwt): gate key reuse on the claim field, not on the claim value

Requiring the claim value to equal the resolved user_id skipped reuse for users matched through the sso_user_id or case-insensitive email fallback, whose stored user_id differs from the JWT sub. That is the lookup LIT-5378 asks for. Reuse is now allowed when the virtual key claim is the user_id or user_email JWT field, globally or for the token's issuer, which still keeps shared claims such as azp or client_id on the mint path

* fix(jwt): let an issuer's own user field replace the global one when gating key reuse

An issuer that identifies users by uid no longer treats the global sub field as a user identity claim, so a shared sub under that issuer mints instead of reusing a personal key

* test(jwt): make the flag-off test fail when the flag no longer gates key reuse

The flag-off test used a config where sub was not a user identity claim, so deleting the flag check still passed. Configure user_id_jwt_field=sub so only the flag keeps the lookup off, and drop test docstrings

* chore(lint): drop mutable-ok suppressions that LIT013 flags as no-ops

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Mrinal Chanshetty <mchanshetty@Mrinals-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 12:11:33 -07:00
yuneng-jiang
7d50a31eb5
test(e2e): move live-provider legacy tests into tests/e2e (#44120)
* test(e2e): move live-provider legacy tests into tests/e2e

Port legacy tests that exercise real providers into the tests/e2e suites that own them, using the harness (/model/new plus deferred cleanup) and asserting on what the caller receives. Delete legacy tests already covered at equal or stronger strength by e2e, integration or unit tests, and drop the now empty ocr_testing CircleCI job

* test(e2e): address review on the live-provider test move

Assert the SSE error frame a client actually receives when a post_call guardrail blocks a stream, and require a tool call for every requested city before checking the answer. Restore the OCR matrix and its CircleCI job, the Claude Agent SDK streaming test, and test_async_create_batch, since their SDK-level and callback assertions have no equivalent in tests/e2e

* test(e2e): accept both guardrail block shapes on a blocked stream

A post_call block before the first chunk reaches the client as HTTP 400 with either a JSON error body or a single SSE error frame, depending on whether the block surfaced as an exception or an error chunk. Assert the policy message is present and the blocked output is absent in both

* test(realtime): restore direct SDK realtime tests against OpenAI

The e2e realtime tests go through the proxy and the remaining SDK tests either mock the upstream or assert less, so keep the direct litellm._arealtime tests with and without intent, and TestOpenAIRealtime::test_realtime_connection, in place

* test: make realtime and Nova stream checks deterministic

The direct SDK realtime tests now fail on a refused connection instead of skipping. The with-intent test asserts OpenAI rejects the exact intent value sent, which only happens when the intent is forwarded. The Nova /v1/messages stream test asserts stream structure, stop reason and usage instead of model wording

* test(realtime): own intent forwarding with a unit test instead of a live rejection

Assert litellm._arealtime passes the intent query param into the OpenAI realtime websocket URL, which is the behavior LiteLLM owns, and drop the live test that depended on OpenAI's rejection wording
2026-10-02 00:02:18 -07:00
yuneng-jiang
4b1d9bf148
test(e2e): keep 1ms-timeout deployments off the provider cache (#44082)
The timeout reliability tests rely on a 1ms deadline the real backend always
misses. With E2E_PROVIDER_CACHE on, the deployment pointed at the cache edge,
and its healthy sibling in the same test had already recorded a response for
the same canonical request, so the edge answered from Redis inside the 1ms read
window. Build 342 of litellm-e2e saw test_timeout_trips_cooldown_then_recovers
get a 200 from the timing-out deployment itself, with a recording made about
12 hours earlier. Both timeout helpers now register on the live provider path,
which PROVIDER_CACHE.md reserves for tests that need real provider timing
2026-10-01 15:54:02 -07:00
devin-ai-integration[bot]
008fcb4fe3
feat(tool-policies): show the user who owns the key that discovered a tool (#43892)
* feat(tool-policies): show the user who owns the key that discovered a tool

GET /v1/tool/list and GET /v1/tool/{tool_name} resolve the discovering key's owner from the verification token and user tables at response time and return it as a nullable user field. The Tool Policies page adds a User column that shows alias, then email, then ID, with the same cell the Virtual Keys page uses. Keys without an owner, deleted owners, and rows without a key hash show no user, and a database failure in the owner lookup keeps the tools listed with user null

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tool-policies): bound the owner lookup with chunked membership queries

The key-by-token and user-by-id lookups behind the tool rows' user field
put every distinct key hash into one IN list. BaseRepository gains
find_many_in, which runs the repository's chunked membership query and
converts the rows like find_many does, and the owner lookup uses it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(tool-policies): cover the owner column across the tool routes and the dashboard

Integration cells for the direct, detail and filtered tool routes, owners without alias or email, deleted owners and keys, keyless and unknown-key historical rows, more keys than one membership chunk, repeated reads, two-worker reads during discovery and a failed owner lookup. A Playwright cell drives the bundled Tool Policies page against the live proxy and follows the owner link

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 20:34:57 +00:00
yuneng-jiang
0da00d4b2e
test(e2e): bill Sail windows that synchronous calls can still use (#44058)
Sail now rejects completion_window "flex" on synchronous requests with a
400 saying flex is only for background responses or Batch work. The chat
flex case and the responses flex case have failed on every scheduled
litellm-e2e run in builds 337, 340 and 341. The chat cases keep balanced and
auto, and the responses case sends a caller metadata.completion_window of
balanced, so both still prove the window reaches Sail and the bill uses
that window's distinct rates
2026-10-01 12:54:26 -07:00
devin-ai-integration[bot]
2cfa5ec126
test(proxy): delete the legacy proxy test tree and shard tests/unit/proxy by glob (#44018)
* test(proxy): delete the legacy proxy test tree and serve the redirect test from loopback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): exercise the shard check directly for unit_selection-owned children

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): serve the redirect test from respx instead of a socket

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): credit shard ownership only to unit flags wired in gha

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): implement the wired-flag shard crediting the tests assert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): split the root proxy test files into their own unit shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): point the rate-limit skip reason at the usage-based-routing-v2 RPM tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 11:40:45 -07:00
devin-ai-integration[bot]
a76b59db9f
test(proxy): move middleware, spend_tracking, pass_through, common_utils and root proxy tests into tests/unit/proxy (#44015)
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:23:31 +00:00
devin-ai-integration[bot]
bfd3f39dca
feat(s3_v2): add s3_partition_granularity option for hourly S3 folders (#43748)
* feat(s3_v2): add s3_partition_granularity option for hourly S3 folders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover s3 v2 partition granularity across surfaces, settings and chaos

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover previous_response_id history rebuilt from an hourly cold storage object

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): reuse the cold storage key only when s3_v2 owns cold storage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): cover hour rollover, postgres outage, in-flight switches, key/team vars and real S3 layout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(liccheck): authorize libfaketime, the GPLv2 dev-only clock the s3 rollover integration test preloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): wait for the rejected-request cell's payloads by id, not by line count

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): declare the postgres outage cell's models in config and trip the relay on burst ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): drop the libfaketime hour rollover cell and its dev dependency

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): deselect the s3_v2 live e2e on the stage-mirror stack

The stage-mirror config enables no s3_v2 callback, so every test in test_s3_log_e2e.py fails its readiness check there. The file keeps running in the Buildkite e2e lane, which configures s3_v2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): declare the sink outage burst models in config

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(s3_v2): read cold storage metadata without an empty dict default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for the proxy to reconnect before the postgres outage recovery request

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-01 11:00:21 -07:00
yuneng-jiang
6ca90b927c
test(ci): repair stale tests and flaky CI infrastructure (#43983)
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden

#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key

* test(ui): give the auto-router threshold save wait room for the availability debounce

#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed

* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default

#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against

* test(e2e-ui): wait for the call-id search before hovering the logs row

The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time

* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off

The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support

* test(integration): read the agent 365 guardrail status by its own name in spend logs

The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict

* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check

gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler

* test(e2e-ui): fill the create-tag fields inside the dialog

#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step

* test(integration): run integration proxies with the CI license

Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments

* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5

Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving

codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI

* test(unit): join the session-minting thread before collecting the handler

asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference

* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early

owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.

uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds

* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build

The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there

* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections

The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry

* ci: move the unit-test uv cache split into a composite action

check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes

* test(unit): point tiktoken at the bundled cache for every unit test

The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.

* test(integration): answer model discovery probes in the hosted_vllm wire tests

The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.
2026-10-01 17:46:43 +00:00
ryan-crabbe-berri
ae05f7d2c1
test(e2e): typed per-test metadata for the e2e suite (#42044)
* feat(e2e): give e2e tests typed metadata for what they drive

@meta(Subject(domain, route, providers, models, capabilities, mode)) declares
what a test is about with closed enums, and each field lands in the JUnit report
as a property. The quota_management suites are the first to declare it.

* docs(e2e): say e2e_metadata avoids litellm, not that it is stdlib-only

It already imports pydantic and pytest, both of which the suite needs to collect. The rule that matters is no litellm import

* test(e2e): declare models through the constant each test drives

43 @meta declarations in quota_management typed the model name out again, so changing the call would leave the coverage report naming the old model. Each file now has one constant used by both, and a guard fails on any model written as a string literal in @meta

* refactor(e2e): set route only when the endpoint is what the test checks

A budget or rate-limit test whose chat call only triggers the block now leaves route unset, since its steps already name the call. Tests of an endpoint keep it: budget CRUD, key creation, spend reporting reads, and the per-endpoint spend tests for chat, messages, embeddings, batches and health. The two /spend/logs tests tagged chat_completions are now spend_reporting

* refactor(e2e): build the declared properties without mutating a list

subject_properties seeded a list and grew it with append and extend. It now flattens one tuple per field, and the plural-name table is a read-only mapping

* fix(e2e): tag each spend-route probe with the endpoint it checks

The breadth test gave all 33 probes spend_reporting, so /key/list, /user/list, /team/list, /organization/list and /customer/list counted as spend reporting. Each case now carries its own route, with organization and customer management added to Route
2026-09-30 21:03:21 -07:00
ryan-crabbe-berri
424bfd8758
feat(e2e): record each e2e test's steps, starting with ProxyClient (#42393)
* feat(e2e): record each e2e test's steps, starting with ProxyClient

@step on a harness method records a plain-English line for every call, in
order, as repeated JUnit step properties. Labels are templates filled from the
call's parameters, like "Generate a virtual key with models: claude-haiku-4-5
and rpm limit: 3", and secret request fields are marked Field(repr=False) so
they never print. ProxyClient and the rate-limit QuotaClient carry steps first;
the other harnesses follow one area at a time. The recorder and JUnit tests run
in the Code Quality workflow's test_e2e_metadata step.

* docs(e2e): rewrite the recorded test steps guide in plain language

* fix(e2e): keep logging callback credentials out of recorded steps

* fix(e2e): mask the run's credentials in every recorded step

* fix(e2e): attach steps before the oauth failure snapshot

The failed setup or call report of an mcp_oauth_live test copied user_properties before the steps were attached, so it carried no steps. Every setup and call report now takes its properties after the steps attach

* fix(e2e): name the saved credential in its recorded step

The create_credential label read credential_info, which defaults to {} and is never set by the live callers, so the step printed nothing after 'for'. It now reads the required credential_name, and a guard fails on any label that reads a field with a default
2026-09-30 19:33:53 -07:00
yuneng-jiang
e61733b170
test(e2e): align completion, SAIL, and spend-log fixtures with supported contracts (#43902) 2026-09-30 14:21:22 -07:00
devin-ai-integration[bot]
f285229b51
fix(proxy): delete large teams without per-member transaction fan-out (#42998)
* fix(proxy): delete large teams without per-member transaction fan-out

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): evict email-only member caches and reset team members metric on delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep new delete-team literals within the LIT002 ceiling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve deleted-team member ids before the locked delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup

`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.

* fix(repositories): slice find_by_emails into bounded IN statements

The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.

* fix(proxy): delete a team once when /team/delete repeats its id

The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.

* test(integration): audit cells for /team/delete on large, legacy and concurrent teams

Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.

Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.

Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.

* test(integration): pin each chaos outage to a live /team/delete

The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 13:49:18 -07:00
devin-ai-integration[bot]
8afabe81f1
fix(ui): surface x-litellm-call-id in Logs search, table and drawer (#42436)
* fix(ui): surface x-litellm-call-id in Logs search, table and drawer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate api types for spend logs search description

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): drop redundant comments from the call id logs helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(e2e): format logs call id helper and spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep one id per Logs row, move x-litellm-call-id to hover and drawer

The Request ID cell shows only request_id again. When the row's litellm_call_id
differs, the cell tooltip lists it as x-litellm-call-id with its own copy button,
and the drawer header labels the second line x-litellm-call-id: instead of the
call id caption. Stacking two ids in every row made the column noisy for the
common case where the viewer only needs the row they searched for.

* test(e2e): cover the Request ID tooltip hover and copy path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: poll the clipboard after the tooltip copy and drop a jsdom aside

The e2e read navigator.clipboard right after the click, so a slow async write
could fail the check even though copy works. The unit test's fireEvent choice
(jsdom has no layout, so a real pointer move off the trigger closes the tooltip
before the click lands) is documented here instead of inline.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 02:16:56 +00:00
yuneng-jiang
d098b02ed9
fix(auth): give UI/CLI session tokens their own AES-GCM context and header-safe shape (#43790)
Some checks are pending
Unit Tests / misc (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* refactor(auth): bind UI/CLI session tokens to their own AES-GCM context

UI and CLI session tokens are now always encrypted with AES-256-GCM and a
fixed session associated-data value, and the session-token check only accepts
AES-GCM values carrying that same value. Stored secrets keep their current
encryption and decrypt unchanged, so nothing needs migrating.

encrypt_value_helper and decrypt_value_helper take an optional aad. XSalsa20
cannot bind associated data, so an AAD-bound value is always written as
AES-256-GCM, and an AAD-bound decrypt refuses the legacy format.

Session tokens issued before the upgrade stop validating, so UI and CLI users
sign in once more after upgrading.

* test(e2e): cover real SSO login through the dashboard and the lite CLI

Adds two specs under tests/e2e/ui/oidc, run by playwright.oidc.config.ts
against a live Keycloak stack. The dashboard spec checks that the SSO
session authorizes the Virtual Keys and Models data requests. The CLI
spec runs a real lite login in an isolated HOME with the keyring
disabled, then lists models and sends one chat completion with the
stored session. The main Playwright config now ignores oidc/.

* fix(auth): encode UI/CLI session tokens as unpadded base64url

Session tokens carried the v2:gcm: storage prefix and base64 padding. Basic-auth parsers split on the first colon and browsers reject ':' and '=' in WebSocket subprotocols, so Langfuse pass-through and the realtime playground could not use them

Tokens are now plain unpadded base64url, the same header-safe shape as any bearer token

* fix(auth): prefix UI/CLI session tokens with litellm_login_

A prefix-less token starts with sk- about once in 262,144 logins and is then routed as a virtual key, so that login gets a 401. The prefix also makes session tokens easy to spot in logs

The prefix doubles as the token's AES-GCM associated data, so the visible kind and the encrypted kind cannot disagree

---------

Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-29 18:42:24 -07:00
devin-ai-integration[bot]
e814532033
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): satisfy type-discipline and strict ruff budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): stamp the served service_tier on every Responses bridge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): cover anthropic and responses served-tier billing paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): bill disconnects through the router's anthropic stream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: apply ruff format to the anthropic stream changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(coverage): ignore delegating properties the ast scan cannot see

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: keep the cast-ok reasons on the cast call line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(streaming): parameterize delegated chunks and messages types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): follow the anthropic pass_through rename after merging main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drain the logging worker between response cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): keep the served service_tier on streamed chunks and bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): type the served service_tier chunk without a loose kwargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill the served service_tier over the requested one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): drop explanatory comment from the tier resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-29 12:54:17 -07:00
yucheng-berri
336c7c0849
test(integration): sweep proxy logs, metrics, a Datadog intake and the Logs drawer for credential canaries (#43306)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): sweep proxy logs, metrics, a gzip Datadog intake and the Logs drawer for credential canaries

* test(e2e): treat an unset prompt-storage setting as unset and restore it

* test(integration): name the Datadog sink slot G1d

* test(e2e): search the Logs page for base64 forms of the deployment key

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): pass the resolved deployment id to the Datadog route sweep

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown
2026-09-29 17:57:28 +00:00
devin-ai-integration[bot]
f4a217d005
fix(ui): rename All Models tab to Deployed Models and model filters to All Proxy Models (#43638)
* fix(ui): rename All Models tab to Deployed Models and view filter to All Proxy Models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): use All Proxy Models label for the public model name filter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): rename ALL_MODELS_VIEW constant to ALL_PROXY_MODELS_VIEW

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 17:59:47 -07:00
yuneng-jiang
303434d573
test(e2e): report batch cleanup leftovers as a plain UserWarning (#43405)
The leftover warning used a class defined in a test-directory module. The xdist controller cannot import it, so an uncaught leftover warning crashed the whole e2e run. Same change as #43391 on rc/1.103.0
2026-09-27 03:26:23 +00:00
devin-ai-integration[bot]
b396b0b724
feat(e2e): read management routes back from the control plane replicas (#43373)
* feat(e2e): read management routes back from the control plane replicas

The suite's management read-backs (/key/info, /team/info and friends)
polled the same replica list as the data plane. On a componentized stack
whose LITELLM_PROXY_REPLICA_URLS names the gateway pods directly, that
list answers those routes 404, since a gateway pod trims the management
routes at startup. A new LITELLM_CONTROL_PLANE_REPLICA_URLS names the
addresses a management read-back polls instead: an exported list wins,
and when it is unset the old rule stands, the data-plane replicas while
the control plane shares the suite's base URL and the control-plane base
alone once it is split.

build_proxy_client takes the list as control_replica_urls and
read_back_everywhere picks its replicas per path, the way the rest of the
client already does.

* fix(e2e): derive the control replicas of a client built for another proxy

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-26 18:05:55 -07:00
devin-ai-integration[bot]
eea1d0f269
fix(responses): stream guardrail pre-call block as SSE with a typed output item (#42507)
* fix(responses): stream guardrail pre-call block as SSE with a typed output item

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): import blocked usage helper from the guardrail utils module

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): drop narrating docstrings and poll without rebinding

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover pre-call guardrail block on /v1/responses stream and json

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for responses guardrail block contract

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): observe upstream on the recorded chat route for responses denial cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): tidy responses denial audit cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for worker count to recover after SIGKILL

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): require a replacement worker after SIGKILL

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the blocked response test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-26 18:05:21 -07:00
devin-ai-integration[bot]
69ad004015
refactor(anthropic): rename experimental_pass_through to pass_through (#43329)
* refactor(anthropic): rename experimental_pass_through to pass_through

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): point compact patch targets at renamed pass_through path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 13:00:50 -07:00
devin-ai-integration[bot]
3a6744cd02
feat(sail): add Sail as a provider with service_tier mapped to its completion window (#42840)
Register Sail (providers.json, LlmProviders.SAIL, OpenAI-compatible lists,
ProviderConfigManager) for chat, Responses and /v1/messages, and add its 12
models to both cost maps with asap, balanced and flex price columns.

Sail picks speed and price with metadata.completion_window and rejects
service_tier, so the Sail chat and Responses configs translate the tier:
default and priority to asap, flex to flex, balanced to balanced, auto to no
window. Billing prices the window that was sent. A tier Sail has no window
for, or a window or tier set where billing cannot see it (request metadata,
extra_body), is a 400 unless drop_params is set.

Add balanced to ServiceTier and its _balanced price columns to the model
info types, the Rust catalog and the dashboard schema. A transform_extra_body
hook on the chat and Responses base configs, which returns extra_body
unchanged by default, lets Sail keep the window when a caller also sends
extra_body.metadata. Sail is listed in the Add Model form and model picker.

Co-authored-by: shrey kharbanda <shrey@berri.ai>
2026-09-26 12:57:48 -07:00
yuneng-jiang
e53e67ede5
test(e2e): assert only litellm-owned batch behavior and move the blank S3 env pin to an integration test (#43321) 2026-09-26 11:17:03 -07:00
yuneng-jiang
14f4c34c61
fix(ci): stop stale CI reds, keep unit tests off the host env, retry CyberArk policy conflicts (#43294)
* fix(ci): stop five stale or flaky CI reds and retry CyberArk policy-load conflicts

The Langfuse redaction unit test exports to a local OTLP capture instead of
polling Langfuse Cloud through a recorded lookup. The passthrough worker-kill
test only requires spend rows for requests the surviving worker served. The
spend-routes sweep treats the intentional /spend/capture_rate 503 as expected.
CyberArk retries a 409 policy load in Python, Rust and the e2e Conjur helper
instead of reading it as "variable exists". The integration egress guard now
matches the script's own cgroup, so it no longer blocks the CircleCI agent,
which runs as the same user.

* fix(ci): keep the policy-load backoff typed as float

* fix(ci): retry CyberArk policy loads without blocking the event loop and tighten the worker-kill and Langfuse tests

* fix(secrets): load CyberArk policy one request at a time per manager

* test(secrets): pin that non-conflict CyberArk policy failures are not retried

* test(unit): run tests/unit with only an allowlisted host environment

CircleCI's unit job inherits every project env var, so real provider keys,
REDIS_HOST, DATABASE_URL and AWS or Azure credentials reached tests that
assume none are set. Locally, litellm's import-time load_dotenv did the same
from any .env up the tree. The unit conftest now drops every variable outside
a small allowlist and disables dotenv before litellm is imported.

* test(e2e): name a failed search and the stuck batch status instead of misattributing them

The websearch session test read an empty web_search_tool_result_error block as a
successful search, so a failing search tool surfaced as a session billing bug.
The batch cancellation timeout now reports the last status the proxy returned.

* fix(ci): scrub the host environment per unit test instead of for the whole pytest process

GHA shards run tests/unit next to other suites in one process, so the import-time
scrub deleted MCP_TEST_PEER_PYTHON before tests/mcp_tests read it and the MCP
upstream fell back to the SDK2 interpreter. The two websearch tests that called
OpenAI and Perplexity live are removed: tests/unit no longer sees their keys.

* fix(ci): scrub only the host variables present before litellm is imported

The per-test scrub also deleted TIKTOKEN_CACHE_DIR, which litellm sets at import to
its bundled encodings, so tokenizer paths tried to download them and hit the
socket guard. The prisma setup test now passes its own database URL instead of
reading one another test leaked into the process environment.

* fix(ci): stop the order-dependent unit reds and settle logging tasks on their own queue

LoggingWorker marked a task done on whichever queue was current when the callback
finished, so a callback that outlived an event-loop change raised "task_done()
called too many times" or undercounted the new loop's queue. It now settles the
queue the task came from.

The rest are test isolation fixes for failures that only appeared when another
file ran first on the same xdist worker: a replaced user_api_key_cache, breaker
metrics unregistered by prometheus tests, semantic_router's health-check filter on
uvicorn.access, logging tasks carried over from bedrock tests, a Router-written
model_cost entry, and a stray post captured by the langflow test. The token
counter check now asserts bounded chunking instead of wall-clock time.

* test(e2e/ui): wait for the logout redirect before visiting a protected page

Logout revokes the session server-side before clearing cookies and navigating, so an immediate page.goto either ran with the cookie still set or was aborted by the logout redirect (net::ERR_ABORTED).

* test(unit): restore the prometheus metrics config per test and settle logs carried from earlier tests in the a2a cost tests

* test(router): pin the router clock in the usage counter tests so a minute rollover cannot empty the read

* test(e2e/ui): wait for logout to clear the token cookie instead of for a login redirect

* test(integration/mcp): answer the model-info probe another test's proxy sends to the model double
2026-09-26 09:25:13 -07:00
devin-ai-integration[bot]
a942c343ab
fix(e2e): resolve the blank-S3 gateway repo root from the litellm package location (#42911)
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 23:40:55 -07:00
devin-ai-integration[bot]
3ebf6a1fd8
test(e2e): accept the otel cost write as a linked root trace (#42931)
* test(e2e): accept the otel cost write as a linked root trace

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): window and fail-closed the linked otel trace read-back

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): poll the otel read-back without recursion

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 22:40:45 -07:00
Refael Iliaguyev
e065a2575b
fix(proxy): send a real error event when a /v1/messages stream fails (#41826)
* fix(proxy): send a real error event when a /v1/messages stream fails

When a stream failed halfway through, the proxy wrote the error as a plain
`data: {"error": ...}` line with no `event:` in front of it. Anthropic clients
pick stream events by that name, so they skip the line and the request looks
like it simply stopped with nothing in it.

Write the failure as an `event: error` frame with Anthropic's own payload, and
take the error type from the status code

* fix(proxy): use the shared Anthropic error mapping for the stream error frame

The first pass added a third copy of the status to error-type table, and it
disagreed with the documented one: 529 came out as `api_error` rather than
`overloaded_error`, and 413 as `invalid_request_error` rather than
`request_too_large`, which hides the two failures a client can actually act on.

Drop that copy and put the frame builder next to the table litellm already
keeps in anthropic_interface/exceptions. The bridged adapter path was building
the same frame inline, so it uses the shared one now too

* fix(proxy): seal a torn SSE frame before the /v1/messages error event

An upstream that drops mid-frame leaves the client inside an open event,
so the error frame that follows is glued onto the torn data line and the
Anthropic SDK raises a JSON decode error instead of an APIStatusError.
Close the open frame with a ping event the SDK skips before writing the
error event, and add the e2e stream-cut edge with Bedrock, Anthropic
boundary, and Anthropic mid-frame legs.

* fix(proxy): keep the SSE tail unchanged on a chunk that is not text

A serializer that hands a dict or model object through as-is has no bytes
the frame tail can learn from, so advance_sse_tail leaves it alone instead
of slicing it.

* fix(proxy): answer a /v1/messages stream that fails before its first byte as a JSON error carrying its status

* test: move the Anthropic error frame tests into tests/unit

* test(e2e): cut the upstream stream only after content has been relayed

* test(e2e): carry split SSE lines across chunks and always tear a data line mid-frame

* test(e2e): find the next data line across a chunk boundary before tearing it

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-25 15:36:35 -07:00
devin-ai-integration[bot]
601c75a475
fix(proxy): record streamed /v1/responses container ownership before the response.completed frame (#43140)
* test(e2e): cover Azure code_interpreter container files by native id with a service-account key

* test(e2e): require the code_interpreter tool, skip at collection, and scope the container call timeout

* fix(e2e): fail the containers suite when the Azure credentials are missing instead of skipping

* fix(proxy): record streamed /v1/responses container ownership before the response.completed frame

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-25 13:45:43 -07:00
devin-ai-integration[bot]
2701e2008b
fix(ui): group cost optimization cache leakage by model group (#43008)
* test(ui): cover cost optimization cache leakage grouping by model group

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): group cost optimization cache leakage by model group

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 23:51:51 -07:00
devin-ai-integration[bot]
9915df6875
test(e2e): cover Azure code_interpreter container files by native id with a service-account key (#43122)
* test(e2e): cover Azure code_interpreter container files by native id with a service-account key

* test(e2e): require the code_interpreter tool, skip at collection, and scope the container call timeout

* fix(e2e): fail the containers suite when the Azure credentials are missing instead of skipping

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 22:16:18 -07:00
devin-ai-integration[bot]
020e5dee9b
fix(anthropic): keep the replayed prefix byte-stable for preserved thinking on chat completions (#42630)
* feat(anthropic): placement policy for mid-conversation system messages

Pure functions over the OpenAI-format message list: split off the leading
system run, keep later system messages as role=system at a placement Anthropic
accepts on models flagged supports_mid_conversation_system (after a user turn,
before an assistant turn or the end, never adjacent), and convert them to user
turns in place elsewhere, keeping tool_result first in a merged user turn.

* fix(anthropic): keep mid-conversation system out of the chat completions system prompt

translate_system_message hoisted every role=system message, at any index, into
the top-level system block. On a conversation carrying a mid-session reminder
that rewrites the cached prefix, so the provider re-bills the whole history at
cache-write pricing on every turn (#36559). #36968 fixed this on /v1/messages;
the chat completions path, shared by first-party Anthropic, Vertex, Azure AI
and Bedrock Invoke, still hoisted.

Only the leading system run becomes the system prompt now. Later system
messages go through the placement policy, and anthropic_messages_pt emits a
system message instead of rejecting the role. The caller's message list is no
longer mutated. Tests pin the two-turn prefix invariant across all four chat
configs and both flag states.

* refactor(anthropic): single-source the converted system note

The /v1/messages pass-through and the chat completions path must prefix a
converted system turn with the same operator note.

* test(e2e): prove the prompt cache survives a mid-conversation system reminder on chat completions

Same priming and assertions as the /v1/messages cases, through
/v1/chat/completions with OpenAI-format messages, for first-party Anthropic
and Bedrock Invoke on a flagged (Opus 4.8) and an unflagged (Haiku 4.5) model.
The reminder sits between the assistant turn and the next user turn, the shape
OpenAI-style agent frameworks send, which is the placement the chat path has
to translate.

* test(anthropic): cover the cache_control rebuild shapes and type the test helpers

Codecov flagged the 5m ttl branch and the empty-system path of the wire
builder; both now have a test. Greptile asked for full typing on the new
test helpers.

* refactor(anthropic): read the mid-conversation flag through a public supports_ helper

supports_mid_conversation_system joins the other supports_* helpers in
litellm.utils, so the chat transformation stops importing the private
_supports_factory.

* chore(typing): declare the mid-conversation type aliases with TypeAlias

The Final sweep tightened LIT010, which exempts TypeAlias declarations but
counts a bare alias assignment as an unannotated binding.

* fix(anthropic): let add_code_execution_tool take the pass-through message union

The translator now emits role=system inside messages for models that accept it,
so anthropic_messages_pt returns the pass-through union. add_code_execution_tool
still declared the narrower user/assistant union while only ever reading
content, so upstream's strip_advisor_blocks_from_messages call in between made
the mismatch visible to the type checker.

* fix(bedrock): keep mid-conversation system messages in place on converse path

* fix: ruff format + multi tool_result order + regression test

* fix: satisfy type-discipline gate + update osv ignore for mlflow PYSEC-2026-3865

* fix(bedrock): restore role narrowing in hoisted system loop for basedpyright budget

* test(bedrock): cover mid-conversation system conversion branches

- non-dict guard in _opens_with_tool_result
- in-place conversion without tool context
- str/list cache_control preservation in mid-conversation path
- drop unreachable non-system guard in hoisted loop

* Place type-discipline suppressions on the lines the gate scans

* Narrow hoisted loop to system role so basedpyright sees the right TypedDict

* fix(anthropic): place mid-conversation system runs by their neighbours only

A run after an assistant turn now slides behind the user turn that
immediately follows it, and a run that ends the array or precedes an
assistant turn becomes a user turn in place. No later message can move
an earlier run, so a client that replays the conversation with more
turns appended sends a byte-identical prefix and preserved thinking
blocks keep their binding

* refactor(bedrock): share the converted system note with the anthropic module

Converse imports CONVERTED_SYSTEM_NOTE instead of carrying its own copy
of the same text, and the reordering helpers lose their comments

* test: pin the replayed request prefix across preserved-thinking turns

One test per audited feature, through the real entrypoint: the chat
transformations for anthropic, bedrock invoke, vertex and converse, the
modify_params dummy tool result, dotprompt with unchanged variables, and
Presidio masking against an in-process fake. Each serializes system,
tools and the earlier messages of turn N and N+1 and asserts they match.
The e2e mid-conversation system test imports its content blocks from
models.py again and is marked provider_live

* fix(anthropic): move mid-conversation system placement into prompt_templates

The prompt factory imported the placement helper from the Anthropic provider
package, whose common_utils reads a factory constant at import time, so loading
the factory first raised ImportError. The module now sits next to
anthropic_messages_pt and every consumer imports core utils

A user turn with content [] or None puts no block on the wire, so a system run
anchored to it landed first in messages or behind an assistant turn. Such a run
now converts in place; empty strings and empty text blocks still anchor because
the factory fills them with a placeholder

* fix(anthropic): anchor system messages only on user turns that reach the wire

* fix(bedrock): type the converse system-message helpers over the message TypedDicts

* fix(anthropic): read replayed pydantic messages in the Converse helpers and convert a system run whose assistant follower sends nothing

A history that replays the previous turn as the litellm.Message object
was invisible to the Converse system-message helpers, so a mid-conversation
system stayed between a tool call and its result or reached Converse as
role: system. The helpers now read fields through the shared
message_field and parts_of accessors and drop the local role predicate.

Flagged placement anchored a system run on any assistant follower, but
anthropic_messages_pt drops an assistant turn that puts no block on the
wire (content None, an empty list, an unsigned thinking part), so the
system landed directly before the next user turn, which Anthropic
rejects. Such a run now converts in place. An empty or whitespace text
turn still anchors, since the converter pads it with a placeholder.

* fix(anthropic): treat bridged encrypted reasoning as a vanishing assistant turn for system placement

An assistant turn whose only blocks carry Responses API encrypted reasoning is
dropped by anthropic_messages_pt, so a mid-conversation system run anchored
before it landed directly before the next user turn. The unsignable-thinking
predicate now lives in common_utils and both the factory and the placement
policy consult it.

* fix(anthropic): let an inline thinking part hide separate thinking_blocks in system placement

anthropic_messages_pt skips an assistant turn's separate thinking_blocks as soon
as its content list carries an inline thinking or redacted_thinking part, so a
turn whose inline part is unsigned puts nothing on the wire even when the
separate block is signed. The placement policy now mirrors that rule.

---------

Co-authored-by: Shifat Islam Santo <shifatislamsanto764@gmail.com>
Co-authored-by: ege-arhan <egearhany@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 22:01:20 -07:00
devin-ai-integration[bot]
57bead9842
test(e2e): pin end-user and tag attribution from Codex-style headers on /v1/responses (#43093)
* test(e2e): pin end-user and tag attribution from Codex-style headers on /v1/responses

Codex CLI has no body field for the end user, so its config.toml
http_headers attach x-litellm-customer-id or x-litellm-end-user-id plus
x-litellm-tags to every /v1/responses call. The proxy already honors
those headers on the Responses route, but nothing in the e2e stack
pinned it. The new case sends that exact wire shape with each standard
customer header and fails unless the spend row carries the end user,
the tags, the aresponses call type, and a nonzero cost.

* test(e2e): assert the customer's /customer/info total matches the header-attributed Responses row

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 18:33:49 -07:00
devin-ai-integration[bot]
dc83a9c979
docs(e2e): carve harness tests out of the no-unit-tests hard rule (#43076)
* docs(e2e): carve harness tests out of the no-unit-tests hard rule

The Hard Rules bullet in tests/e2e/AGENTS.md banned unit tests of any
kind, the harness's own included, while the same file's claude_code/
and load/ entries, CONTRIBUTING.md, and the harness conftest all
describe markerless harness tests that run without a proxy. Reword the
rule to keep the product-feature and no-mocks bans, name the harness
trees as the one exception with the standard they are judged by, and
put the harness sentence back in the marker paragraph so the two docs
agree

* docs(e2e): name env vars set through monkeypatch as inputs, not patches

The rule banned monkeypatching anywhere under tests/e2e while its harness
carve-out named root-level test_*.py files that set env vars through
pytest's monkeypatch fixture. Say the ban is about patching code and that
an env var set that way is an input, so the examples and the rule agree

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 16:43:34 -07:00
devin-ai-integration[bot]
7faeb15ff3
fix(e2e): skip unpublished npm versions in the Claude Code PR-gate resolver (#43053)
npm keeps an unpublished version's timestamp in the packument's time map but drops it from versions, so the resolver could hand npm install a version it refuses with ETARGET. Only versions still present in versions are candidates now.

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 15:17:31 -07:00
devin-ai-integration[bot]
b72a030501
test: take keys out of the legacy proxy, enterprise and mcp unit tests before moving them (#42901)
* ci: fix the litellm-tests unit job with sysmon coverage, an env allowlist and coverage upload on failure

* test: replace key-dependent proxy, enterprise and mcp unit tests with synthetic values and integration and e2e coverage

* test: drop key reads at the legacy proxy, enterprise and mcp paths and wire the gemini pass-through split

* ci: fail the unit shard when circleci tests split errors

* test: drop restating comments from the gemini pass-through split

* ci: exit the unit shard cleanly when circleci tests split assigns it no files

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-09-24 14:56:52 -07:00
yuneng-jiang
2be2d68ac3
test(e2e/ui): hide the LiteAdmin button in the shared admin session (#43033)
#42443 pins a floating LiteAdmin button to the bottom-right corner, where it covers the logs page's next-page control. Flip the per-user Hide LiteAdmin switch during global setup so every spec reusing the admin storage state loads with the button hidden.
2026-09-24 14:48:35 -07:00
devin-ai-integration[bot]
86ba4fc16f
feat(compat-matrix): resolve and install the Claude Code CLI per run (#43038)
* feat(compat-matrix): resolve and install the Claude Code CLI per run

* chore(compat-matrix): drop the installer header comment

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 14:35:03 -07:00
devin-ai-integration[bot]
1fbd1e9ce9
fix(bedrock): route unmapped openai family model ids to converse (#42713)
* fix(bedrock): route unmapped openai family model ids to converse

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(bedrock): rename the e2e openai family backend constant

global.openai.gpt-6-sol has a cost-map row now, so the constant no longer
names an unmapped model

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 13:08:27 -07:00
devin-ai-integration[bot]
c9e8a04139
feat(vertex): native batch JSONL passthrough with cost tracking (#42810)
* feat(vertex): native batch JSONL passthrough with cost tracking

Add a per-request `passthrough=true` multipart field on `POST /v1/files`
(and the same kwarg on `litellm.create_file`) that uploads a native
Vertex AI batch JSONL to the deployment's GCS bucket unchanged, so rows
using `googleSearch` and other Gemini-only features run as written and
the output, `groundingMetadata` included, comes back untouched.

Passthrough is sticky through the GCS object path
(`litellm-vertex-files/passthrough/...`), so batch create and output
retrieval inherit it without new state. Native output rows are costed
from their `usageMetadata` with the deployment's model and model_info,
in the polling and retrieve paths and for the existing global
`disable_vertex_batch_output_transformation` flag, which billed $0
before.

The proxy requires the target to resolve to vertex_ai deployments only,
refuses `passthrough` with a non-batch purpose, a non-default
`target_storage`, or pre-call guardrails, and validates native rows on
`request` instead of the OpenAI batch keys.

* refactor(vertex): keep native batch row pricing inside the Vertex adapter

Moves native Vertex batch row detection, response parsing, and per-row
pricing from litellm/batches/batch_utils.py into
litellm/llms/vertex_ai/batches/transformation.py, so batch_utils only
aggregates the rows it gets back. Adds tests/test_litellm/files to the
misc unit shard so the new test directory is claimed by a shard.

* fix(files): say what a passthrough batch upload takes when a row is not native

The missing-key 400 listed bare key names, so an OpenAI-shaped row under
passthrough=true read "Each line must be a JSON object with keys request".
The batch line shape now carries its own hint, and the passthrough one says
a passthrough upload takes native Vertex batch rows with a request key

* fix(batches): bill native Vertex embedding batch rows on the native cost path

A native Vertex output row whose response holds an embedding was validated as a
generateContent response, so the documented tokenCount-only shape counted as a failed
row. Price embedding rows from their own usage (promptTokenCount, else tokenCount) with
the helper the transformed embeddings path already used, and drop the prompt-details
helper nothing calls anymore.

* fix(batches): keep modality batch rates on native Vertex embedding rows

An embedding row that carries usageMetadata was billed from promptTokenCount alone, so
its promptTokensDetails no longer reached the audio, image, and video batch rates the
way it did before the native cost path. Run every row with usageMetadata through the
Gemini usage parser and keep the flat tokenCount fallback for embedding rows without it.

* fix(batches): price native Vertex batch rows by modelVersion under a wildcard deployment

A `vertex_ai/*` deployment hands the batch cost path `*` as the deployment model, which
no cost map resolves, so every native (passthrough or flag-on) row was billed at $0. A
wildcard deployment model now defers to the row's own `modelVersion`, the way the
transformed path already prices by the row's `model`.

Also moves the native passthrough tests under tests/test_litellm, the tree codecov
reads, and covers the raw upload chunking, the embedding output translation, the
unpriceable-row path, and the flag-on dispatch.

* fix(batches): keep explicit deployment prices for native Vertex rows without a modelVersion

Under a wildcard deployment a native batch row that carries no modelVersion (an embedding
row, or a generateContent row Vertex returned without one) was billed at $0 even when the
deployment's model_info sets explicit batch prices, because the cost calculator was never
called. The row now falls back to the wildcard name, which the cost calculator prices from
the explicit model_info, and only a row with neither a modelVersion nor a deployment model
is billed at $0 with the warning

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 12:35:34 -07:00
yuneng-jiang
259f954f56
test(e2e/ui): give the logout specs their own admin session (#42930)
#42463 made Logout revoke the dashboard session key on the server. Both
logout specs ran on the shared ADMIN_STORAGE_PATH session that globalSetup
mints once, so clicking Logout revoked the key every later admin spec reuses.
The CircleCI run is serial, and from the auth/ folder on, every admin-session
spec failed with "Invalid proxy server token passed" (80 failures, up from 7)
while the internal-user, internal-viewer and team-admin specs kept passing.

Each logout spec now starts from an empty storage state and logs in through
the login page, so the session it revokes is its own. The login steps live in
a shared logInThroughLoginPage helper next to the other onboarding helpers.
2026-09-24 11:55:07 -07:00
devin-ai-integration[bot]
12960f3edf
fix(ui): pass is_proxy_admin for proxy admins on the models page team drill-in (#43003)
* fix(ui): pass is_proxy_admin for proxy admins on the models page team drill-in

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): drop explanatory comments from the models page team drill-in tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): exclude view-only admins from is_proxy_admin on the models page team drill-in

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): add browser integration contract for the team guardrail kill switch on the models page

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 11:54:52 -07:00
devin-ai-integration[bot]
8477fe4108
fix(proxy): list key and team model aliases in GET /v1/models (#42908)
* fix(proxy): list key and team model aliases in GET /v1/models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep alias listing helpers within the type discipline budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover alias rows on GET /v1/models and /v1/models/{id}

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): apply team then key aliases like chat completions and keep the alias as the retrieved id

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): apply key aliases twice like chat completions and skip only malformed alias entries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): apply the global model_alias_map between the key alias passes like chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): list only the caller's own aliases and never rewrite a listed model id on retrieval

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy): ruff format model_info alias lookup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): hide undiscoverable names from model retrieval so an alias named like one resolves to its target

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep undiscoverable models retrievable by id while excluding them from the alias guard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): pass an immutable name sequence into the model_info alias guard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): type the model list alias test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): annotate the new alias listing test fixtures and helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 06:41:13 -07:00
devin-ai-integration[bot]
1175559c39
feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00