The streaming hook is registered globally, so it runs on every
streaming response the proxy serves. It drained the whole stream into
a list before checking whether shunt was armed for the request, so
every unarmed request's full response sat in memory for nothing and
enough concurrent long streams could exhaust a worker with shunt
switched off everywhere. Resolve the config first and pass an unarmed
stream straight through.
The regression test asserts the interleaving rather than the chunks,
since buffer-then-replay returns the right chunks either way: the hook
must yield its first chunk while the source still has more to produce.
Inject _worker_text's processor factory and send function instead of
monkeypatching methods onto ProxyBaseLLMRequestProcessing in tests,
per the repo's dependency-injection rule. Typing send as
Awaitable[Awaitable[...]] also encodes route_request's two-await
contract, so the missing second await this PR fixed earlier would now
be a type error rather than a runtime 502.
The two route handlers and _worker_config had no unit coverage, only
the live run. Adds tests for the parts that own real behavior: the
worker model comes from the marker's config rather than anything the
caller names, a router without shunt armed is refused, team and tags
scope the lookup, bulk_read pairs each upload with its own filename,
and code_write strips the fences a chat model wraps code in.
Verified by mutation: swapping bulk_read's worker model, dropping
strip_code_fences, and reversing the filename pairing each fail a
test. Module coverage goes 80% to 98%, worker.py to 100%.
route_request resolves the deployment and hands back the provider
coroutine unawaited, so awaiting it once yielded a coroutine, not a
response. Every /v1/bulk_read and /v1/code_write call 502'd with
"worker model returned no completion". Caught by running the generated
command against a live proxy; the unit test missed it because the fake
route_request returned a ModelResponse directly instead of a coroutine.
Fix the second await and the fake, so the test now fails without it.
_read_upload_text checked the incoming remaining_total_bytes but used
min(_MAX_UPLOAD_BYTES_PER_FILE, ...) for the read size, so a zero or
negative LITELLM_SHUNT_MAX_UPLOAD_BYTES_PER_FILE made the read size
negative while the budget was still positive. UploadFile.read treats
that as "read the whole file", turning the bound into no bound. Check
the computed limit instead, which covers both inputs, and read the
three limits with get_env_int_in_range so an out-of-range override
warns and falls back rather than silently disabling the cap.
Restore test_settings.py to its committed state. Rewriting it dropped
the ENABLE_TOOL_SEARCH assertions, the no-mutation regression, and the
comment explaining why every model tier needs its own env override.
None of that was related to this PR.
The worker call went straight to llm_router.acompletion, so it skipped
everything common_processing_pre_call_logic does for /chat/completions:
registered guardrails, rate limiters, and budget checks. A caller already
over budget or rate limited could keep spending through these endpoints.
Route through ProxyBaseLLMRequestProcessing + route_request instead, and
let _handle_llm_api_exception map failures, so a rejected request also
releases its rate-limit reservation rather than leaking it.
Also bound upload reads. bulk_read and code_write read each UploadFile
without a size limit, so a large multipart body was buffered whole. Add
per-file, aggregate, and file-count budgets, with each read capped by
what the aggregate budget has left.
_read_upload_text rejects a non-positive remaining budget before calling
.read() rather than letting min() produce a non-positive limit:
UploadFile.read treats a negative size as "read the whole file", which
would silently defeat the budget for any caller that passes one.
Rename test_endpoints.py to test_shunt_worker_endpoints.py: pytest
imports test modules by basename with no __init__.py present, so it
collided with credential_endpoints/test_endpoints.py.
Three fixes from this review round, plus a CI shard registration and a comment-density pass.
A caller admitted through JWT or a custom auth path has no DB-backed key hash
(UserAPIKeyAuth.api_key is None or some other non-sk- value), which crashed the capability
token mint instead of leaving the tool_use untouched. _mint_caller_capability_token now
returns None for that case, and _endpoints_for_request treats it the same as an unreachable
base URL.
A resolved master-key caller carried the real master key as its own api_key, which
_worker_text later placed in the outbound request's metadata["user_api_key"] -- reachable by
any raw-metadata logging callback. Now substitutes LITELLM_PROXY_MASTER_KEY_ALIAS there,
matching what normal master-key auth already does for exactly this reason.
The worker call went straight to llm_router.acompletion, skipping every registered rate-limit
and budget callback (they run as async_pre_call_hook, which only proxy_logging_obj.pre_call_hook
walks). A caller already over budget or rate-limited could keep spending through this endpoint
indefinitely. _worker_text now calls pre_call_hook first and lets a block propagate.
Registers tests/test_litellm/proxy/shunt_endpoints in test-unit.yml's proxy-endpoints shard;
CI's shard-coverage assertion failed without it since the directory held tests but named no
owning shard.
Trims several docstrings/module comments in the shunt modules down to the non-obvious "why"
they were justified by, cutting repetition and one stale field name a rename had left behind.
Replaces the ANTHROPIC_AUTH_TOKEN/ANTHROPIC_API_KEY environment-variable read with a
short-lived token the proxy mints itself, so the generated command authenticates to
/v1/bulk_read and /v1/code_write without depending on the calling client's shell holding
either variable. That dependency only ever held for Claude Code; any other client (Cursor, a
custom agent) would have sent an empty bearer and 401'd.
The token is a sealed grant (litellm/proxy/guardrails/shunt_capability_token.py) built on the
proxy's own encrypt_value_helper, the same primitive the gateway's OAuth flow already seals
values with. It carries a reference to the caller's key hash rather than the key itself, and
a two-minute expiry rather than a single-use guard: an agent retrying a timed-out Bash command
must still authenticate, and a single-use claim would turn that ordinary retry into a
permanent 401. The worst a replay inside the window can do is spend the caller's own already-
budgeted quota on a request they already made.
Carried in the generated command's Authorization header, never a URL query string: every
other sealed token in this proxy already avoids query strings, since they routinely end up in
access logs.
/v1/bulk_read and /v1/code_write now authenticate exclusively via this token instead of the
normal user_api_key_auth path, since nothing but a shunt-generated command should ever call
them. Master-key callers (UserAPIKeyAuth.api_key holds a stable alias rather than a DB-backed
hash for that case) carry the real master key in the grant instead, compared directly at the
endpoint.
Shunt is three litellm_params fields layered onto any complexity router, not
its own routing strategy. Presenting it as a peer of anthropic_family/
openai_family/etc. in the template dropdown made it look mutually exclusive
with a model-family choice when the two were never actually in conflict, and
picking it meant swapping in shunt's own hardcoded tier list.
Removes the "shunt" entry from autorouter_presets.json and the partial-fit
preset machinery that existed only to support it (isPartialFitPreset,
PARTIAL_FIT_PRESET_KEYS, hasNoUsableModelsAtAll, dropUnresolvedTierEntries,
shuntStateFromPreset). The Advanced: Shunt section on the create form is
unaffected and stays off by default: a caller composes it with whichever
model-family preset (or Custom Configuration) they actually want.
Both Greptile and veria-ai flagged these as auto-approving too much: Claude Code
matches an allow rule's text up to its first wildcard with no host awareness, so
"Bash(curl -sS -F question=*)" approves a curl to any destination, not just this
proxy. A prompt-injected repo instruction could get a local file uploaded
somewhere without the user ever seeing a permission prompt.
lite autoroute up already prompts for genuinely new commands and lets a user
persist their own approval once they've seen it, which is the safer place for
that decision to live. Removes SHUNT_BASH_ALLOW_RULES and the settings.json
permissions merge entirely; the static-token env merge that up already did is
unaffected.
Addresses the review findings on this PR.
The generated commands escaped only double quotes, so a model-supplied path, question, spec,
reference, or target containing $(...) or backticks was command-substituted and ran on the
developer's machine. Every interpolated value now goes through shlex.quote, and the size-check
line reports the path with printf instead of a double-quoted echo. Confirmed against a real
bash: the payload used to create its marker file, and no longer does.
The commands also carried the caller's Authorization header verbatim, which put the key in the
model's response, the conversation history, and the next upstream turn. They now reference
${ANTHROPIC_AUTH_TOKEN:-$ANTHROPIC_API_KEY} and the shell resolves it locally, so the secret
never leaves the client. That also fixes clients authenticating with x-api-key, which were
skipped entirely because only the authorization header was read.
Files were uploaded as paths[] while the endpoint binds them under paths. FastAPI matches the
form name exactly, so every delegated read returned 422 and the feature never actually worked.
The worker endpoints resolved the marker with no tags, so a marker armed only under a tag
returned 400 even though the rewrite had matched it; the tags now ride along in the query
string. They also sent only a team id to the router, so worker spend was not attributed to the
calling key. They now reuse the proxy's own key-metadata builder.
An unset worker model could fall through to an empty string. It now falls back to the SIMPLE
tier, then the default model, and refuses to arm rather than naming no model at all.
A caller that already defines its own bulk_read or code_write got a duplicate definition
injected and its own tool calls rewritten into shunt's curl. Shunt now leaves those requests
alone in all three hooks.
The allow rules only matched commands starting with curl, but the bounded read starts with the
wc -l size check, so the main large-file path still prompted. Added a rule for that shape.
Registers the new routes in the gateway allowlist, which the component-coverage test requires.
Ports the shunt technique for Claude Code
(https://engineering.atspotify.com/2026/9/portal-by-spotify-cut-my-claude-code-token-usage-by-90)
into the auto router, server-side, so it works for any client with no plugin to install.
shunt intercepts large file reads and boilerplate generation at the client with PreToolUse
hooks and hands them to a cheap worker model. This makes the same decision on the proxy: a
new always-on ShuntGuardrail injects bulk_read and code_write tool definitions pre-call, then
post-call rewrites large Read/Bash/bulk_read/code_write tool_use blocks into a Bash command
carrying shunt's own wc -l conditional, so small files still get read directly and only large
ones are delegated. Two new endpoints, /v1/bulk_read and /v1/code_write, run the worker call
through llm_router.acompletion with shunt's verbatim system prompts, so worker spend is
tracked against the calling key and team like any other request.
The guardrail arms per request off an auto-router marker's own litellm_params
(auto_router_shunt_min_lines and the two worker-model fields), the same shape
auto_router_compression already uses, so no guardrails: config entry is needed. It stays
inert when the resolved model has no shunt config.
In the UI, "Shunt" is one entry in the existing preset dropdown. Its tiers resolve through
fallback chains against whatever models the proxy actually has, so it greys out only when
there are no chat models at all, never merely because the named models are absent, and all
three settings stay editable under Advanced.
Streaming buffers the whole response before rewriting, matching tool_permission.py, because
input_json_delta fragments split mid-token. An unparseable stream passes through untouched:
shunt is an optimization, not a safety control, so a request should never fail because the
rewrite could not run.
The CLI's Bash allow rules only reduce prompts for the generated commands. Claude Code
matches rule text before the first wildcard with no host-aware matching, so they cannot pin
the destination, and the docs recommend a PreToolUse hook where a real boundary is needed.
* fix(router): keep provider response headers on streaming chat completions
The Router re-wraps a deployment's CustomStreamWrapper in FallbackStreamWrapper
(and its sync twin) so a mid-stream failure can fail over. Neither wrapper
forwarded `_response_headers`, so every streaming chat completion handed the
proxy's callbacks and its response-header builder a wrapper with no provider
headers, and a successful mid-stream fallback still published the failed
deployment's identity, `x-request-id` and rate limit counters.
Forward `_response_headers` into both wrappers, repoint the wrapper at the
deployment that served the stream once a fallback takes over, and rebuild the
proxy's response headers from that deployment while `create_response` still has
the first chunk buffered.
* fix(router): follow a nested fallback to the deployment that served the stream
A fallback the router picks is itself a fallback-aware wrapper, and it only
repoints at its own fallback once it yields, so reading its hidden params at
selection time named a deployment that produced no output. Re-read them when
the first fallback item arrives, which is still before the proxy commits
response headers.
Also addresses review feedback: the streaming header builder reads self.data
instead of taking a coarse request_data parameter, and the new test recorder
local is Final.
* test(router): cover the fallback header adoption helper directly
The router_code_coverage gate wants every router.py function named in a
router test, and this also pins the weak-reference behavior: a wrapper
collected mid-stream must not break the generator still draining it.
* refactor(proxy): take a read-only mapping for the model-id lookup
_get_model_id_from_response only reads its request payload, so a Mapping
says what it needs and the two metadata hops are narrowed instead of
assumed to be dicts.
* test: drop mutable recorder locals and routine comments from the new tests
An AsyncMock await_count and an asyncio.Event say the same thing as a
list and a dict that the test mutates.
* chore(router): justify the two rebinds in the fallback loops
Both are the one-shot re-read that follows a nested fallback, so they get
the repo's rebind-ok note like the rest of the file.
* fix(ui): show inherited MCP servers on the internal-user editor and flag access groups with no members
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): consult the unfiltered access group registry before calling a group empty
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(skills): semantic search over the LiteLLM-hosted skill registry
Adds GET /v1/skills?query= (custom_llm_provider=litellm_proxy) and a
skill_search MCP virtual tool, ranking the caller's accessible skills by
semantic similarity, mirroring the A2A agent registry search (LIT-6309).
Also fixes a pre-existing bug where create_skill() dropped description and
instructions for the litellm_proxy provider, which left every LiteLLM-hosted
skill with no searchable text.
* fix(mcp): coerce skill_search top_k instead of raising 500 on malformed input
The MCP-REST skill_search dispatch validated raw tool arguments through a
pydantic model directly, so a non-numeric top_k raised a ValidationError
that the endpoint's catch-all turned into an HTTP 500. Mirrors the
agent_search branch's tolerant coerce_top_k handling instead.
* fix(skills): enforce key limits on search embeddings and bound the semantic index
Semantic search embeddings now run the same pre_call_hook the /embeddings
route runs, so key rate limits, budgets and guardrails apply before the
embedding model is called. The shared SemanticTextIndex caps cached vectors
and evicts the least recently searched entries, and each skill's embedded
text is capped so one skill cannot inflate the embedding batch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(skills): surface proxy 429s from search embeddings instead of a 503
ProxyRateLimitError is also an OpenAIError, so the search engine was folding
a key rate limit into skill_search_unavailable. Proxy HTTPExceptions now
propagate so the caller gets the same 429 the /embeddings route returns
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(skills): import assert_never from typing_extensions for Python 3.10
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(skills): embed the request as the pre-call hooks returned it, not the original text
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(skills): keep the litellm_proxy provider check for GET /v1/skills?query= inside llms/
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(skills): move the GET /v1/skills?query= endpoint tests under tests/test_litellm/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): hide the Create Vector Store flow from non proxy admins
The vector stores page rendered the Create Vector Store tab, the
+ Add Vector Store button and a GET /credentials call for every role,
while the proxy only lets proxy admins call POST /vector_store/new and
GET /credentials. Internal users landed on the create form and got an
Only proxy admin error toast. Gate all three on isProxyAdminRole and
default everyone else to the Manage tab, matching the Indexes tab and
the Add Model gating.
Resolves LIT-7131
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): exclude view-only admin sessions from the vector store create flow
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(ui): add key-scoped auto-router usage tab
GET /auto_router/benchmarks takes an optional api_key filter, applied in the
rollup aggregate on the primary key's leading column. Proxy admins get a
separate Auto-router usage tab on key detail pages with spend, baseline,
savings, tier routing, cache metrics and the existing router selector
* fix(ui): share key analytics date range
* feat(proxy): resolve root_path per request from SERVER_ROOT_PATHS
One deployment can encode exactly one client-visible URL path prefix
today: SERVER_ROOT_PATH is a scalar stamped onto the app at startup, so
a pod fronting several ingress prefixes 404s every prefix but one before
any handler runs, and MCP OAuth discovery can emit only one prefix's
URLs (RFC 9728 section 3 exact-match fails for the rest).
Add an opt-in outermost ASGI middleware that matches the request path
against a configured prefix list (SERVER_ROOT_PATHS, comma-separated) on
a segment boundary and sets scope["root_path"] for that request only.
Everything downstream is stock Starlette: route matching strips
root_path so routes stay registered root-relative, and request.base_url
re-includes it, so the discovery documents' resource and the 401
challenges' resource_metadata land under the prefix the client actually
called — with no discovery-builder changes.
LazyFeatureMiddleware now strips the scope root_path (falling back to
the cached SERVER_ROOT_PATH scalar) before feature prefix matching, so
lazily-registered routers — the MCP OAuth discovery router among them —
load under per-request prefixes.
Follow-up to the routing discussion on #35226; composes with, but does
not depend on, #35576.
* fix(proxy): import Sequence from collections.abc (ruff UP035 strict-budget gate)
* review(greptile): trim implementation commentary; fixture-own MCP registry state in tests
Addresses both P2s from the first Greptile pass:
- per_request_root_path_middleware.py (and the related _lazy_features /
proxy_server comments) cut down to the constraints the code cannot
express, per repo comment guidance
- the new discovery tests no longer clear/repopulate the shared MCP
registry inline; a fixture snapshots it, hands the test an empty
registry, and restores it afterwards so no state leaks between cases
* fix(lint): mutable-ok marker on the prefix accumulator (LIT002 type-discipline gate)
* fix(proxy): tie 401 challenges and get_custom_url to the per-request root_path
The per-request root_path middleware sets scope["root_path"] to the
prefix the client actually called, but the OAuth 401 challenges
(raise_user_oauth_challenge / raise_token_exchange_challenge) still
built their resource_metadata from SERVER_ROOT_PATH. On a pod fronting
several prefixes, the challenge advertised a discovery URL under a
different prefix than the discovery document served — the two
disagreed on where the resource metadata lives, and a strict RFC 9728
client refused the challenge. Route the challenges through a small
ContextVar the middleware populates so they read the same effective
root_path Starlette resolves the request under.
The same accessor fixes get_custom_url: when a request lives under a
SERVER_ROOT_PATHS-matched prefix, request.base_url already carries it,
so appending the SERVER_ROOT_PATH scalar on top produced e.g.
/tenant-a/legacy/sso/callback — a path that does not exist. Reading
the per-request prefix instead (and relying on join_paths's tail-dedup)
keeps SSO login/callback URLs under one prefix — the one the request
actually arrived on.
Fallback: outside a request (module-load-time UI URL builders,
background tasks) the ContextVar is unset and the accessor reads
SERVER_ROOT_PATH, matching get_server_root_path() so scalar-only
deployments are byte-identical.
* fix(mcp): challenge URL under per-request prefix must route, and mock parity
Two follow-ups to the review fix that made the 401 challenge use the
per-request root_path:
1. oauth_protected_resource_path must pick the URL structure that
actually routes for the mechanism in use:
- The scalar SERVER_ROOT_PATH deployment registers the well-known
routes with the prefix INSERTED (via well_known_root_suffix at
import time), matching RFC 8414 §3. The challenge URL must use the
same insertion or a client fetching it 404s.
- The per-request SERVER_ROOT_PATHS deployment can't register routes
per prefix; PerRequestRootPathMiddleware strips the prefix from
scope["path"] and the router matches the un-inserted route. The
URL must place the prefix BEFORE .well-known so the strip leaves a
matching path.
The previous fix used the insertion form for both, which 404'd the
discovery fetch on the per-request path — the discovery doc and the
challenge would then disagree on where the resource metadata lives,
the very failure the review flagged. End-to-end verified: the URL
the challenge advertises routes and the doc's `resource` field
equals the URL the client originally called (RFC 9728 §3).
2. get_request_root_path now delegates its fallback through
get_server_root_path() instead of reading the env directly, so every
existing `monkeypatch.setattr("litellm.proxy.utils.get_server_root_path"`
test override keeps working. This unstubbed the mock on the /v2/login
test that failed on the last CI run.
Plus the lint budget: annotate the local accumulator Final, tag the
scope["root_path"] rewrite as an intentional ASGI-contract mutation,
tag the reused `path`/`root_path` rebinds in LazyFeatureMiddleware, and
add reason strings to the two new PLC0415 lazy-import noqas.
* test(mcp): pin the reviewer's expected end-state — challenge URL routes, resource matches called URL
End-to-end regression test that mounts the discoverable router + the
per-request root_path middleware, hits an MCP endpoint that raises
raise_user_oauth_challenge, fetches the resource_metadata URL the
challenge advertises, and checks the returned document's `resource`
equals the URL the client originally called (RFC 9728 §3 exact match).
Covers /tenant-a, /tenant-b, and the unprefixed path on the same app so
a regression on any prefix — challenge URL 404s, or doc emits a
different prefix than the client called — fails at this test rather
than in a strict MCP client's discovery.
---------
Co-authored-by: gym-cmd <186399764+gym-cmd@users.noreply.github.com>
* fix(spend_logs): keep partition DDL transactions alive for their statement timeout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style: ruff format changed files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lint): avoid dict-literal kwargs and keep cast-ok on the cast line
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lint): cast at the call site instead of widening PrismaClient.tx
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(spend_logs): require partition tx timeout to strictly exceed statement bound
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): prove the virtual key lifecycle on every replica
Walks one virtual key through create, read, partial update, clear, enforce
and delete against a live proxy and database, reading every write back on
every gateway replica.
The management suite already had single write-then-read tests for keys, but
none of them proved that a partial /key/update leaves the untouched fields
alone, that an explicit null clears a field, or that a write is visible on
more than the one gateway that took it.
Adds read_back_everywhere to the shared ProxyClient: it polls a GET path on
every URL in PROXY_REPLICA_URLS until each replica's parsed body satisfies
the caller's predicate, and fails naming the replica that never converged.
The CLEAR sentinel in the e2e models makes an explicit JSON null expressible
in a body the transport otherwise strips of None fields.
Documents /key/update's merge patch semantics on the endpoint docstring.
* test(e2e): prove key revocation and field preservation on every replica
Applies the findings from an adversarial review of the first commit.
The delete step only checked that chat was refused on the gateway that took
the write, so it would have passed while a sibling gateway kept serving the
deleted key. It now serves one call from every replica first, so each has the
key cached and the delete has something to revoke everywhere, then polls every
replica for the refusal.
The file also carried its own poll loop that tested the deadline before
attempting, so it gave up one attempt early and skipped the attempt landing
exactly on the deadline. It now shares the harness helper, which is generic
over the polled value rather than over a parsed body, so the same loop covers
both the info read-back and the chat refusal.
The model the enforcement step registers now carries a unique marker in its
alias, matching every other deployment this suite creates, so concurrent runs
never share one model group.
The docstring sentence claimed an explicit null clears any field. It does not:
the metadata-backed fields merge into stored metadata, where a null is a silent
no-op, and only the key's own columns clear. Regenerating the dashboard types
picks up the corrected text.
* fix(e2e): delete a deployment that never becomes servable
Registering a model posts /model/new and then waits for every replica to list
it. When that wait timed out the deployment already existed in the database but
its id had never been returned, so no caller could delete it and the row
outlived the run. It is now deleted before the failure propagates.
Found by review on the key lifecycle suite, whose module fixture registers a
deployment this way, but every caller of the shared helper had the same
exposure.
* docs(e2e): drop the duplicated notes from the lifecycle docstrings
The delete method restated what the warm-up helper already explains, and the
module restated the merge patch rule that the endpoint and the request model
both document.
* test(e2e/ui): cover member role and budget edits, member permission delegation, and team guardrail removal
Three Playwright specs for the Teams flows enterprise customers hit most, each
owning its fixtures and proving the mutation through a read-back rather than a
toast.
- teamMemberEdit: an admin edits a member's team role and per-member budget,
and both survive a reload of the Members table
- memberPermissions: a plain member is refused /key/generate for their team,
a team admin grants it on the Member Permissions tab, and the member then
creates a team key that serves a real completion
- teamGuardrailRemoval: clearing a team's only guardrail on the Settings tab
really clears it, and traffic the guardrail refused starts serving again
* test(e2e/ui): make the new team specs safe to run in parallel
Fixture ids came from Date.now(), so two repeats starting in the same
millisecond minted the same user id: one got a 409 and the loser's teardown
deleted the user the other was still signed in as. Ids now carry a random
suffix.
Also move the member-permissions setup inside the cleanup-protected block so a
half-finished setup cannot leak a team, and close both browser contexts the
test opens.
* test(ui): pin wire contracts for key, model and MCP server forms
Add vitest cases that pin what the key edit, key create, model edit and
MCP server edit forms put on the wire: an edited field reaches the
request with its new value, a cleared field reaches it as an explicit
null, and the dirty-only body is pinned as an expected failure until
each form moves to pickDirty. Model edit also pins the cost-map-derived
model_info fields as an expected failure.
KeyEditView hands a cleared max_budget to KeyInfoView as an empty
string and handleKeyUpdate maps it to null, so the null is pinned at
the /key/update boundary in key_info_view.test.tsx and the KeyEditView
case is an expected failure. buildEditServerPayload passes a cleared
description through as an empty string, so that case is an expected
failure too.
* test(ui): split masked model_info pins and retarget the create tracker
The model_info expected-failure case held three assertions, and it.fails
stops at the first one, so a later revamp that fixed max_input_tokens
while leaving mode leaking would still report an expected failure. Split
it into one case per pinned field group so each flips on its own.
The key create tracker asserted a body of only key_alias, which a create
can never send: key_type, user_id, duration and metadata are always
mounted. Retarget it at the real over-send, which is the Optional
Settings section adding fifteen undefined-valued keys when the user opens
it without filling anything in.
* fix(mcp): apply key and team guardrails to MCP tool calls
Guardrails attached to a virtual key or team were only enforced on LLM
routes. The synthetic request built for MCP tool call guardrail hooks
carried no guardrails in its metadata, so a guardrail with default_on
false never ran on tools/call even when the key explicitly listed it.
Resolve key, team, and project guardrails onto the synthetic request
with the same helper the chat path uses.
* fix(mcp): pass project metadata through without a mutable default
* fix(mcp): mark the request dict parameter mutable-ok with a reason
* test(mcp): explain the premium_user patch and tighten the helper docstring
* ci: limit Rust workflows to Rust directory changes
* ci: run Rust checks when their workflow changes
* ci: report Rust wheels only for successful Rust changes
* ci: keep Rust wheel reports in the workflow summary
* ci: group Rust lint and validation jobs
* ci: keep Rust job names distinct from required lint and test checks
* ci: drop the unused Python setup from the Rust lint job
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>