The chat translation layer emits each message content part as its own texts
entry, but the model concatenates adjacent parts (a multimodal message's text
parts join with no separator), so a blocked phrase split across two segments was
seen whole by neither per-segment scan and slipped through. Extend the existing
detection-only overlap approach with a window spanning each adjacent prompt-text
junction, bounded to the ordered prompt texts via WindowConfig.text_segment_count
so tool-call args and tool/function definitions are not falsely joined. The
tuning knobs move into a frozen WindowConfig to keep evaluate_segments within the
argument-count budget.
Reduce PLR0913 by collapsing get_or_create_client and resolve_credentials
arg lists into frozen ClientBuildSpec / CredentialConfig dataclasses, drop
ANN401 by narrowing Any to object/Callable on the SDK-facing helpers, remove
a redundant UP037 string annotation, and suppress the four net-new
reportMissingTypeStubs from the optional wonderfence_sdk imports with
pyright: ignore (the repo disables type: ignore for basedpyright).
MASK verdicts on legacy functions[].description now write the redacted value
back into request_data["functions"] via the same path-based mechanism used for
inputs["tools"]. Previously the path was discarded and MASK was silently
ignored, leaving the original unredacted description forwarded to the model.
Also update the test that asserted the old "detection only, no mask" behaviour
to reflect the corrected semantics (DETECT still logs without mutating; only
MASK writes back).
The deprecated top-level functions[] request parameter is forwarded to providers
(litellm converts it to tools only later, during the LLM call, after the
guardrail runs), so blocked content in functions[].description or nested
parameter descriptions reached the model unscanned. Each functions[] entry is
shaped like a tool's function object, so its descriptions are now extracted
(reusing the tool-definition walker) and evaluated as request-side segments.
Read from request_data because the chat translation layer surfaces tools but not
functions in inputs. Detection only: BLOCK raises and DETECT logs; functions has
no inputs write-back path so it is not masked (matching the precedent set by the
cisco_ai_defense guardrail, which scans both tools and functions for detection).
Regression tests: BLOCK on a function description, BLOCK on a nested parameter
description, a functions-only request still scanned, and detection-without-mask
leaving request_data["functions"] untouched; the BLOCK cases fail on prior code.
The basedpyright reportMissingParameterType gate flagged the lone unannotated
parameter in the new module (**kwargs on WonderFenceGuardrail.__init__), one
over the codebase ceiling. Annotate it **kwargs: Any; within the file's
any-discipline budget. No behavior change.
The rebase onto latest litellm_internal_staging pulled in newer governance gates
that the Alice WonderFence module (new to the base) tripped:
- UP045/UP006/UP035: use built-in generics and `X | None` (the repo targets
Python >=3.10) instead of typing.Optional/List/Dict/Tuple, removing the
net-new strict-rule violations that exceeded the codebase ceiling.
- recursive_detector: rewrite the tool-definition description walker iteratively
(explicit stack) instead of recursively; unbounded recursion over
caller-supplied tool schemas is a stack-overflow/DoS risk anyway.
- any-discipline: record per-file Any baselines for the module in
any-discipline-budget.json, matching how every other guardrail provider is
budgeted (SDK-interop and JSON traversal inherently surface Any).
The PR title was also lowercased to satisfy the Conventional Commits subject
check. No behavior change; 100 unit tests still pass.
The chat translation layer forwards caller-supplied inputs["tools"] to the model
verbatim, but apply_guardrail only scanned message text and tool-call arguments,
so blocked content in tools[].function.description or nested parameter
descriptions reached the model unevaluated. Extract every description string
from each tool definition (top-level and recursively through the parameters JSON
schema) as a request-side segment, evaluate it alongside the others, and write a
MASK verdict back to the originating slot via its path. BLOCK on any tool-def
segment blocks the request; the empty-input early return accounts for tool defs
too. Tool definitions are scanned by default like other request content;
operators who don't want their tool schemas evaluated can scope them out
upstream.
Regression tests: BLOCK on a tool description, BLOCK on a nested parameter
description, MASK written back to function.description in place, a tools-only
request still scanned, and path round-tripping for tool_definition_segments.
Splitting an oversized segment into disjoint <=MAX_PROMPT_CHARS chunks let a
blocked phrase straddle a boundary so neither chunk saw it whole. Multi-chunk
segments now also evaluate a detection-only window spanning each boundary (last
N + first N chars, N=CHUNK_OVERLAP_CHARS, clamped to max_chars/2), feeding
BLOCK/DETECT so a phrase up to ~2N chars can't slip through the split. Masking
still uses the disjoint chunks so the lossless rejoin holds; a boundary window
that flags maskable content surfaces as DETECT since it can't be redacted
across chunks. Single-chunk segments add no extra calls. Regression test: a
phrase straddling the boundary blocks with overlap and evades with overlap=0.
resolve_api_key / resolve_app_id guarded each source with a bare truthiness
check, so a non-string metadata.alice_wonderfence_api_key / _app_id (list, dict,
number) was returned as-is. It then reached the SDK/cache, raised a type error,
and with fail_open=True the broad handler in apply_guardrail swallowed it and
returned the request unscanned. Validate every source (request, key, team,
default) as a non-empty string; an invalid type is ignored, so app_id with no
other source raises WonderFenceMissingSecrets -> HTTP 500, which is never
fail-open. Regression tests cover the resolver level and an apply_guardrail
fail_open=True path that must 500 without calling the SDK.
The tool-call DETECT branch in apply_verdicts (the symmetric counterpart to the
text-side DETECT path) had no test, leaving two lines uncovered. Add a
regression asserting a DETECT verdict on a tool-call argument passes through
without blocking or mutating the arguments; this would catch a mutation that
turned tool-call DETECT into a block or mask. processing.py and the
alice_wonderfence package are now at 100% line coverage.
The stash recovery fell back to any sibling alice_wonderfence instance's stash
when the current instance had none. With two instances on one request, a strict
instance (allow_request_metadata_override=False) could inherit a permissive
sibling's caller-supplied request-body credentials, scanning under credentials
it would itself reject. Remove the fallback and fail closed when this
instance's own stash is absent.
The stash is now stored under a per-guardrail attribute
(_alice_wonderfence_resolved__<name>) rather than a shared dict keyed by name,
so the isolation is structural: there is no sibling slot to read. This only
affects multi-instance during_call-only configs (pre_call stashes each
instance's own, and key/team credentials re-resolve in post_call without a
stash), so common single-instance and pre_call configs are unchanged.
Regression tests: a strict reader with a permissive writer sibling fails closed
instead of borrowing, and recover_resolved returns None for a name that never
stashed even when a sibling did. Both fail on the prior implementation.
The post_call bridge stashed the resolved (api_key, app_id) under
logging_obj.model_call_details, which LiteLLM forwards verbatim as kwargs to
every success/failure callback and logging exporter; the redaction layer only
scrubs message input/output and known StandardLoggingPayload fields, not
arbitrary custom keys, so a tenant-specific WonderFence api_key leaked into
logs. Move the stash to a private instance attribute on the same logging_obj.
It is request scoped and visible across the pre/during/post hooks and the
asyncio.gather task boundary exactly as before (same object passed by
reference), but it is not part of the kwargs dict handed to callbacks.
Tests use a real LiteLLMLoggingObj (not a Mock, whose attribute auto-creation
would hide whether the attribute is genuinely settable/readable) and assert the
api_key never appears in model_call_details; that assertion fails on the prior
implementation. The post_call bridge tests now run against the real object too.
tool_calls reach the model (request side, from assistant messages) and the
client (response side, model-generated), and the translation layer threads them
through inputs["tool_calls"] and writes mutations back, but apply_guardrail only
looked at inputs["texts"]. Disallowed content placed in
tool_calls[].function.arguments therefore went unscanned. The early return also
skipped requests whose only content was a tool call (empty texts).
Each tool-call argument string is now evaluated as a segment alongside the text
segments through the same WonderFence call; BLOCK raises, MASK rewrites
inputs["tool_calls"][i]["function"]["arguments"] in place (the translation layer
writes it back), DETECT logs. The empty-texts early return now also accounts for
tool-call args. Regression tests cover request/response BLOCK on tool args, MASK
write-back, and the tool-calls-without-texts case; all fail on the prior code.
A caller can send `metadata` as a non-object value (string, list). The old
`caller or litellm_md or {}` returned that non-dict verbatim, so resolve_api_key
/ resolve_app_id then called `.get()` on it and raised; with `fail_open=True`
the guardrail swallowed the error and skipped scanning entirely, and the
proxy-injected `litellm_metadata` admin pins were dropped on routes like
/v1/responses where the caller bucket is separate. Coerce each bucket to {}
when it is not a dict before merging, so a malformed caller `metadata` can
neither bypass scanning nor shadow admin-pinned credentials.
Added regression tests (non-dict caller metadata: get_metadata preserves the
litellm_metadata admin pin; resolve_* succeeds from the pin instead of raising).
Filtering the request side back down to user-role messages reopened the same
class of bypass for non-user content: disallowed text placed in a system,
assistant (prefill), or tool message went unscanned while still reaching the
model. Evaluate every segment the translation layer hands us in inputs["texts"]
instead of re-filtering by role; whether a role is included is already governed
upstream by skip_system_message_in_guardrail / skip_tool_message_in_guardrail,
so the hook should not hardcode its own role policy. This also removes the
role-mapping replay and its count-mismatch fallback entirely.
example_config sets skip_system_message_in_guardrail: true so admin-controlled
system prompts are excluded by default, which avoids false positives while still
scanning the caller-controllable assistant and tool segments.
The messages array is fully caller-controlled and unverified, so placing
disallowed content in an earlier user turn and ending with a benign message
let it reach the model unscanned; only the last consecutive user block was
evaluated. Now every user-role message is evaluated on its own (each chunked
to the WonderFence prompt limit), all calls fan out in parallel under a
concurrency cap, and verdicts are aggregated per message with BLOCK > MASK >
DETECT precedence. Masking writes back only through texts, matching what the
chat translation layer reads.
The chunk + parallel-evaluate + aggregate logic lives in one replaceable unit
(chunked_evaluation.py) that is WonderFence-agnostic via an injected evaluate
callable.
get_metadata used `metadata or litellm_metadata`, which short-circuits: a
truthy caller-supplied `metadata` bucket meant `litellm_metadata` was never
read. On LITELLM_METADATA_ROUTES (e.g. /v1/responses) proxy-injected admin
pins land in `litellm_metadata` while the caller bucket is `metadata`, so a
caller could shadow the admin pins and fall through to their own
request-metadata override (only exploitable with
allow_request_metadata_override=True — the trusted-gateway case where pinning
must hold).
Merge both buckets with proxy-injected litellm_metadata winning on collision.
Admin pins (nested under user_api_key_metadata / user_api_key_team_metadata)
can no longer be shadowed; the caller's top-level request-override
alice_wonderfence_* keys don't collide and still survive the merge.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Caller-supplied metadata.alice_wonderfence_app_id / alice_wonderfence_api_key
no longer outrank admin-pinned key/team metadata. Adds
allow_request_metadata_override (default False) as an explicit opt-in for
trusted-gateway deployments — even when enabled, key/team metadata still
wins. Closes the high-severity precedence inversion flagged on PR #26901
(review comment r3226452019).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replace stray app_name="test-app" with comment noting app_id is per-request
via metadata.alice_wonderfence_app_id, matching example_config.yaml.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
OpenAI chat translation populates both `structured_messages` and `texts`
on guardrail input but reads back only `texts` after apply_guardrail
returns. MASK was writing only to `structured_messages` when that was
the analyzed source, so the unmasked `texts` slot won downstream and
the original prompt reached the LLM while the response header still
claimed the guardrail applied.
MASK now also overwrites `texts[-1]` whenever `texts` is populated,
keeping both slots consistent.
* fix(docker): bake non_root prisma engines at /opt/prisma so migrations run offline for any uid
The non_root image baked the prisma CLI and engines under /app/.cache and used
the CLI's default (library) engine mode. Prisma stopped baking the library
engine, so `prisma migrate deploy` fell back to downloading it at startup,
which needs network egress and a writable cache. Under an arbitrary non-root
uid (OpenShift restricted-v2), an air-gapped network, or a readOnlyRootFilesystem,
that download fails and the proxy starts on an empty schema while every DB
endpoint returns 500. The migration entrypoint exits 0 on that failure, so a
default-uid `docker run` with network never surfaced it
Bake to /opt/prisma, a fixed world-readable path no cache mount shadows, and
pin PRISMA_CLI_PATH plus PRISMA_CLI_QUERY_ENGINE_TYPE=binary so the baked binary
engine is used directly, matching Dockerfile and Dockerfile.database. A
build-time guard asserts the binary query engine is present, so a future prisma
change that stops baking it fails the image build instead of silently degrading
migrations
Adds docker/test_offline_migration.sh, run from image-scan, which migrates a
fresh Postgres with no egress as a non-root uid and asserts the schema was
created, the case a default-uid `docker run` with network cannot catch
* test(docker): move the offline migration check into a gated pytest and stop pinning XDG_CACHE_HOME at the read-only bake
The offline migration check lived in docker/ as a shell script. It now lives in
tests/proxy_migration_tests/ as a pytest gated on LITELLM_IMAGE, matching the
sibling schema-migration test gated on DATABASE_URL, and image-scan invokes it
with pytest instead of bash. It also asserts the migration entrypoint's exit
code alongside the table count, so a crash or a container-startup failure fails
loudly rather than only surfacing as a low table count
Runtime XDG_CACHE_HOME pointed at /opt/prisma/.cache, which is baked a+rX with
no write, so any XDG-aware library writing a cache at runtime would be denied
for every uid. Leave it unset so it falls back to $HOME/.cache (/app/.cache,
created here and owned by the runtime uid), matching Dockerfile and
Dockerfile.database which never pin XDG at runtime. A second test guards against
a future edit pointing a cache or home var back at the read-only bake
* test(ui): pin search-tools info view behavior before shadcn migration
Rewrite the markup-coupled copy-button assertions in SearchToolView to
role queries plus lucide icon-state, and add a role/text characterization
suite for SearchConnectionTest, which had none. Both are green against the
current antd/Tremor components so they can act as an unedited regression
net across the migration.
* refactor(ui): migrate search-tools info view to shadcn
Port the search-tools detail view and its two helpers off antd and Tremor
onto the installed shadcn primitives plus token utilities:
- SearchToolView (the info page reached by clicking a tool) now uses
ui/button, ui/card and a plain CSS-grid header instead of Tremor
Card/Grid/Title/Text and antd Button
- SearchToolTester swaps antd Input/Button/Spin and Tremor Card/Title for
ui/input, ui/button and UiLoadingSpinner, with no inline styles
- SearchConnectionTest swaps antd Button/Divider/Typography and the inline
keyframe spinner for ui/button, ui/separator and UiLoadingSpinner
Markup only; no behavior change. The list page (SearchTools) and its create
and edit forms stay on antd because they are Form-bearing and blocked until
the forms migration. Icons move from antd and heroicons to lucide. The
retired antd no-restricted-imports suppressions are pruned from the
baseline.
* fix(ui): drop redundant vertical padding in SearchToolTester card
The shadcn Card already applies py-6 and gap-6 to its flex children, so
the pt-6/pb-6/mb-6 added during the migration stacked on top of it and
roughly doubled the vertical whitespace. Keep only px-6 (Card has no
horizontal padding) and let the Card own the vertical rhythm, which
restores the original 24px spacing.
* style(ui): format SearchConnectionTest test file with prettier
* fix(organization): persist cleared fields on /organization/update
Clearing an org field (the Metadata box or a TPM/RPM/max_budget limit) via PATCH /organization/update looked like it saved but reverted on refresh; the partial-update merge could not tell a cleared field from an untouched one and dropped every clear
The endpoint now decides SET vs CLEAR vs UNTOUCHED purely from which keys the raw request body carried, via a pure build_organization_update_plan. Budget nulls flow to update_budget (null clears via exclude_unset), metadata is replace-when-sent (written as {} for the non-nullable Json column), and a budget write on an org with no budget_id creates and links a budget row. This removes the exclude_none dump, both "if v is not None" filters, and the additive _update_dictionary merge
Resolves LIT-3664
* feat(organization): add RESTful PATCH /v2/organization/{organization_id}
Adds a v2 organization-update endpoint with a deterministic partial-update contract, and reverts the v1 /organization/update changes so its public behavior stays untouched
On v2 a field present in the request body is written (null/[]/{} clears, a value sets) and an omitted field is left untouched; presence is read from model_fields_set. Clearing a TPM/RPM/max_budget limit or the metadata now persists instead of being dropped as if it were never sent. Metadata is replace-when-sent and written as {} when cleared, since the org metadata Json column is non-nullable. Budget nulls flow to update_budget, and an org with no budget row gets one created and linked. The endpoint is hidden from the public Swagger docs via include_in_schema=False, and stays typed in the generated dashboard schema
Resolves LIT-3664
* test(organization): cover v2 auth guard, negative budget, and object_permission
Adds v2 endpoint tests that were missing: the real _verify_org_access path rejects a non-admin caller with 403 and writes nothing, a negative max_budget is rejected with 400 before any DB access, and a sent object_permission is passed to the upsert helper with its id linked onto the org write
Refs LIT-3664
* fix(organization): 400 on null-clear of required org fields; drop dead budget upsert
organization_alias and models are non-nullable columns, so a v2 request clearing them with null hit a 500 (NOT NULL violation) and could partially apply the budget half of the request first; the endpoint now returns a 400 with a clear message. Also removes the unreachable "create a budget when the org has none" branch from _apply_organization_budget_updates, since budget_id is a non-nullable FK and every org already has one, so the endpoint no longer needs to link a newly-created budget id
Refs LIT-3664
* fix(organization): let v2 clear object permissions when sent as null
Sending object_permission: null now detaches the org's permission by setting the nullable object_permission_id to null, instead of being a silent no-op, so the endpoint honors its documented "null clears" contract and an admin can actually revoke vector-store/MCP access. Sending a value still merges as before
Refs LIT-3664
* fix(organization): make v2 PATCH atomic, strict, and 422-consistent
Tighten the PATCH /v2/organization/{id} endpoint against standard HTTP
PATCH (RFC 5789 / RFC 7396 JSON Merge Patch) semantics:
- Apply the budget-row and org-row writes in one prisma transaction so a
failure between them can no longer half-apply the patch (RFC 5789 requires
a PATCH to apply atomically). The budget write is inlined as a tx-aware
call mirroring the team-member budget path rather than the standalone
update_budget route handler
- Set extra="forbid" on OrganizationUpdateRequestV2 so an unknown or
misspelled key is a 422 instead of a silently dropped no-op; the contract
is presence-driven, so swallowing unknown keys is unsafe
- Return 422 (not 400) for the hand-rolled field validations (negative
budgets, null-clear of required organization_alias/models, invalid
model_max_budget) so every validation failure matches the 422 that
pydantic already returns for bad values
- Document the per-field clear tokens accurately: null clears budget limits
and metadata, [] clears models, and organization_alias cannot be cleared
Tests cover the single-transaction write path, unknown-field rejection, the
422 status changes, and the budget_reset_at recompute.
* fix(organization): reject empty object_permission on v2 PATCH instead of silently keeping grants
object_permission is a nested merge field on PATCH /v2/organization/{id}: a
sent object merges into the existing permission row (updating one grant list
without touching the others), and null detaches it. An empty {} therefore
merged nothing and left every existing vector-store/MCP grant in place, so an
admin who sent {"object_permission": {}} to strip access silently kept it.
Reject a present-but-empty object_permission with a 422 that points the caller
at null, mirroring how the endpoint already rejects a null clear of the
required organization_alias/models. This keeps merge semantics for non-empty
payloads and does not affect the Admin UI, which only ever sends a fully
populated object or omits the field.
* fix(organization): JSON-serialize model_max_budget on the v2 budget write
model_max_budget is a Json column on the budget table. Route the budget-row
write through jsonify_object so a dict value is serialized the same way
new_budget and the org-row metadata write already do it, keeping every Json
column on this endpoint written consistently.
Raw dicts already round-trip (update_budget writes them unserialized), so this
is not a correctness fix so much as making the one Json column on the budget
path follow the same serialization as the rest of the file. Added a test that
a patched model_max_budget reaches the budget write JSON-serialized.
* refactor(organization): trim v2 docstrings and consolidate planner tests
Trim the verbose docstrings on the v2 endpoint, request model, and the two
pure helpers to the essential contract, and drop a stale line that still
referenced update_budget's exclude_unset (the budget write is inlined now).
Collapse the nine per-case planner tests into one parametrized test asserting
exact budget/org split per body, and fold the two model-validation rejection
cases into one parametrized test. Same 36 test cases run; the planner
assertions get stronger (exact-equality instead of presence/absence) and the
test additions shrink by ~85 lines.
* refactor(organization): inline the v2 update planner into the endpoint
Fold the OrganizationUpdatePlan dataclass and build_organization_update_plan
helper into update_organization_v2. The budget-vs-org split is a few dict
comprehensions built in one shot, so the extra type plus builder was more
ceremony than the job needed. Drops the now-unused dataclass/AbstractSet
imports and the isolated planner unit tests; the split is exercised end-to-end
by the endpoint tests.
* fix(organization): run v2 object permission upsert inside the update transaction
prepare_object_permission_upsert splits the shared helper's read-and-merge
step from its write so the v2 endpoint can upsert the permission row on the
same prisma transaction as the budget and org writes. Previously the upsert
ran before the transaction, so a rolled-back org write left merged grants
live on the permission row the org still pointed at. The upsert record now
pins object_permission_id, since the column's @default(uuid()) would
otherwise mint a fresh-create id different from the one linked on the org.
v1 and the team/key callers of handle_update_object_permission_common keep
their existing behavior
* fix(lint): keep the v2 org PR within the strict-rule budget
The strict gate flagged the PR's new code after the base merge: 11 UP045
Optional fields and a typing.List on OrganizationUpdateRequestV2, Dict
annotations in the new upsert helper and the TypeAdapter, and a B008 from
the v2 endpoint's Depends default. The model and helper now use pipe
unions and builtin generics, and the endpoint takes its auth dependency
via Annotated, which avoids the call-in-default pattern B008 targets
* fix(routes): expose /v2/organization on the backend component allowlist
The component-split coverage test requires every app route on a component;
the new v2 org PATCH belongs with the other management endpoints on the
backend, alongside the existing /v2/key and /v2/team prefixes
* fix(organization): clear budget_reset_at when budget_duration is cleared via v2 PATCH
* feat(ui): give each Models + Endpoints tab its own path
* refactor(ui): decompose Models + Endpoints into per-tab pages with URL-driven detail
Dissolve the 488-line ModelsAndEndpointsView monolith into one page per tab
under the models-and-endpoints route, with a persistent layout owning the
header, cost banner, tab bar and refresh. Each tab page owns only its own
state; shared lists come from a small useModelDashboardData hook.
Replace the stateful model/team drill-in (setSelectedModelId/setSelectedTeamId
full-page takeover) with real URL navigation: ?model=<id> and ?team=<id> render
ModelInfoView/TeamInfoView from the layout, so a model or team detail view is
now shareable, bookmarkable and back-button friendly. Removes the empty
placeholder pages from the first commit.
Swap the tab bar off phased-out tremor onto antd Tabs.
* fix(ui): render model tab panels standalone instead of Tremor TabPanel
AllModelsTab, ModelRetrySettingsTab and PriceDataManagementTab rooted their
render in a Tremor <TabPanel>, which only renders inside a Tremor <TabGroup>.
After the decomposition these panels live under antd Tabs / as route pages with
no such ancestor, so All Models (and the other two) rendered blank. Root them in
a plain container instead.
The existing component tests mocked @tremor/react (stubbing TabPanel to render
children), which hid this; add a regression test that renders with real Tremor
and asserts the content is visible standalone.
Also type visibleSlugs/TAB_LABELS with the canonical ModelTabSlug so a tab added
without a matching label is a compile error.
* fix(ui): make model/team drill-in navigation work under the /ui static mount
The drill-in close (Back to Models) and open were no-ops: the dashboard is a
static export served under /ui, a prefix the Next router (basePath "") does not
know, so a router.push to the current pathname with only the query changed is
deduped and never re-renders. Drive the ?model=/?team= overlay via real browser
navigation (window.location) so open and close reliably work; verified live.
Also address review feedback: gate the tab-permission redirect on teams/uiSettings
having loaded so a team admin hard-loading /add is not bounced to the base before
their membership resolves, and memoize getProviderFromModel on modelCostMapData so
the health tab's provider labels refresh when the cost map loads.
* fix(ui): use window.location.replace for the tab-permission redirect
router.replace is unreliable under the /ui static mount (same class of Next-router
issue that broke the drill-in back button), so the forbidden-tab redirect could
fail to fire. Use window.location.replace, which keeps the no-history semantics of
a permission redirect and is deterministic. Redirect stays gated on teams/uiSettings
having loaded.
* fix(ui): drive model/team drill-in with history.pushState for client-side nav
Switch the ?model=/?team= overlay navigation from window.location.assign to
window.history.pushState, which Next's App Router observes. This keeps navigation
client-side (no full page reload, React Query cache preserved) while still working
for the same-path query-only change that router.push cannot do under the /ui static
mount. Open, close (Back to Models) and browser Back are all verified in the built
UI. Adds unit coverage for the open/close/read behavior.
Extract the inline sync streaming post/decode into make_sync_call so it can be
exercised with an injected client, mirroring make_async_call, and add a sync
regression test that each token is forwarded after exactly one pulled frame.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Mirror the sagemaker_chat fix on the native sagemaker/ streaming path: the
sync and async completion handlers read the invocations-response-stream body
with iter_bytes(chunk_size=1024) / aiter_bytes(chunk_size=1024), so httpx
withholds bytes until 1024 accumulate and tokens arrive in gap-then-burst
waves. Drop the fixed chunk size so each decoded event is forwarded as its
bytes arrive.
Also add a boundary-agnostic decoder test proving frames reassemble correctly
regardless of where transport reads split the stream.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): include type in tool_choice disable_parallel_tool_use config for Converse
* fix(bedrock): let parallel_tool_calls-derived disable flag win over raw tool_choice value
* chore(ui): add filename, size, JSX-handler, prefer-const, and antd lint rules
Wires up five error-level ESLint rules on the dashboard, grandfathering every
current offender into eslint-suppressions.json so the gate only bites new code
and ratchets down as files are fixed
- local/filename-pascal-case: new local rule requiring PascalCase .tsx names,
exempting Next.js reserved files (page, layout, route, ...) and test/spec files
(239 grandfathered)
- max-lines: 800 lines over src/**, excluding tests, src/data, and generated
schema.d.ts (20 grandfathered)
- local/no-complex-jsx-arrow: new local rule flagging inline JSX arrow handlers
with block bodies over two statements; each failure is a small extract-to-named
-handler refactor (65 grandfathered)
- prefer-const: flipped from off to error (103 grandfathered)
- no-restricted-imports: added antd to the phase-out ban alongside tremor, and
pointed both messages at shadcn/ui primitives (405 antd import sites grandfathered)
Both new local rules ship with RuleTester coverage
* fix(ui): preserve secondary extensions in filename-pascal-case suggestion
The suggestion text built the rename from only the head segment, so a
multi-dot file like my-component.utils.tsx was told to become
MyComponent.tsx instead of MyComponent.utils.tsx. Rebuild it from the
PascalCased head plus the untouched remaining segments, and add tests
covering multi-dot filenames and the hyphenated Next.js reserved names
(global-error, apple-icon, opengraph-image, twitter-image)
The UI vitest suite is CPU-bound; move it to a 16-core larger runner and raise vitest fork concurrency from 4 to 14 (leaving headroom for the coordinator, jsdom, and the OS) so the full suite and PR-scoped runs finish faster.
The huggingface embedding test fixture reloaded
litellm.llms.custom_httpx.http_handler, creating a new HTTPHandler class
object. llm_http_handler keeps the class captured at import time, so any
test running later in the same process that injects a client built from
the reloaded class fails the isinstance check and the mock is silently
discarded, causing a real network call. Under pytest-xdist loadscope this
surfaced as a deterministic failure of
test_accept_header_in_completion_request_jwt whenever an unrelated PR
shifted worker distribution.
Also removes the same reload pattern from the vertex rerank integration
test (both were previously removed in a6df01caec and resurrected by a
merge conflict resolution) and hardens the agentcore victim test by
dropping the bare except that swallowed the real error
Resolves LIT-4581
A true_passthrough MCP server created without the at-creation auth step
has no stored client_id, and the tools-page browser flow supplies none,
so GET /v1/mcp/server/oauth/{id}/authorize dead-ended on a 400
missing_client_id. The client-forwarded-token modes forbid the gateway
from persisting an OAuth client, so client acquisition moves into the one
chokepoint every caller crosses: the authorize endpoint.
resolve_ephemeral_dcr_client owns the whole mint policy (mode gate,
authorization-url precondition, required S256 PKCE, redirect trust, then
a TTL-deduped, per-server single-flighted RFC 7591 mint). The minted
client rides the encrypted OAuth state; /callback seals it with the
upstream code and server_id into an llm_ptcode_ gateway code, and
redeem_passthrough_authorization_code recovers it at the token endpoint
(server binding plus required code_verifier) to authenticate the upstream
exchange. Nothing is persisted; every value rides the encrypted blobs, so
it works across replicas.
Client acquisition is one predicate applied across the whole auth-mode
matrix: the gateway mints for a clientless authorize iff true_passthrough
(any dcr_bridge) or oauth_delegate-and-not-dcr_bridge, and the UI
gatewayMintsClientFor mirrors that set exactly so the browser pre-registers
a client through the dcr_bridge front door only for the cells the gateway
does not mint (the interactive oauth_delegate dcr_bridge sign-in and the
legacy oauth2 passthrough). A minted flow runs the bridge short-circuit
arm; the relay front door stays for external clients that present their
own client_id. Both sides are pinned against the same truth table
(test_resolve_ephemeral_dcr_client_mint_set_is_exact and the
gatewayMintsClientFor matrix test) so no mode can silently diverge. The
authorization_code hook and M2M/token-exchange modes are unchanged.
* test(e2e): cover key max_budget blocks on personal, team, and team-member keys
* refactor(e2e): convert budget enforcement cases to the resources-fixture pattern
The E2ECase class pattern existed only in this file; every other suite uses
plain pytest tests with the resources fixture. Rewrites the nine cases as two
spec classes and removes the now-dead E2ECase protocol and run_case driver
from lifecycle.py
Moves the dashboard's next pin from 16.2.6 to the latest 16.2.x patch and bumps eslint-config-next to match. Regenerating the lock also healed in explicit bundled-dependency records under @tailwindcss/oxide-wasm32-wasi