missingReferencedModels read complexityRouterConfig.classifier_llm_config and embeddingModel unconditionally, but buildComplexityRouterConfig only includes classifier_llm_config when classifierType is "llm" and embedding_model when semanticMatchingEnabled is on. A dormant selection left over from a toggle no longer in effect (embeddingModel still set after turning semantic matching off, or a classifier_llm_config seeded with model: "" before a caller picks one) was being checked and reported as a missing model that would never actually be submitted, wrongly disabling the button.
Gate both fields on the same conditions buildComplexityRouterConfig itself uses. Also hardened getRequiredModels to filter out empty-string models, not just null/undefined, so a not-yet-chosen classifier model can't be misread as a real reference even when its field is legitimately included.
Deleting the preset-specific submit-time re-check removed the one place that verified availability at all, leaving a real gap: a caller whose access narrows after a model already entered a tier (via a preset or picked by hand) had nothing catching a now-unavailable model before creation, and the backend does not validate this either.
Generalized autorouter_presets.ts's getRequiredModelsInPreset/getMissingModelsInPreset into config-shaped getRequiredModels/getMissingModels (the preset-specific versions are now thin wrappers), so the same accessor works on a preset's bundled config or on complexityRouterConfig as actually built. submitBlockedReason now also checks the config's referenced models against availableModelSet, which is reactively kept current by the existing useQuery, no new fetch or async gap involved. This checks the config that will actually be submitted regardless of whether it arrived via a preset or Custom, closing the gap Bugbot's first finding pointed at directly ("verifies preset, not config").
A preset was being treated as a persistent identity that had to be re-verified against its own bundled model list at submit time (verifyPresetStillAvailable), separately from the tiers a caller actually built. That's wrong: handlePresetChange only ever prefills complexityRouterConfig once, and everything after that is edited exactly like Custom. Tier edits after applying a preset left selectedPreset unchanged, so the submit-time check was verifying the preset's original models, not the config actually being submitted; it could block a valid customized config or silently accept a manually-added model it never checked.
Deleted verifyPresetStillAvailable and the whole submit-time preset-recheck (with it, the async loading gap that made every round of the token-consistency findings possible in the first place - there is no longer an async step between the button click and form.validateFields). The tier selects already only ever offer models from modelInfo, so submitBlockedReason (tiers + keyword rules) is a complete, accurate check regardless of whether the config came from a preset or from scratch.
Also dropped the requirement to explicitly choose a template before submitting: a caller who fills in tiers manually without ever touching the Template dropdown is functionally identical to one who clicked "Custom Configuration" first, so gating on that distinction was friction with no safety benefit. Removed the decorative required asterisk and the inline "select a template" error along with it.
Fixed a latent test-order bug surfaced by this cleanup: several tests override getMissingTiersError with mockReturnValue(null), which vi.clearAllMocks() does not undo (it clears call history, not the implementation), so the override was leaking into whichever test ran next. beforeEach now restores the real implementation explicitly.
Net -198 LOC across the two files.
Staging landed #35705 (LIT-5133) in the same window, which added a declarative submitBlockedReason (missing tiers or an empty keyword rule) that disables the Add Auto Router button with a Tooltip explaining why. That touched the same button and validation area as this branch's async preset-availability check.
Merged both: the button is disabled by submitBlockedReason (synchronous, tier/keyword completeness) and additionally shows a loading state while verifyPresetStillAvailable runs (asynchronous, preset-model-availability). One of staging's new keyword-rule tests needed the same "select Custom Configuration first" fix already applied to the team and session-affinity tests, since it never touches the template selector and collided with the required-template guard.
Reading accessTokenRef.current independently in verifyPresetStillAvailable and again at the create call still left a residual gap: two separate reads separated by an await can observe two different tokens if one rotates in between, so verification and creation could still end up representing different callers. Chasing "freshest at every read" can never fully close that, since there's always a window between any two reads.
The actual invariant needed is internal consistency for one submission, not maximum freshness at each step: capture accessToken once, via the plain closure that already exists, and thread that same value through both the verification fetch and the create call. Deletes the ref and its syncing effect entirely; a plain prop closure already guarantees two reads within the same function invocation return the same value, no ref needed.
"Add keyword rule" seeds a row with no keywords, and the only check that a
rule carried one lived inside getSemanticConfigError, which returns early
when semantic keyword matching is off. Off is the default, so an unfilled
row fell through to serializeKeywordTierRules and was discarded on the way
to the payload; the create reported success and the rule was gone.
The row now reports the gap itself and the submit is withheld while one is
outstanding, on the create form and the edit modal alike, both reading
emptyKeywordTierRuleIndexes so the row named and the row marked cannot
differ. Enter commits a typed keyword: the dropdown is kept closed, which
left antd nothing for Enter to select, and submitting was what used to
supply the blur that saved the word.
The backend already refused such a rule, but only when the router built the
deployment, so a caller that sent one anyway got the row written, dropped on
reload, and a 500. The management write paths now parse the incoming
complexity_router_config with the router's own ComplexityRouterConfig, judged
on the config alone so a patch that writes one without naming a model is
covered too, and reject it with a 400 having persisted nothing.
* fix(bedrock): stop forwarding no-op toolSpec.strict to Converse
`strict: false` is the Chat Completions default, so sending it to Bedrock
Converse communicates nothing the provider does not already assume, while
Bedrock rejects the key by presence rather than by value: any Claude model
routed through its Anthropic-compatible validator 400s with
`tools.0.custom.strict: Extra inputs are not permitted`.
The existing `bedrock_converse_supports_strict_tools` gate only protects
models whose `model_prices_and_context_window.json` entry carries the flag,
which makes every newly released Claude model broken by default until someone
adds it. That is a losing race for a field that carries no information when
false, and it is unrecoverable from the client side on `/v1/responses`, where
the Responses to Chat Completions bridge stamps `strict: false` onto every
function tool even when the caller never sent one. `drop_params` cannot help
there because the caller never supplied the param.
Drop the key when falsy instead. `strict: true` still honors the per-model
gate, so models that accept strict schemas keep the behavior they have today
and the flag keeps doing its job for the values that actually mean something.
* fix(bedrock): flag Claude Sonnet 5 as rejecting toolSpec.strict
Bedrock routes Sonnet 5 through the Anthropic-compatible validator that
rejects `toolSpec.strict`, but its six pricing-map entries never got
`bedrock_converse_supports_strict_tools: false`, so the gate fell back to
forwarding for Anthropic models and every tool call carrying `strict: true`
400'd. Verified live in us-east-1: before this, `strict: true` against
`us.anthropic.claude-sonnet-5` returns
`tools.0.custom.strict: Extra inputs are not permitted`; after, it returns a
real tool call.
Measured the rest of the family the same way rather than trusting the map:
Sonnet 4.5, Sonnet 4.6 and Haiku 4.5 all accept `strict: true`, and Opus 4.8
already carries the flag. Sonnet 5 was the only entry where the map disagreed
with the provider, so it is the only one changed here.
Same shape as the Opus 4.7/4.8 and Sonnet 4 fixes before it.
The ref fix made router creation use whichever token is current when the create call fires, but the preceding availability check still went through the query's own refetch, which stays bound to whatever token was current when that render's useQuery was set up. If the token rotated in between, that split verification from a stale caller's model list against creation under a different one, meaning the caller actually creating the router never had its own access checked. Fetch directly against accessTokenRef in the verification step too, so both calls agree on the same live identity instead of each independently chasing "freshest."
submitRecommendedRouter awaits a network round trip (the fresh preset re-check) before calling handleAddAutoRouterSubmit, so the accessToken it closed over at click time could be stale by the time that call fires if the token rotates during the wait. Read it through a ref kept in sync via useEffect instead, so the create call always uses whatever token is current when it actually runs.
* ci(circleci): install a pinned Rust toolchain on the Linux jobs
The cimg/python images have no Rust toolchain, so every Linux job that
runs `uv sync` or `uv build` builds litellm-rust through maturin with no
cargo on PATH. maturin's puccinialin helper then fetches rustup-init from
the unversioned /rustup/dist/ path with no checksum and provisions a
floating `stable` toolchain, so the compiler a job builds with drifts
with whatever upstream published that day. uv hides build-backend output
on a successful sync, so none of this shows up in the job log.
Add an install_rust command that mirrors the Windows job: download a
pinned rustup 1.28.2, verify its SHA-256 against rust-lang's published
sidecar, install toolchain 1.97.1 with the minimal profile, and export
~/.cargo/bin through BASH_ENV. Run it after install_uv in every job that
builds the workspace; upload-coverage only runs `uv tool run coverage`
and is left alone.
Net download cost is unchanged, since puccinialin was already pulling a
rustup and a toolchain in each of these jobs.
* test(ci): guard that no CircleCI job builds the workspace without a pinned Rust
A green CI run does not notice the gap this closes: uv suppresses
build-backend output on a successful sync, so a job that syncs with no
cargo on PATH silently gets maturin's own unpinned rustup and a floating
toolchain, and the log looks identical either way.
Pin the invariant statically instead. Every job and reusable command is
walked in step order, and reaching a `uv sync` / `uv build` without a
Rust toolchain provisioned first is a failure. install_rust and the
Windows job's inline pinned install both satisfy it, so a new job that
forgets one is named in the assertion message at PR time. Separate cases
cover install_rust's own pins: a versioned /rustup/archive/ URL, a
SHA-256 verified before the installer is executed, and an exact
toolchain version rather than a channel name.
* ci(circleci): provision Rust for base_sdk_install
base_sdk_install landed on staging while this branch was open. It runs
`uv build --wheel` on cimg/python:3.12 behind install_uv alone, so it
built the bridge with maturin's own unpinned rustup. The guardrail added
here caught it on the merge result, which is the case it exists for.
* feat(team): custom metadata validation hook for team create and update
Operators can point general_settings.custom_team_metadata_validate at an
async Python function that validates team metadata before /team/new,
POST /team/update, and PATCH /team/{team_id} commit their writes. The
hook receives the metadata that will actually be written (the merged
result on PATCH) plus the stored metadata and requester context, and
fails closed: a rejected value returns the function's own message as a
400 while any exception or timeout blocks the write with a configurable
generic message as a 503. Premium-gated like enforced_params.
* fix(team): validate metadata before model alias writes and strip system keys from validator input
Review follow-ups on the team metadata validation hook: run the validator
before the model_aliases table insert so a rejected create leaves no
orphaned model rows, strip system-managed keys from existing_metadata so
the validator sees symmetric input on both fields, and accept class
instances exposing an async __call__ as validators. Adds a three-way
validator implementation matrix (allowlist function, HTTP-service-backed
function, immutability-enforcing class instance) driven through the real
create, update, and patch endpoints, including an HTTP stub service and
outage coverage.
* test(team): run the metadata validation matrix against the DB-backed proxy in CI
Adds the validator matrix to the proxy_store_model_in_db_tests CircleCI
job so every scenario runs full e2e against a Postgres-backed proxy. The
proxy config registers a dispatching validator that routes each request
to one of the three implementations via a metadata key and accepts
anything that does not opt in, keeping the rest of the suite unaffected.
CI starts a stand-in cost center service on the host for the HTTP-backed
implementation, reached from the container via host.docker.internal, and
the outage path targets a closed port to prove the fail-closed 503
without stopping services.
* feat(ui): edit team metadata as key-value pairs in team create and edit forms
The team create and edit forms asked for metadata as a raw JSON blob in a
textarea buried under Additional Settings. Both forms now render a key-value
pair editor directly under the TPM/RPM limit fields, backed by a shared
MetadataKeyValueFields component. Values round-trip losslessly: non-string
values display as JSON and parse back to their typed form on save, and
JSON-ambiguous strings are quoted so their type survives the trip. The edit
form hides UI-managed keys (logging, guardrails, model rate limits, etc.)
that dedicated controls already own and re-add on save.
* fix(ui): explain typed JSON parsing in the team metadata help text
* feat(team): schema-driven metadata fields from team_metadata_schema config
* refactor(team): render schema metadata fields as locked key-value rows, drop allowed_values
* refactor(team): schema fields reduce to key and label, tag-rendered keys, clean rejection toasts
* refactor(ui): prepopulate declared metadata keys as ordinary key-value rows
* fix(team): let non-admin dashboard users read the team metadata schema
* test(proxy): pin timeout wiring, boundary, and error-message contracts for team metadata validation
* fix(proxy): use pooled async httpx client in the e2e team metadata validator example
* refactor(team): satisfy staging lint ratchets inherited by the merge
The backend never validates an auto-router's referenced model names against the caller's access at creation time (POST /model/new only checks whether the caller can create a model at all, not whether they can use the specific models a complexity_router's tiers name). The frontend's cached availability check is the only thing that catches this, and it can go stale: nothing invalidates an already-selected preset, and a failed background refetch deliberately keeps trusting the cache (by design, so a passive hiccup doesn't wrongly block a still-valid preset). That combination meant a caller whose access narrowed at exactly the wrong moment could still create a router referencing models they no longer have.
Extracted the fix into verifyPresetStillAvailable: a small, named async check that forces a fresh fetch right before creating the router, independent of whatever's cached, and only for the preset path (Custom never claimed this guarantee). Doing this invisibly inside the click handler would leave the button looking unresponsive for a real network round trip, so it's surfaced the same way this file already surfaces async submit work: a loading state on the button itself, matching the existing Test Connection convention.
Extracting the check into its own function also pulled submitRecommendedRouter back under the complexity budget threshold it had just crossed.
* fix(ui): hide guardrail review buttons from non-admin users
The team guardrail submissions list rendered Approve/Reject buttons for
non-admin users even though the backend correctly rejected the calls.
Thread userRole from the page through GuardrailsPanel into
TeamGuardrailsTab and gate the row-card and detail-panel review buttons
on isAdmin so the UI matches the backend authorization.
Defense in depth only — the backend remains the source of truth and is
double-gated at both the route admin check and the explicit endpoint
role check.
Refs LIT-2494
* refactor(ui): read userRole from useAuthorized hook instead of prop drilling
Drop the userRole prop chain through GuardrailsPage → GuardrailsPanel →
TeamGuardrailsTab. Each component reads userRole directly from the
useAuthorized hook, matching the pattern used elsewhere in the dashboard.
Tests now mock useAuthorized per case (the same pattern as
top_key_view.test.tsx) instead of passing userRole as a prop.
Refs LIT-2494
* fix(ui): drop userRole prop on GuardrailsPanel call site in src/app/page.tsx
Missed in the earlier refactor — GuardrailsPanel no longer accepts
userRole as a prop (reads from useAuthorized hook), so callers must
not pass it. The build was failing in production type-check.
Refs LIT-2494
* fix(ui): gate guardrail forward-key toggle and header editors on proxy admin
* refactor(ui): remove dead app_admin case from user role formatting
handlePresetChange builds a fresh ComplexityRouterConfigValue from the preset's config on every selection; four fields added to that payload since the object literal was written (session_affinity, classifier_context_window_size, classifier_context_per_turn_chars, classifier_context_include_assistant_turns) were never added to it. Since setComplexityRouterConfig replaces the whole value rather than merging, applying a preset silently dropped all four to undefined regardless of what the preset specified or what the caller had set manually; for session_affinity specifically that resolved to false at submit time via the existing destructuring default, so both bundled presets (already false) masked it. Verified directly: toggling session affinity on and then applying a preset reverted the switch to off and submitted false, not the caller's prior choice.
Thread all four fields through from the preset's config, matching how every other optional field here is already carried over. Extended the existing falsy-preset test to assert session_affinity survives when a preset sets it true, and added a dedicated regression for the reported flow: toggle on, then apply a preset, and confirm the preset's own value wins over the stale manual edit rather than being silently discarded.
react-query keeps the last successful list around when a later refetch fails: data stays populated, but isError flips true. presetAvailability read isError alone, so once a caller had a good cached list, any subsequent refetch hiccup (window refocus, a manual retry that itself fails, etc.) made every preset unverifiable again, wrongly blocked an already-selected preset at the new submit-time check, and showed a "models are no longer available" toast for models that were, per the cache, still there. Only treat the state as unverifiable when there has never been a successful fetch (data is still undefined); otherwise keep trusting the cached list, matching how react-query itself treats stale-but-valid data.
handlePresetChange only ever applies a preset that was verified available at selection time, but that guarantee could go stale by submit time if the caller's model access narrowed in between (a token change re-keys the model query without clearing the selection, since clearing it would erase in-progress Custom edits too). Re-run the same presetAvailability check at the submit boundary instead of trusting state gathered earlier, so a stale preset can no longer create a router referencing models the current caller doesn't have.
Rebased onto litellm_internal_staging, which landed session_affinity as a required ComplexityRouterConfigPayload field (#35714) after this branch forked. Add it to both bundled presets so the type cast in autorouter_presets.ts holds, and drive the two new session-affinity tests through the Custom template path so they clear the template-required check added earlier in this PR.
* feat(guardrails/rubrik): prompt moderation, response-text blocking, streaming buffer, failure logging (#34019)
* feat(guardrails/rubrik): add prompt moderation, response-text blocking, streaming buffer, failure logging
- Add `pre_call` prompt moderation via `/v1/before_prompt/openai/v1` webhook:
structured messages are flattened and sent before the LLM is called; blocked
prompts surface a `ModifyResponseException` with the refusal text.
- Extend `post_call` response moderation to cover assistant text in addition to
tool calls; text blocks (wholesale replacement) are distinguished from
tool-block explanations (appended) via `startswith` diffing.
- Add `streaming_end_of_stream_only = True` and `streaming_buffer_until_moderated = True`
so streamed responses are withheld until end-of-stream moderation passes
(requires litellm >= BerriAI/litellm#31389; older versions fall back to
detect-only).
- Add `_MalformedToolBlockingResponseError` for structurally invalid service
responses; `_guarded` logs at CRITICAL so operators notice misconfiguration.
- Add `max_queue_size = 10_000`, `_enforce_max_queue_size`, and drop-oldest
backpressure so a webhook outage cannot grow the retry queue unboundedly.
- Add `flush_queue` override that snapshots once for both send and drain,
preventing duplicate delivery on concurrent flush calls.
- Make `_log_batch_to_rubrik` re-raise on error so `flush_queue` preserves
undelivered events for the next retry.
- Add `async_post_call_failure_hook` to log blocked requests
(`ModifyResponseException`) with a best-effort fallback payload for prompt
blocks (where no `standard_logging_object` exists yet).
- Add `_correlation_id` / `_apply_correlation_id` / `_prepend_system_prompt`
helpers; `_prepare_log_payload` now applies them for all providers (not just
Anthropic) so every log correlates by `litellm_call_id`.
- Add `get_supported_event_hooks` classmethod advertising `[pre_call, post_call]`.
- Use dedicated `httpx.AsyncClient` (`moderation_client`) for webhook calls
with explicit pool limits, separate from the shared logging client.
- Drop module-level `rubrik_handler` singleton (inappropriate for a library).
- Update `initialize_guardrail` docstring to explain `pre_call` vs `post_call` mode.
- Update tests: rename `tool_blocking_client` → `moderation_client`,
`tool_blocking_endpoint` → `response_moderation_endpoint`, `_flush_task` →
`_periodic_flush_task`; migrate `TestExtractBlockedTools` to
`TestExtractResponseBlock` for the new combined text+tool block API; add
tests for prompt moderation, text blocking, streaming flags, and failure
payload construction.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* test(guardrails/rubrik): add tests to reach 100% coverage
50 new tests across 18 classes covering previously-untested paths:
- Prompt moderation: passthrough, block, no-messages skip, message
flattening (content-list → string), payload construction with
tools/user/correlation_key/litellm_call_id fallback, refusal extraction
- async_post_call_failure_hook: non-matching exception no-op, missing
stash warning, valid stash → enqueue, AttributeError in payload build,
flush exception handling
- Block payload building: standard_logging_object present vs fallback
path, missing start_time
- async_log_success_event: _rubrik_blocked=True skip path
- aclose: task cancel + moderation_client.aclose()
- Edge cases: sampling rate clamp warning, unknown input_type passthrough,
empty-inputs early return, model_call_details warning, _stash_block_context,
duck-typed tool-call normalization, request_data["tools"] preference over
optional_params, system-prompt exception handler, flush-at-batch-size,
enqueue exception swallowing, queue empty/lock-None guards, non-dict JSON
response TypeError
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): use get_async_httpx_client, ruff format
- Replace bare httpx.AsyncClient with get_async_httpx_client (required
by ensure_async_clients_test; avoids per-request client creation)
- aclose() calls close() (AsyncHTTPHandler interface, not aclose())
- ruff format on rubrik.py and guardrail_hooks/rubrik/__init__.py
- Update 3 tests for AsyncHTTPHandler type (isinstance check, close())
osv-scan and documentation CI failures are pre-existing on the base
branch and unrelated to this PR.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): fix UP006 strict ruff violation
get_supported_event_hooks return type used List[...] (UP006) instead of
list[...]. Replace with the built-in generic and remove the now-unused
List import from typing.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): fix 3 reportArgumentType basedpyright violations
Use `# pyright: ignore[reportArgumentType]` (not `# type: ignore`) to
suppress the three errors basedpyright reports in --outputjson mode:
- convert_content_list_to_str call (dict vs AllMessageValues)
- _apply_correlation_id call (StandardLoggingPayload vs dict[str, Any])
- _prepend_system_prompt call (same)
Also tighten _apply_correlation_id and _prepend_system_prompt signatures
from bare `dict` to `dict[str, Any]`.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): don't close shared HTTP client in aclose()
moderation_client and async_httpx_client both come from LiteLLM's global
HTTP-client cache (get_async_httpx_client keys on llm_provider + params).
Two RubrikLogger instances with the same parameters share the same
underlying AsyncHTTPHandler object. Calling close() in aclose() closed
the shared connection pool for all instances, breaking any subsequent
moderation request on other loggers.
aclose() now only cancels the periodic flush task and lets LiteLLM
manage the shared client lifecycle. Tests updated to assert close() is
NOT called.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): use Counter for duplicate tool-call ID detection
Set-based comparison lost ID multiplicity: two original tool calls with
the same ID both appeared "allowed" even when the service returned only
one (e.g. one allowed + one prohibited sharing an ID). Replace with
Counter so returned_id_counts[id] >= required_id_counts[id] must hold
for every ID. Matches the approach in the original _extract_blocked_tools.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): respect default_on=true when omitted from config
LitellmParams.__init__ converts an omitted default_on to False before
initialize_guardrail receives it, so litellm_params.default_on is always
bool and never None. The is-None guard in RubrikLogger.__init__ therefore
never fired on the proxy path, leaving prompt/response moderation inactive
for any config that omitted default_on.
Fix: read the raw guardrail dict (before LitellmParams coercion) to
distinguish an explicit `default_on: false` from the absent-means-True
default. When the key is absent from the raw config, default_on=True is
used; when it is explicitly set (either True or False), that value wins.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* style: ruff format rubrik.py after Counter import addition
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): detect ID-less tool call removal; fix UP045
ID-less tool calls (tc.id is falsy) were excluded from required_id_counts,
so the Counter comparison never caught their removal. Add a cardinality
check (len(returned) < len(original)) that fires on any removal regardless
of ID presence, combined with the Counter check for duplicate-ID attacks.
Also fix 5 UP045 violations (Optional[X] → X | None) introduced by our
new code against the daily-branch baseline.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): filter optional_params through ModelParamHelper in fallback payload
_build_fallback_payload forwarded the raw optional_params dict as
model_parameters. optional_params can contain extra_headers, api_key,
and other upstream provider credentials that must not reach the Rubrik
webhook. The normal standard_logging_object path already filters through
ModelParamHelper.get_standard_logging_model_parameters(), which
allowlists only safe LLM API parameters. Apply the same filter here.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): scope failure hook by guardrail_name; moderate text-completions
Guard async_post_call_failure_hook by guardrail_name so multiple Rubrik
instances don't cross-log: the failure hook is called for every registered
callback; without the check the first instance pops the stash and the
originating instance finds None and silently skips logging. Now each
instance only handles blocks raised by itself.
Also moderate /v1/completions prompts: _moderate_prompt returned early
when structured_messages was absent. For text-completion requests litellm
supplies inputs["texts"] with no structured_messages. Added a fallback
that synthesises a user-message from texts so the before_prompt webhook
can evaluate text-completion prompts.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(lint): add reason comments to pyright: ignore suppressions
type-discipline budget requires each # pyright: ignore[...] to carry an
explanatory comment. Add reasons to the three bare suppressions on lines
483, 651, 652.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): include tool-call arguments in prompt moderation
_flatten_messages_for_moderation only sent the content field, silently
dropping tool_calls[].function.arguments and function_call.arguments.
An attacker could embed prohibited text in tool-call arguments inside
assistant history turns and bypass prompt moderation entirely.
Now collects all attacker-controlled text per message: text content via
convert_content_list_to_str, plus all tool_calls[].function.arguments
and the deprecated function_call.arguments, joined with newlines before
being sent to the before_prompt webhook.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): tighten append detection to prevent prefix bypass
startswith(sent_content) allowed any replacement whose text shares the
original as a prefix (e.g. "Hello" → "Hello, blocked.") to be classified
as a tool-block append rather than a text block, bypassing detection.
Use startswith(f"{sent_content}\n\n") to require the exact two-newline
separator the webhook uses between original text and appended tool-block
explanations. Also add `returned_content != sent_content` to text_blocked
so an unchanged passthrough is never classified as a block.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* fix(guardrails/rubrik): default_on=False when omitted (follow existing pattern)
Remove the custom raw-dict lookup that was defaulting default_on to True
when omitted from the guardrail config. Follow the standard litellm
convention: omitted resolves to False (users must explicitly opt in with
default_on: true).
- initialize_guardrail: pass litellm_params.default_on directly
- RubrikLogger.__init__: is-None guard defaults to False not True
- Test updated to assert the correct False default
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* chore(rubrik): keep the ported guardrail within staging lint budgets
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore: credit the original author of the rubrik guardrail work
Co-authored-by: Joseph Barker <156112794+seph-barker@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore: keep this mirror PR's diff limited to the rubrik files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Joseph Barker <156112794+seph-barker@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A failed fetchAvailableModels call disabled every preset with no way to recover short of closing and reopening the modal, even for a transient error. Surface the failure next to the template selector with a Retry action that re-runs the same useQuery.
The Template selector carried a manual required asterisk with nothing behind it: submitting with no template chosen fell through to the unrelated missing-tiers error instead of naming the actual problem. Add an explicit check ahead of the existing tier/classifier/semantic validation and an inline hint under the selector, reusing the showValidationErrors flag the rest of the form already uses.
sample_spec duplicated anthropic_family's exact model list under a different label, kept only as shape documentation; that duplication could drift silently if the real preset's models changed without a matching edit. Delete it and drop the now-unneeded filter in autorouter_presets.ts.
availableModelSet was rebuilt on every render while presets right above it was already memoized; wrap it in useMemo for consistency and to stop recomputing it on unrelated re-renders.
The model list drives preset availability and is scoped to accessToken. Loading it with a useState + useEffect + ignore-flag loader coupled the fetch lifecycle to manual resets, and each token-change edge (stale list, out-of-order resolution, config erased on reset) was handled by adding another line to that effect. This replaces the whole loader with useQuery keyed on accessToken, matching how EditFallbacks and the rest of the dashboard fetch model data.
react-query owns the race surface: a caller switch is a new query key, so the previous caller's list is never read for the new caller and out-of-order resolutions are discarded by key. There is no reset-on-token-change anymore, so a token change can no longer erase the user's in-progress configuration; only the model list is token-scoped, and preset availability recomputes from it. isLoading and isError replace the hand-rolled load-state union.
Net effect deletes the two model-related useState hooks, the loader effect, and the ignore guard. Tests updated to drive the query mock and assert re-gating on caller switch; mutation-checked (removing accessToken from the query key, skipping the loading gate, and swapping the nullish match_threshold prefill each turn a test red).
The prior fix for token-change races reset ALL configuration when the token changed, erasing user-edited tiers, keywords, semantic settings, and adaptive config alongside the preset selection and model-verification state. This was a regression: only the preset-tied state is token-scoped; user-entered config survives a token change.
Revised: on token change, reset only the preset selection (setSelectedPreset), model load state, and the cached model list. The user's manually-edited complexity-router config, keywords, and adaptive settings persist. This closes the availability race without erasing user input.
Updated test assertions to match the corrected design: "clears the preset selection" and "ignores a stale in-flight fetch".
* fix(proxy): redact credential headers from request logging copies
clean_headers preserves an Anthropic subscription OAuth token, and other
client-supplied provider credentials, so they can be forwarded upstream. The
same dict was also stored as proxy_server_request["headers"] and
metadata["headers"], so those credentials reached every logging callback and
the SpendLogs proxy_server_request column that the Admin UI logs page renders.
Build the observability facing copies through redact_credential_headers, and
drop the transport-only keys (provider_specific_header, headers, api_key) from
the request body snapshot since they have to keep the real values.
* fix(proxy): use the redacted header copy in the request debug log
The stdout secret filter matches Bearer and sk- shaped values, so an MCP auth
token printed by the request-header debug line survived it in cleartext.
* fix(proxy): resolve the configured MCP auth header name through the secret manager
get_secret_str also consults a configured secret manager, so a deployment that
stores the header name there now gets that header masked too. Drops the added
comments in favour of a named constant.
* perf(proxy): resolve the MCP auth header name once per process
get_secret_str issues a blocking secret-manager SDK call when one is configured,
and configured_credential_header_names runs on every proxied request.
* fix(proxy): read the MCP auth header name live, cache only the secret manager
The config reloader rewrites os.environ on an interval and after /config/update,
and MCPRequestHandler resolves the same setting per request, so caching the env
lookup left a renamed header logged in the clear until the process restarted.
Only the blocking secret-manager call stays cached.
* refactor(proxy): narrow header redaction to the reported credential set
Drops the MCP header-name resolution, its per-request config and secret-manager
lookups, and the x-mcp- prefix rule. Those cover a separate credential family
than the one this ticket reports and carried their own config-reload staleness
surface; they belong in their own change.
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The Pretty view only parsed the Chat Completions shape (messages /
choices[0].message), so any spend log storing the Responses API shape
(input / output) rendered an empty Input card and the literal text
"No response data available" even though the row held the full request
and response. This also hit plain /v1/chat/completions callers, because
litellm may route those over the Responses bridge and then store the
upstream Responses-shaped body.
Parsing now branches on a tagged union covering both shapes, which also
replaces the any-typed key sniffing and the role guessing it relied on.
Adds 61 unit tests on the modules #35694 extracted: 46 on the payload
builder, 15 on the OAuth redirect snapshot. They run in 9ms against 240s
for the 77 full-render tests they partly replace. Nine of nine mutants
were killed when the extracted logic was deliberately broken, so the
speed does not come at the cost of signal.
Deletes six cases across four blocks that rendered the whole modal to
assert one payload key belonging to a field they never touched. Every
test that proves a form field reaches the right payload key stays; those
cover field to form value to payload, which a unit test cannot reach.
Replaces "should not render when user is not an admin", which asserted
the admin title was absent and so passed for the wrong reason: the modal
does render for a non-admin, retitled. registerMCPServer was mocked but
never asserted anywhere, leaving the whole non-admin submission path
uncovered. It now drives a real submit and asserts the call lands there
and never on createMCPServer.
Renames the slow file to CreateMCPServer.integration.test.tsx and
documents the three tiers in the dashboard CLAUDE.md. No production code
changes.
Team-scoped DD credentials (dd_api_key, dd_site) set via POST /team/{id}/callback were silently dropped because _request_blocked_callback_params blocks them from standard_callback_dynamic_params. The security block is correct for request-level injection, but team callback_vars are admin-configured and trusted.
Store the raw init kwargs on the Logging instance and read dd_* params from there in _process_dynamic_callback_list instead of from standard_callback_dynamic_params.
Adds an integration test that exercises the full Logging.__init__ flow with team callback_vars to prevent regression.
Co-authored-by: Aanchal Khandelwal <aan2210khandelwal@gmail.com>
* feat(ui): expose an Auto-Router session affinity toggle
session_affinity on ComplexityRouterConfig defaults to True, and neither the
create form nor the edit modal ever emitted the key, so every auto-router built
in the UI silently pinned each session to its first turn's model for an hour
with no way to see or change that.
Adds an "Advanced: Session Affinity" switch to both surfaces, defaulted on to
match the backend field. Both paths now write the key explicitly instead of
falling through to the backend default, so a stored config states what the
router actually does. A stored config with the key absent hydrates as on, since
those routers are running with affinity enabled today; showing them as off would
report the opposite of reality and persist it on the next save.
* feat(complexity_router): default session affinity off and expose it in the UI
session_affinity defaulted to True and the Auto-Router UI never emitted the
key, so every router built there silently pinned each session to whatever model
its first turn classified into for an hour, refreshed on every hit. There was
no way to see that from the UI and no way to change it without hand-editing
config.yaml.
The default flips to False, so every turn is classified on its own merits and
lands on the cheapest adequate tier. Pinning is now opt-in.
The toggle added in the previous commit follows the field: it renders off, and
both the create tab and the edit modal keep writing the key explicitly, so a
stored config states what the router does instead of inheriting a default that
can move under it.
Behavior change for existing routers: those created before this have no
session_affinity key stored, so they pick up the new default and start
reclassifying every turn. That gives up the provider prompt cache the pin was
preserving, and a multi-turn session can now change model between turns. Set
session_affinity: true to keep the old behavior.
Key and team `router_settings.model_group_alias` was accepted, persisted and
echoed back by `/key/info`, but never applied at request time, so the request
ran on the group the caller asked for. `route_request` forwards only the
settings the Router accepts as per-request kwargs, and `model_group_alias` is
not one of them: the Router resolves aliases from its own instance attribute,
which holds the global config map and is shared across requests.
Resolve the alias in the proxy instead, alongside the existing model-alias
rewrites and ahead of the pre-call hooks, so per-model limits and guardrails
key off the group that actually serves the request. Authorize the alias target
before the rewrite; model access was checked against the requested group, so a
key whose alias points at a group it cannot call gets the usual 403 rather than
being quietly served it.
Resolves LIT-4879
Google Cloud has renamed Vertex AI RAG Engine to "RAG Engine" and
Vertex AI Search to "Agent Search" in its console. Users following our
setup instructions hit a naming mismatch when they cross-reference the
GCP console. Keep "Vertex AI" as the primary term (the generic new
names would make our provider UI ambiguous) and surface the new names
as secondary asides only where users leave the UI for the console.
Resolves LIT-3081
PR #35492 was authored before the ruff sweep removed Union from the typing
imports in litellm/llms/openai/common_utils.py, so the merge landed an
annotation referencing Union without an import. The annotation is evaluated at
class-definition time, so importing litellm raises NameError and every test
shard on litellm_internal_staging fails at collection. Rewrites the annotation
(and the same latent one in openai.py) as httpx.Client | httpx.AsyncClient |
None, matching the file's PEP 604 style, so no typing import is needed at all.
Pulls four modules out of the 1398-line create component, which drops to
896 lines. No behavior changes: CreateMCPServer.test.tsx is untouched and
all 77 of its tests pass against the refactored component, which is the
review contract for this PR.
createServerPayload.ts is a pure form-values-to-payload function whose
failures are a tagged union instead of inline notification calls, so the
transformation is reachable without a DOM. createOAuthUiState.ts owns the
snapshot that survives the OAuth authorize redirect, keeping every
presence guard the inline version had. AwsSigV4Fields and
OpenApiByokFields are the two largest JSX blocks, moved verbatim so they
can be diffed as moves.
The create/edit setToken divergence, the mcpLogoImg export, and the
untyped form-values bag are left alone on purpose; each is a behavior or
cross-file change that does not belong in a move.
Pure rename, no behavior change. create_mcp_server.tsx and its test move
to CreateMCPServer, the two importers and one stale e2e comment follow,
and the local/filename-pascal-case suppression drops now that the file
passes the rule on its own.
The rename is scoped to this one component rather than the whole
directory because three PRs are currently open against its snake_case
siblings; the rest can follow once those land.
An evicted client was left for the garbage collector, but every OpenAI/Azure
SDK client is a reference cycle, so nothing freed the client or its pooled TCP
connections until a generational sweep ran. Driving 2000 azure calls through
the official image with no forced collection, live clients and open sockets
climbed from 202 to 1361 while the cache stayed at its 200-entry bound, and RSS
grew 279 MB to 456 MB against a TLS upstream.
Closing on eviction is what caused the earlier 'Cannot send a request, as the
client has been closed' regression, so an evicted client litellm created is now
closed only once a grace window has passed, by which point any request that was
already holding it has finished. A client the caller supplied is never closed,
since litellm does not own its lifecycle.
Resolves LIT-4883
gitpython arrives transitively through mlflow-skinny, which accepts
>=3.1.9,<4, so this is a lock-only move with no pyproject change.
Relocked with `uv lock --upgrade-package gitpython`; gitpython is the only
package whose version changed. `uv sync --all-groups --all-extras` and
tests/test_litellm/integrations/test_mlflow.py pass on the result.
The dashboard pins both packages exactly in `overrides`, so the lockfile
stays on whatever those pins say. Move brace-expansion from 5.0.8 to 5.0.9
and postcss from 8.5.22 to 8.5.23, both upstream patch releases, and
regenerate the lockfile.
`npm ci`, `next build`, and the 5888-test vitest suite all pass on the
updated lockfile.
* feat(teams): apply default organization to new teams from default team settings
Adds organization_id to DefaultTeamSSOParams so proxy admins can pick a
default organization in Default Team Settings. new_team applies it before
org validation whenever a team is created without an explicit
organization_id, so API, Admin UI, SCIM, SSO, and team upsert creations
all inherit it and go through the same existence and org-limit checks.
Explicit organization selections win and existing teams are untouched.
The default is validated at save time (PATCH /update/default_team_settings
returns 400 for an unknown org) and at create time, where a missing org now
surfaces as a clean 400 instead of a 500 by routing OrganizationNotFoundError
into the previously dead org_table None guard.
The Admin UI Default Team Settings tab gets a Default Organization row
backed by the shared OrganizationDropdown.
* fix(teams): validate org limits against final team state including defaults
Applies default_team_params and the legacy max_budget fallback before the
organization validation block, so _check_org_team_limits sees the values the
team will actually be persisted with. Also loads the org's budget table in
the lookup; without include_budget_table every budget comparison in
_check_org_team_limits was skipped because litellm_budget_table was None.
* test(proxy_behavior): pin org team limits as enforced on /team/new
The dead-code pins existed to turn red when include_budget_table went
live; that happened, so the scenarios now assert the 400 rejections plus
within-cap acceptance, and the unknown-org pin asserts the handler's 400
instead of the surfaced 500.
Adds a Stream responses checkbox (default on) to the playground Model
Settings popover. When unchecked, chat completions and responses API
requests are sent with stream: false and the full reply renders at
once. The non-streamed result is replayed through the existing
streaming handlers as synthesized chunks/events so MCP events, vector
store results, usage and response ids behave identically in both
modes. TTFT is suppressed when not streaming; total latency now also
reported for the responses API. The toggle is scoped to the chat and
responses endpoints, persists via sessionStorage, and is isolated from
the simplified Agent Builder chat.
Resolves LIT-3251
* fix(proxy): backfill null user_email on existing users during JWT auth
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): guard mapped-key email backfill and make null update atomic
Resolve Greptile review on the JWT user_email backfill:
- only backfill when the mapped virtual-key owner is the JWT principal, so a
mismatched admin-created mapping cannot write one user's email onto another
- make the best-effort mapped-key enrichment non-fatal so a database outage on
a cached-key request no longer fails otherwise-valid authentication
- persist the backfill with an atomic null-guarded update_many so concurrent
writers cannot overwrite an already-populated email
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep cache coherent when a concurrent backfill wins the null-email update
* fix(proxy): cache DB-persisted email after JWT backfill, not the proposed value
Resolve the Greptile finding that a successful null-guarded backfill could
cache this request's proposed email even if a concurrent ordinary user update
wrote a different email first. The helper now always re-reads the row after the
atomic update and refreshes the cache from the value the database holds, so
cache-hit auth and attribution stay consistent with the persisted record.
Annotate the Prisma and model_copy dict literals to keep the LIT002 budget within its ceiling.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>