create_batch routes a model-embedded or unified file id on the model
bound to that file and ignores the top-level model, so deriving the
provider skip from the top-level model first let a caller point model at
a skip-listed provider while the file routed a rate-limited one, skipping
counter enforcement. Resolve the routing model from the file binding
first, matching the batch endpoint.
The per-model skip matched skip_batch_input_file_rate_limiting_for_models
against the model bound to the input file id. That model comes from
decode_model_from_file_id / the unified file id, both unsigned base64 the
caller fully controls, so a caller could re-encode an accessible provider
file id with a skip-listed model while the JSONL still routes rate-limited
body.model entries and bypass the batch RPM/TPM counters. The models a batch
actually runs are its JSONL body.model entries, which cannot be known without
reading the file, so no caller-influenced model identifier can safely gate a
skip.
Remove the per-model skip entirely. The provider skip stays because the
provider is resolved from trusted deployment credentials and the batch is
constrained to run on that provider; the global disable and
no-applicable-limits skips stay because they do not depend on caller input.
The per-model skip resolved its model from _get_batch_routing_model, which
prefers the client-supplied top-level model field. That field only selects
routing credentials; the models a batch actually runs are the body.model
entries in the input JSONL. An unrestricted key could therefore name a
skip-listed deployment at the top level while routing a different,
same-provider model through the file, skipping the download, token count
and rate-limit reservation to bypass batch RPM/TPM limits.
Match the per-model skip against the file-bound model only (model-embedded
file id or unified managed file target), which is fixed when the file is
created and reflects the model the batch runs. The provider skip keeps using
the routing model since an admin opting out of a whole provider already
accepts any of that provider's models.
The litellm_metadata.skip_batch_input_file_rate_limiting flag was read
straight from the request body, so any caller whose key had unrestricted
model access could send it and skip the input-file download, token count,
and RPM/TPM reservation, bypassing their batch rate limits. Skip decisions
now derive only from server-controlled general_settings.
The spoofed-provider test configured empty descriptors, so the no-limits
shortcut skipped the file fetch and the assertion only proved the provider
allow-list did not short-circuit before descriptor evaluation. Give the key an
applicable rate limit so the only thing that can prevent the fetch is the
provider skip, then assert afile_content is awaited and the counters are
incremented; the spoofed custom_llm_provider must not skip processing.
Also cover the wildcard / all-proxy-models plus access_group_ids combination in
the model-access predicate so the wildcard-wins behavior is locked down.
Resolve the batch provider from router deployment credentials instead of
the user-supplied custom_llm_provider request field, so an unrestricted
key cannot spoof a skip-listed provider to bypass batch rate limiting.
Strengthen the provider-skip test to assert the file download and
descriptor work were short-circuited, and add a test that a spoofed
provider still falls through to rate-limit evaluation.
Add unit tests for the batch rate limiter's new skip/routing helpers so
the diff's patch coverage no longer depends on the CircleCI batches job,
whose coverage upload is blocked when an unrelated Bedrock integration
test aborts the run. Covers _get_batch_routing_model, _matches_skip_list,
_key_requires_batch_model_access_check, _has_applicable_batch_rate_limits,
_should_skip_batch_input_file_processing, _resolve_batch_input_file_fetch_params,
the descriptor-reuse path of _check_and_increment_batch_counters, and the
non-bytes file content guard in count_input_file_usage.
* fix(proxy): map stripped batch body.model to proxy alias for auth
replace_model_in_jsonl rewrites JSONL body.model to the provider id before
upload; batch file access checks must resolve that id back to model_name
so keys granted the proxy alias are not rejected with 403.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(proxy): surface resolved proxy alias in batch file 403 detail
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(mcp): resolve key.access_group_ids → MCP servers (ungated)
A teamless virtual key whose unified access_group_ids grant an
MCP-granting access group now sees and can call that server instead of
getting an empty list / 403. The key path previously read only the
legacy object_permission; this folds key.access_group_ids into the
key's base scope (ungated), mirroring can_key_call_model's fallback.
The gated assigned_*-checked override is unchanged.
* fix(mcp): expand name/alias entries in key access-group servers
access_mcp_server_ids may hold server names/aliases, not just ids. The
new ungated key path now runs them through expand_permission_list at the
source, so the early-return and union branches both surface resolved ids
— matching the legacy object_permission path and the gated extras path.
* test(e2e): cover navbar Logout flow as proxy admin
The Logout button under the navbar User dropdown was an uncovered
manual-QA step. This test signs in as admin, opens the dropdown,
clicks Logout, then navigates to a protected page and asserts the
redirect to /ui/login — proving the session was cleared.
* test(e2e): fix logout dropdown trigger and account-menu selector
The button never rendered the literal text "User" (it shows initials +
display name), and the antd Dropdown uses trigger={["click"]}, so the
synthetic mouseover/mouseenter never opened the popup. Open it with a real
click on the button's aria-label ("Account menu — ...").
* test(e2e): cover Internal User key modal, team info, key page
Three previously-uncovered manual-QA paths for the Internal User role:
- Create Key modal — confirm the team dropdown is populated with the
user's teams (verifying the role-scoped UI flow exists).
- Team info page — confirm the Settings/Members tabs are hidden for a
regular team member; only the read-only tabs render.
- Virtual Keys page — confirm the proxy's internal litellm-dashboard
team keys never leak into an internal user's table.
* test(e2e): share clickTeamId helper, strengthen key-filter assertion
Address review feedback on the Internal User e2e spec:
- Extract clickTeamId into helpers/navigation.ts; import in both
internalUser and teams specs instead of duplicating it.
- Anchor the litellm-dashboard absence check on the user's own seeded
key so it cannot pass vacuously against an empty table.
- Drop redundant dismissFeedbackPopup calls (navigateToPage already
dismisses internally).
* test(e2e): cover Internal Viewer nav, key, and team-info gating
Three previously-uncovered manual-QA paths for the Internal Viewer role:
- Nav only renders the read-only sections; admin-only items
(Internal Users, Organizations, Models + Endpoints) stay hidden.
- Virtual Keys page hides Create New Key, and the key detail view
hides Regenerate / Reset Spend / Delete actions.
- Team info page hides Members and Settings tabs for the viewer.
* test(e2e): scope viewer nav to sidebar, strengthen tab assertions
Address review feedback on the Internal Viewer e2e spec:
- Scope the nav test to the sidebar complementary landmark and match
items by link role + accessible name. The prior CSS nav, aside
selector grabbed the top bar (the sidebar is a complementary
landmark, not a <nav> tag), so the assertions never hit the real
nav links.
- Land via navigateToPage so the networkidle wait settles the
role-gated nav before asserting.
- Assert the Virtual Keys tab is visible (was only commented).
- Use toHaveCount(0) for hidden team tabs to match the nav block;
tabs are conditionally rendered, not CSS-hidden.
- Drop redundant dismissFeedbackPopup calls (navigateToPage already
dismisses internally).
* fix(anthropic): don't inject output_config.effort=xhigh on models without xhigh
The legacy-thinking translator on the /v1/messages route mapped any
thinking.budget_tokens >= 24000 to effort=xhigh and injected it into
output_config without checking model support. Claude Code's default
thinking budget (31999) hit this bucket, so Sonnet 4.6 (and Opus 4.6)
on Bedrock/Vertex started returning
400 output_config.effort: Input should be 'low', 'medium', 'high' or 'max'
Gate the xhigh choice on _supports_effort_level(model, "xhigh"), the
same capability check the reasoning_effort path already uses. Models
that advertise xhigh (Opus 4.7) keep it; everything else falls to high.
Fixes#29282
* test(anthropic): pin Opus 4.6 in legacy-thinking xhigh-clamp regression test
Opus 4.6 (bare, bedrock/invoke, vertex_ai) has supports_adaptive_thinking
but no supports_xhigh_reasoning_effort, so it hits the same clamping path as
Sonnet 4.6. It was named in the PR scope but lacked a pinned regression
guard; add the three variants to the parametrize list.
* docs: replace generated CLAUDE.md with hand-written guidance, remove AGENTS.md
Swap the auto-generated CLAUDE.md for a concise hand-written version that captures how we actually want agents to work in this repo: minimal comments, simplicity first, meaningful tests with a high mutation kill rate, PRs based off litellm_internal_staging rather than main, and curl against a live proxy as proof of fix instead of pasted pytest output. Remove AGENTS.md so there is one source of truth for agent guidance. The customer and company name confidentiality policy, along with the MCP available_on_public_internet note, are carried over from the previous CLAUDE.md.
* fix: further clarify communication guidelines
* docs: point GEMINI.md at CLAUDE.md instead of duplicating guidance
Replace the standalone GEMINI.md copy, which had already drifted from the new CLAUDE.md, with a one-line pointer so Gemini reads the same single source of truth.
* docs: simplify PR template test checklist item
Replace the rigid "at least 1 test is a hard requirement" checklist line with "I have added meaningful tests", which matches the testing guidance in CLAUDE.md, and tidy a comma into a semicolon in the scope-isolation item.
* docs: point AGENTS.md at CLAUDE.md instead of deleting it
Keep AGENTS.md so tools that read it still resolve guidance, but collapse it to the same one-line pointer to CLAUDE.md used by GEMINI.md, keeping a single source of truth.
* fix: make AI-generated rules more concise
* fix: spelling
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* fix: make the .env usage more careful
* docs: restore MCP available_on_public_internet note to CLAUDE.md
The PR description states this note was carried over verbatim from the
previous CLAUDE.md, but it was dropped in the rewrite. Restore it so the
file matches the description and the team guidance is not lost.
* docs: restore browser storage and CI supply-chain safety notes to CLAUDE.md
These security-relevant rules were dropped in the rewrite. Restore the
sessionStorage-over-localStorage (XSS) guidance and the CI supply-chain
rules (no curl|bash, pin versions, verify checksums) so agents editing UI
or CI code are still steered away from those pitfalls.
* docs: move area-specific guidance into nested CLAUDE.md files
The MCP, browser-storage, and CI supply-chain notes are scoped to
particular parts of the tree, so move each into a nested CLAUDE.md that
Claude Code loads on demand when those files are touched: the MCP note
under the mcp_server gateway, the browser-storage rule under the UI
dashboard, and the CI supply-chain rules under .circleci. Keeps the root
CLAUDE.md focused on general guidance while the area notes surface where
they are relevant.
* docs: keep CI supply-chain note in root CLAUDE.md
CI guidance applies beyond .circleci (it also covers downloads in GitHub
workflows and any CI script), and CI work does not reliably touch a single
subtree, so a nested file under .circleci would not surface it dependably.
Keep it in the always-loaded root instead. The MCP and browser-storage
notes stay nested where they map cleanly to one area of the tree.
* fix: make it clear we prefer httpOnly
* chore: make ci rule more concise
* chore: make concise
Fix formatting and punctuation in MCP note.
* fix: don't include Claude attribution
---------
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* test(e2e): cover Team Admin view + member + key flows
Adds a new spec exercising the previously-uncovered team-admin manual-QA
items: viewing all team keys (including other members'), adding a member,
removing a member, and creating a team key with All Team Models. Also
seeds a dedicated invitee user so the add-member test can run in parallel
with the proxy-admin invite test without colliding on the team roster.
* test(e2e): harden team-admin member specs per review feedback
Address Greptile feedback on the Team Admin spec:
- locate the delete action via getByTestId("delete-member") instead of
the fragile svg/img .last() selector
- match the seeded removable member by user_id (members_with_roles stores
no email, so the roster renders user_id)
- assert exact success-toast strings rather than broad regexes that could
match unrelated "success" text
* fix(guardrails): restore disable_global_guardrails persistence for keys
The per-key/team "Disable Global Guardrails" toggle silently stopped
working after #17042, which removed `disable_global_guardrails` from the
key/team request models and from the premium metadata allowlist. Without
those, the UI's top-level field was dropped by pydantic and never folded
into key `metadata`, so the runtime gate always read False and global
default_on guardrails kept running.
Restore the request-model fields (KeyRequestBase, NewTeamRequest,
UpdateTeamRequest) and the `LiteLLM_ManagementEndpoint_MetadataFields_Premium`
entry so the flag is promoted into metadata again. Because the key edit
form always submits the flag (false by default), guard the UI so it is
only sent when it actually changed (edit) or is enabled (create) — this
keeps the premium gate on enabling intact while not 403-ing non-premium
users who edit unrelated key fields, mirroring how guardrails/tags are
already stripped.
* test(guardrails): cover disable_global_guardrails toggle-off + clarify premium field comment
Add a prepare_metadata_fields case asserting `disable_global_guardrails: False`
overwrites an existing `True`, and rewrite the PREMIUM_METADATA_FIELDS comment to
explain why boolean premium fields are excluded from the empty-value strip loop.
The OSS-staging sync (d52fbfb45) overwrote the Bedrock batch model's
s3_bucket_name and aws_batch_role_arn with public-safe placeholders
(account 123456789012 / *_EXAMPLE role). The e2e_openai_endpoints CI job
runs the proxy with AWS account 941277531214 credentials, so on file
upload test_bedrock_batches_api failed with:
NoSuchBucket: The specified bucket does not exist
<BucketName>litellm-proxy-123456789012</BucketName>
Restore the real resources that live in account 941277531214 (verified
to exist) — the same values tests/batches_tests/test_bedrock_files_and_batches.py
already references.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(anthropic): add Claude Opus 4.8 and prune reasoning-effort flags
Register claude-opus-4-8 across the anthropic/bedrock/vertex/azure cost-map
entries, BEDROCK_CONVERSE_MODELS, and the setup-wizard provider list.
Prune two reasoning-effort fields from the cost map:
- Drop supports_minimal_reasoning_effort from the Claude fleet (58 entries).
"minimal" is not a real Anthropic effort level (the API accepts only
low/medium/high/xhigh/max), so LiteLLM degrades it to "low" regardless;
the flag was inert and misleading on Anthropic.
- Remove tool_use_system_prompt_tokens everywhere (103 entries). It is not in
the ModelInfo type and is read by no production code.
Update the affected config/schema tests; the reasoning-effort registry tests
now assert the Claude fleet omits supports_minimal.
* fix(anthropic): recognize output_config effort after minimal-flag prune
Pruning supports_minimal_reasoning_effort from the Claude fleet removed the
only "supports effort param" marker from 11 Opus 4.5 / mythos-preview map
entries that lack supports_output_config. _model_supports_effort_param then
returned False for them, so output_config was wrongly dropped under
drop_params=True -- regressing
test_anthropic_model_supports_effort_param_recognizes_supporting_models for
claude-opus-4-5-20251101 and the mythos preview.
- _model_supports_effort_param now treats supports_output_config as a
sufficient signal, matching the bedrock-invoke call sites that already
check supports_output_config OR a reasoning-effort flag. Shared map lookup
extracted into _supports_model_capability.
- Add supports_output_config: true to the 11 Opus 4.5 / mythos entries that
lost their only marker, restoring prior effort-forwarding behavior without
re-adding the inert minimal flag.
Updates the gollem_go_agent_framework example to the current Go release.
Clears stale Go stdlib advisories reported by osv-scanner against the
older 1.25.1 directive. No source changes; the single pinned dependency
(gollem v0.1.0) is backward compatible.
The Redis-backed VCR layer was recording and replaying the Google
OAuth2/STS token-mint call. The replayed ya29.* access token is
long-expired, but its recorded expires_in keeps credentials.expired
False, so litellm never refreshes it and sends the stale token to a live
Vertex/Gemini endpoint, which returns 401 ACCESS_TOKEN_EXPIRED. This
broke live partner-model tests whose completion call is not itself
cassette-backed (e.g. test_vertex_ai_llama_tool_calling).
Force credential-exchange hosts to pass through live (never recorded,
never replayed) by returning None from before_record_request, mirroring
the existing telemetry passthrough, so a fresh token is minted each run.
Regression from #28826, which added OAuth-token matcher tolerance plus
TTL-refresh-on-read so a stale token episode matched and never expired.
* fix(deps): bump vulnerable proxy dependencies (starlette/fastapi, granian, pyarrow, semantic-router)
Resolve known CVEs flagged by osv-scanner/grype against uv.lock. All bumped
versions verified to resolve, install, and pass the proxy auth/route/middleware
unit suites (717 tests) plus an import smoke on the new stack.
- starlette 0.50.0 -> 1.1.0 (CVE-2026-48710 "BadHost", GHSA-86qp-5c8j-p5mr):
versions <1.0.1 reconstruct request.url from the unvalidated Host header,
poisoning request.url.path. Required raising fastapi 0.124.4 -> 0.136.3,
which dropped fastapi's starlette<0.51.0 cap; an explicit starlette>=1.0.1
floor blocks regression to a vulnerable transitive resolution. The proxy's
own auth already reads scope["path"] via get_request_route, but the locked
starlette still flagged in container scanners and left other request.url
consumers exposed.
- granian 2.5.7 -> 2.7.4 (CVE-2026-42544, unauthenticated DoS via WebSocket
subprotocol header panic; CVE-2026-42545, WSGI response-header-panic DoS).
granian is a selectable proxy server (proxy_cli).
- pyarrow 22.0.0 -> 23.0.1 (CVE-2026-25087 / PYSEC-2026-113).
- semantic-router 0.1.12 -> 0.1.15: 0.1.12 was yanked (CVE-2026-42208 — its
unbounded litellm pin could resolve a credential-exfiltrating litellm==1.82.8
wheel).
Not fixable by bump: diskcache 5.6.3 (CVE-2025-69872, unsafe pickle
deserialization) has no upstream fix and is left pinned; exploiting it requires
write access to the local cache directory.
Relock side effect: sse-starlette 3.4.2 -> 3.4.4.
* deps: relax exact pins in optional extras to compatible ranges
The proxy/optional extras exact-pinned every dependency, which (1) forces
downstream `pip install litellm[proxy]` consumers into version lockstep and
(2) blocks them from pulling transitive security patches without forking — the
structural cause behind needing a litellm release to clear the starlette CVE in
the previous commit.
Convert the ordinary extras deps to `>=current,<next_major` ranges, mirroring
the core [project].dependencies style. Reproducibility for litellm's own
Docker/CI is unaffected: images install via `uv sync --frozen`, and the lock
re-resolves to the identical versions (no locked version changed).
Kept exact-pinned:
- litellm-proxy-extras, litellm-enterprise — litellm's own sub-packages,
versioned in lockstep with the release.
- opentelemetry-api/sdk/exporter-otlp — must resolve to matching versions.
- grpcio — supply-chain-pinned to a vetted, aged release.
Also corrects the stale comment claiming the extras are exact-pinned for Docker
reproducibility (the images use the lock, not these pins).
* fix(ci): resolve license-check lookup version from the floor for ranged deps
check_licenses.py derived the PyPI lookup version with
`next(iter(req.specifier))`, which returns an arbitrary specifier clause. For
a range like `>=0.12.1,<1.0` it picked the upper bound (`1.0`) — a version
that doesn't exist on PyPI — so the license lookup 404'd and the package was
flagged as having an unknown license.
The previous commit's switch from exact pins to ranges exposed this for
soundfile, pyroscope-io, redisvl, diskcache, and mlflow (the ranged deps not
already in liccheck.ini's allowlist). Prefer a lower-bound/exact version (a
real released version) for the lookup.
* fix(proxy): set strict_content_type=False on the FastAPI app
Starlette 1.0 / FastAPI 0.13x flipped the default to strict_content_type=True,
which refuses to parse a JSON request body when the client omits the
Content-Type header. The proxy previously accepted those requests, so the
fastapi/starlette bump in this PR would silently break clients that don't send
a Content-Type. Restore the prior lenient behavior explicitly.
Co-authored-by: stuxf <70670632+stuxf@users.noreply.github.com>
* feat(helm): split per-component ServiceAccounts for gateway, backend, and UI
Replace the single shared serviceAccount with three separate serviceAccounts
(gateway, backend, ui) so operators can attach different IRSA / Workload
Identity annotations per component without granting data-plane credentials
to the UI pod.
Key changes:
- values.yaml: rename serviceAccount → serviceAccounts with gateway/backend/ui
sub-keys; UI defaults to automount: false
- _helpers.tpl: replace litellm.serviceAccountName with three component-scoped
helpers (litellm.gateway/backend/ui.serviceAccountName)
- serviceaccount.yaml: create up to three separate ServiceAccount objects with
component labels and per-SA automountServiceAccountToken
- gateway/backend deployments: use their respective SA helpers
- ui deployment: use litellm.ui.serviceAccountName + explicit
automountServiceAccountToken: false on the pod spec so the projected token
is absent even when the SA itself allows it
- migrations-job: share the backend SA (both need DB write access)
Resolves LIT-3171
https://claude.ai/code/session_01QPy362WnjmEpeNuJaPUqmF
* fix(helm): enforce automountServiceAccountToken on all pod specs; fix leading --- in serviceaccount.yaml
- gateway/backend deployments: add explicit automountServiceAccountToken on
the pod spec so serviceAccounts.*.automount is honoured regardless of
whether the SA is chart-created or operator-supplied (previously the flag
only took effect on the SA object when create: true, creating an asymmetry
with the UI which already enforced it at pod-spec level)
- serviceaccount.yaml: use a $prev sentinel to emit --- only between
documents, preventing a leading --- when gateway SA is skipped but
backend or ui SA is created (avoids lint/GitOps warnings from strict
YAML parsers and tools like ArgoCD)
https://claude.ai/code/session_01QPy362WnjmEpeNuJaPUqmF
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(vertex-ai): pass litellm_params to validate_environment in video handlers and implement video edit for Veo
- Pass litellm_params to validate_environment in 11 video handler call sites
(remix, create_character, get_character, edit, extension, delete) so
DB-stored Vertex AI credentials are used instead of falling back to ADC
- Implement transform_video_edit_request/response for VertexAI: fetches
source video via fetchPredictOperation then submits a new
predictLongRunning request with the video bytes/gcsUri + edit prompt
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(vertex-ai): hoist fetchPredictOperation into handlers to avoid blocking event loop
- Add get_video_edit_prefetch_params() to BaseVideoConfig (returns None)
- VertexAI overrides it to return the fetchPredictOperation URL/body
- Both sync and async video_edit handlers call this and use their shared
httpx client for the fetch, passing the result as prefetched_source_data
- transform_video_edit_request is now a pure transform with no HTTP calls
- Fix extra_body.pop() mutation by working on a shallow copy
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(vertex-ai): include prefetch call inside _handle_error try/except block
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(videos): add prefetched_source_data param to all transform_video_edit_request overrides
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(video_edit): keep transform/pre_call outside try so validation errors propagate
Move transform_video_edit_request and logging_obj.pre_call outside the
try/except that wraps HTTP calls in (async_)video_edit_handler so that
ValueError validation errors (e.g. 'source video not complete yet') are
not silently wrapped as 500s by _handle_error. The prefetch HTTP call
keeps its own try/except so its errors are still mapped through the
provider's error handler. Matches the pattern used by
video_extension_handler and video_remix_handler.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* refactor(vertex_ai): delegate get_video_edit_prefetch_params to status retrieve
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* Fix varia review
* fix(video_edit): route transform errors through _handle_error
Wrap transform_video_edit_request and pre_call in the same try/except
as the HTTP call in sync and async handlers so validation failures
(e.g. source video not complete) return typed LiteLLM exceptions.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* fix(proxy): enforce tag budgets for key-level tags
Merge API key metadata.tags into request_data before _tag_max_budget_check
so per-tag budgets apply when tags are set on the key at creation time.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(auth): avoid false reject for key-inherited tags
Run reject_clientside_metadata_tags before key-tag injection, then inject key metadata tags immediately before tag budget checks so key tags still enforce budgets without being treated as client-supplied tags.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(guardrails): wire apply_guardrail into proxy logging callbacks
Route /apply_guardrail through pre/post proxy hooks and LiteLLM success/failure handlers so Langfuse and OTEL integrations receive input/output on guardrail-only requests.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(guardrails): fix Greptile review comments on apply_guardrail logging
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(apply_guardrail): preserve original exception and capture modified response
- Capture return value from post_call_success_hook so callback-modified
responses propagate to the caller.
- Wrap success/failure logging calls in defensive try/except so logging
infrastructure failures don't replace the user-visible response or mask
the original guardrail exception.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* Fix mypy
* fix(apply_guardrail): isolate failure logging and use post-hook response for logging
- Split async_failure_handler and post_call_failure_hook into independent
try/except blocks so a callback bug in one does not silently skip the
other.
- Build response_for_logging inside _emit_guardrail_success_logs after
post_call_success_hook runs, so logged data matches the response the
caller actually receives when the hook modifies the response.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* fix(apply_guardrail): fix black formatting and update tests for fastapi_request param
- Run black on guardrail_endpoints.py to fix CI formatting check
- Add _mock_proxy_logging() helper to enterprise guardrail tests to patch
proxy-server globals imported at call time
- Pass fastapi_request=Mock() in all direct apply_guardrail test calls
to match updated function signature
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(guardrails): use transformed exception from post_call_failure_hook in apply_guardrail
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* fix(guardrails): isolate sync/async logging handlers in apply_guardrail
Separate each logging handler call into its own try/except so a failure
in the async handler does not silently skip the sync handler submission
(and vice versa). Matches the docstring's defensive intent.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* fix(apply_guardrail): guard transformed_exception with isinstance check
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(guardrails): mock proxy globals in not_found test and share apply_guardrail logging fixture
- Add proxy-server global mocks to test_apply_guardrail_not_found so the
failure-path post_call_failure_hook call doesn't touch the real proxy
logging singleton.
- Extract the duplicated _mock_proxy_logging context manager out of the
two enterprise apply_guardrail test files into a shared conftest fixture
so the helper stays in one place.
* fix(guardrails): use update_messages to keep logging obj in sync
Co-authored-by: Yassin Kortam <yassin@berri.ai>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat: support goal mode for claude on bedrock
* fix failing lint test
* addressing greptile comments
* fixing failed test
* address greptile: copy output_config and warn on dropped converse format
* fix(bedrock): skip redundant output_config normalization on Converse reasoning_effort path
When reasoning_effort is mapped via _handle_reasoning_effort_parameter, the
resulting output_config is already normalized via
normalize_bedrock_opus_output_config_effort. Mark it as normalized so
_prepare_request_params can skip the redundant call (and the associated
get_model_info lookup) on every request.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* test(reasoning-effort-grid): reflect Bedrock opus-4-6 xhigh→max clamping
* fix(bedrock): stop leaking output_config marker and message-content mutation
* fix(bedrock): guard effort key access in normalize_bedrock_opus_output_config_effort
Defensively check that 'effort' is a valid key in _BEDROCK_OUTPUT_CONFIG_EFFORT_ORDER
before indexing, to prevent a KeyError if the hardcoded guard tuple ever drifts from
the order dict's keys.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* fix(bedrock): drop dead second clause in effort normalization guard
The 'effort not in _BEDROCK_OUTPUT_CONFIG_EFFORT_ORDER' check is
unreachable once 'effort not in ("xhigh", "max")' has been ruled out,
since both literals are present in the order dict. Keep the literal
membership check and let the dict lookups below speak for themselves.
* fix(bedrock): clamp output_config.effort against ceiling for any known value
The early return when effort was not 'xhigh'/'max' meant a ceiling of
'low' or 'medium' would silently forward an out-of-range value. Gate on
the known effort ordering instead so the ceiling comparison runs for
every recognized effort.
* test(grid_spec): use _CAPS_OPUS_4_7 for non-Bedrock opus-4-6 entries
claude-opus-4-6 now declares supports_xhigh_reasoning_effort in the model
map, so production accepts xhigh on Azure AI and Vertex AI routes. Update
those grid_spec entries to match production capabilities so expected()
predicts 200 for xhigh instead of 400.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* test(grid_spec): revert xhigh caps for non-Bedrock opus-4-6
azure_ai/claude-opus-4-6 and vertex_ai/claude-opus-4-6 do not declare
supports_xhigh_reasoning_effort in model_prices_and_context_window.json.
Azure AI upstream rejects xhigh with HTTP 400 ("Supported levels: high,
low, max, medium"). Restore _CAPS_4_6 so the grid predicts 400 for
xhigh, matching production capabilities.
* fix: stop advertising xhigh effort on Opus 4.5/4.6
Only Opus 4.7 supports the xhigh reasoning effort level. Remove the
supports_xhigh_reasoning_effort flag from every Opus 4.5 and Opus 4.6
entry (direct Anthropic, Bedrock, and regional variants) in both model
catalog files.
On the direct Anthropic path there is no effort clamp, so flagging 4.5/4.6
as xhigh-capable caused litellm to forward xhigh to a model that rejects it
(and made get_model_info misreport the capability). xhigh now correctly
degrades to high / raises on those models.
Bedrock graceful degradation for Claude Code goal mode is unaffected: it
relies solely on the bedrock_output_config_effort_ceiling clamp (4.5->high,
4.6->max, 4.7->xhigh), which runs before validation, so xhigh requests to
older Bedrock Opus models are still silently lowered rather than rejected.
Update effort-gating tests to reflect that 4.5/4.6 no longer accept xhigh.
* fix: clamp xhigh effort on Bedrock Invoke /v1/messages instead of rejecting
Claude Code "goal mode" sends output_config.effort=xhigh over the Anthropic
/v1/messages API, which routes Bedrock models through
AmazonAnthropicClaudeMessagesConfig. That path validated effort against the
model's native capability and raised 400 for xhigh on Opus 4.6, while the
chat-completions paths (Converse + Invoke) already clamp xhigh to the model's
bedrock_output_config_effort_ceiling. That asymmetry broke goal mode on the
exact API surface Claude Code uses.
Apply the same ceiling clamp on the messages path before the shared effort
gate runs, so xhigh degrades to max on Opus 4.6 (and stays xhigh on 4.7).
Scoped to adaptive-thinking models and to models that declare a ceiling, so
Sonnet 4.6 (no ceiling) and Opus 4.5 (budget mode) are unaffected and still
reject xhigh.
* fix(bedrock): preserve user output_config when applying reasoning_effort
- Converse path: merge mapped effort into existing output_config via
setdefault instead of overwriting it, matching the Anthropic Messages
path. Prevents user-supplied output_config.format from being silently
dropped when reasoning_effort is also provided.
- tests: clear _get_local_model_cost_map lru_cache in the autouse
fixture alongside get_bedrock_response_stream_shape to avoid stale
cache leakage between tests.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* fix(bedrock): pre-clamp reasoning_effort for chat invoke; correct test caps
- Add _clamp_adaptive_reasoning_effort_for_bedrock to AmazonAnthropicClaudeConfig
so raw reasoning_effort=xhigh degrades to the model's bedrock effort ceiling
before AnthropicConfig.map_openai_params converts it to output_config.
Mirrors converse path (_handle_reasoning_effort_parameter) and messages path
(_clamp_adaptive_reasoning_effort_for_bedrock) so the three Bedrock paths
are consistent.
- grid_spec: restore caps=_CAPS_4_6 for Bedrock converse/invoke Opus 4.6 entries
so the test reflects the model's actual JSON capabilities. Teach expected()
to bypass the xhigh/max cap check when bedrock_effort_ceiling will clamp
the wire effort, so the test still passes for Bedrock's graceful degradation
contract without lying about native model caps.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
---------
Co-authored-by: Dennis Henry <dennis.henry@okta.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Admin-configured skips (disable_batch_input_file_rate_limiting,
skip_batch_input_file_rate_limiting_for_models/_for_providers) and the
no-applicable-rate-limits short-circuit previously bypassed
_enforce_batch_file_model_access. A key with a restricted model
allowlist could therefore submit a batch JSONL referencing models
outside its allowlist whenever any of these skip paths fired, and the
provider-skip path was attacker-controllable via the request body's
custom_llm_provider field. Hoist the model-access guard to the top so
restricted keys always have their JSONL validated regardless of which
skip would otherwise apply.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
The skip_batch_input_file_rate_limiting flag in litellm_metadata is
user-controllable for batch requests (request-body metadata lands in
litellm_metadata via LITELLM_METADATA_ROUTES). Honoring it
unconditionally also skipped _enforce_batch_file_model_access, letting
a restricted key submit a JSONL referencing models outside its
allowlist. Only honor the metadata-based skip when the key has no
model allowlist to enforce.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Skip expensive pre-read of batch input files when no batch limits apply and model allowlist checks are not required, and decode model-embedded file IDs before file-content fetches to prevent upstream 404s.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(e2e): cover AI Hub make-public flow and public model_hub_table
Three previously-uncovered manual-QA paths land in one spec:
- Admin opens "Select Models to Make Public", advances through the
multi-step modal, and verifies the success toast.
- AI Hub tab strip exposes Model Hub / Agent Hub / MCP Hub / Skill Hub
— note the manual-QA "Claude Code Plugin Marketplace" label was
renamed to Skill Hub; the test pins the current name.
- Anonymous /ui/model_hub_table loads with the master key as `?key=`
and renders the Model Hub tab. Agent Hub / MCP Hub tabs are
conditional on public data and are not asserted here.
* test(e2e): harden AI Hub make-public + public hub assertions
Address Greptile review:
- Make-public test now asserts "Select All (N)" with N>=1 before clicking,
so a missing-seed-data run surfaces immediately instead of timing out
on the disabled Next button or the success toast.
- Public model_hub_table test dismisses the feedback popup before the
tab visibility assertion, matching the ordering used by navigateToPage
so a popup race can't mask the tab mid-evaluation.
* docs(e2e): explain admin vs public AI Hub tab asymmetry
Greptile flagged the all-4-tabs assertion as a potential CI flake,
inferring from the public-page comment that Agent Hub / MCP Hub might
be data-conditional in the admin view too. They aren't — ModelHubTable
renders all four tabs unconditionally for admins. Document the asymmetry
inline so future readers (and future review passes) don't re-derive it.
* test(e2e): cover add-MCP-server flow via discovery → custom form
The "Add MCP server" manual-QA step was uncovered. This adds a test
that opens the discovery modal, jumps into the custom-server form,
fills name + Streamable HTTP transport + a placeholder URL + None
auth, submits, and verifies both the success toast and the new row.
* test(e2e): apply greptile fixes to MCP add-server test
- Anchor the auth-type Select via its enclosing Collapse panel
("Authentication") instead of the placeholder text. The Form.Item has
no label prop, so the previous `hasText: /auth type/i` filter was
matching via "Select auth type" placeholder copy — fragile.
- Document the intentional lack of teardown, matching the pattern used
in addModel.spec.ts: the e2e runner discards the DB per invocation.
Addresses Greptile P2s on PR #29070.
* test(e2e): scope MCP row assertion to the servers table
Scope the post-create row lookup to `table tbody` so the form modal's
`server_name` input — which still holds the timestamped value during
its close animation — can't satisfy the assertion before the server
actually lands in the list.
* docs(e2e): note MCP coverage scope and link to tracker
This spec only smoke-tests the happy-path Streamable HTTP + None auth
flow. Add a top-of-file comment pointing at E2E_COVERAGE.md so future
contributors can see what's still uncovered (other transports, all
auth types, edit/delete, BYOK, tool list/call, access groups).
* fix(containers): record ownership for service-account keys + fix Prisma Json field serialization
- Track containers created implicitly via /v1/responses by extracting container IDs
from the response output and calling record_container_owner for each one, so
subsequent file-API calls from the same service account pass ownership checks.
- Fix DataError: Prisma Python requires Json fields to be JSON strings; serialize
file_object with json.dumps() before insert/update in LiteLLM_ManagedObjectTable.
- Add collect_container_ids_from_responses_response utility to responses/utils.py
that walks all output item shapes (code_interpreter_call, message annotations).
- Tests: two new cases covering the responses-tracking path and the end-to-end
record-then-assert flow for service accounts with team scope.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(containers): swallow all exceptions in ownership hook; tighten file_object_json type to str
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(containers): parse file_object JSON string in existing ownership test
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: container ownership recording bugs
- Remove unreachable _aresponses_websocket from route_type set in
base_process_llm_request; the WebSocket endpoint never flows through
base_process_llm_request, so this branch was dead code that gave a
false impression of coverage.
- Drop the HTTPException re-raise in record_container_owners_from_responses_response
so per-container failures (including HTTP 403/500 from conflicting
ownership rows) no longer abort the batch and skip recording for the
remaining container IDs in the same response.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* fix(containers): record ownership for streaming /v1/responses too
Streaming /v1/responses returns through the select_data_generator
branch in base_process_llm_request and bypasses the non-streaming
ownership tail, so code-interpreter containers created mid-stream
were never written to LiteLLM_ManagedObjectTable. Follow-up file API
calls would then 403.
Wrap the SSE generator so container ownership is recorded once the
upstream iterator finishes assembling completed_response. Also covers
the background-polling path, which loops body_iterator end-to-end.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
* test(e2e): cover Team-BYOK add-model flow as proxy admin
The team-only model + team assignment was an uncovered manual-QA path.
This adds a premium-gated test that toggles Team-BYOK, picks the seeded
E2E Team CRUD, submits, and verifies the model lands in All Models with
the team alias attached.
* test(e2e): apply greptile fixes to Team-BYOK test
- Add the 2s networkidle settle that the sibling addModel tests use —
networkidle fires before the All Models table finishes re-rendering,
so the search input was racing with the render.
- Assert on `models-results-count` before inspecting the table body so
an empty search result fails with a clear "expected results count"
message instead of timing out on a missing row.
Addresses Greptile P2s on PR #29068.
* test(e2e): harden Team-BYOK test against flake and stale state
- Add before/after cleanup that deletes any Cohere model already scoped
to e2e-team-crud via /v2/model/info + /model/delete, so Playwright
retries and local reruns don't accumulate rows.
- Pick the team from the dropdown by role/option name instead of a
global getByText match — avoids matching a previously-rendered tag
elsewhere in the form.
- Scope the "created successfully" assertion to .ant-notification so a
stale toast from an earlier test in the same browser context can't
vacuously satisfy it.
- Tighten the All Models assertion: require a single row that contains
BOTH the cohere model name AND the e2e-team-crud alias, so the
team-less wildcard from the sibling "Add wildcard route" test can't
satisfy the check.
* test(e2e): cover add-fallback flow in Router Settings as proxy admin
The Router Settings → Fallbacks → Add Fallbacks flow was an uncovered
manual-QA path. This adds a test that opens the modal, picks a primary
+ fallback from the seeded mock models, saves, and verifies both render
in the fallback table.
* fix(e2e): make router-fallback test idempotent and pick antd options by text
- Match `.ant-select-item-option` by text instead of `getByTitle(...)` —
FallbackGroupConfig uses `options=` (not <Select.Option> children), so
no `title` attribute is emitted and the title-based selector hangs.
- Add before/after hooks that wipe any fallback for fake-openai-gpt-4 via
/config/update so retries and local reruns don't trip on leftover state.
- Tighten the success assertion to a single tbody row containing BOTH the
primary and the fallback names — pre-existing rows can no longer
vacuously satisfy the check.
- Fix the stale "Three tabs" comment to "Four tabs".
Addresses Greptile P2s on PR #29069.
* fix(e2e): keyboard-select fallback models + correct cleanup endpoint
- Replace mouse-based option clicks with click-to-focus + type + Enter.
FallbackGroupConfig's Selects use `options=` and a custom
getPopupContainer, so locating options via `.ant-select-dropdown`
hit several races: DOM-clicks left antd's popup state stale (the
primary popup then intercepted the fallback click), `getByRole`
matched always-mounted hidden options, and pointer stability fought
the open animation. Typing into the showSearch input narrows the
listbox to one option and Enter selects it cleanly.
- Assert on dialog-side state changes (the active tab adopts the
primary model name; the chain helper shows "1/10 used") instead of
popup contents — these reflect the actual selection landing.
- Cleanup helper now hits /get/config/callbacks (the real endpoint;
/get/callbacks returns 404), so the before/after reset actually
clears prior router_settings.fallbacks state.
* test(ui): e2e cover team model edit + admin identity in navbar
Adds two Playwright tests as part of the manual-QA → e2e migration:
"Edit team model selection" exercises the Settings tab Models multi-select
+ Save Changes flow on a seeded team, and the existing login test now
opens the User dropdown and asserts the role and User ID render — guarding
against regressions where login succeeds but the auth context is empty.
Resolves LIT-3093
* test(ui): restore seeded models in team-edit test so retries don't fail
The 'Edit team model selection' test removed fake-anthropic-claude from
E2E_TEAM_CRUD_ID without restoring it. CI runs with retries: 2 and the seed
script runs once before the suite, so a flake on this test would fail the
retry at the "tag is visible" assertion. Wrap the test in try/finally and
restore the seeded models via /team/update before and after.
* test(e2e): fail loudly if team/update restore call fails
Surfaces the real cause when the master key is wrong or the proxy is
unreachable, instead of silently leaving the team in a stale state and
failing later on the visibility assertion.
* fix(e2e): match navbar account button by aria-label, not non-existent "User" text
The previous trigger filter (hasText: /^User$/) didn't match the rendered
UserDropdown button — its text is the displayName ("Account" for the
master-key admin, an email for SSO users), never "User". The evaluate
call then timed out after 15s in CI. Use the stable aria-label prefix
the component always emits, and click directly since the dropdown is
configured trigger=["click"] (the synthetic hover was unnecessary).
* fix(mcp): resolve team.access_group_ids → MCP servers
A virtual key whose team has an MCP-granting access group attached
via /v1/access_group now sees that server through /v1/mcp/server (and
can call tools on it) instead of getting an empty list. The runtime
already resolves the key's unified access_group_ids; this adds the
symmetric resolution on the team side, mirroring the model-side
pattern in can_team_access_model — the group being on the team is
itself the gate, so no assigned_team_ids re-check is needed.
Resolves#27657
* chore(mcp): address greptile review on team access-group resolver
Forward already-imported prisma_client / user_api_key_cache /
proxy_logging_obj to _get_mcp_server_ids_from_access_groups so it
skips its lazy re-import path. Update test docstring + assertions to
reflect that the resolver is invoked with [] (and short-circuits
without DB access) rather than skipped entirely.
* refactor(ui): extract auth state into AuthContext
Move auth state (token, userID, userRole, accessToken, premiumUser, userEmail,
disabledPersonalKeyCreation, showSSOBanner) out of src/app/page.tsx into a
new AuthProvider at src/contexts/AuthContext.tsx. Wrapped at the root layout
so login/onboarding/dashboard routes all have access via useAuth().
Day 1 foundation for the App Router migration: migrated (dashboard)/X/page.tsx
route entry points won't have a parent passing props, so shared auth state
must live in a context they can read from.
Sub-components are unchanged — they still receive accessToken/userID/userRole
as props from page.tsx (which now reads them from useAuth()). Only the
page.tsx → top-level-page-component handoff is de-drilled; deeper prop
drilling is left for the per-page migration to address.
Net change: -86 lines from page.tsx (state + two effects moved), +5 in
layout.tsx (provider wrap), new AuthContext.tsx (~140 lines), test update
to wrap CreateKeyPage in AuthProvider.
Fixes LIT-3366
Part of LIT-3128
* fix(ui): await getUiConfig before clearing authLoading
The AuthContext refactor flipped authLoading to false synchronously on mount
while letting getUiConfig() run fire-and-forget. On SERVER_ROOT_PATH deployments
this races the unauthenticated login-redirect effect: the redirect fires with
proxyBaseUrl still at its module-init value, sending users to /ui/login instead
of {SERVER_ROOT_PATH}/ui/login.
Restores the original sequencing inside AuthProvider's mount effect and adds a
Playwright spec wired into the existing SERVER_ROOT_PATH workflow matrix. The
spec delays the config endpoint via page.route() to make the race deterministic
across CI runners.
Adds a rule to CLAUDE.md and AGENTS.md instructing AI coding agents to
pause and ask the user before introducing any third-party organization
name that does not already appear in the repository. Names already
established in the codebase (existing LLM providers, etc.) remain fine
to use without prompting.