Commit graph

37946 commits

Author SHA1 Message Date
yuneng-jiang
b021d5c109
Merge pull request #26525 from BerriAI/litellm_rag_aws_endpoint_cleanup
[Fix] broaden RAG ingestion credential cleanup to AWS endpoint/identity fields
2026-04-25 15:13:42 -07:00
yuneng-jiang
1dee006423
Merge pull request #26520 from BerriAI/litellm_feat-team-my-user-tab
[Feat] Add "My User" tab to team info page
2026-04-25 15:05:05 -07:00
Yuneng Jiang
c364c51115
style: black formatting 2026-04-25 14:47:54 -07:00
Shivam Rawat
7e57b15de2
fix(guardrails): team-level guardrails and global policy guardrails can run together (#26466)
* fix(guardrails): apply team-level guardrails alongside global policy guardrails

Two bugs prevented team-direct guardrails from being automatically applied
when using a team-scoped API key:

1. Auth caching: `valid_token.team_metadata` was never refreshed from the
   freshly-fetched team object at the "Check 6" step in
   `_user_api_key_auth_builder`. Guardrails added to a team after the key
   was first cached were therefore invisible to `move_guardrails_to_metadata`.
   Fix: propagate `_team_obj.metadata` → `valid_token.team_metadata` after
   every "Check 6" team fetch (user_api_key_auth.py).

2. Guardrail execution: `get_guardrail_from_metadata` checked
   `data["litellm_metadata"]` before `data["metadata"]`. When a request
   carried a non-empty `litellm_metadata` without a "guardrails" key, the
   merged guardrail list written to `data["metadata"]` by
   `move_guardrails_to_metadata` was shadowed and the guardrail received an
   empty requested-guardrails list (custom_guardrail.py).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix merge conflict

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-25 14:45:06 -07:00
Yuneng Jiang
30cb459d1c
fix: broaden RAG ingestion config cleanup to cover AWS endpoint/identity fields
Mirrors the existing api_base handling in _load_credentials_from_config():
when a vector_store_config references litellm_credential_name and the stored
credential does not itself define aws_sts_endpoint or aws_web_identity_token,
any caller-supplied value on those keys is removed. Keeps endpoint and
identity fields aligned with the credential definition.
2026-04-25 14:34:59 -07:00
Shivam Rawat
5416a7c86e
fix(bedrock guardrail): dedupe post-call log entry when only post_call is configured (#26474)
* fix(bedrock guardrail): dedupe post-call log when only post_call is configured

When a Bedrock guardrail runs with only post_call configured, the post-call
trace section showed the same guardrail twice (one entry per parallel
INPUT/OUTPUT API call). Add skip_logging param to make_bedrock_api_request
and pass it for the INPUT scan so the OUTPUT scan stands as the single
canonical post_call log entry. INPUT exceptions still propagate, so the
input-side blocking behavior is preserved.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(bedrock guardrail): post_call only scans OUTPUT, not INPUT

post_call is the response-validation hook by definition — input scanning
belongs to pre_call / during_call. The previous code ran an extra INPUT
scan in post_call when no pre/during hook was configured, which produced
a duplicate "post-call" entry in the trace and was semantically wrong
for a "post-call" event.

Drops the should_validate_input branch and parallel asyncio.gather in both
async_post_call_success_hook and async_post_call_streaming_iterator_hook
in favor of a single OUTPUT scan. Reverts the now-unneeded skip_logging
parameter on make_bedrock_api_request introduced in the previous commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 14:30:06 -07:00
Ryan Crabbe
79762e4f5c
fix(team): address review feedback on My User tab
- Fall back to email match when looking up the caller in
  members_with_roles — email-onboarded members may have user_id=None on the
  stored entry, which caused a false 404 for valid members. (P1)
- Replace 3 raw Prisma queries with get_team_object / get_team_membership /
  get_user_object so the endpoint reuses the cache + retry layer the rest of
  the proxy uses. (P2)
- Allow internal_user role to reach /team/{team_id}/members/me by adding the
  route to LiteLLMRoutes.self_managed_routes (the handler already enforces
  member-of-team access).
- Return null from the UI fetch on 404 instead of throwing, so a proxy admin
  who isn't a team member sees the existing empty state rather than an error
  string in the always-visible tab. (P2)
- Move the fetch out of networking.tsx into a colocated React Query hook
  (useMyTeamMember) next to MyUserTab; TeamInfo now passes only teamId.
- Tooltip + empty-state copy on Model Scope: drop "(all team models)"
  parenthetical and the redundant tooltip line.
- Tests: build real LiteLLM_TeamMembership / LiteLLM_BudgetTableFull
  fixtures (with created_at) so the Pydantic Union resolves to the Full
  variant; add an assertion that budget_reset_at survives end-to-end; add a
  test for the email-only member match path.
2026-04-25 14:28:27 -07:00
Mateo Wang
319193604c
[Feat] Add azure/gpt-5.5 + azure/gpt-5.5-pro entries (+ dated variants) (#26361)
* feat(azure): add azure/gpt-5.5 + azure/gpt-5.5-pro entries (+ dated variants)

Azure variants of OpenAI's GPT-5.5 family. Microsoft has not yet
shipped GPT-5.5 on Azure OpenAI (latest GA on the Foundry models page
is GPT-5.4 as of 2026-04-24), but adding the entries day-0 mirrors the
established precedent for azure/gpt-5.4* (which were in the cost map
before the Azure rollout) so cost tracking and capability flags work
the moment customers deploy.

Schema follows the existing azure/gpt-5.4* shape:
- Same base/long-context pricing as openai/gpt-5.5*: $5/$30 chat,
  $60/$360 pro per 1M, with priority tier 2x base
- Azure variants drop the flex/batches keys (Azure has no flex tier)
  but keep priority pricing, matching gpt-5.4* precedent
- mode=chat for the thinking model, mode=responses for pro

reasoning_effort capability flags mirror the OpenAI variants exactly
since Azure proxies the same API contract: minimal rejection on both
chat and pro, low/none rejection on pro. Once #26456 (which sets
supports_low_reasoning_effort + minimal=false on openai/gpt-5.5*)
lands, OpenAI and Azure flag profiles align.

Tests pin entry presence + pricing for all four Azure variants and
verify the live-API-derived reasoning_effort flags.

* test: register supports_low_reasoning_effort in cost-map JSON schema

azure/gpt-5.5-pro and azure/gpt-5.5-pro-2026-04-23 added in this branch
carry supports_low_reasoning_effort=false. The strict
'additionalProperties: false' schema in
test_aaamodel_prices_and_context_window_json_is_valid rejected the new
key. Register it alongside the other supports_*_reasoning_effort
entries.

Note: the runtime side of this flag (code that reads it) lands in
#26456. Until that PR merges the flag is inert for both Azure and
OpenAI pro entries, but having the schema accept it lets cost-map
tests pass on either merge order.
2026-04-25 14:19:59 -07:00
yuneng-jiang
4a11362695
Merge pull request #26522 from BerriAI/litellm_yj_apr25
[Infra] Merge dev branch
2026-04-25 14:18:15 -07:00
michelligabriele
ae925baaa1
fix(model_management): refresh in-memory router after POST /model/update (#26427) 2026-04-25 14:10:23 -07:00
michelligabriele
9c00f9776b
fix(content_filter): log guardrail_information on streaming post-call (#26448)
* fix(content_filter): log guardrail_information on streaming post-call

* fix(content_filter): reset detections per chunk in streaming hook
2026-04-25 14:06:38 -07:00
yuneng-jiang
ceed00fc2f
Merge pull request #26513 from BerriAI/litellm_/intelligent-maxwell-35e39d
[Fix] Harden /model/info redaction for plural credential field names
2026-04-25 14:02:31 -07:00
michelligabriele
db8ef44323
fix(key_management): enforce upperbound_key_generate_params on /key/regenerate (#26340) 2026-04-25 13:49:00 -07:00
yuneng-jiang
7705cb39f2
Merge pull request #26512 from BerriAI/litellm_rag_credential_precedence
[Fix] bind RAG ingestion config to stored credential values
2026-04-25 13:43:13 -07:00
Ryan Crabbe
9d71ad4796
[Feat] Add "My User" tab to team info page
Adds a new "My User" tab on the team detail page (between Overview and
Virtual Keys) so non-admin team members can see their own spend, budget,
budget reset date, rate limits, model scope, and team role.

Backend
- New `GET /team/{team_id}/members/me` endpoint that resolves the caller
  from the API key and returns only their own LiteLLM_TeamMembership row
  plus minimal team context (alias, role, email). Returns 404 if the
  caller is not a member of the team. Avoids exposing other members'
  data, which would happen if we filtered `/team/info` client-side.
- New `TeamMemberInfoResponse` Pydantic model.

Frontend
- New `MyUserTab` component (antd) — read-only summary cards.
- New `teamMemberMeCall` helper in networking.tsx.
- Tab is visible to all team members (including non-admins).
2026-04-25 13:42:51 -07:00
ryan-crabbe-berri
9f60b751e1
Merge pull request #26338 from BerriAI/litellm_feat-mcp-server-alias-permissions
feat(mcp): resolve team/key MCP permissions by name or alias
2026-04-25 13:22:04 -07:00
yuneng-jiang
8f4f2a1b30
Merge pull request #26516 from BerriAI/litellm_mcp_oauth_caller_access_check
[Fix] Align MCP OAuth proxy endpoints with per-server access policy
2026-04-25 13:05:26 -07:00
Yuneng Jiang
0b9a7044b7
fix: leave vector_store_config untouched when credential lookup misses
When litellm_credential_name resolves to no values (name not registered
in litellm.credential_list), bail out before touching the resolved
config. Prevents stripping a caller-supplied api_base in the no-match
case.
2026-04-25 13:02:09 -07:00
Yuneng Jiang
7503f14f3f
fix: harden /model/info redaction to cover plural credential field names
Mirrors the existing vertex_credentials handling for the newer
vertex_ai_credentials field, and extends the dynamic masker's sensitive
pattern set to recognize the plural form so other plural-named credential
fields are also covered.
2026-04-25 12:59:12 -07:00
shin-berri
0a166157af
Merge pull request #26511 from BerriAI/cursor/guard-fork-dependencies-eec1
ci: add supply-chain guard to block fork PRs that modify dependencies
2026-04-25 12:54:16 -07:00
yuneng-jiang
449146e815
Merge pull request #24428 from BerriAI/litellm_staging_03_23_2026
Litellm staging 03 23 2026
2026-04-25 12:43:44 -07:00
Yuneng Jiang
7a6ab3870f
[Fix] bind RAG ingestion config to stored credential values
When a vector_store config references a stored credential by name,
overwrite caller-supplied values on the same keys with the credential
values, and drop api_base from the resolved config when the credential
does not define one. Keeps the api_key / api_base pair consistent with
the credential definition.
2026-04-25 12:41:36 -07:00
Yuneng Jiang
91f6661b37
[Fix] Align MCP OAuth proxy endpoints with per-server access policy
Bring `/server/oauth/{server_id}/authorize`, `/token`, and `/register`
in line with `fetch_mcp_server`: the helper that resolves the server now
also applies the per-caller access policy. Admin-view callers are
unrestricted; non-admins must have the server in their allowed-servers
set; servers resolved from the admin-only `/server/oauth/session`
temporary cache reject non-admins.
2026-04-25 12:39:44 -07:00
Chesars
40eeffe089 test(vertex): use valid dimension (128) for multimodalembedding live test
dimensions=1 was silently dropped before #24415 wired the param through
to Vertex. Now that it's forwarded, Vertex rejects it (valid values:
128, 256, 512, 1408). Switch the test to a valid dimension.
2026-04-25 16:10:25 -03:00
Cursor Agent
0bd9213d8d
ci: add supply-chain guard to block fork PRs that modify dependencies
Add a new CI workflow that rejects pull requests from forks when they:
- Modify uv.lock (any change at all)
- Add new dependencies to any pyproject.toml file (root, litellm-proxy-extras, enterprise)

Security properties:
- Uses pull_request (not pull_request_target) so no secrets are exposed
- All action refs pinned to full SHA hashes
- persist-credentials: false on all checkouts
- permissions: {} (no GitHub token permissions)
- No user-controlled input in run: blocks (no script injection)
- Proper TOML parsing via stdlib tomllib (not regex on raw text)
- Only triggers when dependency files are actually changed (paths filter)

Internal PRs (from branches in the canonical repo) skip the job entirely.

Co-authored-by: Krrish Dholakia <krrish-berri-2@users.noreply.github.com>
2026-04-25 18:46:50 +00:00
Cesar Garcia
c8ceafe941
Merge pull request #26510 from BerriAI/litellm_internal_staging
Sync litellm_staging_03_23_2026 with litellm_internal_staging
2026-04-25 15:22:29 -03:00
Chesars
ebe16072f2 Merge remote-tracking branch 'upstream/litellm_internal_staging' into litellm_staging_03_23_2026
# Conflicts:
#	model_prices_and_context_window.json
#	tests/test_litellm/llms/vertex_ai/multimodal_embeddings/test_vertex_ai_multimodal_embedding_transformation.py
2026-04-25 15:16:13 -03:00
Chesars
f118ecc1b6 ci: align CI/workflows with litellm_internal_staging
Sync CI configs with upstream/litellm_internal_staging:
- Migrate from Poetry to uv (PR #25007)
- Pull in zizmor security fixes for workflows
- Add isolated unit test workflows + reusable bases
- Drop redundant matrix workflow and azure-batches workflow
- Drop .circleci/requirements.txt (replaced by uv lockfile)
2026-04-25 15:09:00 -03:00
Chesars
384cfdad47 Revert "Merge pull request #24164 from dongyu-turo/feat/update-bedrock-claude-price-above-200k"
This reverts commit b8189ea1de, reversing
changes made to 19c8f3d565.
2026-04-25 15:04:05 -03:00
Chesars
a4a8d86c2b Revert "Merge pull request #24417 from Chesars/refactor/shared-format-mapping"
This reverts commit 3f3d275e67, reversing
changes made to 498c113933.
2026-04-25 15:03:24 -03:00
yuneng-jiang
bb61f747c3
Merge pull request #26496 from BerriAI/litellm_yj_apr23
[Infra] Merge dev branch
2026-04-25 11:01:16 -07:00
Chesars
3c8677209b Revert "Merge pull request #24420 from Chesars/docs/complete-integrations-landing"
This reverts commit e5f5ffbe15, reversing
changes made to 4399b7614d.
2026-04-25 14:58:25 -03:00
Yuneng Jiang
410b163f95
fix: annotate centralized common_checks gather results for mypy
After ``return_exceptions=True``, gather results are typed
``Any | BaseException``. Mypy could not narrow through
``None if isinstance(x, HTTPException) else x`` ternaries, leaving
BaseException in the type passed to ``common_checks``.

Switch the narrowing checks to ``BaseException`` (HTTPException is the
only one still possible after authz failures were re-raised) and add
explicit ``Optional[...]`` annotations so the typed objects flow into
``common_checks``. No behavior change.
2026-04-25 10:40:15 -07:00
Yuneng Jiang
151d7ab1bc
fix: isolate per-fetch HTTPException in centralized common_checks gate
The asyncio.gather in `_run_centralized_common_checks` ran with
`return_exceptions=False` and a single bare `except HTTPException`
arm, so an HTTPException from any one fetch (the realistic case is a
404 from `get_team_object` when a token references a deleted team)
zeroed out the user, end-user, project, and global-spend contexts in
addition to falling back the team object. That silently skipped the
user budget, end-user budget, and project enforcement passes inside
`common_checks` for the unrelated contexts that had actually fetched
fine.

Switch to `return_exceptions=True` and apply per-fetch fallback
(matches the pre-refactor per-fetch try/except pattern in the builder):

- ProxyException / BudgetExceededError still propagate as authz failures.
- HTTPException on the team fetch reconstructs from the token; on the
  other fetches it nulls only that one context.
- Successful fetches always reach `common_checks` intact.

Adds two unit tests covering the team-404 and user-404 cases to lock
the per-fetch isolation in. Drops the inaccurate `PROXY_ADMIN tokens
short-circuit` claim from the docstring — admin tokens still flow
through `common_checks`; admin status is only honored where the
underlying check exempts it.
2026-04-25 09:58:23 -07:00
Yuneng Jiang
4884b0b611
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_yj_apr23
# Conflicts:
#	litellm/proxy/management_endpoints/key_management_endpoints.py
2026-04-25 09:47:47 -07:00
yuneng-jiang
c05de83f1c
Merge pull request #26490 from BerriAI/litellm_restrict_global_spend_routes
[Fix] Restrict /global/spend/* routes to admin roles
2026-04-25 09:30:36 -07:00
yuneng-jiang
5fae8051c5
Merge pull request #26488 from BerriAI/litellm_sortable_spend_logs_columns
[Feature] UI - Spend Logs: sortable Model and TTFT columns
2026-04-25 09:29:43 -07:00
ryan-crabbe-berri
9839ab7f96
Merge pull request #26004 from BerriAI/litellm_fix-preserve-reserved-metadata
fix: preserve service_account_id in metadata on /key/update
2026-04-25 09:28:47 -07:00
Ryan Crabbe
7eab549190
test: tighten prepare_key_update_data mock and apply black
- test_prepare_key_update_data: replace bare MagicMock with
  MagicMock(spec=LiteLLM_VerificationToken) and explicitly set
  existing_key_row.metadata = {}, so reserved-field reads return real
  values instead of MagicMock-returning-MagicMock. Fixes a regression
  surfaced by the new reserved-metadata preservation logic.
- test_key_management_endpoints.py: black-format-only changes from
  recent edits.
2026-04-25 09:08:55 -07:00
ryan-crabbe-berri
25513b9e43
Merge pull request #26001 from BerriAI/litellm_fix-ui-unconditional-cost-write
fix(ui): stop injecting $0 cost on model edit
2026-04-25 08:50:46 -07:00
ryan-crabbe-berri
419a34a900
Merge pull request #26442 from BerriAI/litellm_feat-restrict-org-admin-permissions
feat: UI setting to disable /key/generate for org admins
2026-04-25 08:50:26 -07:00
Ryan Crabbe
ea5c762b00
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_fix-preserve-reserved-metadata 2026-04-25 08:46:17 -07:00
yuneng-jiang
9426aff475
Merge pull request #26493 from BerriAI/litellm_key_route_caller_checks_addendum
[Fix] Extend caller-permission checks to service-account + tighten raw-body acceptance
2026-04-24 23:40:58 -07:00
Yuneng Jiang
f7f77d0cf4
docs: clarify intent of empty-dict handle_key_type call in regenerate 2026-04-24 23:18:18 -07:00
Yuneng Jiang
5190bd07eb
fix: extend caller-permission checks to service-account + harden raw-body acceptance 2026-04-24 23:18:18 -07:00
yuneng-jiang
ffdcc689f4
Merge pull request #26492 from BerriAI/litellm_key_route_caller_checks
[Fix] Tighten caller-permission checks on key route fields
2026-04-24 23:05:43 -07:00
Yuneng Jiang
2220f3076a
fix: tighten caller-permission checks on key route fields 2026-04-24 22:55:21 -07:00
Yuneng Jiang
2047446546
Scope NULLS LAST to ttft_ms only
The previous version appended NULLS LAST to every ORDER BY, which would
silently change DESC semantics for any nullable sort column added to the
whitelist later. Today the existing sort columns (spend, total_tokens,
startTime, endTime, request_duration_ms, model) are all non-null in the
result set, so the clause is a no-op for them — but the broader form is
misleading.

Apply NULLS LAST only when sorting by ttft_ms (the only column whose
computed expression actually produces NULLs). Update the model test to
assert the clause is absent for non-nullable columns.
2026-04-24 22:52:04 -07:00
Yuneng Jiang
01eee0944c
[Fix] Restrict /global/spend/* routes to admin roles
The routes in `global_spend_tracking_routes` (e.g. /global/spend/report,
/global/spend/teams, /global/spend/keys) return spend aggregated across
every team, customer, and api_key in the proxy. They were included in
`internal_user_routes` and `internal_user_view_only_routes`, so non-admin
roles could read proxy-wide spend.

Drop them from both non-admin route lists. PROXY_ADMIN and
PROXY_ADMIN_VIEW_ONLY access is preserved through their existing branches
in route_checks.py, and the `get_spend_routes` permission opt-in
continues to grant access for keys that need it.

Updates two pre-existing test parametrizations whose expected results
flip from True to False, and adds parametrized coverage over every
route in `global_spend_tracking_routes` for: PROXY_ADMIN_VIEW_ONLY
allowed, INTERNAL_USER blocked, INTERNAL_USER_VIEW_ONLY blocked,
INTERNAL_USER + get_spend_routes permission allowed.
2026-04-24 22:46:07 -07:00
Yuneng Jiang
5c0349a635
[Feature] UI - Spend Logs: sortable Model and TTFT columns
Extend the /spend/logs/ui sort_by whitelist to accept "model" and
"ttft_ms", and wrap the Model and TTFT (s) column headers with the
existing SortableHeader component so users can sort by either.

TTFT has no stored column, so it is computed inline as
(completionStartTime - startTime) milliseconds. Non-streaming rows
(completionStartTime null or equal to endTime) yield NULL and the
ORDER BY uses NULLS LAST so they always sort to the bottom regardless
of direction, matching the existing "-" display in the UI.

Adds parametrized backend tests for sort_by=model and a dedicated test
covering streaming + non-streaming TTFT ordering in both directions.
2026-04-24 22:29:15 -07:00