Commit graph

53398 commits

Author SHA1 Message Date
yucheng
265919f9a9 fix(mcp): annotate the general_settings cast for the type-discipline gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
6adff6ffac test(mcp): cover throttled token exchange surfacing as an outage rather than a sign-in challenge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
f2a876f23f fix(mcp): keep a route hidden from a client ip from rerouting to a case variant of its name
get_mcp_server_answering_to stops at the pass that finds an exact name or id and hides it from client_ip instead of falling through to the case-insensitive and prefix passes, and _scoped_server treats a name the registry knows for some caller but not this one as denied, so connect, discovery and scoped routing all refuse the hidden route instead of serving a public server whose alias only differs by case

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
2172b8b60e fix(mcp): resolve scoped, connect and discovery routes through one exact-first, ip-aware lookup and stop denied names widening to access groups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
f89b92763b fix(mcp): resolve /mcp/{name} routes through one exact-first lookup for connect, discovery and the scoped router
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
23eb6bd8a5 test(mcp): patch the scoped-connect resolver the preemptive challenge reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
674356d3d8 fix(mcp): resolve case-variant scoped connects with the same alias-first priority as the exact name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
bc07570af8 test(mcp): expect the connected route name in the connect-time OBO preflight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
eeadf7bdc1 fix(mcp): name the connected route in OBO rejection challenges and test the initialized Agent 365 guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
06e5d9f7c5 fix(guardrails): return the Entra exchange fallback verdict from one path so CodeQL sees no implicit None
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
be5a5ccf81 fix(guardrails): return explicitly from every Agent 365 preflight exchange outcome
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
f00a2c1c18 fix(mcp): run the single-server admission lookup once for the challenge, sign-in preflight and exchange
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
1bddacfb68 fix(guardrails): stop pointing admins at the removed resource_app_id in the connect-time 503
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
4236a43aa7 feat(mcp): challenge opaque caller bearers at connect and move Agent 365 sign-in onto the fixed production constants
Rework the Agent 365 sign-in provider for the guardrail shape 43189 landed on main: the OBO scope and
resource come from the fixed AGENT_365_PROD_* constants instead of the removed resource_app_id/api_base
fields, and the exchange runs through the shared TokenExchanger so the connect preflight and the tool
call reuse one cached token per caller assertion.

A present but non-JWS bearer is now rejected in preflight_caller_sign_in, so the connect answers 401
with the RFC 9728 challenge instead of letting the call reach tools/call and lose WWW-Authenticate in
the JSON-RPC error. The OBO-only tool-call challenge in operations.py stays narrowed to token_exchange
servers.

Immutable rewrites (tuple, MappingProxyType, explicit None checks) keep the LIT002 total within the
budget without a mutable-ok

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
072898ab19 fix(mcp): satisfy type discipline gate and the merged input-schema key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
c2f85ca2e7 fix(mcp): share the allowed lookup between the sign-in and exchange preflights
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
e6a6e3ec35 style(mcp): format the new connect sign-in preflight tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
0099b96ddc feat(mcp): challenge rejected caller sign-in subjects at connect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
fcfd7b4f14 fix(mcp): name the connected segment in sign-in challenge resource metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
6a71a3b1a0 test(mcp): add integration coverage for the caller sign-in gates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
692086a0ae fix(mcp): restore the raw bearer hook kwarg, exact-name-first resolution, and the base OBO challenge gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
4be107e779 fix(mcp): flatten the sign-in merge without stacked comprehension clauses
LIT014 budgets one for clause per comprehension; chain.from_iterable keeps
the dedupe immutable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
df076ce24c fix(mcp): keep the exact-name fallback in the challenge resolver and reformat
get_mcp_server_answering_to now falls back to get_mcp_server_by_name when
no published prefix form matches, preserving the exact-name lookup the
preemptive path had before the router-equivalent resolver. Applies ruff
format to caller_sign_in.py and agent_365.py per the lint gate.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
4cadb402e5 feat(mcp): fold gateway sign-in into a caller sign-in contract on the challenge path
Replace the parallel gateway sign-in provider registry with
CallerSignInProvider, merged into the existing oauth2_token_exchange
challenge: gates key off caller_sign_in_for() returning non-None, the
resolver answers by server_id, case-insensitive name, and short prefix
like the router, keeps_caller_authorization covers OBO servers, and the
Agent 365 guardrail exchanges through an injected TokenExchanger.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
9c68ca4412 fix(mcp): keep the stored BYOK credential for catalog identity only on tools/list
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
Listing used the resolved stored credential both to key the caller's catalog slot and as the
upstream transport header, so REST api_key and bearer_token listings sent the user's secret
instead of the server's static token and the MCPJWTSigner gate went quiet. The upstream client
and the signer gate now read the caller-supplied mcp_auth_header for every auth type, exactly
as before the catalog existed, and the stored credential only names the slot tools/call reads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 01:59:51 +00:00
yucheng
0e4de07c65 test(mcp): oauth2 BYOK listing sends the minted token, not the stored secret, through the real proxy
Integration cell for the listing fix: a client_credentials BYOK server with a stored user credential,
one tools/list as that user, the peer must see a live minted bearer and one /token mint. Red at the
pre-fix tip (zero mints, stored secret upstream), green at the fixed head and at the merge base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 00:40:49 +00:00
yucheng
de92a1103c fix(mcp): keep oauth2 listing on the minted or signed credential, not the stored BYOK secret
The listing helper that keys the per-caller catalog by the stored BYOK credential also handed that
credential to the upstream client, which on an oauth2 server short-circuited the client_credentials
mint and the MCPJWTSigner gate. Split the two: the catalog identity keeps the stored credential so
tools/call finds the caller's slot, while an oauth2 server's tools/list sends only the per-request
header, letting the M2M mint or signed JWT proceed as on main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 23:49:02 +00:00
yucheng
c8870f0080 chore(mcp): merge main into listed-tool metadata branch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 18:15:26 +00:00
joshua-berri
79756cbb9b
feat(agents): enforce authoritative agent permissions (#43721)
* feat(agents): authoritative permissions

* fix: enforce authoritative managed agent permissions

* fix(agents): only consult the identity store for managed targets

is_agent_allowed entered the identity-store path whenever a prisma client
was configured, so an ordinary agent paired with an internal user returned
503 instead of 200. Classify the target from the registry first and fall
back to the store only when the registry has no entry, so an unmanaged
target never depends on the store being reachable.

* fix(agents): gate the managed path on an admitted policy object

Ten call sites branched on `managed_agent_policy is not None`, which any
MagicMock attribute satisfies, so the managed path fired on unmanaged
subjects and died in Pydantic validation as a 503. Route every check
through a shared helper that requires a real AgentResponse.

* test(mcp): stub the writer replica the fresh-policy reads use

reload_admitted_user now passes check_db_only through to get_user_object,
so the user row is read from writer_db. Point the mocks at the replica the
code actually reads and give each parametrized case its own user id.

* fix(agents): cap a managed agent at the invoking team's agents

resolve_agent_access returned the managed policy's grants before the
agent_caller ceiling was applied, so a managed agent acting on behalf of a
user reached agents that user's team was never granted. Intersect with the
caller ceiling the unmanaged path already honours.

* fix(agents): restore token narrowing and scope the private-access suppressions

The managed-model check lost its valid_token narrowing when it moved to the
shared helper. Make the caller-access resolver public rather than reaching
into it from module scope, and give each remaining private access a reason.

* docs(agents): drop the comment claiming admins skip the A2A permission check

The check has never had an admin bypass on this path, so the comment
described behaviour the code does not implement.

* test(proxy): stub the writer reads and restore the MCP manager singleton

Fresh-policy user lookups read writer_db, so the team and rest-endpoint
mocks stubbed a replica the code no longer reads, and the dashboard
session fake still had the pre-kwarg signature. The manager reload also
rebound global_mcp_server_manager in every MCP module without restoring
it, leaking an empty manager into later files.

* style: sort imports under the litellm package ruff config

* fix(mcp): cap a managed agent's servers and tools at the invoking caller

managed_agent_servers and managed_agent_tools returned the agent's own
grants without the agent_caller ceiling the unmanaged resolvers apply, so
a managed agent reached MCP servers and tools the echoed caller could not.
Call the existing ceiling helpers on both axes.

* refactor(mcp): return the caller-capped tools without an interim list

The ceiling helper already returns a sequence, so materializing it into a
list added a mutable collection for nothing. Sort at the return sites
instead, which also makes the tool order stable across both branches.

* fix(agents): preserve actor ceilings during managed target checks

* fix(agents): keep managed permission ceilings authoritative

* fix(mcp): fail closed on authoritative caller team outages

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-30 11:11:37 -07:00
ryan-crabbe-berri
0736a14326
fix(ui): right-align money and count columns across tables (#37889)
* fix(ui): right-align money and count columns across tables

Make numeric the one alignment token for the three shared table
wrappers (DataTable, MemberTable, SimpleTable) via NUMERIC_CELL_CLASS,
and flag every money, cost, and bare-count column that was still
left-aligned. Raw ui/table usages that render spend or budgets get the
same class on their header and cell.

* test(ui): render the organizations alignment test with providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): query alignment cells by role instead of DOM traversal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 10:59:53 -07:00
yucheng-berri
7204942756
fix(proxy): keep request-body credentials out of stored spend-log requests (#43635)
* fix(proxy): keep request-body aws credentials out of stored spend-log requests

* fix(proxy): redact every credential-named request-body field in stored spend-log requests

Replace the hard-coded AWS key check in the spend-log request-body sanitizer with
SensitiveDataMasker's key classification, so Azure, Vertex, watsonx, OCI, GigaChat,
Gemini and header credentials are redacted too. Proxy-stamped key identity metadata
is kept.

* fix(proxy): keep request identifiers named like keys in stored spend-log requests

* refactor(proxy): drop the AWS-only snapshot exclusion now that spend-log redaction is name-based

* refactor(proxy): use SensitiveDataMasker's key classification without an exclusion list

* refactor(proxy): always redact credential-named fields in stored spend-log payloads
2026-09-30 17:31:48 +00:00
berriai-litellm-provider-info-sync[bot]
025292e75b
feat(pricing): add vertex_ai gemini-3.8 flash tts rows (#43876)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 10:20:57 -07:00
yucheng-berri
e5c74cb2a6
fix(cost_calculator): stop copying optional_params into response hidden params (#43637)
* fix(cost_calculator): stop copying optional_params into response hidden params

* test(cost_calculator): assert the stored spend-log request and logging payload carry no forwarded credentials
2026-09-30 10:17:21 -07:00
devin-ai-integration[bot]
2ed9761921
fix(router): strip encrypted reasoning the pinned deployment cannot decrypt (#43781)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 09:56:27 -07:00
devin-ai-integration[bot]
82d8b3797c
fix(proxy): attribute completed batch cost rows to /batches in daily activity (#43870)
* fix(proxy): attribute completed batch cost rows to /batches in daily activity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): wait for priced batch tokens before asserting team endpoint activity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 16:16:52 +00:00
berriai-litellm-provider-info-sync[bot]
efdccd8811
chore(cost-map): add fireworks priority prices for ember-1, nemotron and glm 5.3 us rows (#43811)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 09:08:16 -07:00
berriai-litellm-provider-info-sync[bot]
971e60660b
chore(cost-map): add openai gpt-image-2.5 batch prices from the pricing page (#43869)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 09:06:31 -07:00
devin-ai-integration[bot]
b80052839e
refactor(rust): centralize bridge execution wrappers (#43871)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 15:54:27 +00:00
devin-ai-integration[bot]
9dda4d895f
fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates (#43764)
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
2026-09-30 07:32:16 -07:00
devin-ai-integration[bot]
b781d157d7
chore(model_prices): add Gemini Veo, Mistral and Azure Claude 4.5 deprecation dates (#43857) 2026-09-30 07:26:17 -07:00
devin-ai-integration[bot]
b370996b9d
test(router): settle the shared logging worker before recording shadow callbacks (#43847)
* test(router): settle the shared logging worker before recording shadow callbacks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): always stop the shared logging worker after settling it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 04:44:43 -07:00
devin-ai-integration[bot]
6bc17f98d7
refactor: clean up fresh tech debt from 2026-09-29 (#43830)
* refactor: clean up fresh tech debt from 2026-09-29

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(routing): pin usage-based routing Redis reads through the proxy and SDK

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 02:40:37 -07:00
yucheng
001fcad930 test(mcp): align listed-tool slot tests with per-caller keying
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 09:31:42 +00:00
yucheng
7dee160b9f fix(mcp): key OpenAPI listed-tool entries per caller so tools/call reads its own guarded listing
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 09:08:00 +00:00
yucheng
2e1c6bfc70 fix(mcp): hand tools/call hooks the exact catalog entry tools/list served
get_listed_tool re-applied the admin description override on top of the cached listing, so a
guardrail-masked description was restored to its original wording at call time, and the OpenAPI /
local-registry call path built its metadata from the registry instead of the guarded caller catalog.
Both paths now return the cached entry as served, falling back to the registry only when no listing
was recorded

Adds tests/integration/mcp/test_mcp_listed_tool_metadata.py (red on the prior head for the two
regressions, red on the merge base for the feature, green on this head)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 08:32:47 +00:00
shrey-berri
04fa760bf2
fix(bedrock): add beta header for output config in message (#43778) 2026-09-30 00:50:09 -07:00
devin-ai-integration[bot]
314ff111e5
fix(router): carry per-request routing reads on context variables instead of public method kwargs (#43814)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 07:34:06 +00:00
devin-ai-integration[bot]
d79600987e
perf(router): honour the cooldown read interval in the routing prefetch (#43815)
Resolves LIT-9043

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 00:29:20 -07:00
shrey-berri
6cf51383bf
fix(params): filter internal traceback flag from provider requests (#43783) 2026-09-30 00:02:41 -07:00
yuneng-jiang
d02ff435bf
test(bedrock): accept regional aliases that inherit Converse routing (#43785)
* test(bedrock): accept regional aliases that inherit Converse routing

* test(bedrock): cover regional alias metadata independently of catalog
2026-09-29 23:30:06 -07:00