Commit graph

53461 commits

Author SHA1 Message Date
Yucheng He
86c3b910a6 fix(mcp): route a scoped name to the caller's granted server before the registry's pick
A key granted only the server named `docs` was refused with 403 on
/mcp/docs when an ungranted server held `docs` as its alias, and a key
granted only `github` was refused on /mcp/GITHUB when an ungranted
`GitHub` existed: the scoped router took the registry-wide pick and
answered "denied" whenever that pick was not among the caller's servers.

The router now runs the registry's own pass order (exact alias,
server_name, name, server_id; the same case-insensitively; prefix forms)
over the caller's granted servers first, through
get_mcp_server_answering_to(among=...), with IP hiding applied at the
pass that found the server. A name hidden from the client IP stays denied
before any grant lookup, and a name the registry places only on an
ungranted server stays denied rather than being retried as an access
group.

The connect preflight reuses the router's selection for a single scoped
name as the server it challenges, signs in and exchanges for, so a 401
names the granted server; an ungranted caller keeps the registry pick and
the downstream 403, and the no-key path is unchanged. Unauthenticated
RFC 9728 discovery has no grant list and stays on the registry pick.

Tests: the `only_b` assertion in
test_scoped_router_selects_the_server_the_connect_preflight_resolves and
the no-IP `internal` assertion in
test_scoped_router_hides_a_private_server_from_an_external_ip_like_the_connect_preflight
now expect the granted server, which is what the base branch returned for
both shapes; the unentitled-key exchange test's fixture becomes the empty
selection the router returns for such a key; listing tests that stub the
manager now point its lookup at a real empty manager so the `among` pass
runs. New tests cover the grant-first router, the manager `among` pass
order and IP hiding, and the connect preflight under an alias collision.
2026-10-01 18:00:25 -07:00
yucheng
608fff7b08 Merge remote-tracking branch 'origin/litellm_mcp_listed_tool_metadata' into litellm_mcp_gateway_sign_in_provider_pr
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
2026-10-01 17:24:33 +00:00
yucheng
52a517cff7 Merge remote-tracking branch 'origin/main' into litellm_mcp_listed_tool_metadata 2026-10-01 17:14:28 +00:00
devin-ai-integration[bot]
39e31958f8
test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy (#43998)
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): stub HIBP through respx by disabling the aiohttp transport

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): share the httpx transport fixture across proxy unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): restore proxy globals without a missing-value sentinel

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): isolate the mcp server manager per test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 10:11:45 -07:00
moe-berri
d9f73245be
feat(lens): track worker spend through virtual keys (#43989)
* feat(lens): bill worker analysis through virtual keys

* fix(lens): pin the verified worker image and add setup proof

* fix(lens): preserve network checks and redact billed analysis logs

* test(lens): preserve legacy worker result submission during upgrade

* fix(lens): enforce trusted worker IPs and restore coverage uploads

* docs(lens): explain trusted proxy requirements for worker allowlists

* fix(lens): yield to worker disconnects after the synthetic body
2026-10-01 09:47:30 -07:00
devin-ai-integration[bot]
6f123b7083
test(proxy): migrate DB and Redis backed proxy tests into tests/integration (#43996)
* test(proxy): migrate DB and Redis backed proxy tests into tests/integration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): drop a type suppression comment from the key metadata integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): scope integration test cleanup to owned rows and wait for backend stats flush

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): seed NULL cache_hit and bound recovery reads from below

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 09:23:36 -07:00
shrey-berri
d96477abce
fix(proxy): preserve decision request bodies under token limits (#43920) 2026-10-01 09:23:33 -07:00
shrey-berri
c2ae483782
fix(bedrock): add beta for thinking display updates (#43832) 2026-10-01 09:21:10 -07:00
devin-ai-integration[bot]
4b06d04334
test(anthropic): native /v1/messages reasoning integration tests built on a captured Claude Code request (#43361)
* test(integration): group /v1/messages contracts under tests/integration/messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): make ci coverage census collect nested test dirs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): nest /v1/messages contracts under messages_endpoint/providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): replay a real Claude Code /v1/messages request through the native wire

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drop legacy covers marker from claude code wire test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): inline the Claude Code request instead of a json fixture

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drop the legacy covers marker from the new contract

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): add Claude Code /v1/messages customer-journey matrix (native + responses bridge) (#43386)

* test(anthropic): add Claude Code /v1/messages customer-journey matrix (native + responses bridge)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): strengthen bot-flagged assertions in the Claude Code matrix

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): type the usage mapping parameter in the shared builders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drop low-priority Claude Code error and count_tokens tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): send the full 24-tool Claude Code request and pin upstream headers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drop responses bridge Claude Code tests to keep this PR Anthropic direct only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): fix duplicate WebSearch tool, drop mutation in stream builders, ignore pings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): assert upstream request order in multi-turn Claude Code tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): name Claude Code wire tests by behavior and move provider-agnostic ones to routing and streaming

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): sort imports after moving the Claude Code fixture into _support

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): group anthropic messages tests into feature subfolders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): drop the pre-move anthropic test paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): cover native reasoning translation, response and pricing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): narrow PR to new reasoning tests, restore moved files and drop non-reasoning tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): scope reasoning tests to reasoning and cover betas, thinking usage and streamed pricing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): write reasoning cases as literal sent and received fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): require stopped stream blocks and check upstream model on switch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 09:12:53 -07:00
berriai-litellm-provider-info-sync[bot]
3a11192f68
fix(cost-map): reprice fireworks deepseek v4.1 flash to the 2026-10-01 pricing update (#44024)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-01 08:16:30 -07:00
devin-ai-integration[bot]
b4bb2a77a2
test: inject the HIBP client and the MCP loop clock so two backend tests stop flaking (#44007)
* test(proxy): inject the HIBP client into the breached-password update test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): drive optional-discovery deadlines with an injected loop clock

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: bound MCP deadline checks, support Python 3.10, pin HIBP URL

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 05:16:49 -07:00
devin-ai-integration[bot]
5107f205a0
refactor: clean up fresh tech debt from 2026-09-30 (#43993)
* refactor: clean up fresh tech debt from 2026-09-30

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor: keep agent tracing route comment in LiteLLMRoutes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor: move agent tracing route comment above the trace routes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor: drop route comment that duplicates the trace handler docstring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 04:46:59 -07:00
devin-ai-integration[bot]
0980f756bd
fix(guardrails): scan Responses API input in Azure Text Moderation (#43965)
* fix(guardrails): scan Responses API input in Azure Text Moderation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): log Azure Text Moderation prompts at debug and cover streamed Responses blocking

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 22:58:03 -07:00
devin-ai-integration[bot]
19842da059
fix(guardrails): scan Responses API input in Azure Prompt Shield (#43786)
* fix(guardrails): scan Responses API input in Azure Prompt Shield and Text Moderation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): tolerate unmodeled Responses input items in Azure prompt extraction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): pick Azure prompt source by call type so a messages stub cannot hide Responses input

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): tighten Azure Content Safety endpoint test types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): suppress Azure cast lint violations with cast-ok reasons

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): audit Azure content safety across endpoints

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): isolate worker-kill audit rig and cover during_call on chat

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(guardrails): shorten Azure cast-ok reasons to fit the line limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): inline spend row count in the concurrency audit cell

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): assert caller-observed outcomes in Azure call type unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): reuse the existing text moderation response helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): assert no duplicate rows instead of exact row count after worker kill

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): poll worker-kill spend rows to settle before the duplicate check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep Azure Text Moderation on messages only so this PR stays Prompt Shield scoped

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-30 22:48:20 -07:00
moe-berri
6d7d183a80
feat(lens): investigate sampled traces and retain batch results (#43942)
* fix(lens): parallelize scan analysis with bounded concurrency

* feat(lens): investigate sampled activity and preserve scan results

* fix(lens): pin the compatible investigation worker image

* fix(lens): report incomplete reviews and simplify setup validation

* fix(lens): stabilize large investigations and preserve incomplete results

* fix(lens): preserve bounded readers and distinguish counterexamples

* fix(lens): pin compatible worker and verify batched grouping cost

* fix(lens): exclude counterexamples from finding recurrence

* feat(lens): show completed scan duration in results and history

* fix(lens): fold batch selection into results navigation
2026-09-30 22:30:04 -07:00
yuneng-jiang
2eb2bf130b
fix(proxy): restore pre-config-wins handling of pass-through endpoints (#43962)
* fix(proxy): restore pre-config-wins handling of pass-through endpoints

Config-wins (#41779) made general_settings.pass_through_endpoints a config-owned key. The DB reader then got the config list back as if it were DB rows, re-registered each entry without forward_headers on every DB sync, and the stripped copy won the route lookup, so a config pass-through with forward_headers: true stopped forwarding Authorization. UI create, update and delete of pass-throughs were also rejected while the config declared any.

This puts pass-throughs back on their pre-#41779 path: the settings store no longer lets the config own the key, the config list is captured env-resolved at load_config, each DB sync merges DB entries with config entries on paths the DB does not declare, and /config/field/info reads the stored rows only. A UI pass-through write re-applies that merge immediately so the config entries stay served until the next sync.

* fix(proxy): keep config pass-throughs in every reload of the merged list

get_config now returns DB pass-throughs plus config ones on other paths,
each DB sync republishes that merged list, and /config/field/info reads
pass_through_endpoints from the DB row so a UI write never drops stored
entries when models are not stored in the DB

* fix(proxy): keep serving pass-throughs while the config file reloads

load_yaml cleared the runtime pass-through list, so auth: false routes
answered 401 while get_config awaited the database

* fix(proxy): read stored pass-throughs from the writer before a UI write

A lagging read replica could return an older list, and the UI create and
edit flows write the whole field back

* fix(proxy): apply config file pass-through auth changes on reload

The kept runtime list was merged as if it were DB entries, so an edited
config entry on the same path was dropped. Merge the stored DB row with
the fresh config instead, and give the field-info test mock a writer

* fix(proxy): keep pass-throughs served while a DB sync reads the database

get_config resets the stored DB rows before reading them again, which
cleared the served pass-through list and made auth: false routes answer
401 for the length of the read

* refactor(proxy): move the settings store reload out of the loop

basedpyright rejects a Final variable assigned inside a loop
2026-09-30 22:09:29 -07:00
berriai-litellm-provider-info-sync[bot]
2b19ddb7a3
fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (#43916)
Some checks failed
LiteLLM Rust / rust-lint (push) Has been cancelled
LiteLLM Rust / rust-test (push) Has been cancelled
LiteLLM Rust / rust-wheel (push) Has been cancelled
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 21:59:59 -07:00
berriai-litellm-provider-info-sync[bot]
ef6aa4ad66
chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (#43898)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 21:37:44 -07:00
berriai-litellm-provider-info-sync[bot]
9b8ddb0982
chore(cost-map): add fireworks inkling priority prices from the prices api (#43949)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-09-30 21:36:41 -07:00
devin-ai-integration[bot]
a308a8e579
feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (#43134)
* feat(guardrails): honor litellm_params.timeout in every HTTP guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): accept timeout kwarg in presidio and responses-handler post stubs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): bound hiddenlayer startup jwt call by configured timeout, drop akto from timeout coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): narrow hiddenlayer startup auth timeout without cast

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): bound hiddenlayer jwt refresh by configured timeout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep provider timeout defaults when unset and bound only rubrik moderation calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): cover model_armor and run timeout probes concurrently

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): match sink calls to the exact guardrail name

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 21:18:47 -07:00
ryan-crabbe-berri
ae05f7d2c1
test(e2e): typed per-test metadata for the e2e suite (#42044)
* feat(e2e): give e2e tests typed metadata for what they drive

@meta(Subject(domain, route, providers, models, capabilities, mode)) declares
what a test is about with closed enums, and each field lands in the JUnit report
as a property. The quota_management suites are the first to declare it.

* docs(e2e): say e2e_metadata avoids litellm, not that it is stdlib-only

It already imports pydantic and pytest, both of which the suite needs to collect. The rule that matters is no litellm import

* test(e2e): declare models through the constant each test drives

43 @meta declarations in quota_management typed the model name out again, so changing the call would leave the coverage report naming the old model. Each file now has one constant used by both, and a guard fails on any model written as a string literal in @meta

* refactor(e2e): set route only when the endpoint is what the test checks

A budget or rate-limit test whose chat call only triggers the block now leaves route unset, since its steps already name the call. Tests of an endpoint keep it: budget CRUD, key creation, spend reporting reads, and the per-endpoint spend tests for chat, messages, embeddings, batches and health. The two /spend/logs tests tagged chat_completions are now spend_reporting

* refactor(e2e): build the declared properties without mutating a list

subject_properties seeded a list and grew it with append and extend. It now flattens one tuple per field, and the plural-name table is a read-only mapping

* fix(e2e): tag each spend-route probe with the endpoint it checks

The breadth test gave all 33 probes spend_reporting, so /key/list, /user/list, /team/list, /organization/list and /customer/list counted as spend reporting. Each case now carries its own route, with organization and customer management added to Route
2026-09-30 21:03:21 -07:00
devin-ai-integration[bot]
321be01877
fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (#43973)
* fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch

The post-call Redis batch only goes out after every success callback finishes or the 1s deadline, so an
identical request sent right after the first response missed the cache and went to the provider again.
Response-cache writes go straight to Redis again, the counters, rate limits, TPM and slot releases keep
riding the post-call batch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): re-export DualCache explicitly instead of through a noqa

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(caching): keep the DualCache re-export as a reasoned noqa for the strict ruff gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 20:39:48 -07:00
Arnold Gálovics
f67caac8d4
feat(ui): filter tags by name and description on the Tag Management page (#42949) 2026-09-30 20:16:08 -07:00
devin-ai-integration[bot]
ae60fd1b2f
feat(providers): add Cortecs as an OpenAI-compatible provider (#43872)
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: markoarnauto <7702545+markoarnauto@users.noreply.github.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 19:49:24 -07:00
ryan-crabbe-berri
424bfd8758
feat(e2e): record each e2e test's steps, starting with ProxyClient (#42393)
* feat(e2e): record each e2e test's steps, starting with ProxyClient

@step on a harness method records a plain-English line for every call, in
order, as repeated JUnit step properties. Labels are templates filled from the
call's parameters, like "Generate a virtual key with models: claude-haiku-4-5
and rpm limit: 3", and secret request fields are marked Field(repr=False) so
they never print. ProxyClient and the rate-limit QuotaClient carry steps first;
the other harnesses follow one area at a time. The recorder and JUnit tests run
in the Code Quality workflow's test_e2e_metadata step.

* docs(e2e): rewrite the recorded test steps guide in plain language

* fix(e2e): keep logging callback credentials out of recorded steps

* fix(e2e): mask the run's credentials in every recorded step

* fix(e2e): attach steps before the oauth failure snapshot

The failed setup or call report of an mcp_oauth_live test copied user_properties before the steps were attached, so it carried no steps. Every setup and call report now takes its properties after the steps attach

* fix(e2e): name the saved credential in its recorded step

The create_credential label read credential_info, which defaults to {} and is never set by the live callers, so the step printed nothing after 'for'. It now reads the required credential_name, and a guard fails on any label that reads a field with a default
2026-09-30 19:33:53 -07:00
yuneng-jiang
c168199e33
test(ci): repair stale tests and move retired OpenAI text-completion fixtures (#43958)
* test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups

Request fakes now carry the scope a real Starlette request has, the GCS pub/sub
spend-log golden gains the agent identity keys from #43722, the auto-router
session tests follow the baseline_models contract from #43348, and the
Interactions spec checks resolve the create body and resource paths from the
live spec instead of hardcoded names

* test(ci): move retired OpenAI text-completion fixtures to live vehicles

OpenAI still serves native /v1/completions on the gpt-5.4 family, so the
single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt
batches and echo with logprobs now 500 on every OpenAI model, so those cases
keep the same text-completion-openai transport pointed at Fireworks, which
documents both. The optional-params test asserts the request body actually
sent instead of a success callback whose assertions were swallowed

* test(ci): use a serverless Fireworks model for the text-completion batch and echo cases

gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not
deployed; glm-5p3-flash is listed as serverless

* test(ci): skip the ROI calculator repository listing in the security route sweep

GET /roi-calculator/repositories (#43669) lists repositories from the configured
GitHub API, api.github.com by default, so the S2 sweep's GET of every route made
the owned proxy reach an external host and failed the egress check in 31
integration-security tests. It joins /get/latest_release_info in the deny list
2026-09-30 19:19:59 -07:00
devin-ai-integration[bot]
0c515ed7a8
feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (#43063)
* feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token

Stamp metadata.used_client_oauth_token where the proxy decides to forward a
client's Anthropic OAuth token, carry it through StandardLoggingMetadata into
the spend log row, add a used_client_oauth_token filter to /spend/logs/ui, and
surface it on the Logs page as a Credential filter and drawer field. The token
itself never reaches the log

* fix(proxy): carry used_client_oauth_token onto failure spend rows for litellm_metadata routes

* fix(proxy): resolve used_client_oauth_token against the provider the call was sent to

* fix(proxy): keep the proxy's used_client_oauth_token stamp on failure rows and move the resolver under llms/anthropic

* fix(logging): read used_client_oauth_token from the proxy-stamped metadata slot

On routes that carry proxy metadata in litellm_metadata, metadata is the
caller's own body field, and merge_litellm_metadata lets it win. Resolve the
flag from litellm_metadata when the proxy stamped it there so a caller cannot
set it in the standard logging payload

* fix(spend-logs): read used_client_oauth_token from the bucket the route stamped

A guardrail on the unified path adds litellm_metadata to a chat request after
the proxy stamped metadata, so both spend row writers read the new bucket and
stored null. The success row now resolves the flag the same way the callback
payload does, and the failure row picks the bucket from the request route.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-30 19:17:01 -07:00
yucheng
265919f9a9 fix(mcp): annotate the general_settings cast for the type-discipline gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
6adff6ffac test(mcp): cover throttled token exchange surfacing as an outage rather than a sign-in challenge
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
f2a876f23f fix(mcp): keep a route hidden from a client ip from rerouting to a case variant of its name
get_mcp_server_answering_to stops at the pass that finds an exact name or id and hides it from client_ip instead of falling through to the case-insensitive and prefix passes, and _scoped_server treats a name the registry knows for some caller but not this one as denied, so connect, discovery and scoped routing all refuse the hidden route instead of serving a public server whose alias only differs by case

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
2172b8b60e fix(mcp): resolve scoped, connect and discovery routes through one exact-first, ip-aware lookup and stop denied names widening to access groups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
f89b92763b fix(mcp): resolve /mcp/{name} routes through one exact-first lookup for connect, discovery and the scoped router
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
23eb6bd8a5 test(mcp): patch the scoped-connect resolver the preemptive challenge reads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
674356d3d8 fix(mcp): resolve case-variant scoped connects with the same alias-first priority as the exact name
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
bc07570af8 test(mcp): expect the connected route name in the connect-time OBO preflight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
eeadf7bdc1 fix(mcp): name the connected route in OBO rejection challenges and test the initialized Agent 365 guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
06e5d9f7c5 fix(guardrails): return the Entra exchange fallback verdict from one path so CodeQL sees no implicit None
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
be5a5ccf81 fix(guardrails): return explicitly from every Agent 365 preflight exchange outcome
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
f00a2c1c18 fix(mcp): run the single-server admission lookup once for the challenge, sign-in preflight and exchange
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
1bddacfb68 fix(guardrails): stop pointing admins at the removed resource_app_id in the connect-time 503
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
4236a43aa7 feat(mcp): challenge opaque caller bearers at connect and move Agent 365 sign-in onto the fixed production constants
Rework the Agent 365 sign-in provider for the guardrail shape 43189 landed on main: the OBO scope and
resource come from the fixed AGENT_365_PROD_* constants instead of the removed resource_app_id/api_base
fields, and the exchange runs through the shared TokenExchanger so the connect preflight and the tool
call reuse one cached token per caller assertion.

A present but non-JWS bearer is now rejected in preflight_caller_sign_in, so the connect answers 401
with the RFC 9728 challenge instead of letting the call reach tools/call and lose WWW-Authenticate in
the JSON-RPC error. The OBO-only tool-call challenge in operations.py stays narrowed to token_exchange
servers.

Immutable rewrites (tuple, MappingProxyType, explicit None checks) keep the LIT002 total within the
budget without a mutable-ok

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
072898ab19 fix(mcp): satisfy type discipline gate and the merged input-schema key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
c2f85ca2e7 fix(mcp): share the allowed lookup between the sign-in and exchange preflights
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
e6a6e3ec35 style(mcp): format the new connect sign-in preflight tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
0099b96ddc feat(mcp): challenge rejected caller sign-in subjects at connect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
fcfd7b4f14 fix(mcp): name the connected segment in sign-in challenge resource metadata
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
6a71a3b1a0 test(mcp): add integration coverage for the caller sign-in gates
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
692086a0ae fix(mcp): restore the raw bearer hook kwarg, exact-name-first resolution, and the base OBO challenge gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
4be107e779 fix(mcp): flatten the sign-in merge without stacked comprehension clauses
LIT014 budgets one for clause per comprehension; chain.from_iterable keeps
the dedupe immutable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00
yucheng
df076ce24c fix(mcp): keep the exact-name fallback in the challenge resolver and reformat
get_mcp_server_answering_to now falls back to get_mcp_server_by_name when
no published prefix form matches, preserving the exact-name lookup the
preemptive path had before the router-equivalent resolver. Applies ruff
format to caller_sign_in.py and agent_365.py per the lint gate.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 02:01:55 +00:00