Commit graph

14064 commits

Author SHA1 Message Date
yuneng-jiang
c168199e33
test(ci): repair stale tests and move retired OpenAI text-completion fixtures (#43958)
* test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups

Request fakes now carry the scope a real Starlette request has, the GCS pub/sub
spend-log golden gains the agent identity keys from #43722, the auto-router
session tests follow the baseline_models contract from #43348, and the
Interactions spec checks resolve the create body and resource paths from the
live spec instead of hardcoded names

* test(ci): move retired OpenAI text-completion fixtures to live vehicles

OpenAI still serves native /v1/completions on the gpt-5.4 family, so the
single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt
batches and echo with logprobs now 500 on every OpenAI model, so those cases
keep the same text-completion-openai transport pointed at Fireworks, which
documents both. The optional-params test asserts the request body actually
sent instead of a success callback whose assertions were swallowed

* test(ci): use a serverless Fireworks model for the text-completion batch and echo cases

gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not
deployed; glm-5p3-flash is listed as serverless

* test(ci): skip the ROI calculator repository listing in the security route sweep

GET /roi-calculator/repositories (#43669) lists repositories from the configured
GitHub API, api.github.com by default, so the S2 sweep's GET of every route made
the owned proxy reach an external host and failed the egress check in 31
integration-security tests. It joins /get/latest_release_info in the deny list
2026-09-30 19:19:59 -07:00
devin-ai-integration[bot]
0c515ed7a8
feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (#43063)
* feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token

Stamp metadata.used_client_oauth_token where the proxy decides to forward a
client's Anthropic OAuth token, carry it through StandardLoggingMetadata into
the spend log row, add a used_client_oauth_token filter to /spend/logs/ui, and
surface it on the Logs page as a Credential filter and drawer field. The token
itself never reaches the log

* fix(proxy): carry used_client_oauth_token onto failure spend rows for litellm_metadata routes

* fix(proxy): resolve used_client_oauth_token against the provider the call was sent to

* fix(proxy): keep the proxy's used_client_oauth_token stamp on failure rows and move the resolver under llms/anthropic

* fix(logging): read used_client_oauth_token from the proxy-stamped metadata slot

On routes that carry proxy metadata in litellm_metadata, metadata is the
caller's own body field, and merge_litellm_metadata lets it win. Resolve the
flag from litellm_metadata when the proxy stamped it there so a caller cannot
set it in the standard logging payload

* fix(spend-logs): read used_client_oauth_token from the bucket the route stamped

A guardrail on the unified path adds litellm_metadata to a chat request after
the proxy stamped metadata, so both spend row writers read the new bucket and
stored null. The success row now resolves the flag the same way the callback
payload does, and the failure row picks the bucket from the request route.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-30 19:17:01 -07:00
devin-ai-integration[bot]
a3a7650569
fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (#43956)
* fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): patch the shared proxy logger directly in the straiker api_version test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): keep the straiker stray-version block marker separate from the shared block marker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): drop redundant comments on the straiker api_version integration tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 18:20:57 -07:00
devin-ai-integration[bot]
6997223068
fix(grayswan): send request conversation and tool calls to post-call monitor (#43770)
* fix(grayswan): send request conversation and tool calls to post-call monitor

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(grayswan): tighten post-call context typing and wire test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grayswan): resolve post-call surface from request route before call_type

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grayswan): omit tools from post-call monitor when request context is empty

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(grayswan): apply ruff format to post-call context changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grayswan): merge response text and tool calls into one assistant monitor message

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grayswan): only merge tool calls into the response text for single-choice responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): audit post-call context across endpoints, modes and outages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): share the upstream model probe reply across audit responders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): assert the full generic guardrail body and kill a real serving worker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): normalize the client user agent in the generic body assert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): normalize accept-encoding in generic body assertion

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): keep volatile header placeholders only when the header is present

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): capture monitor calls immutably in the unit test client

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(grayswan): type the test helper parameters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 17:52:53 -07:00
yuneng-jiang
f553c80cd3
test(proxy): scope user_api_key_auth overrides in proxy_server tests (#43952)
TestPriceDataReloadAPI, TestPriceDataReloadIntegration and TestInvitationEndpoints set
app.dependency_overrides[user_api_key_auth] and never removed it. Under xdist the shared
proxy app kept the override, disabling auth for later tests on the same worker and failing
test_harness_smoke.py::test_auth_as_cleans_up_on_exit. Set it through monkeypatch.setitem
so it is undone at teardown

A per-test leak check over every tests/test_litellm/proxy file that touches
dependency_overrides found these 20 tests as the only leakers; it reports none after this change
2026-10-01 00:27:53 +00:00
moe-berri
6fd9334751
feat(lens): analyze agent activity with a separate worker (#43889)
* feat(tracing): bring current ingestion prerequisite onto main

Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.

* feat(lens): add trace analysis and standalone worker

* fix(lens): clarify review limits and finalize main integration

* fix(lens): simplify worker setup and show the next check

* fix(lens): simplify analyzer setup and resolve integration failures

* fix(lens): preserve durations and evidence from later trace reads

* fix(lens): trust server context for internal analysis exclusion

* fix(lens): pin reviewed analyzer image and verify request inclusion

* test(lens): select time units before entering custom duration

* test(lens): allow the standalone analyzer lifetime HTTP client

* test(lens): run analyzer tests in active proxy coverage shard
2026-09-30 22:42:09 +00:00
ryan-crabbe-berri
632b69b5c8
refactor(proxy): answer every team access check with TeamAccess.allows (#43364)
* refactor(proxy): route every team-admin decision through auth/team_access.py

Move the six team-admin helpers out of common_utils, team_endpoints and
key_management_endpoints into litellm/proxy/auth/team_access.py under public
names, and point every management route and helper at them. The key routes
keep checking team admin before org admin, so a team admin whose user row is
gone still passes as before. Status codes and bodies are unchanged, which the
223-case team-admin matrix confirms at the merge base and at the tip

common_utils keeps `_is_user_team_admin` as an alias because the published
litellm-enterprise 0.1.71 wheel still imports it from there

* refactor(proxy): answer every team access check with TeamAccess.allows

Replace the six helpers in auth/team_access.py with one resolver in
litellm/proxy/management/teams/access.py. Each route passes the roles it
accepts (TEAM_OR_ORG_ADMIN or TEAM_ADMIN_ONLY), and /team/update and
/team/info rank roles through strongest_role so org admin still outranks
team admin there

The org lookup moves behind an OrgRoles protocol, implemented by
PrismaOrgRoles in management/users/service.py, and get_team_access in
management/teams/dependencies.py is the only place that reads proxy_server
globals. _check_key_admin_access keeps its name and body from main

Routes that checked org admin first now read the roster first, so a team
admin whose org lookup errors now passes on /team/delete, /team/block,
/team/unblock, member reset_spend and reset_budget, and the team callback
routes. No allowed caller is denied
2026-09-30 15:27:33 -07:00
joshua-berri
52b9fa2ba1
feat(agents): add identity registration and dashboard controls (#43723)
* feat(agents): identity registration and dashboard

* fix(agents): preserve retired identity ownership

* fix(agents): preserve configuration during identity updates

* fix(agents): retain intentional card edits in the dashboard

* fix: remove mutable agent identity registration constructions

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-30 15:21:21 -07:00
ryan-crabbe-berri
f39c811d34
fix(packaging): keep wheel paths under Windows MAX_PATH for Store Python (#43903)
* fix(packaging): keep wheel paths under Windows MAX_PATH for Store Python

pip install litellm fails on Microsoft Store Python because its user
site-packages is already 134 chars plus the profile name, and the content
filter guardrail ships YAML five directories deep under
litellm/proxy/guardrails/guardrail_hooks/litellm_content_filter/. The
existing wheel guard assumed a 100-char install prefix, so it never saw it.

Move categories/ and policy_templates/ to
litellm/proxy/guardrails/content_filter_data/ and drop the benchmark
fixtures from the wheel. Old category_file paths keep resolving because
the resolver only keys on the trailing categories/<file> or
policy_templates/<file> suffix.

Derive the guard's worst-case prefix from the Store Python site-packages
path with a 15-char profile name (149), fail files at 260 and directories
at 248 (CreateDirectoryW), and fix the off-by-one that let a 260-char
path through.

Fixes #43851

* ci: run the Windows wheel install guard on pull requests

The two Windows jobs live in CircleCI, which never runs on pull requests,
so nothing installs the wheel on Windows before merge. Add a GitHub Actions
job on windows-latest that builds the wheel and runs the guard.

Two things make the run deterministic instead of image dependent. The job
turns the LongPathsEnabled registry key off first, because runner images
ship with it on and python.exe is long-path aware, so a 300-char path would
install fine. The guard installs with pip instead of uv, because uv writes
files from Rust, which switches to extended-length paths on its own and can
never hit MAX_PATH.

* fix(guardrails): keep the old content filter package dir as a category search root

Deployments that copied their own category YAML into
guardrail_hooks/litellm_content_filter/ before the data move would have had
that file rejected by the new directory jail and missing from by-name loads,
inherit_from lookups, the UI category listing and the category YAML endpoint.
Every lookup now searches the bundled data dir first and the old package dir
second, with the bundled copy winning on a name clash.

* fix(guardrails): resolve category files through safe_join

By-name category lookups and the suffix search in the category_file resolver now go through safe_join, so a name or suffix that would escape its data root never reaches the filesystem. The LITELLM_CONTENT_FILTER_ALLOW_EXTERNAL_PATHS opt-out keeps its unjailed search. Clears the two CodeQL path-injection findings on the new lookup code.

* fix(guardrails): keep symlinked category files loadable by name

By-name category lookups resolved symlinks through safe_join, so a category file symlinked into the categories folder from elsewhere stopped loading. Those lookups now only reject names that leave the folder lexically and return the link untouched, matching how by-name loads behaved before the data move. The category_file resolver keeps its realpath jail as before.

* fix(guardrails): keep the category viewer inside the category folders

GET /guardrails/ui/category_yaml/{name} hands raw file contents to any valid key, and on main it refused a symlink whose target left the categories folder. The previous commit let by-name lookups follow symlinks again, which also let the viewer read whatever a symlink in a legacy categories folder pointed at. The viewer now checks the found file's real path against every categories folder it searches and answers 400 as before, while the guardrail's own by-name loads keep following symlinks

The roots come in through a FastAPI dependency so the check is testable against a temp folder, and the content filter's realpath containment moves to path_utils.is_within so both surfaces share it. The test that patched os.path.commonpath covered a branch that no longer exists and goes with it

* ci: drop the Windows wheel install job from pull requests

The job took about 13 minutes on every PR to guard an edge case. The
guard still runs its path-length check on Linux in base_sdk_install and
on Windows in the CircleCI windows_release_wheel job.
2026-09-30 22:20:36 +00:00
yujonglee
629c2b5808
feat(tracing): store spend in ClickHouse automatically (#43928) 2026-09-30 22:17:43 +00:00
devin-ai-integration[bot]
e662772ad1
fix(proxy): register a UI-configured arize callback next to otel under OTel v2 (#43906)
Callbacks saved from the Admin UI reach litellm through _add_custom_logger_callback_to_specific_event, which skipped the freshly built logger whenever a callback of the same exact class was already registered. With LITELLM_OTEL_V2 enabled every OTel preset (otel, arize, ...) is an OpenTelemetryV2, so a UI-added arize was treated as a duplicate of the yaml otel callback and never attached, and no trace ever reached Arize

The exists check now compares the exact class and the logger's callback_name, so presets that share a class register side by side while a true re-registration of the same preset is still skipped

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 14:58:40 -07:00
devin-ai-integration[bot]
fc8f3a26bb
fix(proxy): relay Azure passthrough body model groups through the router (#43896)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 14:57:31 -07:00
yuneng-jiang
a6f6c64b6e
feat(ui): adopt the new LiteLLM logo and monogram (#43913)
* feat(ui): adopt the new LiteLLM logo and monogram

Swap the bundled admin UI logos for the new brand assets: the primary
logo in blue for light mode and white for dark mode, and the monogram for
the collapsed sidebar, favicons, and the built-in guardrail cards.

/get_image gains a variant=monogram query parameter so the collapsed
sidebar can request the monogram while admin-configured UI_LOGO_PATH /
UI_LOGO_PATH_DARK logos still take precedence. The bundled light logo
moves from JPEG to a transparent PNG.

* test(ui): query collapsed sidebar logos by role to stay within the lint budget

* fix: point remaining logo consumers at the new bundled assets

The Rust gateway UI served /get_image from the removed litellm_logo.jpg,
the non-root get_image tests pinned logo.jpg, and a cookbook script read
litellm/proxy/logo.jpg. Point them at the monogram and logo.png.

* fix(mcp): serve the BYOK OAuth page logo from /get_image

The page pointed at /ui/assets/logos/litellm_logo.jpg, which the rebrand
removes from the dashboard sources, so the next UI build would drop it.
/get_image?variant=monogram is always served by the proxy and follows any
admin-configured logo.

* fix(ui): invert the LiteLLM monogram on dark guardrail cards

The blue monogram has a transparent train cut-out, so on a dark card it
read as a muddy blue block. Inverting it yields the brand's white mark,
which the logo guidelines prescribe for dark backgrounds.

* fix(gateway-ui): serve theme and variant aware logos from the dashboard export

The Rust gateway served one monogram for every /get_image request, and the
committed export lacked it, so /get_image returned 404 until the next UI
release build. Pick the full or monogram logo in light or dark from the
query, ship those assets in the dashboard's public dir and the committed
export, and drop two comments that restated asserted paths.
2026-09-30 14:51:06 -07:00
yujonglee
268eb4d6e6
feat(tracing): add OTLP trace ingestion and reads (#43915) 2026-09-30 21:12:29 +00:00
devin-ai-integration[bot]
f285229b51
fix(proxy): delete large teams without per-member transaction fan-out (#42998)
* fix(proxy): delete large teams without per-member transaction fan-out

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): evict email-only member caches and reset team members metric on delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep new delete-team literals within the LIT002 ceiling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve deleted-team member ids before the locked delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup

`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.

* fix(repositories): slice find_by_emails into bounded IN statements

The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.

* fix(proxy): delete a team once when /team/delete repeats its id

The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.

* test(integration): audit cells for /team/delete on large, legacy and concurrent teams

Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.

Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.

Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.

* test(integration): pin each chaos outage to a live /team/delete

The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 13:49:18 -07:00
joshua-berri
405ed414cb
feat(agents): authenticate Entra identities and delegated requests (#43722)
* feat(agents): authenticate Entra identities and delegated requests

* fix(agents): enforce target policy and preserve trusted authentication

* fix(proxy): make inference model selection exhaustive

* fix(agents): authorize targets against current database policy

* fix(proxy): make exhaustive model resolution return explicit

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-30 13:05:22 -07:00
devin-ai-integration[bot]
264b09ac8d
fix(responses): scan and mask top-level instructions with guardrails (#43629)
* fix(responses): scan and mask top-level instructions with guardrails

The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.

Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.

Resolves LIT-8931

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): reject empty guardrail rewrites instead of forwarding raw input

An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the texts-replacing guardrail helper explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): cover skip_system_message_in_guardrail on the live proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system

Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): annotate new guardrail tests with return types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the guardrail test doubles explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 11:44:35 -07:00
devin-ai-integration[bot]
3930c5bab6
fix(proxy): strip caller credentials from websocket passthrough (#43855)
* fix(proxy): strip caller credentials from websocket passthrough

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover configured x-api-key in websocket passthrough credential test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 11:22:25 -07:00
joshua-berri
79756cbb9b
feat(agents): enforce authoritative agent permissions (#43721)
* feat(agents): authoritative permissions

* fix: enforce authoritative managed agent permissions

* fix(agents): only consult the identity store for managed targets

is_agent_allowed entered the identity-store path whenever a prisma client
was configured, so an ordinary agent paired with an internal user returned
503 instead of 200. Classify the target from the registry first and fall
back to the store only when the registry has no entry, so an unmanaged
target never depends on the store being reachable.

* fix(agents): gate the managed path on an admitted policy object

Ten call sites branched on `managed_agent_policy is not None`, which any
MagicMock attribute satisfies, so the managed path fired on unmanaged
subjects and died in Pydantic validation as a 503. Route every check
through a shared helper that requires a real AgentResponse.

* test(mcp): stub the writer replica the fresh-policy reads use

reload_admitted_user now passes check_db_only through to get_user_object,
so the user row is read from writer_db. Point the mocks at the replica the
code actually reads and give each parametrized case its own user id.

* fix(agents): cap a managed agent at the invoking team's agents

resolve_agent_access returned the managed policy's grants before the
agent_caller ceiling was applied, so a managed agent acting on behalf of a
user reached agents that user's team was never granted. Intersect with the
caller ceiling the unmanaged path already honours.

* fix(agents): restore token narrowing and scope the private-access suppressions

The managed-model check lost its valid_token narrowing when it moved to the
shared helper. Make the caller-access resolver public rather than reaching
into it from module scope, and give each remaining private access a reason.

* docs(agents): drop the comment claiming admins skip the A2A permission check

The check has never had an admin bypass on this path, so the comment
described behaviour the code does not implement.

* test(proxy): stub the writer reads and restore the MCP manager singleton

Fresh-policy user lookups read writer_db, so the team and rest-endpoint
mocks stubbed a replica the code no longer reads, and the dashboard
session fake still had the pre-kwarg signature. The manager reload also
rebound global_mcp_server_manager in every MCP module without restoring
it, leaking an empty manager into later files.

* style: sort imports under the litellm package ruff config

* fix(mcp): cap a managed agent's servers and tools at the invoking caller

managed_agent_servers and managed_agent_tools returned the agent's own
grants without the agent_caller ceiling the unmanaged resolvers apply, so
a managed agent reached MCP servers and tools the echoed caller could not.
Call the existing ceiling helpers on both axes.

* refactor(mcp): return the caller-capped tools without an interim list

The ceiling helper already returns a sequence, so materializing it into a
list added a mutable collection for nothing. Sort at the return sites
instead, which also makes the tool order stable across both branches.

* fix(agents): preserve actor ceilings during managed target checks

* fix(agents): keep managed permission ceilings authoritative

* fix(mcp): fail closed on authoritative caller team outages

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-30 11:11:37 -07:00
yucheng-berri
7204942756
fix(proxy): keep request-body credentials out of stored spend-log requests (#43635)
* fix(proxy): keep request-body aws credentials out of stored spend-log requests

* fix(proxy): redact every credential-named request-body field in stored spend-log requests

Replace the hard-coded AWS key check in the spend-log request-body sanitizer with
SensitiveDataMasker's key classification, so Azure, Vertex, watsonx, OCI, GigaChat,
Gemini and header credentials are redacted too. Proxy-stamped key identity metadata
is kept.

* fix(proxy): keep request identifiers named like keys in stored spend-log requests

* refactor(proxy): drop the AWS-only snapshot exclusion now that spend-log redaction is name-based

* refactor(proxy): use SensitiveDataMasker's key classification without an exclusion list

* refactor(proxy): always redact credential-named fields in stored spend-log payloads
2026-09-30 17:31:48 +00:00
devin-ai-integration[bot]
82d8b3797c
fix(proxy): attribute completed batch cost rows to /batches in daily activity (#43870)
* fix(proxy): attribute completed batch cost rows to /batches in daily activity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): wait for priced batch tokens before asserting team endpoint activity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 16:16:52 +00:00
devin-ai-integration[bot]
b71f02dbcf
fix(ui): keep MCP permissions visible after key, team and MCP server saves (#43810)
* fix(ui): keep MCP permissions visible after key, team and MCP server saves

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): type the object_permission include as a prisma TypedDict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): do not block key save confirmation on cache refetch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 22:49:20 -07:00
devin-ai-integration[bot]
61a73c59b0
fix(proxy): look up hashed key names with two spend log rows per key (#43656)
* fix(proxy): look up hashed key names with two spend log rows per key

The spend-log fallback for keys missing from the key table read every row per key to check that all named rows agreed, which passed the 5s statement timeout on busy keys even with the (api_key, startTime) index. Probe only the oldest and newest named row per key, so the lookup stays two index reads per key however much the key logged.

* fix(proxy): cap each spend log name probe at 100 rows per key

* fix(proxy): bound the newest-row probe at where the oldest probe stopped

The newest-row probe now starts at the row where the oldest-row probe gave up, so a key with under 200 rows in the window is read once instead of twice, and the lookup transaction turns bitmap scans off so the planner walks the (api_key, startTime) index instead of every row of a busy key when statistics or the visibility map are stale.

* test(integration): add spend log alias probe cells for the daily activity routes

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-29 20:21:55 -07:00
devin-ai-integration[bot]
8afabe81f1
fix(ui): surface x-litellm-call-id in Logs search, table and drawer (#42436)
* fix(ui): surface x-litellm-call-id in Logs search, table and drawer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate api types for spend logs search description

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): drop redundant comments from the call id logs helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(e2e): format logs call id helper and spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep one id per Logs row, move x-litellm-call-id to hover and drawer

The Request ID cell shows only request_id again. When the row's litellm_call_id
differs, the cell tooltip lists it as x-litellm-call-id with its own copy button,
and the drawer header labels the second line x-litellm-call-id: instead of the
call id caption. Stacking two ids in every row made the column noisy for the
common case where the viewer only needs the row they searched for.

* test(e2e): cover the Request ID tooltip hover and copy path

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: poll the clipboard after the tooltip copy and drop a jsdom aside

The e2e read navigator.clipboard right after the click, so a slow async write
could fail the check even though copy works. The unit test's fireEvent choice
(jsdom has no layout, so a real pointer move off the trigger closes the tooltip
before the click lands) is documented here instead of inline.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 02:16:56 +00:00
yuneng-jiang
d098b02ed9
fix(auth): give UI/CLI session tokens their own AES-GCM context and header-safe shape (#43790)
Some checks are pending
Unit Tests / misc (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* refactor(auth): bind UI/CLI session tokens to their own AES-GCM context

UI and CLI session tokens are now always encrypted with AES-256-GCM and a
fixed session associated-data value, and the session-token check only accepts
AES-GCM values carrying that same value. Stored secrets keep their current
encryption and decrypt unchanged, so nothing needs migrating.

encrypt_value_helper and decrypt_value_helper take an optional aad. XSalsa20
cannot bind associated data, so an AAD-bound value is always written as
AES-256-GCM, and an AAD-bound decrypt refuses the legacy format.

Session tokens issued before the upgrade stop validating, so UI and CLI users
sign in once more after upgrading.

* test(e2e): cover real SSO login through the dashboard and the lite CLI

Adds two specs under tests/e2e/ui/oidc, run by playwright.oidc.config.ts
against a live Keycloak stack. The dashboard spec checks that the SSO
session authorizes the Virtual Keys and Models data requests. The CLI
spec runs a real lite login in an isolated HOME with the keyring
disabled, then lists models and sends one chat completion with the
stored session. The main Playwright config now ignores oidc/.

* fix(auth): encode UI/CLI session tokens as unpadded base64url

Session tokens carried the v2:gcm: storage prefix and base64 padding. Basic-auth parsers split on the first colon and browsers reject ':' and '=' in WebSocket subprotocols, so Langfuse pass-through and the realtime playground could not use them

Tokens are now plain unpadded base64url, the same header-safe shape as any bearer token

* fix(auth): prefix UI/CLI session tokens with litellm_login_

A prefix-less token starts with sk- about once in 262,144 logins and is then routed as a virtual key, so that login gets a 401. The prefix also makes session tokens easy to spot in logs

The prefix doubles as the token's AES-GCM associated data, so the visible kind and the encrypted kind cannot disagree

---------

Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-29 18:42:24 -07:00
joshua-berri
6684256136
feat(agents): add identity storage and validation contracts (#43720)
* feat(agents): identity storage and contracts

* fix(agents): cache positive identity lookups with fresh policy checks

* test(agents): include identity attribution in spend fixture

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-29 17:59:21 -07:00
devin-ai-integration[bot]
13d004fc5a
perf(proxy): refresh auth management objects through the request Redis pipeline (#43776)
Identity objects (key, end user) load through the request MGET and their write-backs, the registry
reads and the management-object SETs ride the request pipeline. A team refresh invalidates its alias
with a pipelined DEL instead of a synchronous DEL plus a duplicate async one, and an MGET miss is
remembered so no per-key GET follows it in the same request.

Resolves LIT-9012

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Co-authored-by: yassin <yassin@berri.ai>
2026-09-29 17:56:31 -07:00
devin-ai-integration[bot]
c129ea4fc9
fix(mcp): scope OpenAPI listings to the exact server prefix and drop upstream OAuth metadata when a server is saved (#43608)
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates

Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys

An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.

The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.

Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep OAuth metadata generations only while a fetch is in flight

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 17:25:39 -07:00
devin-ai-integration[bot]
ffb15f946f
perf(proxy): one request-scoped Redis pipeline for auth, spend, rate-limit and routing reads (#43407)
RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua
scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis
Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth
identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing
read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded
on rejection; local cooldowns win over the prefetch.

The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection)

Resolves LIT-8882

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 16:42:13 -07:00
devin-ai-integration[bot]
f5a1c9f1f1
fix(proxy): recover session key owners from daily spend for usage attribution (#43642)
* fix(proxy): recover daily spend key owners

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): simplify daily spend owner recovery

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy): format daily activity metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover recovered owner metadata merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): bound the daily spend owner lookup with the statement timeout

* test(integration): audit the daily activity key owner fallback on every usage route

Thirty five integration cells under tests/integration/spend cover the daily
spend owner fallback on all nine daily activity routes and /usage/ai/chat:
the happy path per route, the unanimity rules (two users, blank and null
rows, an owner the user table lacks, live and deleted keys with and without
their own user, a spend log alias), a non admin reader, an invalid key, a 5 KB
key, a locked LiteLLM_DailyUserSpend, 300 keys of one team, repeated reads, a
second user landing between reads, a concurrent burst across the unified
endpoints, a killed worker, and a proxy restart

The traffic cells ignore the GET /v1/models call the proxy's five minute token
limit refresh makes to every registered OpenAI compatible deployment, since it
lands on a test's provider wire whenever the refresh instant falls inside the
test

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-29 15:36:09 -07:00
devin-ai-integration[bot]
2d034bb35b
perf(proxy): hold one spend counter batch across admission and across post-call accounting (#43369)
Auth's spend counter MGET scope spans common checks, model budget check and reservation;
reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary
increment pipeline and update_cache uses one batched read. Over-budget reservation counters are
charged one at a time so a rejection never touches the counters after it; post-call counter keys are
derived from ids without validating a UserAPIKeyAuth.

Resolves LIT-8881

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 15:05:33 -07:00
devin-ai-integration[bot]
e814532033
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): satisfy type-discipline and strict ruff budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): stamp the served service_tier on every Responses bridge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): cover anthropic and responses served-tier billing paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): bill disconnects through the router's anthropic stream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: apply ruff format to the anthropic stream changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(coverage): ignore delegating properties the ast scan cannot see

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: keep the cast-ok reasons on the cast call line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(streaming): parameterize delegated chunks and messages types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): follow the anthropic pass_through rename after merging main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drain the logging worker between response cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): keep the served service_tier on streamed chunks and bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): type the served service_tier chunk without a loose kwargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill the served service_tier over the requested one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): drop explanatory comment from the tier resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-29 12:54:17 -07:00
devin-ai-integration[bot]
fb74957ddd
fix(guardrails): enable explicit PANW MCP output scanning (#43109)
* fix(guardrails): declare post_mcp_call for PANW Prisma AIRS

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): exercise post_mcp_call_hook dispatch in PANW tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-catalog): add fal_ai resolution-tiered image cost fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep post_mcp_call opt-in for PANW Prisma AIRS

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-catalog): add fal_ai resolution-tiered image cost keys to cost map schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: joshua <joshua@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-29 12:50:08 -07:00
tin-berri
3b2a447fae
fix(autorouter): compare historical and new savings consistently (#43348)
* fix(autorouter): compare historical and new savings consistently

* fix(autorouter): reject comparisons if request counts changed

* fix(router): restore eligible LLM and classification breakdown

* fix(router): avoid ambiguous baseline labels for partial comparisons
2026-09-29 12:40:19 -07:00
devin-ai-integration[bot]
abc85c2651
fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614)
* feat(cost_calculator): add cost_per_second for chat per-second pricing

Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins

Move Bedrock commitment rows to cost_per_second so they bill once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop legacy per-second fields from chat paths

Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost_calculator): recognize output-only per-second rates

Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:27:14 -07:00
yucheng-berri
b3dcf8208d
test(integration): callback credential canary slots C1-C3 and D5 (#43630)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown

* test(integration): callback credential canary slots C1-C3 and D5

Team callback, team callback_settings, config default_team_settings and key metadata.logging Langfuse secrets, a team Datadog dd_api_key, and request-body Langfuse keys (allow_client_side_credentials) must reach only their sink. Each scenario checks its sink received the canary as auth and that the marker is visible at the stored body, the Logs drawer route and the sink. Adds a unit test that the stored request body snapshot carries no callback parameter.

* test(integration): give the callback sink waits a wider bound

* test(integration): sweep provider requests for callback credentials
2026-09-29 10:49:30 -07:00
Itai Modiano
0c553f0398
feat(guardrails): send a configured gateway_name from noma_v2 to Noma (#43678)
* feat(guardrails): send a configured gateway_name from noma_v2 to Noma

The noma_v2 guardrail accepts a gateway_name param, falling back to the
NOMA_GATEWAY_NAME env var. The value is stripped, and when it is non-empty
it goes out as a top-level gateway_name field on /litellm/guardrail. The
param works for both guardrail: noma_v2 and guardrail: noma with use_v2,
and it is appended after the existing constructor params so positional
callers keep their meaning

* chore(ui): regenerate OpenAPI snapshot and dashboard types for gateway_name

The new noma_v2 gateway_name param shows up in the proxy OpenAPI spec, so
the lazy snapshot and the generated dashboard types need regenerating

* Update litellm/proxy/guardrails/guardrail_hooks/noma/noma_v2.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-29 10:20:40 -07:00
devin-ai-integration[bot]
7f95b5f361
refactor: clean up fresh tech debt from 2026-09-28 (#43674)
* refactor: clean up fresh tech debt from 2026-09-28

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor: group leaderboard rows in one pass and wrap docstring at 120

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 02:15:12 -07:00
fedaeho
85dc7cb62e
fix(proxy): resolve model_group_alias in the zero-cost budget predicate (#43512)
`_is_model_cost_zero()` reads a group's cost through `Router.get_model_group_info()`,
which resolves `model_group_alias`, and then gates that on `_is_cost_explicitly_configured()`,
which scanned `Router.model_list` for an exact `model_name` match. Alias names live only in
`Router.model_group_alias` and are never `model_name` entries, so the scan found nothing and
returned False. That False means "the zero cost was defaulted, not configured" (the sparse
auto-registration gate added for #24770), so a model priced explicitly at 0 had budget
enforced against it when requested through an alias, while the same deployment under its own
name was exempt. Both names route to the same deployment and add nothing to spend.

The two lookups in one function disagreeing is the bug, so they now share one resolution:
`_is_cost_explicitly_configured()` resolves through `Router.get_model_list()`, the same
alias-aware path `get_model_group_info()` takes. That also reaches a deployment which prices
itself through its `model_info` block, whose cost-map entry lands under the deployment id.
`_group_declares_explicit_cost()` was an alias-aware copy of this function, wired only into
`model_has_no_cost_mapping()` and never into the budget path; its body is what
`_is_cost_explicitly_configured()` now carries, and both callers share it so the two cannot
drift apart again.

`_has_ptu_flat_cost()` scanned `model_list` the same way and runs after the gate above, so
resolving one without the other would let an aliased PTU group — explicit zero per-token
price alongside a flat capacity cost — pass as free. It resolves the same way now.

Tests cover the predicate and the request path it feeds: over-budget requests through
`_should_skip_budget_checks()` into `common_checks()` for an aliased free model (allowed) and
an aliased paid model (refused), the predicate for free, paid, PTU, hidden and dangling
aliases, and `model_has_no_cost_mapping()` through an alias so the other caller of the shared
check stays covered.

Unchanged: priced groups (the predicate returns False before the gate), unmapped groups whose
zero cost was defaulted (#24770), hidden aliases and aliases pointing at a nonexistent group
(`get_model_group_info()` returns None for both, so the cost is unknown and budget is
enforced), and non-aliased PTU groups.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:46:33 -07:00
ishaan-berri
118ce3cc91
feat: add model leaderboard page (#43649)
* feat(proxy): add model leaderboard analytics

* feat: add model insights task and range constants

* feat: record task type from task tags in model usage rollup

* feat: serve 365 days of model insights by UTC date

* test: cover task tag resolution in model usage rollup

* test: update model insights range limit test to 365 days

* chore: regenerate dashboard api types for model insights

* feat: add model insights aggregation helpers

* test: cover model insights aggregation helpers

* feat: redesign model leaderboard with stacked bars, treemap and ranking

* test: update model leaderboard view test

* feat: mark model leaderboard as beta in sidebar

* chore: sync schema.prisma copies from root

* fix: only treat task: prefixed tags as model insight tasks

* feat: add metric type for model insights ranking

* fix: rank model insights by selected metric and scope detail queries to ranked deployments

* test: plain tags are not model insight tasks

* test: cover metric ranking, deployment scoping and rollup round trip

* fix: build model insights weeks and halves from the requested date range

* test: cover empty weeks and range-based change comparison

* fix: refetch by metric, show load errors and ignore stale responses

* test: cover metric refetch and error state

* feat: define model insight tasks in a JSON file

* feat: return task labels and categories from model insights

* feat: load model insight tasks from JSON

* refactor: validate rollup task tags against the JSON task list

* feat: serve the task list with model insights

* refactor: drop hardcoded task list from constants

* build: ship model insight tasks JSON in the wheel

* test: cover model insight task JSON

* refactor: take task labels and categories from the API

* test: pass task info to task tile builder

* refactor: color treemap by API-provided category

* test: include tasks in model leaderboard fixture

* fix: make daily model usage migration idempotent

* feat: bound the model insights task query size

* fix: compute task breakdown independent of the chart metric

* test: task breakdown is stable across chart metrics

* chore: regenerate lazy openapi snapshot for model insights

* chore: regenerate dashboard api types for model insights

* fix: keep previous ranking dimmed while a new metric loads

* test: cover stale metric state in model leaderboard

* refactor: drop task row cap constant

* fix: return the full task breakdown instead of a truncated one

* test: task query is not truncated

* feat: add task summary types for model insights

* feat: summarise tasks server-side on a separate model insights endpoint

* test: cover the model insights tasks endpoint

* chore: regenerate lazy openapi snapshot for model insights tasks

* chore: regenerate dashboard api types for model insights tasks

* refactor: drop client-side task aggregation

* test: remove client-side task aggregation tests

* feat: load task breakdown separately from the chart metric

* test: task breakdown is not refetched on chart metric change

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-28 19:40:04 -07:00
devin-ai-integration[bot]
ce25856424
feat(mcp): scan and pin upstream tool descriptions (#43283)
Some checks failed
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
* feat(mcp): scan and pin upstream tool descriptions

Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.

* chore: sync schema.prisma copies from root

* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes

* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending

The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.

* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper

* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift

* refactor(mcp): trim the tool catalog guard docstrings to one line

* test(mcp): cover guarded discovery boundaries and response definitions

* fix(mcp): bound discovery guardrail concurrency per catalog

* fix(mcp): scan tool catalogs in bounded parallel batches

* fix(mcp): hide pinned catalogs from restricted management views

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-28 18:38:49 -07:00
joshua-berri
98c710c411
fix(guardrails): preserve Presidio output selection and restoration (#43401)
* fix(guardrails): preserve Presidio output callback intent and tag selection

* test(guardrails): verify Presidio callback stages after registry updates

* fix(guardrails): preserve standalone Presidio token behavior

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-28 17:34:38 -07:00
Yassin Kortam
7e383c9f6a
revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" (#43377)
* Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)"

* revert: "feat(usage): search keys beyond the top-N usage subset (#42827)" (#43378)

* Revert "feat(usage): search keys beyond the top-N usage subset (#42827)"

* revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595)

* Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table"

* Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596)

Co-authored-by: yassin <yassin@berri.ai>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:47:46 +00:00
devin-ai-integration[bot]
317430db4e
fix(panw_prisma_airs): honor experimental_use_latest_role_message_only on every request shape (#42447)
* fix(panw_prisma_airs): apply experimental_use_latest_role_message_only to every request shape

Explicit true/false now applies to chat completions, Anthropic /v1/messages and /v1/responses alike; unset keeps latest-only for Anthropic and full history otherwise. Text indices are mapped back to their source message by value instead of by count, so Responses instructions, function_call_output and reasoning items no longer derail the alignment and silently rescan the whole history

Co-authored-by: scthornton <scthornton@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(panw_prisma_airs): type latest-message helpers against AllMessageValues

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): require forward and reverse text attribution to agree

A Responses function_call_output whose text equals the latest user turn could claim that turn's slot in a forward-only walk and demote the latest-only scan to an earlier message. Walk both directions and fall back to the full role-filter scan when they disagree

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): pick the latest human turn from messages, not from aligned texts

An image-only latest user turn no longer promotes an earlier user turn into the
latest-only scan; it scans nothing on the request side, as the Anthropic path did
before. A latest user/developer message whose text never reached texts (a trailing
Responses reasoning item) falls back to the role-filter scan instead of narrowing.
Types the test helpers, drops the narrating docstrings and adds regressions for both
shapes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): log when latest-only selection leaves nothing to scan

An image-only latest user turn with experimental_use_latest_role_message_only=true intentionally yields zero scanner calls. Emit a debug line naming the call_id so operators can tell this apart from the guardrail not firing.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): keep Responses reasoning items out of latest-turn selection

The Responses translation handler gives reasoning input items the default user
role, so a reasoning item with text content after the latest prompt was picked
as the latest human turn and the real prompt went unscanned under
experimental_use_latest_role_message_only. Map reasoning items back to their
texts positions from the raw input and exclude them; fall back to the
role-filter scan when the raw items do not account for every text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: scthornton <scthornton@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:16:18 +00:00
tin-berri
6d4ccf7e97
refactor(mcp): consolidate hub publication predicate (#43394) 2026-09-28 13:26:59 -07:00
Yassin Kortam
fe76c2473d
Revert "feat(proxy): server-side Team Usage export beyond the top-N key cap (#42996)" (#43376)
This reverts commit 77eccaca78.
2026-09-28 12:33:18 -07:00
yuneng-jiang
f20c400374
fix(proxy): unregister logging callbacks removed from the stored config (#43428)
* fix(proxy): unregister logging callbacks removed from the stored config

POST /config/callback/delete saved the config and resynced, but the resync only
ever added callbacks, so a deleted callback kept exporting and kept showing in
/get/config/callbacks as read-only on every worker.

ProxyConfig now tracks which callback list entries each DB config sync
registered and unregisters them once the stored config stops listing them.
Callbacks it did not register (YAML, code) are never touched, and a failed
config load skips the sync instead of treating the config as empty.

* refactor(proxy): keep callback sync comprehensions to one for clause

* fix(proxy): restore code-registered callbacks the DB sync replaced

Registering a custom-logger callback from the DB swaps an existing string
entry for a logger instance. Deleting the DB entry then removed the instance
and left the code-registered callback gone. The sync now records the entries
it displaced and puts them back when it unregisters.
2026-09-28 12:22:01 -07:00
devin-ai-integration[bot]
3726ce2cfc
refactor(guardrails): fix agent 365 to the production endpoint and log the opt-in fail_open at error level (#43189)
* feat(guardrails): fail open by default when Agent 365 cannot evaluate and count it in Prometheus

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(tests): ruff format the Prometheus fail-open registry test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(guardrails): add Agent 365 authority host override, fail-open integration test and per-guardrail YAML default

Add `authority_host` to the Agent 365 config (also read from AGENT365_AUTHORITY_HOST, then AZURE_AUTHORITY_HOST) so sovereign clouds and the integration test can point the OBO exchange at a different Entra host.

Add tests/integration/mcp/test_mcp_agent_365_guardrail.py, a real proxy test with Postgres, Redis, a scripted MCP upstream and local Entra and Agent 365 doubles covering the default fail-open, explicit fail-closed and fail-open, Defender Skipped, policy denial, persisted status and Prometheus counter.

Use PrometheusLogger.get_instance for the fail-open metric lookup instead of a hand-rolled callback scan. Clarify the config description: gateway credential failures fail open, caller token failures block.

Extract the dashboard YAML preview into teamGuardrailConfigYaml.ts so the effective per-guardrail default is unit tested and the "default" hint only shows when nothing was set explicitly.

Regenerate the lazy OpenAPI snapshot and schema.d.ts for the new field.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): default a scheme-less Agent 365 authority host to https and treat a null fallback as unset in the YAML preview

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for the Agent 365 fail-open default across entry points, Entra faults, throttling and two workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prove both Agent 365 workers serve and that a killed worker is replaced

Each fresh connection reports its worker pid from /debug/memory/summary and its MCP catalog on the same
connection, so the two-worker readiness wait covers both workers by identity. The kill test now kills a
pid the proxy reported as a worker and waits for a replacement pid, instead of the first psutil child

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): drop the prometheus fail-open counter from the agent 365 guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): keep agent 365 fail closed by default and make fail_open an explicit opt-in

Restores the shared unreachable_fallback default and the sibling guardrail initializers, drops the Admin UI YAML preview that only existed for the per-guardrail default, and reworks the unit and integration tests so the default blocks with HTTP 503 while unreachable_fallback: fail_open lets availability failures through as Unscanned

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): append authority_host after the existing Agent365Guardrail parameters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(guardrails): agent 365 fails open by default and hides the production overrides from the UI form

Agent 365 sits in the runtime path of every MCP tool call, so an Entra or
Agent 365 outage now lets the call through unscanned (logged at error level,
recorded as Unscanned with guardrail_failed_to_respond) instead of blocking it.
unreachable_fallback: fail_closed stays as the opt-in strict mode. Policy
blocks, throttling, 4xx rejections and a rejected caller token still block

The shared unreachable_fallback field becomes nullable so each guardrail owns
its default; every sibling still resolves None to fail_closed and typesafe
keeps failing open

api_base, resource_app_id and agent_id have production defaults and leave the
dashboard form (ui_hidden); they stay available in config.yaml and env. The
authority_host override and its env keys are gone, the OBO exchange always
uses login.microsoftonline.com. The integration suite keeps only the cells
that need no Entra double, the evaluation paths live in unit tests with an
injected handler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate openapi snapshot and schema.d.ts for the nullable unreachable_fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): fix agent 365 to the production endpoint and keep fail_closed as the default

Remove api_base, resource_app_id and agent_id from the Agent 365 config model, their AGENT365_* env fallbacks and the _is_ui_hidden helper: the evaluation URL and the Agent Tools app id are fixed production constants and the agent identity is always the caller's key alias. Revert the fail_open default; unreachable_fallback: fail_open stays an explicit opt-in. Restore the shared unreachable_fallback field, the sibling guardrail initializers and typesafe to main. Move the Entra dependent cells from the subprocess integration suite to unit tests with an injected HTTP handler.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): warn when agent 365 yaml still carries the removed override keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): inject the http handler into the agent 365 initializer instead of assigning it after construction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 12:05:47 -07:00
devin-ai-integration[bot]
703eb4fa68
security(proxy): keep team callback credentials out of the stored request body (#43217)
* security(proxy): keep team callback credentials out of the stored request body

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: allow the security conventional commit type

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 11:16:58 -07:00
devin-ai-integration[bot]
f191e08d67
fix(proxy): log key owner identity on expired key auth failures (#43105)
* fix(proxy): log key owner identity on expired key auth failures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): escape control characters in logged key identity fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): gate key identity in auth failure logs behind log_auth_failure_key_identity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): apply log_auth_failure_key_identity from DB config reloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 08:18:48 -07:00