Commit graph

239 commits

Author SHA1 Message Date
devin-ai-integration[bot]
6b9766fa0c
feat(proxy): add native ROI calculator for gateway spend vs merged PRs (#43669)
* feat(proxy): add native ROI calculator for gateway spend vs merged PRs

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): serialize ROI Prisma inputs with builtin containers

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* style(proxy): format ROI calculator backend files

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix: parse fenced ROI estimates and retain completed reports

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(roi-calculator): correct estimator and dashboard behavior

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* chore(ui): drop next dev generated AGENTS.md block

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): chunk ROI spend user lookup

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ui): show reused ROI estimates after sync

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(security): address ROI CodeQL alerts

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(proxy): make ROI calculator unit tests discoverable

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(ci): run ROI calculator tests in proxy infra shard

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(roi): page repository search and recover polling errors

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* feat(roi): bring scheduled analysis and guided setup into the gateway

* fix(roi): recover interrupted syncs and resolve review findings

* fix(roi): preserve cached estimates across report scope changes

* fix(roi): normalize scheduler timestamps to UTC

* fix(roi): fence cancelled syncs and read reports from writer

* fix(roi): preserve reports during metadata outages

* refactor(roi): isolate outage validation and verify uncached retry

* fix(roi): make scheduled job registration repeatable

* style(roi): format scheduler import

* fix(roi): continue syncing accessible repositories

* fix(roi): preserve reports and identity during upstream outages

* fix(roi): persist refreshed identities for reused estimates

* perf(roi): skip writes for unchanged cached identities

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-authored-by: moe-berri <moe@berri.ai>
2026-09-30 14:16:41 -07:00
devin-ai-integration[bot]
f285229b51
fix(proxy): delete large teams without per-member transaction fan-out (#42998)
* fix(proxy): delete large teams without per-member transaction fan-out

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): evict email-only member caches and reset team members metric on delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep new delete-team literals within the LIT002 ceiling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve deleted-team member ids before the locked delete

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve email-only deleted-team members with one case-insensitive lookup

`_deleted_team_member_user_ids` looked each email-only roster entry up with its own
`find_users_by_email` call inside an unbounded `asyncio.gather`: one exact-match query
per email, so a large roster fanned out against the pool again and a roster email that
differed in case from its user row was missed. Add `UserRepository.find_by_emails`, a
single case-insensitive `in` query, and call it once before the locked delete.
`management_helpers/utils.py` goes back to its main-branch shape since the single-email
helper no longer needs exporting.

* fix(repositories): slice find_by_emails into bounded IN statements

The unbounded-IN lint flagged the case-insensitive email lookup added for
/team/delete cache eviction. chunked_in.find_many_in cannot carry Prisma's
insensitive mode, so the repository slices the deduplicated list into
IN_LIST_CHUNK_SIZE statements itself and concatenates the pages. Empty input
still returns () without a query.

* fix(proxy): delete a team once when /team/delete repeats its id

The audit sent {"team_ids": [T, T]}: main answered 400 "User not found in
team" after deleting the keys and memberships and writing two tombstones,
leaving the team row behind; this branch answered 200 but still wrote the
tombstone, audit row and eviction twice. DeleteTeamRequest now collapses
repeated ids in order, so every later step sees each team once and the
response lists each deleted team once.

* test(integration): audit cells for /team/delete on large, legacy and concurrent teams

Thirty-eight deterministic cells in tests/integration/management/ (the CircleCI
integration-management group) covering the /team/delete happy, sad, edge and chaos rows:
250 members against a pool limit of five on two workers, the advisory-lock wait, email-only
legacy roster entries in every casing, member and team cache eviction on both proxies for
every client and endpoint, the Prometheus gauge, audit rows, malformed and duplicate input,
the route gate, and a worker kill, a Redis outage and a proxy restart mid-burst.

Every cell runs against the real proxy, Postgres and Redis with the scripted upstream; no
component is mocked. On the merge base the rows this fix changes are red (P2028 on the
250-member team, two lock waiters, case-mismatched email lookups, duplicate ids, orphaned
LiteLLM_UserTable.teams references under a concurrent burst); on the tip every cell is green
twice with identical selections.

Two pre-existing behaviours are pinned as observed rather than fixed here: a roster entry with
neither user_id nor user_email answers 500, and the LiteLLM_DeletedTeamTable row is committed
before the locked transaction, so a delete that dies in between leaves a tombstone for a live
team and the retry adds a second.

* test(integration): pin each chaos outage to a live /team/delete

The three chaos cells applied the outage once three deletes had answered, which on a fast
run let the whole burst finish before the worker kill, Redis stop or SIGTERM landed, so the
cells passed without exercising the failure. Each cell now holds the first team's advisory
lock from a test-owned transaction, waits until that team's delete is queued behind it in
Postgres with its request unanswered, applies the outage, and only then releases the lock,
so an in-flight delete meets the failure on every run and both legs. The pinned team's
outcome and the number of deletes answered before the outage are recorded as junit
properties (pinned_delete, answered_before_outage).

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-09-30 13:49:18 -07:00
devin-ai-integration[bot]
657bb777fa
fix(hosted_vllm): keep reasoning_content on replayed assistant messages (#43599)
* fix(hosted_vllm): keep reasoning_content on assistant messages in _transform_messages

vLLM accepts reasoning_content (200 on the wire) and qwen/deepseek/glm
chat templates consume it, so popping it made reasoning models lose
earlier reasoning across tool loops. thinking_blocks is still removed
for vLLM compatibility.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(hosted_vllm): forward replayed reasoning_content only when it is a string

* test(integration): cover hosted_vllm reasoning_content replay across endpoints

* test(integration): require the surviving worker to serve its held requests in the sigkill chaos cell

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
2026-09-30 13:12:28 -07:00
devin-ai-integration[bot]
50f5cc9bbb
feat(otel v2): excluded_services opt-out for datastore spans on tenant destinations (#43278)
* feat(otel): excluded_services opt-out for datastore spans on tenant destinations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep upstream support unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): excluded_services resolves from the otel callback config only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): name the otel callback logger so excluded_services owner lookup matches

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): assert no aux datastore traces reach the tenant sink

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): read bogus-start proxy log from the results dir

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): read only this invocation's bogus-start proxy log

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): assert operator kept db spans over the whole recorded window

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): split operator db-span asserts by trace scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): build the otel logger after preset callbacks and validate the exclusion env at boot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): tolerate a bogus exclusion env when callback config wins

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): keep bogus exclusion env fatal when a preset parses it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): hoist the preset check out of the callback loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): use a rule-scoped pyright suppression

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): log and drop unknown excluded_services instead of failing boot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for the operator spend-writer span before checking the tenant for postgres

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): leave callback init and boot untouched when excluded_services is unset

Read callback_settings.otel.excluded_services directly instead of making the otel callback build its own logger, and drop the new boot-time parse of callback_settings.otel, so a proxy without the setting behaves exactly as on main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): normalize callback_settings excluded_services without rereading OTel env vars

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel v2): log and ignore malformed excluded_services instead of failing startup

Lowercase and trim names, drop non-string items, and add an integration matrix over endpoints, clients, cache hits, destination outages and setting shapes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): cover failed upstream calls in the excluded_services matrix

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel v2): pin operator Langfuse credentials in preset-only excluded_services tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
2026-09-30 12:20:27 -07:00
devin-ai-integration[bot]
264b09ac8d
fix(responses): scan and mask top-level instructions with guardrails (#43629)
* fix(responses): scan and mask top-level instructions with guardrails

The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.

Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.

Resolves LIT-8931

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): reject empty guardrail rewrites instead of forwarding raw input

An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the texts-replacing guardrail helper explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): cover skip_system_message_in_guardrail on the live proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system

Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): annotate new guardrail tests with return types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the guardrail test doubles explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 11:44:35 -07:00
devin-ai-integration[bot]
82d8b3797c
fix(proxy): attribute completed batch cost rows to /batches in daily activity (#43870)
* fix(proxy): attribute completed batch cost rows to /batches in daily activity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): wait for priced batch tokens before asserting team endpoint activity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 16:16:52 +00:00
devin-ai-integration[bot]
9dda4d895f
fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates (#43764)
Some checks are pending
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
2026-09-30 07:32:16 -07:00
devin-ai-integration[bot]
6bc17f98d7
refactor: clean up fresh tech debt from 2026-09-29 (#43830)
* refactor: clean up fresh tech debt from 2026-09-29

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(routing): pin usage-based routing Redis reads through the proxy and SDK

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 02:40:37 -07:00
yuneng-jiang
ba6d6d1a95
test(ci): repair MCP Responses and budget fixtures (#43788)
* test(ci): repair MCP Responses and budget fixtures

* test(auth): verify delegated budget changes persist
2026-09-29 23:29:52 -07:00
devin-ai-integration[bot]
61a73c59b0
fix(proxy): look up hashed key names with two spend log rows per key (#43656)
* fix(proxy): look up hashed key names with two spend log rows per key

The spend-log fallback for keys missing from the key table read every row per key to check that all named rows agreed, which passed the 5s statement timeout on busy keys even with the (api_key, startTime) index. Probe only the oldest and newest named row per key, so the lookup stays two index reads per key however much the key logged.

* fix(proxy): cap each spend log name probe at 100 rows per key

* fix(proxy): bound the newest-row probe at where the oldest probe stopped

The newest-row probe now starts at the row where the oldest-row probe gave up, so a key with under 200 rows in the window is read once instead of twice, and the lookup transaction turns bitmap scans off so the planner walks the (api_key, startTime) index instead of every row of a busy key when statistics or the visibility map are stale.

* test(integration): add spend log alias probe cells for the daily activity routes

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-29 20:21:55 -07:00
devin-ai-integration[bot]
c129ea4fc9
fix(mcp): scope OpenAPI listings to the exact server prefix and drop upstream OAuth metadata when a server is saved (#43608)
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates

Discovery-list cache identity now uses the hashed token instead of the raw
api_key and treats MCPJWTSigner-signed servers as per caller. Server
definition changes also drop the cached upstream OAuth metadata. OpenAPI
listings look tools up under the normalized registry prefix with the
separator, so an overlapping sibling prefix no longer leaks into the list.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys

An upstream metadata fetch that started before a server edit could store its stale reply after
invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch
only stores when the generation it captured before I/O is unchanged.

The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no
reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so
both go back to the merge-base behavior.

Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared
server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep OAuth metadata generations only while a fetch is in flight

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 17:25:39 -07:00
devin-ai-integration[bot]
f5a1c9f1f1
fix(proxy): recover session key owners from daily spend for usage attribution (#43642)
* fix(proxy): recover daily spend key owners

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): simplify daily spend owner recovery

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(proxy): format daily activity metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover recovered owner metadata merge

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): bound the daily spend owner lookup with the statement timeout

* test(integration): audit the daily activity key owner fallback on every usage route

Thirty five integration cells under tests/integration/spend cover the daily
spend owner fallback on all nine daily activity routes and /usage/ai/chat:
the happy path per route, the unanimity rules (two users, blank and null
rows, an owner the user table lacks, live and deleted keys with and without
their own user, a spend log alias), a non admin reader, an invalid key, a 5 KB
key, a locked LiteLLM_DailyUserSpend, 300 keys of one team, repeated reads, a
second user landing between reads, a concurrent burst across the unified
endpoints, a killed worker, and a proxy restart

The traffic cells ignore the GET /v1/models call the proxy's five minute token
limit refresh makes to every registered OpenAI compatible deployment, since it
lands on a test's provider wire whenever the refresh instant falls inside the
test

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-29 15:36:09 -07:00
devin-ai-integration[bot]
e814532033
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): satisfy type-discipline and strict ruff budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): stamp the served service_tier on every Responses bridge chunk

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): cover anthropic and responses served-tier billing paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(service-tier): bill disconnects through the router's anthropic stream wrapper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: apply ruff format to the anthropic stream changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(coverage): ignore delegating properties the ast scan cannot see

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style: keep the cast-ok reasons on the cast call line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(streaming): keep service_tier on OpenAI-compatible parsed chunks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(streaming): parameterize delegated chunks and messages types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(tests): follow the anthropic pass_through rename after merging main

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(anthropic): drain the logging worker between response cache tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): keep the served service_tier on streamed chunks and bill it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(databricks): type the served service_tier chunk without a loose kwargs dict

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill the served service_tier over the requested one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): drop explanatory comment from the tier resolution

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: kerry <kerry@berri.ai>
2026-09-29 12:54:17 -07:00
devin-ai-integration[bot]
abc85c2651
fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614)
* feat(cost_calculator): add cost_per_second for chat per-second pricing

Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins

Move Bedrock commitment rows to cost_per_second so they bill once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop legacy per-second fields from chat paths

Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost_calculator): recognize output-only per-second rates

Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(pricing): cover cost_per_second and legacy per-second aliases through the proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:27:14 -07:00
yucheng-berri
336c7c0849
test(integration): sweep proxy logs, metrics, a Datadog intake and the Logs drawer for credential canaries (#43306)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): sweep proxy logs, metrics, a gzip Datadog intake and the Logs drawer for credential canaries

* test(e2e): treat an unset prompt-storage setting as unset and restore it

* test(integration): name the Datadog sink slot G1d

* test(e2e): search the Logs page for base64 forms of the deployment key

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): pass the resolved deployment id to the Datadog route sweep

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown
2026-09-29 17:57:28 +00:00
yucheng-berri
cede93e826
test(integration): request-path credential canary slots D1-D4 (#43307)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): request-path credential canary slots D1-D4

* test(integration): read the Logs drawer and spend-log filter for failed request rows

* test(integration): check the marker in each failed row's spend-log filter; run header slots on chat-family routes

* test(integration): run D2 on embeddings again; only the client-header slot runs on chat-family routes

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): pass the slot deployment's model_info id to the route sweep

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown
2026-09-29 10:52:13 -07:00
yucheng-berri
b3dcf8208d
test(integration): callback credential canary slots C1-C3 and D5 (#43630)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown

* test(integration): callback credential canary slots C1-C3 and D5

Team callback, team callback_settings, config default_team_settings and key metadata.logging Langfuse secrets, a team Datadog dd_api_key, and request-body Langfuse keys (allow_client_side_credentials) must reach only their sink. Each scenario checks its sink received the canary as auth and that the marker is visible at the stored body, the Logs drawer route and the sink. Adds a unit test that the stored request body snapshot carries no callback parameter.

* test(integration): give the callback sink waits a wider bound

* test(integration): sweep provider requests for callback credentials
2026-09-29 10:49:30 -07:00
devin-ai-integration[bot]
d46304900f
fix(router): stream /v1/messages lifecycle frames live when no fallback can take over (#43600)
* fix(router): stream anthropic messages lifecycle frames live when no fallback can take over

The /v1/messages streaming wrapper buffered message_start and
content_block_start until the first content_block_delta and dropped
pings behind buffered frames unconditionally, even for requests no
fallback could ever recover. With adaptive thinking on Bedrock or
Vertex the client saw no bytes for the whole thinking pass and hit
read timeouts.

Buffering now applies only while a fallback can still take over
(generic or refusal chain resolving), and a ping is always forwarded
live since it carries no lifecycle and keeps the connection alive.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): mirror every dispatcher fallback path in the anthropic stream gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): skip already-tried order levels in the anthropic stream gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(router): keep a transport-split ping behind buffered lifecycle frames instead of forwarding its head live

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(router): credit the #39566 branch this fix supersedes

Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Radu Swigler <radu.porumba@gmail.com>
2026-09-29 09:22:05 -07:00
devin-ai-integration[bot]
9dfa42dcde
refactor(types): replace Any with proven types in 7 files (#43704)
* refactor(types): replace Any with proven types in 11 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert Any changes that broke existing callers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): drop prompt factory helper wrappers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover typing sweep surfaces

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): tighten sweep audit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 06:12:58 -07:00
devin-ai-integration[bot]
39d14bd855
feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers (#43641)
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): drive the router request test through an httpx MockTransport

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(fireworks_ai): assert tool definitions reach the router upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 19:08:16 -07:00
yucheng-berri
3572d359a1
test(integration): stored-config credential canary slots (#43309)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): stored-config credential canary slots

Add canary slots for credentials the proxy holds in its env, config or
database: virtual key raw value, master key, deployment api_key via
/model/new, named credentials, AWS secret key, Vertex service-account
JSON and its minted token, team model_config credential overrides,
config guardrail api_key, and sink credentials from env.

* test(integration): resolve the config guardrail id and require detail routes to return 200

* test(integration): drop repeated timeout comments

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown

* test(integration): sweep config-deployment routes with the real model id and use the rig admin for the master-key slot

* test(integration): mark the configure-hook config edits as intended

* test(integration): drop suppression markers that suppress nothing

* Use claude-haiku-4-5 for the Bedrock stored-credential test model
2026-09-28 17:23:30 -07:00
yucheng-berri
db9c307e1a
test(integration): credential canary slots for MCP and pass-through credentials (#43308)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): credential canary slots for MCP and pass-through credentials

Adds slots F1 (MCP static auth), F2 (per-user MCP OAuth token), F2E (per-user
MCP env var), F3 (x-mcp client auth header), H1 (pass-through credential header),
H2 (vector store api_key) and H2S (search tool api_key) to the credential canary
suite. The OAuth double gains an optional mint hook so a test can choose the
issued access token.

* test(integration): wait for MCP spend rows by call type

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): canary MCP and pass-through slots pass resolved ids

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown
2026-09-28 17:10:51 -07:00
yucheng-berri
dab2deb5ed
test(integration): credential canary suite harness (#43300)
* test(integration): credential canary suite harness

Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.

* test(integration): widen canary route sweep and harden the rig

Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.

* test(integration): descend into any decoded value that can still hold an encoded canary

* test(integration): bound canary decoding by depth and decoded bytes

* test(integration): scope log-table and spend-log reads to the scenario window

* test(integration): sweep spend-log rows in the scenario date window

* test(integration): keep spend-log date window summarized

* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot

* test(integration): expect 404 from the caller-scoped team membership route

* test(integration): use the rig's own master key and expect 404 from submission lookups

* test(integration): check the overridden rig key without assuming the default key is unknown
2026-09-28 16:12:08 -07:00
Yassin Kortam
7e383c9f6a
revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" (#43377)
* Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)"

* revert: "feat(usage): search keys beyond the top-N usage subset (#42827)" (#43378)

* Revert "feat(usage): search keys beyond the top-N usage subset (#42827)"

* revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595)

* Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table"

* Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596)

Co-authored-by: yassin <yassin@berri.ai>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:47:46 +00:00
devin-ai-integration[bot]
317430db4e
fix(panw_prisma_airs): honor experimental_use_latest_role_message_only on every request shape (#42447)
* fix(panw_prisma_airs): apply experimental_use_latest_role_message_only to every request shape

Explicit true/false now applies to chat completions, Anthropic /v1/messages and /v1/responses alike; unset keeps latest-only for Anthropic and full history otherwise. Text indices are mapped back to their source message by value instead of by count, so Responses instructions, function_call_output and reasoning items no longer derail the alignment and silently rescan the whole history

Co-authored-by: scthornton <scthornton@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(panw_prisma_airs): type latest-message helpers against AllMessageValues

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): require forward and reverse text attribution to agree

A Responses function_call_output whose text equals the latest user turn could claim that turn's slot in a forward-only walk and demote the latest-only scan to an earlier message. Walk both directions and fall back to the full role-filter scan when they disagree

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): pick the latest human turn from messages, not from aligned texts

An image-only latest user turn no longer promotes an earlier user turn into the
latest-only scan; it scans nothing on the request side, as the Anthropic path did
before. A latest user/developer message whose text never reached texts (a trailing
Responses reasoning item) falls back to the role-filter scan instead of narrowing.
Types the test helpers, drops the narrating docstrings and adds regressions for both
shapes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): log when latest-only selection leaves nothing to scan

An image-only latest user turn with experimental_use_latest_role_message_only=true intentionally yields zero scanner calls. Emit a debug line naming the call_id so operators can tell this apart from the guardrail not firing.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(panw_prisma_airs): keep Responses reasoning items out of latest-turn selection

The Responses translation handler gives reasoning input items the default user
role, so a reasoning item with text content after the latest prompt was picked
as the latest human turn and the real prompt went unscanned under
experimental_use_latest_role_message_only. Map reasoning items back to their
texts positions from the raw input and exclude them; fall back to the
role-filter scan when the raw items do not account for every text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: scthornton <scthornton@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:16:18 +00:00
Yassin Kortam
fe76c2473d
Revert "feat(proxy): server-side Team Usage export beyond the top-N key cap (#42996)" (#43376)
This reverts commit 77eccaca78.
2026-09-28 12:33:18 -07:00
devin-ai-integration[bot]
3726ce2cfc
refactor(guardrails): fix agent 365 to the production endpoint and log the opt-in fail_open at error level (#43189)
* feat(guardrails): fail open by default when Agent 365 cannot evaluate and count it in Prometheus

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(tests): ruff format the Prometheus fail-open registry test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(guardrails): add Agent 365 authority host override, fail-open integration test and per-guardrail YAML default

Add `authority_host` to the Agent 365 config (also read from AGENT365_AUTHORITY_HOST, then AZURE_AUTHORITY_HOST) so sovereign clouds and the integration test can point the OBO exchange at a different Entra host.

Add tests/integration/mcp/test_mcp_agent_365_guardrail.py, a real proxy test with Postgres, Redis, a scripted MCP upstream and local Entra and Agent 365 doubles covering the default fail-open, explicit fail-closed and fail-open, Defender Skipped, policy denial, persisted status and Prometheus counter.

Use PrometheusLogger.get_instance for the fail-open metric lookup instead of a hand-rolled callback scan. Clarify the config description: gateway credential failures fail open, caller token failures block.

Extract the dashboard YAML preview into teamGuardrailConfigYaml.ts so the effective per-guardrail default is unit tested and the "default" hint only shows when nothing was set explicitly.

Regenerate the lazy OpenAPI snapshot and schema.d.ts for the new field.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): default a scheme-less Agent 365 authority host to https and treat a null fallback as unset in the YAML preview

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for the Agent 365 fail-open default across entry points, Entra faults, throttling and two workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): prove both Agent 365 workers serve and that a killed worker is replaced

Each fresh connection reports its worker pid from /debug/memory/summary and its MCP catalog on the same
connection, so the two-worker readiness wait covers both workers by identity. The kill test now kills a
pid the proxy reported as a worker and waits for a replacement pid, instead of the first psutil child

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): drop the prometheus fail-open counter from the agent 365 guardrail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): keep agent 365 fail closed by default and make fail_open an explicit opt-in

Restores the shared unreachable_fallback default and the sibling guardrail initializers, drops the Admin UI YAML preview that only existed for the per-guardrail default, and reworks the unit and integration tests so the default blocks with HTTP 503 while unreachable_fallback: fail_open lets availability failures through as Unscanned

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): append authority_host after the existing Agent365Guardrail parameters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(guardrails): agent 365 fails open by default and hides the production overrides from the UI form

Agent 365 sits in the runtime path of every MCP tool call, so an Entra or
Agent 365 outage now lets the call through unscanned (logged at error level,
recorded as Unscanned with guardrail_failed_to_respond) instead of blocking it.
unreachable_fallback: fail_closed stays as the opt-in strict mode. Policy
blocks, throttling, 4xx rejections and a rejected caller token still block

The shared unreachable_fallback field becomes nullable so each guardrail owns
its default; every sibling still resolves None to fail_closed and typesafe
keeps failing open

api_base, resource_app_id and agent_id have production defaults and leave the
dashboard form (ui_hidden); they stay available in config.yaml and env. The
authority_host override and its env keys are gone, the OBO exchange always
uses login.microsoftonline.com. The integration suite keeps only the cells
that need no Entra double, the evaluation paths live in unit tests with an
injected handler

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate openapi snapshot and schema.d.ts for the nullable unreachable_fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): fix agent 365 to the production endpoint and keep fail_closed as the default

Remove api_base, resource_app_id and agent_id from the Agent 365 config model, their AGENT365_* env fallbacks and the _is_ui_hidden helper: the evaluation URL and the Agent Tools app id are fixed production constants and the agent identity is always the caller's key alias. Revert the fail_open default; unreachable_fallback: fail_open stays an explicit opt-in. Restore the shared unreachable_fallback field, the sibling guardrail initializers and typesafe to main. Move the Entra dependent cells from the subprocess integration suite to unit tests with an injected HTTP handler.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): warn when agent 365 yaml still carries the removed override keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): inject the http handler into the agent 365 initializer instead of assigning it after construction

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 12:05:47 -07:00
ryan-crabbe-berri
6c34c6e7ee
test(integration): pin org-admin status codes in the team-admin matrix (#43592)
Run every route in the team-admin matrix again on a team that belongs to an
organization, as an admin of that organization and as an admin of another
organization. 90 new cases pin what org admins get today, so collapsing the
team-admin helpers into one gate can prove parity for org admins too

Scenario gains organization() and org_member() helpers that clean up the
organization, its budget row and the member users
2026-09-28 19:04:05 +00:00
devin-ai-integration[bot]
7aba77197d
feat(otel): add SigNoz preset for OpenTelemetry v2 (#43296)
Some checks are pending
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
* feat(otel): add SigNoz preset for OpenTelemetry v2

Adds the signoz callback (OTLP/HTTP exporter, GenAI vocabulary, key and team level dynamic ingestion endpoint and key) as an OpenTelemetry v2 preset, with the preset factory accepting the allow_missing_credentials kwarg the V2 registry always passes so construction no longer falls back silently to legacy OpenTelemetry. Ships the deterministic tests/integration/observability/test_signoz_delivery.py audit suite

Absorbs the work from https://github.com/BerriAI/litellm/pull/38206

Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(otel): drop explanatory comments from the SigNoz preset

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(types): keep signoz dynamic param lines within ruff format width

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(signoz): assert the missing-endpoint boot path directly instead of in an except block

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate schema.d.ts for the signoz health service

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): allowlist SigNoz key/team endpoints and route keyless collectors without the operator key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): terminate the SigNoz shutdown cell before the flush and drop test docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(otel): keep the shared tenant routing untouched and require an ingestion key for SigNoz key/team endpoints

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): warn about a keyless SigNoz team endpoint from the header resolver so the shared cache actually reaches it

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 18:15:45 -07:00
devin-ai-integration[bot]
eea1d0f269
fix(responses): stream guardrail pre-call block as SSE with a typed output item (#42507)
* fix(responses): stream guardrail pre-call block as SSE with a typed output item

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): import blocked usage helper from the guardrail utils module

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): drop narrating docstrings and poll without rebinding

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover pre-call guardrail block on /v1/responses stream and json

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for responses guardrail block contract

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): observe upstream on the recorded chat route for responses denial cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): tidy responses denial audit cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for worker count to recover after SIGKILL

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): require a replacement worker after SIGKILL

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the blocked response test helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-26 18:05:21 -07:00
devin-ai-integration[bot]
849f3037b4
fix(langtrace): deliver spans to app.langtrace.ai/api/trace with x-api-key (#43322)
* test(langtrace): integration test for the built-in callback wire (path, x-api-key)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(langtrace): deliver spans to app.langtrace.ai/api/trace with x-api-key

The built-in langtrace callback posted to the dead host langtrace.ai, sent
the key as api_key instead of x-api-key, and let the OTLP endpoint
normalizer append /v1/traces to the complete /api/trace path, so every
export returned 404. Default the host to https://app.langtrace.ai, honor
LANGTRACE_API_HOST for self-hosted servers, pass the key as an exporter
header instead of a process-wide env var, and keep the /api/trace path
unchanged for traces on the langtrace callback only

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(langtrace): keep LANGTRACE_API_HOST that already ends in /api/trace

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(langtrace): deterministic audit inventory for the built-in callback (surfaces, failures, endpoints, chaos)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(langtrace): build the repeated-request body once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(langtrace): outage cell asserts at-most-once delivery, not span loss

The OTLP HTTP exporter reposts once on ConnectionError and the batch
processor may still be flushing the previous burst when the sink closes,
so whether the outage burst is lost or delivered after revival depends
on timing. The invariant is no duplicate and recovery on the same port

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(langtrace): delivered spans must carry the prompt in the gen_ai.content.prompt event

The upstream echoes the marker into the completion, so a whole-span
match alone would still pass if the prompt event disappeared

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(langtrace): disable model info refresh so the scripted upstream only sees completion requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:31:52 -07:00
shrey-berri
e494723105
fix(params): validate stream_chunk_size once, before any provider call (#43222)
* fix(params): validate stream_chunk_size once and carry it as typed control options

Checks stream_chunk_size at the top of completion() and acompletion(), accepts
digit strings, returns a 400 naming the param unless drop_params is set, and
stores the checked value under _litellm_control. Bedrock Converse and Invoke
read it from litellm_params; the Bedrock-only checker and the dead Invoke pops
are gone. Owned-kwarg filtering now runs through one helper everywhere.

Refs LIT-8317

* test(bedrock): drop tests for the removed stream_chunk_size_from helper

Refs LIT-8317

* fix(params): check stream_chunk_size before the MCP gateway branch

Refs LIT-8317

* fix(params): return assert_never in the exhaustive control-options match

Refs LIT-8317

* fix(params): address council review of the control options change

Read all_litellm_params live so names registered after import stay
LiteLLM-owned, make litellm_params a required keyword on the stream
wrapper hooks, give digit strings and ints the same 18-digit range,
share the default-chunking test table, test the Responses bridge through
litellm.responses, and revert formatting-only churn in existing tests.

Refs LIT-8317

* fix(params): address the second council review of control options

Keep the Responses bridge on its original all_litellm_params forwarding,
narrow _int_from_decimal_string inline so it type-checks, bound nested
huge ints in the error message, store _litellm_control only when a value
is set, simplify the parser to its single field, drop the one-caller
wrapper, and tighten the tests.

Refs LIT-8317

* fix(params): keep the 18-digit length check on stream_chunk_size strings

A 19-character string with leading zeros such as 0000000000000000001 would
otherwise pass as 1, although the rule and the error message say at most
18 digits.

Refs LIT-8317

* test(params): tidy control options tests after council sign-off

Move the Responses bridge test into the existing bridge test file, drop the
rebind test that pinned an implementation detail, assert through
stored_control_options instead of the storage key, and cover
drop_params="true" through Bedrock streaming.

Refs LIT-8317

* test(params): wrap a chunking test row that went past 120 characters

Refs LIT-8317
2026-09-26 23:01:20 +00:00
devin-ai-integration[bot]
f12f7b5a03
test(integration): group /v1/messages contracts under tests/integration/messages_endpoint (#43352)
* test(integration): group /v1/messages contracts under tests/integration/messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): make ci coverage census collect nested test dirs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): nest /v1/messages contracts under messages_endpoint/providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:00:32 -07:00
yuneng-jiang
635a718ba1
ci: cut CircleCI wall time without loosening test isolation (#43347)
* ci: cut CircleCI wall time without loosening test isolation

* fix(ci): parse integration split files that follow --results

The CircleCI machine image ships Python 3.12.2, whose argparse leaves the
files positional empty when it follows an option and another positional, so
every extensions node exited with 'unrecognized arguments'. Reproduced on
3.12.2; parse_intermixed_args selects the files on 3.12.2, 3.12.13 and 3.13

* test(ci): resolve command references in the Rust toolchain guard

The Windows rustup install moved into the install_windows_toolchain command,
which the guard only recognized for install_rust. It now accepts any command
that installs a pinned rustup and reads the Windows toolchain pin from it

* ci: cache the Windows release cargo build from main

windows_release_wheel rebuilt every dependency with fat LTO on each run. It now
restores the release target and cargo registry saved by main's scheduled run,
drops the workspace crates' fingerprints so they always rebuild from the
checked-out source, and still runs the full LTO link

* ci: run the Windows release wheel build on windows.xlarge

The fat-LTO release build is the slowest job in the pipeline; more cores
speed up the dependency compile ahead of the final link

* ci: skip the Windows fingerprint cleanup when the cargo cache missed

On a cold cache the release fingerprint directory does not exist, and the
CircleCI PowerShell wrapper failed the step on the suppressed not-found error
2026-09-26 15:34:53 -07:00
shrey-berri
9413b82477
fix(params): keep _litellm_* kwargs out of provider request bodies by construction (#43221)
Kwargs LiteLLM code introduces for its own use were only kept out of provider
bodies if someone also listed them in all_litellm_params. Undeclared ones went
into extra_body or optional_params, reached the provider, and the provider
rejected the request. is_litellm_owned_kwarg in types/utils.py now defines
LiteLLM-owned once: a registered name, or any name starting with
INTERNAL_KWARG_PREFIX from litellm/constants.py. Every filter that builds
provider params from kwargs uses it: chat completion, transcription,
embedding, image generation and edit, search and video, ElevenLabs text to
speech, and the Bedrock batch mapper. The two untyped shared filters now take
Mapping[str, object]

The stream_chunk_size wire test becomes test_internal_params_wire.py. It also
sends an undeclared _litellm_ kwarg and asserts that no _litellm_ key reaches
any of the six provider bodies, while extra_body passthrough keeps working

Refs LIT-8318, LIT-8319
2026-09-26 15:20:09 -07:00
devin-ai-integration[bot]
1474ea53e6
feat(proxy): add maximum_daily_tag_spend_retention_period cleanup setting (#39221)
* feat(proxy): add maximum_daily_tag_spend_retention_period cleanup setting

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate schema.d.ts for new retention setting

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): rebase daily tag spend retention onto the run-budgeted cleanup job

Reworks the cleanup on top of the refactored SpendLogCleanup: the daily tag spend table is pruned through the shared batched delete with a text cutoff on the indexed ISO date column, the setting is picked up by /config/update and the scheduler registration, and an integration test proves rows older than the period are pruned while the cutoff day and unset retention are left alone

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): schedule the cleanup job when a retention db row lands before the side effects run

A config reload applies the db row to the SettingsStore before _update_general_settings snapshots the previous retention values, so the before/after compare saw no change and a retention period first set through /config/update never scheduled the cleanup job. Also reschedule when the job is missing but a retention period is set

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): accept list-valued top-level keys in the base integration proxy config

The shared tests/integration/proxy_config.yaml now carries list-valued top-level keys, so the retention config helper validates only the mapping it merges into. Also drops a SQL-shape assertion from the unit test in favor of the behavioral cutoff-day check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover runtime update, invalid value, independent horizons and worker loss for daily tag spend retention

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): restore the shared retention setting, capture seeded days once and kill a listening worker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): retry a failed cleanup schedule only when its settings change

_apply_retention_settings rescheduled whenever retention was set and no job existed, so an unparseable cleanup cron was retried on every config reload. Remember the last attempted retention, cron and interval tuple and retry only when it differs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): give daily tag spend retention tests a 240s timeout

Each node boots a proxy and waits for a whole-minute cleanup cron tick, so the global 90s pytest-timeout can expire during teardown on a slow runner, as integration-accounting did on pipeline 90302

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): reschedule cleanup when only the cron or interval changes at runtime

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): reschedule cleanup when the first db sync changes only the cron or interval

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(ui): add text input for String general settings so retention periods can be set from the Admin UI

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): record a cleanup schedule attempt only after it did not raise

Records _last_cleanup_schedule_attempt after _reschedule_spend_log_cleanup_job returns, so a transient add_job error is retried on the next config sync while an invalid cron, which is caught and logged inside the reschedule, is still attempted once per settings value

Also adds --num_workers 2 to the dev proxy command in AGENTS.md as requested on the PR

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: revert unrelated AGENTS.md dev command change

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Revert "docs: revert unrelated AGENTS.md dev command change"

This reverts commit 047706a623.

* Revert "fix(proxy): record a cleanup schedule attempt only after it did not raise"

This reverts commit 678f7c72b4.

* Revert "feat(ui): add text input for String general settings so retention periods can be set from the Admin UI"

This reverts commit 24e49d71d7.

* fix(proxy): record a cleanup schedule attempt only after it did not raise

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): validate the cleanup schedule before swapping the job and leave startup registration to the startup block

_reschedule_spend_log_cleanup_job builds the new trigger first and only touches the live job once it parsed, so an invalid cron or interval (including a non string value) keeps the previous schedule running instead of removing it. An error raised while rescheduling is logged and retried on the next sync, so it no longer stops the rest of the general settings sync. _apply_retention_settings skips the job-missing path while the scheduler is still stopped, so the startup block is the only registration before start and the cross-replica stagger it applies to pending jobs survives

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): skip cleanup rescheduling while the scheduler is stopped and retry a failed replacement

The stopped-scheduler guard only covered the missing-job path, so the first DB sync (which runs before the startup block) still registered the cleanup job whenever the DB schedule differed from yaml, and startup then replaced it. Every runtime path now defers to the startup block while the scheduler is stopped.

A raised add_job that was replacing a live job was never retried because the live job kept wants_job == has_job; the sync now remembers the failure and retries on the next sync until the schedule is applied.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop redundant docstring on _spend_log_cleanup_trigger

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): schedule DB-only retention at boot and log overflowing cleanup intervals once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): drop explanatory comment from startup cleanup block

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): reject a non-string cleanup cron at startup and drop legacy covers markers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): reschedule spend log cleanup when the reload path already applied a DB cron or interval edit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): assert cleanup scheduling on a real paused scheduler instead of mock call counts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-09-26 15:07:56 -07:00
devin-ai-integration[bot]
e47b1f2a3f
fix(s3_v2): upload fresh events first, drop terminal failures and hour-old retries by default, opt-in adaptive concurrency (#43022)
* fix(s3_v2): drop terminal upload failures, bound retries per flush and enforce the queue cap at enqueue

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): keep retrying credential-rotation 403s, only AccessDenied-style errors are terminal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(env_keys): exclude DEFAULT_S3_MAX_FLUSH_ATTEMPTS as an internal tuning var

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): retry every 5xx, warn on first queue overflow, validate the flush budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): read the queue cap defensively so un-initialized loggers still enqueue

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): drop the getattr in _enqueue and tighten the retry tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): keep the constructor flush budget when the callback override is invalid

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(s3_v2): adapt per-object upload concurrency to sink latency and throttling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(s3_v2): tidy adaptive limiter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(s3_v2): wake one waiter per released upload slot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(s3_v2): make the enqueue queue cap configurable with s3_max_queue_size

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): audit cells for cache hits, coded 403 and callback modes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): retry bucket-wide failures by default, age-budget requeues and make terminal drops and adaptive concurrency opt-in

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(s3_v2): suppress the missing-waiter ValueError explicitly in the adaptive limiter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): count oldest events trimmed after a failed flush as callback failures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(s3_v2): move the mutable-ok marker onto the list literal it suppresses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): report post-flush overflow drops once and grow adaptive concurrency above the floor before asserting back-off

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): fail the SlowDown back-off test when the measured window sees no PUTs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): restore base retry defaults, opt-in age budget, no enqueue cap, back off outside the limiter slot

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): hoist the default no-op upload slot to a module constant

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): drop unused mutable-ok suppressions on queue appends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): keep the retry queue oldest-first and prioritise fresh events at upload time

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(s3_v2): keep the mutable-ok marker on the queue list literal

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): build request bodies inside the upload slot and keep the sync retry set at base parity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): drop wall-clock sleeps from the unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): rebuild the request body inside the slot on every retry attempt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): anchor the backoff window on the first observed failure and tighten shard assertions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): default the upload slot to the logger limiter so monkeypatched doubles keep working

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): clear ambient AWS env credentials so the rotating profile signs the sync retry test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): shrink the linear send-batch perf test to 2k/8k elements

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): fix stale batch sizes in the perf test assert message

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): keep async in-call retries on the base 403/500/503 set

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): hold the upload slot across retries, restore the bool upload contract, and fail safe on bool config

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): mark dropped uploads by element identity so a shared key cannot mask a retryable sibling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): make the per-flush drop lookup constant time

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): take the upload slot in the caller like base, build the body once per attempt loop

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): match base retry, logging and hook behaviour unless the new options are opted in

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): assert the wire key in the init-bypassed sync upload test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): assert the signed headers and wire key in the init-bypassed sync upload test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): default the upload limiter at class level instead of reading it with getattr

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): drop the duplicate annotations that redeclare the class-level counters

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(s3_v2): drop terminal-failed uploads by default and bound retry age to one hour

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): fall back to the configured retry age on invalid values

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): drive retry-age tests from a fixed clock

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): audit cells for retry-age opt-out and 429 single-put parity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 14:58:28 -07:00
joshua-berri
3743c8563e
fix(mcp): adopt shared server resolution and caller authorization (#43263)
* test(mcp): characterize server resolution and authorization

* refactor(mcp): extract shared server resolution

* fix(mcp): adopt shared resolution in management endpoints

* fix(mcp): scope credential metadata resolution outside loop

* test(mcp): pin catalog isolation and batched credential permissions

* fix(mcp): restrict catalog detail and batch credential permissions

* test(mcp): enforce identity isolation in database fixtures

* test(mcp): name resolution tests by behavior

* test(mcp): describe detail access assertion failures

* chore: keep agent naming discipline local

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-26 14:37:32 -07:00
ryan-crabbe-berri
96c008f420
ci: fail on new unbounded SQL IN lists and add a Prisma chunking helper (#42629)
* ci: warn on SQL IN lists with no written bound

Postgres caps a prepared statement at 32,767 bind parameters and a
membership filter binds one per value, so an IN list built from table
data breaks once the table outgrows the cap. That is how the budget reset
job froze every due budget (LIT-7535, #40564).

check_unbounded_in_lists.py reports every Prisma "in" / "not_in" filter
whose value has no fixed size and every raw SQL literal that splices a
list in after "IN (", unless the line carries "# bounded-ok: <reason>".
It only warns for now: the output is the inventory for RCA action item
AI-1, and it exits 0.

* ci: decide a constant IN list by its module binding, not its casing

An ALL_CAPS name imported or filled at runtime is as unbounded as any
other, so a name now passes only when the module binds it once to a
value of fixed size. Adds Final to the locals a loop does not forbid.

* ci: only a frozen module value makes an IN list constant

A module list bound once could still grow through append or extend, so
a name now counts as fixed only when it is bound to a tuple, frozenset
or constant. Trims the module docstring to what a reader needs.

* ci: chunk Prisma IN lists with a shared helper and fail on new unbounded ones

Add litellm.repositories.bounded_in: find_many_in, count_in, update_many_in
and delete_many_in split a deduplicated value list into 5,000-value chunks,
AND each chunk with the caller's where, run them in order (a transaction
handle works) and combine the results. Writes take a required atomicity
argument, and a where that already filters the chunked field is refused.

check_unbounded_in_lists.py now fails CI on any finding missing from
unbounded_in_baseline.txt and on any stale baseline entry, so the baseline
only shrinks. Entries are keyed by path, enclosing scope, kind, field and
occurrence, not line numbers. The helper module is exempt, a constant
spread into a frozen tuple counts as fixed, and messages point at the
helper for "in" and at an array parameter for "not_in" and raw SQL.

A real-Postgres integration test shows a raw 40,000-value filter rejected
for too many bind variables while the helpers handle it.

* refactor: rename bounded_in to chunked_in and let callers pick a chunk size

The helper module is litellm.repositories.chunked_in, and its unit and
integration tests, the checker's exemption path and its finding messages
follow the new name. The `# bounded-ok` marker is unchanged.

find_many_in, count_in, update_many_in and delete_many_in take a
keyword-only chunk_size, defaulting to IN_LIST_CHUNK_SIZE (5,000). A value
below 1 or above MAX_IN_LIST_CHUNK_SIZE (30,000) raises ValueError before
any query, which leaves the rest of the filter headroom under Postgres's
32,767 bind-parameter cap.

* refactor: flatten chunked_in's stacked comprehensions with chain.from_iterable

LIT014 (#42650) caps a comprehension at one for and one if clause. The four nested walks in the helper now chain their iterables instead, with the same order and results.

* refactor: recover user details with find_many_in, sending chunks as lists

_details_for_user_ids reads users through find_many_in instead of a raw
"in" filter, so its lookup stays under the bind-parameter cap for any
number of recovered keys. Up to 5,000 ids it still sends one find_many
with the same where dict, and a PrismaError from any chunk is still
logged and treated as no details.

The helper now sends each chunk as a list, so a chunked filter equals
the dict a hand-written call would send and a migrated call site's
existing assertions keep passing.

The site's baseline entry is gone.

* ci: skip functional TypedDict field maps in the unbounded IN list check

The dict passed as the field map of TypedDict("Name", {...}), or as its fields= keyword, names fields: an "in" or "notIn" key there is a type, not a filter. Only that dict is skipped, for TypedDict, typing.TypedDict and typing_extensions.TypedDict; a filter nested in a field value or passed to any other call is still reported. The two types/proxy/management_endpoints/team_endpoints.py entries leave the baseline, which is now 156.

* fix: refuse an update_many_in whose data writes the chunked field

Chunks run one after another, so an update that sets the chunked field can move a row into a later chunk, which updates it again and counts it twice: values ["old", "new"] with chunk_size=1 and data={"id": "new"} does exactly that. update_many_in now raises ChunkedFieldWriteError before any query when data has the chunked field as a top-level key, in any form, including Prisma operators such as {"set": ...}.

* docs: cut the unbounded IN list checker's docstring to what it flags and how to clear it

It now says what is reported, the three ways to clear a finding, and how the baseline and --update-baseline work, in 11 lines. The per-shape detail lives in the tests.

* ci: key an unbounded IN list finding by its filtered expression too

A baseline key of path, scope, kind, field and occurrence let a PR delete
a baselined filter and add a different unbounded one on the same field in
the same function, and the new one took over the old key. The key now
also carries the filtered expression's source, whitespace-normalized
(the Prisma value, or a raw-SQL `IN (...)` slot), so that swap reads as
one new and one stale entry and fails the run. The same expression
re-added in the same function is still the same finding.

Every baseline entry is rewritten in the new form; the 156 findings are
unchanged, and only occurrence indexes renumber where one field had
several different expressions.
2026-09-26 13:40:44 -07:00
joshua-berri
0300fc4ab2
test(mcp): pin server resolution and authorization behavior (#43261)
* test(mcp): characterize server resolution and authorization

* test(mcp): pin catalog isolation and batched credential permissions

* test(mcp): enforce identity isolation in database fixtures

* test(mcp): name resolution tests by behavior

* test(mcp): describe detail access assertion failures

* chore: keep agent naming discipline local

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-26 12:59:11 -07:00
ryan-crabbe-berri
41070b1363
test(integration): pin the team-admin status-code matrix across every management route (#43249)
* test(integration): pin the team-admin status-code matrix across every management door

Every endpoint that admits a team admin today is called as a proxy admin, an admin of the
target team, a plain member, an admin of another team and a teamless user, and the current
status code is asserted per actor. The matrix is the parity check for collapsing the five
team-admin helpers into one shared gate and for the later default-off permission flip.

* test(integration): pin the permission-enabled team-admin doors in the gate matrix

Adds three doors that run with team_admin_editable_team_fields granting max_budget, projects and member_key_budgets, so the enabled path is pinned alongside the default-off one. Hoists the ui_settings toggle from test_warmed_policy into the shared client so both files use one helper

* test(integration): grant each permitted door only the permission it needs

A door now names its single grant instead of every permission at once, so a gate that checks the wrong permission for a route turns that door red

* test(integration): rewrite the team-admin matrix rows as request plus expected codes

Each row now names the route it calls and the code each caller gets, and creates the member, key, model, callback or invitation it acts on through plain helpers on the shared team. Drops the Need, Target, World and Door types and the prepare step that seeded fixtures by enum.
2026-09-26 11:32:24 -07:00
yuneng-jiang
e53e67ede5
test(e2e): assert only litellm-owned batch behavior and move the blank S3 env pin to an integration test (#43321) 2026-09-26 11:17:03 -07:00
yuneng-jiang
14f4c34c61
fix(ci): stop stale CI reds, keep unit tests off the host env, retry CyberArk policy conflicts (#43294)
* fix(ci): stop five stale or flaky CI reds and retry CyberArk policy-load conflicts

The Langfuse redaction unit test exports to a local OTLP capture instead of
polling Langfuse Cloud through a recorded lookup. The passthrough worker-kill
test only requires spend rows for requests the surviving worker served. The
spend-routes sweep treats the intentional /spend/capture_rate 503 as expected.
CyberArk retries a 409 policy load in Python, Rust and the e2e Conjur helper
instead of reading it as "variable exists". The integration egress guard now
matches the script's own cgroup, so it no longer blocks the CircleCI agent,
which runs as the same user.

* fix(ci): keep the policy-load backoff typed as float

* fix(ci): retry CyberArk policy loads without blocking the event loop and tighten the worker-kill and Langfuse tests

* fix(secrets): load CyberArk policy one request at a time per manager

* test(secrets): pin that non-conflict CyberArk policy failures are not retried

* test(unit): run tests/unit with only an allowlisted host environment

CircleCI's unit job inherits every project env var, so real provider keys,
REDIS_HOST, DATABASE_URL and AWS or Azure credentials reached tests that
assume none are set. Locally, litellm's import-time load_dotenv did the same
from any .env up the tree. The unit conftest now drops every variable outside
a small allowlist and disables dotenv before litellm is imported.

* test(e2e): name a failed search and the stuck batch status instead of misattributing them

The websearch session test read an empty web_search_tool_result_error block as a
successful search, so a failing search tool surfaced as a session billing bug.
The batch cancellation timeout now reports the last status the proxy returned.

* fix(ci): scrub the host environment per unit test instead of for the whole pytest process

GHA shards run tests/unit next to other suites in one process, so the import-time
scrub deleted MCP_TEST_PEER_PYTHON before tests/mcp_tests read it and the MCP
upstream fell back to the SDK2 interpreter. The two websearch tests that called
OpenAI and Perplexity live are removed: tests/unit no longer sees their keys.

* fix(ci): scrub only the host variables present before litellm is imported

The per-test scrub also deleted TIKTOKEN_CACHE_DIR, which litellm sets at import to
its bundled encodings, so tokenizer paths tried to download them and hit the
socket guard. The prisma setup test now passes its own database URL instead of
reading one another test leaked into the process environment.

* fix(ci): stop the order-dependent unit reds and settle logging tasks on their own queue

LoggingWorker marked a task done on whichever queue was current when the callback
finished, so a callback that outlived an event-loop change raised "task_done()
called too many times" or undercounted the new loop's queue. It now settles the
queue the task came from.

The rest are test isolation fixes for failures that only appeared when another
file ran first on the same xdist worker: a replaced user_api_key_cache, breaker
metrics unregistered by prometheus tests, semantic_router's health-check filter on
uvicorn.access, logging tasks carried over from bedrock tests, a Router-written
model_cost entry, and a stray post captured by the langflow test. The token
counter check now asserts bounded chunking instead of wall-clock time.

* test(e2e/ui): wait for the logout redirect before visiting a protected page

Logout revokes the session server-side before clearing cookies and navigating, so an immediate page.goto either ran with the cookie still set or was aborted by the logout redirect (net::ERR_ABORTED).

* test(unit): restore the prometheus metrics config per test and settle logs carried from earlier tests in the a2a cost tests

* test(router): pin the router clock in the usage counter tests so a minute rollover cannot empty the read

* test(e2e/ui): wait for logout to clear the token cookie instead of for a login redirect

* test(integration/mcp): answer the model-info probe another test's proxy sends to the model double
2026-09-26 09:25:13 -07:00
yuneng-jiang
dd63637322
test(integration): run the Langfuse DB-callback test on its own scratch database (#43288)
* test(integration): run the Langfuse DB-callback test on its own scratch database

The test from #43282 wrote success_callback=langfuse and the LANGFUSE_* env
into the shared integration LiteLLM_Config. The suite's long-running gateway
reloads that table and only ever adds callbacks, so it kept exporting to the
test's closed Langfuse fake for the rest of the shard even after the rows were
restored. The owned proxy now gets a scratch database, which also removes the
snapshot/restore code. scratch_database moves into _support/database.py so
test_cache_and_quota and this test share one copy, and the stock-config guard
now checks the callback settings instead of the raw YAML text.

* test(integration): include failure_callback in the stock-config Langfuse guard
2026-09-26 00:15:56 -07:00
devin-ai-integration[bot]
99655b6f86
test: finish the non-proxy half of tests/test_litellm (#43281)
* test: move key-gated tests/test_litellm SDK tests into tests/llm_translation and drop empty folders

* test: make token counter and health check unit tests run offline

* ci: point unit shards, rust path filter, Makefile and docs at tests/unit

* docs: fix stale test_litellm run paths in moved llm_translation tests

* fix: correct databricks e2e sys.path depth and contributing example path

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 22:43:41 -07:00
devin-ai-integration[bot]
e11c3f5815
test(integration): port langfuse callbacks-in-db coverage to the local harness (#43282)
* test(integration): port langfuse callbacks-in-db coverage to the local harness

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): assert the langfuse db rows after the exported span

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 21:54:11 -07:00
devin-ai-integration[bot]
d08746feb1
feat(proxy): email alerts at configured percentages of a team member budget (#42665)
* feat(proxy): email alerts at configured percentages of a team member budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(alerting): label team member budget crossings as team member budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(auth): cover the team member alert dispatch from _check_team_member_budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(email): drop the emoji from the team member budget alert template

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): ignore team member alert thresholds outside 1 to 100 on both the backend and the dashboard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): bound team member alert threshold key length before int parsing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): drop the legacy covers marker from the team member alert test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(team): reject malformed team_member_max_budget_alert_emails on team writes

Thresholds outside 1-100, non-list recipients, and invalid emails now return 422 on
/team/new, /team/update and PATCH /team/{id} instead of being stored and silently
ignored. The value is stored canonically. Read-side LiteLLM_TeamTable is unchanged,
and the PATCH body stays a raw merge patch so a null threshold still deletes it.

* fix(auth): enforce and alert on team member budgets only in common_checks

The builder re-checked the team member budget inline before common_checks ran the same
check, so one request that crossed a team_member_max_budget_alert_emails threshold
dispatched two alerts. Drop the inline check; common_checks is the single authorization
point and already covers per-member rows, the team default member budget, zero-cost
skips and the cross-pod spend counter. Its 422 message now uses the TeamMember=user:team
form the builder and budget reservation already returned.

* Revert "fix(team): reject malformed team_member_max_budget_alert_emails on team writes"

This reverts commit 703e754b46.

* fix(alerting): keep BaseBudgetAlertType.get_event_message zero-arg

Requiring user_info broke existing callers and out-of-tree subclasses. The team member
label now comes from SlackAlerting.budget_alerts, so the interface and its Readme are
unchanged from main.

* fix(mcp): keep team member budget enforcement on the MCP OAuth auth dependency

The MCP OAuth dependency stops at _user_api_key_auth_builder and never reaches common_checks, so removing the builder's inline member budget check would have let over-budget members through there. Enforce it explicitly for that caller.

* fix(auth): keep main's team member budget enforcement, alert once per request

Restore the builder's team member budget check and 422 message exactly as on main and drop the MCP-only gate. The builder sends the member alert only on the request it rejects; common_checks sends it for requests that get past the builder, so no request alerts twice.

* test(integration): read team member alert deliveries without a shared accumulator

* test(integration): match team member alert deliveries by subject so other alerts cannot race the count

* refactor(proxy): build the team member alert threshold config without mutable collections

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): collapse the alert recipient isinstance checks into one call

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): read the SMTP sink through lock-guarded snapshots and assert the exact deliveries

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 02:42:45 +00:00
joshua-berri
6b7688869e
feat(mcp): configure protocol versions and capability discovery (#43169)
* feat(mcp): configure protocol versions and capability discovery

* test(mcp): return SDK initialization result in REST pagination fixture

* test(mcp): arm cancellation deadlines after TCP calls start

* fix(mcp): avoid serialized discovery and listing spend logs

* test(mcp): isolate default protocol header policy

* fix(mcp): honor preview protocol pins and refresh migrated tests

* fix(mcp): preserve edited protocol pins in saved OAuth previews

* fix(mcp): retain saved protocol pins when previews omit versions

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-09-25 13:13:52 -07:00
devin-ai-integration[bot]
c19ce71bcd
fix(presidio): mask streamed /v1/messages output when the first upstream read is a keepalive, a data-less ping, or a split utf8 character (#43023)
* test(presidio): cover first-frame utf8 split, comment keepalive and data-less ping in streaming output masking

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(presidio): classify the streaming output shape on a frame with a data line and tolerate a split utf8 boundary

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(presidio): relay leading data-less sse frames before classifying the stream shape

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 00:48:42 -07:00
devin-ai-integration[bot]
2701e2008b
fix(ui): group cost optimization cache leakage by model group (#43008)
* test(ui): cover cost optimization cache leakage grouping by model group

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): group cost optimization cache leakage by model group

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 23:51:51 -07:00