mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-03 02:22:24 +00:00
14048 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
264b09ac8d
|
fix(responses): scan and mask top-level instructions with guardrails (#43629)
* fix(responses): scan and mask top-level instructions with guardrails The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present. Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment. Resolves LIT-8931 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): reject empty guardrail rewrites instead of forwarding raw input An explicit texts=[] answer from a guardrail now fails the count check and raises UnappliableRequestRewrite like any other misaligned rewrite; only a missing texts key means no rewrite. Types the out-param as dict[str, object] and adds integration coverage for instructions blocking, masking, empty instructions, tool loops, latest-only and concurrent workers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): type the texts-replacing guardrail helper explicitly Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): honor skip_system_message_in_guardrail for instructions and system input items Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): cover skip_system_message_in_guardrail on the live proxy Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system Trust a guardrail's structured_messages_cover_full_request claim only when it returns as many rows as the full normalized request, otherwise merge the scoped rows back so skipped instructions and system items survive the write-back. Make PANW's Responses reasoning alignment skip-aware so latest-only still picks the latest user turn when system content is excluded from texts. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): annotate new guardrail tests with return types Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): type the guardrail test doubles explicitly Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
3930c5bab6
|
fix(proxy): strip caller credentials from websocket passthrough (#43855)
* fix(proxy): strip caller credentials from websocket passthrough Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): cover configured x-api-key in websocket passthrough credential test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: oliver <oliver@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
79756cbb9b
|
feat(agents): enforce authoritative agent permissions (#43721)
* feat(agents): authoritative permissions * fix: enforce authoritative managed agent permissions * fix(agents): only consult the identity store for managed targets is_agent_allowed entered the identity-store path whenever a prisma client was configured, so an ordinary agent paired with an internal user returned 503 instead of 200. Classify the target from the registry first and fall back to the store only when the registry has no entry, so an unmanaged target never depends on the store being reachable. * fix(agents): gate the managed path on an admitted policy object Ten call sites branched on `managed_agent_policy is not None`, which any MagicMock attribute satisfies, so the managed path fired on unmanaged subjects and died in Pydantic validation as a 503. Route every check through a shared helper that requires a real AgentResponse. * test(mcp): stub the writer replica the fresh-policy reads use reload_admitted_user now passes check_db_only through to get_user_object, so the user row is read from writer_db. Point the mocks at the replica the code actually reads and give each parametrized case its own user id. * fix(agents): cap a managed agent at the invoking team's agents resolve_agent_access returned the managed policy's grants before the agent_caller ceiling was applied, so a managed agent acting on behalf of a user reached agents that user's team was never granted. Intersect with the caller ceiling the unmanaged path already honours. * fix(agents): restore token narrowing and scope the private-access suppressions The managed-model check lost its valid_token narrowing when it moved to the shared helper. Make the caller-access resolver public rather than reaching into it from module scope, and give each remaining private access a reason. * docs(agents): drop the comment claiming admins skip the A2A permission check The check has never had an admin bypass on this path, so the comment described behaviour the code does not implement. * test(proxy): stub the writer reads and restore the MCP manager singleton Fresh-policy user lookups read writer_db, so the team and rest-endpoint mocks stubbed a replica the code no longer reads, and the dashboard session fake still had the pre-kwarg signature. The manager reload also rebound global_mcp_server_manager in every MCP module without restoring it, leaking an empty manager into later files. * style: sort imports under the litellm package ruff config * fix(mcp): cap a managed agent's servers and tools at the invoking caller managed_agent_servers and managed_agent_tools returned the agent's own grants without the agent_caller ceiling the unmanaged resolvers apply, so a managed agent reached MCP servers and tools the echoed caller could not. Call the existing ceiling helpers on both axes. * refactor(mcp): return the caller-capped tools without an interim list The ceiling helper already returns a sequence, so materializing it into a list added a mutable collection for nothing. Sort at the return sites instead, which also makes the tool order stable across both branches. * fix(agents): preserve actor ceilings during managed target checks * fix(agents): keep managed permission ceilings authoritative * fix(mcp): fail closed on authoritative caller team outages --------- Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
7204942756
|
fix(proxy): keep request-body credentials out of stored spend-log requests (#43635)
* fix(proxy): keep request-body aws credentials out of stored spend-log requests * fix(proxy): redact every credential-named request-body field in stored spend-log requests Replace the hard-coded AWS key check in the spend-log request-body sanitizer with SensitiveDataMasker's key classification, so Azure, Vertex, watsonx, OCI, GigaChat, Gemini and header credentials are redacted too. Proxy-stamped key identity metadata is kept. * fix(proxy): keep request identifiers named like keys in stored spend-log requests * refactor(proxy): drop the AWS-only snapshot exclusion now that spend-log redaction is name-based * refactor(proxy): use SensitiveDataMasker's key classification without an exclusion list * refactor(proxy): always redact credential-named fields in stored spend-log payloads |
||
|
|
82d8b3797c
|
fix(proxy): attribute completed batch cost rows to /batches in daily activity (#43870)
* fix(proxy): attribute completed batch cost rows to /batches in daily activity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): wait for priced batch tokens before asserting team endpoint activity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
b71f02dbcf
|
fix(ui): keep MCP permissions visible after key, team and MCP server saves (#43810)
* fix(ui): keep MCP permissions visible after key, team and MCP server saves Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): type the object_permission include as a prisma TypedDict Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ui): do not block key save confirmation on cache refetch Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: ryan <ryan@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
61a73c59b0
|
fix(proxy): look up hashed key names with two spend log rows per key (#43656)
* fix(proxy): look up hashed key names with two spend log rows per key The spend-log fallback for keys missing from the key table read every row per key to check that all named rows agreed, which passed the 5s statement timeout on busy keys even with the (api_key, startTime) index. Probe only the oldest and newest named row per key, so the lookup stays two index reads per key however much the key logged. * fix(proxy): cap each spend log name probe at 100 rows per key * fix(proxy): bound the newest-row probe at where the oldest probe stopped The newest-row probe now starts at the row where the oldest-row probe gave up, so a key with under 200 rows in the window is read once instead of twice, and the lookup transaction turns bitmap scans off so the planner walks the (api_key, startTime) index instead of every row of a busy key when statistics or the visibility map are stale. * test(integration): add spend log alias probe cells for the daily activity routes --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
8afabe81f1
|
fix(ui): surface x-litellm-call-id in Logs search, table and drawer (#42436)
* fix(ui): surface x-litellm-call-id in Logs search, table and drawer Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(ui): regenerate api types for spend logs search description Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): drop redundant comments from the call id logs helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(e2e): format logs call id helper and spec Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ui): keep one id per Logs row, move x-litellm-call-id to hover and drawer The Request ID cell shows only request_id again. When the row's litellm_call_id differs, the cell tooltip lists it as x-litellm-call-id with its own copy button, and the drawer header labels the second line x-litellm-call-id: instead of the call id caption. Stacking two ids in every row made the column noisy for the common case where the viewer only needs the row they searched for. * test(e2e): cover the Request ID tooltip hover and copy path Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: poll the clipboard after the tooltip copy and drop a jsdom aside The e2e read navigator.clipboard right after the click, so a slow async write could fail the check even though copy works. The unit test's fireEvent choice (jsdom has no layout, so a real pointer move off the trigger closes the tooltip before the click lands) is documented here instead of inline. --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: ryan-crabbe-berri <ryan@berri.ai> |
||
|
|
d098b02ed9
|
fix(auth): give UI/CLI session tokens their own AES-GCM context and header-safe shape (#43790)
Some checks are pending
Unit Tests / misc (push) Waiting to run
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* refactor(auth): bind UI/CLI session tokens to their own AES-GCM context UI and CLI session tokens are now always encrypted with AES-256-GCM and a fixed session associated-data value, and the session-token check only accepts AES-GCM values carrying that same value. Stored secrets keep their current encryption and decrypt unchanged, so nothing needs migrating. encrypt_value_helper and decrypt_value_helper take an optional aad. XSalsa20 cannot bind associated data, so an AAD-bound value is always written as AES-256-GCM, and an AAD-bound decrypt refuses the legacy format. Session tokens issued before the upgrade stop validating, so UI and CLI users sign in once more after upgrading. * test(e2e): cover real SSO login through the dashboard and the lite CLI Adds two specs under tests/e2e/ui/oidc, run by playwright.oidc.config.ts against a live Keycloak stack. The dashboard spec checks that the SSO session authorizes the Virtual Keys and Models data requests. The CLI spec runs a real lite login in an isolated HOME with the keyring disabled, then lists models and sends one chat completion with the stored session. The main Playwright config now ignores oidc/. * fix(auth): encode UI/CLI session tokens as unpadded base64url Session tokens carried the v2:gcm: storage prefix and base64 padding. Basic-auth parsers split on the first colon and browsers reject ':' and '=' in WebSocket subprotocols, so Langfuse pass-through and the realtime playground could not use them Tokens are now plain unpadded base64url, the same header-safe shape as any bearer token * fix(auth): prefix UI/CLI session tokens with litellm_login_ A prefix-less token starts with sk- about once in 262,144 logins and is then routed as a virtual key, so that login gets a 401. The prefix also makes session tokens easy to spot in logs The prefix doubles as the token's AES-GCM associated data, so the visible kind and the encrypted kind cannot disagree --------- Co-authored-by: ryan-crabbe-berri <ryan@berri.ai> |
||
|
|
6684256136
|
feat(agents): add identity storage and validation contracts (#43720)
* feat(agents): identity storage and contracts * fix(agents): cache positive identity lookups with fresh policy checks * test(agents): include identity attribution in spend fixture --------- Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
13d004fc5a
|
perf(proxy): refresh auth management objects through the request Redis pipeline (#43776)
Identity objects (key, end user) load through the request MGET and their write-backs, the registry reads and the management-object SETs ride the request pipeline. A team refresh invalidates its alias with a pipelined DEL instead of a synchronous DEL plus a duplicate async one, and an MGET miss is remembered so no per-key GET follows it in the same request. Resolves LIT-9012 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yassin <yassin@berri.ai> |
||
|
|
c129ea4fc9
|
fix(mcp): scope OpenAPI listings to the exact server prefix and drop upstream OAuth metadata when a server is saved (#43608)
* fix(mcp): key discovery caches per caller correctly and drop stale caches on server updates Discovery-list cache identity now uses the hashed token instead of the raw api_key and treats MCPJWTSigner-signed servers as per caller. Server definition changes also drop the cached upstream OAuth metadata. OpenAPI listings look tools up under the normalized registry prefix with the separator, so an overlapping sibling prefix no longer leaks into the list. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep the discovery cache digest call unchanged so CodeQL matches the existing alert Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): guard OAuth metadata cache writes with a per-server generation and drop unproven per-caller discovery keys An upstream metadata fetch that started before a server edit could store its stale reply after invalidate_oauth_metadata_cache ran. Invalidation now bumps a per-server generation and the fetch only stores when the generation it captured before I/O is unchanged. The MCPJWTSigner-based per-caller discovery classification and the api_key to token key change had no reproduction (the signer only injects on tools/list, and UserAPIKeyAuth hashes api_key in place), so both go back to the merge-base behavior. Integration coverage under tests/integration/mcp: overlapping OpenAPI aliases, a config-declared server name with a space, OAuth metadata refetch after a save, and the in-flight stale-write race Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep OAuth metadata generations only while a fetch is in flight Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): count queued OAuth metadata fetchers so invalidation survives lock handoff Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mcp): keep a held OAuth metadata lock registered even when no fetcher slot claims it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mcp): prove a peer worker drops stale upstream OAuth metadata after a save elsewhere Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
ffb15f946f
|
perf(proxy): one request-scoped Redis pipeline for auth, spend, rate-limit and routing reads (#43407)
RedisBatch: one pipeline per Redis backend for independently declared operations (MGET, GET, Lua scripts, INCRBYFLOAT, SET, DEL), a future per operation so each owner keeps its own fallback, Redis Cluster hash-slot fallback. A request-scoped batch middleware shares that pipeline across the auth identity reads and write-back, the spend counter MGET, the rate limiter Lua groups and the routing read. A rate-limit denial stands when another pipelined group fails; every pipelined group is refunded on rejection; local cooldowns win over the prefetch. The routing prefetch failure log line strips request line breaks (CodeQL py/log-injection) Resolves LIT-8882 Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
f5a1c9f1f1
|
fix(proxy): recover session key owners from daily spend for usage attribution (#43642)
* fix(proxy): recover daily spend key owners Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): simplify daily spend owner recovery Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(proxy): format daily activity metadata Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): cover recovered owner metadata merge Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): bound the daily spend owner lookup with the statement timeout * test(integration): audit the daily activity key owner fallback on every usage route Thirty five integration cells under tests/integration/spend cover the daily spend owner fallback on all nine daily activity routes and /usage/ai/chat: the happy path per route, the unanimity rules (two users, blank and null rows, an owner the user table lacks, live and deleted keys with and without their own user, a spend log alias), a non admin reader, an invalid key, a 5 KB key, a locked LiteLLM_DailyUserSpend, 300 keys of one team, repeated reads, a second user landing between reads, a concurrent burst across the unified endpoints, a killed worker, and a proxy restart The traffic cells ignore the GET /v1/models call the proxy's five minute token limit refresh makes to every registered OpenAI compatible deployment, since it lands on a test's provider wire whenever the refresh instant falls inside the test --------- Co-authored-by: jesus <jesus@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
2d034bb35b
|
perf(proxy): hold one spend counter batch across admission and across post-call accounting (#43369)
Auth's spend counter MGET scope spans common checks, model budget check and reservation; reservation increments go out as one pipeline; post-call reconcile adjustments ride the ordinary increment pipeline and update_cache uses one batched read. Over-budget reservation counters are charged one at a time so a rejection never touches the counters after it; post-call counter keys are derived from ids without validating a UserAPIKeyAuth. Resolves LIT-8881 Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
e814532033
|
fix(streaming): keep the served service_tier on streamed chunks and spend rows (#42870)
* fix(streaming): keep the provider's served service_tier on streamed chunks and spend rows Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(streaming): satisfy type-discipline and strict ruff budgets Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(streaming): stamp the served service_tier on every Responses bridge chunk Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(anthropic-adapter): expose streamed chunks so disconnects bill partial spend Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(service-tier): cover anthropic and responses served-tier billing paths Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(anthropic-adapter): return a chunks-exposing stream so disconnects bill partial spend Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(service-tier): bill disconnects through the router's anthropic stream wrapper Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style: apply ruff format to the anthropic stream changes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(coverage): ignore delegating properties the ast scan cannot see Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style: keep the cast-ok reasons on the cast call line Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover served service_tier billing for streamed chat and messages, complete and disconnected Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(anthropic-cache): delegate chunks/messages/model through the messages stream cache writer Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(streaming): keep service_tier on OpenAI-compatible parsed chunks Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(streaming): parameterize delegated chunks and messages types Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(tests): follow the anthropic pass_through rename after merging main Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): drain the logging worker between response cache tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(spend): cover azure, databricks, responses bridge and gemini served tiers in the stream billing integration test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(databricks): keep the served service_tier on streamed chunks and bill it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(databricks): type the served service_tier chunk without a loose kwargs dict Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost): bill the served service_tier over the requested one Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(cost): drop explanatory comment from the tier resolution Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: kerry <kerry@berri.ai> |
||
|
|
fb74957ddd
|
fix(guardrails): enable explicit PANW MCP output scanning (#43109)
* fix(guardrails): declare post_mcp_call for PANW Prisma AIRS Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(guardrails): exercise post_mcp_call_hook dispatch in PANW tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(model-catalog): add fal_ai resolution-tiered image cost fields Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): keep post_mcp_call opt-in for PANW Prisma AIRS Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(model-catalog): add fal_ai resolution-tiered image cost keys to cost map schema Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: joshua <joshua@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
3b2a447fae
|
fix(autorouter): compare historical and new savings consistently (#43348)
* fix(autorouter): compare historical and new savings consistently * fix(autorouter): reject comparisons if request counts changed * fix(router): restore eligible LLM and classification breakdown * fix(router): avoid ambiguous baseline labels for partial comparisons |
||
|
|
abc85c2651
|
fix(cost_calculator): bill chat per-second pricing once with a new cost_per_second field (#43614)
* feat(cost_calculator): add cost_per_second for chat per-second pricing Keep legacy input_cost_per_second and output_cost_per_second as aliases for chat, completion, embedding and responses. When both legacy fields are set, input_cost_per_second wins Move Bedrock commitment rows to cost_per_second so they bill once Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(cost_calculator): drop legacy per-second fields from chat paths Keep Azure chat token pricing generic and update inert Voxtral rates and SageMaker examples Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost_calculator): recognize output-only per-second rates Include output_cost_per_second when checking whether a deployment cost entry has pricing so output-only legacy aliases remain attached to the deployment during cost selection Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(pricing): cover cost_per_second and legacy per-second aliases through the proxy Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(cost_calculator): drop output_cost_per_second as a chat per-second alias Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(cost_calculator): restore output_cost_per_second as a chat per-second fallback Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cost-map): keep input_cost_per_second on bedrock commitment rows for older clients Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
b3dcf8208d
|
test(integration): callback credential canary slots C1-C3 and D5 (#43630)
* test(integration): credential canary suite harness Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix. * test(integration): widen canary route sweep and harden the rig Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy. * test(integration): descend into any decoded value that can still hold an encoded canary * test(integration): bound canary decoding by depth and decoded bytes * test(integration): scope log-table and spend-log reads to the scenario window * test(integration): sweep spend-log rows in the scenario date window * test(integration): keep spend-log date window summarized * test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot * test(integration): expect 404 from the caller-scoped team membership route * test(integration): use the rig's own master key and expect 404 from submission lookups * test(integration): check the overridden rig key without assuming the default key is unknown * test(integration): callback credential canary slots C1-C3 and D5 Team callback, team callback_settings, config default_team_settings and key metadata.logging Langfuse secrets, a team Datadog dd_api_key, and request-body Langfuse keys (allow_client_side_credentials) must reach only their sink. Each scenario checks its sink received the canary as auth and that the marker is visible at the stored body, the Logs drawer route and the sink. Adds a unit test that the stored request body snapshot carries no callback parameter. * test(integration): give the callback sink waits a wider bound * test(integration): sweep provider requests for callback credentials |
||
|
|
0c553f0398
|
feat(guardrails): send a configured gateway_name from noma_v2 to Noma (#43678)
* feat(guardrails): send a configured gateway_name from noma_v2 to Noma The noma_v2 guardrail accepts a gateway_name param, falling back to the NOMA_GATEWAY_NAME env var. The value is stripped, and when it is non-empty it goes out as a top-level gateway_name field on /litellm/guardrail. The param works for both guardrail: noma_v2 and guardrail: noma with use_v2, and it is appended after the existing constructor params so positional callers keep their meaning * chore(ui): regenerate OpenAPI snapshot and dashboard types for gateway_name The new noma_v2 gateway_name param shows up in the proxy OpenAPI spec, so the lazy snapshot and the generated dashboard types need regenerating * Update litellm/proxy/guardrails/guardrail_hooks/noma/noma_v2.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> --------- Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> |
||
|
|
7f95b5f361
|
refactor: clean up fresh tech debt from 2026-09-28 (#43674)
* refactor: clean up fresh tech debt from 2026-09-28 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor: group leaderboard rows in one pass and wrap docstring at 120 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
85dc7cb62e
|
fix(proxy): resolve model_group_alias in the zero-cost budget predicate (#43512)
`_is_model_cost_zero()` reads a group's cost through `Router.get_model_group_info()`, which resolves `model_group_alias`, and then gates that on `_is_cost_explicitly_configured()`, which scanned `Router.model_list` for an exact `model_name` match. Alias names live only in `Router.model_group_alias` and are never `model_name` entries, so the scan found nothing and returned False. That False means "the zero cost was defaulted, not configured" (the sparse auto-registration gate added for #24770), so a model priced explicitly at 0 had budget enforced against it when requested through an alias, while the same deployment under its own name was exempt. Both names route to the same deployment and add nothing to spend. The two lookups in one function disagreeing is the bug, so they now share one resolution: `_is_cost_explicitly_configured()` resolves through `Router.get_model_list()`, the same alias-aware path `get_model_group_info()` takes. That also reaches a deployment which prices itself through its `model_info` block, whose cost-map entry lands under the deployment id. `_group_declares_explicit_cost()` was an alias-aware copy of this function, wired only into `model_has_no_cost_mapping()` and never into the budget path; its body is what `_is_cost_explicitly_configured()` now carries, and both callers share it so the two cannot drift apart again. `_has_ptu_flat_cost()` scanned `model_list` the same way and runs after the gate above, so resolving one without the other would let an aliased PTU group — explicit zero per-token price alongside a flat capacity cost — pass as free. It resolves the same way now. Tests cover the predicate and the request path it feeds: over-budget requests through `_should_skip_budget_checks()` into `common_checks()` for an aliased free model (allowed) and an aliased paid model (refused), the predicate for free, paid, PTU, hidden and dangling aliases, and `model_has_no_cost_mapping()` through an alias so the other caller of the shared check stays covered. Unchanged: priced groups (the predicate returns False before the gate), unmapped groups whose zero cost was defaulted (#24770), hidden aliases and aliases pointing at a nonexistent group (`get_model_group_info()` returns None for both, so the cost is unknown and budget is enforced), and non-aliased PTU groups. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
118ce3cc91
|
feat: add model leaderboard page (#43649)
* feat(proxy): add model leaderboard analytics * feat: add model insights task and range constants * feat: record task type from task tags in model usage rollup * feat: serve 365 days of model insights by UTC date * test: cover task tag resolution in model usage rollup * test: update model insights range limit test to 365 days * chore: regenerate dashboard api types for model insights * feat: add model insights aggregation helpers * test: cover model insights aggregation helpers * feat: redesign model leaderboard with stacked bars, treemap and ranking * test: update model leaderboard view test * feat: mark model leaderboard as beta in sidebar * chore: sync schema.prisma copies from root * fix: only treat task: prefixed tags as model insight tasks * feat: add metric type for model insights ranking * fix: rank model insights by selected metric and scope detail queries to ranked deployments * test: plain tags are not model insight tasks * test: cover metric ranking, deployment scoping and rollup round trip * fix: build model insights weeks and halves from the requested date range * test: cover empty weeks and range-based change comparison * fix: refetch by metric, show load errors and ignore stale responses * test: cover metric refetch and error state * feat: define model insight tasks in a JSON file * feat: return task labels and categories from model insights * feat: load model insight tasks from JSON * refactor: validate rollup task tags against the JSON task list * feat: serve the task list with model insights * refactor: drop hardcoded task list from constants * build: ship model insight tasks JSON in the wheel * test: cover model insight task JSON * refactor: take task labels and categories from the API * test: pass task info to task tile builder * refactor: color treemap by API-provided category * test: include tasks in model leaderboard fixture * fix: make daily model usage migration idempotent * feat: bound the model insights task query size * fix: compute task breakdown independent of the chart metric * test: task breakdown is stable across chart metrics * chore: regenerate lazy openapi snapshot for model insights * chore: regenerate dashboard api types for model insights * fix: keep previous ranking dimmed while a new metric loads * test: cover stale metric state in model leaderboard * refactor: drop task row cap constant * fix: return the full task breakdown instead of a truncated one * test: task query is not truncated * feat: add task summary types for model insights * feat: summarise tasks server-side on a separate model insights endpoint * test: cover the model insights tasks endpoint * chore: regenerate lazy openapi snapshot for model insights tasks * chore: regenerate dashboard api types for model insights tasks * refactor: drop client-side task aggregation * test: remove client-side task aggregation tests * feat: load task breakdown separately from the chart metric * test: task breakdown is not refetched on chart metric change --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
ce25856424
|
feat(mcp): scan and pin upstream tool descriptions (#43283)
Some checks failed
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
* feat(mcp): scan and pin upstream tool descriptions
Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.
* chore: sync schema.prisma copies from root
* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes
* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending
The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.
* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper
* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift
* refactor(mcp): trim the tool catalog guard docstrings to one line
* test(mcp): cover guarded discovery boundaries and response definitions
* fix(mcp): bound discovery guardrail concurrency per catalog
* fix(mcp): scan tool catalogs in bounded parallel batches
* fix(mcp): hide pinned catalogs from restricted management views
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
|
||
|
|
98c710c411
|
fix(guardrails): preserve Presidio output selection and restoration (#43401)
* fix(guardrails): preserve Presidio output callback intent and tag selection * test(guardrails): verify Presidio callback stages after registry updates * fix(guardrails): preserve standalone Presidio token behavior --------- Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
7e383c9f6a
|
revert: "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" (#43377)
* Revert "feat(usage): search team keys beyond the top-N in the Team usage view (#42857)" * revert: "feat(usage): search keys beyond the top-N usage subset (#42827)" (#43378) * Revert "feat(usage): search keys beyond the top-N usage subset (#42827)" * revert: "feat(proxy): add LiteLLM_DailyGlobalSpend key-free rollup for the usage dashboard (#41324)" (#43595) * Revert "Merge pull request #41324 from BerriAI/litellm_daily_global_spend_table" * Revert "Merge pull request #41293 from BerriAI/litellm_usage_key_free_aggregate_split" (#43596) Co-authored-by: yassin <yassin@berri.ai> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
317430db4e
|
fix(panw_prisma_airs): honor experimental_use_latest_role_message_only on every request shape (#42447)
* fix(panw_prisma_airs): apply experimental_use_latest_role_message_only to every request shape Explicit true/false now applies to chat completions, Anthropic /v1/messages and /v1/responses alike; unset keeps latest-only for Anthropic and full history otherwise. Text indices are mapped back to their source message by value instead of by count, so Responses instructions, function_call_output and reasoning items no longer derail the alignment and silently rescan the whole history Co-authored-by: scthornton <scthornton@gmail.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(panw_prisma_airs): type latest-message helpers against AllMessageValues Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(panw_prisma_airs): require forward and reverse text attribution to agree A Responses function_call_output whose text equals the latest user turn could claim that turn's slot in a forward-only walk and demote the latest-only scan to an earlier message. Walk both directions and fall back to the full role-filter scan when they disagree Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(panw_prisma_airs): pick the latest human turn from messages, not from aligned texts An image-only latest user turn no longer promotes an earlier user turn into the latest-only scan; it scans nothing on the request side, as the Anthropic path did before. A latest user/developer message whose text never reached texts (a trailing Responses reasoning item) falls back to the role-filter scan instead of narrowing. Types the test helpers, drops the narrating docstrings and adds regressions for both shapes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(panw_prisma_airs): log when latest-only selection leaves nothing to scan An image-only latest user turn with experimental_use_latest_role_message_only=true intentionally yields zero scanner calls. Emit a debug line naming the call_id so operators can tell this apart from the guardrail not firing. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(panw_prisma_airs): keep Responses reasoning items out of latest-turn selection The Responses translation handler gives reasoning input items the default user role, so a reasoning item with text content after the latest prompt was picked as the latest human turn and the real prompt went unscanned under experimental_use_latest_role_message_only. Map reasoning items back to their texts positions from the raw input and exclude them; fall back to the role-filter scan when the raw items do not account for every text Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: scthornton <scthornton@gmail.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
6d4ccf7e97
|
refactor(mcp): consolidate hub publication predicate (#43394) | ||
|
|
fe76c2473d
|
Revert "feat(proxy): server-side Team Usage export beyond the top-N key cap (#42996)" (#43376)
This reverts commit
|
||
|
|
f20c400374
|
fix(proxy): unregister logging callbacks removed from the stored config (#43428)
* fix(proxy): unregister logging callbacks removed from the stored config POST /config/callback/delete saved the config and resynced, but the resync only ever added callbacks, so a deleted callback kept exporting and kept showing in /get/config/callbacks as read-only on every worker. ProxyConfig now tracks which callback list entries each DB config sync registered and unregisters them once the stored config stops listing them. Callbacks it did not register (YAML, code) are never touched, and a failed config load skips the sync instead of treating the config as empty. * refactor(proxy): keep callback sync comprehensions to one for clause * fix(proxy): restore code-registered callbacks the DB sync replaced Registering a custom-logger callback from the DB swaps an existing string entry for a logger instance. Deleting the DB entry then removed the instance and left the code-registered callback gone. The sync now records the entries it displaced and puts them back when it unregisters. |
||
|
|
3726ce2cfc
|
refactor(guardrails): fix agent 365 to the production endpoint and log the opt-in fail_open at error level (#43189)
* feat(guardrails): fail open by default when Agent 365 cannot evaluate and count it in Prometheus Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(tests): ruff format the Prometheus fail-open registry test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(guardrails): add Agent 365 authority host override, fail-open integration test and per-guardrail YAML default Add `authority_host` to the Agent 365 config (also read from AGENT365_AUTHORITY_HOST, then AZURE_AUTHORITY_HOST) so sovereign clouds and the integration test can point the OBO exchange at a different Entra host. Add tests/integration/mcp/test_mcp_agent_365_guardrail.py, a real proxy test with Postgres, Redis, a scripted MCP upstream and local Entra and Agent 365 doubles covering the default fail-open, explicit fail-closed and fail-open, Defender Skipped, policy denial, persisted status and Prometheus counter. Use PrometheusLogger.get_instance for the fail-open metric lookup instead of a hand-rolled callback scan. Clarify the config description: gateway credential failures fail open, caller token failures block. Extract the dashboard YAML preview into teamGuardrailConfigYaml.ts so the effective per-guardrail default is unit tested and the "default" hint only shows when nothing was set explicitly. Regenerate the lazy OpenAPI snapshot and schema.d.ts for the new field. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): default a scheme-less Agent 365 authority host to https and treat a null fallback as unset in the YAML preview Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit cells for the Agent 365 fail-open default across entry points, Entra faults, throttling and two workers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): prove both Agent 365 workers serve and that a killed worker is replaced Each fresh connection reports its worker pid from /debug/memory/summary and its MCP catalog on the same connection, so the two-worker readiness wait covers both workers by identity. The kill test now kills a pid the proxy reported as a worker and waits for a replacement pid, instead of the first psutil child Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): drop the prometheus fail-open counter from the agent 365 guardrail Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): keep agent 365 fail closed by default and make fail_open an explicit opt-in Restores the shared unreachable_fallback default and the sibling guardrail initializers, drops the Admin UI YAML preview that only existed for the per-guardrail default, and reworks the unit and integration tests so the default blocks with HTTP 503 while unreachable_fallback: fail_open lets availability failures through as Unscanned Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): append authority_host after the existing Agent365Guardrail parameters Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(guardrails): agent 365 fails open by default and hides the production overrides from the UI form Agent 365 sits in the runtime path of every MCP tool call, so an Entra or Agent 365 outage now lets the call through unscanned (logged at error level, recorded as Unscanned with guardrail_failed_to_respond) instead of blocking it. unreachable_fallback: fail_closed stays as the opt-in strict mode. Policy blocks, throttling, 4xx rejections and a rejected caller token still block The shared unreachable_fallback field becomes nullable so each guardrail owns its default; every sibling still resolves None to fail_closed and typesafe keeps failing open api_base, resource_app_id and agent_id have production defaults and leave the dashboard form (ui_hidden); they stay available in config.yaml and env. The authority_host override and its env keys are gone, the OBO exchange always uses login.microsoftonline.com. The integration suite keeps only the cells that need no Entra double, the evaluation paths live in unit tests with an injected handler Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(ui): regenerate openapi snapshot and schema.d.ts for the nullable unreachable_fallback Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): fix agent 365 to the production endpoint and keep fail_closed as the default Remove api_base, resource_app_id and agent_id from the Agent 365 config model, their AGENT365_* env fallbacks and the _is_ui_hidden helper: the evaluation URL and the Agent Tools app id are fixed production constants and the agent identity is always the caller's key alias. Revert the fail_open default; unreachable_fallback: fail_open stays an explicit opt-in. Restore the shared unreachable_fallback field, the sibling guardrail initializers and typesafe to main. Move the Entra dependent cells from the subprocess integration suite to unit tests with an injected HTTP handler. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(guardrails): warn when agent 365 yaml still carries the removed override keys Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(guardrails): inject the http handler into the agent 365 initializer instead of assigning it after construction Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
703eb4fa68
|
security(proxy): keep team callback credentials out of the stored request body (#43217)
* security(proxy): keep team callback credentials out of the stored request body Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci: allow the security conventional commit type Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: oliver <oliver@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
f191e08d67
|
fix(proxy): log key owner identity on expired key auth failures (#43105)
* fix(proxy): log key owner identity on expired key auth failures Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): escape control characters in logged key identity fields Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(proxy): gate key identity in auth failure logs behind log_auth_failure_key_identity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): apply log_auth_failure_key_identity from DB config reloads Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: mrinal <mrinal@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
491d454826
|
fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking (#43414)
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking
Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.
Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.
(cherry picked from commit
|
||
|
|
2101c860c2
|
fix(vertex_ai): make Gemma fake streams work with traced Responses (#43147)
* test(vertex_ai): reproduce traced Gemma Responses stream failure * fix(vertex_ai): wrap Gemma fake streams for Responses tracing * test(vertex_ai): cover Gemma traced streams and usage options * test(vertex_ai): inject gemma test deps and assert hidden usage accounting Replace class-level patches in the Vertex AI shard test with the provider's documented dependency-injection seams (httpx.MockTransport client + credential cache), and pin the default/omit-usage trace behavior: LiteLLM still accounts all tokens; ddtrace's metric is absent by design, asserted rather than silent. Mutation-checked: commenting out CustomStreamWrapper chunk accumulation turns the new assertions red; restoring them turns green. * test(vertex_ai): drop explanatory comment from usage-option assertions |
||
|
|
013d5fa015
|
feat(cli): reuse saved agent setup and add reconfigure (#43392) | ||
|
|
7aba77197d
|
feat(otel): add SigNoz preset for OpenTelemetry v2 (#43296)
Some checks are pending
Unit Tests: Documentation Validation / documentation (push) Waiting to run
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
* feat(otel): add SigNoz preset for OpenTelemetry v2 Adds the signoz callback (OTLP/HTTP exporter, GenAI vocabulary, key and team level dynamic ingestion endpoint and key) as an OpenTelemetry v2 preset, with the preset factory accepting the allow_missing_credentials kwarg the V2 registry always passes so construction no longer falls back silently to legacy OpenTelemetry. Ships the deterministic tests/integration/observability/test_signoz_delivery.py audit suite Absorbs the work from https://github.com/BerriAI/litellm/pull/38206 Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(otel): drop explanatory comments from the SigNoz preset Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(types): keep signoz dynamic param lines within ruff format width Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(signoz): assert the missing-endpoint boot path directly instead of in an except block Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(ui): regenerate schema.d.ts for the signoz health service Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): allowlist SigNoz key/team endpoints and route keyless collectors without the operator key Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(otel): terminate the SigNoz shutdown cell before the flush and drop test docstrings Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(otel): keep the shared tenant routing untouched and require an ingestion key for SigNoz key/team endpoints Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(otel): warn about a keyless SigNoz team endpoint from the header resolver so the shared cache actually reaches it Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
eea1d0f269
|
fix(responses): stream guardrail pre-call block as SSE with a typed output item (#42507)
* fix(responses): stream guardrail pre-call block as SSE with a typed output item Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): import blocked usage helper from the guardrail utils module Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(e2e): drop narrating docstrings and poll without rebinding Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover pre-call guardrail block on /v1/responses stream and json Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): audit cells for responses guardrail block contract Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): observe upstream on the recorded chat route for responses denial cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): tidy responses denial audit cells Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): wait for worker count to recover after SIGKILL Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): require a replacement worker after SIGKILL Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(responses): type the blocked response test helpers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: yucheng <yucheng@berri.ai> |
||
|
|
5d777c16d9
|
fix(mcp): align hub publication status and controls (#43241)
* fix(mcp): align hub publication status and controls * refactor(mcp): keep hub visibility guard outside table rendering --------- Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
7244040908
|
fix(mcp): report reachability without stored credentials (#43240) | ||
|
|
1474ea53e6
|
feat(proxy): add maximum_daily_tag_spend_retention_period cleanup setting (#39221)
* feat(proxy): add maximum_daily_tag_spend_retention_period cleanup setting Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore(ui): regenerate schema.d.ts for new retention setting Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(proxy): rebase daily tag spend retention onto the run-budgeted cleanup job Reworks the cleanup on top of the refactored SpendLogCleanup: the daily tag spend table is pruned through the shared batched delete with a text cutoff on the indexed ISO date column, the setting is picked up by /config/update and the scheduler registration, and an integration test proves rows older than the period are pruned while the cutoff day and unset retention are left alone Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): schedule the cleanup job when a retention db row lands before the side effects run A config reload applies the db row to the SettingsStore before _update_general_settings snapshots the previous retention values, so the before/after compare saw no change and a retention period first set through /config/update never scheduled the cleanup job. Also reschedule when the job is missing but a retention period is set Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): accept list-valued top-level keys in the base integration proxy config The shared tests/integration/proxy_config.yaml now carries list-valued top-level keys, so the retention config helper validates only the mapping it merges into. Also drops a SQL-shape assertion from the unit test in favor of the behavioral cutoff-day check Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): cover runtime update, invalid value, independent horizons and worker loss for daily tag spend retention Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): restore the shared retention setting, capture seeded days once and kill a listening worker Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): retry a failed cleanup schedule only when its settings change _apply_retention_settings rescheduled whenever retention was set and no job existed, so an unparseable cleanup cron was retried on every config reload. Remember the last attempted retention, cron and interval tuple and retry only when it differs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(integration): give daily tag spend retention tests a 240s timeout Each node boots a proxy and waits for a whole-minute cleanup cron tick, so the global 90s pytest-timeout can expire during teardown on a slow runner, as integration-accounting did on pipeline 90302 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): reschedule cleanup when only the cron or interval changes at runtime Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): reschedule cleanup when the first db sync changes only the cron or interval Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(ui): add text input for String general settings so retention periods can be set from the Admin UI Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): record a cleanup schedule attempt only after it did not raise Records _last_cleanup_schedule_attempt after _reschedule_spend_log_cleanup_job returns, so a transient add_job error is retried on the next config sync while an invalid cron, which is caught and logged inside the reschedule, is still attempted once per settings value Also adds --num_workers 2 to the dev proxy command in AGENTS.md as requested on the PR Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs: revert unrelated AGENTS.md dev command change Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Revert "docs: revert unrelated AGENTS.md dev command change" This reverts commit |
||
|
|
e47b1f2a3f
|
fix(s3_v2): upload fresh events first, drop terminal failures and hour-old retries by default, opt-in adaptive concurrency (#43022)
* fix(s3_v2): drop terminal upload failures, bound retries per flush and enforce the queue cap at enqueue Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): keep retrying credential-rotation 403s, only AccessDenied-style errors are terminal Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(env_keys): exclude DEFAULT_S3_MAX_FLUSH_ATTEMPTS as an internal tuning var Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): retry every 5xx, warn on first queue overflow, validate the flush budget Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): read the queue cap defensively so un-initialized loggers still enqueue Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): drop the getattr in _enqueue and tighten the retry tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): keep the constructor flush budget when the callback override is invalid Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(s3_v2): adapt per-object upload concurrency to sink latency and throttling Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(s3_v2): tidy adaptive limiter Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * perf(s3_v2): wake one waiter per released upload slot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(s3_v2): make the enqueue queue cap configurable with s3_max_queue_size Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): audit cells for cache hits, coded 403 and callback modes Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): retry bucket-wide failures by default, age-budget requeues and make terminal drops and adaptive concurrency opt-in Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(s3_v2): suppress the missing-waiter ValueError explicitly in the adaptive limiter Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): count oldest events trimmed after a failed flush as callback failures Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(s3_v2): move the mutable-ok marker onto the list literal it suppresses Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): report post-flush overflow drops once and grow adaptive concurrency above the floor before asserting back-off Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): fail the SlowDown back-off test when the measured window sees no PUTs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): restore base retry defaults, opt-in age budget, no enqueue cap, back off outside the limiter slot Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): hoist the default no-op upload slot to a module constant Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): drop unused mutable-ok suppressions on queue appends Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): keep the retry queue oldest-first and prioritise fresh events at upload time Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(s3_v2): keep the mutable-ok marker on the queue list literal Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): build request bodies inside the upload slot and keep the sync retry set at base parity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): drop wall-clock sleeps from the unit tests Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): rebuild the request body inside the slot on every retry attempt Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): anchor the backoff window on the first observed failure and tighten shard assertions Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): default the upload slot to the logger limiter so monkeypatched doubles keep working Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): clear ambient AWS env credentials so the rotating profile signs the sync retry test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): shrink the linear send-batch perf test to 2k/8k elements Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): fix stale batch sizes in the perf test assert message Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): keep async in-call retries on the base 403/500/503 set Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): hold the upload slot across retries, restore the bool upload contract, and fail safe on bool config Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): mark dropped uploads by element identity so a shared key cannot mask a retryable sibling Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): make the per-flush drop lookup constant time Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): take the upload slot in the caller like base, build the body once per attempt loop Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): match base retry, logging and hook behaviour unless the new options are opted in Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): assert the wire key in the init-bypassed sync upload test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): assert the signed headers and wire key in the init-bypassed sync upload test Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): default the upload limiter at class level instead of reading it with getattr Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): drop the duplicate annotations that redeclare the class-level counters Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * feat(s3_v2): drop terminal-failed uploads by default and bound retry age to one hour Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(s3_v2): fall back to the configured retry age on invalid values Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): drive retry-age tests from a fixed clock Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(s3_v2): audit cells for retry-age opt-out and 429 single-put parity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yucheng <yucheng@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
3743c8563e
|
fix(mcp): adopt shared server resolution and caller authorization (#43263)
* test(mcp): characterize server resolution and authorization * refactor(mcp): extract shared server resolution * fix(mcp): adopt shared resolution in management endpoints * fix(mcp): scope credential metadata resolution outside loop * test(mcp): pin catalog isolation and batched credential permissions * fix(mcp): restrict catalog detail and batch credential permissions * test(mcp): enforce identity isolation in database fixtures * test(mcp): name resolution tests by behavior * test(mcp): describe detail access assertion failures * chore: keep agent naming discipline local --------- Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
26bf575f15
|
feat(guardrails): scan retrieved vector store chunks with the request's pre-call guardrails (#43271)
* feat(guardrails): scan retrieved vector store chunks with the request's pre-call guardrails Vector store retrieval runs inside acompletion after the proxy's pre-call guardrails have already seen the request, so a retrieved chunk carrying an injection reached the prompt unscanned. Each retrieved context message now goes through every pre-call guardrail the request is subject to before it is injected: a block raises the same 400 the guardrail gives for user text, a masking guardrail rewrites the context, and a guardrail that fails while scanning fails the request instead of injecting the chunk unscanned * fix(guardrails): return a guardrail block unmapped from exception_type so the Responses API surfaces the guardrail's own 400 * fix(guardrails): build the deployment hooks' identity from stamped metadata only Top-level user_api_key_* fields in a request body are client controlled, so the pre-call, chunk scan, and post-call deployment hooks now take UserAPIKeyAuth from the metadata the proxy stamped, and the chunk scanner returns or raises on every branch. * fix(guardrails): block route verdicts on retrieved chunks, keep guardrail verdicts out of router retries and fallbacks, and scan chunks against the client's request * fix(guardrails): keep the merged guardrail list when scanning chunks against the client's request The scan request laid the client's kept body over the deployment kwargs, so a client that sent its own top-level guardrails list shadowed the merged metadata.guardrails list and a key or team guardrail skipped the chunk scan. The kwargs now win and the keys the proxy relocates into metadata are dropped from the body's contribution. * fix(guardrails): strip the deployment's guardrail keys from the chunk scan request so merged team guardrails still run * test(guardrails): type the vector store scan test doubles * test(guardrails): type the scan double's request data --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |
||
|
|
96c008f420
|
ci: fail on new unbounded SQL IN lists and add a Prisma chunking helper (#42629)
* ci: warn on SQL IN lists with no written bound Postgres caps a prepared statement at 32,767 bind parameters and a membership filter binds one per value, so an IN list built from table data breaks once the table outgrows the cap. That is how the budget reset job froze every due budget (LIT-7535, #40564). check_unbounded_in_lists.py reports every Prisma "in" / "not_in" filter whose value has no fixed size and every raw SQL literal that splices a list in after "IN (", unless the line carries "# bounded-ok: <reason>". It only warns for now: the output is the inventory for RCA action item AI-1, and it exits 0. * ci: decide a constant IN list by its module binding, not its casing An ALL_CAPS name imported or filled at runtime is as unbounded as any other, so a name now passes only when the module binds it once to a value of fixed size. Adds Final to the locals a loop does not forbid. * ci: only a frozen module value makes an IN list constant A module list bound once could still grow through append or extend, so a name now counts as fixed only when it is bound to a tuple, frozenset or constant. Trims the module docstring to what a reader needs. * ci: chunk Prisma IN lists with a shared helper and fail on new unbounded ones Add litellm.repositories.bounded_in: find_many_in, count_in, update_many_in and delete_many_in split a deduplicated value list into 5,000-value chunks, AND each chunk with the caller's where, run them in order (a transaction handle works) and combine the results. Writes take a required atomicity argument, and a where that already filters the chunked field is refused. check_unbounded_in_lists.py now fails CI on any finding missing from unbounded_in_baseline.txt and on any stale baseline entry, so the baseline only shrinks. Entries are keyed by path, enclosing scope, kind, field and occurrence, not line numbers. The helper module is exempt, a constant spread into a frozen tuple counts as fixed, and messages point at the helper for "in" and at an array parameter for "not_in" and raw SQL. A real-Postgres integration test shows a raw 40,000-value filter rejected for too many bind variables while the helpers handle it. * refactor: rename bounded_in to chunked_in and let callers pick a chunk size The helper module is litellm.repositories.chunked_in, and its unit and integration tests, the checker's exemption path and its finding messages follow the new name. The `# bounded-ok` marker is unchanged. find_many_in, count_in, update_many_in and delete_many_in take a keyword-only chunk_size, defaulting to IN_LIST_CHUNK_SIZE (5,000). A value below 1 or above MAX_IN_LIST_CHUNK_SIZE (30,000) raises ValueError before any query, which leaves the rest of the filter headroom under Postgres's 32,767 bind-parameter cap. * refactor: flatten chunked_in's stacked comprehensions with chain.from_iterable LIT014 (#42650) caps a comprehension at one for and one if clause. The four nested walks in the helper now chain their iterables instead, with the same order and results. * refactor: recover user details with find_many_in, sending chunks as lists _details_for_user_ids reads users through find_many_in instead of a raw "in" filter, so its lookup stays under the bind-parameter cap for any number of recovered keys. Up to 5,000 ids it still sends one find_many with the same where dict, and a PrismaError from any chunk is still logged and treated as no details. The helper now sends each chunk as a list, so a chunked filter equals the dict a hand-written call would send and a migrated call site's existing assertions keep passing. The site's baseline entry is gone. * ci: skip functional TypedDict field maps in the unbounded IN list check The dict passed as the field map of TypedDict("Name", {...}), or as its fields= keyword, names fields: an "in" or "notIn" key there is a type, not a filter. Only that dict is skipped, for TypedDict, typing.TypedDict and typing_extensions.TypedDict; a filter nested in a field value or passed to any other call is still reported. The two types/proxy/management_endpoints/team_endpoints.py entries leave the baseline, which is now 156. * fix: refuse an update_many_in whose data writes the chunked field Chunks run one after another, so an update that sets the chunked field can move a row into a later chunk, which updates it again and counts it twice: values ["old", "new"] with chunk_size=1 and data={"id": "new"} does exactly that. update_many_in now raises ChunkedFieldWriteError before any query when data has the chunked field as a top-level key, in any form, including Prisma operators such as {"set": ...}. * docs: cut the unbounded IN list checker's docstring to what it flags and how to clear it It now says what is reported, the three ways to clear a finding, and how the baseline and --update-baseline work, in 11 lines. The per-shape detail lives in the tests. * ci: key an unbounded IN list finding by its filtered expression too A baseline key of path, scope, kind, field and occurrence let a PR delete a baselined filter and add a different unbounded one on the same field in the same function, and the new one took over the old key. The key now also carries the filtered expression's source, whitespace-normalized (the Prisma value, or a raw-SQL `IN (...)` slot), so that swap reads as one new and one stale entry and fails the run. The same expression re-added in the same function is still the same finding. Every baseline entry is rewritten in the new form; the 156 findings are unchanged, and only occurrence indexes renumber where one field had several different expressions. |
||
|
|
40297e6268
|
refactor(mcp): add shared server resolver without changing callers (#43262)
* test(mcp): characterize server resolution and authorization * refactor(mcp): extract shared server resolution * test(mcp): pin catalog isolation and batched credential permissions * test(mcp): enforce identity isolation in database fixtures * test(mcp): name resolution tests by behavior * test(mcp): describe detail access assertion failures * chore: keep agent naming discipline local --------- Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
69ad004015
|
refactor(anthropic): rename experimental_pass_through to pass_through (#43329)
* refactor(anthropic): rename experimental_pass_through to pass_through Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(anthropic): point compact patch targets at renamed pass_through path Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: kerry <kerry@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0300fc4ab2
|
test(mcp): pin server resolution and authorization behavior (#43261)
* test(mcp): characterize server resolution and authorization * test(mcp): pin catalog isolation and batched credential permissions * test(mcp): enforce identity isolation in database fixtures * test(mcp): name resolution tests by behavior * test(mcp): describe detail access assertion failures * chore: keep agent naming discipline local --------- Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com> |
||
|
|
dfb5d905ea
|
fix(guardrails): block private destinations in custom code http_request and bound guardrail execution time (#43280)
* fix(guardrails): block private destinations in custom code http_request and bound guardrail execution time * fix(guardrails): keep startup fail-closed on a custom code compile error and report a load timeout on the test endpoint A compile failure is no longer a ValueError, so a config-file custom code guardrail that does not compile stops the proxy at startup as it did before, while POST /guardrails catches it by name and still rolls back. The admin test endpoint reports a module-level timeout as an execution timeout instead of a compile error, a caller-supplied Host header is stripped from http_* requests while validation is on, and GET keeps the shared client's connect timeout. * test(guardrails): cover the http_request methods, header passthrough and cancellation paths --------- Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com> |