* feat(proxy): add model leaderboard analytics
* feat: add model insights task and range constants
* feat: record task type from task tags in model usage rollup
* feat: serve 365 days of model insights by UTC date
* test: cover task tag resolution in model usage rollup
* test: update model insights range limit test to 365 days
* chore: regenerate dashboard api types for model insights
* feat: add model insights aggregation helpers
* test: cover model insights aggregation helpers
* feat: redesign model leaderboard with stacked bars, treemap and ranking
* test: update model leaderboard view test
* feat: mark model leaderboard as beta in sidebar
* chore: sync schema.prisma copies from root
* fix: only treat task: prefixed tags as model insight tasks
* feat: add metric type for model insights ranking
* fix: rank model insights by selected metric and scope detail queries to ranked deployments
* test: plain tags are not model insight tasks
* test: cover metric ranking, deployment scoping and rollup round trip
* fix: build model insights weeks and halves from the requested date range
* test: cover empty weeks and range-based change comparison
* fix: refetch by metric, show load errors and ignore stale responses
* test: cover metric refetch and error state
* feat: define model insight tasks in a JSON file
* feat: return task labels and categories from model insights
* feat: load model insight tasks from JSON
* refactor: validate rollup task tags against the JSON task list
* feat: serve the task list with model insights
* refactor: drop hardcoded task list from constants
* build: ship model insight tasks JSON in the wheel
* test: cover model insight task JSON
* refactor: take task labels and categories from the API
* test: pass task info to task tile builder
* refactor: color treemap by API-provided category
* test: include tasks in model leaderboard fixture
* fix: make daily model usage migration idempotent
* feat: bound the model insights task query size
* fix: compute task breakdown independent of the chart metric
* test: task breakdown is stable across chart metrics
* chore: regenerate lazy openapi snapshot for model insights
* chore: regenerate dashboard api types for model insights
* fix: keep previous ranking dimmed while a new metric loads
* test: cover stale metric state in model leaderboard
* refactor: drop task row cap constant
* fix: return the full task breakdown instead of a truncated one
* test: task query is not truncated
* feat: add task summary types for model insights
* feat: summarise tasks server-side on a separate model insights endpoint
* test: cover the model insights tasks endpoint
* chore: regenerate lazy openapi snapshot for model insights tasks
* chore: regenerate dashboard api types for model insights tasks
* refactor: drop client-side task aggregation
* test: remove client-side task aggregation tests
* feat: load task breakdown separately from the chart metric
* test: task breakdown is not refetched on chart metric change
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): drive the router request test through an httpx MockTransport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): assert tool definitions reach the router upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(codeowners): drop UI and migration code owners
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(codeowners): drop CODEOWNERS self-owner line
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost-map): add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(cost-map): keep bedrock_mantle claude 5.5 change to cost map rows only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(mcp): scan and pin upstream tool descriptions
Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.
* chore: sync schema.prisma copies from root
* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes
* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending
The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.
* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper
* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift
* refactor(mcp): trim the tool catalog guard docstrings to one line
* test(mcp): cover guarded discovery boundaries and response definitions
* fix(mcp): bound discovery guardrail concurrency per catalog
* fix(mcp): scan tool catalogs in bounded parallel batches
* fix(mcp): hide pinned catalogs from restricted management views
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* fix(ui): rename All Models tab to Deployed Models and view filter to All Proxy Models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): use All Proxy Models label for the public model name filter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): rename ALL_MODELS_VIEW constant to ALL_PROXY_MODELS_VIEW
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): stored-config credential canary slots
Add canary slots for credentials the proxy holds in its env, config or
database: virtual key raw value, master key, deployment api_key via
/model/new, named credentials, AWS secret key, Vertex service-account
JSON and its minted token, team model_config credential overrides,
config guardrail api_key, and sink credentials from env.
* test(integration): resolve the config guardrail id and require detail routes to return 200
* test(integration): drop repeated timeout comments
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* test(integration): sweep config-deployment routes with the real model id and use the rig admin for the master-key slot
* test(integration): mark the configure-hook config edits as intended
* test(integration): drop suppression markers that suppress nothing
* Use claude-haiku-4-5 for the Bedrock stored-credential test model
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): credential canary slots for MCP and pass-through credentials
Adds slots F1 (MCP static auth), F2 (per-user MCP OAuth token), F2E (per-user
MCP env var), F3 (x-mcp client auth header), H1 (pass-through credential header),
H2 (vector store api_key) and H2S (search tool api_key) to the credential canary
suite. The OAuth double gains an optional mint hook so a test can choose the
issued access token.
* test(integration): wait for MCP spend rows by call type
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): canary MCP and pass-through slots pass resolved ids
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* ci: sync the weekly release cycle with Linear releases
* tmp: dry-run trigger
* tmp: backfill 1.105.0 from the rc/1.104.0 cut
* ci: drop the temporary branch trigger used to verify the Linear sync
* ci: fail on a broken rc-branch lookup and never move a just-cut release back to main
* tmp: dry-run trigger
* ci: drop the temporary branch trigger again
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* feat(providers): add Prism provider
* fix(providers): complete Prism registration
* feat(providers): expose Prism responses and messages
* feat(providers): add DeepSeek V4.1 Flash to Prism
* test(providers): exercise Prism endpoint requests
* fix(providers): align Prism pricing and limits with the live catalog
deepseek-v4.1-flash bills 0.17/0.63 USD per 1M input/output tokens and takes image input;
deepseek-v4-flash bills 0.17/0.21 and caps output at 384000 tokens, per GET /v1/models
* test(prism): assert cost-map invariants instead of pinning catalog facts
* test(prism): derive the asserted model list from the cost map instead of pinning it
* test(prism): capture requests through respx instead of appending to a list and swapping the client transport
---------
Co-authored-by: rajitkhanna <rajitskhanna@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
* fix(panw_prisma_airs): apply experimental_use_latest_role_message_only to every request shape
Explicit true/false now applies to chat completions, Anthropic /v1/messages and /v1/responses alike; unset keeps latest-only for Anthropic and full history otherwise. Text indices are mapped back to their source message by value instead of by count, so Responses instructions, function_call_output and reasoning items no longer derail the alignment and silently rescan the whole history
Co-authored-by: scthornton <scthornton@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(panw_prisma_airs): type latest-message helpers against AllMessageValues
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): require forward and reverse text attribution to agree
A Responses function_call_output whose text equals the latest user turn could claim that turn's slot in a forward-only walk and demote the latest-only scan to an earlier message. Walk both directions and fall back to the full role-filter scan when they disagree
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): pick the latest human turn from messages, not from aligned texts
An image-only latest user turn no longer promotes an earlier user turn into the
latest-only scan; it scans nothing on the request side, as the Anthropic path did
before. A latest user/developer message whose text never reached texts (a trailing
Responses reasoning item) falls back to the role-filter scan instead of narrowing.
Types the test helpers, drops the narrating docstrings and adds regressions for both
shapes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): log when latest-only selection leaves nothing to scan
An image-only latest user turn with experimental_use_latest_role_message_only=true intentionally yields zero scanner calls. Emit a debug line naming the call_id so operators can tell this apart from the guardrail not firing.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): keep Responses reasoning items out of latest-turn selection
The Responses translation handler gives reasoning input items the default user
role, so a reasoning item with text content after the latest prompt was picked
as the latest human turn and the real prompt went unscanned under
experimental_use_latest_role_message_only. Map reasoning items back to their
texts positions from the raw input and exclude them; fall back to the
role-filter scan when the raw items do not account for every text
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: scthornton <scthornton@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(provider): accept 2xx status in handleResponse for access group create
handleResponse (resource_team.go) only accepted exactly HTTP 200, but
POST /v1/access_group (and its /v1/unified_access_group alias) legitimately
answers 201 Created. litellm_access_group and litellm_unified_access_group
create both succeeded on the proxy and failed in the provider, leaving the
group out of state and forcing an import to recover on the next apply's
409 for the now-duplicate name.
Same fix and shape as #40723, which widened this exact check in
sendRequest/handleAPIResponse/handleMCPAPIResponse for mcp_server, model,
key and organization_member. handleResponse is the one shared status-check
helper that fix didn't reach -- it's a different function in a different
file (resource_team.go, not client.go/utils.go), so this is fully
independent of that PR and can land before, after, or alongside it with
no conflict.
handleResponse is also used by agent, budget, guardrail, organization,
prompt, search_tool, tag, team, team_block, key_block, team_member(_add)
and user -- all unaffected in practice, since every one of their own
endpoints already answers exactly 200. Widening the check costs them
nothing and only changes behavior for the two resources that were
actually broken.
Verified: go test ./... passes, including a new
TestHandleResponseAcceptsFullSuccessRange table test covering
200/201/202/204 (accepted) and 400/404/409/500 (still rejected), mirroring
#40723's own TestHandleAPIResponseAcceptsFullSuccessRange.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix: address Greptile review on PR 42461
- CHANGELOG: narrowed the fix's scope to unified_access_group only.
litellm_access_group (legacy) calls /access_group/new, a completely
different, unrelated endpoint (model_access_group_management_endpoints.py)
that already returns 200 -- it was never affected. I'd wrongly assumed
both resources shared the same /v1/access_group route; they don't.
- resource_team_test.go: removed the preamble comment above the new test,
which restated what the table test already shows -- against repository
guidance (AGENTS.md) that reserves comments for complex logic, tool
inputs, or TODOs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* chore: retrigger CI
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(proxy): unregister logging callbacks removed from the stored config
POST /config/callback/delete saved the config and resynced, but the resync only
ever added callbacks, so a deleted callback kept exporting and kept showing in
/get/config/callbacks as read-only on every worker.
ProxyConfig now tracks which callback list entries each DB config sync
registered and unregisters them once the stored config stops listing them.
Callbacks it did not register (YAML, code) are never touched, and a failed
config load skips the sync instead of treating the config as empty.
* refactor(proxy): keep callback sync comprehensions to one for clause
* fix(proxy): restore code-registered callbacks the DB sync replaced
Registering a custom-logger callback from the DB swaps an existing string
entry for a logger instance. Deleting the DB entry then removed the instance
and left the code-registered callback gone. The sync now records the entries
it displaced and puts them back when it unregisters.
* refactor(rust): extract litellm-host-native as the shared Rust host driver
Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): interrupt the machine when the in-process stream consumer fails
Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): separate the machine contract from coroutine execution
* auth update
* refactor(rust): use standard flow control for host requests
* style(rust): keep host driver imports formatted
* chores
* mostly relocation
* refactor(rust): separate interceptors from queued observers
* refactor(rust): centralize legacy callback mappings and lifecycle
* docs: define Python host boundaries and migration plan
* refactor: enforce Python host and bridge boundaries
* refactor(rust): separate operations from callback composition
* refactor(rust): compose SDK policy through call hooks
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
test_ssl_verify_unit.py inserted tests/unit at the front of sys.path, so any later import of litellm_proxy_extras resolved to the tests/unit/litellm_proxy_extras test package. Whenever the CircleCI shard split collected that file before test_litellm_proxy_extras_logging.py, collection failed with ModuleNotFoundError. test_gemini_session_leak.py had the same insert for its own directory
* feat(anthropic): add Claude Sonnet 5.5
Adds the anthropic cost map entry for claude-sonnet-5-5 mirroring
claude-sonnet-5 pricing and capabilities, with prompt_cache_min_tokens
at 512, thinking_always_on (thinking cannot be disabled on this model),
and supports_forced_tool_use false (tool_choice required/named returns
400 upstream). Omits thinking cache preservation, same as Opus 5.5, and
registers the model in the setup wizard provider list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys
Sets prompt_cache_min_tokens 512, thinking_always_on, and
supports_forced_tool_use false on every anthropic, bedrock, vertex_ai,
and azure_ai Sonnet 5.5 key, dropping the thinking cache preservation
flag cloned from Sonnet 5. Removes unpublished deprecation dates on
azure_ai and vertex_ai, renames the OpenRouter key to the live
anthropic/claude-sonnet-5.5 id and drops its batch variant, and removes
the aihubmix, deepinfra, and databricks keys for vendors that do not
list the model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop vendor-absence assertions for Sonnet 5.5 keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(guardrails): fail open by default when Agent 365 cannot evaluate and count it in Prometheus
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(tests): ruff format the Prometheus fail-open registry test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(guardrails): add Agent 365 authority host override, fail-open integration test and per-guardrail YAML default
Add `authority_host` to the Agent 365 config (also read from AGENT365_AUTHORITY_HOST, then AZURE_AUTHORITY_HOST) so sovereign clouds and the integration test can point the OBO exchange at a different Entra host.
Add tests/integration/mcp/test_mcp_agent_365_guardrail.py, a real proxy test with Postgres, Redis, a scripted MCP upstream and local Entra and Agent 365 doubles covering the default fail-open, explicit fail-closed and fail-open, Defender Skipped, policy denial, persisted status and Prometheus counter.
Use PrometheusLogger.get_instance for the fail-open metric lookup instead of a hand-rolled callback scan. Clarify the config description: gateway credential failures fail open, caller token failures block.
Extract the dashboard YAML preview into teamGuardrailConfigYaml.ts so the effective per-guardrail default is unit tested and the "default" hint only shows when nothing was set explicitly.
Regenerate the lazy OpenAPI snapshot and schema.d.ts for the new field.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): default a scheme-less Agent 365 authority host to https and treat a null fallback as unset in the YAML preview
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit cells for the Agent 365 fail-open default across entry points, Entra faults, throttling and two workers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): prove both Agent 365 workers serve and that a killed worker is replaced
Each fresh connection reports its worker pid from /debug/memory/summary and its MCP catalog on the same
connection, so the two-worker readiness wait covers both workers by identity. The kill test now kills a
pid the proxy reported as a worker and waits for a replacement pid, instead of the first psutil child
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): drop the prometheus fail-open counter from the agent 365 guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): keep agent 365 fail closed by default and make fail_open an explicit opt-in
Restores the shared unreachable_fallback default and the sibling guardrail initializers, drops the Admin UI YAML preview that only existed for the per-guardrail default, and reworks the unit and integration tests so the default blocks with HTTP 503 while unreachable_fallback: fail_open lets availability failures through as Unscanned
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): append authority_host after the existing Agent365Guardrail parameters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(guardrails): agent 365 fails open by default and hides the production overrides from the UI form
Agent 365 sits in the runtime path of every MCP tool call, so an Entra or
Agent 365 outage now lets the call through unscanned (logged at error level,
recorded as Unscanned with guardrail_failed_to_respond) instead of blocking it.
unreachable_fallback: fail_closed stays as the opt-in strict mode. Policy
blocks, throttling, 4xx rejections and a rejected caller token still block
The shared unreachable_fallback field becomes nullable so each guardrail owns
its default; every sibling still resolves None to fail_closed and typesafe
keeps failing open
api_base, resource_app_id and agent_id have production defaults and leave the
dashboard form (ui_hidden); they stay available in config.yaml and env. The
authority_host override and its env keys are gone, the OBO exchange always
uses login.microsoftonline.com. The integration suite keeps only the cells
that need no Entra double, the evaluation paths live in unit tests with an
injected handler
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate openapi snapshot and schema.d.ts for the nullable unreachable_fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): fix agent 365 to the production endpoint and keep fail_closed as the default
Remove api_base, resource_app_id and agent_id from the Agent 365 config model, their AGENT365_* env fallbacks and the _is_ui_hidden helper: the evaluation URL and the Agent Tools app id are fixed production constants and the agent identity is always the caller's key alias. Revert the fail_open default; unreachable_fallback: fail_open stays an explicit opt-in. Restore the shared unreachable_fallback field, the sibling guardrail initializers and typesafe to main. Move the Entra dependent cells from the subprocess integration suite to unit tests with an injected HTTP handler.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): warn when agent 365 yaml still carries the removed override keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): inject the http handler into the agent 365 initializer instead of assigning it after construction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Run every route in the team-admin matrix again on a team that belongs to an
organization, as an admin of that organization and as an admin of another
organization. 90 new cases pin what org admins get today, so collapsing the
team-admin helpers into one gate can prove parity for org admins too
Scenario gains organization() and org_member() helpers that clean up the
organization, its budget row and the member users
* security(proxy): keep team callback credentials out of the stored request body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: allow the security conventional commit type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): replace Any with proven types in 8 files
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(types): keep email logger untyped where its alert types differ
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(rust): add litellm-db and litellm-db-testing workspace scaffolding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(db-testing): apply the real Prisma migrations in a test and drop the sort mutation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* build(rust): package the gateway container
* ci: exempt the gateway Dockerfile from the CI coverage gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>