* refactor: clean up fresh tech debt from 2026-09-28
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor: group leaderboard rows in one pass and wrap docstring at 120
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
`_is_model_cost_zero()` reads a group's cost through `Router.get_model_group_info()`,
which resolves `model_group_alias`, and then gates that on `_is_cost_explicitly_configured()`,
which scanned `Router.model_list` for an exact `model_name` match. Alias names live only in
`Router.model_group_alias` and are never `model_name` entries, so the scan found nothing and
returned False. That False means "the zero cost was defaulted, not configured" (the sparse
auto-registration gate added for #24770), so a model priced explicitly at 0 had budget
enforced against it when requested through an alias, while the same deployment under its own
name was exempt. Both names route to the same deployment and add nothing to spend.
The two lookups in one function disagreeing is the bug, so they now share one resolution:
`_is_cost_explicitly_configured()` resolves through `Router.get_model_list()`, the same
alias-aware path `get_model_group_info()` takes. That also reaches a deployment which prices
itself through its `model_info` block, whose cost-map entry lands under the deployment id.
`_group_declares_explicit_cost()` was an alias-aware copy of this function, wired only into
`model_has_no_cost_mapping()` and never into the budget path; its body is what
`_is_cost_explicitly_configured()` now carries, and both callers share it so the two cannot
drift apart again.
`_has_ptu_flat_cost()` scanned `model_list` the same way and runs after the gate above, so
resolving one without the other would let an aliased PTU group — explicit zero per-token
price alongside a flat capacity cost — pass as free. It resolves the same way now.
Tests cover the predicate and the request path it feeds: over-budget requests through
`_should_skip_budget_checks()` into `common_checks()` for an aliased free model (allowed) and
an aliased paid model (refused), the predicate for free, paid, PTU, hidden and dangling
aliases, and `model_has_no_cost_mapping()` through an alias so the other caller of the shared
check stays covered.
Unchanged: priced groups (the predicate returns False before the gate), unmapped groups whose
zero cost was defaulted (#24770), hidden aliases and aliases pointing at a nonexistent group
(`get_model_group_info()` returns None for both, so the cost is unknown and budget is
enforced), and non-aliased PTU groups.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude Code attaches output_config to mid-conversation system messages and
sends the per-turn-control-2026-07-01 beta with it. The Vertex beta map
dropped that beta, so Vertex rejected the body with
'messages.N.output_config: Extra inputs are not permitted'.
Forward the beta for vertex_ai, the way azure_ai already does, and add it on
the Vertex Messages path whenever a message carries output_config.
* fix(otel): send cache and reasoning tokens in langfuse usage_details
The OTel V2 Langfuse mapper only sent input, output and total, so cache reads, cache writes and reasoning tokens never reached Langfuse. Emit them as input_cached_tokens, input_cache_creation and output_reasoning_tokens, and send input/output net of those buckets so Langfuse does not price the same tokens twice.
Fixes#43542
* fix(otel): drop redundant comments from the usage_details change
* feat(proxy): add model leaderboard analytics
* feat: add model insights task and range constants
* feat: record task type from task tags in model usage rollup
* feat: serve 365 days of model insights by UTC date
* test: cover task tag resolution in model usage rollup
* test: update model insights range limit test to 365 days
* chore: regenerate dashboard api types for model insights
* feat: add model insights aggregation helpers
* test: cover model insights aggregation helpers
* feat: redesign model leaderboard with stacked bars, treemap and ranking
* test: update model leaderboard view test
* feat: mark model leaderboard as beta in sidebar
* chore: sync schema.prisma copies from root
* fix: only treat task: prefixed tags as model insight tasks
* feat: add metric type for model insights ranking
* fix: rank model insights by selected metric and scope detail queries to ranked deployments
* test: plain tags are not model insight tasks
* test: cover metric ranking, deployment scoping and rollup round trip
* fix: build model insights weeks and halves from the requested date range
* test: cover empty weeks and range-based change comparison
* fix: refetch by metric, show load errors and ignore stale responses
* test: cover metric refetch and error state
* feat: define model insight tasks in a JSON file
* feat: return task labels and categories from model insights
* feat: load model insight tasks from JSON
* refactor: validate rollup task tags against the JSON task list
* feat: serve the task list with model insights
* refactor: drop hardcoded task list from constants
* build: ship model insight tasks JSON in the wheel
* test: cover model insight task JSON
* refactor: take task labels and categories from the API
* test: pass task info to task tile builder
* refactor: color treemap by API-provided category
* test: include tasks in model leaderboard fixture
* fix: make daily model usage migration idempotent
* feat: bound the model insights task query size
* fix: compute task breakdown independent of the chart metric
* test: task breakdown is stable across chart metrics
* chore: regenerate lazy openapi snapshot for model insights
* chore: regenerate dashboard api types for model insights
* fix: keep previous ranking dimmed while a new metric loads
* test: cover stale metric state in model leaderboard
* refactor: drop task row cap constant
* fix: return the full task breakdown instead of a truncated one
* test: task query is not truncated
* feat: add task summary types for model insights
* feat: summarise tasks server-side on a separate model insights endpoint
* test: cover the model insights tasks endpoint
* chore: regenerate lazy openapi snapshot for model insights tasks
* chore: regenerate dashboard api types for model insights tasks
* refactor: drop client-side task aggregation
* test: remove client-side task aggregation tests
* feat: load task breakdown separately from the chart metric
* test: task breakdown is not refetched on chart metric change
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(fireworks_ai): route and list the auto, auto-instant and firerouter routers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): drive the router request test through an httpx MockTransport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fireworks_ai): let custom firerouter/<models> IDs inherit the firerouter row's capabilities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): integration coverage for router short names forwarding tool_choice and reasoning_effort
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): assert tool definitions reach the router upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(cost-map): add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* revert(cost-map): keep bedrock_mantle claude 5.5 change to cost map rows only
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(mcp): scan and pin upstream tool descriptions
Run every discovered MCP tool's description and input schema through the
pre_mcp_call guardrails before a listing reaches the client, drop the tools
a guardrail blocks, and serve the guardrail's masked text otherwise. Add
POST and DELETE /v1/mcp/server/{server_id}/pin so an admin can freeze a
server's tool names and descriptions; the gateway serves the pinned catalog
and raises a Slack alert with the diff when the upstream drifts.
* chore: sync schema.prisma copies from root
* fix(mcp): pin input schemas, scan before pinning, admin-only pin writes
* fix(mcp): apply overrides and the pin before the discovery scan, dedupe alerts before sending
The guardrail scan now runs on the text the client is about to see: description overrides are applied first, the pinned catalog next, and the scan last, so a masked pinned or override description is served masked and a pinned tool keeps serving its pinned text while the upstream's text is poisoned. The alert signature is recorded before the send and dropped only when that send fails, so a recovery during a slow send is never undone. A tool whose scan payload cannot be built is hidden alone instead of failing the listing. apply_tool_overrides shrinks to apply_display_name_overrides and the MagicMock servers in the MCP tests carry pinned_tools=None.
* fix(mcp): snapshot the pin through the REST module's unpinned catalog helper
* fix(mcp): pin the raw upstream catalog so an override never hides upstream description drift
* refactor(mcp): trim the tool catalog guard docstrings to one line
* test(mcp): cover guarded discovery boundaries and response definitions
* fix(mcp): bound discovery guardrail concurrency per catalog
* fix(mcp): scan tool catalogs in bounded parallel batches
* fix(mcp): hide pinned catalogs from restricted management views
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
* fix(ui): rename All Models tab to Deployed Models and view filter to All Proxy Models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(ui): use All Proxy Models label for the public model name filter
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): rename ALL_MODELS_VIEW constant to ALL_PROXY_MODELS_VIEW
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): stored-config credential canary slots
Add canary slots for credentials the proxy holds in its env, config or
database: virtual key raw value, master key, deployment api_key via
/model/new, named credentials, AWS secret key, Vertex service-account
JSON and its minted token, team model_config credential overrides,
config guardrail api_key, and sink credentials from env.
* test(integration): resolve the config guardrail id and require detail routes to return 200
* test(integration): drop repeated timeout comments
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* test(integration): sweep config-deployment routes with the real model id and use the rig admin for the master-key slot
* test(integration): mark the configure-hook config edits as intended
* test(integration): drop suppression markers that suppress nothing
* Use claude-haiku-4-5 for the Bedrock stored-credential test model
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): credential canary slots for MCP and pass-through credentials
Adds slots F1 (MCP static auth), F2 (per-user MCP OAuth token), F2E (per-user
MCP env var), F3 (x-mcp client auth header), H1 (pass-through credential header),
H2 (vector store api_key) and H2S (search tool api_key) to the credential canary
suite. The OAuth double gains an optional mint hook so a test can choose the
issued access token.
* test(integration): wait for MCP spend rows by call type
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): canary MCP and pass-through slots pass resolved ids
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* test(integration): credential canary suite harness
Adds tests/integration/security with canary generation and search, sweeps over the database, GET routes, client responses, sink doubles and Redis, an owned proxy rig, a sweep sensitivity self-test and the config deployment api_key slot. Registers the security group in run.py, the manifest and the CircleCI integration matrix.
* test(integration): widen canary route sweep and harden the rig
Enumerate lazily registered feature routers, call parameterized routes with placeholder ids, fail on routes that return no response, skip provider pass-through routes, add an explicit admin-only route allowance, let the sink double use a configurable token, inflate gzip members anywhere in a blob, sweep Redis before the route walk, and trap outbound connections from the owned proxy.
* test(integration): descend into any decoded value that can still hold an encoded canary
* test(integration): bound canary decoding by depth and decoded bytes
* test(integration): scope log-table and spend-log reads to the scenario window
* test(integration): sweep spend-log rows in the scenario date window
* test(integration): keep spend-log date window summarized
* test(integration): resolve deployment ids, scope paginated log lists, key allowances by slot
* test(integration): expect 404 from the caller-scoped team membership route
* test(integration): use the rig's own master key and expect 404 from submission lookups
* test(integration): check the overridden rig key without assuming the default key is unknown
* feat(providers): add Prism provider
* fix(providers): complete Prism registration
* feat(providers): expose Prism responses and messages
* feat(providers): add DeepSeek V4.1 Flash to Prism
* test(providers): exercise Prism endpoint requests
* fix(providers): align Prism pricing and limits with the live catalog
deepseek-v4.1-flash bills 0.17/0.63 USD per 1M input/output tokens and takes image input;
deepseek-v4-flash bills 0.17/0.21 and caps output at 384000 tokens, per GET /v1/models
* test(prism): assert cost-map invariants instead of pinning catalog facts
* test(prism): derive the asserted model list from the cost map instead of pinning it
* test(prism): capture requests through respx instead of appending to a list and swapping the client transport
---------
Co-authored-by: rajitkhanna <rajitskhanna@gmail.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: ryan <ryan@berri.ai>
* fix(panw_prisma_airs): apply experimental_use_latest_role_message_only to every request shape
Explicit true/false now applies to chat completions, Anthropic /v1/messages and /v1/responses alike; unset keeps latest-only for Anthropic and full history otherwise. Text indices are mapped back to their source message by value instead of by count, so Responses instructions, function_call_output and reasoning items no longer derail the alignment and silently rescan the whole history
Co-authored-by: scthornton <scthornton@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(panw_prisma_airs): type latest-message helpers against AllMessageValues
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): require forward and reverse text attribution to agree
A Responses function_call_output whose text equals the latest user turn could claim that turn's slot in a forward-only walk and demote the latest-only scan to an earlier message. Walk both directions and fall back to the full role-filter scan when they disagree
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): pick the latest human turn from messages, not from aligned texts
An image-only latest user turn no longer promotes an earlier user turn into the
latest-only scan; it scans nothing on the request side, as the Anthropic path did
before. A latest user/developer message whose text never reached texts (a trailing
Responses reasoning item) falls back to the role-filter scan instead of narrowing.
Types the test helpers, drops the narrating docstrings and adds regressions for both
shapes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): log when latest-only selection leaves nothing to scan
An image-only latest user turn with experimental_use_latest_role_message_only=true intentionally yields zero scanner calls. Emit a debug line naming the call_id so operators can tell this apart from the guardrail not firing.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): keep Responses reasoning items out of latest-turn selection
The Responses translation handler gives reasoning input items the default user
role, so a reasoning item with text content after the latest prompt was picked
as the latest human turn and the real prompt went unscanned under
experimental_use_latest_role_message_only. Map reasoning items back to their
texts positions from the raw input and exclude them; fall back to the
role-filter scan when the raw items do not account for every text
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: scthornton <scthornton@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): unregister logging callbacks removed from the stored config
POST /config/callback/delete saved the config and resynced, but the resync only
ever added callbacks, so a deleted callback kept exporting and kept showing in
/get/config/callbacks as read-only on every worker.
ProxyConfig now tracks which callback list entries each DB config sync
registered and unregisters them once the stored config stops listing them.
Callbacks it did not register (YAML, code) are never touched, and a failed
config load skips the sync instead of treating the config as empty.
* refactor(proxy): keep callback sync comprehensions to one for clause
* fix(proxy): restore code-registered callbacks the DB sync replaced
Registering a custom-logger callback from the DB swaps an existing string
entry for a logger instance. Deleting the DB entry then removed the instance
and left the code-registered callback gone. The sync now records the entries
it displaced and puts them back when it unregisters.
* refactor(rust): extract litellm-host-native as the shared Rust host driver
Move service and hook dispatch out of host-http into a Driver that owns the
machine and Rust handlers, returning at completion or a stream boundary and
holding the demand reply until the consumer advances. Move the in-process
runner onto the same driver. host-http now layers encoding, SSE, body polling
and lifecycle observation over it. host-python keeps driving litellm-host
directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(rust): interrupt the machine when the in-process stream consumer fails
Restores the pre-refactor interruption path for StreamConsumer errors via
Driver::fail and ports the generic run lifecycle tests into host-native.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(rust): separate the machine contract from coroutine execution
* auth update
* refactor(rust): use standard flow control for host requests
* style(rust): keep host driver imports formatted
* chores
* mostly relocation
* refactor(rust): separate interceptors from queued observers
* refactor(rust): centralize legacy callback mappings and lifecycle
* docs: define Python host boundaries and migration plan
* refactor: enforce Python host and bridge boundaries
* refactor(rust): separate operations from callback composition
* refactor(rust): compose SDK policy through call hooks
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
test_ssl_verify_unit.py inserted tests/unit at the front of sys.path, so any later import of litellm_proxy_extras resolved to the tests/unit/litellm_proxy_extras test package. Whenever the CircleCI shard split collected that file before test_litellm_proxy_extras_logging.py, collection failed with ModuleNotFoundError. test_gemini_session_leak.py had the same insert for its own directory
* feat(anthropic): add Claude Sonnet 5.5
Adds the anthropic cost map entry for claude-sonnet-5-5 mirroring
claude-sonnet-5 pricing and capabilities, with prompt_cache_min_tokens
at 512, thinking_always_on (thinking cannot be disabled on this model),
and supports_forced_tool_use false (tool_choice required/named returns
400 upstream). Omits thinking cache preservation, same as Opus 5.5, and
registers the model in the setup wizard provider list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(model_prices): correct Claude Sonnet 5.5 capabilities and provider keys
Sets prompt_cache_min_tokens 512, thinking_always_on, and
supports_forced_tool_use false on every anthropic, bedrock, vertex_ai,
and azure_ai Sonnet 5.5 key, dropping the thinking cache preservation
flag cloned from Sonnet 5. Removes unpublished deprecation dates on
azure_ai and vertex_ai, renames the OpenRouter key to the live
anthropic/claude-sonnet-5.5 id and drops its batch variant, and removes
the aihubmix, deepinfra, and databricks keys for vendors that do not
list the model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(anthropic): drop vendor-absence assertions for Sonnet 5.5 keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(guardrails): fail open by default when Agent 365 cannot evaluate and count it in Prometheus
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(tests): ruff format the Prometheus fail-open registry test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(guardrails): add Agent 365 authority host override, fail-open integration test and per-guardrail YAML default
Add `authority_host` to the Agent 365 config (also read from AGENT365_AUTHORITY_HOST, then AZURE_AUTHORITY_HOST) so sovereign clouds and the integration test can point the OBO exchange at a different Entra host.
Add tests/integration/mcp/test_mcp_agent_365_guardrail.py, a real proxy test with Postgres, Redis, a scripted MCP upstream and local Entra and Agent 365 doubles covering the default fail-open, explicit fail-closed and fail-open, Defender Skipped, policy denial, persisted status and Prometheus counter.
Use PrometheusLogger.get_instance for the fail-open metric lookup instead of a hand-rolled callback scan. Clarify the config description: gateway credential failures fail open, caller token failures block.
Extract the dashboard YAML preview into teamGuardrailConfigYaml.ts so the effective per-guardrail default is unit tested and the "default" hint only shows when nothing was set explicitly.
Regenerate the lazy OpenAPI snapshot and schema.d.ts for the new field.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): default a scheme-less Agent 365 authority host to https and treat a null fallback as unset in the YAML preview
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit cells for the Agent 365 fail-open default across entry points, Entra faults, throttling and two workers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): prove both Agent 365 workers serve and that a killed worker is replaced
Each fresh connection reports its worker pid from /debug/memory/summary and its MCP catalog on the same
connection, so the two-worker readiness wait covers both workers by identity. The kill test now kills a
pid the proxy reported as a worker and waits for a replacement pid, instead of the first psutil child
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): drop the prometheus fail-open counter from the agent 365 guardrail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): keep agent 365 fail closed by default and make fail_open an explicit opt-in
Restores the shared unreachable_fallback default and the sibling guardrail initializers, drops the Admin UI YAML preview that only existed for the per-guardrail default, and reworks the unit and integration tests so the default blocks with HTTP 503 while unreachable_fallback: fail_open lets availability failures through as Unscanned
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): append authority_host after the existing Agent365Guardrail parameters
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(guardrails): agent 365 fails open by default and hides the production overrides from the UI form
Agent 365 sits in the runtime path of every MCP tool call, so an Entra or
Agent 365 outage now lets the call through unscanned (logged at error level,
recorded as Unscanned with guardrail_failed_to_respond) instead of blocking it.
unreachable_fallback: fail_closed stays as the opt-in strict mode. Policy
blocks, throttling, 4xx rejections and a rejected caller token still block
The shared unreachable_fallback field becomes nullable so each guardrail owns
its default; every sibling still resolves None to fail_closed and typesafe
keeps failing open
api_base, resource_app_id and agent_id have production defaults and leave the
dashboard form (ui_hidden); they stay available in config.yaml and env. The
authority_host override and its env keys are gone, the OBO exchange always
uses login.microsoftonline.com. The integration suite keeps only the cells
that need no Entra double, the evaluation paths live in unit tests with an
injected handler
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate openapi snapshot and schema.d.ts for the nullable unreachable_fallback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): fix agent 365 to the production endpoint and keep fail_closed as the default
Remove api_base, resource_app_id and agent_id from the Agent 365 config model, their AGENT365_* env fallbacks and the _is_ui_hidden helper: the evaluation URL and the Agent Tools app id are fixed production constants and the agent identity is always the caller's key alias. Revert the fail_open default; unreachable_fallback: fail_open stays an explicit opt-in. Restore the shared unreachable_fallback field, the sibling guardrail initializers and typesafe to main. Move the Entra dependent cells from the subprocess integration suite to unit tests with an injected HTTP handler.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): warn when agent 365 yaml still carries the removed override keys
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): inject the http handler into the agent 365 initializer instead of assigning it after construction
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Run every route in the team-admin matrix again on a team that belongs to an
organization, as an admin of that organization and as an admin of another
organization. 90 new cases pin what org admins get today, so collapsing the
team-admin helpers into one gate can prove parity for org admins too
Scenario gains organization() and org_member() helpers that clean up the
organization, its budget row and the member users
* security(proxy): keep team callback credentials out of the stored request body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: allow the security conventional commit type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
completion() imported vertexai only to check that the package exists. Partner
models are reached with an authenticated httpx client and never use that SDK,
the same reasoning count_tokens in this file already follows (#28084). The
import loads all of google-cloud-aiplatform on the first request of every
process and made a google-auth-only install fail with a 400
* fix(responses): emit the reasoning item on streaming /v1/responses for signature-only thinking
Anthropic models return thinking blocks with empty text and the reasoning carried in the
signature: Claude Fable 5.1 and Claude Opus 5.5 by default, and Bedrock adaptive thinking
with or without an effort. On streaming /v1/responses the chat->Responses bridge opened a
reasoning output item only on reasoning_content text
(LiteLLMCompletionStreamingIterator._ensure_output_item_for_chunk), and
ChunkProcessor.get_combined_thinking_content kept an assembled thinking block only when it
had thinking text. Such a response emitted no reasoning item mid-stream and none in
response.completed, so a streaming Responses client could not replay the reasoning even
though the reasoning tokens were billed. Non-streaming /v1/responses was unaffected.
Open the reasoning item when the delta carries a signed or redacted thinking block, and
keep a signed block through stream assembly even when its thinking text is empty.
Unsigned text-only fragments are still dropped. The reasoning-text path is unchanged.
(cherry picked from commit bc9b6f8a5c)
* test(vertex_ai): move orphaned gemma streaming tests into the llm-vertex-ai shard
PR #43147 left a copy of the Gemma streaming tests under
tests/test_litellm/llms, a tree no CI shard claims, which broke
assert-ci-coverage and assert-shard-coverage on main. Fold the two
streaming tests into the existing tests/unit/llms/vertex_ai file so the
llm-vertex-ai shard runs them
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Chloe Lu <chloe.lxd@gmail.com>
Co-authored-by: Krrish Dholakia <krrishdholakia@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(vertex_ai): reproduce traced Gemma Responses stream failure
* fix(vertex_ai): wrap Gemma fake streams for Responses tracing
* test(vertex_ai): cover Gemma traced streams and usage options
* test(vertex_ai): inject gemma test deps and assert hidden usage accounting
Replace class-level patches in the Vertex AI shard test with the
provider's documented dependency-injection seams (httpx.MockTransport
client + credential cache), and pin the default/omit-usage trace
behavior: LiteLLM still accounts all tokens; ddtrace's metric is
absent by design, asserted rather than silent.
Mutation-checked: commenting out CustomStreamWrapper chunk accumulation
turns the new assertions red; restoring them turns green.
* test(vertex_ai): drop explanatory comment from usage-option assertions
* fix(gemini): forward seed to the Gemini API instead of rejecting it
The gemini/ provider left seed out of its supported params, so requests with seed
failed with UnsupportedParamsError, or lost the seed silently when drop_params was on.
The Gemini API accepts generationConfig.seed and the inherited mapping already
translates it, so adding it to the allowlist is enough
* test(gemini): assert the forwarded seed without mutating shared state
* fix(tools): salvage concatenated JSON tool call arguments
* fix(tools): harden concatenated tool-call salvage for review findings
Skip non-dict JSON during split so salvage cannot emit empty tool calls.
Collapse srvtoolu_ expansions to the first object so server results stay paired.
Allocate __concat_n ids that cannot collide with sibling tool call ids.
Propagate cache_control onto every expanded Anthropic tool_use block.
Rename the XML invoke loop variable so the key-leak gate no longer flags {args}
* test(tools): cover concat id bump and srvtoolu array keep
Only collapse srvtoolu_ when concatenated salvage expanded; a valid JSON
array argument stays one server tool input
* revert(anthropic): drop concat expansion from pass-through adapter
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* revert(tools): keep concat salvage out of request-side tool converters
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* fix(tools): expand strictly salvaged concatenated tool arguments in normalized tool calls
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* fix(tools): retain at most the salvage cap while validating concatenated arguments
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* test(tools): assert concat sibling ids unique after sanitization
A sibling id that only collides after colon-to-underscore sanitization must force the next concat suffix
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
* refactor(tools): drop unused strict mode from split_concatenated_json_objects
Strict mode had no production caller. Rejection cases now sit on salvage, and split matches upstream main
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
---------
Co-authored-by: Techboy bebop <kumarpriyanshu09@users.noreply.github.com>
The leftover warning used a class defined in a test-directory module. The xdist controller cannot import it, so an uncaught leftover warning crashed the whole e2e run. Same change as #43391 on rc/1.103.0
_is_unsignable_thinking_block() only checked block["signature"], so a
thinking block with a valid-looking signature but empty (or
whitespace-only) thinking text sailed through _drop_unsignable_thinking_blocks
and into anthropic_messages_pt(). Anthropic rejects that with:
400 messages.N.content.M.thinking: each thinking block must contain thinking
This is reachable whenever a thinking_blocks history item gets replayed
through this Anthropic-shaped request path (e.g. a non-Anthropic reasoning
turn with no summary text), the same class of bug PR #36033 fixed on the
Responses adapter's own separate code path.
Now the signature check runs first (unsigned blocks are still dropped, same
as before), then an additional check drops the block if `thinking` is
missing, not a string, or strips to empty. redacted_thinking blocks are
untouched since they don't have type == "thinking".
* fix(vertex_ai): consider tools when validating context caching min tokens
Pass tools to is_prompt_caching_valid_prompt in both sync and async
check_and_create_cache before popping them into the cachedContents
request body. This allows agent-shaped requests with heavy tool schemas
and small message histories to reach the minimum token threshold and
benefit from prompt caching.
Fixes#42804
* test(vertex_ai): avoid doubles on internal code and assert tools in cache payload
* feat(otel): add SigNoz preset for OpenTelemetry v2
Adds the signoz callback (OTLP/HTTP exporter, GenAI vocabulary, key and team level dynamic ingestion endpoint and key) as an OpenTelemetry v2 preset, with the preset factory accepting the allow_missing_credentials kwarg the V2 registry always passes so construction no longer falls back silently to legacy OpenTelemetry. Ships the deterministic tests/integration/observability/test_signoz_delivery.py audit suite
Absorbs the work from https://github.com/BerriAI/litellm/pull/38206
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(otel): drop explanatory comments from the SigNoz preset
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(types): keep signoz dynamic param lines within ruff format width
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(signoz): assert the missing-endpoint boot path directly instead of in an except block
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(ui): regenerate schema.d.ts for the signoz health service
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): allowlist SigNoz key/team endpoints and route keyless collectors without the operator key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel): terminate the SigNoz shutdown cell before the flush and drop test docstrings
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(otel): keep the shared tenant routing untouched and require an ingestion key for SigNoz key/team endpoints
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(otel): warn about a keyless SigNoz team endpoint from the header resolver so the shared cache actually reaches it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Nagesh Bansal <nageshbansal59@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(e2e): read management routes back from the control plane replicas
The suite's management read-backs (/key/info, /team/info and friends)
polled the same replica list as the data plane. On a componentized stack
whose LITELLM_PROXY_REPLICA_URLS names the gateway pods directly, that
list answers those routes 404, since a gateway pod trims the management
routes at startup. A new LITELLM_CONTROL_PLANE_REPLICA_URLS names the
addresses a management read-back polls instead: an exported list wins,
and when it is unset the old rule stands, the data-plane replicas while
the control plane shares the suite's base URL and the control-plane base
alone once it is split.
build_proxy_client takes the list as control_replica_urls and
read_back_everywhere picks its replicas per path, the way the rest of the
client already does.
* fix(e2e): derive the control replicas of a client built for another proxy
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(responses): stream guardrail pre-call block as SSE with a typed output item
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): import blocked usage helper from the guardrail utils module
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): drop narrating docstrings and poll without rebinding
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover pre-call guardrail block on /v1/responses stream and json
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit cells for responses guardrail block contract
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): observe upstream on the recorded chat route for responses denial cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): tidy responses denial audit cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): wait for worker count to recover after SIGKILL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): require a replacement worker after SIGKILL
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(responses): type the blocked response test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>