Commit graph

54023 commits

Author SHA1 Message Date
devin-ai-integration[bot]
488594e03f
refactor(proxy): inject one UTC clock read per operation into gateway tracking, PTU rollup and Mavvrik export (#45520)
* fix(proxy): read the UTC clock once per operation in gateway tracking, PTU rollup and Mavvrik export

* refactor(proxy): default the injected clocks to get_utc_datetime instead of three private copies

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-09 00:15:38 -07:00
yuneng-jiang
a7ff709f3d
fix(ui): let the model picker remove selections that are no longer available (#43944)
* fix(ui): let ModelSelect deselect selections that are no longer offered

A selected model that is no longer served (removed from config, deleted,
or outside the org ceiling) never appeared in the dropdown, so once it sat
past the fifth chip nothing could remove it. List such selections in an
Unavailable group ahead of the offered options so they can be found and
unchecked.

* fix(ui): skip the Unavailable group when no live models were loaded

A failed or empty model list made every selection look unavailable. Only
flag selections as unavailable when there is a live model list to compare
against, and cover ordering with several unavailable selections.

* fix(ui): only suppress the Unavailable group when the model list failed to load

Keying the guard on the context-filtered list hid the group when an
organization ceiling excluded every selection, which brought the dead end
back for team forms. Key it on the proxy model list having loaded.

* fix(ui): keep other selections when removing one beside a special option

Removing an unavailable model while a special option stayed selected
collapsed the whole selection to the special option, dropping the other
saved values. Collapse only when a special option is newly picked. Also
skip the Unavailable group while an org team's model ceiling is unknown,
since the offered list is empty for that reason alone.
2026-10-09 00:00:26 -07:00
devin-ai-integration[bot]
0d17f954c0
feat(credentials): add display_name and make credential_name immutable on PATCH (#43148)
* feat(credentials): add credential_alias and make credential_name immutable on PATCH

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): query credential rows by role to stay under the no-node-access budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): expect credential_alias in load_credential_list dump

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): drive the credential SearchSelect by placeholder and option roles

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): pass the decrypted CredentialItem to update_db_credential during master key rotation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): trim credential_alias in CredentialModal so whitespace-only input clears the alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(credentials): replace credential_alias with display_name and reject edits to config credentials

credential_name stays the immutable reference key. display_name is a nullable, trimmed, max 255
character label set on POST or PATCH (omit keeps it, null clears it, blank is a 400). Reads return
display_name plus an in-memory source tag (db or config), and PATCH or DELETE on a config-defined
credential answers 400 instead of a misleading 404

* feat(ui): show credential display names everywhere and lock config credentials

The credentials table shows the display name with the credential name beneath it, badges config
credentials and disables their edit and delete actions. The edit modal adds an editable display
name next to the read-only credential name and only sends it when it changed. Every credential
picker and reference (add model, model info, models table, vector store form and info) shows the
label while still submitting credential_name

* feat(cli): set display names on credentials and show them in the list

* refactor(credentials): keep the new credential types inside the lint gate budgets

* refactor(cli): print credential command JSON through one helper

* fix(credentials): treat an empty credential_name on PATCH as omitted and pin the 404 for vanished rows

A blank credential_name never renamed anything, so PATCH accepts it again instead of answering 400. New tests pin a trimmed display_name on PATCH, the repository carrying display_name, and a 404 (not the config-owned 400) for a DB credential another worker already deleted. The hydration helper no longer copies display_name, since none of its callers read it, and the CLI update passes display_name straight through because the exactly-one check already makes it None when clearing.

* fix(ui): keep a model's credential when its picker text is emptied and skip the admin-only list for other roles

Clearing the search text in the model edit form's credential picker used to submit null, silently detaching the model's credential on save. None is now the only way to clear it. The models table fetches /credentials only for proxy admins, since everyone else got a 403 on each page load and falls back to the raw name anyway. Also fixes a type error in the re-use credential dialog and adds tests for the display name surviving a provider switch, the 255 character limit, the request payload, and the vector store None choice.

* fix(client): percent-encode the credential name when updating its display name

A name with ? or # was cut short in the URL, so the PATCH landed on a different credential or 404'd.

* fix(credentials): answer 405 with Allow: GET for PATCH and DELETE on config credentials

A config-defined credential exists (GET returns it) but is read-only through the API, which is what 405 Method Not Allowed means. 400 described a malformed request, and 404 would claim the credential does not exist.

* fix(proxy): ignore a display_name set on config credential_list entries so non-string values cannot fail boot

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 23:49:56 -07:00
devin-ai-integration[bot]
5c6ea040b0
test: update stale MCP call_tool and usage card assertions, add timeout headroom to request-log index boot test (#45523)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 23:44:46 -07:00
devin-ai-integration[bot]
145d382089
ci: run only integration jobs on PR CircleCI pipelines and move router and guardrails suites to GHA (#45509)
Restrict 26 build_and_test jobs to main with job-level branch filters, delete litellm_router_unit_testing in favor of a router-unit-tests GHA shard, and narrow guardrails_testing to the license-dependent test while the rest of tests/guardrails_tests runs in a guardrails-tests GHA shard

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 22:58:55 -07:00
devin-ai-integration[bot]
d34ad35281
fix(anthropic): count a leading system run through count_tokens' system parameter (#45463)
* fix(anthropic): count a leading system run through count_tokens' system parameter

Anthropic's count_tokens rejects role "system" at the head of messages, so
/v1/responses/input_tokens with instructions, and /v1/messages/count_tokens
with a system-role message, fell back to the local tokenizer. The shared
Anthropic count_tokens transformation now lifts the leading run of system
messages into the top-level system parameter, the way the chat path sends
it, after any system the caller set. Anthropic direct, Azure AI Anthropic,
and Bedrock Mantle share that transformation.

* test(integration): cover the count_tokens leading-system lift across Anthropic, Azure AI and Bedrock Mantle

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 22:34:25 -07:00
yuneng-jiang
72ab736863
test(routing): wait on deployment registration instead of boot model-info traffic (#45267) 2026-10-08 21:31:53 -07:00
devin-ai-integration[bot]
4b975e6f49
ci: render lint, unit and smoke checks as <tier> / <job> with one collector per tier (#45480)
* ci: restructure lint, unit and smoke workflows

* ci: preserve source formatting in tier workflows

* ci: retire per-shard coverage flags and tighten the tier workflow guards

* test(ci): freeze the workflow startup safety models

* ci: keep the current required check names running until the ruleset moves to the tier collectors

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-09 04:28:41 +00:00
devin-ai-integration[bot]
50ab9a3211
fix(token_counter): price a base64 PDF document per page instead of as one image (#45301)
* fix(token_counter): price a base64 PDF document per page instead of as one image

A `document` or `file` block carrying inline PDF bytes was priced like a single
image (85 tokens), so when a provider's count-tokens endpoint rejected the
model (Bedrock Opus) the local fallback answered 116 for a 12-page PDF the
provider then billed at 35941 input tokens. The counter now reads the PDF with
pypdf and prices each page as its extracted text plus the image Anthropic
renders it to (1568 px long edge, 1.15 MP, 750 pixels per token), falling back
to the old image pricing when pypdf is missing or the bytes are not a readable PDF.

* refactor(token_counter): count PDF pages as they are read

Sum each page's text and rendered-image tokens straight from the pypdf
reader instead of materializing a page list first, keep the fallback
to image pricing atomic when a page cannot be read, and annotate the
new tests' locals as Final

* test(integration): audit cells for page-priced PDF documents in count_tokens, pre-call checks and spend

* test(integration): release the held peer when the concurrent budget wait times out

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 21:20:46 -07:00
berriai-litellm-provider-info-sync[bot]
4df5006f29
fix(bedrock): add the claude-sonnet-4-5 EOL date from the Bedrock model card (#45496)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 20:53:52 -07:00
joshua-berri
04f0ea8c33
feat(mcp): translate input requests and bind continuations (#45464)
Some checks failed
Unit Tests / misc (push) Blocked by required conditions
Unit Tests / misc-dirs (push) Blocked by required conditions
Unit Tests / proxy-auth (push) Blocked by required conditions
Unit Tests / proxy-hooks-client (push) Blocked by required conditions
Unit Tests / proxy-infra (push) Blocked by required conditions
Unit Tests / proxy-infra-root (push) Blocked by required conditions
Unit Tests / proxy-server (push) Blocked by required conditions
Unit Tests / unit (push) Blocked by required conditions
Unit Tests / responses-caching-types (push) Blocked by required conditions
Unit Tests / Lens Python 3.10 (push) Waiting to run
Unit Tests / assert-shard-coverage (push) Waiting to run
Unit Tests / auth-checks (push) Blocked by required conditions
Unit Tests / budgets (push) Blocked by required conditions
Unit Tests / custom-logging (push) Blocked by required conditions
Unit Tests / db-and-spend (push) Blocked by required conditions
Unit Tests / endpoints-and-responses (push) Blocked by required conditions
Unit Tests / guardrails-hooks (push) Blocked by required conditions
Unit Tests / jwt-and-keys (push) Blocked by required conditions
Unit Tests / key-generation (push) Blocked by required conditions
Unit Tests / logging-misc (push) Blocked by required conditions
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Code Quality Checks / code-quality (push) Has been cancelled
Unit Tests: Documentation Validation / documentation (push) Has been cancelled
Code Quality Checks / python-310-import-smoke (push) Has been cancelled
CI Coverage / assert-ci-coverage (push) Has been cancelled
Helm unit test / unit-test (push) Has been cancelled
Lens Worker Image / lens-worker-image (amd64, ubuntu-latest) (push) Has been cancelled
Lens Worker Image / lens-worker-image (arm64, ubuntu-24.04-arm) (push) Has been cancelled
UI Unit Tests / ui-unit-tests (push) Has been cancelled
Lens Worker Image / Publish Lens development index (push) Has been cancelled
* feat(mcp): translate input requests and bind resumable continuations

* fix(mcp): keep continuations bound to their original upstream

* fix(mcp): sync advertised protocol API schema

* test(mcp): run interaction regressions in GitHub CI

* test(mcp): set source paths for interaction proxy processes

* test(mcp): verify salt guidance across interaction carriers

* fix(mcp): preserve interaction authentication and elicitation policy

* fix(mcp): bound preflight bodies and preserve challenge routes

* test(mcp): type interaction regression fixtures

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-08 20:10:10 -07:00
devin-ai-integration[bot]
0ae62c5d02
feat(ui): add the evaluation mode to the Add Model form (#45481)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 19:53:27 -07:00
berriai-litellm-provider-info-sync[bot]
ea43e27485
fix(bedrock): take gpt-6.1-sol context window from the Bedrock model card (#45488)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 19:50:47 -07:00
devin-ai-integration[bot]
deaf88e1a4
fix(ci): import seed_tracing_fixtures from the pytest scripts path in rust trace tests (#45485)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 19:41:22 -07:00
berriai-litellm-provider-info-sync[bot]
2c29e9c360
fix(bedrock): add gpt-6.1-sol ultrafast tier prices from the Bedrock model card (#45482)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 19:28:25 -07:00
devin-ai-integration[bot]
fe1e8d1182
fix(router): match deployment pricing ids against the cost map only within the deployment's provider (#45472)
* fix(router): match deployment pricing ids against the cost map only within the deployment's provider

A deployment model_info.id that equals another provider's catalog key
(e.g. baseten/zai-org/glm-5.2 on an openai-compatible api_base) was merged
into that provider's built-in row by register_model, keeping
litellm_provider=baseten on the entry. _check_provider_match then rejected
the row at request time and the deployment billed $0. register_model now
takes a keyword-only custom_llm_provider used to scope
_get_builtin_model_info_for_registration and
_resolve_builtin_model_cost_entry, and the router passes the deployment's
provider through. The stored entry stays provider-less for
non-colliding ids.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(utils): pass custom_llm_provider straight to the registration lookup

Drop the per-entry lookup-provider expression in register_model per review:
the kwarg alone scopes _get_builtin_model_info_for_registration, and
_resolve_builtin_model_cost_entry keeps its main signature and caller.
_register_custom_pricing_for_request passes the provider through so
router-originated per-request registrations get the same scoped lookup.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(utils): drop docstring and simplify set restore in per-request collision test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(router): type the colliding-id deployment fixture

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 19:20:21 -07:00
joshua-berri
911aff2752
refactor(sdk): separate core AWS and tokenizer dependencies (#44447)
* refactor(sdk): separate core AWS and tokenizer dependencies

* fix(sdk): preserve runtime tokenizer alias compatibility

* test(sdk): compare tokenizer installs through existing entry point

* fix(sdk): preserve bearer headers and optional dependency interfaces

* fix(sdk): retain safe tokenizer fallback diagnostics

* test(sdk): inject tokenizer dependency for fallback diagnostics

* fix(core): preserve AWS dependency errors with retries

* fix(core): centralize optional AWS dependency handling

* refactor(core): reuse optional import helper for Invoke streams

* test(core): compare complete shared dependency requirements

* fix(core): keep tokenizer logging import compatible with main

* fix(core): preserve native token decoding and Responses dependency errors

* test(core): isolate optional dependency import failures

* fix(sdk): retain typed exception message access

* test(sdk): scope HTTP verification environment changes

* fix(sdk): preserve bearer request typing across Bedrock callers

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-08 19:07:57 -07:00
devin-ai-integration[bot]
79c3de46cf
fix(bedrock): remove the bare openai.gpt-6.1-sol cost-map row that AWS cannot invoke (#45479)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 18:52:02 -07:00
devin-ai-integration[bot]
390bea6a53
fix(proxy): save file details for every batch output file so they list and retrieve (#41761)
* fix(proxy): register bedrock batch output files with a file object so they list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): guard output file size when content is missing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover managed batch output file listings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): persist metadata for batch output files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover batch output file listing and retrieval end to end

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve provider metadata for managed batch files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): update managed batch output registration mocks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): inject fake prisma client in managed batch output file tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): retry provider file details and refresh fallback batch output entries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover fallback batch output file details refresh after provider recovers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover model-name managed batch file retrieval

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(managed_files): bound the batch output file lookup so a slow provider cannot stall GET /v1/batches

* fix(managed_files): skip the provider lookup for a fallback entry written moments ago

* fix(proxy): safely refresh managed batch file details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep batch listing provider-free

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): return metadata for empty S3 objects

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): strengthen batch fallback regression assertions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(proxy): describe provider reads and DB writes for batch output files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): sanitize caller-derived ids in managed file logs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(proxy): simplify batch and file endpoint docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate API types from proxy OpenAPI spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(test): resolve managed files lint violations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): save refreshed managed file details through ManagedFileRepository

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): classify retrieved file purpose by its own bucket prefix

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs(proxy): move batch and file endpoint behavior notes to litellm-docs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): resolve strict lint violations in file retrieval

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(bedrock): mark retrieval request metadata handoffs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): resolve managed batch type-check diagnostics

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): avoid duplicate final response bindings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Yucheng He <yucheng@berri.ai>
2026-10-08 18:51:45 -07:00
joshua-berri
f48d837cd2
feat(sdk): build core from an independent packaging manifest (#44340)
* feat(sdk): build core from an independent packaging manifest

* test(sdk): compare rebuilt core payload without generated SBOM identity

* fix(sdk): retain native build configuration in the core sdist

* test(sdk): collect coverage from the core build entry point

* fix(ci): isolate core packaging coverage configuration

* fix(test): identify installed core metadata on Python 3.10

* fix(packaging): preserve Git ignore rules in core staging

* fix(packaging): support source-only core builds

* fix(packaging): reject overlapping SDK distributions

* fix(packaging): align core requirements with current main

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-08 18:40:14 -07:00
berriai-litellm-provider-info-sync[bot]
70fd2c19cf
feat(bedrock): add twelvelabs pegasus 1.5 inference profiles (#45477)
* feat(bedrock): add twelvelabs pegasus 1.5 inference profiles

Price-Sync: litellm-providers

* test(bedrock): whitelist twelvelabs pegasus 1.5 invoke profiles

---------

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
Co-authored-by: Kerry Lu <klu@berri.ai>
2026-10-08 18:32:50 -07:00
devin-ai-integration[bot]
c5c5154c4b
fix(cli): pin pi compat flags so lite pi stops sending store to Anthropic models (#40739)
* fix(cli): pin pi compat flags for gateway models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cli): annotate pi compatibility field

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cli): clarify pi compatibility suppression

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cli): drop the stale mutable-ok suppression on the pi compat block

---------

Co-authored-by: jesus <jesus@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 18:30:27 -07:00
berriai-litellm-provider-info-sync[bot]
591637d475
chore(cost-map): sync openrouter prices from the models API (#45469)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 17:54:22 -07:00
devin-ai-integration[bot]
8a176c0f0a
fix(router): drop the encrypted reasoning a fallback hop's target cannot decrypt (#45393)
* fix(router): drop the encrypted reasoning a fallback hop's target cannot decrypt

An order-based or configured fallback hop replayed the failed provider's
encrypted reasoning items to the next deployment, which answered 400
(Bedrock Mantle: invalid encrypted reasoning; OpenAI:
invalid_encrypted_content), so every multi-turn Responses fallback for
Codex-style clients failed. The hop now drops the encrypted reasoning
its target cannot decrypt and keeps each item's readable summary. With
encrypted_content_affinity on, the pin narrows to the hop's target
order instead of emptying it, so the hop reaches the next order
instead of failing with no deployments available.

* fix(router): keep encrypted reasoning a same-boundary fallback hop can decrypt

Unmarked encrypted reasoning on a hop is attributed to the deployment that just
failed, read from the retry breadcrumb, so a hop to a deployment on the same
api_base and api_key keeps it and a cross-provider hop still drops it. The hop
tests script the upstream at the httpx boundary instead of doubling the handler,
and the router coverage script lists the two hop helpers with their tests

* fix(router): read the hop's failed deployment from its own metadata bucket and carry it into the Responses mid-stream snapshot

* test(router): use a real Router without the origin deployment in the hop strip test

* test(router): cover the hop strip when no failed deployment is known

* fix(proxy): drop the router's fallback hop state keys from the client body

* fix(proxy): keep a request's max_fallbacks cap, drop only the hop state keys

* test(integration): audit cells for the fallback hop encrypted reasoning strip

Forty-five checked-in cells under tests/integration/routing prove the hop strips the previous deployment's encrypted reasoning on /v1/responses, /v1/chat/completions and /v1/messages (httpx, OpenAI and Anthropic SDKs, sync and async, streaming and not), that the affinity pin yields to the hop, that a client-sent fallback_depth, _target_order and attempted_targets never move or strip a request, and that a concurrent burst, an order-1 outage and a killed worker keep every request stripped and logged once. Every call goes through a lane pinned to one worker that already lists the deployments it needs, because the peer worker learns a /model/new row through the config-sync resync up to sixteen seconds later

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 17:30:01 -07:00
yucheng-berri
51fea825cc
feat(guardrails): support logging_only mode for the Akto guardrail (#45461)
* feat(guardrails): support logging_only mode for the Akto guardrail

* test(guardrails): wait for the spend row before reading Akto calls and cover mixed modes

* test(guardrails): cover Akto logging_only on MCP tool calls

* test(guardrails): type the Akto logging_only unit tests and inject the HTTP handler

* test(guardrails): cover failing Akto replies, provider failures and a mixed outage burst under logging_only

* test(guardrails): cover an unreachable Akto with fail_open under logging_only
2026-10-08 17:28:08 -07:00
devin-ai-integration[bot]
1a42c5abde
ci(unit): add unit passed collector job and fold proxy-db shards into test-unit.yml (#45466)
* ci(unit): add unit passed collector job and fold proxy-db shards into test-unit.yml

* test(ci): cover the unit passed gate's success, failure, cancelled and skipped results

* test(ci): run the unit passed gate test from the code quality workflow

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 17:24:24 -07:00
berriai-litellm-provider-info-sync[bot]
acab4761a2
fix(cost-map): update baseten DeepSeek-V4.1-Flash-Fast cache read price and max output (#45468)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-08 17:23:42 -07:00
yujonglee
1bc85fad40
refactor(python-bridge): ship a signature base and read the resolved call (#45450)
* refactor(python-bridge): ship a signature base and read the resolved call

NativeCall carries base (positionals by name plus signature defaults) instead
of the fully bound dict. resolved lays kwargs over base, which is what bound
held, so every pre-hook read and every route host keeps seeing the same values.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(bedrock): build the transcription NativeCall with an empty base

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(dispatch): read the resolved call instead of bound

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(host-python): rename effective to effective_py_args and note the shallow copy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 17:20:29 -07:00
yujonglee
abee1c14e7
refactor(python-bridge): derive NativeCall extraction (#45449)
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 17:20:28 -07:00
devin-ai-integration[bot]
f45407b4c3
test(streaming): expect the unwrapped provider error in bridged /v1/messages error frames (#45465)
#44800 re-raises the provider's own error from the chat adapter, so a bridged /v1/messages stream's error frame now reads litellm.RateLimitError without the litellm.MidStreamFallbackError: prefix that #44989's tests pinned shortly before #44800 merged. The three tests now expect the provider error text once and no sentinel, matching the chat_limited case and every other assertion in the two files

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 17:01:36 -07:00
ishaan-berri
6b8dcec64e
feat(ui): add Cmd+K command palette with key search (#45456)
* feat(ui): add Cmd+K command palette with header search box

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ui): refine command palette selection and results

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* feat(ui): add subtle Cmd+K discovery hint

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ui): address command palette review findings

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 23:34:33 +00:00
devin-ai-integration[bot]
6a2d3edce3
fix(bedrock): split the <reasoning> tag for gpt-oss only on native Chat Completions (#44432)
* fix(bedrock): split the <reasoning> tag for gpt-oss only on native Chat Completions

The native Chat Completions route moved a leading <reasoning>...</reasoning>
block into reasoning_content for every model, while only gpt-oss writes its
reasoning inline. A GPT 5.6 or Grok answer that starts with a literal
<reasoning> tag lost that text, streaming and non-streaming alike. Both paths
now split only when the model id is gpt-oss.

* fix(bedrock): drive the inline reasoning split from a cost-map flag

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 16:32:07 -07:00
devin-ai-integration[bot]
805bb6888b
test: make 38 legacy live base-class translation tests offline (#45351)
* test: make 38 legacy live base-class translation tests offline

* test: tighten offline base-class translation test assertions

* test: cover remaining nova invoke and xai base copies offline

* test: assert outbound bodies, restore dropped providers and fix shared tool dict mutation in base chat translation tests

* test: assert outbound thinking bodies and streamed thinking blocks across anthropic and bedrock converse

* test: restore volcengine/voyage max_retries and openai timestamp_granularities coverage, assert rerank id/meta/cost and scope the interactions transport fixture

* test: call the raw provider helper for thinking stream cases

* test: flatten thinking blocks before collecting signatures

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 23:18:25 +00:00
devin-ai-integration[bot]
4f62bbfd8b
fix(bedrock): remove the bare xai.grok-4.7 cost-map row that AWS cannot invoke (#45460)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 23:02:31 +00:00
devin-ai-integration[bot]
f09408dd37
test: replace 61 live logging and otel tests with offline unit and integration coverage (#45362)
* test: replace 61 live logging and otel tests with offline unit and integration coverage

* test: restore datadog formatting, deliver datadog logs and redis failures over the wire, assert full router hook payloads, tighten stream usage and otel checks

Restores the pre-existing test_datadog.py lines the branch had reflowed. Datadog success, failure and redis-failure replacements now assert the gzip body posted to the intake, with redis failing through a real cache on a closed local port. Router hook sequence recorder checks the legacy field types and asserts concrete payloads, exact streaming and fallback sequences. Stream usage asserts the default include_usage request body and the redacted messages value. Otel asserts response id and token counts.

* test: assert budget envelope figures, guardrail inspection and exact prometheus samples per request

* test: isolate the prometheus latency test on its own deployment so counts do not depend on order

* test: cover timed slack delivery, keep the redis failure test off the network, and make new payload types read-only

* test: drain the redis test's logging and scope otel span checks to the test's own trace

* test: drive the periodic slack flush without a wall-clock interval

* test: scope router hook events to the test, script the db clock, split a nested comprehension

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 15:55:57 -07:00
devin-ai-integration[bot]
8bae65e0b6
test(ui-e2e): update usage page selectors after the #45221 redesign (#45347)
* test(ui-e2e): update usage page selectors after the #45221 redesign

* test(ui-e2e): drop the top keys locator comment

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-08 15:55:22 -07:00
devin-ai-integration[bot]
308c42a4a2
fix(otel): emit OpenInference tool calls and metadata on Arize OTel v2 spans (#43698)
* fix(otel): emit OpenInference tool calls and metadata on Arize OTel v2 spans

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): shed OpenInference output tool calls individually under the span attribute budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): avoid mutation in Arize OTel v2 integration helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): audit Arize OTel v2 OpenInference spans across endpoints, modes and chaos

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): wait for each exported span before the next request in Arize OTel v2 audit tests

* test(otel): make Arize OTel v2 audit absence and outage checks deterministic

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): collect Arize OTel v2 outage spans through an in-order sentinel

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(otel): add skipped BUG cells for pre-existing Arize OTel v2 gaps

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): repair Arize OTel v2 regressions from #43698 (linear fit, metadata slot, repr tool args) (#44488)

* fix(otel): keep Arize OTel v2 regressions in check — O(n) fit, metadata slot, repr tool args

Three regressions from #43698's OpenInference tool-call/metadata emission:

1. Metadata evicted indexed message attributes: the new `metadata` key
   competed for the 128-attribute span budget, and the fit sheds whole
   message groups BEFORE `span.set_attribute`, so the SDK's dropped
   counter stayed 0 — an invisible eviction (live A/B: input-message
   attributes 86 -> 84). Two-part fix: the fit pins `metadata` behind
   every message group (it sheds only once all indexed messages are
   gone), and the span budget no longer charges pre-set attributes the
   mappers overwrite in place — a boundary-opened LLM span already
   carries keys like `gen_ai.request.model`, so the old accounting
   reserved slots the fit could never spend. Live: 86 input-message
   attributes with the metadata attribute riding alongside.

2. Quadratic shed on long prompts: `_message_shed_groups` rescanned the
   full group map once per message (measured on a real acompletion:
   0.032/0.128/0.478/1.910s at 1000/2000/4000/8000 messages vs
   0.007/0.010/0.022/0.029s at base). Index the tool-call groups once by
   (family, message index): the fit is linear again (0.008/0.007/0.014/
   0.031s, same rig).

3. Malformed Python tool arguments lost the whole span: provider
   adapters and `model_construct` responses hand over raw objects, and
   `json.dumps` raises on tuple-keyed dicts (TypeError) and cycles
   (ValueError) before the span is exported. Serialize with a repr
   fallback; both cases now export with a readable arguments attribute.

The attribute budget change affects every boundary-opened LLM-call span
(strictly more attributes retained, never fewer); the mapper changes
only touch the OpenInference vocabulary.

* refactor(otel): build the tool-call group index in one shot

Review follow-up: the dict.setdefault/append seeding in
_tool_call_groups_by_message violated the no-mutation coding convention
(AGENTS.md: build values in one shot with comprehensions or generators
wrapped in tuple()/MappingProxyType()). Rebuild it as a sorted groupby
comprehension; randomized parity harness confirms the shed order is
byte-identical to the seeded version (400 trials).

Also pin the overflow corner Greptile asked about: a pre-set
indexed-message key the fit sheds keeps its earlier value in place, so
the span total can never exceed the SDK limit (new emitter test).

* fix(otel): key the groupby with an explicit tuple to keep basedpyright at budget

The slice-keyed groupby (group[:2]) widened the key to tuple[str | int],
adding one reportGeneralTypeIssues over the codebase ceiling. Key by the
explicit (family, message index) pair instead; shed order unchanged
(300-trial randomized parity harness).

* fix(otel): read pre-set span keys through a helper typed for both runtime shapes

The SDK annotates ReadableSpan.attributes as a Mapping, but an ended span
hands back a tuple of pairs, so the inline isinstance branch narrowed to
Never and pushed reportGeneralTypeIssues one over the codebase ceiling.
Extract _carried_keys with the runtime union declared on the parameter;
behavior unchanged.

* test(otel): drop the wall-clock bound from the long-prompt attribute fit test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(otel): annotate to_openai_dict as Mapping after the rebase onto main

Main's type-discipline budget tightened since the branch point; the plain
dict return annotation is the one violation the rebased branch adds.
Callers only serialize the result, so the read-only view is accurate.

* chore(otel): drop the restating docstring from to_openai_dict

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 15:53:07 -07:00
devin-ai-integration[bot]
5d207d85ee
fix(proxy): accept team-scoped models by their public name on POST /fallback (#45455)
* fix(proxy): accept team-scoped models by their public name on POST /fallback

create_fallback validated the primary and fallback models against the
router's stored model names only, so a model created through POST
/model/new with model_info.team_id was rejected with a 404 unless the
caller used the generated model_name_<team_id>_<uuid> name. The
endpoint now also accepts the team public model names, which request
time fallback matching already keys on, and lists both kinds of names
in the 404's available_models.

* fix(proxy): read fallback rules fresh and clear the config cache after a fallback write

A second POST or DELETE /fallback within the 60 s config cache TTL started from
a cached copy of router_settings and dropped every rule stored since that copy
was taken, by any instance. Both endpoints now evict the cached row before the
read and invalidate it after the upsert

* test(proxy): type the fallback endpoint tests and prove a team request fails over by public name

* fix(proxy): keep fallback writes working when the Redis config cache is down and type the stored settings read

* fix(proxy): keep the non-standard fallback shapes the router accepts on writes and resolve the config cache at call time

* test(proxy): mark the config cache outage test's result as Final

* fix(proxy): replace a same-key fallback rule in place and read a null rule list as empty

* test(proxy): let the stored router settings fixture carry a null rule list

* test(proxy): audit cells for fallback rules by team public name

* test(proxy): delete the fallback rules the audit cells save

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 22:38:38 +00:00
devin-ai-integration[bot]
5133485009
fix(ui): link model access group chips to the access group filter (#45402)
* fix(ui): link model access group chips to the access group filter

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format access group chip link changes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): mock access group hook in affected suites

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep access group chips unlinked until the group lookup resolves

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep cached access group names when a refetch fails

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: nate <nate@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 15:32:33 -07:00
devin-ai-integration[bot]
577d74c1ce
fix(bedrock): count tokens on bedrock-mantle when bedrock-runtime cannot count a Claude model (#45317) 2026-10-08 15:26:49 -07:00
devin-ai-integration[bot]
3822947b0d
fix(rust): add the inline-tools-2026-09-15 beta to AnthropicBeta (#45439)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 15:25:17 -07:00
yuneng-jiang
b1eae8855e
fix(ui): show the Add Model picker once the model catalog loads after a provider is picked (#45426)
* fix(ui): show the Add Model picker once the model catalog loads after a provider is picked

* fix(ui): show the Add Model picker when a name typed before the catalog loaded was cleared
2026-10-08 15:12:25 -07:00
devin-ai-integration[bot]
406514fcaf
feat(spend_logs): configure which metadata fields are stored in LiteLLM_SpendLogs (#44659)
* feat(spend_logs): configure which metadata fields are stored in LiteLLM_SpendLogs

Adds general_settings.spend_logs_metadata_fields with mutually exclusive include and exclude lists. The filter runs on a copy of the row right before it is queued for Postgres, so daily spend rollups, budgets and callbacks still see every key.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_logs): keep excluded auto-router savings keys out of published spend log metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend_logs): cover metadata retention across endpoints, failures, cache hits, batches, runtime updates and auto-router publication

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate schema.d.ts for spend_logs_metadata_fields

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_logs): filter metadata at DB write so guardrail usage sees full rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend_logs): poll guardrail daily metrics instead of reading once

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend_logs): drop timeout comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(spend_logs): read spend_logs_metadata_fields through typed general settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 14:51:44 -07:00
devin-ai-integration[bot]
0e48048bd5
chore(decisions): remove the System One converters and OpenAI spec types left dead by #45214 (#45442)
#45214 routed both decision formats through the shared decisions IR in
litellm/llms/base_llm/decisions/transformation.py and litellm/types/decisions.py.
That left litellm/llms/base_llm/decisions/systemone.py (to_system_one_request,
question_keys, to_decisions_response, the SystemOne* models and their adapter)
and litellm/types/openai_decisions.py with no importer outside their own unit
tests, so both modules and both test files go.

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-08 14:34:06 -07:00
yujonglee
0a98c7be10
refactor(rust): derive strum VariantArray and string conversions (#45434)
* refactor(rust): derive strum VariantArray and string conversions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): use strum conversions directly with explicit spellings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(rust): use rstest values for key management systems

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 21:24:49 +00:00
moe-berri
4ed0267b5a
fix(lens): prevent progress updates from starving analysis budget reservations (#45432)
* fix(lens): serialize analysis budget reservations with progress updates

* test(lens): verify budget reservations block competing progress writes
2026-10-08 21:08:55 +00:00
yujonglee
0721cffab2
refactor(python-bridge): take NativeCall directly and fold routes into per-route folders (#45413)
* refactor(python-bridge): take NativeCall directly and fold routes into per-route folders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(python-bridge): use rstest for updated tests and keep embedding's stub parameter name

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 13:48:28 -07:00
devin-ai-integration[bot]
85a3869dfd
feat(guardrails): per-mode stream_scope with bedrock stream and pass-through fixes (#43801)
* feat(guardrails): run each mode only on streaming, non-streaming, or both

Add stream_scope so a rail can target streaming inference, non-streaming inference, or both per pre, during, and post mode

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): honor stream_scope in pipelines and dashboard types

Pipeline steps skipped the stream_scope filter, direct construction ignored mixed-case maps, and schema.d.ts was stale.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): skip unmatched stream_scope steps instead of allowing

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): keep stored stream_scope keys for modes not on screen

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): format stream_scope helpers and update fork MCP unit tests

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(mcp): keep hang cancellation tests from timing out during setup

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): honor path-defined streaming for stream_scope

Passthrough routes like Gemini streamGenerateContent decide streaming from the URL, so stamp that onto hook data before guardrails run.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(rust): copy AnthropicModelCapabilities instead of cloning

Clippy treats clone-on-Copy as an error, which failed rust-lint on the messages request tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): trust only server stream classification

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(guardrails): keep streaming marker through deepcopy

scan_raw_request snapshots copy each field, so a plain object() marker would lose identity and skip streaming-only rails.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cost-map): drop duplicate perceptron-mk1.5 row

Two main cost-map PRs both added the OpenRouter model, so the merge left a second key that CI rejects.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(rust): expect native transcription 429 as RustUpstreamError

HTTP status errors from native routes map through route_error_to_pyerr, so the wheel SIGINT child was dying on an outdated RuntimeError check and never reached the hang probe.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(tests): follow Google Interactions OpenAPI without hardcoded names

The live spec dropped CreateModelInteractionParams and renamed the item path to {interactionsId}. Misc CI failed because the compliance tests still looked those names up as literals.

* test(guardrails): reproduce stream_scope bedrock and passthrough field gaps

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): cover invalid stored stream_scope reads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): classify bedrock stream actions and keep caller is_streaming_request

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(guardrails): make stream_scope_allows public and drop mutable builds

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): classify pass-through and Bedrock stream scopes

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): harden stream classification and validation

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): add stream scope integration audit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): avoid mutating passthrough custom body

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): run stream scope audit without enterprise license

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): keep stored scope restart cell on one worker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): scope stream scope audit sink assertions to the rail under test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): assert logging_only scope absence behind an ordered barrier rail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): assert one logging_only scan per phase after the barrier rail

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): set request scope on pass-through stream fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): tolerate invalid YAML stream_scope in v1 guardrails list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: merge main into litellm_guardrail_stream_scope_fixes

Update pass-through pre-call test callbacks for main's endpoint_type argument

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): add request paths to pass-through fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): cover websocket pass-through stream scope and LIT-9050 outage spend rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(lint): remove unused type discipline suppressions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(guardrails): script the vertex live upstream in the websocket stream scope test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): keep stream marker json-serializable and restore pass-through helper names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lint): allow required Bedrock action re-export

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): strip the stream marker from pass-through payloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(guardrails): drop stream marker by value in scans and snapshots, plain-tuple stream scope state

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): render tag-scoped guardrail modes read-only in the custom code editor

A guardrail whose litellm_params.mode is the tag-scoped dict {tags, default}
crashed the Custom Code editor on open: normalizeMode wrapped the dict into
the mode array and StreamScopeFields rendered it as a React child (error #31,
whole dashboard unmounted). Treat a non-string non-array mode as no editable
modes, show formatGuardrailMode(mode) in a disabled input (read-only, matching
the guardrail info view), and keep mode/stream_scope out of the update payload
for such guardrails.

* fix(ui): resolve merge fallout in guardrails components

Deduplicate toModeArray import after the merge, and move the read-only
guardrail details block back into GuardrailReadOnlyDetails (now rendering
the shared mode/logging-only rows plus the stream-scope detail) so
guardrail_info.tsx stays under the 800-line lint budget. Guardrails UI
suite: 298 passed.

* fix(ui): drop duplicate toModeArray import reintroduced by merge

* chore(pass-through): document the deliberate in-place marker strip as mutable-ok

The clear/update on _parsed_body is load-bearing: rebinding to a fresh
mapping instead breaks 75 pass-through tests because the marker-free body
must propagate through the caller's request dict so downstream guardrail
scans and snapshots never observe the server streaming marker.

* fix(guardrails): move stream_scope after timeout in CustomGuardrail init

Inserting stream_scope before the existing timeout parameter shifted the
positional slot of timeout, so positional callers constructed with their
timeout bound to stream_scope (ValueError) and timeout silently None.
Restores the base parameter order; keyword callers are unaffected.

* fix(guardrails): typing pass for the lint gates

stream_scope leaves the declared constructor parameters (restoring the
base positional surface; keyword construction unchanged), unknown config
values crossing the new stream-scope code paths get typed locals or
cast-ok boundaries, and the passthrough payload literals are annotated.
All three lint gates pass against current main; the guardrail suites are
unchanged (236+5019 passing; one known anyio-driver failure pre-existing
on base).

* style: sort cast imports for the ruff gate

* chore(pass-through): drop unused BEDROCK_STREAMING_ACTIONS re-export

The streaming check now uses is_bedrock_streaming_endpoint; nothing in
the repo imports the name from this module.

---------

Co-authored-by: Shivi Jain <mobile.350017@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: gabriele <gabriele@berri.ai>
2026-10-08 13:46:53 -07:00
devin-ai-integration[bot]
7ec2a94cd8
fix(ci): drop the publicly known master key prefix from the voyage routing test (#45431)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 13:41:03 -07:00
yuneng-jiang
a7b04a8028
test(decisions): post System One bodies to /v1/systemone in the translation bases (#45327) 2026-10-08 13:23:43 -07:00