Commit graph

53693 commits

Author SHA1 Message Date
dependabot[bot]
801aba88ed
chore(deps): bump langgraph-sdk from 0.4.2 to 0.4.4 (#44713)
Bumps [langgraph-sdk](https://github.com/langchain-ai/langgraph) from 0.4.2 to 0.4.4.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/0.4.2...0.4.4)

---
updated-dependencies:
- dependency-name: langgraph-sdk
  dependency-version: 0.4.4
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-05 17:35:38 -07:00
devin-ai-integration[bot]
d7e7f4071d
test(rust_bridge): expect PartRow start_time and end_time as required LENS_CONTENT fields (#44719)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 00:14:41 +00:00
devin-ai-integration[bot]
169af2f883
ci: split slow unit shards and build the Rust bridge once per run (#44622)
* ci: split slow unit shards and build the Rust bridge once per run

* ci: key the Rust bridge cache on source files only

* ci: keep the unit setup ceiling unchanged with the shared Rust bridge

* ci: fall back to the Cargo cache when the Rust bridge artifact is missing

* ci: keep reruns on enterprise-routing for the prompt caching flake

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 23:43:01 +00:00
moe-berri
c29b42a32b
fix(lens): preserve span timestamps in investigation evidence (#44702)
* fix(lens): preserve span timestamps in investigation evidence

* style(lens): format chronology regression test
2026-10-05 16:42:42 -07:00
nate-berri
371b527d60
fix(build): rebuild the Rust bridge when uv sync sees Rust sources change (#44708)
uv only rebuilds the editable litellm package when its cache keys change, and
the default keys are pyproject.toml, setup.py and setup.cfg. Editing or
switching to a branch with different Rust code left the old
litellm/rust_bridge/_native.abi3.so installed. Key the build on the Rust
toolchain pin, Cargo config, lockfile, workspace manifest and every file
under litellm-rust/crates, since crates embed .sql, .json and .jinja files at
compile time.

Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:42:22 -07:00
moe-berri
c561372a02
fix(lens): show investigation findings for agent traces (#44696)
* fix(lens): show investigation findings for agent traces

Replace child tool-error counts in the trace table with distinct investigation findings. Keep unassessed traces separate from completed clean investigations and use the root status for failure filters and timeline counts

* fix(lens): stabilize findings updates and repair UI checks

* fix(lens): allow viewer findings reads and index trace lookups

* test(lens): cover viewer findings reads with Postgres
2026-10-05 16:27:54 -07:00
devin-ai-integration[bot]
65b0557f80
test(integration): point the scratch upgraded proxy's read replica at the scratch database (#44613)
* test(integration): point the scratch upgraded proxy's read replica at the scratch database

* test(integration): check the scratch upgraded proxy's reader role is connected to the scratch database

* test(integration): assert every configured proxy role holds a scratch connection without a mode branch

* test(integration): skip backend workers without a role in the scratch connection scan

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 23:06:36 +00:00
moe-berri
d7210f95e6
fix(lens): persist final coverage with source diagnostics (#44682) 2026-10-05 15:53:49 -07:00
devin-ai-integration[bot]
68d9b8bbb8
fix(router): retry a /v1/messages stream the provider drops before the first content chunk (#44276)
* fix(router): retry a /v1/messages stream the provider drops before the first content chunk

A /v1/messages stream that the upstream closed before any content reached the
client answered an error event after a single attempt, so the router's
num_retries never applied to that drop. The pre-content failure is now retried
within the model group before the fallback chain runs, with the budget resolved
the way a failure raised before the stream opened resolves it: a retry policy
that names the error class, then the request's num_retries, then the
deployment's, then the router's. A drop after content reached the client keeps
surfacing the provider's error after one attempt.

Fixes #44238

* fix(router): hand a retry's non-retriable error to the fallback chain and type the retry helpers

A retry that failed before its stream opened with an error no retry covers raised straight to the
client, skipping a fallback the first attempt would have used. assert_never now comes from
typing_extensions so the router imports on Python 3.10, and the retry helpers read their kwargs
through typed narrowing instead of Mapping[str, Any]

* fix(router): cast the untyped router fallback defaults the stream retry gate reads

The retry gate passed the router's fallback attributes, declared without element types, to the
typed request override helper, which basedpyright counted as new unknown-argument errors

* fix(router): consult context_window_fallbacks when a retried /v1/messages stream overflows

A retry attempt raising ContextWindowExceededError reached the fallback chain inside its
mid-stream envelope, so only the regular fallbacks list matched. The fallback attempt now
unwraps it the way it unwraps a content policy error. The new router helpers are covered for
the router code coverage check with two direct-call tests and named covering tests

* fix(router): retry a 408 raised by a /v1/messages retry and honor deployment num_retries before the stream opens

* fix(router): attribute a retried /v1/messages stream to the deployment that served it and bound the retry-policy hold

* fix(router): retry /v1/messages error frames under their retry-policy class and keep the first drop's committed budget

An `event: error` frame that arrives before the first content delta now raises the exception class the pre-stream mapping gives an HTTP answer with the same status (429 RateLimitError, 500 and 529 InternalServerError, 503 ServiceUnavailableError, 504 Timeout), so a retry policy's per-class budget governs it the way it governs the error before the stream opened. The status the client sees is unchanged

A retry that lands on a sibling deployment keeps the budget the first drop committed to, read back from the request's attempted_retries and max_retries, instead of recomputing it from the new deployment's num_retries, matching the pre-stream retry loop

* refactor(anthropic): keep the error-frame exception mapping under llms and type the retry test helper

The status-to-exception mapping an `event: error` frame gets before the retry policy is consulted now lives next to the Anthropic error status map in llms/anthropic/common_utils.py, with its own unit test, and the two-deployment retry test helper takes explicit typed parameters instead of a bare dict and untyped kwargs

* refactor(anthropic): map an error frame's status with explicit returns on every path

* fix(router): map stream error frames through the pre-stream exception mapping

An overloaded `event: error` frame on a /v1/messages stream now raises the InternalServerError a 529 answer maps to, built by exception_type from the frame's own body, so one retry policy class governs the error before and after the first byte; a failed fallback after such a frame answers 500 like every other litellm path instead of the frame map's 503

A model_group_retry_policy that does not parse (a non-integer budget, an entry that is not a mapping) no longer fails every healthy stream of that group before its first attempt: the stream runs with no policy and the plain num_retries budget, with a warning naming the group

* fix(router): forward an error frame nothing can take over for as the provider sent it

A pre-content error frame whose class the retry policy grants no retry, with no fallback configured, raised an HTTP error only on the first attempt while the same frame after exhausted retries reached the client verbatim. Both now pass through as sent, the way the merge base forwarded every frame.

* test(integration): audit /v1/messages pre-content retry across routes and budgets

Adds the /audit cells for the pre-content stream retry: the native Anthropic route
(drops and error frames before content, HTTP rejections before the stream opens, SDK
sync and async, after-content and non-retriable controls, budget exhaustion, cache
twin, spend row and headers), the chat and responses bridges, the generic routes
(responses, chat, vllm pass-through, Gemini generateContent, fine-tuning jobs list),
owned two-worker proxies for router-level budgets, retry policies and fallbacks, and
two chaos cells (a worker killed mid burst, an outage on every first attempt). Shared
helpers for scripted Anthropic SSE upstreams and OpenAI-compatible wire replies live
in tests/integration/_support

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 22:53:39 +00:00
devin-ai-integration[bot]
7bad8de067
chore: move PR template to PULL_REQUEST_TEMPLATE/general.md, add rust.md (#44693)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:47:39 +00:00
devin-ai-integration[bot]
d9b1d57f07
fix(proxy): let a listed team alias win over a same-named key alias in the customer model check (#44677)
* fix(proxy): let a listed team alias win over a same-named key alias in the customer model check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): check each requested name in a plain loop in can_customer_access_model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): only let a listed team alias skip the customer check when its target is live

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): move the per-name customer alias check into a local function

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:43:16 +00:00
devin-ai-integration[bot]
a1a42768c1
fix(spend-tracking): stop caching failed spend-log metadata lookups as confirmed misses (#43560)
* fix(spend-tracking): stop caching failed spend-log metadata lookups as confirmed misses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(spend-tracking): share the short-lived miss cache write

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover key alias recovery after spend log lookup failures across usage routes

* test(spend): bound outage alias lookups per miss window instead of a fixed count

* fix(spend-tracking): treat any spend-log lookup failure as a short-lived miss

The Prisma client raises a plain AttributeError when the database drops
the connection mid-query, so the PrismaError catch let it through and
the whole usage call answered 500. Any failure now keeps the 30 second
backoff only, and the integration proxy patches its test entitlement at
import so uvicorn's spawned workers inherit it

* test(integration): audit spend-log metadata recovery under timeouts and dropped connections

Cover the daily activity routes, the usage AI chat, the Vantage and
CloudZero dry runs and exports under a locked spend-log table and under
a database connection dropped mid-lookup, on a two-worker proxy, with
the recovery after the outage asserted through the proxy's own miss TTL.

Add a dropped_connection_relay that closes only the connection whose
bytes carry a trigger, so a cell can drop the one connection the
recovery query runs on while the rest of the pool keeps serving. Rewrite
the sweep and JWT cells for the merged main: the export route reads
metadata by SQL join and never calls the recovery, the search routes
answer key rows and find deleted keys by alias, and the daily-spend
owner recovery names the user while the alias stays blank. The sweep
cell now times out a second lookup under the same lock, which pins the
keys blank on the merge base and recovers on this branch.

* test(integration): match a dropped-connection trigger split across two reads

The dropped-connection relay checked each TCP read on its own, so a SQL
marker that straddled two reads never tripped it and the outage cells
would run without the outage they meant to exercise. Carry the tail of
the previous read into the next check, as the held-statement relay
already does, and pin that with a unit test that splits the trigger
across two writes.

* test(integration): scan relay triggers through an in-process helper

The dropped-connection relay now matches its SQL trigger through a TriggerScanner that carries the previous read's tail, and the unit test exercises that scanner directly instead of opening loopback sockets, which tests/unit forbids. The relay's end to end behavior stays covered by the integration cells

---------

Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 22:36:25 +00:00
devin-ai-integration[bot]
b9251dafad
refactor(rust): derive string enum serde through strum and serde_with (#44675)
* docs(rust): document string enum serde conversions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): derive string enum serde through strum and serde_with

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:34:34 -07:00
devin-ai-integration[bot]
8b7b42f3aa
test(integration): basic translation cases for the openai route (#44667)
* test(integration): openai-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): ignore x-stainless headers in translation runner

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): skip openai chat completions basic cases on LIT-9235

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:32:50 -07:00
devin-ai-integration[bot]
6a92a49c32
test(integration): basic translation cases for the gemini route (#44662)
* test(integration): gemini-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): raise gemini basic cases to 1024 output tokens

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:32:26 -07:00
devin-ai-integration[bot]
74ad633e53
refactor(rust): rename litellm-framing crate to litellm-framer (#44683)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:32:00 +00:00
devin-ai-integration[bot]
a6b7f62760
test(integration): assert the Messages API health probe on Mantle Claude (#44612)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 15:28:02 -07:00
devin-ai-integration[bot]
4314f3ce0c
test(integration): basic translation cases for the azure route (#44672)
* test(integration): azure-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): use the three-line form for the azure chat LIT-9235 skip

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:24:42 +00:00
tin-berri
9062fd3931
feat(proxy): embed enterprise LiteAdmin MCP in LiteLLM images (#44610) 2026-10-05 15:24:34 -07:00
devin-ai-integration[bot]
739192ec19
test(integration): pin the Claude-tokenizer recount for streamed claude-sonnet-5 no-usage cases (#44643)
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:17:12 +00:00
berriai-litellm-provider-info-sync[bot]
8c15631836
fix(bedrock): add priority and flex prices for Grok 4.3, 4.6, 4.7 and Kimi K3 (#44670)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-05 15:12:37 -07:00
berriai-litellm-provider-info-sync[bot]
489962cc22
fix(vertex-ai): add deprecation_date to gemini-3.1-flash-lite-image (#44639)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-05 15:10:49 -07:00
moe-berri
b69d744993
feat(lens): analyze trace workspaces with confined Python and compaction (#44640)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* feat(lens): integrate confined workspace analysis with live reviews

* fix(lens): synchronize confined Python process monitoring

* Update review.md

* fix(lens): allow mixed context capacities and correct review assertions

* fix(lens): retrieve evidence on demand and isolate failed reviews

* fix(lens): isolate incomplete evidence reads from peer reviews

* test(lens): await trace status filter option

* test(lens): wait for reclaimed review state to settle

* fix(lens): recover from incomplete cross-session evidence

---------

Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
2026-10-05 22:06:48 +00:00
devin-ai-integration[bot]
7123484c28
test(integration): basic translation cases for the openai_responses route (#44663)
* test(integration): openai_responses-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): skip openai_responses-route chat completions basic cases until LIT-9196 is fixed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:06:10 -07:00
devin-ai-integration[bot]
d23a0636b4
test(integration): azure_ai-route basic translation cases on messages, chat completions and responses (#44658)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:01:55 -07:00
ryan-crabbe-berri
034664001f
fix(ui): send null instead of $0 when a team's member default budget is cleared (#44644)
Clearing Default Budget (USD) under Team Member Settings sent Number("") = 0, which
turned the shared member default into a $0 cap and blocked every member still on
the default. Opening Team Member Settings on a team whose default has no dollar cap
did the same through Number(null).

Team settings numeric fields now go through one shared numberOrNull helper, which
the team admin settings form already used
2026-10-05 15:01:12 -07:00
moe-berri
fe24be2e3d
fix(lens): make tool steps and conversations readable (#44645)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* Update review.md

* fix(lens): make tool steps and conversations readable

* fix(lens): address trace rendering review and test failures

* fix(lens): preserve conversations with incomplete tool calls

* test(lens): retain failed tool styling coverage

* fix(lens): keep tool metadata in accessible result groups

* refactor(lens): build stable agent labels without mutation

* perf(lens): group and sort agent labels without repeated scans

* test(lens): await trace status filter option

---------

Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
2026-10-05 22:00:57 +00:00
devin-ai-integration[bot]
7dc5b73f93
feat(ui): declare shared search operators and flush queries on blur (#44665)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:00:34 -07:00
devin-ai-integration[bot]
ee92183319
fix(sso): let CLI and Claude Code gateway sign-in through on DISABLE_ADMIN_UI nodes (#44620)
* fix(sso): let CLI and Claude Code gateway sign-in through on DISABLE_ADMIN_UI nodes

* test(proxy): cover CLI SSO sign-in on a UI-disabled node

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 21:59:50 +00:00
devin-ai-integration[bot]
a3e15774ad
feat(proxy): limit which models an end user can call (#43904)
* feat(proxy): add models column to the end user table

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): enforce the end user models allowlist in model access checks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): accept and return models on the customer endpoints

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover the customer models allowlist

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): resolve team aliases before the end user model check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:58:03 +00:00
ishaan-berri
a3a52466a5
feat(lens): restore compact navigation with live investigation review (#44472)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* Update review.md

* test(lens): await trace status filter option

---------

Co-authored-by: moe-berri <moe@berri.ai>
2026-10-05 14:50:43 -07:00
devin-ai-integration[bot]
76ffec9238
test(integration): run the anthropic /v1/responses basic translation cases now that the bridge echoes request params (#44615)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:40:29 +00:00
devin-ai-integration[bot]
d3be5c3ab7
test(integration): bedrock_converse-route basic translation cases on messages, chat completions and responses (#44646)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:39:34 +00:00
devin-ai-integration[bot]
f0415ee033
fix(ui): explain why team member reset spend is unavailable instead of hiding it (#44629)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 14:31:04 -07:00
devin-ai-integration[bot]
cad87a900f
fix(cost): bill per-second transcription models outside chat modes (#44458)
* fix(cost): bill per-second transcription models outside chat modes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): restore request-time billing for per-second transcription models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): bill audio length for per-second transcription models when known

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 13:36:29 -07:00
berriai-litellm-provider-info-sync[bot]
9cdedf81cd
fix(gemini): add priority tier audio input price to three Gemini rows (#44632)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-05 12:29:26 -07:00
michelligabriele
5dd77eb8ca
fix(openai/realtime): drop model from upstream URL for intent=transcription (#43854)
* fix(openai/realtime): drop model from upstream URL for intent=transcription

* fix(openai/realtime): keep model on the upstream URL for non-OpenAI hosts

A custom api_base on the openai provider can be a gateway that routes on
the model query param, so the transcription model drop now applies only
to OpenAI's own hosts. xAI's URL is unchanged again

* test(integration): audit realtime transcription upstream URL on a custom api_base

Twenty-two integration cells cover transcription and conversation sessions on every realtime route, the OpenAI SDK sync and async clients, malformed and duplicated query params, unauthenticated and refused upgrades, an unknown model, idempotent spend logging, a twenty-session burst, an upstream outage mid burst, a worker kill, and a proxy restart, each asserting the exact query pairs the upstream received. The scripted upstream now records every websocket upgrade as an observation

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 19:14:56 +00:00
devin-ai-integration[bot]
45e7be1abc
feat(tracing)!: return only data from SQL queries (#44609)
* feat(tracing): add generated SQL response contract

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): return data-only trace SQL responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): format trace response schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 12:11:59 -07:00
ryan-crabbe-berri
9247826cdd
refactor(proxy): rename management/teams/access.py to authz.py (#44624)
access already means access groups in this codebase; authz names what the
module decides (who may act on a team). Pure rename: every importer now uses
authz, and access.py stays as a re-export because the published
litellm-enterprise 0.1.73 wheel imports is_team_admin from it.
2026-10-05 19:11:20 +00:00
devin-ai-integration[bot]
e2c106101c
fix(harness): drop the unused Any import that fails ruff on main (#44627)
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 19:08:47 +00:00
devin-ai-integration[bot]
7625b8e788
chore(harness): drop the unused Any import left in harness options (#44625)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 12:01:59 -07:00
moyai-devin-berriai[bot]
75b45e39c9
fix(proxy): enforce internal-user model creation prohibition (#44438)
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
2026-10-05 11:32:05 -07:00
devin-ai-integration[bot]
f81a3f5243
fix(responses): report truncated bridged output as incomplete and echo request params (#44460)
* fix Bedrock json mode bug

* fix finish reason when max tokens hit

* fix hardcoded completion for streaming responses

* fix: echo Responses API request params onto bridged response

* refactor(responses): drop explanatory comments and build null finish_reason test without mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the terminal event and Bedrock mock helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep truncated WebSocket turns in history and drop invalid echoed params

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mrinal Chanshetty <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 11:27:12 -07:00
devin-ai-integration[bot]
14913548d1
test(integration): anthropic-route basic translation cases for six Claude models on messages, chat completions and responses (#44607)
* test(integration): add anthropic-route /v1/messages basic cases for haiku-4-5, sonnet-5, opus-4-8 and sonnet-5-5

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): add anthropic-route /v1/chat/completions and /v1/responses basic cases with generated-field matchers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): skip anthropic-route /v1/responses basic cases on LIT-9231 and expect the echoed request params

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): write LiteLLM-generated ids and timestamps as mock.ANY and drop matchers.py

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 18:24:01 +00:00
moe-berri
a99bccacea
fix(ui): restore inline Lens onboarding and responsive layout (#44604)
* fix(ui): improve gateway layouts on mobile

* fix(ui): limit mobile cleanup to navigation and header

* fix(ui): restore inline Lens onboarding and responsive layout

* fix(ui): smooth Lens tab and panel connections

* fix(ui): keep trace time controls within narrow panels
2026-10-05 11:10:10 -07:00
Mateo Wang
d80f8c28ca
refactor(types): replace Any with proven types in 137 files (#44478)
* refactor(types): prove runtime types at harness, search, rag and client boundaries

Replace Any with adapter-validated types in the litellm.agent() harness, the
search provider transformations, RAG ingestion and query, the vector store
pre-call hook and registry, the galileo and opik logging integrations and the
proxy client CLI. Each boundary gets unit tests for well-formed and malformed
payloads.

* chore(typing): prove types at more provider boundaries and restore search transformations

Second pass over a2a, embedding, rerank, image, audio and small provider
modules. Search transformations go back to their previous form because
validating their response bodies would change the proxy status for
malformed upstream bodies from 400 to 500.

* refactor(types): prove types at logging, files, rerank, image, audio and management boundaries

Replace Any with validated or annotation-only types in 37 more files: logging
integrations, token counters, provider files/rerank/image generation/audio
transcription transformations, pass-through logging handlers and management
endpoints. No proxy HTTP status or error type changes.

* refactor(types): prove types at repository, spend, files and router boundaries

Replace Any with repository table accessors, validated mappings and
annotation-only types in 33 more files: Prisma repositories, the enterprise
batch and responses cost checkers, budget reservation, files endpoints,
management endpoints, the policy registry, the adaptive and complexity routers
and the secret managers. No proxy HTTP status or error type changes.

* test(types): run the aiohttp transformation test in-process and cover repository row conversion

The aiohttp chat transformation test no longer starts a server. It feeds the
transformation a response whose json() returns the body under test.

The proxy unit shards now exercise stored model rows whose params are JSON
strings and the object permission create and update paths.
2026-10-05 11:01:45 -07:00
moe-berri
3b2ed83152
fix(ui): make mobile sidebar and top bar responsive (#44603)
* fix(ui): improve gateway layouts on mobile

* fix(ui): limit mobile cleanup to navigation and header

* fix(ui): share cached settings with mobile navigation
2026-10-05 10:56:08 -07:00
devin-ai-integration[bot]
8364f88cbb
ci: run unit selections from GHA test-path and drop the CircleCI unit jobs (#44461)
* ci: run unit selections from GHA test-path and drop the CircleCI unit jobs

* ci: keep existing shard token order and header note

* ci: throwaway, drop tests/unit/repositories from the unit shard to show assert-ci-coverage fails

* ci: revert throwaway assert-ci-coverage check

* ci: stop crediting --ignore paths as invoked in assert_ci_coverage

* ci: install the caching, extra_proxy and proxy-runtime extras in the GHA unit sync

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 10:49:16 -07:00
devin-ai-integration[bot]
9b6a6a0b71
refactor(tracing): generate existing HTTP request models from Rust schemas (#44591)
* test(tracing): pin HTTP request compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): generate existing HTTP request models from Rust schemas

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): bind trace query params to generated request models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 48 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(tracing): read the trace list clock without a thread-pool dispatch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(traces): share the trace page-size bounds between schema and reader

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): type trace request queries against the generated OpenAPI schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: encode the trace contract boundary in AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format trace request aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 17:37:52 +00:00
Ninad Phalak
d946706744
feat(guardrails): add llm shield pii redaction and rehydration guardrail (#42645)
* feat(guardrails): add llm shield pii redaction and rehydration guardrail

LLM Shield is a self-hosted PII gateway. This adds it as a guardrail so a
proxy operator can redact personal data out of outbound requests and have
the original values restored in the model's reply.

The substitution is reversible, which is the difference from a masking
guardrail. Outbound text is replaced with placeholders held in a session
vault inside the operator's own LLM Shield deployment, and the reply is
restored before it reaches the caller, so the end user still sees real
values while the provider never received them.

Streaming responses are restored incrementally. LLM Shield holds back only
the trailing characters that could still turn out to be part of a
placeholder, so tokens are forwarded as they arrive rather than the whole
response being collected first. A placeholder split across two chunks is
never emitted in fragments.

The integration talks to LLM Shield over HTTP and adds no dependency.

Notes for reviewers:

- The guardrail sets use_native_lifecycle_hooks, since redaction and
  restoration need the native pre-call, post-call and streaming hooks
  rather than the unified path.
- Per-request state lives on the request dict, never on the guardrail
  instance, because the proxy registers a single instance process-wide.
  The streaming carry-over is a local of the generator for the same reason.
- Every failure blocks the request. A redaction guardrail that fails open
  would send the exact data it exists to protect to the provider.

* feat(ui): list llm shield in the guardrail garden

Adds the card, preset and logo so operators can pick LLM Shield from the
guardrails page the same way as the other partner guardrails.

* docs(guardrails): add llm shield example config

Shows both modes on one entry. Listing only pre_call redacts the request
and then hands the placeholders back to the end user, so the test asserts
both hooks are enabled.

* feat(ui): use the llm shield brand mark for the guardrail logo

* fix(guardrails): restore llm shield values in anthropic replies

The /v1/messages reply is a plain dict with a content block list and no
choices, so it fell through the restore path and went back to the caller
still carrying placeholders. The request was redacted correctly, which is
what made this easy to miss.

Found by running all three endpoints against a live provider; the mocked
tests all passed because they only built the OpenAI shape. Adds tests for
the message shape and for leaving non-text blocks alone.

* docs(guardrails): correct the llm shield start command

* fix(guardrails): redact every request shape and restore every reply shape

Three gaps, all of which let an enabled guardrail hand data to the provider
or hand placeholders to the caller.

Requests only walked `messages`. The Responses API `input` and tool call
`arguments` went out untouched. Measured against a live provider: a request
sent through `/v1/responses` reached the model with the real address in it
while the guardrail reported as enabled. Request traversal now covers chat
content (string and multimodal), tool call arguments, and `input` as a bare
string or a list of items.

Fixing that exposed the matching gap on the way back: the Responses API reply
carries `output` items rather than `choices`, so it returned to the caller
still holding placeholders. It now gets its own walk, handling text blocks as
dicts or objects.

The dashboard preset seeded only pre_call, so a guardrail created from the UI
would redact the request and return the placeholders to the user. Presets can
now seed both modes; the form already normalised either shape.

Adds tests for each request shape, for both Responses API reply forms, and
replaces a test that had asserted the `input` bypass as correct behaviour.

* fix(guardrails): narrow the stream delta before writing to it

basedpyright could not prove the delta was non-None on the write path, and
reportOptionalMemberAccess has a zero budget. The guard is also clearer than
relying on the text check to imply it.

* fix(guardrails): mint the vault id instead of trusting the caller's

The vault id was taken from caller-supplied session metadata, and every
caller shares one LLM Shield key. Someone who knew or guessed another
caller's session id could send a placeholder, have the model echo it back,
and get that caller's plaintext restored into their own reply.

Vault ids are now minted per request behind a per-process prefix, so a
caller cannot name a vault this process uses. Redaction mints, restoration
reads back, and a reply whose id does not match is left holding its
placeholders rather than resolved against some other vault.

Also covers two more request fields that were reaching the provider intact:
the Responses API `instructions`, and the legacy `function_call.arguments`
alongside `tool_calls`.

The collectors move to module level, which drops the traversal back under
the complexity limit and lets the code carry its own explanation instead of
the comments that were restating it.

* fix(guardrails): drop Final from a loop-assigned local

basedpyright rejects a Final assigned inside a loop, and
reportGeneralTypeIssues sits one over its budget ceiling.

* fix(guardrails): redact completion prompts and responses tool items

Two more provider-bound request shapes were reaching the model intact while
the guardrail reported as enabled.

/v1/completions carries its text in a top-level `prompt`, which the
traversal never looked at. It is handled as a string and as the array form,
where each entry is rewritten in place.

Responses input items hold tool data outside `content`: a function_call item
in `arguments`, a function_call_output item in `output`. Both are now
collected alongside the item's content.

Adds a test per shape.

* fix(guardrails): redact the anthropic system prompt and string-array input

Two more provider-bound shapes, found by walking the request types rather
than waiting for them to be reported.

/v1/messages carries its system prompt at the top level, as a string or a
list of text blocks. It is one of the endpoints this guardrail claims to
cover, and a system prompt is a natural place to put a customer's details.

`input` as an array of bare strings, the embeddings and moderations shape,
was skipped because the loop only handled item dicts.

Verified against a live provider: a system prompt holding an address now
reaches the model as a stand-in and is restored in the reply.

* fix(guardrails): narrow prompt and input to a list before iterating

Guarding with a conditional iterable left the value un-narrowed, so passing
it on was an argument-type error and the element checks read as unreachable.
An early return narrows it properly and reads better.

* fix(guardrails): restore every streaming choice, not just the first

Streaming rehydration read and rewrote choices[0] only, so with n>1 every
later choice went back to the caller still holding its placeholders.

Each choice is its own token stream, so the sliding window is now tracked
per choice index rather than once per stream. A single shared window would
have been worse than the bug: it would splice the characters held back for
one choice onto the next one's delta.

The final flush walks every choice the same way, and the two helpers that
only ever looked at choices[0] are gone.

Adds a test that both choices come back restored, and one that each choice
gets its own window handed back rather than its neighbour's.

* refactor(guardrails): name the guardrail llm_shield_proxy throughout

The integration was called llm_shield in code, llm-shield in the example
config, and LLM Shield in the dashboard, while the product and its PyPI
package are both llm-shield-proxy. An operator who saw the guardrail in
LiteLLM could not tell what to install.

One identifier now: llm_shield_proxy for the enum value, module, directory,
class, config model, logo and environment variables, with LLM Shield Proxy
as the display name. That matches `pip install llm-shield-proxy`.

Renames only; no behaviour change.

* feat(guardrails): redact the participant name on a message

`name` on a user or assistant turn identifies a person and was going to the
provider intact. The proxy this integrates with already redacts it, so the
integration was the weaker of the two.

On a tool or function turn the same field carries the function's name, which
has to arrive unchanged or the call stops routing. That case is skipped, and
a test asserts the value is never even sent to the shield.

* fix(guardrails): flush every held choice, and cover tool results and suffix

Three review findings.

The trailing flush walked the last chunk's choices, so a choice that finished
earlier and stopped appearing lost whatever text was still held for it and its
answer was truncated. It is now driven by the windows themselves and emits one
chunk per choice, synthesising the choice when the terminal chunk omits it.
That was data loss, not just under-redaction.

An Anthropic tool_result carries its own content, as a string or as further
blocks, and only each part's `text` was being collected. Handled recursively;
image and audio parts still fall through untouched.

The legacy completions `suffix` is forwarded to providers that support it and
was never collected. Note the placement: it has to be gathered before the
string-prompt early return, which is what the new test pins.

* fix(guardrails): walk nested tool results iteratively, with a depth bound

CI flagged _collect_content as recursive. It was, and worse, it was unbounded:
a tool_result nests its own content, the nesting is caller controlled, and the
descent had nothing to stop it. That is a JSON bomb, not a style issue.

Now an explicit queue with a depth bound of 8. Real payloads nest one or two
deep. The queue is walked in document order because the shield maps its replies
back by position, so collection order is part of the contract.

* fix(guardrails): redact Responses PromptObject variables

A Responses request can send `prompt` as a PromptObject rather than a string.
Its `variables` are substituted into the stored prompt on the provider side, so
they are caller text, and the dict shape was falling through untouched.

`id` and `version` pick which stored prompt to run and are left unchanged.

* test(guardrails): assert the depth bound instead of only reaching the end

The depth test asserted nothing, so it passed whether or not the bound held,
and the test-quality gate counted it as a zero-assert test. It now sends a
shallow value alongside a 200-deep chain and asserts the shallow one is
collected while the value past the bound is not.

* fix(guardrails): keep system-prompt values out of the restored reply

Redaction put every span of a request into one vault, and the reply was restored
against that same vault. System prompts are written by the application and the
caller never sees them, so a caller who got the model to echo a placeholder back
had its plaintext restored into their own reply -- a way to read a system prompt
they were never shown.

Server-authored spans now go into a vault of their own: system and developer
turns, Anthropic's top-level `system`, and the Responses API `instructions`.
Its id is deliberately never stored, so nothing restores against it. The reply
is restored against the caller's vault alone, and an echoed placeholder from a
system prompt comes back as the placeholder.

Values the caller also wrote themselves are unaffected -- they are in the
caller's vault too, and still restore. The extra round trip happens only when a
request actually carries server-authored text.

* style(guardrails): satisfy ruff format and annotate the new tests

`ruff format` wanted the widened `_collect_responses_fields` signature on one
line, and the three tests added with the split-vault fix needed return
annotations to keep ANN201 level with the base.

* fix(guardrails): restore tool calls in the LLM Shield guardrail

The request walk redacted a tool call's `arguments` -- plus the legacy `function_call`,
Anthropic `tool_use.input` leaves and the Responses API's `function_call` /
`function_call_output` fields -- while the response walk restored only `message.content`.
A placeholder therefore reached the caller inside a tool call, and nothing raised.

This is the same change as the out-of-tree example adapter this file is copied from, kept
body-identical on purpose: the response side now collects every restorable span in one
positional rehydrate batch, streaming keeps a window per (choice index, tool-call index)
and flushes each into the chunk carrying the finish_reason, and `apply_guardrail` restores
`inputs["tool_calls"]` on the response side. The declared limit on restoring values inside
a JSON string is documented in the module.

* fix(guardrails): import copy, keep the vault id off the provider, drop recursion

Three defects Greptile and veria-ai found on the reopened PR, all real:

- `copy.deepcopy` was called in `apply_guardrail` with no `import copy`, a
  guaranteed NameError on every response carrying tool calls. It landed on
  2026-09-13, ten days after the review that rated this branch safe, and no test
  reached it: every tool-call test covered the request side. Adds the import and
  a regression test on the response side.
- The vault session id was stored in `metadata`, which is forwarded to the
  provider on /v1/responses. A provider holding the placeholders and the session
  id can call the shield's rehydrate endpoint and read back the plaintext this
  guardrail exists to withhold. Moves it to `litellm_metadata`, which is not
  forwarded, and reads it back from there only.
- `_collect_json_leaves` recursed over model-controlled JSON; the repo's
  recursive_detector gate rejects that. Rewritten with an explicit stack, same
  depth bound.

52 tests pass. ruff format, ruff-strict and check_type_discipline all clean, with
LIT counts identical to the merge base.

* fix(guardrails): build llm_shield_proxy stream deltas without new mutable literals

The lint job's LIT002 budget gate failed on this PR: the file added 11
mutable-collection constructions and the tree sits at its limit. Build the
index-only tool-call continuation in one helper, keep read-only inputs as
tuples, and annotate the lists the delta and texts fields require.

Adds tests for the two tool-call flush paths the refactor touches, which
had no coverage: held arguments landing in the finish_reason chunk next to
that chunk's own fragment, and the trailing flush of a stream that ends
without a finish_reason.

* fix(guardrails): drop Final from loop-body locals in llm_shield_proxy

basedpyright rejects Final on a name assigned inside a loop, and the eleven
such locals put reportGeneralTypeIssues over its budget (112/101). The LIT010
Final rule already exempts loop-body assignments, so the annotations go.

* feat(guardrails): restore llm_shield_proxy placeholders on native streams

Anthropic /v1/messages and /v1/responses streams have no `choices`, so the
streaming hook passed them through with placeholders still in them. Both
are now restored incrementally, with the same per-stream windows as chat:

- /v1/messages arrives as raw SSE. Frames are cut at event boundaries,
  text_delta and input_json_delta are restored per block index, and held
  text is emitted as one more delta ahead of content_block_stop. Signed
  thinking deltas, frames from other endpoints and non-SSE raw streams
  pass through unchanged.
- /v1/responses events are restored per item and part. Held text goes out
  as a copy of the stream's last delta before its .done event, and the
  events that repeat the reply (.done, content_part.done, output_item.done,
  response.completed) are restored in full.

The request side now also redacts Anthropic tool_use inputs and Responses
reasoning summaries, and sends tool and function descriptions (including
parameter schema descriptions) and the user / safety_identifier fields to
the non-restorable vault, like system prompts. Tool results stay
restorable: the model reads them to answer, so restoring them returns what
the caller would have seen without the guardrail.

* fix(guardrails): redact llm_shield_proxy predicted outputs and output schemas

`prediction.content` is the caller's own draft of the reply, so it is
redacted into the caller vault and restored with the reply. The
descriptions in a structured-output schema (Chat
response_format.json_schema, Responses text.format) are application
authored like tool schemas, so they go to the non-restorable vault.

* fix(guardrails): fail closed on deep llm_shield_proxy requests, widen coverage

- Request walks no longer skip what lies past their depth bound. Content
  nested past it, and tool inputs or schemas past the new JSON bound, now
  block the request instead of reaching the provider unredacted. The old
  depth test asserted the skip; it now asserts the block.
- Tool and output schemas are walked by their JSON Schema structure, and
  give up `title`, `examples` and `default` as well as `description`.
  `enum` and `const` still go out as sent.
- Responses events are matched by shape: any `*.delta` with a string delta
  is a token stream, and any `*.done` restores every non-identifier text
  field plus the `part` or `item` it repeats. This covers
  reasoning_summary_part.done and MCP arguments, and future families.
  Audio deltas are left alone.
- An SSE stream whose first chunk ends partway through a field name
  (`b"eve"`) is no longer taken for a non-SSE stream.

* fix(guardrails): scan llm_shield_proxy schemas by default

The schema walk collected an allowlist of keywords, so any keyword it did
not list -- draft-07 `dependencies`, `$comment`, vendor `x-` extensions --
went to the provider in clear. Invert it: every string is collected except
under keywords whose value must go out verbatim (types, formats, patterns,
references, required lists, enum, const). Name -> subschema maps still
treat their keys as property names, so a property called `type` is
walked, not skipped.

* fix(guardrails): redact llm_shield_proxy schema enum and const values

`enum` and `const` were skipped by the schema walk, so a value holding PII
went to the provider in clear. They now go to the caller's vault rather
than the non-restorable one: the model emits the stand-in in its tool
arguments or structured output, and restoring the reply turns it back into
the value the schema allows, so the call still routes.

* fix(guardrails): redact llm_shield_proxy web search user locations

Web search forwards the user's approximate location, and its free-text
`city` and `region` fields can hold an address. Collect them into the
non-restorable vault, from Chat `web_search_options.user_location` and
from the `user_location` of Responses and Anthropic web-search tools.

* fix(guardrails): drop unused llm_shield_proxy suppressions

Upstream added LIT013 (a *-ok marker that suppresses nothing) and LIT014
(at most one for and one if per comprehension). Remove the 34 markers
that no longer suppress anything and flatten the finished streams with
itertools.chain.from_iterable.

* fix(guardrails): type the llm_shield_proxy request and reply walks

Narrowing with isinstance(x, dict) leaves keys and values unknown, so
every call that passed a narrowed value counted against the
reportUnknownArgumentType budget. Parse into dict[str, object] and
list[object] once, in _as_object and _as_array, type the carry keys and
accumulators, and bind writers with functools.partial instead of lambdas.
The shield's batch reply is now also checked to hold only strings.

* fix(guardrails): keep restored llm_shield_proxy replies out of the cache, widen coverage

Addresses the open veria-ai and Cursor Bugbot findings on #42645.

- Restore a copy of the reply and of each stream chunk, never LiteLLM's own object.
  LiteLLM caches and logs that object, and placeholders are numbered per request, so
  two callers' redacted requests can share a cache key: restoring in place cached one
  caller's plaintext for the next. The deployment hook no longer restores either,
  since LiteLLM caches what it returns; the proxy's post-call hook restores
  model-level guardrails after the cache write.
- Restore /v1/completions replies, streamed and not, which carry `choice.text`.
- Redact Responses replay fields the reply side already restores: tool output sent
  as input_text parts, custom_tool_call `input`, code_interpreter_call `code`.
- Redact typed Responses prompt variables (`{"type": "input_text", "text": ...}`).
- Put Responses system and developer input items in the non-restorable vault, like
  their Chat counterparts.
- Expose LLMShieldProxyGuardrailConfigModel through get_config_model, so the
  dashboard can collect the Shield URL and key.

* fix(guardrails): redact llm_shield_proxy plain-text document blocks

An Anthropic document block carries text inline, in a text source's `data` or a
content source's `content`, and that text reached the provider unredacted. Collect
both, plus the block's `title` and `context`; base64, URL and file sources pass
untouched.

* fix(guardrails): redact llm_shield_proxy extra_body overrides

LiteLLM merges extra_body over the transformed request just before sending, so text
placed there (input, messages, system, ...) replaced the redacted field on the wire.
Walk extra_body with the same collectors as the request, keeping the caller /
application split.

* test(guardrails): import InMemoryCache directly in the llm_shield_proxy cache test

litellm keeps a deprecated module-level `caching` bool, so `litellm.caching.caching`
resolves to that bool once an earlier test in the same worker has set it, and the test
failed with AttributeError depending on test order.

* fix(guardrails): restore llm_shield_proxy replies for model-level use outside the proxy

71e68fd stopped the deployment post-call hook from restoring, so the response cache
never holds restored plaintext. Inside the proxy that is right: the proxy's post-call
hook restores after the cache write. But with model-level `guardrails` on the SDK,
the deployment hooks are the only redact and restore steps, so callers got
placeholders back.

When the deployment pre-call hook is the one that redacts, it now records the
request's vault id and marks the request no-cache / no-store; the deployment
post-call hook restores only when that record matches. The cache key there is built
from the redacted request and a cache hit skips the post-call hook, so a cached reply
could neither be restored nor safely shared. Proxy requests carry no record and keep
restoring in the proxy's post-call hook, after the cache write.

* fix(guardrails): don't repeat usage in llm_shield_proxy end-of-stream flush chunks

With n>=2 and stream_options.include_usage, the end-of-stream flush copies the last
chunk the stream carried, which is the one holding usage, so each synthetic flush
chunk repeated it and a consumer summing usage chunks counted the request twice.
The copy now drops `usage`, matching a normal mid-stream chunk. Reported by
@yucheng-berri.

* fix(guardrails): keep restored llm_shield_proxy values out of telemetry, refuse SDK streams

- The post-call restore hook no longer goes through log_guardrail_information,
  which recorded its whole return value, the restored reply, as guardrail_response.
  That field is exported to traces even with message logging turned off.
- A model-level stream outside the proxy is refused once redacted. Nothing restores
  an SDK stream, and its cache writer reads the request from before the deployment
  hook, so it also got cached despite the no-store bypass.
- Drop a narrating comment, and keep example_config.yaml to config only; the
  how-to lives in the docs PR.

* refactor(guardrails): split llm_shield_proxy into payload, request walk and stream modules

The module had grown past 1,700 lines. Shared payload types and helpers move to
payload.py, the request walk to request_walk.py and the stream restorers to
stream_restorers.py; llm_shield_proxy.py keeps the guardrail class. No behaviour
change.

* style(guardrails): drop routine comments from llm_shield_proxy

AGENTS.md keeps source comments to tool directives and genuinely complex logic;
the rationale stays in the docstrings.
2026-10-05 10:33:56 -07:00