Commit graph

6085 commits

Author SHA1 Message Date
ishaan-berri
6764b9af12
feat(lens): type to filter the traces agent dropdown (#44745)
* feat(lens): type to filter the traces agent dropdown

* test(lens): cover typing, Enter and clear in the agent filter
2026-10-05 18:36:22 -07:00
moe-berri
b58e2d7175
fix(lens): bound result recovery and preserve partial results (#44692)
* fix(lens): restate response contract during model repair

* fix(lens): separate instructions and recover rejected results

* fix(lens): correct loop type annotations and checks

* fix(lens): preserve access to prior findings after compaction

* fix(lens): cap result retries and preserve partial completion
2026-10-05 18:24:53 -07:00
tin-berri
464fe5bd90
feat(ui): configure cache-aware auto routing (#43396)
* feat(ui): configure cache-aware auto routing

* style(ui): keep cache routing config within line limit
2026-10-05 18:24:24 -07:00
moe-berri
4cc442f5ee
feat(lens): show investigation names in findings table (#44753) 2026-10-06 01:23:50 +00:00
devin-ai-integration[bot]
56a581a536
fix(deps): bump source-map-js, smol-toml, mako, multidict and werkzeug for OSV advisories (#44728)
* fix(deps): bump source-map-js to 1.2.2 for GHSA-68fv-2mgg-jv7q

* fix(deps): bump mako, multidict, werkzeug and smol-toml for OSV advisories

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-06 00:57:04 +00:00
moe-berri
2190003bb8
fix(lens): reconstruct native coding agent conversations (#44711)
* fix(lens): reconstruct native coding agent conversations

* fix(lens): display timestamps used for conversation ordering

* fix(lens): preserve coding trace identity and message provenance

* fix(lens): complete capture checks after trace pagination

* fix(lens): show loaded replies during trace pagination
2026-10-05 17:51:05 -07:00
moyai-devin-berriai[bot]
6e75f28289
feat(ui): support native decisions endpoint in decision playground (#44664)
* feat(ui): support native decisions endpoint in decision playground

* fix(ui): preserve decision playground defaults and test late cancellation results

* test(ui): remove redundant cancellation test comment

* Update ui/litellm-dashboard/src/app/(dashboard)/playground/components/systemOneUI/SystemOneUI.integration.test.tsx

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-10-06 00:48:10 +00:00
moe-berri
c29b42a32b
fix(lens): preserve span timestamps in investigation evidence (#44702)
* fix(lens): preserve span timestamps in investigation evidence

* style(lens): format chronology regression test
2026-10-05 16:42:42 -07:00
moe-berri
c561372a02
fix(lens): show investigation findings for agent traces (#44696)
* fix(lens): show investigation findings for agent traces

Replace child tool-error counts in the trace table with distinct investigation findings. Keep unassessed traces separate from completed clean investigations and use the root status for failure filters and timeline counts

* fix(lens): stabilize findings updates and repair UI checks

* fix(lens): allow viewer findings reads and index trace lookups

* test(lens): cover viewer findings reads with Postgres
2026-10-05 16:27:54 -07:00
moe-berri
b69d744993
feat(lens): analyze trace workspaces with confined Python and compaction (#44640)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* feat(lens): integrate confined workspace analysis with live reviews

* fix(lens): synchronize confined Python process monitoring

* Update review.md

* fix(lens): allow mixed context capacities and correct review assertions

* fix(lens): retrieve evidence on demand and isolate failed reviews

* fix(lens): isolate incomplete evidence reads from peer reviews

* test(lens): await trace status filter option

* test(lens): wait for reclaimed review state to settle

* fix(lens): recover from incomplete cross-session evidence

---------

Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
2026-10-05 22:06:48 +00:00
ryan-crabbe-berri
034664001f
fix(ui): send null instead of $0 when a team's member default budget is cleared (#44644)
Clearing Default Budget (USD) under Team Member Settings sent Number("") = 0, which
turned the shared member default into a $0 cap and blocked every member still on
the default. Opening Team Member Settings on a team whose default has no dollar cap
did the same through Number(null).

Team settings numeric fields now go through one shared numberOrNull helper, which
the team admin settings form already used
2026-10-05 15:01:12 -07:00
moe-berri
fe24be2e3d
fix(lens): make tool steps and conversations readable (#44645)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* Update review.md

* fix(lens): make tool steps and conversations readable

* fix(lens): address trace rendering review and test failures

* fix(lens): preserve conversations with incomplete tool calls

* test(lens): retain failed tool styling coverage

* fix(lens): keep tool metadata in accessible result groups

* refactor(lens): build stable agent labels without mutation

* perf(lens): group and sort agent labels without repeated scans

* test(lens): await trace status filter option

---------

Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
2026-10-05 22:00:57 +00:00
devin-ai-integration[bot]
7dc5b73f93
feat(ui): declare shared search operators and flush queries on blur (#44665)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:00:34 -07:00
devin-ai-integration[bot]
a3e15774ad
feat(proxy): limit which models an end user can call (#43904)
* feat(proxy): add models column to the end user table

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): enforce the end user models allowlist in model access checks

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): accept and return models on the customer endpoints

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover the customer models allowlist

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): resolve team aliases before the end user model check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:58:03 +00:00
ishaan-berri
a3a52466a5
feat(lens): restore compact navigation with live investigation review (#44472)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* Update review.md

* test(lens): await trace status filter option

---------

Co-authored-by: moe-berri <moe@berri.ai>
2026-10-05 14:50:43 -07:00
devin-ai-integration[bot]
f0415ee033
fix(ui): explain why team member reset spend is unavailable instead of hiding it (#44629)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 14:31:04 -07:00
devin-ai-integration[bot]
45e7be1abc
feat(tracing)!: return only data from SQL queries (#44609)
* feat(tracing): add generated SQL response contract

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): return data-only trace SQL responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(rust): format trace response schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 12:11:59 -07:00
moyai-devin-berriai[bot]
75b45e39c9
fix(proxy): enforce internal-user model creation prohibition (#44438)
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
2026-10-05 11:32:05 -07:00
moe-berri
a99bccacea
fix(ui): restore inline Lens onboarding and responsive layout (#44604)
* fix(ui): improve gateway layouts on mobile

* fix(ui): limit mobile cleanup to navigation and header

* fix(ui): restore inline Lens onboarding and responsive layout

* fix(ui): smooth Lens tab and panel connections

* fix(ui): keep trace time controls within narrow panels
2026-10-05 11:10:10 -07:00
moe-berri
3b2ed83152
fix(ui): make mobile sidebar and top bar responsive (#44603)
* fix(ui): improve gateway layouts on mobile

* fix(ui): limit mobile cleanup to navigation and header

* fix(ui): share cached settings with mobile navigation
2026-10-05 10:56:08 -07:00
devin-ai-integration[bot]
9b6a6a0b71
refactor(tracing): generate existing HTTP request models from Rust schemas (#44591)
* test(tracing): pin HTTP request compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): generate existing HTTP request models from Rust schemas

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): bind trace query params to generated request models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 48 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(tracing): read the trace list clock without a thread-pool dispatch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(traces): share the trace page-size bounds between schema and reader

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): type trace request queries against the generated OpenAPI schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: encode the trace contract boundary in AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format trace request aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 17:37:52 +00:00
Ninad Phalak
d946706744
feat(guardrails): add llm shield pii redaction and rehydration guardrail (#42645)
* feat(guardrails): add llm shield pii redaction and rehydration guardrail

LLM Shield is a self-hosted PII gateway. This adds it as a guardrail so a
proxy operator can redact personal data out of outbound requests and have
the original values restored in the model's reply.

The substitution is reversible, which is the difference from a masking
guardrail. Outbound text is replaced with placeholders held in a session
vault inside the operator's own LLM Shield deployment, and the reply is
restored before it reaches the caller, so the end user still sees real
values while the provider never received them.

Streaming responses are restored incrementally. LLM Shield holds back only
the trailing characters that could still turn out to be part of a
placeholder, so tokens are forwarded as they arrive rather than the whole
response being collected first. A placeholder split across two chunks is
never emitted in fragments.

The integration talks to LLM Shield over HTTP and adds no dependency.

Notes for reviewers:

- The guardrail sets use_native_lifecycle_hooks, since redaction and
  restoration need the native pre-call, post-call and streaming hooks
  rather than the unified path.
- Per-request state lives on the request dict, never on the guardrail
  instance, because the proxy registers a single instance process-wide.
  The streaming carry-over is a local of the generator for the same reason.
- Every failure blocks the request. A redaction guardrail that fails open
  would send the exact data it exists to protect to the provider.

* feat(ui): list llm shield in the guardrail garden

Adds the card, preset and logo so operators can pick LLM Shield from the
guardrails page the same way as the other partner guardrails.

* docs(guardrails): add llm shield example config

Shows both modes on one entry. Listing only pre_call redacts the request
and then hands the placeholders back to the end user, so the test asserts
both hooks are enabled.

* feat(ui): use the llm shield brand mark for the guardrail logo

* fix(guardrails): restore llm shield values in anthropic replies

The /v1/messages reply is a plain dict with a content block list and no
choices, so it fell through the restore path and went back to the caller
still carrying placeholders. The request was redacted correctly, which is
what made this easy to miss.

Found by running all three endpoints against a live provider; the mocked
tests all passed because they only built the OpenAI shape. Adds tests for
the message shape and for leaving non-text blocks alone.

* docs(guardrails): correct the llm shield start command

* fix(guardrails): redact every request shape and restore every reply shape

Three gaps, all of which let an enabled guardrail hand data to the provider
or hand placeholders to the caller.

Requests only walked `messages`. The Responses API `input` and tool call
`arguments` went out untouched. Measured against a live provider: a request
sent through `/v1/responses` reached the model with the real address in it
while the guardrail reported as enabled. Request traversal now covers chat
content (string and multimodal), tool call arguments, and `input` as a bare
string or a list of items.

Fixing that exposed the matching gap on the way back: the Responses API reply
carries `output` items rather than `choices`, so it returned to the caller
still holding placeholders. It now gets its own walk, handling text blocks as
dicts or objects.

The dashboard preset seeded only pre_call, so a guardrail created from the UI
would redact the request and return the placeholders to the user. Presets can
now seed both modes; the form already normalised either shape.

Adds tests for each request shape, for both Responses API reply forms, and
replaces a test that had asserted the `input` bypass as correct behaviour.

* fix(guardrails): narrow the stream delta before writing to it

basedpyright could not prove the delta was non-None on the write path, and
reportOptionalMemberAccess has a zero budget. The guard is also clearer than
relying on the text check to imply it.

* fix(guardrails): mint the vault id instead of trusting the caller's

The vault id was taken from caller-supplied session metadata, and every
caller shares one LLM Shield key. Someone who knew or guessed another
caller's session id could send a placeholder, have the model echo it back,
and get that caller's plaintext restored into their own reply.

Vault ids are now minted per request behind a per-process prefix, so a
caller cannot name a vault this process uses. Redaction mints, restoration
reads back, and a reply whose id does not match is left holding its
placeholders rather than resolved against some other vault.

Also covers two more request fields that were reaching the provider intact:
the Responses API `instructions`, and the legacy `function_call.arguments`
alongside `tool_calls`.

The collectors move to module level, which drops the traversal back under
the complexity limit and lets the code carry its own explanation instead of
the comments that were restating it.

* fix(guardrails): drop Final from a loop-assigned local

basedpyright rejects a Final assigned inside a loop, and
reportGeneralTypeIssues sits one over its budget ceiling.

* fix(guardrails): redact completion prompts and responses tool items

Two more provider-bound request shapes were reaching the model intact while
the guardrail reported as enabled.

/v1/completions carries its text in a top-level `prompt`, which the
traversal never looked at. It is handled as a string and as the array form,
where each entry is rewritten in place.

Responses input items hold tool data outside `content`: a function_call item
in `arguments`, a function_call_output item in `output`. Both are now
collected alongside the item's content.

Adds a test per shape.

* fix(guardrails): redact the anthropic system prompt and string-array input

Two more provider-bound shapes, found by walking the request types rather
than waiting for them to be reported.

/v1/messages carries its system prompt at the top level, as a string or a
list of text blocks. It is one of the endpoints this guardrail claims to
cover, and a system prompt is a natural place to put a customer's details.

`input` as an array of bare strings, the embeddings and moderations shape,
was skipped because the loop only handled item dicts.

Verified against a live provider: a system prompt holding an address now
reaches the model as a stand-in and is restored in the reply.

* fix(guardrails): narrow prompt and input to a list before iterating

Guarding with a conditional iterable left the value un-narrowed, so passing
it on was an argument-type error and the element checks read as unreachable.
An early return narrows it properly and reads better.

* fix(guardrails): restore every streaming choice, not just the first

Streaming rehydration read and rewrote choices[0] only, so with n>1 every
later choice went back to the caller still holding its placeholders.

Each choice is its own token stream, so the sliding window is now tracked
per choice index rather than once per stream. A single shared window would
have been worse than the bug: it would splice the characters held back for
one choice onto the next one's delta.

The final flush walks every choice the same way, and the two helpers that
only ever looked at choices[0] are gone.

Adds a test that both choices come back restored, and one that each choice
gets its own window handed back rather than its neighbour's.

* refactor(guardrails): name the guardrail llm_shield_proxy throughout

The integration was called llm_shield in code, llm-shield in the example
config, and LLM Shield in the dashboard, while the product and its PyPI
package are both llm-shield-proxy. An operator who saw the guardrail in
LiteLLM could not tell what to install.

One identifier now: llm_shield_proxy for the enum value, module, directory,
class, config model, logo and environment variables, with LLM Shield Proxy
as the display name. That matches `pip install llm-shield-proxy`.

Renames only; no behaviour change.

* feat(guardrails): redact the participant name on a message

`name` on a user or assistant turn identifies a person and was going to the
provider intact. The proxy this integrates with already redacts it, so the
integration was the weaker of the two.

On a tool or function turn the same field carries the function's name, which
has to arrive unchanged or the call stops routing. That case is skipped, and
a test asserts the value is never even sent to the shield.

* fix(guardrails): flush every held choice, and cover tool results and suffix

Three review findings.

The trailing flush walked the last chunk's choices, so a choice that finished
earlier and stopped appearing lost whatever text was still held for it and its
answer was truncated. It is now driven by the windows themselves and emits one
chunk per choice, synthesising the choice when the terminal chunk omits it.
That was data loss, not just under-redaction.

An Anthropic tool_result carries its own content, as a string or as further
blocks, and only each part's `text` was being collected. Handled recursively;
image and audio parts still fall through untouched.

The legacy completions `suffix` is forwarded to providers that support it and
was never collected. Note the placement: it has to be gathered before the
string-prompt early return, which is what the new test pins.

* fix(guardrails): walk nested tool results iteratively, with a depth bound

CI flagged _collect_content as recursive. It was, and worse, it was unbounded:
a tool_result nests its own content, the nesting is caller controlled, and the
descent had nothing to stop it. That is a JSON bomb, not a style issue.

Now an explicit queue with a depth bound of 8. Real payloads nest one or two
deep. The queue is walked in document order because the shield maps its replies
back by position, so collection order is part of the contract.

* fix(guardrails): redact Responses PromptObject variables

A Responses request can send `prompt` as a PromptObject rather than a string.
Its `variables` are substituted into the stored prompt on the provider side, so
they are caller text, and the dict shape was falling through untouched.

`id` and `version` pick which stored prompt to run and are left unchanged.

* test(guardrails): assert the depth bound instead of only reaching the end

The depth test asserted nothing, so it passed whether or not the bound held,
and the test-quality gate counted it as a zero-assert test. It now sends a
shallow value alongside a 200-deep chain and asserts the shallow one is
collected while the value past the bound is not.

* fix(guardrails): keep system-prompt values out of the restored reply

Redaction put every span of a request into one vault, and the reply was restored
against that same vault. System prompts are written by the application and the
caller never sees them, so a caller who got the model to echo a placeholder back
had its plaintext restored into their own reply -- a way to read a system prompt
they were never shown.

Server-authored spans now go into a vault of their own: system and developer
turns, Anthropic's top-level `system`, and the Responses API `instructions`.
Its id is deliberately never stored, so nothing restores against it. The reply
is restored against the caller's vault alone, and an echoed placeholder from a
system prompt comes back as the placeholder.

Values the caller also wrote themselves are unaffected -- they are in the
caller's vault too, and still restore. The extra round trip happens only when a
request actually carries server-authored text.

* style(guardrails): satisfy ruff format and annotate the new tests

`ruff format` wanted the widened `_collect_responses_fields` signature on one
line, and the three tests added with the split-vault fix needed return
annotations to keep ANN201 level with the base.

* fix(guardrails): restore tool calls in the LLM Shield guardrail

The request walk redacted a tool call's `arguments` -- plus the legacy `function_call`,
Anthropic `tool_use.input` leaves and the Responses API's `function_call` /
`function_call_output` fields -- while the response walk restored only `message.content`.
A placeholder therefore reached the caller inside a tool call, and nothing raised.

This is the same change as the out-of-tree example adapter this file is copied from, kept
body-identical on purpose: the response side now collects every restorable span in one
positional rehydrate batch, streaming keeps a window per (choice index, tool-call index)
and flushes each into the chunk carrying the finish_reason, and `apply_guardrail` restores
`inputs["tool_calls"]` on the response side. The declared limit on restoring values inside
a JSON string is documented in the module.

* fix(guardrails): import copy, keep the vault id off the provider, drop recursion

Three defects Greptile and veria-ai found on the reopened PR, all real:

- `copy.deepcopy` was called in `apply_guardrail` with no `import copy`, a
  guaranteed NameError on every response carrying tool calls. It landed on
  2026-09-13, ten days after the review that rated this branch safe, and no test
  reached it: every tool-call test covered the request side. Adds the import and
  a regression test on the response side.
- The vault session id was stored in `metadata`, which is forwarded to the
  provider on /v1/responses. A provider holding the placeholders and the session
  id can call the shield's rehydrate endpoint and read back the plaintext this
  guardrail exists to withhold. Moves it to `litellm_metadata`, which is not
  forwarded, and reads it back from there only.
- `_collect_json_leaves` recursed over model-controlled JSON; the repo's
  recursive_detector gate rejects that. Rewritten with an explicit stack, same
  depth bound.

52 tests pass. ruff format, ruff-strict and check_type_discipline all clean, with
LIT counts identical to the merge base.

* fix(guardrails): build llm_shield_proxy stream deltas without new mutable literals

The lint job's LIT002 budget gate failed on this PR: the file added 11
mutable-collection constructions and the tree sits at its limit. Build the
index-only tool-call continuation in one helper, keep read-only inputs as
tuples, and annotate the lists the delta and texts fields require.

Adds tests for the two tool-call flush paths the refactor touches, which
had no coverage: held arguments landing in the finish_reason chunk next to
that chunk's own fragment, and the trailing flush of a stream that ends
without a finish_reason.

* fix(guardrails): drop Final from loop-body locals in llm_shield_proxy

basedpyright rejects Final on a name assigned inside a loop, and the eleven
such locals put reportGeneralTypeIssues over its budget (112/101). The LIT010
Final rule already exempts loop-body assignments, so the annotations go.

* feat(guardrails): restore llm_shield_proxy placeholders on native streams

Anthropic /v1/messages and /v1/responses streams have no `choices`, so the
streaming hook passed them through with placeholders still in them. Both
are now restored incrementally, with the same per-stream windows as chat:

- /v1/messages arrives as raw SSE. Frames are cut at event boundaries,
  text_delta and input_json_delta are restored per block index, and held
  text is emitted as one more delta ahead of content_block_stop. Signed
  thinking deltas, frames from other endpoints and non-SSE raw streams
  pass through unchanged.
- /v1/responses events are restored per item and part. Held text goes out
  as a copy of the stream's last delta before its .done event, and the
  events that repeat the reply (.done, content_part.done, output_item.done,
  response.completed) are restored in full.

The request side now also redacts Anthropic tool_use inputs and Responses
reasoning summaries, and sends tool and function descriptions (including
parameter schema descriptions) and the user / safety_identifier fields to
the non-restorable vault, like system prompts. Tool results stay
restorable: the model reads them to answer, so restoring them returns what
the caller would have seen without the guardrail.

* fix(guardrails): redact llm_shield_proxy predicted outputs and output schemas

`prediction.content` is the caller's own draft of the reply, so it is
redacted into the caller vault and restored with the reply. The
descriptions in a structured-output schema (Chat
response_format.json_schema, Responses text.format) are application
authored like tool schemas, so they go to the non-restorable vault.

* fix(guardrails): fail closed on deep llm_shield_proxy requests, widen coverage

- Request walks no longer skip what lies past their depth bound. Content
  nested past it, and tool inputs or schemas past the new JSON bound, now
  block the request instead of reaching the provider unredacted. The old
  depth test asserted the skip; it now asserts the block.
- Tool and output schemas are walked by their JSON Schema structure, and
  give up `title`, `examples` and `default` as well as `description`.
  `enum` and `const` still go out as sent.
- Responses events are matched by shape: any `*.delta` with a string delta
  is a token stream, and any `*.done` restores every non-identifier text
  field plus the `part` or `item` it repeats. This covers
  reasoning_summary_part.done and MCP arguments, and future families.
  Audio deltas are left alone.
- An SSE stream whose first chunk ends partway through a field name
  (`b"eve"`) is no longer taken for a non-SSE stream.

* fix(guardrails): scan llm_shield_proxy schemas by default

The schema walk collected an allowlist of keywords, so any keyword it did
not list -- draft-07 `dependencies`, `$comment`, vendor `x-` extensions --
went to the provider in clear. Invert it: every string is collected except
under keywords whose value must go out verbatim (types, formats, patterns,
references, required lists, enum, const). Name -> subschema maps still
treat their keys as property names, so a property called `type` is
walked, not skipped.

* fix(guardrails): redact llm_shield_proxy schema enum and const values

`enum` and `const` were skipped by the schema walk, so a value holding PII
went to the provider in clear. They now go to the caller's vault rather
than the non-restorable one: the model emits the stand-in in its tool
arguments or structured output, and restoring the reply turns it back into
the value the schema allows, so the call still routes.

* fix(guardrails): redact llm_shield_proxy web search user locations

Web search forwards the user's approximate location, and its free-text
`city` and `region` fields can hold an address. Collect them into the
non-restorable vault, from Chat `web_search_options.user_location` and
from the `user_location` of Responses and Anthropic web-search tools.

* fix(guardrails): drop unused llm_shield_proxy suppressions

Upstream added LIT013 (a *-ok marker that suppresses nothing) and LIT014
(at most one for and one if per comprehension). Remove the 34 markers
that no longer suppress anything and flatten the finished streams with
itertools.chain.from_iterable.

* fix(guardrails): type the llm_shield_proxy request and reply walks

Narrowing with isinstance(x, dict) leaves keys and values unknown, so
every call that passed a narrowed value counted against the
reportUnknownArgumentType budget. Parse into dict[str, object] and
list[object] once, in _as_object and _as_array, type the carry keys and
accumulators, and bind writers with functools.partial instead of lambdas.
The shield's batch reply is now also checked to hold only strings.

* fix(guardrails): keep restored llm_shield_proxy replies out of the cache, widen coverage

Addresses the open veria-ai and Cursor Bugbot findings on #42645.

- Restore a copy of the reply and of each stream chunk, never LiteLLM's own object.
  LiteLLM caches and logs that object, and placeholders are numbered per request, so
  two callers' redacted requests can share a cache key: restoring in place cached one
  caller's plaintext for the next. The deployment hook no longer restores either,
  since LiteLLM caches what it returns; the proxy's post-call hook restores
  model-level guardrails after the cache write.
- Restore /v1/completions replies, streamed and not, which carry `choice.text`.
- Redact Responses replay fields the reply side already restores: tool output sent
  as input_text parts, custom_tool_call `input`, code_interpreter_call `code`.
- Redact typed Responses prompt variables (`{"type": "input_text", "text": ...}`).
- Put Responses system and developer input items in the non-restorable vault, like
  their Chat counterparts.
- Expose LLMShieldProxyGuardrailConfigModel through get_config_model, so the
  dashboard can collect the Shield URL and key.

* fix(guardrails): redact llm_shield_proxy plain-text document blocks

An Anthropic document block carries text inline, in a text source's `data` or a
content source's `content`, and that text reached the provider unredacted. Collect
both, plus the block's `title` and `context`; base64, URL and file sources pass
untouched.

* fix(guardrails): redact llm_shield_proxy extra_body overrides

LiteLLM merges extra_body over the transformed request just before sending, so text
placed there (input, messages, system, ...) replaced the redacted field on the wire.
Walk extra_body with the same collectors as the request, keeping the caller /
application split.

* test(guardrails): import InMemoryCache directly in the llm_shield_proxy cache test

litellm keeps a deprecated module-level `caching` bool, so `litellm.caching.caching`
resolves to that bool once an earlier test in the same worker has set it, and the test
failed with AttributeError depending on test order.

* fix(guardrails): restore llm_shield_proxy replies for model-level use outside the proxy

71e68fd stopped the deployment post-call hook from restoring, so the response cache
never holds restored plaintext. Inside the proxy that is right: the proxy's post-call
hook restores after the cache write. But with model-level `guardrails` on the SDK,
the deployment hooks are the only redact and restore steps, so callers got
placeholders back.

When the deployment pre-call hook is the one that redacts, it now records the
request's vault id and marks the request no-cache / no-store; the deployment
post-call hook restores only when that record matches. The cache key there is built
from the redacted request and a cache hit skips the post-call hook, so a cached reply
could neither be restored nor safely shared. Proxy requests carry no record and keep
restoring in the proxy's post-call hook, after the cache write.

* fix(guardrails): don't repeat usage in llm_shield_proxy end-of-stream flush chunks

With n>=2 and stream_options.include_usage, the end-of-stream flush copies the last
chunk the stream carried, which is the one holding usage, so each synthetic flush
chunk repeated it and a consumer summing usage chunks counted the request twice.
The copy now drops `usage`, matching a normal mid-stream chunk. Reported by
@yucheng-berri.

* fix(guardrails): keep restored llm_shield_proxy values out of telemetry, refuse SDK streams

- The post-call restore hook no longer goes through log_guardrail_information,
  which recorded its whole return value, the restored reply, as guardrail_response.
  That field is exported to traces even with message logging turned off.
- A model-level stream outside the proxy is refused once redacted. Nothing restores
  an SDK stream, and its cache writer reads the request from before the deployment
  hook, so it also got cached despite the no-store bypass.
- Drop a narrating comment, and keep example_config.yaml to config only; the
  how-to lives in the docs PR.

* refactor(guardrails): split llm_shield_proxy into payload, request walk and stream modules

The module had grown past 1,700 lines. Shared payload types and helpers move to
payload.py, the request walk to request_walk.py and the stream restorers to
stream_restorers.py; llm_shield_proxy.py keeps the guardrail class. No behaviour
change.

* style(guardrails): drop routine comments from llm_shield_proxy

AGENTS.md keeps source comments to tool directives and genuinely complex logic;
the rationale stays in the docstrings.
2026-10-05 10:33:56 -07:00
devin-ai-integration[bot]
2d82915084
fix(mcp): bind OAuth clients to their upstream issuer (#37777)
* feat(mcp): advertise the SDK's latest spec revision and validate the RFC 9207 iss

MCPSpecVersion stopped at 2025-06-18 while the pinned SDK negotiates 2025-11-25, and the version LiteLLM puts on its own outbound initialize was a hardcoded historical member. Add the missing revision, name the highest revision we speak once, and pin it to the SDK's LATEST_PROTOCOL_VERSION with a test so the two cannot drift apart silently.

/authorize now seals the issuer it sent the user to into the OAuth state, and /callback holds the authorization response's RFC 9207 iss against it, refusing to forward a code that came back from an authorization server we never sent the user to. An absent iss, an unanchored server row and a state minted before the seal all keep their current behavior.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): keep params, query and fragment significant in issuer comparison

The shared canonicalizer drops all three, so two issuers differing only outside the path compared equal and a response from another tenant's authorization server would have continued through the flow.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(mcp): refresh generated API snapshots

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover OAuth client isolation and lint checks

* fix(mcp): preserve registered clients in the existing save payload

* test(mcp): cover optional OAuth registration metadata

* fix(mcp): preserve compatible OAuth registrations across edits

* fix(mcp): retain OAuth state through pending authorization

* fix(mcp): guard pending OAuth at form submission

* fix(mcp): discard canceled OAuth edit snapshots

* test(mcp): preserve complete OAuth registration assertions

* refactor(mcp): construct OAuth credential updates without mutation

* fix(mcp): simplify issuer binding and reject unverifiable callbacks

* fix(mcp): preserve replacement clients and pending redirect bindings

* fix(mcp): retain clients with replacement authentication methods

* fix(mcp): preserve cached clients and pin manual OAuth issuers

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-05 10:25:47 -07:00
moe-berri
eeb192d4a7
chore!: retire the integrated ROI calculator (#44477)
* chore!: retire the integrated ROI calculator

* chore(ui): remove unused ROI demo notice
2026-10-05 10:02:19 -07:00
devin-ai-integration[bot]
21881c5711
refactor(ui): extract shared timeline and time-range controls (#44584)
Move the timeline renderer to shared/timeline/Timeline taking buckets,
a selected window, and callbacks as props, and TimeRangeControls to
shared/timeline. Lens keeps bucketRuns as the adapter that converts
loaded traces into buckets, preserving behavior without the histogram
endpoint.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 16:11:52 +00:00
devin-ai-integration[bot]
adb59a7af8
fix(ui): remove Top models by task card from Model Leaderboard (#44502)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-05 08:42:19 -07:00
devin-ai-integration[bot]
9cc15e9320
feat(ui): add persistent columns and loading skeletons to Lens runs (#44579)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:03:37 +00:00
devin-ai-integration[bot]
8f6546df9f
refactor(ui): route dashboard URL state through nuqs parsers (#44537)
* docs(ui): add url-state agent skill

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(ui): migrate dashboard URL state to nuqs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): cover chat URL id sync with the real chat shell provider

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 22:49:51 -07:00
devin-ai-integration[bot]
5a5fd92d39
fix(dashboard): route dev API calls on Accept and fail fast on non-JSON 2xx (#44528)
* fix(dashboard): route dev API calls on Accept and fail fast on non-JSON 2xx

The next dev rewrite that sends API calls to the proxy keyed on
Content-Type: application/json, which openapi-fetch rightly omits on a
bodyless GET, so GET /lens fell through to the Lens page and returned
HTML. Route on Accept: application/json instead, send it from both HTTP
clients, turn a non-JSON 2xx into a non-retryable ApiError in the typed
client, and stop react-query from retrying ApiError below 500.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(dashboard): accept a JSON body served without a JSON content type

Test fakes and some servers hand back JSON as text/plain, so the typed
client only rejects a 2xx whose body does not parse as JSON. The
system_one request test expects the Accept header the legacy client now
sends.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 17:17:48 -07:00
devin-ai-integration[bot]
461a58c40a
refactor(lens): storage-independent trace reads, shared keyset pager, typed read failures (#44422)
* feat(lens): own trace reads behind a cached TraceStore port

Move storage-independent trace reads into litellm-traces-cache behind a
TraceStore port that ClickHouse implements. One keyset pager drives the
span, list span and spend reads, and a run list batch reads spend once.

Trace opens, pages and list summaries share one resolved read per trace
in an in-process cache with single-flight loading. Live traces and reads
with unknown spend expire after 5s, quiet traces after 10 minutes, failed
reads are never cached, and the accepted list page size is remembered per
scope.

Trace read failures map to their own status and code (400, 409, 413, 503
with Retry-After), and the trace drawer retries temporary failures while
offering only a refresh for changed or oversized traces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(lens): seed large profiles with server-side copies and long sessions

Replay one copy through the proxy, then copy it inside ClickHouse and
PostgreSQL with INSERT ... SELECT, rewriting trace, span and call IDs so
every copy keeps its own spend. Add three long single-trace sessions for
drawer paging and the oversized read path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chores

* style(lens): float the investigation setup badge on the tab edge

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens): restructure the trace drawer and polish its layout

Split the 547-line TraceDrawer into run/, tree/, span/, content/ and
conversation/ modules. Step rows now sit on one line with colored span
family tiles, and the per-row timing bar moved into an optional Waterfall
layout with a time axis. The steps and details panes are separated by the
shadcn Resizable handle, with the split remembered per orientation.

Span payloads go through one pure classifier (payloadView) that picks
messages, a tool result, a nested field tree or text. JSON-encoded field
values unfold into a tree, prose renders as markdown, repr and tracebacks
stay monospace, and every section offers a Raw view. LangChain's
serialized messages now render as conversation cards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* typesafety

* wip

* fmt

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 12:39:47 -07:00
devin-ai-integration[bot]
e1d16f51d1
refactor(ui): share CopyButton between Lens traces and logs (#44513)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 10:00:49 -07:00
devin-ai-integration[bot]
1459e00430
refactor(ui): rename view_logs to logs and split request, audit and detail (#44505)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 10:00:49 -07:00
devin-ai-integration[bot]
05f1c73a3c
refactor(ui): move TraceView into components/lens/traces (#44501)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 10:00:48 -07:00
devin-ai-integration[bot]
9a5e828310
feat(ui): inline Lens settings tab and investigation editor (#44479)
* feat(ui): move worker status into the Lens notch and New investigation into the list toolbar

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): make the Lens notch entry a settings gear that houses the worker section

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): replace the Lens worker modal with an inline Settings tab

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): section the Lens settings tab with tracing status and worker cards

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): let Lens settings sections span the full card width

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): replace the investigation setup modal with an inline side-by-side editor

New, edit, and duplicate now take over the Investigations tab body: matching
activity on the left, every setting on the right, with no wizard steps or modal

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): step the inline investigation setup vertically with traces alongside

Setup now sits on the left as three progressive steps (activity, criteria,
run) that collapse to a summary once done and reopen on click. Matching
activity stays on the right for every step. The editor gets a back control
and the Investigations notch shows a New, Editing, or Duplicate badge while
the editor is open

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): page the matching activity preview with useInfiniteQuery as it scrolls

Replace the Previous/Next offset buttons with the same infinite query and
near-tail prefetch the traces list uses, so the preview keeps loaded runs
and its title while the next page arrives

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): make the investigation step field map exhaustive over the form schema

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): always show the Settings tab label in the Lens mode switch

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): collapse Lens worker cards into compact status rows

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): keep the Lens Settings tab icon-only in every state

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* wip

* fix(ui): tick the Lens worker health dot so an expired heartbeat goes stale

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): regroup Lens settings, model and api layers

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): dedupe Lens formatting helpers

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): poll Lens once, drive the interval from data, settle mutations before invalidating

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): let Lens leaves fetch their own data

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): bind drawer and trace shortcuts through react-hotkeys-hook

One useShortcut hook replaces the three hand-rolled keydown listeners in SidePanel,
the trace step tree and the log drawer, with a layer option deciding which keys a
pane claims from the panel around it. The span tree footer now renders ShortcutHints
from what is actually bound instead of hand-typed kbd text.

* refactor(ui): extract Inspector from SidePanel

Inspector.Root owns the open item, J/K stepping, Escape and full screen;
Inspector.Row marks a list entry with aria-selected and data-state and toggles
it on click or Enter/Space; Inspector.Panel is the resizable side panel with the
exit animation and click-outside rules. The runs table and section compose these
parts directly, so RunDrawer and the SidePanel prop bag go away.

* refactor(ui): model the Lens worker screen as a tagged union and slot in its ready action

workerScreen() decides between list, form and install from the worker rows,
the registration result and the edit target, so WorkerSettings switches on
one value and each card owns its own copy. The post-install CTA is now a
ReactNode slot that LensWorkspace fills instead of an onReady callback
threaded through LensSettings and WorkerSettings. Clipboard copy state lives
in WorkerInstall as mutations, and the styled settings leaves export Props
types, set data-slot and accept native element props.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): compose the Lens setup stepper from SetupStep children

Each step's heading, summary and fields now live together in one
SetupStep instead of four parallel structures keyed by index, and the
last-step spacing comes from CSS rather than a passed index. The mode
prop is now required since InvestigationsView always passes it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): keep the Lens Settings panel mounted so a pending worker install survives tab switches

The settings TabsContent unmounted WorkerSettings whenever another tab was
active, dropping the one-time worker token shown during install. The panel
now uses keepMounted, and the workspace test registers a worker, switches
tabs and back, then follows the connected worker into the first
investigation through the slotted CTA.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ui): cover the Lens worker install waiting-to-connected transition

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): give the Lens activity preview a grouped contract and a structural debounce

useMatchingActivity now owns its return types (scope options, preview
status, page and optional manual selection) instead of borrowing them
from the components it feeds, and the preview takes those groups plus
the section attributes. The clear-selection action moves into the
preview footer, RunList becomes a RunRow leaf, and ScopeFields drops
the unused nameField and id props now that MetadataFilters calls useId
itself. The preview scope settles through a hashKey-based
useDebouncedValue instead of JSON round-tripping into state, and the
loading title follows isPlaceholderData since the query keeps previous
data. RunFields and AnalysisModelField take the analysis models and the
model gate as two objects instead of seven flat props.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): select the Lens investigations screen with a pure tagged union

investigationScreen maps the list query and the route to one of loading,
failed, welcome, list, detail, setup or missing, so the view can switch
instead of juggling mutually exclusive booleans. The status model gains
activeJob, carries connected inside Readiness and folds the activity probe
into one ActivityCheck value for the welcome page

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ui): cover the Lens preview footer clear action and the preview debounce

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): let Lens investigation leaves own their URL slice and express intent

InvestigationsView renders the screen union and owns every write through
useInvestigationActions, so leaves receive on* handlers instead of the API
writer. useInvestigationResults becomes useRunSnapshot; FindingsTab,
HistoryTab, RunPicker and RequestEvidenceSheet read their own nuqs slice
and run their own queries. The finding sheet becomes an Inspector side
panel (FindingDetails) keyed per finding, with the trace and request
evidence sheets grouped in EvidenceSheets. WatchAllBanner owns its
mutation, the welcome page takes the readiness and activity values, the
progress sampler records on the wall clock outside render, and run
history invalidates when the list reports a scheduler-started job

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ui): cover Lens investigation intents, pause, cancel, history refresh and the finding panel

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): keep the inspector open behind sheet overlays and mount HotkeysProvider from a client wrapper

* feat(ui): open Lens runs and evidence in the inspector panel

A quote's original trace step or logged request now stacks inside the finding
panel behind a back link, keeping the finding and its feedback draft mounted.
The detail Runs tab and the setup activity preview open runs in the same panel
with J/K stepping, so the TraceSheet and RequestEvidenceSheet modals are gone.
Picking a different finding or run clears any stacked evidence from the URL.

* feat(ui): open Lens investigations in the inspector panel beside the list

The investigations list stays on screen and a row opens its investigation in the
side panel, so J/K walk investigations and their open findings in display order
and the selected row carries the same highlight as runs. The panel body is the
former detail page; findings and runs opened inside it nest their own inspector,
which claims the keys from the one around it while open. Opening an investigation
and peeking at a finding now replace each other in the URL.

* fix(ui): run the Lens notch border along the tab pill and flag only a disconnected worker

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): show the shortcut hints in every inspector panel

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): extract the Lens dot field into composable DotFieldRoot and DotFieldCanvas

Move the dot grid model and canvas painter out of TracesTimeline into
components/lens/dotField so other Lens surfaces can reuse it. The root
owns layout and context; overlays compose as children. agoLabel moves to
lens/model/format and the unused columnTop helper is dropped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): drop the manual refresh button from the runs time controls

Live polls and range changes refetch, so the button only cleared the zoom, which Escape already does

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): extract the Lens run search into a composable SearchBox primitive

Move the query parser, glob matcher and autocomplete out of runSearch into
components/lens/search, generic over a QueryLanguage (field specs plus
what free text searches). SearchBox.Root owns the ProseMirror state, menu
and keyboard; SearchBox.Input and SearchBox.Suggestions compose under it.
Clause highlighting becomes a ProseMirror plugin built from the language.
RunSearch now only declares the run fields and composes the parts, so the
investigations tab can define its own language and reuse the same box.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(ui): label the runs range by preset while Live and pin it once paused

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): split the Lens search language from where its data lives

QueryLanguage is now pure vocabulary (keys, groups, icons). Reading
fields off loaded items moves to a ClientIndex consumed by a separate
evaluator, and value suggestions come from an injectable ValueSource, so
a server-backed runs list can plug in a facet lookup while the
investigations tab keeps filtering in memory. The parsed query serializes
to a typed SearchQuery (text terms plus eq/neq/glob/nglob filters) that
the client evaluator consumes today and a server can consume later. The
suggestion menu shows a loading row while a source is still answering.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* style(ui): format the Lens SearchBox and its test with the project prettier config

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): mirror the Lens run filters as trace SQL with a copyable curl in the search footer

The suggestions footer gains a slot, and SearchBox.ApiHint fills it with the
API equivalent of the typed query: a dialect chip, a one-line preview and a
Copy as curl button. The runs box translates each filter to a predicate over
the agent_traces_by_key rollup, bounded to the range the list shows, and
copies a POST to /v1/traces/query. Any other list can plug its own translate
into the same part.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): show the Lens introduction as a first-visit dialog with a typed don't-show-again

The guided setup no longer replaces the Lens tabs. It opens in a dialog on
the first visit of a session or from ?setup=lens, with a close and a
"Don't show this again" checkbox in its top-right corner. The header
"Set up Lens" button is gone. Dismissal state lives in a new schema-validated
web storage helper (src/lib/storage.ts) that reads through
useSyncExternalStore, so server renders see the fallback and other tabs stay
in sync; only Lens uses it for now.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): keep only Copy as curl in the Lens run search footer

Drop the SQL chip and predicate preview; SearchBox.ApiHint becomes
SearchBox.CopyCommand, which takes the command for the current query.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(ui): align the Lens investigations list with the traces list

Use the shared query SearchBox with investigation fields (name, agent, status, schedule), match the traces toolbar, and drop the count footer and inner padding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): let Lens settings bring back the introduction after don't show again

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): share one InspectorTable between the traces list and the Lens investigations tree

Compose TanStack Table, react-virtual and the shadcn table cells into InspectorTable parts (Root, Grid, Header, Body, Row, Indent). Investigations get findings as real sub-rows with TanStack expansion instead of a hand-rolled flattener.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): break Lens import cycles and move shared pieces out of lens

Search and the dot field go to components/shared, run search and the preview
button go to view_logs where they are consumed. Lens api, services and demo
live under data/, all URL state in route.ts, storage keys in storage.ts, and
the session frame styles become cva variants.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): one Lens readiness source and one onboarding flow

Readiness is computed once in model/readiness and read through
useLensReadiness, replacing useLensSetup, status.readiness and the welcome
screen's own checks. The Investigations welcome now renders the same
onboarding steps as the introduction dialog, with permissions and actions
coming from an OnboardingProvider instead of props passed down four levels.
StepIndicator and StateMessage are shared lens components, and the step
panels are labelled accordion regions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): read the Lens token from services and use semantic status colors

Lens services carry the access token they were built for, so trace evidence,
readiness and onboarding read it from context instead of a prop threaded
through six components. List and history invalidation lives in one data
hook. Status colors use the success, warning and destructive tokens, and
template-literal class names go through cn.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): split Lens demo fixtures from the fake demo APIs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): show Lens check history as a dot timeline

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chores

* fix(ui): clear stale Lens evidence on run change and keep read-only users off Settings

Also names inline option objects to bring local/no-large-inline-object-arg back under budget.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 07:20:14 +00:00
moe-berri
2936307661
feat(lens): guide setup through the first investigation (#44475)
* feat(lens): guide setup through the first investigation

* fix(lens): restore the onboarding reference visuals

* fix(lens): compact onboarding and animate gateway flow

* feat(lens): refine onboarding motion and linked examples

* feat(lens): turn the LED swarm into organized dot groups

* fix(lens): make the LED dot flow visibly animate

* feat(lens): refine the swarm scale palette and motion

* fix(lens): preserve investigations during activity refresh errors
2026-10-03 20:16:41 -07:00
devin-ai-integration[bot]
5ddcc45b3a
feat(ui): share trace drawer as a closable SidePanel and polish Lens (#44473)
Some checks failed
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / caching-local (push) Waiting to run
Unit Tests / core-utils (push) Waiting to run
Unit Tests / enterprise-package (push) Waiting to run
Unit Tests / enterprise-routing (push) Waiting to run
Unit Tests / integrations (push) Waiting to run
Unit Tests / All Other Providers (push) Waiting to run
Unit Tests / Vertex AI (push) Waiting to run
Unit Tests / mcp-integration (push) Waiting to run
Unit Tests / misc (push) Waiting to run
Unit Tests / proxy-auth (push) Waiting to run
Unit Tests / proxy-endpoints (push) Waiting to run
Unit Tests / proxy-extras (push) Waiting to run
Unit Tests / proxy-infra (push) Waiting to run
Unit Tests / proxy-infra-root (push) Waiting to run
Unit Tests / proxy-server (push) Waiting to run
Unit Tests / responses-caching-types (push) Waiting to run
GitHub Actions Security Analysis / zizmor (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
* feat(ui): extract trace drawer into a shared SidePanel that closes on outside press

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): center Lens mode switch in a notch joined to the content card

Larger Traces/Investigations switch, a subtle dot when an investigation is running or queued, and a bigger Lens title with a docs link.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): stronger Lens frame border and header spacing

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): align Lens notch fillet with the notch border

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): centered Lens loading/error states and proxy JSON calls in dev

The dev server answered GET /lens with the Lens page HTML because the UI route shadowed the proxy fallback rewrite. JSON API requests now go to the proxy before page routes, and a non-JSON success body raises a readable ApiError.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): theme-scale type and shape in trace views, aligned pane bars

Add local/no-arbitrary-design-value, scoped to TraceView and Lens, banning
arbitrary font size, tracking, leading, radius, border and CSS property
values. Map the Figma-export values onto the theme scale and replace hex
colors with info/destructive tokens.

Add PaneBar, a fixed-height bordered row, and build the step tree and span
detail headers from it so their borders line up across the split.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): restore AgentTracesSection emptied in e7c3092571

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): move Set up tracing into the empty runs state

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): let the runs table gate the setup CTA on an empty range

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): mark Lens demo mode with a blue toggle and frame instead of a banner

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): draw Lens notch corners with CSS borders and thicken the demo frame

The SVG corner strokes did not snap to the same device pixels as the tab and
card borders, leaving a visible offset at the join.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): drop the redundant Lens timeline header

The status, run counts, truncated agent legend and range span all repeated the Live toggle, runs table and range picker

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): keep drawer shortcuts out of open menus and use usehooks-ts for timers and observers

SidePanel J/K/Esc now yields to menus and listboxes, not just dialogs.
The step tree shortcut footer wraps instead of clipping in narrow columns.
Replace hand-rolled timeout, keydown, media query and ResizeObserver effects
with usehooks-ts, and drop routine doc comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): offer tracing setup when filters hide every run

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): replace Lens runs filters with a single ProseMirror query box

Agent and status dropdowns are gone. One query box (react-prosemirror) takes
free text plus key:value clauses (-key:value, key:*glob*) over name, agent,
status, model, input and trace_id, with field and value autocomplete. The
editor emits after a 150ms pause so typing no longer re-renders the runs view
per keystroke, and the URL keeps only q.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): keep the open trace on screen while the next one loads

Switching to an unvisited trace remounted the panel and flashed a loading skeleton.
The drawer now keeps the previous trace visible, dimmed and inert, until the new one arrives.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): virtualize the Lens runs table and drop its status footer

The footer's run count only tracked how many pages had loaded and
"Updated just now" never changed, so it carried no signal. The zoom
clear button moves onto the timeline. Rows now render through
@tanstack/react-virtual so scrolling deep into a range stays cheap

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): load traces with Suspense and show the previous run via useDeferredValue

Replaces keepPreviousData with the React pattern for showing stale content while fresh content loads.
Load failures go through an error boundary that retries the query on reset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ui): open investigation details from list rows instead of the edit dialog

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 02:20:14 +00:00
devin-ai-integration[bot]
c80e6e2474
test(ui): wait for step search value to settle in TraceDrawer test (#44470)
* test(ui): wait for step search value to settle in TraceDrawer test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): dispatch search typing and nav keys deterministically in TraceDrawer test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 18:25:01 -07:00
devin-ai-integration[bot]
62fb808d4b
feat: improve Lens dev seeding and live UI (#44468)
* feat: seed Lens dev with configurable load profiles

* feat: run Lens UI live through the dev launcher

* fix: verify Lens UI startup before seeding

* fix(ui): render run timestamps on one compact line

The agent runs table printed the long locale form with timezone, which
wrapped to two lines per row. Use a fixed-width 24h form with
milliseconds that matches the timeline axis, keep the long form in the
hover title, and show the timezone once in the column header

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): use the root query client for the Lens demo

Drop the demo's nested QueryClient. Cache keys are already partitioned by
scope, and the root client now skips retries on 4xx ApiErrors, which covers
the demo's not-in-demo and read-only rejections and live 4xx alike.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: drop unused synthetic_spend reference from query help

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(traces): drop the deeplite fixtures and map every SDK to a logo

The deeplite captures predate the example repository and carried
synthetic spend rows, which leaked a fixture-only column into the
query help SQL and pinned tests to its shape. Replace them with the
SDK captures in the ClickHouse round trip and query API tests, and
derive the seed tenant lookup from the capture metadata

The runs table only knew the two Anthropic framework slugs. Register
the slugs the normalizer emits for LangChain, LangGraph, Deep Agents,
CrewAI, Google ADK, LlamaIndex, OpenAI Agents, Pydantic AI, Strands,
Vercel AI SDK, Codex, Cursor and Copilot, and fall back to the generic
agent glyph when a run has no known framework

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(ui): inject the Lens sample through data sources instead of demo checks

Components no longer ask whether they are in the demo. The traces source
carries live and handoff, navigation state comes from a URL or memory
route, and the preview action comes from context instead of onDemo props

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): infinite scroll for the agent runs list

Replace the Load more button with a sentinel that fetches the next
cursor page as the list nears its end. Placeholder rows hold the tail
while more runs exist, and a failed page stops auto-loading until Retry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ui): keep the whole Lens view in the URL

Lens navigation now lives entirely in query params through nuqs: the
sample session (demo=true, with a Demo data switch in the header), the
open run (trace, trace_ref), the selected step, view and detail section
(span, view, span_tab) and the list filters and range (q, agent, status,
hours). Any Lens view is a shareable link and the back button walks runs

RunView takes its selection injected: the drawer feeds it URL state and
the investigations evidence sheet keeps a local one, so a finding's
original run never writes step ids into the URL. Leaving the sample
session clears every Lens key except the tab so sample ids never point
at live data

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* style(traces): cargo fmt captures tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-04 01:03:11 +00:00
tin-berri
6370104c53
feat(enterprise): bundle LiteAdmin Slack with native gateway login (#44444)
* feat(enterprise): bundle LiteAdmin Slack worker with native gateway login

* fix(enterprise): preserve gateway prefixes during native Slack linking

* fix(enterprise): retain native Slack linking on the admin backend

* fix(enterprise): reuse shared native Slack connection services
2026-10-03 17:28:18 -07:00
moe-berri
4b67a2b845
feat(roi): default people and branch lists to matched accounts (#44465)
* feat(roi): show matched people by default in contributor lists

* fix(roi): keep matched filter tabs readable on narrow screens

* fix(roi): retain spend-only users and support older browsers
2026-10-03 17:24:18 -07:00
devin-ai-integration[bot]
0ed1c08f02
feat(anthropic): workload identity federation and pluggable identity sources (#44448)
* feat(anthropic): workload identity federation and pluggable identity sources

Backend half of #38818 (internal copy of the fork PR #38013), rebuilt as one
commit on top of litellm_internal_staging without the dashboard changes.

Deployments on anthropic/ without a static api_key can exchange an OIDC
workload assertion for a short-lived sk-ant-oat01 token through a shared
RFC 7523 JWT-bearer engine. The assertion comes from a mounted token file,
an env token, a LiteLLM-signed issuer, or Keycloak, chosen per deployment,
per named credential, or through ANTHROPIC_IDENTITY_SOURCE. The federation
fields are server-owned: refused inline in request bodies and on
POST /model/new, proxy-admin only on credentials, and the token exchange
is pinned to api.anthropic.com unless LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS
adds a host. GET /credentials/{name}/jwks exports the public key set of a
LiteLLM-signed credential for the Claude Console.

The OpenAI federation trio from #39613 rides along on the backend side with
the same server-owned handling.

Fixes #28607
Resolves LIT-6107

Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>

* fix(anthropic): let batch-result downloads mint from deployment params and accept host:port allowlist entries

The files handler enabled workload identity on batch-result downloads but never received the
deployment's litellm_params, so a deployment authenticating through a named credential could only
mint from process-wide env vars. It now threads litellm_params through to the auth header the way
the batch retrieve path already does.

LITELLM_ANTHROPIC_WIF_ALLOWED_HOSTS entries written as host:port were read by urlsplit as a scheme,
so the allowlist kept the raw entry while the exchange compared bare hostnames and refused the
gateway. Entries are now parsed as network locations whether or not they carry a scheme.

* fix(types): move the WIF kwargs key sets to a leaf module so the kwargs funnel imports without a cycle

* test(anthropic): pin case-insensitive matching of WIF exchange-host allowlist entries

* fix(anthropic): end workload identity federation errors without a period so the router suffix reads cleanly

* fix(proxy): decrypt stored litellm_params before the WIF write gate

* fix(proxy): hide WIF secret references from /health output

* fix(proxy): keep the proxy error shape on credential endpoint refusals

* fix(proxy): hide identity token file paths from /health output

* fix(anthropic): rename the federation workspace param so Bedrock's anthropic_workspace_id keeps working

The Bedrock Claude Platform route already reads anthropic_workspace_id from
optional_params, so banning that spelling as a server-owned federation
parameter broke a pre-existing client capability. The federation field is now
anthropic_federation_workspace_id (env ANTHROPIC_FEDERATION_WORKSPACE_ID),
which restores the base branch's behavior for Bedrock callers, drops the
Bedrock-specific hint from the refusal message, and deletes the unconditional
ban constant that no longer had a reader

* fix(auth): share one exchanged token across workers reading the same assertion

Anthropic accepts each identity assertion exactly once, so two uvicorn
workers reading the same token file both minting from it means the second
exchange is denied with jti_reused. Minted tokens now land in a per-user
0700 cache directory guarded by a file lock, so workers on the same host
reuse one exchange until the token expires or the assertion rotates. A 401
is only retried when the re-read assertion actually differs, and the denial
hint explains jti_reused. LITELLM_TOKEN_EXCHANGE_CACHE_DIR moves the cache
and an empty value disables it

* fix: keep anthropic federation from being shadowed or leaked

An empty or whitespace-only ANTHROPIC_API_KEY counted as set, so a federated
deployment sent an empty x-api-key on every call instead of minting a token.
Blank values now read as unset, and a real static key on a federated deployment
logs once that it outranks federation and nothing is being federated.

The exchange-host allowlist matched hostnames only, so a second process on
another port of an allowed host was trusted with the workload's identity token.
An entry that names a port now trusts that port alone, while a bare host still
trusts every port.

The shared token store exists so the workers reading one projected token file do
not each spend its single-use jti. A source that mints its own assertion per
exchange shares nothing with another worker, so it no longer writes a live token
to disk for a lookup that can never hit.

* fix: unlink a staged token file a failed write leaves behind

The 401 denial hint now also says federation ignores ANTHROPIC_WORKSPACE_ID, which the Bedrock Claude platform provider already reads.

* refactor: move anthropic jwks derivation behind a provider-owned tagged union

* fix: unlink the staged token file when its write fails at close

A buffered write only reaches the disk when the handle closes, so a full disk surfaces at close and left the staging file behind holding a usable token.

* fix(anthropic): close the staging descriptor before writing the shared token file

* fix(wif): judge federation writes by what they set, not what is stored

The admin gate read the stored deployment, so a team admin lost edit, delete
and Test Connection on any deployment carrying federation params. It now
returns early unless the submitted fields touch the federation surface, and a
Test Connection probe that points the deployment at its own api_base is still
refused, with the 403 no longer wrapped into a 500

The rest of the same review pass: POST /model/new refuses only a blocking
value of `blocked`, so a client that always sends `blocked: false` is not
turned away; a request body can no longer pick which federated identity to
mint as by naming a stored credential; an advisory refresh the executor
refuses disarms the entry instead of wedging the identity until the follower
timeout; the static-key shadow warning resolves its env fallback inside the
cache instead of once per request; credential writes drop nulls before
storing them; the token exchange validates the endpoint URL before reading an
assertion and keeps refusing redirects across a client heal; /health hides
every server-owned federation field from non-admins; and the async create_file
and create_batch paths say which setting is missing when the provider resolves
no URL

* fix(proxy): let a deployment write name a federated credential

reject_federated_credential_reference runs from is_request_body_safe, which
pre_db_read_auth_checks calls on every route, so it also fired on POST
/model/new, /model/update, /model/{id}/update and /health/test_connection. A
proxy admin could no longer attach a federated credential to a deployment over
the API or the Admin UI, leaving a static config.yaml entry as the only way to
configure the feature the rejection told the caller to go configure, and
_reject_non_admin_wif_write never got to make the call it exists to make.

is_request_body_safe now takes the route and skips only the credential-reference
check on the routes that reach can_user_make_model_call. Federation fields typed
inline into a body stay refused everywhere, and a call naming a federated
credential still cannot pick the identity it mints as.

* refactor(proxy): derive health display policy from the federation key sets

The health check module hand-copied the five workload identity fields whose
value is a credential, so a shared proxy surface named provider-specific
parameters and a newly added secret-bearing field would have gone on being
displayed until someone remembered both places

WIF_SECRET_BEARING_KEYS now sits beside the key sets it splits out of,
types/utils derives secret_bearing_wif_litellm_params from it, and the health
layer splats that tuple the same way it already splats the admin-only one

* fix(anthropic_wif): treat blank identity-source fields as unset

* test(proxy): classify the federation params in the credential slot registry

main's registry test (#43298) now fails the build for any credential-named
deployment param without a classification. The five federation fields that
carry a token, a token file path, or a signing or client secret reference are
Unplanted, matching WIF_SECRET_BEARING_KEYS; the four remaining Keycloak
settings name a URL, a client id, an auth method, or a scope and are NotSecret

* fix(anthropic_wif): declare federation params as owned connection leaves and chart their metrics

Register the 18 Anthropic and 3 OpenAI federation params as frozen
ConnectionSettings leaves so the owned-kwarg registry, the kwargs funnel
and the request-body ban list read one declaration. Pass the deployment
api_base through to the count-tokens handler instead of a pre-suffixed
URL, which doubled the /count_tokens path on main's prompt-cache
predictor. Add the five litellm_anthropic_wif_* families to the
all-metrics Grafana dashboard.

* fix(credentials): gate PATCH on WIF fields resolved from model_id

The credential PATCH handler checked server-owned workload identity
federation fields only on the values the caller sent, while a body that
named a deployment through model_id had its credential values resolved
after that check. A non-admin could therefore copy a federated
deployment's WIF fields onto an ordinary credential. Resolve the incoming
values first and run the non-admin gate on them, matching the POST path

* fix(anthropic): count tokens with ANTHROPIC_AUTH_TOKEN through the shared auth header

Count-tokens walked its own credential ladder: a static key, else skip minting when
ANTHROPIC_AUTH_TOKEN is set, else mint a federated token. With only the auth token set it
forwarded nothing and the proxy silently fell back to its local tokenizer while chat on the
same deployment authenticated with that token. The handler now takes the auth header that
AnthropicModelInfo.aget_auth_header resolves, the same ladder chat, files, batches and skills
use, and merges the oauth beta a minted or consumer token carries with the token-counting beta

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: derhornspieler <15236687+derhornspieler@users.noreply.github.com>
Co-authored-by: mateo-berri <happymvw@gmail.com>
2026-10-03 17:08:30 -07:00
devin-ai-integration[bot]
f0eda6d2a6
fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API (#44419)
* fix(health): probe Bedrock Mantle Claude deployments over the Anthropic Messages API

Bedrock Mantle serves Claude ids only on /anthropic/v1/messages, but health
checks probed every chat-mode deployment over /v1/chat/completions, so a
bedrock_mantle Claude deployment showed unhealthy while real /v1/messages
traffic to it succeeded

Add an anthropic_messages health check mode and make it the default for
bedrock_mantle Claude models. An explicit model_info.mode still wins, and
/health/test_connection and the Add Model form accept the new mode

* fix(health): resolve the test connection mode from the deployment when the request omits it

The Admin UI model page sent the mode /model/info had filled in from the cost
map back as the probe mode, so Test Connection on a Bedrock Mantle Claude
deployment still went over chat completions. The page now forwards only the
row's id, and /health/test_connection resolves a missing mode the way /health
does: the stored model_info.mode, then the mode the provider requires, then the
cost map.

* fix(health): resolve an omitted ahealth_check mode the way the proxy does

* fix(health): test connection honors a stored mode only for the stored model and rejects a non-string mode

A request that selects a stored deployment and sends a different litellm_params.model now resolves the probe mode from that model instead of the stored model_info.mode. A litellm_params.mode that is not a string answers 400 instead of 500. The Bedrock Mantle rule that Claude models are probed over the Messages API moves into the provider package.

* fix(health): shape test connection probe params for the model the request probes

A request that selects a stored deployment by id and overrides the model
resolved its probe mode from the overridden model but still injected
max_tokens from the stored mode, so an embedding override of an
anthropic_messages deployment failed with a Mistral 422 extra_forbidden

* fix(health): report an early ahealth_check failure as itself, not as a missing mode

With the mode resolved automatically when the caller omits it, a failure
before that resolution (no model, a non-string model, a provider that does
not resolve) was wrapped as "Missing mode", a hint that pointed at the wrong
fix and dropped raw_request_typed_dict from the result. Every failure now
returns the same shape.

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:05:47 -07:00
moe-berri
cb17588276
feat(lens): coordinate worker releases and bundled installs (#44428)
* feat(lens): coordinate worker versions and bundled installs

* test(lens): exercise bundled Compose startup and restart in CI

* fix(lens): refund failed model requests without a response

* test(lens): verify trace persistence in the bundled stack

* fix(lens): align Helm images and isolate Compose storage

* fix(lens): reject worker builds without release identity

* fix(lens): encode Compose credentials and normalize worker versions

* fix(lens): refuse worker recommendations for unidentified builds
2026-10-03 16:33:45 -07:00
yuneng-jiang
427158eb5b
feat(ui): show invitation and reset password links in a copyable field (#44454)
* feat(ui): show invitation and reset password links in a copyable field

Put the link in a read-only input with a Copy button beside it, stack the
User ID and link labels above their values, and focus Copy on open so the
field shows the start of the URL. Copy now goes through the shared
copyToClipboard helper, which falls back to a selection copy where the
Clipboard API is unavailable.

* fix(ui): keep focus on the copy control after a fallback clipboard copy

The execCommand fallback focused a temporary textarea and removed it, so
focus fell to the page body and a second Enter on Copy did nothing.
Restore focus to the element that had it. Move the rendered dialog
tests to the integration tier.
2026-10-03 16:23:13 -07:00
moe-berri
1a7023366f
feat(roi): measure shipping velocity, quality, and recorded spend (#44426)
* feat(ui): prototype observed engineering ROI dashboard

* feat(roi): replace effort estimates with measured repository metrics

* fix(roi): finish connection recovery and generated API contracts

* fix(roi): show merged changes before accounts are linked

* fix(roi): preserve selected report tab across refreshes

* fix(roi): recover app authorization and keep detail values readable

* fix(roi): reuse the shared OAuth HTTP client

* feat(roi): combine providers and compare equal reporting periods

* docs: explain ROI metrics for first-time readers

* fix(roi): preserve connections and scheduled reports during setup

* ci(roi): assign database contracts to the active Postgres shard

* fix(roi): preserve issue counts and normalized connections

* fix(ui): compact ROI dashboard header and metrics

* fix(ui): show ROI repository count with expandable list

* fix(ui): wrap ROI controls within narrow panels

* fix(roi): restore sample report preview and simplify setup
2026-10-03 23:07:34 +00:00
devin-ai-integration[bot]
fe683ea139
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models

GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.

* fix(proxy): offer a Codex service tier only when every deployment of the model lists it

* fix(codex-catalog): an invalid service_tiers value offers no tier for the model

* fix(codex-catalog): read service tiers off the deployments the key's team can route to

A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them

The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default

* test(codex-catalog): drop the redundant module docstring and sort the imports

* test(integration): add the Codex catalog audit cells and the multi-worker convergence note

* test(integration): clean up every catalog test model and answer the refresh GET

* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers

Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.

* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut

The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns

The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 20:42:55 +00:00
ishaan-berri
dd31692282
feat(lens): always-on investigations with findings and investigations tables (#44418)
* feat(lens): record run steps, trigger and exact run windows on investigations

* feat(lens): scan only traces since the last run and keep a capped step log

* feat(lens): log each analysis model call with its model, tokens and cost

* feat(lens): accept agent and time window on run now and add turn-all-on

* test(lens): cover new-traces-only windows, manual runs and the step cap

* test(lens): cover run now overrides and turning paused investigations on

* chore(ui): regenerate api types for lens run steps and run now options

* feat(lens): group open findings into one row per problem with agent filters

* test(lens): cover the findings table grouping, filters and schedule labels

* feat(lens): add a findings table across all investigations

* feat(lens): show a live step feed with the model behind each call

* feat(lens): offer turning all paused investigations on

* feat(lens): let run now pick an agent and time window

* test(lens): cover run now request building

* feat(lens): show the step feed and run now dialog on an investigation

* feat(lens): open run now choices instead of running immediately

* feat(lens): open findings first and peek a finding without leaving the table

* feat(lens): fold investigation actions into the findings toolbar

* feat(lens): show each investigation's schedule and open findings

* feat(lens): name the agent on a finding

* feat(lens): keep new investigations watching every 15 minutes by default

* feat(lens): show the watch schedule outside advanced options

* feat(lens): send run now options and turn-all-on from the dashboard

* test(lens): give demo runs steps and a trigger

* feat(lens): let the findings table fill the screen

* test(lens): add steps and trigger to progress fixtures

* test(lens): add steps and trigger to status fixtures

* test(lens): cover the default watch schedule in setup

* test(lens): cover run now choices from an investigation

* test(lens): open saved investigations from the manage view

* feat(lens): use one tab bar for traces, findings and investigations

* feat(lens): place page actions on the lens tab row

* feat(lens): drop the nested tabs and edit investigations in place

* feat(lens): show investigations as a table with run now and edit

* feat(lens): name each findings row for screen readers

* test(lens): open saved investigation links on findings

* test(lens): reach findings and investigations from the top tabs

* fix(lens): mark run now jobs manual and keep them from moving the scheduled scan

* fix(lens): keep run now since-last-run windows even with an agent override

* test(lens): cover that manual runs never skip scheduled traces

* test(lens): cover run now windows with agent and lookback overrides

* fix(lens): group findings without Map.groupBy and expose sampled runs

* fix(lens): open older findings and review every merged copy from one row

* fix(lens): hide edit and run now from read-only viewers

* chore(lens): drop restating comments from the findings table

* chore(lens): drop restating comments from the step feed

* chore(lens): drop restating comments from the paused banner

* chore(lens): drop restating comments from header actions

* chore(lens): drop restating comments from run now

* test(lens): cover merged findings and read-only investigation rows

* fix(lens): record a model step even when the response has no usage

* test(lens): cover model steps with and without reported usage

* fix(lens): keep a merged finding open when one of its updates fails

* refactor(lens): accept update results from the investigations view

* refactor(lens): accept update results in investigation actions

* test(lens): cover retrying a merged finding after a failed update
2026-10-03 20:27:02 +00:00
moe-berri
af36e5c693
fix(lens): preserve full trace access and expose investigation failures (#44406)
* fix(lens): preserve full trace access and expose investigation failures

* chore(lens): sync worker registration schema

* fix(lens): support durations without a configured maximum

* fix(lens): expose every page of fetched investigation evidence

* fix(lens): preserve repeated trace content and interrupt cancelled runs

* fix(lens): retry transient heartbeat failures during analysis
2026-10-03 11:52:55 -07:00
moe-berri
50190134c3
fix(lens): batch run reads and reset trace pagination (#44398)
* fix(lens): batch run reads and reset trace pagination

* fix(lens): scope batched list spend to each run
2026-10-03 18:43:23 +00:00
tin-berri
4732647de2
feat(ui): make LiteAdmin enterprise-only (#44399) 2026-10-03 11:07:54 -07:00