Commit graph

6057 commits

Author SHA1 Message Date
Ishaan Jaff
09ac1663b8
feat(lens): add a live trace results drawer with readable spans 2026-10-03 17:04:40 -07:00
Ishaan Jaff
b4ef33b7a1
feat(lens): track active jobs before their first review 2026-10-03 17:04:37 -07:00
Ishaan Jaff
d16b0d7132
feat(lens): derive strip status, honest issue counts and drawer focus from a job 2026-10-03 17:00:31 -07:00
Ishaan Jaff
af1c2f1a95
feat(lens): format review span previews as readable messages 2026-10-03 17:00:29 -07:00
Ishaan Jaff
616c59c8d0
fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size 2026-10-03 16:49:55 -07:00
Ishaan Jaff
1fb70ff6be
feat(lens): collapse the live run to an ambient line with show work 2026-10-03 16:49:53 -07:00
Ishaan Jaff
d2e96c41f2
feat(lens): add a reading ticker line and replay for finished runs 2026-10-03 16:49:50 -07:00
Ishaan Jaff
ac4d11abbc
refactor(lens): restyle the live run as the native progress panel 2026-10-03 16:44:49 -07:00
Ishaan Jaff
25c417d8ba
fix(lens): show the live run only for real reviews and keep fixtures test-only 2026-10-03 16:44:45 -07:00
Ishaan Jaff
cac908783a
feat(lens): stream large review backlogs at 150ms or less and list newest first 2026-10-03 16:42:51 -07:00
Ishaan Jaff
d77bdf8fa4
feat(lens): show the live run on investigations and open it from run now 2026-10-03 16:36:40 -07:00
Ishaan Jaff
97225a7143
feat(lens): add live run layout with queue, reading trace and conclusions 2026-10-03 16:36:38 -07:00
Ishaan Jaff
5a68f71bfa
feat(lens): add sample reviews for previewing the live run 2026-10-03 16:36:34 -07:00
Ishaan Jaff
53275517cb
feat(lens): pick the analysis model and slow single-review pacing 2026-10-03 16:36:32 -07:00
Ishaan Jaff
e8f49b18ca
feat(lens): add live review playback model 2026-10-03 16:23:29 -07:00
Ishaan Jaff
a3bd0a6613
feat(lens): type job reviews and fill them in lens fixtures 2026-10-03 16:23:28 -07:00
Ishaan Jaff
010d046404
chore(ui): regenerate api types for lens job reviews 2026-10-03 16:23:22 -07:00
devin-ai-integration[bot]
fe683ea139
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models

GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.

* fix(proxy): offer a Codex service tier only when every deployment of the model lists it

* fix(codex-catalog): an invalid service_tiers value offers no tier for the model

* fix(codex-catalog): read service tiers off the deployments the key's team can route to

A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them

The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default

* test(codex-catalog): drop the redundant module docstring and sort the imports

* test(integration): add the Codex catalog audit cells and the multi-worker convergence note

* test(integration): clean up every catalog test model and answer the refresh GET

* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers

Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.

* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut

The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns

The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 20:42:55 +00:00
ishaan-berri
dd31692282
feat(lens): always-on investigations with findings and investigations tables (#44418)
* feat(lens): record run steps, trigger and exact run windows on investigations

* feat(lens): scan only traces since the last run and keep a capped step log

* feat(lens): log each analysis model call with its model, tokens and cost

* feat(lens): accept agent and time window on run now and add turn-all-on

* test(lens): cover new-traces-only windows, manual runs and the step cap

* test(lens): cover run now overrides and turning paused investigations on

* chore(ui): regenerate api types for lens run steps and run now options

* feat(lens): group open findings into one row per problem with agent filters

* test(lens): cover the findings table grouping, filters and schedule labels

* feat(lens): add a findings table across all investigations

* feat(lens): show a live step feed with the model behind each call

* feat(lens): offer turning all paused investigations on

* feat(lens): let run now pick an agent and time window

* test(lens): cover run now request building

* feat(lens): show the step feed and run now dialog on an investigation

* feat(lens): open run now choices instead of running immediately

* feat(lens): open findings first and peek a finding without leaving the table

* feat(lens): fold investigation actions into the findings toolbar

* feat(lens): show each investigation's schedule and open findings

* feat(lens): name the agent on a finding

* feat(lens): keep new investigations watching every 15 minutes by default

* feat(lens): show the watch schedule outside advanced options

* feat(lens): send run now options and turn-all-on from the dashboard

* test(lens): give demo runs steps and a trigger

* feat(lens): let the findings table fill the screen

* test(lens): add steps and trigger to progress fixtures

* test(lens): add steps and trigger to status fixtures

* test(lens): cover the default watch schedule in setup

* test(lens): cover run now choices from an investigation

* test(lens): open saved investigations from the manage view

* feat(lens): use one tab bar for traces, findings and investigations

* feat(lens): place page actions on the lens tab row

* feat(lens): drop the nested tabs and edit investigations in place

* feat(lens): show investigations as a table with run now and edit

* feat(lens): name each findings row for screen readers

* test(lens): open saved investigation links on findings

* test(lens): reach findings and investigations from the top tabs

* fix(lens): mark run now jobs manual and keep them from moving the scheduled scan

* fix(lens): keep run now since-last-run windows even with an agent override

* test(lens): cover that manual runs never skip scheduled traces

* test(lens): cover run now windows with agent and lookback overrides

* fix(lens): group findings without Map.groupBy and expose sampled runs

* fix(lens): open older findings and review every merged copy from one row

* fix(lens): hide edit and run now from read-only viewers

* chore(lens): drop restating comments from the findings table

* chore(lens): drop restating comments from the step feed

* chore(lens): drop restating comments from the paused banner

* chore(lens): drop restating comments from header actions

* chore(lens): drop restating comments from run now

* test(lens): cover merged findings and read-only investigation rows

* fix(lens): record a model step even when the response has no usage

* test(lens): cover model steps with and without reported usage

* fix(lens): keep a merged finding open when one of its updates fails

* refactor(lens): accept update results from the investigations view

* refactor(lens): accept update results in investigation actions

* test(lens): cover retrying a merged finding after a failed update
2026-10-03 20:27:02 +00:00
moe-berri
af36e5c693
fix(lens): preserve full trace access and expose investigation failures (#44406)
* fix(lens): preserve full trace access and expose investigation failures

* chore(lens): sync worker registration schema

* fix(lens): support durations without a configured maximum

* fix(lens): expose every page of fetched investigation evidence

* fix(lens): preserve repeated trace content and interrupt cancelled runs

* fix(lens): retry transient heartbeat failures during analysis
2026-10-03 11:52:55 -07:00
moe-berri
50190134c3
fix(lens): batch run reads and reset trace pagination (#44398)
* fix(lens): batch run reads and reset trace pagination

* fix(lens): scope batched list spend to each run
2026-10-03 18:43:23 +00:00
tin-berri
4732647de2
feat(ui): make LiteAdmin enterprise-only (#44399) 2026-10-03 11:07:54 -07:00
ishaan-berri
f0abe1bea1
feat(lens): live dot field timeline and full-screen traces view (#44390)
* feat(lens): add dot field layout and live status helpers for the traces timeline

* test(lens): cover dot field layout, agent colors and live status

* feat(lens): draw the traces timeline as a live dot field

* feat(lens): add the sweep animation for the live traces timeline

* feat(lens): put the lens tabs in a compact header and fill the screen with traces
2026-10-03 10:56:48 -07:00
moe-berri
ad8babae33
fix(lens): paginate trace reads within ClickHouse limits (#44384) 2026-10-03 10:38:40 -07:00
devin-ai-integration[bot]
8b1990b4bc
feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers (#44236)
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register typesafe as a provider so Jev deployments load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move provider endpoints under llms and validate proxy bodies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add Cloudflare Clef and Strands Decider backends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register decisions routes for managed agents and gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(decisions): use raw regex for cloudflare missing account match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): avoid cast in Cloudflare response unwrapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): default model, evaluation health probe, short Cloudflare names

The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.

* fix(decisions): let health_check_params override the evaluation probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit the decisions endpoint across providers, limits, health and chaos

Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).

The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.

* fix(decisions): send env API keys to a configured api_base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register the routes through the lazy feature registry

The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.

The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.

* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell

* fix(proxy): let a config pass-through beat a lazily registered route in eager mode

With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:38:38 +00:00
devin-ai-integration[bot]
e340e546e2
feat(traces): tracing development seed (#44363)
* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* wip

* wip

* wip

* chore(trace): checkpoint ongoing Rust migration

* refactor(trace): group Python bridge under trace package

* refactor(traces): read span conventions through a Convention trait

Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.

The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.

* feat(trace): export Rust-owned wire schemas and enforce contract bounds

* fix(trace): bound quoted counts in ClickHouse wire schemas

* feat(trace): generate Python wire contracts with datamodel-code-generator

* test(trace): validate migrated callers and generated contracts at the native boundary

* refactor(traces): rename normalization convention to format

* fix(traces): reconcile spend evidence and preserve unknown costs

* feat(traces): normalize additional telemetry formats

* test(traces): cover captured normalization fixtures

* refactor(traces): isolate SDK normalization rules

* feat(tracing): seed all trace exports for local dashboard

* fix(clickhouse): preserve custom LiteLLM request metadata

* docs(traces): define normalization module boundaries

* docs(traces): define resolution and OTLP boundaries

* fix(ui): normalize nullable trace message names

* refactor(traces): split resolver modules and cover resolution behavior

* test(traces): replace normalization snapshots with behavior assertions

* fix(ui): align dashboard API contracts with generated types

* refactor(traces): type normalization and storage boundaries

* fix(traces): seed captured SDK spend and preserve provider identities

* wip

* test(traces): verify guide discovery and content ordering

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:59:51 +00:00
devin-ai-integration[bot]
e92bd50de1
refactor(ui): inject Lens backends as services instead of a fake HTTP client (#44369)
The interactive Lens demo used to build a fetch shim that encoded in-memory
fixtures as HTTP responses so the shared ApiClient could decode them again,
and every trace view branched on demo vs live to pick a URL. Lens and the
trace views now depend on two small service interfaces, LensApi and
TracesApi, with named operations. The live layer wraps the existing HTTP
calls, the demo layer reads fixtures directly, and a React context provides
whichever one the session runs on. Without a provider the hooks fall back to
the live implementation, so the live app and existing tests are unchanged.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 09:11:08 +00:00
devin-ai-integration[bot]
09313bcf1b
refactor(ui): reorganize Lens dashboard components (#44354)
* refactor(ui): reorganize Lens dashboard components

* refactor(ui): align Lens forms with react-hook-form

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep Lens submit errors out of form validity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): reduce Lens lint budget usage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): use a query key factory for Lens queries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 08:29:16 +00:00
moe-berri
cb00dbecd7
feat(ui): improve trace inspection and ROI estimation (#44351)
* feat(lens): simplify trace inspection in the gateway drawer

* feat(lens): add a conversation view for full traces

* test(lens): keep normalized message fixtures type safe

* refactor(lens): make conversation view read like a chat

* refactor(lens): use a quiet trace view menu

* refactor(lens): make trace view tabs explicit

* refactor(lens): restore compact trace view switch

* fix(lens): preserve complete conversation history and tool types

* feat(lens): add trace full-screen and close controls

* fix(lens): show forwarded answers and agent errors once

* fix(lens): reset full screen when closing a trace

* fix(roi): estimate linked authors and clarify model selection

* fix: preserve trace errors and ROI results across partial failures

* fix(roi): correct pagination variable typing

* fix(ui): place loaded root failures in conversation order

* fix(roi): read estimator recommendations from model catalog

* revert: remove catalog-driven ROI recommendations
2026-10-03 01:18:16 -07:00
tin-berri
8e32d4568c
feat(ui): move LiteAdmin into the header with a docked side panel (#44293)
The floating bottom-right LiteAdmin button covered page controls such as
the Logs pagination buttons, and Playground had to hide it entirely.
Render the trigger as a pill in the header tools ahead of Docs and open
LiteAdmin as a panel docked beside the content column, which narrows the
page instead of covering it. Add a Cmd/Ctrl+J toggle and drop the
Playground override.

The Logs and trace drawers treated Cmd+J as a plain J and advanced the
selection, so they now share RunDrawer's rule that letter shortcuts
yield to modified presses and typing.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 00:00:10 -07:00
devin-ai-integration[bot]
0c86d6bfc9
refactor(dashboard): migrate to zod 4 and openai 6 (#44345)
* refactor(dashboard): migrate to zod 4 and openai 6

Bump the dashboard to real zod 4.6.5 and openai 6.49.0 so every module
imports from bare "zod" instead of "zod/v4". Ports the ten files that
still used the zod 3 API (error params, record, passthrough, strict,
email/date validators, union discriminator codes) and adapts the
LiteAdmin tool schemas and form plumbing where openai 6 and zod 4
changed behaviour.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(dashboard): prettier format zod 4 schema files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(dashboard): restore system one missing state message under zod 4

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:53:41 +00:00
moe-berri
caee45fed4
feat(roi): add GitLab sources and branch cost attribution (#44324)
* feat(roi): support GitLab and tagged branch costs

* fix(roi): count tagged branches independently of estimation status

* test(roi): capture live GitHub and GitLab report validation

* fix(roi): open estimate details at the start

* fix(roi): clarify cost views and unify report layout

* feat(roi): showcase per-PR costs in the sample report

* fix(roi): separate report tabs and preserve branch cost attribution

* fix(roi): preserve demo previews and align progress spacing

* fix(roi): isolate demo loading and parallelize fork lookups

Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression

* fix(roi): separate demo and live loading states

Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers

* fix(roi): ignore refreshes from a previous source

* fix: trust gateway context for ROI estimator exclusion

* fix: preserve historical ROI estimator exclusion
2026-10-03 06:01:43 +00:00
ishaan-berri
80a2f4d8a8
feat(lens): add preset watch-for checks to investigation setup (#44313)
* feat(lens): add preset watch-for checks for common agent failures

* feat(lens): add keyboard-driven watch-for picker with lens dot animation

* feat(lens): use the watch-for picker in investigation setup

* feat(lens): show preset checks by name in the criteria tab

* test(lens): cover saving and editing watch-for presets

* feat(lens): shorten watch-for summaries and start with three presets on

* feat(lens): lay out watch-for presets as toggle tiles with a clear add-your-own button

* feat(lens): open a custom check from the watch-for picker

* test(lens): cover watch-for tiles and the add-your-own button

* fix(lens): draw the selected tile border inside the tile so the dialog edge cannot clip it
2026-10-02 21:06:04 -07:00
ishaan-berri
2ccb7ed06e
feat(lens): issue briefs with problem, user goal, outcome and test cases (#44311)
* refactor(lens): move analysis prompts into markdown files

* feat(lens): ask the investigator for a scoped agent fix brief with two options

* test(lens): cover the agent fix brief through investigation and merges

* chore(ui): regenerate api types for the lens fix brief

* feat(ui): build copyable lens fix prompts

* feat(ui): show the lens fix brief with copy buttons for claude code and codex

* test(ui): cover copying a lens fix option

* refactor(lens): replace the fix options with a plain issue brief

* feat(lens): ask for problem, user goal, outcome and test cases without prescribing code changes

* test(lens): cover the issue brief through investigation and merges

* chore(ui): regenerate api types for the lens issue brief

* refactor(ui): drop the lens fix prompt builders

* feat(ui): add a lens issue brief panel

* feat(ui): show the lens issue brief in the finding drawer

* test(ui): cover the lens issue brief and the legacy fallback

* feat(ui): render a lens issue brief as a markdown document

* test(ui): pin the lens issue brief markdown layout

* feat(ui): show the issue brief as a copyable file with claude code and codex buttons

* feat(ui): pass the finding title into the issue brief

* test(ui): cover copying the issue brief for claude code and codex

* feat(ui): bold the input and expected labels in lens test cases

* test(ui): pin the bold test case labels in the issue brief

* feat(ui): render the issue brief as formatted markdown

* test(ui): cover the rendered issue brief sections and raw markdown copy
2026-10-03 03:00:29 +00:00
tin-berri
f63d989ff9
feat: add Bespoke Nimble gateway and OSS classifier support (#44246)
* feat: add Bespoke Nimble gateway and OSS classifier support

* feat: accept Ollama's nimble model name for the Bespoke provider

* test: exempt the POST-only bespoke decisions route from the all-methods check

test_pass_through_routes_support_all_methods requires every built-in
pass-through route to accept every HTTP method unless it is listed in
PROTOCOL_CONSTRAINED_PASS_THROUGH_ROUTES. /bespoke/v1/systemone is
POST-only like /laya/v1/systemone, so the test failed at this branch
and passed at the merge base. List it alongside Laya.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-02 19:37:08 -07:00
yujonglee
677205b3f5
fix(proxy): share ownership permissions for spend logs and traces (#44239)
* refactor(proxy): extract shared spend log read policy

* test(proxy): use named bindings for spend scope regression

* test(proxy): reuse existing spend log query harness

* test(proxy): cover spend log permission lookup adoption

* chore(proxy): relocate existing spend query baseline

* refactor(proxy): make scope query returns explicit

* refactor(proxy): inject deferred log permission lookup

* test(proxy): cover teamless management compatibility lookup

* refactor(proxy): compose user and team log grants

* refactor(proxy): share generic authorization composition

* refactor(proxy): compose trace read permissions

* refactor(proxy): centralize spend and trace authorization

* refactor(proxy): strengthen spend and trace scope types

* refactor(proxy): flatten log read scope into owned logs

Replace the AnyOf grant tree with a flat OwnedLogs(user_id, team_ids) scope,
and OwnedTraces(logs, api_key_hash) for traces, since every consumer flattened
the tree back into that shape.

A caller with no user id now gets an empty scope instead of matching ownerless
rows through Prisma's IS NULL. The dead request_id guard in ui_view_spend_logs
is removed, and the management facets inject the log team lookup and reuse
read_scope_sql instead of the list shim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(proxy): run spend scope tests through one SQLite emulator

Replace the string-matching payload emulator and the hand-rolled Prisma where
interpreter with one SQLite helper that runs the real scope SQL. Session scope
tests now go through the endpoint, including the no-user caller that must not
match ownerless rows. Drop duplicated lookup-failure and trace mapping cases.

load_permitted_log_team_ids returns no teams without a database instead of
relying on the resolver's broad except.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(proxy): unify log and trace ownership permissions

* test(tracing): align fixtures with ownership read scopes

* refactor(tracing): align query scopes with row ownership

* refactor(spend): make ownership SQL predicates explicit

* test(spend): validate ownership SQL against PostgreSQL

* docs(traces): drop key-row visibility from query help guide

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): reach the empty-memberships branch in team lookup test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate dashboard API types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 01:38:54 +00:00
devin-ai-integration[bot]
4b9f9903f3
refactor(ui): compose dashboard pages with shared layouts (#44306)
* refactor(ui): compose logs tabs directly in the route

* feat(ui): share composable dashboard page layouts

* refactor(ui): compose page header and logs toolbar from parts

PageHeader drops its icon/title/subtitle/primaryAction/tabs/utilities props and the
leadingControls render prop in favor of PageHeaderTitle, PageHeaderDescription and
PageHeaderControls that each wrap one element and forward native props.

LogsTableToolbar's 15 props collapse into one LogsTimeRange value plus composable
LogsToolbar, LogsTimeRangePicker and LogsToolbarSwitch parts assembled in the panel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ui): express DataTable layout classes as cva variants

Replaces the hand-rolled class-pair constants with boolean cva variants,
which also brings DataTable back under the complexity budget.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 01:35:34 +00:00
ishaan-berri
b28ce93d2b
feat(lens): show investigation progress as one staged bar with time left (#44301)
* feat(lens): compute overall investigation progress, speed and time left

* test(lens): cover overall progress, speed and time left estimates

* feat(lens): show investigation progress as one staged bar with time left

* feat(lens): track per-stage counts and durations for the progress readout

* test(lens): cover stage durations and short time left labels

* feat(lens): restyle investigation progress as a terminal-style readout
2026-10-03 01:28:22 +00:00
moe-berri
f498176a27
feat(lens): add sample previews and improve setup and worker feedback (#44268)
* feat(lens): add interactive traces and investigations demo

* fix(lens): limit sample previews to setup screens

* fix(lens): simplify tracing setup and align sample previews

* fix(ui): unify Lens and ROI demo notices

* fix(lens): distinguish preparation from zero selected runs

* fix(lens): report incompatible workers and address setup review
2026-10-02 17:27:13 -07:00
devin-ai-integration[bot]
22bd49e231
fix(ui): make the Lens traces refresh button always clickable (#44252)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-02 16:54:13 -07:00
moe-berri
aef0a53837
fix(lens): default to traces and restore closed span details (#44260) 2026-10-02 16:05:00 -07:00
devin-ai-integration[bot]
485ad76635
chore(ui): untrack tsconfig.tsbuildinfo and gitignore *.tsbuildinfo (#44261)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:49:47 +00:00
yujonglee
688d791fa0
feat(traces): type queries and align read access with log visibility (#44228)
* wip

* wip

* test(traces): separate root status from diagnostic error counts

* test(traces): cover normalization precedence and fallbacks

* chore(cache): remove stray comments from trace PR

* test(traces): name lens test for shared query path

* fix(traces): place query implementation before test module

* test(traces): use unified read scope in migration tests

* ci(rust): allow feature checks to finish

* ci(mcp): allow dependency resolution to finish

* fix(traces): preserve key visibility and safe spend attribution
2026-10-02 21:55:42 +00:00
devin-ai-integration[bot]
ba75a588c9
fix(ui): make model leaderboard chart bars wide and readable (#44249)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-02 21:25:58 +00:00
ishaan-berri
481a403090
feat(tracing): support claude agent sdk traces with agent name, logo and chat content (#44248)
* feat(traces): add Framework column to otel_traces

* feat(traces): pass span events to normalizers and add framework field

* feat(traces): add Claude Code and Agent SDK span normalizer

* feat(traces): decode events before normalizing and apply tool span names

* feat(traces): list distinct frameworks per trace

* feat(traces): return span framework in trace spans query

* test(traces): add scrubbed Claude Agent SDK OTLP fixtures

* test(traces): cover Claude Agent SDK normalization from real exports

* test(traces): assert trace list frameworks stay scoped per trace

* feat(tracing): validate framework in native normalized spans

* feat(tracing): add framework to Span and frameworks to TraceSummary

* feat(tracing): store normalized framework on span rows

* feat(tracing): surface span framework and trace frameworks

* test(tracing): cover framework aggregation in trace summaries

* test(tracing): decode Claude Agent SDK rows with framework and tool args

* chore(ui): regenerate API types for trace frameworks

* feat(ui): add trace framework registry for Claude Agent SDK and Claude Code

* feat(ui): show SDK logo and label in the runs list Agent column

* feat(ui): show SDK logo and label in the run header

* test(ui): cover SDK label and logo in the runs list

* test(ui): cover SDK label and logo in the run header

* feat(tracing): show the agent's final answer as claude agent span output

* feat(tracing): name claude code agents after their otel service

* test(tracing): cover claude code agent naming from the service

* fix(tracing): mark the span row framework field read-only

* test(tracing): scrub host os details from the claude sdk fixture

* test(tracing): scrub host os details from the detailed claude sdk fixture

* fix(ui): hide the decorative sdk logo from screen readers

* feat(ui): show the agent name with the sdk logo in the runs list

* feat(ui): show the agent name with the sdk logo in the run header

* test(ui): cover agent names beside the sdk logo in the runs list

* test(ui): cover the agent name in the run header
2026-10-02 21:24:56 +00:00
yuneng-jiang
626357549f
fix(ui): shrink the sidebar logo so it stops outweighing page titles (#44247)
* fix(ui): shrink the sidebar logo so it stops outweighing page titles

At h-7 the wordmark's capitals render about 21.5px tall, taller and heavier
than the 24px page titles (about 17px capitals). h-5 brings them to about
15px, between the 13px nav labels and the page title.

* fix(ui): keep the collapsed sidebar monogram at 28px
2026-10-02 14:24:35 -07:00
devin-ai-integration[bot]
0238ec9721
fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key (#43978)
* fix(scim): apply path-less group PATCH ops instead of storing them under an empty metadata key

A path-less add/replace op (RFC 7644 3.5.2, what Okta Push Groups sends on a
rename) carries a partial Group resource. Each of its attributes now applies as
if sent with that path, so displayName updates the team alias and externalId
and members get their usual handling, and the pushed attributes merge into the
scim_data snapshot the PUT path already writes. A path-less remove or a
path-less op without an object value is rejected with a 400. Any group PATCH
drops an empty metadata key an earlier push left behind, and the Admin UI
metadata form skips an empty key so an affected team can save its settings.

* fix(scim): let a later path op win over an earlier path-less value in the group snapshot

* fix(scim): type the stored team metadata before the JSON object check

* test(scim): run the real group transformation in the path-less replace test

* test(scim): assert the renamed group comes back from the path-less replace

* test(scim): audit the path-less group PATCH on the live proxy

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 14:20:04 -07:00
devin-ai-integration[bot]
c8cd885251
feat(ui): add System One (Jev) tab to the playground (#44043)
* feat(ui): add System One (Jev) playground tab

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): validate System One inputs and refresh request context

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): validate System One response payloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): fall back to requested model for System One

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): use useMutation and zod schemas for the System One tab, mark it Beta

Replace the hand-written request and response guards with zod schemas, which also
provide the types. Send requests through useMutation instead of manual loading,
error and race-guard state. Unknown spec fields now pass through to the upstream
model, and validation errors point at the exact offending key

* feat(ui): flag the System One tab as a TypeSafe-only beta

* feat(ui): highlighted JSON editor for the System One tab

Line numbers, JSON syntax highlighting through the existing react-syntax-highlighter
dependency, a valid or issue-count status badge, and a compact path plus message issue
list replace the bare textarea and stacked alerts. Answer card type badges now sit on
the header row

* fix(ui): let the System One results scroll to the bottom

The tab panel was viewport height but sat below the tab bar, so its bottom was cut off.
On wide screens the editor and results now scroll independently, answers render above
the question breakdown, and the raw response no longer nests its own scroll area

* feat(ui): color the model, state and questions blocks in the System One editor

Tints each top-level request block in the JSON editor and marks the matching breakdown sections with the same color, so it is clear which part of the payload feeds which panel

* feat(ui): wrap long lines in the System One JSON editor

Long state strings no longer need horizontal scrolling. Each line renders as its own row with its number and block color, so wrapped lines keep their line number and the caret stays aligned

* fix(ui): remove horizontal scrolling from the System One tab

Long unbroken text in the state, question ids, choice labels and the raw response now wraps instead of widening its box

* refactor(ui): replace System One presets with one example and a reset button

The tab now starts with a single product review example that uses all three question types, and Reset example restores it after editing

* refactor(ui): use an issue triage request as the System One example

* fix(ui): type the System One line renderer from exported props

rendererProps is not exported by the react-syntax-highlighter types, which broke the dashboard build

* fix(ui): address System One review feedback

Highlights the score level nearest a fractional calibrated score, keeps extra noul criteria fields in the sent payload, and clears an answer when the request key changes

* refactor(ui): parse System One root blocks without mutation and preview all noul criteria

The root-block finder is now a tokenizer plus a pure reduce, and the question preview lists every noul criterion that will be sent

* refactor(ui): group System One playground files into components and lib

Drops the repeated SystemOne prefix from the inner component files and moves the pure logic (schemas, example, payload validation, root block parsing) into lib/. Tests stay colocated with their files, matching the rest of the dashboard. No behavior change

* feat(ui): link the decision models discussion from the System One beta notice

* feat(ui): ask for decision model feedback in the System One beta notice

* feat(ui): make the decision model feedback text the discussion link

---------

Co-authored-by: ryan <ryan@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 20:10:50 +00:00
tin-berri
b1e0e9e84b
feat(ui): select Laya for OSS classification (#43768) 2026-10-02 12:17:55 -07:00
moe-berri
2584721ca3
fix(lens): preserve framework agent names and GenAI message content (#44218)
* fix(lens): use recorded agent identities across framework traces

* fix(lens): tighten agent identity and bound trace lookups

* style(tracing): wrap framework agent identity test case
2026-10-02 11:58:32 -07:00