Commit graph

1142 commits

Author SHA1 Message Date
moe-berri
e9cfba2c17
feat(lens): isolate ingestion and investigations in a Rust service (#45148)
* feat(lens): isolate trace storage and investigation in a Rust service

* fix(lens): include Rust sources in the image build context

* feat(lens): wire service setup, scoped delivery receipts and lease attempts

* fix(lens): complete service routing and reject stale investigation results

* fix(lens): retry key propagation and validate isolated Compose setup

* fix(lens): seed through isolated ingestion and preserve upstream queue fixes

* chore: sync schema.prisma copies from root

* fix(lens): bind nullable due timestamps as text for Prisma

* chore(ui): remove stale lint suppressions

* fix(lens): address CI failures and review findings

* refactor(lens): remove retired Python worker and run evaluations in Rust

* fix(lens): reuse control connections and satisfy review checks

* test(lens): install and upgrade both Helm charts on Kubernetes

* test(lens): run connection reuse coverage as an integration test

* fix(ui): upgrade Next.js to 16.3.8 security release

* fix(lens): fence stale attempts and preserve reviewed evidence

* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"

This reverts commit 2f79a51b25.

* fix(lens): stop failed investigations and stream history excerpts

* test(lens): cover model tool and result contracts

* test(lens): fix retired routes and reuse installation build artifacts

* test(lens): use portable grep in Helm installation smoke

* test(lens): wait for migrations before forwarding Helm services

* fix(lens): keep failed evidence reads retryable

* fix(lens): preserve sandbox output during process exit

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-10-07 17:18:58 -07:00
devin-ai-integration[bot]
c38505d265
chore(codeowners): replace kerry-berri with kerrylu-berri (#45173)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 17:02:37 -07:00
ryan-crabbe-berri
048a1500df
test(e2e): move the harness self-tests out of tests/e2e (#45172)
* test(e2e): move the harness self-tests out of tests/e2e

The nightly Buildkite run copies tests/e2e into the runner image and runs
bare pytest, so the 672 tests of the harness itself (fixture parsing, JUnit
properties, the stack lock, the load aggregators, the Claude Code driver)
counted as e2e tests on the status page even though none of them reaches a
proxy. They now live in tests/e2e_harness, mirroring the tests/e2e layout,
and run in the GitHub Actions lint job and the CircleCI
provider_replay_harness job instead

* fix(ci): point the providers replay controls at tests/e2e_harness

The providers integration job still selected the four replay-control
tests under tests/e2e/test_provider_edge.py, so pytest exited before
they ran. The raw-HTTP check's file walk also drops to one loop per
comprehension

* style(tests): mark the raw-HTTP check's bindings Final
2026-10-07 22:20:02 +00:00
devin-ai-integration[bot]
d474b433cf
ci(codeql): run the default suite on full scans and security-extended on PRs (#45149)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-07 13:24:28 -07:00
ishaan-berri
5fb23eabd8
feat(lens): add datasets built from real traces (#44765)
* feat(lens): add dataset case and size limits

* feat(lens): add dataset, case and build models

* feat(lens): build dataset cases from traces, findings and text

* feat(lens): store dataset revisions insert-only

* feat(lens): add dataset routes for build, save, export and eval cases

* feat(lens): mount the dataset router before lens routes

* feat(lens): add LiteLLM_LensDataset table

* feat(lens): add LiteLLM_LensDataset table to proxy schema

* feat(lens): add LiteLLM_LensDataset table to extras schema

* feat(lens): add migration that creates the dataset table

* test(lens): cover dataset case building, dedupe and limits

* test(lens): cover dataset revisions, conflicts and eval cases

* chore(ui): regenerate API types for lens datasets

* feat(lens): add dataset UI types

* feat(lens): add datasets API client

* feat(lens): add dataset query and mutation hooks

* feat(lens): add case selection and expected edit logic

* test(lens): cover case selection and expected edits

* feat(lens): add the add to dataset dialog

* test(lens): cover saving picked cases from the dialog

* feat(lens): add datasets list

* feat(lens): add dataset detail with revisions and export

* test(lens): cover editing, revisions and export in datasets tab

* feat(lens): expose datasets on the lens API

* feat(lens): add in-memory datasets for demo mode

* feat(lens): wire demo datasets into the demo lens API

* feat(lens): add optional lens API hook

* feat(lens): add optional onboarding hook

* feat(lens): add datasets tab and dataset routing

* feat(lens): show the datasets tab

* feat(lens): add to dataset from the trace header

* feat(lens): add a single turn to a dataset from a step

* feat(lens): add finding evidence to a dataset

* docs(lens): add datasets screenshots for the PR

* test(lens): cover dataset revision storage against Postgres

* test(lens): cover dataset trace paging, findings and route errors

* test(lens): cover dataset build fallbacks and no_content skips

* fix(lens): register the dataset table for postgres span names

* feat(lens): add invalid skip reason for unparseable case lines

* fix(lens): keep valid JSONL cases, provenance and size limits; stop reading past the case cap

* test(lens): cover malformed JSONL, re-import provenance, size fields and early cap

* chore(ui): regenerate API types for the invalid skip reason

* feat(lens): show text for the invalid skip reason

* feat(lens): render datasets in the same inspector table as investigations

* test(lens): open a dataset by clicking its table row

* docs(lens): update the datasets list screenshot

* style(lens): format the datasets table

* fix(lens): normalize JSON text span input and output into messages when building cases

* test(lens): cover JSON text span normalization for dataset cases

* Revert "fix(lens): normalize JSON text span input and output into messages when building cases"

This reverts commit 29c3e34962fa12427dbd516c08f2599db997a905.

* fix(traces): normalize agent assistant summaries into UI messages

* feat(lens): add dataset case view helpers built on the trace parsers

* test(lens): cover dataset case view helpers

* feat(lens): show dataset cases in an inspector table

* feat(lens): open a dataset case in a side panel with trace message cards

* feat(lens): rebuild the dataset page header and layout

* feat(lens): keep the open dataset case in the URL

* test(lens): drive dataset edits through the case table and panel

* docs(lens): update dataset view screenshots

* Revert "test(lens): cover JSON text span normalization for dataset cases"

This reverts commit 30229ae8ae90265c09d755c43c1f1eeaa66050c7.

* fix(lens): import StateMessage from shared in dataset detail

* fix(lens): import StateMessage from shared in datasets list

* fix(lens): give the datasets migration a unique timestamp after review checkpoints
2026-10-06 23:04:43 +00:00
devin-ai-integration[bot]
acd2a7ccc5
fix(lens): use async-timeout on python 3.10 for budget reservation timeouts (#44911)
* fix(lens): use async-timeout on python 3.10 for budget reservation timeouts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(deps): keep uv.lock diff to the async-timeout entry

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(lens): cover real request deadline expiry in reserved_budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lens): finish Python 3.10 timeout coverage and dependency checks

* test(lens): control event-loop time for deadline regressions

* test(tracing): include priced call count in trace fixture

* fix(lens): limit timeout compatibility changes to PR scope

---------

Co-authored-by: Moe Khalil <moe@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 14:19:37 -07:00
devin-ai-integration[bot]
5cdebded90
fix(ci): skip the generated dashboard bundle in the master key guard (#44890)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 18:15:09 +00:00
devin-ai-integration[bot]
837c6a7481
fix(security): remove the publicly known master key from the repo (#44718)
* fix(security): hash the publicly known master key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: replace weak master key examples and regenerate artifacts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: replace weak key fixtures with generated test keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: generate master keys for proxy startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve lens dev key entropy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: restore proxy key compatibility in scrub examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: scrub merged SSO fixture and refresh dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ci): stabilize test keys and metadata collection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: drop the rebuilt dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:55:24 -07:00
moe-berri
b58e2d7175
fix(lens): bound result recovery and preserve partial results (#44692)
* fix(lens): restate response contract during model repair

* fix(lens): separate instructions and recover rejected results

* fix(lens): correct loop type annotations and checks

* fix(lens): preserve access to prior findings after compaction

* fix(lens): cap result retries and preserve partial completion
2026-10-05 18:24:53 -07:00
devin-ai-integration[bot]
169af2f883
ci: split slow unit shards and build the Rust bridge once per run (#44622)
* ci: split slow unit shards and build the Rust bridge once per run

* ci: key the Rust bridge cache on source files only

* ci: keep the unit setup ceiling unchanged with the shared Rust bridge

* ci: fall back to the Cargo cache when the Rust bridge artifact is missing

* ci: keep reruns on enterprise-routing for the prompt caching flake

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 23:43:01 +00:00
devin-ai-integration[bot]
7bad8de067
chore: move PR template to PULL_REQUEST_TEMPLATE/general.md, add rust.md (#44693)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:47:39 +00:00
tin-berri
9062fd3931
feat(proxy): embed enterprise LiteAdmin MCP in LiteLLM images (#44610) 2026-10-05 15:24:34 -07:00
moe-berri
b69d744993
feat(lens): analyze trace workspaces with confined Python and compaction (#44640)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* feat(lens): integrate confined workspace analysis with live reviews

* fix(lens): synchronize confined Python process monitoring

* Update review.md

* fix(lens): allow mixed context capacities and correct review assertions

* fix(lens): retrieve evidence on demand and isolate failed reviews

* fix(lens): isolate incomplete evidence reads from peer reviews

* test(lens): await trace status filter option

* test(lens): wait for reclaimed review state to settle

* fix(lens): recover from incomplete cross-session evidence

---------

Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
2026-10-05 22:06:48 +00:00
devin-ai-integration[bot]
8364f88cbb
ci: run unit selections from GHA test-path and drop the CircleCI unit jobs (#44461)
* ci: run unit selections from GHA test-path and drop the CircleCI unit jobs

* ci: keep existing shard token order and header note

* ci: throwaway, drop tests/unit/repositories from the unit shard to show assert-ci-coverage fails

* ci: revert throwaway assert-ci-coverage check

* ci: stop crediting --ignore paths as invoked in assert_ci_coverage

* ci: install the caching, extra_proxy and proxy-runtime extras in the GHA unit sync

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 10:49:16 -07:00
devin-ai-integration[bot]
9b6a6a0b71
refactor(tracing): generate existing HTTP request models from Rust schemas (#44591)
* test(tracing): pin HTTP request compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): generate existing HTTP request models from Rust schemas

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(tracing): bind trace query params to generated request models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 48 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* perf(tracing): read the trace list clock without a thread-pool dispatch

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(traces): share the trace page-size bounds between schema and reader

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): type trace request queries against the generated OpenAPI schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: encode the trace contract boundary in AGENTS.md

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format trace request aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 17:37:52 +00:00
devin-ai-integration[bot]
902736bfe7
ci: move Postgres, MCP and Redis suites to CircleCI integration (#44453)
* ci: move Postgres, MCP and Redis suites to CircleCI integration

* ci: throwaway, drop tests/proxy_behavior from its CircleCI job to show assert-ci-coverage fails

* ci: revert throwaway assert-ci-coverage check

* ci: keep the e2e helpers the gate tests still use

* ci: move the roi-database Postgres shard to CircleCI integration

* ci: run redis-compat without CircleCI's Azure and cassette env, cover postgres_suite test_path

* ci: match the GitHub env for the moved Postgres and Redis jobs

* ci: unset provider keys in the CircleCI MCP job and drop unused e2e-stack helpers

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 09:33:14 -07:00
moe-berri
8520626e7a
fix(lens): preserve approved worker digests and harden its image (#44467)
* fix(lens): pin worker dependencies and support approved image digests

* fix(lens): include locked dependencies and release identity in build context

* fix(lens): select the dev worker package for SHA-tagged charts
2026-10-03 17:54:17 -07:00
moe-berri
cb17588276
feat(lens): coordinate worker releases and bundled installs (#44428)
* feat(lens): coordinate worker versions and bundled installs

* test(lens): exercise bundled Compose startup and restart in CI

* fix(lens): refund failed model requests without a response

* test(lens): verify trace persistence in the bundled stack

* fix(lens): align Helm images and isolate Compose storage

* fix(lens): reject worker builds without release identity

* fix(lens): encode Compose credentials and normalize worker versions

* fix(lens): refuse worker recommendations for unidentified builds
2026-10-03 16:33:45 -07:00
moe-berri
1a7023366f
feat(roi): measure shipping velocity, quality, and recorded spend (#44426)
* feat(ui): prototype observed engineering ROI dashboard

* feat(roi): replace effort estimates with measured repository metrics

* fix(roi): finish connection recovery and generated API contracts

* fix(roi): show merged changes before accounts are linked

* fix(roi): preserve selected report tab across refreshes

* fix(roi): recover app authorization and keep detail values readable

* fix(roi): reuse the shared OAuth HTTP client

* feat(roi): combine providers and compare equal reporting periods

* docs: explain ROI metrics for first-time readers

* fix(roi): preserve connections and scheduled reports during setup

* ci(roi): assign database contracts to the active Postgres shard

* fix(roi): preserve issue counts and normalized connections

* fix(ui): compact ROI dashboard header and metrics

* fix(ui): show ROI repository count with expandable list

* fix(ui): wrap ROI controls within narrow panels

* fix(roi): restore sample report preview and simplify setup
2026-10-03 23:07:34 +00:00
devin-ai-integration[bot]
8b1990b4bc
feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers (#44236)
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register typesafe as a provider so Jev deployments load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move provider endpoints under llms and validate proxy bodies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add Cloudflare Clef and Strands Decider backends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register decisions routes for managed agents and gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(decisions): use raw regex for cloudflare missing account match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): avoid cast in Cloudflare response unwrapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): default model, evaluation health probe, short Cloudflare names

The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.

* fix(decisions): let health_check_params override the evaluation probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit the decisions endpoint across providers, limits, health and chaos

Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).

The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.

* fix(decisions): send env API keys to a configured api_base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register the routes through the lazy feature registry

The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.

The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.

* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell

* fix(proxy): let a config pass-through beat a lazily registered route in eager mode

With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:38:38 +00:00
devin-ai-integration[bot]
e340e546e2
feat(traces): tracing development seed (#44363)
* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* wip

* wip

* wip

* chore(trace): checkpoint ongoing Rust migration

* refactor(trace): group Python bridge under trace package

* refactor(traces): read span conventions through a Convention trait

Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.

The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.

* feat(trace): export Rust-owned wire schemas and enforce contract bounds

* fix(trace): bound quoted counts in ClickHouse wire schemas

* feat(trace): generate Python wire contracts with datamodel-code-generator

* test(trace): validate migrated callers and generated contracts at the native boundary

* refactor(traces): rename normalization convention to format

* fix(traces): reconcile spend evidence and preserve unknown costs

* feat(traces): normalize additional telemetry formats

* test(traces): cover captured normalization fixtures

* refactor(traces): isolate SDK normalization rules

* feat(tracing): seed all trace exports for local dashboard

* fix(clickhouse): preserve custom LiteLLM request metadata

* docs(traces): define normalization module boundaries

* docs(traces): define resolution and OTLP boundaries

* fix(ui): normalize nullable trace message names

* refactor(traces): split resolver modules and cover resolution behavior

* test(traces): replace normalization snapshots with behavior assertions

* fix(ui): align dashboard API contracts with generated types

* refactor(traces): type normalization and storage boundaries

* fix(traces): seed captured SDK spend and preserve provider identities

* wip

* test(traces): verify guide discovery and content ordering

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:59:51 +00:00
moe-berri
caee45fed4
feat(roi): add GitLab sources and branch cost attribution (#44324)
* feat(roi): support GitLab and tagged branch costs

* fix(roi): count tagged branches independently of estimation status

* test(roi): capture live GitHub and GitLab report validation

* fix(roi): open estimate details at the start

* fix(roi): clarify cost views and unify report layout

* feat(roi): showcase per-PR costs in the sample report

* fix(roi): separate report tabs and preserve branch cost attribution

* fix(roi): preserve demo previews and align progress spacing

* fix(roi): isolate demo loading and parallelize fork lookups

Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression

* fix(roi): separate demo and live loading states

Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers

* fix(roi): ignore refreshes from a previous source

* fix: trust gateway context for ROI estimator exclusion

* fix: preserve historical ROI estimator exclusion
2026-10-03 06:01:43 +00:00
devin-ai-integration[bot]
6d8434f940
fix(proxy): return 4xx instead of 500 for missing required params, invalid pagination and unknown ids (#43787)
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: run search_endpoints tests in proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(llm_http_handler): keep provider error text when re-raising mapped errors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): allow promptless image edits and default search models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): default missing image edit image to None

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): build image edit defaults without mutating request data

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): inject a fake router for the search default model test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): cover provider error status on vector store and file lookup handlers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): keep the lookup handler raise block to a single statement

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover missing required body params and provider lookup status codes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bind spend-row request id with partial to satisfy B023

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): only reject non-positive page_size on vector store list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): remove unreachable fine-tuning body validation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover streaming anthropic messages reaching the upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): count only provider calls when asserting missing params never reach the upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve merge-base request compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve interaction completion model defaults

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-02 22:48:53 -07:00
yujonglee
688d791fa0
feat(traces): type queries and align read access with log visibility (#44228)
* wip

* wip

* test(traces): separate root status from diagnostic error counts

* test(traces): cover normalization precedence and fallbacks

* chore(cache): remove stray comments from trace PR

* test(traces): name lens test for shared query path

* fix(traces): place query implementation before test module

* test(traces): use unified read scope in migration tests

* ci(rust): allow feature checks to finish

* ci(mcp): allow dependency resolution to finish

* fix(traces): preserve key visibility and safe spend attribution
2026-10-02 21:55:42 +00:00
devin-ai-integration[bot]
b21e44cbf9
feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key (#42375)
* test(e2e): jwt auto_register map-existing-key repro

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): exclude blocked keys from auto_register_map_existing_key reuse

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(jwt): route existing-key lookup through VerificationTokenRepository

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): stop requiring LITELLM_SALT_KEY for the owned JWT gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): gate the owned JWT gateway tests behind E2E_OWNED_GATEWAY

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(jwt): only reuse keys that can call LLM routes in auto_register_map_existing_key

Skip Admin UI session keys and keys whose allowed_routes restrict them to
anything other than llm_api_routes (management, read_only, password-reset
sessions). Mapping a JWT to one of those left the user with 401s or 403s on
every LLM call, since the mapping persists.

* fix(jwt): scope auto_register_map_existing_key reuse to the JWT-resolved team

Only reuse a key whose team_id matches the team auth_builder resolved for
the JWT (no team matches no team), so a personal key can no longer bypass
the resolved team's model and budget limits.

With the flag on, the first JWT request now falls through to the same
virtual-key checks later mapped requests get, instead of returning early,
so a reused key's own limits apply from request one rather than 200 then
403. Flag off keeps the early return unchanged.

* fix(jwt): keep the early return when no master key is set

Without a master key the generic virtual-key path returns a bare
INTERNAL_USER object, so falling through on the first auto-registered
request dropped the key's team, models and budgets. Only fall through when
a master key is configured.

Tests now assert the reused key per team rather than the query shape, and
cover the flag-off early return and the no-master-key case.

* test(jwt): assert on race-loser's returned key, not only mocks (TQ002)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(jwt): close the auto_register_map_existing_key race, shared-claim and expiry holes

A key auto_register just minted is never adopted by a concurrent request, so the race loser's cleanup can no longer delete a key another request mapped and cascade its mapping away (503, user left with no key)

Reuse only happens when the claim value is the JWT-resolved user_id. A shared claim such as azp or client_id falls back to minting, so one user can no longer land on another user's personal key and budget

Only keys that never expire are reused, so an expiring key can no longer pin the claim to a permanent 401

Integration tests on a real proxy and Postgres cover all three. The race test holds the first mapping insert in a Postgres relay, so the interleaving is forced rather than timed. The where-clause shape unit tests are replaced by these, since only a real database proves the filter

* test(e2e): create the reused key in the team the JWT resolves to

The flag only reuses a key in the JWT-resolved team, and this identity's groups claim resolves to its team, so a teamless key was never eligible and the test could not pass

* test(integration): match the held statement across TCP reads

The relay looked for the trigger inside one read, so an insert split across two reads was never held and the race test would fail waiting for it. It now matches one exact trigger over a window that keeps the end of the previous read

* fix(jwt): gate key reuse on the claim field, not on the claim value

Requiring the claim value to equal the resolved user_id skipped reuse for users matched through the sso_user_id or case-insensitive email fallback, whose stored user_id differs from the JWT sub. That is the lookup LIT-5378 asks for. Reuse is now allowed when the virtual key claim is the user_id or user_email JWT field, globally or for the token's issuer, which still keeps shared claims such as azp or client_id on the mint path

* fix(jwt): let an issuer's own user field replace the global one when gating key reuse

An issuer that identifies users by uid no longer treats the global sub field as a user identity claim, so a shared sub under that issuer mints instead of reusing a personal key

* test(jwt): make the flag-off test fail when the flag no longer gates key reuse

The flag-off test used a config where sub was not a user identity claim, so deleting the flag check still passed. Configure user_id_jwt_field=sub so only the flag keeps the lookup off, and drop test docstrings

* chore(lint): drop mutable-ok suppressions that LIT013 flags as no-ops

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Mrinal Chanshetty <mchanshetty@Mrinals-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 12:11:33 -07:00
ishaan-berri
0f6a06a6b9
feat: add litellm.agent() to run claude code, codex, opencode and deep agents through the ai gateway (#43885)
* feat(harness): add litellm/__init__.py

* feat(harness): add litellm/constants.py

* feat(harness): add litellm/harness/__init__.py

* feat(harness): add litellm/harness/adapters/__init__.py

* feat(harness): add litellm/harness/adapters/base.py

* feat(harness): add litellm/harness/adapters/claude_code.py

* feat(harness): add litellm/harness/adapters/codex.py

* feat(harness): add litellm/harness/adapters/opencode.py

* feat(harness): add litellm/harness/endpoint.py

* feat(harness): add litellm/harness/errors.py

* feat(harness): add litellm/harness/options.py

* feat(harness): add litellm/harness/runtime.py

* feat(harness): add litellm/harness/sandbox/__init__.py

* feat(harness): add litellm/harness/sandbox/base.py

* feat(harness): add litellm/harness/sandbox/docker.py

* feat(harness): add litellm/harness/sandbox/local.py

* feat(harness): add litellm/harness/sandbox/snapshot.py

* feat(harness): add litellm/harness/sync.py

* feat(harness): add litellm/harness/types.py

* feat(harness): add litellm/sandbox/__init__.py

* feat(harness): add README.md

* feat(harness): add tests/harness_e2e/__init__.py

* feat(harness): add tests/harness_e2e/conftest.py

* feat(harness): add tests/harness_e2e/test_harness_e2e.py

* feat(harness): add tests/test_litellm/harness/__init__.py

* feat(harness): add tests/test_litellm/harness/adapters/__init__.py

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/api_error.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/max_turns.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/resume_turn.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/structured_output.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/claude_code/success_tools.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/reasoning.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/structured_output.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn_failed.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn1_bash.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/codex/turn2_resume_apply_patch.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/api_error.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/endpoint_requests.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/readonly_denied_bash.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/turn1_write_read.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/fixtures/opencode/turn2_session_skill.jsonl

* feat(harness): add tests/test_litellm/harness/adapters/test_claude_code.py

* feat(harness): add tests/test_litellm/harness/adapters/test_codex.py

* feat(harness): add tests/test_litellm/harness/adapters/test_opencode.py

* feat(harness): add tests/test_litellm/harness/core_fakes.py

* feat(harness): add tests/test_litellm/harness/sandbox/__init__.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_docker.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_local.py

* feat(harness): add tests/test_litellm/harness/sandbox/test_snapshot.py

* feat(harness): add tests/test_litellm/harness/test_endpoint.py

* feat(harness): add tests/test_litellm/harness/test_init.py

* feat(harness): add tests/test_litellm/harness/test_runtime.py

* feat(harness): add tests/test_litellm/harness/test_sync.py

* feat(harness): add tests/test_litellm/harness/test_types.py

* test(harness): use word recall in stream e2e test

* refactor(harness): update litellm/__init__.py

* refactor(harness): update litellm/constants.py

* refactor(harness): update litellm/harness/__init__.py

* refactor(harness): remove litellm/harness/adapters/__init__.py

* refactor(harness): update litellm/harness/context.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/__init__.py

* refactor(harness): update litellm/harness/handlers/base.py

* refactor(harness): update litellm/harness/handlers/cli_handler.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/docker.py

* refactor(harness): update litellm/harness/sandbox/local.py

* refactor(harness): update litellm/harness/sync.py

* refactor(harness): update litellm/harness/types.py

* refactor(harness): update litellm/llms/base_llm/harness/__init__.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/__init__.py

* refactor(harness): update litellm/llms/claude_code/harness/__init__.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/__init__.py

* refactor(harness): update litellm/llms/codex/harness/__init__.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/__init__.py

* refactor(harness): update litellm/llms/deepagents/harness/__init__.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/__init__.py

* refactor(harness): update litellm/llms/opencode/harness/__init__.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/utils.py

* refactor(harness): update README.md

* refactor(harness): update tests/harness_e2e/conftest.py

* refactor(harness): update tests/harness_e2e/test_harness_e2e.py

* refactor(harness): remove tests/test_litellm/harness/__init__.py

* refactor(harness): remove tests/test_litellm/harness/adapters/__init__.py

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/api_error.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/max_turns.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/resume_turn.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/structured_output.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/claude_code/success_tools.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/reasoning.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/structured_output.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn_failed.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn1_bash.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/codex/turn2_resume_apply_patch.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/api_error.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/endpoint_requests.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/readonly_denied_bash.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/turn1_write_read.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/fixtures/opencode/turn2_session_skill.jsonl

* refactor(harness): remove tests/test_litellm/harness/adapters/test_claude_code.py

* refactor(harness): remove tests/test_litellm/harness/adapters/test_codex.py

* refactor(harness): remove tests/test_litellm/harness/adapters/test_opencode.py

* refactor(harness): remove tests/test_litellm/harness/core_fakes.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/__init__.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_docker.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_local.py

* refactor(harness): remove tests/test_litellm/harness/sandbox/test_snapshot.py

* refactor(harness): remove tests/test_litellm/harness/test_endpoint.py

* refactor(harness): remove tests/test_litellm/harness/test_init.py

* refactor(harness): remove tests/test_litellm/harness/test_runtime.py

* refactor(harness): remove tests/test_litellm/harness/test_sync.py

* refactor(harness): remove tests/test_litellm/harness/test_types.py

* refactor(harness): update tests/unit/harness/__init__.py

* refactor(harness): update tests/unit/harness/core_fakes.py

* refactor(harness): update tests/unit/harness/handlers/__init__.py

* refactor(harness): update tests/unit/harness/handlers/test_deepagents_handler.py

* refactor(harness): update tests/unit/harness/sandbox/__init__.py

* refactor(harness): update tests/unit/harness/sandbox/test_docker.py

* refactor(harness): update tests/unit/harness/sandbox/test_local.py

* refactor(harness): update tests/unit/harness/sandbox/test_snapshot.py

* refactor(harness): update tests/unit/harness/test_endpoint.py

* refactor(harness): update tests/unit/harness/test_init.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/harness/test_sync.py

* refactor(harness): update tests/unit/harness/test_types.py

* refactor(harness): update tests/unit/llms/claude_code/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/api_error.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/max_turns.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/resume_turn.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/structured_output.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/success_tools.jsonl

* refactor(harness): update tests/unit/llms/claude_code/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/codex/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/reasoning.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/structured_output.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn1_bash.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn2_resume_apply_patch.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/turn_failed.jsonl

* refactor(harness): update tests/unit/llms/codex/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/deepagents/__init__.py

* refactor(harness): update tests/unit/llms/deepagents/harness/__init__.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/opencode/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/api_error.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/endpoint_requests.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/readonly_denied_bash.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/turn1_write_read.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/turn2_session_skill.jsonl

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* ci: allowlist tests/harness_e2e, which needs live runtimes and a gateway

* fix(harness): bound the turn event queue

* fix(harness): bound the turn event queue with backpressure

* style: sort imports in utils

* ci: exclude agent-harness config folders from provider docs check

* refactor(harness): update tests/harness_e2e/conftest.py

* refactor(harness): update tests/harness_e2e/test_harness_e2e.py

* refactor(harness): update tests/unit/harness/core_fakes.py

* refactor(harness): update tests/unit/harness/handlers/test_deepagents_handler.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/harness/test_sync.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/options.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/snapshot.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/sandbox/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/test_transformation.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_sandbox_backend_symlinks.py

* refactor(harness): update tests/unit/llms/opencode/harness/test_transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update tests/code_coverage_tests/recursive_detector.py

* refactor(harness): update tests/unit/llms/base_llm/harness/__init__.py

* refactor(harness): update tests/unit/llms/claude_code/harness/fixtures/__init__.py

* refactor(harness): update tests/unit/llms/codex/harness/fixtures/__init__.py

* refactor(harness): update tests/unit/llms/opencode/harness/fixtures/__init__.py

* refactor(harness): update litellm/harness/__init__.py

* refactor(harness): update litellm/harness/context.py

* refactor(harness): update litellm/harness/endpoint.py

* refactor(harness): update litellm/harness/handlers/__init__.py

* refactor(harness): update litellm/harness/handlers/base.py

* refactor(harness): update litellm/harness/handlers/cli_handler.py

* refactor(harness): update litellm/harness/handlers/deepagents_handler.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/harness/sandbox/__init__.py

* refactor(harness): update litellm/harness/sandbox/base.py

* refactor(harness): update litellm/harness/sandbox/docker.py

* refactor(harness): update litellm/harness/sandbox/local.py

* refactor(harness): update litellm/harness/sandbox/snapshot.py

* refactor(harness): update litellm/harness/sync.py

* refactor(harness): update litellm/harness/types.py

* refactor(harness): update litellm/llms/base_llm/harness/transformation.py

* refactor(harness): update litellm/llms/base_llm/harness/utils.py

* refactor(harness): update litellm/llms/claude_code/harness/transformation.py

* refactor(harness): update litellm/llms/codex/harness/transformation.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update litellm/llms/deepagents/harness/transformation.py

* refactor(harness): update litellm/llms/opencode/harness/transformation.py

* refactor(harness): update litellm/types/llms/custom_http.py

* refactor(harness): update tests/unit/harness/test_endpoint.py

* refactor(harness): update litellm/harness/runtime.py

* refactor(harness): update litellm/llms/deepagents/harness/sandbox_backend.py

* refactor(harness): update tests/unit/harness/test_init.py

* refactor(harness): update tests/unit/harness/test_runtime.py

* refactor(harness): update tests/unit/llms/deepagents/harness/test_sandbox_backend_symlinks.py
2026-10-01 22:27:49 +00:00
devin-ai-integration[bot]
2cfa5ec126
test(proxy): delete the legacy proxy test tree and shard tests/unit/proxy by glob (#44018)
* test(proxy): delete the legacy proxy test tree and serve the redirect test from loopback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): exercise the shard check directly for unit_selection-owned children

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): serve the redirect test from respx instead of a socket

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): credit shard ownership only to unit flags wired in gha

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): implement the wired-flag shard crediting the tests assert

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): split the root proxy test files into their own unit shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): point the rate-limit skip reason at the usage-based-routing-v2 RPM tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 11:40:45 -07:00
devin-ai-integration[bot]
a76b59db9f
test(proxy): move middleware, spend_tracking, pass_through, common_utils and root proxy tests into tests/unit/proxy (#44015)
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:23:31 +00:00
devin-ai-integration[bot]
24584d3d3d
test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy (#44012)
* test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): keep tuple identity in proxy state restore and fix misc target paths

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 18:14:24 +00:00
devin-ai-integration[bot]
25109a523b
test(proxy): move utils, agent_endpoints and endpoint tests into tests/unit/proxy (#44006)
* test(proxy): move utils, agent_endpoints and endpoint tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): package moved unit test directories

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): exclude proxy-db-owned files from the misc target

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): drop the redundant fixture docstrings in the proxy conftest

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 11:06:42 -07:00
moe-berri
259c166ef6
refactor(lens)!: rename internal engine code and API (#44034)
* chore(lens): remove deployment screenshots

* refactor(lens)!: rename internal engine package and API

* fix(lens): pin worker image for renamed API

* test(lens): cover fresh and populated rename migrations

* fix(lens): protect db-push upgrades and restore routing and CI

* fix(lens): resolve migration tables across schemas and include database driver
2026-10-01 11:02:14 -07:00
devin-ai-integration[bot]
bfd3f39dca
feat(s3_v2): add s3_partition_granularity option for hourly S3 folders (#43748)
* feat(s3_v2): add s3_partition_granularity option for hourly S3 folders

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover s3 v2 partition granularity across surfaces, settings and chaos

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover previous_response_id history rebuilt from an hourly cold storage object

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(s3_v2): reuse the cold storage key only when s3_v2 owns cold storage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): cover hour rollover, postgres outage, in-flight switches, key/team vars and real S3 layout

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(liccheck): authorize libfaketime, the GPLv2 dev-only clock the s3 rollover integration test preloads

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): wait for the rejected-request cell's payloads by id, not by line count

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): declare the postgres outage cell's models in config and trip the relay on burst ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): drop the libfaketime hour rollover cell and its dev dependency

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): deselect the s3_v2 live e2e on the stage-mirror stack

The stage-mirror config enables no s3_v2 callback, so every test in test_s3_log_e2e.py fails its readiness check there. The file keeps running in the Buildkite e2e lane, which configures s3_v2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(s3_v2): declare the sink outage burst models in config

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(s3_v2): read cold storage metadata without an empty dict default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for the proxy to reconnect before the postgres outage recovery request

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-01 11:00:21 -07:00
devin-ai-integration[bot]
73072b8643
test(proxy): move management_endpoints, management_helpers and guardrails tests into tests/unit/proxy (#44003)
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): stub HIBP through respx by disabling the aiohttp transport

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): share the httpx transport fixture across proxy unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): restore proxy globals without a missing-value sentinel

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): isolate the mcp server manager per test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): move management_endpoints, management_helpers and guardrails tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): reuse the shared httpx transport fixture in moved proxy tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): stub outbound HTTP and package moved test dirs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): restore the config server hostname in the mcp resolution test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): pin the completion tokenizer model in the straiker screening test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 10:52:03 -07:00
yuneng-jiang
6ca90b927c
test(ci): repair stale tests and flaky CI infrastructure (#43983)
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden

#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key

* test(ui): give the auto-router threshold save wait room for the availability debounce

#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed

* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default

#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against

* test(e2e-ui): wait for the call-id search before hovering the logs row

The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time

* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off

The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support

* test(integration): read the agent 365 guardrail status by its own name in spend logs

The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict

* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check

gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler

* test(e2e-ui): fill the create-tag fields inside the dialog

#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step

* test(integration): run integration proxies with the CI license

Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments

* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5

Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving

codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI

* test(unit): join the session-minting thread before collecting the handler

asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference

* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early

owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.

uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds

* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build

The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there

* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections

The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry

* ci: move the unit-test uv cache split into a composite action

check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes

* test(unit): point tiktoken at the bundled cache for every unit test

The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.

* test(integration): answer model discovery probes in the hosted_vllm wire tests

The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.
2026-10-01 17:46:43 +00:00
devin-ai-integration[bot]
39e31958f8
test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy (#43998)
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): stub HIBP through respx by disabling the aiohttp transport

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): share the httpx transport fixture across proxy unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): restore proxy globals without a missing-value sentinel

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): isolate the mcp server manager per test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 10:11:45 -07:00
moe-berri
d9f73245be
feat(lens): track worker spend through virtual keys (#43989)
* feat(lens): bill worker analysis through virtual keys

* fix(lens): pin the verified worker image and add setup proof

* fix(lens): preserve network checks and redact billed analysis logs

* test(lens): preserve legacy worker result submission during upgrade

* fix(lens): enforce trusted worker IPs and restore coverage uploads

* docs(lens): explain trusted proxy requirements for worker allowlists

* fix(lens): yield to worker disconnects after the synthetic body
2026-10-01 09:47:30 -07:00
moe-berri
6d7d183a80
feat(lens): investigate sampled traces and retain batch results (#43942)
* fix(lens): parallelize scan analysis with bounded concurrency

* feat(lens): investigate sampled activity and preserve scan results

* fix(lens): pin the compatible investigation worker image

* fix(lens): report incomplete reviews and simplify setup validation

* fix(lens): stabilize large investigations and preserve incomplete results

* fix(lens): preserve bounded readers and distinguish counterexamples

* fix(lens): pin compatible worker and verify batched grouping cost

* fix(lens): exclude counterexamples from finding recurrence

* feat(lens): show completed scan duration in results and history

* fix(lens): fold batch selection into results navigation
2026-09-30 22:30:04 -07:00
ryan-crabbe-berri
424bfd8758
feat(e2e): record each e2e test's steps, starting with ProxyClient (#42393)
* feat(e2e): record each e2e test's steps, starting with ProxyClient

@step on a harness method records a plain-English line for every call, in
order, as repeated JUnit step properties. Labels are templates filled from the
call's parameters, like "Generate a virtual key with models: claude-haiku-4-5
and rpm limit: 3", and secret request fields are marked Field(repr=False) so
they never print. ProxyClient and the rate-limit QuotaClient carry steps first;
the other harnesses follow one area at a time. The recorder and JUnit tests run
in the Code Quality workflow's test_e2e_metadata step.

* docs(e2e): rewrite the recorded test steps guide in plain language

* fix(e2e): keep logging callback credentials out of recorded steps

* fix(e2e): mask the run's credentials in every recorded step

* fix(e2e): attach steps before the oauth failure snapshot

The failed setup or call report of an mcp_oauth_live test copied user_properties before the steps were attached, so it carried no steps. Every setup and call report now takes its properties after the steps attach

* fix(e2e): name the saved credential in its recorded step

The create_credential label read credential_info, which defaults to {} and is never set by the live callers, so the step printed nothing after 'for'. It now reads the required credential_name, and a guard fails on any label that reads a field with a default
2026-09-30 19:33:53 -07:00
moe-berri
6fd9334751
feat(lens): analyze agent activity with a separate worker (#43889)
* feat(tracing): bring current ingestion prerequisite onto main

Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.

* feat(lens): add trace analysis and standalone worker

* fix(lens): clarify review limits and finalize main integration

* fix(lens): simplify worker setup and show the next check

* fix(lens): simplify analyzer setup and resolve integration failures

* fix(lens): preserve durations and evidence from later trace reads

* fix(lens): trust server context for internal analysis exclusion

* fix(lens): pin reviewed analyzer image and verify request inclusion

* test(lens): select time units before entering custom duration

* test(lens): allow the standalone analyzer lifetime HTTP client

* test(lens): run analyzer tests in active proxy coverage shard
2026-09-30 22:42:09 +00:00
devin-ai-integration[bot]
6b9766fa0c
feat(proxy): add native ROI calculator for gateway spend vs merged PRs (#43669)
* feat(proxy): add native ROI calculator for gateway spend vs merged PRs

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): serialize ROI Prisma inputs with builtin containers

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* style(proxy): format ROI calculator backend files

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix: parse fenced ROI estimates and retain completed reports

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(roi-calculator): correct estimator and dashboard behavior

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* chore(ui): drop next dev generated AGENTS.md block

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(proxy): chunk ROI spend user lookup

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(ui): show reused ROI estimates after sync

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(security): address ROI CodeQL alerts

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(proxy): make ROI calculator unit tests discoverable

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* test(ci): run ROI calculator tests in proxy infra shard

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* fix(roi): page repository search and recover polling errors

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* feat(roi): bring scheduled analysis and guided setup into the gateway

* fix(roi): recover interrupted syncs and resolve review findings

* fix(roi): preserve cached estimates across report scope changes

* fix(roi): normalize scheduler timestamps to UTC

* fix(roi): fence cancelled syncs and read reports from writer

* fix(roi): preserve reports during metadata outages

* refactor(roi): isolate outage validation and verify uncached retry

* fix(roi): make scheduled job registration repeatable

* style(roi): format scheduler import

* fix(roi): continue syncing accessible repositories

* fix(roi): preserve reports and identity during upstream outages

* fix(roi): persist refreshed identities for reused estimates

* perf(roi): skip writes for unchanged cached identities

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-authored-by: moe-berri <moe@berri.ai>
2026-09-30 14:16:41 -07:00
yujonglee
268eb4d6e6
feat(tracing): add OTLP trace ingestion and reads (#43915) 2026-09-30 21:12:29 +00:00
devin-ai-integration[bot]
3e21e5e348
chore(codeowners): drop UI, migration, and CODEOWNERS self owners (#43653)
* chore(codeowners): drop UI and migration code owners

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(codeowners): drop CODEOWNERS self-owner line

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 01:48:22 +00:00
yuneng-jiang
3913be6b2a
ci: sync the weekly release cycle with Linear releases (#43636)
* ci: sync the weekly release cycle with Linear releases

* tmp: dry-run trigger

* tmp: backfill 1.105.0 from the rc/1.104.0 cut

* ci: drop the temporary branch trigger used to verify the Linear sync

* ci: fail on a broken rc-branch lookup and never move a just-cut release back to main

* tmp: dry-run trigger

* ci: drop the temporary branch trigger again
2026-09-28 17:00:56 -07:00
devin-ai-integration[bot]
703eb4fa68
security(proxy): keep team callback credentials out of the stored request body (#43217)
* security(proxy): keep team callback credentials out of the stored request body

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: allow the security conventional commit type

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 11:16:58 -07:00
devin-ai-integration[bot]
438bffc26e
build(rust): package the gateway container (#43471)
* build(rust): package the gateway container

* ci: exempt the gateway Dockerfile from the CI coverage gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 16:49:39 -07:00
devin-ai-integration[bot]
268e8bb735
refactor(rust): share anthropic types, request helpers, and streaming contracts across crates (#43426)
* refactor(rust): standardize Azure Messages module path

* docs(rust): define shared types crate boundaries

* refactor(rust): share request helpers and type Anthropic blocks

* docs(rust): format shared type invariants as bullets

* test(rust): parameterize repeated cases with rstest

* refactor(rust): move Responses transform result into llms

* fix(anthropic): validate chat and batch responses

* docs(rust): clarify API format ownership boundaries

* docs: clarify Rust error message construction

* refactor(auth): keep shared Rust errors provider-neutral

* refactor(rust): separate format contracts from provider policy

* fix(rust): type Anthropic chat response text collection

* fix(rust): pass audio secret sources through hosts

* fix(rust): unblock batch lint and OCR error assertions

* test(rust): assert response failures at the adapter boundary

* refactor(rust): declare error messages with typed context

* wip

* fix(rust): adapt Bedrock error details

* style(rust): cargo fmt bedrock audio transcription

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): adapt tests and dead code to typed error details

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): keep converse error contracts and read env secrets without litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(rust): raise the native wheel size gate to 45 MB

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(rust): tolerate missing usage in converse responses on the transcription route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 14:53:12 -07:00
devin-ai-integration[bot]
f12f7b5a03
test(integration): group /v1/messages contracts under tests/integration/messages_endpoint (#43352)
* test(integration): group /v1/messages contracts under tests/integration/messages

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): make ci coverage census collect nested test dirs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): nest /v1/messages contracts under messages_endpoint/providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:00:32 -07:00
devin-ai-integration[bot]
f8870b64e9
docs(pr-template): add the backport-stable label only for a P0 regression (#43351)
* docs(pr-template): add the backport-stable label only for a P0 regression

* docs(pr-template): keep a narrow security regression eligible for backport-stable

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-26 15:14:14 -07:00
ryan-crabbe-berri
96c008f420
ci: fail on new unbounded SQL IN lists and add a Prisma chunking helper (#42629)
* ci: warn on SQL IN lists with no written bound

Postgres caps a prepared statement at 32,767 bind parameters and a
membership filter binds one per value, so an IN list built from table
data breaks once the table outgrows the cap. That is how the budget reset
job froze every due budget (LIT-7535, #40564).

check_unbounded_in_lists.py reports every Prisma "in" / "not_in" filter
whose value has no fixed size and every raw SQL literal that splices a
list in after "IN (", unless the line carries "# bounded-ok: <reason>".
It only warns for now: the output is the inventory for RCA action item
AI-1, and it exits 0.

* ci: decide a constant IN list by its module binding, not its casing

An ALL_CAPS name imported or filled at runtime is as unbounded as any
other, so a name now passes only when the module binds it once to a
value of fixed size. Adds Final to the locals a loop does not forbid.

* ci: only a frozen module value makes an IN list constant

A module list bound once could still grow through append or extend, so
a name now counts as fixed only when it is bound to a tuple, frozenset
or constant. Trims the module docstring to what a reader needs.

* ci: chunk Prisma IN lists with a shared helper and fail on new unbounded ones

Add litellm.repositories.bounded_in: find_many_in, count_in, update_many_in
and delete_many_in split a deduplicated value list into 5,000-value chunks,
AND each chunk with the caller's where, run them in order (a transaction
handle works) and combine the results. Writes take a required atomicity
argument, and a where that already filters the chunked field is refused.

check_unbounded_in_lists.py now fails CI on any finding missing from
unbounded_in_baseline.txt and on any stale baseline entry, so the baseline
only shrinks. Entries are keyed by path, enclosing scope, kind, field and
occurrence, not line numbers. The helper module is exempt, a constant
spread into a frozen tuple counts as fixed, and messages point at the
helper for "in" and at an array parameter for "not_in" and raw SQL.

A real-Postgres integration test shows a raw 40,000-value filter rejected
for too many bind variables while the helpers handle it.

* refactor: rename bounded_in to chunked_in and let callers pick a chunk size

The helper module is litellm.repositories.chunked_in, and its unit and
integration tests, the checker's exemption path and its finding messages
follow the new name. The `# bounded-ok` marker is unchanged.

find_many_in, count_in, update_many_in and delete_many_in take a
keyword-only chunk_size, defaulting to IN_LIST_CHUNK_SIZE (5,000). A value
below 1 or above MAX_IN_LIST_CHUNK_SIZE (30,000) raises ValueError before
any query, which leaves the rest of the filter headroom under Postgres's
32,767 bind-parameter cap.

* refactor: flatten chunked_in's stacked comprehensions with chain.from_iterable

LIT014 (#42650) caps a comprehension at one for and one if clause. The four nested walks in the helper now chain their iterables instead, with the same order and results.

* refactor: recover user details with find_many_in, sending chunks as lists

_details_for_user_ids reads users through find_many_in instead of a raw
"in" filter, so its lookup stays under the bind-parameter cap for any
number of recovered keys. Up to 5,000 ids it still sends one find_many
with the same where dict, and a PrismaError from any chunk is still
logged and treated as no details.

The helper now sends each chunk as a list, so a chunked filter equals
the dict a hand-written call would send and a migrated call site's
existing assertions keep passing.

The site's baseline entry is gone.

* ci: skip functional TypedDict field maps in the unbounded IN list check

The dict passed as the field map of TypedDict("Name", {...}), or as its fields= keyword, names fields: an "in" or "notIn" key there is a type, not a filter. Only that dict is skipped, for TypedDict, typing.TypedDict and typing_extensions.TypedDict; a filter nested in a field value or passed to any other call is still reported. The two types/proxy/management_endpoints/team_endpoints.py entries leave the baseline, which is now 156.

* fix: refuse an update_many_in whose data writes the chunked field

Chunks run one after another, so an update that sets the chunked field can move a row into a later chunk, which updates it again and counts it twice: values ["old", "new"] with chunk_size=1 and data={"id": "new"} does exactly that. update_many_in now raises ChunkedFieldWriteError before any query when data has the chunked field as a top-level key, in any form, including Prisma operators such as {"set": ...}.

* docs: cut the unbounded IN list checker's docstring to what it flags and how to clear it

It now says what is reported, the three ways to clear a finding, and how the baseline and --update-baseline work, in 11 lines. The per-shape detail lives in the tests.

* ci: key an unbounded IN list finding by its filtered expression too

A baseline key of path, scope, kind, field and occurrence let a PR delete
a baselined filter and add a different unbounded one on the same field in
the same function, and the new one took over the old key. The key now
also carries the filtered expression's source, whitespace-normalized
(the Prisma value, or a raw-SQL `IN (...)` slot), so that swap reads as
one new and one stale entry and fails the run. The same expression
re-added in the same function is still the same finding.

Every baseline entry is rewritten in the new form; the 156 findings are
unchanged, and only occurrence indexes renumber where one field had
several different expressions.
2026-09-26 13:40:44 -07:00
devin-ai-integration[bot]
99655b6f86
test: finish the non-proxy half of tests/test_litellm (#43281)
* test: move key-gated tests/test_litellm SDK tests into tests/llm_translation and drop empty folders

* test: make token counter and health check unit tests run offline

* ci: point unit shards, rust path filter, Makefile and docs at tests/unit

* docs: fix stale test_litellm run paths in moved llm_translation tests

* fix: correct databricks e2e sys.path depth and contributing example path

---------

Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 22:43:41 -07:00