* feat(lens): isolate trace storage and investigation in a Rust service
* fix(lens): include Rust sources in the image build context
* feat(lens): wire service setup, scoped delivery receipts and lease attempts
* fix(lens): complete service routing and reject stale investigation results
* fix(lens): retry key propagation and validate isolated Compose setup
* fix(lens): seed through isolated ingestion and preserve upstream queue fixes
* chore: sync schema.prisma copies from root
* fix(lens): bind nullable due timestamps as text for Prisma
* chore(ui): remove stale lint suppressions
* fix(lens): address CI failures and review findings
* refactor(lens): remove retired Python worker and run evaluations in Rust
* fix(lens): reuse control connections and satisfy review checks
* test(lens): install and upgrade both Helm charts on Kubernetes
* test(lens): run connection reuse coverage as an integration test
* fix(ui): upgrade Next.js to 16.3.8 security release
* fix(lens): fence stale attempts and preserve reviewed evidence
* Revert "fix(ui): upgrade Next.js to 16.3.8 security release"
This reverts commit 2f79a51b25.
* fix(lens): stop failed investigations and stream history excerpts
* test(lens): cover model tool and result contracts
* test(lens): fix retired routes and reuse installation build artifacts
* test(lens): use portable grep in Helm installation smoke
* test(lens): wait for migrations before forwarding Helm services
* fix(lens): keep failed evidence reads retryable
* fix(lens): preserve sandbox output during process exit
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* test(e2e): move the harness self-tests out of tests/e2e
The nightly Buildkite run copies tests/e2e into the runner image and runs
bare pytest, so the 672 tests of the harness itself (fixture parsing, JUnit
properties, the stack lock, the load aggregators, the Claude Code driver)
counted as e2e tests on the status page even though none of them reaches a
proxy. They now live in tests/e2e_harness, mirroring the tests/e2e layout,
and run in the GitHub Actions lint job and the CircleCI
provider_replay_harness job instead
* fix(ci): point the providers replay controls at tests/e2e_harness
The providers integration job still selected the four replay-control
tests under tests/e2e/test_provider_edge.py, so pytest exited before
they ran. The raw-HTTP check's file walk also drops to one loop per
comprehension
* style(tests): mark the raw-HTTP check's bindings Final
* feat(lens): add dataset case and size limits
* feat(lens): add dataset, case and build models
* feat(lens): build dataset cases from traces, findings and text
* feat(lens): store dataset revisions insert-only
* feat(lens): add dataset routes for build, save, export and eval cases
* feat(lens): mount the dataset router before lens routes
* feat(lens): add LiteLLM_LensDataset table
* feat(lens): add LiteLLM_LensDataset table to proxy schema
* feat(lens): add LiteLLM_LensDataset table to extras schema
* feat(lens): add migration that creates the dataset table
* test(lens): cover dataset case building, dedupe and limits
* test(lens): cover dataset revisions, conflicts and eval cases
* chore(ui): regenerate API types for lens datasets
* feat(lens): add dataset UI types
* feat(lens): add datasets API client
* feat(lens): add dataset query and mutation hooks
* feat(lens): add case selection and expected edit logic
* test(lens): cover case selection and expected edits
* feat(lens): add the add to dataset dialog
* test(lens): cover saving picked cases from the dialog
* feat(lens): add datasets list
* feat(lens): add dataset detail with revisions and export
* test(lens): cover editing, revisions and export in datasets tab
* feat(lens): expose datasets on the lens API
* feat(lens): add in-memory datasets for demo mode
* feat(lens): wire demo datasets into the demo lens API
* feat(lens): add optional lens API hook
* feat(lens): add optional onboarding hook
* feat(lens): add datasets tab and dataset routing
* feat(lens): show the datasets tab
* feat(lens): add to dataset from the trace header
* feat(lens): add a single turn to a dataset from a step
* feat(lens): add finding evidence to a dataset
* docs(lens): add datasets screenshots for the PR
* test(lens): cover dataset revision storage against Postgres
* test(lens): cover dataset trace paging, findings and route errors
* test(lens): cover dataset build fallbacks and no_content skips
* fix(lens): register the dataset table for postgres span names
* feat(lens): add invalid skip reason for unparseable case lines
* fix(lens): keep valid JSONL cases, provenance and size limits; stop reading past the case cap
* test(lens): cover malformed JSONL, re-import provenance, size fields and early cap
* chore(ui): regenerate API types for the invalid skip reason
* feat(lens): show text for the invalid skip reason
* feat(lens): render datasets in the same inspector table as investigations
* test(lens): open a dataset by clicking its table row
* docs(lens): update the datasets list screenshot
* style(lens): format the datasets table
* fix(lens): normalize JSON text span input and output into messages when building cases
* test(lens): cover JSON text span normalization for dataset cases
* Revert "fix(lens): normalize JSON text span input and output into messages when building cases"
This reverts commit 29c3e34962fa12427dbd516c08f2599db997a905.
* fix(traces): normalize agent assistant summaries into UI messages
* feat(lens): add dataset case view helpers built on the trace parsers
* test(lens): cover dataset case view helpers
* feat(lens): show dataset cases in an inspector table
* feat(lens): open a dataset case in a side panel with trace message cards
* feat(lens): rebuild the dataset page header and layout
* feat(lens): keep the open dataset case in the URL
* test(lens): drive dataset edits through the case table and panel
* docs(lens): update dataset view screenshots
* Revert "test(lens): cover JSON text span normalization for dataset cases"
This reverts commit 30229ae8ae90265c09d755c43c1f1eeaa66050c7.
* fix(lens): import StateMessage from shared in dataset detail
* fix(lens): import StateMessage from shared in datasets list
* fix(lens): give the datasets migration a unique timestamp after review checkpoints
* fix(lens): use async-timeout on python 3.10 for budget reservation timeouts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(deps): keep uv.lock diff to the async-timeout entry
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(lens): cover real request deadline expiry in reserved_budget
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lens): finish Python 3.10 timeout coverage and dependency checks
* test(lens): control event-loop time for deadline regressions
* test(tracing): include priced call count in trace fixture
* fix(lens): limit timeout compatibility changes to PR scope
---------
Co-authored-by: Moe Khalil <moe@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(lens): restate response contract during model repair
* fix(lens): separate instructions and recover rejected results
* fix(lens): correct loop type annotations and checks
* fix(lens): preserve access to prior findings after compaction
* fix(lens): cap result retries and preserve partial completion
* ci: split slow unit shards and build the Rust bridge once per run
* ci: key the Rust bridge cache on source files only
* ci: keep the unit setup ceiling unchanged with the shared Rust bridge
* ci: fall back to the Cargo cache when the Rust bridge artifact is missing
* ci: keep reruns on enterprise-routing for the prompt caching flake
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* feat(lens): add per-trace review models to jobs and progress
* feat(lens): append worker reviews to the job, capped, and count every review
* feat(lens): report a review with reasoning for each screened trace
* chore(ui): regenerate api types for lens job reviews
* feat(lens): type job reviews and fill them in lens fixtures
* feat(lens): add live review playback model
* feat(lens): pick the analysis model and slow single-review pacing
* feat(lens): add sample reviews for previewing the live run
* feat(lens): add live run layout with queue, reading trace and conclusions
* feat(lens): show the live run on investigations and open it from run now
* feat(lens): stream large review backlogs at 150ms or less and list newest first
* fix(lens): show the live run only for real reviews and keep fixtures test-only
* refactor(lens): restyle the live run as the native progress panel
* fix(lens): retry contended investigation updates with jittered backoff
* feat(lens): add a reading ticker line and replay for finished runs
* feat(lens): collapse the live run to an ambient line with show work
* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size
* feat(lens): format review span previews as readable messages
* feat(lens): derive strip status, honest issue counts and drawer focus from a job
* feat(lens): track active jobs before their first review
* feat(lens): add a live trace results drawer with readable spans
* feat(lens): put the live strip under the progress bar and drop the inline panel
* feat(lens): add an ambient live strip that opens the drawer
* fix(lens): wait out provider rate limits and retry model calls four times
* style(lens): format repository contention tests
* feat(lens): read review spans as a conversation timeline
Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.
* fix(lens): list recorded agents in the run now dialog
The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.
* feat(lens): pace live playback so each trace stays readable
Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.
* feat(lens): keep the live run ambient until View run is clicked
The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.
* feat(lens): show the live run as a two-pane trace and conclusions view
Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.
* fix(lens): run several investigations per worker and poll every two seconds
* feat(lens): add worker slot and poll interval settings
* feat(lens): add list summaries and an incremental review filter
* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately
* test(lens): cover list summaries, review polling and review access
* feat(lens): explain why a queued investigation is waiting
Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.
* feat(lens): show the queue reason and what the worker is doing in the live strip
The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.
* feat(lens): add a review page model carrying the total reviewed count
* fix(lens): page live reviews by index so out-of-order reviews are never skipped
* feat(lens): take an index cursor on the reviews endpoint
* test(lens): cover index cursors across out-of-order and rolled-over reviews
* chore(ui): regenerate api types for the lens reviews endpoint
* feat(lens): page job reviews by index cursor
Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.
* feat(ui): add a lens reviews query that polls the index cursor while live
* fix(lens): feed the live run from the reviews endpoint and keep View run open
LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.
* fix(lens): split live conclusions into issues and patterns
A check could show up twice with the same label, once as an issue and once as a pattern.
* fix(lens): group live conclusions by check with short labels
There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.
* feat(lens): simplify View run to traces and conclusions
The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.
* refactor(lens): drop client-side replay in favour of real in-flight rows
Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.
* feat(lens): show what the worker is reading and make View run obvious
Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.
* feat(lens): sum up a finished live run with time taken
doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.
* feat(lens): slide one model rectangle over the traces being read
A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.
* feat(lens): add in-flight runs to jobs and worker progress
* feat(lens): store in-flight runs from progress and clear them when a job ends
* refactor(lens): route progress, cancel and results through shared job transitions
* feat(lens): report each run as in flight when its review starts
* feat(lens): send in-flight runs with worker progress
* test(lens): cover in-flight runs across progress, old workers and terminal states
* test(lens): cover in-flight reporting under original run ids
* chore(ui): regenerate api types for lens in-flight runs
* feat(lens): model live reading lanes from in-flight runs and reviews
* feat(lens): show a now reading stage that types each trace's reasoning
* feat(lens): put the now reading stage above the trace list in View run
* fix(lens): resolve the analysis provider logo from the model catalog
* fix(lens): give demo jobs an empty in-flight list
* style(lens): format endpoint tests
* refactor(lens): name the run now handler in investigations view
* refactor(lens): name now reading conditions
* refactor(lens): name inline objects in the live run
* style(lens): format live run files
* fix(lens): keep worker settings inside the standalone worker package
* refactor(lens): keep update retry settings next to the repository
* fix(lens): start review history over when a run is reclaimed
* chore(lens): drop the unused review fixture
* refactor(lens): remove dead live helpers and use generated in-flight types
* fix(lens): keep polling a finished run until its last reviews arrive
* perf(lens): tick fast only while reasoning is typing
* fix(lens): isolate retried reviews and finding identities
* fix(lens): space the model name in run summary
* feat(lens): integrate confined workspace analysis with live reviews
* fix(lens): synchronize confined Python process monitoring
* Update review.md
* fix(lens): allow mixed context capacities and correct review assertions
* fix(lens): retrieve evidence on demand and isolate failed reviews
* fix(lens): isolate incomplete evidence reads from peer reviews
* test(lens): await trace status filter option
* test(lens): wait for reclaimed review state to settle
* fix(lens): recover from incomplete cross-session evidence
---------
Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
* ci: run unit selections from GHA test-path and drop the CircleCI unit jobs
* ci: keep existing shard token order and header note
* ci: throwaway, drop tests/unit/repositories from the unit shard to show assert-ci-coverage fails
* ci: revert throwaway assert-ci-coverage check
* ci: stop crediting --ignore paths as invoked in assert_ci_coverage
* ci: install the caching, extra_proxy and proxy-runtime extras in the GHA unit sync
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* test(tracing): pin HTTP request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(tracing): generate existing HTTP request models from Rust schemas
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(tracing): bind trace query params to generated request models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci(rust): raise the native wheel size gate to 48 MB
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* perf(tracing): read the trace list clock without a thread-pool dispatch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(traces): share the trace page-size bounds between schema and reader
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(ui): type trace request queries against the generated OpenAPI schema
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs: encode the trace contract boundary in AGENTS.md
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(ui): format trace request aliases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: move Postgres, MCP and Redis suites to CircleCI integration
* ci: throwaway, drop tests/proxy_behavior from its CircleCI job to show assert-ci-coverage fails
* ci: revert throwaway assert-ci-coverage check
* ci: keep the e2e helpers the gate tests still use
* ci: move the roi-database Postgres shard to CircleCI integration
* ci: run redis-compat without CircleCI's Azure and cassette env, cover postgres_suite test_path
* ci: match the GitHub env for the moved Postgres and Redis jobs
* ci: unset provider keys in the CircleCI MCP job and drop unused e2e-stack helpers
---------
Co-authored-by: yuneng <yuneng@berri.ai>
* fix(lens): pin worker dependencies and support approved image digests
* fix(lens): include locked dependencies and release identity in build context
* fix(lens): select the dev worker package for SHA-tagged charts
* feat(ui): prototype observed engineering ROI dashboard
* feat(roi): replace effort estimates with measured repository metrics
* fix(roi): finish connection recovery and generated API contracts
* fix(roi): show merged changes before accounts are linked
* fix(roi): preserve selected report tab across refreshes
* fix(roi): recover app authorization and keep detail values readable
* fix(roi): reuse the shared OAuth HTTP client
* feat(roi): combine providers and compare equal reporting periods
* docs: explain ROI metrics for first-time readers
* fix(roi): preserve connections and scheduled reports during setup
* ci(roi): assign database contracts to the active Postgres shard
* fix(roi): preserve issue counts and normalized connections
* fix(ui): compact ROI dashboard header and metrics
* fix(ui): show ROI repository count with expandable list
* fix(ui): wrap ROI controls within narrow panels
* fix(roi): restore sample report preview and simplify setup
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register typesafe as a provider so Jev deployments load
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(decisions): move provider endpoints under llms and validate proxy bodies
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add Cloudflare Clef and Strands Decider backends
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register decisions routes for managed agents and gateway
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(decisions): use raw regex for cloudflare missing account match
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): avoid cast in Cloudflare response unwrapping
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): default model, evaluation health probe, short Cloudflare names
The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.
* fix(decisions): let health_check_params override the evaluation probe
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): audit the decisions endpoint across providers, limits, health and chaos
Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).
The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.
* fix(decisions): send env API keys to a configured api_base
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(decisions): register the routes through the lazy feature registry
The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.
The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.
* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell
* fix(proxy): let a config pass-through beat a lazily registered route in eager mode
With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.
---------
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* feat(dev): seed linked tracing and spend fixtures
* chore(dev): use OpenAI model in tracing config
* chore(dev): align tracing credentials with UI E2E
* fix(dev): update fixture seeder query scope
* wip
* wip
* wip
* chore(trace): checkpoint ongoing Rust migration
* refactor(trace): group Python bridge under trace package
* refactor(traces): read span conventions through a Convention trait
Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.
The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.
* feat(trace): export Rust-owned wire schemas and enforce contract bounds
* fix(trace): bound quoted counts in ClickHouse wire schemas
* feat(trace): generate Python wire contracts with datamodel-code-generator
* test(trace): validate migrated callers and generated contracts at the native boundary
* refactor(traces): rename normalization convention to format
* fix(traces): reconcile spend evidence and preserve unknown costs
* feat(traces): normalize additional telemetry formats
* test(traces): cover captured normalization fixtures
* refactor(traces): isolate SDK normalization rules
* feat(tracing): seed all trace exports for local dashboard
* fix(clickhouse): preserve custom LiteLLM request metadata
* docs(traces): define normalization module boundaries
* docs(traces): define resolution and OTLP boundaries
* fix(ui): normalize nullable trace message names
* refactor(traces): split resolver modules and cover resolution behavior
* test(traces): replace normalization snapshots with behavior assertions
* fix(ui): align dashboard API contracts with generated types
* refactor(traces): type normalization and storage boundaries
* fix(traces): seed captured SDK spend and preserve provider identities
* wip
* test(traces): verify guide discovery and content ordering
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(roi): support GitLab and tagged branch costs
* fix(roi): count tagged branches independently of estimation status
* test(roi): capture live GitHub and GitLab report validation
* fix(roi): open estimate details at the start
* fix(roi): clarify cost views and unify report layout
* feat(roi): showcase per-PR costs in the sample report
* fix(roi): separate report tabs and preserve branch cost attribution
* fix(roi): preserve demo previews and align progress spacing
* fix(roi): isolate demo loading and parallelize fork lookups
Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression
* fix(roi): separate demo and live loading states
Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers
* fix(roi): ignore refreshes from a previous source
* fix: trust gateway context for ROI estimator exclusion
* fix: preserve historical ROI estimator exclusion
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run search_endpoints tests in proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(llm_http_handler): keep provider error text when re-raising mapped errors
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): allow promptless image edits and default search models
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): default missing image edit image to None
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(proxy): build image edit defaults without mutating request data
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): inject a fake router for the search default model test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on vector store and file lookup handlers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): keep the lookup handler raise block to a single statement
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover missing required body params and provider lookup status codes
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): bind spend-row request id with partial to satisfy B023
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): only reject non-positive page_size on vector store list
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): remove unreachable fine-tuning body validation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover streaming anthropic messages reaching the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): count only provider calls when asserting missing params never reach the upstream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve merge-base request compatibility
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): preserve interaction completion model defaults
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* wip
* wip
* test(traces): separate root status from diagnostic error counts
* test(traces): cover normalization precedence and fallbacks
* chore(cache): remove stray comments from trace PR
* test(traces): name lens test for shared query path
* fix(traces): place query implementation before test module
* test(traces): use unified read scope in migration tests
* ci(rust): allow feature checks to finish
* ci(mcp): allow dependency resolution to finish
* fix(traces): preserve key visibility and safe spend attribution
* test(e2e): jwt auto_register map-existing-key repro
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(jwt): auto_register_map_existing_key maps JWT to the user's existing virtual key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(jwt): exclude blocked keys from auto_register_map_existing_key reuse
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(jwt): route existing-key lookup through VerificationTokenRepository
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): stop requiring LITELLM_SALT_KEY for the owned JWT gateway
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): gate the owned JWT gateway tests behind E2E_OWNED_GATEWAY
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(jwt): only reuse keys that can call LLM routes in auto_register_map_existing_key
Skip Admin UI session keys and keys whose allowed_routes restrict them to
anything other than llm_api_routes (management, read_only, password-reset
sessions). Mapping a JWT to one of those left the user with 401s or 403s on
every LLM call, since the mapping persists.
* fix(jwt): scope auto_register_map_existing_key reuse to the JWT-resolved team
Only reuse a key whose team_id matches the team auth_builder resolved for
the JWT (no team matches no team), so a personal key can no longer bypass
the resolved team's model and budget limits.
With the flag on, the first JWT request now falls through to the same
virtual-key checks later mapped requests get, instead of returning early,
so a reused key's own limits apply from request one rather than 200 then
403. Flag off keeps the early return unchanged.
* fix(jwt): keep the early return when no master key is set
Without a master key the generic virtual-key path returns a bare
INTERNAL_USER object, so falling through on the first auto-registered
request dropped the key's team, models and budgets. Only fall through when
a master key is configured.
Tests now assert the reused key per team rather than the query shape, and
cover the flag-off early return and the no-master-key case.
* test(jwt): assert on race-loser's returned key, not only mocks (TQ002)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(jwt): close the auto_register_map_existing_key race, shared-claim and expiry holes
A key auto_register just minted is never adopted by a concurrent request, so the race loser's cleanup can no longer delete a key another request mapped and cascade its mapping away (503, user left with no key)
Reuse only happens when the claim value is the JWT-resolved user_id. A shared claim such as azp or client_id falls back to minting, so one user can no longer land on another user's personal key and budget
Only keys that never expire are reused, so an expiring key can no longer pin the claim to a permanent 401
Integration tests on a real proxy and Postgres cover all three. The race test holds the first mapping insert in a Postgres relay, so the interleaving is forced rather than timed. The where-clause shape unit tests are replaced by these, since only a real database proves the filter
* test(e2e): create the reused key in the team the JWT resolves to
The flag only reuses a key in the JWT-resolved team, and this identity's groups claim resolves to its team, so a teamless key was never eligible and the test could not pass
* test(integration): match the held statement across TCP reads
The relay looked for the trigger inside one read, so an insert split across two reads was never held and the race test would fail waiting for it. It now matches one exact trigger over a window that keeps the end of the previous read
* fix(jwt): gate key reuse on the claim field, not on the claim value
Requiring the claim value to equal the resolved user_id skipped reuse for users matched through the sso_user_id or case-insensitive email fallback, whose stored user_id differs from the JWT sub. That is the lookup LIT-5378 asks for. Reuse is now allowed when the virtual key claim is the user_id or user_email JWT field, globally or for the token's issuer, which still keeps shared claims such as azp or client_id on the mint path
* fix(jwt): let an issuer's own user field replace the global one when gating key reuse
An issuer that identifies users by uid no longer treats the global sub field as a user identity claim, so a shared sub under that issuer mints instead of reusing a personal key
* test(jwt): make the flag-off test fail when the flag no longer gates key reuse
The flag-off test used a config where sub was not a user identity claim, so deleting the flag check still passed. Configure user_id_jwt_field=sub so only the flag keeps the lookup off, and drop test docstrings
* chore(lint): drop mutable-ok suppressions that LIT013 flags as no-ops
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Mrinal Chanshetty <mchanshetty@Mrinals-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test(proxy): delete the legacy proxy test tree and serve the redirect test from loopback
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): exercise the shard check directly for unit_selection-owned children
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): serve the redirect test from respx instead of a socket
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): credit shard ownership only to unit flags wired in gha
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): implement the wired-flag shard crediting the tests assert
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): split the root proxy test files into their own unit shard
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): point the rate-limit skip reason at the usage-based-routing-v2 RPM tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move proxy_server, _experimental and db tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): keep tuple identity in proxy state restore and fix misc target paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move utils, agent_endpoints and endpoint tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved unit test directories
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): exclude proxy-db-owned files from the misc target
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): drop the redundant fixture docstrings in the proxy conftest
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(lens): remove deployment screenshots
* refactor(lens)!: rename internal engine package and API
* fix(lens): pin worker image for renamed API
* test(lens): cover fresh and populated rename migrations
* fix(lens): protect db-push upgrades and restore routing and CI
* fix(lens): resolve migration tables across schemas and include database driver
* feat(s3_v2): add s3_partition_granularity option for hourly S3 folders
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover s3 v2 partition granularity across surfaces, settings and chaos
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): cover previous_response_id history rebuilt from an hourly cold storage object
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(s3_v2): reuse the cold storage key only when s3_v2 owns cold storage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): cover hour rollover, postgres outage, in-flight switches, key/team vars and real S3 layout
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(liccheck): authorize libfaketime, the GPLv2 dev-only clock the s3 rollover integration test preloads
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): wait for the rejected-request cell's payloads by id, not by line count
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): declare the postgres outage cell's models in config and trip the relay on burst ids
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): drop the libfaketime hour rollover cell and its dev dependency
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(e2e): deselect the s3_v2 live e2e on the stage-mirror stack
The stage-mirror config enables no s3_v2 callback, so every test in test_s3_log_e2e.py fails its readiness check there. The file keeps running in the Buildkite e2e lane, which configures s3_v2
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(s3_v2): declare the sink outage burst models in config
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(s3_v2): read cold storage metadata without an empty dict default
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): wait for the proxy to reconnect before the postgres outage recovery request
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub HIBP through respx by disabling the aiohttp transport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): share the httpx transport fixture across proxy unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore proxy globals without a missing-value sentinel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): isolate the mcp server manager per test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): move management_endpoints, management_helpers and guardrails tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): reuse the shared httpx transport fixture in moved proxy tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub outbound HTTP and package moved test dirs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore the config server hostname in the mcp resolution test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): pin the completion tokenizer model in the straiker screening test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(ci): add used_client_oauth_token to the GCS pub/sub spend-log golden
#43063 stamps used_client_oauth_token into spend-log metadata, so
test_async_gcs_pub_sub_v1 failed on main with an extra metadata key
* test(ui): give the auto-router threshold save wait room for the availability debounce
#42625 keeps Save disabled while a 300ms-debounced availability check runs.
This test waits for Save right after the change, so the whole debounce lands
inside waitFor's 1s default and it times out under CI load. It is the
recurring UI Unit Tests failure on main since #42625 landed
* test(e2e): expect no pricing tier on bills for streamed calls OpenAI served at default
#42870 added both the rule that a served default or standard tier bills at
base pricing and records no service_tier, and streamed tests expecting the
row to record 'default'. They have failed on every scheduled litellm-e2e run
since. The tests now map the served tier to the pricing basis the bill must
record and check input is billed at that basis's rate; the messages case
registers custom rates so the rate check has something to compare against
* test(e2e-ui): wait for the call-id search before hovering the logs row
The row the spec hovers is already on the unfiltered first page, so it was
found before the search request returned. The search response then
re-rendered the table under the mouse, and the Base UI tooltip never opened.
Reproduced with Playwright against a local proxy: hovering right after the
fill never shows the tooltip, hovering after the search response shows the
call id every time
* test(e2e): run the Together structured-output case on the hybrid Qwen with reasoning off
The case picked the cheapest Together row flagged supports_response_schema.
DeepSeek-V4-Flash-0731 hit its cost-map deprecation date on 2026-09-29, so the
pick moved to GLM-5.3-Flash, a reasoning-only model that spends the 1024-token
budget thinking and returns content=None. Qwen3.5-9B is the pinned hybrid model
the reasoning_effort=none case already exercises, and Together lists it with
structured output support
* test(integration): read the agent 365 guardrail status by its own name in spend logs
The MCP shard runs under xdist against one database, and a sibling file creates a
default_on pre_mcp_call content filter there. The owned proxy reloads DB guardrails, so
that filter's 'success' entry could land first in guardrail_information and the test
read it instead of the agent 365 verdict
* test(unit): ignore asyncio's leaked-task records in the budget limiter push-failure log check
gc.collect() inside the caplog window can collect a pending task an earlier test left
on a closed loop, and asyncio logs 'Task was destroyed but it is pending' into this
test's records. The check still counts every LiteLLM logger, and unretrieved task
exceptions on this loop still go through the asserted exception handler
* test(e2e-ui): fill the create-tag fields inside the dialog
#42949 added 'Filter by tag name' and 'Filter by description' inputs to the Tag
Management page, so page-wide getByLabel('Tag Name') and getByLabel('Description')
match two elements and Playwright's strict mode fails the create step
* test(integration): run integration proxies with the CI license
Multi-worker proxies start each uvicorn worker in a fresh process, so every
worker reads the license from its environment. Forward LITELLM_LICENSE into the
proxy and test runner environments
* ci: save GitHub Actions caches only from main and bump codecov-action to 5.5.5
Every pull request saved its own uv, maturin, Rust and Prisma caches, about
4.5 GB per PR, so the repository's 10 GB cache budget evicted main's entries
within minutes. Pull request jobs then missed every cache, downloaded all
dependencies from PyPI and hit the install step timeouts. Pull requests now
restore only, and main keeps the caches warm for them. test-linting and
check-ui-api-types run only on pull requests and keep saving
codecov-action 5.5.4 imports its signing key from the deleted codecovsecurity
keybase account, so every upload failed signature verification. 5.5.5 reads it
from codecovsecops; the key ID matches the one signing the current CLI
* test(unit): join the session-minting thread before collecting the handler
asyncio.to_thread resumes the test as soon as the worker sets its result,
while the pool thread can still hold the work item and through it the
handler. gc.collect() then cannot finalize the handler and the session stays
open. A pool that shuts down before the test continues drops that reference
* test(integration): relaunch owned proxies that lose their port, expire idle gateway connections early
owned_proxy_process released its reserved port and the proxy bound it only
after full startup, so another xdist worker or an outgoing connection could
take it first and the proxy exited with 'address already in use'. The launch
now retries on a fresh port when that happens and stops every failed attempt.
uvicorn closes idle keep-alive connections after 5 seconds and httpx expired
them at the same 5 seconds, so a request sent right at that mark could reuse a
socket the server was closing and get 'Connection reset by peer'. Gateway
clients now drop idle connections after 2 seconds
* ci(circleci): give the base SDK wheel build the same 30 minute no-output window as the Windows build
The release profile builds with fat LTO and one codegen unit, so the final
link of litellm-cache-s3 runs silently for minutes. Successful builds take
711 to 749 seconds, right at the default 10 minute no-output limit, and about
30% of recent runs were killed there
* test(integration): model the budget-reset database outage as 10 seconds instead of 5 refused connections
The proxy retries the database about every 30 seconds and each retry opens
roughly one connection, so a 5-connection outage took 3 to 4 retries to clear
and recovery landed between 60 and 90 seconds, straddling the test's 80 second
reset window. A fixed 10 second outage still refuses the immediate reconnect
and recovers on the next retry
* ci: move the unit-test uv cache split into a composite action
check_workflow_startup_safety sums every setup step's timeout, so the save and
restore variants each counted 5 minutes although only one runs. One composite
step keeps the setup ceiling at 35 minutes
* test(unit): point tiktoken at the bundled cache for every unit test
The rust_bridge tokenizer tests loaded o200k_base before any test in their
xdist worker had imported default_encoding, so tiktoken fell back to the
temp cache and tried to download under pytest-socket. Move the session
fixture from litellm_core_utils/conftest.py to the root unit conftest.
* test(integration): answer model discovery probes in the hosted_vllm wire tests
The router's periodic upstream model info refresh sends GET /v1/models to
hosted_vllm deployments, so a wire server that is live during a refresh
sees an extra request. Answer the probe with an empty model list and leave
it out of the provider-call assertions, matching the responses bridge
tests.
* test(proxy): move auth, hooks, policy_engine and client tests into tests/unit/proxy
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): stub HIBP through respx by disabling the aiohttp transport
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): share the httpx transport fixture across proxy unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): restore proxy globals without a missing-value sentinel
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): package moved dirs and stub the login breach check at the HTTP boundary
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): isolate the mcp server manager per test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* feat(e2e): record each e2e test's steps, starting with ProxyClient
@step on a harness method records a plain-English line for every call, in
order, as repeated JUnit step properties. Labels are templates filled from the
call's parameters, like "Generate a virtual key with models: claude-haiku-4-5
and rpm limit: 3", and secret request fields are marked Field(repr=False) so
they never print. ProxyClient and the rate-limit QuotaClient carry steps first;
the other harnesses follow one area at a time. The recorder and JUnit tests run
in the Code Quality workflow's test_e2e_metadata step.
* docs(e2e): rewrite the recorded test steps guide in plain language
* fix(e2e): keep logging callback credentials out of recorded steps
* fix(e2e): mask the run's credentials in every recorded step
* fix(e2e): attach steps before the oauth failure snapshot
The failed setup or call report of an mcp_oauth_live test copied user_properties before the steps were attached, so it carried no steps. Every setup and call report now takes its properties after the steps attach
* fix(e2e): name the saved credential in its recorded step
The create_credential label read credential_info, which defaults to {} and is never set by the live callers, so the step printed nothing after 'for'. It now reads the required credential_name, and a guard fails on any label that reads a field with a default
* feat(tracing): bring current ingestion prerequisite onto main
Port the prerequisite implementation from BerriAI/litellm#43915 at 5aacd57455 so Lens does not depend on the retired tracing stack.
* feat(lens): add trace analysis and standalone worker
* fix(lens): clarify review limits and finalize main integration
* fix(lens): simplify worker setup and show the next check
* fix(lens): simplify analyzer setup and resolve integration failures
* fix(lens): preserve durations and evidence from later trace reads
* fix(lens): trust server context for internal analysis exclusion
* fix(lens): pin reviewed analyzer image and verify request inclusion
* test(lens): select time units before entering custom duration
* test(lens): allow the standalone analyzer lifetime HTTP client
* test(lens): run analyzer tests in active proxy coverage shard
* chore(codeowners): drop UI and migration code owners
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(codeowners): drop CODEOWNERS self-owner line
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: sync the weekly release cycle with Linear releases
* tmp: dry-run trigger
* tmp: backfill 1.105.0 from the rc/1.104.0 cut
* ci: drop the temporary branch trigger used to verify the Linear sync
* ci: fail on a broken rc-branch lookup and never move a just-cut release back to main
* tmp: dry-run trigger
* ci: drop the temporary branch trigger again
* security(proxy): keep team callback credentials out of the stored request body
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: allow the security conventional commit type
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: oliver <oliver@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* build(rust): package the gateway container
* ci: exempt the gateway Dockerfile from the CI coverage gate
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): group /v1/messages contracts under tests/integration/messages
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): make ci coverage census collect nested test dirs
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(integration): nest /v1/messages contracts under messages_endpoint/providers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* docs(pr-template): add the backport-stable label only for a P0 regression
* docs(pr-template): keep a narrow security regression eligible for backport-stable
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* ci: warn on SQL IN lists with no written bound
Postgres caps a prepared statement at 32,767 bind parameters and a
membership filter binds one per value, so an IN list built from table
data breaks once the table outgrows the cap. That is how the budget reset
job froze every due budget (LIT-7535, #40564).
check_unbounded_in_lists.py reports every Prisma "in" / "not_in" filter
whose value has no fixed size and every raw SQL literal that splices a
list in after "IN (", unless the line carries "# bounded-ok: <reason>".
It only warns for now: the output is the inventory for RCA action item
AI-1, and it exits 0.
* ci: decide a constant IN list by its module binding, not its casing
An ALL_CAPS name imported or filled at runtime is as unbounded as any
other, so a name now passes only when the module binds it once to a
value of fixed size. Adds Final to the locals a loop does not forbid.
* ci: only a frozen module value makes an IN list constant
A module list bound once could still grow through append or extend, so
a name now counts as fixed only when it is bound to a tuple, frozenset
or constant. Trims the module docstring to what a reader needs.
* ci: chunk Prisma IN lists with a shared helper and fail on new unbounded ones
Add litellm.repositories.bounded_in: find_many_in, count_in, update_many_in
and delete_many_in split a deduplicated value list into 5,000-value chunks,
AND each chunk with the caller's where, run them in order (a transaction
handle works) and combine the results. Writes take a required atomicity
argument, and a where that already filters the chunked field is refused.
check_unbounded_in_lists.py now fails CI on any finding missing from
unbounded_in_baseline.txt and on any stale baseline entry, so the baseline
only shrinks. Entries are keyed by path, enclosing scope, kind, field and
occurrence, not line numbers. The helper module is exempt, a constant
spread into a frozen tuple counts as fixed, and messages point at the
helper for "in" and at an array parameter for "not_in" and raw SQL.
A real-Postgres integration test shows a raw 40,000-value filter rejected
for too many bind variables while the helpers handle it.
* refactor: rename bounded_in to chunked_in and let callers pick a chunk size
The helper module is litellm.repositories.chunked_in, and its unit and
integration tests, the checker's exemption path and its finding messages
follow the new name. The `# bounded-ok` marker is unchanged.
find_many_in, count_in, update_many_in and delete_many_in take a
keyword-only chunk_size, defaulting to IN_LIST_CHUNK_SIZE (5,000). A value
below 1 or above MAX_IN_LIST_CHUNK_SIZE (30,000) raises ValueError before
any query, which leaves the rest of the filter headroom under Postgres's
32,767 bind-parameter cap.
* refactor: flatten chunked_in's stacked comprehensions with chain.from_iterable
LIT014 (#42650) caps a comprehension at one for and one if clause. The four nested walks in the helper now chain their iterables instead, with the same order and results.
* refactor: recover user details with find_many_in, sending chunks as lists
_details_for_user_ids reads users through find_many_in instead of a raw
"in" filter, so its lookup stays under the bind-parameter cap for any
number of recovered keys. Up to 5,000 ids it still sends one find_many
with the same where dict, and a PrismaError from any chunk is still
logged and treated as no details.
The helper now sends each chunk as a list, so a chunked filter equals
the dict a hand-written call would send and a migrated call site's
existing assertions keep passing.
The site's baseline entry is gone.
* ci: skip functional TypedDict field maps in the unbounded IN list check
The dict passed as the field map of TypedDict("Name", {...}), or as its fields= keyword, names fields: an "in" or "notIn" key there is a type, not a filter. Only that dict is skipped, for TypedDict, typing.TypedDict and typing_extensions.TypedDict; a filter nested in a field value or passed to any other call is still reported. The two types/proxy/management_endpoints/team_endpoints.py entries leave the baseline, which is now 156.
* fix: refuse an update_many_in whose data writes the chunked field
Chunks run one after another, so an update that sets the chunked field can move a row into a later chunk, which updates it again and counts it twice: values ["old", "new"] with chunk_size=1 and data={"id": "new"} does exactly that. update_many_in now raises ChunkedFieldWriteError before any query when data has the chunked field as a top-level key, in any form, including Prisma operators such as {"set": ...}.
* docs: cut the unbounded IN list checker's docstring to what it flags and how to clear it
It now says what is reported, the three ways to clear a finding, and how the baseline and --update-baseline work, in 11 lines. The per-shape detail lives in the tests.
* ci: key an unbounded IN list finding by its filtered expression too
A baseline key of path, scope, kind, field and occurrence let a PR delete
a baselined filter and add a different unbounded one on the same field in
the same function, and the new one took over the old key. The key now
also carries the filtered expression's source, whitespace-normalized
(the Prisma value, or a raw-SQL `IN (...)` slot), so that swap reads as
one new and one stale entry and fails the run. The same expression
re-added in the same function is still the same finding.
Every baseline entry is rewritten in the new form; the 156 findings are
unchanged, and only occurrence indexes renumber where one field had
several different expressions.
* test: move key-gated tests/test_litellm SDK tests into tests/llm_translation and drop empty folders
* test: make token counter and health check unit tests run offline
* ci: point unit shards, rust path filter, Makefile and docs at tests/unit
* docs: fix stale test_litellm run paths in moved llm_translation tests
* fix: correct databricks e2e sys.path depth and contributing example path
---------
Co-authored-by: yuneng <yuneng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>