Commit graph

53720 commits

Author SHA1 Message Date
berriai-litellm-provider-info-sync[bot]
e366e72502
feat(bedrock): add glm 5.3 cross-region rows and nova 2.5 sonic (#44710)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-05 20:25:24 -07:00
devin-ai-integration[bot]
ab3a59fe81
fix(chatgpt,github_copilot): refuse device-code login inside an event loop or worker thread (#39585)
* fix(chatgpt,github_copilot): refuse device-code login when an event loop is running

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(github_copilot): drop stray whitespace change

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(chatgpt): bound token refresh timeout and drop placeholder assignment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(chatgpt,github_copilot): keep the token file path out of the event-loop 401 message

* fix(auth): refuse device-code login from worker threads too

/v1/messages runs its handler in an executor thread, where the
running-loop check never fires, so a chatgpt or github_copilot model
still started the interactive device-code login there and the request
hung for up to 15 minutes. The guard now also requires the main thread,
so the login only runs where a human can actually answer it.

* test(chatgpt): keep authenticator tests out of the real token directory

* fix(chatgpt): keep the 5 second connect timeout and the operator's request_timeout on the token refresh call

* test(integration): cover the device-code login guard on the proxy and the SDK

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 03:20:07 +00:00
devin-ai-integration[bot]
64cddd6e13
test(integration): basic translation cases for the bedrock_mantle route (#44750)
* test(integration): bedrock_mantle-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): rename translation runner run to assert_translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): use assert_translation and drop the LIT-9196 skips in the bedrock_mantle basic cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 03:17:34 +00:00
devin-ai-integration[bot]
a45c6b8f65
test(integration): keep scripted upstream connections open past the proxy's keepalive (#44695)
* test(integration): re-pin claude-sonnet-5 stream recount tokens and outlast proxy keepalive in scripted upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): derive scripted upstream keepalive from AIOHTTP_KEEPALIVE_TIMEOUT

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): hardcode scripted upstream keepalive to avoid importing litellm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 03:13:06 +00:00
ishaan-berri
2fd5725c04
feat(lens): add copy link button to trace header (#44749)
* feat(lens): add traceShareUrl helper for shareable trace links

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens): add copy link button to trace header

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens): cover copy link on trace header

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 03:07:56 +00:00
devin-ai-integration[bot]
402fa63366
feat(proxy): cap batch file records, daily batch uploads, and per-file downloads (#43632)
* feat(proxy): cap batch file records, daily batch uploads, and per-file downloads

Adds three opt-in limits for batch jobs, each settable in general_settings as a
per-key default and overridable in key or team metadata by a proxy admin:

max_batch_file_records rejects a purpose=batch upload with more request lines
than allowed with a 413 before it reaches the provider.

max_batch_file_uploads_per_day counts accepted batch uploads per key and per
team in a UTC day and returns 429 with Retry-After once the count is used.

max_file_downloads_per_minute counts GET /v1/files/{id}/content per key and per
team for each file in a one-minute window and returns 429 with Retry-After.

The existing admin-only guard for batch_enqueued_token_limit now covers all four
metadata keys.

* fix(proxy): keep file usage counters in their own store and gate batch limits on user creation

* fix(proxy): count keyless JWT callers per user, require positive file caps, and take the upload slot after request validation

* refactor(proxy): end file usage cap describers with an explicit return after the match

* fix(proxy): keep file usage counters when more than 200 are live without Redis

The file usage counter store used a default in-memory cache, which holds 200
entries and evicts the one that expires soonest. Without Redis, a caller got a
fresh per-file download allowance after touching about 200 other file ids in
the same minute, and a key got a fresh daily upload allowance once about 200
other keys had uploaded that day. The store now tracks up to 20,000 live
counters per worker, the same bound the login throttle uses

* test(proxy): move the file usage cap tests into the directory the proxy shard runs

Main's shard coverage check found tests/unit/proxy/openai_files_endpoints
claimed by no shard, so its tests would not run in CI. The file moves next to
the other files endpoint tests in tests/unit/proxy/openai_files_endpoint, which
the proxy-endpoints shard already runs

* fix(proxy): declare the file usage counters as rate limit calls

Main's redis producer gate requires every module that writes a shared cache to name its key family, and the file usage counters wrote theirs without one.

* test(files): audit batch file usage caps across processes, Redis outages, and config reloads

* test(files): guard the chaos cells against minute boundaries and open Redis breakers

Two chaos cells each failed once in the audit run. The restart check ran
three sequential downloads with no guard against straddling a UTC minute,
and the exact-cap probe after a Redis outage ran while both workers' Redis
circuit breakers were still open (60 s default recovery), so it counted in
per-process memory and the two workers split the cap

Every burst now carries a window guard, the chaos fixture lowers the
breaker recovery to 2 s, and the post-outage check drives a fresh key to its
cap through a one-worker sibling proxy and then expects the two-worker
candidate to refuse the whole burst, which only the shared Redis count can
produce, polled until the breakers close

* test(files): release the held uploads when the killed-worker cell fails early

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 03:06:36 +00:00
devin-ai-integration[bot]
12dcce8db2
refactor(tests): extract Rust cache test split (#44780)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 20:03:00 -07:00
nate-berri
922ea71fe5
fix(vertex_ai): apply regional endpoint uplift on image generation cost path (#44679)
* fix(vertex_ai): apply regional endpoint uplift on image generation cost path

completion_cost had vertex_location but never passed it to the image
generation cost router, so a regional_endpoint_uplift_multiplier on a
Vertex image row would be ignored. No image row carries the multiplier
yet, so nothing is misbilled today. Pass the location through to the
Vertex image calculator for both the token-based price and the
per-image fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(vertex_ai): cover regional image cost through the proxy logging path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(images): hand image_edit's vertex_location to its cost resolver

* test(integration): audit Vertex image regional uplift billing

* test(integration): require every spend row after a proxy restart

* test(integration): reject extra spend rows after a proxy restart

---------

Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 20:02:22 -07:00
devin-ai-integration[bot]
c1d639afff
test(integration): basic translation cases for the vertex_ai route (#44751)
* test(integration): vertex_ai-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): rename translation runner run to assert_translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): use assert_translation in vertex_ai-route basic translation cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 20:00:19 -07:00
devin-ai-integration[bot]
4a963a6c0b
fix(proxy): judge the free-model budget waiver by the group an alias routes to (#44638)
Some checks failed
Unit Tests / Build the Rust bridge (push) Waiting to run
Unit Tests / caching-local (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests / core-utils (push) Blocked by required conditions
Unit Tests / enterprise-managed-files (push) Blocked by required conditions
Unit Tests / enterprise-package (push) Blocked by required conditions
Unit Tests / enterprise-routing (push) Blocked by required conditions
Unit Tests / integrations (push) Blocked by required conditions
Unit Tests / OpenAI and Meta Providers (push) Blocked by required conditions
Unit Tests / All Other Providers (push) Blocked by required conditions
Unit Tests / Vertex AI (push) Blocked by required conditions
Unit Tests / misc (push) Blocked by required conditions
Unit Tests / misc-dirs (push) Blocked by required conditions
Unit Tests / proxy-auth (push) Blocked by required conditions
Unit Tests / proxy-endpoints (push) Blocked by required conditions
Unit Tests / proxy-extras (push) Blocked by required conditions
Unit Tests / proxy-feature-endpoints (push) Blocked by required conditions
Unit Tests / proxy-hooks-client (push) Blocked by required conditions
Unit Tests / proxy-infra (push) Blocked by required conditions
Unit Tests / proxy-infra-root (push) Blocked by required conditions
Unit Tests / proxy-server (push) Blocked by required conditions
Unit Tests / responses-caching-types (push) Blocked by required conditions
Unit Tests / unit (push) Blocked by required conditions
GitHub Actions Security Analysis / zizmor (push) Waiting to run
VS Code Extension / vscode-extension (push) Has been cancelled
* fix(proxy): judge the free-model budget waiver by the group an alias routes to

A hidden model_group_alias can reuse the name of a real model_name. The
router serves that name from the alias target, but two of the three checks
behind the free-model budget waiver read the deployments of both the alias
target and the real model that shares the name.

So an over-budget key was served through an alias whose name belongs to a
model with an explicit $0 price when the target is $0 only by the cost map,
which the same key is refused on by its own name. The mirror case refused a
free target because the shadowed name belongs to a PTU-priced deployment.

Resolve the alias once and have the explicit-cost and PTU checks read the
routed group, the same group the price check already reads.

* test(proxy): cover the plain alias form of a shadowed free model name

The same wrong verdict exists for a plain string alias, so the unpriced-target case now runs for both alias shapes.

* fix(proxy): judge an alias chain's budget waiver by the deployments it is served from

The explicit-price and PTU checks read the alias target through
Router.get_model_list(), which follows a second alias hop when the target is
itself an alias key. The router never takes that hop, so an alias chain was
judged by a deployment the request never reaches. The checks now take the
deployments named after the routed group, or the wildcard deployment serving
it when none carries its name.

* fix(proxy): refuse the budget waiver when an alias chain is served by a priced wildcard route

* test(integration): cover the shadowing alias budget gate end to end

Adds the audit cells for a hidden alias whose name shadows an explicitly
free group: streamed SDK refusals on chat, responses and messages,
embeddings, the free wildcard and mixed-group paths, per-model budgets,
JWT and custom auth callers, cache hits, alias removal under traffic,
provider failure and fallback, a concurrent outage burst, and a worker
kill on an owned two-worker proxy

* test(integration): cite the cost-map rows the shadowing alias cells rely on

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 02:44:35 +00:00
devin-ai-integration[bot]
ff5084687c
fix(ci): namespace claude session ids in tracing seeds and allowlist /v1/logs on backend (#44761)
* fix(ci): namespace claude session ids in tracing seeds and allowlist /v1/logs on backend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): allowlist /v1/logs on the gateway alongside /v1/traces

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 19:39:12 -07:00
devin-ai-integration[bot]
bc2df7cdcd
fix(gemini): stop replaying thinking block signatures to Gemini (#44661)
* fix(gemini): stop replaying thinking block signatures to Gemini

A thinking block's signature has no provenance, and LiteLLM never fills it
from a Gemini response (Google signs text and functionCall parts, which ride
provider_specific_fields and the tool call id), so a Claude signature replayed
through a mixed model group reached Gemini as a thoughtSignature and Google
answered 400 Invalid thought signature on every later Gemini-served turn. The
same replay also sent the thinking text a second time as a plain text part.
The thinking text now goes out once, as the thought part built from
reasoning_content, and no part is built from thinking_blocks

* test(gemini): type the parts helper and split its comprehension

* test(gemini): cover thinking signature replay on the integration rig

Two integration files from the audit of the foreign thought signature fix: 62 wire cells asserting the model turn Google receives on chat, messages and responses across gemini and vertex_ai, streaming and not, SDK and httpx clients, the sad shapes of thinking_blocks, context caching through cachedContents, and 3 chaos cells (a concurrent burst across endpoints, upstream stream drops, a worker SIGKILL mid burst)

* test(gemini): read the integration salt from the environment

The wire test decrypted Responses ids with a literal salt; tests/integration/_support/process.py boots the proxy with LITELLM_SALT_KEY when it is set, so the test now reads the same variable with the same default

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 02:31:53 +00:00
devin-ai-integration[bot]
7cdb3d3760
test(integration): basic translation cases for the bedrock_invoke route (#44752)
* test(integration): bedrock_invoke-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): rename translation runner run to assert_translation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): call assert_translation in the bedrock_invoke basic cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 19:31:46 -07:00
devin-ai-integration[bot]
2156d6c3c7
test(e2e): assert the sibling-replica cooldown through the router (#44706)
* test(e2e): assert the sibling-replica cooldown through the router

* test(e2e): warm the cooldown reads concurrently so every pod's read lands just before the trip

* test(e2e): send the trip right behind the warm so every pod's cooldown read is pinned to it

* test(e2e): warm every pod with a canned-answer group and trip only after every warm call answered

* test(e2e): trim the sibling cell's module docstring to what the design needs

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 02:12:57 +00:00
devin-ai-integration[bot]
061d60b04c
feat(model_prices): add chatgpt subscription rows for the gpt-6 family (#44758)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 19:10:26 -07:00
moe-berri
d232b008bc
test(lens): use current worker protocol in billing integration (#44760) 2026-10-05 18:39:38 -07:00
ishaan-berri
6764b9af12
feat(lens): type to filter the traces agent dropdown (#44745)
* feat(lens): type to filter the traces agent dropdown

* test(lens): cover typing, Enter and clear in the agent filter
2026-10-05 18:36:22 -07:00
devin-ai-integration[bot]
6f5ca84f69
feat(prometheus): cap series per metric for every labeled metric (#44420)
* feat(prometheus): cap series per metric for every labeled metric

Add prometheus_metrics_max_series_per_metric: per metric and per worker process, the first N label
sets keep a series of their own. Counters and histograms record every later label set on one series
whose labels are all "other", so totals stay exact, and gauges skip it. The cap holds with multiple
workers because it never needs to remove a series.

Add prometheus_metrics_ttl_seconds: a series idle for that long is removed and its slot is freed.
The prometheus client cannot remove a series in multi-process mode, so the TTL is ignored there with
a startup warning.

Both settings are off by default. The end_user caps are unchanged.

* fix(prometheus): share the series cap across workers of one proxy instance

Workers writing to one PROMETHEUS_MULTIPROC_DIR now agree on which label
sets get a series through an append-only admissions file per metric, so a
merged scrape stays at the cap plus `other` instead of growing with every
worker and every worker restart. The two fallback counters now pass their
label names as a keyword so the cap and prometheus_exclude_labels apply to
them, admission and child creation happen under one lock, the test fixture
restores the shared registry, and the `other` label value lives in
constants.py.

* test(prometheus): check emitted labels instead of wrapper types, close the admission match

The exclude-labels test now emits through the spend and provider budget
metrics and checks the scrape keeps all their labels. The admission match
arms end in assert_never so the match is exhaustive.

* fix(prometheus): return the exhaustive-match fallback so every admission arm returns

* fix(prometheus): pick the series tracker with isinstance so every path of _admits returns

* fix(prometheus): skip an admissions line a worker could only write part of

* fix(prometheus): frame each admissions record with newlines so a cut-off record cannot swallow the next

A record a worker could only write part of used to merge with the next worker's record, and both were skipped for one request. Each record is now written between two newlines, so the fragment is a line of its own. The clock fixture in the series tests starts from a constant instead of reading the real clock

* fix(prometheus): ignore a non-positive series cap or TTL with a warning instead of failing the logger

A cap or TTL of 0 or less raised at logger init. The proxy logs that as a non-blocking error and keeps serving, so the result was a running proxy with no Prometheus metrics at all. The setting is now ignored with a startup warning naming it, the same rule the end_user cap already follows for a non-positive value

* fix(prometheus): start the series cap over on a one-worker restart and audit it live

A proxy with one worker and an operator-set PROMETHEUS_MULTIPROC_DIR now drops litellm's admission files at boot, so a restart frees every slot there the way it already does with several workers. A cap or TTL that is not a number greater than 0 (a bool, a non-numeric string, an empty value) is ignored with the startup warning instead of breaking the logger

The integration cells drive the cap on every endpoint through the OpenAI and Anthropic SDKs and raw httpx, streaming and not, plus gauges, cache hits, failures, both workers of one instance, the TTL on one worker and its warning on two, ignored settings, excluded labels on the fallback counters, a null cap, /config/update, a concurrent burst scraped mid-flight, a provider outage, a killed worker, and restarts with one and two workers

* fix(prometheus): wipe an operator-set multiprocess directory on a one-worker boot too

* fix(prometheus): leave the multiprocess directory alone on a setup-only run

A run with --skip_server_startup starts no worker, so it no longer creates or
wipes PROMETHEUS_MULTIPROC_DIR. Wiping there deleted the samples of a proxy
already running against the same directory

* fix(prometheus): free the capped series slots when a gateway or backend container restarts

The component image entrypoint starts uvicorn without the proxy CLI and wiped only the .db sample files at container start, so the admitted-series files of the previous container survived an in-place restart. Every label set seen after the restart was then counted on `other` once the previous container had filled the cap

* test(prometheus): cover a setup-only run and a gateway image restart under the cap

Two integration cells from the audit: a `--skip_server_startup` run pointed at a live
two-worker proxy's operator directory leaves its samples alone, and the gateway image
(`docker/component_entrypoint.sh` running `python -m gateway.launch`) restarted on a kept
PROMETHEUS_MULTIPROC_DIR starts the cap over. The burst cells now wait for every counter
they assert on, since the request and failure counters of one call increment at different
points of the logging callback

* test(prometheus): prove the cap reaches the fallback counters in the X1 cell

* fix(prometheus): ignore a cleanup interval that is not a number of at least 0

A string or negative prometheus_metrics_cleanup_interval_seconds reached the
series tracker unvalidated, so the first labeled emit with a TTL on raised
TypeError inside the callback and recorded no series. The interval is now
validated the way the cap and the TTL are: an invalid value is ignored with a
warning and the default 60 seconds applies. The I2 integration cell drives a
string interval through a live proxy and reads the warning from its log

* test(prometheus): give the restart cells the boot budget of their siblings

C4 and C5 boot two proxies each and hit the file's 240 s budget on a loaded
box; C3 and D1 already carry 420 s

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 01:28:44 +00:00
moe-berri
b58e2d7175
fix(lens): bound result recovery and preserve partial results (#44692)
* fix(lens): restate response contract during model repair

* fix(lens): separate instructions and recover rejected results

* fix(lens): correct loop type annotations and checks

* fix(lens): preserve access to prior findings after compaction

* fix(lens): cap result retries and preserve partial completion
2026-10-05 18:24:53 -07:00
tin-berri
464fe5bd90
feat(ui): configure cache-aware auto routing (#43396)
* feat(ui): configure cache-aware auto routing

* style(ui): keep cache routing config within line limit
2026-10-05 18:24:24 -07:00
moe-berri
4cc442f5ee
feat(lens): show investigation names in findings table (#44753) 2026-10-06 01:23:50 +00:00
devin-ai-integration[bot]
655baa50be
test(integration): rename translation runner run to assert_translation (#44741)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 01:08:44 +00:00
devin-ai-integration[bot]
56a581a536
fix(deps): bump source-map-js, smol-toml, mako, multidict and werkzeug for OSV advisories (#44728)
* fix(deps): bump source-map-js to 1.2.2 for GHSA-68fv-2mgg-jv7q

* fix(deps): bump mako, multidict, werkzeug and smol-toml for OSV advisories

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-06 00:57:04 +00:00
devin-ai-integration[bot]
3242bdfed2
fix(responses): honor request cache controls on chat completions bridged to the Responses API (#44676)
* fix(responses): honor request cache controls on chat completions bridged to the Responses API

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): assert spend and cache-hit status on bridged no-cache rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): unskip LIT-9196 openai_responses basic translation cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): restore azure attribution check on bridged no-cache rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): exercise bridged cache controls through a real local cache instead of patched responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 17:52:44 -07:00
moe-berri
2190003bb8
fix(lens): reconstruct native coding agent conversations (#44711)
* fix(lens): reconstruct native coding agent conversations

* fix(lens): display timestamps used for conversation ordering

* fix(lens): preserve coding trace identity and message provenance

* fix(lens): complete capture checks after trace pagination

* fix(lens): show loaded replies during trace pagination
2026-10-05 17:51:05 -07:00
moyai-devin-berriai[bot]
6e75f28289
feat(ui): support native decisions endpoint in decision playground (#44664)
* feat(ui): support native decisions endpoint in decision playground

* fix(ui): preserve decision playground defaults and test late cancellation results

* test(ui): remove redundant cancellation test comment

* Update ui/litellm-dashboard/src/app/(dashboard)/playground/components/systemOneUI/SystemOneUI.integration.test.tsx

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-10-06 00:48:10 +00:00
anniejeng
aa4852f579
feat: add Reka as an OpenAI-compatible provider (#44278)
* Add Reka as an OpenAI-compatible provider

* Add Reka supported endpoints

* fix: remove stray fragment after reka entry in provider_endpoints_support.json

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(reka): register provider identity and cover routing, credentials, and bridged endpoints

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-05 17:47:23 -07:00
dependabot[bot]
801aba88ed
chore(deps): bump langgraph-sdk from 0.4.2 to 0.4.4 (#44713)
Bumps [langgraph-sdk](https://github.com/langchain-ai/langgraph) from 0.4.2 to 0.4.4.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/0.4.2...0.4.4)

---
updated-dependencies:
- dependency-name: langgraph-sdk
  dependency-version: 0.4.4
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-05 17:35:38 -07:00
devin-ai-integration[bot]
d7e7f4071d
test(rust_bridge): expect PartRow start_time and end_time as required LENS_CONTENT fields (#44719)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 00:14:41 +00:00
devin-ai-integration[bot]
169af2f883
ci: split slow unit shards and build the Rust bridge once per run (#44622)
* ci: split slow unit shards and build the Rust bridge once per run

* ci: key the Rust bridge cache on source files only

* ci: keep the unit setup ceiling unchanged with the shared Rust bridge

* ci: fall back to the Cargo cache when the Rust bridge artifact is missing

* ci: keep reruns on enterprise-routing for the prompt caching flake

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 23:43:01 +00:00
moe-berri
c29b42a32b
fix(lens): preserve span timestamps in investigation evidence (#44702)
* fix(lens): preserve span timestamps in investigation evidence

* style(lens): format chronology regression test
2026-10-05 16:42:42 -07:00
nate-berri
371b527d60
fix(build): rebuild the Rust bridge when uv sync sees Rust sources change (#44708)
uv only rebuilds the editable litellm package when its cache keys change, and
the default keys are pyproject.toml, setup.py and setup.cfg. Editing or
switching to a branch with different Rust code left the old
litellm/rust_bridge/_native.abi3.so installed. Key the build on the Rust
toolchain pin, Cargo config, lockfile, workspace manifest and every file
under litellm-rust/crates, since crates embed .sql, .json and .jinja files at
compile time.

Co-authored-by: Nate Armstrong <narmstrong@Nates-MacBook-Pro.local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:42:22 -07:00
moe-berri
c561372a02
fix(lens): show investigation findings for agent traces (#44696)
* fix(lens): show investigation findings for agent traces

Replace child tool-error counts in the trace table with distinct investigation findings. Keep unassessed traces separate from completed clean investigations and use the root status for failure filters and timeline counts

* fix(lens): stabilize findings updates and repair UI checks

* fix(lens): allow viewer findings reads and index trace lookups

* test(lens): cover viewer findings reads with Postgres
2026-10-05 16:27:54 -07:00
devin-ai-integration[bot]
65b0557f80
test(integration): point the scratch upgraded proxy's read replica at the scratch database (#44613)
* test(integration): point the scratch upgraded proxy's read replica at the scratch database

* test(integration): check the scratch upgraded proxy's reader role is connected to the scratch database

* test(integration): assert every configured proxy role holds a scratch connection without a mode branch

* test(integration): skip backend workers without a role in the scratch connection scan

---------

Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 23:06:36 +00:00
moe-berri
d7210f95e6
fix(lens): persist final coverage with source diagnostics (#44682) 2026-10-05 15:53:49 -07:00
devin-ai-integration[bot]
68d9b8bbb8
fix(router): retry a /v1/messages stream the provider drops before the first content chunk (#44276)
* fix(router): retry a /v1/messages stream the provider drops before the first content chunk

A /v1/messages stream that the upstream closed before any content reached the
client answered an error event after a single attempt, so the router's
num_retries never applied to that drop. The pre-content failure is now retried
within the model group before the fallback chain runs, with the budget resolved
the way a failure raised before the stream opened resolves it: a retry policy
that names the error class, then the request's num_retries, then the
deployment's, then the router's. A drop after content reached the client keeps
surfacing the provider's error after one attempt.

Fixes #44238

* fix(router): hand a retry's non-retriable error to the fallback chain and type the retry helpers

A retry that failed before its stream opened with an error no retry covers raised straight to the
client, skipping a fallback the first attempt would have used. assert_never now comes from
typing_extensions so the router imports on Python 3.10, and the retry helpers read their kwargs
through typed narrowing instead of Mapping[str, Any]

* fix(router): cast the untyped router fallback defaults the stream retry gate reads

The retry gate passed the router's fallback attributes, declared without element types, to the
typed request override helper, which basedpyright counted as new unknown-argument errors

* fix(router): consult context_window_fallbacks when a retried /v1/messages stream overflows

A retry attempt raising ContextWindowExceededError reached the fallback chain inside its
mid-stream envelope, so only the regular fallbacks list matched. The fallback attempt now
unwraps it the way it unwraps a content policy error. The new router helpers are covered for
the router code coverage check with two direct-call tests and named covering tests

* fix(router): retry a 408 raised by a /v1/messages retry and honor deployment num_retries before the stream opens

* fix(router): attribute a retried /v1/messages stream to the deployment that served it and bound the retry-policy hold

* fix(router): retry /v1/messages error frames under their retry-policy class and keep the first drop's committed budget

An `event: error` frame that arrives before the first content delta now raises the exception class the pre-stream mapping gives an HTTP answer with the same status (429 RateLimitError, 500 and 529 InternalServerError, 503 ServiceUnavailableError, 504 Timeout), so a retry policy's per-class budget governs it the way it governs the error before the stream opened. The status the client sees is unchanged

A retry that lands on a sibling deployment keeps the budget the first drop committed to, read back from the request's attempted_retries and max_retries, instead of recomputing it from the new deployment's num_retries, matching the pre-stream retry loop

* refactor(anthropic): keep the error-frame exception mapping under llms and type the retry test helper

The status-to-exception mapping an `event: error` frame gets before the retry policy is consulted now lives next to the Anthropic error status map in llms/anthropic/common_utils.py, with its own unit test, and the two-deployment retry test helper takes explicit typed parameters instead of a bare dict and untyped kwargs

* refactor(anthropic): map an error frame's status with explicit returns on every path

* fix(router): map stream error frames through the pre-stream exception mapping

An overloaded `event: error` frame on a /v1/messages stream now raises the InternalServerError a 529 answer maps to, built by exception_type from the frame's own body, so one retry policy class governs the error before and after the first byte; a failed fallback after such a frame answers 500 like every other litellm path instead of the frame map's 503

A model_group_retry_policy that does not parse (a non-integer budget, an entry that is not a mapping) no longer fails every healthy stream of that group before its first attempt: the stream runs with no policy and the plain num_retries budget, with a warning naming the group

* fix(router): forward an error frame nothing can take over for as the provider sent it

A pre-content error frame whose class the retry policy grants no retry, with no fallback configured, raised an HTTP error only on the first attempt while the same frame after exhausted retries reached the client verbatim. Both now pass through as sent, the way the merge base forwarded every frame.

* test(integration): audit /v1/messages pre-content retry across routes and budgets

Adds the /audit cells for the pre-content stream retry: the native Anthropic route
(drops and error frames before content, HTTP rejections before the stream opens, SDK
sync and async, after-content and non-retriable controls, budget exhaustion, cache
twin, spend row and headers), the chat and responses bridges, the generic routes
(responses, chat, vllm pass-through, Gemini generateContent, fine-tuning jobs list),
owned two-worker proxies for router-level budgets, retry policies and fallbacks, and
two chaos cells (a worker killed mid burst, an outage on every first attempt). Shared
helpers for scripted Anthropic SSE upstreams and OpenAI-compatible wire replies live
in tests/integration/_support

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 22:53:39 +00:00
devin-ai-integration[bot]
7bad8de067
chore: move PR template to PULL_REQUEST_TEMPLATE/general.md, add rust.md (#44693)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:47:39 +00:00
devin-ai-integration[bot]
d9b1d57f07
fix(proxy): let a listed team alias win over a same-named key alias in the customer model check (#44677)
* fix(proxy): let a listed team alias win over a same-named key alias in the customer model check

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): check each requested name in a plain loop in can_customer_access_model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): only let a listed team alias skip the customer check when its target is live

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): move the per-name customer alias check into a local function

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:43:16 +00:00
devin-ai-integration[bot]
a1a42768c1
fix(spend-tracking): stop caching failed spend-log metadata lookups as confirmed misses (#43560)
* fix(spend-tracking): stop caching failed spend-log metadata lookups as confirmed misses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(spend-tracking): share the short-lived miss cache write

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): cover key alias recovery after spend log lookup failures across usage routes

* test(spend): bound outage alias lookups per miss window instead of a fixed count

* fix(spend-tracking): treat any spend-log lookup failure as a short-lived miss

The Prisma client raises a plain AttributeError when the database drops
the connection mid-query, so the PrismaError catch let it through and
the whole usage call answered 500. Any failure now keeps the 30 second
backoff only, and the integration proxy patches its test entitlement at
import so uvicorn's spawned workers inherit it

* test(integration): audit spend-log metadata recovery under timeouts and dropped connections

Cover the daily activity routes, the usage AI chat, the Vantage and
CloudZero dry runs and exports under a locked spend-log table and under
a database connection dropped mid-lookup, on a two-worker proxy, with
the recovery after the outage asserted through the proxy's own miss TTL.

Add a dropped_connection_relay that closes only the connection whose
bytes carry a trigger, so a cell can drop the one connection the
recovery query runs on while the rest of the pool keeps serving. Rewrite
the sweep and JWT cells for the merged main: the export route reads
metadata by SQL join and never calls the recovery, the search routes
answer key rows and find deleted keys by alias, and the daily-spend
owner recovery names the user while the alias stays blank. The sweep
cell now times out a second lookup under the same lock, which pins the
keys blank on the merge base and recovers on this branch.

* test(integration): match a dropped-connection trigger split across two reads

The dropped-connection relay checked each TCP read on its own, so a SQL
marker that straddled two reads never tripped it and the outage cells
would run without the outage they meant to exercise. Carry the tail of
the previous read into the next check, as the held-statement relay
already does, and pin that with a unit test that splits the trigger
across two writes.

* test(integration): scan relay triggers through an in-process helper

The dropped-connection relay now matches its SQL trigger through a TriggerScanner that carries the previous read's tail, and the unit test exercises that scanner directly instead of opening loopback sockets, which tests/unit forbids. The relay's end to end behavior stays covered by the integration cells

---------

Co-authored-by: gabriele <gabriele@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 22:36:25 +00:00
devin-ai-integration[bot]
b9251dafad
refactor(rust): derive string enum serde through strum and serde_with (#44675)
* docs(rust): document string enum serde conversions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(rust): derive string enum serde through strum and serde_with

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:34:34 -07:00
devin-ai-integration[bot]
8b7b42f3aa
test(integration): basic translation cases for the openai route (#44667)
* test(integration): openai-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): ignore x-stainless headers in translation runner

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): skip openai chat completions basic cases on LIT-9235

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:32:50 -07:00
devin-ai-integration[bot]
6a92a49c32
test(integration): basic translation cases for the gemini route (#44662)
* test(integration): gemini-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): raise gemini basic cases to 1024 output tokens

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 15:32:26 -07:00
devin-ai-integration[bot]
74ad633e53
refactor(rust): rename litellm-framing crate to litellm-framer (#44683)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:32:00 +00:00
devin-ai-integration[bot]
a6b7f62760
test(integration): assert the Messages API health probe on Mantle Claude (#44612)
Co-authored-by: yuneng <yuneng@berri.ai>
2026-10-05 15:28:02 -07:00
devin-ai-integration[bot]
4314f3ce0c
test(integration): basic translation cases for the azure route (#44672)
* test(integration): azure-route basic translation cases on messages, chat completions and responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): use the three-line form for the azure chat LIT-9235 skip

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:24:42 +00:00
tin-berri
9062fd3931
feat(proxy): embed enterprise LiteAdmin MCP in LiteLLM images (#44610) 2026-10-05 15:24:34 -07:00
devin-ai-integration[bot]
739192ec19
test(integration): pin the Claude-tokenizer recount for streamed claude-sonnet-5 no-usage cases (#44643)
Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 22:17:12 +00:00
berriai-litellm-provider-info-sync[bot]
8c15631836
fix(bedrock): add priority and flex prices for Grok 4.3, 4.6, 4.7 and Kimi K3 (#44670)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-05 15:12:37 -07:00
berriai-litellm-provider-info-sync[bot]
489962cc22
fix(vertex-ai): add deprecation_date to gemini-3.1-flash-lite-image (#44639)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-05 15:10:49 -07:00
moe-berri
b69d744993
feat(lens): analyze trace workspaces with confined Python and compaction (#44640)
* feat(lens): add per-trace review models to jobs and progress

* feat(lens): append worker reviews to the job, capped, and count every review

* feat(lens): report a review with reasoning for each screened trace

* chore(ui): regenerate api types for lens job reviews

* feat(lens): type job reviews and fill them in lens fixtures

* feat(lens): add live review playback model

* feat(lens): pick the analysis model and slow single-review pacing

* feat(lens): add sample reviews for previewing the live run

* feat(lens): add live run layout with queue, reading trace and conclusions

* feat(lens): show the live run on investigations and open it from run now

* feat(lens): stream large review backlogs at 150ms or less and list newest first

* fix(lens): show the live run only for real reviews and keep fixtures test-only

* refactor(lens): restyle the live run as the native progress panel

* fix(lens): retry contended investigation updates with jittered backoff

* feat(lens): add a reading ticker line and replay for finished runs

* feat(lens): collapse the live run to an ambient line with show work

* fix(ui): crop the cerebras logo viewBox to its mark so it reads at icon size

* feat(lens): format review span previews as readable messages

* feat(lens): derive strip status, honest issue counts and drawer focus from a job

* feat(lens): track active jobs before their first review

* feat(lens): add a live trace results drawer with readable spans

* feat(lens): put the live strip under the progress bar and drop the inline panel

* feat(lens): add an ambient live strip that opens the drawer

* fix(lens): wait out provider rate limits and retry model calls four times

* style(lens): format repository contention tests

* feat(lens): read review spans as a conversation timeline

Turns spans into the user's ask, tool calls with args and results, and the agent's reply, dropping system prompts. Also handles a preview cut that lands inside the Output header.

* fix(lens): list recorded agents in the run now dialog

The run now agent field used a native datalist, whose suggestions do not show inside the modal dialog, so the agent list looked empty even though /lens/agents returned names. Use the same Combobox as investigation setup.

* feat(lens): pace live playback so each trace stays readable

Every trace now stays up for at least 1.5s. A backlog is cleared by skipping to the newest few instead of flickering through them. Conclusions count traces per check and kind, and new helpers cover share bars, group filters and flashes.

* feat(lens): keep the live run ambient until View run is clicked

The drawer no longer opens on Run or when entering a running investigation. LiveRun takes reviews as a prop so it can move to a dedicated reviews endpoint.

* feat(lens): show the live run as a two-pane trace and conclusions view

Left pane: the trace being reviewed as a readable timeline, followed by Lens's reasoning and the verdict. Right pane: ranked conclusion groups with share bars, plus a trace list you can filter.

* fix(lens): run several investigations per worker and poll every two seconds

* feat(lens): add worker slot and poll interval settings

* feat(lens): add list summaries and an incremental review filter

* perf(lens): strip reviews and run attributes from the lens list and serve reviews separately

* test(lens): cover list summaries, review polling and review access

* feat(lens): explain why a queued investigation is waiting

Works out whether no worker is connected, the worker is busy (with its running investigations and an estimated start time), or it is just being picked up.

* feat(lens): show the queue reason and what the worker is doing in the live strip

The progress header and the strip replace "Queued for your worker" with the concrete reason. While waiting, the strip lists the busy worker's investigations; click one to open it.

* feat(lens): add a review page model carrying the total reviewed count

* fix(lens): page live reviews by index so out-of-order reviews are never skipped

* feat(lens): take an index cursor on the reviews endpoint

* test(lens): cover index cursors across out-of-order and rolled-over reviews

* chore(ui): regenerate api types for the lens reviews endpoint

* feat(lens): page job reviews by index cursor

Adds api.reviews for GET /lens/{id}/runs/{job}/reviews?after=N, with a demo implementation. appendPage adds pages in arrival order and keeps the latest 200. liveJob now keys off reviewed, since the list no longer carries reviews.

* feat(ui): add a lens reviews query that polls the index cursor while live

* fix(lens): feed the live run from the reviews endpoint and keep View run open

LiveRun now gets its reviews from useJobReviews instead of the list, which no longer carries them. View run stays clickable while a run is queued or running, and before the first trace the opened view says what the worker is doing.

* fix(lens): split live conclusions into issues and patterns

A check could show up twice with the same label, once as an issue and once as a pattern.

* fix(lens): group live conclusions by check with short labels

There is now one group per check_id: issue traces are the main count and pattern traces a secondary note, so there are no duplicate red and grey cards. A long instruction falls back to the humanized check id. Adds briefReasoning and traceRows for the simplified trace list, and drops helpers nothing uses.

* feat(lens): simplify View run to traces and conclusions

The left pane is the trace list. A soft highlighter carrying the provider and model slides to the trace being reviewed, and clicking a row shows just Lens's reasoning and verdicts. The right pane keeps one conclusion card per check.

* refactor(lens): drop client-side replay in favour of real in-flight rows

Removes the playback reducer and its pacing. liveRows lists the traces the worker is reading, from job.reading, followed by completed reviews newest first, keyed by execution_id so a trace keeps its row when it finishes.

* feat(lens): show what the worker is reading and make View run obvious

Each trace in flight gets a highlighted row with the model and a live timer, and becomes its completed row in place. Completed rows show the real review time. View run is an outline button next to the progress line, and clicking anywhere on the strip opens it too.

* feat(lens): sum up a finished live run with time taken

doneLine reads like "Reviewed 30 traces in 31s with", measured from when reading started.

* feat(lens): slide one model rectangle over the traces being read

A single rounded rectangle carrying the provider logo and model wraps the real in-flight rows from job.reading. It translates and resizes over 250ms as traces finish in place. Before job.reading arrives it sits on a top slot showing the honest progress line, and when the run completes it fades out over 400ms. Rows have a fixed height and stable execution_id keys, so polls don't cause jumps or flicker.

* feat(lens): add in-flight runs to jobs and worker progress

* feat(lens): store in-flight runs from progress and clear them when a job ends

* refactor(lens): route progress, cancel and results through shared job transitions

* feat(lens): report each run as in flight when its review starts

* feat(lens): send in-flight runs with worker progress

* test(lens): cover in-flight runs across progress, old workers and terminal states

* test(lens): cover in-flight reporting under original run ids

* chore(ui): regenerate api types for lens in-flight runs

* feat(lens): model live reading lanes from in-flight runs and reviews

* feat(lens): show a now reading stage that types each trace's reasoning

* feat(lens): put the now reading stage above the trace list in View run

* fix(lens): resolve the analysis provider logo from the model catalog

* fix(lens): give demo jobs an empty in-flight list

* style(lens): format endpoint tests

* refactor(lens): name the run now handler in investigations view

* refactor(lens): name now reading conditions

* refactor(lens): name inline objects in the live run

* style(lens): format live run files

* fix(lens): keep worker settings inside the standalone worker package

* refactor(lens): keep update retry settings next to the repository

* fix(lens): start review history over when a run is reclaimed

* chore(lens): drop the unused review fixture

* refactor(lens): remove dead live helpers and use generated in-flight types

* fix(lens): keep polling a finished run until its last reviews arrive

* perf(lens): tick fast only while reasoning is typing

* fix(lens): isolate retried reviews and finding identities

* fix(lens): space the model name in run summary

* feat(lens): integrate confined workspace analysis with live reviews

* fix(lens): synchronize confined Python process monitoring

* Update review.md

* fix(lens): allow mixed context capacities and correct review assertions

* fix(lens): retrieve evidence on demand and isolate failed reviews

* fix(lens): isolate incomplete evidence reads from peer reviews

* test(lens): await trace status filter option

* test(lens): wait for reclaimed review state to settle

* fix(lens): recover from incomplete cross-session evidence

---------

Co-authored-by: Ishaan Jaff <ishaan@berri.ai>
2026-10-05 22:06:48 +00:00