Commit graph

53585 commits

Author SHA1 Message Date
joshua-berri
0ea166c160
fix(mcp): preserve upstream tool schemas and parameter headers (#44425)
* fix(mcp): preserve tool schemas and modern parameter headers

* fix(mcp): allow bounded cold schema worker startup

* fix(mcp): align catalog deadlines and bound schema traversal

---------

Co-authored-by: Joshua Valluru <326636767+joshua-berri@users.noreply.github.com>
2026-10-03 15:00:43 -07:00
devin-ai-integration[bot]
98337c9334
fix(lens): release budget reservations when the analysis model call fails (#44431)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:39:27 -07:00
devin-ai-integration[bot]
41a3781d4e
fix(proxy): treat Postgres connection exhaustion as backpressure, not poison rows (#44266)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:28:52 -07:00
devin-ai-integration[bot]
9a4b1951a2
fix(proxy): treat Postgres connection exhaustion as DB unavailable, not poison spend-log rows (#44270)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:28:52 -07:00
devin-ai-integration[bot]
e9bf2cfd01
refactor(types): replace Any with proven types in 9 files (#44389)
* refactor(types): replace Any with proven types in 13 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep provider error paths for malformed prefetch and poll JSON

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's login body parsing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's Copilot auth and budget alert typing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit cells for Any sweep 20261003_2

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): load audit video deployments at proxy start

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): keep main's video-edit prefetch handling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): fix audit chaos Responses cells

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): probe every route after audit worker kill

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 21:03:16 +00:00
devin-ai-integration[bot]
fe683ea139
feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models (#44136)
* feat(proxy): serve a Codex-native model catalog with per-model service_tiers from /v1/models

GET /v1/models and /models answer Codex CLI's catalog fetch (the request
carrying its client_version query parameter) with Codex's own
{"models": [...]} shape: a model Codex knows keeps the metadata of its
bundled 0.159.3 catalog (vendored), any other model gets Codex's fallback
entry, and model_info.service_tiers becomes each entry's service tiers so
Codex offers them as slash commands that send service_tier upstream.
Without the parameter the OpenAI list shape is unchanged. The CLI's
litellm agents codex catalog shares the same builder.

* fix(proxy): offer a Codex service tier only when every deployment of the model lists it

* fix(codex-catalog): an invalid service_tiers value offers no tier for the model

* fix(codex-catalog): read service tiers off the deployments the key's team can route to

A tier is offered to Codex only when every deployment of the model name a
request from the key's team can route to lists it, so another team's
deployment of the name and a deployment an admin paused via model_info.blocked
no longer withhold or add tiers for requests that never reach them

The catalog's always-null fields are annotated NoneType so the module imports
under pydantic 2.12.0 on Python 3.14, the lowest pin the MCP resolve job
installs, which rejects a None annotation with a None default

* test(codex-catalog): drop the redundant module docstring and sort the imports

* test(integration): add the Codex catalog audit cells and the multi-worker convergence note

* test(integration): clean up every catalog test model and answer the refresh GET

* fix(proxy): keep tiered models under Codex's catalog cut and resolve alias tiers

Under Codex's 1 MiB catalog limit the entries offering a service tier are kept
ahead of those offering none, each group in model_list order, with every kept
entry at its listing position, so the model an operator configured tiers for
survives a wide key's long listing. A model_group_alias row reads its target's
deployments, so it carries the target's tiers and stock metadata under the
alias name.

* fix(proxy): pick Codex catalog metadata per team and skip entries too large for the cut

The upstream model that selects Codex's stock entry was read off the first deployment of a name
without checking the key's team, so a team whose requests route to a different deployment could be
handed another team's prompt, reasoning levels, and tiers. The upstream model and the tiers now come
from the same team-aware selection routing uses, and a caller with no team reads the deployments no
team owns

The byte cut kept a prefix of the tier-first order, so one entry larger than the whole limit emptied
the catalog. An entry too large for the bytes left is now passed over and the smaller ones after it
are still kept

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 20:42:55 +00:00
Misbah Syed
fd14e51ecf
feat(docker): add a Windows quickstart in PowerShell, and a styled terminal for both quickstarts (#44310)
* feat(docker): Windows quickstart in PowerShell, and a styled terminal for both quickstarts

scripts/quickstart.ps1 does what scripts/quickstart.sh does, for Windows:
  powershell -c "irm https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/quickstart.ps1 | iex"
It asks the same two questions, prints the same lines, writes the same .env
(UTF-8 without a byte order mark, LF endings, readable only by the current
Windows account), and runs in Windows PowerShell 5.1 and PowerShell 7.

Both scripts now give a person at a terminal colors, a check mark per step, a
spinner while Docker starts, and a framed summary. Agents, CI, log files, and
NO_COLOR get the same lines as plain text.

* feat(docker): support podman and rancher desktop in the quickstart scripts

Both quickstarts hard-required the docker CLI. They now pick the first
available engine among docker, podman, and nerdctl (Rancher Desktop in
containerd mode; its dockerd mode already provides a docker CLI), route
every invocation through it, and tailor the start hint (podman machine
start) and the printed stop/logs/volume commands to that engine

* fix(docker): test port availability by binding instead of connecting

On a WSL2-backed engine (Podman, Rancher Desktop), the Windows localhost
relay swallows connection refusals on closed ports, so every connect
waits out its 2-second timeout and Test-PortFree reported ports 4000 to
4099 all taken on a machine with none of them in use. Binding the port
answers instantly and accurately

* fix(quickstart): address review findings

- The PowerShell script names the project with the same POSIX cksum as the
  shell script, so the earlier-install check sees the database from either
  script.
- Both scripts stop when git tracks .env in the install folder.
- .env is written to a temp file with owner-only permissions and moved into
  place, so a failed write never leaves a partial file.
- Errors stay plain text when stderr is redirected.
- The spinners remove their temp files on exit and Ctrl+C, and the shell
  spinner no longer runs date on every frame.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(quickstart): offer to copy the admin password to the clipboard

After the summary, the quickstarts ask "What next?": copy the admin
password, open the admin UI, or finish. The password is piped to the
clipboard (pbcopy, wl-copy, xclip, xsel, clip.exe, or Set-Clipboard), so
it never appears on screen or in the process list. Over SSH, or with no
clipboard, the option is left out and the summary points to .env as
before. The PowerShell menu now honours Ctrl+C.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(quickstart): write .env through mktemp, let Ctrl+C cancel the PowerShell menu, drop redundant comments

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(quickstart): remove the in-progress .env temp file when interrupted

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Mubashir Osmani <mubashir.osmani777@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:35:32 -07:00
ishaan-berri
dd31692282
feat(lens): always-on investigations with findings and investigations tables (#44418)
* feat(lens): record run steps, trigger and exact run windows on investigations

* feat(lens): scan only traces since the last run and keep a capped step log

* feat(lens): log each analysis model call with its model, tokens and cost

* feat(lens): accept agent and time window on run now and add turn-all-on

* test(lens): cover new-traces-only windows, manual runs and the step cap

* test(lens): cover run now overrides and turning paused investigations on

* chore(ui): regenerate api types for lens run steps and run now options

* feat(lens): group open findings into one row per problem with agent filters

* test(lens): cover the findings table grouping, filters and schedule labels

* feat(lens): add a findings table across all investigations

* feat(lens): show a live step feed with the model behind each call

* feat(lens): offer turning all paused investigations on

* feat(lens): let run now pick an agent and time window

* test(lens): cover run now request building

* feat(lens): show the step feed and run now dialog on an investigation

* feat(lens): open run now choices instead of running immediately

* feat(lens): open findings first and peek a finding without leaving the table

* feat(lens): fold investigation actions into the findings toolbar

* feat(lens): show each investigation's schedule and open findings

* feat(lens): name the agent on a finding

* feat(lens): keep new investigations watching every 15 minutes by default

* feat(lens): show the watch schedule outside advanced options

* feat(lens): send run now options and turn-all-on from the dashboard

* test(lens): give demo runs steps and a trigger

* feat(lens): let the findings table fill the screen

* test(lens): add steps and trigger to progress fixtures

* test(lens): add steps and trigger to status fixtures

* test(lens): cover the default watch schedule in setup

* test(lens): cover run now choices from an investigation

* test(lens): open saved investigations from the manage view

* feat(lens): use one tab bar for traces, findings and investigations

* feat(lens): place page actions on the lens tab row

* feat(lens): drop the nested tabs and edit investigations in place

* feat(lens): show investigations as a table with run now and edit

* feat(lens): name each findings row for screen readers

* test(lens): open saved investigation links on findings

* test(lens): reach findings and investigations from the top tabs

* fix(lens): mark run now jobs manual and keep them from moving the scheduled scan

* fix(lens): keep run now since-last-run windows even with an agent override

* test(lens): cover that manual runs never skip scheduled traces

* test(lens): cover run now windows with agent and lookback overrides

* fix(lens): group findings without Map.groupBy and expose sampled runs

* fix(lens): open older findings and review every merged copy from one row

* fix(lens): hide edit and run now from read-only viewers

* chore(lens): drop restating comments from the findings table

* chore(lens): drop restating comments from the step feed

* chore(lens): drop restating comments from the paused banner

* chore(lens): drop restating comments from header actions

* chore(lens): drop restating comments from run now

* test(lens): cover merged findings and read-only investigation rows

* fix(lens): record a model step even when the response has no usage

* test(lens): cover model steps with and without reported usage

* fix(lens): keep a merged finding open when one of its updates fails

* refactor(lens): accept update results from the investigations view

* refactor(lens): accept update results in investigation actions

* test(lens): cover retrying a merged finding after a failed update
2026-10-03 20:27:02 +00:00
devin-ai-integration[bot]
f445e466b4
refactor(traces): extract snapshot cache (#44424)
Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 13:03:30 -07:00
devin-ai-integration[bot]
671748067d
fix(model_prices): registry audit 2026-10-03, add cohere embed v5, grok-imagine-video-1.5-lite and openrouter gpt-image rows (#44376)
* fix(model_prices): registry audit 2026-10-03, add cohere embed v5 and absorb openrouter gpt-image rows

Co-authored-by: tinysolver <iam.tinysolver@gmail.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): add grok-imagine-video-1.5-lite and gemini 3.5 transcribe limits, drop deprecation-only date

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: tinysolver <iam.tinysolver@gmail.com>
2026-10-03 12:53:07 -07:00
devin-ai-integration[bot]
d306d6d70b
feat(traces): render curated SQL examples from shared files (#44388)
* feat(traces): render curated SQL examples from shared files

* fix(traces): filter trace-summary example after grouping so boundary-spanning traces keep full totals

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(traces): order failed-spans example by timestamp so newest failures survive the limit

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 12:40:01 -07:00
devin-ai-integration[bot]
1b98748528
fix(bedrock): honor the per-request timeout on Converse and Invoke streaming (internal copy of #38210) (#44134)
* fix(bedrock): propagate timeout to streaming requests

* test(bedrock): prove streaming fails at the request timeout against a slow upstream

* test(bedrock): simulate the slow upstream in process instead of over a local socket

* test(bedrock): audit the Converse and Invoke stream timeout on the proxy

Two integration files drive the per-request timeout on Bedrock streams
through the real proxy against an owned wire peer: the wire file covers
every surface (chat, messages, responses, invoke, pass-through), the
sad, edge and precedence rows, and the chaos file covers bursts, a
dropping upstream, a killed worker and a proxy stopped mid-burst.

The wire peer gains Reply.drop_connection so a cell can close the
socket before any response, and the harness's graceful stop grace is
now INTEGRATION_PROXY_STOP_SECONDS (default unchanged at 30), since a
two-worker supervisor's interpreter finalization takes longer than that
on a loaded box.

* test(bedrock): pin the fallback audit cell to one proxy worker

The fallback cell created both deployments through /model/new on one
worker and sent the chat request to the other, whose registry
read-through loads only the requested model, so the fallback target
was unknown there until the periodic DB poll. The cell now warms the
fallback model and sends the request over one keep-alive client, so
one TCP connection stays with one uvicorn worker, and it expects the
fallback upstream to see both requests.

---------

Co-authored-by: Sainyam Kapoor <hello@sainyam.me>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 19:29:15 +00:00
ishaan-berri
9e9c29f404
feat: add make lens-dev for one-command lens local dev (#44413)
* feat(lens): add scripts/lens_dev.sh for one-command lens local dev

* feat: add make lens-dev target

* chore: gitignore .lens-dev local state

* docs(lens): mention make lens-dev in the developer note

* fix(lens): resolve a relative LENS_DEV_CONFIG against the caller's cwd

* fix(lens): random lens-dev master key, verify reused services, drop inherited REDIS_*

* test(lens): cover lens-dev worker token, master key, env and cleanup

* fix(lens): skip compose postgres when LENS_DEV_DATABASE_URL is set

* test(lens): external LENS_DEV_DATABASE_URL never starts compose postgres
2026-10-03 19:10:04 +00:00
moe-berri
af36e5c693
fix(lens): preserve full trace access and expose investigation failures (#44406)
* fix(lens): preserve full trace access and expose investigation failures

* chore(lens): sync worker registration schema

* fix(lens): support durations without a configured maximum

* fix(lens): expose every page of fetched investigation evidence

* fix(lens): preserve repeated trace content and interrupt cancelled runs

* fix(lens): retry transient heartbeat failures during analysis
2026-10-03 11:52:55 -07:00
moe-berri
50190134c3
fix(lens): batch run reads and reset trace pagination (#44398)
* fix(lens): batch run reads and reset trace pagination

* fix(lens): scope batched list spend to each run
2026-10-03 18:43:23 +00:00
berriai-litellm-provider-info-sync[bot]
9fe6442172
fix(azure): add azure_ai/flux.2-pro input token limit from the models sold directly page (#44405)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 11:40:44 -07:00
devin-ai-integration[bot]
a57cdfddf8
fix(ci): give the pass-through auth regression test a real FastAPI app (#44403)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 18:31:48 +00:00
devin-ai-integration[bot]
a849093d89
test(cost): pin OpenAI reported web search count with mixed actions (#44414)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 11:31:45 -07:00
devin-ai-integration[bot]
6c32384d8c
fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3 (#44292)
* fix(bedrock): stop emitting Converse cachePoint blocks for Kimi K3

Bedrock prices Kimi K3 cache reads through implicit caching but Converse
rejects the explicit cachePoint marker ("This model doesn't support the
cachePoint field"), so any cache_control on the request answered 400.
Mark the three K3 rows supports_prompt_cache_breakpoint: false and have
bedrock_model_accepts_cache_points honor that flag before falling back to
supports_prompt_caching, keeping cached-token pricing intact.

* test(bedrock): assert cache points per request section

* fix(bedrock): honor a deployment's cache breakpoint flag for unmapped models

* fix(bedrock): read a converse-routed deployment's cache breakpoint flag

* refactor(bedrock): look up cache breakpoint flags by key

* test(bedrock): add the Kimi K3 cache point wire audit

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 18:11:46 +00:00
tin-berri
4732647de2
feat(ui): make LiteAdmin enterprise-only (#44399) 2026-10-03 11:07:54 -07:00
ishaan-berri
f0abe1bea1
feat(lens): live dot field timeline and full-screen traces view (#44390)
* feat(lens): add dot field layout and live status helpers for the traces timeline

* test(lens): cover dot field layout, agent colors and live status

* feat(lens): draw the traces timeline as a live dot field

* feat(lens): add the sweep animation for the live traces timeline

* feat(lens): put the lens tabs in a compact header and fill the screen with traces
2026-10-03 10:56:48 -07:00
moe-berri
ad8babae33
fix(lens): paginate trace reads within ClickHouse limits (#44384) 2026-10-03 10:38:40 -07:00
devin-ai-integration[bot]
8b1990b4bc
feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers (#44236)
* feat(decisions): add unified /v1/decisions endpoint for Jev-compatible providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register typesafe as a provider so Jev deployments load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(decisions): move provider endpoints under llms and validate proxy bodies

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add Cloudflare Clef and Strands Decider backends

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register decisions routes for managed agents and gateway

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(decisions): use raw regex for cloudflare missing account match

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): avoid cast in Cloudflare response unwrapping

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): default model, evaluation health probe, short Cloudflare names

The proxy validates only state and questions, so a request without a
model falls through to the configured default model like every other
route. Health checks probe evaluation-mode deployments through the
Decisions API instead of failing with an unsupported mode, and
cloudflare/clef and cloudflare/clef-flash get cost-map rows so the short
names resolve a mode and a price. The registry no longer claims typed
decisions for a provider with no backend.

* fix(decisions): let health_check_params override the evaluation probe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): audit the decisions endpoint across providers, limits, health and chaos

Adds the /v1/decisions audit cells: one wire contract per provider (path, key, body and cost-map billing), the gateway-only fields and tags, the sad paths (invalid bodies, unknown model, key checks, api_base in the body, upstream 401/429/500, a 200 without answers, an unreachable upstream), the two evaluation-mode health probes, and three chaos cells (a mixed-failure burst over both routes, a worker SIGKILL mid-burst, an upstream outage and restart on the same port).

The PR's cost case read the upstream observations through the gateway, which answers 404 for that path; it now reads them from the upstream URL. The owned proxy harness takes extra CLI arguments, and its graceful stop waits as long as a worker boot may take, since a worker still starting honors SIGTERM only once it is up and the 30 second wait forced a cleanup under load.

* fix(decisions): send env API keys to a configured api_base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(decisions): add zero-cost evaluation cost-map entry for Strands Decider

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(decisions): register the routes through the lazy feature registry

The Decisions router was included at import, ahead of the config and DB
pass-through endpoints, so a pass-through configured at /v1/decisions
was skipped and answered 400 as an unknown Decisions provider. The
routes now register through LAZY_FEATURES, which splices them in after
every eager route, so a pass-through at /v1/decisions keeps its route
while /decisions still serves natively. The lazy OpenAPI snapshot carries
the two paths so the schema shows them before the first call.

The audit cells add the env-key egress to a configured api_base, the
client api_base opt-in shared with chat, the pass-through precedence on
an owned proxy, and the Strands evaluation health check resolved from
the cost map. The integration config exports the Perplexity env key the
first cell needs.

* fix(decisions): keep the Cloudflare api_base message in its transformation and read the audit upstream once per cell

* fix(proxy): let a config pass-through beat a lazily registered route in eager mode

With LITELLM_DISABLE_LAZY_ROUTES set the decisions routes are registered at
startup, so SafeRouteAdder treated a config pass-through at exactly
/v1/decisions as already registered and dropped it. In lazy mode a pass-through
created through the API after the first native call was skipped the same way.
Routes a lazy feature owns no longer count as registered, and a route added at
one of their paths is placed ahead of them, the precedence lazy mode gives a
config pass-through when the feature has not loaded yet.

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-03 17:38:38 +00:00
devin-ai-integration[bot]
b024950353
feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop (#44391)
* feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

* refactor(harness): name per-tool spec FunctionTool

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
2026-10-03 17:33:01 +00:00
berriai-litellm-provider-info-sync[bot]
231a46e40b
feat(azure): add azure_ai/kimi-k2-thinking from Azure Kimi pricing page (#44382)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 10:04:59 -07:00
devin-ai-integration[bot]
a5e9f275ea
fix(proxy): keep the database error when the log_db_metrics failure hook raises (#44383)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 16:38:15 +00:00
devin-ai-integration[bot]
4c6c84afb7
perf(proxy): stop prompt-cache eligibility from tokenizing the whole conversation (#44221)
is_prompt_caching_valid_prompt ran the full Python token_counter over every message to compare against the deployment's prompt cache minimum, 500 to 1000 ms at 440k to 740k tokens on every request that reaches the prompt_caching pre-call check, Rust on or off. messages_reach_token_count does the same arithmetic as token_counter(...) >= threshold and stops at the first message that reaches the threshold. Groups with one healthy deployment skip the prefix hash and pin lookup, which cannot change the result for them

Four fixed name span events make the pre-LLM phases measurable with OTel v2: litellm.request.body_received (with body_bytes) once per body read before parsing, on the JSON, binary and form branches, body_parsed, pre_call_completed, and deployment_selected emitted once per pick inside Router.async_get_available_deployment and get_available_deployment with attempt, reason and model group, so every router surface, retry and fallback is covered. Measured locally on /v1/chat/completions, /v1/messages and /v1/responses at 440k tokens with Rust on and off against a fake upstream

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
797353f13a
fix(otel): name postgres service spans by operation and table (#44240)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:28 -07:00
devin-ai-integration[bot]
564d236985
fix(otel): nest cache spans under their operation and name service spans by purpose (#44150)
Response cache reads and writes open cache.get llm_response and cache.set llm_response phase spans with their Redis spans nested underneath, on the Python path and on the native Rust path, and deployment selection runs inside a route {model_group} phase so the cooldown, usage and model-id reads the router issues nest under it before chat {model}. The autorouter classifier call nests under that route phase as well and carries its typed internal origin on litellm.request.purpose, so it is told apart from the provider attempt. Service spans are named {service}.{verb} {target} from a low-cardinality key family the producer declares (llm_response, auth_objects, spend_counters, router_cooldowns, claude_code_session_router_binding, rate_limits, pod_lock, budget_reset, ...) instead of the raw method or a per-request pipeline length; a pipeline flush is targeted by the one family its ops share or by mixed with the sorted families on litellm.redis.families, a batch op keeps the family it was declared under whichever pipeline or standalone read settles it, and the ambient family labels Redis spans only, never the DB write-back a task spawned inside that context performs later. The raw method stays on litellm.service.call_type and on the Prometheus and Datadog labels. Caller attribution is carried across asyncio task boundaries on a ContextVar so forwarder-only chains no longer surface, the raw cache key is dropped from Redis span metadata, pipeline op counts land as an integer attribute, every call_type the Redis cache layer emits maps to a verb, and a scan over litellm/ and enterprise/ fails when a Redis producer, batch reservation included, declares no key family.

A V2 logger built for a key or team logging entry while the operator's V2 logger is already registered keeps only the exporters its own preset contributed, whether or not the operator holds credentials for that backend, so every chat span no longer reaches the operator's collector twice. A span the success callback has to open itself, with no pre-call carrier, starts at the provider handoff (api_call_start_time) instead of the logging object's creation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
devin-ai-integration[bot]
53d2ab6b05
fix(proxy): emit postgres service spans only on real DB reads in auth cache helpers (#44148)
log_db_metrics wrapped whole cache-first auth helpers and always emitted a ServiceTypes.DB success event, so in-memory cache hits showed up as postgres <fn> spans and DB service metrics. The decorator now installs a ContextVar witness that _TrackedPrismaEngine marks on every Prisma query and transaction call, and the DB event is emitted only when the witness was marked. Real reads keep their existing call_type names, the failure path and the PROXY batch-write branch are unchanged, and Redis instrumentation is untouched.

A decorated helper that reaches Prisma only through another decorated helper (get_key_object -> get_object_permission, get_team_object_by_alias -> get_object_permission, get_tag_object -> get_tag_objects_batch) used to emit two events for one query. The inner wrapper now marks its witness as reported when it emits a success or DB failure event, and only unreported activity is handed up to the enclosing witness, so the inner event is the one that survives. An outer helper that also queries Prisma directly or through undecorated callees still gets its own event.

Tests: get_user_object and get_org_object cache hits emit no DB event; a get_user_object miss through the generated Prisma client emits exactly one postgres get_user_object event; decorator-level tests cover nested calls emitting only the inner event, outer calls with their own query, inner non-DB failures, bounded lookups and sibling-request isolation.

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:20:27 -07:00
berriai-litellm-provider-info-sync[bot]
66a422ea50
fix(azure): add MAI-Image max_input_tokens from models sold directly page (#44375)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-03 08:49:09 -07:00
devin-ai-integration[bot]
4ece6c9fb8
build(deps): suppress unfixed braces GHSA-vfj7-8cjw-p6xm to clear osv-scan (#44347)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 07:55:14 -07:00
devin-ai-integration[bot]
e768ad55ce
refactor(types): replace Any with proven types in 4 files (#44370)
* refactor(types): replace Any with proven types in 8 files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert unproven email logger protocol and iterator annotation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(types): revert unproven deepagents and http handler annotations

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 03:52:28 -07:00
devin-ai-integration[bot]
e340e546e2
feat(traces): tracing development seed (#44363)
* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* feat(dev): seed linked tracing and spend fixtures

* chore(dev): use OpenAI model in tracing config

* chore(dev): align tracing credentials with UI E2E

* fix(dev): update fixture seeder query scope

* wip

* wip

* wip

* chore(trace): checkpoint ongoing Rust migration

* refactor(trace): group Python bridge under trace package

* refactor(traces): read span conventions through a Convention trait

Each span format (Claude Code, LangSmith, OpenInference, gen_ai) now lives under
normalize/convention/ as a unit struct implementing Convention, owning both its
detection and its extraction. Precedence is one ordered registry instead of an
if-chain in mod.rs that reached into each module differently.

The modules now share one way to read attributes: present() for the first
non-empty key and Payload for a text that also reports the key it consumed,
replacing three different idioms and the &mut Vec threaded through payload
readers. Instrumentation::adjust returns a new Extraction instead of mutating
one, with each SDK rule as its own function, and the LangChain middleware
suffix list exists once.

* feat(trace): export Rust-owned wire schemas and enforce contract bounds

* fix(trace): bound quoted counts in ClickHouse wire schemas

* feat(trace): generate Python wire contracts with datamodel-code-generator

* test(trace): validate migrated callers and generated contracts at the native boundary

* refactor(traces): rename normalization convention to format

* fix(traces): reconcile spend evidence and preserve unknown costs

* feat(traces): normalize additional telemetry formats

* test(traces): cover captured normalization fixtures

* refactor(traces): isolate SDK normalization rules

* feat(tracing): seed all trace exports for local dashboard

* fix(clickhouse): preserve custom LiteLLM request metadata

* docs(traces): define normalization module boundaries

* docs(traces): define resolution and OTLP boundaries

* fix(ui): normalize nullable trace message names

* refactor(traces): split resolver modules and cover resolution behavior

* test(traces): replace normalization snapshots with behavior assertions

* fix(ui): align dashboard API contracts with generated types

* refactor(traces): type normalization and storage boundaries

* fix(traces): seed captured SDK spend and preserve provider identities

* wip

* test(traces): verify guide discovery and content ordering

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 09:59:51 +00:00
devin-ai-integration[bot]
d260765652
refactor: clean up fresh tech debt from 2026-10-02 (#44362)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 02:12:03 -07:00
devin-ai-integration[bot]
e92bd50de1
refactor(ui): inject Lens backends as services instead of a fake HTTP client (#44369)
The interactive Lens demo used to build a fetch shim that encoded in-memory
fixtures as HTTP responses so the shared ApiClient could decode them again,
and every trace view branched on demo vs live to pick a URL. Lens and the
trace views now depend on two small service interfaces, LensApi and
TracesApi, with named operations. The live layer wraps the existing HTTP
calls, the demo layer reads fixtures directly, and a React context provides
whichever one the session runs on. Without a provider the hooks fall back to
the live implementation, so the live app and existing tests are unchanged.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 09:11:08 +00:00
devin-ai-integration[bot]
09313bcf1b
refactor(ui): reorganize Lens dashboard components (#44354)
* refactor(ui): reorganize Lens dashboard components

* refactor(ui): align Lens forms with react-hook-form

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): keep Lens submit errors out of form validity

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): reduce Lens lint budget usage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): use a query key factory for Lens queries

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 08:29:16 +00:00
moe-berri
cb00dbecd7
feat(ui): improve trace inspection and ROI estimation (#44351)
* feat(lens): simplify trace inspection in the gateway drawer

* feat(lens): add a conversation view for full traces

* test(lens): keep normalized message fixtures type safe

* refactor(lens): make conversation view read like a chat

* refactor(lens): use a quiet trace view menu

* refactor(lens): make trace view tabs explicit

* refactor(lens): restore compact trace view switch

* fix(lens): preserve complete conversation history and tool types

* feat(lens): add trace full-screen and close controls

* fix(lens): show forwarded answers and agent errors once

* fix(lens): reset full screen when closing a trace

* fix(roi): estimate linked authors and clarify model selection

* fix: preserve trace errors and ROI results across partial failures

* fix(roi): correct pagination variable typing

* fix(ui): place loaded root failures in conversation order

* fix(roi): read estimator recommendations from model catalog

* revert: remove catalog-driven ROI recommendations
2026-10-03 01:18:16 -07:00
devin-ai-integration[bot]
5724117116
fix(responses): merge bridged tool calls into the same choice as the text (#44346)
* revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295)

This reverts commit ca1994e403.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): merge bridged tool calls into the same choice as the text

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 07:00:52 +00:00
tin-berri
8e32d4568c
feat(ui): move LiteAdmin into the header with a docked side panel (#44293)
The floating bottom-right LiteAdmin button covered page controls such as
the Logs pagination buttons, and Playground had to hide it entirely.
Render the trigger as a pill in the header tools ahead of Docs and open
LiteAdmin as a panel docked beside the content column, which narrows the
page instead of covering it. Add a Cmd/Ctrl+J toggle and drop the
Playground override.

The Logs and trace drawers treated Cmd+J as a plain J and advanced the
selection, so they now share RunDrawer's rule that letter shortcuts
yield to modified presses and typing.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 00:00:10 -07:00
devin-ai-integration[bot]
0c86d6bfc9
refactor(dashboard): migrate to zod 4 and openai 6 (#44345)
* refactor(dashboard): migrate to zod 4 and openai 6

Bump the dashboard to real zod 4.6.5 and openai 6.49.0 so every module
imports from bare "zod" instead of "zod/v4". Ports the ten files that
still used the zod 3 API (error params, record, passthrough, strict,
email/date validators, union discriminator codes) and adapts the
LiteAdmin tool schemas and form plumbing where openai 6 and zod 4
changed behaviour.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(dashboard): prettier format zod 4 schema files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(dashboard): restore system one missing state message under zod 4

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:53:41 +00:00
devin-ai-integration[bot]
d96e56c76f
revert(responses): revert "fix(responses): keep gpt-5.4/5.5 tool calls on chat and merge bridged tool calls into one choice" (#44295) (#44344)
This reverts commit ca1994e403.

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:47:11 +00:00
devin-ai-integration[bot]
7b432d78d2
fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model (#44341)
* fix(proxy): stamp the client alias on a copy of each streamed chunk so pricing sees the deployment model

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): streamed alias matching a capability rule bills the deployment price

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(spend): assert every streamed chunk carries the client alias

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): log the client alias on the priced streamed response, the same as non-streamed

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:44:52 +00:00
devin-ai-integration[bot]
a76ba8c01e
revert(cost): revert "fix(cost): price rule-only model names at the deployment's rate" (#44144) (#44335)
This reverts commit f9a32ffcb5.

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 06:09:00 +00:00
devin-ai-integration[bot]
0fce5bccbc
fix(model-prices): mark azure us/eu responses-only models as mode responses (#44323)
* fix(model-prices): mark azure us/eu responses-only models as mode responses

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model-prices): keep azure us/eu o3-deep-research on chat, which Azure lists as chat capable

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 23:05:55 -07:00
moe-berri
caee45fed4
feat(roi): add GitLab sources and branch cost attribution (#44324)
* feat(roi): support GitLab and tagged branch costs

* fix(roi): count tagged branches independently of estimation status

* test(roi): capture live GitHub and GitLab report validation

* fix(roi): open estimate details at the start

* fix(roi): clarify cost views and unify report layout

* feat(roi): showcase per-PR costs in the sample report

* fix(roi): separate report tabs and preserve branch cost attribution

* fix(roi): preserve demo previews and align progress spacing

* fix(roi): isolate demo loading and parallelize fork lookups

Preserve active sync status when source changes finish saving, keep live reports available when demo requests fail, and cover each review regression

* fix(roi): separate demo and live loading states

Clear the demo URL on fallback, wait for live requests on exit, and retain request errors until the corresponding operation recovers

* fix(roi): ignore refreshes from a previous source

* fix: trust gateway context for ROI estimator exclusion

* fix: preserve historical ROI estimator exclusion
2026-10-03 06:01:43 +00:00
yucheng-berri
12d4b75b7a
fix(search_tools): encrypt search tool litellm_params at rest (#43631)
* fix(search_tools): encrypt search tool litellm_params at rest

Encrypt every string value of a search tool's litellm_params on create and
update and decrypt on every DB read, so legacy plaintext rows load unchanged.
Include the table in master key rotation, LITELLM_MIGRATE_FROM_MASTER_KEY and
the migrate-encryption scan.

* fix(search_tools): keep edits made while the master key rotates

Write each rotated search tool row only if it still holds the litellm_params
that were read, and re-read and rotate it again if it was edited in between,
so a PUT that lands during /key/regenerate is not overwritten.

* fix(search_tools): retry rotation writes until the row stops changing

Rotate a search tool row again for as long as it keeps being edited instead of
giving up after five attempts, and stop with a warning only when the conditional
write fails on an unchanged row. Build the decrypted read result without
mutating it in place.

* refactor(search_tools): rotate edited rows in a loop, drop the step comment

Retry the conditional rotation write in a loop instead of recursion so sustained
edits cannot deepen the call stack, drop the step comment on the rotation call,
and stop mutating local state in the rotation tests.

* test(search_tools): drop the rotation test docstring

* Store search tool params as written when no encryption key is configured

* Rotate search tools under the salt key, keep non-ciphertext values and loaded tools that do not decrypt

* Treat a search tool as undecryptable only when its provider is ciphertext-length

* Drop suppressions the type discipline gate on main now reports as unused

* Show the loaded search tool in the admin list and info views when its DB params do not decrypt

* Keep the DB row's other fields when the admin views substitute loaded params
2026-10-02 22:53:18 -07:00
yucheng-berri
5d42cb7cfa
fix(guardrails): encrypt guardrail litellm_params secrets at rest (#43627)
* fix(guardrails): encrypt guardrail litellm_params secrets at rest

* fix(guardrails): keep salt-key encryption on master key rotation and retry rows edited mid-rotation

- rotate guardrail params under LITELLM_SALT_KEY when set, matching the key reads decrypt with
- re-read and retry a row whose updated_at moved during rotation, up to GUARDRAIL_ROTATION_ATTEMPTS
- build decrypted Guardrail rows and the rotation count without mutating locals

* refactor(guardrails): retry guardrail rotation by bounded recursion instead of a rebound cursor

- each attempt re-reads the row and recurses with attempts_left - 1, so no loop variable is rebound
- cover the give-up path after GUARDRAIL_ROTATION_ATTEMPTS writes

* test(guardrails): drive the real guardrail rotator from the master key rotation test

- inject an encrypted guardrail row through the prisma client instead of replacing the GuardrailRegistry method
- assert the written params decrypt under the new master key

* Annotate guardrail param encryption collections for type-discipline gate

* Type guardrail param recursion through validated JSON containers

* Type guardrail registry test helpers and drop section comment

* Reject client-supplied encrypted values in guardrail litellm_params

* Allow depth-bounded contains_encrypted_marker in the recursion detector

* Keep a loaded guardrail when its DB params do not decrypt with the current key

* Apply other DB edits while keeping loaded values that do not decrypt, including PATCH models

* Keep the loaded guardrail when an undecryptable param has no loaded value

* Drop suppressions the type discipline gate on main now reports as unused

* Assert what the reinitialized guardrail holds after an edit to an undecryptable one

* Drive the rotation sync tests through a registered guardrail instead of patching reinitialize

* Type the rotation test helpers and drop the new test docstrings

* fix(guardrails): refuse to approve a submission whose params do not decrypt

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:49:41 -07:00
devin-ai-integration[bot]
6d8434f940
fix(proxy): return 4xx instead of 500 for missing required params, invalid pagination and unknown ids (#43787)
* fix(proxy): return 400 instead of 500 for missing required params and invalid pagination

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: run search_endpoints tests in proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): return 4xx for missing required params across all LLM routes and propagate provider status on lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(llm_http_handler): keep provider error text when re-raising mapped errors

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): allow promptless image edits and default search models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): default missing image edit image to None

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(proxy): build image edit defaults without mutating request data

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): keep image edit defaults within type-discipline budget and give request mocks a scope

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): inject a fake router for the search default model test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): cover provider error status on vector store and file lookup handlers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): keep the lookup handler raise block to a single statement

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(llms): cover provider error status on eval, eval run, skill and vector store file content lookups

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover missing required body params and provider lookup status codes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: run tests/unit/proxy/search_endpoints in the proxy-endpoints shard

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): bind spend-row request id with partial to satisfy B023

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): only reject non-positive page_size on vector store list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): remove unreachable fine-tuning body validation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover streaming anthropic messages reaching the upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): count only provider calls when asserting missing params never reach the upstream

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve merge-base request compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): preserve interaction completion model defaults

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): retry model read-through before rejecting params a DB-only deployment may default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng <yucheng@berri.ai>
2026-10-02 22:48:53 -07:00
devin-ai-integration[bot]
8efb4a21f6
fix(vector_stores): return managed file ids from vector store file list (#43800)
* fix(vector_stores): return managed file ids from vector store file list

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(vector_stores): cover managed file list route

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(vector_stores): only map round-trippable managed ids and index flat file ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy-extras): build managed file gin index concurrently

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(proxy-extras): move the managed file gin index migration after main's newest

* fix(vector_stores): satisfy lint gate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(proxy): drop the stale no-index note on the raw file-id guard

* test(vector_stores): cover managed file ids on the vector store file list end to end

Integration cells for GET /v1/vector_stores/{vs}/files mapping provider file ids back to
the caller's owner-scoped managed ids and decoding managed after and before cursors: raw
httpx, the OpenAI SDK sync and async pagers, the three credential routing modes, the owner
filter branches, raw and unmappable cursors, provider errors, duplicate and non-string ids,
a provider outage mid-burst, a worker SIGKILL mid-burst, and the GIN index migration applied
by the migration entrypoint and by db push

---------

Co-authored-by: shivam <shivam@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-02 21:28:50 -07:00