Commit graph

53798 commits

Author SHA1 Message Date
Simon
48c88bae04
fix(responses): preserve prompt cache reuse in chat bridge (#42281)
* fix(responses): preserve prompt cache breakpoints in chat bridge

* fix(responses): preserve multimodal cache breakpoints

* fix(cache): retain implicit lookup with injected breakpoints

* test(cache): expect implicit responses lookup

* test(cache): expect implicit chat lookup
2026-10-06 17:50:20 -07:00
moyai-devin-berriai[bot]
5dd3ff1714
feat(mistral): add Mistral Large 4 (Le Chonk) support (#44870)
Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
2026-10-07 00:35:19 +00:00
berriai-litellm-provider-info-sync[bot]
a5a86cdda7
chore(cost-map): sync openrouter prices from the models API (#44975)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 17:20:09 -07:00
devin-ai-integration[bot]
47afd1ba0b
test: fix shared-provider discovery, Codex catalog size and generated master key mismatches in CircleCI suites (#44905)
* test(integration): isolate Codex catalog and provider discovery fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): keep model discovery constant import-safe

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: align test keys and CI env with generated master keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 00:11:25 +00:00
devin-ai-integration[bot]
c3a23fe499
refactor: expose core private helpers under public names (#44871)
* refactor: expose core private symbols with compatibility aliases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* types: narrow core migration diagnostics

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve runtime behavior in core symbol migration

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve core private usage migration behavior

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve private value rebinding compatibility

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: restore optional imports and cover public helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: exempt router property from call coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: align recursive detector ignore names

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-07 00:05:27 +00:00
moyai-devin-berriai[bot]
b9ab1eee1f
fix(proxy): stop logging license values during verification (#44956)
* fix(proxy): stop logging license values during verification

* fix(tests): address license logging review feedback

---------

Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
2026-10-06 17:05:07 -07:00
ryan-crabbe-berri
0b0fdedd1e
test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata (#44949)
* test(e2e): add enum values, auto-discovering label gates and secret hiding for e2e metadata

* docs(e2e): name every markerless harness test file that carries no Subject

* test(e2e): keep the step discovery comprehensions to one for clause
2026-10-06 17:03:35 -07:00
yucheng-berri
cb76270bd9
fix(terraform): keep unconfigured allowed_routes plan-known and unsent (#44487)
* feat(terraform): expose key type on virtual keys

* docs(terraform): remove in-tree key type docs

* fix(terraform): preserve server-derived key routes

* fix(terraform): keep unconfigured key routes plan-known and unsent

Two regressions from exposing key_type on litellm_key:

1. Marking allowed_routes Computed makes an omitted attribute unknown at
   plan time ("known only after apply"), so any plan that consumes it
   before the key exists fails, e.g.
   for_each = toset(coalesce(litellm_key.x.allowed_routes, [])).
   Computed is dropped again; server-derived routes still land in state
   through reads, and a DiffSuppressFunc keyed on the raw config keeps a
   config that never declares the attribute from showing a perpetual
   removal diff against those routes (a config that shrinks the list or
   sets it still diffs).

2. mapResourceDataToKey copies allowed_routes unconditionally and
   UpdateKey sends it when non-empty, so once reads materialize the
   server's routes into state, every update re-asserts them: an
   alias-only rename POSTs allowed_routes (the pre-key_type provider
   sent none), and with stale state (-refresh=false) it silently
   overwrites routes managed outside Terraform. Updates now omit the
   field whenever the raw config does not declare it.

The key_type flow is unchanged: create still sends key_type, the proxy
presets the routes, reads materialize them into state, and plans stay
drift-free.

* fix(terraform): reject allowed_routes alongside a presetting key_type

The proxy derives allowed_routes from the key_type preset and overwrites
whatever the request declared, so a config combining the two could never
match what gets stored: the key came back with the preset routes and
drifted against the declared list on every plan. A CustomizeDiff now
fails the plan with an actionable message when a presetting key_type
(llm_api, management, read_only) is combined with allowed_routes.
key_type "default" presets nothing and keeps declared routes.

* fix(terraform): scope key_type route rejection to create-shaped plans

/key/update stores an explicit allowed_routes verbatim and never reapplies
the key_type preset, so an existing or imported typed key can manage its
routes in place. Only plans that create a key (fresh, or a replacement
that changes key_type) still reject the combination, because there the
preset always overwrites the declared list. A replacement forced by
another ForceNew attribute converges on the next apply, which re-sends
the declared routes.

* fix(terraform): restore declared routes on typed key creation

/key/generate replaces a declared allowed_routes with the key_type
preset while /key/update stores the list verbatim, so any create that
carries both (a fresh key, or a replacement forced by key_type or
another ForceNew attribute) used to leave the key holding the preset
instead of the declared routes until a second apply. When the generate
response does not match the declared list, create now follows up with an
update that re-sends the full create payload against the new key hash,
so the first apply already stores the declared routes. This also
replaces the plan-time rejection of the combination: every config shape
now converges, and existing typed keys keep managing routes in place as
before.

* fix(terraform): delete the key when a route restore fails at create

If /key/generate succeeds but the restore update is rejected, the key
exists server-side while terraform holds no state for it: an active key
with the type preset would be orphaned and a retried apply would mint
another one. The restore failure path now deletes the created key, and a
delete that also fails names the key hash in the error so an operator
can remove it manually.

* fix(terraform): make the route restore surgical and keep supplied keys

Two sharp edges on the create-time route restore:

- Re-sending the full create payload rewrote fields the config never
  declared: /key/update is a merge patch, so the empty metadata and
  model_rpm_limit/model_tpm_limit maps the restored struct carried would
  clear server-applied values such as team-inherited rate limits. The
  restore now sends only the routes plus the two fields /key/update
  requires non-null (permissions, model_max_budget); every other stored
  value is kept.
- /key/generate upserts a config-supplied key value, so a restore
  failure on such a key must not delete it: it may be an existing
  credential that predates this apply. The compensating delete now runs
  only for proxy-minted keys, and the error names the hash either way.

* fix(terraform): echo stored permissions and budgets in route restore

The surgical restore body carried empty permissions and model_max_budget
objects, and /key/update writes fields that are present: a key created
with declared permissions or model budgets next to a presetting key_type
and allowed_routes lost them on the first apply. The restore now echoes
the values /key/generate just stored (falling back to the configured
values when the response omits them), so the only field the restore ever
changes is allowed_routes.

* test(terraform): pin echoed budgets in the route restore

Adds the nonempty model_max_budget case Greptile asked for (the restore
must echo the stored map, never clear it) and drops a comment that
restated its own line.

* test(terraform): assert the declared budget reaches key generation

The budget echo case fed the raw config a malformed JSON string (a
template leftover), so nothing verified the declared budget actually
reached /key/generate. The config now carries the valid JSON and the
generate payload is asserted to match it.

* chore(terraform): trim the restore test preface to the proxy facts

---------

Co-authored-by: Roman Soletskyi <roman@mistral.ai>
2026-10-06 16:50:45 -07:00
ishaan-berri
d8bc2b78e4
fix(lens): show the first user message as the run input (#44958)
* fix(traces): use the first user message for the run input preview

* test(traces): cover first user message as the input preview

* test(traces): check fixture previews against the first user message

* fix(lens): show the whole input preview on one line in the runs table

* test(lens): cover multi-line input previews in the runs table
2026-10-06 16:30:52 -07:00
tin-berri
9a3f000c9b
fix(cli): show full-session auto-router cost comparison (#44959) 2026-10-06 16:24:14 -07:00
ishaan-berri
3c79db6122
feat(lens-ui): replace the conversation view with a thread view (#44947)
* feat(lens-ui): group a trace conversation into prompt, work and reply turns

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): cover thread turns, repeated prompts and subagent work

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens-ui): export conversation step renderers for reuse

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): add a Thread view with a folded Worked bar per turn

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): prove the Thread tab folds work and opens the exact step

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): add thread to the trace view routes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): add a Thread tab to the run header

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): render the Thread view in the run body

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(lens-ui): let the thread view retry steps that failed to load

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): cover retrying a failed step in the thread view

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens-ui): move step and message renderers into ConversationParts

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): show run errors and capture warnings in the thread view

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): cover failed tools, run errors, warnings and subagents in thread

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(lens-ui): skip the missing replies warning when an answer was recorded

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): cover the missing replies warning with a recorded answer

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens-ui): drop conversation from the trace view routes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(lens-ui): give the Steps and Thread tabs icons and drop Conversation

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(lens-ui): stop rendering the conversation view in the run body

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(lens-ui): switch the workspace drawer test to the Thread tab

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 23:08:27 +00:00
ishaan-berri
5fb23eabd8
feat(lens): add datasets built from real traces (#44765)
* feat(lens): add dataset case and size limits

* feat(lens): add dataset, case and build models

* feat(lens): build dataset cases from traces, findings and text

* feat(lens): store dataset revisions insert-only

* feat(lens): add dataset routes for build, save, export and eval cases

* feat(lens): mount the dataset router before lens routes

* feat(lens): add LiteLLM_LensDataset table

* feat(lens): add LiteLLM_LensDataset table to proxy schema

* feat(lens): add LiteLLM_LensDataset table to extras schema

* feat(lens): add migration that creates the dataset table

* test(lens): cover dataset case building, dedupe and limits

* test(lens): cover dataset revisions, conflicts and eval cases

* chore(ui): regenerate API types for lens datasets

* feat(lens): add dataset UI types

* feat(lens): add datasets API client

* feat(lens): add dataset query and mutation hooks

* feat(lens): add case selection and expected edit logic

* test(lens): cover case selection and expected edits

* feat(lens): add the add to dataset dialog

* test(lens): cover saving picked cases from the dialog

* feat(lens): add datasets list

* feat(lens): add dataset detail with revisions and export

* test(lens): cover editing, revisions and export in datasets tab

* feat(lens): expose datasets on the lens API

* feat(lens): add in-memory datasets for demo mode

* feat(lens): wire demo datasets into the demo lens API

* feat(lens): add optional lens API hook

* feat(lens): add optional onboarding hook

* feat(lens): add datasets tab and dataset routing

* feat(lens): show the datasets tab

* feat(lens): add to dataset from the trace header

* feat(lens): add a single turn to a dataset from a step

* feat(lens): add finding evidence to a dataset

* docs(lens): add datasets screenshots for the PR

* test(lens): cover dataset revision storage against Postgres

* test(lens): cover dataset trace paging, findings and route errors

* test(lens): cover dataset build fallbacks and no_content skips

* fix(lens): register the dataset table for postgres span names

* feat(lens): add invalid skip reason for unparseable case lines

* fix(lens): keep valid JSONL cases, provenance and size limits; stop reading past the case cap

* test(lens): cover malformed JSONL, re-import provenance, size fields and early cap

* chore(ui): regenerate API types for the invalid skip reason

* feat(lens): show text for the invalid skip reason

* feat(lens): render datasets in the same inspector table as investigations

* test(lens): open a dataset by clicking its table row

* docs(lens): update the datasets list screenshot

* style(lens): format the datasets table

* fix(lens): normalize JSON text span input and output into messages when building cases

* test(lens): cover JSON text span normalization for dataset cases

* Revert "fix(lens): normalize JSON text span input and output into messages when building cases"

This reverts commit 29c3e34962fa12427dbd516c08f2599db997a905.

* fix(traces): normalize agent assistant summaries into UI messages

* feat(lens): add dataset case view helpers built on the trace parsers

* test(lens): cover dataset case view helpers

* feat(lens): show dataset cases in an inspector table

* feat(lens): open a dataset case in a side panel with trace message cards

* feat(lens): rebuild the dataset page header and layout

* feat(lens): keep the open dataset case in the URL

* test(lens): drive dataset edits through the case table and panel

* docs(lens): update dataset view screenshots

* Revert "test(lens): cover JSON text span normalization for dataset cases"

This reverts commit 30229ae8ae90265c09d755c43c1f1eeaa66050c7.

* fix(lens): import StateMessage from shared in dataset detail

* fix(lens): import StateMessage from shared in datasets list

* fix(lens): give the datasets migration a unique timestamp after review checkpoints
2026-10-06 23:04:43 +00:00
Sam Mobach
292fcb1c65
fix(proxy): resolve oidc/ pass-through credentials on every request (#44577)
* fix(proxy): resolve oidc/ pass-through credentials on every request

A pass-through credential (a use_in_pass_through deployment's api_key or
the provider's env var, e.g. TYPESAFE_API_KEY) was sent as a literal
string, so an `oidc/...` reference such as
`oidc/file//var/run/secrets/<name>/token` ended up on the wire as
`Bearer oidc/file/...`, and a rotating projected Kubernetes service
account token could not be used for pass-through auth.

Resolve credentials that start with `oidc/` through get_secret_str() on
every get_credentials() call, so the file is re-read and rotations apply
without a restart. oidc/file/ keeps its credential-directory allowlist;
other credentials are returned unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(proxy): only resolve local oidc/ pass-through credentials inline

get_credentials() runs synchronously inside async pass-through routes, so
resolving oidc/google/, oidc/github/ etc. on a cache miss would make a
blocking HTTP call to the identity provider (timeout up to 600s) on the
event loop. Limit per-request resolution to oidc/file/, oidc/env/ and
oidc/env_path/, which need no network I/O; network-backed references are
left unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(proxy): surface unreadable oidc/ pass-through credentials as 401

A missing or disallowed token file raised FileNotFoundError or
OidcPathNotAllowedError out of the route, which the proxy turned into a
generic 500 on every request. The resolver now raises an HTTPException
401 that names the credential reference and the reason, matching how the
pass-through routes already report an unset key.

oidc/env_path/ is no longer resolved inline because it reads any file
named by an environment variable without the oidc/file/ credential
directory allowlist; it is forwarded unchanged like the network-backed
references.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-10-06 15:55:25 -07:00
moyai-devin-berriai[bot]
2cee61626d
feat(lens): add Copy for agent to investigation details (#44945)
* feat(lens): add Copy for agent to investigation details

* Update ui/litellm-dashboard/src/components/lens/investigations/agentHandoff.ts

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* refactor(lens): build agent handoff selection without reassignment

---------

Co-authored-by: moyai-devin-berriai[bot] <336287033+moyai-devin-berriai[bot]@users.noreply.github.com>
Co-authored-by: moe-berri <moe@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-10-06 15:40:52 -07:00
tin-berri
a193c67347
feat(otel): trace auto-router configuration and classifier failures (#44926) 2026-10-06 15:40:34 -07:00
devin-ai-integration[bot]
2477635213
fix(proxy): bound daily spend rollup row-lock waits with lock_timeout and requeue 55P03 (#44450)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 17:38:55 -05:00
moe-berri
d181bc7b80
fix(lens): refresh open traces without claiming session completion (#44900)
* fix(lens): refresh open traces without claiming session completion

* fix(lens): cancel paused refreshes and distinguish refresh failures

* style(lens): format live refresh regressions

* fix(lens): preserve manual reads when pausing live updates

* fix(tracing): preserve optional provider evidence in fixture replay

* fix(tracing): keep copied provider identities consistent

* fix(lens): refresh resumed native sessions and retain paging

* fix(lens): serialize conversation paging with refresh

* fix(lens): refresh recorded content with trace details

* fix(lens): refresh content using resolved trace references

* fix(lens): preserve content through refresh failures

* fix(lens): serialize conversation paging with content refresh
2026-10-06 15:31:45 -07:00
devin-ai-integration[bot]
191967c207
fix(cost): price batch image output tokens at the batch image rate (#44897)
* fix(cost): price batch image completion tokens at image batch rate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): cover default vertex batch output transformation for image cost

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): sync model prices schema

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): simplify batch completion cost with rate helpers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): keep global batch pricing fallback for image-only deployment rates

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(test): rename connect tunnel helper to https redirect

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(cost): merge deployment batch rates over global pricing

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): put flash-image batch image rate on the GA row

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(cost): drop redundant comment in batch fallback

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(cost): use _batch_or_half for the batch image rate

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost): price nano banana 2.1 fixtures off the real cost map rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(cost): use the vertex 4k image token count in the batch cost test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 22:18:28 +00:00
yuneng-jiang
7566f9bc54
bump: litellm-enterprise 0.1.73 -> 0.1.74, litellm-proxy-extras 0.4.105 -> 0.4.106, litellm 1.105.0 -> 1.106.0 (#44942) 2026-10-06 15:06:47 -07:00
devin-ai-integration[bot]
e7243f1b55
fix(model_prices): drop /v1/batch from gemini/gemini-nano-banana-2.1 (#44940)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 21:47:52 +00:00
devin-ai-integration[bot]
3fec4705d7
fix(projects): show Projects to team and org admins and scope it correctly (#41325)
* fix(ui): show Projects nav to team and org admins

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): satisfy frontend lint budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format leftnav regression tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(projects): let org admins see projects of every team in their orgs

/project/list and /project/info only knew proxy admins and team members, so an
org admin saw projects only for teams they had personally joined. Both now use
the same team access check as /team/info.

* fix(ui): one shared Projects access rule for nav, data and actions

The Projects query skipped global org_admin users, the page-visibility picker
could never offer Projects, and New/Edit showed to users the backend would
reject. Nav and queries now share one rule, Projects is selectable in the
allowlist, New/Edit follow team_admin_editable_team_fields, and the project
modal only lists teams the user administers.

* fix(projects): record IN-list bounds for project visibility filters

Both filters are bounded by one caller's team memberships and admin orgs. Also
cover a project whose team was deleted.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-10-06 14:33:31 -07:00
berriai-litellm-provider-info-sync[bot]
53eaa0c357
feat(azure): add azure_ai/grok-4.7 and update grok-4.6 input price (#44937)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 14:24:41 -07:00
Hektor Jacynycz Garcia
157d96dcd0
fix(logging): log spend once for large non-streaming requests on the chat to Responses bridge (#44508)
* fix(logging): claim the async success dedup flag before the first await

Nested @client wrappers on the chat -> Responses bridge schedule two async
success tasks on one logging object. The dedup flag was checked at the top
of _async_success_handler_body but only set after several awaits; the
worker-thread base64 offload for messages >= 256 KiB yields there, so the
second task passed the same check and the request was logged twice
(two SpendLogs rows, key/team/daily spend counted twice).

Fixes #44500

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(logging): serialise async success handlers per logging object instead of claiming early

Claiming the dedup flag before the first await (previous commit) let a first
handler that is cancelled before its callbacks run (e.g. by the logging
worker's per-coroutine timeout) suppress the second one, losing the request's
success accounting. Serialise non-streaming handlers on the same Logging
object with an asyncio.Lock kept in a module-level WeakKeyDictionary: the
second handler waits, then skips if the first logged, or logs itself if the
first raised or was cancelled first. Streaming chunk logging is not locked.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(integration): pin one spend row and one key charge for large bridged chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): wait for the large request's spend row before asserting

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): use fixed timestamps in async success dedup tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(logging): trim the async success dedup lock comment

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): reset background interaction dedupe under the success dedup lock

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): use a fixed start time in the interactions logging helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: claim the async success log at schedule time instead of locking the handler

Nested @client wrappers (chat over the Responses bridge, Anthropic Messages over the chat adapter) exit with one shared Logging object and each queued their own async success handler. The handler checked has_logged_async_success before its awaited work and set it after, so both could pass the check and the request was logged and billed twice.

_schedule_async_success_logging now claims the log on the Logging object synchronously through claim_async_success_log, and a later wrapper returns without enqueuing. There is no await between the check and the decision, so no lock is needed. The per-object asyncio.Lock, the WeakKeyDictionary that held it and the locked block in async_log_background_interaction_completion are gone.

Tests: the regression test moves to tests/unit/test_utils.py and drives two _dispatch_success_logging exits with one Logging object through the logging worker, asserting one log with the inner result. It fails without the claim. The cancelled-handler test is dropped since a second wrapper no longer retries. The background interaction completion test stays.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(logging): pin the background completion event order and name the innermost wrapper in the claim docstrings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): send one bridged request and assert the inner Responses result is the logged one

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 21:23:19 +00:00
moe-berri
7f5499a751
perf(lens): reduce rendering on live investigations (#44935)
* perf(lens): avoid redundant investigation rendering

* fix(lens): keep history polling after run fetch failures
2026-10-06 14:19:52 -07:00
devin-ai-integration[bot]
acd2a7ccc5
fix(lens): use async-timeout on python 3.10 for budget reservation timeouts (#44911)
* fix(lens): use async-timeout on python 3.10 for budget reservation timeouts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(deps): keep uv.lock diff to the async-timeout entry

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(lens): cover real request deadline expiry in reserved_budget

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(lens): finish Python 3.10 timeout coverage and dependency checks

* test(lens): control event-loop time for deadline regressions

* test(tracing): include priced call count in trace fixture

* fix(lens): limit timeout compatibility changes to PR scope

---------

Co-authored-by: Moe Khalil <moe@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 14:19:37 -07:00
moe-berri
a3e74a5223
fix(tracing): preserve optional provider evidence in fixture replay (#44930)
* fix(tracing): preserve optional provider evidence in fixture replay

* fix(tracing): keep copied provider identities consistent
2026-10-06 14:00:40 -07:00
moe-berri
fe29c55d11
fix(lens): preserve chronological order when grouping trace steps (#44895)
* fix(lens): preserve chronological order when grouping trace steps

* test(lens): cover independent paging of repeated trace groups
2026-10-06 13:51:58 -07:00
moe-berri
5086fb3038
fix(tracing): retain native logs and headless tool results (#44894)
* fix(tracing): retain native logs and headless tool results

* fix(tracing): bound uncorrelated log annotations after decoding

* fix(tracing): canonicalize absent native log context
2026-10-06 13:51:38 -07:00
moe-berri
8f041a8f06
fix(lens): copy supported trace and span content requests (#44902) 2026-10-06 13:51:22 -07:00
tin-berri
2cecc7e33f
fix(auto-router): separate tuning and select chained heuristic (#44928) 2026-10-06 13:43:53 -07:00
devin-ai-integration[bot]
abd3af422d
fix(caching): never store or serve a chat completion with no choices (#44709)
* fix(caching): never store or serve a chat completion with no choices

A provider response with empty choices was written to the response cache and served on every identical request until the TTL ended, with no provider call in between. The cache now skips storing such a response and treats an already stored one as a miss, so the next request goes back to the provider and its answer replaces the entry.

* fix(caching): skip responses with no output on the Responses API and Anthropic Messages too

* fix(caching): skip streams and stored entries that carry no output

A chat or text completion stream whose chunks carried no choice is closed
by the stream wrapper with one empty choice of its own, so the assembled
response passed the choices check and was cached. The assembled stream is
now judged on its content: a stream with no text, tool call, or other
output in any choice is never stored, on the async and sync writers alike.
The Responses API stream writer and the Anthropic Messages stream writer
apply the same no-output check before storing.

A stored entry with no output read through the worker memory tier is now
evicted from that tier on the miss, so the next read reaches Redis where
the refill lands; the text completion and messages writers only write to
Redis, and the memory copy otherwise kept missing until its own TTL.

* test(integration): response cache cells for answers without output

Deterministic cells for the response cache on every unified endpoint,
streamed and not, through the OpenAI and Anthropic SDKs and raw httpx,
plus the sync SDK paths, stale entries, malformed answers, per-request
TTLs, cache delete, and chaos (Redis stopped or paused mid burst, a
worker killed, in-memory cache mode). The scripted upstream counts only
POSTs as deployment calls, since the proxy's boot-time GET /v1/models
discovery of a config deployment is not one.

* test(caching): pin the stored entry timestamp in the worker-copy test

* test(integration): drop the restating comments from the chaos cells

* test(integration): close the breaker on the first call after the Redis restart

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-06 20:07:30 +00:00
ishaan-berri
5ed7ec8511
fix(lens): price agent traces by joining gen_ai.response.id to spend logs (#44738)
* fix(lens): price agent traces by joining gen_ai.response.id to spend logs

Trace spend now joins each model call to spend_logs on one key: the
span's response id (gen_ai.response.id or the id the normalizers read
from OpenInference/LangChain output) against spend_logs.response_id or
the upstream id embedded in a managed resp_ id. The litellm.call_id and
traceparent transport join paths and the per-row ownership gate are
removed; the spend SQL still restricts rows to what the reader can see.

A run with some unpriced calls now reports the sum of its priced calls
plus priced_calls, instead of an unknown total.

* test(lens): cover response id spend join and partial trace totals

* chore(lens): regenerate trace types for priced_calls

* feat(lens): show partial run cost as a lower bound with priced call count

* fix(lens): treat litellm.call_id as the same assigned call id for spend joins

The id LiteLLM assigned to a call is either the response id it returned
(gen_ai.response.id -> spend_logs.response_id) or its gateway call id
(litellm.call_id -> spend_logs.litellm_call_id). Both are exact ids the
gateway mints and logs, so the join stays one rule. Transport span
matching and the ownership gate stay removed.

* test(lens): cover litellm.call_id spend joins and restore captured totals

* feat(lens): link each priced model call to its spend log

Spans gain spend_log_request_id, the spend_logs.request_id the call was
priced from, and spend_match, which says whether a model call matched or
why not (no assigned id on the span, no spend log with that id, or an
ambiguous match). A model call span is priced from the same ids as the
run total, so its cost and the total agree.

* test(lens): cover spend log links on model call spans

* chore(lens): regenerate trace types for spend log links

* feat(lens): open the matched spend log from an LLM step

An LLM step's header now shows a Spend log chip with the matched
request id and cost; clicking it opens the request log drawer over the
run, fetched by the exact spend_logs.request_id instead of the span's
own response id. Unpriced steps say why (no assigned id on the span, or
no spend log with it). Tree rows show each model call's cost, and a
partial run cost shows its priced call count inline.

* test(lens): cover the spend log link and unmatched cost reasons

* feat(lens): show the spend log link as a bordered LiteLLM Spend Log button

* feat(lens): add a back link from the spend log drawer to the agent trace

* feat(lens): label the spend log back link Back to Lens trace with the Lens icon

* fix(lens): ignore assigned ids that name no spend log when pricing a call

An id that names no row no longer vetoes the call, so a span carrying
both a response id and a call id still prices from a spend row logged
before litellm_call_id existed. An id naming two or more rows makes the
call ambiguous, and the match reason comes from the same per-id result,
so a single matched row with no cost is reported as matched.

* fix(lens): price a trace only from spend logs in its own team

A reader with several teams could see the same assigned id in another
team's spend log; only rows from the trace's team now price it. The
user and key ownership gate stays removed.

* perf(lens): resolve each model call's spend once per trace

Model call matches are computed once when the trace is resolved and
looked up by span index, instead of scanning the model call list for
every span and walking the graph again for spans, agents and the run
total.

* fix(lens): hide a step's Cost fact only when its spend log link shows the cost

* chore(lens): drop narrative doc comments from the spend join

* fix(lens): price a model call only when its ids agree on one spend log per span

* fix(lens): keep pricing spend logs written before litellm_call_id by their request id

* fix(lens): price every attempt a model call's ids name when they agree
2026-10-06 19:55:26 +00:00
devin-ai-integration[bot]
282d733fb6
feat(auth): deny search tools by default when search_tool_deny_by_default is set (#44490)
Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 12:50:26 -07:00
devin-ai-integration[bot]
abc543e701
feat(proxy): add opt-in vector_store_deny_by_default for least-privilege vector store access (#44244)
* feat(proxy): add opt-in vector_store_deny_by_default for standalone virtual keys

Adds general_settings.vector_store_deny_by_default (typed bool, default false). When enabled, a virtual key
with no team must list the requested vector store in its object_permission.vector_stores; no permission
record, null, or an empty list is denied with key_vector_store_access_denied. Omitted or false keeps the
existing behavior, including nonempty allowlist enforcement. The master key is unchanged in both modes.
Team keys and keyless callers are deferred to later increments of LIT-6035

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): require key and team vector store grants for team keys under vector_store_deny_by_default

With the flag enabled, a virtual key on a team needs both its own grant and its team's grant for every requested vector store. A missing permission record, an empty list or an unresolved team grants nothing. Dashboard session keys and the master key keep their existing behavior, and flag-off behavior is unchanged

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat(proxy): require user or team vector store grants for keyless requests under vector_store_deny_by_default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(proxy): cover vector_store_deny_by_default through a real proxy with key, team and JWT identities

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): regenerate schema.d.ts for vector_store_deny_by_default

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(auth): cover path and file_search vector store ids under vector_store_deny_by_default

Strict mode now reads vector_store_ids and tools[].vector_store_ids from the request body without needing a vector store registry, so /v1/vector_stores/{id}/search and Responses file_search are checked. User grants load through the object permission cache, and the proxy admin user rebuild keeps object_permission_id

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): broadcast user entitlement cache eviction to every worker

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(auth): reuse VectorStoreRegistry id extraction for vector_store_deny_by_default

Strict mode now uses get_vector_store_ids_to_run on an empty registry when none is loaded, instead of a parallel set of request-shape helpers, and vector_store_access_check documents the policy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): add strict vector store audit cells for routes, SDKs, workers and concurrency

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(integration): consolidate vector_store_deny_by_default coverage to core cases

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(proxy): reject invalid vector_store_deny_by_default at config load and return 400 for malformed vector_store_ids

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(auth): read vector store key and team grants through the object permission cache

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(auth): return typed permission rows in request flow vector store tests

Co-authored-by: mrinal <mrinal@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(auth): validate vector store ids and tools as immutable sequences

Co-authored-by: mrinal <mrinal@berri.ai>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 12:50:25 -07:00
devin-ai-integration[bot]
26a9b02f7b
feat(otel): let team and key Arize callbacks choose the OTLP transport (#44492)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 14:46:50 -05:00
devin-ai-integration[bot]
b6d587d23e
refactor(rust_bridge): remove rule-gated native secret-manager selection (#44906)
* refactor(rust_bridge): remove rule-gated native secret-manager selection

Mirror the cache treatment: SecretManagerRule/SecretManagerContext and the resolve_native_* binding plumbing are gone. The Rust bridge now selects the native backend from the explicitly configured client (capture_secret_manager / _SecretManagerRuntime.from_client), and get_secret_from_manager is the plain Python handler path.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(deps): bump sharp to 0.35.5 for GHSA-wq5f-xc86-pv6w

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 19:14:55 +00:00
moe-berri
fb554d34f3
fix(lens): paginate visible conversation entries (#44878)
* fix(lens): paginate visible conversation entries

* fix(lens): stop pagination at exhausted subagents

* refactor(lens): derive pending branches without mutation

* fix(lens): bound pending ancestry work for deep traces
2026-10-06 12:10:19 -07:00
moe-berri
5147aefa8b
fix(lens): reuse trace reviews and consolidate findings across runs (#44778)
* fix(lens): reuse trace reviews and consolidate findings across runs

* chore: sync schema.prisma copies from root

* fix(lens): preserve partial reviews and bound budget admission

* fix(lens): resolve CI regressions and clarify reused runs

* test(lens): exercise review reuse in worker container smoke

* fix(lens): show actual reviews when a run finishes

* fix(lens): distinguish reuse plans and stop blocked scans

* fix(lens): retain partial findings when model requests stop

* docs(lens): clarify budget edits during active runs

* fix(lens): publish findings only after reconciliation completes

* fix(lens): guide insufficient-budget runs to budget settings

* fix(lens): renew budget holds and bound admission waits

* fix(lens): preserve provider errors during budget cleanup

* fix(ci): update PgBouncer and sharp for current builds

* fix(lens): serialize settlement and fence review checkpoints

* fix(lens): defer generated Prisma client type import

* test(lens): verify checkpoint and progress rollback together

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-10-06 12:09:56 -07:00
berriai-litellm-provider-info-sync[bot]
fa3b0a6d95
feat(vertex-ai): add gemini-nano-banana-2.1 and Nano Banana 2 priority pricing (#44892)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 11:52:07 -07:00
devin-ai-integration[bot]
de984d6ebc
fix(model_prices): correct supported_endpoints on gemini and vertex image rows (#44891)
* fix(model_prices): correct supported_endpoints on gemini and vertex image rows

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(model_prices): drop images/generations from gemini/nano-banana-pro-preview

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 11:26:27 -07:00
devin-ai-integration[bot]
5cdebded90
fix(ci): skip the generated dashboard bundle in the master key guard (#44890)
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 18:15:09 +00:00
berriai-litellm-provider-info-sync[bot]
d9f8dbe23a
chore(prices): add gemini/gemini-nano-banana-2.1 from the Gemini pricing page (#44868)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 18:06:35 +00:00
devin-ai-integration[bot]
837c6a7481
fix(security): remove the publicly known master key from the repo (#44718)
* fix(security): hash the publicly known master key

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: replace weak master key examples and regenerate artifacts

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: replace weak key fixtures with generated test keys

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: generate master keys for proxy startup

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve lens dev key entropy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: restore proxy key compatibility in scrub examples

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: scrub merged SSO fixture and refresh dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ci): stabilize test keys and metadata collection

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: drop the rebuilt dashboard bundle

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mateo <mateo@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:55:24 -07:00
berriai-litellm-provider-info-sync[bot]
ab61410a39
chore(pricing): add azure model-router and whisper rows (#44872)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 10:54:50 -07:00
devin-ai-integration[bot]
e2c55daaaf
fix(azure): add model router flat fee to azure provider cost tracking (#44876)
* fix(azure): add model router flat fee to azure provider cost tracking

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(azure): price the azure router fee from one canonical entry

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(azure): reuse the azure_ai model router fee path for azure

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(azure): drop model router fee unit tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:47:26 -07:00
devin-ai-integration[bot]
096b20b6db
chore(model_prices): remove malformed, duplicate and decommissioned palm cost map entries (#44880)
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 17:33:54 +00:00
devin-ai-integration[bot]
090a4c3f24
fix(ui): read MCP submission rules from bare-array /config/list response (#44648)
* fix(ui): read MCP submission rules from bare-array /config/list response

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(mcp): cover submission rules contract and dashboard preload

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): disable MCP submission rules editor until rules load

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format MCP submission integration test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): restore prior MCP submission rules after the rules spec

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(e2e): move MCP submission rules setup and cleanup into fixtures

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(ui): use Promise.withResolvers in MCP rules loading test

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): clear lint warnings in MCPSubmissionsTab

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): drop redundant JSX comments in MCPSubmissionsTab

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: mrinal <mrinal@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 10:09:46 -07:00
devin-ai-integration[bot]
878ba39e7c
refactor(rust): extract inference-testing crate (#44873)
Move the shared test helpers out of litellm-inference's test-support
feature into a publish = false litellm-inference-testing crate used only
as a dev dependency by the format crates.

Also drop the dead src/constants.rs (OPENAI_DEFAULT_API_BASE had no
users) and declare the litellm-http/litellm-llms test-support features
on the crates that actually use them instead of relying on feature
unification through litellm-inference's dev-dependencies.

Co-authored-by: Yujong Lee <yujong@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 16:33:06 +00:00
devin-ai-integration[bot]
eb385cca3e
fix(ui): restore key activity search to the top of the tab and add model activity search (#44521)
* fix(ui): restore key activity search and add model activity search

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore(ui): drop rebuilt dashboard bundle from the usage search change

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* style(ui): format usage search tests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): only show model no-match when the range has models

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Mubashir Osmani <mubashir@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-06 09:27:22 -07:00
berriai-litellm-provider-info-sync[bot]
bd23e6fc3d
fix(cost-map): update together_ai Kimi-K3 and Qwen3.8-Flash prices to published rates (#44864)
Price-Sync: litellm-providers

Co-authored-by: berriai-litellm-provider-info-sync[bot] <328147090+berriai-litellm-provider-info-sync[bot]@users.noreply.github.com>
2026-10-06 09:03:40 -07:00