Commit graph

4348 commits

Author SHA1 Message Date
mubashir1osmani
77fef603b5 style(ui): size chat suggestions to match the composer width
Use a full-width 3-column grid above the textarea with slightly larger
chip padding and type so prompts stay readable
2026-08-06 16:03:46 -07:00
mubashir1osmani
1d6e42b425 style(ui): tighten playground chat suggestion chips into a 3-col grid
Keep the three prompts on one row with a narrower max width and smaller
padding/type so they sit cleanly above the composer
2026-08-06 16:02:53 -07:00
mubashir1osmani
50c66ab99d test(ui): cover interactions MCP tool block serialization
Assert the playground interactions client forwards buildMcpToolBlocks
when servers are selected and omits tools otherwise
2026-08-06 15:51:28 -07:00
mubashir1osmani
e99ef9d5dd fix(ui): make playground MCP work on all tool-capable endpoints
Route chat, responses, anthropic, and interactions through the shared
buildMcpToolBlocks helper so MCP server_url/label resolution is correct
for any model. Chat was labeling every server as "litellm" and preferring
aliases; responses used absolute /mcp URLs that broke the all-servers path
2026-08-06 15:51:28 -07:00
mubashir1osmani
609f7beddf fix(ui): show MCP servers in playground multi-select
Anchor MultiSelect popup to chips, forward refs on ComboboxChips, harden
MCP list response parsing, and load servers with the active key so
existing MCP servers appear in the dropdown
2026-08-06 15:51:28 -07:00
mubashir1osmani
ba339cac0c fix(ui): restore minimal realtime playground client
Drop the custom cancel/queue/dual-buffer state machine and the duplicate
OPEN_AI_REALTIME_VOICES list. Reuse OPEN_AI_VOICE_SELECT_OPTIONS and the
original thin WebSocket event handler (append deltas, response.done
fallback) with the shared ChatComposer shell
2026-08-06 15:51:28 -07:00
mubashir1osmani
ca6de79a8d fix(ui): stop realtime responses from stalling or splitting bubbles
Text turns use text-only modalities with VAD off, and voice re-enables
audio+VAD only while recording so residual VAD cannot race a second
response. Stream text and transcript into one draft, always update the
last assistant bubble (even past status lines), and unlock with a cancel
timeout so the UI cannot stick mid-response forever
2026-08-06 15:42:04 -07:00
mubashir1osmani
d4ac065b26 fix(ui): queue realtime pending messages and drop new comments
Keep every text submission while a response is cancelled, not only the
latest. Remove newly added source comments called out in review
2026-08-06 15:36:03 -07:00
mubashir1osmani
0fe2b137a2 fix(ui): prevent concurrent realtime responses and garbled transcripts
Track in-progress responses so text/mic turns cannot race response.create,
prefer a single transcript stream per turn, and queue sends until the active
response finishes or is cancelled
2026-08-06 15:35:27 -07:00
mubashir1osmani
6f2555441c feat(ui): restyle realtime playground with shared chat composer
Move RealtimePlayground off Ant Design onto shadcn controls and the
shared ChatComposer, use the realtime-safe voice list, and reuse the
same composer for Compare message input
2026-08-06 15:35:27 -07:00
mubashir1osmani
48de145fd5 fix(ui): surface model load failures and align request key with models
Throw when both model endpoints fail so ChatUI can show an error instead of
an empty list. Use the same resolved key for model listing and chat requests,
and drop the fetch_models JSDoc
2026-08-06 15:35:24 -07:00
mubashir1osmani
34c18ca2b2 fix(ui): show human labels on playground SDK type select
Match the SelectValue children pattern so OpenAI SDK / Azure SDK
render instead of the raw openai/azure values
2026-08-06 15:35:03 -07:00
mubashir1osmani
7ff6f26b93 fix(ui): show human labels on playground virtual key source select
Base UI SelectValue renders the raw value unless children supply the
label, so map session/custom to Current UI Session and Virtual Key
2026-08-06 15:35:03 -07:00
mubashir1osmani
65203b2920 fix(ui): load playground models for virtual keys via /v1/models
Prefer the key-scoped OpenAI models list when the Virtual Key source is
selected, enrich with mode from model_group/info, and debounce custom key
input so models appear for the key's access set
2026-08-06 15:35:02 -07:00
mubashir1osmani
3018fb7dc1 fix(ui): size chat composer textarea with CSS field-sizing
Drop direct el.style.height mutation in favor of field-sizing:content
2026-08-06 15:34:42 -07:00
mubashir1osmani
d050060cd4 style(ui): strengthen playground chat composer border and shadow
Make the shared chat input stand out with a fuller border, layered
shadow, and a slightly stronger focus ring
2026-08-06 15:34:22 -07:00
mubashir1osmani
74b400d5be feat(ui): adopt vercel-style chat composer for playground
Replace the compact single-line input with a PromptInput-style composer:
taller auto-growing textarea, rounded card shell, footer tools, and
stop button while a request is in flight
2026-08-06 15:34:22 -07:00
mubashir1osmani
d635f44683 fix(ui): exclude unknown model modes from playground endpoint filters
Modes outside ModelMode (batch, rerank, ocr, etc.) must not collapse to
chat-compatible, or conversational endpoints surface unusable models
2026-08-06 15:34:13 -07:00
mubashir1osmani
4b2781ccb3 fix(ui): restore playground model filtering by endpoint
Bring back the prior Chat model dropdown filter (including chat models
on responses/anthropic/interactions and image models on image_edits), and
map mode realtime so the realtime endpoint only lists compatible models
2026-08-06 15:33:58 -07:00
mubashir1osmani
5102b9c0d8 fix(ui): ignore stale playground model loads on key switch
Cancel in-flight model fetches when the key or source changes so an older
response cannot overwrite modelInfo. Drop the inverted endpoint-filter
assertion; filtering coverage lands in the next stack PR
2026-08-06 15:33:30 -07:00
mubashir1osmani
bb5d9a199a feat(ui): migrate ChatUI off Ant Design and Tremor
Replace ChatUI cards, inputs, dialogs, popovers, MCP selects, uploads,
tooltips, and icons with shadcn/Base UI and Lucide. Update ChatUI tests
to drive searchable combobox controls instead of Ant Design selectors
2026-08-06 14:33:39 -07:00
mubashir1osmani
ea654d11a5 feat(ui): migrate playground tabs and message icons to shadcn
Replace Tremor playground page tabs with Base UI tabs and swap remaining
message bubble and attachment renderer icons to Lucide
2026-08-06 14:26:15 -07:00
mubashir1osmani
f5d98c0b8c feat(ui): migrate playground chat controls toward shadcn
Continue the Playground Chat Ant Design/Tremor migration: shared MultiSelect,
upload validation with semantic file inputs, collapsible message widgets, and
AdditionalModelSettings on Base UI controls
2026-08-06 14:25:11 -07:00
Mateo Wang
0acca3e86a
Merge pull request #24548 from mpcusack-altos/fix/bedrock-batch-credential-fields
fix(router): include Bedrock batch/S3 fields and model in deployment credentials
2026-08-05 22:39:39 -07:00
tin-berri
34fc8d2ee7
fix: expired-miss share over all measured turns + cost-optimization tab labels (#36037)
* fix(ui): make the expired-miss stat row a focusable tooltip trigger

* fix: auto-router expired-miss percentage and cost-optimization tab labels

- change expired-miss percentage denominator from return-to-tier misses to
  all measured turns (same_model + first_visit + return_to_tier). when
  auto-routers flip tiers rapidly within TTL, return-to-tier turns become
  hits and disappear from the miss count; the old metric reported only the
  rare failure population. the new metric contextualizes that population as
  a share of overall coverage
- rename usage tab from 'Usage' to 'Overall'
- rename auto-router-usage tab from 'Auto-Router Usage' to 'Auto-Router'
- update component and unit tests to match new semantics
2026-08-05 22:34:55 -07:00
tin-berri
86890654c5
fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day (#36051)
* fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day

* fix(proxy): gate the current-UTC-day extension behind an opt-in param sent by the cost optimization dashboard

* fix(ui): label cost optimization savings dates as UTC days
2026-08-05 22:33:54 -07:00
Michael Cusack
3d275d97fe fix(router): return model and Bedrock batch fields in deployment credentials
get_deployment_credentials_with_provider dropped s3_region_name,
s3_encryption_key_id, and aws_batch_role_arn because
CredentialLiteLLMParams never declared them, and it never returned the
deployment's model, so proxy batch creation against Bedrock failed with
"LiteLLM doesn't support custom_llm_provider=bedrock for 'create_batch'"
or "AWS IAM role ARN is required" (#25104)

Provider-only file and batch calls keep their no-model contract:
get_team_provider_credentials strips the model key so a provider-scoped
request is not pinned to an arbitrary matching deployment
2026-08-05 21:49:31 -07:00
Mateo Wang
ba91768146
Merge pull request #35925 from BerriAI/litellm_tier_aware_reasoning_token_cost
fix(cost): bill reasoning tokens at the service tier output rate
2026-08-05 21:09:55 -07:00
tin-berri
7c621b3141
fix(auto-router): accept every reminder marker pair a harness emits (#36029)
* fix(auto-router): accept every reminder marker pair a harness emits

reminder_markers held one (open, close) pair, so a harness that wraps
injected context differently per agent type only got the slice of traffic
using the configured envelope stripped. Every other agent type kept hitting
the original bug: its reminder-only turn never stripped to empty, won
"newest human ask", and the harness blob got classified in place of the
real question, choosing the tier and therefore the spend.

The field now takes a list of ReminderMarkerPair, following the
KeywordTierRule pattern already in this file so each pair validates itself
and errors point at reminder_markers.N.close rather than a bare index.

Blocks from different pairs can nest, which the gap construction could not
handle: resuming the kept text at an inner block's end walks back inside
the enclosing block and leaks its remainder. Running the block ends through
a maximum collapses nested and overlapping spans without a separate merge
pass, and stays linear in block count, which a fold over a growing tuple
of merged spans would not.

A single pair's ends already increase, so the maximum is the identity and
the default path is byte-identical: verified against the shipped function
over 200k generated inputs, and every existing reminder test passes
unchanged. The prior single-pair config shape is rejected loudly at
startup and at /model/new rather than silently stripping nothing.

* docs(auto-router): document reminder_markers in the complexity router README

* chore(ui): regenerate dashboard API types for the reminder_markers shape

---------

Co-authored-by: Abhimanyu Kapur <38531241+akapur99@users.noreply.github.com>
2026-08-05 21:03:36 -07:00
mateo-berri
1b30b1bc20 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_tier_aware_reasoning_token_cost
# Conflicts:
#	litellm/types/utils.py
2026-08-05 20:45:34 -07:00
tin-berri
ece652f6a7
feat(ui): add the auto-router usage tab to cost optimization (#35995) 2026-08-06 02:15:35 +00:00
Yuneng Jiang
6e434c926b
Merge branch 'litellm_internal_staging' into litellm_dead_locals_5_8 2026-08-05 18:07:23 -07:00
Yuneng Jiang
ce5c4c1bf9
refactor(ui): drop dead locals and unused React state across the dashboard
Removes declarations nothing reads, along with the writes that fed them, so
the remaining code says what it actually does.

Where a declaration was dead but its initializer had a real effect, the call
survives and only the binding goes: spies stay installed, renders still run,
and every awaited request keeps its await. Pure computations are deleted
whole rather than left as statements that build a value and throw it away.

Dead useState pairs are removed outright instead of being elided to
const [, setX], which would keep a hook and every write to a value nothing
reads. Three chains turned out to be dead end to end and are removed with
their fetches: the tool detail team list, the Teams MCP access group load,
and the user dashboard proxy settings load.

ColumnMeta's declaration merging in columnMeta.ts and view_logs/table.tsx is
a false positive; TypeScript requires those type parameters to match the
upstream signature exactly, so both get a scoped suppression instead.
2026-08-05 18:07:17 -07:00
yuneng-jiang
624fa11d71
Merge pull request #36025 from BerriAI/litellm_dead_locals_3_tests_and_destructures
refactor(ui): drop unreferenced locals from tests and narrow destructures
2026-08-05 17:54:12 -07:00
Yuneng Jiang
888f911133
refactor(ui): drop unreferenced locals from tests and narrow destructures
Third and fourth slices of the sweep, combined because they raise nearly the
same question and neither changes what runs.

Nine test files plus one source file lose symbols whose only mention was
their own declaration. Ten more narrow a destructure to the keys actually
read, so `const { accessToken, userRole, userId: userID, premiumUser } =
useAuthorized()` keeps only `accessToken`. Aliases are preserved as written.

ignoreRestSiblings stays on so the omit idiom `const { tags, ...rest } =
metadata` is left alone; dropping `tags` there would fold it back into rest.

ToolDetail is held back again. Its unread binding only looks like a plain
deletion on the first pass, because the dead useMemo still reads it; one more
pass exposes a useQuery that issues a real request. That belongs with the
slices that get QA'd.

Part of LIT-5162.
2026-08-05 17:46:39 -07:00
yuneng-jiang
c2c795fad5
Merge pull request #35821 from BerriAI/litellm_dead_locals_2_components
refactor(ui): drop unreferenced locals from shared dashboard components
2026-08-05 17:44:38 -07:00
ryan-crabbe-berri
f2690aa60e
fix(ui): opening a project now pushes ?project= so back and deep links work (#36001)
* fix(ui): drive project detail selection from the ?project= url param

Opening a project kept selectedProjectId in useState, so the URL never changed; the detail view could not be linked or reloaded and browser Back skipped past the Projects page entirely.

Selection now lives in the ?project= query param via nuqs with history: push, matching how Teams, Organizations and Virtual Keys already work.

* fix(ui): project detail close replaces history to match the other detail pages

Adopts the close semantics from PR #36013 so browser Back after an
in-page close leaves the Projects page instead of reopening the
dismissed detail; the close test now pins the replace mode
2026-08-05 17:31:22 -07:00
Yuneng Jiang
80627c0477
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_dead_locals_2_components
# Conflicts:
#	ui/litellm-dashboard/src/components/add_model/add_auto_router_tab.tsx
2026-08-05 17:29:54 -07:00
yuneng-jiang
1b059f472d
Merge pull request #35819 from BerriAI/litellm_dead_locals_1_app_routes
refactor(ui): drop unreferenced locals from dashboard route components
2026-08-05 17:28:15 -07:00
ryan-crabbe-berri
1dab133d24
fix(ui): link project page keys to their virtual key detail (#36002) 2026-08-05 17:26:13 -07:00
ryan-crabbe-berri
6a2e4e6c36
fix(ui): sync projects list page index to ?page= so back and reload keep the page (#36003)
* fix(ui): sync projects list page index to ?page= so back and reload keep the page

Paging the Projects list only moved TanStack's internal page index, so the URL
never changed: reload dropped you on page 1, browser Back left the page entirely,
and the page could not be shared.

The page index now comes from a nuqs ?page= query state with history: "push".
Pagination stays controlled off that value and the footer writes the URL
directly, because TanStack resets its page index whenever the data array
identity changes; letting it own the state would clear a deep-linked page as
soon as the projects query resolved. A page outside the current row set falls
back to page 1, which covers both a hand-typed ?page=99 and a search that
narrows the list below the current page.

* fix(ui): carry page_size in the url so restored history entries show the same rows

Greptile flagged that a history entry restoring ?page=N under a changed
local page size displays different projects than it originally showed.
Page size now rides the same query string via useQueryStates, size
changes reset the page inside a single history entry, and values outside
the offered options fall back to the default
2026-08-05 17:25:42 -07:00
tin-berri
55e666a05f
feat(complexity_router): report LLM classifier cost per request via routing_decision and x-litellm-classifier-cost header (#36015) 2026-08-05 16:27:32 -07:00
tin-berri
265945dfcd
feat(ui): match auto-router preset models against deployments' underlying model IDs (#35972)
The bundled presets only became selectable when an admin's public model_group
names matched the preset's hardcoded model names. model_name is admin-arbitrary,
so renamed deployments (my-claude-fast, bedrock-opus) left both presets greyed
out. Resolve preset models against each deployment's litellm_params.model and
model_info.base_model from /v2/model/info via a normalized ID join, and prefill
the admin's registered group names. Resolves LIT-5225
2026-08-05 15:24:54 -07:00
Abhimanyu Kapur
c76882b51b
fix(auto-router): stop the embedding model's context window from failing long requests (#35956)
* fix(auto-router): stop the embedding model's context window from failing long requests

The auto-router embeds the last user message to pick a model and sent it to the
embedding model unbounded. Embedding models carry 512 to 8k token windows while the
chat models they route to carry 200k+, so any prompt over the encoder's window failed
at the routing step with a 400 the destination model would never have raised.

Cut every doc to a character cap inside LiteLLMRouterEncoder, which is the one choke
point the auto-router, complexity-router, semantic guard and MCP tool filter all share.
Default 2000 chars, roughly 500 tokens, which fits even a 512-token self-hosted encoder,
overridable per deployment with auto_router_max_input_chars and globally with
DEFAULT_MAX_EMBEDDING_INPUT_CHARS.

Truncation alone cannot cover provider-side batch and byte limits, so any failure of
the route call now falls back to the auto-router's default model instead of propagating.
That path also fixes two latent bugs: a no-match left the auto-router alias in place as
the model name, which fails downstream with "Unmapped LLM provider" rather than reaching
default_model, and an empty route list raised IndexError.

Fixes #17869
Fixes #20277

* fix(auto-router): make the embedding input cap opt-in so guards still see whole prompts

Defaulting the cap inside the shared encoder truncated every consumer, not just the
auto-router. The semantic guard builds the same encoder, so its pre-call check would
have classified only the first 2000 characters while the full message still reached the
model, which a benign opener in front of an injection payload walks straight past. The
MCP tool filter and complexity router were silently narrowed the same way.

The encoder now defaults to sending docs whole and cuts only when a caller passes
max_input_chars. The auto-router is the only caller that does, so guard, MCP filter and
complexity-router behaviour is unchanged from before this branch.

DEFAULT_MAX_EMBEDDING_INPUT_CHARS becomes DEFAULT_AUTO_ROUTER_MAX_INPUT_CHARS, since it
is now specific to the auto-router, and drops its env override: the per-deployment
auto_router_max_input_chars already covers it, and every env var in constants.py has to
be documented, which is what broke the documentation and code-quality checks.

Also drops the added comments and the redundant type: ignore that review flagged.

* test(auto-router): cover the max_input_chars wiring from litellm_params

Nothing asserted that auto_router_max_input_chars on the deployment reaches the
AutoRouter that embeds prompts. Dropping the wiring left every test green while the cap
silently reverted to the default, so an operator with a 512-token embedding model could
not lower it and every long prompt would fall back to the default model instead of
being routed.

* test(auto-router): cover the populated route-choice list branch

The route layer can hand back a list, and picking its first element is where the
IndexError lived: the empty case was covered but the populated one was not, so the
branch that reads route_choice[0].name could be deleted with every test still green.
2026-08-05 14:47:40 -07:00
ryan-crabbe-berri
2dc49a913c
refactor(ui): replace hand-rolled query-param routing with nuqs (#35871)
* refactor(ui): replace hand-rolled query-param routing with nuqs

The dashboard carried five copies of the same pushState-based detail
routing hook plus a shared navigateWithParams helper, each with its own
plumbing test and a copy-pasted reactive useSearchParams mock in
component tests. nuqs provides the same shallow history-API routing
behind useQueryState/useQueryStates, so the key, team and org hooks are
deleted in favor of inline useQueryState at their single consumers,
while the models and logs hooks keep their interfaces but drop their
hand-rolled internals. Component tests now mount NuqsTestingAdapter
(via renderWithProviders or locally) instead of patching window.history,
and URL assertions go through onUrlUpdate spies that can additionally
distinguish push from replace, which the old window.location checks
could not

* test(ui): assert browser back closes the log drawer after in-drawer selection

Greptile flagged that the nuqs port of the switching-logs test stopped
at asserting emitted push and replace modes. The test now replays those
recorded modes against a history stack and performs the back step, so a
regression to push-on-select or broken URL-derived drawer state fails
the test instead of passing silently
2026-08-05 13:29:58 -07:00
tin-berri
32deaff015
feat(spend): rebuild the auto-router benchmarks backend as a per-session rollup (#35910)
Folds every successful auto-routed request into LiteLLM_AutoRouterSession with one
conditional upsert at spend-write time, classifying each turn (same model, first
visit, return to tier, out of order) against the row's own columns so nothing is
read before the write. The upsert's placeholders and argument tuple both derive
from the transaction dataclass's own field order, so the SQL and the call site
cannot drift apart. GET /auto_router/benchmarks aggregates the rollup, grouped
by the full (router, type) identity, and never scans LiteLLM_SpendLogs. A turn's
cache interaction is derived once from its usage record (savings.py owns the
extraction; compute_savings_spend derives cache reads from usage_object itself),
hits are counted order-independently so the overall hit rate matches its covered
denominator, caller-chosen session ids are bounded before entering the primary
key, and a poisoned statement drops only its own session's remaining turns.
Return misses inside the recorded TTL are named for what the telemetry shows
(within_ttl) rather than a presumed cause, since a provider can evict early.
Savings ride each router's derived baseline by default, so the response carries
no deployment-wide baseline label. Rollup retention has its own
maximum_autorouter_session_retention_period setting, pattern-identical to the
spend-logs knob and running in the same cleanup job on its own cutoff. Every
drain trigger sizes the queues through one owner and the enqueue honors
disable_spend_logs beside the tool-usage queue it mirrors.
2026-08-05 20:06:32 +00:00
Abhimanyu Kapur
b8df48cd7f
feat(auto-router): let operators replace the LLM classifier's system prompt (#35855)
* feat(auto-router): let operators replace the LLM classifier's system prompt

The complexity router's LLM classifier has always sent one built-in rubric, so the
router could only ever grade difficulty. Operators can now supply their own system
prompt, which replaces the rubric outright and repurposes the same tier machinery for
whatever taxonomy the prompt defines, data sensitivity being the obvious case.

Replacement is total: neither the rubric nor its closing line is appended, since both
describe grading difficulty over a "current message" and a prompt grading something
else is entitled to contradict them. That closing paragraph is also the classifier's
prompt-injection defense, so the config field and the dashboard editor both warn that
a replacement omitting it lets a caller ask for a tier and get it.

The heuristic fallback still scores complexity, which is meaningless for a repurposed
taxonomy, so classifier_fallback now chooses between the heuristic scorer and routing
straight to default_model. The default_model path bypasses tier pools, the adaptive
bandit, and escalation, because no tier was decided and the point of that fallback is
a known destination. It reports itself as default_model_fallback in the spend logs.

The dashboard's prompt editor prefills from a new
/auto_router/classifier/default_prompt endpoint rather than a copy of the rubric in
the frontend, and stores no override when the draft matches the default, so later
rubric improvements still reach every router that never customized it.

Tier names stay SIMPLE/MEDIUM/COMPLEX/REASONING; a custom prompt redefines what they
mean, not what they are called.

* fix(complexity-router): don't let the default_model classifier fallback bypass routing plugins

* fix(complexity-router): don't pin a session to the default model after a classifier failure

* fix(complexity-router): omit the tier from a default-model-fallback routing decision

The classifier never answered, so no tier was decided. The record reported the
tier whose pool happens to hold default_model, which reads in the spend log and
the UI as if the request was classified. Matches how default_fallback already
records a route that no tier produced.

* fix(proxy): allowlist /auto_router/ on the UI backend component

The new GET /auto_router/classifier/default_prompt is a UI-consumed management
route, so it belongs on the control plane. Without the prefix it was exposed by
neither component and test_gateway_plus_backend_covers_full_app failed.

* docs(ui): reword the classifier prompt disclaimer

Frames the closing paragraph as a strong recommendation rather than a
description of what gets dropped, names prompt injection explicitly, and
notes the tier names stay fixed regardless of their display names.

* fix(complexity-router): stop logging a fabricated tier on the plugin fallback path

The classifier-failed fallback resolves a tier so the routing-plugin pipeline has a
pool to filter, but nothing about the request produced that tier. The non-plugin
short-circuit already dropped it from the logged decision; the plugin path still
reported it, so a spend log claimed a classification the request never received.
Record the pool as a plugin-filtered-pool signal instead.

Also name the real problem when the resolved tier has no models at all: that raised
"No candidate models left after routing-plugin filtering" and sent operators hunting
for a policy plugin that never narrowed anything.
2026-08-05 19:48:11 +00:00
Yassin Kortam
09dd167b5a
feat(sgr): make the gateway middleware the source of truth for successful requests (#35717)
SGR has had two independent definitions. The admin UI derived it from
SpendLogs, so it counted what litellm's logging callbacks observed and could
attribute and price. BillableRequestMetricsMiddleware counted what the proxy
actually answered at the ASGI edge, but only exported to OTLP for enterprise
metering. The two disagree by design in places, and the SpendLogs figure goes
quiet whenever spend logging is disabled or the callbacks are bypassed.

This adds LiteLLM_DailyGatewayRequests, written by the middleware, and points
the dashboard's Successful Requests tile at it.

Requests fold into an in-memory map at record time rather than going through a
queue like the spend path. A count is a pure aggregate, and every dimension of
the key is chosen by the proxy from a closed set: the date, the category, and a
route that the classifier maps to one of a fixed list of strings rather than
passing the raw path through. Nothing a caller sends can add a key, so the fold
and the table are bounded by (days x categories x routes) however much traffic
arrives; the spend queue blocks once full, which is not acceptable in the
response path. A scheduler job drains it on the existing batch interval, and a
failed flush merges its counts back so a database blip undercounts nothing.

The middleware previously returned early when no billing recorder was
injected, which is the unlicensed case. The new sink is not license-gated, so
that early return now requires both sinks to be absent. The billing recorder
keeps its 2xx-only gate; the sink takes every status so failed_requests is
real. The sink is not told which deployment served the request, unlike the
billing recorder. That id is a sha256 over litellm_params, credentials
included, so a caller who puts a credential in the request body mints a fresh
one per distinct value. No configuration is needed for that: api_base and
base_url are on _BANNED_REQUEST_BODY_PARAMS and need allow_client_side_
credentials, but api_key is not on that list, and both reach the same
_handle_clientside_credential branch. The read endpoint aggregates the
dimension away regardless, so the key is better off without it.

The new table carries no key, user or team dimension, so /gateway/daily/activity
is restricted to proxy admin roles and the per-key and per-model breakdowns
keep reading the daily spend tables. The old path is left running and marked
with TODOs.

A fetched result carries the range key it was fetched for, and the render
selects it only when that key matches the range on screen. Both the gateway
counts and the spend aggregate go through that rule: the request tiles read the
first and fall through to the second, so stamping only one of them would leave
the tile showing a superseded range by the other route.

The paginated pages behind that aggregate are reached through a failure flag,
so the flag is stamped too. A flag left over from the previous range would let
those pages through while a new range is in flight, which is the same defect
one fallback further down.
2026-08-05 12:40:47 -07:00
Mateo Wang
332ec6c17a
Merge pull request #35926 from BerriAI/litellm_remove_types_ruff_exclusion
chore(lint): remove litellm/types from the ruff lint exclusion
2026-08-05 12:35:02 -07:00
tin-berri
d3d30353aa
refactor(ui): remove the three dashboard lint-budget violations added by #35893 (#35960)
PR #35929 zeroed the eslint budget headroom while #35893 added UI code in parallel, so staging went over budget by one complexity violation and two no-large-inline-object-arg violations, failing frontend-lint on every UI-touching PR until #35964 reverted the ratchet. This removes the three violations at the source so the budgets can ratchet back down: the submit-blocked-reason chain in add_auto_router_tab moves to a module-level helper, taking the component arrow from complexity 21 to 18, and the two four-property object literals in build_complexity_router_config.test.ts move into named variables. No behavior change; the touched suites pass (101 tests)
2026-08-05 11:51:24 -07:00