Commit graph

4353 commits

Author SHA1 Message Date
Abhimanyu Kapur
43265a5292 fix: three review findings — cap bypass, start race, request-path DB read
Zero-estimate cap bypass (cursor): a key quiet during the estimate
lookback gets cost_estimate 0.0, which the spend cap treated the same
as 'no estimate' and left uncapped — a later traffic spike on exactly
that job would bill until ends_at. Only a NULL estimate (rows predating
estimates) is uncapped now; $0 still gets the $1 floor.

Concurrent-start race (cursor): the find_first-then-create check passes
on both sides of a race, giving one key two active jobs and double
judge spend. A partial unique index (api_key_id WHERE status IN
(pending, running)) — raw SQL in the unshipped migration, since
schema.prisma cannot express partial indexes — makes the DB the
arbiter; the losing create surfaces as the same 409 as the advisory
check. The old find_many('desc') + reversed() insertion in the logger
cache already prefers the newest job for any legacy duplicates.

Request-path DB read (greptile): an expired job snapshot awaited
find_many inside the success callback, so every N seconds one request
per pod paid a synchronous Prisma read. The lookup is now sync-only:
it serves the current snapshot and kicks a detached refresh task when
stale. Cost: a cold pod's first ~1 refresh-window of samples are
missed (acceptable for a sampled eval); stale-if-error semantics keep
a DB blip from disabling the feature.

Also from review: a collapsed previous-job row said 'no verdicts' for
jobs with thousands of verdicts, because the list endpoint omits
results by design — it now says 'view results' when completed_count>0.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 17:02:49 -07:00
Abhimanyu Kapur
029be4d89e fix(ui): judge model field is a real selector, not a pre-filled default
The judge model box was a free-text input pre-filled with
anthropic/claude-sonnet-5, implying it was sent as-is when nothing was
actually defaulted client-side (the backend's own default only applies
if the field is omitted entirely). Replace it with a searchable combobox
backed by the model cost map (litellm_provider + mode === "chat"),
starting empty and requiring an explicit pick.

Recommends three models spanning different providers — anthropic/
claude-sonnet-5, openai/gpt-4o, gemini/gemini-2.5-pro — pinned to the
top of the list with a 'Recommended' badge, all backend-verified to
resolve correctly via litellm.get_llm_provider/cost_per_token.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 14:00:24 -07:00
Abhimanyu Kapur
d3020685cf Merge remote-tracking branch 'origin/litellm_internal_staging' into shadow-eval-pre-adoption 2026-08-08 13:17:18 -07:00
Abhimanyu Kapur
cc61e12e46 feat: jump-to-shadow-eval button; fix: block shadow/judge calls past budget
Adds a "Shadow eval" button next to the Auto-router usage heading that
smooth-scrolls to the shadow eval section, so it's reachable without
scrolling past the benchmarks body first.

Also fixes a PR review comment (veria-ai): shadow and judge calls ran
outside the normal auth path, so they never went through
reserve_budget_for_request and could push an already-exhausted key or
team further over budget before their own spend was even recorded.
_key_or_team_is_over_budget reads the same cross-pod spend counters
that path reserves against (via the existing get_current_spend) and
skips the shadow/judge pair outright when the shadowed key or its team
is already at or over budget. This is a read-time check, not a
reservation — appropriate for a best-effort background measurement
task, not a billed user request — so it narrows the window rather than
closing it against concurrent bursts, which the response comment
explains.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 13:17:00 -07:00
yuneng-jiang
97a59c8c90
Merge pull request #36293 from BerriAI/litellm_fix_circleci_88641_outdated_tests
test: repair stale CircleCI contracts
2026-08-08 13:08:22 -07:00
tin-berri
e35ee4e5fa
feat(router): independent, default-on deployment affinity for the auto-router (#36146) 2026-08-08 13:02:29 -07:00
Yuneng Jiang
1a40a67394
fix: stabilize generated user role ordering 2026-08-08 12:54:23 -07:00
Abhimanyu Kapur
9f9ae2b718 feat: time-bound shadow eval jobs with a real start form
A shadow eval samples ongoing traffic, so a job without an end date keeps
billing judge calls until someone remembers to stop it — and the upfront
estimate silently priced exactly one week regardless. Jobs now take a
duration_days (1-30, default 7): the start endpoint stamps ends_at, the
estimate scales trailing volume to the requested window, and the logger
completes a job past its window through the same guarded update + cache
eviction path as the spend cap (generalized into _finalize_job). The
existing shadow eval migration is amended in place since it has not
shipped anywhere yet.

The start form no longer asks anyone to paste a key hash: the key is a
type-to-search combobox backed by /key/list alias substring search that
submits the token, the auto-router is a filter-as-you-type combobox fed
by the configured auto-router deployments, and duration is a select.
Active job cards show when the job will end.

The judge model field is now labelled as such, with guidance: judging
two answers blind needs solid comprehension and reliable JSON, not
frontier reasoning — a mid-tier model (Claude Sonnet / GPT-4o class) is
recommended, nano/mini-class judges give unreliable verdicts, and
frontier reasoning models add cost without changing outcomes. Same
guidance mirrored into the API field description.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 12:36:34 -07:00
devin-ai-integration[bot]
cfd64d45a8
fix(ui): show team BYOK models in team fallback settings (#36241)
* fix(ui): show team BYOK models in team fallback settings

Team router settings loaded fallback options from /model_group/info, which resolves models without a team, so a team's own BYOK deployments were never selectable in its own fallback config. Load the team-scoped listing when a team id is present.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(ui): ignore stale team model responses in router settings

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor(ui): use react-query for fallback model listing in router settings accordion

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: ryan-crabbe-berri <ryan@berri.ai>
2026-08-08 19:28:57 +00:00
Yuneng Jiang
0d7f7c689a
test: repair stale CircleCI contracts 2026-08-08 12:19:29 -07:00
Abhimanyu Kapur
28878e877b fix: attribute shadow eval spend to the shadowed key's budget
Shadow and judge calls were fired with no caller identity on their metadata.
The proxy's cost callback requires user_api_key/_team_id/... to log spend and
apply budget checks, and silently drops the entry without them — so an
admin-enabled eval billed real provider spend that landed on no key, no team,
and no budget counter, invisible to every limit the shadowed key is normally
subject to.

Extract the identity-forwarding rules the auto-router classifier already
implements into a shared litellm/litellm_core_utils/internal_call_metadata.py:
forward the caller's identity subset, strip the parent's budget reservation
(top-level and the copy nested in user_api_key_auth) so a sub-call can't
finalize a reservation that belongs to the parent, and stamp the sub-call's
origin. The classifier now uses this module instead of its own copy.

Wire both shadow eval call sites (_call_router_shadow, _call_judge) through
it, and add a per-job spend cap (job cost_actual >= max(3x the quoted
estimate, $0.50)) so a bad estimate or a traffic spike can't turn a quoted
eval into a runaway bill; a capped job self-completes and is evicted from
cache so it can't be resurrected by an in-flight request.

Also surface api_key_id/team_id on GetShadowEvalJobResponse so an admin
running several jobs can tell which key's traffic a given win rate belongs
to, and regenerate the dashboard's OpenAPI types for the new fields.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 11:45:55 -07:00
Abhimanyu Kapur
0b94603ca8 fix: widen judge output budget and surface prior shadow eval jobs
Three fixes from the live end-to-end run:

- The judge ran with max_tokens=200, which truncated roughly 12% of
  verdicts mid-JSON so they were lost to failed_count. Raise it to a
  named JUDGE_MAX_OUTPUT_TOKENS=500 and price the upfront cost estimate
  off the same constant, so the estimate can't silently drift from what
  the judge is actually allowed to emit.

- The UI only ever rendered the newest job, so starting a new eval hid
  the results of a populated older one. Prior jobs are now listed in a
  collapsible 'Previous evaluations' card, each expandable to its own
  per-tier results.

- Move the shadow eval section above the benchmarks body: pre-adoption
  keys have no router sessions, so it was buried under an empty state.
  It stays outside BenchmarksBody so it survives that early return.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 11:01:58 -07:00
Abhimanyu Kapur
1da3bdc5cc fix: default shadow eval cost_actual to 0 so judge spend accumulates
The column was declared Float? with no default, so every row started NULL.
The verdict writer increments it, and NULL + x is NULL in Postgres, meaning
judge spend never accumulated and the UI always showed no spend.

Makes the column non-null with a default of 0 across all three schema copies
and the migration, and tightens the response model to a plain float.
2026-08-08 09:40:13 -07:00
Abhimanyu Kapur
aba922ff40 Fix CI lint and test failures
- Remove unused imports (datetime, timezone, ModelResponse) flagged by ruff.
- Update test_cost_tracking_adds_two_callbacks_when_prisma_set to expect 2
  callbacks on litellm.callbacks (not 1): ShadowEvalLogger now registers
  alongside _ProxyDBLogger in cost_tracking(). Test name already said 'two',
  now it actually tests for the correct count.
- Format ShadowEvalSection.tsx/.test.tsx per prettier.

Tests: 94/94 passing (lifecycle, shadow-eval, auto-router endpoints).
Lint: ruff + prettier all clean.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 09:30:40 -07:00
Abhimanyu Kapur
2092b809d3 Merge origin/litellm_internal_staging into shadow-eval-pre-adoption
Resolves two real conflicts (the PR's actual base branch is
litellm_internal_staging, not main):

- litellm/types/management_endpoints/auto_router_endpoints.py: kept both
  the Mapping and Literal imports, both used by pre-existing types.
- tests/e2e/proxy_client.py: kept upstream's more detailed create_model()
  docstring covering multi-replica propagation.

Everything else auto-merged cleanly. Regenerated the OpenAPI schema
(schema.d.ts) to pick up upstream's tier_turns addition to the auto-router
benchmarks response.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-07 21:18:17 -07:00
Abhimanyu Kapur
0b96dccd07 shadow-eval: fix production issues & add comprehensive tests
- Registry check now scans all pre-routing strategy registries (auto, complexity,
  adaptive, quality), not just auto_routers. Fixes 400 on adoption for
  complexity-router or adaptive-router users.

- Split auth into _require_admin_viewer (GET) and _require_admin_writer
  (start/stop); view-only admins can no longer initiate paid work (judge calls).

- request_count UPDATE buffering: in-memory counter flushed every 10s instead
  of one UPDATE per request. High-traffic keys now cost one DB op per flush
  interval, not per request.

- Default judge model: anthropic/claude-sonnet-5 (was unmapped claude-3-5-sonnet).
  Cost estimation now prices correctly; fallback is no longer needed.

- completion_cost error handling: try/except around litellm.completion_cost() so
  unmapped judge models don't crash the verdict write.

- UI: ShadowEvalSection now always renders (pre-adoption keys have no router
  sessions yet but still show the start form). Added judge_model parameter to
  the start form. Fixed accessToken undefined in AutoRouterBenchmarksTab.

Tests:
- test_shadow_eval_logger.py (26 tests): sampling, verdict parsing, unmasking,
  skip logic, metadata isolation.
- ShadowEvalSection.test.tsx (7 tests): form, active job display, per-tier
  results, low-sample flagging, completed job handling.
- All existing auto-router endpoint tests (22) and component tests (96) pass.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-07 20:58:07 -07:00
mateo
f668c10609 chore(ui): regenerate dashboard api types for ModelInfo pricing fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-08 03:25:09 +00:00
Abhimanyu Kapur
a50a391066 feat: complete pre-adoption shadow eval (backend + UI)
Backend:
- ShadowEvalLogger: background task on async_log_success_event, deterministic sampling, blind pairwise judge, verdict persistence + job counter updates
- Endpoints: POST /auto_router/shadow_eval/start (cost estimate), GET /job_id (per-tier results), GET (list), POST /job_id/stop
- Schema + migrations: LiteLLM_ShadowEvalJob, LiteLLM_ShadowEvalVerdict tables
- Types: StartShadowEvalRequest, GetShadowEvalJobResponse, ShadowEvalResult, ShadowEvalTierResult
- Proxy wiring: logger registered in cost_tracking()

UI:
- useShadowEval.ts: hooks for start/stop mutations + job query with live polling
- ShadowEvalSection.tsx: consent-gate form, job status badge, per-tier results table
- Integrated into AutoRouterBenchmarksTab alongside existing benchmarks
- OpenAPI schema regenerated to include shadow_eval endpoints

Tests: unit tests for logger sampling/unmasking/verdict-parsing
Linting: all files python3.11 syntax-check + tsx lint ready

This ships the full pre-adoption evaluation flow:
1. User starts job: specifies key, router, sampling %, gets upfront cost estimate
2. Sampled requests duplicated through router, judged blind, verdicts persisted
3. Dashboard shows per-tier win rates, cost tracking, live progress
4. Ready for Tyler/Tinder/Access Group pre-launch sign-off

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-07 20:14:51 -07:00
mateo
96a8b7f488 chore(ui): regenerate dashboard api types for tier_turns
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-08 02:09:53 +00:00
ryan-crabbe-berri
2a9aac7004
fix(ui): let access groups be a team's only model source, with hover provenance (#36234)
* feat(proxy): return per-group model provenance on /team/info

/team/info now carries access_group_details, one entry per resolved access
group with its id, name, and model list, so the UI can attribute each
inherited model to the group granting it. The batch resolver returns the
access group rows keyed by id instead of a stringly dict of lists, and the
team member budget helper returns a copy instead of mutating its parameter.
Type discipline and basedpyright budgets ratchet down accordingly.

* feat(ui): allow group-only teams and show model provenance on hover

Team create and edit no longer require a model selection: an empty
selection is saved as the no-default-models sentinel, never as a bare
empty list, since an empty team model list means unrestricted access.
The team info Models card now renders every badge with a hover tooltip
naming how the team got that model: directly, via named access groups,
or both, and group-granted badges stay visible when the direct list is
empty or a sentinel.

* refactor(proxy): dedupe access group ids and return copies instead of mutating

Duplicate access_group_ids no longer amplify the /team/info response: ids
collapse order-preserving before provenance is built, pinned by a regression
test. The resolver returns a model_copy rather than mutating its parameter,
and the team create call sends a new object instead of reassigning
formValues.models. Budgets ratchet down further with the mutation removal.
2026-08-07 17:45:50 -07:00
ryan-crabbe-berri
8db2fbaad0
feat(ui): show user email or alias in usage data export (#36232)
* feat(ui): resolve user email/alias in usage export instead of raw user id

* test(ui): cover email/alias resolution in usage export data builders

* chore(ui): drop explanatory comment per repo comment policy
2026-08-07 22:53:59 +00:00
mateo
2f70f4edba build(deps): bump nanoid to 3.3.17 in the dashboard lockfile
GHSA-2v37-7h3g-55p8 (CVE-2026-67213) rates 8.2 against nanoid 3.3.16 and reds osv-scan on every PR into staging. Custom generators can loop indefinitely when size is zero, so a caller that passes through a zero size hangs the process.

nanoid is transitive through the dashboard's toolchain and 3.3.17 is a patch release that published 08-03, so it already clears the .npmrc min-release-age=3 guard. The lock diff is the version, resolved url, and integrity hash for that one package.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-07 21:30:09 +00:00
mateo
493eb9054e chore(ui): regenerate schema.d.ts for the /key/info docstring update
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-08-07 19:13:12 +00:00
ryan-crabbe-berri
527dc0a8bb
feat(proxy): add apply_user_budget_to_team_keys opt-in (#36102)
* feat(proxy): add apply_user_budget_to_team_keys opt-in

PR #32005 made a user's personal max_budget apply to their team-scoped keys
too, and PR #35271 reverted the whole thing (behavior plus the
skip_user_budget_on_team_key opt-out) because that flipped the default for
everyone. This brings the behavior back the other way round: default is
unchanged, and general_settings.apply_user_budget_to_team_keys opts a
deployment into charging the key owner's personal budget on team keys.

The flag reaches all three personal-budget gates so an opted-in deployment
enforces consistently: the read-time check in common_checks, the optimistic
reservation counter in _get_budget_counters, and the _PROXY_MaxBudgetLimiter
pre-call hook. It is also in the /config/list allowed args and, unlike the
reverted flag, in the _update_general_settings propagation allowlist, so the
Admin UI General Settings toggle actually takes effect at runtime; an explicit
YAML value still wins over the DB value on reload.

get_config_list's allowed_args moves to a module-level frozen mapping of
field name to type string, dropping 18 LIT002 violations and rebuilding one
less dict per request.

* style(proxy): drop explanatory comments from the budget flag paths
2026-08-07 15:40:13 +00:00
yuneng-jiang
4e5495e1bd
Merge pull request #36147 from BerriAI/litellm_bump_js_yaml_4_3_1
build(deps): bump h2 to 4.4.1 and js-yaml to 4.3.1
2026-08-06 19:10:07 -07:00
ryan-crabbe-berri
429a5dc430
fix(ui): allow clearing a key's budget reset from the Edit Key form (#36140) 2026-08-06 18:46:28 -07:00
Yuneng Jiang
0253154780
build(deps-dev): bump js-yaml to 4.3.1
Closes GHSA-5p4m-2wfm-xmqj (CVSS 7.5), flagged by osv-scan against
ui/litellm-dashboard/package-lock.json. js-yaml is pinned by an exact
npm override, so the override and the lock move together.

Dev-only dependency: js-yaml reaches the tree through eslintrc, knip
and @redocly/openapi-core, none of which ship in the built dashboard.

4.3.1 published 2026-07-31, clear of the 3-day min-release-age cooldown.
2026-08-06 18:14:26 -07:00
tin-berri
7da891a42a
fix(ui): match auto-router preset models against wildcard-expanded model groups (#36111) 2026-08-06 17:47:05 -07:00
Mateo Wang
0acca3e86a
Merge pull request #24548 from mpcusack-altos/fix/bedrock-batch-credential-fields
fix(router): include Bedrock batch/S3 fields and model in deployment credentials
2026-08-05 22:39:39 -07:00
tin-berri
34fc8d2ee7
fix: expired-miss share over all measured turns + cost-optimization tab labels (#36037)
* fix(ui): make the expired-miss stat row a focusable tooltip trigger

* fix: auto-router expired-miss percentage and cost-optimization tab labels

- change expired-miss percentage denominator from return-to-tier misses to
  all measured turns (same_model + first_visit + return_to_tier). when
  auto-routers flip tiers rapidly within TTL, return-to-tier turns become
  hits and disappear from the miss count; the old metric reported only the
  rare failure population. the new metric contextualizes that population as
  a share of overall coverage
- rename usage tab from 'Usage' to 'Overall'
- rename auto-router-usage tab from 'Auto-Router Usage' to 'Auto-Router'
- update component and unit tests to match new semantics
2026-08-05 22:34:55 -07:00
tin-berri
86890654c5
fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day (#36051)
* fix(proxy): include today's UTC bucket when a daily activity range ends at the caller's current day

* fix(proxy): gate the current-UTC-day extension behind an opt-in param sent by the cost optimization dashboard

* fix(ui): label cost optimization savings dates as UTC days
2026-08-05 22:33:54 -07:00
Michael Cusack
3d275d97fe fix(router): return model and Bedrock batch fields in deployment credentials
get_deployment_credentials_with_provider dropped s3_region_name,
s3_encryption_key_id, and aws_batch_role_arn because
CredentialLiteLLMParams never declared them, and it never returned the
deployment's model, so proxy batch creation against Bedrock failed with
"LiteLLM doesn't support custom_llm_provider=bedrock for 'create_batch'"
or "AWS IAM role ARN is required" (#25104)

Provider-only file and batch calls keep their no-model contract:
get_team_provider_credentials strips the model key so a provider-scoped
request is not pinned to an arbitrary matching deployment
2026-08-05 21:49:31 -07:00
Mateo Wang
ba91768146
Merge pull request #35925 from BerriAI/litellm_tier_aware_reasoning_token_cost
fix(cost): bill reasoning tokens at the service tier output rate
2026-08-05 21:09:55 -07:00
tin-berri
7c621b3141
fix(auto-router): accept every reminder marker pair a harness emits (#36029)
* fix(auto-router): accept every reminder marker pair a harness emits

reminder_markers held one (open, close) pair, so a harness that wraps
injected context differently per agent type only got the slice of traffic
using the configured envelope stripped. Every other agent type kept hitting
the original bug: its reminder-only turn never stripped to empty, won
"newest human ask", and the harness blob got classified in place of the
real question, choosing the tier and therefore the spend.

The field now takes a list of ReminderMarkerPair, following the
KeywordTierRule pattern already in this file so each pair validates itself
and errors point at reminder_markers.N.close rather than a bare index.

Blocks from different pairs can nest, which the gap construction could not
handle: resuming the kept text at an inner block's end walks back inside
the enclosing block and leaks its remainder. Running the block ends through
a maximum collapses nested and overlapping spans without a separate merge
pass, and stays linear in block count, which a fold over a growing tuple
of merged spans would not.

A single pair's ends already increase, so the maximum is the identity and
the default path is byte-identical: verified against the shipped function
over 200k generated inputs, and every existing reminder test passes
unchanged. The prior single-pair config shape is rejected loudly at
startup and at /model/new rather than silently stripping nothing.

* docs(auto-router): document reminder_markers in the complexity router README

* chore(ui): regenerate dashboard API types for the reminder_markers shape

---------

Co-authored-by: Abhimanyu Kapur <38531241+akapur99@users.noreply.github.com>
2026-08-05 21:03:36 -07:00
mateo-berri
1b30b1bc20 Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_tier_aware_reasoning_token_cost
# Conflicts:
#	litellm/types/utils.py
2026-08-05 20:45:34 -07:00
tin-berri
ece652f6a7
feat(ui): add the auto-router usage tab to cost optimization (#35995) 2026-08-06 02:15:35 +00:00
Yuneng Jiang
6e434c926b
Merge branch 'litellm_internal_staging' into litellm_dead_locals_5_8 2026-08-05 18:07:23 -07:00
Yuneng Jiang
ce5c4c1bf9
refactor(ui): drop dead locals and unused React state across the dashboard
Removes declarations nothing reads, along with the writes that fed them, so
the remaining code says what it actually does.

Where a declaration was dead but its initializer had a real effect, the call
survives and only the binding goes: spies stay installed, renders still run,
and every awaited request keeps its await. Pure computations are deleted
whole rather than left as statements that build a value and throw it away.

Dead useState pairs are removed outright instead of being elided to
const [, setX], which would keep a hook and every write to a value nothing
reads. Three chains turned out to be dead end to end and are removed with
their fetches: the tool detail team list, the Teams MCP access group load,
and the user dashboard proxy settings load.

ColumnMeta's declaration merging in columnMeta.ts and view_logs/table.tsx is
a false positive; TypeScript requires those type parameters to match the
upstream signature exactly, so both get a scoped suppression instead.
2026-08-05 18:07:17 -07:00
yuneng-jiang
624fa11d71
Merge pull request #36025 from BerriAI/litellm_dead_locals_3_tests_and_destructures
refactor(ui): drop unreferenced locals from tests and narrow destructures
2026-08-05 17:54:12 -07:00
Yuneng Jiang
888f911133
refactor(ui): drop unreferenced locals from tests and narrow destructures
Third and fourth slices of the sweep, combined because they raise nearly the
same question and neither changes what runs.

Nine test files plus one source file lose symbols whose only mention was
their own declaration. Ten more narrow a destructure to the keys actually
read, so `const { accessToken, userRole, userId: userID, premiumUser } =
useAuthorized()` keeps only `accessToken`. Aliases are preserved as written.

ignoreRestSiblings stays on so the omit idiom `const { tags, ...rest } =
metadata` is left alone; dropping `tags` there would fold it back into rest.

ToolDetail is held back again. Its unread binding only looks like a plain
deletion on the first pass, because the dead useMemo still reads it; one more
pass exposes a useQuery that issues a real request. That belongs with the
slices that get QA'd.

Part of LIT-5162.
2026-08-05 17:46:39 -07:00
yuneng-jiang
c2c795fad5
Merge pull request #35821 from BerriAI/litellm_dead_locals_2_components
refactor(ui): drop unreferenced locals from shared dashboard components
2026-08-05 17:44:38 -07:00
ryan-crabbe-berri
f2690aa60e
fix(ui): opening a project now pushes ?project= so back and deep links work (#36001)
* fix(ui): drive project detail selection from the ?project= url param

Opening a project kept selectedProjectId in useState, so the URL never changed; the detail view could not be linked or reloaded and browser Back skipped past the Projects page entirely.

Selection now lives in the ?project= query param via nuqs with history: push, matching how Teams, Organizations and Virtual Keys already work.

* fix(ui): project detail close replaces history to match the other detail pages

Adopts the close semantics from PR #36013 so browser Back after an
in-page close leaves the Projects page instead of reopening the
dismissed detail; the close test now pins the replace mode
2026-08-05 17:31:22 -07:00
Yuneng Jiang
80627c0477
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_dead_locals_2_components
# Conflicts:
#	ui/litellm-dashboard/src/components/add_model/add_auto_router_tab.tsx
2026-08-05 17:29:54 -07:00
yuneng-jiang
1b059f472d
Merge pull request #35819 from BerriAI/litellm_dead_locals_1_app_routes
refactor(ui): drop unreferenced locals from dashboard route components
2026-08-05 17:28:15 -07:00
ryan-crabbe-berri
1dab133d24
fix(ui): link project page keys to their virtual key detail (#36002) 2026-08-05 17:26:13 -07:00
ryan-crabbe-berri
6a2e4e6c36
fix(ui): sync projects list page index to ?page= so back and reload keep the page (#36003)
* fix(ui): sync projects list page index to ?page= so back and reload keep the page

Paging the Projects list only moved TanStack's internal page index, so the URL
never changed: reload dropped you on page 1, browser Back left the page entirely,
and the page could not be shared.

The page index now comes from a nuqs ?page= query state with history: "push".
Pagination stays controlled off that value and the footer writes the URL
directly, because TanStack resets its page index whenever the data array
identity changes; letting it own the state would clear a deep-linked page as
soon as the projects query resolved. A page outside the current row set falls
back to page 1, which covers both a hand-typed ?page=99 and a search that
narrows the list below the current page.

* fix(ui): carry page_size in the url so restored history entries show the same rows

Greptile flagged that a history entry restoring ?page=N under a changed
local page size displays different projects than it originally showed.
Page size now rides the same query string via useQueryStates, size
changes reset the page inside a single history entry, and values outside
the offered options fall back to the default
2026-08-05 17:25:42 -07:00
tin-berri
55e666a05f
feat(complexity_router): report LLM classifier cost per request via routing_decision and x-litellm-classifier-cost header (#36015) 2026-08-05 16:27:32 -07:00
tin-berri
265945dfcd
feat(ui): match auto-router preset models against deployments' underlying model IDs (#35972)
The bundled presets only became selectable when an admin's public model_group
names matched the preset's hardcoded model names. model_name is admin-arbitrary,
so renamed deployments (my-claude-fast, bedrock-opus) left both presets greyed
out. Resolve preset models against each deployment's litellm_params.model and
model_info.base_model from /v2/model/info via a normalized ID join, and prefill
the admin's registered group names. Resolves LIT-5225
2026-08-05 15:24:54 -07:00
Abhimanyu Kapur
c76882b51b
fix(auto-router): stop the embedding model's context window from failing long requests (#35956)
* fix(auto-router): stop the embedding model's context window from failing long requests

The auto-router embeds the last user message to pick a model and sent it to the
embedding model unbounded. Embedding models carry 512 to 8k token windows while the
chat models they route to carry 200k+, so any prompt over the encoder's window failed
at the routing step with a 400 the destination model would never have raised.

Cut every doc to a character cap inside LiteLLMRouterEncoder, which is the one choke
point the auto-router, complexity-router, semantic guard and MCP tool filter all share.
Default 2000 chars, roughly 500 tokens, which fits even a 512-token self-hosted encoder,
overridable per deployment with auto_router_max_input_chars and globally with
DEFAULT_MAX_EMBEDDING_INPUT_CHARS.

Truncation alone cannot cover provider-side batch and byte limits, so any failure of
the route call now falls back to the auto-router's default model instead of propagating.
That path also fixes two latent bugs: a no-match left the auto-router alias in place as
the model name, which fails downstream with "Unmapped LLM provider" rather than reaching
default_model, and an empty route list raised IndexError.

Fixes #17869
Fixes #20277

* fix(auto-router): make the embedding input cap opt-in so guards still see whole prompts

Defaulting the cap inside the shared encoder truncated every consumer, not just the
auto-router. The semantic guard builds the same encoder, so its pre-call check would
have classified only the first 2000 characters while the full message still reached the
model, which a benign opener in front of an injection payload walks straight past. The
MCP tool filter and complexity router were silently narrowed the same way.

The encoder now defaults to sending docs whole and cuts only when a caller passes
max_input_chars. The auto-router is the only caller that does, so guard, MCP filter and
complexity-router behaviour is unchanged from before this branch.

DEFAULT_MAX_EMBEDDING_INPUT_CHARS becomes DEFAULT_AUTO_ROUTER_MAX_INPUT_CHARS, since it
is now specific to the auto-router, and drops its env override: the per-deployment
auto_router_max_input_chars already covers it, and every env var in constants.py has to
be documented, which is what broke the documentation and code-quality checks.

Also drops the added comments and the redundant type: ignore that review flagged.

* test(auto-router): cover the max_input_chars wiring from litellm_params

Nothing asserted that auto_router_max_input_chars on the deployment reaches the
AutoRouter that embeds prompts. Dropping the wiring left every test green while the cap
silently reverted to the default, so an operator with a 512-token embedding model could
not lower it and every long prompt would fall back to the default model instead of
being routed.

* test(auto-router): cover the populated route-choice list branch

The route layer can hand back a list, and picking its first element is where the
IndexError lived: the empty case was covered but the populated one was not, so the
branch that reads route_choice[0].name could be deleted with every test still green.
2026-08-05 14:47:40 -07:00
ryan-crabbe-berri
2dc49a913c
refactor(ui): replace hand-rolled query-param routing with nuqs (#35871)
* refactor(ui): replace hand-rolled query-param routing with nuqs

The dashboard carried five copies of the same pushState-based detail
routing hook plus a shared navigateWithParams helper, each with its own
plumbing test and a copy-pasted reactive useSearchParams mock in
component tests. nuqs provides the same shallow history-API routing
behind useQueryState/useQueryStates, so the key, team and org hooks are
deleted in favor of inline useQueryState at their single consumers,
while the models and logs hooks keep their interfaces but drop their
hand-rolled internals. Component tests now mount NuqsTestingAdapter
(via renderWithProviders or locally) instead of patching window.history,
and URL assertions go through onUrlUpdate spies that can additionally
distinguish push from replace, which the old window.location checks
could not

* test(ui): assert browser back closes the log drawer after in-drawer selection

Greptile flagged that the nuqs port of the switching-logs test stopped
at asserting emitted push and replace modes. The test now replays those
recorded modes against a history stack and performs the back step, so a
regression to push-on-select or broken URL-derived drawer state fails
the test instead of passing silently
2026-08-05 13:29:58 -07:00