mirror of
https://github.com/BerriAI/litellm.git
synced 2026-08-28 05:25:59 +00:00
* feat(ui): configure the auto router's heuristic scorer from the Admin UI The complexity router has always read tier_boundaries, token_thresholds and dimension_weights from its config, and /model/new already persists them, but the dashboard had no control for any of the three, so tuning the scorer meant editing config.yaml by hand. Adds an "Advanced scoring" panel to the classification section, shown whenever the scorer actually runs: on a heuristic router, and on an LLM classifier that falls back to the heuristic. An untouched knob is omitted from the payload, so a router keeps tracking the shipped defaults instead of freezing today's numbers. The three keys join MANAGED_COMPLEXITY_ROUTER_KEYS, so the edit modal now rebuilds them from form state rather than carrying the stored copy through. That makes hydration load-bearing, and it hydrates an absent knob to undefined rather than to the defaults, so an untouched save cannot pin a router that was tracking them. The "How Classification Works" card now reads the configured boundaries instead of hardcoding 0.15 / 0.35 / 0.60, which would otherwise start lying the moment an operator changed them. * test(complexity_router): pin the dashboard scorer defaults against config.py The Admin UI keeps its own copy of the boundary, threshold and weight defaults to prefill its controls. The copy is display only, since an untouched knob is omitted from the payload, so drift shows a stale placeholder rather than pinning a router. Nothing caught that drift before, and a blank or dead control is worse, so the two copies and the dimension key set are pinned against each other here. * fix(ui): surface out-of-order scorer thresholds as an error, not a hint Boundaries that decrease make the tiers between them unreachable, which silently changes where traffic goes, so amber body text undersold it. Saving stays allowed: a router configured this way in config.yaml would otherwise become uneditable in the UI for every unrelated change. * fix(test): search the default-model picker instead of trusting option order The pinned model is appended after every model the presets contribute, and that list has reached 11, so the option fell outside the virtualized dropdown's rendered slice and the two default-model-pin cases failed on staging. CI only runs them when this file is touched, which is why they went unnoticed. Searching for the model filters the list to it, so the cases no longer depend on how long the preset list grows. * revert(test): drop the dashboard scorer defaults parity test It parsed TypeScript from Python with a hand-rolled brace matcher and a numeric literal regex, which is not a mechanism this repo should carry: two review rounds went into fixing the parser rather than the feature. The UI copy of the defaults is display only, since an untouched knob is omitted from the payload, so drift shows a stale placeholder and cannot pin a router. * fix(ui): clamp the scorer inputs and drive the panel from one group spec min and max are inert attributes on a text input, so the fields accepted a weight of 999, a boundary of -50, and Infinity, and persisted them into the router config. Values are now clamped on commit and non-finite input is refused. The three sections were near copies of each other, so they now render from a single group spec, which also removes the triplicated warning logic. Moves the scorer constants and types into heuristic_scoring_knobs, the leaf module. Reading them back through ComplexityRouterConfig was a cycle, so the top-level DIMENSION_KEYS.map in the panel ran while the constant was still undefined and every test importing it failed to collect. * feat(ui): serve the scorer defaults from the proxy instead of mirroring them The dashboard kept its own copy of DEFAULT_TIER_BOUNDARIES, DEFAULT_TOKEN_THRESHOLDS and DEFAULT_DIMENSION_WEIGHTS to prefill the Advanced scoring controls. Two copies of one fact, and the earlier attempt to police the gap parsed TypeScript from a Python test, which was worse than the problem. GET /public/complexity_router/scorer_defaults now returns them, following the /public/providers/fields pattern: a typed response model, the dashboard fetching it through a react-query hook next to useProviderFields. The controls and the "How Classification Works" card both read that, so a recalibration of the defaults can no longer leave the form stating numbers the router stopped using. The dimension set now comes from the proxy too, so a dimension added backend-side renders without a dashboard change, under its raw key until it is given a label. Hydration keeps a stored dict exactly as stored rather than filling it from a local copy, since the backend already defaults any key omitted at scoring time. * fix(types): type the scorer defaults response as Mapping, not dict LIT001 gates mutable collections in annotations, and the three dict fields tripped it. Mapping is what the codebase already uses for a read-only map on a response model, and the endpoint hands the config constants over directly rather than copying them into a fresh dict, which would have traded the LIT001 hit for a LIT002 one. * test(ui): stub the scorer defaults request for the auto-router tree The Advanced scoring panel and the classification card read the shipped defaults over the network, so every render of that tree in a test paid for a request jsdom cannot serve. That was enough to push the slowest default-model-pin case past its 30s timeout on CI, where the suite runs 14 forks in parallel. One fixture in tests/mocks, pulled in by a single vi.mock line per test file, rather than the same stub pasted into each of the seven that render the tree. * fix(ui): tell a failed scorer-defaults load apart from a slow one The panel read only the query's data, so a permanent failure was indistinguishable from a request still in flight and it sat on "Loading the shipped defaults..." for good. It now branches on the query state: pending says loading, an error says so and offers a retry, and the values the router already overrides stay visible and editable either way. Two more places had the same flaw. The classification card silently dropped the tier ranges it used to always show, and now says they could not be loaded. The weight total was summed over whatever keys were present, so a failed load made it state a total built from the overrides alone; a total is only shown when the dimension set is known. |
||
|---|---|---|
| .. | ||
| complexityScorerDefaults.ts | ||