* feat(complexity_router): calibrate the classifier rubric with worked examples
The built-in rubric stated its tier boundaries as prose alone, and prose
calibrated to consumer chat puts "non-trivial code, multi-step technical work"
at the top of the scale. That is the median request in developer and agent
traffic, so ordinary engineering read as top-tier and the router paid for the
most expensive model on it.
Adds calibration examples to the rubric, selected by a new
classifier_llm_config.rubric preset. The agentic preset (now the default)
anchors routine installs, builds, multi-file edits, and standard debugging at
MEDIUM; the chat preset omits those anchors for deployments serving only
conversational traffic. Both share the same tier criteria, the trust-boundary
paragraph, and the context-window closing line, so this moves where the
boundary sits without changing the taxonomy.
Both presets render byte-identical to the strings a prompt sweep scored, and a
test pins that, so the measured accuracy describes what a router sends.
* feat(ui): pick the classifier rubric preset on an auto-router
Adds a Rubric dropdown to the auto-router's classification panel, so the
agentic and chat presets are selectable rather than config-file only. The
prompt editor prefills from the selected preset, since prefilling agentic text
for a router on chat would show examples its classifier never receives.
The picker is disabled while a custom prompt is set, and the payload builder
drops the preset in that case: a custom prompt is the classifier's whole system
role, so the backend rejects the two together. The builder records the default
preset explicitly, so a later change to which preset is default cannot silently
move an existing router.
* fix(complexity_router): mark an unchosen rubric preset with None, not model_fields_set
The mutual-exclusion check read model_fields_set to tell an explicit preset
from the default. That flag does not survive serialization, and this config is
dumped and handed straight back to ComplexityRouter by /auto_router/test_routing,
where a dump re-states every field. So a custom-prompt classifier saved fine and
then failed validation on preview, rejecting on the second pass what it accepted
on the first.
The preset is now optional, with None meaning the default, matching how None
already means the built-in rubric for system_prompt on the same model. The
default lives in one place, DEFAULT_RUBRIC_PRESET, resolved where the prompt is
assembled. The dashboard stops sending a copy of the default it displays, so a
router nobody configured follows the default rather than pinning today's value,
and UI-built routers behave the same as hand-written config.
Regenerates schema.d.ts, which was left stale by an earlier description edit.
* feat(complexity_router): grandfather existing routers onto the uncalibrated rubric
An unset preset now means LEGACY, the rubric exactly as it shipped before
calibration examples existed, so upgrading cannot move the tier decisions or the
bill of a router that is already running. Config-file routers get this for free
since they name no preset, and a stored config that never had one reads the same
way.
New routers still get the calibrated rubric: switching a classifier to LLM
stamps the agentic preset, because a classifier being configured for the first
time has no prior tier behaviour to preserve. The picker offers legacy so an
existing router's state is representable and opening the form cannot silently
upgrade it.
Each preset is pinned byte-identical to the text the prompt sweep scored,
legacy included, which is what proves an existing router's prompt did not move.
Also collapses the preset data from a NamedTuple with group wrappers and
per-preset frozensets into plain text blocks in a MappingProxyType, matching how
the tier criteria next to it are already stored: 21 lines of prompt text no
longer cost 190 lines of constructors. Tiers are format placeholders so
tier_labels still reach the examples.
* refactor(complexity_router): name the field classification_rubric
`rubric` alone did not say what it selects, and the field sits beside
`system_prompt`, which genuinely is the whole classification prompt. The name
now says which of the two an operator is reaching for: the rubric the built-in
prompt is assembled from, not the prompt itself.
Renames the config field, the query param, the enum, and the dashboard label to
match, and moves the preset text to classification_rubrics.py.
* test(ui): set the preset the mutual-exclusion case is meant to drop
The rename left classification_classification_rubric in the custom-prompt case,
so its input never carried a preset and the assertion held for the wrong reason:
it proved an absent preset stays absent, not that a set one is dropped. A
normalizer that forwards the preset whenever one is set passed with the typo and
fails without it.
tsc reports the typo as TS2353; the earlier sweep grepped for the source file
and not the test, so it went unseen.
* test(ui): scope the role-gate assertions to each page's own endpoint
The memory, workflows, and guardrails-monitor page tests asserted that a denied
role fires no request at all. Their names, and the assertion on the very next
line, say the intent is narrower: the page must not fetch its own data.
Resolving whether a caller is an org admin goes through /organization/list for
every role, since deciding org-admin-for-any-org needs the list, and the route
scopes rows per caller. That legitimate request fails a blanket no-fetch
assertion, so all three files went red on staging for a reason unrelated to
what they test.
Drops the blanket assertion and keeps the scoped one. Bypassing the gate in
memory/page.tsx still fails five tests, so the narrower assertion continues to
catch a genuinely broken gate.
* fix(complexity_router): document that an unset rubric keeps the legacy prompt
The field said 'Leave unset for agentic' while an omitted rubric resolves to
LEGACY, so the OpenAPI schema an operator reads promised calibrated routing
where they got the uncalibrated one.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
System prompts (harnesses, tools, framework boilerplate) are session-wide
constants identical across all requests. Scoring them saturates keyword-match
signals and produces false-positive high-complexity classifications on
trivial utterances like 'hi', routing them to expensive models (sonnet/opus)
instead of tier-1 haiku. A real ~1.6KB CLI-agent harness alone supplied
5 codePresence + 2 technicalTerms matches, overshadowing user signal.
Rescope four scoring dimensions (codePresence, technicalTerms, simpleIndicators,
multiStepPatterns) from full_text (system + user) to user_text (user only).
reasoningMarkers was already scoped this way. This returns 0.63 of the weight
budget to text that actually varies per-request.
Now that every dimension scores user_text only, _score_keyword_match's
disclosable_text param is redundant -- it existed solely to let the signal
name terms matched in the caller's own message while withholding terms
matched only in the (invisible-to-the-caller) system prompt. With no more
system-prompt text in scope, text and disclosable_text were identical at
every call site, so the param is dropped and the function collapses to a
single text argument.
Add mutation-proven regression test: trivial 'hi' message with realistic
Claude Code agent system prompt now routes to haiku tier-1 (not sonnet).
- Unfixed: haiku -> sonnet (bug)
- Fixed: haiku -> haiku (correct)
Invert three pre-existing assertions in TestSignalsNeverQuoteTheSystemPrompt
to capture the corrected behavior: system-prompt-only terms produce no signal.
Co-authored-by: Claude <noreply@anthropic.com>
POST /azure_ai/indexes carries no index name, so
get_azure_ai_search_index_from_endpoint returns None,
is_vector_store_index never matches any segment, and the request falls
through to the generic Azure passthrough on the proxy's own
AZURE_API_BASE and AZURE_API_KEY without ever reaching
is_allowed_to_call_vector_store_endpoint. A non-admin could therefore
create a Search index whenever AZURE_API_BASE points at the Search
service.
The earlier lifecycle commit made this look covered. Its test asserts
that POST /indexes?api-version=... is refused with "Only proxy admins can
create", but it calls the permission gate directly, and that gate is
exactly what the route skips for a path with no index name, so the guard
was verified in isolation while the route stayed open.
Gate the service-level create on the route itself, before the segment
loop, with assert_proxy_admin_for_vector_store_index_management. Scope it
to POST on a path whose last segment is indexes, mirroring the
endswith("/indexes") branch the lifecycle helper already uses, so the
managed-index paths and ordinary Azure OpenAI passthrough traffic are
untouched.
Add route-level tests: a non-admin is refused with the admin-only message
and never reaches the passthrough handler, an admin still creates, and the
new predicate is parametrized over the service-level, per-index, and
non-Search paths.
The endpoint map covered document reads through the ("GET", "/indexes/")
entry plus POST /docs/search, which left Azure's remaining POST query
endpoints unclassified. POST /docs/suggest, POST /docs/autocomplete, and
POST /analyze matched neither list, so the permission gate resolved
permission_type to None and raised 403 before the caller's
allowed_vector_store_indexes grant was consulted; a non-admin team with a
read grant on the index still could not call them.
Add the three as reads. They are query endpoints that never mutate the
index, so a read grant is the right gate, and each needs its own literal
entry because the write entry also matches on POST.
Keep every pattern literal rather than a {placeholder} template: the
matcher falls back to the substring before a {, which for these routes is
always /indexes/, and reads are matched before writes, so a templated
read would shadow the /docs/index write and let a read-only team upload.
Extend the regression tests to the full non-lifecycle read surface
(stats, GET-form search, $count, point lookup, and both forms of suggest
and autocomplete, plus analyze), asserting a read grant reaches all of
them and a write-only grant reaches none.
The Azure passthrough scanned every URL segment for one matching a registered
index, authorized against that, then forwarded the original path. A caller with
a grant on a managed index named e.g. "index" or "docs" could send
POST /azure_ai/indexes/{victim}/docs/index: the scan matched the trailing
segment and authorized on the caller's own index while Azure applied the batch
write to {victim} on the same Search service, enabling cross-index document
uploads or deletions.
Resolve the index positionally from the /indexes/{name} segment and require
that exact name to be the one authorized and credentialed, so the authorized
index and the physical target can never diverge. Add a pure helper plus
regression tests covering positional extraction and the route-level cross-index
attack.
The service-level index-create guard checked normalized.endswith("/indexes")
without stripping the query string, so Azure's real create request
POST /indexes?api-version=... was never classified as a lifecycle request and
fell through to the generic permission check instead of the explicit admin-only
guard. Strip the query string before the suffix check, mirroring how the
PUT/DELETE index paths already tolerate a trailing ?.
Add the POST create path to the lifecycle regression parametrize so a non-admin
team with a write grant is denied with the clear admin-only message.
The Azure AI Search vector store config declared its write endpoint as
`PUT /docs` and its read endpoints as only `/docs/search`. The passthrough
permission gate (`is_allowed_to_call_vector_store_endpoint`) derives a
read/write permission type by matching the request route against those
lists, and a route matching neither resolves to `None` and raises a 403
before the caller's `allowed_vector_store_indexes` grant is ever checked.
Two real Azure routes fell through that gap for non-admins: document
upload/merge/delete is `POST /docs/index` (not `PUT /docs`), and get
index details is `GET /indexes/{name}` (no `/docs/search` suffix). So a
team with a valid write or read grant still got 403 on upload and on
reading index details, while admins slipped through because they skip the
gate entirely.
Correct the map: read is any GET under `/indexes/` (get details, stats,
count, and the GET form of search) plus `POST /docs/search`; write is
`POST /docs/index`. Index lifecycle (create/update/delete the index
itself) stays proxy-admin only because it is handled first by the
separate lifecycle check on POST/PUT/DELETE/PATCH, so this does not let a
team create or delete indexes.
Add regression tests that exercise the real AzureAIVectorStoreConfig map:
a write-granted team may upload, a read-granted team may search and get
index details, a team missing the matching grant is still denied, and a
team cannot manage index lifecycle even with a write grant.
Resolve the authorized own-user and permitted-team predicates once and add
regression coverage for explicit-user intersection, unfiltered team scope,
and team lookup failure fallback.
Co-Authored-By: Codex
Add a bounded spend-log user facet for the Request Logs picker and
intersect explicit user filters with the caller's own and permitted-team
scope.
Co-Authored-By: Codex
Address review feedback on #36762:
- Only use the parsed 5m/1h split when it fully accounts for
cacheWriteInputTokens; an unrecognized ttl or missing entry now falls
back to the aggregate (previous behavior) instead of silently
understating cost.
- Mark TypedDict fields ReadOnly (AWS response data, never constructed
by us) to satisfy the repo's type-discipline lint gate.
- Trim comments and add Final to locals per repo style.
Co-Authored-By: pi (Claude/GPT via @earendil-works/pi-coding-agent) <noreply@earendil.works>
AmazonConverseConfig._transform_usage only read the aggregate
cacheWriteInputTokens field, so cache_creation_token_details was always
unset for Bedrock Converse responses. calculate_cache_writing_cost bills
the whole cache-write count at the 5m rate whenever that field is None,
so 1-hour TTL cache writes on the standard Bedrock chat path were always
undercounted, even though Bedrock returns the 5m/1h split in
usage.cacheDetails.
Parse cacheDetails (when present) into CacheCreationTokenDetails so the
correct rate applies to each portion. No cacheDetails in the response
(older models/regions) keeps the previous behavior.
Fixes#36760
Co-Authored-By: pi (Claude/GPT via @earendil-works/pi-coding-agent) <noreply@earendil.works>
langfuse_* request headers land in metadata as strings, but the trace path reads
mask_input/mask_output with a bare truthiness check and iterates update_trace_keys
directly. A header saying mask_input: false redacted the payload it was asked to
keep, and update_trace_keys was walked one character at a time so every requested
key silently failed to match
* fix(guardrails): scan and re-emit raw Anthropic SSE streams in the bedrock post-call hook
* fix(guardrails): keep upstream id and model on a blocked Anthropic stream
* fix(guardrails): deliver a blocked Anthropic stream as an error frame
* fix(guardrails): deliver an unscannable Anthropic stream as an error frame
* fix(guardrails): emit the guardrail block detail as JSON in the stream error frame
* fix(guardrails): deliver an Anthropic block through the shared block-SSE builder
* fix(guardrails): keep the shared SSE assembler behavior-identical for existing callers
* fix(guardrails): keep the stream error message a string and drop an unreachable branch
* chore(guardrails): drop a comment that repeated its own docstring
* fix(guardrails): let bedrock service failures keep their status instead of framing them as blocks
* fix(guardrails): key the streamed block decision on status, not detail shape
InvokeGuardrailChecks details a Mapping on its 500 for an unparseable response,
so a detail-shape test read that outage as a policy block and framed it as a 200
guardrail_error. Both block sites raise 400, so gate on the status too.
* refactor(guardrails): narrow the SSE error-frame helper to the input it actually takes
Both callers pass a string, so the Mapping overload and its json.dumps branch
were unreachable. Folds the block branch's narrative comment into the rebind
suppressions that already carry a reason.
* test(interactions): follow Google spec drift replacing Turn with typed steps
* test(interactions): send step and content-list input to the live Gemini API
* fix(langfuse): emit otel trace version and release on the keys langfuse v4 reads
The langfuse_otel exporter wrote version to langfuse.generation.version and
langfuse.trace.version, and release to langfuse.trace.release. Langfuse v4
recognizes neither, so both landed in the generic span attribute bag and every
trace reported version and release as null. v4 has a single langfuse.version
key, lifted to the trace when it sits on the root span, plus langfuse.release.
Also routes the otel v2 preset's per-request headers through the shared builder
so key-scoped and team-scoped exports carry x-langfuse-ingestion-version like
the other three exporter paths already do.
* fix(langfuse): give trace_version precedence over version on the shared v4 key
Matches the documented contract in docs/observability/langfuse_integration.md
and the legacy langfuse SDK callback, which both treat trace_version as the
authoritative trace version with version as its fallback.
cost_per_token reads litellm.model_cost, which CI loads from the
model_prices_and_context_window.json on main, so the new entries were
missing until merge. Use the same local_model_cost_map fixture the
sonnet 5 metadata test uses to force the branch's own backup map.
Pre-routing now reads the request's tags on every request with a registered
strategy, including the single-strategy case that used to short-circuit before
looking at tags. Metadata is request-controlled, so a caller that sends
`litellm_metadata` (or `tags`) as a string or any other non-dict shape crashed
tag lookup with an AttributeError instead of routing untagged.
* fix: never price a strategy-router alias
A strategy-router alias (auto_router/complexity_router/<name>) is never the
deployment that gets called or billed, but custom pricing configured on it was
being treated as real pricing in two places:
- registered in litellm.model_cost under the alias deployment id, so an
explicit zero made _is_cost_explicitly_configured() report the group as a
genuinely free model and every budget check was skipped, while the request
routed to a paid deployment and accrued real spend
- copied onto request_kwargs by the alias-params merge, so the routed
deployment got re-registered at the alias price and the request billed 0.0
Both are fixed at the writer, so config, /model/new and price-map reload all
take the same path
Co-Authored-By: Claude <noreply@anthropic.com>
* chore: annotate filtered cost-map copy for the mutable-collection gate
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(proxy): serialize model reconciles so concurrent writes stop evicting each other
A model write is a read-modify-write of the shared `llm_router` global: read the
db into a snapshot, then make the router match that snapshot. Nothing serialized
it, so two of them interleaving was not a lost update but an eviction --
_delete_deployment removes every live deployment absent from the snapshot it was
handed, so the request holding the older snapshot reconciles the newer request's
model straight back out of the router. The row survives in the db, which is what
makes it easy to miss: the pod simply stops serving a model it was told to serve
until some later reload happens to put it back.
clear_cache compounds it. It deletes every db model from the router before
reloading them, so for the width of that reload the pod serves none of them --
and any concurrent write sampling the router in that window sees the hole.
Fix is one lock (MODEL_RECONCILE_LOCK) held across both, so each reconcile reads
the db and applies it atomically and no stale snapshot can evict a newer model.
clear_cache holds it across wipe+reload and calls the already-locked
_add_deployment_locked, since asyncio.Lock is not reentrant and routing back
through the public add_deployment would deadlock the pod's whole model-write
path.
The verdict needed the same treatment. raise_if_reload_degraded_serving compared
a desired-set read during the reload against a router snapshot taken after it,
so a neighbouring reconcile's in-flight wipe was reported to the caller as
collateral damage from its own reload -- a 500 on a create that had in fact
succeeded. Reconciles now return a ReconcileOutcome carrying both the desired set
and the post-reconcile serving state, captured before the lock is released, and
the verdict judges against that. Omitting live_after keeps the old live re-read,
which stays correct for the no-reconcile-ran case.
Found by running the e2e suite with pytest-xdist at 8 workers: three unrelated
tests failed together on "Previously served model id(s) [...] are also no longer
being served by this pod", which is this. Serial runs concurrent enough to hit it
are rare, which is why 78 minutes of sequential e2e never surfaced it -- but any
customer provisioning models in parallel (terraform, CI) is in exactly this race.
test_reconciles_serialize_so_no_stale_snapshot_can_evict fails with 5 == 1
without the lock.
* fix(tests): return a ReconcileOutcome from the PTU test's add_deployment mock
test_ptu_model_settings.py stubs proxy_config.add_deployment with
AsyncMock(return_value=None). Now that add_deployment returns a
ReconcileOutcome, add_new_model reads .still_desired off that None and
the two PTU gate tests fail with "'NoneType' object has no attribute
'still_desired'".
Return ReconcileOutcome(still_desired=None, live_after=None), matching
the other reconcile mocks. Both fields None means no reconcile state was
captured, so the serving verdict falls back to reading the router live,
which is what the test's mock_router already drives — the PTU assertions
are unchanged.
Two sibling test files were updated for this in the parent commit; this
one was missed because the local env cannot collect four modules under
tests/test_litellm/proxy (prisma generate artifacts), so the full shard
only ran in CI.
Also applies ruff format to proxy_server.py: the new add_deployment
wrapper's single call fits on one line under the project's line length.
* fix(proxy): lock the delete evictions and stop clear_cache wiping deployments
Two follow-ups to MODEL_RECONCILE_LOCK, both found by review.
1. delete_model and delete_team_models evict from llm_router directly,
outside the lock. The db row is gone by then, but a reconcile that
snapshotted the db BEFORE the delete still lists that id as desired and
upserts the deployment straight back, so the pod keeps serving a model
the database no longer has until some later reconcile notices. Taking
the lock orders the eviction after any in-flight reconcile's re-add.
Both new tests fail without the lock ("did not wait for
MODEL_RECONCILE_LOCK") and pass with it.
2. clear_cache no longer wipes deployments. It used to delete_deployment()
every db model before the reload restored them, which left the router
serving ZERO db models for the entire width of the reload -- every
inference request landing in that window fell into a real hole, and
serializing reconciles made the aggregate outage additive rather than
overlapping. The wipe was also redundant: _delete_deployment evicts
exactly the ids the db no longer lists, and upsert_deployment
pops-and-re-adds a deployment whose params changed while no-opping one
that did not, so the reconcile converges to the same state on its own.
Every mutation is visible to that comparison (blocked, and updated_at
for premium, are written into model_info).
The auto-router pops are NOT redundant and stay: they are keyed by
model_name, which no deployment-id reconcile touches.
The new tests patch their own lock rather than contending the module-level
one: asyncio.Lock binds to the event loop of its first contended acquire
and raises on every other loop after that, which would poison the next
asyncio test in the process. The proxy has a single event loop for its
lifetime so this is test-only, but it is a trap worth naming for whoever
writes the next concurrency test here.
* fix(proxy): scope the clear_cache wipe to auto-router deployments
Review caught a regression in the previous commit. Dropping the wipe
entirely stranded every db-backed auto-router on the pod.
The strategy registries (auto_routers, complexity_routers,
adaptive_routers, quality_routers) are keyed by model_name, which no
deployment-id reconcile touches, so clear_cache pops them and relies on
the reload to rebuild them. But the rebuild only happens on the ADD path:
Router.upsert_deployment returns early when a deployment is unchanged and
never reaches add_deployment -> _add_deployment ->
init_auto_router_deployment, which is what repopulates them. With the wipe
gone the deployment was always unchanged, so the pop was permanent: ANY
unrelated model write -- a team admin patching one team-owned model --
left every db-backed auto, complexity, adaptive and quality router
unroutable across tenants until a restart.
Restore the wipe for exactly the auto_router/* db deployments, whose
strategy entries are the ones being popped. Deleting them forces upsert
down the add path so both the deployment and its strategy entry come back.
Ordinary db models stay un-wiped, which is the point of the previous
commit: wiping them un-served every db model for the width of the reload,
and the reconcile converges without it.
test_clear_cache_wipes_auto_routers_but_leaves_ordinary_db_models pins
both halves against each other, since fixing either one naively breaks the
other. Both clear_cache tests fail with the pop-without-delete version.
* refactor(clear_cache): fold auto-router wipe into the classification pass
The auto-router scoping added in 5deddfd introduced two new mutable-collection
constructions, pushing LIT002 five over its budget ceiling.
Rather than suppress, do the work in the single pass that already walks
current_models: detect and delete the auto_router/* db deployments while
classifying, accumulating names into a set that replaces the old
db_router_deployments comprehension. Net-zero LIT002, same behaviour.
Comment updated to describe where the wipe actually happens now.