litellm/litellm/router_strategy
devin-ai-integration[bot] cd1107aac4
fix(router): parse the classifier verdict out of surrounding prose instead of falling to the default tier (#43215)
* fix(router): parse the classifier verdict out of prose and fences on every parse path

The complexity router's classifier parsers only tolerated a bare JSON object
(labeled tier) or a leading Markdown fence (capability, LLM V2), so a
json_object classifier that writes its verdict fenced and then explains it in
markdown, which Bedrock Haiku 4.5 does on nearly every Claude Code request,
failed validation and every request fell to the fallback tier.

All three parse paths now extract the first complete JSON object from the
reply with json.JSONDecoder.raw_decode, whatever prose or fence surrounds it,
and a reply that still fails validation is logged with the pydantic field
problems and the raw reply text, withheld when the request turns off message
logging. The capability failure reason names the exception type like the
labeled path does instead of interpolating str(e), which for a ValidationError
carried the whole reply as input_value and for TimeoutError was empty.

* fix(router): stop rejecting an LLM V2 forecast over a long explanation field

LLMV2Verdict capped crux and each forecast's likely_failure at 512 characters
through the ShortText alias, so a verdict whose explanation ran long failed
validation and the request fell to the capable tier, even though nothing
downstream reads either field. Five of nineteen real Claude Code replies from
Bedrock Haiku 4.5 tripped the cap. Both fields keep the strip and non-empty
constraints and lose the length cap; the operator-set calibration version keeps
ShortText.

* fix(complexity_router): withhold the rejected classifier reply under every message-logging opt-out and survive undecodable replies

* fix(complexity_router): withhold the rejected classifier reply when the redaction decision cannot be made

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-25 16:18:51 -07:00
..
adaptive_router Merge remote-tracking branch 'origin/main' into litellm_decrease_anys_opus5_r5 2026-09-21 12:57:38 -07:00
auto_router refactor(types): replace Any with proven types in 34 files 2026-09-20 10:16:48 +00:00
complexity_router fix(router): parse the classifier verdict out of surrounding prose instead of falling to the default tier (#43215) 2026-09-25 16:18:51 -07:00
quality_router refactor(typing): replace Any with proven types in 65 backend files 2026-09-02 09:11:36 +00:00
base_routing_strategy.py fix(proxy): stop leaking periodic tasks on every DB config reload (#42784) 2026-09-24 15:42:36 -05:00
budget_limiter.py fix(router): count provider budget spend on every API surface (#38172) 2026-09-25 14:06:49 -07:00
lar1_routing.py feat(lint): enforce Final on locals and freeze function parameters (LIT010, LIT011) 2026-08-04 12:54:39 -07:00
least_busy.py chore: merge main into fix/batch-retrieve-model-group 2026-09-15 07:02:59 -07:00
lowest_cost.py chore: merge main into fix/batch-retrieve-model-group 2026-09-15 07:02:59 -07:00
lowest_latency.py chore: merge main into fix/batch-retrieve-model-group 2026-09-15 07:02:59 -07:00
lowest_tpm_rpm.py fix(router): keep batch retrieves out of routing strategy state 2026-09-06 01:02:53 -07:00
lowest_tpm_rpm_v2.py fix(batches): stamp the model group on the proxy's model-encoded retrieve path 2026-09-06 03:03:45 -07:00
savings_baseline.py fix(spend): remove the proxy-wide autorouter savings baseline override (#38700) 2026-08-28 14:53:27 -07:00
simple_shuffle.py fix(router): honor team and key provider weights 2026-09-14 23:31:52 -07:00
tag_based_routing.py feat(lint): add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones (#42793) 2026-09-23 17:50:09 -07:00