Commit graph

265 commits

Author SHA1 Message Date
daqiangganjun
8327cd6d47
fix(router): count provider budget spend on every API surface (#38172)
* fix(router): count provider budget spend on every API surface

RouterBudgetLimiting read custom_llm_provider from litellm_params, which only
chat completions populates. Responses, anthropic_messages, embedding and rerank
calls raised inside the success callback before any spend was recorded, so those
budgets never moved and a ceiling made up mostly of that traffic was never hit.

Read the provider from the standard logging payload, which every surface fills
in. Dropping the raise also stops one missing field from taking the deployment
and tag budgets down with it.

* chore(router): drop the inline comment and type the budget limiter test helper

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-25 14:06:49 -07:00
yuneng-jiang
f6882246d4
test: move tests/test_litellm root and small trees into tests/unit (#43186)
* ci: run the unit_selection.sh shard files on every event instead of only fork pull requests

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: rename fork-flag to unit-flag now that it applies on every event

* test: move tests/test_litellm root and small trees into tests/unit

Pure renames, no content changes. Follow-up commits in this PR fix
references, merge the three files that already existed in tests/unit,
keep live-provider tests in tests/test_litellm and wire CI.

* test: carry tests/test_litellm conftest isolation into tests/unit

Callback lists, routing fallbacks, cached HTTP clients, logger state, AWS,
proxy-URL and keychain env, and session-end client cleanup now reset for
unit tests too. The environment isolation owns its MonkeyPatch so a test's
own monkeypatch is undone before the model-cost teardown runs.

* test: merge, split and prune the moved root and small-tree tests

Merge batches/test_batch_utils.py and the chat_completions and messages
dispatch tests into the files that already existed in tests/unit. Keep
the live Gemini interactions tests, the async image-fetch format test and
the OpenAI embedding scorer test in tests/test_litellm since they need
real network or keys. Put test_router.py under tests/unit/test_router so
the existing package no longer shadows it. Delete eight tests the audit
found superseded by stronger ones kept in this move.

* ci: run the moved root and small-tree tests under their legacy flags

Add the misc and responses-caching-types flags to unit_selection.sh and
CircleCI, extend enterprise-routing and mcp-integration, and point the
legacy GHA shards, Makefile, redis-compat workflow, merge smoke manifest
and change classifier at the new paths.

* test: make the new tests/unit directories packages

tests/unit/test_package_layout.py requires every directory to carry an
__init__.py, and without one the moved and retained
test_litellm_responses_bridge.py modules collide on import.

* test: scope the unit socket block to tests/unit in shared sessions

The GHA shards collect the legacy test-path and the unit selection in one
pytest session. The unit conftest's loopback-only block leaked into legacy
modules that reach the network at import. The legacy conftest now lifts the
restriction at collect and setup time, and the unit conftest re-applies it
when collecting its own modules.

* test: give the shard-script tests their own GITHUB_OUTPUT

They only passed where the runner set it. The CircleCI unit job's env
allowlist drops it, so the script's redirect failed there.

* test: point the router and module-deletion checks at tests/unit

router_code_coverage and code_qa_check_tests only searched tests/test_litellm,
so the moved router tests no longer counted. The two silent-experiment tests
the audit deleted were the only direct callers of those methods; they are
replaced with tests that assert the forwarded shadow request and the
recursion guard.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 11:30:43 -07:00
devin-ai-integration[bot]
ccf866c801
fix(router): await budget redis pipeline before sync reads (internal copy of #32618) (#43125)
* fix(router): await budget redis pipeline before sync reads

* refactor(router): remove superseded Redis flush helper

* fix(router): preserve concurrent spend during Redis synchronization

* fix(router): type per-key spend totals without loop Final bindings

* fix(router): log Redis failures before cancellable cleanup

* fix(router): finalize Redis batches before propagating cancellation

* test(router): reproduce cancellation while Redis cleanup is blocked

* test: align budget hotpath checks with awaited Redis flush

* fix(budgets): finish Redis flush after cancellation while queued

* test(budgets): consolidate Redis regressions in mapped tests

* fix(budgets): coalesce failed Redis increments by key

* test(budgets): assert spend behavior instead of batch state

* perf(router): sum pending budget spend by key once per sync

---------

Co-authored-by: Emerson Gomes <emerson.gomes@thalesgroup.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-09-24 22:27:13 -07:00
devin-ai-integration[bot]
fc87a06f00
fix(proxy): stop leaking periodic tasks on every DB config reload (#42784)
Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-24 15:42:36 -05:00
tin-berri
3b9c9e0523
feat(router): add group-scoped priority routing strategy (#42378)
* feat(router): add group-scoped priority routing strategy

* fix(router): satisfy priority routing type-discipline checks
2026-09-22 13:09:38 -07:00
tin-berri
5a8c4f48e4
feat(router): native compact-to-fit across conversation APIs (#42074)
* feat(router): native compact-to-fit across conversation APIs

* fix(router): preserve compaction admission and shared client boundaries

* fix(router): honor compaction fit fallbacks and router-scoped access

* fix(router): charge compaction usage to caller token limits

* test(http): keep FastAPI inside proxy tests

* fix(router): check compactor capacity before skipping escalation
2026-09-21 22:52:29 -07:00
tin-berri
275c0c4d96
Merge pull request #42057 from BerriAI/litellm_classifier_forecast_cards
feat(ui): show Capability and FUSE v2 routing forecasts
2026-09-21 17:55:25 -07:00
tin-berri
094a60bb9c
Merge pull request #41872 from BerriAI/litellm_context_escalation_opt_in
fix(router): make context-window escalation opt-in
2026-09-21 17:03:33 -07:00
yuneng-jiang
8d4ef24496
Merge pull request #41795 from BerriAI/litellm_wt_0918_5836
test(router): cover legacy lowest TPM selection
2026-09-21 16:57:19 -07:00
Tin Chi Lo
9221109d18 chore(router): resolve merge conflict with main 2026-09-21 16:52:47 -07:00
Tin Chi Lo
7f997a420d chore(ui): resolve routing forecast merge conflict 2026-09-21 12:57:57 -07:00
Tin Chi Lo
71465fc2f7 chore(router): resolve merge conflict with main 2026-09-21 11:29:48 -07:00
Moe Khalil
bb46e8b774 chore(auto-router): merge main and preserve JEV configuration
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-21 17:16:18 +00:00
Tin Chi Lo
a233ba910d feat(auto-router): configure heuristic v2 success threshold 2026-09-21 08:53:13 -07:00
Moe Khalil
8505d10017 chore(auto-router): merge main into JEV launch branch
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 14:53:05 +00:00
yuneng
78a751c049 test: migrate phase 15 legacy tests to tests/unit
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 11:55:43 +00:00
Tin Chi Lo
7c79c7efad feat(ui): show Capability and FUSE v2 routing forecasts 2026-09-19 17:04:14 -07:00
Moe Khalil
e77c154c36 chore(auto-router): merge main into JEV launch branch
Some checks failed
LiteLLM Rust / rust-lint (push) Waiting to run
LiteLLM Rust / rust-test (push) Waiting to run
LiteLLM Rust / rust-wheel (push) Waiting to run
Terraform Provider / gofmt, vet, build, test (push) Has been cancelled
Terraform Provider / Provider endpoints vs proxy OpenAPI schema (push) Has been cancelled
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-19 23:28:15 +00:00
tin-berri
252a0f1eac
Merge pull request #41617 from BerriAI/litellm_fuse_model_profile_presets
feat(router): add maintained Fuse model and harness presets
2026-09-19 16:17:19 -07:00
tin-berri
cc2c0d0f66
Merge pull request #42001 from BerriAI/litellm_heuristic_v2_scores
fix(auto-router): show heuristic v2 score estimates in routing details
2026-09-19 16:16:59 -07:00
Tin Chi Lo
ed40241d26 fix(proxy): estimate auto-router baseline costs from durable cache history 2026-09-19 12:44:47 -07:00
Tin Chi Lo
dde73968cf fix(auto-router): show heuristic v2 score estimates in routing details 2026-09-19 12:19:43 -07:00
Tin Chi Lo
4d659135b6 fix(router): preserve unavailable Fuse presets 2026-09-19 11:48:40 -07:00
Tin Chi Lo
e9109ddf4a fix(router): make context-window escalation opt-in 2026-09-19 10:01:43 -07:00
Moe Khalil
e0b2c51144 fix(auto-router): validate JEV usage and clear stale context
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 23:41:20 +00:00
Moe Khalil
8e5f43f458 fix(auto-router): preserve JEV accounting and context bounds
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 22:09:27 +00:00
Moe Khalil
86e079d7a8 feat(auto-router): integrate JEV context and usage accounting
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-18 21:38:15 +00:00
Yuneng Jiang
066cc1883a
test(router): cover legacy lowest TPM selection 2026-09-18 02:39:58 -07:00
ryan
a9ad1bbad8 Merge remote-tracking branch 'origin/main' into litellm_routing_groups_atomic_validation 2026-09-18 09:04:26 +00:00
mateo-berri
853fd04bd5 test(router): price the unpriced Jev cost test off a model the registry never ships 2026-09-17 17:18:03 -07:00
mateo-berri
1e7e5b695f fix(router): reject a blank Jev api_key so it cannot pair with a caller-chosen api_base 2026-09-17 16:38:54 -07:00
mateo-berri
7f581f6bc7 fix(router): keep the TypeSafe key off caller-chosen Jev endpoints 2026-09-17 16:09:31 -07:00
mateo
0a66328663 fix(router): validate Jev classifier probabilities
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:31:08 +00:00
mateo
9b7fcd0480 feat(router): add TypeSafe Jev as a complexity router classifier
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-17 17:05:07 +00:00
Tin Chi Lo
2fa115db2b feat(router): add maintained Fuse model and harness presets 2026-09-17 12:53:15 -04:00
ryan
9afac68995 test(router): cover _register_router_selector and _replace_routing_groups directly
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:43:04 +00:00
ryan
9f990c4f86 fix(router): validate routing_groups at save time and keep invalid DB groups from blocking SSO load
Overlapping routing_groups persisted from the Admin UI raised inside
Router._init_routing_groups during the DB config reconcile, which skipped
loading SSO, guardrails and the other DB-backed settings while leaving the
proxy healthy. /config/update now returns 400 for overlapping models,
duplicate names, the reserved default name and unknown strategies before
writing, the Router builds every group selector before replacing its state
so a rejected update keeps the previous groups routing, and the proxy applies
routing_groups separately from the other router settings so an already
persisted invalid value is logged and skipped instead of aborting the
reconcile. The Admin UI modal blocks picking a model another group owns.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-16 01:25:55 +00:00
Tin Chi Lo
39cf1f302d feat(router): apply entitlement limits to forecast classifiers 2026-09-15 16:17:00 -07:00
Tin Chi Lo
9352d24863 fix(router): accept fenced Fuse classifier verdicts 2026-09-15 13:11:56 -07:00
Tin Chi Lo
56b20525f5 fix(router): honor Fuse task context and fallback policy 2026-09-15 13:11:56 -07:00
Tin Chi Lo
92bece2baa fix(router): expose exact Fuse v2 forecast metadata 2026-09-15 13:11:56 -07:00
Tin Chi Lo
902b2e7ef8 feat(router): add experimental joint LLM V2 classifier 2026-09-15 13:11:55 -07:00
tin-berri
d07b2e87d2
Merge pull request #41270 from BerriAI/litellm_capability_classifier_pr
feat(router): add capability classifier as Fuse foundation
2026-09-15 13:04:21 -07:00
Yassin Kortam
1ca4579375
Merge pull request #41178 from BerriAI/litellm_request_override_selector_callbacks
fix(router): bind per-request routing_strategy override selectors to the request's callbacks
2026-09-15 12:39:20 -07:00
Tin Chi Lo
cadb7ee44d fix(router): preserve native encrypted capability tasks 2026-09-15 12:05:30 -07:00
Tin Chi Lo
896f35c751 fix(router): extract capability tasks with request scoped markers 2026-09-15 11:52:20 -07:00
Tin Chi Lo
e62f0d0376 fix(router): reject unknown capability policy fields 2026-09-15 11:43:04 -07:00
Tin Chi Lo
2da9bbfc0f chore: merge main into capability classifier copy 2026-09-15 11:24:01 -07:00
tin-berri
3ad9a7f336
Merge pull request #41174 from BerriAI/litellm_tier_model_affinity
fix(router): preserve session model choice within each complexity tier
2026-09-15 09:54:53 -07:00
Tin Chi Lo
81340439fc fix(router): preserve session model choice within each complexity tier 2026-09-15 00:09:17 -07:00