diff --git a/cookbook/litellm_router/fusion_models.md b/cookbook/litellm_router/fusion_models.md
deleted file mode 100644
index 635d8830ba6..00000000000
--- a/cookbook/litellm_router/fusion_models.md
+++ /dev/null
@@ -1,157 +0,0 @@
-# Fusion models (beta)
-
-Fusion models expose several LiteLLM model groups as one model. For each client call, every panel model receives the
-same canonical conversation and runs in parallel. An aggregator then synthesizes their work into the sole response
-returned to the client
-
-Fusion operates at the model layer. Coding agents, research loops, chat applications, and tool-using workflows keep
-their current control flow and use the Fusion model name anywhere they would name one model
-
-```text
-client or harness
- |
- | one model request
- v
-panel A ----\
-panel B -----+--> aggregator --> one response or tool call
-panel C ----/
-```
-
-If the aggregator returns a tool call, the existing harness executes it, appends the result to its conversation, and
-calls the Fusion model again. That next call starts a new panel round. Fusion does not execute tools, retain private
-panel transcripts, create subagents, or replace the harness
-
-## Create a Fusion model
-
-In the dashboard, open **Models & Endpoints**, select **Fusion Models (Beta)**, and choose **Add Fusion Model**. Select
-two to six existing model groups for the panel and one existing model group as the aggregator. LiteLLM rejects nested
-Fusion models
-
-The equivalent `config.yaml` entry is:
-
-```yaml
-model_list:
- - model_name: panel/one
- litellm_params:
- model: openai/gpt-5
-
- - model_name: panel/two
- litellm_params:
- model: anthropic/claude-sonnet-4-5-20250929
-
- - model_name: fusion-aggregator
- litellm_params:
- model: openai/gpt-5
-
- - model_name: fusion/coding
- litellm_params:
- model: fusion_router
- fusion_router_config:
- panel_models:
- - panel/one
- - panel/two
- aggregator_model: fusion-aggregator
- min_successful_panelists: 2
- panel_timeout_seconds: 120
- max_candidate_chars: 12000
- on_quorum_failure: fail
-```
-
-Clients call `fusion/coding` like a regular model. The beta supports Chat Completions, Responses, and Anthropic
-Messages, including async streaming. Panels finish before the aggregator starts, so the first streamed token arrives
-during aggregation
-
-## Presets and settings
-
-The dashboard offers two behavior presets:
-
-- **Quality First** sets `on_quorum_failure: fail`. The request fails when fewer than `min_successful_panelists` panel
- calls succeed, preserving the configured quality floor
-- **High Availability** sets `on_quorum_failure: aggregator_only`. When the panel misses quorum, the aggregator receives
- the original request without partial candidates and answers alone
-
-Advanced settings stay limited to the controls that affect one Fusion round:
-
-| Setting | Meaning | Bounds |
-| --- | --- | --- |
-| `min_successful_panelists` | Successful panel responses required before synthesis | 1 to panel size |
-| `panel_timeout_seconds` | Deadline applied to each panel call | More than 0, at most 600 seconds |
-| `max_candidate_chars` | Text copied from each candidate into the synthesis request | 1,000 to 50,000 characters |
-
-The default uses Quality First with a quorum of two, a 120-second panel deadline, and 12,000 characters per candidate.
-LiteLLM runs every configured panel member because this feature optimizes answer quality rather than call cost
-
-## Tools and active work
-
-LiteLLM sends client-defined function schemas to every panel member. A panel can reason about available actions and
-propose a function name and arguments. LiteLLM serializes those proposals as untrusted advice, discards their call IDs,
-and never returns or executes them
-
-The aggregator receives the original tools and has sole authority to emit a tool call. It creates the call and arguments
-after considering the panel. LiteLLM withholds provider-hosted tools such as hosted web search from panel members because
-those tools execute inside the provider. The aggregator retains them
-
-A coding or active-task harness follows this loop:
-
-1. The harness sends its transcript, context, and tool schemas to `fusion/coding`
-2. Panel members propose answers, edits, commands, or tool use in isolation
-3. The aggregator synthesizes one response or tool call
-4. The harness executes the call, records the result, and invokes `fusion/coding` again
-
-Research applications use the same loop. Fusion improves the model decision on each call while the application owns
-browsing, citation collection, retries, approvals, and its completion criteria
-
-## Conversation history and compaction
-
-Every panel member receives the complete message list supplied on that call. The aggregator receives that list plus one
-developer message containing bounded panel candidates. LiteLLM inserts the developer message after leading system and
-developer instructions so the insertion preserves assistant/tool adjacency in the transcript
-
-Only the aggregator output enters client-visible history. Panel outputs last for one Fusion round and are not replayed
-on later turns. The canonical conversation therefore matches the transcript a client would retain for one model, and
-every later panel round sees it. Replaying panel reasoning would multiply context use and create conflicting histories
-
-Fusion does not compact across turns. If the client or harness summarizes or truncates the canonical conversation,
-every panel member and the aggregator see that compacted transcript on the next call. Anthropic Messages context
-management runs before the same Fusion core
-
-## Failure, health, and observability
-
-Panel calls fail independently. A provider error, timeout, streaming response where Fusion expected a complete
-candidate, or empty response counts as one failed panelist. Quality First requires a healthy aggregator and enough
-healthy panel dependencies to meet quorum. High Availability requires a healthy aggregator and still reports each
-panel's health without taking the virtual model down
-
-Each child provider call keeps its LiteLLM logging and spend record. Panel calls include
-`internal_call_origin: fusion_panel`; the aggregator remains the authoritative call for the parent request. A successful
-Fusion round makes one billable call per panelist plus one aggregator call
-
-## Beta boundaries
-
-- Fusion supports `n=1`. LiteLLM rejects multiple returned choices because Fusion must produce one authoritative result
-- A Fusion model cannot serve as a panel member or aggregator for another Fusion model
-- Responses background jobs are unsupported
-- Synchronous Python streaming through `Router.completion` and `Router.responses` is unsupported. Use their async
- counterparts. Proxy streaming uses the async paths
-- Embeddings, image generation, audio, batch jobs, and other non-conversational endpoints bypass Fusion. The feature
- covers conversational model calls and tool loops rather than every LiteLLM API type
-
-These boundaries keep each model call deterministic: one parallel panel round runs before one aggregation, and the
-caller retains task lifecycle control
-
-## Design assumptions
-
-The implementation starts with five testable assumptions:
-
-1. The aggregator should synthesize the panel's work. Its prompt permits combining, correcting, rejecting, or replacing
- candidates and preserving supported minority observations
-2. Aggregating on every model call gives operators one predictable policy. The first beta has no hidden cadence or
- conductor-owned state
-3. The canonical client transcript is the sole durable history. Private panel histories would create divergent agents,
- which belongs in a harness
-4. Function-tool awareness helps coding and active work, while one aggregator retains execution authority
-5. Quality is the primary optimization target. LiteLLM exposes cost and latency as consequences rather than routing
- inputs
-
-Evaluations can vary panel composition, aggregator choice, quorum, and failure preset without changing the API contract
-or the harness under test
diff --git a/litellm/proxy/spend_tracking/budget_reservation.py b/litellm/proxy/spend_tracking/budget_reservation.py
index 91d2ece7a51..c3dd919064b 100644
--- a/litellm/proxy/spend_tracking/budget_reservation.py
+++ b/litellm/proxy/spend_tracking/budget_reservation.py
@@ -1012,7 +1012,7 @@ def estimate_request_max_cost(
input_token_counts: Mapping[str, int] | None = None,
) -> float | None:
estimates = [
- _estimate_request_max_cost_for_model(
+ _estimate_request_model_max_cost(
request_body=request_body,
route=route,
model=model_name,
@@ -1027,6 +1027,64 @@ def estimate_request_max_cost(
return max(cast(list[float], estimates))
+def _estimate_request_model_max_cost(
+ request_body: dict,
+ route: str,
+ model: str,
+ llm_router: Router | None,
+ input_tokens: int | None = None,
+) -> float | None:
+ """Estimate one selectable model, expanding Fusion into every billable child call."""
+ registered_model_name: Final = (
+ llm_router._get_model_from_alias(model=model) or model # pyright: ignore[reportPrivateUsage] # admission must price the routed group
+ if llm_router is not None
+ else model
+ )
+ fusion_router: Final = (
+ llm_router.fusion_routers.get(registered_model_name) if llm_router is not None else None
+ )
+ if fusion_router is None:
+ return _estimate_request_max_cost_for_model(
+ request_body=request_body,
+ route=route,
+ model=model,
+ llm_router=llm_router,
+ input_tokens=input_tokens,
+ )
+
+ panel_estimates: Final = tuple(
+ _estimate_request_max_cost_for_model(
+ request_body=request_body,
+ route=route,
+ model=panel_model,
+ llm_router=llm_router,
+ )
+ for panel_model in fusion_router.config.panel_models
+ )
+ original_aggregator_tokens: Final = _count_input_tokens(
+ request_body=request_body,
+ model=fusion_router.config.aggregator_model,
+ )
+ # Four tokens per bounded character plus fixed protocol headroom safely covers
+ # candidate serialization without materializing a synthetic prompt at admission.
+ candidate_token_ceiling: Final = (
+ 4 * fusion_router.config.max_candidate_chars * len(fusion_router.config.panel_models)
+ ) + 1024
+ aggregator_input_tokens: Final = (
+ original_aggregator_tokens + candidate_token_ceiling if original_aggregator_tokens is not None else None
+ )
+ aggregator_estimate: Final = _estimate_request_max_cost_for_model(
+ request_body=request_body,
+ route=route,
+ model=fusion_router.config.aggregator_model,
+ llm_router=llm_router,
+ input_tokens=aggregator_input_tokens,
+ )
+ child_estimates: Final = (*panel_estimates, aggregator_estimate)
+ known_estimates: Final = tuple(estimate for estimate in child_estimates if estimate is not None)
+ return sum(known_estimates) if known_estimates else None
+
+
def estimate_request_input_cost(
request_body: dict,
route: str,
diff --git a/tests/test_litellm/proxy/test_budget_reservation.py b/tests/test_litellm/proxy/test_budget_reservation.py
index b8fb6170d34..a762929b0e8 100644
--- a/tests/test_litellm/proxy/test_budget_reservation.py
+++ b/tests/test_litellm/proxy/test_budget_reservation.py
@@ -1048,6 +1048,60 @@ async def test_should_reserve_tiered_pricing_cost(spend_counter_state):
await release_budget_reservation(reservation)
+def test_fusion_reservation_sums_panels_and_candidate_inflated_aggregator() -> None:
+ router = Router(
+ model_list=[
+ {
+ "model_name": "panel-a",
+ "litellm_params": {"model": "openai/panel-a", "api_key": "fake"},
+ },
+ {
+ "model_name": "panel-b",
+ "litellm_params": {"model": "openai/panel-b", "api_key": "fake"},
+ },
+ {
+ "model_name": "aggregator",
+ "litellm_params": {"model": "openai/aggregator", "api_key": "fake"},
+ },
+ {
+ "model_name": "fusion/test",
+ "litellm_params": {
+ "model": "fusion_router",
+ "fusion_router_config": {
+ "panel_models": ["panel-a", "panel-b"],
+ "aggregator_model": "aggregator",
+ "max_candidate_chars": 1000,
+ },
+ },
+ },
+ ]
+ )
+ request_body = {
+ "model": "fusion/test",
+ "messages": [{"role": "user", "content": "hello"}],
+ "max_tokens": 10,
+ }
+
+ def child_estimate(*, model: str, input_tokens: int | None = None, **_: object) -> float:
+ if model == "aggregator":
+ assert input_tokens is not None
+ assert input_tokens >= 9000
+ return 3.0
+ return {"panel-a": 1.0, "panel-b": 2.0}[model]
+
+ with patch(
+ "litellm.proxy.spend_tracking.budget_reservation._estimate_request_max_cost_for_model",
+ side_effect=child_estimate,
+ ):
+ estimated = estimate_request_max_cost(
+ request_body=request_body,
+ route="/chat/completions",
+ llm_router=router,
+ )
+
+ assert estimated == pytest.approx(6.0)
+
+
def test_tiered_reservation_is_all_or_nothing_with_output_tier_from_input_length():
"""Dashscope tiered pricing is all-or-nothing: the tier is chosen by the total
input tokens and every token (input and output) is billed at that tier's rate.
@@ -2264,6 +2318,7 @@ async def test_should_reserve_all_budgeted_counters(spend_counter_state):
proxy_logging_obj=proxy_logging_obj,
)
+ assert reservation is not None
assert (
counter_cache.in_memory_cache.get_cache(key="spend:key:key-budget-all") == 0.3
)
diff --git a/tests/test_litellm/test_fusion_router.py b/tests/test_litellm/test_fusion_router.py
index 5ddf7e02591..6cd5ead7e8d 100644
--- a/tests/test_litellm/test_fusion_router.py
+++ b/tests/test_litellm/test_fusion_router.py
@@ -1,7 +1,7 @@
import asyncio
import json
from collections.abc import Mapping
-from typing import Final
+from typing import Final, cast
import pytest
@@ -157,6 +157,8 @@ async def test_panel_gets_only_function_schemas_and_aggregator_owns_tool_call()
aggregator_call = completion.calls[-1]
assert aggregator_call["tools"] == [function_tool, hosted_tool]
assert aggregator_call["tool_choice"] == hosted_tool_choice
+ aggregator_metadata = cast(Mapping[str, object], aggregator_call["litellm_metadata"])
+ assert aggregator_metadata["user_api_key_budget_reservation"] == {"id": "must-not-propagate"}
aggregator_messages = aggregator_call["messages"]
assert isinstance(aggregator_messages, list)
instruction = str(aggregator_messages[0]["content"])
@@ -306,6 +308,14 @@ async def test_router_registers_and_executes_fusion_deployment() -> None:
assert response.choices[0].message.content == "Final"
deployment = router.get_deployment(model_id=router.model_list[-1]["model_info"]["id"])
assert deployment is not None
+
+ router._unregister_fusion_router_for_deployment( # pyright: ignore[reportPrivateUsage] # regression covers registry lifecycle
+ deployment
+ )
+ assert "fusion/test" not in router.fusion_routers
+ router.init_fusion_router_deployment(deployment)
+ assert "fusion/test" in router.fusion_routers
+
router.delete_deployment(id=deployment.model_info.id)
assert "fusion/test" not in router.fusion_routers
@@ -341,8 +351,12 @@ async def test_router_responses_api_bridges_through_the_same_fusion_model() -> N
router = Router(model_list=_router_model_list())
response = await router.aresponses(model="fusion/test", input="Answer")
+ direct_response = await router._fusion_aware_aresponses( # pyright: ignore[reportPrivateUsage] # regression covers the Fusion bridge
+ model="fusion/test", input="Answer"
+ )
assert response.output[0].content[0].text == "Final"
+ assert direct_response.output[0].content[0].text == "Final"
with pytest.raises(litellm.BadRequestError, match="Background Responses"):
await router.aresponses(model="fusion/test", input="Answer", background=True)
@@ -361,17 +375,27 @@ async def test_router_anthropic_messages_bridges_through_the_same_fusion_model()
messages=[{"role": "user", "content": "Answer"}],
max_tokens=256,
)
+ direct_response = await router._fusion_aware_aanthropic_messages( # pyright: ignore[reportPrivateUsage] # regression covers the Fusion bridge
+ model="fusion/test",
+ messages=[{"role": "user", "content": "Answer"}],
+ max_tokens=256,
+ )
assert response["content"][0]["text"] == "Final"
assert alias_response["content"][0]["text"] == "Final"
+ assert direct_response["content"][0]["text"] == "Final"
def test_sync_responses_api_supports_nonstreaming_fusion() -> None:
router = Router(model_list=_router_model_list())
response = router.responses(model="fusion/test", input="Answer")
+ direct_response = router._fusion_aware_responses( # pyright: ignore[reportPrivateUsage] # regression covers the Fusion bridge
+ model="fusion/test", input="Answer"
+ )
assert response.output[0].content[0].text == "Final"
+ assert direct_response.output[0].content[0].text == "Final"
with pytest.raises(litellm.BadRequestError, match="Synchronous Responses streaming"):
router.responses(model="fusion/test", input="Answer", stream=True)
diff --git a/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/FusionModelsPanel.test.tsx b/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/FusionModelsPanel.test.tsx
index f1615b1dae4..f2f36827d41 100644
--- a/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/FusionModelsPanel.test.tsx
+++ b/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/FusionModelsPanel.test.tsx
@@ -62,7 +62,14 @@ const existingDeployment = {
const renderPanel = (createScope: "unscoped-ok" | "team-required" = "unscoped-ok") =>
render(
- ,
+ ,
);
describe("FusionModelsPanel", () => {
diff --git a/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/FusionModelsPanel.tsx b/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/FusionModelsPanel.tsx
index 0b0fb98ba35..720b4d3b1cc 100644
--- a/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/FusionModelsPanel.tsx
+++ b/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/FusionModelsPanel.tsx
@@ -34,6 +34,7 @@ interface FusionModelsPanelProps {
accessToken: string;
userRole: string;
userID: string | null;
+ isViewOnly: boolean;
teams: Team[] | null;
createScope: ModelWriteScope;
}
@@ -293,7 +294,14 @@ function FusionModelDialog({
);
}
-export function FusionModelsPanel({ accessToken, userRole, userID, teams, createScope }: FusionModelsPanelProps) {
+export function FusionModelsPanel({
+ accessToken,
+ userRole,
+ userID,
+ isViewOnly,
+ teams,
+ createScope,
+}: FusionModelsPanelProps) {
const { data: deployments, isLoading } = useFusionRouters();
const availableModels = usePlainModelGroups();
const invalidateFusionRouters = useInvalidateFusionRouters();
@@ -305,7 +313,7 @@ export function FusionModelsPanel({ accessToken, userRole, userID, teams, create
const modelOptions = useMemo(() => Array.from(availableModels).sort(), [availableModels]);
const canModify = (deployment: AutoRouterDeployment) =>
- canModifyModel({ userRole, userID }, teams, {
+ canModifyModel({ userRole, userID, isViewOnly }, teams, {
teamId: deployment.model_info?.team_id,
isDbModel: deployment.model_info?.db_model === true,
});
diff --git a/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/fusionModelConfig.test.ts b/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/fusionModelConfig.test.ts
index 5133782fc12..004d31b77fe 100644
--- a/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/fusionModelConfig.test.ts
+++ b/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/components/FusionModels/fusionModelConfig.test.ts
@@ -55,19 +55,20 @@ describe("Fusion model configuration", () => {
});
it("parses stored configs defensively and supplies stable defaults", () => {
- expect(
- parseFusionConfig({
- panel_models: ["a", "a", "b", 4],
- aggregator_model: "judge",
- on_quorum_failure: "aggregator_only",
- }),
- ).toEqual({
+ const expectedConfig = {
panel_models: ["a", "b"],
aggregator_model: "judge",
min_successful_panelists: 2,
panel_timeout_seconds: 120,
max_candidate_chars: 12000,
on_quorum_failure: "aggregator_only",
- });
+ };
+ expect(
+ parseFusionConfig({
+ panel_models: ["a", "a", "b", 4],
+ aggregator_model: "judge",
+ on_quorum_failure: "aggregator_only",
+ }),
+ ).toEqual(expectedConfig);
});
});
diff --git a/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/panels/FusionModelsTabPanel.tsx b/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/panels/FusionModelsTabPanel.tsx
index 2c7daba6b00..018fe870a33 100644
--- a/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/panels/FusionModelsTabPanel.tsx
+++ b/ui/litellm-dashboard/src/app/(dashboard)/models-and-endpoints/panels/FusionModelsTabPanel.tsx
@@ -9,12 +9,12 @@ import { modelCreationScope } from "@/utils/modelPermissions";
import { FusionModelsPanel } from "../components/FusionModels/FusionModelsPanel";
export default function FusionModelsTabPanel() {
- const { accessToken, userRole, userId: userID } = useAuthorized();
+ const { accessToken, userRole, userId: userID, isViewOnly } = useAuthorized();
const { data: teams } = useTeams();
const { data: uiSettings } = useUISettings();
const isInternalUser = userRole != null && internalUserRoles.includes(userRole);
const scope = modelCreationScope(
- { userRole, userID },
+ { userRole, userID, isViewOnly },
{
teams: teams ?? null,
disabledForInternalUsers: isInternalUser && uiSettings?.values?.disable_model_add_for_internal_users === true,
@@ -26,6 +26,7 @@ export default function FusionModelsTabPanel() {
accessToken={accessToken}
userRole={userRole ?? ""}
userID={userID ?? null}
+ isViewOnly={isViewOnly}
teams={teams ?? null}
createScope={scope}
/>