Merge pull request #26092 from BerriAI/litellm_claude_code_any_advisor_tutorial
Some checks failed
Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (key-generation, tests/proxy_unit_tests/test_key_generate_prisma.py, 30, 0) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 30, 8) (push) Has been cancelled
Unit Tests: Security / security (push) Has been cancelled

docs: add Claude Code any-advisor tutorial and fix streaming logging
This commit is contained in:
Sameer Kankute 2026-04-22 17:30:23 +05:30 committed by GitHub
commit 0f16133adb
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
14 changed files with 1147 additions and 160 deletions

View file

@ -1,6 +1,6 @@
---
slug: advisor-tool-chat-completions
title: "Advisor Tool (SDK + Proxy)"
title: "[Beta] Advisor Tool (SDK + Proxy)"
date: 2026-04-14T19:30:00
authors:
- sameer
@ -14,7 +14,7 @@ hide_table_of_contents: false
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
# Advisor Tool
# Advisor Tool (Beta)
LiteLLM now supports the Anthropic advisor tool across `chat/completions` and `messages` APIs (SDK + proxy).

View file

@ -1,7 +1,7 @@
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
# Advisor Tool
# Advisor Tool (Beta)
LiteLLM supports the Anthropic advisor tool across `chat/completions` and `messages` APIs (SDK + proxy).
@ -728,6 +728,102 @@ tools=[
---
## Remapping the advisor model
Some clients (e.g. Claude Code) hardcode the advisor tool's `model` field — you cannot change what the client sends. If you still want the advisor sub-call to hit a different model (for cost, availability, or routing reasons), remap the advisor model using `model_group_alias` on the router.
When the advisor tool's `model` resolves through `model_group_alias` to a **non-native Anthropic advisor model**, LiteLLM automatically takes over the orchestration loop — even when the executor is direct Anthropic — and routes the advisor sub-call through the router. The client keeps seeing the original alias in every response surface (`iterations[].model`), so the remap stays opaque to the caller.
```yaml showLineNumbers title="config.yaml — remap claude-opus-4-7 advisor to o3"
model_list:
- model_name: o3
litellm_params:
model: openai/o3
api_key: os.environ/OPENAI_API_KEY
- model_name: claude-sonnet
litellm_params:
model: anthropic/claude-sonnet-4-6
api_key: os.environ/ANTHROPIC_API_KEY
router_settings:
model_group_alias:
claude-opus-4-7: o3
```
With the config above, a client request that includes:
```json
{
"type": "advisor_20260301",
"name": "advisor",
"model": "claude-opus-4-7"
}
```
will:
1. Run the executor against Anthropic as usual.
2. When the executor calls the advisor, route the sub-call through the router to `openai/o3` using the `o3` deployment's credentials.
3. Emit `iterations[].model == "claude-opus-4-7"` in the response so the client never sees `o3`.
:::info When does the remap trigger?
Only when the resolved model is **not** a native Anthropic advisor (currently `claude-opus-4-6` and `claude-opus-4-7`). If you alias one native advisor model to another (e.g. `claude-opus-4-7 -> claude-opus-4-6`), Anthropic's server-side advisor still handles the request.
:::
### Claude Code quickstart: use any advisor model
If you are using Claude Code and want to run the advisor on a non-Claude model (for example `openai/o3`, Gemini, Bedrock, etc.), use this pattern:
1. Keep Claude Code's advisor tool unchanged:
```json
{
"type": "advisor_20260301",
"name": "advisor",
"model": "claude-opus-4-7"
}
```
2. Map that model name to your actual advisor deployment in LiteLLM:
```yaml showLineNumbers title="config.yaml — Claude Code advisor alias"
model_list:
- model_name: my-real-advisor
litellm_params:
model: openai/o3
api_key: os.environ/OPENAI_API_KEY
- model_name: claude-sonnet
litellm_params:
model: anthropic/claude-sonnet-4-6
api_key: os.environ/ANTHROPIC_API_KEY
router_settings:
model_group_alias:
claude-opus-4-7: my-real-advisor
```
3. Send requests through `/v1/messages` as usual from Claude Code.
What happens at runtime:
- Claude Code sends `model: "claude-opus-4-7"` in the advisor tool.
- LiteLLM resolves it to `my-real-advisor` and routes the sub-call to `openai/o3`.
- Claude Code still sees `claude-opus-4-7` in response-visible fields (alias stays opaque).
:::tip Troubleshooting (Claude Code + non-native advisor)
- If you see `Invalid value: 'thinking'` from OpenAI, upgrade to a LiteLLM build that includes advisor sub-call message translation for non-Anthropic providers.
- If advisor output appears blank in streamed UI, use a build with `advisor_tool_result` text included in `content_block_start` for fake-streamed advisor responses.
- If spend logs are missing for streamed advisor calls, use a build with deferred logging support for non-`CustomStreamWrapper` anthropic streams.
:::
---
## Additional resources
- [Anthropic Advisor Tool Documentation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool)

View file

@ -0,0 +1,166 @@
# Claude Code: Any Advisor Model via LiteLLM (Beta)
This tutorial shows how to use **Claude Code's advisor tool** with **any model/provider** (OpenAI, Gemini, Bedrock, etc.) by routing through LiteLLM.
Claude Code sends advisor tools with a fixed Anthropic model name (for example `claude-opus-4-7`). LiteLLM can remap that to your real advisor model using `model_group_alias`.
<br />
<iframe width="840" height="500" src="https://www.loom.com/embed/db5bc6156c584e0998c7624821bd5272" frameborder="0" webkitallowfullscreen mozallowfullscreen allowfullscreen></iframe>
## Prerequisites
- [Claude Code](https://docs.anthropic.com/en/docs/claude-code/overview) installed
- LiteLLM proxy installed (`uv tool install 'litellm[proxy]'`)
- API keys for:
- your **executor** model (for example Anthropic Sonnet)
- your **advisor** model (for example OpenAI `o3`)
## Step 1: Create `config.yaml`
Create a LiteLLM config where:
1. your executor model is configured in `model_list`
2. your real advisor model is configured in `model_list`
3. `router_settings.model_group_alias` remaps Claude Code's advisor model name to your real advisor deployment
```yaml showLineNumbers title="config.yaml"
model_list:
# Executor model (the main model Claude Code runs on)
- model_name: claude-sonnet
litellm_params:
model: anthropic/claude-sonnet-4-6
api_key: os.environ/ANTHROPIC_API_KEY
# Real advisor model (can be any provider)
- model_name: my-advisor
litellm_params:
model: openai/o3
api_key: os.environ/OPENAI_API_KEY
router_settings:
model_group_alias:
# Claude Code sends this in advisor_20260301 tool model
claude-opus-4-7: my-advisor
litellm_settings:
drop_params: true
```
Set env vars:
```bash
export ANTHROPIC_API_KEY="your-anthropic-key"
export OPENAI_API_KEY="your-openai-key"
export LITELLM_MASTER_KEY="sk-1234"
```
## Step 2: Start LiteLLM Proxy
```bash showLineNumbers title="Run LiteLLM Proxy"
litellm --config /path/to/config.yaml
```
Expected startup endpoint:
```bash
# RUNNING on http://0.0.0.0:4000
```
## Step 3: Point Claude Code to LiteLLM
Configure Claude Code to call your LiteLLM proxy:
```bash showLineNumbers title="Claude Code environment variables"
export ANTHROPIC_BASE_URL="http://0.0.0.0:4000"
export ANTHROPIC_AUTH_TOKEN="$LITELLM_MASTER_KEY"
export ANTHROPIC_MODEL="claude-sonnet"
```
Then launch Claude Code:
```bash
claude
```
## Step 4: Call advisor from Claude Code
In Claude Code, ask for an advisor run (for example):
```text
Can you call advisor as integration test and confirm it works?
```
Claude Code will send an advisor tool like:
```json
{
"type": "advisor_20260301",
"name": "advisor",
"model": "claude-opus-4-7"
}
```
LiteLLM will remap `claude-opus-4-7` -> `my-advisor` -> `openai/o3` and run the advisor loop.
## Step 5: Verify it is using your advisor model
Check proxy logs for advisor sub-calls:
```bash
rg "advisor_sub_call|litellm.acompletion\\(|openai/o3|my-advisor" proxy_server.log
```
You should see:
- the outer `/v1/messages` call on your executor (`claude-sonnet`)
- advisor sub-call routed to your mapped model (`openai/o3`)
## Internal flow
```mermaid
sequenceDiagram
participant CC as Claude Code
participant LL as LiteLLM Proxy
participant EX as Executor Model
participant AD as Real Advisor Model
CC->>LL: /v1/messages + advisor_20260301 model=claude-opus-4-7
LL->>LL: model_group_alias claude-opus-4-7 -> my-advisor
LL->>EX: Executor call
EX-->>LL: advisor tool_use
LL->>AD: Advisor sub-call (my-advisor)
AD-->>LL: advisor result
LL->>EX: inject advisor result + continue
EX-->>LL: final answer
LL-->>CC: final response (client-facing advisor alias unchanged)
```
## Troubleshooting
### Why is my advisor model not found?
- Ensure alias target (`my-advisor`) exists as a `model_name` in `model_list`
- Confirm with:
```bash
curl http://0.0.0.0:4000/v1/models -H "Authorization: Bearer $LITELLM_MASTER_KEY"
```
### Why do I get `Invalid value: 'thinking'` with a non-Anthropic advisor?
- Upgrade LiteLLM to a version that includes advisor sub-call message translation for non-Anthropic providers
### Why is advisor output blank in streamed UI?
- Upgrade LiteLLM to a version where `advisor_tool_result` includes text in `content_block_start` for fake-stream iterator responses
### Why are spend/log rows missing for streamed advisor calls?
- Upgrade LiteLLM to a version that adds deferred logging support for non-`CustomStreamWrapper` anthropic streams
## Related docs
- [Advisor Tool Reference](/docs/completion/anthropic_advisor_tool)
- [Use Claude Code with Non-Anthropic Models](/docs/tutorials/claude_non_anthropic_models)
- [Forward Client Headers](/docs/proxy/forward_client_headers)

View file

@ -147,6 +147,7 @@ const sidebars = {
"tutorials/claude_code_websearch",
"tutorials/claude_mcp",
"tutorials/claude_non_anthropic_models",
"tutorials/claude_code_any_advisor_model",
"tutorials/claude_code_plugin_marketplace",
"tutorials/claude_code_beta_headers",
]

View file

@ -1478,6 +1478,21 @@ class Logging(LiteLLMLoggingBaseClass):
if cache_hit is True:
return 0.0
# If an orchestrator (e.g. advisor tool loop in /v1/messages) has
# already aggregated cost into self.cost_breakdown before this path
# runs again (typically on the streaming @client wrapper's
# update_response_metadata pass with a FakeAnthropicMessagesStream
# iterator), preserve the breakdown and return the aggregated total.
# Falling through to litellm.response_cost_calculator would recompute
# with zero usage (stream not yet consumed) and call
# _store_cost_breakdown_in_logging_obj, wiping additional_costs.
if (
self.cost_breakdown is not None
and self.cost_breakdown.get("total_cost") is not None
and self.cost_breakdown["total_cost"] > 0
):
return self.cost_breakdown["total_cost"]
if isinstance(result, BaseModel) and hasattr(result, "_hidden_params"):
hidden_params = getattr(result, "_hidden_params", {})
if (

View file

@ -146,29 +146,29 @@ class FakeAnthropicMessagesStreamIterator:
)
elif block_type == "advisor_tool_result":
advisor_content = block_dict.get("content") or {}
advisor_text = ""
if isinstance(advisor_content, dict):
advisor_text = advisor_content.get("text", "") or ""
elif isinstance(advisor_content, str):
advisor_text = advisor_content
# Keep advisor result payload fully populated in content_block_start.
# Anthropic tool_result-like blocks are treated as complete in start
# events (no follow-up delta required), and Claude Code renders the
# advisor panel from this payload.
content_block_start = {
"type": "content_block_start",
"index": index,
"content_block": {
"type": "advisor_tool_result",
"tool_use_id": block_dict.get("tool_use_id"),
"content": {"type": "advisor_result", "text": ""},
"content": {"type": "advisor_result", "text": advisor_text},
},
}
chunks.append(
f"event: content_block_start\ndata: {json.dumps(content_block_start)}\n\n".encode()
)
advisor_content = block_dict.get("content") or {}
advisor_text = advisor_content.get("text", "") if isinstance(advisor_content, dict) else ""
if advisor_text:
content_block_delta = {
"type": "content_block_delta",
"index": index,
"delta": {"type": "advisor_result_delta", "text": advisor_text},
}
chunks.append(
f"event: content_block_delta\ndata: {json.dumps(content_block_delta)}\n\n".encode()
)
content_block_stop = {"type": "content_block_stop", "index": index}
chunks.append(

View file

@ -16,18 +16,25 @@ How it works:
import asyncio
import uuid
from typing import Any, AsyncIterator, Dict, List, Optional, Union, cast
from typing import Any, AsyncIterator, Dict, List, Optional, Union
import litellm
import litellm.constants as _c
from litellm._internal_context import is_internal_call
from litellm._logging import verbose_logger
from litellm.llms.anthropic.common_utils import strip_advisor_blocks_from_messages
from litellm.llms.anthropic.experimental_pass_through.adapters.handler import (
LiteLLMMessagesToCompletionTransformationHandler,
)
from litellm.types.llms.anthropic_messages.anthropic_response import (
AnthropicMessagesResponse,
AnthropicUsageIteration,
)
from litellm.types.llms.openai import AllMessageValues
from litellm.types.llms.anthropic import ANTHROPIC_ADVISOR_TOOL_TYPE
from litellm.utils import (
resolve_proxy_model_alias_to_litellm_model,
supports_native_advisor_tool,
)
ADVISOR_MAX_USES: int = _c.ADVISOR_MAX_USES
ADVISOR_TOOL_DESCRIPTION: str = _c.ADVISOR_TOOL_DESCRIPTION
@ -49,16 +56,24 @@ class AdvisorOrchestrationHandler(MessagesInterceptor):
) -> bool:
if not tools:
return False
has_advisor = any(t.get("type") == ANTHROPIC_ADVISOR_TOOL_TYPE for t in tools)
if not has_advisor:
advisor_tools = [
t for t in tools if t.get("type") == ANTHROPIC_ADVISOR_TOOL_TYPE
]
if not advisor_tools:
return False
# Direct Anthropic /messages: the API handles advisor_20260301 natively;
# do not run the LiteLLM orchestration loop here.
# Direct Anthropic /messages: the API handles advisor_20260301 natively
# *only* when the tool's model resolves to a native Anthropic advisor
# model. When an operator remaps the tool's model via model_group_alias
# to a non-native model (e.g. claude-opus-4-7 -> o3) we must take over
# the loop here so the sub-call is routed through litellm.
if custom_llm_provider == "anthropic":
for advisor_tool in advisor_tools:
if not _advisor_tool_uses_native_anthropic_model(advisor_tool):
return True
return False
return True
async def handle(
async def handle( # noqa: PLR0915
self,
*,
model: str,
@ -82,15 +97,24 @@ class AdvisorOrchestrationHandler(MessagesInterceptor):
raise ValueError(
f"handle() called but no {ANTHROPIC_ADVISOR_TOOL_TYPE} tool found in tools list"
)
advisor_model: str = advisor_tool.get("model") or ""
if not advisor_model:
advisor_model = _resolve_default_advisor_model()
if not advisor_model:
advisor_model_alias: str = advisor_tool.get("model") or ""
if not advisor_model_alias:
advisor_model_alias = _resolve_default_advisor_model()
if not advisor_model_alias:
raise ValueError(
"No advisor model specified. Either:\n"
" 1. Set 'default_advisor_model' in advisor_interception_params in your proxy config YAML, or\n"
" 2. Include a 'model' field in the advisor tool definition."
)
# Resolve the tool's ``model`` (which may be a proxy model_group_alias
# like ``claude-opus-4-7`` pointing at ``o3``) to the actual underlying
# litellm model for the sub-call. Keep the alias separate so every
# client-visible surface (iterations[].model) continues to show the
# original name the caller sent — the remap is opaque to the caller.
resolved_advisor_model: str = (
resolve_proxy_model_alias_to_litellm_model(advisor_model_alias)
or advisor_model_alias
)
_raw_max_uses = advisor_tool.get("max_uses")
max_uses: int = (
ADVISOR_MAX_USES if _raw_max_uses is None else int(_raw_max_uses)
@ -129,10 +153,13 @@ class AdvisorOrchestrationHandler(MessagesInterceptor):
# their inner @client wrappers (aresponses/acompletion) still populate
# ``custom_llm_provider``, ``api_base``, ``model_id`` on this shared
# model_call_details via ``update_environment_variables`` — fields the
# outer anthropic_messages path never sets on its own. We mark the
# sub-calls ``_is_litellm_internal_call=True`` so their @client skips
# emitting a duplicate log row; only the outer call emits a single
# aggregated entry.
# outer anthropic_messages path never sets on its own.
#
# NOTE: ``_is_litellm_internal_call`` in kwargs is not sufficient to
# suppress @client logging; wrapper_async checks the ContextVar
# ``is_internal_call``. Keep the kwarg for compatibility, but also set
# the ContextVar around the orchestration loop so nested sub-calls do
# not emit separate proxy billing rows.
litellm_logging_obj = kwargs.get("litellm_logging_obj", None)
kwargs["_is_litellm_internal_call"] = True
iteration = 0
@ -141,110 +168,138 @@ class AdvisorOrchestrationHandler(MessagesInterceptor):
advisor_first_call_cost: float = 0.0
advisor_subcall_cost: float = 0.0
while True:
# --- Executor call (always non-streaming) ---
executor_response: AnthropicMessagesResponse = await _call_messages_handler(
model=model,
messages=current_messages,
tools=executor_tools,
stream=False,
max_tokens=max_tokens,
custom_llm_provider=custom_llm_provider,
metadata={
**metadata_base,
"advisor_sub_call": False,
"parent_request_id": parent_request_id,
},
**kwargs,
)
executor_cost = _get_response_cost(executor_response, model=model)
iterations.append(
_build_iteration_entry(
response=executor_response, iteration_type="message"
)
)
advisor_use_block = _find_advisor_tool_use(executor_response)
if advisor_use_block is None:
# No more advisor calls — this is the final response.
# Inject advisor_tool_result blocks to match Anthropic native format.
_inject_advisor_blocks_into_response(
executor_response, advisor_interactions
)
total_cost = (
advisor_first_call_cost + advisor_subcall_cost + executor_cost
)
_finalize_orchestrated_response(
response=executor_response,
iterations=iterations,
total_cost=total_cost,
final_executor_cost=executor_cost,
advisor_first_call_cost=advisor_first_call_cost,
advisor_subcall_cost=advisor_subcall_cost,
litellm_logging_obj=litellm_logging_obj,
)
if stream:
return FakeAnthropicMessagesStreamIterator(executor_response)
return executor_response
# Executor response triggered another advisor call → count it as a
# "first/intermediate" executor turn. Only the terminating turn is
# treated as the base response.
advisor_first_call_cost += executor_cost
iteration += 1
if iteration > max_uses:
raise AdvisorMaxIterationsError(
f"Advisor orchestration loop exceeded max_uses={max_uses}. "
"Increase max_uses in the advisor tool definition or cap the request."
_prev_internal = is_internal_call.get()
is_internal_call.set(True)
try:
while True:
# --- Executor call (always non-streaming) ---
executor_response: AnthropicMessagesResponse = await _call_messages_handler(
model=model,
messages=current_messages,
tools=executor_tools,
stream=False,
max_tokens=max_tokens,
custom_llm_provider=custom_llm_provider,
metadata={
**metadata_base,
"advisor_sub_call": False,
"parent_request_id": parent_request_id,
},
**kwargs,
)
# --- Build advisor context ---
advisor_messages = _build_advisor_context(
current_messages, executor_response, advisor_use_block
)
# --- Advisor sub-call (always non-streaming, no tools) ---
advisor_response: AnthropicMessagesResponse = await _call_advisor_with_router(
model=advisor_model,
messages=advisor_messages,
max_tokens=max_tokens,
metadata={
**metadata_base,
"advisor_sub_call": True,
"parent_request_id": parent_request_id,
},
api_key=advisor_api_key,
api_base=advisor_api_base,
)
advisor_call_cost = _get_response_cost(advisor_response, model=advisor_model)
advisor_subcall_cost += advisor_call_cost
iterations.append(
_build_iteration_entry(
response=advisor_response,
iteration_type="advisor_message",
model=advisor_model,
executor_cost = _get_response_cost(executor_response, model=model)
iterations.append(
_build_iteration_entry(
response=executor_response, iteration_type="message"
)
)
)
advisor_text = _extract_response_text(advisor_response)
advisor_use_block = _find_advisor_tool_use(executor_response)
# Record the interaction for later injection into the final response.
advisor_interactions.append({
"tool_use_id": advisor_use_block.get("id", f"srvtoolu_{uuid.uuid4().hex[:24]}"),
"advisor_text": advisor_text,
})
if advisor_use_block is None:
# No more advisor calls — this is the final response.
# Inject advisor_tool_result blocks to match Anthropic native format.
_inject_advisor_blocks_into_response(
executor_response, advisor_interactions
)
total_cost = (
advisor_first_call_cost + advisor_subcall_cost + executor_cost
)
_finalize_orchestrated_response(
response=executor_response,
iterations=iterations,
total_cost=total_cost,
final_executor_cost=executor_cost,
advisor_first_call_cost=advisor_first_call_cost,
advisor_subcall_cost=advisor_subcall_cost,
litellm_logging_obj=litellm_logging_obj,
)
if stream:
# The outer ``@client`` async wrapper skips
# ``_client_async_logging_helper`` for streaming
# requests — it assumes a ``CustomStreamWrapper`` will
# fire logging on iteration. ``FakeAnthropicMessagesStreamIterator``
# is a plain iterator over pre-built SSE bytes and
# does not know about the logging obj, so nothing fires
# the proxy log row. We've already aggregated the full
# response into ``executor_response`` (same dict shape
# the non-streaming path uses for logging), so fire
# the success handler ourselves with that dict before
# wrapping — this mirrors the non-streaming flow and
# avoids double-logging (the @client path is skipped).
_fire_async_success_logging(
litellm_logging_obj=litellm_logging_obj,
result=executor_response,
)
return FakeAnthropicMessagesStreamIterator(executor_response)
return executor_response
# --- Inject advisor result and continue loop ---
current_messages = _inject_advisor_turn(
current_messages,
executor_response,
advisor_use_block,
advisor_text,
)
# Executor response triggered another advisor call → count it as a
# "first/intermediate" executor turn. Only the terminating turn is
# treated as the base response.
advisor_first_call_cost += executor_cost
iteration += 1
if iteration > max_uses:
raise AdvisorMaxIterationsError(
f"Advisor orchestration loop exceeded max_uses={max_uses}. "
"Increase max_uses in the advisor tool definition or cap the request."
)
# --- Build advisor context ---
advisor_messages = _build_advisor_context(
current_messages, executor_response, advisor_use_block
)
# --- Advisor sub-call (always non-streaming, no tools) ---
# Use the resolved model so router routing / cost lookup hit the
# real underlying deployment; the alias is kept only for the
# client-visible iteration entry below.
advisor_response: AnthropicMessagesResponse = await _call_advisor_with_router(
model=resolved_advisor_model,
messages=advisor_messages,
max_tokens=max_tokens,
metadata={
**metadata_base,
"advisor_sub_call": True,
"parent_request_id": parent_request_id,
},
api_key=advisor_api_key,
api_base=advisor_api_base,
)
advisor_call_cost = _get_response_cost(
advisor_response, model=resolved_advisor_model
)
advisor_subcall_cost += advisor_call_cost
iterations.append(
_build_iteration_entry(
response=advisor_response,
iteration_type="advisor_message",
model=advisor_model_alias,
)
)
advisor_text = _extract_response_text(advisor_response)
# Record the interaction for later injection into the final response.
advisor_interactions.append({
"tool_use_id": advisor_use_block.get(
"id", f"srvtoolu_{uuid.uuid4().hex[:24]}"
),
"advisor_text": advisor_text,
})
# --- Inject advisor result and continue loop ---
current_messages = _inject_advisor_turn(
current_messages,
executor_response,
advisor_use_block,
advisor_text,
)
finally:
is_internal_call.set(_prev_internal)
# ---------------------------------------------------------------------------
@ -260,6 +315,31 @@ def _resolve_default_advisor_model() -> str:
return params.get("default_advisor_model", "") or ""
def _advisor_tool_uses_native_anthropic_model(advisor_tool: Dict) -> bool:
"""
Return True iff the advisor tool's ``model`` (after proxy alias resolution)
is a native Anthropic advisor model.
Used by :class:`AdvisorOrchestrationHandler.can_handle` to decide whether
to let Anthropic's server-side advisor handle the tool or to intercept it
and route the sub-call through LiteLLM.
"""
advisor_model = advisor_tool.get("model") or _resolve_default_advisor_model()
if not advisor_model:
# No model specified — let native Anthropic handle it (or fail there
# with its own error). This path should not be hit in practice because
# advisor_interception_params enforces a default upstream.
return True
resolved_model = (
resolve_proxy_model_alias_to_litellm_model(advisor_model) or advisor_model
)
if resolved_model.startswith("anthropic/"):
resolved_model = resolved_model.split("/", 1)[1]
return supports_native_advisor_tool(
model=resolved_model, custom_llm_provider="anthropic"
)
_SYNTHETIC_ADVISOR_TOOL_NAME = "consult_advisor"
@ -445,6 +525,55 @@ def _finalize_orchestrated_response(
)
def _fire_async_success_logging(
litellm_logging_obj: Any,
result: Any,
) -> None:
"""
Manually enqueue the ``async_success_handler`` for a streaming advisor
response.
The outer ``@client`` async wrapper only calls
``_client_async_logging_helper`` for non-streaming results; streaming
results are expected to log from inside a ``CustomStreamWrapper``. Our
synthetic :class:`FakeAnthropicMessagesStreamIterator` has no logging
hook, so without this helper the proxy UI would never get a row for
streaming advisor calls. We already built the aggregated response dict
(same shape the non-streaming path logs from), so we can fire logging
exactly once here with the same arguments the non-streaming path uses.
"""
if litellm_logging_obj is None:
return
try:
import datetime as _dt
from litellm.litellm_core_utils.logging_worker import GLOBAL_LOGGING_WORKER
start_time = getattr(litellm_logging_obj, "start_time", None) or _dt.datetime.now()
end_time = _dt.datetime.now()
GLOBAL_LOGGING_WORKER.ensure_initialized_and_enqueue(
async_coroutine=litellm_logging_obj.async_success_handler(
result=result, start_time=start_time, end_time=end_time
)
)
try:
litellm_logging_obj.handle_sync_success_callbacks_for_async_calls(
result=result,
start_time=start_time,
end_time=end_time,
)
except Exception as sync_cb_error:
verbose_logger.debug(
"AdvisorOrchestration: sync success callbacks failed: %s",
str(sync_cb_error),
)
except Exception as logging_error:
verbose_logger.debug(
"AdvisorOrchestration: failed to fire async success logging: %s",
str(logging_error),
)
def _find_advisor_tool_use(response: Any) -> Optional[Dict]:
"""Return the first tool_use block whose name matches our synthetic advisor."""
content = response.get("content") if isinstance(response, dict) else []
@ -609,6 +738,12 @@ def _build_advisor_context(
tool_use blocks are excluded because Anthropic requires tool_use to be
immediately followed by tool_result not the advisor question.
Messages stay in Anthropic ``/v1/messages`` shape here; ``_call_advisor_with_router``
runs the same ``LiteLLMMessagesToCompletionTransformationHandler`` path used when
a client calls the messages endpoint with a non-Anthropic model, so provider
translation (including interleaved ``thinking`` blocks) matches the rest of
the stack.
"""
question = (advisor_use_block.get("input") or {}).get("question") or (
"Please provide guidance on the current task."
@ -774,18 +909,26 @@ async def _call_advisor_with_router(
if api_base is not None:
kwargs["api_base"] = api_base
openai_messages: List[AllMessageValues] = cast(List[AllMessageValues], messages)
# Same translation path as ``/v1/messages`` → non-Anthropic model: Anthropic
# request shape → Chat Completions kwargs for the target provider.
(
completion_kwargs,
_tool_name_mapping,
) = LiteLLMMessagesToCompletionTransformationHandler._prepare_completion_kwargs(
max_tokens=max_tokens,
messages=messages,
model=model,
metadata=metadata,
stream=False,
extra_kwargs=kwargs,
)
# Inner advisor call is always a plain completion (no tools).
completion_kwargs["tools"] = None
openai_response = None
if llm_router is not None:
try:
openai_response = await llm_router.acompletion(
model=model,
messages=openai_messages,
tools=None,
max_tokens=max_tokens,
**kwargs,
)
openai_response = await llm_router.acompletion(**completion_kwargs)
except Exception:
verbose_logger.debug(
"AdvisorOrchestration: Router call for advisor model '%s' failed, "
@ -794,12 +937,6 @@ async def _call_advisor_with_router(
)
if openai_response is None:
openai_response = await _litellm.acompletion(
model=model,
messages=openai_messages,
tools=None,
max_tokens=max_tokens,
**kwargs,
)
openai_response = await _litellm.acompletion(**completion_kwargs)
return _openai_response_to_anthropic_dict(openai_response)

View file

@ -17,7 +17,10 @@ from litellm.types.llms.anthropic_messages.anthropic_response import (
)
from litellm.types.llms.anthropic_tool_search import get_tool_search_beta_header
from litellm.types.router import GenericLiteLLMParams
from litellm.utils import resolve_proxy_model_alias_to_litellm_model
from litellm.utils import (
resolve_proxy_model_alias_to_litellm_model,
supports_native_advisor_tool,
)
from ...common_utils import (
AnthropicError,
@ -36,6 +39,14 @@ def _normalize_anthropic_advisor_tool_models(tools: List[Dict]) -> List[Dict]:
Anthropic expects advisor tool model values like ``claude-opus-4-6``.
Proxy alias names (e.g. ``claude_opus``) and provider-prefixed values
(e.g. ``anthropic/claude-opus-4-6``) are converted.
Defensive guard: if the alias resolves to a model Anthropic's native
advisor tool does not support (e.g. ``claude-opus-4-7`` -> ``o3`` via
``model_group_alias``), leave the original alias string in place rather
than forwarding the unsupported model to Anthropic. In that case
``AdvisorOrchestrationHandler`` is responsible for intercepting the
request and running the loop through litellm; this branch only runs if
the interceptor was somehow bypassed.
"""
normalized_tools: List[Dict] = []
for tool in tools:
@ -49,11 +60,21 @@ def _normalize_anthropic_advisor_tool_models(tools: List[Dict]) -> List[Dict]:
updated_tool = dict(tool)
advisor_model = updated_tool.get("model")
if isinstance(advisor_model, str) and advisor_model.strip():
resolved = resolve_proxy_model_alias_to_litellm_model(advisor_model.strip())
canonical_model = resolved or advisor_model.strip()
original_model = advisor_model.strip()
resolved = resolve_proxy_model_alias_to_litellm_model(original_model)
canonical_model = resolved or original_model
if canonical_model.startswith("anthropic/"):
canonical_model = canonical_model.split("/", 1)[1]
updated_tool["model"] = canonical_model
# Only substitute the resolved/canonical value if Anthropic
# natively supports it as an advisor model. Otherwise keep the
# caller's original alias so we never leak a non-Anthropic model
# name (e.g. ``o3``) into the Anthropic request body.
if supports_native_advisor_tool(
model=canonical_model, custom_llm_provider="anthropic"
):
updated_tool["model"] = canonical_model
else:
updated_tool["model"] = original_model
normalized_tools.append(updated_tool)
return normalized_tools

View file

@ -9094,6 +9094,7 @@
"supports_assistant_prefill": false,
"supports_computer_use": true,
"supports_function_calling": true,
"supports_native_advisor_tool": true,
"supports_pdf_input": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
@ -9126,6 +9127,7 @@
"supports_assistant_prefill": false,
"supports_computer_use": true,
"supports_function_calling": true,
"supports_native_advisor_tool": true,
"supports_pdf_input": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
@ -9158,6 +9160,7 @@
"supports_assistant_prefill": false,
"supports_computer_use": true,
"supports_function_calling": true,
"supports_native_advisor_tool": true,
"supports_pdf_input": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
@ -9190,6 +9193,7 @@
"supports_assistant_prefill": false,
"supports_computer_use": true,
"supports_function_calling": true,
"supports_native_advisor_tool": true,
"supports_pdf_input": true,
"supports_prompt_caching": true,
"supports_reasoning": true,

View file

@ -1171,6 +1171,37 @@ class ProxyBaseLLMRequestProcessing:
logging_obj._on_deferred_stream_complete = _on_deferred_stream_complete # type: ignore[union-attr]
# Fallback for streaming responses that are NOT CustomStreamWrapper
# (e.g. FakeAnthropicMessagesStreamIterator used by advisor
# orchestration). These responses bypass CSW's internal deferred
# callback wiring, so without this fallback the outer proxy log row
# is never emitted.
if (
self._is_streaming_response(response)
and not isinstance(response, CustomStreamWrapper)
and getattr(logging_obj, "_on_deferred_stream_complete", None)
is None
):
_captured_data = self.data
_captured_user_api_key_dict = user_api_key_dict
_captured_logging_obj = logging_obj
async def _on_deferred_stream_complete_non_csw(
assembled_response, cache_hit
):
await ProxyBaseLLMRequestProcessing._run_deferred_stream_guardrails(
captured_data=_captured_data,
captured_user_api_key_dict=_captured_user_api_key_dict,
captured_logging_obj=_captured_logging_obj,
assembled_response=assembled_response,
cache_hit=cache_hit,
)
logging_obj._on_deferred_stream_complete = _on_deferred_stream_complete_non_csw # type: ignore[union-attr]
# _fire_deferred_stream_logging only fires when args are present.
# Non-CSW iterators never set these, so seed default args here.
logging_obj._deferred_stream_complete_args = (None, None) # type: ignore[union-attr]
if route_type == "allm_passthrough_route":
# Check if response is an async generator
if self._is_streaming_response(response):

View file

@ -9158,6 +9158,7 @@
"supports_assistant_prefill": false,
"supports_computer_use": true,
"supports_function_calling": true,
"supports_native_advisor_tool": true,
"supports_pdf_input": true,
"supports_prompt_caching": true,
"supports_reasoning": true,
@ -9190,6 +9191,7 @@
"supports_assistant_prefill": false,
"supports_computer_use": true,
"supports_function_calling": true,
"supports_native_advisor_tool": true,
"supports_pdf_input": true,
"supports_prompt_caching": true,
"supports_reasoning": true,

View file

@ -0,0 +1,413 @@
"""
Tests for advisor tool model aliasing on the /v1/messages path.
Scenario: an operator remaps the advisor tool's model via
``model_group_alias`` (e.g. ``claude-opus-4-7 -> o3``) because the client
(Claude Code) hardcodes the advisor ``model`` field to ``claude-opus-4-7``.
Required behaviour:
* ``AdvisorOrchestrationHandler.can_handle`` intercepts the request when
the executor is direct Anthropic but the advisor tool resolves to a
non-native advisor model.
* ``handle()`` dispatches the advisor sub-call with the *resolved* model,
but every client-visible surface (``iterations[].model``) keeps the
original alias so the remap is opaque to the caller.
* ``_normalize_anthropic_advisor_tool_models`` never forwards a
non-Anthropic model to the native API it leaves the alias untouched
if the resolved model is not natively supported.
"""
from typing import Dict
from unittest.mock import AsyncMock, patch
import pytest
ADVISOR_TOOL_ALIAS = {
"type": "advisor_20260301",
"name": "advisor",
"model": "claude-opus-4-7",
}
MESSAGES = [
{"role": "user", "content": "Write a Python function that checks if a number is prime."}
]
def _make_text_response(text: str, model: str = "openai/o3") -> Dict:
return {
"id": "msg_test",
"type": "message",
"role": "assistant",
"model": model,
"content": [{"type": "text", "text": text}],
"stop_reason": "end_turn",
"usage": {"input_tokens": 10, "output_tokens": 20},
}
def _make_advisor_tool_use_response(
question: str = "How should I approach this?",
tool_id: str = "toolu_advisor_01",
model: str = "claude-opus-4-7",
) -> Dict:
return {
"id": "msg_test",
"type": "message",
"role": "assistant",
"model": model,
"content": [
{
"type": "tool_use",
"id": tool_id,
"name": "consult_advisor",
"input": {"question": question},
}
],
"stop_reason": "tool_use",
"usage": {"input_tokens": 10, "output_tokens": 15},
}
# ---------------------------------------------------------------------------
# 1. can_handle: alias -> non-native forces interception even on anthropic
# ---------------------------------------------------------------------------
def test_can_handle_alias_to_non_native_intercepts_on_anthropic():
"""
When the advisor tool's model (``claude-opus-4-7``) aliases to a
non-Anthropic model (``o3``), the handler must intercept even though the
executor provider is direct Anthropic Anthropic's native advisor tool
can't run ``o3`` for us.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor import (
AdvisorOrchestrationHandler,
)
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.resolve_proxy_model_alias_to_litellm_model",
return_value="openai/o3",
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.supports_native_advisor_tool",
return_value=False,
):
h = AdvisorOrchestrationHandler()
assert h.can_handle([ADVISOR_TOOL_ALIAS], "anthropic") is True
def test_can_handle_alias_to_native_still_defers_to_anthropic():
"""
When the advisor tool's model aliases to a still-native Anthropic model
(e.g. someone maps ``claude-opus-4-7 -> claude-opus-4-6``), the native
Anthropic server-side advisor can still handle it we must not
intercept.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor import (
AdvisorOrchestrationHandler,
)
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.resolve_proxy_model_alias_to_litellm_model",
return_value="anthropic/claude-opus-4-6",
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.supports_native_advisor_tool",
return_value=True,
):
h = AdvisorOrchestrationHandler()
assert h.can_handle([ADVISOR_TOOL_ALIAS], "anthropic") is False
def test_can_handle_non_anthropic_executor_always_intercepts():
"""
Non-Anthropic executors always need orchestration regardless of the
advisor tool's resolved model — no behaviour change from before.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor import (
AdvisorOrchestrationHandler,
)
h = AdvisorOrchestrationHandler()
assert h.can_handle([ADVISOR_TOOL_ALIAS], "openai") is True
assert h.can_handle([ADVISOR_TOOL_ALIAS], "bedrock") is True
# ---------------------------------------------------------------------------
# 2. handle(): alias vs resolved separation
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_handle_uses_resolved_model_for_subcall_and_alias_for_iterations():
"""
The advisor sub-call receives the *resolved* model (``openai/o3``) so
routing and cost lookup hit the real deployment, while every
client-visible ``iterations[].model`` entry of type ``advisor_message``
keeps the original alias (``claude-opus-4-7``). The alias must never
leak to the sub-call, and the resolved name must never leak to the
response.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor import (
AdvisorOrchestrationHandler,
)
advisor_tool_use_resp = _make_advisor_tool_use_response(
question="What algorithm should I use?",
tool_id="toolu_01",
)
advisor_advice_resp = _make_text_response(
"Use a sieve for large n, trial division for small n.",
model="openai/o3",
)
final_resp = _make_text_response(
"def is_prime(n): ...",
model="claude-opus-4-7",
)
executor_call_count = 0
async def mock_messages(model, messages, tools, stream, max_tokens, **kwargs):
nonlocal executor_call_count
executor_call_count += 1
if executor_call_count == 1:
return advisor_tool_use_resp
return final_resp
advisor_mock = AsyncMock(return_value=advisor_advice_resp)
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor._call_messages_handler",
side_effect=mock_messages,
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor._call_advisor_with_router",
advisor_mock,
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.resolve_proxy_model_alias_to_litellm_model",
return_value="openai/o3",
):
h = AdvisorOrchestrationHandler()
result = await h.handle(
model="claude-opus-4-6",
messages=MESSAGES,
tools=[ADVISOR_TOOL_ALIAS],
stream=False,
max_tokens=512,
custom_llm_provider="anthropic",
)
assert advisor_mock.await_count == 1
advisor_call_kwargs = advisor_mock.await_args.kwargs
assert advisor_call_kwargs["model"] == "openai/o3", (
"Advisor sub-call must use the resolved router model, not the alias"
)
usage = result.get("usage", {})
iterations = usage.get("iterations", [])
advisor_iterations = [
it for it in iterations if it.get("type") == "advisor_message"
]
assert len(advisor_iterations) == 1
assert advisor_iterations[0]["model"] == "claude-opus-4-7", (
"iterations[].model must preserve the client-facing alias"
)
# Resolved model must never appear anywhere in the iterations surface.
for it in iterations:
assert it.get("model") != "openai/o3"
@pytest.mark.asyncio
async def test_handle_without_alias_still_works():
"""
When ``resolve_proxy_model_alias_to_litellm_model`` returns ``""`` (no
alias configured), ``handle()`` must fall back to using the tool's
original model for both the sub-call and the iteration entry. Nothing
regresses for users who don't configure an alias.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor import (
AdvisorOrchestrationHandler,
)
advisor_tool = {
"type": "advisor_20260301",
"name": "advisor",
"model": "openai/gpt-4o-mini",
}
advisor_tool_use_resp = _make_advisor_tool_use_response()
advisor_advice_resp = _make_text_response("advice", model="openai/gpt-4o-mini")
final_resp = _make_text_response("final")
executor_call_count = 0
async def mock_messages(model, messages, tools, stream, max_tokens, **kwargs):
nonlocal executor_call_count
executor_call_count += 1
return advisor_tool_use_resp if executor_call_count == 1 else final_resp
advisor_mock = AsyncMock(return_value=advisor_advice_resp)
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor._call_messages_handler",
side_effect=mock_messages,
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor._call_advisor_with_router",
advisor_mock,
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.resolve_proxy_model_alias_to_litellm_model",
return_value="",
):
h = AdvisorOrchestrationHandler()
result = await h.handle(
model="openai/gpt-4o-mini",
messages=MESSAGES,
tools=[advisor_tool],
stream=False,
max_tokens=512,
custom_llm_provider="openai",
)
assert advisor_mock.await_args.kwargs["model"] == "openai/gpt-4o-mini"
advisor_iterations = [
it
for it in result["usage"]["iterations"]
if it.get("type") == "advisor_message"
]
assert advisor_iterations[0]["model"] == "openai/gpt-4o-mini"
# ---------------------------------------------------------------------------
# 3. _normalize_anthropic_advisor_tool_models defensive guard
# ---------------------------------------------------------------------------
def test_normalize_leaves_alias_when_resolved_model_is_non_native():
"""
Defensive guard: if the alias resolves to a non-Anthropic advisor model,
the normalizer must leave the caller's original alias in place rather
than substituting the unsupported model into the Anthropic request body.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.transformation import (
_normalize_anthropic_advisor_tool_models,
)
tools = [
{
"type": "advisor_20260301",
"name": "advisor",
"model": "claude-opus-4-7",
}
]
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.transformation.resolve_proxy_model_alias_to_litellm_model",
return_value="openai/o3",
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.transformation.supports_native_advisor_tool",
return_value=False,
):
normalized = _normalize_anthropic_advisor_tool_models(tools)
assert normalized[0]["model"] == "claude-opus-4-7", (
"Normalizer must not push the non-native resolved model to Anthropic"
)
def test_normalize_strips_anthropic_prefix_when_resolved_model_is_native():
"""
Regression: for the classic path (alias resolves to a native Anthropic
model), the normalizer still strips the ``anthropic/`` prefix so the
Anthropic API receives a bare model name.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.transformation import (
_normalize_anthropic_advisor_tool_models,
)
tools = [
{
"type": "advisor_20260301",
"name": "advisor",
"model": "claude_opus",
}
]
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.transformation.resolve_proxy_model_alias_to_litellm_model",
return_value="anthropic/claude-opus-4-6",
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.transformation.supports_native_advisor_tool",
return_value=True,
):
normalized = _normalize_anthropic_advisor_tool_models(tools)
assert normalized[0]["model"] == "claude-opus-4-6"
# ---------------------------------------------------------------------------
# 4. Advisor sub-call uses the same /v1/messages → completion translation path
# ---------------------------------------------------------------------------
def test_prepare_completion_kwargs_moves_thinking_out_of_content():
"""
Advisor sub-calls must use ``LiteLLMMessagesToCompletionTransformationHandler``
(same as non-Anthropic ``/v1/messages``), so interleaved ``thinking`` blocks
become OpenAI-shaped messages never raw ``content[].type == "thinking"``.
"""
from litellm.llms.anthropic.experimental_pass_through.adapters.handler import (
LiteLLMMessagesToCompletionTransformationHandler,
)
messages = [
{"role": "user", "content": [{"type": "text", "text": "hi"}]},
{
"role": "assistant",
"content": [
{"type": "thinking", "thinking": "secret reasoning", "signature": "s"},
{"type": "redacted_thinking", "data": "redacted"},
{"type": "text", "text": "hello"},
],
},
]
completion_kwargs, _ = (
LiteLLMMessagesToCompletionTransformationHandler._prepare_completion_kwargs(
max_tokens=100,
messages=messages,
model="openai/gpt-5-nano",
stream=False,
)
)
for msg in completion_kwargs["messages"]:
content = msg.get("content")
if isinstance(content, list):
for part in content:
if isinstance(part, dict):
assert part.get("type") != "thinking"
assert part.get("type") != "redacted_thinking"
def test_build_advisor_context_preserves_string_content_and_plain_messages():
"""
Messages with plain string content or only supported block types must be
passed through unchanged.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor import (
_build_advisor_context,
)
messages = [
{"role": "user", "content": "plain string user"},
{"role": "assistant", "content": [{"type": "text", "text": "plain reply"}]},
]
executor_response = {"content": [{"type": "text", "text": "draft"}]}
advisor_use_block = {
"type": "tool_use",
"name": "advisor",
"input": {"question": "advise"},
}
result = _build_advisor_context(messages, executor_response, advisor_use_block)
assert result[0] == {"role": "user", "content": "plain string user"}
assert result[1] == {
"role": "assistant",
"content": [{"type": "text", "text": "plain reply"}],
}

View file

@ -74,16 +74,24 @@ def test_can_handle_edge_cases():
)
h = AdvisorOrchestrationHandler()
assert h.can_handle([ADVISOR_TOOL], "openai")
assert h.can_handle([ADVISOR_TOOL], "bedrock")
assert h.can_handle([ADVISOR_TOOL], "gemini")
assert not h.can_handle([ADVISOR_TOOL], "anthropic")
assert not h.can_handle([], "openai")
assert not h.can_handle(None, "openai")
assert not h.can_handle([{"type": "function", "name": "bash"}], "openai")
# provider=None: unknown → should intercept (treat as non-native)
assert h.can_handle([ADVISOR_TOOL], None)
# Ensure this edge-case test is deterministic regardless of any proxy-level
# model_group_alias configured by other tests.
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.resolve_proxy_model_alias_to_litellm_model",
return_value="",
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.supports_native_advisor_tool",
return_value=True,
):
assert h.can_handle([ADVISOR_TOOL], "openai")
assert h.can_handle([ADVISOR_TOOL], "bedrock")
assert h.can_handle([ADVISOR_TOOL], "gemini")
assert not h.can_handle([ADVISOR_TOOL], "anthropic")
assert not h.can_handle([], "openai")
assert not h.can_handle(None, "openai")
assert not h.can_handle([{"type": "function", "name": "bash"}], "openai")
# provider=None: unknown → should intercept (treat as non-native)
assert h.can_handle([ADVISOR_TOOL], None)
# ---------------------------------------------------------------------------
@ -102,9 +110,16 @@ async def test_anthropic_native_interceptor_skipped():
)
h = AdvisorOrchestrationHandler()
assert not h.can_handle(
[ADVISOR_TOOL], "anthropic"
), "Interceptor must NOT trigger for anthropic provider"
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.resolve_proxy_model_alias_to_litellm_model",
return_value="",
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor.supports_native_advisor_tool",
return_value=True,
):
assert not h.can_handle(
[ADVISOR_TOOL], "anthropic"
), "Interceptor must NOT trigger for anthropic provider"
# ---------------------------------------------------------------------------
@ -300,6 +315,61 @@ async def test_loop_streaming_wraps_response():
assert "message_start" in first
@pytest.mark.asyncio
async def test_loop_streaming_advisor_block_start_contains_text():
"""
Regression: when advisor orchestration is streamed via FakeAnthropicMessagesStreamIterator,
``advisor_tool_result`` must carry the full advisor text in content_block_start.
Claude Code renders the advisor panel from that payload.
"""
from litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor import (
AdvisorOrchestrationHandler,
)
executor_first = _make_advisor_tool_use_response(
question="Please confirm integration status.", tool_id="toolu_advisor_123"
)
advisor_response = _make_text_response("Integration test: working correctly.")
executor_final = _make_text_response("All set.")
with patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor._call_messages_handler",
new_callable=AsyncMock,
side_effect=[executor_first, executor_final],
), patch(
"litellm.llms.anthropic.experimental_pass_through.messages.interceptors.advisor._call_advisor_with_router",
new_callable=AsyncMock,
return_value=advisor_response,
):
h = AdvisorOrchestrationHandler()
stream_iter = await h.handle(
model="openai/gpt-4o-mini",
messages=MESSAGES,
tools=[ADVISOR_TOOL],
stream=True,
max_tokens=512,
custom_llm_provider="openai",
)
chunks = []
async for chunk in stream_iter:
chunks.append(chunk.decode() if isinstance(chunk, bytes) else str(chunk))
advisor_start_events = [
c
for c in chunks
if '"type": "content_block_start"' in c
and '"type": "advisor_tool_result"' in c
]
assert advisor_start_events, "Expected advisor_tool_result content_block_start event"
assert (
'"text": "Integration test: working correctly."' in advisor_start_events[0]
)
# advisor_tool_result should be complete in content_block_start (no extra delta needed)
assert not any('"type": "advisor_result_delta"' in c for c in chunks)
# ---------------------------------------------------------------------------
# 7. Multi-turn: prior advisor blocks replaced with text in history
# ---------------------------------------------------------------------------

View file

@ -0,0 +1,31 @@
import litellm
from litellm.router_utils.provider_account_fallback_errors import (
is_provider_account_fallback_eligible_error,
)
def test_credit_balance_message_detected():
err = litellm.BadRequestError(
message="AnthropicException - Your credit balance is too low to access the Anthropic API. Please go to Plans & Billing to upgrade or purchase credits.",
model="claude-opus-4-6",
llm_provider="anthropic",
)
assert is_provider_account_fallback_eligible_error(err) is True
def test_openai_billing_hard_limit_detected():
err = litellm.BadRequestError(
message='{"error":{"code":"billing_hard_limit_reached","message":"Billing"}}',
model="gpt-4",
llm_provider="openai",
)
assert is_provider_account_fallback_eligible_error(err) is True
def test_generic_validation_400_not_detected():
err = litellm.BadRequestError(
message="invalid maxOutputTokens",
model="gemini-pro",
llm_provider="vertex_ai",
)
assert is_provider_account_fallback_eligible_error(err) is False