From c03488c661f11c439df95105cd41a2a82876471c Mon Sep 17 00:00:00 2001 From: Sameer Kankute Date: Wed, 15 Apr 2026 17:38:25 +0530 Subject: [PATCH] docs(advisor): revamp advisor tool docs for cross-provider support - Clearly separate the two tool formats: advisor_20260301 (native/proxy) vs litellm_advisor function tool (chat completions with callback) - Answer "how do I configure the stronger model?" inline where the user first encounters each format - Add supported providers table with native vs orchestration distinction - Show AdvisorInterceptionLogger setup is required for litellm_advisor function format in SDK usage - Update proxy config example to use model_list deployment names - Document provider_specific_fields.advisor_tool_results for chat completions - Document server_tool_use + advisor_tool_result response blocks for messages - Add max_uses, cost, and streaming behavior notes - Remove tool_choice="required" from all examples (causes forced loops) Made-with: Cursor --- .../advisor_tool_chat_completions/index.md | 586 +++++++++++------- .../docs/completion/anthropic_advisor_tool.md | 54 +- 2 files changed, 406 insertions(+), 234 deletions(-) diff --git a/docs/my-website/blog/advisor_tool_chat_completions/index.md b/docs/my-website/blog/advisor_tool_chat_completions/index.md index 77792c267ec..cbe9b7dfe12 100644 --- a/docs/my-website/blog/advisor_tool_chat_completions/index.md +++ b/docs/my-website/blog/advisor_tool_chat_completions/index.md @@ -20,50 +20,75 @@ LiteLLM now supports the Anthropic advisor tool across `chat/completions` and `m Use the advisor tool to let an executor model call a stronger advisor model during generation. For non-Anthropic providers, LiteLLM runs the advisor orchestration loop automatically. -For updates and changes after this post on advisor, see the [latest Advisor Tool docs](/docs/completion/anthropic_advisor_tool). +For updates after this post see the [latest Advisor Tool docs](/docs/completion/anthropic_advisor_tool). :::info Beta -The advisor tool is in beta. Include `anthropic-beta: advisor-tool-2026-03-01` in your requests — LiteLLM adds this automatically when it detects the advisor tool in your `tools` array. +The advisor tool is in beta. LiteLLM adds the required `anthropic-beta: advisor-tool-2026-03-01` header automatically when it detects the advisor tool in your `tools` array. ::: -## Supported Providers +--- -| Provider | Chat Completions API | Messages API | Notes | -|----------|---------------------|--------------|-------| -| **Anthropic API** | ✅ | ✅ | Native — runs server-side | +## Two tool formats + +There are two ways to specify the advisor tool. Which one to use depends on your executor provider and setup. + +### 1. Anthropic native format (`advisor_20260301`) + +```json +{ + "type": "advisor_20260301", + "name": "advisor", + "model": "claude-opus-4-6" +} +``` + +The `model` field is **required** and specifies the advisor. Use this format when: +- Your executor is an Anthropic model and the advisor is `claude-opus-4-6` - Anthropic handles the advisor call natively, server-side. +- Your executor is any non-Anthropic model (OpenAI, Gemini, etc.) via the **Messages API** - LiteLLM's built-in interception converts this automatically. + +### 2. OpenAI function format (`litellm_advisor`) + +```json +{ + "type": "function", + "function": { + "name": "litellm_advisor", + "description": "Consult a stronger advisor model.", + "parameters": { + "type": "object", + "properties": { "question": { "type": "string" } }, + "required": ["question"] + } + } +} +``` + +This format does **not** carry a `model` field. The advisor model comes from your `AdvisorInterceptionLogger` setup or proxy config — see below. Use this format when calling through the **Chat Completions API** and you cannot send custom tool types (e.g. using a plain OpenAI client against the proxy). + +:::warning You must configure the advisor model + +Sending `litellm_advisor` as a bare function tool without setting up `AdvisorInterceptionLogger` (or the proxy `advisor_interception_params`) does nothing useful — the provider treats it as a regular custom tool and returns a `tool_use` response your code has to handle manually. Always pair it with the setup below. + +::: + +--- + +## Supported providers + +| Provider | Chat Completions API | Messages API | Mode | +|----------|---------------------|--------------|------| +| **Anthropic** (executor + advisor = Opus 4.6) | ✅ | ✅ | Native server-side | +| **Anthropic** (executor) + **any other advisor** | ✅ | ✅ | LiteLLM orchestration loop | | **OpenAI / Azure OpenAI** | ✅ | ✅ | LiteLLM orchestration loop | | **Amazon Bedrock** | ✅ | ✅ | LiteLLM orchestration loop | -| **Google Vertex AI** | ✅ | ✅ | LiteLLM orchestration loop | +| **Google Vertex AI / Gemini** | ✅ | ✅ | LiteLLM orchestration loop | | **Groq / Mistral / others** | ✅ | ✅ | LiteLLM orchestration loop | -For non-Anthropic providers, LiteLLM implements the advisor loop itself. +**Native path:** Executor is Anthropic and advisor is `claude-opus-4-6` → Anthropic runs the advisor inference server-side. No LiteLLM orchestration involved. -- **Messages API** (`litellm.anthropic.messages.create/acreate`): built-in interception in the messages handler -- **Chat Completions API** (`litellm.completion/acompletion`): enable `AdvisorInterceptionLogger` to convert advisor tools + run the loop - -When a request arrives with an `advisor_20260301` tool and a non-Anthropic provider, LiteLLM translates the advisor tool into a regular function tool the provider understands, then runs an orchestration loop: - -![Advisor Orchestration Flow](/img/advisor_orchestration_flow.svg) - -**What LiteLLM does for you:** - -- Strips `advisor_20260301` from the outgoing request — the provider only sees a standard function tool named `advisor` -- When the executor calls it, intercepts before the result reaches you, runs the advisor sub-call, and injects the advice -- Strips any `advisor_tool_result` / `server_tool_use` blocks from message history on re-send so non-Anthropic providers never see Anthropic-specific types -- Wraps the final response in an SSE stream if you requested `stream=True` -- Enforces `max_uses` as a hard cap — `AdvisorMaxIterationsError` is raised if exceeded; `max_uses=0` disables the advisor entirely - -## Model Compatibility - -The executor and advisor models must form a valid pair. Currently the only supported advisor model is `claude-opus-4-6`. - -| Executor | Advisor | -|----------|---------| -| `claude-haiku-4-5-20251001` | `claude-opus-4-6` | -| `claude-sonnet-4-6` | `claude-opus-4-6` | -| `claude-opus-4-6` | `claude-opus-4-6` | +**Orchestration path:** Everything else → LiteLLM intercepts the executor's tool call, runs the advisor as a sub-call using the credentials you configured, injects the advice, and continues. The advisor can be any provider. --- @@ -72,9 +97,73 @@ The executor and advisor models must form a valid pair. Currently the only suppo -#### Basic Example (Anthropic-native executor) +### Configuring the advisor model (SDK) -```python showLineNumbers title="Advisor Tool — litellm.completion()" +Register `AdvisorInterceptionLogger` in `litellm.callbacks` and set `default_advisor_model`. This is what routes advisor sub-calls to the right model and credentials. + +```python showLineNumbers title="SDK setup — register AdvisorInterceptionLogger" +import litellm +from litellm.integrations.advisor_interception import AdvisorInterceptionLogger + +litellm.callbacks = [ + AdvisorInterceptionLogger( + # Any provider LiteLLM supports. Use full model string for direct calls, + # or a model_name from your model_list when using the proxy router. + default_advisor_model="openai/gpt-4o", + # Optional: limit interception to specific executor providers. + # Remove this line to intercept for all providers. + enabled_providers=["anthropic", "openai"], + ) +] +``` + +`default_advisor_model` is used when the tool definition has no `model` field (i.e. the `litellm_advisor` function format). If you pass the `advisor_20260301` native format with an explicit `model` field, that takes precedence. + +--- + +### Anthropic executor (any advisor) + +```python showLineNumbers title="Anthropic executor + OpenAI advisor" +import asyncio +import litellm +from litellm.integrations.advisor_interception import AdvisorInterceptionLogger + +litellm.callbacks = [ + AdvisorInterceptionLogger(default_advisor_model="openai/gpt-4o") +] + +async def main(): + response = await litellm.acompletion( + model="anthropic/claude-sonnet-4-6", + messages=[ + {"role": "user", "content": "Build a concurrent worker pool in Go with graceful shutdown."} + ], + tools=[ + { + "type": "function", + "function": { + "name": "litellm_advisor", + "description": "Consult a stronger advisor model.", + "parameters": { + "type": "object", + "properties": {"question": {"type": "string"}}, + "required": ["question"], + }, + }, + } + ], + max_tokens=4096, + ) + print(response.choices[0].message.content) + +asyncio.run(main()) +``` + +LiteLLM detects the `litellm_advisor` function tool, converts it to a provider-compatible tool, intercepts the tool call in the response, calls `openai/gpt-4o` as the advisor, and injects the advice before returning the final answer. + +**To use Anthropic's native advisor path** (Anthropic handles advisor inference server-side), use the `advisor_20260301` format with `model: "claude-opus-4-6"` — no callback needed: + +```python showLineNumbers title="Anthropic-native path (executor + advisor both Anthropic)" import litellm response = litellm.completion( @@ -84,179 +173,183 @@ response = litellm.completion( ], tools=[ { - "type": "function", - "function": { - "name": "litellm_advisor", - "description": "Consult a stronger advisor model.", - "parameters": { - "type": "object", - "properties": { - "question": {"type": "string"} - }, - "required": ["question"], - }, - }, + "type": "advisor_20260301", + "name": "advisor", + "model": "claude-opus-4-6", # advisor model — required } ], max_tokens=4096, ) - print(response.choices[0].message.content) ``` -#### Non-Anthropic Executor (Chat Completions interception) +--- -```python showLineNumbers title="Advisor Tool with OpenAI executor via chat-completions" +### Non-Anthropic executor + +```python showLineNumbers title="OpenAI executor + OpenAI advisor" import asyncio import litellm -from litellm.integrations.advisor_interception import ( - AdvisorInterceptionLogger, - get_litellm_advisor_tool, -) +from litellm.integrations.advisor_interception import AdvisorInterceptionLogger -litellm.callbacks = [AdvisorInterceptionLogger(enabled_providers=["openai"])] +litellm.callbacks = [ + AdvisorInterceptionLogger(default_advisor_model="openai/gpt-4o") +] async def main(): response = await litellm.acompletion( - model="gpt-5.4-mini", + model="openai/gpt-4o-mini", messages=[ - {"role": "user", "content": "Build a concurrent worker pool in Go with graceful shutdown."} + {"role": "user", "content": "Design a rate limiter for a distributed API gateway."} ], - # You can still use Anthropic-native advisor tool format. - tools=[get_litellm_advisor_tool(model="claude-opus-4-6")], - max_tokens=4096, + tools=[ + { + "type": "function", + "function": { + "name": "litellm_advisor", + "description": "Consult a stronger advisor model.", + "parameters": { + "type": "object", + "properties": {"question": {"type": "string"}}, + "required": ["question"], + }, + }, + } + ], + max_tokens=2048, ) print(response.choices[0].message.content) asyncio.run(main()) ``` -::::note +You can also pass the advisor model directly in the tool definition using the native format — this overrides `default_advisor_model`: -`AdvisorInterceptionLogger` converts advisor tool definitions to provider-compatible function tools for non-Anthropic chat-completions providers and runs the advisor sub-call loop server-side. +```python showLineNumbers title="Advisor model set per-request in tool definition" +from litellm.integrations.advisor_interception import get_litellm_advisor_tool -:::: - -#### With Optional Parameters - -```python showLineNumbers title="Advisor Tool with max_uses and caching" -import litellm - -response = litellm.completion( - model="anthropic/claude-sonnet-4-6", - messages=[ - {"role": "user", "content": "Build a REST API with authentication in Python."} - ], - tools=[ - { - "type": "advisor_20260301", - "name": "advisor", - "model": "claude-opus-4-6", - "max_uses": 3, # cap advisor calls per request - "caching": {"type": "ephemeral", "ttl": "5m"}, # enable for 3+ calls per conversation - } - ], - max_tokens=4096, -) +tools=[ + get_litellm_advisor_tool( + model="openai/gpt-4o", # overrides default_advisor_model for this request + max_uses=2, + ) +] ``` -#### Streaming +--- + +### Streaming ```python showLineNumbers title="Streaming with Advisor Tool" +import asyncio import litellm +from litellm.integrations.advisor_interception import AdvisorInterceptionLogger -response = litellm.completion( - model="anthropic/claude-sonnet-4-6", - messages=[ - {"role": "user", "content": "Implement a distributed rate limiter."} - ], - tools=[ - { - "type": "advisor_20260301", - "name": "advisor", - "model": "claude-opus-4-6", - } - ], - max_tokens=4096, - stream=True, -) +litellm.callbacks = [ + AdvisorInterceptionLogger(default_advisor_model="openai/gpt-4o") +] -for chunk in response: - if chunk.choices[0].delta.content: - print(chunk.choices[0].delta.content, end="") +async def main(): + response = await litellm.acompletion( + model="openai/gpt-4o-mini", + messages=[ + {"role": "user", "content": "Implement a distributed rate limiter."} + ], + tools=[ + { + "type": "function", + "function": { + "name": "litellm_advisor", + "description": "Consult a stronger advisor model.", + "parameters": { + "type": "object", + "properties": {"question": {"type": "string"}}, + "required": ["question"], + }, + }, + } + ], + max_tokens=4096, + stream=True, + ) + + async for chunk in response: + if chunk.choices[0].delta.content: + print(chunk.choices[0].delta.content, end="") + +asyncio.run(main()) ``` :::note Streaming behavior -The advisor sub-inference does not stream. The executor's stream pauses while the advisor runs, then the full advisor result arrives in a single event. Executor output resumes streaming afterward. - -::: - -#### Multi-Turn Conversation - -```python showLineNumbers title="Multi-Turn with Advisor Tool" -import litellm - -tools = [ - { - "type": "advisor_20260301", - "name": "advisor", - "model": "claude-opus-4-6", - } -] - -messages = [ - {"role": "user", "content": "Build a concurrent worker pool in Go with graceful shutdown."} -] - -response = litellm.completion( - model="anthropic/claude-sonnet-4-6", - messages=messages, - tools=tools, - max_tokens=4096, -) - -# Append the full response (includes server_tool_use + advisor_tool_result blocks) -messages.append({"role": "assistant", "content": response.choices[0].message.content}) - -# Continue the conversation — keep the same tools array -messages.append({"role": "user", "content": "Now add a max-in-flight limit of 10."}) - -response2 = litellm.completion( - model="anthropic/claude-sonnet-4-6", - messages=messages, - tools=tools, - max_tokens=4096, -) -``` - -:::tip Auto-strip on follow-up turns - -LiteLLM automatically strips `advisor_tool_result` blocks from message history when the advisor tool is not present in the current request. This prevents the Anthropic 400 error that would otherwise occur. +The advisor sub-inference does not stream. When the executor calls the advisor tool, the stream pauses, the advisor runs to completion, and its output is injected before the executor resumes streaming. ::: -#### Proxy Configuration +### Configuring the advisor model (Proxy) + +Add the advisor as a named deployment in `model_list` and reference it in `advisor_interception_params`. The proxy router resolves the correct credentials automatically. ```yaml showLineNumbers title="config.yaml" model_list: + # Advisor model — can be any provider + - model_name: my-advisor + litellm_params: + model: openai/gpt-4o + api_key: os.environ/OPENAI_API_KEY + + # Or use an Anthropic model as advisor + # - model_name: my-advisor + # litellm_params: + # model: anthropic/claude-opus-4-6 + # api_key: os.environ/ANTHROPIC_API_KEY + + # Executor models - model_name: claude-sonnet litellm_params: model: anthropic/claude-sonnet-4-6 api_key: os.environ/ANTHROPIC_API_KEY + + - model_name: gpt-4o-mini + litellm_params: + model: openai/gpt-4o-mini + api_key: os.environ/OPENAI_API_KEY + + - model_name: gemini-flash + litellm_params: + model: vertex_ai/gemini-2.5-flash + vertex_project: my-project + vertex_location: us-central1 + +litellm_settings: + callbacks: ["advisor_interception"] # use callbacks, not success_callback + advisor_interception_params: + # Must be a model_name from model_list above. + # The router uses this to pick the right deployment + credentials. + default_advisor_model: "my-advisor" ``` -#### Client Request via Proxy +:::info -```python showLineNumbers title="Advisor Tool via AI Gateway" +- Use `callbacks`, not `success_callback`. The advisor hooks run through `litellm.callbacks`. +- `default_advisor_model` must match a `model_name` from `model_list`. This is how the proxy resolves the correct API key and deployment for the advisor sub-call. +- You can still override it per-request by passing `model` in an `advisor_20260301` tool definition. + +::: + +--- + +### Client request — native advisor format + +```python showLineNumbers title="Advisor via proxy (advisor_20260301 format)" from openai import OpenAI client = OpenAI( api_key="your-litellm-proxy-key", - base_url="http://0.0.0.0:4000/v1" + base_url="http://0.0.0.0:4000/v1", ) response = client.chat.completions.create( @@ -268,18 +361,19 @@ response = client.chat.completions.create( { "type": "advisor_20260301", "name": "advisor", - "model": "claude-opus-4-6", + "model": "my-advisor", # matches model_name in config.yaml } ], max_tokens=4096, ) +print(response.choices[0].message.content) ``` -#### Client Request via Proxy (OpenAI-compatible function tool) +### Client request — OpenAI function format -Use this format when your chat-completions client sends OpenAI-style tools. +Use this when your client cannot send custom `type` values (e.g. plain OpenAI SDK). The proxy uses `default_advisor_model` from config. -```python showLineNumbers title="Proxy Chat Completions with litellm_advisor" +```python showLineNumbers title="Advisor via proxy (litellm_advisor function format)" from openai import OpenAI client = OpenAI( @@ -288,9 +382,9 @@ client = OpenAI( ) response = client.chat.completions.create( - model="gemini-flash", + model="gpt-4o-mini", messages=[ - {"role": "user", "content": "Call advisor once, then answer in one line: integration ok."} + {"role": "user", "content": "Design a fault-tolerant task queue in Python."} ], tools=[ { @@ -300,26 +394,18 @@ response = client.chat.completions.create( "description": "Consult a stronger advisor model.", "parameters": { "type": "object", - "properties": { - "question": {"type": "string"} - }, + "properties": {"question": {"type": "string"}}, "required": ["question"], }, }, } ], - max_tokens=512, + max_tokens=2048, ) print(response.choices[0].message.content) ``` -::::note - -For non-Anthropic chat-completions providers behind proxy, this OpenAI-compatible -`litellm_advisor` function tool is the recommended request shape. -The advisor model defaults to `claude-opus-4-6` unless overridden by your integration config. - -:::: +The proxy intercepts the `litellm_advisor` tool call, calls `my-advisor` (from config), injects the result, and returns the final answer — your client only sees the finished response. @@ -331,9 +417,11 @@ The advisor model defaults to `claude-opus-4-6` unless overridden by your integr -#### Basic Example +The Messages API (`litellm.anthropic.messages`) has built-in interception — no callback registration needed. Pass the `advisor_20260301` tool with the `model` field and LiteLLM handles the rest. -```python showLineNumbers title="Advisor Tool — litellm.anthropic.messages" +#### Anthropic executor — native path + +```python showLineNumbers title="Advisor Tool — Messages API, Anthropic native" import asyncio import litellm @@ -347,7 +435,7 @@ async def main(): { "type": "advisor_20260301", "name": "advisor", - "model": "claude-opus-4-6", + "model": "claude-opus-4-6", # Anthropic runs this natively } ], max_tokens=4096, @@ -357,6 +445,36 @@ async def main(): asyncio.run(main()) ``` +#### Non-Anthropic executor — LiteLLM orchestration loop + +When the executor is not Anthropic, or when the advisor model is not Claude Opus 4.6, LiteLLM runs the loop itself. The `model` field in the tool definition is the advisor — it can be any provider. + +```python showLineNumbers title="Advisor Tool — Messages API, OpenAI executor" +import asyncio +import litellm + +async def main(): + response = await litellm.anthropic.messages.acreate( + model="openai/gpt-4o", + messages=[ + {"role": "user", "content": "Implement a Python LRU cache with O(1) get and put."} + ], + tools=[ + { + "type": "advisor_20260301", + "name": "advisor", + "model": "openai/gpt-4o-mini", # advisor model — any provider works + "max_uses": 2, + } + ], + max_tokens=1024, + custom_llm_provider="openai", + ) + print(response["content"][0]["text"]) + +asyncio.run(main()) +``` + #### Streaming ```python showLineNumbers title="Messages API Streaming with Advisor Tool" @@ -396,24 +514,16 @@ asyncio.run(main()) -#### Proxy Configuration +Use the same `config.yaml` shown in the Chat Completions proxy tab. The `advisor_interception_params` config applies to both APIs. -```yaml showLineNumbers title="config.yaml" -model_list: - - model_name: claude-sonnet - litellm_params: - model: anthropic/claude-sonnet-4-6 - api_key: os.environ/ANTHROPIC_API_KEY -``` - -#### Client Request via Proxy (Anthropic SDK) +#### Client request — Anthropic SDK ```python showLineNumbers title="Advisor Tool via AI Gateway (Anthropic SDK)" import anthropic client = anthropic.Anthropic( api_key="your-litellm-proxy-key", - base_url="http://0.0.0.0:4000" + base_url="http://0.0.0.0:4000", ) response = client.beta.messages.create( @@ -427,92 +537,97 @@ response = client.beta.messages.create( { "type": "advisor_20260301", "name": "advisor", - "model": "claude-opus-4-6", + "model": "my-advisor", # model_name from config.yaml } ], ) print(response) ``` -#### Non-Anthropic Provider (LiteLLM orchestration loop) - -```python showLineNumbers title="Advisor Tool with OpenAI executor" -import asyncio -import litellm - -async def main(): - # executor: openai/gpt-4.1-mini | advisor: claude-opus-4-6 - # LiteLLM runs the orchestration loop automatically - response = await litellm.anthropic.messages.acreate( - model="openai/gpt-4.1-mini", - messages=[ - {"role": "user", "content": "Implement a Python LRU cache with O(1) get and put."} - ], - tools=[ - { - "type": "advisor_20260301", - "name": "advisor", - "model": "claude-opus-4-6", - "max_uses": 3, - } - ], - max_tokens=1024, - custom_llm_provider="openai", - ) - # Final response is clean — no advisor tool_use blocks - print(response["content"][0]["text"]) - -asyncio.run(main()) -``` - --- -## Response Structure +## Response structure -A successful advisor call returns `server_tool_use` and `advisor_tool_result` blocks in the assistant content: +### Messages API -```json title="Response with advisor blocks" +Both native and orchestration paths return `server_tool_use` and `advisor_tool_result` blocks in the assistant content: + +```json title="Messages API response" { "role": "assistant", "content": [ { "type": "text", - "text": "Let me consult the advisor on this." + "text": "Here is the implementation:" }, { "type": "server_tool_use", "id": "srvtoolu_abc123", - "name": "advisor", - "input": {} + "name": "advisor" }, { "type": "advisor_tool_result", "tool_use_id": "srvtoolu_abc123", "content": { "type": "advisor_result", - "text": "Use a channel-based coordination pattern. The tricky part is draining in-flight work during shutdown: close the input channel first, then wait on a WaitGroup..." + "text": "Use a channel-based coordination pattern..." } }, { "type": "text", - "text": "Here's the implementation using a channel-based coordination pattern..." + "text": "Here's the full implementation..." } ] } ``` -Pass the full assistant content, including advisor blocks, back on subsequent turns. LiteLLM handles this automatically through `provider_specific_fields`. +### Chat Completions API + +For chat completions, the advisor blocks are in `provider_specific_fields` on the response message: + +```python title="Accessing advisor results from chat completions" +response = await litellm.acompletion(...) + +message = response.choices[0].message +print(message.content) # final answer + +# Advisor trace — available when the advisor was called +psf = message.provider_specific_fields or {} +for block in psf.get("advisor_tool_results", []): + if block["type"] == "advisor_tool_result": + print("Advisor said:", block["content"]["text"]) +``` + +```json title="provider_specific_fields structure" +{ + "advisor_tool_results": [ + { + "type": "server_tool_use", + "id": "call_abc123", + "name": "advisor" + }, + { + "type": "advisor_tool_result", + "tool_use_id": "call_abc123", + "content": { + "type": "advisor_result", + "text": "Use a channel-based coordination pattern..." + } + } + ] +} +``` --- -## Cost Control +## Cost control -Advisor calls run as a separate sub-inference billed at the advisor model's rates. Usage is reported in `usage.iterations[]`: +Advisor calls run as separate sub-inferences billed at the advisor model's rates. Usage is reported in `usage.iterations[]` (Messages API) or accumulated in `usage` (Chat Completions): -```json title="Usage with advisor sub-inference" +```json title="Messages API usage with advisor sub-inference" { "usage": { "input_tokens": 412, @@ -539,9 +654,22 @@ Advisor calls run as a separate sub-inference billed at the advisor model's rate } ``` -Top-level `usage` reflects executor tokens only. Advisor tokens appear in `iterations` entries with `type: "advisor_message"` and are billed at Opus rates. +Use `max_uses` in the tool definition to cap how many times the advisor can be called per request: -## Additional Resources +```python +tools=[ + { + "type": "advisor_20260301", + "name": "advisor", + "model": "my-advisor", + "max_uses": 2, # raise AdvisorMaxIterationsError after 2 advisor calls + } +] +``` + +--- + +## Additional resources - [Anthropic Advisor Tool Documentation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool) - [LiteLLM Tool Calling Guide](https://docs.litellm.ai/docs/completion/function_call) diff --git a/docs/my-website/docs/completion/anthropic_advisor_tool.md b/docs/my-website/docs/completion/anthropic_advisor_tool.md index bc16e4681a7..b98fcad644b 100644 --- a/docs/my-website/docs/completion/anthropic_advisor_tool.md +++ b/docs/my-website/docs/completion/anthropic_advisor_tool.md @@ -42,7 +42,9 @@ When a request arrives with an `advisor_20260301` tool and a non-Anthropic provi ## Model Compatibility -The executor and advisor models must form a valid pair. Currently the only supported advisor model is `claude-opus-4-6`. +The advisor model is fully configurable. You can use any model deployed in your proxy as the advisor — it does not need to be Anthropic. + +For **Anthropic-native** requests (where Anthropic runs the advisor server-side), the executor and advisor must form a valid Anthropic pair: | Executor | Advisor | |----------|---------| @@ -50,6 +52,8 @@ The executor and advisor models must form a valid pair. Currently the only suppo | `claude-sonnet-4-6` | `claude-opus-4-6` | | `claude-opus-4-6` | `claude-opus-4-6` | +For **non-Anthropic** executors (where LiteLLM orchestrates the advisor loop), you can use any model as the advisor — including OpenAI, Vertex AI, Bedrock, etc. + --- ## Chat Completions API @@ -226,14 +230,43 @@ LiteLLM automatically strips `advisor_tool_result` blocks from message history w #### Proxy Configuration +Configure the advisor model as a deployment in your `model_list` and reference it in `advisor_interception_params`. This ensures the advisor sub-calls use the correct credentials and go through the proxy's deployment routing. + ```yaml showLineNumbers title="config.yaml" model_list: + # The advisor model + - model_name: advisor-model + litellm_params: + model: anthropic/claude-sonnet-4-20250514 + api_key: os.environ/ANTHROPIC_API_KEY + + # Executor models - model_name: claude-sonnet litellm_params: model: anthropic/claude-sonnet-4-6 api_key: os.environ/ANTHROPIC_API_KEY + + - model_name: gemini-flash + litellm_params: + model: vertex_ai/gemini-2.5-flash + vertex_project: my-project + vertex_location: us-central1 + +litellm_settings: + callbacks: ["advisor_interception"] + advisor_interception_params: + # Must match a model_name from model_list — the router resolves + # the correct deployment and credentials automatically. + default_advisor_model: "advisor-model" ``` +:::info Important + +- Use `callbacks`, not `success_callback`. The advisor interception hooks run through `litellm.callbacks`. +- The `default_advisor_model` value must be a `model_name` from your `model_list`. The proxy router resolves it to the correct deployment with the correct API key. This means you can use any provider as your advisor model — not just Anthropic. + +::: + #### Client Request via Proxy ```python showLineNumbers title="Advisor Tool via AI Gateway" @@ -253,7 +286,7 @@ response = client.chat.completions.create( { "type": "advisor_20260301", "name": "advisor", - "model": "claude-opus-4-6", + "model": "advisor-model", } ], max_tokens=4096, @@ -262,7 +295,7 @@ response = client.chat.completions.create( #### Client Request via Proxy (OpenAI-compatible function tool) -Use this format when your chat-completions client sends OpenAI-style tools. +Use this format when your chat-completions client sends OpenAI-style tools. The proxy uses the `default_advisor_model` from your config. ```python showLineNumbers title="Proxy Chat Completions with litellm_advisor" from openai import OpenAI @@ -302,7 +335,7 @@ print(response.choices[0].message.content) For non-Anthropic chat-completions providers behind proxy, this OpenAI-compatible `litellm_advisor` function tool is the recommended request shape. -The advisor model defaults to `claude-opus-4-6` unless overridden by your integration config. +The advisor model is determined by the `default_advisor_model` in your `advisor_interception_params` config. :::: @@ -383,12 +416,23 @@ asyncio.run(main()) #### Proxy Configuration +Use the same config shown in the Chat Completions proxy tab. The `advisor_interception_params` config applies to both APIs. + ```yaml showLineNumbers title="config.yaml" model_list: + - model_name: advisor-model + litellm_params: + model: anthropic/claude-sonnet-4-20250514 + api_key: os.environ/ANTHROPIC_API_KEY - model_name: claude-sonnet litellm_params: model: anthropic/claude-sonnet-4-6 api_key: os.environ/ANTHROPIC_API_KEY + +litellm_settings: + callbacks: ["advisor_interception"] + advisor_interception_params: + default_advisor_model: "advisor-model" ``` #### Client Request via Proxy (Anthropic SDK) @@ -412,7 +456,7 @@ response = client.beta.messages.create( { "type": "advisor_20260301", "name": "advisor", - "model": "claude-opus-4-6", + "model": "advisor-model", } ], )