mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-11 22:51:28 +00:00
docs(advisor): revamp advisor tool docs for cross-provider support
- Clearly separate the two tool formats: advisor_20260301 (native/proxy) vs litellm_advisor function tool (chat completions with callback) - Answer "how do I configure the stronger model?" inline where the user first encounters each format - Add supported providers table with native vs orchestration distinction - Show AdvisorInterceptionLogger setup is required for litellm_advisor function format in SDK usage - Update proxy config example to use model_list deployment names - Document provider_specific_fields.advisor_tool_results for chat completions - Document server_tool_use + advisor_tool_result response blocks for messages - Add max_uses, cost, and streaming behavior notes - Remove tool_choice="required" from all examples (causes forced loops) Made-with: Cursor
This commit is contained in:
parent
1f3fc7b5bb
commit
c03488c661
2 changed files with 406 additions and 234 deletions
|
|
@ -20,50 +20,75 @@ LiteLLM now supports the Anthropic advisor tool across `chat/completions` and `m
|
|||
|
||||
Use the advisor tool to let an executor model call a stronger advisor model during generation. For non-Anthropic providers, LiteLLM runs the advisor orchestration loop automatically.
|
||||
|
||||
For updates and changes after this post on advisor, see the [latest Advisor Tool docs](/docs/completion/anthropic_advisor_tool).
|
||||
For updates after this post see the [latest Advisor Tool docs](/docs/completion/anthropic_advisor_tool).
|
||||
|
||||
:::info Beta
|
||||
|
||||
The advisor tool is in beta. Include `anthropic-beta: advisor-tool-2026-03-01` in your requests — LiteLLM adds this automatically when it detects the advisor tool in your `tools` array.
|
||||
The advisor tool is in beta. LiteLLM adds the required `anthropic-beta: advisor-tool-2026-03-01` header automatically when it detects the advisor tool in your `tools` array.
|
||||
|
||||
:::
|
||||
|
||||
## Supported Providers
|
||||
---
|
||||
|
||||
| Provider | Chat Completions API | Messages API | Notes |
|
||||
|----------|---------------------|--------------|-------|
|
||||
| **Anthropic API** | ✅ | ✅ | Native — runs server-side |
|
||||
## Two tool formats
|
||||
|
||||
There are two ways to specify the advisor tool. Which one to use depends on your executor provider and setup.
|
||||
|
||||
### 1. Anthropic native format (`advisor_20260301`)
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6"
|
||||
}
|
||||
```
|
||||
|
||||
The `model` field is **required** and specifies the advisor. Use this format when:
|
||||
- Your executor is an Anthropic model and the advisor is `claude-opus-4-6` - Anthropic handles the advisor call natively, server-side.
|
||||
- Your executor is any non-Anthropic model (OpenAI, Gemini, etc.) via the **Messages API** - LiteLLM's built-in interception converts this automatically.
|
||||
|
||||
### 2. OpenAI function format (`litellm_advisor`)
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "litellm_advisor",
|
||||
"description": "Consult a stronger advisor model.",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": { "question": { "type": "string" } },
|
||||
"required": ["question"]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This format does **not** carry a `model` field. The advisor model comes from your `AdvisorInterceptionLogger` setup or proxy config — see below. Use this format when calling through the **Chat Completions API** and you cannot send custom tool types (e.g. using a plain OpenAI client against the proxy).
|
||||
|
||||
:::warning You must configure the advisor model
|
||||
|
||||
Sending `litellm_advisor` as a bare function tool without setting up `AdvisorInterceptionLogger` (or the proxy `advisor_interception_params`) does nothing useful — the provider treats it as a regular custom tool and returns a `tool_use` response your code has to handle manually. Always pair it with the setup below.
|
||||
|
||||
:::
|
||||
|
||||
---
|
||||
|
||||
## Supported providers
|
||||
|
||||
| Provider | Chat Completions API | Messages API | Mode |
|
||||
|----------|---------------------|--------------|------|
|
||||
| **Anthropic** (executor + advisor = Opus 4.6) | ✅ | ✅ | Native server-side |
|
||||
| **Anthropic** (executor) + **any other advisor** | ✅ | ✅ | LiteLLM orchestration loop |
|
||||
| **OpenAI / Azure OpenAI** | ✅ | ✅ | LiteLLM orchestration loop |
|
||||
| **Amazon Bedrock** | ✅ | ✅ | LiteLLM orchestration loop |
|
||||
| **Google Vertex AI** | ✅ | ✅ | LiteLLM orchestration loop |
|
||||
| **Google Vertex AI / Gemini** | ✅ | ✅ | LiteLLM orchestration loop |
|
||||
| **Groq / Mistral / others** | ✅ | ✅ | LiteLLM orchestration loop |
|
||||
|
||||
For non-Anthropic providers, LiteLLM implements the advisor loop itself.
|
||||
**Native path:** Executor is Anthropic and advisor is `claude-opus-4-6` → Anthropic runs the advisor inference server-side. No LiteLLM orchestration involved.
|
||||
|
||||
- **Messages API** (`litellm.anthropic.messages.create/acreate`): built-in interception in the messages handler
|
||||
- **Chat Completions API** (`litellm.completion/acompletion`): enable `AdvisorInterceptionLogger` to convert advisor tools + run the loop
|
||||
|
||||
When a request arrives with an `advisor_20260301` tool and a non-Anthropic provider, LiteLLM translates the advisor tool into a regular function tool the provider understands, then runs an orchestration loop:
|
||||
|
||||

|
||||
|
||||
**What LiteLLM does for you:**
|
||||
|
||||
- Strips `advisor_20260301` from the outgoing request — the provider only sees a standard function tool named `advisor`
|
||||
- When the executor calls it, intercepts before the result reaches you, runs the advisor sub-call, and injects the advice
|
||||
- Strips any `advisor_tool_result` / `server_tool_use` blocks from message history on re-send so non-Anthropic providers never see Anthropic-specific types
|
||||
- Wraps the final response in an SSE stream if you requested `stream=True`
|
||||
- Enforces `max_uses` as a hard cap — `AdvisorMaxIterationsError` is raised if exceeded; `max_uses=0` disables the advisor entirely
|
||||
|
||||
## Model Compatibility
|
||||
|
||||
The executor and advisor models must form a valid pair. Currently the only supported advisor model is `claude-opus-4-6`.
|
||||
|
||||
| Executor | Advisor |
|
||||
|----------|---------|
|
||||
| `claude-haiku-4-5-20251001` | `claude-opus-4-6` |
|
||||
| `claude-sonnet-4-6` | `claude-opus-4-6` |
|
||||
| `claude-opus-4-6` | `claude-opus-4-6` |
|
||||
**Orchestration path:** Everything else → LiteLLM intercepts the executor's tool call, runs the advisor as a sub-call using the credentials you configured, injects the advice, and continues. The advisor can be any provider.
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -72,9 +97,73 @@ The executor and advisor models must form a valid pair. Currently the only suppo
|
|||
<Tabs>
|
||||
<TabItem value="chat-completions-sdk" label="SDK">
|
||||
|
||||
#### Basic Example (Anthropic-native executor)
|
||||
### Configuring the advisor model (SDK)
|
||||
|
||||
```python showLineNumbers title="Advisor Tool — litellm.completion()"
|
||||
Register `AdvisorInterceptionLogger` in `litellm.callbacks` and set `default_advisor_model`. This is what routes advisor sub-calls to the right model and credentials.
|
||||
|
||||
```python showLineNumbers title="SDK setup — register AdvisorInterceptionLogger"
|
||||
import litellm
|
||||
from litellm.integrations.advisor_interception import AdvisorInterceptionLogger
|
||||
|
||||
litellm.callbacks = [
|
||||
AdvisorInterceptionLogger(
|
||||
# Any provider LiteLLM supports. Use full model string for direct calls,
|
||||
# or a model_name from your model_list when using the proxy router.
|
||||
default_advisor_model="openai/gpt-4o",
|
||||
# Optional: limit interception to specific executor providers.
|
||||
# Remove this line to intercept for all providers.
|
||||
enabled_providers=["anthropic", "openai"],
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
`default_advisor_model` is used when the tool definition has no `model` field (i.e. the `litellm_advisor` function format). If you pass the `advisor_20260301` native format with an explicit `model` field, that takes precedence.
|
||||
|
||||
---
|
||||
|
||||
### Anthropic executor (any advisor)
|
||||
|
||||
```python showLineNumbers title="Anthropic executor + OpenAI advisor"
|
||||
import asyncio
|
||||
import litellm
|
||||
from litellm.integrations.advisor_interception import AdvisorInterceptionLogger
|
||||
|
||||
litellm.callbacks = [
|
||||
AdvisorInterceptionLogger(default_advisor_model="openai/gpt-4o")
|
||||
]
|
||||
|
||||
async def main():
|
||||
response = await litellm.acompletion(
|
||||
model="anthropic/claude-sonnet-4-6",
|
||||
messages=[
|
||||
{"role": "user", "content": "Build a concurrent worker pool in Go with graceful shutdown."}
|
||||
],
|
||||
tools=[
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "litellm_advisor",
|
||||
"description": "Consult a stronger advisor model.",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {"question": {"type": "string"}},
|
||||
"required": ["question"],
|
||||
},
|
||||
},
|
||||
}
|
||||
],
|
||||
max_tokens=4096,
|
||||
)
|
||||
print(response.choices[0].message.content)
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
LiteLLM detects the `litellm_advisor` function tool, converts it to a provider-compatible tool, intercepts the tool call in the response, calls `openai/gpt-4o` as the advisor, and injects the advice before returning the final answer.
|
||||
|
||||
**To use Anthropic's native advisor path** (Anthropic handles advisor inference server-side), use the `advisor_20260301` format with `model: "claude-opus-4-6"` — no callback needed:
|
||||
|
||||
```python showLineNumbers title="Anthropic-native path (executor + advisor both Anthropic)"
|
||||
import litellm
|
||||
|
||||
response = litellm.completion(
|
||||
|
|
@ -84,179 +173,183 @@ response = litellm.completion(
|
|||
],
|
||||
tools=[
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "litellm_advisor",
|
||||
"description": "Consult a stronger advisor model.",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"question": {"type": "string"}
|
||||
},
|
||||
"required": ["question"],
|
||||
},
|
||||
},
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6", # advisor model — required
|
||||
}
|
||||
],
|
||||
max_tokens=4096,
|
||||
)
|
||||
|
||||
print(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
#### Non-Anthropic Executor (Chat Completions interception)
|
||||
---
|
||||
|
||||
```python showLineNumbers title="Advisor Tool with OpenAI executor via chat-completions"
|
||||
### Non-Anthropic executor
|
||||
|
||||
```python showLineNumbers title="OpenAI executor + OpenAI advisor"
|
||||
import asyncio
|
||||
import litellm
|
||||
from litellm.integrations.advisor_interception import (
|
||||
AdvisorInterceptionLogger,
|
||||
get_litellm_advisor_tool,
|
||||
)
|
||||
from litellm.integrations.advisor_interception import AdvisorInterceptionLogger
|
||||
|
||||
litellm.callbacks = [AdvisorInterceptionLogger(enabled_providers=["openai"])]
|
||||
litellm.callbacks = [
|
||||
AdvisorInterceptionLogger(default_advisor_model="openai/gpt-4o")
|
||||
]
|
||||
|
||||
async def main():
|
||||
response = await litellm.acompletion(
|
||||
model="gpt-5.4-mini",
|
||||
model="openai/gpt-4o-mini",
|
||||
messages=[
|
||||
{"role": "user", "content": "Build a concurrent worker pool in Go with graceful shutdown."}
|
||||
{"role": "user", "content": "Design a rate limiter for a distributed API gateway."}
|
||||
],
|
||||
# You can still use Anthropic-native advisor tool format.
|
||||
tools=[get_litellm_advisor_tool(model="claude-opus-4-6")],
|
||||
max_tokens=4096,
|
||||
tools=[
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "litellm_advisor",
|
||||
"description": "Consult a stronger advisor model.",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {"question": {"type": "string"}},
|
||||
"required": ["question"],
|
||||
},
|
||||
},
|
||||
}
|
||||
],
|
||||
max_tokens=2048,
|
||||
)
|
||||
print(response.choices[0].message.content)
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
::::note
|
||||
You can also pass the advisor model directly in the tool definition using the native format — this overrides `default_advisor_model`:
|
||||
|
||||
`AdvisorInterceptionLogger` converts advisor tool definitions to provider-compatible function tools for non-Anthropic chat-completions providers and runs the advisor sub-call loop server-side.
|
||||
```python showLineNumbers title="Advisor model set per-request in tool definition"
|
||||
from litellm.integrations.advisor_interception import get_litellm_advisor_tool
|
||||
|
||||
::::
|
||||
|
||||
#### With Optional Parameters
|
||||
|
||||
```python showLineNumbers title="Advisor Tool with max_uses and caching"
|
||||
import litellm
|
||||
|
||||
response = litellm.completion(
|
||||
model="anthropic/claude-sonnet-4-6",
|
||||
messages=[
|
||||
{"role": "user", "content": "Build a REST API with authentication in Python."}
|
||||
],
|
||||
tools=[
|
||||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
"max_uses": 3, # cap advisor calls per request
|
||||
"caching": {"type": "ephemeral", "ttl": "5m"}, # enable for 3+ calls per conversation
|
||||
}
|
||||
],
|
||||
max_tokens=4096,
|
||||
)
|
||||
tools=[
|
||||
get_litellm_advisor_tool(
|
||||
model="openai/gpt-4o", # overrides default_advisor_model for this request
|
||||
max_uses=2,
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
#### Streaming
|
||||
---
|
||||
|
||||
### Streaming
|
||||
|
||||
```python showLineNumbers title="Streaming with Advisor Tool"
|
||||
import asyncio
|
||||
import litellm
|
||||
from litellm.integrations.advisor_interception import AdvisorInterceptionLogger
|
||||
|
||||
response = litellm.completion(
|
||||
model="anthropic/claude-sonnet-4-6",
|
||||
messages=[
|
||||
{"role": "user", "content": "Implement a distributed rate limiter."}
|
||||
],
|
||||
tools=[
|
||||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
}
|
||||
],
|
||||
max_tokens=4096,
|
||||
stream=True,
|
||||
)
|
||||
litellm.callbacks = [
|
||||
AdvisorInterceptionLogger(default_advisor_model="openai/gpt-4o")
|
||||
]
|
||||
|
||||
for chunk in response:
|
||||
if chunk.choices[0].delta.content:
|
||||
print(chunk.choices[0].delta.content, end="")
|
||||
async def main():
|
||||
response = await litellm.acompletion(
|
||||
model="openai/gpt-4o-mini",
|
||||
messages=[
|
||||
{"role": "user", "content": "Implement a distributed rate limiter."}
|
||||
],
|
||||
tools=[
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "litellm_advisor",
|
||||
"description": "Consult a stronger advisor model.",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {"question": {"type": "string"}},
|
||||
"required": ["question"],
|
||||
},
|
||||
},
|
||||
}
|
||||
],
|
||||
max_tokens=4096,
|
||||
stream=True,
|
||||
)
|
||||
|
||||
async for chunk in response:
|
||||
if chunk.choices[0].delta.content:
|
||||
print(chunk.choices[0].delta.content, end="")
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
:::note Streaming behavior
|
||||
|
||||
The advisor sub-inference does not stream. The executor's stream pauses while the advisor runs, then the full advisor result arrives in a single event. Executor output resumes streaming afterward.
|
||||
|
||||
:::
|
||||
|
||||
#### Multi-Turn Conversation
|
||||
|
||||
```python showLineNumbers title="Multi-Turn with Advisor Tool"
|
||||
import litellm
|
||||
|
||||
tools = [
|
||||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
}
|
||||
]
|
||||
|
||||
messages = [
|
||||
{"role": "user", "content": "Build a concurrent worker pool in Go with graceful shutdown."}
|
||||
]
|
||||
|
||||
response = litellm.completion(
|
||||
model="anthropic/claude-sonnet-4-6",
|
||||
messages=messages,
|
||||
tools=tools,
|
||||
max_tokens=4096,
|
||||
)
|
||||
|
||||
# Append the full response (includes server_tool_use + advisor_tool_result blocks)
|
||||
messages.append({"role": "assistant", "content": response.choices[0].message.content})
|
||||
|
||||
# Continue the conversation — keep the same tools array
|
||||
messages.append({"role": "user", "content": "Now add a max-in-flight limit of 10."})
|
||||
|
||||
response2 = litellm.completion(
|
||||
model="anthropic/claude-sonnet-4-6",
|
||||
messages=messages,
|
||||
tools=tools,
|
||||
max_tokens=4096,
|
||||
)
|
||||
```
|
||||
|
||||
:::tip Auto-strip on follow-up turns
|
||||
|
||||
LiteLLM automatically strips `advisor_tool_result` blocks from message history when the advisor tool is not present in the current request. This prevents the Anthropic 400 error that would otherwise occur.
|
||||
The advisor sub-inference does not stream. When the executor calls the advisor tool, the stream pauses, the advisor runs to completion, and its output is injected before the executor resumes streaming.
|
||||
|
||||
:::
|
||||
|
||||
</TabItem>
|
||||
<TabItem value="chat-completions-proxy" label="Proxy">
|
||||
|
||||
#### Proxy Configuration
|
||||
### Configuring the advisor model (Proxy)
|
||||
|
||||
Add the advisor as a named deployment in `model_list` and reference it in `advisor_interception_params`. The proxy router resolves the correct credentials automatically.
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
model_list:
|
||||
# Advisor model — can be any provider
|
||||
- model_name: my-advisor
|
||||
litellm_params:
|
||||
model: openai/gpt-4o
|
||||
api_key: os.environ/OPENAI_API_KEY
|
||||
|
||||
# Or use an Anthropic model as advisor
|
||||
# - model_name: my-advisor
|
||||
# litellm_params:
|
||||
# model: anthropic/claude-opus-4-6
|
||||
# api_key: os.environ/ANTHROPIC_API_KEY
|
||||
|
||||
# Executor models
|
||||
- model_name: claude-sonnet
|
||||
litellm_params:
|
||||
model: anthropic/claude-sonnet-4-6
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
|
||||
- model_name: gpt-4o-mini
|
||||
litellm_params:
|
||||
model: openai/gpt-4o-mini
|
||||
api_key: os.environ/OPENAI_API_KEY
|
||||
|
||||
- model_name: gemini-flash
|
||||
litellm_params:
|
||||
model: vertex_ai/gemini-2.5-flash
|
||||
vertex_project: my-project
|
||||
vertex_location: us-central1
|
||||
|
||||
litellm_settings:
|
||||
callbacks: ["advisor_interception"] # use callbacks, not success_callback
|
||||
advisor_interception_params:
|
||||
# Must be a model_name from model_list above.
|
||||
# The router uses this to pick the right deployment + credentials.
|
||||
default_advisor_model: "my-advisor"
|
||||
```
|
||||
|
||||
#### Client Request via Proxy
|
||||
:::info
|
||||
|
||||
```python showLineNumbers title="Advisor Tool via AI Gateway"
|
||||
- Use `callbacks`, not `success_callback`. The advisor hooks run through `litellm.callbacks`.
|
||||
- `default_advisor_model` must match a `model_name` from `model_list`. This is how the proxy resolves the correct API key and deployment for the advisor sub-call.
|
||||
- You can still override it per-request by passing `model` in an `advisor_20260301` tool definition.
|
||||
|
||||
:::
|
||||
|
||||
---
|
||||
|
||||
### Client request — native advisor format
|
||||
|
||||
```python showLineNumbers title="Advisor via proxy (advisor_20260301 format)"
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
api_key="your-litellm-proxy-key",
|
||||
base_url="http://0.0.0.0:4000/v1"
|
||||
base_url="http://0.0.0.0:4000/v1",
|
||||
)
|
||||
|
||||
response = client.chat.completions.create(
|
||||
|
|
@ -268,18 +361,19 @@ response = client.chat.completions.create(
|
|||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
"model": "my-advisor", # matches model_name in config.yaml
|
||||
}
|
||||
],
|
||||
max_tokens=4096,
|
||||
)
|
||||
print(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
#### Client Request via Proxy (OpenAI-compatible function tool)
|
||||
### Client request — OpenAI function format
|
||||
|
||||
Use this format when your chat-completions client sends OpenAI-style tools.
|
||||
Use this when your client cannot send custom `type` values (e.g. plain OpenAI SDK). The proxy uses `default_advisor_model` from config.
|
||||
|
||||
```python showLineNumbers title="Proxy Chat Completions with litellm_advisor"
|
||||
```python showLineNumbers title="Advisor via proxy (litellm_advisor function format)"
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
|
|
@ -288,9 +382,9 @@ client = OpenAI(
|
|||
)
|
||||
|
||||
response = client.chat.completions.create(
|
||||
model="gemini-flash",
|
||||
model="gpt-4o-mini",
|
||||
messages=[
|
||||
{"role": "user", "content": "Call advisor once, then answer in one line: integration ok."}
|
||||
{"role": "user", "content": "Design a fault-tolerant task queue in Python."}
|
||||
],
|
||||
tools=[
|
||||
{
|
||||
|
|
@ -300,26 +394,18 @@ response = client.chat.completions.create(
|
|||
"description": "Consult a stronger advisor model.",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"question": {"type": "string"}
|
||||
},
|
||||
"properties": {"question": {"type": "string"}},
|
||||
"required": ["question"],
|
||||
},
|
||||
},
|
||||
}
|
||||
],
|
||||
max_tokens=512,
|
||||
max_tokens=2048,
|
||||
)
|
||||
print(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
::::note
|
||||
|
||||
For non-Anthropic chat-completions providers behind proxy, this OpenAI-compatible
|
||||
`litellm_advisor` function tool is the recommended request shape.
|
||||
The advisor model defaults to `claude-opus-4-6` unless overridden by your integration config.
|
||||
|
||||
::::
|
||||
The proxy intercepts the `litellm_advisor` tool call, calls `my-advisor` (from config), injects the result, and returns the final answer — your client only sees the finished response.
|
||||
|
||||
</TabItem>
|
||||
</Tabs>
|
||||
|
|
@ -331,9 +417,11 @@ The advisor model defaults to `claude-opus-4-6` unless overridden by your integr
|
|||
<Tabs>
|
||||
<TabItem value="messages-sdk" label="SDK">
|
||||
|
||||
#### Basic Example
|
||||
The Messages API (`litellm.anthropic.messages`) has built-in interception — no callback registration needed. Pass the `advisor_20260301` tool with the `model` field and LiteLLM handles the rest.
|
||||
|
||||
```python showLineNumbers title="Advisor Tool — litellm.anthropic.messages"
|
||||
#### Anthropic executor — native path
|
||||
|
||||
```python showLineNumbers title="Advisor Tool — Messages API, Anthropic native"
|
||||
import asyncio
|
||||
import litellm
|
||||
|
||||
|
|
@ -347,7 +435,7 @@ async def main():
|
|||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
"model": "claude-opus-4-6", # Anthropic runs this natively
|
||||
}
|
||||
],
|
||||
max_tokens=4096,
|
||||
|
|
@ -357,6 +445,36 @@ async def main():
|
|||
asyncio.run(main())
|
||||
```
|
||||
|
||||
#### Non-Anthropic executor — LiteLLM orchestration loop
|
||||
|
||||
When the executor is not Anthropic, or when the advisor model is not Claude Opus 4.6, LiteLLM runs the loop itself. The `model` field in the tool definition is the advisor — it can be any provider.
|
||||
|
||||
```python showLineNumbers title="Advisor Tool — Messages API, OpenAI executor"
|
||||
import asyncio
|
||||
import litellm
|
||||
|
||||
async def main():
|
||||
response = await litellm.anthropic.messages.acreate(
|
||||
model="openai/gpt-4o",
|
||||
messages=[
|
||||
{"role": "user", "content": "Implement a Python LRU cache with O(1) get and put."}
|
||||
],
|
||||
tools=[
|
||||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "openai/gpt-4o-mini", # advisor model — any provider works
|
||||
"max_uses": 2,
|
||||
}
|
||||
],
|
||||
max_tokens=1024,
|
||||
custom_llm_provider="openai",
|
||||
)
|
||||
print(response["content"][0]["text"])
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
#### Streaming
|
||||
|
||||
```python showLineNumbers title="Messages API Streaming with Advisor Tool"
|
||||
|
|
@ -396,24 +514,16 @@ asyncio.run(main())
|
|||
</TabItem>
|
||||
<TabItem value="messages-proxy" label="Proxy">
|
||||
|
||||
#### Proxy Configuration
|
||||
Use the same `config.yaml` shown in the Chat Completions proxy tab. The `advisor_interception_params` config applies to both APIs.
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
model_list:
|
||||
- model_name: claude-sonnet
|
||||
litellm_params:
|
||||
model: anthropic/claude-sonnet-4-6
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
```
|
||||
|
||||
#### Client Request via Proxy (Anthropic SDK)
|
||||
#### Client request — Anthropic SDK
|
||||
|
||||
```python showLineNumbers title="Advisor Tool via AI Gateway (Anthropic SDK)"
|
||||
import anthropic
|
||||
|
||||
client = anthropic.Anthropic(
|
||||
api_key="your-litellm-proxy-key",
|
||||
base_url="http://0.0.0.0:4000"
|
||||
base_url="http://0.0.0.0:4000",
|
||||
)
|
||||
|
||||
response = client.beta.messages.create(
|
||||
|
|
@ -427,92 +537,97 @@ response = client.beta.messages.create(
|
|||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
"model": "my-advisor", # model_name from config.yaml
|
||||
}
|
||||
],
|
||||
)
|
||||
print(response)
|
||||
```
|
||||
|
||||
#### Non-Anthropic Provider (LiteLLM orchestration loop)
|
||||
|
||||
```python showLineNumbers title="Advisor Tool with OpenAI executor"
|
||||
import asyncio
|
||||
import litellm
|
||||
|
||||
async def main():
|
||||
# executor: openai/gpt-4.1-mini | advisor: claude-opus-4-6
|
||||
# LiteLLM runs the orchestration loop automatically
|
||||
response = await litellm.anthropic.messages.acreate(
|
||||
model="openai/gpt-4.1-mini",
|
||||
messages=[
|
||||
{"role": "user", "content": "Implement a Python LRU cache with O(1) get and put."}
|
||||
],
|
||||
tools=[
|
||||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
"max_uses": 3,
|
||||
}
|
||||
],
|
||||
max_tokens=1024,
|
||||
custom_llm_provider="openai",
|
||||
)
|
||||
# Final response is clean — no advisor tool_use blocks
|
||||
print(response["content"][0]["text"])
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
</Tabs>
|
||||
|
||||
---
|
||||
|
||||
## Response Structure
|
||||
## Response structure
|
||||
|
||||
A successful advisor call returns `server_tool_use` and `advisor_tool_result` blocks in the assistant content:
|
||||
### Messages API
|
||||
|
||||
```json title="Response with advisor blocks"
|
||||
Both native and orchestration paths return `server_tool_use` and `advisor_tool_result` blocks in the assistant content:
|
||||
|
||||
```json title="Messages API response"
|
||||
{
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Let me consult the advisor on this."
|
||||
"text": "Here is the implementation:"
|
||||
},
|
||||
{
|
||||
"type": "server_tool_use",
|
||||
"id": "srvtoolu_abc123",
|
||||
"name": "advisor",
|
||||
"input": {}
|
||||
"name": "advisor"
|
||||
},
|
||||
{
|
||||
"type": "advisor_tool_result",
|
||||
"tool_use_id": "srvtoolu_abc123",
|
||||
"content": {
|
||||
"type": "advisor_result",
|
||||
"text": "Use a channel-based coordination pattern. The tricky part is draining in-flight work during shutdown: close the input channel first, then wait on a WaitGroup..."
|
||||
"text": "Use a channel-based coordination pattern..."
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Here's the implementation using a channel-based coordination pattern..."
|
||||
"text": "Here's the full implementation..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Pass the full assistant content, including advisor blocks, back on subsequent turns. LiteLLM handles this automatically through `provider_specific_fields`.
|
||||
### Chat Completions API
|
||||
|
||||
For chat completions, the advisor blocks are in `provider_specific_fields` on the response message:
|
||||
|
||||
```python title="Accessing advisor results from chat completions"
|
||||
response = await litellm.acompletion(...)
|
||||
|
||||
message = response.choices[0].message
|
||||
print(message.content) # final answer
|
||||
|
||||
# Advisor trace — available when the advisor was called
|
||||
psf = message.provider_specific_fields or {}
|
||||
for block in psf.get("advisor_tool_results", []):
|
||||
if block["type"] == "advisor_tool_result":
|
||||
print("Advisor said:", block["content"]["text"])
|
||||
```
|
||||
|
||||
```json title="provider_specific_fields structure"
|
||||
{
|
||||
"advisor_tool_results": [
|
||||
{
|
||||
"type": "server_tool_use",
|
||||
"id": "call_abc123",
|
||||
"name": "advisor"
|
||||
},
|
||||
{
|
||||
"type": "advisor_tool_result",
|
||||
"tool_use_id": "call_abc123",
|
||||
"content": {
|
||||
"type": "advisor_result",
|
||||
"text": "Use a channel-based coordination pattern..."
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Cost Control
|
||||
## Cost control
|
||||
|
||||
Advisor calls run as a separate sub-inference billed at the advisor model's rates. Usage is reported in `usage.iterations[]`:
|
||||
Advisor calls run as separate sub-inferences billed at the advisor model's rates. Usage is reported in `usage.iterations[]` (Messages API) or accumulated in `usage` (Chat Completions):
|
||||
|
||||
```json title="Usage with advisor sub-inference"
|
||||
```json title="Messages API usage with advisor sub-inference"
|
||||
{
|
||||
"usage": {
|
||||
"input_tokens": 412,
|
||||
|
|
@ -539,9 +654,22 @@ Advisor calls run as a separate sub-inference billed at the advisor model's rate
|
|||
}
|
||||
```
|
||||
|
||||
Top-level `usage` reflects executor tokens only. Advisor tokens appear in `iterations` entries with `type: "advisor_message"` and are billed at Opus rates.
|
||||
Use `max_uses` in the tool definition to cap how many times the advisor can be called per request:
|
||||
|
||||
## Additional Resources
|
||||
```python
|
||||
tools=[
|
||||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "my-advisor",
|
||||
"max_uses": 2, # raise AdvisorMaxIterationsError after 2 advisor calls
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Additional resources
|
||||
|
||||
- [Anthropic Advisor Tool Documentation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool)
|
||||
- [LiteLLM Tool Calling Guide](https://docs.litellm.ai/docs/completion/function_call)
|
||||
|
|
|
|||
|
|
@ -42,7 +42,9 @@ When a request arrives with an `advisor_20260301` tool and a non-Anthropic provi
|
|||
|
||||
## Model Compatibility
|
||||
|
||||
The executor and advisor models must form a valid pair. Currently the only supported advisor model is `claude-opus-4-6`.
|
||||
The advisor model is fully configurable. You can use any model deployed in your proxy as the advisor — it does not need to be Anthropic.
|
||||
|
||||
For **Anthropic-native** requests (where Anthropic runs the advisor server-side), the executor and advisor must form a valid Anthropic pair:
|
||||
|
||||
| Executor | Advisor |
|
||||
|----------|---------|
|
||||
|
|
@ -50,6 +52,8 @@ The executor and advisor models must form a valid pair. Currently the only suppo
|
|||
| `claude-sonnet-4-6` | `claude-opus-4-6` |
|
||||
| `claude-opus-4-6` | `claude-opus-4-6` |
|
||||
|
||||
For **non-Anthropic** executors (where LiteLLM orchestrates the advisor loop), you can use any model as the advisor — including OpenAI, Vertex AI, Bedrock, etc.
|
||||
|
||||
---
|
||||
|
||||
## Chat Completions API
|
||||
|
|
@ -226,14 +230,43 @@ LiteLLM automatically strips `advisor_tool_result` blocks from message history w
|
|||
|
||||
#### Proxy Configuration
|
||||
|
||||
Configure the advisor model as a deployment in your `model_list` and reference it in `advisor_interception_params`. This ensures the advisor sub-calls use the correct credentials and go through the proxy's deployment routing.
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
model_list:
|
||||
# The advisor model
|
||||
- model_name: advisor-model
|
||||
litellm_params:
|
||||
model: anthropic/claude-sonnet-4-20250514
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
|
||||
# Executor models
|
||||
- model_name: claude-sonnet
|
||||
litellm_params:
|
||||
model: anthropic/claude-sonnet-4-6
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
|
||||
- model_name: gemini-flash
|
||||
litellm_params:
|
||||
model: vertex_ai/gemini-2.5-flash
|
||||
vertex_project: my-project
|
||||
vertex_location: us-central1
|
||||
|
||||
litellm_settings:
|
||||
callbacks: ["advisor_interception"]
|
||||
advisor_interception_params:
|
||||
# Must match a model_name from model_list — the router resolves
|
||||
# the correct deployment and credentials automatically.
|
||||
default_advisor_model: "advisor-model"
|
||||
```
|
||||
|
||||
:::info Important
|
||||
|
||||
- Use `callbacks`, not `success_callback`. The advisor interception hooks run through `litellm.callbacks`.
|
||||
- The `default_advisor_model` value must be a `model_name` from your `model_list`. The proxy router resolves it to the correct deployment with the correct API key. This means you can use any provider as your advisor model — not just Anthropic.
|
||||
|
||||
:::
|
||||
|
||||
#### Client Request via Proxy
|
||||
|
||||
```python showLineNumbers title="Advisor Tool via AI Gateway"
|
||||
|
|
@ -253,7 +286,7 @@ response = client.chat.completions.create(
|
|||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
"model": "advisor-model",
|
||||
}
|
||||
],
|
||||
max_tokens=4096,
|
||||
|
|
@ -262,7 +295,7 @@ response = client.chat.completions.create(
|
|||
|
||||
#### Client Request via Proxy (OpenAI-compatible function tool)
|
||||
|
||||
Use this format when your chat-completions client sends OpenAI-style tools.
|
||||
Use this format when your chat-completions client sends OpenAI-style tools. The proxy uses the `default_advisor_model` from your config.
|
||||
|
||||
```python showLineNumbers title="Proxy Chat Completions with litellm_advisor"
|
||||
from openai import OpenAI
|
||||
|
|
@ -302,7 +335,7 @@ print(response.choices[0].message.content)
|
|||
|
||||
For non-Anthropic chat-completions providers behind proxy, this OpenAI-compatible
|
||||
`litellm_advisor` function tool is the recommended request shape.
|
||||
The advisor model defaults to `claude-opus-4-6` unless overridden by your integration config.
|
||||
The advisor model is determined by the `default_advisor_model` in your `advisor_interception_params` config.
|
||||
|
||||
::::
|
||||
|
||||
|
|
@ -383,12 +416,23 @@ asyncio.run(main())
|
|||
|
||||
#### Proxy Configuration
|
||||
|
||||
Use the same config shown in the Chat Completions proxy tab. The `advisor_interception_params` config applies to both APIs.
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
model_list:
|
||||
- model_name: advisor-model
|
||||
litellm_params:
|
||||
model: anthropic/claude-sonnet-4-20250514
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
- model_name: claude-sonnet
|
||||
litellm_params:
|
||||
model: anthropic/claude-sonnet-4-6
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
|
||||
litellm_settings:
|
||||
callbacks: ["advisor_interception"]
|
||||
advisor_interception_params:
|
||||
default_advisor_model: "advisor-model"
|
||||
```
|
||||
|
||||
#### Client Request via Proxy (Anthropic SDK)
|
||||
|
|
@ -412,7 +456,7 @@ response = client.beta.messages.create(
|
|||
{
|
||||
"type": "advisor_20260301",
|
||||
"name": "advisor",
|
||||
"model": "claude-opus-4-6",
|
||||
"model": "advisor-model",
|
||||
}
|
||||
],
|
||||
)
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue