chore: update docs

This commit is contained in:
Sameer Kankute 2026-04-14 18:33:51 +05:30
parent 7f23230f90
commit fdef44a58b
No known key found for this signature in database
3 changed files with 300 additions and 73 deletions

View file

@ -1,10 +1,11 @@
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
# Advisor Tool
Pair a faster executor model with a higher-intelligence advisor model that provides strategic guidance mid-generation.
LiteLLM now supports the Anthropic advisor tool across `chat/completions` and `messages` APIs (SDK + proxy).
The advisor tool lets a fast, lower-cost executor model (Sonnet or Haiku) consult a high-intelligence advisor model (Opus 4.6) mid-generation. The advisor reads the full conversation and produces a plan or course correction — typically 400–700 text tokens — and the executor continues with the task.
This pattern is well-suited for long-horizon agentic workloads (coding agents, computer use, multi-step research) where most turns are mechanical but having an excellent plan is crucial. You get close to advisor-solo quality while the bulk of token generation happens at executor-model rates.
Use the advisor tool to let an executor model call a stronger advisor model during generation. For non-Anthropic providers, LiteLLM runs the advisor orchestration loop automatically.
:::info Beta
@ -22,32 +23,14 @@ The advisor tool is in beta. Include `anthropic-beta: advisor-tool-2026-03-01` i
| **Google Vertex AI** | ✅ | ✅ | LiteLLM orchestration loop |
| **Groq / Mistral / others** | ✅ | ✅ | LiteLLM orchestration loop |
## How it works (LiteLLM native orchestration)
For non-Anthropic providers, LiteLLM implements the advisor loop itself.
For non-Anthropic providers, LiteLLM implements the advisor loop itself. The API you call is identical — LiteLLM handles everything transparently.
- **Messages API** (`litellm.anthropic.messages.create/acreate`): built-in interception in the messages handler
- **Chat Completions API** (`litellm.completion/acompletion`): enable `AdvisorInterceptionLogger` to convert advisor tools + run the loop
When a request arrives with an `advisor_20260301` tool and a non-Anthropic provider, `AdvisorOrchestrationHandler` intercepts it. It translates the advisor tool into a regular function tool the provider understands, then runs an orchestration loop:
When a request arrives with an `advisor_20260301` tool and a non-Anthropic provider, LiteLLM translates the advisor tool into a regular function tool the provider understands, then runs an orchestration loop:
```mermaid
flowchart TD
A["Your request\ntools: advisor_20260301\nmodel: e.g. openai/gpt-4.1-mini"] --> B["AdvisorOrchestrationHandler\ntranslates advisor → regular fn tool"]
B --> C["EXECUTOR CALL\nopenai / bedrock / vertex / etc."]
C --> D{"executor calls\nadvisor tool?"}
D -->|"yes — tool_use\nname=advisor"| E{"max_uses\nexceeded?"}
E -->|no| F["ADVISOR SUB-CALL\nclaude-opus-4-6\nfull transcript forwarded\nno tools"]
F --> G["Inject advice as\ntool_result into history"]
G --> C
E -->|yes| H["AdvisorMaxIterationsError"]
D -->|"no — end_turn\nor other stop reason"| I["Clean final response\nno advisor blocks in output"]
```
![Advisor Orchestration Flow](/img/advisor_orchestration_flow.svg)
**What LiteLLM does for you:**
@ -71,9 +54,10 @@ The executor and advisor models must form a valid pair. Currently the only suppo
## Chat Completions API
### SDK Usage
<Tabs>
<TabItem value="chat-completions-sdk" label="SDK">
#### Basic Example
#### Basic Example (Anthropic-native executor)
```python showLineNumbers title="Advisor Tool — litellm.completion()"
import litellm
@ -85,9 +69,18 @@ response = litellm.completion(
],
tools=[
{
"type": "advisor_20260301",
"name": "advisor",
"model": "claude-opus-4-6",
"type": "function",
"function": {
"name": "litellm_advisor",
"description": "Consult a stronger advisor model.",
"parameters": {
"type": "object",
"properties": {
"question": {"type": "string"}
},
"required": ["question"],
},
},
}
],
max_tokens=4096,
@ -96,6 +89,39 @@ response = litellm.completion(
print(response.choices[0].message.content)
```
#### Non-Anthropic Executor (Chat Completions interception)
```python showLineNumbers title="Advisor Tool with OpenAI executor via chat-completions"
import asyncio
import litellm
from litellm.integrations.advisor_interception import (
AdvisorInterceptionLogger,
get_litellm_advisor_tool,
)
litellm.callbacks = [AdvisorInterceptionLogger(enabled_providers=["openai"])]
async def main():
response = await litellm.acompletion(
model="gpt-5.4-mini",
messages=[
{"role": "user", "content": "Build a concurrent worker pool in Go with graceful shutdown."}
],
# You can still use Anthropic-native advisor tool format.
tools=[get_litellm_advisor_tool(model="claude-opus-4-6")],
max_tokens=4096,
)
print(response.choices[0].message.content)
asyncio.run(main())
```
::::note
`AdvisorInterceptionLogger` converts advisor tool definitions to provider-compatible function tools for non-Anthropic chat-completions providers and runs the advisor sub-call loop server-side.
::::
#### With Optional Parameters
```python showLineNumbers title="Advisor Tool with max_uses and caching"
@ -195,7 +221,8 @@ LiteLLM automatically strips `advisor_tool_result` blocks from message history w
:::
### AI Gateway Usage
</TabItem>
<TabItem value="chat-completions-proxy" label="Proxy">
#### Proxy Configuration
@ -233,11 +260,61 @@ response = client.chat.completions.create(
)
```
#### Client Request via Proxy (OpenAI-compatible function tool)
Use this format when your chat-completions client sends OpenAI-style tools.
```python showLineNumbers title="Proxy Chat Completions with litellm_advisor"
from openai import OpenAI
client = OpenAI(
api_key="your-litellm-proxy-key",
base_url="http://0.0.0.0:4000/v1",
)
response = client.chat.completions.create(
model="gemini-flash",
messages=[
{"role": "user", "content": "Call advisor once, then answer in one line: integration ok."}
],
tools=[
{
"type": "function",
"function": {
"name": "litellm_advisor",
"description": "Consult a stronger advisor model.",
"parameters": {
"type": "object",
"properties": {
"question": {"type": "string"}
},
"required": ["question"],
},
},
}
],
max_tokens=512,
)
print(response.choices[0].message.content)
```
::::note
For non-Anthropic chat-completions providers behind proxy, this OpenAI-compatible
`litellm_advisor` function tool is the recommended request shape.
The advisor model defaults to `claude-opus-4-6` unless overridden by your integration config.
::::
</TabItem>
</Tabs>
---
## Messages API
### SDK Usage
<Tabs>
<TabItem value="messages-sdk" label="SDK">
#### Basic Example
@ -301,7 +378,8 @@ async def main():
asyncio.run(main())
```
### AI Gateway Usage
</TabItem>
<TabItem value="messages-proxy" label="Proxy">
#### Proxy Configuration
@ -372,6 +450,9 @@ async def main():
asyncio.run(main())
```
</TabItem>
</Tabs>
---
## Response Structure
@ -445,44 +526,6 @@ Advisor calls run as a separate sub-inference billed at the advisor model's rate
Top-level `usage` reflects executor tokens only. Advisor tokens appear in `iterations` entries with `type: "advisor_message"` and are billed at Opus rates.
**Tips:**
- Enable `caching` on the tool definition only when you expect 3+ advisor calls per conversation; it costs more than it saves below that threshold.
- Use `max_uses` to cap advisor calls per request. Once reached, the executor continues without further advice.
- For conversation-level caps, count advisor calls client-side. When you reach your limit, remove the advisor tool from `tools`.
---
## Recommended System Prompt
For coding and agent tasks, Anthropic recommends prepending these blocks to your system prompt for consistent advisor timing and optimal cost/quality:
```text title="Timing guidance (prepend to system prompt)"
You have access to an `advisor` tool backed by a stronger reviewer model. It takes NO parameters — when you call advisor(), your entire conversation history is automatically forwarded. They see the task, every tool call you've made, every result you've seen.
Call advisor BEFORE substantive work — before writing, before committing to an interpretation, before building on an assumption. If the task requires orientation first (finding files, fetching a source, seeing what's there), do that, then call advisor. Orientation is not substantive work. Writing, editing, and declaring an answer are.
Also call advisor:
- When you believe the task is complete. BEFORE this call, make your deliverable durable: write the file, save the result, commit the change.
- When stuck — errors recurring, approach not converging, results that don't fit.
- When considering a change of approach.
On tasks longer than a few steps, call advisor at least once before committing to an approach and once before declaring done. On short reactive tasks where the next action is dictated by tool output you just read, you don't need to keep calling.
```
```text title="Advice weight guidance (add after timing block)"
Give the advice serious weight. If you follow a step and it fails empirically, or you have primary-source evidence that contradicts a specific claim, adapt. A passing self-test is not evidence the advice is wrong.
If you've already retrieved data pointing one way and the advisor points another: don't silently switch. Surface the conflict in one more advisor call — "I found X, you suggest Y, which constraint breaks the tie?"
```
To reduce advisor output length by 35–45% without losing quality, add:
```text title="Cost reduction (optional, add before timing block)"
The advisor should respond in under 100 words and use enumerated steps, not explanations.
```
---
## Additional Resources
- [Anthropic Advisor Tool Documentation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool)

View file

@ -0,0 +1,88 @@
<svg width="1200" height="760" viewBox="0 0 1200 760" xmlns="http://www.w3.org/2000/svg" role="img" aria-labelledby="title desc">
<title id="title">LiteLLM Advisor Orchestration Flow</title>
<desc id="desc">Flow for non-Anthropic providers: request conversion, executor call, advisor loop, and final response.</desc>
<defs>
<marker id="arrow" markerWidth="12" markerHeight="12" refX="10" refY="6" orient="auto">
<path d="M0,0 L12,6 L0,12 Z" fill="#475467" />
</marker>
<style>
.title { font: 700 30px -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif; fill: #111827; }
.subtitle { font: 500 16px -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif; fill: #475467; }
.box-title { font: 700 17px -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif; fill: #111827; }
.box-text { font: 500 14px -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif; fill: #344054; }
.decision-text { font: 700 15px -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif; fill: #111827; text-anchor: middle; }
.label { font: 600 13px -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif; fill: #475467; }
.stroke { stroke: #98A2B3; stroke-width: 2; }
.flow { fill: none; stroke: #475467; stroke-width: 2.5; marker-end: url(#arrow); }
</style>
</defs>
<rect x="0" y="0" width="1200" height="760" fill="#F8FAFC" />
<text x="52" y="58" class="title">Advisor Tool Orchestration (Non-Anthropic Providers)</text>
<text x="52" y="84" class="subtitle">LiteLLM converts advisor tools and runs the advisor loop server-side.</text>
<!-- Request -->
<rect x="70" y="130" width="300" height="110" rx="14" fill="#FFFFFF" class="stroke" />
<text x="92" y="164" class="box-title">Incoming Request</text>
<text x="92" y="190" class="box-text">tools: advisor_20260301</text>
<text x="92" y="212" class="box-text">model: openai/gpt-4.1-mini (example)</text>
<!-- Conversion -->
<rect x="470" y="130" width="360" height="130" rx="14" fill="#ECF3FF" class="stroke" />
<text x="492" y="164" class="box-title">AdvisorInterceptionLogger</text>
<text x="492" y="190" class="box-text">Converts advisor tool -&gt; standard function tool</text>
<text x="492" y="212" class="box-text">Stores advisor config (model, max_uses, creds)</text>
<text x="492" y="234" class="box-text">Disables raw provider stream while loop runs</text>
<!-- Executor call -->
<rect x="900" y="130" width="240" height="110" rx="14" fill="#FFFFFF" class="stroke" />
<text x="922" y="164" class="box-title">Executor Call</text>
<text x="922" y="190" class="box-text">OpenAI / Gemini / Bedrock / etc.</text>
<text x="922" y="212" class="box-text">Returns tool call or final text</text>
<!-- Decision 1 -->
<polygon points="1020,300 1120,380 1020,460 920,380" fill="#FFF7E6" class="stroke" />
<text x="1020" y="374" class="decision-text">Advisor</text>
<text x="1020" y="394" class="decision-text">tool called?</text>
<!-- Decision 2 -->
<polygon points="720,300 820,380 720,460 620,380" fill="#FFF7E6" class="stroke" />
<text x="720" y="374" class="decision-text">max_uses</text>
<text x="720" y="394" class="decision-text">exceeded?</text>
<!-- Advisor sub-call -->
<rect x="550" y="530" width="340" height="100" rx="14" fill="#ECFDF3" class="stroke" />
<text x="572" y="562" class="box-title">Advisor Sub-call</text>
<text x="572" y="586" class="box-text">model: claude-opus-4-6</text>
<text x="572" y="608" class="box-text">advisor response injected as tool result</text>
<!-- Error -->
<rect x="125" y="530" width="250" height="100" rx="14" fill="#FEF3F2" class="stroke" />
<text x="147" y="564" class="box-title">AdvisorMaxIterationsError</text>
<text x="147" y="588" class="box-text">Hard cap when max_uses is reached</text>
<!-- Final -->
<rect x="895" y="530" width="250" height="100" rx="14" fill="#EEF4FF" class="stroke" />
<text x="917" y="564" class="box-title">Final Response</text>
<text x="917" y="588" class="box-text">Clean assistant output</text>
<text x="917" y="610" class="box-text">No advisor-only blocks leaked</text>
<!-- Arrows -->
<path d="M370 185 L470 185" class="flow" />
<path d="M830 185 L900 185" class="flow" />
<path d="M1020 240 L1020 298" class="flow" />
<path d="M920 380 L822 380" class="flow" />
<text x="865" y="366" class="label">yes</text>
<path d="M620 380 L250 380 L250 528" class="flow" />
<text x="430" y="366" class="label">yes</text>
<path d="M720 460 L720 530" class="flow" />
<text x="734" y="496" class="label">no</text>
<path d="M720 630 L720 690 L860 690 L860 380 L920 380" class="flow" />
<path d="M1020 460 L1020 530" class="flow" />
<text x="1034" y="496" class="label">no</text>
</svg>

After

Width:  |  Height:  |  Size: 4.7 KiB

View file

@ -1,5 +1,6 @@
import pytest
import litellm
from litellm.integrations.advisor_interception.handler import AdvisorInterceptionLogger
from litellm.integrations.advisor_interception.tools import (
LITELLM_ADVISOR_TOOL_NAME,
@ -148,3 +149,98 @@ async def test_should_run_chat_completion_agentic_loop_detects_legacy_function_c
assert should_run is True
assert len(tools_dict["advisor_calls"]) == 1
assert tools_dict["advisor_calls"][0]["question"] == "Can you confirm advisor path?"
@pytest.mark.asyncio
async def test_run_chat_completion_agentic_loop_aggregates_subcall_costs(monkeypatch):
logger = AdvisorInterceptionLogger(enabled_providers=["openai"])
initial_response = ModelResponse(
id="initial",
choices=[
Choices(
finish_reason="tool_calls",
index=0,
message=Message(
role="assistant",
content=None,
tool_calls=[
ChatCompletionMessageToolCall(
id="call_abc",
type="function",
function=Function(
name=LITELLM_ADVISOR_TOOL_NAME,
arguments='{"question":"Need advisor guidance"}',
),
)
],
),
)
],
model="gpt-4o-mini",
object="chat.completion",
created=123,
)
initial_response._hidden_params["response_cost"] = 1.0
advisor_subcall_response = ModelResponse(
id="advisor-subcall",
choices=[
Choices(
finish_reason="stop",
index=0,
message=Message(
role="assistant",
content="Advisor says this looks good.",
),
)
],
model="claude-opus-4-6",
object="chat.completion",
created=124,
)
advisor_subcall_response._hidden_params["response_cost"] = 0.3
final_response = ModelResponse(
id="final",
choices=[
Choices(
finish_reason="stop",
index=0,
message=Message(
role="assistant",
content="integration ok.",
),
)
],
model="gpt-4o-mini",
object="chat.completion",
created=125,
)
final_response._hidden_params["response_cost"] = 0.7
calls = {"count": 0}
async def mock_acompletion(*args, **kwargs):
calls["count"] += 1
if calls["count"] == 1:
return advisor_subcall_response
if calls["count"] == 2:
return final_response
raise AssertionError("Unexpected extra acompletion call")
monkeypatch.setattr(litellm, "acompletion", mock_acompletion)
response = await logger.async_run_chat_completion_agentic_loop(
tools={"advisor_config": {"advisor_model": "claude-opus-4-6", "max_uses": 3}},
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Test"}],
response=initial_response,
optional_params={"tools": [get_litellm_advisor_tool_openai()], "max_tokens": 256},
logging_obj=None,
stream=False,
kwargs={"litellm_call_id": "cost-loop-1", "custom_llm_provider": "openai"},
)
assert calls["count"] == 2
assert response is final_response
assert response._hidden_params["response_cost"] == pytest.approx(2.0)