mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-10 03:28:53 +00:00
Merge branch 'BerriAI:main' into main
This commit is contained in:
commit
23b737d2a4
410 changed files with 4540 additions and 1631 deletions
|
|
@ -327,6 +327,81 @@ curl --location 'http://0.0.0.0:4000/v1/messages' \
|
|||
</TabItem>
|
||||
</Tabs>
|
||||
|
||||
## Usage - Azure Anthropic (Azure Foundry Claude)
|
||||
|
||||
LiteLLM funnels Azure Claude deployments through the `azure_ai/` provider so Claude Opus models on Azure Foundry keep working with Tool Search, Effort, streaming, and the rest of the advanced feature set. Point `AZURE_AI_API_BASE` to `https://<resource>.services.ai.azure.com/anthropic` (LiteLLM appends `/v1/messages` automatically) and authenticate with `AZURE_AI_API_KEY` or an Azure AD token.
|
||||
|
||||
<Tabs>
|
||||
<TabItem value="sdk" label="LiteLLM Python SDK">
|
||||
|
||||
```python
|
||||
import os
|
||||
from litellm import completion
|
||||
|
||||
# Configure Azure credentials
|
||||
os.environ["AZURE_AI_API_KEY"] = "your-azure-ai-api-key"
|
||||
os.environ["AZURE_AI_API_BASE"] = "https://my-resource.services.ai.azure.com/anthropic"
|
||||
|
||||
response = completion(
|
||||
model="azure_ai/claude-opus-4-1",
|
||||
messages=[{"role": "user", "content": "Explain how Azure Anthropic hosts Claude Opus differently from the public Anthropic API."}],
|
||||
max_tokens=1200,
|
||||
temperature=0.7,
|
||||
stream=True,
|
||||
)
|
||||
|
||||
for chunk in response:
|
||||
if chunk.choices[0].delta.content:
|
||||
print(chunk.choices[0].delta.content, end="", flush=True)
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
<TabItem value="proxy" label="LiteLLM Proxy">
|
||||
|
||||
**1. Set environment variables**
|
||||
|
||||
```bash
|
||||
export AZURE_AI_API_KEY="your-azure-ai-api-key"
|
||||
export AZURE_AI_API_BASE="https://my-resource.services.ai.azure.com/anthropic"
|
||||
```
|
||||
|
||||
**2. Configure the proxy**
|
||||
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: claude-4-azure
|
||||
litellm_params:
|
||||
model: azure_ai/claude-opus-4-1
|
||||
api_key: os.environ/AZURE_AI_API_KEY
|
||||
api_base: os.environ/AZURE_AI_API_BASE
|
||||
```
|
||||
|
||||
**3. Start LiteLLM**
|
||||
|
||||
```bash
|
||||
litellm --config /path/to/config.yaml
|
||||
```
|
||||
|
||||
**4. Test the Azure Claude route**
|
||||
|
||||
```bash
|
||||
curl --location 'http://0.0.0.0:4000/chat/completions' \
|
||||
--header 'Content-Type: application/json' \
|
||||
--header 'Authorization: Bearer $LITELLM_KEY' \
|
||||
--data '{
|
||||
"model": "claude-4-azure",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "How do I use Claude Opus 4 via Azure Anthropic in LiteLLM?"
|
||||
}
|
||||
],
|
||||
"max_tokens": 1024
|
||||
}'
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
</Tabs>
|
||||
|
||||
|
||||
## Tool Search {#tool-search}
|
||||
|
|
|
|||
22
docs/my-website/docs/projects/openai-agents.md
Normal file
22
docs/my-website/docs/projects/openai-agents.md
Normal file
|
|
@ -0,0 +1,22 @@
|
|||
|
||||
# OpenAI Agents SDK
|
||||
|
||||
The [OpenAI Agents SDK](https://github.com/openai/openai-agents-python) is a lightweight framework for building multi-agent workflows.
|
||||
It includes an official LiteLLM extension that lets you use any of the 100+ supported providers (Anthropic, Gemini, Mistral, Bedrock, etc.)
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.models.litellm_model import LitellmModel
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="You are a helpful assistant.",
|
||||
model=LitellmModel(model="provider/model-name")
|
||||
)
|
||||
|
||||
result = Runner.run_sync(agent, "your_prompt_here")
|
||||
print("Result:", result.final_output)
|
||||
```
|
||||
|
||||
- [GitHub](https://github.com/openai/openai-agents-python)
|
||||
- [LiteLLM Extension Docs](https://openai.github.io/openai-agents-python/ref/extensions/litellm/)
|
||||
|
|
@ -16,7 +16,7 @@ Azure Foundry supports the following Claude models:
|
|||
| Property | Details |
|
||||
|-------|-------|
|
||||
| Description | Claude models deployed via Microsoft Azure Foundry. Uses the same API as Anthropic's Messages API but with Azure authentication. |
|
||||
| Provider Route on LiteLLM | `azure/` (add this prefix to Claude model names - e.g. `azure/claude-sonnet-4-5`) |
|
||||
| Provider Route on LiteLLM | `azure_ai/` (add this prefix to Claude model names - e.g. `azure_ai/claude-sonnet-4-5`) |
|
||||
| Provider Doc | [Azure Foundry Claude Models ↗](https://learn.microsoft.com/en-us/azure/ai-services/foundry-models/claude) |
|
||||
| API Endpoint | `https://<resource-name>.services.ai.azure.com/anthropic/v1/messages` |
|
||||
| Supported Endpoints | `/chat/completions`, `/anthropic/v1/messages`|
|
||||
|
|
@ -68,7 +68,7 @@ os.environ["AZURE_API_BASE"] = "https://<resource-name>.services.ai.azure.com/an
|
|||
|
||||
# Make a completion request
|
||||
response = completion(
|
||||
model="azure/claude-sonnet-4-5",
|
||||
model="azure_ai/claude-sonnet-4-5",
|
||||
messages=[
|
||||
{"role": "user", "content": "What are 3 things to visit in Seattle?"}
|
||||
],
|
||||
|
|
@ -85,7 +85,7 @@ print(response)
|
|||
import litellm
|
||||
|
||||
response = litellm.completion(
|
||||
model="azure/claude-sonnet-4-5",
|
||||
model="azure_ai/claude-sonnet-4-5",
|
||||
api_base="https://<resource-name>.services.ai.azure.com/anthropic",
|
||||
api_key="your-azure-api-key",
|
||||
messages=[
|
||||
|
|
@ -101,7 +101,7 @@ response = litellm.completion(
|
|||
import litellm
|
||||
|
||||
response = litellm.completion(
|
||||
model="azure/claude-sonnet-4-5",
|
||||
model="azure_ai/claude-sonnet-4-5",
|
||||
api_base="https://<resource-name>.services.ai.azure.com/anthropic",
|
||||
azure_ad_token="your-azure-ad-token",
|
||||
messages=[
|
||||
|
|
@ -117,7 +117,7 @@ response = litellm.completion(
|
|||
from litellm import completion
|
||||
|
||||
response = completion(
|
||||
model="azure/claude-sonnet-4-5",
|
||||
model="azure_ai/claude-sonnet-4-5",
|
||||
messages=[
|
||||
{"role": "user", "content": "Write a short story"}
|
||||
],
|
||||
|
|
@ -136,7 +136,7 @@ for chunk in response:
|
|||
from litellm import completion
|
||||
|
||||
response = completion(
|
||||
model="azure/claude-sonnet-4-5",
|
||||
model="azure_ai/claude-sonnet-4-5",
|
||||
messages=[
|
||||
{"role": "user", "content": "What's the weather in Seattle?"}
|
||||
],
|
||||
|
|
@ -181,7 +181,7 @@ export AZURE_API_BASE="https://<resource-name>.services.ai.azure.com/anthropic"
|
|||
model_list:
|
||||
- model_name: claude-sonnet-4-5
|
||||
litellm_params:
|
||||
model: azure/claude-sonnet-4-5
|
||||
model: azure_ai/claude-sonnet-4-5
|
||||
api_base: https://<resource-name>.services.ai.azure.com/anthropic
|
||||
api_key: os.environ/AZURE_API_KEY
|
||||
```
|
||||
|
|
@ -331,7 +331,7 @@ os.environ["AZURE_API_BASE"] = "https://my-resource.services.ai.azure.com/anthro
|
|||
|
||||
# Make a request
|
||||
response = completion(
|
||||
model="azure/claude-sonnet-4-5",
|
||||
model="azure_ai/claude-sonnet-4-5",
|
||||
messages=[
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": "Explain quantum computing in simple terms."}
|
||||
|
|
@ -358,7 +358,7 @@ Or pass it directly:
|
|||
|
||||
```python
|
||||
response = completion(
|
||||
model="azure/claude-sonnet-4-5",
|
||||
model="azure_ai/claude-sonnet-4-5",
|
||||
api_base="https://<resource-name>.services.ai.azure.com/anthropic",
|
||||
# ...
|
||||
)
|
||||
|
|
|
|||
209
docs/my-website/docs/providers/publicai.md
Normal file
209
docs/my-website/docs/providers/publicai.md
Normal file
|
|
@ -0,0 +1,209 @@
|
|||
import Tabs from '@theme/Tabs';
|
||||
import TabItem from '@theme/TabItem';
|
||||
|
||||
# PublicAI
|
||||
|
||||
## Overview
|
||||
|
||||
| Property | Details |
|
||||
|-------|-------|
|
||||
| Description | PublicAI provides large language models including essential models like the swiss-ai apertus model. |
|
||||
| Provider Route on LiteLLM | `publicai/` |
|
||||
| Link to Provider Doc | [PublicAI ↗](https://platform.publicai.co/) |
|
||||
| Base URL | `https://platform.publicai.co/` |
|
||||
| Supported Operations | [`/chat/completions`](#sample-usage) |
|
||||
|
||||
<br />
|
||||
<br />
|
||||
|
||||
https://platform.publicai.co/
|
||||
|
||||
**We support ALL PublicAI models, just set `publicai/` as a prefix when sending completion requests**
|
||||
|
||||
## Required Variables
|
||||
|
||||
```python showLineNumbers title="Environment Variables"
|
||||
os.environ["PUBLICAI_API_KEY"] = "" # your PublicAI API key
|
||||
```
|
||||
|
||||
You can overwrite the base url with:
|
||||
|
||||
```
|
||||
os.environ["PUBLICAI_API_BASE"] = "https://platform.publicai.co/v1"
|
||||
```
|
||||
|
||||
## Usage - LiteLLM Python SDK
|
||||
|
||||
### Non-streaming
|
||||
|
||||
```python showLineNumbers title="PublicAI Non-streaming Completion"
|
||||
import os
|
||||
import litellm
|
||||
from litellm import completion
|
||||
|
||||
os.environ["PUBLICAI_API_KEY"] = "" # your PublicAI API key
|
||||
|
||||
messages = [{"content": "Hello, how are you?", "role": "user"}]
|
||||
|
||||
# PublicAI call
|
||||
response = completion(
|
||||
model="publicai/swiss-ai/apertus-8b-instruct",
|
||||
messages=messages
|
||||
)
|
||||
|
||||
print(response)
|
||||
```
|
||||
|
||||
### Streaming
|
||||
|
||||
```python showLineNumbers title="PublicAI Streaming Completion"
|
||||
import os
|
||||
import litellm
|
||||
from litellm import completion
|
||||
|
||||
os.environ["PUBLICAI_API_KEY"] = "" # your PublicAI API key
|
||||
|
||||
messages = [{"content": "Hello, how are you?", "role": "user"}]
|
||||
|
||||
# PublicAI call with streaming
|
||||
response = completion(
|
||||
model="publicai/swiss-ai/apertus-8b-instruct",
|
||||
messages=messages,
|
||||
stream=True
|
||||
)
|
||||
|
||||
for chunk in response:
|
||||
print(chunk)
|
||||
```
|
||||
|
||||
## Usage - LiteLLM Proxy
|
||||
|
||||
Add the following to your LiteLLM Proxy configuration file:
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
model_list:
|
||||
- model_name: swiss-ai-apertus-8b
|
||||
litellm_params:
|
||||
model: publicai/swiss-ai/apertus-8b-instruct
|
||||
api_key: os.environ/PUBLICAI_API_KEY
|
||||
|
||||
- model_name: swiss-ai-apertus-70b
|
||||
litellm_params:
|
||||
model: publicai/swiss-ai/apertus-70b-instruct
|
||||
api_key: os.environ/PUBLICAI_API_KEY
|
||||
```
|
||||
|
||||
Start your LiteLLM Proxy server:
|
||||
|
||||
```bash showLineNumbers title="Start LiteLLM Proxy"
|
||||
litellm --config config.yaml
|
||||
|
||||
# RUNNING on http://0.0.0.0:4000
|
||||
```
|
||||
|
||||
<Tabs>
|
||||
<TabItem value="openai-sdk" label="OpenAI SDK">
|
||||
|
||||
```python showLineNumbers title="PublicAI via Proxy - Non-streaming"
|
||||
from openai import OpenAI
|
||||
|
||||
# Initialize client with your proxy URL
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:4000", # Your proxy URL
|
||||
api_key="your-proxy-api-key" # Your proxy API key
|
||||
)
|
||||
|
||||
# Non-streaming response
|
||||
response = client.chat.completions.create(
|
||||
model="swiss-ai-apertus-8b",
|
||||
messages=[{"role": "user", "content": "hello from litellm"}]
|
||||
)
|
||||
|
||||
print(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
```python showLineNumbers title="PublicAI via Proxy - Streaming"
|
||||
from openai import OpenAI
|
||||
|
||||
# Initialize client with your proxy URL
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:4000", # Your proxy URL
|
||||
api_key="your-proxy-api-key" # Your proxy API key
|
||||
)
|
||||
|
||||
# Streaming response
|
||||
response = client.chat.completions.create(
|
||||
model="swiss-ai-apertus-8b",
|
||||
messages=[{"role": "user", "content": "hello from litellm"}],
|
||||
stream=True
|
||||
)
|
||||
|
||||
for chunk in response:
|
||||
if chunk.choices[0].delta.content is not None:
|
||||
print(chunk.choices[0].delta.content, end="")
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
|
||||
<TabItem value="litellm-sdk" label="LiteLLM SDK">
|
||||
|
||||
```python showLineNumbers title="PublicAI via Proxy - LiteLLM SDK"
|
||||
import litellm
|
||||
|
||||
# Configure LiteLLM to use your proxy
|
||||
response = litellm.completion(
|
||||
model="litellm_proxy/swiss-ai-apertus-8b",
|
||||
messages=[{"role": "user", "content": "hello from litellm"}],
|
||||
api_base="http://localhost:4000",
|
||||
api_key="your-proxy-api-key"
|
||||
)
|
||||
|
||||
print(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
```python showLineNumbers title="PublicAI via Proxy - LiteLLM SDK Streaming"
|
||||
import litellm
|
||||
|
||||
# Configure LiteLLM to use your proxy with streaming
|
||||
response = litellm.completion(
|
||||
model="litellm_proxy/swiss-ai-apertus-8b",
|
||||
messages=[{"role": "user", "content": "hello from litellm"}],
|
||||
api_base="http://localhost:4000",
|
||||
api_key="your-proxy-api-key",
|
||||
stream=True
|
||||
)
|
||||
|
||||
for chunk in response:
|
||||
if hasattr(chunk.choices[0], 'delta') and chunk.choices[0].delta.content is not None:
|
||||
print(chunk.choices[0].delta.content, end="")
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
|
||||
<TabItem value="curl" label="cURL">
|
||||
|
||||
```bash showLineNumbers title="PublicAI via Proxy - cURL"
|
||||
curl http://localhost:4000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer your-proxy-api-key" \
|
||||
-d '{
|
||||
"model": "swiss-ai-apertus-8b",
|
||||
"messages": [{"role": "user", "content": "hello from litellm"}]
|
||||
}'
|
||||
```
|
||||
|
||||
```bash showLineNumbers title="PublicAI via Proxy - cURL Streaming"
|
||||
curl http://localhost:4000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer your-proxy-api-key" \
|
||||
-d '{
|
||||
"model": "swiss-ai-apertus-8b",
|
||||
"messages": [{"role": "user", "content": "hello from litellm"}],
|
||||
"stream": true
|
||||
}'
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
</Tabs>
|
||||
|
||||
For more detailed information on using the LiteLLM Proxy, see the [LiteLLM Proxy documentation](../providers/litellm_proxy).
|
||||
|
|
@ -576,6 +576,8 @@ router_settings:
|
|||
| GENERIC_USER_PROVIDER_ATTRIBUTE | Attribute specifying the user's provider
|
||||
| GENERIC_USER_ROLE_ATTRIBUTE | Attribute specifying the user's role
|
||||
| GENERIC_USERINFO_ENDPOINT | Endpoint to fetch user information in generic OAuth
|
||||
| GENERIC_LOGGER_ENDPOINT | Endpoint URL for the Generic Logger callback to send logs to
|
||||
| GENERIC_LOGGER_HEADERS | JSON string of headers to include in Generic Logger callback requests
|
||||
| GEMINI_API_BASE | Base URL for Gemini API. Default is https://generativelanguage.googleapis.com
|
||||
| GALILEO_BASE_URL | Base URL for Galileo platform
|
||||
| GALILEO_PASSWORD | Password for Galileo authentication
|
||||
|
|
|
|||
|
|
@ -7,9 +7,38 @@ import TabItem from '@theme/TabItem';
|
|||
LiteLLM provides the LiteLLM Tool Permission Guardrail that lets you control which **tool calls** a model is allowed to invoke, using configurable allow/deny rules. This offers fine-grained, provider-agnostic control over tool execution (e.g., OpenAI Chat Completions `tool_calls`, Anthropic Messages `tool_use`, MCP tools).
|
||||
|
||||
## Quick Start
|
||||
### 1. Define Guardrails on your LiteLLM config.yaml
|
||||
|
||||
Define your guardrails under the `guardrails` section
|
||||
### LiteLLM UI
|
||||
|
||||
#### Step 1: Select Tool Permission Guardrail
|
||||
|
||||
Open the LiteLLM Dashboard, click **Add New Guardrail**, and choose **LiteLLM Tool Permission Guardrail**. This loads the rule builder UI.
|
||||
|
||||
<Image img={require('../../../img/create_guard_tool_permission.png')} alt="Configure tool permission guardrail in LiteLLM UI" />
|
||||
|
||||
#### Step 2: Define Regex Rules
|
||||
|
||||
1. Click **Add Rule**.
|
||||
2. Enter a unique Rule ID.
|
||||
3. Provide a regex for the tool name (e.g., `^mcp__github_.*$`).
|
||||
4. Optionally add a regex for tool type (e.g., `^function$`).
|
||||
5. Pick **Allow** or **Deny**.
|
||||
|
||||
<Image img={require('../../../img/create_rule_tool_permission.png')} alt="Configure tool permission guardrail in LiteLLM UI" />
|
||||
|
||||
#### Step 3: Restrict Tool Arguments (Optional)
|
||||
|
||||
Select **+ Restrict tool arguments** to attach regex validations to nested paths (dot + `[]` notation). This enforces that sensitive parameters (such as `arguments.to[]`) conform to pre-approved formats.
|
||||
|
||||
#### Step 4: Choose Defaults & Actions
|
||||
|
||||
- Set the fallback decision (`default_action`) for tools that do not hit any rule.
|
||||
- Decide how disallowed tools behave: **Block** halts the request, **Rewrite** strips forbidden tools and returns an error message inside the response.
|
||||
- Customize `violation_message_template` if you want branded error copy.
|
||||
- Save the guardrail.
|
||||
|
||||
### LiteLLM Config.yaml Setup
|
||||
|
||||
```yaml
|
||||
guardrails:
|
||||
- guardrail_name: "tool-permission-guardrail"
|
||||
|
|
@ -21,16 +50,17 @@ guardrails:
|
|||
tool_name: "Bash"
|
||||
decision: "allow"
|
||||
- id: "allow_github_mcp"
|
||||
tool_name: "mcp__github_*"
|
||||
tool_name: "^mcp__github_.*$"
|
||||
decision: "allow"
|
||||
- id: "allow_aws_documentation"
|
||||
tool_name: "mcp__aws-documentation_*_documentation"
|
||||
tool_name: "^mcp__aws-documentation_.*_documentation$"
|
||||
decision: "allow"
|
||||
- id: "deny_read_commands"
|
||||
tool_name: "Read"
|
||||
decision: "Deny"
|
||||
decision: "deny"
|
||||
- id: "mail-domain"
|
||||
tool_name: "send_email"
|
||||
tool_name: "^send_email$"
|
||||
tool_type: "^function$"
|
||||
decision: "allow"
|
||||
allowed_param_patterns:
|
||||
"to[]": "^.+@berri\\.ai$"
|
||||
|
|
@ -44,7 +74,8 @@ guardrails:
|
|||
|
||||
```yaml
|
||||
- id: "unique_rule_id" # Unique identifier for the rule
|
||||
tool_name: "pattern" # Tool name or pattern to match
|
||||
tool_name: "^regex$" # Regex for tool name (optional, at least one of name/type required)
|
||||
tool_type: "^function$" # Regex for tool type (optional)
|
||||
decision: "allow" # "allow" or "deny"
|
||||
allowed_param_patterns: # Optional - regex map for argument paths (dot + [] notation)
|
||||
"path.to[].field": "^regex$"
|
||||
|
|
|
|||
250
docs/my-website/docs/proxy/pass_through_guardrails.md
Normal file
250
docs/my-website/docs/proxy/pass_through_guardrails.md
Normal file
|
|
@ -0,0 +1,250 @@
|
|||
# Guardrails on Pass-Through Endpoints
|
||||
|
||||
import Image from '@theme/IdealImage';
|
||||
|
||||
## Overview
|
||||
|
||||
| Property | Details |
|
||||
|----------|---------|
|
||||
| Description | Enable guardrail execution on LiteLLM pass-through endpoints with opt-in activation and automatic inheritance from org/team/key levels |
|
||||
| Supported Guardrails | All LiteLLM guardrails (Bedrock, Aporia, Lakera, etc.) |
|
||||
| Default Behavior | Guardrails are **disabled** on pass-through endpoints unless explicitly enabled |
|
||||
|
||||
## Quick Start
|
||||
|
||||
You can configure guardrails on pass-through endpoints either via the **UI** (recommended) or **config file**.
|
||||
|
||||
### Using the UI
|
||||
|
||||
#### 1. Navigate to Pass-Through Endpoints
|
||||
|
||||
Go to **Models + Endpoints** → Click **+ Add Pass-Through Endpoint**
|
||||
|
||||
<Image img={require('../../img/pt_guard1.png')} alt="Add guardrails to pass-through endpoint" />
|
||||
|
||||
Scroll to the **Guardrails** section and select which guardrails to enforce.
|
||||
|
||||
:::tip Default Behavior
|
||||
By default, you don't need to specify fields - LiteLLM will JSON dump the entire request/response payload and send it to the guardrail.
|
||||
:::
|
||||
|
||||
#### 2. Target Specific Fields (Optional)
|
||||
|
||||
<Image img={require('../../img/pt_guard2.png')} alt="Configure field-level targeting" />
|
||||
|
||||
To check only specific fields instead of the entire payload:
|
||||
|
||||
1. Select your guardrails
|
||||
2. In **Field Targeting (Optional)**, specify fields for each guardrail
|
||||
3. Use the quick-add buttons (`+ query`, `+ documents[*]`) or type custom JSONPath expressions
|
||||
4. **Request Fields (pre_call)**: Fields to check before sending to target API
|
||||
5. **Response Fields (post_call)**: Fields to check in the response from target API
|
||||
|
||||
**Example**: In the screenshot above, we set `query` as a request field, so only the `query` field is sent to the guardrail instead of the entire request.
|
||||
|
||||
---
|
||||
|
||||
### Using Config File
|
||||
|
||||
#### 1. Define guardrails and pass-through endpoint
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
guardrails:
|
||||
- guardrail_name: "pii-guard"
|
||||
litellm_params:
|
||||
guardrail: bedrock
|
||||
mode: pre_call
|
||||
guardrailIdentifier: "your-guardrail-id"
|
||||
guardrailVersion: "1"
|
||||
|
||||
general_settings:
|
||||
pass_through_endpoints:
|
||||
- path: "/v1/rerank"
|
||||
target: "https://api.cohere.com/v1/rerank"
|
||||
headers:
|
||||
Authorization: "bearer os.environ/COHERE_API_KEY"
|
||||
guardrails:
|
||||
pii-guard:
|
||||
```
|
||||
|
||||
#### 2. Start proxy
|
||||
|
||||
```bash
|
||||
litellm --config config.yaml
|
||||
```
|
||||
|
||||
#### 3. Test request
|
||||
|
||||
```bash
|
||||
curl -X POST "http://localhost:4000/v1/rerank" \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer sk-1234" \
|
||||
-d '{
|
||||
"model": "rerank-english-v3.0",
|
||||
"query": "What is the capital of France?",
|
||||
"documents": ["Paris is the capital of France."]
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Opt-In Behavior
|
||||
|
||||
| Configuration | Behavior |
|
||||
|--------------|----------|
|
||||
| `guardrails` not set | No guardrails execute (default) |
|
||||
| `guardrails` set | All org/team/key + pass-through guardrails execute |
|
||||
|
||||
When guardrails are enabled, the system collects and executes:
|
||||
- Org-level guardrails
|
||||
- Team-level guardrails
|
||||
- Key-level guardrails
|
||||
- Pass-through specific guardrails
|
||||
|
||||
---
|
||||
|
||||
|
||||
## How It Works
|
||||
|
||||
The diagram below shows what happens when a client makes a request to `/special/rerank` - a pass-through endpoint configured with guardrails in your `config.yaml`.
|
||||
|
||||
When guardrails are configured on a pass-through endpoint:
|
||||
1. **Pre-call guardrails** run on the request before forwarding to the target API
|
||||
2. If `request_fields` is specified (e.g., `["query"]`), only those fields are sent to the guardrail. Otherwise, the entire request payload is evaluated.
|
||||
3. The request is forwarded to the target API only if guardrails pass
|
||||
4. **Post-call guardrails** run on the response from the target API
|
||||
5. If `response_fields` is specified (e.g., `["results[*].text"]`), only those fields are evaluated. Otherwise, the entire response is checked.
|
||||
|
||||
:::info
|
||||
If the `guardrails` block is omitted or empty in your pass-through endpoint config, the request skips the guardrail flow entirely and goes directly to the target API.
|
||||
:::
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant Client
|
||||
box rgb(200, 220, 255) LiteLLM Proxy
|
||||
participant PassThrough as Pass-through Endpoint
|
||||
participant Guardrails
|
||||
end
|
||||
participant Target as Target API (Cohere, etc.)
|
||||
|
||||
Client->>PassThrough: POST /special/rerank
|
||||
Note over PassThrough,Guardrails: Collect passthrough + org/team/key guardrails
|
||||
PassThrough->>Guardrails: Run pre_call (request_fields or full payload)
|
||||
Guardrails-->>PassThrough: ✓ Pass / ✗ Block
|
||||
PassThrough->>Target: Forward request
|
||||
Target-->>PassThrough: Response
|
||||
PassThrough->>Guardrails: Run post_call (response_fields or full payload)
|
||||
Guardrails-->>PassThrough: ✓ Pass / ✗ Block
|
||||
PassThrough-->>Client: Return response (or error)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Field-Level Targeting
|
||||
|
||||
Target specific JSON fields instead of the entire request/response payload.
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
guardrails:
|
||||
- guardrail_name: "pii-detection"
|
||||
litellm_params:
|
||||
guardrail: bedrock
|
||||
mode: pre_call
|
||||
guardrailIdentifier: "pii-guard-id"
|
||||
guardrailVersion: "1"
|
||||
|
||||
- guardrail_name: "content-moderation"
|
||||
litellm_params:
|
||||
guardrail: bedrock
|
||||
mode: post_call
|
||||
guardrailIdentifier: "content-guard-id"
|
||||
guardrailVersion: "1"
|
||||
|
||||
general_settings:
|
||||
pass_through_endpoints:
|
||||
- path: "/v1/rerank"
|
||||
target: "https://api.cohere.com/v1/rerank"
|
||||
headers:
|
||||
Authorization: "bearer os.environ/COHERE_API_KEY"
|
||||
guardrails:
|
||||
pii-detection:
|
||||
request_fields: ["query", "documents[*].text"]
|
||||
content-moderation:
|
||||
response_fields: ["results[*].text"]
|
||||
```
|
||||
|
||||
### Field Options
|
||||
|
||||
| Field | Description |
|
||||
|-------|-------------|
|
||||
| `request_fields` | JSONPath expressions for input (pre_call) |
|
||||
| `response_fields` | JSONPath expressions for output (post_call) |
|
||||
| Neither specified | Guardrail runs on entire payload |
|
||||
|
||||
### JSONPath Examples
|
||||
|
||||
| Expression | Matches |
|
||||
|------------|---------|
|
||||
| `query` | Single field named `query` |
|
||||
| `documents[*].text` | All `text` fields in `documents` array |
|
||||
| `messages[*].content` | All `content` fields in `messages` array |
|
||||
|
||||
---
|
||||
|
||||
## Configuration Examples
|
||||
|
||||
### Single guardrail on entire payload
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
guardrails:
|
||||
- guardrail_name: "pii-detection"
|
||||
litellm_params:
|
||||
guardrail: bedrock
|
||||
mode: pre_call
|
||||
guardrailIdentifier: "your-id"
|
||||
guardrailVersion: "1"
|
||||
|
||||
general_settings:
|
||||
pass_through_endpoints:
|
||||
- path: "/v1/rerank"
|
||||
target: "https://api.cohere.com/v1/rerank"
|
||||
guardrails:
|
||||
pii-detection:
|
||||
```
|
||||
|
||||
### Multiple guardrails with mixed settings
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
guardrails:
|
||||
- guardrail_name: "pii-detection"
|
||||
litellm_params:
|
||||
guardrail: bedrock
|
||||
mode: pre_call
|
||||
guardrailIdentifier: "pii-id"
|
||||
guardrailVersion: "1"
|
||||
|
||||
- guardrail_name: "content-moderation"
|
||||
litellm_params:
|
||||
guardrail: bedrock
|
||||
mode: post_call
|
||||
guardrailIdentifier: "content-id"
|
||||
guardrailVersion: "1"
|
||||
|
||||
- guardrail_name: "prompt-injection"
|
||||
litellm_params:
|
||||
guardrail: lakera
|
||||
mode: pre_call
|
||||
api_key: os.environ/LAKERA_API_KEY
|
||||
|
||||
general_settings:
|
||||
pass_through_endpoints:
|
||||
- path: "/v1/rerank"
|
||||
target: "https://api.cohere.com/v1/rerank"
|
||||
guardrails:
|
||||
pii-detection:
|
||||
request_fields: ["input", "query"]
|
||||
content-moderation:
|
||||
prompt-injection:
|
||||
request_fields: ["messages[*].content"]
|
||||
```
|
||||
BIN
docs/my-website/img/pt_guard1.png
Normal file
BIN
docs/my-website/img/pt_guard1.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 769 KiB |
BIN
docs/my-website/img/pt_guard2.png
Normal file
BIN
docs/my-website/img/pt_guard2.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 548 KiB |
|
|
@ -419,7 +419,8 @@ const sidebars = {
|
|||
]
|
||||
},
|
||||
"pass_through/vllm",
|
||||
"proxy/pass_through"
|
||||
"proxy/pass_through",
|
||||
"proxy/pass_through_guardrails"
|
||||
]
|
||||
},
|
||||
"rag_ingest",
|
||||
|
|
@ -621,6 +622,7 @@ const sidebars = {
|
|||
"providers/ovhcloud",
|
||||
"providers/perplexity",
|
||||
"providers/petals",
|
||||
"providers/publicai",
|
||||
"providers/predibase",
|
||||
"providers/recraft",
|
||||
"providers/replicate",
|
||||
|
|
@ -814,9 +816,10 @@ const sidebars = {
|
|||
"Learn how to deploy + call models from different providers on LiteLLM",
|
||||
slug: "/project",
|
||||
},
|
||||
items: [
|
||||
items: [
|
||||
"projects/smolagents",
|
||||
"projects/mini-swe-agent",
|
||||
"projects/openai-agents",
|
||||
"projects/Docq.AI",
|
||||
"projects/PDL",
|
||||
"projects/OpenInterpreter",
|
||||
|
|
|
|||
|
|
@ -555,6 +555,7 @@ deepgram_models: Set = set()
|
|||
elevenlabs_models: Set = set()
|
||||
dashscope_models: Set = set()
|
||||
moonshot_models: Set = set()
|
||||
publicai_models: Set = set()
|
||||
v0_models: Set = set()
|
||||
morph_models: Set = set()
|
||||
lambda_ai_models: Set = set()
|
||||
|
|
@ -781,6 +782,8 @@ def add_known_models():
|
|||
dashscope_models.add(key)
|
||||
elif value.get("litellm_provider") == "moonshot":
|
||||
moonshot_models.add(key)
|
||||
elif value.get("litellm_provider") == "publicai":
|
||||
publicai_models.add(key)
|
||||
elif value.get("litellm_provider") == "v0":
|
||||
v0_models.add(key)
|
||||
elif value.get("litellm_provider") == "morph":
|
||||
|
|
@ -899,6 +902,7 @@ model_list = list(
|
|||
| elevenlabs_models
|
||||
| dashscope_models
|
||||
| moonshot_models
|
||||
| publicai_models
|
||||
| v0_models
|
||||
| morph_models
|
||||
| lambda_ai_models
|
||||
|
|
@ -992,6 +996,7 @@ models_by_provider: dict = {
|
|||
"heroku": heroku_models,
|
||||
"dashscope": dashscope_models,
|
||||
"moonshot": moonshot_models,
|
||||
"publicai": publicai_models,
|
||||
"v0": v0_models,
|
||||
"morph": morph_models,
|
||||
"lambda_ai": lambda_ai_models,
|
||||
|
|
@ -1120,7 +1125,7 @@ from .llms.openrouter.chat.transformation import OpenrouterConfig
|
|||
from .llms.datarobot.chat.transformation import DataRobotConfig
|
||||
from .llms.anthropic.chat.transformation import AnthropicConfig
|
||||
from .llms.anthropic.common_utils import AnthropicModelInfo
|
||||
from .llms.azure.anthropic.transformation import AzureAnthropicConfig
|
||||
from .llms.azure_ai.anthropic.transformation import AzureAnthropicConfig
|
||||
from .llms.groq.stt.transformation import GroqSTTConfig
|
||||
from .llms.anthropic.completion.transformation import AnthropicTextConfig
|
||||
from .llms.triton.completion.transformation import TritonConfig
|
||||
|
|
@ -1370,6 +1375,7 @@ from .llms.nebius.chat.transformation import NebiusConfig
|
|||
from .llms.wandb.chat.transformation import WandbConfig
|
||||
from .llms.dashscope.chat.transformation import DashScopeChatConfig
|
||||
from .llms.moonshot.chat.transformation import MoonshotChatConfig
|
||||
from .llms.publicai.chat.transformation import PublicAIChatConfig
|
||||
from .llms.docker_model_runner.chat.transformation import DockerModelRunnerChatConfig
|
||||
from .llms.v0.chat.transformation import V0ChatConfig
|
||||
from .llms.oci.chat.transformation import OCIChatConfig
|
||||
|
|
|
|||
|
|
@ -384,6 +384,7 @@ LITELLM_CHAT_PROVIDERS = [
|
|||
"nebius",
|
||||
"dashscope",
|
||||
"moonshot",
|
||||
"publicai",
|
||||
"v0",
|
||||
"heroku",
|
||||
"oci",
|
||||
|
|
@ -526,6 +527,7 @@ openai_compatible_endpoints: List = [
|
|||
"api.studio.nebius.ai/v1",
|
||||
"https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
|
||||
"https://api.moonshot.ai/v1",
|
||||
"https://platform.publicai.co/v1",
|
||||
"https://api.v0.dev/v1",
|
||||
"https://api.morphllm.com/v1",
|
||||
"https://api.lambda.ai/v1",
|
||||
|
|
@ -571,6 +573,7 @@ openai_compatible_providers: List = [
|
|||
"nebius",
|
||||
"dashscope",
|
||||
"moonshot",
|
||||
"publicai",
|
||||
"v0",
|
||||
"morph",
|
||||
"lambda_ai",
|
||||
|
|
@ -593,6 +596,7 @@ openai_text_completion_compatible_providers: List = (
|
|||
"nebius",
|
||||
"dashscope",
|
||||
"moonshot",
|
||||
"publicai",
|
||||
"v0",
|
||||
"lambda_ai",
|
||||
"hyperbolic",
|
||||
|
|
|
|||
|
|
@ -22,17 +22,16 @@ def _is_non_openai_azure_model(model: str) -> bool:
|
|||
return False
|
||||
|
||||
|
||||
def _is_azure_anthropic_model(model: str) -> Optional[str]:
|
||||
def _is_azure_claude_model(model: str) -> bool:
|
||||
"""
|
||||
Check if a model name contains 'claude' (case-insensitive).
|
||||
Used to detect Claude models that need Anthropic-specific handling.
|
||||
"""
|
||||
try:
|
||||
model_parts = model.split("/", 1)
|
||||
if len(model_parts) > 1:
|
||||
model_name = model_parts[1].lower()
|
||||
# Check if model name contains claude
|
||||
if "claude" in model_name or model_name.startswith("claude"):
|
||||
return model_parts[1] # Return model name without "azure/" prefix
|
||||
model_lower = model.lower()
|
||||
return "claude" in model_lower or model_lower.startswith("claude")
|
||||
except Exception:
|
||||
pass
|
||||
return None
|
||||
return False
|
||||
|
||||
|
||||
def handle_cohere_chat_model_custom_llm_provider(
|
||||
|
|
@ -136,11 +135,6 @@ def get_llm_provider( # noqa: PLR0915
|
|||
# AZURE AI-Studio Logic - Azure AI Studio supports AZURE/Cohere
|
||||
# If User passes azure/command-r-plus -> we should send it to cohere_chat/command-r-plus
|
||||
if model.split("/", 1)[0] == "azure":
|
||||
# Check if it's an Azure Anthropic model (claude models)
|
||||
azure_anthropic_model = _is_azure_anthropic_model(model)
|
||||
if azure_anthropic_model:
|
||||
custom_llm_provider = "azure_anthropic"
|
||||
return azure_anthropic_model, custom_llm_provider, dynamic_api_key, api_base
|
||||
if _is_non_openai_azure_model(model):
|
||||
custom_llm_provider = "openai"
|
||||
return model, custom_llm_provider, dynamic_api_key, api_base
|
||||
|
|
@ -258,6 +252,9 @@ def get_llm_provider( # noqa: PLR0915
|
|||
elif endpoint == "api.moonshot.ai/v1":
|
||||
custom_llm_provider = "moonshot"
|
||||
dynamic_api_key = get_secret_str("MOONSHOT_API_KEY")
|
||||
elif endpoint == "platform.publicai.co/v1":
|
||||
custom_llm_provider = "publicai"
|
||||
dynamic_api_key = get_secret_str("PUBLICAI_API_KEY")
|
||||
elif endpoint == "https://api.v0.dev/v1":
|
||||
custom_llm_provider = "v0"
|
||||
dynamic_api_key = get_secret_str("V0_API_KEY")
|
||||
|
|
@ -759,6 +756,13 @@ def _get_openai_compatible_provider_info( # noqa: PLR0915
|
|||
) = litellm.MoonshotChatConfig()._get_openai_compatible_provider_info(
|
||||
api_base, api_key
|
||||
)
|
||||
elif custom_llm_provider == "publicai":
|
||||
(
|
||||
api_base,
|
||||
dynamic_api_key,
|
||||
) = litellm.PublicAIChatConfig()._get_openai_compatible_provider_info(
|
||||
api_base, api_key
|
||||
)
|
||||
elif custom_llm_provider == "docker_model_runner":
|
||||
(
|
||||
api_base,
|
||||
|
|
|
|||
|
|
@ -374,6 +374,9 @@ class Logging(LiteLLMLoggingBaseClass):
|
|||
# Init Caching related details
|
||||
self.caching_details: Optional[CachingDetails] = None
|
||||
|
||||
# Passthrough endpoint guardrails config for field targeting
|
||||
self.passthrough_guardrails_config: Optional[Dict[str, Any]] = None
|
||||
|
||||
self.model_call_details: Dict[str, Any] = {
|
||||
"litellm_trace_id": litellm_trace_id,
|
||||
"litellm_call_id": litellm_call_id,
|
||||
|
|
|
|||
|
|
@ -41,6 +41,7 @@ class AnthropicMessagesHandler(BaseTranslation):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input messages by applying guardrails to text content.
|
||||
|
|
@ -145,6 +146,7 @@ class AnthropicMessagesHandler(BaseTranslation):
|
|||
self,
|
||||
response: "AnthropicMessagesResponse",
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output response by applying guardrails to text content.
|
||||
|
|
|
|||
|
|
@ -7,7 +7,6 @@ from typing import TYPE_CHECKING, Callable, Union
|
|||
|
||||
import httpx
|
||||
|
||||
import litellm
|
||||
from litellm.llms.anthropic.chat.handler import AnthropicChatCompletion
|
||||
from litellm.llms.custom_httpx.http_handler import (
|
||||
AsyncHTTPHandler,
|
||||
|
|
@ -55,7 +54,6 @@ class AzureAnthropicChatCompletion(AnthropicChatCompletion):
|
|||
Completion method that uses Azure authentication instead of Anthropic's x-api-key.
|
||||
All other logic is the same as AnthropicChatCompletion.
|
||||
"""
|
||||
from litellm.utils import ProviderConfigManager
|
||||
|
||||
optional_params = copy.deepcopy(optional_params)
|
||||
stream = optional_params.pop("stream", None)
|
||||
|
|
@ -64,8 +62,10 @@ class AzureAnthropicChatCompletion(AnthropicChatCompletion):
|
|||
_is_function_call = False
|
||||
messages = copy.deepcopy(messages)
|
||||
|
||||
# Use AzureAnthropicConfig instead of AnthropicConfig
|
||||
headers = AzureAnthropicConfig().validate_environment(
|
||||
# Use AzureAnthropicConfig for both azure_anthropic and azure_ai Claude models
|
||||
config = AzureAnthropicConfig()
|
||||
|
||||
headers = config.validate_environment(
|
||||
api_key=api_key,
|
||||
headers=headers,
|
||||
model=model,
|
||||
|
|
@ -74,15 +74,6 @@ class AzureAnthropicChatCompletion(AnthropicChatCompletion):
|
|||
litellm_params=litellm_params,
|
||||
)
|
||||
|
||||
config = ProviderConfigManager.get_provider_chat_config(
|
||||
model=model,
|
||||
provider=litellm.types.utils.LlmProviders(custom_llm_provider),
|
||||
)
|
||||
if config is None:
|
||||
raise ValueError(
|
||||
f"Provider config not found for model: {model} and provider: {custom_llm_provider}"
|
||||
)
|
||||
|
||||
data = config.transform_request(
|
||||
model=model,
|
||||
messages=messages,
|
||||
|
|
@ -183,7 +174,7 @@ class AzureAnthropicChatCompletion(AnthropicChatCompletion):
|
|||
return CustomStreamWrapper(
|
||||
completion_stream=completion_stream,
|
||||
model=model,
|
||||
custom_llm_provider="azure_anthropic",
|
||||
custom_llm_provider="azure_ai",
|
||||
logging_obj=logging_obj,
|
||||
_response_headers=process_anthropic_headers(response_headers),
|
||||
)
|
||||
|
|
@ -21,7 +21,7 @@ class AzureAnthropicConfig(AnthropicConfig):
|
|||
|
||||
@property
|
||||
def custom_llm_provider(self) -> Optional[str]:
|
||||
return "azure_anthropic"
|
||||
return "azure_ai"
|
||||
|
||||
def validate_environment(
|
||||
self,
|
||||
|
|
@ -94,3 +94,29 @@ class AzureAnthropicConfig(AnthropicConfig):
|
|||
|
||||
return headers
|
||||
|
||||
def transform_request(
|
||||
self,
|
||||
model: str,
|
||||
messages: List[AllMessageValues],
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
headers: dict,
|
||||
) -> dict:
|
||||
"""
|
||||
Transform request using parent AnthropicConfig, then remove extra_body if present.
|
||||
Azure Anthropic doesn't support extra_body parameter.
|
||||
"""
|
||||
# Call parent transform_request
|
||||
data = super().transform_request(
|
||||
model=model,
|
||||
messages=messages,
|
||||
optional_params=optional_params,
|
||||
litellm_params=litellm_params,
|
||||
headers=headers,
|
||||
)
|
||||
|
||||
# Remove extra_body if present (Azure Anthropic doesn't support it)
|
||||
data.pop("extra_body", None)
|
||||
|
||||
return data
|
||||
|
||||
|
|
@ -1,8 +1,9 @@
|
|||
from abc import ABC, abstractmethod
|
||||
from typing import TYPE_CHECKING, Any
|
||||
from typing import TYPE_CHECKING, Any, Optional
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.integrations.custom_guardrail import CustomGuardrail
|
||||
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
|
||||
|
||||
|
||||
class BaseTranslation(ABC):
|
||||
|
|
@ -11,6 +12,7 @@ class BaseTranslation(ABC):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional["LiteLLMLoggingObj"] = None,
|
||||
) -> Any:
|
||||
pass
|
||||
|
||||
|
|
@ -19,5 +21,6 @@ class BaseTranslation(ABC):
|
|||
self,
|
||||
response: Any,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional["LiteLLMLoggingObj"] = None,
|
||||
) -> Any:
|
||||
pass
|
||||
|
|
|
|||
|
|
@ -1,5 +1,5 @@
|
|||
"""
|
||||
Legacy /v1/embedding transformation logic for Bedrock Cohere.
|
||||
Legacy /v1/embedding transformation logic for Bedrock Cohere.
|
||||
"""
|
||||
|
||||
from typing import Any, List, Optional, Union
|
||||
|
|
@ -123,7 +123,13 @@ class CohereEmbeddingConfig:
|
|||
"""
|
||||
embeddings = response_json["embeddings"]
|
||||
output_data = []
|
||||
is_embeddings_by_type = response_json.get("response_type") == "embeddings_by_type"
|
||||
is_embeddings_by_type = (
|
||||
response_json.get("response_type") == "embeddings_by_type"
|
||||
)
|
||||
|
||||
if isinstance(embeddings, dict):
|
||||
is_embeddings_by_type = True
|
||||
|
||||
if is_embeddings_by_type:
|
||||
for embedding_type in embeddings:
|
||||
for idx, embedding in enumerate(embeddings[embedding_type]):
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ This module provides guardrail translation support for the rerank endpoint.
|
|||
The handler processes only the 'query' parameter for guardrails.
|
||||
"""
|
||||
|
||||
from typing import TYPE_CHECKING, Any
|
||||
from typing import TYPE_CHECKING, Any, Optional
|
||||
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
|
||||
|
|
@ -34,6 +34,7 @@ class CohereRerankHandler(BaseTranslation):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input query by applying guardrails.
|
||||
|
|
@ -68,6 +69,7 @@ class CohereRerankHandler(BaseTranslation):
|
|||
self,
|
||||
response: "RerankResponse",
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output response - not applicable for rerank.
|
||||
|
|
|
|||
|
|
@ -42,6 +42,7 @@ class OpenAIChatCompletionsHandler(BaseTranslation):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input messages by applying guardrails to text content.
|
||||
|
|
@ -148,6 +149,7 @@ class OpenAIChatCompletionsHandler(BaseTranslation):
|
|||
self,
|
||||
response: "ModelResponse",
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output response by applying guardrails to text content.
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ This module provides guardrail translation support for OpenAI's text completion
|
|||
The handler processes the 'prompt' parameter for guardrails.
|
||||
"""
|
||||
|
||||
from typing import TYPE_CHECKING, Any
|
||||
from typing import TYPE_CHECKING, Any, Optional
|
||||
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
|
||||
|
|
@ -32,6 +32,7 @@ class OpenAITextCompletionHandler(BaseTranslation):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input prompt by applying guardrails to text content.
|
||||
|
|
@ -100,6 +101,7 @@ class OpenAITextCompletionHandler(BaseTranslation):
|
|||
self,
|
||||
response: "TextCompletionResponse",
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output response by applying guardrails to completion text.
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ This module provides guardrail translation support for OpenAI's image generation
|
|||
The handler processes the 'prompt' parameter for guardrails.
|
||||
"""
|
||||
|
||||
from typing import TYPE_CHECKING, Any
|
||||
from typing import TYPE_CHECKING, Any, Optional
|
||||
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
|
||||
|
|
@ -31,6 +31,7 @@ class OpenAIImageGenerationHandler(BaseTranslation):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input prompt by applying guardrails to text content.
|
||||
|
|
@ -72,6 +73,7 @@ class OpenAIImageGenerationHandler(BaseTranslation):
|
|||
self,
|
||||
response: "ImageResponse",
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output response - typically not needed for image generation.
|
||||
|
|
|
|||
|
|
@ -56,6 +56,7 @@ class OpenAIResponsesHandler(BaseTranslation):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input by applying guardrails to text content.
|
||||
|
|
@ -177,6 +178,7 @@ class OpenAIResponsesHandler(BaseTranslation):
|
|||
self,
|
||||
response: "ResponsesAPIResponse",
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output response by applying guardrails to text content.
|
||||
|
|
|
|||
|
|
@ -238,6 +238,24 @@ class OpenAIResponsesAPIConfig(BaseResponsesAPIConfig):
|
|||
event_pydantic_model = OpenAIResponsesAPIConfig.get_event_model_class(
|
||||
event_type=event_type
|
||||
)
|
||||
# Defensive: Some OpenAI-compatible providers may send `error.code: null`.
|
||||
# Pydantic will raise a ValidationError when it expects a string but gets None.
|
||||
# Coalesce a None `error.code` to a stable default string so streaming
|
||||
# iteration does not crash (see issue report). This keeps behavior similar
|
||||
# to previous fixes (coalesce before validation) and lets higher-level
|
||||
# handlers still receive an `ErrorEvent` object.
|
||||
try:
|
||||
error_obj = parsed_chunk.get("error")
|
||||
if isinstance(error_obj, dict) and error_obj.get("code") is None:
|
||||
# Preserve other fields, but ensure `code` is a non-null string
|
||||
parsed_chunk = dict(parsed_chunk)
|
||||
parsed_chunk["error"] = dict(error_obj)
|
||||
parsed_chunk["error"]["code"] = "unknown_error"
|
||||
except Exception:
|
||||
# If anything unexpected happens here, fall back to attempting
|
||||
# instantiation and let higher-level handlers manage errors.
|
||||
verbose_logger.debug("Failed to coalesce error.code in parsed_chunk")
|
||||
|
||||
return event_pydantic_model(**parsed_chunk)
|
||||
|
||||
@staticmethod
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ This module provides guardrail translation support for OpenAI's text-to-speech e
|
|||
The handler processes the 'input' text parameter (output is audio, so no text to guardrail).
|
||||
"""
|
||||
|
||||
from typing import TYPE_CHECKING, Any
|
||||
from typing import TYPE_CHECKING, Any, Optional
|
||||
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
|
||||
|
|
@ -30,6 +30,7 @@ class OpenAITextToSpeechHandler(BaseTranslation):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input text by applying guardrails.
|
||||
|
|
@ -72,6 +73,7 @@ class OpenAITextToSpeechHandler(BaseTranslation):
|
|||
self,
|
||||
response: "HttpxBinaryResponseContent",
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output - not applicable for text-to-speech.
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ This module provides guardrail translation support for OpenAI's audio transcript
|
|||
The handler processes the output transcribed text (input is audio, so no text to guardrail).
|
||||
"""
|
||||
|
||||
from typing import TYPE_CHECKING, Any
|
||||
from typing import TYPE_CHECKING, Any, Optional
|
||||
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
|
||||
|
|
@ -30,6 +30,7 @@ class OpenAIAudioTranscriptionHandler(BaseTranslation):
|
|||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input - not applicable for audio transcription.
|
||||
|
|
@ -54,6 +55,7 @@ class OpenAIAudioTranscriptionHandler(BaseTranslation):
|
|||
self,
|
||||
response: "TranscriptionResponse",
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional[Any] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output transcription by applying guardrails to transcribed text.
|
||||
|
|
|
|||
12
litellm/llms/pass_through/__init__.py
Normal file
12
litellm/llms/pass_through/__init__.py
Normal file
|
|
@ -0,0 +1,12 @@
|
|||
"""
|
||||
Pass-Through Endpoint Guardrail Translation
|
||||
|
||||
This module exists here (under litellm/llms/) so it can be auto-discovered by
|
||||
load_guardrail_translation_mappings() which scans for guardrail_translation
|
||||
directories under litellm/llms/.
|
||||
|
||||
The main passthrough endpoint implementation is in:
|
||||
litellm/proxy/pass_through_endpoints/
|
||||
|
||||
See guardrail_translation/README.md for more details.
|
||||
"""
|
||||
41
litellm/llms/pass_through/guardrail_translation/README.md
Normal file
41
litellm/llms/pass_through/guardrail_translation/README.md
Normal file
|
|
@ -0,0 +1,41 @@
|
|||
# Pass-Through Endpoint Guardrail Translation
|
||||
|
||||
## Why This Exists Here
|
||||
|
||||
This module is located under `litellm/llms/` (instead of with the main passthrough code) because:
|
||||
|
||||
1. **Auto-discovery**: The `load_guardrail_translation_mappings()` function in `litellm/llms/__init__.py` scans for `guardrail_translation/` directories under `litellm/llms/`
|
||||
2. **Consistency**: All other guardrail translation handlers follow this pattern (e.g., `openai/chat/guardrail_translation/`, `anthropic/chat/guardrail_translation/`)
|
||||
|
||||
## Main Passthrough Implementation
|
||||
|
||||
The main passthrough endpoint implementation is in:
|
||||
|
||||
```
|
||||
litellm/proxy/pass_through_endpoints/
|
||||
├── pass_through_endpoints.py # Core passthrough routing logic
|
||||
├── passthrough_guardrails.py # Guardrail collection and field targeting
|
||||
├── jsonpath_extractor.py # JSONPath field extraction utility
|
||||
└── ...
|
||||
```
|
||||
|
||||
## What This Handler Does
|
||||
|
||||
The `PassThroughEndpointHandler` enables guardrails to run on passthrough endpoint requests by:
|
||||
|
||||
1. **Field Targeting**: Extracts specific fields from the request/response using JSONPath expressions configured in `request_fields` / `response_fields`
|
||||
2. **Full Payload Fallback**: If no field targeting is configured, processes the entire payload
|
||||
3. **Config Access**: Uses `get_passthrough_guardrails_config()` / `set_passthrough_guardrails_config()` helpers to access the passthrough guardrails configuration stored in request metadata
|
||||
|
||||
## Example Config
|
||||
|
||||
```yaml
|
||||
passthrough_endpoints:
|
||||
- path: "/v1/rerank"
|
||||
target: "https://api.cohere.com/v1/rerank"
|
||||
guardrails:
|
||||
bedrock-pre-guard:
|
||||
request_fields: ["query", "documents[*].text"]
|
||||
response_fields: ["results[*].text"]
|
||||
```
|
||||
|
||||
15
litellm/llms/pass_through/guardrail_translation/__init__.py
Normal file
15
litellm/llms/pass_through/guardrail_translation/__init__.py
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
"""Pass-Through Endpoint guardrail translation handler."""
|
||||
|
||||
from litellm.llms.pass_through.guardrail_translation.handler import (
|
||||
PassThroughEndpointHandler,
|
||||
)
|
||||
from litellm.types.utils import CallTypes
|
||||
|
||||
guardrail_translation_mappings = {
|
||||
CallTypes.pass_through: PassThroughEndpointHandler,
|
||||
}
|
||||
|
||||
__all__ = [
|
||||
"guardrail_translation_mappings",
|
||||
"PassThroughEndpointHandler",
|
||||
]
|
||||
165
litellm/llms/pass_through/guardrail_translation/handler.py
Normal file
165
litellm/llms/pass_through/guardrail_translation/handler.py
Normal file
|
|
@ -0,0 +1,165 @@
|
|||
"""
|
||||
Pass-Through Endpoint Message Handler for Unified Guardrails
|
||||
|
||||
This module provides a handler for passthrough endpoint requests.
|
||||
It uses the field targeting configuration from litellm_logging_obj
|
||||
to extract specific fields for guardrail processing.
|
||||
"""
|
||||
|
||||
from typing import TYPE_CHECKING, Any, List, Optional
|
||||
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
|
||||
from litellm.proxy._types import PassThroughGuardrailSettings
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.integrations.custom_guardrail import CustomGuardrail
|
||||
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
|
||||
|
||||
|
||||
class PassThroughEndpointHandler(BaseTranslation):
|
||||
"""
|
||||
Handler for processing passthrough endpoint requests with guardrails.
|
||||
|
||||
Uses passthrough_guardrails_config from litellm_logging_obj
|
||||
to determine which fields to extract for guardrail processing.
|
||||
"""
|
||||
|
||||
def _get_guardrail_settings(
|
||||
self,
|
||||
litellm_logging_obj: Optional["LiteLLMLoggingObj"],
|
||||
guardrail_name: Optional[str],
|
||||
) -> Optional[PassThroughGuardrailSettings]:
|
||||
"""
|
||||
Get the guardrail settings for a specific guardrail from logging_obj.
|
||||
"""
|
||||
from litellm.proxy.pass_through_endpoints.passthrough_guardrails import (
|
||||
PassthroughGuardrailHandler,
|
||||
)
|
||||
|
||||
if litellm_logging_obj is None:
|
||||
return None
|
||||
|
||||
passthrough_config = getattr(
|
||||
litellm_logging_obj, "passthrough_guardrails_config", None
|
||||
)
|
||||
if not passthrough_config or not guardrail_name:
|
||||
return None
|
||||
|
||||
return PassthroughGuardrailHandler.get_settings(
|
||||
passthrough_config, guardrail_name
|
||||
)
|
||||
|
||||
def _extract_text_for_guardrail(
|
||||
self,
|
||||
data: dict,
|
||||
field_expressions: Optional[List[str]],
|
||||
) -> str:
|
||||
"""
|
||||
Extract text from data for guardrail processing.
|
||||
|
||||
If field_expressions provided, extracts only those fields.
|
||||
Otherwise, returns the full payload as JSON.
|
||||
"""
|
||||
from litellm.proxy.pass_through_endpoints.jsonpath_extractor import (
|
||||
JsonPathExtractor,
|
||||
)
|
||||
|
||||
if field_expressions:
|
||||
text = JsonPathExtractor.extract_fields(
|
||||
data=data,
|
||||
jsonpath_expressions=field_expressions,
|
||||
)
|
||||
verbose_proxy_logger.debug(
|
||||
"PassThroughEndpointHandler: Extracted targeted fields: %s",
|
||||
text[:200] if text else None,
|
||||
)
|
||||
return text
|
||||
|
||||
# Use entire payload, excluding internal fields
|
||||
from litellm.litellm_core_utils.safe_json_dumps import safe_dumps
|
||||
|
||||
payload_to_check = {
|
||||
k: v
|
||||
for k, v in data.items()
|
||||
if not k.startswith("_") and k not in ("metadata", "litellm_logging_obj")
|
||||
}
|
||||
verbose_proxy_logger.debug(
|
||||
"PassThroughEndpointHandler: Using full payload for guardrail"
|
||||
)
|
||||
return safe_dumps(payload_to_check)
|
||||
|
||||
async def process_input_messages(
|
||||
self,
|
||||
data: dict,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional["LiteLLMLoggingObj"] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process input by applying guardrails to targeted fields or full payload.
|
||||
"""
|
||||
guardrail_name = guardrail_to_apply.guardrail_name
|
||||
verbose_proxy_logger.debug(
|
||||
"PassThroughEndpointHandler: Processing input for guardrail=%s",
|
||||
guardrail_name,
|
||||
)
|
||||
|
||||
# Get field targeting settings for this guardrail
|
||||
settings = self._get_guardrail_settings(litellm_logging_obj, guardrail_name)
|
||||
field_expressions = settings.request_fields if settings else None
|
||||
|
||||
# Extract text to check
|
||||
text_to_check = self._extract_text_for_guardrail(data, field_expressions)
|
||||
|
||||
if not text_to_check:
|
||||
verbose_proxy_logger.debug(
|
||||
"PassThroughEndpointHandler: No text to check, skipping guardrail"
|
||||
)
|
||||
return data
|
||||
|
||||
# Apply guardrail
|
||||
await guardrail_to_apply.apply_guardrail(
|
||||
text=text_to_check,
|
||||
request_data=data,
|
||||
)
|
||||
|
||||
return data
|
||||
|
||||
async def process_output_response(
|
||||
self,
|
||||
response: Any,
|
||||
guardrail_to_apply: "CustomGuardrail",
|
||||
litellm_logging_obj: Optional["LiteLLMLoggingObj"] = None,
|
||||
) -> Any:
|
||||
"""
|
||||
Process output response by applying guardrails to targeted fields.
|
||||
"""
|
||||
if not isinstance(response, dict):
|
||||
verbose_proxy_logger.debug(
|
||||
"PassThroughEndpointHandler: Response is not a dict, skipping"
|
||||
)
|
||||
return response
|
||||
|
||||
guardrail_name = guardrail_to_apply.guardrail_name
|
||||
verbose_proxy_logger.debug(
|
||||
"PassThroughEndpointHandler: Processing output for guardrail=%s",
|
||||
guardrail_name,
|
||||
)
|
||||
|
||||
# Get field targeting settings for this guardrail
|
||||
settings = self._get_guardrail_settings(litellm_logging_obj, guardrail_name)
|
||||
field_expressions = settings.response_fields if settings else None
|
||||
|
||||
# Extract text to check
|
||||
text_to_check = self._extract_text_for_guardrail(response, field_expressions)
|
||||
|
||||
if not text_to_check:
|
||||
return response
|
||||
|
||||
# Apply guardrail
|
||||
await guardrail_to_apply.apply_guardrail(
|
||||
text=text_to_check,
|
||||
request_data=response,
|
||||
)
|
||||
|
||||
return response
|
||||
114
litellm/llms/publicai/chat/transformation.py
Normal file
114
litellm/llms/publicai/chat/transformation.py
Normal file
|
|
@ -0,0 +1,114 @@
|
|||
"""
|
||||
Translates from OpenAI's `/v1/chat/completions` to PublicAI's `/v1/chat/completions`
|
||||
"""
|
||||
|
||||
from typing import Any, Coroutine, List, Literal, Optional, Tuple, Union, overload
|
||||
|
||||
from litellm.litellm_core_utils.prompt_templates.common_utils import (
|
||||
handle_messages_with_content_list_to_str_conversion,
|
||||
)
|
||||
from litellm.secret_managers.main import get_secret_str
|
||||
from litellm.types.llms.openai import AllMessageValues
|
||||
|
||||
from ...openai.chat.gpt_transformation import OpenAIGPTConfig
|
||||
|
||||
|
||||
class PublicAIChatConfig(OpenAIGPTConfig):
|
||||
@overload
|
||||
def _transform_messages(
|
||||
self, messages: List[AllMessageValues], model: str, is_async: Literal[True]
|
||||
) -> Coroutine[Any, Any, List[AllMessageValues]]:
|
||||
...
|
||||
|
||||
@overload
|
||||
def _transform_messages(
|
||||
self,
|
||||
messages: List[AllMessageValues],
|
||||
model: str,
|
||||
is_async: Literal[False] = False,
|
||||
) -> List[AllMessageValues]:
|
||||
...
|
||||
|
||||
def _transform_messages(
|
||||
self, messages: List[AllMessageValues], model: str, is_async: bool = False
|
||||
) -> Union[List[AllMessageValues], Coroutine[Any, Any, List[AllMessageValues]]]:
|
||||
"""
|
||||
PublicAI does not support content in list format.
|
||||
"""
|
||||
messages = handle_messages_with_content_list_to_str_conversion(messages)
|
||||
if is_async:
|
||||
return super()._transform_messages(
|
||||
messages=messages, model=model, is_async=True
|
||||
)
|
||||
else:
|
||||
return super()._transform_messages(
|
||||
messages=messages, model=model, is_async=False
|
||||
)
|
||||
|
||||
def _get_openai_compatible_provider_info(
|
||||
self, api_base: Optional[str], api_key: Optional[str]
|
||||
) -> Tuple[Optional[str], Optional[str]]:
|
||||
api_base = (
|
||||
api_base
|
||||
or get_secret_str("PUBLICAI_API_BASE")
|
||||
or "https://platform.publicai.co/v1"
|
||||
) # type: ignore
|
||||
dynamic_api_key = api_key or get_secret_str("PUBLICAI_API_KEY")
|
||||
return api_base, dynamic_api_key
|
||||
|
||||
def get_complete_url(
|
||||
self,
|
||||
api_base: Optional[str],
|
||||
api_key: Optional[str],
|
||||
model: str,
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
stream: Optional[bool] = None,
|
||||
) -> str:
|
||||
"""
|
||||
If api_base is not provided, use the default PublicAI /chat/completions endpoint.
|
||||
"""
|
||||
if not api_base:
|
||||
api_base = "https://platform.publicai.co/v1"
|
||||
|
||||
if not api_base.endswith("/chat/completions"):
|
||||
api_base = f"{api_base}/chat/completions"
|
||||
|
||||
return api_base
|
||||
|
||||
def get_supported_openai_params(self, model: str) -> list:
|
||||
"""
|
||||
Get the supported OpenAI params for PublicAI models
|
||||
|
||||
PublicAI limitations:
|
||||
- functions parameter is not supported (use tools instead)
|
||||
"""
|
||||
excluded_params: List[str] = ["functions"]
|
||||
|
||||
base_openai_params = super().get_supported_openai_params(model=model)
|
||||
final_params: List[str] = []
|
||||
for param in base_openai_params:
|
||||
if param not in excluded_params:
|
||||
final_params.append(param)
|
||||
|
||||
return final_params
|
||||
|
||||
def map_openai_params(
|
||||
self,
|
||||
non_default_params: dict,
|
||||
optional_params: dict,
|
||||
model: str,
|
||||
drop_params: bool,
|
||||
) -> dict:
|
||||
"""
|
||||
Map OpenAI parameters to PublicAI parameters
|
||||
"""
|
||||
supported_openai_params = self.get_supported_openai_params(model)
|
||||
for param, value in non_default_params.items():
|
||||
if param == "max_completion_tokens":
|
||||
optional_params["max_tokens"] = value
|
||||
elif param in supported_openai_params:
|
||||
optional_params[param] = value
|
||||
|
||||
return optional_params
|
||||
|
||||
|
|
@ -121,5 +121,10 @@ class SambanovaConfig(OpenAIGPTConfig):
|
|||
SambaNova API doesn't support content as a list - only string content.
|
||||
This converts content lists like [{"type": "text", "text": "..."}] to strings.
|
||||
"""
|
||||
async def _async_transform():
|
||||
return handle_messages_with_content_list_to_str_conversion(messages)
|
||||
|
||||
if is_async:
|
||||
return _async_transform()
|
||||
messages = handle_messages_with_content_list_to_str_conversion(messages)
|
||||
return messages
|
||||
|
|
|
|||
208
litellm/main.py
208
litellm/main.py
|
|
@ -58,9 +58,11 @@ from litellm import ( # type: ignore
|
|||
get_litellm_params,
|
||||
get_optional_params,
|
||||
)
|
||||
|
||||
# Logging is imported lazily when needed to avoid loading litellm_logging at import time
|
||||
if TYPE_CHECKING:
|
||||
from litellm.litellm_core_utils.litellm_logging import Logging
|
||||
|
||||
from litellm.constants import (
|
||||
DEFAULT_MOCK_RESPONSE_COMPLETION_TOKEN_COUNT,
|
||||
DEFAULT_MOCK_RESPONSE_PROMPT_TOKEN_COUNT,
|
||||
|
|
@ -107,7 +109,6 @@ from litellm.utils import (
|
|||
ProviderConfigManager,
|
||||
Usage,
|
||||
_get_model_info_helper,
|
||||
get_requester_metadata,
|
||||
add_provider_specific_params_to_optional_params,
|
||||
async_mock_completion_streaming_obj,
|
||||
convert_to_model_response_object,
|
||||
|
|
@ -120,6 +121,7 @@ from litellm.utils import (
|
|||
get_optional_params_embeddings,
|
||||
get_optional_params_image_gen,
|
||||
get_optional_params_transcription,
|
||||
get_requester_metadata,
|
||||
get_secret,
|
||||
get_standard_openai_params,
|
||||
mock_completion_streaming_obj,
|
||||
|
|
@ -154,11 +156,11 @@ from .litellm_core_utils.prompt_templates.factory import (
|
|||
)
|
||||
from .litellm_core_utils.streaming_chunk_builder_utils import ChunkProcessor
|
||||
from .llms.anthropic.chat import AnthropicChatCompletion
|
||||
from .llms.azure.anthropic.handler import AzureAnthropicChatCompletion
|
||||
from .llms.azure.audio_transcriptions import AzureAudioTranscription
|
||||
from .llms.azure.azure import AzureChatCompletion, _check_dynamic_azure_params
|
||||
from .llms.azure.chat.o_series_handler import AzureOpenAIO1ChatCompletion
|
||||
from .llms.azure.completion.handler import AzureTextCompletion
|
||||
from .llms.azure_ai.anthropic.handler import AzureAnthropicChatCompletion
|
||||
from .llms.azure_ai.embed import AzureAIEmbedding
|
||||
from .llms.bedrock.chat import BedrockConverseLLM, BedrockLLM
|
||||
from .llms.bedrock.embed.embedding import BedrockEmbedding
|
||||
|
|
@ -1686,57 +1688,109 @@ def completion( # type: ignore # noqa: PLR0915
|
|||
elif custom_llm_provider == "azure_ai":
|
||||
from litellm.llms.azure_ai.common_utils import AzureFoundryModelInfo
|
||||
|
||||
api_base = AzureFoundryModelInfo.get_api_base(api_base)
|
||||
# set API KEY
|
||||
api_key = AzureFoundryModelInfo.get_api_key(api_key)
|
||||
|
||||
headers = headers or litellm.headers
|
||||
|
||||
if extra_headers is not None:
|
||||
optional_params["extra_headers"] = extra_headers
|
||||
|
||||
## FOR COHERE
|
||||
if "command-r" in model: # make sure tool call in messages are str
|
||||
messages = stringify_json_tool_call_content(messages=messages)
|
||||
|
||||
## COMPLETION CALL
|
||||
try:
|
||||
response = base_llm_http_handler.completion(
|
||||
# Check if this is a Claude model - route to Azure Anthropic handler
|
||||
model_lower = model.lower()
|
||||
if "claude" in model_lower:
|
||||
# Use Azure Anthropic handler for Claude models
|
||||
api_base = AzureFoundryModelInfo.get_api_base(api_base)
|
||||
if api_base is None:
|
||||
raise ValueError(
|
||||
"Azure Anthropic requests require an api_base. "
|
||||
"Set `api_base` or the AZURE_AI_API_BASE env var."
|
||||
)
|
||||
api_key = AzureFoundryModelInfo.get_api_key(api_key)
|
||||
|
||||
# Ensure the URL ends with /v1/messages for Anthropic
|
||||
if api_base:
|
||||
api_base = api_base.rstrip("/")
|
||||
if not api_base.endswith("/v1/messages"):
|
||||
if "/anthropic" in api_base:
|
||||
parts = api_base.split("/anthropic", 1)
|
||||
api_base = parts[0] + "/anthropic"
|
||||
else:
|
||||
api_base = api_base + "/anthropic"
|
||||
api_base = api_base + "/v1/messages"
|
||||
|
||||
response = azure_anthropic_chat_completions.completion(
|
||||
model=model,
|
||||
messages=messages,
|
||||
headers=headers,
|
||||
model_response=model_response,
|
||||
api_key=api_key,
|
||||
api_base=api_base,
|
||||
acompletion=acompletion,
|
||||
logging_obj=logging,
|
||||
custom_prompt_dict=litellm.custom_prompt_dict,
|
||||
model_response=model_response,
|
||||
print_verbose=print_verbose,
|
||||
optional_params=optional_params,
|
||||
litellm_params=litellm_params,
|
||||
shared_session=shared_session,
|
||||
timeout=timeout, # type: ignore
|
||||
client=client, # pass AsyncOpenAI, OpenAI client
|
||||
custom_llm_provider=custom_llm_provider,
|
||||
logger_fn=logger_fn,
|
||||
encoding=encoding,
|
||||
stream=stream,
|
||||
)
|
||||
except Exception as e:
|
||||
## LOGGING - log the original exception returned
|
||||
logging.post_call(
|
||||
input=messages,
|
||||
api_key=api_key,
|
||||
original_response=str(e),
|
||||
additional_args={"headers": headers},
|
||||
logging_obj=logging,
|
||||
headers=headers,
|
||||
timeout=timeout,
|
||||
client=client,
|
||||
custom_llm_provider=custom_llm_provider,
|
||||
)
|
||||
raise e
|
||||
if optional_params.get("stream", False) or acompletion is True:
|
||||
## LOGGING
|
||||
logging.post_call(
|
||||
input=messages,
|
||||
api_key=api_key,
|
||||
original_response=response,
|
||||
)
|
||||
response = response
|
||||
else:
|
||||
# Non-Claude models use standard Azure AI flow
|
||||
api_base = AzureFoundryModelInfo.get_api_base(api_base)
|
||||
# set API KEY
|
||||
api_key = AzureFoundryModelInfo.get_api_key(api_key)
|
||||
|
||||
if optional_params.get("stream", False):
|
||||
## LOGGING
|
||||
logging.post_call(
|
||||
input=messages,
|
||||
api_key=api_key,
|
||||
original_response=response,
|
||||
additional_args={"headers": headers},
|
||||
)
|
||||
headers = headers or litellm.headers
|
||||
|
||||
if extra_headers is not None:
|
||||
optional_params["extra_headers"] = extra_headers
|
||||
|
||||
## FOR COHERE
|
||||
if "command-r" in model: # make sure tool call in messages are str
|
||||
messages = stringify_json_tool_call_content(messages=messages)
|
||||
|
||||
## COMPLETION CALL
|
||||
try:
|
||||
response = base_llm_http_handler.completion(
|
||||
model=model,
|
||||
messages=messages,
|
||||
headers=headers,
|
||||
model_response=model_response,
|
||||
api_key=api_key,
|
||||
api_base=api_base,
|
||||
acompletion=acompletion,
|
||||
logging_obj=logging,
|
||||
optional_params=optional_params,
|
||||
litellm_params=litellm_params,
|
||||
shared_session=shared_session,
|
||||
timeout=timeout, # type: ignore
|
||||
client=client, # pass AsyncOpenAI, OpenAI client
|
||||
custom_llm_provider=custom_llm_provider,
|
||||
encoding=encoding,
|
||||
stream=stream,
|
||||
)
|
||||
except Exception as e:
|
||||
## LOGGING - log the original exception returned
|
||||
logging.post_call(
|
||||
input=messages,
|
||||
api_key=api_key,
|
||||
original_response=str(e),
|
||||
additional_args={"headers": headers},
|
||||
)
|
||||
raise e
|
||||
|
||||
if optional_params.get("stream", False):
|
||||
## LOGGING
|
||||
logging.post_call(
|
||||
input=messages,
|
||||
api_key=api_key,
|
||||
original_response=response,
|
||||
additional_args={"headers": headers},
|
||||
)
|
||||
elif (
|
||||
custom_llm_provider == "text-completion-openai"
|
||||
or "ft:babbage-002" in model
|
||||
|
|
@ -2359,70 +2413,6 @@ def completion( # type: ignore # noqa: PLR0915
|
|||
original_response=response,
|
||||
)
|
||||
response = response
|
||||
elif custom_llm_provider == "azure_anthropic":
|
||||
# Azure Anthropic uses same API as Anthropic but with Azure authentication
|
||||
api_key = (
|
||||
api_key
|
||||
or litellm.azure_key
|
||||
or litellm.api_key
|
||||
or get_secret("AZURE_API_KEY")
|
||||
or get_secret("AZURE_OPENAI_API_KEY")
|
||||
)
|
||||
custom_prompt_dict = custom_prompt_dict or litellm.custom_prompt_dict
|
||||
# Azure Foundry endpoint format: https://<resource-name>.services.ai.azure.com/anthropic/v1/messages
|
||||
api_base = (
|
||||
api_base
|
||||
or litellm.api_base
|
||||
or get_secret("AZURE_API_BASE")
|
||||
)
|
||||
|
||||
if api_base is None:
|
||||
raise ValueError(
|
||||
"Missing Azure API Base - Please set `api_base` or `AZURE_API_BASE` environment variable. "
|
||||
"Expected format: https://<resource-name>.services.ai.azure.com/anthropic"
|
||||
)
|
||||
|
||||
# Ensure the URL ends with /v1/messages
|
||||
api_base = api_base.rstrip("/")
|
||||
if api_base.endswith("/v1/messages"):
|
||||
pass
|
||||
elif api_base.endswith("/anthropic/v1/messages"):
|
||||
pass
|
||||
else:
|
||||
if "/anthropic" in api_base:
|
||||
parts = api_base.split("/anthropic", 1)
|
||||
api_base = parts[0] + "/anthropic"
|
||||
else:
|
||||
api_base = api_base + "/anthropic"
|
||||
api_base = api_base + "/v1/messages"
|
||||
|
||||
response = azure_anthropic_chat_completions.completion(
|
||||
model=model,
|
||||
messages=messages,
|
||||
api_base=api_base,
|
||||
acompletion=acompletion,
|
||||
custom_prompt_dict=litellm.custom_prompt_dict,
|
||||
model_response=model_response,
|
||||
print_verbose=print_verbose,
|
||||
optional_params=optional_params,
|
||||
litellm_params=litellm_params,
|
||||
logger_fn=logger_fn,
|
||||
encoding=encoding, # for calculating input/output tokens
|
||||
api_key=api_key,
|
||||
logging_obj=logging,
|
||||
headers=headers,
|
||||
timeout=timeout,
|
||||
client=client,
|
||||
custom_llm_provider=custom_llm_provider,
|
||||
)
|
||||
if optional_params.get("stream", False) or acompletion is True:
|
||||
## LOGGING
|
||||
logging.post_call(
|
||||
input=messages,
|
||||
api_key=api_key,
|
||||
original_response=response,
|
||||
)
|
||||
response = response
|
||||
elif custom_llm_provider == "nlp_cloud":
|
||||
nlp_cloud_key = (
|
||||
api_key
|
||||
|
|
@ -6236,9 +6226,9 @@ async def ahealth_check(
|
|||
"x-ms-region": str,
|
||||
}
|
||||
"""
|
||||
from litellm.litellm_core_utils.health_check_helpers import HealthCheckHelpers
|
||||
from litellm.litellm_core_utils.cached_imports import get_litellm_logging_class
|
||||
|
||||
from litellm.litellm_core_utils.health_check_helpers import HealthCheckHelpers
|
||||
|
||||
# Use cached import helper to lazy-load Logging class (only loads when function is called)
|
||||
Logging = get_litellm_logging_class()
|
||||
|
||||
|
|
|
|||
|
|
@ -1151,7 +1151,7 @@
|
|||
},
|
||||
"azure/claude-haiku-4-5": {
|
||||
"input_cost_per_token": 1e-06,
|
||||
"litellm_provider": "azure_anthropic",
|
||||
"litellm_provider": "azure_ai",
|
||||
"max_input_tokens": 200000,
|
||||
"max_output_tokens": 64000,
|
||||
"max_tokens": 64000,
|
||||
|
|
@ -1169,7 +1169,7 @@
|
|||
},
|
||||
"azure/claude-opus-4-1": {
|
||||
"input_cost_per_token": 1.5e-05,
|
||||
"litellm_provider": "azure_anthropic",
|
||||
"litellm_provider": "azure_ai",
|
||||
"max_input_tokens": 200000,
|
||||
"max_output_tokens": 32000,
|
||||
"max_tokens": 32000,
|
||||
|
|
@ -1187,7 +1187,7 @@
|
|||
},
|
||||
"azure/claude-sonnet-4-5": {
|
||||
"input_cost_per_token": 3e-06,
|
||||
"litellm_provider": "azure_anthropic",
|
||||
"litellm_provider": "azure_ai",
|
||||
"max_input_tokens": 200000,
|
||||
"max_output_tokens": 64000,
|
||||
"max_tokens": 64000,
|
||||
|
|
@ -21570,6 +21570,116 @@
|
|||
"mode": "chat",
|
||||
"output_cost_per_token": 2.8e-07
|
||||
},
|
||||
"publicai/swiss-ai/apertus-8b-instruct": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 8192,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 8192,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"publicai/swiss-ai/apertus-70b-instruct": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 8192,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 8192,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"publicai/aisingapore/Gemma-SEA-LION-v4-27B-IT": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 8192,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 8192,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"publicai/BSC-LT/salamandra-7b-instruct-tools-16k": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 16384,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 16384,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"publicai/BSC-LT/ALIA-40b-instruct_Q8_0": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 8192,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 8192,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"publicai/allenai/Olmo-3-7B-Instruct": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 32768,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 32768,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"publicai/aisingapore/Qwen-SEA-LION-v4-32B-IT": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 32768,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 32768,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"publicai/allenai/Olmo-3-7B-Think": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 32768,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 32768,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_reasoning": true
|
||||
},
|
||||
"publicai/allenai/Olmo-3-32B-Think": {
|
||||
"input_cost_per_token": 0.0,
|
||||
"litellm_provider": "publicai",
|
||||
"max_input_tokens": 32768,
|
||||
"max_output_tokens": 4096,
|
||||
"max_tokens": 32768,
|
||||
"mode": "chat",
|
||||
"output_cost_per_token": 0.0,
|
||||
"source": "https://platform.publicai.co/docs",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_reasoning": true
|
||||
},
|
||||
"qwen.qwen3-coder-480b-a35b-v1:0": {
|
||||
"input_cost_per_token": 2.2e-07,
|
||||
"litellm_provider": "bedrock_converse",
|
||||
|
|
|
|||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -1 +0,0 @@
|
|||
"use strict";(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[3665],{84566:function(e,t,s){s.d(t,{GH$:function(){return l}});var c=s(2265);let l=({color:e="currentColor",size:t=24,className:s,...l})=>c.createElement("svg",{viewBox:"0 0 24 24",xmlns:"http://www.w3.org/2000/svg",width:t,height:t,fill:e,...l,className:"remixicon "+(s||"")},c.createElement("path",{d:"M12 22C6.47715 22 2 17.5228 2 12C2 6.47715 6.47715 2 12 2C17.5228 2 22 6.47715 22 12C22 17.5228 17.5228 22 12 22ZM12 20C16.4183 20 20 16.4183 20 12C20 7.58172 16.4183 4 12 4C7.58172 4 4 7.58172 4 12C4 16.4183 7.58172 20 12 20ZM11.0026 16L6.75999 11.7574L8.17421 10.3431L11.0026 13.1716L16.6595 7.51472L18.0737 8.92893L11.0026 16Z"}))}}]);
|
||||
|
|
@ -0,0 +1 @@
|
|||
"use strict";(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[3665],{84566:function(e,t,s){s.d(t,{GH$:function(){return l}});var c=s(2265);let l=({color:e="currentColor",size:t=24,className:s,...l})=>c.createElement("svg",{viewBox:"0 0 24 24",xmlns:"http://www.w3.org/2000/svg",width:t,height:t,fill:e,...l,className:"remixicon "+(s||"")},c.createElement("path",{d:"M4 12C4 7.58172 7.58172 4 12 4C16.4183 4 20 7.58172 20 12C20 16.4183 16.4183 20 12 20C7.58172 20 4 16.4183 4 12ZM12 2C6.47715 2 2 6.47715 2 12C2 17.5228 6.47715 22 12 22C17.5228 22 22 17.5228 22 12C22 6.47715 17.5228 2 12 2ZM17.4571 9.45711L16.0429 8.04289L11 13.0858L8.20711 10.2929L6.79289 11.7071L11 15.9142L17.4571 9.45711Z"}))}}]);
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Reference in a new issue