Merge branch 'BerriAI:main' into main

This commit is contained in:
abbas jafari 2025-11-28 15:28:52 +01:00 • committed by GitHub
commit 23b737d2a4
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
410 changed files with 4540 additions and 1631 deletions

View file

@ -327,6 +327,81 @@ curl --location 'http://0.0.0.0:4000/v1/messages' \
</TabItem>
</Tabs>
## Usage - Azure Anthropic (Azure Foundry Claude)
LiteLLM funnels Azure Claude deployments through the `azure_ai/` provider so Claude Opus models on Azure Foundry keep working with Tool Search, Effort, streaming, and the rest of the advanced feature set. Point `AZURE_AI_API_BASE` to `https://<resource>.services.ai.azure.com/anthropic` (LiteLLM appends `/v1/messages` automatically) and authenticate with `AZURE_AI_API_KEY` or an Azure AD token.
<Tabs>
<TabItem value="sdk" label="LiteLLM Python SDK">
```python
import os
from litellm import completion
# Configure Azure credentials
os.environ["AZURE_AI_API_KEY"] = "your-azure-ai-api-key"
os.environ["AZURE_AI_API_BASE"] = "https://my-resource.services.ai.azure.com/anthropic"
response = completion(
model="azure_ai/claude-opus-4-1",
messages=[{"role": "user", "content": "Explain how Azure Anthropic hosts Claude Opus differently from the public Anthropic API."}],
max_tokens=1200,
temperature=0.7,
stream=True,
)
for chunk in response:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
```
</TabItem>
<TabItem value="proxy" label="LiteLLM Proxy">
**1. Set environment variables**
```bash
export AZURE_AI_API_KEY="your-azure-ai-api-key"
export AZURE_AI_API_BASE="https://my-resource.services.ai.azure.com/anthropic"
```
**2. Configure the proxy**
```yaml
model_list:
- model_name: claude-4-azure
litellm_params:
model: azure_ai/claude-opus-4-1
api_key: os.environ/AZURE_AI_API_KEY
api_base: os.environ/AZURE_AI_API_BASE
```
**3. Start LiteLLM**
```bash
litellm --config /path/to/config.yaml
```
**4. Test the Azure Claude route**
```bash
curl --location 'http://0.0.0.0:4000/chat/completions' \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer $LITELLM_KEY' \
--data '{
"model": "claude-4-azure",
"messages": [
{
"role": "user",
"content": "How do I use Claude Opus 4 via Azure Anthropic in LiteLLM?"
}
],
"max_tokens": 1024
}'
```
</TabItem>
</Tabs>
## Tool Search {#tool-search}

View file

@ -0,0 +1,22 @@
# OpenAI Agents SDK
The [OpenAI Agents SDK](https://github.com/openai/openai-agents-python) is a lightweight framework for building multi-agent workflows.
It includes an official LiteLLM extension that lets you use any of the 100+ supported providers (Anthropic, Gemini, Mistral, Bedrock, etc.)
```python
from agents import Agent, Runner
from agents.extensions.models.litellm_model import LitellmModel
agent = Agent(
name="Assistant",
instructions="You are a helpful assistant.",
model=LitellmModel(model="provider/model-name")
)
result = Runner.run_sync(agent, "your_prompt_here")
print("Result:", result.final_output)
```
- [GitHub](https://github.com/openai/openai-agents-python)
- [LiteLLM Extension Docs](https://openai.github.io/openai-agents-python/ref/extensions/litellm/)

View file

@ -16,7 +16,7 @@ Azure Foundry supports the following Claude models:
| Property | Details |
|-------|-------|
| Description | Claude models deployed via Microsoft Azure Foundry. Uses the same API as Anthropic's Messages API but with Azure authentication. |
| Provider Route on LiteLLM | `azure/` (add this prefix to Claude model names - e.g. `azure/claude-sonnet-4-5`) |
| Provider Route on LiteLLM | `azure_ai/` (add this prefix to Claude model names - e.g. `azure_ai/claude-sonnet-4-5`) |
| Provider Doc | [Azure Foundry Claude Models ↗](https://learn.microsoft.com/en-us/azure/ai-services/foundry-models/claude) |
| API Endpoint | `https://<resource-name>.services.ai.azure.com/anthropic/v1/messages` |
| Supported Endpoints | `/chat/completions`, `/anthropic/v1/messages`|
@ -68,7 +68,7 @@ os.environ["AZURE_API_BASE"] = "https://<resource-name>.services.ai.azure.com/an
# Make a completion request
response = completion(
model="azure/claude-sonnet-4-5",
model="azure_ai/claude-sonnet-4-5",
messages=[
{"role": "user", "content": "What are 3 things to visit in Seattle?"}
],
@ -85,7 +85,7 @@ print(response)
import litellm
response = litellm.completion(
model="azure/claude-sonnet-4-5",
model="azure_ai/claude-sonnet-4-5",
api_base="https://<resource-name>.services.ai.azure.com/anthropic",
api_key="your-azure-api-key",
messages=[
@ -101,7 +101,7 @@ response = litellm.completion(
import litellm
response = litellm.completion(
model="azure/claude-sonnet-4-5",
model="azure_ai/claude-sonnet-4-5",
api_base="https://<resource-name>.services.ai.azure.com/anthropic",
azure_ad_token="your-azure-ad-token",
messages=[
@ -117,7 +117,7 @@ response = litellm.completion(
from litellm import completion
response = completion(
model="azure/claude-sonnet-4-5",
model="azure_ai/claude-sonnet-4-5",
messages=[
{"role": "user", "content": "Write a short story"}
],
@ -136,7 +136,7 @@ for chunk in response:
from litellm import completion
response = completion(
model="azure/claude-sonnet-4-5",
model="azure_ai/claude-sonnet-4-5",
messages=[
{"role": "user", "content": "What's the weather in Seattle?"}
],
@ -181,7 +181,7 @@ export AZURE_API_BASE="https://<resource-name>.services.ai.azure.com/anthropic"
model_list:
- model_name: claude-sonnet-4-5
litellm_params:
model: azure/claude-sonnet-4-5
model: azure_ai/claude-sonnet-4-5
api_base: https://<resource-name>.services.ai.azure.com/anthropic
api_key: os.environ/AZURE_API_KEY
```
@ -331,7 +331,7 @@ os.environ["AZURE_API_BASE"] = "https://my-resource.services.ai.azure.com/anthro
# Make a request
response = completion(
model="azure/claude-sonnet-4-5",
model="azure_ai/claude-sonnet-4-5",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in simple terms."}
@ -358,7 +358,7 @@ Or pass it directly:
```python
response = completion(
model="azure/claude-sonnet-4-5",
model="azure_ai/claude-sonnet-4-5",
api_base="https://<resource-name>.services.ai.azure.com/anthropic",
# ...
)

View file

@ -0,0 +1,209 @@
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
# PublicAI
## Overview
| Property | Details |
|-------|-------|
| Description | PublicAI provides large language models including essential models like the swiss-ai apertus model. |
| Provider Route on LiteLLM | `publicai/` |
| Link to Provider Doc | [PublicAI ↗](https://platform.publicai.co/) |
| Base URL | `https://platform.publicai.co/` |
| Supported Operations | [`/chat/completions`](#sample-usage) |
<br />
<br />
https://platform.publicai.co/
**We support ALL PublicAI models, just set `publicai/` as a prefix when sending completion requests**
## Required Variables
```python showLineNumbers title="Environment Variables"
os.environ["PUBLICAI_API_KEY"] = "" # your PublicAI API key
```
You can overwrite the base url with:
```
os.environ["PUBLICAI_API_BASE"] = "https://platform.publicai.co/v1"
```
## Usage - LiteLLM Python SDK
### Non-streaming
```python showLineNumbers title="PublicAI Non-streaming Completion"
import os
import litellm
from litellm import completion
os.environ["PUBLICAI_API_KEY"] = "" # your PublicAI API key
messages = [{"content": "Hello, how are you?", "role": "user"}]
# PublicAI call
response = completion(
model="publicai/swiss-ai/apertus-8b-instruct",
messages=messages
)
print(response)
```
### Streaming
```python showLineNumbers title="PublicAI Streaming Completion"
import os
import litellm
from litellm import completion
os.environ["PUBLICAI_API_KEY"] = "" # your PublicAI API key
messages = [{"content": "Hello, how are you?", "role": "user"}]
# PublicAI call with streaming
response = completion(
model="publicai/swiss-ai/apertus-8b-instruct",
messages=messages,
stream=True
)
for chunk in response:
print(chunk)
```
## Usage - LiteLLM Proxy
Add the following to your LiteLLM Proxy configuration file:
```yaml showLineNumbers title="config.yaml"
model_list:
- model_name: swiss-ai-apertus-8b
litellm_params:
model: publicai/swiss-ai/apertus-8b-instruct
api_key: os.environ/PUBLICAI_API_KEY
- model_name: swiss-ai-apertus-70b
litellm_params:
model: publicai/swiss-ai/apertus-70b-instruct
api_key: os.environ/PUBLICAI_API_KEY
```
Start your LiteLLM Proxy server:
```bash showLineNumbers title="Start LiteLLM Proxy"
litellm --config config.yaml
# RUNNING on http://0.0.0.0:4000
```
<Tabs>
<TabItem value="openai-sdk" label="OpenAI SDK">
```python showLineNumbers title="PublicAI via Proxy - Non-streaming"
from openai import OpenAI
# Initialize client with your proxy URL
client = OpenAI(
base_url="http://localhost:4000", # Your proxy URL
api_key="your-proxy-api-key" # Your proxy API key
)
# Non-streaming response
response = client.chat.completions.create(
model="swiss-ai-apertus-8b",
messages=[{"role": "user", "content": "hello from litellm"}]
)
print(response.choices[0].message.content)
```
```python showLineNumbers title="PublicAI via Proxy - Streaming"
from openai import OpenAI
# Initialize client with your proxy URL
client = OpenAI(
base_url="http://localhost:4000", # Your proxy URL
api_key="your-proxy-api-key" # Your proxy API key
)
# Streaming response
response = client.chat.completions.create(
model="swiss-ai-apertus-8b",
messages=[{"role": "user", "content": "hello from litellm"}],
stream=True
)
for chunk in response:
if chunk.choices[0].delta.content is not None:
print(chunk.choices[0].delta.content, end="")
```
</TabItem>
<TabItem value="litellm-sdk" label="LiteLLM SDK">
```python showLineNumbers title="PublicAI via Proxy - LiteLLM SDK"
import litellm
# Configure LiteLLM to use your proxy
response = litellm.completion(
model="litellm_proxy/swiss-ai-apertus-8b",
messages=[{"role": "user", "content": "hello from litellm"}],
api_base="http://localhost:4000",
api_key="your-proxy-api-key"
)
print(response.choices[0].message.content)
```
```python showLineNumbers title="PublicAI via Proxy - LiteLLM SDK Streaming"
import litellm
# Configure LiteLLM to use your proxy with streaming
response = litellm.completion(
model="litellm_proxy/swiss-ai-apertus-8b",
messages=[{"role": "user", "content": "hello from litellm"}],
api_base="http://localhost:4000",
api_key="your-proxy-api-key",
stream=True
)
for chunk in response:
if hasattr(chunk.choices[0], 'delta') and chunk.choices[0].delta.content is not None:
print(chunk.choices[0].delta.content, end="")
```
</TabItem>
<TabItem value="curl" label="cURL">
```bash showLineNumbers title="PublicAI via Proxy - cURL"
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-proxy-api-key" \
-d '{
"model": "swiss-ai-apertus-8b",
"messages": [{"role": "user", "content": "hello from litellm"}]
}'
```
```bash showLineNumbers title="PublicAI via Proxy - cURL Streaming"
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-proxy-api-key" \
-d '{
"model": "swiss-ai-apertus-8b",
"messages": [{"role": "user", "content": "hello from litellm"}],
"stream": true
}'
```
</TabItem>
</Tabs>
For more detailed information on using the LiteLLM Proxy, see the [LiteLLM Proxy documentation](../providers/litellm_proxy).

View file

@ -576,6 +576,8 @@ router_settings:
| GENERIC_USER_PROVIDER_ATTRIBUTE | Attribute specifying the user's provider
| GENERIC_USER_ROLE_ATTRIBUTE | Attribute specifying the user's role
| GENERIC_USERINFO_ENDPOINT | Endpoint to fetch user information in generic OAuth
| GENERIC_LOGGER_ENDPOINT | Endpoint URL for the Generic Logger callback to send logs to
| GENERIC_LOGGER_HEADERS | JSON string of headers to include in Generic Logger callback requests
| GEMINI_API_BASE | Base URL for Gemini API. Default is https://generativelanguage.googleapis.com
| GALILEO_BASE_URL | Base URL for Galileo platform
| GALILEO_PASSWORD | Password for Galileo authentication

View file

@ -7,9 +7,38 @@ import TabItem from '@theme/TabItem';
LiteLLM provides the LiteLLM Tool Permission Guardrail that lets you control which **tool calls** a model is allowed to invoke, using configurable allow/deny rules. This offers fine-grained, provider-agnostic control over tool execution (e.g., OpenAI Chat Completions `tool_calls`, Anthropic Messages `tool_use`, MCP tools).
## Quick Start
### 1. Define Guardrails on your LiteLLM config.yaml
Define your guardrails under the `guardrails` section
### LiteLLM UI
#### Step 1: Select Tool Permission Guardrail
Open the LiteLLM Dashboard, click **Add New Guardrail**, and choose **LiteLLM Tool Permission Guardrail**. This loads the rule builder UI.
<Image img={require('../../../img/create_guard_tool_permission.png')} alt="Configure tool permission guardrail in LiteLLM UI" />
#### Step 2: Define Regex Rules
1. Click **Add Rule**.
2. Enter a unique Rule ID.
3. Provide a regex for the tool name (e.g., `^mcp__github_.*$`).
4. Optionally add a regex for tool type (e.g., `^function$`).
5. Pick **Allow** or **Deny**.
<Image img={require('../../../img/create_rule_tool_permission.png')} alt="Configure tool permission guardrail in LiteLLM UI" />
#### Step 3: Restrict Tool Arguments (Optional)
Select **+ Restrict tool arguments** to attach regex validations to nested paths (dot + `[]` notation). This enforces that sensitive parameters (such as `arguments.to[]`) conform to pre-approved formats.
#### Step 4: Choose Defaults & Actions
- Set the fallback decision (`default_action`) for tools that do not hit any rule.
- Decide how disallowed tools behave: **Block** halts the request, **Rewrite** strips forbidden tools and returns an error message inside the response.
- Customize `violation_message_template` if you want branded error copy.
- Save the guardrail.
### LiteLLM Config.yaml Setup
```yaml
guardrails:
- guardrail_name: "tool-permission-guardrail"
@ -21,16 +50,17 @@ guardrails:
tool_name: "Bash"
decision: "allow"
- id: "allow_github_mcp"
tool_name: "mcp__github_*"
tool_name: "^mcp__github_.*$"
decision: "allow"
- id: "allow_aws_documentation"
tool_name: "mcp__aws-documentation_*_documentation"
tool_name: "^mcp__aws-documentation_.*_documentation$"
decision: "allow"
- id: "deny_read_commands"
tool_name: "Read"
decision: "Deny"
decision: "deny"
- id: "mail-domain"
tool_name: "send_email"
tool_name: "^send_email$"
tool_type: "^function$"
decision: "allow"
allowed_param_patterns:
"to[]": "^.+@berri\\.ai$"
@ -44,7 +74,8 @@ guardrails:
```yaml
- id: "unique_rule_id" # Unique identifier for the rule
tool_name: "pattern" # Tool name or pattern to match
tool_name: "^regex$" # Regex for tool name (optional, at least one of name/type required)
tool_type: "^function$" # Regex for tool type (optional)
decision: "allow" # "allow" or "deny"
allowed_param_patterns: # Optional - regex map for argument paths (dot + [] notation)
"path.to[].field": "^regex$"

View file

@ -0,0 +1,250 @@
# Guardrails on Pass-Through Endpoints
import Image from '@theme/IdealImage';
## Overview
| Property | Details |
|----------|---------|
| Description | Enable guardrail execution on LiteLLM pass-through endpoints with opt-in activation and automatic inheritance from org/team/key levels |
| Supported Guardrails | All LiteLLM guardrails (Bedrock, Aporia, Lakera, etc.) |
| Default Behavior | Guardrails are **disabled** on pass-through endpoints unless explicitly enabled |
## Quick Start
You can configure guardrails on pass-through endpoints either via the **UI** (recommended) or **config file**.
### Using the UI
#### 1. Navigate to Pass-Through Endpoints
Go to **Models + Endpoints** → Click **+ Add Pass-Through Endpoint**
<Image img={require('../../img/pt_guard1.png')} alt="Add guardrails to pass-through endpoint" />
Scroll to the **Guardrails** section and select which guardrails to enforce.
:::tip Default Behavior
By default, you don't need to specify fields - LiteLLM will JSON dump the entire request/response payload and send it to the guardrail.
:::
#### 2. Target Specific Fields (Optional)
<Image img={require('../../img/pt_guard2.png')} alt="Configure field-level targeting" />
To check only specific fields instead of the entire payload:
1. Select your guardrails
2. In **Field Targeting (Optional)**, specify fields for each guardrail
3. Use the quick-add buttons (`+ query`, `+ documents[*]`) or type custom JSONPath expressions
4. **Request Fields (pre_call)**: Fields to check before sending to target API
5. **Response Fields (post_call)**: Fields to check in the response from target API
**Example**: In the screenshot above, we set `query` as a request field, so only the `query` field is sent to the guardrail instead of the entire request.
---
### Using Config File
#### 1. Define guardrails and pass-through endpoint
```yaml showLineNumbers title="config.yaml"
guardrails:
- guardrail_name: "pii-guard"
litellm_params:
guardrail: bedrock
mode: pre_call
guardrailIdentifier: "your-guardrail-id"
guardrailVersion: "1"
general_settings:
pass_through_endpoints:
- path: "/v1/rerank"
target: "https://api.cohere.com/v1/rerank"
headers:
Authorization: "bearer os.environ/COHERE_API_KEY"
guardrails:
pii-guard:
```
#### 2. Start proxy
```bash
litellm --config config.yaml
```
#### 3. Test request
```bash
curl -X POST "http://localhost:4000/v1/rerank" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-1234" \
-d '{
"model": "rerank-english-v3.0",
"query": "What is the capital of France?",
"documents": ["Paris is the capital of France."]
}'
```
---
## Opt-In Behavior
| Configuration | Behavior |
|--------------|----------|
| `guardrails` not set | No guardrails execute (default) |
| `guardrails` set | All org/team/key + pass-through guardrails execute |
When guardrails are enabled, the system collects and executes:
- Org-level guardrails
- Team-level guardrails
- Key-level guardrails
- Pass-through specific guardrails
---
## How It Works
The diagram below shows what happens when a client makes a request to `/special/rerank` - a pass-through endpoint configured with guardrails in your `config.yaml`.
When guardrails are configured on a pass-through endpoint:
1. **Pre-call guardrails** run on the request before forwarding to the target API
2. If `request_fields` is specified (e.g., `["query"]`), only those fields are sent to the guardrail. Otherwise, the entire request payload is evaluated.
3. The request is forwarded to the target API only if guardrails pass
4. **Post-call guardrails** run on the response from the target API
5. If `response_fields` is specified (e.g., `["results[*].text"]`), only those fields are evaluated. Otherwise, the entire response is checked.
:::info
If the `guardrails` block is omitted or empty in your pass-through endpoint config, the request skips the guardrail flow entirely and goes directly to the target API.
:::
```mermaid
sequenceDiagram
participant Client
box rgb(200, 220, 255) LiteLLM Proxy
participant PassThrough as Pass-through Endpoint
participant Guardrails
end
participant Target as Target API (Cohere, etc.)
Client->>PassThrough: POST /special/rerank
Note over PassThrough,Guardrails: Collect passthrough + org/team/key guardrails
PassThrough->>Guardrails: Run pre_call (request_fields or full payload)
Guardrails-->>PassThrough: ✓ Pass / ✗ Block
PassThrough->>Target: Forward request
Target-->>PassThrough: Response
PassThrough->>Guardrails: Run post_call (response_fields or full payload)
Guardrails-->>PassThrough: ✓ Pass / ✗ Block
PassThrough-->>Client: Return response (or error)
```
---
## Field-Level Targeting
Target specific JSON fields instead of the entire request/response payload.
```yaml showLineNumbers title="config.yaml"
guardrails:
- guardrail_name: "pii-detection"
litellm_params:
guardrail: bedrock
mode: pre_call
guardrailIdentifier: "pii-guard-id"
guardrailVersion: "1"
- guardrail_name: "content-moderation"
litellm_params:
guardrail: bedrock
mode: post_call
guardrailIdentifier: "content-guard-id"
guardrailVersion: "1"
general_settings:
pass_through_endpoints:
- path: "/v1/rerank"
target: "https://api.cohere.com/v1/rerank"
headers:
Authorization: "bearer os.environ/COHERE_API_KEY"
guardrails:
pii-detection:
request_fields: ["query", "documents[*].text"]
content-moderation:
response_fields: ["results[*].text"]
```
### Field Options
| Field | Description |
|-------|-------------|
| `request_fields` | JSONPath expressions for input (pre_call) |
| `response_fields` | JSONPath expressions for output (post_call) |
| Neither specified | Guardrail runs on entire payload |
### JSONPath Examples
| Expression | Matches |
|------------|---------|
| `query` | Single field named `query` |
| `documents[*].text` | All `text` fields in `documents` array |
| `messages[*].content` | All `content` fields in `messages` array |
---
## Configuration Examples
### Single guardrail on entire payload
```yaml showLineNumbers title="config.yaml"
guardrails:
- guardrail_name: "pii-detection"
litellm_params:
guardrail: bedrock
mode: pre_call
guardrailIdentifier: "your-id"
guardrailVersion: "1"
general_settings:
pass_through_endpoints:
- path: "/v1/rerank"
target: "https://api.cohere.com/v1/rerank"
guardrails:
pii-detection:
```
### Multiple guardrails with mixed settings
```yaml showLineNumbers title="config.yaml"
guardrails:
- guardrail_name: "pii-detection"
litellm_params:
guardrail: bedrock
mode: pre_call
guardrailIdentifier: "pii-id"
guardrailVersion: "1"
- guardrail_name: "content-moderation"
litellm_params:
guardrail: bedrock
mode: post_call
guardrailIdentifier: "content-id"
guardrailVersion: "1"
- guardrail_name: "prompt-injection"
litellm_params:
guardrail: lakera
mode: pre_call
api_key: os.environ/LAKERA_API_KEY
general_settings:
pass_through_endpoints:
- path: "/v1/rerank"
target: "https://api.cohere.com/v1/rerank"
guardrails:
pii-detection:
request_fields: ["input", "query"]
content-moderation:
prompt-injection:
request_fields: ["messages[*].content"]
```

Binary file not shown.

After

Width:  |  Height:  |  Size: 769 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 548 KiB

View file

@ -419,7 +419,8 @@ const sidebars = {
]
},
"pass_through/vllm",
"proxy/pass_through"
"proxy/pass_through",
"proxy/pass_through_guardrails"
]
},
"rag_ingest",
@ -621,6 +622,7 @@ const sidebars = {
"providers/ovhcloud",
"providers/perplexity",
"providers/petals",
"providers/publicai",
"providers/predibase",
"providers/recraft",
"providers/replicate",
@ -814,9 +816,10 @@ const sidebars = {
"Learn how to deploy + call models from different providers on LiteLLM",
slug: "/project",
},
items: [
items: [
"projects/smolagents",
"projects/mini-swe-agent",
"projects/openai-agents",
"projects/Docq.AI",
"projects/PDL",
"projects/OpenInterpreter",

View file

@ -555,6 +555,7 @@ deepgram_models: Set = set()
elevenlabs_models: Set = set()
dashscope_models: Set = set()
moonshot_models: Set = set()
publicai_models: Set = set()
v0_models: Set = set()
morph_models: Set = set()
lambda_ai_models: Set = set()
@ -781,6 +782,8 @@ def add_known_models():
dashscope_models.add(key)
elif value.get("litellm_provider") == "moonshot":
moonshot_models.add(key)
elif value.get("litellm_provider") == "publicai":
publicai_models.add(key)
elif value.get("litellm_provider") == "v0":
v0_models.add(key)
elif value.get("litellm_provider") == "morph":
@ -899,6 +902,7 @@ model_list = list(
| elevenlabs_models
| dashscope_models
| moonshot_models
| publicai_models
| v0_models
| morph_models
| lambda_ai_models
@ -992,6 +996,7 @@ models_by_provider: dict = {
"heroku": heroku_models,
"dashscope": dashscope_models,
"moonshot": moonshot_models,
"publicai": publicai_models,
"v0": v0_models,
"morph": morph_models,
"lambda_ai": lambda_ai_models,
@ -1120,7 +1125,7 @@ from .llms.openrouter.chat.transformation import OpenrouterConfig
from .llms.datarobot.chat.transformation import DataRobotConfig
from .llms.anthropic.chat.transformation import AnthropicConfig
from .llms.anthropic.common_utils import AnthropicModelInfo
from .llms.azure.anthropic.transformation import AzureAnthropicConfig
from .llms.azure_ai.anthropic.transformation import AzureAnthropicConfig
from .llms.groq.stt.transformation import GroqSTTConfig
from .llms.anthropic.completion.transformation import AnthropicTextConfig
from .llms.triton.completion.transformation import TritonConfig
@ -1370,6 +1375,7 @@ from .llms.nebius.chat.transformation import NebiusConfig
from .llms.wandb.chat.transformation import WandbConfig
from .llms.dashscope.chat.transformation import DashScopeChatConfig
from .llms.moonshot.chat.transformation import MoonshotChatConfig
from .llms.publicai.chat.transformation import PublicAIChatConfig
from .llms.docker_model_runner.chat.transformation import DockerModelRunnerChatConfig
from .llms.v0.chat.transformation import V0ChatConfig
from .llms.oci.chat.transformation import OCIChatConfig

View file

@ -384,6 +384,7 @@ LITELLM_CHAT_PROVIDERS = [
"nebius",
"dashscope",
"moonshot",
"publicai",
"v0",
"heroku",
"oci",
@ -526,6 +527,7 @@ openai_compatible_endpoints: List = [
"api.studio.nebius.ai/v1",
"https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
"https://api.moonshot.ai/v1",
"https://platform.publicai.co/v1",
"https://api.v0.dev/v1",
"https://api.morphllm.com/v1",
"https://api.lambda.ai/v1",
@ -571,6 +573,7 @@ openai_compatible_providers: List = [
"nebius",
"dashscope",
"moonshot",
"publicai",
"v0",
"morph",
"lambda_ai",
@ -593,6 +596,7 @@ openai_text_completion_compatible_providers: List = (
"nebius",
"dashscope",
"moonshot",
"publicai",
"v0",
"lambda_ai",
"hyperbolic",

View file

@ -22,17 +22,16 @@ def _is_non_openai_azure_model(model: str) -> bool:
return False
def _is_azure_anthropic_model(model: str) -> Optional[str]:
def _is_azure_claude_model(model: str) -> bool:
"""
Check if a model name contains 'claude' (case-insensitive).
Used to detect Claude models that need Anthropic-specific handling.
"""
try:
model_parts = model.split("/", 1)
if len(model_parts) > 1:
model_name = model_parts[1].lower()
# Check if model name contains claude
if "claude" in model_name or model_name.startswith("claude"):
return model_parts[1] # Return model name without "azure/" prefix
model_lower = model.lower()
return "claude" in model_lower or model_lower.startswith("claude")
except Exception:
pass
return None
return False
def handle_cohere_chat_model_custom_llm_provider(
@ -136,11 +135,6 @@ def get_llm_provider( # noqa: PLR0915
# AZURE AI-Studio Logic - Azure AI Studio supports AZURE/Cohere
# If User passes azure/command-r-plus -> we should send it to cohere_chat/command-r-plus
if model.split("/", 1)[0] == "azure":
# Check if it's an Azure Anthropic model (claude models)
azure_anthropic_model = _is_azure_anthropic_model(model)
if azure_anthropic_model:
custom_llm_provider = "azure_anthropic"
return azure_anthropic_model, custom_llm_provider, dynamic_api_key, api_base
if _is_non_openai_azure_model(model):
custom_llm_provider = "openai"
return model, custom_llm_provider, dynamic_api_key, api_base
@ -258,6 +252,9 @@ def get_llm_provider( # noqa: PLR0915
elif endpoint == "api.moonshot.ai/v1":
custom_llm_provider = "moonshot"
dynamic_api_key = get_secret_str("MOONSHOT_API_KEY")
elif endpoint == "platform.publicai.co/v1":
custom_llm_provider = "publicai"
dynamic_api_key = get_secret_str("PUBLICAI_API_KEY")
elif endpoint == "https://api.v0.dev/v1":
custom_llm_provider = "v0"
dynamic_api_key = get_secret_str("V0_API_KEY")
@ -759,6 +756,13 @@ def _get_openai_compatible_provider_info( # noqa: PLR0915
) = litellm.MoonshotChatConfig()._get_openai_compatible_provider_info(
api_base, api_key
)
elif custom_llm_provider == "publicai":
(
api_base,
dynamic_api_key,
) = litellm.PublicAIChatConfig()._get_openai_compatible_provider_info(
api_base, api_key
)
elif custom_llm_provider == "docker_model_runner":
(
api_base,

View file

@ -374,6 +374,9 @@ class Logging(LiteLLMLoggingBaseClass):
# Init Caching related details
self.caching_details: Optional[CachingDetails] = None
# Passthrough endpoint guardrails config for field targeting
self.passthrough_guardrails_config: Optional[Dict[str, Any]] = None
self.model_call_details: Dict[str, Any] = {
"litellm_trace_id": litellm_trace_id,
"litellm_call_id": litellm_call_id,

View file

@ -41,6 +41,7 @@ class AnthropicMessagesHandler(BaseTranslation):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process input messages by applying guardrails to text content.
@ -145,6 +146,7 @@ class AnthropicMessagesHandler(BaseTranslation):
self,
response: "AnthropicMessagesResponse",
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process output response by applying guardrails to text content.

View file

@ -7,7 +7,6 @@ from typing import TYPE_CHECKING, Callable, Union
import httpx
import litellm
from litellm.llms.anthropic.chat.handler import AnthropicChatCompletion
from litellm.llms.custom_httpx.http_handler import (
AsyncHTTPHandler,
@ -55,7 +54,6 @@ class AzureAnthropicChatCompletion(AnthropicChatCompletion):
Completion method that uses Azure authentication instead of Anthropic's x-api-key.
All other logic is the same as AnthropicChatCompletion.
"""
from litellm.utils import ProviderConfigManager
optional_params = copy.deepcopy(optional_params)
stream = optional_params.pop("stream", None)
@ -64,8 +62,10 @@ class AzureAnthropicChatCompletion(AnthropicChatCompletion):
_is_function_call = False
messages = copy.deepcopy(messages)
# Use AzureAnthropicConfig instead of AnthropicConfig
headers = AzureAnthropicConfig().validate_environment(
# Use AzureAnthropicConfig for both azure_anthropic and azure_ai Claude models
config = AzureAnthropicConfig()
headers = config.validate_environment(
api_key=api_key,
headers=headers,
model=model,
@ -74,15 +74,6 @@ class AzureAnthropicChatCompletion(AnthropicChatCompletion):
litellm_params=litellm_params,
)
config = ProviderConfigManager.get_provider_chat_config(
model=model,
provider=litellm.types.utils.LlmProviders(custom_llm_provider),
)
if config is None:
raise ValueError(
f"Provider config not found for model: {model} and provider: {custom_llm_provider}"
)
data = config.transform_request(
model=model,
messages=messages,
@ -183,7 +174,7 @@ class AzureAnthropicChatCompletion(AnthropicChatCompletion):
return CustomStreamWrapper(
completion_stream=completion_stream,
model=model,
custom_llm_provider="azure_anthropic",
custom_llm_provider="azure_ai",
logging_obj=logging_obj,
_response_headers=process_anthropic_headers(response_headers),
)

View file

@ -21,7 +21,7 @@ class AzureAnthropicConfig(AnthropicConfig):
@property
def custom_llm_provider(self) -> Optional[str]:
return "azure_anthropic"
return "azure_ai"
def validate_environment(
self,
@ -94,3 +94,29 @@ class AzureAnthropicConfig(AnthropicConfig):
return headers
def transform_request(
self,
model: str,
messages: List[AllMessageValues],
optional_params: dict,
litellm_params: dict,
headers: dict,
) -> dict:
"""
Transform request using parent AnthropicConfig, then remove extra_body if present.
Azure Anthropic doesn't support extra_body parameter.
"""
# Call parent transform_request
data = super().transform_request(
model=model,
messages=messages,
optional_params=optional_params,
litellm_params=litellm_params,
headers=headers,
)
# Remove extra_body if present (Azure Anthropic doesn't support it)
data.pop("extra_body", None)
return data

View file

@ -1,8 +1,9 @@
from abc import ABC, abstractmethod
from typing import TYPE_CHECKING, Any
from typing import TYPE_CHECKING, Any, Optional
if TYPE_CHECKING:
from litellm.integrations.custom_guardrail import CustomGuardrail
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
class BaseTranslation(ABC):
@ -11,6 +12,7 @@ class BaseTranslation(ABC):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional["LiteLLMLoggingObj"] = None,
) -> Any:
pass
@ -19,5 +21,6 @@ class BaseTranslation(ABC):
self,
response: Any,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional["LiteLLMLoggingObj"] = None,
) -> Any:
pass

View file

@ -1,5 +1,5 @@
"""
Legacy /v1/embedding transformation logic for Bedrock Cohere.
Legacy /v1/embedding transformation logic for Bedrock Cohere.
"""
from typing import Any, List, Optional, Union
@ -123,7 +123,13 @@ class CohereEmbeddingConfig:
"""
embeddings = response_json["embeddings"]
output_data = []
is_embeddings_by_type = response_json.get("response_type") == "embeddings_by_type"
is_embeddings_by_type = (
response_json.get("response_type") == "embeddings_by_type"
)
if isinstance(embeddings, dict):
is_embeddings_by_type = True
if is_embeddings_by_type:
for embedding_type in embeddings:
for idx, embedding in enumerate(embeddings[embedding_type]):

View file

@ -5,7 +5,7 @@ This module provides guardrail translation support for the rerank endpoint.
The handler processes only the 'query' parameter for guardrails.
"""
from typing import TYPE_CHECKING, Any
from typing import TYPE_CHECKING, Any, Optional
from litellm._logging import verbose_proxy_logger
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
@ -34,6 +34,7 @@ class CohereRerankHandler(BaseTranslation):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process input query by applying guardrails.
@ -68,6 +69,7 @@ class CohereRerankHandler(BaseTranslation):
self,
response: "RerankResponse",
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process output response - not applicable for rerank.

View file

@ -42,6 +42,7 @@ class OpenAIChatCompletionsHandler(BaseTranslation):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process input messages by applying guardrails to text content.
@ -148,6 +149,7 @@ class OpenAIChatCompletionsHandler(BaseTranslation):
self,
response: "ModelResponse",
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process output response by applying guardrails to text content.

View file

@ -5,7 +5,7 @@ This module provides guardrail translation support for OpenAI's text completion
The handler processes the 'prompt' parameter for guardrails.
"""
from typing import TYPE_CHECKING, Any
from typing import TYPE_CHECKING, Any, Optional
from litellm._logging import verbose_proxy_logger
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
@ -32,6 +32,7 @@ class OpenAITextCompletionHandler(BaseTranslation):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process input prompt by applying guardrails to text content.
@ -100,6 +101,7 @@ class OpenAITextCompletionHandler(BaseTranslation):
self,
response: "TextCompletionResponse",
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process output response by applying guardrails to completion text.

View file

@ -5,7 +5,7 @@ This module provides guardrail translation support for OpenAI's image generation
The handler processes the 'prompt' parameter for guardrails.
"""
from typing import TYPE_CHECKING, Any
from typing import TYPE_CHECKING, Any, Optional
from litellm._logging import verbose_proxy_logger
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
@ -31,6 +31,7 @@ class OpenAIImageGenerationHandler(BaseTranslation):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process input prompt by applying guardrails to text content.
@ -72,6 +73,7 @@ class OpenAIImageGenerationHandler(BaseTranslation):
self,
response: "ImageResponse",
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process output response - typically not needed for image generation.

View file

@ -56,6 +56,7 @@ class OpenAIResponsesHandler(BaseTranslation):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process input by applying guardrails to text content.
@ -177,6 +178,7 @@ class OpenAIResponsesHandler(BaseTranslation):
self,
response: "ResponsesAPIResponse",
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process output response by applying guardrails to text content.

View file

@ -238,6 +238,24 @@ class OpenAIResponsesAPIConfig(BaseResponsesAPIConfig):
event_pydantic_model = OpenAIResponsesAPIConfig.get_event_model_class(
event_type=event_type
)
# Defensive: Some OpenAI-compatible providers may send `error.code: null`.
# Pydantic will raise a ValidationError when it expects a string but gets None.
# Coalesce a None `error.code` to a stable default string so streaming
# iteration does not crash (see issue report). This keeps behavior similar
# to previous fixes (coalesce before validation) and lets higher-level
# handlers still receive an `ErrorEvent` object.
try:
error_obj = parsed_chunk.get("error")
if isinstance(error_obj, dict) and error_obj.get("code") is None:
# Preserve other fields, but ensure `code` is a non-null string
parsed_chunk = dict(parsed_chunk)
parsed_chunk["error"] = dict(error_obj)
parsed_chunk["error"]["code"] = "unknown_error"
except Exception:
# If anything unexpected happens here, fall back to attempting
# instantiation and let higher-level handlers manage errors.
verbose_logger.debug("Failed to coalesce error.code in parsed_chunk")
return event_pydantic_model(**parsed_chunk)
@staticmethod

View file

@ -5,7 +5,7 @@ This module provides guardrail translation support for OpenAI's text-to-speech e
The handler processes the 'input' text parameter (output is audio, so no text to guardrail).
"""
from typing import TYPE_CHECKING, Any
from typing import TYPE_CHECKING, Any, Optional
from litellm._logging import verbose_proxy_logger
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
@ -30,6 +30,7 @@ class OpenAITextToSpeechHandler(BaseTranslation):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process input text by applying guardrails.
@ -72,6 +73,7 @@ class OpenAITextToSpeechHandler(BaseTranslation):
self,
response: "HttpxBinaryResponseContent",
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process output - not applicable for text-to-speech.

View file

@ -5,7 +5,7 @@ This module provides guardrail translation support for OpenAI's audio transcript
The handler processes the output transcribed text (input is audio, so no text to guardrail).
"""
from typing import TYPE_CHECKING, Any
from typing import TYPE_CHECKING, Any, Optional
from litellm._logging import verbose_proxy_logger
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
@ -30,6 +30,7 @@ class OpenAIAudioTranscriptionHandler(BaseTranslation):
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process input - not applicable for audio transcription.
@ -54,6 +55,7 @@ class OpenAIAudioTranscriptionHandler(BaseTranslation):
self,
response: "TranscriptionResponse",
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional[Any] = None,
) -> Any:
"""
Process output transcription by applying guardrails to transcribed text.

View file

@ -0,0 +1,12 @@
"""
Pass-Through Endpoint Guardrail Translation
This module exists here (under litellm/llms/) so it can be auto-discovered by
load_guardrail_translation_mappings() which scans for guardrail_translation
directories under litellm/llms/.
The main passthrough endpoint implementation is in:
litellm/proxy/pass_through_endpoints/
See guardrail_translation/README.md for more details.
"""

View file

@ -0,0 +1,41 @@
# Pass-Through Endpoint Guardrail Translation
## Why This Exists Here
This module is located under `litellm/llms/` (instead of with the main passthrough code) because:
1. **Auto-discovery**: The `load_guardrail_translation_mappings()` function in `litellm/llms/__init__.py` scans for `guardrail_translation/` directories under `litellm/llms/`
2. **Consistency**: All other guardrail translation handlers follow this pattern (e.g., `openai/chat/guardrail_translation/`, `anthropic/chat/guardrail_translation/`)
## Main Passthrough Implementation
The main passthrough endpoint implementation is in:
```
litellm/proxy/pass_through_endpoints/
├── pass_through_endpoints.py # Core passthrough routing logic
├── passthrough_guardrails.py # Guardrail collection and field targeting
├── jsonpath_extractor.py # JSONPath field extraction utility
└── ...
```
## What This Handler Does
The `PassThroughEndpointHandler` enables guardrails to run on passthrough endpoint requests by:
1. **Field Targeting**: Extracts specific fields from the request/response using JSONPath expressions configured in `request_fields` / `response_fields`
2. **Full Payload Fallback**: If no field targeting is configured, processes the entire payload
3. **Config Access**: Uses `get_passthrough_guardrails_config()` / `set_passthrough_guardrails_config()` helpers to access the passthrough guardrails configuration stored in request metadata
## Example Config
```yaml
passthrough_endpoints:
- path: "/v1/rerank"
target: "https://api.cohere.com/v1/rerank"
guardrails:
bedrock-pre-guard:
request_fields: ["query", "documents[*].text"]
response_fields: ["results[*].text"]
```

View file

@ -0,0 +1,15 @@
"""Pass-Through Endpoint guardrail translation handler."""
from litellm.llms.pass_through.guardrail_translation.handler import (
PassThroughEndpointHandler,
)
from litellm.types.utils import CallTypes
guardrail_translation_mappings = {
CallTypes.pass_through: PassThroughEndpointHandler,
}
__all__ = [
"guardrail_translation_mappings",
"PassThroughEndpointHandler",
]

View file

@ -0,0 +1,165 @@
"""
Pass-Through Endpoint Message Handler for Unified Guardrails
This module provides a handler for passthrough endpoint requests.
It uses the field targeting configuration from litellm_logging_obj
to extract specific fields for guardrail processing.
"""
from typing import TYPE_CHECKING, Any, List, Optional
from litellm._logging import verbose_proxy_logger
from litellm.llms.base_llm.guardrail_translation.base_translation import BaseTranslation
from litellm.proxy._types import PassThroughGuardrailSettings
if TYPE_CHECKING:
from litellm.integrations.custom_guardrail import CustomGuardrail
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
class PassThroughEndpointHandler(BaseTranslation):
"""
Handler for processing passthrough endpoint requests with guardrails.
Uses passthrough_guardrails_config from litellm_logging_obj
to determine which fields to extract for guardrail processing.
"""
def _get_guardrail_settings(
self,
litellm_logging_obj: Optional["LiteLLMLoggingObj"],
guardrail_name: Optional[str],
) -> Optional[PassThroughGuardrailSettings]:
"""
Get the guardrail settings for a specific guardrail from logging_obj.
"""
from litellm.proxy.pass_through_endpoints.passthrough_guardrails import (
PassthroughGuardrailHandler,
)
if litellm_logging_obj is None:
return None
passthrough_config = getattr(
litellm_logging_obj, "passthrough_guardrails_config", None
)
if not passthrough_config or not guardrail_name:
return None
return PassthroughGuardrailHandler.get_settings(
passthrough_config, guardrail_name
)
def _extract_text_for_guardrail(
self,
data: dict,
field_expressions: Optional[List[str]],
) -> str:
"""
Extract text from data for guardrail processing.
If field_expressions provided, extracts only those fields.
Otherwise, returns the full payload as JSON.
"""
from litellm.proxy.pass_through_endpoints.jsonpath_extractor import (
JsonPathExtractor,
)
if field_expressions:
text = JsonPathExtractor.extract_fields(
data=data,
jsonpath_expressions=field_expressions,
)
verbose_proxy_logger.debug(
"PassThroughEndpointHandler: Extracted targeted fields: %s",
text[:200] if text else None,
)
return text
# Use entire payload, excluding internal fields
from litellm.litellm_core_utils.safe_json_dumps import safe_dumps
payload_to_check = {
k: v
for k, v in data.items()
if not k.startswith("_") and k not in ("metadata", "litellm_logging_obj")
}
verbose_proxy_logger.debug(
"PassThroughEndpointHandler: Using full payload for guardrail"
)
return safe_dumps(payload_to_check)
async def process_input_messages(
self,
data: dict,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional["LiteLLMLoggingObj"] = None,
) -> Any:
"""
Process input by applying guardrails to targeted fields or full payload.
"""
guardrail_name = guardrail_to_apply.guardrail_name
verbose_proxy_logger.debug(
"PassThroughEndpointHandler: Processing input for guardrail=%s",
guardrail_name,
)
# Get field targeting settings for this guardrail
settings = self._get_guardrail_settings(litellm_logging_obj, guardrail_name)
field_expressions = settings.request_fields if settings else None
# Extract text to check
text_to_check = self._extract_text_for_guardrail(data, field_expressions)
if not text_to_check:
verbose_proxy_logger.debug(
"PassThroughEndpointHandler: No text to check, skipping guardrail"
)
return data
# Apply guardrail
await guardrail_to_apply.apply_guardrail(
text=text_to_check,
request_data=data,
)
return data
async def process_output_response(
self,
response: Any,
guardrail_to_apply: "CustomGuardrail",
litellm_logging_obj: Optional["LiteLLMLoggingObj"] = None,
) -> Any:
"""
Process output response by applying guardrails to targeted fields.
"""
if not isinstance(response, dict):
verbose_proxy_logger.debug(
"PassThroughEndpointHandler: Response is not a dict, skipping"
)
return response
guardrail_name = guardrail_to_apply.guardrail_name
verbose_proxy_logger.debug(
"PassThroughEndpointHandler: Processing output for guardrail=%s",
guardrail_name,
)
# Get field targeting settings for this guardrail
settings = self._get_guardrail_settings(litellm_logging_obj, guardrail_name)
field_expressions = settings.response_fields if settings else None
# Extract text to check
text_to_check = self._extract_text_for_guardrail(response, field_expressions)
if not text_to_check:
return response
# Apply guardrail
await guardrail_to_apply.apply_guardrail(
text=text_to_check,
request_data=response,
)
return response

View file

@ -0,0 +1,114 @@
"""
Translates from OpenAI's `/v1/chat/completions` to PublicAI's `/v1/chat/completions`
"""
from typing import Any, Coroutine, List, Literal, Optional, Tuple, Union, overload
from litellm.litellm_core_utils.prompt_templates.common_utils import (
handle_messages_with_content_list_to_str_conversion,
)
from litellm.secret_managers.main import get_secret_str
from litellm.types.llms.openai import AllMessageValues
from ...openai.chat.gpt_transformation import OpenAIGPTConfig
class PublicAIChatConfig(OpenAIGPTConfig):
@overload
def _transform_messages(
self, messages: List[AllMessageValues], model: str, is_async: Literal[True]
) -> Coroutine[Any, Any, List[AllMessageValues]]:
...
@overload
def _transform_messages(
self,
messages: List[AllMessageValues],
model: str,
is_async: Literal[False] = False,
) -> List[AllMessageValues]:
...
def _transform_messages(
self, messages: List[AllMessageValues], model: str, is_async: bool = False
) -> Union[List[AllMessageValues], Coroutine[Any, Any, List[AllMessageValues]]]:
"""
PublicAI does not support content in list format.
"""
messages = handle_messages_with_content_list_to_str_conversion(messages)
if is_async:
return super()._transform_messages(
messages=messages, model=model, is_async=True
)
else:
return super()._transform_messages(
messages=messages, model=model, is_async=False
)
def _get_openai_compatible_provider_info(
self, api_base: Optional[str], api_key: Optional[str]
) -> Tuple[Optional[str], Optional[str]]:
api_base = (
api_base
or get_secret_str("PUBLICAI_API_BASE")
or "https://platform.publicai.co/v1"
) # type: ignore
dynamic_api_key = api_key or get_secret_str("PUBLICAI_API_KEY")
return api_base, dynamic_api_key
def get_complete_url(
self,
api_base: Optional[str],
api_key: Optional[str],
model: str,
optional_params: dict,
litellm_params: dict,
stream: Optional[bool] = None,
) -> str:
"""
If api_base is not provided, use the default PublicAI /chat/completions endpoint.
"""
if not api_base:
api_base = "https://platform.publicai.co/v1"
if not api_base.endswith("/chat/completions"):
api_base = f"{api_base}/chat/completions"
return api_base
def get_supported_openai_params(self, model: str) -> list:
"""
Get the supported OpenAI params for PublicAI models
PublicAI limitations:
- functions parameter is not supported (use tools instead)
"""
excluded_params: List[str] = ["functions"]
base_openai_params = super().get_supported_openai_params(model=model)
final_params: List[str] = []
for param in base_openai_params:
if param not in excluded_params:
final_params.append(param)
return final_params
def map_openai_params(
self,
non_default_params: dict,
optional_params: dict,
model: str,
drop_params: bool,
) -> dict:
"""
Map OpenAI parameters to PublicAI parameters
"""
supported_openai_params = self.get_supported_openai_params(model)
for param, value in non_default_params.items():
if param == "max_completion_tokens":
optional_params["max_tokens"] = value
elif param in supported_openai_params:
optional_params[param] = value
return optional_params

View file

@ -121,5 +121,10 @@ class SambanovaConfig(OpenAIGPTConfig):
SambaNova API doesn't support content as a list - only string content.
This converts content lists like [{"type": "text", "text": "..."}] to strings.
"""
async def _async_transform():
return handle_messages_with_content_list_to_str_conversion(messages)
if is_async:
return _async_transform()
messages = handle_messages_with_content_list_to_str_conversion(messages)
return messages

View file

@ -58,9 +58,11 @@ from litellm import ( # type: ignore
get_litellm_params,
get_optional_params,
)
# Logging is imported lazily when needed to avoid loading litellm_logging at import time
if TYPE_CHECKING:
from litellm.litellm_core_utils.litellm_logging import Logging
from litellm.constants import (
DEFAULT_MOCK_RESPONSE_COMPLETION_TOKEN_COUNT,
DEFAULT_MOCK_RESPONSE_PROMPT_TOKEN_COUNT,
@ -107,7 +109,6 @@ from litellm.utils import (
ProviderConfigManager,
Usage,
_get_model_info_helper,
get_requester_metadata,
add_provider_specific_params_to_optional_params,
async_mock_completion_streaming_obj,
convert_to_model_response_object,
@ -120,6 +121,7 @@ from litellm.utils import (
get_optional_params_embeddings,
get_optional_params_image_gen,
get_optional_params_transcription,
get_requester_metadata,
get_secret,
get_standard_openai_params,
mock_completion_streaming_obj,
@ -154,11 +156,11 @@ from .litellm_core_utils.prompt_templates.factory import (
)
from .litellm_core_utils.streaming_chunk_builder_utils import ChunkProcessor
from .llms.anthropic.chat import AnthropicChatCompletion
from .llms.azure.anthropic.handler import AzureAnthropicChatCompletion
from .llms.azure.audio_transcriptions import AzureAudioTranscription
from .llms.azure.azure import AzureChatCompletion, _check_dynamic_azure_params
from .llms.azure.chat.o_series_handler import AzureOpenAIO1ChatCompletion
from .llms.azure.completion.handler import AzureTextCompletion
from .llms.azure_ai.anthropic.handler import AzureAnthropicChatCompletion
from .llms.azure_ai.embed import AzureAIEmbedding
from .llms.bedrock.chat import BedrockConverseLLM, BedrockLLM
from .llms.bedrock.embed.embedding import BedrockEmbedding
@ -1686,57 +1688,109 @@ def completion( # type: ignore # noqa: PLR0915
elif custom_llm_provider == "azure_ai":
from litellm.llms.azure_ai.common_utils import AzureFoundryModelInfo
api_base = AzureFoundryModelInfo.get_api_base(api_base)
# set API KEY
api_key = AzureFoundryModelInfo.get_api_key(api_key)
headers = headers or litellm.headers
if extra_headers is not None:
optional_params["extra_headers"] = extra_headers
## FOR COHERE
if "command-r" in model: # make sure tool call in messages are str
messages = stringify_json_tool_call_content(messages=messages)
## COMPLETION CALL
try:
response = base_llm_http_handler.completion(
# Check if this is a Claude model - route to Azure Anthropic handler
model_lower = model.lower()
if "claude" in model_lower:
# Use Azure Anthropic handler for Claude models
api_base = AzureFoundryModelInfo.get_api_base(api_base)
if api_base is None:
raise ValueError(
"Azure Anthropic requests require an api_base. "
"Set `api_base` or the AZURE_AI_API_BASE env var."
)
api_key = AzureFoundryModelInfo.get_api_key(api_key)
# Ensure the URL ends with /v1/messages for Anthropic
if api_base:
api_base = api_base.rstrip("/")
if not api_base.endswith("/v1/messages"):
if "/anthropic" in api_base:
parts = api_base.split("/anthropic", 1)
api_base = parts[0] + "/anthropic"
else:
api_base = api_base + "/anthropic"
api_base = api_base + "/v1/messages"
response = azure_anthropic_chat_completions.completion(
model=model,
messages=messages,
headers=headers,
model_response=model_response,
api_key=api_key,
api_base=api_base,
acompletion=acompletion,
logging_obj=logging,
custom_prompt_dict=litellm.custom_prompt_dict,
model_response=model_response,
print_verbose=print_verbose,
optional_params=optional_params,
litellm_params=litellm_params,
shared_session=shared_session,
timeout=timeout, # type: ignore
client=client, # pass AsyncOpenAI, OpenAI client
custom_llm_provider=custom_llm_provider,
logger_fn=logger_fn,
encoding=encoding,
stream=stream,
)
except Exception as e:
## LOGGING - log the original exception returned
logging.post_call(
input=messages,
api_key=api_key,
original_response=str(e),
additional_args={"headers": headers},
logging_obj=logging,
headers=headers,
timeout=timeout,
client=client,
custom_llm_provider=custom_llm_provider,
)
raise e
if optional_params.get("stream", False) or acompletion is True:
## LOGGING
logging.post_call(
input=messages,
api_key=api_key,
original_response=response,
)
response = response
else:
# Non-Claude models use standard Azure AI flow
api_base = AzureFoundryModelInfo.get_api_base(api_base)
# set API KEY
api_key = AzureFoundryModelInfo.get_api_key(api_key)
if optional_params.get("stream", False):
## LOGGING
logging.post_call(
input=messages,
api_key=api_key,
original_response=response,
additional_args={"headers": headers},
)
headers = headers or litellm.headers
if extra_headers is not None:
optional_params["extra_headers"] = extra_headers
## FOR COHERE
if "command-r" in model: # make sure tool call in messages are str
messages = stringify_json_tool_call_content(messages=messages)
## COMPLETION CALL
try:
response = base_llm_http_handler.completion(
model=model,
messages=messages,
headers=headers,
model_response=model_response,
api_key=api_key,
api_base=api_base,
acompletion=acompletion,
logging_obj=logging,
optional_params=optional_params,
litellm_params=litellm_params,
shared_session=shared_session,
timeout=timeout, # type: ignore
client=client, # pass AsyncOpenAI, OpenAI client
custom_llm_provider=custom_llm_provider,
encoding=encoding,
stream=stream,
)
except Exception as e:
## LOGGING - log the original exception returned
logging.post_call(
input=messages,
api_key=api_key,
original_response=str(e),
additional_args={"headers": headers},
)
raise e
if optional_params.get("stream", False):
## LOGGING
logging.post_call(
input=messages,
api_key=api_key,
original_response=response,
additional_args={"headers": headers},
)
elif (
custom_llm_provider == "text-completion-openai"
or "ft:babbage-002" in model
@ -2359,70 +2413,6 @@ def completion( # type: ignore # noqa: PLR0915
original_response=response,
)
response = response
elif custom_llm_provider == "azure_anthropic":
# Azure Anthropic uses same API as Anthropic but with Azure authentication
api_key = (
api_key
or litellm.azure_key
or litellm.api_key
or get_secret("AZURE_API_KEY")
or get_secret("AZURE_OPENAI_API_KEY")
)
custom_prompt_dict = custom_prompt_dict or litellm.custom_prompt_dict
# Azure Foundry endpoint format: https://<resource-name>.services.ai.azure.com/anthropic/v1/messages
api_base = (
api_base
or litellm.api_base
or get_secret("AZURE_API_BASE")
)
if api_base is None:
raise ValueError(
"Missing Azure API Base - Please set `api_base` or `AZURE_API_BASE` environment variable. "
"Expected format: https://<resource-name>.services.ai.azure.com/anthropic"
)
# Ensure the URL ends with /v1/messages
api_base = api_base.rstrip("/")
if api_base.endswith("/v1/messages"):
pass
elif api_base.endswith("/anthropic/v1/messages"):
pass
else:
if "/anthropic" in api_base:
parts = api_base.split("/anthropic", 1)
api_base = parts[0] + "/anthropic"
else:
api_base = api_base + "/anthropic"
api_base = api_base + "/v1/messages"
response = azure_anthropic_chat_completions.completion(
model=model,
messages=messages,
api_base=api_base,
acompletion=acompletion,
custom_prompt_dict=litellm.custom_prompt_dict,
model_response=model_response,
print_verbose=print_verbose,
optional_params=optional_params,
litellm_params=litellm_params,
logger_fn=logger_fn,
encoding=encoding, # for calculating input/output tokens
api_key=api_key,
logging_obj=logging,
headers=headers,
timeout=timeout,
client=client,
custom_llm_provider=custom_llm_provider,
)
if optional_params.get("stream", False) or acompletion is True:
## LOGGING
logging.post_call(
input=messages,
api_key=api_key,
original_response=response,
)
response = response
elif custom_llm_provider == "nlp_cloud":
nlp_cloud_key = (
api_key
@ -6236,9 +6226,9 @@ async def ahealth_check(
"x-ms-region": str,
}
"""
from litellm.litellm_core_utils.health_check_helpers import HealthCheckHelpers
from litellm.litellm_core_utils.cached_imports import get_litellm_logging_class
from litellm.litellm_core_utils.health_check_helpers import HealthCheckHelpers
# Use cached import helper to lazy-load Logging class (only loads when function is called)
Logging = get_litellm_logging_class()

View file

@ -1151,7 +1151,7 @@
},
"azure/claude-haiku-4-5": {
"input_cost_per_token": 1e-06,
"litellm_provider": "azure_anthropic",
"litellm_provider": "azure_ai",
"max_input_tokens": 200000,
"max_output_tokens": 64000,
"max_tokens": 64000,
@ -1169,7 +1169,7 @@
},
"azure/claude-opus-4-1": {
"input_cost_per_token": 1.5e-05,
"litellm_provider": "azure_anthropic",
"litellm_provider": "azure_ai",
"max_input_tokens": 200000,
"max_output_tokens": 32000,
"max_tokens": 32000,
@ -1187,7 +1187,7 @@
},
"azure/claude-sonnet-4-5": {
"input_cost_per_token": 3e-06,
"litellm_provider": "azure_anthropic",
"litellm_provider": "azure_ai",
"max_input_tokens": 200000,
"max_output_tokens": 64000,
"max_tokens": 64000,
@ -21570,6 +21570,116 @@
"mode": "chat",
"output_cost_per_token": 2.8e-07
},
"publicai/swiss-ai/apertus-8b-instruct": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 8192,
"max_output_tokens": 4096,
"max_tokens": 8192,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true
},
"publicai/swiss-ai/apertus-70b-instruct": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 8192,
"max_output_tokens": 4096,
"max_tokens": 8192,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true
},
"publicai/aisingapore/Gemma-SEA-LION-v4-27B-IT": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 8192,
"max_output_tokens": 4096,
"max_tokens": 8192,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true
},
"publicai/BSC-LT/salamandra-7b-instruct-tools-16k": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 16384,
"max_output_tokens": 4096,
"max_tokens": 16384,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true
},
"publicai/BSC-LT/ALIA-40b-instruct_Q8_0": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 8192,
"max_output_tokens": 4096,
"max_tokens": 8192,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true
},
"publicai/allenai/Olmo-3-7B-Instruct": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 32768,
"max_output_tokens": 4096,
"max_tokens": 32768,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true
},
"publicai/aisingapore/Qwen-SEA-LION-v4-32B-IT": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 32768,
"max_output_tokens": 4096,
"max_tokens": 32768,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true
},
"publicai/allenai/Olmo-3-7B-Think": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 32768,
"max_output_tokens": 4096,
"max_tokens": 32768,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true,
"supports_reasoning": true
},
"publicai/allenai/Olmo-3-32B-Think": {
"input_cost_per_token": 0.0,
"litellm_provider": "publicai",
"max_input_tokens": 32768,
"max_output_tokens": 4096,
"max_tokens": 32768,
"mode": "chat",
"output_cost_per_token": 0.0,
"source": "https://platform.publicai.co/docs",
"supports_function_calling": true,
"supports_tool_choice": true,
"supports_reasoning": true
},
"qwen.qwen3-coder-480b-a35b-v1:0": {
"input_cost_per_token": 2.2e-07,
"litellm_provider": "bedrock_converse",

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View file

@ -1 +0,0 @@
"use strict";(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[3665],{84566:function(e,t,s){s.d(t,{GH$:function(){return l}});var c=s(2265);let l=({color:e="currentColor",size:t=24,className:s,...l})=>c.createElement("svg",{viewBox:"0 0 24 24",xmlns:"http://www.w3.org/2000/svg",width:t,height:t,fill:e,...l,className:"remixicon "+(s||"")},c.createElement("path",{d:"M12 22C6.47715 22 2 17.5228 2 12C2 6.47715 6.47715 2 12 2C17.5228 2 22 6.47715 22 12C22 17.5228 17.5228 22 12 22ZM12 20C16.4183 20 20 16.4183 20 12C20 7.58172 16.4183 4 12 4C7.58172 4 4 7.58172 4 12C4 16.4183 7.58172 20 12 20ZM11.0026 16L6.75999 11.7574L8.17421 10.3431L11.0026 13.1716L16.6595 7.51472L18.0737 8.92893L11.0026 16Z"}))}}]);

View file

@ -0,0 +1 @@
"use strict";(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[3665],{84566:function(e,t,s){s.d(t,{GH$:function(){return l}});var c=s(2265);let l=({color:e="currentColor",size:t=24,className:s,...l})=>c.createElement("svg",{viewBox:"0 0 24 24",xmlns:"http://www.w3.org/2000/svg",width:t,height:t,fill:e,...l,className:"remixicon "+(s||"")},c.createElement("path",{d:"M4 12C4 7.58172 7.58172 4 12 4C16.4183 4 20 7.58172 20 12C20 16.4183 16.4183 20 12 20C7.58172 20 4 16.4183 4 12ZM12 2C6.47715 2 2 6.47715 2 12C2 17.5228 6.47715 22 12 22C17.5228 22 22 17.5228 22 12C22 6.47715 17.5228 2 12 2ZM17.4571 9.45711L16.0429 8.04289L11 13.0858L8.20711 10.2929L6.79289 11.7071L11 15.9142L17.4571 9.45711Z"}))}}]);

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

Some files were not shown because too many files have changed in this diff Show more