Merge branch 'BerriAI:main' into main

This commit is contained in:
AnilAren 2025-06-03 14:35:47 +05:30 • committed by GitHub
commit 4bcedc6108
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
272 changed files with 9006 additions and 1475 deletions

View file

@ -79,7 +79,7 @@ jobs:
pip install "pytest-retry==1.6.3"
pip install "pytest-asyncio==0.21.1"
pip install "pytest-cov==5.0.0"
pip install mypy
pip install "mypy==1.15.0"
pip install "google-generativeai==0.3.2"
pip install "google-cloud-aiplatform==1.43.0"
pip install pyarrow
@ -1158,6 +1158,7 @@ jobs:
pip install "google-cloud-aiplatform==1.43.0"
pip install "mlflow==2.17.2"
pip install "anthropic==0.52.0"
pip install "blockbuster==1.5.24"
# Run pytest and generate JUnit XML report
- setup_litellm_enterprise_pip
- run:

View file

@ -7,7 +7,7 @@ on:
jobs:
test:
runs-on: ubuntu-latest
timeout-minutes: 8
timeout-minutes: 15
steps:
- uses: actions/checkout@v4

View file

@ -78,8 +78,9 @@ curl http://localhost:4000/v1/batches \
**Create File for Batch Completion**
```python
from litellm
import litellm
import os
import asyncio
os.environ["OPENAI_API_KEY"] = "sk-.."
@ -97,8 +98,9 @@ print("Response from creating file=", file_obj)
**Create Batch Request**
```python
from litellm
import litellm
import os
import asyncio
create_batch_response = await litellm.acreate_batch(
completion_window="24h",

View file

@ -33,11 +33,11 @@ cd litellm/ui/litellm-dashboard
npm run dev
# starts on http://0.0.0.0:3000/ui
# starts on http://0.0.0.0:3000
```
## 3. Go to local UI
```
http://0.0.0.0:3000/ui
```bash
http://0.0.0.0:3000
```

View file

@ -13,7 +13,7 @@ Here are the core requirements for any PR submitted to LiteLLM
## **Contributor License Agreement (CLA)**
Before contributing code to LiteLLM, you must sign our [Contributor License Agreement (CLA)](<(https://cla-assistant.io/BerriAI/litellm)>). This is a legal requirement for all contributions to be merged into the main repository. The CLA helps protect both you and the project by clearly defining the terms under which your contributions are made.
Before contributing code to LiteLLM, you must sign our [Contributor License Agreement (CLA)](https://cla-assistant.io/BerriAI/litellm). This is a legal requirement for all contributions to be merged into the main repository. The CLA helps protect both you and the project by clearly defining the terms under which your contributions are made.
**Important:** We strongly recommend reviewing and signing the CLA before starting work on your contribution to avoid any delays in the PR process. You can find the CLA [here](https://cla-assistant.io/BerriAI/litellm) and sign it through our CLA management system when you submit your first PR.

View file

@ -14,7 +14,8 @@ LiteLLM provides image editing functionality that maps to OpenAI's `/images/edit
| Fallbacks | ✅ | Works between supported models |
| Loadbalancing | ✅ | Works between supported models |
| Supported operations | Create image edits | |
| Supported LiteLLM Versions | 1.63.8+ | |
| Supported LiteLLM SDK Versions | 1.63.8+ | |
| Supported LiteLLM Proxy Versions | 1.71.1+ | |
| Supported LLM providers | **OpenAI** | Currently only `openai` is supported |
## Usage

View file

@ -0,0 +1,202 @@
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
# Bedrock Agents
Call Bedrock Agents in the OpenAI Request/Response format.
| Property | Details |
|----------|---------|
| Description | Amazon Bedrock Agents use the reasoning of foundation models (FMs), APIs, and data to break down user requests, gather relevant information, and efficiently complete tasks. |
| Provider Route on LiteLLM | `bedrock/agent/{AGENT_ID}/{ALIAS_ID}` |
| Provider Doc | [AWS Bedrock Agents ↗](https://aws.amazon.com/bedrock/agents/) |
## Quick Start
### Model Format to LiteLLM
To call a bedrock agent through LiteLLM, you need to use the following model format to call the agent.
Here the `model=bedrock/agent/` tells LiteLLM to call the bedrock `InvokeAgent` API.
```shell showLineNumbers title="Model Format to LiteLLM"
bedrock/agent/{AGENT_ID}/{ALIAS_ID}
```
**Example:**
- `bedrock/agent/L1RT58GYRW/MFPSBCXYTW`
- `bedrock/agent/ABCD1234/LIVE`
You can find these IDs in your AWS Bedrock console under Agents.
### LiteLLM Python SDK
```python showLineNumbers title="Basic Agent Completion"
import litellm
# Make a completion request to your Bedrock Agent
response = litellm.completion(
model="bedrock/agent/L1RT58GYRW/MFPSBCXYTW", # agent/{AGENT_ID}/{ALIAS_ID}
messages=[
{
"role": "user",
"content": "Hi, I need help with analyzing our Q3 sales data and generating a summary report"
}
],
)
print(response.choices[0].message.content)
print(f"Response cost: ${response._hidden_params['response_cost']}")
```
```python showLineNumbers title="Streaming Agent Responses"
import litellm
# Stream responses from your Bedrock Agent
response = litellm.completion(
model="bedrock/agent/L1RT58GYRW/MFPSBCXYTW",
messages=[
{
"role": "user",
"content": "Can you help me plan a marketing campaign and provide step-by-step execution details?"
}
],
stream=True,
)
for chunk in response:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
```
### LiteLLM Proxy
#### 1. Configure your model in config.yaml
<Tabs>
<TabItem value="config-yaml" label="config.yaml">
```yaml showLineNumbers title="LiteLLM Proxy Configuration"
model_list:
- model_name: bedrock-agent-1
litellm_params:
model: bedrock/agent/L1RT58GYRW/MFPSBCXYTW
aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
aws_region_name: us-west-2
- model_name: bedrock-agent-2
litellm_params:
model: bedrock/agent/AGENT456/ALIAS789
aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
aws_region_name: us-east-1
```
</TabItem>
</Tabs>
#### 2. Start the LiteLLM Proxy
```bash showLineNumbers title="Start LiteLLM Proxy"
litellm --config config.yaml
```
#### 3. Make requests to your Bedrock Agents
<Tabs>
<TabItem value="curl" label="Curl">
```bash showLineNumbers title="Basic Agent Request"
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "bedrock-agent-1",
"messages": [
{
"role": "user",
"content": "Analyze our customer data and suggest retention strategies"
}
]
}'
```
```bash showLineNumbers title="Streaming Agent Request"
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "bedrock-agent-2",
"messages": [
{
"role": "user",
"content": "Create a comprehensive social media strategy for our new product"
}
],
"stream": true
}'
```
</TabItem>
<TabItem value="openai-sdk" label="OpenAI Python SDK">
```python showLineNumbers title="Using OpenAI SDK with LiteLLM Proxy"
from openai import OpenAI
# Initialize client with your LiteLLM proxy URL
client = OpenAI(
base_url="http://localhost:4000",
api_key="your-litellm-api-key"
)
# Make a completion request to your agent
response = client.chat.completions.create(
model="bedrock-agent-1",
messages=[
{
"role": "user",
"content": "Help me prepare for the quarterly business review meeting"
}
]
)
print(response.choices[0].message.content)
```
```python showLineNumbers title="Streaming with OpenAI SDK"
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:4000",
api_key="your-litellm-api-key"
)
# Stream agent responses
stream = client.chat.completions.create(
model="bedrock-agent-2",
messages=[
{
"role": "user",
"content": "Walk me through launching a new feature beta program"
}
],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content is not None:
print(chunk.choices[0].delta.content, end="")
```
</TabItem>
</Tabs>
## Further Reading
- [AWS Bedrock Agents Documentation](https://aws.amazon.com/bedrock/agents/)
- [LiteLLM Authentication to Bedrock](https://docs.litellm.ai/docs/providers/bedrock#boto3---authentication)

View file

@ -371,6 +371,7 @@ router_settings:
| DD_API_KEY | API key for Datadog integration
| DD_SITE | Site URL for Datadog (e.g., datadoghq.com)
| DD_SOURCE | Source identifier for Datadog logs
| DD_TRACER_STREAMING_CHUNK_YIELD_RESOURCE | Resource name for Datadog tracing of streaming chunk yields. Default is "streaming.chunk.yield"
| DD_ENV | Environment identifier for Datadog logs. Only supported for `datadog_llm_observability` callback
| DD_SERVICE | Service identifier for Datadog logs. Defaults to "litellm-server"
| DD_VERSION | Version identifier for Datadog logs. Defaults to "unknown"
@ -406,11 +407,14 @@ router_settings:
| DEFAULT_REPLICATE_GPU_PRICE_PER_SECOND | Default price per second for Replicate GPU. Default is 0.001400
| DEFAULT_REPLICATE_POLLING_DELAY_SECONDS | Default delay in seconds for Replicate polling. Default is 1
| DEFAULT_REPLICATE_POLLING_RETRIES | Default number of retries for Replicate polling. Default is 5
| DEFAULT_S3_BATCH_SIZE | Default batch size for S3 logging. Default is 512
| DEFAULT_S3_FLUSH_INTERVAL_SECONDS | Default flush interval for S3 logging. Default is 10
| DEFAULT_SLACK_ALERTING_THRESHOLD | Default threshold for Slack alerting. Default is 300
| DEFAULT_SOFT_BUDGET | Default soft budget for LiteLLM proxy keys. Default is 50.0
| DEFAULT_TRIM_RATIO | Default ratio of tokens to trim from prompt end. Default is 0.75
| DIRECT_URL | Direct URL for service endpoint
| DISABLE_ADMIN_UI | Toggle to disable the admin UI
| DISABLE_AIOHTTP_TRANSPORT | Flag to disable aiohttp transport. When this is set to True, litellm will use httpx instead of aiohttp. **Default is False**
| DISABLE_SCHEMA_UPDATE | Toggle to disable schema updates
| DOCS_DESCRIPTION | Description text for documentation pages
| DOCS_FILTERED | Flag indicating filtered documentation
@ -642,7 +646,6 @@ router_settings:
| UPSTREAM_LANGFUSE_PUBLIC_KEY | Public key for upstream Langfuse authentication
| UPSTREAM_LANGFUSE_RELEASE | Release version identifier for upstream Langfuse
| UPSTREAM_LANGFUSE_SECRET_KEY | Secret key for upstream Langfuse authentication
| USE_AIOHTTP_TRANSPORT | Flag to enable aiohttp transport. This is a feature flag for the new aiohttp transport. **Default is False**
| USE_AWS_KMS | Flag to enable AWS Key Management Service for encryption
| USE_PRISMA_MIGRATE | Flag to use prisma migrate instead of prisma db push. Recommended for production environments.
| WEBHOOK_URL | URL for receiving webhooks from external services

View file

@ -0,0 +1,59 @@
# UI - Custom Root Path
💥 Use this when you want to serve LiteLLM on a custom base url path like `https://localhost:4000/api/v1`
## Usage
### 1. Set `SERVER_ROOT_PATH` in your .env
👉 Set `SERVER_ROOT_PATH` in your .env and this will be set as your server root path
```
export SERVER_ROOT_PATH="/api/v1"
```
### 2. Run the Proxy
```shell
litellm proxy --config /path/to/config.yaml
```
After running the proxy you can access it on `http://0.0.0.0:4000/api/v1/` (since we set `SERVER_ROOT_PATH="/api/v1"`)
### 3. Reserve the `/litellm` path
LiteLLM uses the `/litellm` path to discover the custom root path. So you need to reserve this path in your proxy.
If you are running the UI, it will query the `/litellm/.well-known/litellm-ui-config` endpoint to get the UI configuration.
So you need to reserve the `/litellm` path in your proxy.
You can see the results with:
```bash
curl http://0.0.0.0:4000/litellm/.well-known/litellm-ui-config
```
Expected result:
```json
{
"server_root_path": "/api/v1",
...
}
```
### 4. Verify Running on correct path
<Image img={require('../../img/custom_root_path.png')} />
**That's it**, that's all you need to run the proxy on a custom root path
## Demo
[Here's a demo video](https://drive.google.com/file/d/1zqAxI0lmzNp7IJH1dxlLuKqX2xi3F_R3/view?usp=sharing) of running the proxy on a custom root path

View file

@ -619,101 +619,8 @@ docker pull ghcr.io/berriai/litellm-non_root:main-stable
### 1. Custom server root path (Proxy base url)
💥 Use this when you want to serve LiteLLM on a custom base url path like `https://localhost:4000/api/v1`
Refer to [Custom Root Path](./custom_root_ui) for more details.
:::info
In a Kubernetes deployment, it's possible to utilize a shared DNS to host multiple applications by modifying the virtual service
:::
Customize the root path to eliminate the need for employing multiple DNS configurations during deployment.
Step 1.
👉 Set `SERVER_ROOT_PATH` in your .env and this will be set as your server root path
```
export SERVER_ROOT_PATH="/api/v1"
```
**Step 2** (If you want the Proxy Admin UI to work with your root path you need to use this dockerfile)
- Use the dockerfile below (it uses litellm as a base image)
- 👉 Set `UI_BASE_PATH=$SERVER_ROOT_PATH/ui` in the Dockerfile, example `UI_BASE_PATH=/api/v1/ui`
Dockerfile
```shell
# Use the provided base image
FROM ghcr.io/berriai/litellm:main-latest
# Set the working directory to /app
WORKDIR /app
# Install Node.js and npm (adjust version as needed)
RUN apt-get update && apt-get install -y nodejs npm
# Copy the UI source into the container
COPY ./ui/litellm-dashboard /app/ui/litellm-dashboard
# Set an environment variable for UI_BASE_PATH
# This can be overridden at build time
# set UI_BASE_PATH to "<your server root path>/ui"
# 👇👇 Enter your UI_BASE_PATH here
ENV UI_BASE_PATH="/api/v1/ui"
# Build the UI with the specified UI_BASE_PATH
WORKDIR /app/ui/litellm-dashboard
RUN npm install
RUN UI_BASE_PATH=$UI_BASE_PATH npm run build
# Create the destination directory
RUN mkdir -p /app/litellm/proxy/_experimental/out
# Move the built files to the appropriate location
# Assuming the build output is in ./out directory
RUN rm -rf /app/litellm/proxy/_experimental/out/* && \
mv ./out/* /app/litellm/proxy/_experimental/out/
# Switch back to the main app directory
WORKDIR /app
# Make sure your entrypoint.sh is executable
RUN chmod +x ./docker/entrypoint.sh
# Expose the necessary port
EXPOSE 4000/tcp
# Override the CMD instruction with your desired command and arguments
# only use --detailed_debug for debugging
CMD ["--port", "4000", "--config", "config.yaml"]
```
**Step 3** build this Dockerfile
```shell
docker build -f Dockerfile -t litellm-prod-build . --progress=plain
```
**Step 4. Run Proxy with `SERVER_ROOT_PATH` set in your env **
```shell
docker run \
-v $(pwd)/proxy_config.yaml:/app/config.yaml \
-p 4000:4000 \
-e LITELLM_LOG="DEBUG"\
-e SERVER_ROOT_PATH="/api/v1"\
-e DATABASE_URL=postgresql://<user>:<password>@<host>:<port>/<dbname> \
-e LITELLM_MASTER_KEY="sk-1234"\
litellm-prod-build \
--config /app/config.yaml
```
After running the proxy you can access it on `http://0.0.0.0:4000/api/v1/` (since we set `SERVER_ROOT_PATH="/api/v1"`)
**Step 5. Verify Running on correct path**
<Image img={require('../../img/custom_root_path.png')} />
**That's it**, that's all you need to run the proxy on a custom root path
### 2. SSL Certification

View file

@ -45,12 +45,12 @@ Setup your config.yaml with your azure model.
```yaml
model_list:
- model_name: gpt-3.5-turbo
- model_name: gpt-4o
litellm_params:
model: azure/my_azure_deployment
api_base: os.environ/AZURE_API_BASE
api_key: "os.environ/AZURE_API_KEY"
api_version: "2024-07-01-preview" # [OPTIONAL] litellm uses the latest azure api_version by default
api_version: "2025-01-01-preview" # [OPTIONAL] litellm uses the latest azure api_version by default
```
---
@ -127,15 +127,15 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer sk-1234' \
-d '{
"model": "gpt-3.5-turbo",
"model": "gpt-4o",
"messages": [
{
"role": "system",
"content": "You are a helpful math tutor. Guide the user through the solution step by step."
"content": "You are an LLM named gpt-4o"
},
{
"role": "user",
"content": "how can I solve 8x + 7 = -23"
"content": "what is your name?"
}
]
}'
@ -145,28 +145,63 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
```bash
{
"id": "chatcmpl-2076f062-3095-4052-a520-7c321c115c68",
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "I am gpt-3.5-turbo",
"role": "assistant",
"tool_calls": null,
"function_call": null
}
}
],
"created": 1724962831,
"model": "gpt-3.5-turbo",
"object": "chat.completion",
"system_fingerprint": null,
"usage": {
"completion_tokens": 20,
"prompt_tokens": 10,
"total_tokens": 30
"id": "chatcmpl-BcO8tRQmQV6Dfw6onqMufxPkLLkA8",
"created": 1748488967,
"model": "gpt-4o-2024-11-20",
"object": "chat.completion",
"system_fingerprint": "fp_ee1d74bde0",
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "My name is **gpt-4o**! How can I assist you today?",
"role": "assistant",
"tool_calls": null,
"function_call": null,
"annotations": []
}
}
],
"usage": {
"completion_tokens": 19,
"prompt_tokens": 28,
"total_tokens": 47,
"completion_tokens_details": {
"accepted_prediction_tokens": 0,
"audio_tokens": 0,
"reasoning_tokens": 0,
"rejected_prediction_tokens": 0
},
"prompt_tokens_details": {
"audio_tokens": 0,
"cached_tokens": 0
}
},
"service_tier": null,
"prompt_filter_results": [
{
"prompt_index": 0,
"content_filter_results": {
"hate": {
"filtered": false,
"severity": "safe"
},
"self_harm": {
"filtered": false,
"severity": "safe"
},
"sexual": {
"filtered": false,
"severity": "safe"
},
"violence": {
"filtered": false,
"severity": "safe"
}
}
}
]
}
```
@ -191,12 +226,12 @@ Track Spend, and control model access via virtual keys for the proxy
```yaml
model_list:
- model_name: gpt-3.5-turbo
- model_name: gpt-4o
litellm_params:
model: azure/my_azure_deployment
api_base: os.environ/AZURE_API_BASE
api_key: "os.environ/AZURE_API_KEY"
api_version: "2024-07-01-preview" # [OPTIONAL] litellm uses the latest azure api_version by default
api_version: "2025-01-01-preview" # [OPTIONAL] litellm uses the latest azure api_version by default
general_settings:
master_key: sk-1234
@ -276,7 +311,7 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer sk-12...' \
-d '{
"model": "gpt-3.5-turbo",
"model": "gpt-4o",
"messages": [
{
"role": "system",
@ -312,7 +347,7 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer sk-12...' \
-d '{
"model": "gpt-3.5-turbo",
"model": "gpt-4o",
"messages": [
{
"role": "system",
@ -331,7 +366,7 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
```bash
{
"error": {
"message": "Max parallel request limit reached. Hit limit for api_key: daa1b272072a4c6841470a488c5dad0f298ff506e1cc935f4a181eed90c182ad. tpm_limit: 100, current_tpm: 29, rpm_limit: 1, current_rpm: 2.",
"message": "LiteLLM Rate Limit Handler for rate limit type = key. Crossed TPM / RPM / Max Parallel Request Limit. current rpm: 1, rpm limit: 1, current tpm: 348, tpm limit: 9223372036854775807, current max_parallel_requests: 0, max_parallel_requests: 9223372036854775807",
"type": "None",
"param": "None",
"code": "429"
@ -371,12 +406,12 @@ You can disable ssl verification with:
```yaml
model_list:
- model_name: gpt-3.5-turbo
- model_name: gpt-4o
litellm_params:
model: azure/my_azure_deployment
api_base: os.environ/AZURE_API_BASE
api_key: "os.environ/AZURE_API_KEY"
api_version: "2024-07-01-preview"
api_version: "2025-01-01-preview"
litellm_settings:
ssl_verify: false # 👈 KEY CHANGE

View file

@ -13,6 +13,7 @@ import TabItem from '@theme/TabItem';
| Supported Entity Types | All Presidio Entity Types |
| Supported Actions | `MASK`, `BLOCK` |
| Supported Modes | `pre_call`, `during_call`, `post_call`, `logging_only` |
| Language Support | Configurable via `presidio_language` parameter (supports multiple languages including English, Spanish, German, etc.) |
## Deployment options
@ -48,6 +49,18 @@ Now select the entity types you want to mask. See the [supported actions here](#
style={{width: '50%', display: 'block', margin: '0'}}
/>
#### 1.3 Set Default Language (Optional)
You can also configure a default language for PII analysis using the `presidio_language` field in the UI. This sets the default language that will be used for all requests unless overridden by a per-request language setting.
**Supported language codes include:**
- `en` - English (default)
- `es` - Spanish
- `de` - German
If not specified, English (`en`) will be used as the default language.
</TabItem>
@ -67,6 +80,7 @@ guardrails:
litellm_params:
guardrail: presidio # supported values: "aporia", "bedrock", "lakera", "presidio"
mode: "pre_call"
presidio_language: "en" # optional: set default language for PII analysis
```
Set the following env vars
@ -380,6 +394,86 @@ print(response)
</Tabs>
### Set default `language` in config.yaml
You can configure a default language for PII analysis in your YAML configuration using the `presidio_language` parameter. This language will be used for all requests unless overridden by a per-request language setting.
```yaml title="Default Language Configuration" showLineNumbers
model_list:
- model_name: gpt-3.5-turbo
litellm_params:
model: openai/gpt-3.5-turbo
api_key: os.environ/OPENAI_API_KEY
guardrails:
- guardrail_name: "presidio-german"
litellm_params:
guardrail: presidio
mode: "pre_call"
presidio_language: "de" # Default to German for PII analysis
pii_entities_config:
CREDIT_CARD: "MASK"
EMAIL_ADDRESS: "MASK"
PERSON: "MASK"
- guardrail_name: "presidio-spanish"
litellm_params:
guardrail: presidio
mode: "pre_call"
presidio_language: "es" # Default to Spanish for PII analysis
pii_entities_config:
CREDIT_CARD: "MASK"
PHONE_NUMBER: "MASK"
```
#### Supported Language Codes
Presidio supports multiple languages for PII detection. Common language codes include:
- `en` - English (default)
- `es` - Spanish
- `de` - German
For a complete list of supported languages, refer to the [Presidio documentation](https://microsoft.github.io/presidio/analyzer/languages/).
#### Language Precedence
The language setting follows this precedence order:
1. **Per-request language** (via `guardrail_config.language`) - highest priority
2. **YAML config language** (via `presidio_language`) - medium priority
3. **Default language** (`en`) - lowest priority
**Example with mixed languages:**
```yaml title="Mixed Language Configuration" showLineNumbers
guardrails:
- guardrail_name: "presidio-multilingual"
litellm_params:
guardrail: presidio
mode: "pre_call"
presidio_language: "de" # Default to German
pii_entities_config:
CREDIT_CARD: "MASK"
PERSON: "MASK"
```
```shell title="Override with per-request language" showLineNumbers
curl http://localhost:4000/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-1234" \
-d '{
"model": "gpt-3.5-turbo",
"messages": [
{"role": "user", "content": "Mi tarjeta de crédito es 4111-1111-1111-1111"}
],
"guardrails": ["presidio-multilingual"],
"guardrail_config": {"language": "es"}
}'
```
In this example, the request will use Spanish (`es`) for PII detection even though the guardrail is configured with German (`de`) as the default language.
### Output parsing

View file

@ -1260,7 +1260,7 @@ model_list:
litellm_params:
model: gpt-3.5-turbo
litellm_settings:
success_callback: ["s3"]
success_callback: ["s3_v2"]
s3_callback_params:
s3_bucket_name: logs-bucket-litellm # AWS Bucket Name for S3
s3_region_name: us-west-2 # AWS Region Name for S3
@ -1304,7 +1304,7 @@ You can add the team alias to the object key by setting the `team_alias` in the
```yaml
litellm_settings:
callbacks: ["s3"]
callbacks: ["s3_v2"]
enable_preview_features: true
s3_callback_params:
s3_bucket_name: logs-bucket-litellm

View file

@ -180,6 +180,19 @@ Use this for LLM API Error monitoring and tracking remaining rate limits and tok
| `litellm_llm_api_latency_metric` | Latency (seconds) for just the LLM API call - tracked for labels "model", "hashed_api_key", "api_key_alias", "team", "team_alias", "requested_model", "end_user", "user" |
| `litellm_llm_api_time_to_first_token_metric` | Time to first token for LLM API call - tracked for labels `model`, `hashed_api_key`, `api_key_alias`, `team`, `team_alias` [Note: only emitted for streaming requests] |
## Tracking `end_user` on Prometheus
By default LiteLLM does not track `end_user` on Prometheus. This is done to reduce the cardinality of the metrics from LiteLLM Proxy.
If you want to track `end_user` on Prometheus, you can do the following:
```yaml showLineNumbers title="config.yaml"
litellm_settings:
callbacks: ["prometheus"]
enable_end_user_cost_tracking_prometheus_only: true
```
## [BETA] Custom Metrics
Track custom metrics on prometheus on all events mentioned above.

View file

@ -0,0 +1,81 @@
# Using Anthropic File API with LiteLLM Proxy
## Overview
This tutorial shows how to create and analyze files with Claude-4 on Anthropic via LiteLLM Proxy.
## Prerequisites
- LiteLLM Proxy running
- Anthropic API key
Add the following to your `.env` file:
```
ANTHROPIC_API_KEY=sk-1234
```
## Usage
### 1. Setup config.yaml
```yaml
model_list:
- model_name: claude-opus
litellm_params:
model: anthropic/claude-opus-4-20250514
api_key: os.environ/ANTHROPIC_API_KEY
```
## 2. Create a file
Use the `/anthropic` passthrough endpoint to create a file.
```bash
curl -L -X POST 'http://0.0.0.0:4000/anthropic/v1/files' \
-H 'x-api-key: sk-1234' \
-H 'anthropic-version: 2023-06-01' \
-H 'anthropic-beta: files-api-2025-04-14' \
-F 'file=@"/path/to/your/file.csv"'
```
Expected response:
```json
{
"created_at": "2023-11-07T05:31:56Z",
"downloadable": false,
"filename": "file.csv",
"id": "file-1234",
"mime_type": "text/csv",
"size_bytes": 1,
"type": "file"
}
```
## 3. Analyze the file with Claude-4 via `/chat/completions`
```bash
curl -L -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer $LITELLM_API_KEY' \
-d '{
"model": "claude-opus",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "What is in this sheet?"},
{
"type": "file",
"file": {
"file_id": "file-1234",
"format": "text/csv" # 👈 IMPORTANT: This is the format of the file you want to analyze
}
}
]
}
]
}'
```

Binary file not shown.

Before

Width:  |  Height:  |  Size: 61 KiB

After

Width:  |  Height:  |  Size: 418 KiB

View file

@ -0,0 +1,243 @@
---
title: v1.72.0-stable
slug: v1.72.0-stable
date: 2025-05-31T10:00:00
authors:
- name: Krrish Dholakia
title: CEO, LiteLLM
url: https://www.linkedin.com/in/krish-d/
image_url: https://media.licdn.com/dms/image/v2/D4D03AQGrlsJ3aqpHmQ/profile-displayphoto-shrink_400_400/B4DZSAzgP7HYAg-/0/1737327772964?e=1749686400&v=beta&t=Hkl3U8Ps0VtvNxX0BNNq24b4dtX5wQaPFp6oiKCIHD8
- name: Ishaan Jaffer
title: CTO, LiteLLM
url: https://www.linkedin.com/in/reffajnaahsi/
image_url: https://pbs.twimg.com/profile_images/1613813310264340481/lz54oEiB_400x400.jpg
hide_table_of_contents: false
---
import Image from '@theme/IdealImage';
import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
:::info
The release candidate is live now.
The production release will be live on Wednesday.
:::
## Deploy this version
<Tabs>
<TabItem value="docker" label="Docker">
``` showLineNumbers title="docker run litellm"
docker run
-e STORE_MODEL_IN_DB=True
-p 4000:4000
ghcr.io/berriai/litellm:main-v1.72.0.rc
```
</TabItem>
<TabItem value="pip" label="Pip">
``` showLineNumbers title="pip install litellm"
pip install litellm==1.72.0
```
</TabItem>
</Tabs>
## Key Highlights
LiteLLM v1.72.0-stable.rc is live now. Here are the key highlights of this release:
- **Vector Store Permissions**: Control Vector Store access at the Key, Team, and Organization level.
- **Rate Limiting Sliding Window support**: Improved accuracy for Key/Team/User rate limits with request tracking across minutes.
- **Aiohttp Transport used by default**: Aiohttp transport is now the default transport for LiteLLM networking requests. This gives users 2x higher RPS per instance with a 40ms median latency overhead.
- **Bedrock Agents**: Call Bedrock Agents with `/chat/completions`, `/response` endpoints.
- **Anthropic File API**: Upload and analyze CSV files with Claude-4 on Anthropic via LiteLLM.
- **Prometheus**: End users (`end_user`) will no longer be tracked by default on Prometheus. Tracking end_users on prometheus is now opt-in. This is done to prevent the response from `/metrics` from becoming too large. [Read More](../../docs/proxy/prometheus#tracking-end_user-on-prometheus)
---
## Vector Store Permissions
This release brings support for managing permissions for vector stores by Keys, Teams, Organizations (entities) on LiteLLM. When a request attempts to query a vector store, LiteLLM will block it if the requesting entity lacks the proper permissions.
This is great for use cases that require access to restricted data that you don't want everyone to use.
Over the next week we plan on adding permission management for MCP Servers.
---
## Aiohttp Transport used by default
Aiohttp transport is now the default transport for LiteLLM networking requests. This gives users 2x higher RPS per instance with a 40ms median latency overhead. This has been live on LiteLLM Cloud for a week + gone through alpha users testing for a week.
If you encounter any issues, you can disable using the aiohttp transport in the following ways:
**On LiteLLM Proxy**
Set the `DISABLE_AIOHTTP_TRANSPORT=True` in the environment variables.
```yaml showLineNumbers title="Environment Variable"
export DISABLE_AIOHTTP_TRANSPORT="True"
```
**On LiteLLM Python SDK**
Set the `disable_aiohttp_transport=True` to disable aiohttp transport.
```python showLineNumbers title="Python SDK"
import litellm
litellm.disable_aiohttp_transport = True # default is False, enable this to disable aiohttp transport
result = litellm.completion(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "Hello, world!"}],
)
print(result)
```
---
## New Models / Updated Models
- **[Bedrock](../../docs/providers/bedrock)**
- Video support for Bedrock Converse - [PR](https://github.com/BerriAI/litellm/pull/11166)
- InvokeAgents support as /chat/completions route - [PR](https://github.com/BerriAI/litellm/pull/11239), [Get Started](../../docs/providers/bedrock_agents)
- AI21 Jamba models compatibility fixes - [PR](https://github.com/BerriAI/litellm/pull/11233)
- Fixed duplicate maxTokens parameter for Claude with thinking - [PR](https://github.com/BerriAI/litellm/pull/11181)
- **[Gemini (Google AI Studio + Vertex AI)](https://docs.litellm.ai/docs/providers/gemini)**
- Parallel tool calling support with `parallel_tool_calls` parameter - [PR](https://github.com/BerriAI/litellm/pull/11125)
- All Gemini models now support parallel function calling - [PR](https://github.com/BerriAI/litellm/pull/11225)
- **[VertexAI](../../docs/providers/vertex)**
- codeExecution tool support and anyOf handling - [PR](https://github.com/BerriAI/litellm/pull/11195)
- Vertex AI Anthropic support on /v1/messages - [PR](https://github.com/BerriAI/litellm/pull/11246)
- Thinking, global regions, and parallel tool calling improvements - [PR](https://github.com/BerriAI/litellm/pull/11194)
- Web Search Support [PR](https://github.com/BerriAI/litellm/commit/06484f6e5a7a2f4e45c490266782ed28b51b7db6)
- **[Anthropic](../../docs/providers/anthropic)**
- Thinking blocks on streaming support - [PR](https://github.com/BerriAI/litellm/pull/11194)
- Files API with form-data support on passthrough - [PR](https://github.com/BerriAI/litellm/pull/11256)
- File ID support on /chat/completion - [PR](https://github.com/BerriAI/litellm/pull/11256)
- **[xAI](../../docs/providers/xai)**
- Web Search Support [PR](https://github.com/BerriAI/litellm/commit/06484f6e5a7a2f4e45c490266782ed28b51b7db6)
- **[Google AI Studio](../../docs/providers/gemini)**
- Web Search Support [PR](https://github.com/BerriAI/litellm/commit/06484f6e5a7a2f4e45c490266782ed28b51b7db6)
- **[Mistral](../../docs/providers/mistral)**
- Updated mistral-medium prices and context sizes - [PR](https://github.com/BerriAI/litellm/pull/10729)
- **[Ollama](../../docs/providers/ollama)**
- Tool calls parsing on streaming - [PR](https://github.com/BerriAI/litellm/pull/11171)
- **[Cohere](../../docs/providers/cohere)**
- Swapped Cohere and Cohere Chat provider positioning - [PR](https://github.com/BerriAI/litellm/pull/11173)
- **[Nebius AI Studio](../../docs/providers/nebius)**
- New provider integration - [PR](https://github.com/BerriAI/litellm/pull/11143)
## LLM API Endpoints
- **[Image Edits API](../../docs/image_generation)**
- Azure support for /v1/images/edits - [PR](https://github.com/BerriAI/litellm/pull/11160)
- Cost tracking for image edits endpoint (OpenAI, Azure) - [PR](https://github.com/BerriAI/litellm/pull/11186)
- **[Completions API](../../docs/completion/chat)**
- Codestral latency overhead tracking on /v1/completions - [PR](https://github.com/BerriAI/litellm/pull/10879)
- **[Audio Transcriptions API](../../docs/audio/speech)**
- GPT-4o mini audio preview pricing without date - [PR](https://github.com/BerriAI/litellm/pull/11207)
- Non-default params support for audio transcription - [PR](https://github.com/BerriAI/litellm/pull/11212)
- **[Responses API](../../docs/response_api)**
- Session management fixes for using Non-OpenAI models - [PR](https://github.com/BerriAI/litellm/pull/11254)
## Management Endpoints / UI
- **Vector Stores**
- Permission management for LiteLLM Keys, Teams, and Organizations - [PR](https://github.com/BerriAI/litellm/pull/11213)
- UI display of vector store permissions - [PR](https://github.com/BerriAI/litellm/pull/11277)
- Vector store access controls enforcement - [PR](https://github.com/BerriAI/litellm/pull/11281)
- Object permissions fixes and QA improvements - [PR](https://github.com/BerriAI/litellm/pull/11291)
- **Teams**
- "All proxy models" display when no models selected - [PR](https://github.com/BerriAI/litellm/pull/11187)
- Removed redundant teamInfo call, using existing teamsList - [PR](https://github.com/BerriAI/litellm/pull/11051)
- Improved model tags display on Keys, Teams and Org pages - [PR](https://github.com/BerriAI/litellm/pull/11022)
- **SSO/SCIM**
- Bug fixes for showing SCIM token on UI - [PR](https://github.com/BerriAI/litellm/pull/11220)
- **General UI**
- Fix "UI Session Expired. Logging out" - [PR](https://github.com/BerriAI/litellm/pull/11279)
- Support for forwarding /sso/key/generate to server root path URL - [PR](https://github.com/BerriAI/litellm/pull/11165)
## Logging / Guardrails Integrations
#### Logging
- **[Prometheus](../../docs/proxy/prometheus)**
- End users will no longer be tracked by default on Prometheus. Tracking end_users on prometheus is now opt-in. [PR](https://github.com/BerriAI/litellm/pull/11192)
- **[Langfuse](../../docs/proxy/logging#langfuse)**
- Performance improvements: Fixed "Max langfuse clients reached" issue - [PR](https://github.com/BerriAI/litellm/pull/11285)
- **[Helicone](../../docs/observability/helicone_integration)**
- Base URL support - [PR](https://github.com/BerriAI/litellm/pull/11211)
- **[Sentry](../../docs/proxy/logging#sentry)**
- Added sentry sample rate configuration - [PR](https://github.com/BerriAI/litellm/pull/10283)
#### Guardrails
- **[Bedrock Guardrails](../../docs/proxy/guardrails/bedrock)**
- Streaming support for bedrock post guard - [PR](https://github.com/BerriAI/litellm/pull/11247)
- Auth parameter persistence fixes - [PR](https://github.com/BerriAI/litellm/pull/11270)
- **[Pangea Guardrails](../../docs/proxy/guardrails/pangea)**
- Added Pangea provider to Guardrails hook - [PR](https://github.com/BerriAI/litellm/pull/10775)
## Performance / Reliability Improvements
- **aiohttp Transport**
- Handling for aiohttp.ClientPayloadError - [PR](https://github.com/BerriAI/litellm/pull/11162)
- SSL verification settings support - [PR](https://github.com/BerriAI/litellm/pull/11162)
- Rollback to httpx==0.27.0 for stability - [PR](https://github.com/BerriAI/litellm/pull/11146)
- **Request Limiting**
- Sliding window logic for parallel request limiter v2 - [PR](https://github.com/BerriAI/litellm/pull/11283)
## Bug Fixes
- **LLM API Fixes**
- Added missing request_kwargs to get_available_deployment call - [PR](https://github.com/BerriAI/litellm/pull/11202)
- Fixed calling Azure O-series models - [PR](https://github.com/BerriAI/litellm/pull/11212)
- Support for dropping non-OpenAI params via additional_drop_params - [PR](https://github.com/BerriAI/litellm/pull/11246)
- Fixed frequency_penalty to repeat_penalty parameter mapping - [PR](https://github.com/BerriAI/litellm/pull/11284)
- Fix for embedding cache hits on string input - [PR](https://github.com/BerriAI/litellm/pull/11211)
- **General**
- OIDC provider improvements and audience bug fix - [PR](https://github.com/BerriAI/litellm/pull/10054)
- Removed AzureCredentialType restriction on AZURE_CREDENTIAL - [PR](https://github.com/BerriAI/litellm/pull/11272)
- Prevention of sensitive key leakage to Langfuse - [PR](https://github.com/BerriAI/litellm/pull/11165)
- Fixed healthcheck test using curl when curl not in image - [PR](https://github.com/BerriAI/litellm/pull/9737)
## New Contributors
* [@agajdosi](https://github.com/agajdosi) made their first contribution in [#9737](https://github.com/BerriAI/litellm/pull/9737)
* [@ketangangal](https://github.com/ketangangal) made their first contribution in [#11161](https://github.com/BerriAI/litellm/pull/11161)
* [@Aktsvigun](https://github.com/Aktsvigun) made their first contribution in [#11143](https://github.com/BerriAI/litellm/pull/11143)
* [@ryanmeans](https://github.com/ryanmeans) made their first contribution in [#10775](https://github.com/BerriAI/litellm/pull/10775)
* [@nikoizs](https://github.com/nikoizs) made their first contribution in [#10054](https://github.com/BerriAI/litellm/pull/10054)
* [@Nitro963](https://github.com/Nitro963) made their first contribution in [#11202](https://github.com/BerriAI/litellm/pull/11202)
* [@Jacobh2](https://github.com/Jacobh2) made their first contribution in [#11207](https://github.com/BerriAI/litellm/pull/11207)
* [@regismesquita](https://github.com/regismesquita) made their first contribution in [#10729](https://github.com/BerriAI/litellm/pull/10729)
* [@Vinnie-Singleton-NN](https://github.com/Vinnie-Singleton-NN) made their first contribution in [#10283](https://github.com/BerriAI/litellm/pull/10283)
* [@trashhalo](https://github.com/trashhalo) made their first contribution in [#11219](https://github.com/BerriAI/litellm/pull/11219)
* [@VigneshwarRajasekaran](https://github.com/VigneshwarRajasekaran) made their first contribution in [#11223](https://github.com/BerriAI/litellm/pull/11223)
* [@AnilAren](https://github.com/AnilAren) made their first contribution in [#11233](https://github.com/BerriAI/litellm/pull/11233)
* [@fadil4u](https://github.com/fadil4u) made their first contribution in [#11242](https://github.com/BerriAI/litellm/pull/11242)
* [@whitfin](https://github.com/whitfin) made their first contribution in [#11279](https://github.com/BerriAI/litellm/pull/11279)
* [@hcoona](https://github.com/hcoona) made their first contribution in [#11272](https://github.com/BerriAI/litellm/pull/11272)
* [@keyute](https://github.com/keyute) made their first contribution in [#11173](https://github.com/BerriAI/litellm/pull/11173)
* [@emmanuel-ferdman](https://github.com/emmanuel-ferdman) made their first contribution in [#11230](https://github.com/BerriAI/litellm/pull/11230)
## Demo Instance
Here's a Demo Instance to test changes:
- Instance: https://demo.litellm.ai/
- Login Credentials:
- Username: admin
- Password: sk-1234
## [Git Diff](https://github.com/BerriAI/litellm/releases)

View file

@ -102,6 +102,7 @@ const sidebars = {
items: [
"proxy/ui",
"proxy/admin_ui_sso",
"proxy/custom_root_ui",
"proxy/self_serve",
"proxy/public_teams",
"tutorials/scim_litellm",
@ -328,6 +329,7 @@ const sidebars = {
label: "Bedrock",
items: [
"providers/bedrock",
"providers/bedrock_agents",
"providers/bedrock_vector_store",
]
},
@ -505,6 +507,7 @@ const sidebars = {
items: [
"tutorials/openweb_ui",
"tutorials/openai_codex",
"tutorials/anthropic_file_usage",
"tutorials/msft_sso",
"tutorials/prompt_caching",
"tutorials/tag_management",

Binary file not shown.

Binary file not shown.

View file

@ -1,17 +1,23 @@
from litellm.proxy._types import SpendLogsPayload
from litellm._logging import verbose_proxy_logger
from typing import Optional, List, Union
import json
from litellm.types.utils import ModelResponse, Message
from typing import TYPE_CHECKING, Any, List, Optional, Union, cast
from litellm._logging import verbose_proxy_logger
from litellm.proxy._types import SpendLogsPayload
from litellm.responses.utils import ResponsesAPIRequestUtils
from litellm.types.llms.openai import (
AllMessageValues,
ChatCompletionResponseMessage,
GenericChatCompletionMessage,
ResponseInputParam,
)
from litellm.types.utils import ChatCompletionMessageToolCall
from litellm.responses.utils import ResponsesAPIRequestUtils
from litellm.responses.litellm_completion_transformation.transformation import ChatCompletionSession
from litellm.types.utils import ChatCompletionMessageToolCall, Message, ModelResponse
if TYPE_CHECKING:
from litellm.responses.litellm_completion_transformation.transformation import (
ChatCompletionSession,
)
else:
ChatCompletionSession = Any
class _ENTERPRISE_ResponsesSessionHandler:
@ -22,9 +28,23 @@ class _ENTERPRISE_ResponsesSessionHandler:
"""
Return the chat completion message history for a previous response id
"""
from litellm.responses.litellm_completion_transformation.transformation import LiteLLMCompletionResponsesConfig
all_spend_logs: List[SpendLogsPayload] = await _ENTERPRISE_ResponsesSessionHandler.get_all_spend_logs_for_previous_response_id(previous_response_id)
from litellm.responses.litellm_completion_transformation.transformation import (
ChatCompletionSession,
LiteLLMCompletionResponsesConfig,
)
verbose_proxy_logger.debug(
"inside get_chat_completion_message_history_for_previous_response_id"
)
all_spend_logs: List[
SpendLogsPayload
] = await _ENTERPRISE_ResponsesSessionHandler.get_all_spend_logs_for_previous_response_id(
previous_response_id
)
verbose_proxy_logger.debug(
"found %s spend logs for this response id", len(all_spend_logs)
)
litellm_session_id: Optional[str] = None
if len(all_spend_logs) > 0:
litellm_session_id = all_spend_logs[0].get("session_id")
@ -39,14 +59,16 @@ class _ENTERPRISE_ResponsesSessionHandler:
]
] = []
for spend_log in all_spend_logs:
proxy_server_request: Union[str, dict] = spend_log.get("proxy_server_request") or "{}"
proxy_server_request: Union[str, dict] = (
spend_log.get("proxy_server_request") or "{}"
)
proxy_server_request_dict: Optional[dict] = None
response_input_param: Optional[Union[str, ResponseInputParam]] = None
if isinstance(proxy_server_request, dict):
proxy_server_request_dict = proxy_server_request
else:
proxy_server_request_dict = json.loads(proxy_server_request)
############################################################
# Add Input messages for this Spend Log
############################################################
@ -55,15 +77,17 @@ class _ENTERPRISE_ResponsesSessionHandler:
if isinstance(_response_input_param, str):
response_input_param = _response_input_param
elif isinstance(_response_input_param, dict):
response_input_param = ResponseInputParam(**_response_input_param)
response_input_param = cast(
ResponseInputParam, _response_input_param
)
if response_input_param:
chat_completion_messages = LiteLLMCompletionResponsesConfig.transform_responses_api_input_to_messages(
input=response_input_param,
responses_api_request=proxy_server_request_dict or {}
responses_api_request=proxy_server_request_dict or {},
)
chat_completion_message_history.extend(chat_completion_messages)
############################################################
# Add Output messages for this Spend Log
############################################################
@ -73,17 +97,22 @@ class _ENTERPRISE_ResponsesSessionHandler:
model_response = ModelResponse(**_response_output)
for choice in model_response.choices:
if hasattr(choice, "message"):
chat_completion_message_history.append(choice.message)
verbose_proxy_logger.debug("chat_completion_message_history %s", json.dumps(chat_completion_message_history, indent=4, default=str))
chat_completion_message_history.append(
getattr(choice, "message")
)
verbose_proxy_logger.debug(
"chat_completion_message_history %s",
json.dumps(chat_completion_message_history, indent=4, default=str),
)
return ChatCompletionSession(
messages=chat_completion_message_history,
litellm_session_id=litellm_session_id
litellm_session_id=litellm_session_id,
)
@staticmethod
async def get_all_spend_logs_for_previous_response_id(
previous_response_id: str
previous_response_id: str,
) -> List[SpendLogsPayload]:
"""
Get all spend logs for a previous response id
@ -94,8 +123,17 @@ class _ENTERPRISE_ResponsesSessionHandler:
SELECT session_id FROM spend_logs WHERE response_id = previous_response_id, SELECT * FROM spend_logs WHERE session_id = session_id
"""
from litellm.proxy.proxy_server import prisma_client
decoded_response_id = ResponsesAPIRequestUtils._decode_responses_api_response_id(previous_response_id)
previous_response_id = decoded_response_id.get("response_id", previous_response_id)
verbose_proxy_logger.debug("decoding response id=%s", previous_response_id)
decoded_response_id = (
ResponsesAPIRequestUtils._decode_responses_api_response_id(
previous_response_id
)
)
previous_response_id = decoded_response_id.get(
"response_id", previous_response_id
)
if prisma_client is None:
return []
@ -111,21 +149,12 @@ class _ENTERPRISE_ResponsesSessionHandler:
ORDER BY "endTime" ASC;
"""
spend_logs = await prisma_client.db.query_raw(
query,
previous_response_id
)
spend_logs = await prisma_client.db.query_raw(query, previous_response_id)
verbose_proxy_logger.debug(
"Found the following spend logs for previous response id %s: %s",
previous_response_id,
json.dumps(spend_logs, indent=4, default=str)
json.dumps(spend_logs, indent=4, default=str),
)
return spend_logs

View file

@ -1,6 +1,6 @@
[tool.poetry]
name = "litellm-enterprise"
version = "0.1.6"
version = "0.1.7"
description = "Package for LiteLLM Enterprise features"
authors = ["BerriAI"]
readme = "README.md"
@ -22,7 +22,7 @@ requires = ["poetry-core"]
build-backend = "poetry.core.masonry.api"
[tool.commitizen]
version = "0.1.6"
version = "0.1.7"
version_files = [
"pyproject.toml:version",
"../requirements.txt:litellm-enterprise==",

View file

@ -119,6 +119,7 @@ _custom_logger_compatible_callbacks_literal = Literal[
"resend_email",
"smtp_email",
"deepeval",
"s3_v2",
]
logged_real_time_event_types: Optional[Union[List[str], Literal["*"]]] = None
_known_custom_logger_compatible_callbacks: List = list(
@ -133,7 +134,7 @@ langsmith_batch_size: Optional[int] = None
prometheus_initialize_budget_metrics: Optional[bool] = False
require_auth_for_metrics_endpoint: Optional[bool] = False
argilla_batch_size: Optional[int] = None
datadog_use_v1: Optional[bool] = False # if you want to use v1 datadog logged payload
datadog_use_v1: Optional[bool] = False # if you want to use v1 datadog logged payload.
gcs_pub_sub_use_v1: Optional[
bool
] = False # if you want to use v1 gcs pubsub logged payload
@ -190,6 +191,7 @@ maritalk_key: Optional[str] = None
ai21_key: Optional[str] = None
ollama_key: Optional[str] = None
openrouter_key: Optional[str] = None
datarobot_key: Optional[str] = None
predibase_key: Optional[str] = None
huggingface_key: Optional[str] = None
vertex_project: Optional[str] = None
@ -215,6 +217,7 @@ use_client: bool = False
ssl_verify: Union[str, bool] = True
ssl_certificate: Optional[str] = None
disable_streaming_logging: bool = False
disable_token_counter: bool = False
disable_add_transform_inline_image_block: bool = False
in_memory_llm_clients_cache: LLMClientCache = LLMClientCache()
safe_memory_mode: bool = False
@ -303,7 +306,8 @@ priority_reservation: Optional[Dict[str, float]] = None
######## Networking Settings ########
use_aiohttp_transport: bool = True
use_aiohttp_transport: bool = True # Older variable, aiohttp is now the default. use disable_aiohttp_transport instead.
disable_aiohttp_transport: bool = False # Set this to true to use httpx instead
force_ipv4: bool = False # when True, litellm will force ipv4 for all LLM requests. Some users have seen httpx ConnectionError when using ipv6.
module_level_aclient = AsyncHTTPHandler(
timeout=request_timeout, client_alias="module level aclient"
@ -401,6 +405,7 @@ mistral_chat_models: List = []
text_completion_codestral_models: List = []
anthropic_models: List = []
openrouter_models: List = []
datarobot_models: List = []
vertex_language_models: List = []
vertex_vision_models: List = []
vertex_chat_models: List = []
@ -508,6 +513,8 @@ def add_known_models():
empower_models.append(key)
elif value.get("litellm_provider") == "openrouter":
openrouter_models.append(key)
elif value.get("litellm_provider") == "datarobot":
datarobot_models.append(key)
elif value.get("litellm_provider") == "vertex_ai-text-models":
vertex_text_models.append(key)
elif value.get("litellm_provider") == "vertex_ai-code-text-models":
@ -658,6 +665,7 @@ model_list = (
+ anthropic_models
+ replicate_models
+ openrouter_models
+ datarobot_models
+ huggingface_models
+ vertex_chat_models
+ vertex_text_models
@ -718,6 +726,7 @@ models_by_provider: dict = {
"together_ai": together_ai_models,
"baseten": baseten_models,
"openrouter": openrouter_models,
"datarobot": datarobot_models,
"vertex_ai": vertex_chat_models
+ vertex_text_models
+ vertex_anthropic_models
@ -868,6 +877,7 @@ from .llms.huggingface.embedding.transformation import HuggingFaceEmbeddingConfi
from .llms.oobabooga.chat.transformation import OobaboogaConfig
from .llms.maritalk import MaritalkConfig
from .llms.openrouter.chat.transformation import OpenrouterConfig
from .llms.datarobot.chat.transformation import DataRobotConfig
from .llms.anthropic.chat.transformation import AnthropicConfig
from .llms.anthropic.common_utils import AnthropicModelInfo
from .llms.groq.stt.transformation import GroqSTTConfig
@ -1136,3 +1146,6 @@ disable_hf_tokenizer_download: Optional[
bool
] = None # disable huggingface tokenizer download. Defaults to openai clk100
global_disable_no_log_param: bool = False
### PASSTHROUGH ###
from .passthrough import allm_passthrough_route, llm_passthrough_route

View file

@ -84,6 +84,19 @@ class InMemoryCache(BaseCache):
except Exception:
return False
def _is_key_expired(self, key: str) -> bool:
"""
Check if a specific key is expired
"""
return key in self.ttl_dict and time.time() > self.ttl_dict[key]
def _remove_key(self, key: str) -> None:
"""
Remove a key from both cache_dict and ttl_dict
"""
self.cache_dict.pop(key, None)
self.ttl_dict.pop(key, None)
def evict_cache(self):
"""
Eviction policy:
@ -97,9 +110,8 @@ class InMemoryCache(BaseCache):
"""
for key in list(self.ttl_dict.keys()):
if time.time() > self.ttl_dict[key]:
self.cache_dict.pop(key, None)
self.ttl_dict.pop(key, None)
if self._is_key_expired(key):
self._remove_key(key)
# de-reference the removed item
# https://www.geeksforgeeks.org/diagnosing-and-fixing-memory-leaks-in-python/
@ -153,13 +165,21 @@ class InMemoryCache(BaseCache):
self.set_cache(key, init_value, ttl=ttl)
return value
def evict_element_if_expired(self, key: str) -> bool:
"""
Returns True if the element is expired and removed from the cache
Returns False if the element is not expired
"""
if self._is_key_expired(key):
self._remove_key(key)
return True
return False
def get_cache(self, key, **kwargs):
if key in self.cache_dict:
if key in self.ttl_dict:
if time.time() > self.ttl_dict[key]:
self.cache_dict.pop(key, None)
self.ttl_dict.pop(key, None)
return None
if self.evict_element_if_expired(key):
return None
original_cached_response = self.cache_dict[key]
try:
cached_response = json.loads(original_cached_response)
@ -207,8 +227,7 @@ class InMemoryCache(BaseCache):
pass
def delete_cache(self, key):
self.cache_dict.pop(key, None)
self.ttl_dict.pop(key, None)
self._remove_key(key)
async def async_get_ttl(self, key: str) -> Optional[int]:
"""

View file

@ -4,6 +4,10 @@ from typing import List, Literal
ROUTER_MAX_FALLBACKS = int(os.getenv("ROUTER_MAX_FALLBACKS", 5))
DEFAULT_BATCH_SIZE = int(os.getenv("DEFAULT_BATCH_SIZE", 512))
DEFAULT_FLUSH_INTERVAL_SECONDS = int(os.getenv("DEFAULT_FLUSH_INTERVAL_SECONDS", 5))
DEFAULT_S3_FLUSH_INTERVAL_SECONDS = int(
os.getenv("DEFAULT_S3_FLUSH_INTERVAL_SECONDS", 10)
)
DEFAULT_S3_BATCH_SIZE = int(os.getenv("DEFAULT_S3_BATCH_SIZE", 512))
DEFAULT_MAX_RETRIES = int(os.getenv("DEFAULT_MAX_RETRIES", 2))
DEFAULT_MAX_RECURSE_DEPTH = int(os.getenv("DEFAULT_MAX_RECURSE_DEPTH", 100))
DEFAULT_MAX_RECURSE_DEPTH_SENSITIVE_DATA_MASKER = int(
@ -154,7 +158,10 @@ FIREWORKS_AI_80_B = int(os.getenv("FIREWORKS_AI_80_B", 80))
#### Logging callback constants ####
REDACTED_BY_LITELM_STRING = "REDACTED_BY_LITELM"
MAX_LANGFUSE_INITIALIZED_CLIENTS = int(
os.getenv("MAX_LANGFUSE_INITIALIZED_CLIENTS", 20)
os.getenv("MAX_LANGFUSE_INITIALIZED_CLIENTS", 50)
)
DD_TRACER_STREAMING_CHUNK_YIELD_RESOURCE = os.getenv(
"DD_TRACER_STREAMING_CHUNK_YIELD_RESOURCE", "streaming.chunk.yield"
)
############### LLM Provider Constants ###############
@ -180,6 +187,7 @@ LITELLM_CHAT_PROVIDERS = [
"replicate",
"huggingface",
"together_ai",
"datarobot",
"openrouter",
"vertex_ai",
"vertex_ai_beta",
@ -587,6 +595,7 @@ BEDROCK_INVOKE_PROVIDERS_LITERAL = Literal[
open_ai_embedding_models: List = ["text-embedding-ada-002"]
cohere_embedding_models: List = [
"embed-v4.0",
"embed-english-v3.0",
"embed-english-light-v3.0",
"embed-multilingual-v3.0",

View file

@ -0,0 +1,438 @@
"""
s3 Bucket Logging Integration
async_log_success_event: Processes the event, stores it in memory for DEFAULT_S3_FLUSH_INTERVAL_SECONDS seconds or until DEFAULT_S3_BATCH_SIZE and then flushes to s3
NOTE 1: S3 does not provide a BATCH PUT API endpoint, so we create tasks to upload each element individually
"""
import asyncio
import json
from datetime import datetime
from typing import List, Optional, cast
import litellm
from litellm._logging import print_verbose, verbose_logger
from litellm.constants import DEFAULT_S3_BATCH_SIZE, DEFAULT_S3_FLUSH_INTERVAL_SECONDS
from litellm.integrations.s3 import get_s3_object_key
from litellm.llms.bedrock.base_aws_llm import BaseAWSLLM
from litellm.llms.custom_httpx.http_handler import (
_get_httpx_client,
get_async_httpx_client,
httpxSpecialProvider,
)
from litellm.types.integrations.s3_v2 import s3BatchLoggingElement
from litellm.types.utils import StandardLoggingPayload
from .custom_batch_logger import CustomBatchLogger
class S3Logger(CustomBatchLogger, BaseAWSLLM):
def __init__(
self,
s3_bucket_name: Optional[str] = None,
s3_path: Optional[str] = None,
s3_region_name: Optional[str] = None,
s3_api_version: Optional[str] = None,
s3_use_ssl: bool = True,
s3_verify: Optional[bool] = None,
s3_endpoint_url: Optional[str] = None,
s3_aws_access_key_id: Optional[str] = None,
s3_aws_secret_access_key: Optional[str] = None,
s3_aws_session_token: Optional[str] = None,
s3_aws_session_name: Optional[str] = None,
s3_aws_profile_name: Optional[str] = None,
s3_aws_role_name: Optional[str] = None,
s3_aws_web_identity_token: Optional[str] = None,
s3_aws_sts_endpoint: Optional[str] = None,
s3_flush_interval: Optional[int] = DEFAULT_S3_FLUSH_INTERVAL_SECONDS,
s3_batch_size: Optional[int] = DEFAULT_S3_BATCH_SIZE,
s3_config=None,
s3_use_team_prefix: bool = False,
**kwargs,
):
try:
verbose_logger.debug(
f"in init s3 logger - s3_callback_params {litellm.s3_callback_params}"
)
# IMPORTANT: We use a concurrent limit of 1 to upload to s3
# Files should get uploaded BUT they should not impact latency of LLM calling logic
self.async_httpx_client = get_async_httpx_client(
llm_provider=httpxSpecialProvider.LoggingCallback,
)
self._init_s3_params(
s3_bucket_name=s3_bucket_name,
s3_region_name=s3_region_name,
s3_api_version=s3_api_version,
s3_use_ssl=s3_use_ssl,
s3_verify=s3_verify,
s3_endpoint_url=s3_endpoint_url,
s3_aws_access_key_id=s3_aws_access_key_id,
s3_aws_secret_access_key=s3_aws_secret_access_key,
s3_aws_session_token=s3_aws_session_token,
s3_aws_session_name=s3_aws_session_name,
s3_aws_profile_name=s3_aws_profile_name,
s3_aws_role_name=s3_aws_role_name,
s3_aws_web_identity_token=s3_aws_web_identity_token,
s3_aws_sts_endpoint=s3_aws_sts_endpoint,
s3_config=s3_config,
s3_path=s3_path,
s3_use_team_prefix=s3_use_team_prefix,
)
verbose_logger.debug(f"s3 logger using endpoint url {s3_endpoint_url}")
asyncio.create_task(self.periodic_flush())
self.flush_lock = asyncio.Lock()
verbose_logger.debug(
f"s3 flush interval: {s3_flush_interval}, s3 batch size: {s3_batch_size}"
)
# Call CustomLogger's __init__
CustomBatchLogger.__init__(
self,
flush_lock=self.flush_lock,
flush_interval=s3_flush_interval,
batch_size=s3_batch_size,
)
self.log_queue: List[s3BatchLoggingElement] = []
# Call BaseAWSLLM's __init__
BaseAWSLLM.__init__(self)
except Exception as e:
print_verbose(f"Got exception on init s3 client {str(e)}")
raise e
def _init_s3_params(
self,
s3_bucket_name: Optional[str] = None,
s3_region_name: Optional[str] = None,
s3_api_version: Optional[str] = None,
s3_use_ssl: bool = True,
s3_verify: Optional[bool] = None,
s3_endpoint_url: Optional[str] = None,
s3_aws_access_key_id: Optional[str] = None,
s3_aws_secret_access_key: Optional[str] = None,
s3_aws_session_token: Optional[str] = None,
s3_aws_session_name: Optional[str] = None,
s3_aws_profile_name: Optional[str] = None,
s3_aws_role_name: Optional[str] = None,
s3_aws_web_identity_token: Optional[str] = None,
s3_aws_sts_endpoint: Optional[str] = None,
s3_config=None,
s3_path: Optional[str] = None,
s3_use_team_prefix: bool = False,
):
"""
Initialize the s3 params for this logging callback
"""
litellm.s3_callback_params = litellm.s3_callback_params or {}
# read in .env variables - example os.environ/AWS_BUCKET_NAME
for key, value in litellm.s3_callback_params.items():
if isinstance(value, str) and value.startswith("os.environ/"):
litellm.s3_callback_params[key] = litellm.get_secret(value)
self.s3_bucket_name = (
litellm.s3_callback_params.get("s3_bucket_name") or s3_bucket_name
)
self.s3_region_name = (
litellm.s3_callback_params.get("s3_region_name") or s3_region_name
)
self.s3_api_version = (
litellm.s3_callback_params.get("s3_api_version") or s3_api_version
)
self.s3_use_ssl = (
litellm.s3_callback_params.get("s3_use_ssl", True) or s3_use_ssl
)
self.s3_verify = litellm.s3_callback_params.get("s3_verify") or s3_verify
self.s3_endpoint_url = (
litellm.s3_callback_params.get("s3_endpoint_url") or s3_endpoint_url
)
self.s3_aws_access_key_id = (
litellm.s3_callback_params.get("s3_aws_access_key_id")
or s3_aws_access_key_id
)
self.s3_aws_secret_access_key = (
litellm.s3_callback_params.get("s3_aws_secret_access_key")
or s3_aws_secret_access_key
)
self.s3_aws_session_token = (
litellm.s3_callback_params.get("s3_aws_session_token")
or s3_aws_session_token
)
self.s3_aws_session_name = (
litellm.s3_callback_params.get("s3_aws_session_name") or s3_aws_session_name
)
self.s3_aws_profile_name = (
litellm.s3_callback_params.get("s3_aws_profile_name") or s3_aws_profile_name
)
self.s3_aws_role_name = (
litellm.s3_callback_params.get("s3_aws_role_name") or s3_aws_role_name
)
self.s3_aws_web_identity_token = (
litellm.s3_callback_params.get("s3_aws_web_identity_token")
or s3_aws_web_identity_token
)
self.s3_aws_sts_endpoint = (
litellm.s3_callback_params.get("s3_aws_sts_endpoint") or s3_aws_sts_endpoint
)
self.s3_config = litellm.s3_callback_params.get("s3_config") or s3_config
self.s3_path = litellm.s3_callback_params.get("s3_path") or s3_path
# done reading litellm.s3_callback_params
self.s3_use_team_prefix = (
bool(litellm.s3_callback_params.get("s3_use_team_prefix", False))
or s3_use_team_prefix
)
return
async def async_log_success_event(self, kwargs, response_obj, start_time, end_time):
try:
verbose_logger.debug(
f"s3 Logging - Enters logging function for model {kwargs}"
)
s3_batch_logging_element = self.create_s3_batch_logging_element(
start_time=start_time,
standard_logging_payload=kwargs.get("standard_logging_object", None),
)
if s3_batch_logging_element is None:
raise ValueError("s3_batch_logging_element is None")
verbose_logger.debug(
"\ns3 Logger - Logging payload = %s", s3_batch_logging_element
)
self.log_queue.append(s3_batch_logging_element)
verbose_logger.debug(
"s3 logging: queue length %s, batch size %s",
len(self.log_queue),
self.batch_size,
)
except Exception as e:
verbose_logger.exception(f"s3 Layer Error - {str(e)}")
pass
async def async_upload_data_to_s3(
self, batch_logging_element: s3BatchLoggingElement
):
try:
import hashlib
import requests
from botocore.auth import SigV4Auth
from botocore.awsrequest import AWSRequest
except ImportError:
raise ImportError("Missing boto3 to call bedrock. Run 'pip install boto3'.")
try:
from litellm.litellm_core_utils.asyncify import asyncify
asyncified_get_credentials = asyncify(self.get_credentials)
credentials = await asyncified_get_credentials(
aws_access_key_id=self.s3_aws_access_key_id,
aws_secret_access_key=self.s3_aws_secret_access_key,
aws_session_token=self.s3_aws_session_token,
aws_region_name=self.s3_region_name,
aws_session_name=self.s3_aws_session_name,
aws_profile_name=self.s3_aws_profile_name,
aws_role_name=self.s3_aws_role_name,
aws_web_identity_token=self.s3_aws_web_identity_token,
aws_sts_endpoint=self.s3_aws_sts_endpoint,
)
verbose_logger.debug(
f"s3_v2 logger - uploading data to s3 - {batch_logging_element.s3_object_key}"
)
# Prepare the URL
url = f"https://{self.s3_bucket_name}.s3.{self.s3_region_name}.amazonaws.com/{batch_logging_element.s3_object_key}"
if self.s3_endpoint_url:
url = self.s3_endpoint_url + "/" + batch_logging_element.s3_object_key
# Convert JSON to string
json_string = json.dumps(batch_logging_element.payload)
# Calculate SHA256 hash of the content
content_hash = hashlib.sha256(json_string.encode("utf-8")).hexdigest()
# Prepare the request
headers = {
"Content-Type": "application/json",
"x-amz-content-sha256": content_hash,
"Content-Language": "en",
"Content-Disposition": f'inline; filename="{batch_logging_element.s3_object_download_filename}"',
"Cache-Control": "private, immutable, max-age=31536000, s-maxage=0",
}
req = requests.Request("PUT", url, data=json_string, headers=headers)
prepped = req.prepare()
# Sign the request
aws_request = AWSRequest(
method=prepped.method,
url=prepped.url,
data=prepped.body,
headers=prepped.headers,
)
SigV4Auth(credentials, "s3", self.s3_region_name).add_auth(aws_request)
# Prepare the signed headers
signed_headers = dict(aws_request.headers.items())
# Make the request
response = await self.async_httpx_client.put(
url, data=json_string, headers=signed_headers
)
response.raise_for_status()
except Exception as e:
verbose_logger.exception(f"Error uploading to s3: {str(e)}")
async def async_send_batch(self):
"""
Sends runs from self.log_queue
Returns: None
Raises: Does not raise an exception, will only verbose_logger.exception()
"""
verbose_logger.debug(f"s3_v2 logger - sending batch of {len(self.log_queue)}")
if not self.log_queue:
return
#########################################################
# Flush the log queue to s3
# the log queue can be bounded by DEFAULT_S3_BATCH_SIZE
# see custom_batch_logger.py which triggers the flush
#########################################################
for payload in self.log_queue:
asyncio.create_task(self.async_upload_data_to_s3(payload))
def create_s3_batch_logging_element(
self,
start_time: datetime,
standard_logging_payload: Optional[StandardLoggingPayload],
) -> Optional[s3BatchLoggingElement]:
"""
Helper function to create an s3BatchLoggingElement.
Args:
start_time (datetime): The start time of the logging event.
standard_logging_payload (Optional[StandardLoggingPayload]): The payload to be logged.
s3_path (Optional[str]): The S3 path prefix.
Returns:
Optional[s3BatchLoggingElement]: The created s3BatchLoggingElement, or None if payload is None.
"""
if standard_logging_payload is None:
return None
team_alias = standard_logging_payload["metadata"].get("user_api_key_team_alias")
team_alias_prefix = ""
if (
litellm.enable_preview_features
and self.s3_use_team_prefix
and team_alias is not None
):
team_alias_prefix = f"{team_alias}/"
s3_file_name = (
litellm.utils.get_logging_id(start_time, standard_logging_payload) or ""
)
s3_object_key = get_s3_object_key(
s3_path=cast(Optional[str], self.s3_path) or "",
team_alias_prefix=team_alias_prefix,
start_time=start_time,
s3_file_name=s3_file_name,
)
s3_object_download_filename = (
"time-"
+ start_time.strftime("%Y-%m-%dT%H-%M-%S-%f")
+ "_"
+ standard_logging_payload["id"]
+ ".json"
)
s3_object_download_filename = f"time-{start_time.strftime('%Y-%m-%dT%H-%M-%S-%f')}_{standard_logging_payload['id']}.json"
return s3BatchLoggingElement(
payload=dict(standard_logging_payload),
s3_object_key=s3_object_key,
s3_object_download_filename=s3_object_download_filename,
)
def upload_data_to_s3(self, batch_logging_element: s3BatchLoggingElement):
try:
import hashlib
import requests
from botocore.auth import SigV4Auth
from botocore.awsrequest import AWSRequest
from botocore.credentials import Credentials
except ImportError:
raise ImportError("Missing boto3 to call bedrock. Run 'pip install boto3'.")
try:
verbose_logger.debug(
f"s3_v2 logger - uploading data to s3 - {batch_logging_element.s3_object_key}"
)
credentials: Credentials = self.get_credentials(
aws_access_key_id=self.s3_aws_access_key_id,
aws_secret_access_key=self.s3_aws_secret_access_key,
aws_session_token=self.s3_aws_session_token,
aws_region_name=self.s3_region_name,
)
# Prepare the URL
url = f"https://{self.s3_bucket_name}.s3.{self.s3_region_name}.amazonaws.com/{batch_logging_element.s3_object_key}"
if self.s3_endpoint_url:
url = self.s3_endpoint_url + "/" + batch_logging_element.s3_object_key
# Convert JSON to string
json_string = json.dumps(batch_logging_element.payload)
# Calculate SHA256 hash of the content
content_hash = hashlib.sha256(json_string.encode("utf-8")).hexdigest()
# Prepare the request
headers = {
"Content-Type": "application/json",
"x-amz-content-sha256": content_hash,
"Content-Language": "en",
"Content-Disposition": f'inline; filename="{batch_logging_element.s3_object_download_filename}"',
"Cache-Control": "private, immutable, max-age=31536000, s-maxage=0",
}
req = requests.Request("PUT", url, data=json_string, headers=headers)
prepped = req.prepare()
# Sign the request
aws_request = AWSRequest(
method=prepped.method,
url=prepped.url,
data=prepped.body,
headers=prepped.headers,
)
SigV4Auth(credentials, "s3", self.s3_region_name).add_auth(aws_request)
# Prepare the signed headers
signed_headers = dict(aws_request.headers.items())
httpx_client = _get_httpx_client()
# Make the request
response = httpx_client.put(url, data=json_string, headers=signed_headers)
response.raise_for_status()
except Exception as e:
verbose_logger.exception(f"Error uploading to s3: {str(e)}")

View file

@ -34,7 +34,6 @@ from litellm.types.vector_stores import (
VectorStoreSearchResponse,
VectorStoreSearchResult,
)
from litellm.utils import load_credentials_from_list
if TYPE_CHECKING:
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
@ -258,22 +257,49 @@ class BedrockVectorStore(BaseVectorStore, BaseAWSLLM):
from fastapi import HTTPException
non_default_params = non_default_params or {}
load_credentials_from_list(kwargs=non_default_params)
credentials_dict: Dict[str, Any] = {}
if litellm.vector_store_registry is not None:
credentials_dict = (
litellm.vector_store_registry.get_credentials_for_vector_store(
knowledge_base_id
)
)
credentials = self.get_credentials(
aws_access_key_id=non_default_params.get("aws_access_key_id", None),
aws_secret_access_key=non_default_params.get("aws_secret_access_key", None),
aws_session_token=non_default_params.get("aws_session_token", None),
aws_region_name=non_default_params.get("aws_region_name", None),
aws_session_name=non_default_params.get("aws_session_name", None),
aws_profile_name=non_default_params.get("aws_profile_name", None),
aws_role_name=non_default_params.get("aws_role_name", None),
aws_web_identity_token=non_default_params.get(
"aws_web_identity_token", None
aws_access_key_id=credentials_dict.get(
"aws_access_key_id", non_default_params.get("aws_access_key_id", None)
),
aws_secret_access_key=credentials_dict.get(
"aws_secret_access_key",
non_default_params.get("aws_secret_access_key", None),
),
aws_session_token=credentials_dict.get(
"aws_session_token", non_default_params.get("aws_session_token", None)
),
aws_region_name=credentials_dict.get(
"aws_region_name", non_default_params.get("aws_region_name", None)
),
aws_session_name=credentials_dict.get(
"aws_session_name", non_default_params.get("aws_session_name", None)
),
aws_profile_name=credentials_dict.get(
"aws_profile_name", non_default_params.get("aws_profile_name", None)
),
aws_role_name=credentials_dict.get(
"aws_role_name", non_default_params.get("aws_role_name", None)
),
aws_web_identity_token=credentials_dict.get(
"aws_web_identity_token",
non_default_params.get("aws_web_identity_token", None),
),
aws_sts_endpoint=credentials_dict.get(
"aws_sts_endpoint", non_default_params.get("aws_sts_endpoint", None)
),
aws_sts_endpoint=non_default_params.get("aws_sts_endpoint", None),
)
aws_region_name = self._get_aws_region_name(
optional_params=self.optional_params
aws_region_name = self.get_aws_region_name_for_non_llm_api_calls(
aws_region_name=credentials_dict.get(
"aws_region_name", non_default_params.get("aws_region_name", None)
),
)
# Prepare request data

View file

@ -514,6 +514,14 @@ def _get_openai_compatible_provider_info( # noqa: PLR0915
) = litellm.LlamafileChatConfig()._get_openai_compatible_provider_info(
api_base, api_key
)
elif custom_llm_provider == "datarobot":
# DataRobot is OpenAI compatible.
(
api_base,
dynamic_api_key
) = litellm.DataRobotConfig()._get_openai_compatible_provider_info(
api_base, api_key
)
elif custom_llm_provider == "lm_studio":
# lm_studio is openai compatible, we just need to set this to custom_openai
(

View file

@ -135,6 +135,7 @@ from ..integrations.opik.opik import OpikLogger
from ..integrations.prometheus import PrometheusLogger
from ..integrations.prompt_layer import PromptLayerLogger
from ..integrations.s3 import S3Logger
from ..integrations.s3_v2 import S3Logger as S3V2Logger
from ..integrations.supabase import Supabase
from ..integrations.traceloop import TraceloopLogger
from ..integrations.weights_biases import WeightsBiasesLogger
@ -2699,7 +2700,9 @@ def set_callbacks(callback_list, function_id=None): # noqa: PLR0915
sentry_sdk_instance.init(
dsn=os.environ.get("SENTRY_DSN"),
traces_sample_rate=float(sentry_trace_rate), # type: ignore
sample_rate=float(sentry_sample_rate),
sample_rate=float(
sentry_sample_rate if sentry_sample_rate else 1.0
),
)
capture_exception = sentry_sdk_instance.capture_exception
add_breadcrumb = sentry_sdk_instance.add_breadcrumb
@ -2867,6 +2870,14 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
_gcs_bucket_logger = GCSBucketLogger()
_in_memory_loggers.append(_gcs_bucket_logger)
return _gcs_bucket_logger # type: ignore
elif logging_integration == "s3_v2":
for callback in _in_memory_loggers:
if isinstance(callback, S3V2Logger):
return callback # type: ignore
_s3_v2_logger = S3V2Logger()
_in_memory_loggers.append(_s3_v2_logger)
return _s3_v2_logger # type: ignore
elif logging_integration == "azure_storage":
for callback in _in_memory_loggers:
if isinstance(callback, AzureBlobStorageLogger):
@ -2962,7 +2973,7 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
galileo_logger = GalileoObserve()
_in_memory_loggers.append(galileo_logger)
return galileo_logger # type: ignore
elif logging_integration == "deepeval":
for callback in _in_memory_loggers:
if isinstance(callback, DeepEvalLogger):
@ -2970,7 +2981,7 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
deepeval_logger = DeepEvalLogger()
_in_memory_loggers.append(deepeval_logger)
return deepeval_logger # type: ignore
elif logging_integration == "logfire":
if "LOGFIRE_TOKEN" not in os.environ:
raise ValueError("LOGFIRE_TOKEN not found in environment variables")
@ -3172,6 +3183,10 @@ def get_custom_logger_compatible_class( # noqa: PLR0915
for callback in _in_memory_loggers:
if isinstance(callback, GCSBucketLogger):
return callback
elif logging_integration == "s3_v2":
for callback in _in_memory_loggers:
if isinstance(callback, S3V2Logger):
return callback
elif logging_integration == "azure_storage":
for callback in _in_memory_loggers:
if isinstance(callback, AzureBlobStorageLogger):

View file

@ -532,6 +532,12 @@ def convert_to_model_response_object( # noqa: PLR0915
if finish_reason is None:
# gpt-4 vision can return 'finish_reason' or 'finish_details'
finish_reason = choice.get("finish_details") or "stop"
if (
finish_reason == "stop"
and message.tool_calls
and len(message.tool_calls) > 0
):
finish_reason = "tool_calls"
logprobs = choice.get("logprobs", None)
enhancements = choice.get("enhancements", None)
choice = Choices(

View file

@ -72,13 +72,11 @@ def get_api_base(
_optional_params.vertex_location is not None
and _optional_params.vertex_project is not None
):
from litellm.llms.vertex_ai.vertex_ai_partner_models.main import (
VertexPartnerProvider,
create_vertex_url,
)
from litellm.llms.vertex_ai.vertex_llm_base import VertexBase
from litellm.types.llms.vertex_ai import VertexPartnerProvider
if "claude" in model:
_api_base = create_vertex_url(
_api_base = VertexBase.create_vertex_url(
vertex_location=_optional_params.vertex_location,
vertex_project=_optional_params.vertex_project,
model=model,

View file

@ -6,7 +6,17 @@ import io
import mimetypes
import re
from os import PathLike
from typing import Any, Dict, List, Literal, Mapping, Optional, Union, cast
from typing import (
TYPE_CHECKING,
Any,
Dict,
List,
Literal,
Mapping,
Optional,
Union,
cast,
)
from litellm.types.llms.openai import (
AllMessageValues,
@ -25,6 +35,9 @@ from litellm.types.utils import (
StreamingChoices,
)
if TYPE_CHECKING: # newer pattern to avoid importing pydantic objects on __init__.py
from litellm.types.llms.openai import ChatCompletionImageObject
DEFAULT_USER_CONTINUE_MESSAGE = ChatCompletionUserMessage(
content="Please continue.", role="user"
)
@ -33,6 +46,9 @@ DEFAULT_ASSISTANT_CONTINUE_MESSAGE = ChatCompletionAssistantMessage(
content="Please continue.", role="assistant"
)
if TYPE_CHECKING:
from litellm.litellm_core_utils.litellm_logging import Logging as LoggingClass
def handle_any_messages_to_chat_completion_str_messages_conversion(
messages: Any,
@ -582,3 +598,93 @@ def is_function_call(optional_params: dict) -> bool:
if "functions" in optional_params and optional_params.get("functions"):
return True
return False
def get_file_ids_from_messages(messages: List[AllMessageValues]) -> List[str]:
"""
Gets file ids from messages
"""
file_ids = []
for message in messages:
if message.get("role") == "user":
content = message.get("content")
if content:
if isinstance(content, str):
continue
for c in content:
if c["type"] == "file":
file_object = cast(ChatCompletionFileObject, c)
file_object_file_field = file_object["file"]
file_id = file_object_file_field.get("file_id")
if file_id:
file_ids.append(file_id)
return file_ids
def check_is_function_call(logging_obj: "LoggingClass") -> bool:
from litellm.litellm_core_utils.prompt_templates.common_utils import (
is_function_call,
)
if hasattr(logging_obj, "optional_params") and isinstance(
logging_obj.optional_params, dict
):
if is_function_call(logging_obj.optional_params):
return True
return False
def filter_value_from_dict(dictionary: dict, key: str, depth: int = 0) -> Any:
"""
Filters a value from a dictionary
Goes through the nested dict and removes the key if it exists
"""
from litellm.constants import DEFAULT_MAX_RECURSE_DEPTH
if depth > DEFAULT_MAX_RECURSE_DEPTH:
return dictionary
# Create a copy of keys to avoid modifying dict during iteration
keys = list(dictionary.keys())
for k in keys:
v = dictionary[k]
if k == key:
del dictionary[k]
elif isinstance(v, dict):
filter_value_from_dict(v, key, depth + 1)
elif isinstance(v, list):
for item in v:
if isinstance(item, dict):
filter_value_from_dict(item, key, depth + 1)
return dictionary
def migrate_file_to_image_url(
message: "ChatCompletionFileObject",
) -> "ChatCompletionImageObject":
"""
Migrate file to image_url
"""
from litellm.types.llms.openai import (
ChatCompletionImageObject,
ChatCompletionImageUrlObject,
)
file_id = message["file"].get("file_id")
file_data = message["file"].get("file_data")
format = message["file"].get("format")
if not file_id and not file_data:
raise ValueError("file_id and file_data are both None")
image_url_object = ChatCompletionImageObject(
type="image_url",
image_url=ChatCompletionImageUrlObject(
url=cast(str, file_id or file_data),
),
)
if format and isinstance(image_url_object["image_url"], dict):
image_url_object["image_url"]["format"] = format
return image_url_object

View file

@ -1385,6 +1385,84 @@ def _anthropic_content_element_factory(
return _anthropic_content_element
def select_anthropic_content_block_type_for_file(
format: str,
) -> Literal["document", "image", "container_upload"]:
if format == "application/pdf" or format == "text/plain":
return "document"
elif format in ["image/jpeg", "image/png", "image/gif", "image/webp"]:
return "image"
else:
return "container_upload"
def anthropic_process_openai_file_message(
message: ChatCompletionFileObject,
) -> Union[
AnthropicMessagesDocumentParam,
AnthropicMessagesImageParam,
AnthropicMessagesContainerUploadParam,
]:
file_message = cast(ChatCompletionFileObject, message)
file_data = file_message["file"].get("file_data")
file_id = file_message["file"].get("file_id")
format = file_message["file"].get("format")
if file_data:
image_chunk = convert_to_anthropic_image_obj(
openai_image_url=file_data,
format=format,
)
anthropic_document_param = AnthropicMessagesDocumentParam(
type="document",
source=AnthropicContentParamSource(
type="base64",
media_type=image_chunk["media_type"],
data=image_chunk["data"],
),
)
return anthropic_document_param
elif file_id:
content_block_type = (
select_anthropic_content_block_type_for_file(format)
if format
else "container_upload"
)
return_block_param: Optional[
Union[
AnthropicMessagesDocumentParam,
AnthropicMessagesImageParam,
AnthropicMessagesContainerUploadParam,
]
] = None
if content_block_type == "document":
return_block_param = AnthropicMessagesDocumentParam(
type="document",
source=AnthropicContentParamSourceFileId(
type="file",
file_id=file_id,
),
)
elif content_block_type == "image":
return_block_param = AnthropicMessagesImageParam(
type="image",
source=AnthropicContentParamSourceFileId(
type="file",
file_id=file_id,
),
)
elif content_block_type == "container_upload":
return_block_param = AnthropicMessagesContainerUploadParam(
type="container_upload", file_id=file_id
)
if return_block_param is None:
raise Exception(f"Unable to parse anthropic file message: {message}")
return return_block_param
raise Exception(
f"Either file_data or file_id must be present in the file message: {message}"
)
def anthropic_messages_pt( # noqa: PLR0915
messages: List[AllMessageValues],
model: str,
@ -1489,24 +1567,11 @@ def anthropic_messages_pt( # noqa: PLR0915
elif m.get("type", "") == "document":
user_content.append(cast(AnthropicMessagesDocumentParam, m))
elif m.get("type", "") == "file":
file_message = cast(ChatCompletionFileObject, m)
file_data = file_message["file"].get("file_data")
if file_data:
image_chunk = convert_to_anthropic_image_obj(
openai_image_url=file_data,
format=file_message["file"].get("format"),
user_content.append(
anthropic_process_openai_file_message(
cast(ChatCompletionFileObject, m)
)
anthropic_document_param = (
AnthropicMessagesDocumentParam(
type="document",
source=AnthropicContentParamSource(
type="base64",
media_type=image_chunk["media_type"],
data=image_chunk["data"],
),
)
)
user_content.append(anthropic_document_param)
)
elif isinstance(user_message_types_block["content"], str):
_anthropic_content_text_element: AnthropicMessagesTextParam = {
"type": "text",

View file

@ -1,10 +1,56 @@
"""
This is a cache for LangfuseLoggers.
Langfuse Python SDK initializes a thread for each client.
This ensures we do
1. Proper cleanup of Langfuse initialized clients.
2. Re-use created langfuse clients.
"""
import hashlib
import json
from typing import Any, Optional
import litellm
from litellm.constants import _DEFAULT_TTL_FOR_HTTPX_CLIENTS
from ...caching import InMemoryCache
class LangfuseInMemoryCache(InMemoryCache):
"""
Ensures we do proper cleanup of Langfuse initialized clients.
Langfuse Python SDK initializes a thread for each client, we need to call Langfuse.shutdown() to properly cleanup.
This ensures we do proper cleanup of Langfuse initialized clients.
"""
def _remove_key(self, key: str) -> None:
"""
Override _remove_key in InMemoryCache to ensure we do proper cleanup of Langfuse initialized clients.
LangfuseLoggers consume threads when initalized, this shuts them down when they are expired
Relevant Issue: https://github.com/BerriAI/litellm/issues/11169
"""
from litellm.integrations.langfuse.langfuse import LangFuseLogger
if isinstance(self.cache_dict[key], LangFuseLogger):
_created_langfuse_logger: LangFuseLogger = self.cache_dict[key]
#########################################################
# Clean up Langfuse initialized clients
#########################################################
litellm.initialized_langfuse_clients -= 1
_created_langfuse_logger.Langfuse.flush()
_created_langfuse_logger.Langfuse.shutdown()
#########################################################
# Call parent class to remove key from cache
#########################################################
return super()._remove_key(key)
class DynamicLoggingCache:
"""
Prevent memory leaks caused by initializing new logging clients on each request.
@ -13,7 +59,7 @@ class DynamicLoggingCache:
"""
def __init__(self) -> None:
self.cache = InMemoryCache()
self.cache = LangfuseInMemoryCache(default_ttl=_DEFAULT_TTL_FOR_HTTPX_CLIENTS)
def get_cache_key(self, args: dict) -> str:
args_str = json.dumps(args, sort_keys=True)

View file

@ -1359,6 +1359,7 @@ class CustomStreamWrapper:
print_verbose(f"self.sent_first_chunk: {self.sent_first_chunk}")
## CHECK FOR TOOL USE
if "tool_calls" in completion_obj and len(completion_obj["tool_calls"]) > 0:
if self.is_function_call is True: # user passed in 'functions' param
completion_obj["function_call"] = completion_obj["tool_calls"][0][
@ -1633,7 +1634,8 @@ class CustomStreamWrapper:
if is_async_iterable(self.completion_stream):
async for chunk in self.completion_stream:
if chunk == "None" or chunk is None:
raise Exception
continue # skip None chunks
elif (
self.custom_llm_provider == "gemini"
and hasattr(chunk, "parts")
@ -1642,7 +1644,9 @@ class CustomStreamWrapper:
continue
# chunk_creator() does logging/stream chunk building. We need to let it know its being called in_async_func, so we don't double add chunks.
# __anext__ also calls async_success_handler, which does logging
print_verbose(f"PROCESSED ASYNC CHUNK PRE CHUNK CREATOR: {chunk}")
verbose_logger.debug(
f"PROCESSED ASYNC CHUNK PRE CHUNK CREATOR: {chunk}"
)
processed_chunk: Optional[ModelResponseStream] = self.chunk_creator(
chunk=chunk

View file

@ -362,6 +362,15 @@ def token_counter(
"""
from litellm.utils import convert_list_message_to_dict
#########################################################
# Flag to disable token counter
# We've gotten reports of this consuming CPU cycles,
# exposing this flag to allow users to disable
# it to confirm if this is indeed the issue
#########################################################
if litellm.disable_token_counter is True:
return 0
verbose_logger.debug(
f"messages in token_counter: {messages}, text in token_counter: {text}"
)

View file

@ -18,7 +18,9 @@ from litellm.litellm_core_utils.prompt_templates.factory import anthropic_messag
from litellm.llms.base_llm.base_utils import type_to_response_format_param
from litellm.llms.base_llm.chat.transformation import BaseConfig, BaseLLMException
from litellm.types.llms.anthropic import (
AllAnthropicMessageValues,
AllAnthropicToolsValues,
AnthropicCodeExecutionTool,
AnthropicComputerTool,
AnthropicHostedTools,
AnthropicInputSchema,
@ -530,6 +532,40 @@ class AnthropicConfig(AnthropicModelInfo, BaseConfig):
return anthropic_system_message_list
def add_code_execution_tool(
self,
messages: List[AllAnthropicMessageValues],
tools: List[Union[AllAnthropicToolsValues, Dict]],
) -> List[Union[AllAnthropicToolsValues, Dict]]:
"""if 'container_upload' in messages, add code_execution tool"""
add_code_execution_tool = False
for message in messages:
message_content = message.get("content", None)
if message_content and isinstance(message_content, list):
for content in message_content:
content_type = content.get("type", None)
if content_type == "container_upload":
add_code_execution_tool = True
break
if add_code_execution_tool:
## check if code_execution tool is already in tools
for tool in tools:
tool_type = tool.get("type", None)
if (
tool_type
and isinstance(tool_type, str)
and tool_type.startswith("code_execution")
):
return tools
tools.append(
AnthropicCodeExecutionTool(
name="code_execution",
type="code_execution_20250522",
)
)
return tools
def transform_request(
self,
model: str,
@ -579,6 +615,18 @@ class AnthropicConfig(AnthropicModelInfo, BaseConfig):
message="{}\nReceived Messages={}".format(str(e), messages),
) # don't use verbose_logger.exception, if exception is raised
## Add code_execution tool if container_upload is in messages
_tools = (
cast(
Optional[List[Union[AllAnthropicToolsValues, Dict]]],
optional_params.get("tools"),
)
or []
)
tools = self.add_code_execution_tool(messages=anthropic_messages, tools=_tools)
if len(tools) > 1:
optional_params["tools"] = tools
## Load Config
config = litellm.AnthropicConfig.get_config()
for k, v in config.items():

View file

@ -7,6 +7,9 @@ from typing import Dict, List, Optional, Union
import httpx
import litellm
from litellm.litellm_core_utils.prompt_templates.common_utils import (
get_file_ids_from_messages,
)
from litellm.llms.base_llm.base_utils import BaseLLMModelInfo
from litellm.llms.base_llm.chat.transformation import BaseLLMException
from litellm.secret_managers.main import get_secret_str
@ -42,6 +45,13 @@ class AnthropicModelInfo(BaseLLMModelInfo):
return False
def is_file_id_used(self, messages: List[AllMessageValues]) -> bool:
"""
Return if {"source": {"type": "file", "file_id": ..}} in message content block
"""
file_ids = get_file_ids_from_messages(messages)
return len(file_ids) > 0
def is_computer_tool_used(
self, tools: Optional[List[AllAnthropicToolsValues]]
) -> bool:
@ -82,6 +92,7 @@ class AnthropicModelInfo(BaseLLMModelInfo):
computer_tool_used: bool = False,
prompt_caching_set: bool = False,
pdf_used: bool = False,
file_id_used: bool = False,
is_vertex_request: bool = False,
user_anthropic_beta_headers: Optional[List[str]] = None,
) -> dict:
@ -90,8 +101,11 @@ class AnthropicModelInfo(BaseLLMModelInfo):
betas.add("prompt-caching-2024-07-31")
if computer_tool_used:
betas.add("computer-use-2024-10-22")
if pdf_used:
betas.add("pdfs-2024-09-25")
# if pdf_used:
# betas.add("pdfs-2024-09-25")
if file_id_used:
betas.add("files-api-2025-04-14")
betas.add("code-execution-2025-05-22")
headers = {
"anthropic-version": anthropic_version or "2023-06-01",
"x-api-key": api_key,
@ -131,6 +145,7 @@ class AnthropicModelInfo(BaseLLMModelInfo):
prompt_caching_set = self.is_cache_control_set(messages=messages)
computer_tool_used = self.is_computer_tool_used(tools=tools)
pdf_used = self.is_pdf_used(messages=messages)
file_id_used = self.is_file_id_used(messages=messages)
user_anthropic_beta_headers = self._get_user_anthropic_beta_headers(
anthropic_beta_header=headers.get("anthropic-beta")
)
@ -139,6 +154,7 @@ class AnthropicModelInfo(BaseLLMModelInfo):
prompt_caching_set=prompt_caching_set,
pdf_used=pdf_used,
api_key=api_key,
file_id_used=file_id_used,
is_vertex_request=optional_params.get("is_vertex_request", False),
user_anthropic_beta_headers=user_anthropic_beta_headers,
)

View file

@ -140,7 +140,7 @@ def anthropic_messages_handler(
)
if anthropic_messages_provider_config is None:
raise ValueError(
f"Anthropic messages provider config not found for model: {model}"
f"Anthropic messages provider config not found for model: {model}, custom_llm_provider: {custom_llm_provider}"
)
if custom_llm_provider is None:
raise ValueError(

View file

@ -1,4 +1,4 @@
from typing import Any, AsyncIterator, Dict, List, Optional
from typing import Any, AsyncIterator, Dict, List, Optional, Tuple
import httpx
@ -50,7 +50,7 @@ class AnthropicMessagesConfig(BaseAnthropicMessagesConfig):
api_base = f"{api_base}/v1/messages"
return api_base
def validate_environment(
def validate_anthropic_messages_environment(
self,
headers: dict,
model: str,
@ -59,14 +59,14 @@ class AnthropicMessagesConfig(BaseAnthropicMessagesConfig):
litellm_params: dict,
api_key: Optional[str] = None,
api_base: Optional[str] = None,
) -> dict:
if "x-api-key" not in headers:
) -> Tuple[dict, Optional[str]]:
if "x-api-key" not in headers and api_key:
headers["x-api-key"] = api_key
if "anthropic-version" not in headers:
headers["anthropic-version"] = DEFAULT_ANTHROPIC_API_VERSION
if "content-type" not in headers:
headers["content-type"] = "application/json"
return headers
return headers, api_base
def transform_anthropic_messages_request(
self,

View file

@ -94,7 +94,7 @@ class AzureAudioTranscription(AzureChatCompletion):
additional_args={"complete_input_dict": data},
original_response=stringified_response,
)
hidden_params = {"model": "whisper-1", "custom_llm_provider": "azure"}
hidden_params = {"model": model, "custom_llm_provider": "azure"}
final_response: TranscriptionResponse = convert_to_model_response_object(response_object=stringified_response, model_response_object=model_response, hidden_params=hidden_params, response_type="audio_transcription") # type: ignore
return final_response
@ -174,7 +174,7 @@ class AzureAudioTranscription(AzureChatCompletion):
},
original_response=stringified_response,
)
hidden_params = {"model": "whisper-1", "custom_llm_provider": "azure"}
hidden_params = {"model": model, "custom_llm_provider": "azure"}
response = convert_to_model_response_object(
_response_headers=headers,
response_object=stringified_response,

View file

@ -18,7 +18,7 @@ else:
class BaseAnthropicMessagesConfig(ABC):
@abstractmethod
def validate_environment(
def validate_anthropic_messages_environment( # use different name because return type is different from base config's validate_environment
self,
headers: dict,
model: str,
@ -27,13 +27,17 @@ class BaseAnthropicMessagesConfig(ABC):
litellm_params: dict,
api_key: Optional[str] = None,
api_base: Optional[str] = None,
) -> dict:
) -> Tuple[dict, Optional[str]]:
"""
OPTIONAL
Validate the environment for the request
Returns:
- headers: dict
- api_base: Optional[str] - If the provider needs to update the api_base, return it here. Otherwise, return None.
"""
return headers
return headers, api_base
@abstractmethod
def get_complete_url(

View file

@ -336,6 +336,36 @@ class BaseAWSLLM:
return aws_region_name
def get_aws_region_name_for_non_llm_api_calls(
self,
aws_region_name: Optional[str] = None,
):
"""
Get the AWS region name for non-llm api calls.
LLM API calls check the model arn and end up using that as the region name.
For non-llm api calls eg. Guardrails, Vector Stores we just need to check the dynamic param or env vars.
"""
if aws_region_name is None:
# check env #
litellm_aws_region_name = get_secret("AWS_REGION_NAME", None)
if litellm_aws_region_name is not None and isinstance(
litellm_aws_region_name, str
):
aws_region_name = litellm_aws_region_name
standard_aws_region_name = get_secret("AWS_REGION", None)
if standard_aws_region_name is not None and isinstance(
standard_aws_region_name, str
):
aws_region_name = standard_aws_region_name
if aws_region_name is None:
aws_region_name = "us-west-2"
return aws_region_name
@tracer.wrap()
def _auth_with_web_identity_token(
self,
@ -527,6 +557,7 @@ class BaseAWSLLM:
api_base: Optional[str],
aws_bedrock_runtime_endpoint: Optional[str],
aws_region_name: str,
endpoint_type: Optional[Literal["runtime", "agent"]] = "runtime",
) -> Tuple[str, str]:
env_aws_bedrock_runtime_endpoint = get_secret("AWS_BEDROCK_RUNTIME_ENDPOINT")
if api_base is not None:
@ -540,7 +571,10 @@ class BaseAWSLLM:
):
endpoint_url = env_aws_bedrock_runtime_endpoint
else:
endpoint_url = f"https://bedrock-runtime.{aws_region_name}.amazonaws.com"
endpoint_url = self._select_default_endpoint_url(
endpoint_type=endpoint_type,
aws_region_name=aws_region_name,
)
# Determine proxy_endpoint_url
if env_aws_bedrock_runtime_endpoint and isinstance(
@ -556,6 +590,19 @@ class BaseAWSLLM:
return endpoint_url, proxy_endpoint_url
def _select_default_endpoint_url(
self, endpoint_type: Optional[Literal["runtime", "agent"]], aws_region_name: str
) -> str:
"""
Select the default endpoint url based on the endpoint type
Default endpoint url is https://bedrock-runtime.{aws_region_name}.amazonaws.com
"""
if endpoint_type == "agent":
return f"https://bedrock-agent-runtime.{aws_region_name}.amazonaws.com"
else:
return f"https://bedrock-runtime.{aws_region_name}.amazonaws.com"
def _get_boto_credentials_from_optional_params(
self, optional_params: dict, model: Optional[str] = None
) -> Boto3CredentialsInfo:

View file

@ -0,0 +1,527 @@
"""
Transformation for Bedrock Invoke Agent
https://docs.aws.amazon.com/bedrock/latest/APIReference/API_agent-runtime_InvokeAgent.html
"""
import base64
import json
import uuid
from typing import TYPE_CHECKING, Any, Dict, List, Optional, Tuple, Union
import httpx
from litellm._logging import verbose_logger
from litellm.litellm_core_utils.prompt_templates.common_utils import (
convert_content_list_to_str,
)
from litellm.llms.base_llm.chat.transformation import BaseConfig, BaseLLMException
from litellm.llms.bedrock.base_aws_llm import BaseAWSLLM
from litellm.llms.bedrock.common_utils import BedrockError
from litellm.types.llms.bedrock_invoke_agents import (
InvokeAgentChunkPayload,
InvokeAgentEvent,
InvokeAgentEventHeaders,
InvokeAgentEventList,
InvokeAgentTrace,
InvokeAgentTracePayload,
InvokeAgentUsage,
)
from litellm.types.llms.openai import AllMessageValues
from litellm.types.utils import Choices, Message, ModelResponse
if TYPE_CHECKING:
from litellm.litellm_core_utils.litellm_logging import Logging as _LiteLLMLoggingObj
LiteLLMLoggingObj = _LiteLLMLoggingObj
else:
LiteLLMLoggingObj = Any
class AmazonInvokeAgentConfig(BaseConfig, BaseAWSLLM):
def __init__(self, **kwargs):
BaseConfig.__init__(self, **kwargs)
BaseAWSLLM.__init__(self, **kwargs)
def get_supported_openai_params(self, model: str) -> List[str]:
"""
This is a base invoke agent model mapping. For Invoke Agent - define a bedrock provider specific config that extends this class.
Bedrock Invoke Agents has 0 OpenAI compatible params
As of May 29th, 2025 - they don't support streaming.
"""
return []
def map_openai_params(
self,
non_default_params: dict,
optional_params: dict,
model: str,
drop_params: bool,
) -> dict:
"""
This is a base invoke agent model mapping. For Invoke Agent - define a bedrock provider specific config that extends this class.
"""
return optional_params
def get_complete_url(
self,
api_base: Optional[str],
api_key: Optional[str],
model: str,
optional_params: dict,
litellm_params: dict,
stream: Optional[bool] = None,
) -> str:
"""
Get the complete url for the request
"""
### SET RUNTIME ENDPOINT ###
aws_bedrock_runtime_endpoint = optional_params.get(
"aws_bedrock_runtime_endpoint", None
) # https://bedrock-runtime.{region_name}.amazonaws.com
endpoint_url, _ = self.get_runtime_endpoint(
api_base=api_base,
aws_bedrock_runtime_endpoint=aws_bedrock_runtime_endpoint,
aws_region_name=self._get_aws_region_name(
optional_params=optional_params, model=model
),
endpoint_type="agent",
)
agent_id, agent_alias_id = self._get_agent_id_and_alias_id(model)
session_id = self._get_session_id(optional_params)
endpoint_url = f"{endpoint_url}/agents/{agent_id}/agentAliases/{agent_alias_id}/sessions/{session_id}/text"
return endpoint_url
def sign_request(
self,
headers: dict,
optional_params: dict,
request_data: dict,
api_base: str,
model: Optional[str] = None,
stream: Optional[bool] = None,
fake_stream: Optional[bool] = None,
) -> Tuple[dict, Optional[bytes]]:
return self._sign_request(
service_name="bedrock",
headers=headers,
optional_params=optional_params,
request_data=request_data,
api_base=api_base,
model=model,
stream=stream,
fake_stream=fake_stream,
)
def _get_agent_id_and_alias_id(self, model: str) -> tuple[str, str]:
"""
model = "agent/L1RT58GYRW/MFPSBCXYTW"
agent_id = "L1RT58GYRW"
agent_alias_id = "MFPSBCXYTW"
"""
# Split the model string by '/' and extract components
parts = model.split("/")
if len(parts) != 3 or parts[0] != "agent":
raise ValueError(
"Invalid model format. Expected format: 'model=agent/AGENT_ID/ALIAS_ID'"
)
return parts[1], parts[2] # Return (agent_id, agent_alias_id)
def _get_session_id(self, optional_params: dict) -> str:
""" """
return optional_params.get("sessionID", None) or str(uuid.uuid4())
def transform_request(
self,
model: str,
messages: List[AllMessageValues],
optional_params: dict,
litellm_params: dict,
headers: dict,
) -> dict:
# use the last message content as the query
query: str = convert_content_list_to_str(messages[-1])
return {
"inputText": query,
"enableTrace": True,
**optional_params,
}
def _parse_aws_event_stream(self, raw_content: bytes) -> InvokeAgentEventList:
"""
Parse AWS event stream format using boto3/botocore's built-in parser.
This is the same approach used in the existing AWSEventStreamDecoder.
"""
try:
from botocore.eventstream import EventStreamBuffer
from botocore.parsers import EventStreamJSONParser
except ImportError:
raise ImportError("boto3/botocore is required for AWS event stream parsing")
events: InvokeAgentEventList = []
parser = EventStreamJSONParser()
event_stream_buffer = EventStreamBuffer()
# Add the entire response to the buffer
event_stream_buffer.add_data(raw_content)
# Process all events in the buffer
for event in event_stream_buffer:
try:
headers = self._extract_headers_from_event(event)
event_type = headers.get("event_type", "")
if event_type == "chunk":
# Handle chunk events specially - they contain decoded content, not JSON
message = self._parse_message_from_event(event, parser)
parsed_event: InvokeAgentEvent = InvokeAgentEvent()
if message:
# For chunk events, create a payload with the decoded content
parsed_event = {
"headers": headers,
"payload": {
"bytes": base64.b64encode(
message.encode("utf-8")
).decode("utf-8")
}, # Re-encode for consistency
}
events.append(parsed_event)
elif event_type == "trace":
# Handle trace events normally - they contain JSON
message = self._parse_message_from_event(event, parser)
if message:
try:
event_data = json.loads(message)
parsed_event = {
"headers": headers,
"payload": event_data,
}
events.append(parsed_event)
except json.JSONDecodeError as e:
verbose_logger.warning(
f"Failed to parse trace event JSON: {e}"
)
else:
verbose_logger.debug(f"Unknown event type: {event_type}")
except Exception as e:
verbose_logger.error(f"Error processing event: {e}")
continue
return events
def _parse_message_from_event(self, event, parser) -> Optional[str]:
"""Extract message content from an AWS event, adapted from AWSEventStreamDecoder."""
try:
response_dict = event.to_response_dict()
verbose_logger.debug(f"Response dict: {response_dict}")
# Use the same response shape parsing as the existing decoder
parsed_response = parser.parse(
response_dict, self._get_response_stream_shape()
)
verbose_logger.debug(f"Parsed response: {parsed_response}")
if response_dict["status_code"] != 200:
decoded_body = response_dict["body"].decode()
if isinstance(decoded_body, dict):
error_message = decoded_body.get("message")
elif isinstance(decoded_body, str):
error_message = decoded_body
else:
error_message = ""
exception_status = response_dict["headers"].get(":exception-type")
error_message = exception_status + " " + error_message
raise BedrockError(
status_code=response_dict["status_code"],
message=(
json.dumps(error_message)
if isinstance(error_message, dict)
else error_message
),
)
if "chunk" in parsed_response:
chunk = parsed_response.get("chunk")
if not chunk:
return None
return chunk.get("bytes").decode()
else:
chunk = response_dict.get("body")
if not chunk:
return None
return chunk.decode()
except Exception as e:
verbose_logger.debug(f"Error parsing message from event: {e}")
return None
def _extract_headers_from_event(self, event) -> InvokeAgentEventHeaders:
"""Extract headers from an AWS event for categorization."""
try:
response_dict = event.to_response_dict()
headers = response_dict.get("headers", {})
# Extract the event-type and content-type headers that we care about
return InvokeAgentEventHeaders(
event_type=headers.get(":event-type", ""),
content_type=headers.get(":content-type", ""),
message_type=headers.get(":message-type", ""),
)
except Exception as e:
verbose_logger.debug(f"Error extracting headers: {e}")
return InvokeAgentEventHeaders(
event_type="", content_type="", message_type=""
)
def _get_response_stream_shape(self):
"""Get the response stream shape for parsing, reusing existing logic."""
try:
# Try to reuse the cached shape from the existing decoder
from litellm.llms.bedrock.chat.invoke_handler import (
get_response_stream_shape,
)
return get_response_stream_shape()
except ImportError:
# Fallback: create our own shape
try:
from botocore.loaders import Loader
from botocore.model import ServiceModel
loader = Loader()
bedrock_service_dict = loader.load_service_model(
"bedrock-runtime", "service-2"
)
bedrock_service_model = ServiceModel(bedrock_service_dict)
return bedrock_service_model.shape_for("ResponseStream")
except Exception as e:
verbose_logger.warning(f"Could not load response stream shape: {e}")
return None
def _extract_response_content(self, events: InvokeAgentEventList) -> str:
"""Extract the final response content from parsed events."""
response_parts = []
for event in events:
headers = event.get("headers", {})
payload = event.get("payload")
event_type = headers.get(
"event_type"
) # Note: using event_type not event-type
if event_type == "chunk" and payload:
# Extract base64 encoded content from chunk events
chunk_payload: InvokeAgentChunkPayload = payload # type: ignore
encoded_bytes = chunk_payload.get("bytes", "")
if encoded_bytes:
try:
decoded_content = base64.b64decode(encoded_bytes).decode(
"utf-8"
)
response_parts.append(decoded_content)
except Exception as e:
verbose_logger.warning(f"Failed to decode chunk content: {e}")
return "".join(response_parts)
def _extract_usage_info(self, events: InvokeAgentEventList) -> InvokeAgentUsage:
"""Extract token usage information from trace events."""
usage_info = InvokeAgentUsage(
inputTokens=0,
outputTokens=0,
model=None,
)
response_model: Optional[str] = None
for event in events:
if not self._is_trace_event(event):
continue
trace_data = self._get_trace_data(event)
if not trace_data:
continue
verbose_logger.debug(f"Trace event: {trace_data}")
# Extract usage from pre-processing trace
self._extract_and_update_preprocessing_usage(
trace_data=trace_data,
usage_info=usage_info,
)
# Extract model from orchestration trace
if response_model is None:
response_model = self._extract_orchestration_model(trace_data)
usage_info["model"] = response_model
return usage_info
def _is_trace_event(self, event: InvokeAgentEvent) -> bool:
"""Check if the event is a trace event."""
headers = event.get("headers", {})
event_type = headers.get("event_type")
payload = event.get("payload")
return event_type == "trace" and payload is not None
def _get_trace_data(self, event: InvokeAgentEvent) -> Optional[InvokeAgentTrace]:
"""Extract trace data from a trace event."""
payload = event.get("payload")
if not payload:
return None
trace_payload: InvokeAgentTracePayload = payload # type: ignore
return trace_payload.get("trace", {})
def _extract_and_update_preprocessing_usage(
self, trace_data: InvokeAgentTrace, usage_info: InvokeAgentUsage
) -> None:
"""Extract usage information from preprocessing trace."""
pre_processing = trace_data.get("preProcessingTrace", {})
if not pre_processing:
return
model_output = pre_processing.get("modelInvocationOutput", {})
if not model_output:
return
metadata = model_output.get("metadata", {})
if not metadata:
return
usage: Optional[Union[InvokeAgentUsage, Dict]] = metadata.get("usage", {})
if not usage:
return
usage_info["inputTokens"] += usage.get("inputTokens", 0)
usage_info["outputTokens"] += usage.get("outputTokens", 0)
def _extract_orchestration_model(
self, trace_data: InvokeAgentTrace
) -> Optional[str]:
"""Extract model information from orchestration trace."""
orchestration_trace = trace_data.get("orchestrationTrace", {})
if not orchestration_trace:
return None
model_invocation = orchestration_trace.get("modelInvocationInput", {})
if not model_invocation:
return None
return model_invocation.get("foundationModel")
def _build_model_response(
self,
content: str,
model: str,
usage_info: InvokeAgentUsage,
model_response: ModelResponse,
) -> ModelResponse:
"""Build the final ModelResponse object."""
# Create the message content
message = Message(content=content, role="assistant")
# Create choices
choice = Choices(finish_reason="stop", index=0, message=message)
# Update model response
model_response.choices = [choice]
model_response.model = usage_info.get("model", model)
# Add usage information if available
if usage_info:
from litellm.types.utils import Usage
usage = Usage(
prompt_tokens=usage_info.get("inputTokens", 0),
completion_tokens=usage_info.get("outputTokens", 0),
total_tokens=usage_info.get("inputTokens", 0)
+ usage_info.get("outputTokens", 0),
)
setattr(model_response, "usage", usage)
return model_response
def transform_response(
self,
model: str,
raw_response: httpx.Response,
model_response: ModelResponse,
logging_obj: LiteLLMLoggingObj,
request_data: dict,
messages: List[AllMessageValues],
optional_params: dict,
litellm_params: dict,
encoding: Any,
api_key: Optional[str] = None,
json_mode: Optional[bool] = None,
) -> ModelResponse:
try:
# Get the raw binary content
raw_content = raw_response.content
verbose_logger.debug(
f"Processing {len(raw_content)} bytes of AWS event stream data"
)
# Parse the AWS event stream format
events = self._parse_aws_event_stream(raw_content)
verbose_logger.debug(f"Parsed {len(events)} events from stream")
# Extract response content from chunk events
content = self._extract_response_content(events)
# Extract usage information from trace events
usage_info = self._extract_usage_info(events)
# Build and return the model response
return self._build_model_response(
content=content,
model=model,
usage_info=usage_info,
model_response=model_response,
)
except Exception as e:
verbose_logger.error(
f"Error processing Bedrock Invoke Agent response: {str(e)}"
)
raise BedrockError(
message=f"Error processing response: {str(e)}",
status_code=raw_response.status_code,
)
def validate_environment(
self,
headers: dict,
model: str,
messages: List[AllMessageValues],
optional_params: dict,
litellm_params: dict,
api_key: Optional[str] = None,
api_base: Optional[str] = None,
) -> dict:
return headers
def get_error_class(
self, error_message: str, status_code: int, headers: Union[dict, httpx.Headers]
) -> BaseLLMException:
return BedrockError(status_code=status_code, message=error_message)
def should_fake_stream(
self,
model: Optional[str],
stream: Optional[bool],
custom_llm_provider: Optional[str] = None,
) -> bool:
return True

View file

@ -402,7 +402,9 @@ class BedrockModelInfo(BaseLLMModelInfo):
return ["us", "eu", "apac"]
@staticmethod
def get_bedrock_route(model: str) -> Literal["converse", "invoke", "converse_like"]:
def get_bedrock_route(
model: str,
) -> Literal["converse", "invoke", "converse_like", "agent"]:
"""
Get the bedrock route for the given model.
"""
@ -414,6 +416,8 @@ class BedrockModelInfo(BaseLLMModelInfo):
return "converse_like"
elif "converse/" in model:
return "converse"
elif "agent/" in model:
return "agent"
elif (
base_model in litellm.bedrock_converse_models
or alt_model in litellm.bedrock_converse_models

View file

@ -38,6 +38,18 @@ class AmazonAnthropicClaude3MessagesConfig(
BaseAnthropicMessagesConfig.__init__(self, **kwargs)
AmazonInvokeConfig.__init__(self, **kwargs)
def validate_anthropic_messages_environment(
self,
headers: dict,
model: str,
messages: List[Any],
optional_params: dict,
litellm_params: dict,
api_key: Optional[str] = None,
api_base: Optional[str] = None,
) -> Tuple[dict, Optional[str]]:
return headers, api_base
def sign_request(
self,
headers: dict,
@ -59,18 +71,6 @@ class AmazonAnthropicClaude3MessagesConfig(
fake_stream=fake_stream,
)
def validate_environment(
self,
headers: dict,
model: str,
messages: List[Any],
optional_params: dict,
litellm_params: dict,
api_key: Optional[str] = None,
api_base: Optional[str] = None,
) -> dict:
return headers
def get_complete_url(
self,
api_base: Optional[str],

View file

@ -5,6 +5,7 @@ from typing import Callable, Dict, Union
import aiohttp
import aiohttp.client_exceptions
import aiohttp.http_exceptions
import httpx
from aiohttp.client import ClientResponse, ClientSession
@ -81,13 +82,20 @@ class AiohttpResponseStream(httpx.AsyncByteStream):
self.CHUNK_SIZE
):
yield chunk
except aiohttp.ClientPayloadError as e:
except (
aiohttp.ClientPayloadError,
aiohttp.client_exceptions.ClientPayloadError,
) as e:
# Handle incomplete transfers more gracefully
# Log the error but don't re-raise if we've already yielded some data
verbose_logger.debug(f"Transfer incomplete, but continuing: {e}")
# If the error is due to incomplete transfer encoding, we can still
# return what we've received so far, similar to how httpx handles it
return
except aiohttp.http_exceptions.TransferEncodingError as e:
# Handle transfer encoding errors gracefully
verbose_logger.debug(f"Transfer encoding error, but continuing: {e}")
return
except Exception:
# For other exceptions, use the normal mapping
with map_aiohttp_exceptions():
@ -203,7 +211,6 @@ class LiteLLMAiohttpTransport(AiohttpTransport):
data=data,
allow_redirects=False,
auto_decompress=False,
compress=False,
timeout=ClientTimeout(
sock_connect=timeout.get("connect"),
sock_read=timeout.get("read"),

View file

@ -505,20 +505,30 @@ class AsyncHTTPHandler:
@staticmethod
def _should_use_aiohttp_transport() -> bool:
"""
This is feature flagged for now and is opt in as we roll out to all users.
AiohttpTransport is the default transport for litellm.
Controlled by either
- litellm.use_aiohttp_transport or os.getenv("USE_AIOHTTP_TRANSPORT") = "True"
Httpx can be used by the following
- litellm.disable_aiohttp_transport = True
- os.getenv("DISABLE_AIOHTTP_TRANSPORT") = "True"
"""
import os
from litellm.secret_managers.main import str_to_bool
#########################################################
# Check if user disabled aiohttp transport
########################################################
if (
str_to_bool(os.getenv("USE_AIOHTTP_TRANSPORT", "False"))
or litellm.use_aiohttp_transport
litellm.disable_aiohttp_transport is True
or str_to_bool(os.getenv("DISABLE_AIOHTTP_TRANSPORT", "False")) is True
):
verbose_logger.debug("Using AiohttpTransport...")
return True
return False
return False
#########################################################
# Default: Use AiohttpTransport
########################################################
verbose_logger.debug("Using AiohttpTransport...")
return True
@staticmethod
def _create_aiohttp_transport(

View file

@ -271,7 +271,6 @@ class BaseLLMHTTPHandler:
):
json_mode: bool = optional_params.pop("json_mode", False)
extra_body: Optional[dict] = optional_params.pop("extra_body", None)
fake_stream = fake_stream or optional_params.pop("fake_stream", False)
provider_config = (
provider_config
@ -284,6 +283,14 @@ class BaseLLMHTTPHandler:
f"Provider config not found for model: {model} and provider: {custom_llm_provider}"
)
fake_stream = (
fake_stream
or optional_params.pop("fake_stream", False)
or provider_config.should_fake_stream(
model=model, custom_llm_provider=custom_llm_provider, stream=stream
)
)
# get config from model, custom llm provider
headers = provider_config.validate_environment(
api_key=api_key,
@ -1090,7 +1097,10 @@ class BaseLLMHTTPHandler:
if provider_specific_header
else {}
)
headers = anthropic_messages_provider_config.validate_environment(
(
headers,
api_base,
) = anthropic_messages_provider_config.validate_anthropic_messages_environment(
headers=extra_headers or {},
model=model,
messages=messages,

View file

@ -0,0 +1,80 @@
"""
Support for OpenAI's `/v1/chat/completions` endpoint.
Calls done in OpenAI/openai.py as DataRobot is openai-compatible.
"""
from typing import Optional, Tuple
from litellm.secret_managers.main import get_secret_str
from ...openai_like.chat.transformation import OpenAILikeChatConfig
class DataRobotConfig(OpenAILikeChatConfig):
@staticmethod
def _resolve_api_key(api_key: Optional[str] = None) -> str:
"""Attempt to ensure that the API key is set, preferring the user-provided key
over the secret manager key (``DATAROBOT_API_TOKEN``).
If both are None, a fake API key is returned for testing.
"""
return api_key or get_secret_str("DATAROBOT_API_TOKEN") or "fake-api-key"
@staticmethod
def _resolve_api_base(api_base: Optional[str] = None) -> Optional[str]:
"""Attempt to ensure that the API base is set, preferring the user-provided key
over the secret manager key (``DATAROBOT_ENDPOINT``).
If both are None, a default Llamafile server URL is returned.
See: https://github.com/Mozilla-Ocho/llamafile/blob/bd1bbe9aabb1ee12dbdcafa8936db443c571eb9d/README.md#L61
"""
api_base = api_base or get_secret_str("DATAROBOT_ENDPOINT")
if api_base is None:
api_base = "https://app.datarobot.com"
# If the api_base is a deployment URL, we do not append the chat completions path
if "api/v2/deployments" not in api_base:
# If the api_base is not a deployment URL, we need to append the chat completions path
if "api/v2/genai/llmgw/chat/completions" not in api_base:
api_base += "/api/v2/genai/llmgw/chat/completions"
# Ensure the url ends with a trailing slash
if not api_base.endswith("/"):
api_base += "/"
return api_base # type: ignore
def _get_openai_compatible_provider_info(
self,
api_base: Optional[str],
api_key: Optional[str]
) -> Tuple[Optional[str], Optional[str]]:
"""Attempts to ensure that the API base and key are set, preferring user-provided values,
before falling back to secret manager values (``DATAROBOT_ENDPOINT`` and ``DATAROBOT_API_TOKEN``
respectively).
If an API key cannot be resolved via either method, a fake key is returned.
"""
api_base = DataRobotConfig._resolve_api_base(api_base)
dynamic_api_key = DataRobotConfig._resolve_api_key(api_key)
return api_base, dynamic_api_key
def get_complete_url(
self,
api_base: Optional[str],
api_key: Optional[str],
model: str,
optional_params: dict,
litellm_params: dict,
stream: Optional[bool] = None,
) -> str:
"""
Get the complete URL for the API call. Datarobot's API base is set to
the complete value, so it does not need to be updated to additionally add
chat completions.
Returns:
str: The complete URL for the API call.
"""
return str(api_base) # type: ignore

View file

@ -186,11 +186,24 @@ class FireworksAIConfig(OpenAIGPTConfig):
"""
Add 'transform=inline' to the url of the image_url
"""
from litellm.litellm_core_utils.prompt_templates.common_utils import (
filter_value_from_dict,
migrate_file_to_image_url,
)
disable_add_transform_inline_image_block = cast(
Optional[bool],
litellm_params.get("disable_add_transform_inline_image_block")
or litellm.disable_add_transform_inline_image_block,
)
## For any 'file' message type with pdf content, move to 'image_url' message type
for message in messages:
if message["role"] == "user":
_message_content = message.get("content")
if _message_content is not None and isinstance(_message_content, list):
for idx, content in enumerate(_message_content):
if content["type"] == "file":
_message_content[idx] = migrate_file_to_image_url(content)
for message in messages:
if message["role"] == "user":
_message_content = message.get("content")
@ -202,6 +215,8 @@ class FireworksAIConfig(OpenAIGPTConfig):
model=model,
disable_add_transform_inline_image_block=disable_add_transform_inline_image_block,
)
filter_value_from_dict(cast(dict, message), "cache_control")
return messages
def get_provider_info(self, model: str) -> ProviderSpecificModelInfo:

View file

@ -6,7 +6,7 @@ from litellm.litellm_core_utils.prompt_templates.factory import (
convert_to_anthropic_image_obj,
)
from litellm.types.llms.openai import AllMessageValues
from litellm.types.llms.vertex_ai import ContentType, PartType
from litellm.types.llms.vertex_ai import ContentType, PartType, SpeechConfig, VoiceConfig, PrebuiltVoiceConfig
from litellm.utils import supports_reasoning
from ...vertex_ai.gemini.transformation import _gemini_convert_messages_with_history
@ -67,6 +67,9 @@ class GoogleAIStudioGeminiConfig(VertexGeminiConfig):
def get_config(cls):
return super().get_config()
def is_model_gemini_audio_model(self, model: str) -> bool:
return "tts" in model
def get_supported_openai_params(self, model: str) -> List[str]:
supported_params = [
"temperature",
@ -84,10 +87,13 @@ class GoogleAIStudioGeminiConfig(VertexGeminiConfig):
"frequency_penalty",
"modalities",
"parallel_tool_calls",
"web_search_options",
]
if supports_reasoning(model):
supported_params.append("reasoning_effort")
supported_params.append("thinking")
if self.is_model_gemini_audio_model(model):
supported_params.append("audio")
return supported_params
def map_openai_params(
@ -97,6 +103,40 @@ class GoogleAIStudioGeminiConfig(VertexGeminiConfig):
model: str,
drop_params: bool,
) -> Dict:
# Handle audio parameter for TTS models
if self.is_model_gemini_audio_model(model):
for param, value in non_default_params.items():
if param == "audio" and isinstance(value, dict):
# Validate audio format - Gemini TTS only supports pcm16
audio_format = value.get("format")
if audio_format is not None and audio_format != "pcm16":
raise ValueError(
f"Unsupported audio format for Gemini TTS models: {audio_format}. "
f"Gemini TTS models only support 'pcm16' format as they return audio data in L16 PCM format. "
f"Please set audio format to 'pcm16'."
)
# Map OpenAI audio parameter to Gemini speech config
speech_config: SpeechConfig = {}
if "voice" in value:
prebuilt_voice_config: PrebuiltVoiceConfig = {
"voiceName": value["voice"]
}
voice_config: VoiceConfig = {
"prebuiltVoiceConfig": prebuilt_voice_config
}
speech_config["voiceConfig"] = voice_config
if speech_config:
optional_params["speechConfig"] = speech_config
# Ensure audio modality is set
if "responseModalities" not in optional_params:
optional_params["responseModalities"] = ["AUDIO"]
elif "AUDIO" not in optional_params["responseModalities"]:
optional_params["responseModalities"].append("AUDIO")
if litellm.vertex_ai_safety_settings is not None:
optional_params["safety_settings"] = litellm.vertex_ai_safety_settings
return super().map_openai_params(

View file

@ -173,7 +173,7 @@ class OllamaConfig(BaseConfig):
if param == "top_p":
optional_params["top_p"] = value
if param == "frequency_penalty":
optional_params["repeat_penalty"] = value
optional_params["frequency_penalty"] = value
if param == "stop":
optional_params["stop"] = value
if param == "response_format" and isinstance(value, dict):

View file

@ -155,7 +155,7 @@ class OpenAIAudioTranscription(OpenAIChatCompletion):
additional_args={"complete_input_dict": data},
original_response=stringified_response,
)
hidden_params = {"model": "whisper-1", "custom_llm_provider": "openai"}
hidden_params = {"model": model, "custom_llm_provider": "openai"}
final_response: TranscriptionResponse = convert_to_model_response_object(response_object=stringified_response, model_response_object=model_response, hidden_params=hidden_params, response_type="audio_transcription") # type: ignore
return final_response
@ -210,7 +210,9 @@ class OpenAIAudioTranscription(OpenAIChatCompletion):
additional_args={"complete_input_dict": data},
original_response=stringified_response,
)
hidden_params = {"model": "whisper-1", "custom_llm_provider": "openai"}
# Extract the actual model from data instead of hardcoding "whisper-1"
actual_model = data.get("model", "whisper-1")
hidden_params = {"model": actual_model, "custom_llm_provider": "openai"}
return convert_to_model_response_object(response_object=stringified_response, model_response_object=model_response, hidden_params=hidden_params, response_type="audio_transcription") # type: ignore
except Exception as e:
## LOGGING

View file

@ -43,7 +43,7 @@ class VertexAIBatchPrediction(VertexLLM):
custom_llm_provider="vertex_ai",
)
default_api_base = self.create_vertex_url(
default_api_base = self.create_vertex_batch_url(
vertex_location=vertex_location or "us-central1",
vertex_project=vertex_project or project_id,
)
@ -117,7 +117,7 @@ class VertexAIBatchPrediction(VertexLLM):
)
return vertex_batch_response
def create_vertex_url(
def create_vertex_batch_url(
self,
vertex_location: str,
vertex_project: str,
@ -145,7 +145,7 @@ class VertexAIBatchPrediction(VertexLLM):
custom_llm_provider="vertex_ai",
)
default_api_base = self.create_vertex_url(
default_api_base = self.create_vertex_batch_url(
vertex_location=vertex_location or "us-central1",
vertex_project=vertex_project or project_id,
)

View file

@ -2,6 +2,7 @@
## httpx client for vertex ai calls
## Initial implementation - covers gemini + image gen calls
import json
import time
import uuid
from copy import deepcopy
from functools import partial
@ -61,6 +62,7 @@ from litellm.types.llms.vertex_ai import (
UsageMetadata,
)
from litellm.types.utils import (
ChatCompletionAudioResponse,
ChatCompletionTokenLogprob,
ChoiceLogprobs,
CompletionTokensDetailsWrapper,
@ -69,7 +71,7 @@ from litellm.types.utils import (
TopLogprob,
Usage,
)
from litellm.utils import CustomStreamWrapper, ModelResponse, supports_reasoning
from litellm.utils import CustomStreamWrapper, ModelResponse, is_base64_encoded, supports_reasoning
from ....utils import _remove_additional_properties, _remove_strict_from_schema
from ..common_utils import VertexAIError, _build_vertex_schema
@ -220,6 +222,7 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
"top_logprobs",
"modalities",
"parallel_tool_calls",
"web_search_options",
]
if supports_reasoning(model):
supported_params.append("reasoning_effort")
@ -251,6 +254,14 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
status_code=400,
)
def _map_web_search_options(self, value: dict) -> Tools:
"""
Base Case: empty dict
Google doesn't support user_location or search_context_size params
"""
return Tools(googleSearch={})
def _map_function(self, value: List[dict]) -> List[Tools]:
gtool_func_declarations = []
googleSearch: Optional[dict] = None
@ -445,6 +456,19 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
response_modalities.append("MODALITY_UNSPECIFIED")
return response_modalities
def validate_parallel_tool_calls(self, value: bool, non_default_params: dict):
tools = non_default_params.get("tools", non_default_params.get("functions"))
num_function_declarations = len(tools) if isinstance(tools, list) else 0
if num_function_declarations > 1:
raise litellm.utils.UnsupportedParamsError(
message=(
"`parallel_tool_calls=False` is not supported by Gemini when multiple tools are "
"provided. Specify a single tool, or set "
"`parallel_tool_calls=True`. If you want to drop this param, set `litellm.drop_params = True` or pass in `(.., drop_params=True)` in the requst - https://docs.litellm.ai/docs/completion/drop_params"
),
status_code=400,
)
def map_openai_params(
self,
non_default_params: Dict,
@ -487,7 +511,9 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
and isinstance(value, list)
and value
):
optional_params["tools"] = self._map_function(value=value)
optional_params = self._add_tools_to_optional_params(
optional_params, self._map_function(value=value)
)
elif param == "tool_choice" and (
isinstance(value, str) or isinstance(value, dict)
):
@ -500,21 +526,7 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
if value is False and not (
drop_params or litellm.drop_params
): # if drop params is True, then we should just ignore this
tools = non_default_params.get(
"tools", non_default_params.get("functions")
)
num_function_declarations = (
len(tools) if isinstance(tools, list) else 0
)
if num_function_declarations > 1:
raise litellm.utils.UnsupportedParamsError(
message=(
"`parallel_tool_calls=False` is not supported when multiple tools are "
"provided for Gemini. Specify a single tool, or set "
"`parallel_tool_calls=True`. If you want to drop this param, set `litellm.drop_params = True` or pass in `(.., drop_params=True)` in the requst - https://docs.litellm.ai/docs/completion/drop_params"
),
status_code=400,
)
self.validate_parallel_tool_calls(value, non_default_params)
else:
optional_params["parallel_tool_calls"] = value
elif param == "seed":
@ -532,7 +544,11 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
elif param == "modalities" and isinstance(value, list):
response_modalities = self.map_response_modalities(value)
optional_params["responseModalities"] = response_modalities
elif param == "web_search_options" and value and isinstance(value, dict):
_tools = self._map_web_search_options(value)
optional_params = self._add_tools_to_optional_params(
optional_params, [_tools]
)
if litellm.vertex_ai_safety_settings is not None:
optional_params["safety_settings"] = litellm.vertex_ai_safety_settings
return optional_params
@ -662,14 +678,30 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
) -> Tuple[Optional[str], Optional[str]]:
content_str: Optional[str] = None
reasoning_content_str: Optional[str] = None
for part in parts:
_content_str = ""
if "text" in part:
_content_str += part["text"]
elif "inlineData" in part: # base64 encoded image
_content_str += "data:{};base64,{}".format(
part["inlineData"]["mimeType"], part["inlineData"]["data"]
)
text_content = part["text"]
# Check if text content is audio data URI - if so, exclude from text content
if text_content.startswith("data:audio") and ";base64," in text_content:
try:
if is_base64_encoded(text_content):
media_type, _ = text_content.split("data:")[1].split(";base64,")
if media_type.startswith("audio/"):
continue
except (ValueError, IndexError):
# If parsing fails, treat as regular text
pass
_content_str += text_content
elif "inlineData" in part:
mime_type = part["inlineData"]["mimeType"]
data = part["inlineData"]["data"]
# Check if inline data is audio - if so, exclude from text content
if mime_type.startswith("audio/"):
continue
_content_str += "data:{};base64,{}".format(mime_type, data)
if len(_content_str) > 0:
if part.get("thought") is True:
if reasoning_content_str is None:
@ -682,6 +714,47 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
return content_str, reasoning_content_str
def _extract_audio_response_from_parts(
self, parts: List[HttpxPartType]
) -> Optional[ChatCompletionAudioResponse]:
"""Extract audio response from parts if present"""
for part in parts:
if "text" in part:
text_content = part["text"]
# Check if text content contains audio data URI
if text_content.startswith("data:audio") and ";base64," in text_content:
try:
if is_base64_encoded(text_content):
media_type, audio_data = text_content.split("data:")[1].split(";base64,")
if media_type.startswith("audio/"):
expires_at = int(time.time()) + (24 * 60 * 60)
transcript = "" # Gemini doesn't provide transcript
return ChatCompletionAudioResponse(
data=audio_data,
expires_at=expires_at,
transcript=transcript
)
except (ValueError, IndexError):
pass
elif "inlineData" in part:
mime_type = part["inlineData"]["mimeType"]
data = part["inlineData"]["data"]
if mime_type.startswith("audio/"):
expires_at = int(time.time()) + (24 * 60 * 60)
transcript = "" # Gemini doesn't provide transcript
return ChatCompletionAudioResponse(
data=data,
expires_at=expires_at,
transcript=transcript
)
return None
def _transform_parts(
self,
parts: List[HttpxPartType],
@ -967,8 +1040,17 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
) = VertexGeminiConfig().get_assistant_content_message(
parts=candidate["content"]["parts"]
)
if content is not None:
audio_response = VertexGeminiConfig()._extract_audio_response_from_parts(
parts=candidate["content"]["parts"]
)
if audio_response is not None:
cast(Dict[str, Any], chat_completion_message)["audio"] = audio_response
chat_completion_message["content"] = None # OpenAI spec
elif content is not None:
chat_completion_message["content"] = content
if reasoning_content is not None:
chat_completion_message["reasoning_content"] = reasoning_content
@ -1185,7 +1267,9 @@ async def make_call(
)
completion_stream = ModelResponseIterator(
streaming_response=response.aiter_lines(), sync_stream=False
streaming_response=response.aiter_lines(),
sync_stream=False,
logging_obj=logging_obj,
)
# LOGGING
logging_obj.post_call(
@ -1223,7 +1307,9 @@ def make_sync_call(
)
completion_stream = ModelResponseIterator(
streaming_response=response.iter_lines(), sync_stream=True
streaming_response=response.iter_lines(),
sync_stream=True,
logging_obj=logging_obj,
)
# LOGGING
@ -1644,11 +1730,19 @@ class VertexLLM(VertexBase):
class ModelResponseIterator:
def __init__(self, streaming_response, sync_stream: bool):
def __init__(
self, streaming_response, sync_stream: bool, logging_obj: LoggingClass
):
from litellm.litellm_core_utils.prompt_templates.common_utils import (
check_is_function_call,
)
self.streaming_response = streaming_response
self.chunk_type: Literal["valid_json", "accumulated_json"] = "valid_json"
self.accumulated_json = ""
self.sent_first_chunk = False
self.logging_obj = logging_obj
self.is_function_call = check_is_function_call(logging_obj)
def chunk_parser(self, chunk: dict) -> GenericStreamingChunk:
try:
@ -1712,11 +1806,23 @@ class ModelResponseIterator:
},
)
returned_chunk = GenericStreamingChunk(
text=text,
tool_use=tool_use,
is_finished=False,
finish_reason=finish_reason,
args: Dict[str, Any] = {
"content": text or None,
"reasoning_content": reasoning_content,
}
if self.is_function_call and tool_use is not None:
args["function_call"] = tool_use["function"]
elif tool_use is not None:
args["tool_calls"] = [tool_use]
returned_chunk = ModelResponseStream(
choices=[
StreamingChoices(
index=0,
delta=Delta(**args),
finish_reason=finish_reason,
)
],
usage=usage,
index=0,
)
@ -1729,7 +1835,7 @@ class ModelResponseIterator:
self.response_iterator = self.streaming_response
return self
def handle_valid_json_chunk(self, chunk: str) -> GenericStreamingChunk:
def handle_valid_json_chunk(self, chunk: str) -> Optional[ModelResponseStream]:
chunk = chunk.strip()
try:
json_chunk = json.loads(chunk)
@ -1747,7 +1853,9 @@ class ModelResponseIterator:
return self.chunk_parser(chunk=json_chunk)
def handle_accumulated_json_chunk(self, chunk: str) -> GenericStreamingChunk:
def handle_accumulated_json_chunk(
self, chunk: str
) -> Optional[ModelResponseStream]:
chunk = litellm.CustomStreamWrapper._strip_sse_data_from_chunk(chunk) or ""
message = chunk.replace("\n\n", "")
@ -1761,16 +1869,9 @@ class ModelResponseIterator:
return self.chunk_parser(chunk=_data)
except json.JSONDecodeError:
# If it's not valid JSON yet, continue to the next event
return GenericStreamingChunk(
text="",
is_finished=False,
finish_reason="",
usage=None,
index=0,
tool_use=None,
)
return None
def _common_chunk_parsing_logic(self, chunk: str) -> GenericStreamingChunk:
def _common_chunk_parsing_logic(self, chunk: str) -> Optional[ModelResponseStream]:
try:
chunk = litellm.CustomStreamWrapper._strip_sse_data_from_chunk(chunk) or ""
if len(chunk) > 0:
@ -1783,15 +1884,7 @@ class ModelResponseIterator:
return self.handle_valid_json_chunk(chunk=chunk)
elif self.chunk_type == "accumulated_json":
return self.handle_accumulated_json_chunk(chunk=chunk)
return GenericStreamingChunk(
text="",
is_finished=False,
finish_reason="",
usage=None,
index=0,
tool_use=None,
)
return None
except Exception:
raise

View file

@ -0,0 +1,96 @@
from typing import Any, Dict, List, Optional, Tuple
import litellm
from litellm.llms.anthropic.experimental_pass_through.messages.transformation import (
AnthropicMessagesConfig,
)
from litellm.secret_managers.main import get_secret_str
from litellm.types.llms.vertex_ai import VertexPartnerProvider
from litellm.types.router import GenericLiteLLMParams
from ....vertex_llm_base import VertexBase
class VertexAIPartnerModelsAnthropicMessagesConfig(AnthropicMessagesConfig, VertexBase):
def validate_anthropic_messages_environment(
self,
headers: dict,
model: str,
messages: List[Any],
optional_params: dict,
litellm_params: dict,
api_key: Optional[str] = None,
api_base: Optional[str] = None,
) -> Tuple[dict, Optional[str]]:
"""
OPTIONAL
Validate the environment for the request
"""
if "Authorization" not in headers:
vertex_ai_project = (
optional_params.pop("vertex_project", None)
or optional_params.pop("vertex_ai_project", None)
or litellm.vertex_project
or get_secret_str("VERTEXAI_PROJECT")
)
vertex_credentials = (
optional_params.pop("vertex_credentials", None)
or optional_params.pop("vertex_ai_credentials", None)
or get_secret_str("VERTEXAI_CREDENTIALS")
)
access_token, project_id = self._ensure_access_token(
credentials=vertex_credentials,
project_id=vertex_ai_project,
custom_llm_provider="vertex_ai",
)
headers["Authorization"] = f"Bearer {access_token}"
api_base = self.get_complete_vertex_url(
custom_api_base=api_base,
vertex_location=optional_params.pop("vertex_location", None),
vertex_project=vertex_ai_project,
project_id=project_id,
partner=VertexPartnerProvider.claude,
stream=optional_params.get("stream", False),
model=model,
)
headers["content-type"] = "application/json"
return headers, api_base
def get_complete_url(
self,
api_base: Optional[str],
api_key: Optional[str],
model: str,
optional_params: dict,
litellm_params: dict,
stream: Optional[bool] = None,
) -> str:
if api_base is None:
raise ValueError(
"api_base is required. Unable to determine the correct api_base for the request."
)
return api_base # no transformation is needed - handled in validate_environment
def transform_anthropic_messages_request(
self,
model: str,
messages: List[Dict],
anthropic_messages_optional_request_params: Dict,
litellm_params: GenericLiteLLMParams,
headers: dict,
) -> Dict:
anthropic_messages_request = super().transform_anthropic_messages_request(
model=model,
messages=messages,
anthropic_messages_optional_request_params=anthropic_messages_optional_request_params,
litellm_params=litellm_params,
headers=headers,
)
anthropic_messages_request["anthropic_version"] = "vertex-2023-10-16"
return anthropic_messages_request

View file

@ -1,12 +1,12 @@
# What is this?
## API Handler for calling Vertex AI Partner Models
from enum import Enum
from typing import Callable, Optional, Union
import httpx # type: ignore
import litellm
from litellm import LlmProviders
from litellm.types.llms.vertex_ai import VertexPartnerProvider
from litellm.utils import ModelResponse
from ...custom_httpx.llm_http_handler import BaseLLMHTTPHandler
@ -15,13 +15,6 @@ from ..vertex_llm_base import VertexBase
base_llm_http_handler = BaseLLMHTTPHandler()
class VertexPartnerProvider(str, Enum):
mistralai = "mistralai"
llama = "llama"
ai21 = "ai21"
claude = "claude"
class VertexAIError(Exception):
def __init__(self, status_code, message):
self.status_code = status_code
@ -35,78 +28,10 @@ class VertexAIError(Exception):
) # Call the base class constructor with the parameters it needs
def create_vertex_url(
vertex_location: str,
vertex_project: str,
partner: VertexPartnerProvider,
stream: Optional[bool],
model: str,
api_base: Optional[str] = None,
) -> str:
"""Return the base url for the vertex partner models"""
api_base = api_base or f"https://{vertex_location}-aiplatform.googleapis.com"
if partner == VertexPartnerProvider.llama:
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/endpoints/openapi/chat/completions"
elif partner == VertexPartnerProvider.mistralai:
if stream:
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/mistralai/models/{model}:streamRawPredict"
else:
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/mistralai/models/{model}:rawPredict"
elif partner == VertexPartnerProvider.ai21:
if stream:
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/publishers/ai21/models/{model}:streamRawPredict"
else:
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/publishers/ai21/models/{model}:rawPredict"
elif partner == VertexPartnerProvider.claude:
if stream:
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/anthropic/models/{model}:streamRawPredict"
else:
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/anthropic/models/{model}:rawPredict"
class VertexAIPartnerModels(VertexBase):
def __init__(self) -> None:
pass
def get_complete_url(
self,
custom_api_base: Optional[str],
vertex_location: Optional[str],
vertex_project: Optional[str],
project_id: str,
partner: VertexPartnerProvider,
stream: Optional[bool],
model: str,
) -> str:
api_base = self.get_api_base(
api_base=custom_api_base, vertex_location=vertex_location
)
default_api_base = create_vertex_url(
vertex_location=vertex_location or "us-central1",
vertex_project=vertex_project or project_id,
partner=partner, # type: ignore
stream=stream,
model=model,
api_base=api_base,
)
if len(default_api_base.split(":")) > 1:
endpoint = default_api_base.split(":")[-1]
else:
endpoint = ""
_, api_base = self._check_custom_proxy(
api_base=custom_api_base,
custom_llm_provider="vertex_ai",
gemini_api_key=None,
endpoint=endpoint,
stream=stream,
auth_header=None,
url=default_api_base,
)
return api_base
def completion(
self,
model: str,
@ -181,7 +106,7 @@ class VertexAIPartnerModels(VertexBase):
else:
raise ValueError(f"Unknown partner model: {model}")
api_base = self.get_complete_url(
api_base = self.get_complete_vertex_url(
custom_api_base=api_base,
vertex_location=vertex_location,
vertex_project=vertex_project,

View file

@ -11,7 +11,7 @@ from typing import TYPE_CHECKING, Any, Dict, Literal, Optional, Tuple
from litellm._logging import verbose_logger
from litellm.litellm_core_utils.asyncify import asyncify
from litellm.llms.custom_httpx.http_handler import AsyncHTTPHandler
from litellm.types.llms.vertex_ai import VERTEX_CREDENTIALS_TYPES
from litellm.types.llms.vertex_ai import VERTEX_CREDENTIALS_TYPES, VertexPartnerProvider
from .common_utils import _get_gemini_url, _get_vertex_url, all_gemini_url_modes
@ -150,6 +150,74 @@ class VertexBase:
else:
return f"https://{self.get_default_vertex_location()}-aiplatform.googleapis.com"
@staticmethod
def create_vertex_url(
vertex_location: str,
vertex_project: str,
partner: VertexPartnerProvider,
stream: Optional[bool],
model: str,
api_base: Optional[str] = None,
) -> str:
"""Return the base url for the vertex partner models"""
api_base = api_base or f"https://{vertex_location}-aiplatform.googleapis.com"
if partner == VertexPartnerProvider.llama:
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/endpoints/openapi/chat/completions"
elif partner == VertexPartnerProvider.mistralai:
if stream:
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/mistralai/models/{model}:streamRawPredict"
else:
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/mistralai/models/{model}:rawPredict"
elif partner == VertexPartnerProvider.ai21:
if stream:
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/publishers/ai21/models/{model}:streamRawPredict"
else:
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/publishers/ai21/models/{model}:rawPredict"
elif partner == VertexPartnerProvider.claude:
if stream:
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/anthropic/models/{model}:streamRawPredict"
else:
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/anthropic/models/{model}:rawPredict"
def get_complete_vertex_url(
self,
custom_api_base: Optional[str],
vertex_location: Optional[str],
vertex_project: Optional[str],
project_id: str,
partner: VertexPartnerProvider,
stream: Optional[bool],
model: str,
) -> str:
api_base = self.get_api_base(
api_base=custom_api_base, vertex_location=vertex_location
)
default_api_base = VertexBase.create_vertex_url(
vertex_location=vertex_location or "us-central1",
vertex_project=vertex_project or project_id,
partner=partner,
stream=stream,
model=model,
api_base=api_base,
)
if len(default_api_base.split(":")) > 1:
endpoint = default_api_base.split(":")[-1]
else:
endpoint = ""
_, api_base = self._check_custom_proxy(
api_base=custom_api_base,
custom_llm_provider="vertex_ai",
gemini_api_key=None,
endpoint=endpoint,
stream=stream,
auth_header=None,
url=default_api_base,
)
return api_base
def refresh_auth(self, credentials: Any) -> None:
from google.auth.transport.requests import (
Request, # type: ignore[import-untyped]

View file

@ -3,6 +3,7 @@ from typing import List, Optional, Tuple
import litellm
from litellm._logging import verbose_logger
from litellm.litellm_core_utils.prompt_templates.common_utils import (
filter_value_from_dict,
strip_name_from_messages,
)
from litellm.secret_managers.main import get_secret_str
@ -44,6 +45,7 @@ class XAIChatConfig(OpenAIGPTConfig):
"top_logprobs",
"top_p",
"user",
"web_search_options",
]
try:
if litellm.supports_reasoning(
@ -66,6 +68,14 @@ class XAIChatConfig(OpenAIGPTConfig):
for param, value in non_default_params.items():
if param == "max_completion_tokens":
optional_params["max_tokens"] = value
elif param == "tools" and value is not None:
tools = []
for tool in value:
tool = filter_value_from_dict(tool, "strict")
if tool is not None:
tools.append(tool)
if len(tools) > 0:
optional_params["tools"] = tools
elif param in supported_openai_params:
if value is not None:
optional_params[param] = value

View file

@ -6,9 +6,21 @@ import litellm
from litellm.llms.base_llm.base_utils import BaseLLMModelInfo
from litellm.secret_managers.main import get_secret_str
from litellm.types.llms.openai import AllMessageValues
from litellm.types.utils import ProviderSpecificModelInfo
class XAIModelInfo(BaseLLMModelInfo):
def get_provider_info(
self,
model: str,
) -> Optional[ProviderSpecificModelInfo]:
"""
Default values all models of this provider support.
"""
return {
"supports_web_search": True,
}
def validate_environment(
self,
headers: dict,

View file

@ -59,6 +59,7 @@ from litellm.constants import (
from litellm.exceptions import LiteLLMUnknownProvider
from litellm.integrations.custom_logger import CustomLogger
from litellm.litellm_core_utils.audio_utils.utils import get_audio_file_for_health_check
from litellm.litellm_core_utils.dd_tracing import tracer
from litellm.litellm_core_utils.health_check_utils import (
_create_health_check_response,
_filter_model_params,
@ -314,6 +315,7 @@ class AsyncCompletions:
return response
@tracer.wrap()
@client
async def acompletion(
model: str,
@ -810,6 +812,7 @@ def mock_completion(
raise Exception("Mock completion response failed - {}".format(e))
@tracer.wrap()
@client
def completion( # type: ignore # noqa: PLR0915
model: str,
@ -2335,6 +2338,26 @@ def completion( # type: ignore # noqa: PLR0915
original_response=response,
additional_args={"headers": headers},
)
elif custom_llm_provider == "datarobot":
response = base_llm_http_handler.completion(
model=model,
messages=messages,
headers=headers,
model_response=model_response,
api_key=api_key,
api_base=api_base,
acompletion=acompletion,
logging_obj=logging,
optional_params=optional_params,
litellm_params=litellm_params,
timeout=timeout, # type: ignore
client=client,
custom_llm_provider=custom_llm_provider,
encoding=encoding,
stream=stream,
provider_config=provider_config,
)
elif custom_llm_provider == "openrouter":
api_base = (
api_base

View file

@ -96,13 +96,7 @@
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_native_streaming": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.03,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.05
}
"supports_native_streaming": true
},
"gpt-4.1-2025-04-14": {
"max_tokens": 32768,
@ -135,13 +129,7 @@
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_native_streaming": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.03,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.05
}
"supports_native_streaming": true
},
"gpt-4.1-mini": {
"max_tokens": 32768,
@ -174,13 +162,7 @@
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_native_streaming": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.025,
"search_context_size_medium": 0.0275,
"search_context_size_high": 0.03
}
"supports_native_streaming": true
},
"gpt-4.1-mini-2025-04-14": {
"max_tokens": 32768,
@ -213,13 +195,7 @@
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_native_streaming": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.025,
"search_context_size_medium": 0.0275,
"search_context_size_high": 0.03
}
"supports_native_streaming": true
},
"gpt-4.1-nano": {
"max_tokens": 32768,
@ -305,13 +281,7 @@
"supports_vision": true,
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.03,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.05
}
"supports_tool_choice": true
},
"watsonx/ibm/granite-3-8b-instruct": {
"max_tokens": 8192,
@ -349,13 +319,7 @@
"supports_vision": true,
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.03,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.05
}
"supports_tool_choice": true
},
"gpt-4o-search-preview": {
"max_tokens": 16384,
@ -527,13 +491,7 @@
"supports_vision": true,
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.025,
"search_context_size_medium": 0.0275,
"search_context_size_high": 0.03
}
"supports_tool_choice": true
},
"gpt-4o-mini-search-preview-2025-03-11": {
"max_tokens": 16384,
@ -553,13 +511,7 @@
"supports_vision": true,
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.025,
"search_context_size_medium": 0.0275,
"search_context_size_high": 0.03
}
"supports_tool_choice": true
},
"gpt-4o-mini-search-preview": {
"max_tokens": 16384,
@ -954,13 +906,7 @@
"supports_vision": true,
"supports_prompt_caching": true,
"supports_system_messages": true,
"supports_tool_choice": true,
"supports_web_search": true,
"search_context_cost_per_query": {
"search_context_size_low": 0.03,
"search_context_size_medium": 0.035,
"search_context_size_high": 0.05
}
"supports_tool_choice": true
},
"gpt-4o-2024-11-20": {
"max_tokens": 16384,
@ -4217,7 +4163,8 @@
"mode": "chat",
"supports_function_calling": true,
"supports_vision": true,
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"xai/grok-2-vision-1212": {
"max_tokens": 32768,
@ -4230,7 +4177,8 @@
"mode": "chat",
"supports_function_calling": true,
"supports_vision": true,
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"xai/grok-2-vision-latest": {
"max_tokens": 32768,
@ -4243,7 +4191,8 @@
"mode": "chat",
"supports_function_calling": true,
"supports_vision": true,
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"xai/grok-2-vision": {
"max_tokens": 32768,
@ -4256,7 +4205,8 @@
"mode": "chat",
"supports_function_calling": true,
"supports_vision": true,
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"xai/grok-3": {
"max_tokens": 131072,
@ -4269,7 +4219,8 @@
"supports_function_calling": true,
"supports_tool_choice": true,
"supports_response_schema": false,
"source": "https://x.ai/api#pricing"
"source": "https://x.ai/api#pricing",
"supports_web_search": true
},
"xai/grok-3-beta": {
"max_tokens": 131072,
@ -4282,7 +4233,8 @@
"supports_function_calling": true,
"supports_tool_choice": true,
"supports_response_schema": false,
"source": "https://x.ai/api#pricing"
"source": "https://x.ai/api#pricing",
"supports_web_search": true
},
"xai/grok-3-fast-beta": {
"max_tokens": 131072,
@ -4295,7 +4247,8 @@
"supports_function_calling": true,
"supports_tool_choice": true,
"supports_response_schema": false,
"source": "https://x.ai/api#pricing"
"source": "https://x.ai/api#pricing",
"supports_web_search": true
},
"xai/grok-3-fast-latest": {
"max_tokens": 131072,
@ -4308,7 +4261,8 @@
"supports_function_calling": true,
"supports_tool_choice": true,
"supports_response_schema": false,
"source": "https://x.ai/api#pricing"
"source": "https://x.ai/api#pricing",
"supports_web_search": true
},
"xai/grok-3-mini-beta": {
"max_tokens": 131072,
@ -4322,7 +4276,8 @@
"supports_tool_choice": true,
"supports_reasoning": true,
"supports_response_schema": false,
"source": "https://x.ai/api#pricing"
"source": "https://x.ai/api#pricing",
"supports_web_search": true
},
"xai/grok-3-mini-fast-beta": {
"max_tokens": 131072,
@ -4336,7 +4291,8 @@
"supports_tool_choice": true,
"supports_reasoning": true,
"supports_response_schema": false,
"source": "https://x.ai/api#pricing"
"source": "https://x.ai/api#pricing",
"supports_web_search": true
},
"xai/grok-3-mini-fast-latest": {
"max_tokens": 131072,
@ -4350,7 +4306,8 @@
"supports_function_calling": true,
"supports_tool_choice": true,
"supports_response_schema": false,
"source": "https://x.ai/api#pricing"
"source": "https://x.ai/api#pricing",
"supports_web_search": true
},
"xai/grok-vision-beta": {
"max_tokens": 8192,
@ -4363,7 +4320,8 @@
"mode": "chat",
"supports_function_calling": true,
"supports_vision": true,
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"xai/grok-2-1212": {
"max_tokens": 131072,
@ -4374,7 +4332,8 @@
"litellm_provider": "xai",
"mode": "chat",
"supports_function_calling": true,
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"xai/grok-2": {
"max_tokens": 131072,
@ -4385,7 +4344,8 @@
"litellm_provider": "xai",
"mode": "chat",
"supports_function_calling": true,
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"xai/grok-2-latest": {
"max_tokens": 131072,
@ -4396,7 +4356,8 @@
"litellm_provider": "xai",
"mode": "chat",
"supports_function_calling": true,
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"deepseek/deepseek-coder": {
"max_tokens": 4096,
@ -6112,7 +6073,8 @@
"text"
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-pro-exp-02-05": {
"max_tokens": 8192,
@ -6152,7 +6114,8 @@
"text"
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-flash-exp": {
"max_tokens": 8192,
@ -6197,7 +6160,8 @@
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
"supports_tool_choice": true,
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-flash-001": {
"max_tokens": 8192,
@ -6232,7 +6196,8 @@
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
"deprecation_date": "2026-02-05",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-flash-thinking-exp": {
"max_tokens": 8192,
@ -6277,7 +6242,8 @@
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
"supports_tool_choice": true,
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-flash-thinking-exp-01-21": {
"max_tokens": 65536,
@ -6322,7 +6288,8 @@
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
"supports_tool_choice": true,
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini/gemini-2.5-pro-exp-03-25": {
"max_tokens": 65535,
@ -6363,7 +6330,8 @@
"supported_output_modalities": [
"text"
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing"
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
"supports_web_search": true
},
"gemini/gemini-2.5-flash-preview-tts": {
"max_tokens": 65535,
@ -6400,7 +6368,8 @@
"supported_output_modalities": [
"audio"
],
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview"
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
"supports_web_search": true
},
"gemini/gemini-2.5-flash-preview-05-20": {
"max_tokens": 65535,
@ -6440,7 +6409,8 @@
"supported_output_modalities": [
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview"
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
"supports_web_search": true
},
"gemini/gemini-2.5-flash-preview-04-17": {
"max_tokens": 65535,
@ -6480,7 +6450,8 @@
"supported_output_modalities": [
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview"
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
"supports_web_search": true
},
"gemini-2.5-flash-preview-05-20": {
"max_tokens": 65535,
@ -6520,7 +6491,8 @@
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.5-flash-preview-04-17": {
"max_tokens": 65535,
@ -6560,7 +6532,8 @@
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-flash": {
"max_tokens": 8192,
@ -6595,7 +6568,8 @@
],
"supports_tool_choice": true,
"source": "https://ai.google.dev/pricing#2_0flash",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-flash-lite": {
"max_input_tokens": 1048576,
@ -6627,7 +6601,8 @@
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
"supports_tool_choice": true,
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-flash-lite-001": {
"max_input_tokens": 1048576,
@ -6660,7 +6635,8 @@
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
"supports_tool_choice": true,
"deprecation_date": "2026-02-25",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.5-pro-preview-05-06": {
"max_tokens": 65535,
@ -6701,7 +6677,8 @@
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.5-pro-preview-03-25": {
"max_tokens": 65535,
@ -6742,7 +6719,8 @@
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.0-flash-preview-image-generation": {
"max_tokens": 8192,
@ -6777,7 +6755,8 @@
],
"supports_tool_choice": true,
"source": "https://ai.google.dev/pricing#2_0flash",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini-2.5-pro-preview-tts": {
"max_tokens": 65535,
@ -6809,7 +6788,8 @@
"audio"
],
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview",
"supports_parallel_function_calling": true
"supports_parallel_function_calling": true,
"supports_web_search": true
},
"gemini/gemini-2.0-pro-exp-02-05": {
"max_tokens": 8192,
@ -6847,7 +6827,8 @@
"supports_pdf_input": true,
"supports_response_schema": true,
"supports_tool_choice": true,
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing"
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
"supports_web_search": true
},
"gemini/gemini-2.0-flash-preview-image-generation": {
"max_tokens": 8192,
@ -6883,7 +6864,8 @@
"image"
],
"supports_tool_choice": true,
"source": "https://ai.google.dev/pricing#2_0flash"
"source": "https://ai.google.dev/pricing#2_0flash",
"supports_web_search": true
},
"gemini/gemini-2.0-flash": {
"max_tokens": 8192,
@ -6919,7 +6901,8 @@
"image"
],
"supports_tool_choice": true,
"source": "https://ai.google.dev/pricing#2_0flash"
"source": "https://ai.google.dev/pricing#2_0flash",
"supports_web_search": true
},
"gemini/gemini-2.0-flash-lite": {
"max_input_tokens": 1048576,
@ -6952,7 +6935,8 @@
"supported_output_modalities": [
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.0-flash-lite"
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.0-flash-lite",
"supports_web_search": true
},
"gemini/gemini-2.0-flash-001": {
"max_tokens": 8192,
@ -6987,7 +6971,8 @@
"text",
"image"
],
"source": "https://ai.google.dev/pricing#2_0flash"
"source": "https://ai.google.dev/pricing#2_0flash",
"supports_web_search": true
},
"gemini/gemini-2.5-pro-preview-tts": {
"max_tokens": 65535,
@ -7020,7 +7005,8 @@
"supported_output_modalities": [
"audio"
],
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview"
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview",
"supports_web_search": true
},
"gemini/gemini-2.5-pro-preview-05-06": {
"max_tokens": 65535,
@ -7056,7 +7042,8 @@
"supported_output_modalities": [
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview"
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview",
"supports_web_search": true
},
"gemini/gemini-2.5-pro-preview-03-25": {
"max_tokens": 65535,
@ -7092,7 +7079,8 @@
"supported_output_modalities": [
"text"
],
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview"
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview",
"supports_web_search": true
},
"gemini/gemini-2.0-flash-exp": {
"max_tokens": 8192,
@ -7138,7 +7126,8 @@
"image"
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"gemini/gemini-2.0-flash-lite-preview-02-05": {
"max_tokens": 8192,
@ -7172,7 +7161,8 @@
"supported_output_modalities": [
"text"
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash-lite"
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash-lite",
"supports_web_search": true
},
"gemini/gemini-2.0-flash-thinking-exp": {
"max_tokens": 8192,
@ -7218,7 +7208,8 @@
"image"
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"gemini/gemini-2.0-flash-thinking-exp-01-21": {
"max_tokens": 8192,
@ -7264,7 +7255,8 @@
"image"
],
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
"supports_tool_choice": true
"supports_tool_choice": true,
"supports_web_search": true
},
"gemini/gemma-3-27b-it": {
"max_tokens": 8192,
@ -8758,6 +8750,20 @@
"notes": "'supports_image_input' is a deprecated field. Use 'supports_embedding_image_input' instead."
}
},
"embed-v4.0": {
"max_tokens": 1024,
"max_input_tokens": 1024,
"input_cost_per_token": 1.2e-07,
"input_cost_per_image": 4.7e-07,
"output_cost_per_token": 0.0,
"litellm_provider": "cohere",
"mode": "embedding",
"supports_image_input": true,
"supports_embedding_image_input": true,
"metadata": {
"notes": "'supports_image_input' is a deprecated field. Use 'supports_embedding_image_input' instead."
}
},
"replicate/meta/llama-2-13b": {
"max_tokens": 4096,
"max_input_tokens": 4096,

View file

@ -0,0 +1,118 @@
This makes it easier to pass through requests to the LLM APIs.
E.g. Route to VLLM's `/classify` endpoint:
## SDK (Basic)
```python
import litellm
response = litellm.llm_passthrough_route(
model="hosted_vllm/papluca/xlm-roberta-base-language-detection",
method="POST",
endpoint="classify",
api_base="http://localhost:8090",
api_key=None,
json={
"model": "swapped-for-litellm-model",
"input": "Hello, world!",
}
)
print(response)
```
## SDK (Router)
```python
import asyncio
from litellm import Router
router = Router(
model_list=[
{
"model_name": "roberta-base-language-detection",
"litellm_params": {
"model": "hosted_vllm/papluca/xlm-roberta-base-language-detection",
"api_base": "http://localhost:8090",
}
}
]
)
request_data = {
"model": "roberta-base-language-detection",
"method": "POST",
"endpoint": "classify",
"api_base": "http://localhost:8090",
"api_key": None,
"json": {
"model": "roberta-base-language-detection",
"input": "Hello, world!",
}
}
async def main():
response = await router.allm_passthrough_route(**request_data)
print(response)
if __name__ == "__main__":
asyncio.run(main())
```
## PROXY
1. Setup config.yaml
```yaml
model_list:
- model_name: roberta-base-language-detection
litellm_params:
model: hosted_vllm/papluca/xlm-roberta-base-language-detection
api_base: http://localhost:8090
```
2. Run the proxy
```bash
litellm proxy --config config.yaml
# RUNNING on http://localhost:4000
```
3. Use the proxy
```bash
curl -X POST http://localhost:4000/vllm/classify \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <your-api-key>" \
-d '{"model": "roberta-base-language-detection", "input": "Hello, world!"}' \
```
# How to add a provider for passthrough
See [VLLMModelInfo](https://github.com/BerriAI/litellm/blob/main/litellm/llms/vllm/common_utils.py) for an example.
1. Inherit from BaseModelInfo
```python
from litellm.llms.base_llm.base_utils import BaseLLMModelInfo
class VLLMModelInfo(BaseLLMModelInfo):
pass
```
2. Register the provider in the ProviderConfigManager.get_provider_model_info
```python
from litellm.utils import ProviderConfigManager
from litellm.types.utils import LlmProviders
provider_config = ProviderConfigManager.get_provider_model_info(
model="my-test-model", provider=LlmProviders.VLLM
)
print(provider_config)
```

View file

@ -0,0 +1,8 @@
from .main import allm_passthrough_route, llm_passthrough_route
from .utils import BasePassthroughUtils
__all__ = [
"allm_passthrough_route",
"llm_passthrough_route",
"BasePassthroughUtils",
]

193
litellm/passthrough/main.py Normal file
View file

@ -0,0 +1,193 @@
"""
This module is used to pass through requests to the LLM APIs.
"""
import asyncio
import contextvars
from functools import partial
from typing import Any, Coroutine, Optional, Union
from urllib.parse import urlencode
import httpx
from httpx._types import CookieTypes, QueryParamTypes, RequestFiles
import litellm
from litellm.litellm_core_utils.get_llm_provider_logic import get_llm_provider
from litellm.llms.custom_httpx.http_handler import AsyncHTTPHandler, HTTPHandler
from litellm.utils import client
from .utils import BasePassthroughUtils
@client
async def allm_passthrough_route(
*,
method: str,
endpoint: str,
custom_llm_provider: Optional[str] = None,
api_base: Optional[str] = None,
api_key: Optional[str] = None,
request_query_params: Optional[dict] = None,
request_headers: Optional[dict] = None,
stream: bool = False,
content: Optional[Any] = None,
data: Optional[dict] = None,
files: Optional[RequestFiles] = None,
json: Optional[Any] = None,
params: Optional[QueryParamTypes] = None,
cookies: Optional[CookieTypes] = None,
client: Optional[Union[HTTPHandler, AsyncHTTPHandler]] = None,
**kwargs,
) -> Union[httpx.Response, Coroutine[Any, Any, httpx.Response]]:
"""
Async: Reranks a list of documents based on their relevance to the query
"""
try:
loop = asyncio.get_event_loop()
kwargs["allm_passthrough_route"] = True
func = partial(
llm_passthrough_route,
method=method,
endpoint=endpoint,
custom_llm_provider=custom_llm_provider,
api_base=api_base,
api_key=api_key,
request_query_params=request_query_params,
request_headers=request_headers,
stream=stream,
content=content,
data=data,
files=files,
json=json,
params=params,
cookies=cookies,
client=client,
**kwargs,
)
ctx = contextvars.copy_context()
func_with_context = partial(ctx.run, func)
init_response = await loop.run_in_executor(None, func_with_context)
if asyncio.iscoroutine(init_response):
response = await init_response
else:
response = init_response
return response
except Exception as e:
raise e
@client
def llm_passthrough_route(
*,
method: str,
endpoint: str,
model: str,
custom_llm_provider: Optional[str] = None,
api_base: Optional[str] = None,
api_key: Optional[str] = None,
request_query_params: Optional[dict] = None,
request_headers: Optional[dict] = None,
allm_passthrough_route: bool = False,
stream: bool = False,
content: Optional[Any] = None,
data: Optional[dict] = None,
files: Optional[RequestFiles] = None,
json: Optional[Any] = None,
params: Optional[QueryParamTypes] = None,
cookies: Optional[CookieTypes] = None,
client: Optional[Union[HTTPHandler, AsyncHTTPHandler]] = None,
**kwargs,
) -> Union[httpx.Response, Coroutine[Any, Any, httpx.Response]]:
"""
Pass through requests to the LLM APIs.
Step 1. Build the request
Step 2. Send the request
Step 3. Return the response
[TODO] Refactor this into a provider-config pattern, once we expand this to non-vllm providers.
"""
if client is None:
if allm_passthrough_route:
client = litellm.module_level_aclient
else:
client = litellm.module_level_client
model, custom_llm_provider, api_key, api_base = get_llm_provider(
model=model,
custom_llm_provider=custom_llm_provider,
api_base=api_base,
api_key=api_key,
)
from litellm.types.utils import LlmProviders
from litellm.utils import ProviderConfigManager
provider_config = ProviderConfigManager.get_provider_model_info(
provider=LlmProviders(custom_llm_provider),
model=model,
)
if provider_config is None:
raise Exception(f"Provider {custom_llm_provider} not found")
base_target_url = provider_config.get_api_base(api_base)
if base_target_url is None:
raise Exception(f"Provider {custom_llm_provider} api base not found")
encoded_endpoint = httpx.URL(endpoint).path
# Ensure endpoint starts with '/' for proper URL construction
if not encoded_endpoint.startswith("/"):
encoded_endpoint = "/" + encoded_endpoint
# Construct the full target URL using httpx
base_url = httpx.URL(base_target_url)
updated_url = base_url.copy_with(path=encoded_endpoint)
if request_query_params:
# Create a new URL with the merged query params
updated_url = updated_url.copy_with(
query=urlencode(request_query_params).encode("ascii")
)
# Add or update query parameters
provider_api_key = provider_config.get_api_key(api_key)
auth_headers = provider_config.validate_environment(
headers={},
model=model,
messages=[],
optional_params={},
litellm_params={},
api_key=provider_api_key,
api_base=base_target_url,
)
headers = BasePassthroughUtils.forward_headers_from_request(
request_headers=request_headers or {},
headers=auth_headers,
forward_headers=False,
)
## SWAP MODEL IN JSON BODY
if json and isinstance(json, dict) and "model" in json:
json["model"] = model
request = client.client.build_request(
method=method,
url=updated_url,
content=content,
data=data,
files=files,
json=json,
params=params,
headers=headers,
cookies=cookies,
)
response = client.client.send(request=request, stream=stream)
return response

View file

@ -0,0 +1,39 @@
from typing import Dict, List, Optional, Union
from urllib.parse import parse_qs
import httpx
class BasePassthroughUtils:
@staticmethod
def get_merged_query_parameters(
existing_url: httpx.URL, request_query_params: Dict[str, Union[str, list]]
) -> Dict[str, Union[str, List[str]]]:
# Get the existing query params from the target URL
existing_query_string = existing_url.query.decode("utf-8")
existing_query_params = parse_qs(existing_query_string)
# parse_qs returns a dict where each value is a list, so let's flatten it
updated_existing_query_params = {
k: v[0] if len(v) == 1 else v for k, v in existing_query_params.items()
}
# Merge the query params, giving priority to the existing ones
return {**request_query_params, **updated_existing_query_params}
@staticmethod
def forward_headers_from_request(
request_headers: dict,
headers: dict,
forward_headers: Optional[bool] = False,
):
"""
Helper to forward headers from original request
"""
if forward_headers is True:
# Header We Should NOT forward
request_headers.pop("content-length", None)
request_headers.pop("host", None)
# Combine request headers with custom headers
headers = {**request_headers, **headers}
return headers

View file

@ -73,13 +73,13 @@ async def get_mcp_servers_by_verificationtoken(
)
)
mcp_servers = []
mcp_servers: Optional[List[str]] = []
if (
verification_token_record is not None
and verification_token_record.object_permission is not None
):
mcp_servers = verification_token_record.object_permission.mcp_servers
return mcp_servers
return mcp_servers or []
async def get_mcp_servers_by_team(
@ -99,10 +99,10 @@ async def get_mcp_servers_by_team(
)
)
mcp_servers = []
mcp_servers: Optional[List[str]] = []
if team_record is not None and team_record.object_permission is not None:
mcp_servers = team_record.object_permission.mcp_servers
return mcp_servers
return mcp_servers or []
async def get_all_mcp_servers_for_user(

View file

@ -72,8 +72,9 @@ class MCPServerManager:
mcp_info = MCPInfo(**_mcp_info)
mcp_info["server_name"] = server_name
mcp_info["description"] = server_config.get("description", None)
server_id = str(uuid.uuid4())
new_server = MCPServer(
server_id=str(uuid.uuid4()),
server_id=server_id,
name=server_name,
url=server_config["url"],
# TODO: utility fn the default values
@ -82,7 +83,6 @@ class MCPServerManager:
auth_type=server_config.get("auth_type", None),
mcp_info=mcp_info,
)
server_id = str(uuid.uuid4())
self.config_mcp_servers[server_id] = new_server
verbose_logger.debug(
f"Loaded MCP Servers: {json.dumps(self.config_mcp_servers, indent=4, default=str)}"

View file

@ -0,0 +1,12 @@
import importlib
def is_mcp_available() -> bool:
"""
Returns True if the MCP module is available, False otherwise
"""
try:
importlib.import_module("mcp")
return True
except ImportError:
return False

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1 @@
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[185],{96443:function(e,n,t){Promise.resolve().then(t.t.bind(t,39974,23)),Promise.resolve().then(t.t.bind(t,2778,23))},2778:function(){},39974:function(e){e.exports={style:{fontFamily:"'__Inter_3373e4', '__Inter_Fallback_3373e4'",fontStyle:"normal"},className:"__className_3373e4"}}},function(e){e.O(0,[919,986,971,117,744],function(){return e(e.s=96443)}),_N_E=e.O()}]);

View file

@ -1 +0,0 @@
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[185],{6580:function(n,e,t){Promise.resolve().then(t.t.bind(t,39974,23)),Promise.resolve().then(t.t.bind(t,2778,23))},2778:function(){},39974:function(n){n.exports={style:{fontFamily:"'__Inter_cf7686', '__Inter_Fallback_cf7686'",fontStyle:"normal"},className:"__className_cf7686"}}},function(n){n.O(0,[919,191,971,117,744],function(){return n(n.s=6580)}),_N_E=n.O()}]);

View file

@ -1 +1 @@
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[418],{11790:function(e,n,u){Promise.resolve().then(u.bind(u,52829))},52829:function(e,n,u){"use strict";u.r(n),u.d(n,{default:function(){return f}});var t=u(57437),s=u(2265),r=u(99376),c=u(92699);function f(){let e=(0,r.useSearchParams)().get("key"),[n,u]=(0,s.useState)(null);return(0,s.useEffect)(()=>{e&&u(e)},[e]),(0,t.jsx)(c.Z,{accessToken:n,publicPage:!0,premiumUser:!1})}}},function(e){e.O(0,[402,313,250,699,971,117,744],function(){return e(e.s=11790)}),_N_E=e.O()}]);
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[418],{21024:function(e,n,u){Promise.resolve().then(u.bind(u,52829))},52829:function(e,n,u){"use strict";u.r(n),u.d(n,{default:function(){return f}});var t=u(57437),s=u(2265),r=u(99376),c=u(92699);function f(){let e=(0,r.useSearchParams)().get("key"),[n,u]=(0,s.useState)(null);return(0,s.useEffect)(()=>{e&&u(e)},[e]),(0,t.jsx)(c.Z,{accessToken:n,publicPage:!0,premiumUser:!1})}}},function(e){e.O(0,[402,313,250,699,971,117,744],function(){return e(e.s=21024)}),_N_E=e.O()}]);

View file

@ -0,0 +1 @@
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[461],{8672:function(e,t,n){Promise.resolve().then(n.bind(n,12011))},12011:function(e,t,n){"use strict";n.r(t),n.d(t,{default:function(){return S}});var o=n(57437),s=n(2265),a=n(99376),c=n(20831),i=n(94789),l=n(12514),r=n(49804),u=n(67101),d=n(84264),m=n(49566),h=n(96761),x=n(84566),f=n(19250),p=n(14474),g=n(13634),k=n(73002),j=n(3914);function S(){let[e]=g.Z.useForm(),t=(0,a.useSearchParams)();(0,j.e)("token");let n=t.get("invitation_id"),[S,w]=(0,s.useState)(null),[Z,_]=(0,s.useState)(""),[b,N]=(0,s.useState)(""),[y,T]=(0,s.useState)(null),[E,v]=(0,s.useState)(""),[C,U]=(0,s.useState)(""),[F,J]=(0,s.useState)(!0);return(0,s.useEffect)(()=>{(0,f.MO)().then(e=>{console.log("ui config in onboarding.tsx:",e),J(!1)})},[]),(0,s.useEffect)(()=>{n&&!F&&(0,f.W_)(n).then(e=>{let t=e.login_url;console.log("login_url:",t),v(t);let n=e.token,o=(0,p.o)(n);U(n),console.log("decoded:",o),w(o.key),console.log("decoded user email:",o.user_email),N(o.user_email),T(o.user_id)})},[n,F]),(0,o.jsx)("div",{className:"mx-auto w-full max-w-md mt-10",children:(0,o.jsxs)(l.Z,{children:[(0,o.jsx)(h.Z,{className:"text-sm mb-5 text-center",children:"\uD83D\uDE85 LiteLLM"}),(0,o.jsx)(h.Z,{className:"text-xl",children:"Sign up"}),(0,o.jsx)(d.Z,{children:"Claim your user account to login to Admin UI."}),(0,o.jsx)(i.Z,{className:"mt-4",title:"SSO",icon:x.GH$,color:"sky",children:(0,o.jsxs)(u.Z,{numItems:2,className:"flex justify-between items-center",children:[(0,o.jsx)(r.Z,{children:"SSO is under the Enterprise Tier."}),(0,o.jsx)(r.Z,{children:(0,o.jsx)(c.Z,{variant:"primary",className:"mb-2",children:(0,o.jsx)("a",{href:"https://forms.gle/W3U4PZpJGFHWtHyA9",target:"_blank",children:"Get Free Trial"})})})]})}),(0,o.jsxs)(g.Z,{className:"mt-10 mb-5 mx-auto",layout:"vertical",onFinish:e=>{console.log("in handle submit. accessToken:",S,"token:",C,"formValues:",e),S&&C&&(e.user_email=b,y&&n&&(0,f.m_)(S,n,y,e.password).then(e=>{let t="/ui/";t+="?login=success",document.cookie="token="+C,console.log("redirecting to:",t);let n=(0,f.zX)();console.log("proxyBaseUrl:",n),n?window.location.href=n+t:window.location.href=t}))},children:[(0,o.jsxs)(o.Fragment,{children:[(0,o.jsx)(g.Z.Item,{label:"Email Address",name:"user_email",children:(0,o.jsx)(m.Z,{type:"email",disabled:!0,value:b,defaultValue:b,className:"max-w-md"})}),(0,o.jsx)(g.Z.Item,{label:"Password",name:"password",rules:[{required:!0,message:"password required to sign up"}],help:"Create a password for your account",children:(0,o.jsx)(m.Z,{placeholder:"",type:"password",className:"max-w-md"})})]}),(0,o.jsx)("div",{className:"mt-10",children:(0,o.jsx)(k.ZP,{htmlType:"submit",children:"Sign Up"})})]})]})})}},3914:function(e,t,n){"use strict";function o(){let e=window.location.hostname,t=["Lax","Strict","None"];["/","/ui"].forEach(n=>{document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,";"),document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; domain=").concat(e,";"),t.forEach(t=>{let o="None"===t?" Secure;":"";document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; SameSite=").concat(t,";").concat(o),document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; domain=").concat(e,"; SameSite=").concat(t,";").concat(o)})}),console.log("After clearing cookies:",document.cookie)}function s(e){let t=document.cookie.split("; ").find(t=>t.startsWith(e+"="));return t?t.split("=")[1]:null}n.d(t,{b:function(){return o},e:function(){return s}})}},function(e){e.O(0,[665,402,899,250,971,117,744],function(){return e(e.s=8672)}),_N_E=e.O()}]);

View file

@ -1 +0,0 @@
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[461],{32922:function(e,t,n){Promise.resolve().then(n.bind(n,12011))},12011:function(e,t,n){"use strict";n.r(t),n.d(t,{default:function(){return S}});var s=n(57437),o=n(2265),a=n(99376),c=n(20831),i=n(94789),l=n(12514),r=n(49804),u=n(67101),m=n(84264),d=n(49566),h=n(96761),x=n(84566),p=n(19250),f=n(14474),k=n(13634),g=n(73002),j=n(3914);function S(){let[e]=k.Z.useForm(),t=(0,a.useSearchParams)();(0,j.e)("token");let n=t.get("invitation_id"),[S,w]=(0,o.useState)(null),[Z,_]=(0,o.useState)(""),[N,b]=(0,o.useState)(""),[T,y]=(0,o.useState)(null),[E,v]=(0,o.useState)(""),[C,U]=(0,o.useState)("");return(0,o.useEffect)(()=>{n&&(0,p.W_)(n).then(e=>{let t=e.login_url;console.log("login_url:",t),v(t);let n=e.token,s=(0,f.o)(n);U(n),console.log("decoded:",s),w(s.key),console.log("decoded user email:",s.user_email),b(s.user_email),y(s.user_id)})},[n]),(0,s.jsx)("div",{className:"mx-auto w-full max-w-md mt-10",children:(0,s.jsxs)(l.Z,{children:[(0,s.jsx)(h.Z,{className:"text-sm mb-5 text-center",children:"\uD83D\uDE85 LiteLLM"}),(0,s.jsx)(h.Z,{className:"text-xl",children:"Sign up"}),(0,s.jsx)(m.Z,{children:"Claim your user account to login to Admin UI."}),(0,s.jsx)(i.Z,{className:"mt-4",title:"SSO",icon:x.GH$,color:"sky",children:(0,s.jsxs)(u.Z,{numItems:2,className:"flex justify-between items-center",children:[(0,s.jsx)(r.Z,{children:"SSO is under the Enterprise Tier."}),(0,s.jsx)(r.Z,{children:(0,s.jsx)(c.Z,{variant:"primary",className:"mb-2",children:(0,s.jsx)("a",{href:"https://forms.gle/W3U4PZpJGFHWtHyA9",target:"_blank",children:"Get Free Trial"})})})]})}),(0,s.jsxs)(k.Z,{className:"mt-10 mb-5 mx-auto",layout:"vertical",onFinish:e=>{console.log("in handle submit. accessToken:",S,"token:",C,"formValues:",e),S&&C&&(e.user_email=N,T&&n&&(0,p.m_)(S,n,T,e.password).then(e=>{let t="/ui/";t+="?login=success",document.cookie="token="+C,console.log("redirecting to:",t),window.location.href=t}))},children:[(0,s.jsxs)(s.Fragment,{children:[(0,s.jsx)(k.Z.Item,{label:"Email Address",name:"user_email",children:(0,s.jsx)(d.Z,{type:"email",disabled:!0,value:N,defaultValue:N,className:"max-w-md"})}),(0,s.jsx)(k.Z.Item,{label:"Password",name:"password",rules:[{required:!0,message:"password required to sign up"}],help:"Create a password for your account",children:(0,s.jsx)(d.Z,{placeholder:"",type:"password",className:"max-w-md"})})]}),(0,s.jsx)("div",{className:"mt-10",children:(0,s.jsx)(g.ZP,{htmlType:"submit",children:"Sign Up"})})]})]})})}},3914:function(e,t,n){"use strict";function s(){let e=window.location.hostname,t=["Lax","Strict","None"];["/","/ui"].forEach(n=>{document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,";"),document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; domain=").concat(e,";"),t.forEach(t=>{let s="None"===t?" Secure;":"";document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; SameSite=").concat(t,";").concat(s),document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; domain=").concat(e,"; SameSite=").concat(t,";").concat(s)})}),console.log("After clearing cookies:",document.cookie)}function o(e){let t=document.cookie.split("; ").find(t=>t.startsWith(e+"="));return t?t.split("=")[1]:null}n.d(t,{b:function(){return s},e:function(){return o}})}},function(e){e.O(0,[665,402,899,250,971,117,744],function(){return e(e.s=32922)}),_N_E=e.O()}]);

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View file

@ -1 +1 @@
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[744],{20169:function(e,n,t){Promise.resolve().then(t.t.bind(t,12846,23)),Promise.resolve().then(t.t.bind(t,19107,23)),Promise.resolve().then(t.t.bind(t,61060,23)),Promise.resolve().then(t.t.bind(t,4707,23)),Promise.resolve().then(t.t.bind(t,80,23)),Promise.resolve().then(t.t.bind(t,36423,23))}},function(e){var n=function(n){return e(e.s=n)};e.O(0,[971,117],function(){return n(54278),n(20169)}),_N_E=e.O()}]);
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[744],{10264:function(e,n,t){Promise.resolve().then(t.t.bind(t,12846,23)),Promise.resolve().then(t.t.bind(t,19107,23)),Promise.resolve().then(t.t.bind(t,61060,23)),Promise.resolve().then(t.t.bind(t,4707,23)),Promise.resolve().then(t.t.bind(t,80,23)),Promise.resolve().then(t.t.bind(t,36423,23))}},function(e){var n=function(n){return e(e.s=n)};e.O(0,[971,117],function(){return n(54278),n(10264)}),_N_E=e.O()}]);

File diff suppressed because one or more lines are too long

View file

@ -1 +1 @@
!function(){"use strict";var e,t,n,r,o,u,i,c,f,a={},l={};function d(e){var t=l[e];if(void 0!==t)return t.exports;var n=l[e]={id:e,loaded:!1,exports:{}},r=!0;try{a[e].call(n.exports,n,n.exports,d),r=!1}finally{r&&delete l[e]}return n.loaded=!0,n.exports}d.m=a,e=[],d.O=function(t,n,r,o){if(n){o=o||0;for(var u=e.length;u>0&&e[u-1][2]>o;u--)e[u]=e[u-1];e[u]=[n,r,o];return}for(var i=1/0,u=0;u<e.length;u++){for(var n=e[u][0],r=e[u][1],o=e[u][2],c=!0,f=0;f<n.length;f++)i>=o&&Object.keys(d.O).every(function(e){return d.O[e](n[f])})?n.splice(f--,1):(c=!1,o<i&&(i=o));if(c){e.splice(u--,1);var a=r();void 0!==a&&(t=a)}}return t},d.n=function(e){var t=e&&e.__esModule?function(){return e.default}:function(){return e};return d.d(t,{a:t}),t},n=Object.getPrototypeOf?function(e){return Object.getPrototypeOf(e)}:function(e){return e.__proto__},d.t=function(e,r){if(1&r&&(e=this(e)),8&r||"object"==typeof e&&e&&(4&r&&e.__esModule||16&r&&"function"==typeof e.then))return e;var o=Object.create(null);d.r(o);var u={};t=t||[null,n({}),n([]),n(n)];for(var i=2&r&&e;"object"==typeof i&&!~t.indexOf(i);i=n(i))Object.getOwnPropertyNames(i).forEach(function(t){u[t]=function(){return e[t]}});return u.default=function(){return e},d.d(o,u),o},d.d=function(e,t){for(var n in t)d.o(t,n)&&!d.o(e,n)&&Object.defineProperty(e,n,{enumerable:!0,get:t[n]})},d.f={},d.e=function(e){return Promise.all(Object.keys(d.f).reduce(function(t,n){return d.f[n](e,t),t},[]))},d.u=function(e){},d.miniCssF=function(e){},d.g=function(){if("object"==typeof globalThis)return globalThis;try{return this||Function("return this")()}catch(e){if("object"==typeof window)return window}}(),d.o=function(e,t){return Object.prototype.hasOwnProperty.call(e,t)},r={},o="_N_E:",d.l=function(e,t,n,u){if(r[e]){r[e].push(t);return}if(void 0!==n)for(var i,c,f=document.getElementsByTagName("script"),a=0;a<f.length;a++){var l=f[a];if(l.getAttribute("src")==e||l.getAttribute("data-webpack")==o+n){i=l;break}}i||(c=!0,(i=document.createElement("script")).charset="utf-8",i.timeout=120,d.nc&&i.setAttribute("nonce",d.nc),i.setAttribute("data-webpack",o+n),i.src=d.tu(e)),r[e]=[t];var s=function(t,n){i.onerror=i.onload=null,clearTimeout(p);var o=r[e];if(delete r[e],i.parentNode&&i.parentNode.removeChild(i),o&&o.forEach(function(e){return e(n)}),t)return t(n)},p=setTimeout(s.bind(null,void 0,{type:"timeout",target:i}),12e4);i.onerror=s.bind(null,i.onerror),i.onload=s.bind(null,i.onload),c&&document.head.appendChild(i)},d.r=function(e){"undefined"!=typeof Symbol&&Symbol.toStringTag&&Object.defineProperty(e,Symbol.toStringTag,{value:"Module"}),Object.defineProperty(e,"__esModule",{value:!0})},d.nmd=function(e){return e.paths=[],e.children||(e.children=[]),e},d.tt=function(){return void 0===u&&(u={createScriptURL:function(e){return e}},"undefined"!=typeof trustedTypes&&trustedTypes.createPolicy&&(u=trustedTypes.createPolicy("nextjs#bundler",u))),u},d.tu=function(e){return d.tt().createScriptURL(e)},d.p="/ui/_next/",i={272:0,919:0,191:0},d.f.j=function(e,t){var n=d.o(i,e)?i[e]:void 0;if(0!==n){if(n)t.push(n[2]);else if(/^(191|272|919)$/.test(e))i[e]=0;else{var r=new Promise(function(t,r){n=i[e]=[t,r]});t.push(n[2]=r);var o=d.p+d.u(e),u=Error();d.l(o,function(t){if(d.o(i,e)&&(0!==(n=i[e])&&(i[e]=void 0),n)){var r=t&&("load"===t.type?"missing":t.type),o=t&&t.target&&t.target.src;u.message="Loading chunk "+e+" failed.\n("+r+": "+o+")",u.name="ChunkLoadError",u.type=r,u.request=o,n[1](u)}},"chunk-"+e,e)}}},d.O.j=function(e){return 0===i[e]},c=function(e,t){var n,r,o=t[0],u=t[1],c=t[2],f=0;if(o.some(function(e){return 0!==i[e]})){for(n in u)d.o(u,n)&&(d.m[n]=u[n]);if(c)var a=c(d)}for(e&&e(t);f<o.length;f++)r=o[f],d.o(i,r)&&i[r]&&i[r][0](),i[r]=0;return d.O(a)},(f=self.webpackChunk_N_E=self.webpackChunk_N_E||[]).forEach(c.bind(null,0)),f.push=c.bind(null,f.push.bind(f))}();
!function(){"use strict";var e,t,n,r,o,u,i,c,f,a={},l={};function d(e){var t=l[e];if(void 0!==t)return t.exports;var n=l[e]={id:e,loaded:!1,exports:{}},r=!0;try{a[e].call(n.exports,n,n.exports,d),r=!1}finally{r&&delete l[e]}return n.loaded=!0,n.exports}d.m=a,e=[],d.O=function(t,n,r,o){if(n){o=o||0;for(var u=e.length;u>0&&e[u-1][2]>o;u--)e[u]=e[u-1];e[u]=[n,r,o];return}for(var i=1/0,u=0;u<e.length;u++){for(var n=e[u][0],r=e[u][1],o=e[u][2],c=!0,f=0;f<n.length;f++)i>=o&&Object.keys(d.O).every(function(e){return d.O[e](n[f])})?n.splice(f--,1):(c=!1,o<i&&(i=o));if(c){e.splice(u--,1);var a=r();void 0!==a&&(t=a)}}return t},d.n=function(e){var t=e&&e.__esModule?function(){return e.default}:function(){return e};return d.d(t,{a:t}),t},n=Object.getPrototypeOf?function(e){return Object.getPrototypeOf(e)}:function(e){return e.__proto__},d.t=function(e,r){if(1&r&&(e=this(e)),8&r||"object"==typeof e&&e&&(4&r&&e.__esModule||16&r&&"function"==typeof e.then))return e;var o=Object.create(null);d.r(o);var u={};t=t||[null,n({}),n([]),n(n)];for(var i=2&r&&e;"object"==typeof i&&!~t.indexOf(i);i=n(i))Object.getOwnPropertyNames(i).forEach(function(t){u[t]=function(){return e[t]}});return u.default=function(){return e},d.d(o,u),o},d.d=function(e,t){for(var n in t)d.o(t,n)&&!d.o(e,n)&&Object.defineProperty(e,n,{enumerable:!0,get:t[n]})},d.f={},d.e=function(e){return Promise.all(Object.keys(d.f).reduce(function(t,n){return d.f[n](e,t),t},[]))},d.u=function(e){},d.miniCssF=function(e){},d.g=function(){if("object"==typeof globalThis)return globalThis;try{return this||Function("return this")()}catch(e){if("object"==typeof window)return window}}(),d.o=function(e,t){return Object.prototype.hasOwnProperty.call(e,t)},r={},o="_N_E:",d.l=function(e,t,n,u){if(r[e]){r[e].push(t);return}if(void 0!==n)for(var i,c,f=document.getElementsByTagName("script"),a=0;a<f.length;a++){var l=f[a];if(l.getAttribute("src")==e||l.getAttribute("data-webpack")==o+n){i=l;break}}i||(c=!0,(i=document.createElement("script")).charset="utf-8",i.timeout=120,d.nc&&i.setAttribute("nonce",d.nc),i.setAttribute("data-webpack",o+n),i.src=d.tu(e)),r[e]=[t];var s=function(t,n){i.onerror=i.onload=null,clearTimeout(p);var o=r[e];if(delete r[e],i.parentNode&&i.parentNode.removeChild(i),o&&o.forEach(function(e){return e(n)}),t)return t(n)},p=setTimeout(s.bind(null,void 0,{type:"timeout",target:i}),12e4);i.onerror=s.bind(null,i.onerror),i.onload=s.bind(null,i.onload),c&&document.head.appendChild(i)},d.r=function(e){"undefined"!=typeof Symbol&&Symbol.toStringTag&&Object.defineProperty(e,Symbol.toStringTag,{value:"Module"}),Object.defineProperty(e,"__esModule",{value:!0})},d.nmd=function(e){return e.paths=[],e.children||(e.children=[]),e},d.tt=function(){return void 0===u&&(u={createScriptURL:function(e){return e}},"undefined"!=typeof trustedTypes&&trustedTypes.createPolicy&&(u=trustedTypes.createPolicy("nextjs#bundler",u))),u},d.tu=function(e){return d.tt().createScriptURL(e)},d.p="/litellm/_next/",i={272:0,919:0,986:0},d.f.j=function(e,t){var n=d.o(i,e)?i[e]:void 0;if(0!==n){if(n)t.push(n[2]);else if(/^(272|919|986)$/.test(e))i[e]=0;else{var r=new Promise(function(t,r){n=i[e]=[t,r]});t.push(n[2]=r);var o=d.p+d.u(e),u=Error();d.l(o,function(t){if(d.o(i,e)&&(0!==(n=i[e])&&(i[e]=void 0),n)){var r=t&&("load"===t.type?"missing":t.type),o=t&&t.target&&t.target.src;u.message="Loading chunk "+e+" failed.\n("+r+": "+o+")",u.name="ChunkLoadError",u.type=r,u.request=o,n[1](u)}},"chunk-"+e,e)}}},d.O.j=function(e){return 0===i[e]},c=function(e,t){var n,r,o=t[0],u=t[1],c=t[2],f=0;if(o.some(function(e){return 0!==i[e]})){for(n in u)d.o(u,n)&&(d.m[n]=u[n]);if(c)var a=c(d)}for(e&&e(t);f<o.length;f++)r=o[f],d.o(i,r)&&i[r]&&i[r][0](),i[r]=0;return d.O(a)},(f=self.webpackChunk_N_E=self.webpackChunk_N_E||[]).forEach(c.bind(null,0)),f.push=c.bind(null,f.push.bind(f))}();

File diff suppressed because one or more lines are too long

View file

@ -1 +0,0 @@
@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/55c55f0601d81cf3-s.woff2) format("woff2");unicode-range:u+0460-052f,u+1c80-1c8a,u+20b4,u+2de0-2dff,u+a640-a69f,u+fe2e-fe2f}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/26a46d62cd723877-s.woff2) format("woff2");unicode-range:u+0301,u+0400-045f,u+0490-0491,u+04b0-04b1,u+2116}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/97e0cb1ae144a2a9-s.woff2) format("woff2");unicode-range:u+1f??}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/581909926a08bbc8-s.woff2) format("woff2");unicode-range:u+0370-0377,u+037a-037f,u+0384-038a,u+038c,u+038e-03a1,u+03a3-03ff}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/df0a9ae256c0569c-s.woff2) format("woff2");unicode-range:u+0102-0103,u+0110-0111,u+0128-0129,u+0168-0169,u+01a0-01a1,u+01af-01b0,u+0300-0301,u+0303-0304,u+0308-0309,u+0323,u+0329,u+1ea0-1ef9,u+20ab}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/6d93bde91c0c2823-s.woff2) format("woff2");unicode-range:u+0100-02ba,u+02bd-02c5,u+02c7-02cc,u+02ce-02d7,u+02dd-02ff,u+0304,u+0308,u+0329,u+1d00-1dbf,u+1e00-1e9f,u+1ef2-1eff,u+2020,u+20a0-20ab,u+20ad-20c0,u+2113,u+2c60-2c7f,u+a720-a7ff}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/a34f9d1faa5f3315-s.p.woff2) format("woff2");unicode-range:u+00??,u+0131,u+0152-0153,u+02bb-02bc,u+02c6,u+02da,u+02dc,u+0304,u+0308,u+0329,u+2000-206f,u+20ac,u+2122,u+2191,u+2193,u+2212,u+2215,u+feff,u+fffd}@font-face{font-family:__Inter_Fallback_cf7686;src:local("Arial");ascent-override:90.49%;descent-override:22.56%;line-gap-override:0.00%;size-adjust:107.06%}.__className_cf7686{font-family:__Inter_cf7686,__Inter_Fallback_cf7686;font-style:normal}

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1 @@
@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/55c55f0601d81cf3-s.woff2) format("woff2");unicode-range:u+0460-052f,u+1c80-1c8a,u+20b4,u+2de0-2dff,u+a640-a69f,u+fe2e-fe2f}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/26a46d62cd723877-s.woff2) format("woff2");unicode-range:u+0301,u+0400-045f,u+0490-0491,u+04b0-04b1,u+2116}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/97e0cb1ae144a2a9-s.woff2) format("woff2");unicode-range:u+1f??}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/581909926a08bbc8-s.woff2) format("woff2");unicode-range:u+0370-0377,u+037a-037f,u+0384-038a,u+038c,u+038e-03a1,u+03a3-03ff}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/df0a9ae256c0569c-s.woff2) format("woff2");unicode-range:u+0102-0103,u+0110-0111,u+0128-0129,u+0168-0169,u+01a0-01a1,u+01af-01b0,u+0300-0301,u+0303-0304,u+0308-0309,u+0323,u+0329,u+1ea0-1ef9,u+20ab}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/8e9860b6e62d6359-s.woff2) format("woff2");unicode-range:u+0100-02ba,u+02bd-02c5,u+02c7-02cc,u+02ce-02d7,u+02dd-02ff,u+0304,u+0308,u+0329,u+1d00-1dbf,u+1e00-1e9f,u+1ef2-1eff,u+2020,u+20a0-20ab,u+20ad-20c0,u+2113,u+2c60-2c7f,u+a720-a7ff}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/e4af272ccee01ff0-s.p.woff2) format("woff2");unicode-range:u+00??,u+0131,u+0152-0153,u+02bb-02bc,u+02c6,u+02da,u+02dc,u+0304,u+0308,u+0329,u+2000-206f,u+20ac,u+2122,u+2191,u+2193,u+2212,u+2215,u+feff,u+fffd}@font-face{font-family:__Inter_Fallback_3373e4;src:local("Arial");ascent-override:90.49%;descent-override:22.56%;line-gap-override:0.00%;size-adjust:107.06%}.__className_3373e4{font-family:__Inter_3373e4,__Inter_Fallback_3373e4;font-style:normal}

Some files were not shown because too many files have changed in this diff Show more