mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-10 03:28:53 +00:00
Merge branch 'BerriAI:main' into main
This commit is contained in:
commit
4bcedc6108
272 changed files with 9006 additions and 1475 deletions
|
|
@ -79,7 +79,7 @@ jobs:
|
|||
pip install "pytest-retry==1.6.3"
|
||||
pip install "pytest-asyncio==0.21.1"
|
||||
pip install "pytest-cov==5.0.0"
|
||||
pip install mypy
|
||||
pip install "mypy==1.15.0"
|
||||
pip install "google-generativeai==0.3.2"
|
||||
pip install "google-cloud-aiplatform==1.43.0"
|
||||
pip install pyarrow
|
||||
|
|
@ -1158,6 +1158,7 @@ jobs:
|
|||
pip install "google-cloud-aiplatform==1.43.0"
|
||||
pip install "mlflow==2.17.2"
|
||||
pip install "anthropic==0.52.0"
|
||||
pip install "blockbuster==1.5.24"
|
||||
# Run pytest and generate JUnit XML report
|
||||
- setup_litellm_enterprise_pip
|
||||
- run:
|
||||
|
|
|
|||
2
.github/workflows/test-litellm.yml
vendored
2
.github/workflows/test-litellm.yml
vendored
|
|
@ -7,7 +7,7 @@ on:
|
|||
jobs:
|
||||
test:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 8
|
||||
timeout-minutes: 15
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
|
|
|||
|
|
@ -78,8 +78,9 @@ curl http://localhost:4000/v1/batches \
|
|||
**Create File for Batch Completion**
|
||||
|
||||
```python
|
||||
from litellm
|
||||
import litellm
|
||||
import os
|
||||
import asyncio
|
||||
|
||||
os.environ["OPENAI_API_KEY"] = "sk-.."
|
||||
|
||||
|
|
@ -97,8 +98,9 @@ print("Response from creating file=", file_obj)
|
|||
**Create Batch Request**
|
||||
|
||||
```python
|
||||
from litellm
|
||||
import litellm
|
||||
import os
|
||||
import asyncio
|
||||
|
||||
create_batch_response = await litellm.acreate_batch(
|
||||
completion_window="24h",
|
||||
|
|
|
|||
|
|
@ -33,11 +33,11 @@ cd litellm/ui/litellm-dashboard
|
|||
|
||||
npm run dev
|
||||
|
||||
# starts on http://0.0.0.0:3000/ui
|
||||
# starts on http://0.0.0.0:3000
|
||||
```
|
||||
|
||||
## 3. Go to local UI
|
||||
|
||||
```
|
||||
http://0.0.0.0:3000/ui
|
||||
```bash
|
||||
http://0.0.0.0:3000
|
||||
```
|
||||
|
|
@ -13,7 +13,7 @@ Here are the core requirements for any PR submitted to LiteLLM
|
|||
|
||||
## **Contributor License Agreement (CLA)**
|
||||
|
||||
Before contributing code to LiteLLM, you must sign our [Contributor License Agreement (CLA)](<(https://cla-assistant.io/BerriAI/litellm)>). This is a legal requirement for all contributions to be merged into the main repository. The CLA helps protect both you and the project by clearly defining the terms under which your contributions are made.
|
||||
Before contributing code to LiteLLM, you must sign our [Contributor License Agreement (CLA)](https://cla-assistant.io/BerriAI/litellm). This is a legal requirement for all contributions to be merged into the main repository. The CLA helps protect both you and the project by clearly defining the terms under which your contributions are made.
|
||||
|
||||
**Important:** We strongly recommend reviewing and signing the CLA before starting work on your contribution to avoid any delays in the PR process. You can find the CLA [here](https://cla-assistant.io/BerriAI/litellm) and sign it through our CLA management system when you submit your first PR.
|
||||
|
||||
|
|
|
|||
|
|
@ -14,7 +14,8 @@ LiteLLM provides image editing functionality that maps to OpenAI's `/images/edit
|
|||
| Fallbacks | ✅ | Works between supported models |
|
||||
| Loadbalancing | ✅ | Works between supported models |
|
||||
| Supported operations | Create image edits | |
|
||||
| Supported LiteLLM Versions | 1.63.8+ | |
|
||||
| Supported LiteLLM SDK Versions | 1.63.8+ | |
|
||||
| Supported LiteLLM Proxy Versions | 1.71.1+ | |
|
||||
| Supported LLM providers | **OpenAI** | Currently only `openai` is supported |
|
||||
|
||||
## Usage
|
||||
|
|
|
|||
202
docs/my-website/docs/providers/bedrock_agents.md
Normal file
202
docs/my-website/docs/providers/bedrock_agents.md
Normal file
|
|
@ -0,0 +1,202 @@
|
|||
import Tabs from '@theme/Tabs';
|
||||
import TabItem from '@theme/TabItem';
|
||||
|
||||
# Bedrock Agents
|
||||
|
||||
Call Bedrock Agents in the OpenAI Request/Response format.
|
||||
|
||||
|
||||
| Property | Details |
|
||||
|----------|---------|
|
||||
| Description | Amazon Bedrock Agents use the reasoning of foundation models (FMs), APIs, and data to break down user requests, gather relevant information, and efficiently complete tasks. |
|
||||
| Provider Route on LiteLLM | `bedrock/agent/{AGENT_ID}/{ALIAS_ID}` |
|
||||
| Provider Doc | [AWS Bedrock Agents ↗](https://aws.amazon.com/bedrock/agents/) |
|
||||
|
||||
## Quick Start
|
||||
|
||||
### Model Format to LiteLLM
|
||||
|
||||
To call a bedrock agent through LiteLLM, you need to use the following model format to call the agent.
|
||||
|
||||
Here the `model=bedrock/agent/` tells LiteLLM to call the bedrock `InvokeAgent` API.
|
||||
|
||||
```shell showLineNumbers title="Model Format to LiteLLM"
|
||||
bedrock/agent/{AGENT_ID}/{ALIAS_ID}
|
||||
```
|
||||
|
||||
**Example:**
|
||||
- `bedrock/agent/L1RT58GYRW/MFPSBCXYTW`
|
||||
- `bedrock/agent/ABCD1234/LIVE`
|
||||
|
||||
You can find these IDs in your AWS Bedrock console under Agents.
|
||||
|
||||
|
||||
### LiteLLM Python SDK
|
||||
|
||||
```python showLineNumbers title="Basic Agent Completion"
|
||||
import litellm
|
||||
|
||||
# Make a completion request to your Bedrock Agent
|
||||
response = litellm.completion(
|
||||
model="bedrock/agent/L1RT58GYRW/MFPSBCXYTW", # agent/{AGENT_ID}/{ALIAS_ID}
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Hi, I need help with analyzing our Q3 sales data and generating a summary report"
|
||||
}
|
||||
],
|
||||
)
|
||||
|
||||
print(response.choices[0].message.content)
|
||||
print(f"Response cost: ${response._hidden_params['response_cost']}")
|
||||
```
|
||||
|
||||
```python showLineNumbers title="Streaming Agent Responses"
|
||||
import litellm
|
||||
|
||||
# Stream responses from your Bedrock Agent
|
||||
response = litellm.completion(
|
||||
model="bedrock/agent/L1RT58GYRW/MFPSBCXYTW",
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Can you help me plan a marketing campaign and provide step-by-step execution details?"
|
||||
}
|
||||
],
|
||||
stream=True,
|
||||
)
|
||||
|
||||
for chunk in response:
|
||||
if chunk.choices[0].delta.content:
|
||||
print(chunk.choices[0].delta.content, end="")
|
||||
```
|
||||
|
||||
|
||||
### LiteLLM Proxy
|
||||
|
||||
#### 1. Configure your model in config.yaml
|
||||
|
||||
<Tabs>
|
||||
<TabItem value="config-yaml" label="config.yaml">
|
||||
|
||||
```yaml showLineNumbers title="LiteLLM Proxy Configuration"
|
||||
model_list:
|
||||
- model_name: bedrock-agent-1
|
||||
litellm_params:
|
||||
model: bedrock/agent/L1RT58GYRW/MFPSBCXYTW
|
||||
aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
|
||||
aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
|
||||
aws_region_name: us-west-2
|
||||
|
||||
- model_name: bedrock-agent-2
|
||||
litellm_params:
|
||||
model: bedrock/agent/AGENT456/ALIAS789
|
||||
aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
|
||||
aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
|
||||
aws_region_name: us-east-1
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
</Tabs>
|
||||
|
||||
#### 2. Start the LiteLLM Proxy
|
||||
|
||||
```bash showLineNumbers title="Start LiteLLM Proxy"
|
||||
litellm --config config.yaml
|
||||
```
|
||||
|
||||
#### 3. Make requests to your Bedrock Agents
|
||||
|
||||
<Tabs>
|
||||
<TabItem value="curl" label="Curl">
|
||||
|
||||
```bash showLineNumbers title="Basic Agent Request"
|
||||
curl http://localhost:4000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer $LITELLM_API_KEY" \
|
||||
-d '{
|
||||
"model": "bedrock-agent-1",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Analyze our customer data and suggest retention strategies"
|
||||
}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
```bash showLineNumbers title="Streaming Agent Request"
|
||||
curl http://localhost:4000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer $LITELLM_API_KEY" \
|
||||
-d '{
|
||||
"model": "bedrock-agent-2",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Create a comprehensive social media strategy for our new product"
|
||||
}
|
||||
],
|
||||
"stream": true
|
||||
}'
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
|
||||
<TabItem value="openai-sdk" label="OpenAI Python SDK">
|
||||
|
||||
```python showLineNumbers title="Using OpenAI SDK with LiteLLM Proxy"
|
||||
from openai import OpenAI
|
||||
|
||||
# Initialize client with your LiteLLM proxy URL
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:4000",
|
||||
api_key="your-litellm-api-key"
|
||||
)
|
||||
|
||||
# Make a completion request to your agent
|
||||
response = client.chat.completions.create(
|
||||
model="bedrock-agent-1",
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Help me prepare for the quarterly business review meeting"
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
print(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
```python showLineNumbers title="Streaming with OpenAI SDK"
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:4000",
|
||||
api_key="your-litellm-api-key"
|
||||
)
|
||||
|
||||
# Stream agent responses
|
||||
stream = client.chat.completions.create(
|
||||
model="bedrock-agent-2",
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Walk me through launching a new feature beta program"
|
||||
}
|
||||
],
|
||||
stream=True
|
||||
)
|
||||
|
||||
for chunk in stream:
|
||||
if chunk.choices[0].delta.content is not None:
|
||||
print(chunk.choices[0].delta.content, end="")
|
||||
```
|
||||
|
||||
</TabItem>
|
||||
</Tabs>
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [AWS Bedrock Agents Documentation](https://aws.amazon.com/bedrock/agents/)
|
||||
- [LiteLLM Authentication to Bedrock](https://docs.litellm.ai/docs/providers/bedrock#boto3---authentication)
|
||||
|
|
@ -371,6 +371,7 @@ router_settings:
|
|||
| DD_API_KEY | API key for Datadog integration
|
||||
| DD_SITE | Site URL for Datadog (e.g., datadoghq.com)
|
||||
| DD_SOURCE | Source identifier for Datadog logs
|
||||
| DD_TRACER_STREAMING_CHUNK_YIELD_RESOURCE | Resource name for Datadog tracing of streaming chunk yields. Default is "streaming.chunk.yield"
|
||||
| DD_ENV | Environment identifier for Datadog logs. Only supported for `datadog_llm_observability` callback
|
||||
| DD_SERVICE | Service identifier for Datadog logs. Defaults to "litellm-server"
|
||||
| DD_VERSION | Version identifier for Datadog logs. Defaults to "unknown"
|
||||
|
|
@ -406,11 +407,14 @@ router_settings:
|
|||
| DEFAULT_REPLICATE_GPU_PRICE_PER_SECOND | Default price per second for Replicate GPU. Default is 0.001400
|
||||
| DEFAULT_REPLICATE_POLLING_DELAY_SECONDS | Default delay in seconds for Replicate polling. Default is 1
|
||||
| DEFAULT_REPLICATE_POLLING_RETRIES | Default number of retries for Replicate polling. Default is 5
|
||||
| DEFAULT_S3_BATCH_SIZE | Default batch size for S3 logging. Default is 512
|
||||
| DEFAULT_S3_FLUSH_INTERVAL_SECONDS | Default flush interval for S3 logging. Default is 10
|
||||
| DEFAULT_SLACK_ALERTING_THRESHOLD | Default threshold for Slack alerting. Default is 300
|
||||
| DEFAULT_SOFT_BUDGET | Default soft budget for LiteLLM proxy keys. Default is 50.0
|
||||
| DEFAULT_TRIM_RATIO | Default ratio of tokens to trim from prompt end. Default is 0.75
|
||||
| DIRECT_URL | Direct URL for service endpoint
|
||||
| DISABLE_ADMIN_UI | Toggle to disable the admin UI
|
||||
| DISABLE_AIOHTTP_TRANSPORT | Flag to disable aiohttp transport. When this is set to True, litellm will use httpx instead of aiohttp. **Default is False**
|
||||
| DISABLE_SCHEMA_UPDATE | Toggle to disable schema updates
|
||||
| DOCS_DESCRIPTION | Description text for documentation pages
|
||||
| DOCS_FILTERED | Flag indicating filtered documentation
|
||||
|
|
@ -642,7 +646,6 @@ router_settings:
|
|||
| UPSTREAM_LANGFUSE_PUBLIC_KEY | Public key for upstream Langfuse authentication
|
||||
| UPSTREAM_LANGFUSE_RELEASE | Release version identifier for upstream Langfuse
|
||||
| UPSTREAM_LANGFUSE_SECRET_KEY | Secret key for upstream Langfuse authentication
|
||||
| USE_AIOHTTP_TRANSPORT | Flag to enable aiohttp transport. This is a feature flag for the new aiohttp transport. **Default is False**
|
||||
| USE_AWS_KMS | Flag to enable AWS Key Management Service for encryption
|
||||
| USE_PRISMA_MIGRATE | Flag to use prisma migrate instead of prisma db push. Recommended for production environments.
|
||||
| WEBHOOK_URL | URL for receiving webhooks from external services
|
||||
|
|
|
|||
59
docs/my-website/docs/proxy/custom_root_ui.md
Normal file
59
docs/my-website/docs/proxy/custom_root_ui.md
Normal file
|
|
@ -0,0 +1,59 @@
|
|||
# UI - Custom Root Path
|
||||
|
||||
💥 Use this when you want to serve LiteLLM on a custom base url path like `https://localhost:4000/api/v1`
|
||||
|
||||
## Usage
|
||||
|
||||
### 1. Set `SERVER_ROOT_PATH` in your .env
|
||||
|
||||
👉 Set `SERVER_ROOT_PATH` in your .env and this will be set as your server root path
|
||||
|
||||
```
|
||||
export SERVER_ROOT_PATH="/api/v1"
|
||||
```
|
||||
|
||||
### 2. Run the Proxy
|
||||
|
||||
```shell
|
||||
litellm proxy --config /path/to/config.yaml
|
||||
```
|
||||
|
||||
After running the proxy you can access it on `http://0.0.0.0:4000/api/v1/` (since we set `SERVER_ROOT_PATH="/api/v1"`)
|
||||
|
||||
|
||||
### 3. Reserve the `/litellm` path
|
||||
|
||||
LiteLLM uses the `/litellm` path to discover the custom root path. So you need to reserve this path in your proxy.
|
||||
|
||||
If you are running the UI, it will query the `/litellm/.well-known/litellm-ui-config` endpoint to get the UI configuration.
|
||||
|
||||
So you need to reserve the `/litellm` path in your proxy.
|
||||
|
||||
You can see the results with:
|
||||
|
||||
```bash
|
||||
curl http://0.0.0.0:4000/litellm/.well-known/litellm-ui-config
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
```json
|
||||
|
||||
{
|
||||
"server_root_path": "/api/v1",
|
||||
...
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
|
||||
### 4. Verify Running on correct path
|
||||
|
||||
<Image img={require('../../img/custom_root_path.png')} />
|
||||
|
||||
**That's it**, that's all you need to run the proxy on a custom root path
|
||||
|
||||
|
||||
## Demo
|
||||
|
||||
[Here's a demo video](https://drive.google.com/file/d/1zqAxI0lmzNp7IJH1dxlLuKqX2xi3F_R3/view?usp=sharing) of running the proxy on a custom root path
|
||||
|
|
@ -619,101 +619,8 @@ docker pull ghcr.io/berriai/litellm-non_root:main-stable
|
|||
|
||||
### 1. Custom server root path (Proxy base url)
|
||||
|
||||
💥 Use this when you want to serve LiteLLM on a custom base url path like `https://localhost:4000/api/v1`
|
||||
Refer to [Custom Root Path](./custom_root_ui) for more details.
|
||||
|
||||
:::info
|
||||
|
||||
In a Kubernetes deployment, it's possible to utilize a shared DNS to host multiple applications by modifying the virtual service
|
||||
|
||||
:::
|
||||
|
||||
Customize the root path to eliminate the need for employing multiple DNS configurations during deployment.
|
||||
|
||||
Step 1.
|
||||
👉 Set `SERVER_ROOT_PATH` in your .env and this will be set as your server root path
|
||||
```
|
||||
export SERVER_ROOT_PATH="/api/v1"
|
||||
```
|
||||
|
||||
**Step 2** (If you want the Proxy Admin UI to work with your root path you need to use this dockerfile)
|
||||
- Use the dockerfile below (it uses litellm as a base image)
|
||||
- 👉 Set `UI_BASE_PATH=$SERVER_ROOT_PATH/ui` in the Dockerfile, example `UI_BASE_PATH=/api/v1/ui`
|
||||
|
||||
Dockerfile
|
||||
|
||||
```shell
|
||||
# Use the provided base image
|
||||
FROM ghcr.io/berriai/litellm:main-latest
|
||||
|
||||
# Set the working directory to /app
|
||||
WORKDIR /app
|
||||
|
||||
# Install Node.js and npm (adjust version as needed)
|
||||
RUN apt-get update && apt-get install -y nodejs npm
|
||||
|
||||
# Copy the UI source into the container
|
||||
COPY ./ui/litellm-dashboard /app/ui/litellm-dashboard
|
||||
|
||||
# Set an environment variable for UI_BASE_PATH
|
||||
# This can be overridden at build time
|
||||
# set UI_BASE_PATH to "<your server root path>/ui"
|
||||
# 👇👇 Enter your UI_BASE_PATH here
|
||||
ENV UI_BASE_PATH="/api/v1/ui"
|
||||
|
||||
# Build the UI with the specified UI_BASE_PATH
|
||||
WORKDIR /app/ui/litellm-dashboard
|
||||
RUN npm install
|
||||
RUN UI_BASE_PATH=$UI_BASE_PATH npm run build
|
||||
|
||||
# Create the destination directory
|
||||
RUN mkdir -p /app/litellm/proxy/_experimental/out
|
||||
|
||||
# Move the built files to the appropriate location
|
||||
# Assuming the build output is in ./out directory
|
||||
RUN rm -rf /app/litellm/proxy/_experimental/out/* && \
|
||||
mv ./out/* /app/litellm/proxy/_experimental/out/
|
||||
|
||||
# Switch back to the main app directory
|
||||
WORKDIR /app
|
||||
|
||||
# Make sure your entrypoint.sh is executable
|
||||
RUN chmod +x ./docker/entrypoint.sh
|
||||
|
||||
# Expose the necessary port
|
||||
EXPOSE 4000/tcp
|
||||
|
||||
# Override the CMD instruction with your desired command and arguments
|
||||
# only use --detailed_debug for debugging
|
||||
CMD ["--port", "4000", "--config", "config.yaml"]
|
||||
```
|
||||
|
||||
**Step 3** build this Dockerfile
|
||||
|
||||
```shell
|
||||
docker build -f Dockerfile -t litellm-prod-build . --progress=plain
|
||||
```
|
||||
|
||||
**Step 4. Run Proxy with `SERVER_ROOT_PATH` set in your env **
|
||||
|
||||
```shell
|
||||
docker run \
|
||||
-v $(pwd)/proxy_config.yaml:/app/config.yaml \
|
||||
-p 4000:4000 \
|
||||
-e LITELLM_LOG="DEBUG"\
|
||||
-e SERVER_ROOT_PATH="/api/v1"\
|
||||
-e DATABASE_URL=postgresql://<user>:<password>@<host>:<port>/<dbname> \
|
||||
-e LITELLM_MASTER_KEY="sk-1234"\
|
||||
litellm-prod-build \
|
||||
--config /app/config.yaml
|
||||
```
|
||||
|
||||
After running the proxy you can access it on `http://0.0.0.0:4000/api/v1/` (since we set `SERVER_ROOT_PATH="/api/v1"`)
|
||||
|
||||
**Step 5. Verify Running on correct path**
|
||||
|
||||
<Image img={require('../../img/custom_root_path.png')} />
|
||||
|
||||
**That's it**, that's all you need to run the proxy on a custom root path
|
||||
|
||||
### 2. SSL Certification
|
||||
|
||||
|
|
|
|||
|
|
@ -45,12 +45,12 @@ Setup your config.yaml with your azure model.
|
|||
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: gpt-3.5-turbo
|
||||
- model_name: gpt-4o
|
||||
litellm_params:
|
||||
model: azure/my_azure_deployment
|
||||
api_base: os.environ/AZURE_API_BASE
|
||||
api_key: "os.environ/AZURE_API_KEY"
|
||||
api_version: "2024-07-01-preview" # [OPTIONAL] litellm uses the latest azure api_version by default
|
||||
api_version: "2025-01-01-preview" # [OPTIONAL] litellm uses the latest azure api_version by default
|
||||
```
|
||||
---
|
||||
|
||||
|
|
@ -127,15 +127,15 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
|
|||
-H 'Content-Type: application/json' \
|
||||
-H 'Authorization: Bearer sk-1234' \
|
||||
-d '{
|
||||
"model": "gpt-3.5-turbo",
|
||||
"model": "gpt-4o",
|
||||
"messages": [
|
||||
{
|
||||
"role": "system",
|
||||
"content": "You are a helpful math tutor. Guide the user through the solution step by step."
|
||||
"content": "You are an LLM named gpt-4o"
|
||||
},
|
||||
{
|
||||
"role": "user",
|
||||
"content": "how can I solve 8x + 7 = -23"
|
||||
"content": "what is your name?"
|
||||
}
|
||||
]
|
||||
}'
|
||||
|
|
@ -145,28 +145,63 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
|
|||
|
||||
```bash
|
||||
{
|
||||
"id": "chatcmpl-2076f062-3095-4052-a520-7c321c115c68",
|
||||
"choices": [
|
||||
{
|
||||
"finish_reason": "stop",
|
||||
"index": 0,
|
||||
"message": {
|
||||
"content": "I am gpt-3.5-turbo",
|
||||
"role": "assistant",
|
||||
"tool_calls": null,
|
||||
"function_call": null
|
||||
}
|
||||
}
|
||||
],
|
||||
"created": 1724962831,
|
||||
"model": "gpt-3.5-turbo",
|
||||
"object": "chat.completion",
|
||||
"system_fingerprint": null,
|
||||
"usage": {
|
||||
"completion_tokens": 20,
|
||||
"prompt_tokens": 10,
|
||||
"total_tokens": 30
|
||||
"id": "chatcmpl-BcO8tRQmQV6Dfw6onqMufxPkLLkA8",
|
||||
"created": 1748488967,
|
||||
"model": "gpt-4o-2024-11-20",
|
||||
"object": "chat.completion",
|
||||
"system_fingerprint": "fp_ee1d74bde0",
|
||||
"choices": [
|
||||
{
|
||||
"finish_reason": "stop",
|
||||
"index": 0,
|
||||
"message": {
|
||||
"content": "My name is **gpt-4o**! How can I assist you today?",
|
||||
"role": "assistant",
|
||||
"tool_calls": null,
|
||||
"function_call": null,
|
||||
"annotations": []
|
||||
}
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"completion_tokens": 19,
|
||||
"prompt_tokens": 28,
|
||||
"total_tokens": 47,
|
||||
"completion_tokens_details": {
|
||||
"accepted_prediction_tokens": 0,
|
||||
"audio_tokens": 0,
|
||||
"reasoning_tokens": 0,
|
||||
"rejected_prediction_tokens": 0
|
||||
},
|
||||
"prompt_tokens_details": {
|
||||
"audio_tokens": 0,
|
||||
"cached_tokens": 0
|
||||
}
|
||||
},
|
||||
"service_tier": null,
|
||||
"prompt_filter_results": [
|
||||
{
|
||||
"prompt_index": 0,
|
||||
"content_filter_results": {
|
||||
"hate": {
|
||||
"filtered": false,
|
||||
"severity": "safe"
|
||||
},
|
||||
"self_harm": {
|
||||
"filtered": false,
|
||||
"severity": "safe"
|
||||
},
|
||||
"sexual": {
|
||||
"filtered": false,
|
||||
"severity": "safe"
|
||||
},
|
||||
"violence": {
|
||||
"filtered": false,
|
||||
"severity": "safe"
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
|
|
@ -191,12 +226,12 @@ Track Spend, and control model access via virtual keys for the proxy
|
|||
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: gpt-3.5-turbo
|
||||
- model_name: gpt-4o
|
||||
litellm_params:
|
||||
model: azure/my_azure_deployment
|
||||
api_base: os.environ/AZURE_API_BASE
|
||||
api_key: "os.environ/AZURE_API_KEY"
|
||||
api_version: "2024-07-01-preview" # [OPTIONAL] litellm uses the latest azure api_version by default
|
||||
api_version: "2025-01-01-preview" # [OPTIONAL] litellm uses the latest azure api_version by default
|
||||
|
||||
general_settings:
|
||||
master_key: sk-1234
|
||||
|
|
@ -276,7 +311,7 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
|
|||
-H 'Content-Type: application/json' \
|
||||
-H 'Authorization: Bearer sk-12...' \
|
||||
-d '{
|
||||
"model": "gpt-3.5-turbo",
|
||||
"model": "gpt-4o",
|
||||
"messages": [
|
||||
{
|
||||
"role": "system",
|
||||
|
|
@ -312,7 +347,7 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
|
|||
-H 'Content-Type: application/json' \
|
||||
-H 'Authorization: Bearer sk-12...' \
|
||||
-d '{
|
||||
"model": "gpt-3.5-turbo",
|
||||
"model": "gpt-4o",
|
||||
"messages": [
|
||||
{
|
||||
"role": "system",
|
||||
|
|
@ -331,7 +366,7 @@ curl -X POST 'http://0.0.0.0:4000/chat/completions' \
|
|||
```bash
|
||||
{
|
||||
"error": {
|
||||
"message": "Max parallel request limit reached. Hit limit for api_key: daa1b272072a4c6841470a488c5dad0f298ff506e1cc935f4a181eed90c182ad. tpm_limit: 100, current_tpm: 29, rpm_limit: 1, current_rpm: 2.",
|
||||
"message": "LiteLLM Rate Limit Handler for rate limit type = key. Crossed TPM / RPM / Max Parallel Request Limit. current rpm: 1, rpm limit: 1, current tpm: 348, tpm limit: 9223372036854775807, current max_parallel_requests: 0, max_parallel_requests: 9223372036854775807",
|
||||
"type": "None",
|
||||
"param": "None",
|
||||
"code": "429"
|
||||
|
|
@ -371,12 +406,12 @@ You can disable ssl verification with:
|
|||
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: gpt-3.5-turbo
|
||||
- model_name: gpt-4o
|
||||
litellm_params:
|
||||
model: azure/my_azure_deployment
|
||||
api_base: os.environ/AZURE_API_BASE
|
||||
api_key: "os.environ/AZURE_API_KEY"
|
||||
api_version: "2024-07-01-preview"
|
||||
api_version: "2025-01-01-preview"
|
||||
|
||||
litellm_settings:
|
||||
ssl_verify: false # 👈 KEY CHANGE
|
||||
|
|
|
|||
|
|
@ -13,6 +13,7 @@ import TabItem from '@theme/TabItem';
|
|||
| Supported Entity Types | All Presidio Entity Types |
|
||||
| Supported Actions | `MASK`, `BLOCK` |
|
||||
| Supported Modes | `pre_call`, `during_call`, `post_call`, `logging_only` |
|
||||
| Language Support | Configurable via `presidio_language` parameter (supports multiple languages including English, Spanish, German, etc.) |
|
||||
|
||||
## Deployment options
|
||||
|
||||
|
|
@ -48,6 +49,18 @@ Now select the entity types you want to mask. See the [supported actions here](#
|
|||
style={{width: '50%', display: 'block', margin: '0'}}
|
||||
/>
|
||||
|
||||
#### 1.3 Set Default Language (Optional)
|
||||
|
||||
You can also configure a default language for PII analysis using the `presidio_language` field in the UI. This sets the default language that will be used for all requests unless overridden by a per-request language setting.
|
||||
|
||||
**Supported language codes include:**
|
||||
- `en` - English (default)
|
||||
- `es` - Spanish
|
||||
- `de` - German
|
||||
|
||||
|
||||
If not specified, English (`en`) will be used as the default language.
|
||||
|
||||
</TabItem>
|
||||
|
||||
|
||||
|
|
@ -67,6 +80,7 @@ guardrails:
|
|||
litellm_params:
|
||||
guardrail: presidio # supported values: "aporia", "bedrock", "lakera", "presidio"
|
||||
mode: "pre_call"
|
||||
presidio_language: "en" # optional: set default language for PII analysis
|
||||
```
|
||||
|
||||
Set the following env vars
|
||||
|
|
@ -380,6 +394,86 @@ print(response)
|
|||
|
||||
</Tabs>
|
||||
|
||||
### Set default `language` in config.yaml
|
||||
|
||||
You can configure a default language for PII analysis in your YAML configuration using the `presidio_language` parameter. This language will be used for all requests unless overridden by a per-request language setting.
|
||||
|
||||
```yaml title="Default Language Configuration" showLineNumbers
|
||||
model_list:
|
||||
- model_name: gpt-3.5-turbo
|
||||
litellm_params:
|
||||
model: openai/gpt-3.5-turbo
|
||||
api_key: os.environ/OPENAI_API_KEY
|
||||
|
||||
guardrails:
|
||||
- guardrail_name: "presidio-german"
|
||||
litellm_params:
|
||||
guardrail: presidio
|
||||
mode: "pre_call"
|
||||
presidio_language: "de" # Default to German for PII analysis
|
||||
pii_entities_config:
|
||||
CREDIT_CARD: "MASK"
|
||||
EMAIL_ADDRESS: "MASK"
|
||||
PERSON: "MASK"
|
||||
|
||||
- guardrail_name: "presidio-spanish"
|
||||
litellm_params:
|
||||
guardrail: presidio
|
||||
mode: "pre_call"
|
||||
presidio_language: "es" # Default to Spanish for PII analysis
|
||||
pii_entities_config:
|
||||
CREDIT_CARD: "MASK"
|
||||
PHONE_NUMBER: "MASK"
|
||||
```
|
||||
|
||||
#### Supported Language Codes
|
||||
|
||||
Presidio supports multiple languages for PII detection. Common language codes include:
|
||||
|
||||
- `en` - English (default)
|
||||
- `es` - Spanish
|
||||
- `de` - German
|
||||
|
||||
For a complete list of supported languages, refer to the [Presidio documentation](https://microsoft.github.io/presidio/analyzer/languages/).
|
||||
|
||||
#### Language Precedence
|
||||
|
||||
The language setting follows this precedence order:
|
||||
|
||||
1. **Per-request language** (via `guardrail_config.language`) - highest priority
|
||||
2. **YAML config language** (via `presidio_language`) - medium priority
|
||||
3. **Default language** (`en`) - lowest priority
|
||||
|
||||
**Example with mixed languages:**
|
||||
|
||||
```yaml title="Mixed Language Configuration" showLineNumbers
|
||||
guardrails:
|
||||
- guardrail_name: "presidio-multilingual"
|
||||
litellm_params:
|
||||
guardrail: presidio
|
||||
mode: "pre_call"
|
||||
presidio_language: "de" # Default to German
|
||||
pii_entities_config:
|
||||
CREDIT_CARD: "MASK"
|
||||
PERSON: "MASK"
|
||||
```
|
||||
|
||||
```shell title="Override with per-request language" showLineNumbers
|
||||
curl http://localhost:4000/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer sk-1234" \
|
||||
-d '{
|
||||
"model": "gpt-3.5-turbo",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Mi tarjeta de crédito es 4111-1111-1111-1111"}
|
||||
],
|
||||
"guardrails": ["presidio-multilingual"],
|
||||
"guardrail_config": {"language": "es"}
|
||||
}'
|
||||
```
|
||||
|
||||
In this example, the request will use Spanish (`es`) for PII detection even though the guardrail is configured with German (`de`) as the default language.
|
||||
|
||||
### Output parsing
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -1260,7 +1260,7 @@ model_list:
|
|||
litellm_params:
|
||||
model: gpt-3.5-turbo
|
||||
litellm_settings:
|
||||
success_callback: ["s3"]
|
||||
success_callback: ["s3_v2"]
|
||||
s3_callback_params:
|
||||
s3_bucket_name: logs-bucket-litellm # AWS Bucket Name for S3
|
||||
s3_region_name: us-west-2 # AWS Region Name for S3
|
||||
|
|
@ -1304,7 +1304,7 @@ You can add the team alias to the object key by setting the `team_alias` in the
|
|||
|
||||
```yaml
|
||||
litellm_settings:
|
||||
callbacks: ["s3"]
|
||||
callbacks: ["s3_v2"]
|
||||
enable_preview_features: true
|
||||
s3_callback_params:
|
||||
s3_bucket_name: logs-bucket-litellm
|
||||
|
|
|
|||
|
|
@ -180,6 +180,19 @@ Use this for LLM API Error monitoring and tracking remaining rate limits and tok
|
|||
| `litellm_llm_api_latency_metric` | Latency (seconds) for just the LLM API call - tracked for labels "model", "hashed_api_key", "api_key_alias", "team", "team_alias", "requested_model", "end_user", "user" |
|
||||
| `litellm_llm_api_time_to_first_token_metric` | Time to first token for LLM API call - tracked for labels `model`, `hashed_api_key`, `api_key_alias`, `team`, `team_alias` [Note: only emitted for streaming requests] |
|
||||
|
||||
## Tracking `end_user` on Prometheus
|
||||
|
||||
By default LiteLLM does not track `end_user` on Prometheus. This is done to reduce the cardinality of the metrics from LiteLLM Proxy.
|
||||
|
||||
If you want to track `end_user` on Prometheus, you can do the following:
|
||||
|
||||
```yaml showLineNumbers title="config.yaml"
|
||||
litellm_settings:
|
||||
callbacks: ["prometheus"]
|
||||
enable_end_user_cost_tracking_prometheus_only: true
|
||||
```
|
||||
|
||||
|
||||
## [BETA] Custom Metrics
|
||||
|
||||
Track custom metrics on prometheus on all events mentioned above.
|
||||
|
|
|
|||
81
docs/my-website/docs/tutorials/anthropic_file_usage.md
Normal file
81
docs/my-website/docs/tutorials/anthropic_file_usage.md
Normal file
|
|
@ -0,0 +1,81 @@
|
|||
# Using Anthropic File API with LiteLLM Proxy
|
||||
|
||||
## Overview
|
||||
|
||||
This tutorial shows how to create and analyze files with Claude-4 on Anthropic via LiteLLM Proxy.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- LiteLLM Proxy running
|
||||
- Anthropic API key
|
||||
|
||||
Add the following to your `.env` file:
|
||||
```
|
||||
ANTHROPIC_API_KEY=sk-1234
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
### 1. Setup config.yaml
|
||||
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: claude-opus
|
||||
litellm_params:
|
||||
model: anthropic/claude-opus-4-20250514
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
```
|
||||
|
||||
## 2. Create a file
|
||||
|
||||
Use the `/anthropic` passthrough endpoint to create a file.
|
||||
|
||||
```bash
|
||||
curl -L -X POST 'http://0.0.0.0:4000/anthropic/v1/files' \
|
||||
-H 'x-api-key: sk-1234' \
|
||||
-H 'anthropic-version: 2023-06-01' \
|
||||
-H 'anthropic-beta: files-api-2025-04-14' \
|
||||
-F 'file=@"/path/to/your/file.csv"'
|
||||
```
|
||||
|
||||
Expected response:
|
||||
|
||||
```json
|
||||
{
|
||||
"created_at": "2023-11-07T05:31:56Z",
|
||||
"downloadable": false,
|
||||
"filename": "file.csv",
|
||||
"id": "file-1234",
|
||||
"mime_type": "text/csv",
|
||||
"size_bytes": 1,
|
||||
"type": "file"
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
## 3. Analyze the file with Claude-4 via `/chat/completions`
|
||||
|
||||
|
||||
```bash
|
||||
curl -L -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
|
||||
-H 'Content-Type: application/json' \
|
||||
-H 'Authorization: Bearer $LITELLM_API_KEY' \
|
||||
-d '{
|
||||
"model": "claude-opus",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "text", "text": "What is in this sheet?"},
|
||||
{
|
||||
"type": "file",
|
||||
"file": {
|
||||
"file_id": "file-1234",
|
||||
"format": "text/csv" # 👈 IMPORTANT: This is the format of the file you want to analyze
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}'
|
||||
```
|
||||
Binary file not shown.
|
Before Width: | Height: | Size: 61 KiB After Width: | Height: | Size: 418 KiB |
243
docs/my-website/release_notes/v1.72.0-stable/index.md
Normal file
243
docs/my-website/release_notes/v1.72.0-stable/index.md
Normal file
|
|
@ -0,0 +1,243 @@
|
|||
---
|
||||
title: v1.72.0-stable
|
||||
slug: v1.72.0-stable
|
||||
date: 2025-05-31T10:00:00
|
||||
authors:
|
||||
- name: Krrish Dholakia
|
||||
title: CEO, LiteLLM
|
||||
url: https://www.linkedin.com/in/krish-d/
|
||||
image_url: https://media.licdn.com/dms/image/v2/D4D03AQGrlsJ3aqpHmQ/profile-displayphoto-shrink_400_400/B4DZSAzgP7HYAg-/0/1737327772964?e=1749686400&v=beta&t=Hkl3U8Ps0VtvNxX0BNNq24b4dtX5wQaPFp6oiKCIHD8
|
||||
- name: Ishaan Jaffer
|
||||
title: CTO, LiteLLM
|
||||
url: https://www.linkedin.com/in/reffajnaahsi/
|
||||
image_url: https://pbs.twimg.com/profile_images/1613813310264340481/lz54oEiB_400x400.jpg
|
||||
|
||||
hide_table_of_contents: false
|
||||
---
|
||||
|
||||
import Image from '@theme/IdealImage';
|
||||
import Tabs from '@theme/Tabs';
|
||||
import TabItem from '@theme/TabItem';
|
||||
|
||||
|
||||
:::info
|
||||
|
||||
The release candidate is live now.
|
||||
|
||||
The production release will be live on Wednesday.
|
||||
|
||||
:::
|
||||
|
||||
## Deploy this version
|
||||
|
||||
<Tabs>
|
||||
<TabItem value="docker" label="Docker">
|
||||
|
||||
``` showLineNumbers title="docker run litellm"
|
||||
docker run
|
||||
-e STORE_MODEL_IN_DB=True
|
||||
-p 4000:4000
|
||||
ghcr.io/berriai/litellm:main-v1.72.0.rc
|
||||
```
|
||||
</TabItem>
|
||||
|
||||
<TabItem value="pip" label="Pip">
|
||||
|
||||
``` showLineNumbers title="pip install litellm"
|
||||
pip install litellm==1.72.0
|
||||
```
|
||||
</TabItem>
|
||||
</Tabs>
|
||||
|
||||
|
||||
## Key Highlights
|
||||
|
||||
LiteLLM v1.72.0-stable.rc is live now. Here are the key highlights of this release:
|
||||
|
||||
- **Vector Store Permissions**: Control Vector Store access at the Key, Team, and Organization level.
|
||||
- **Rate Limiting Sliding Window support**: Improved accuracy for Key/Team/User rate limits with request tracking across minutes.
|
||||
- **Aiohttp Transport used by default**: Aiohttp transport is now the default transport for LiteLLM networking requests. This gives users 2x higher RPS per instance with a 40ms median latency overhead.
|
||||
- **Bedrock Agents**: Call Bedrock Agents with `/chat/completions`, `/response` endpoints.
|
||||
- **Anthropic File API**: Upload and analyze CSV files with Claude-4 on Anthropic via LiteLLM.
|
||||
- **Prometheus**: End users (`end_user`) will no longer be tracked by default on Prometheus. Tracking end_users on prometheus is now opt-in. This is done to prevent the response from `/metrics` from becoming too large. [Read More](../../docs/proxy/prometheus#tracking-end_user-on-prometheus)
|
||||
|
||||
|
||||
---
|
||||
|
||||
## Vector Store Permissions
|
||||
|
||||
This release brings support for managing permissions for vector stores by Keys, Teams, Organizations (entities) on LiteLLM. When a request attempts to query a vector store, LiteLLM will block it if the requesting entity lacks the proper permissions.
|
||||
|
||||
This is great for use cases that require access to restricted data that you don't want everyone to use.
|
||||
|
||||
Over the next week we plan on adding permission management for MCP Servers.
|
||||
|
||||
---
|
||||
## Aiohttp Transport used by default
|
||||
|
||||
Aiohttp transport is now the default transport for LiteLLM networking requests. This gives users 2x higher RPS per instance with a 40ms median latency overhead. This has been live on LiteLLM Cloud for a week + gone through alpha users testing for a week.
|
||||
|
||||
|
||||
If you encounter any issues, you can disable using the aiohttp transport in the following ways:
|
||||
|
||||
**On LiteLLM Proxy**
|
||||
|
||||
Set the `DISABLE_AIOHTTP_TRANSPORT=True` in the environment variables.
|
||||
|
||||
```yaml showLineNumbers title="Environment Variable"
|
||||
export DISABLE_AIOHTTP_TRANSPORT="True"
|
||||
```
|
||||
|
||||
**On LiteLLM Python SDK**
|
||||
|
||||
Set the `disable_aiohttp_transport=True` to disable aiohttp transport.
|
||||
|
||||
```python showLineNumbers title="Python SDK"
|
||||
import litellm
|
||||
|
||||
litellm.disable_aiohttp_transport = True # default is False, enable this to disable aiohttp transport
|
||||
result = litellm.completion(
|
||||
model="openai/gpt-4o",
|
||||
messages=[{"role": "user", "content": "Hello, world!"}],
|
||||
)
|
||||
print(result)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
|
||||
## New Models / Updated Models
|
||||
|
||||
- **[Bedrock](../../docs/providers/bedrock)**
|
||||
- Video support for Bedrock Converse - [PR](https://github.com/BerriAI/litellm/pull/11166)
|
||||
- InvokeAgents support as /chat/completions route - [PR](https://github.com/BerriAI/litellm/pull/11239), [Get Started](../../docs/providers/bedrock_agents)
|
||||
- AI21 Jamba models compatibility fixes - [PR](https://github.com/BerriAI/litellm/pull/11233)
|
||||
- Fixed duplicate maxTokens parameter for Claude with thinking - [PR](https://github.com/BerriAI/litellm/pull/11181)
|
||||
- **[Gemini (Google AI Studio + Vertex AI)](https://docs.litellm.ai/docs/providers/gemini)**
|
||||
- Parallel tool calling support with `parallel_tool_calls` parameter - [PR](https://github.com/BerriAI/litellm/pull/11125)
|
||||
- All Gemini models now support parallel function calling - [PR](https://github.com/BerriAI/litellm/pull/11225)
|
||||
- **[VertexAI](../../docs/providers/vertex)**
|
||||
- codeExecution tool support and anyOf handling - [PR](https://github.com/BerriAI/litellm/pull/11195)
|
||||
- Vertex AI Anthropic support on /v1/messages - [PR](https://github.com/BerriAI/litellm/pull/11246)
|
||||
- Thinking, global regions, and parallel tool calling improvements - [PR](https://github.com/BerriAI/litellm/pull/11194)
|
||||
- Web Search Support [PR](https://github.com/BerriAI/litellm/commit/06484f6e5a7a2f4e45c490266782ed28b51b7db6)
|
||||
- **[Anthropic](../../docs/providers/anthropic)**
|
||||
- Thinking blocks on streaming support - [PR](https://github.com/BerriAI/litellm/pull/11194)
|
||||
- Files API with form-data support on passthrough - [PR](https://github.com/BerriAI/litellm/pull/11256)
|
||||
- File ID support on /chat/completion - [PR](https://github.com/BerriAI/litellm/pull/11256)
|
||||
- **[xAI](../../docs/providers/xai)**
|
||||
- Web Search Support [PR](https://github.com/BerriAI/litellm/commit/06484f6e5a7a2f4e45c490266782ed28b51b7db6)
|
||||
- **[Google AI Studio](../../docs/providers/gemini)**
|
||||
- Web Search Support [PR](https://github.com/BerriAI/litellm/commit/06484f6e5a7a2f4e45c490266782ed28b51b7db6)
|
||||
- **[Mistral](../../docs/providers/mistral)**
|
||||
- Updated mistral-medium prices and context sizes - [PR](https://github.com/BerriAI/litellm/pull/10729)
|
||||
- **[Ollama](../../docs/providers/ollama)**
|
||||
- Tool calls parsing on streaming - [PR](https://github.com/BerriAI/litellm/pull/11171)
|
||||
- **[Cohere](../../docs/providers/cohere)**
|
||||
- Swapped Cohere and Cohere Chat provider positioning - [PR](https://github.com/BerriAI/litellm/pull/11173)
|
||||
- **[Nebius AI Studio](../../docs/providers/nebius)**
|
||||
- New provider integration - [PR](https://github.com/BerriAI/litellm/pull/11143)
|
||||
|
||||
## LLM API Endpoints
|
||||
|
||||
- **[Image Edits API](../../docs/image_generation)**
|
||||
- Azure support for /v1/images/edits - [PR](https://github.com/BerriAI/litellm/pull/11160)
|
||||
- Cost tracking for image edits endpoint (OpenAI, Azure) - [PR](https://github.com/BerriAI/litellm/pull/11186)
|
||||
- **[Completions API](../../docs/completion/chat)**
|
||||
- Codestral latency overhead tracking on /v1/completions - [PR](https://github.com/BerriAI/litellm/pull/10879)
|
||||
- **[Audio Transcriptions API](../../docs/audio/speech)**
|
||||
- GPT-4o mini audio preview pricing without date - [PR](https://github.com/BerriAI/litellm/pull/11207)
|
||||
- Non-default params support for audio transcription - [PR](https://github.com/BerriAI/litellm/pull/11212)
|
||||
- **[Responses API](../../docs/response_api)**
|
||||
- Session management fixes for using Non-OpenAI models - [PR](https://github.com/BerriAI/litellm/pull/11254)
|
||||
|
||||
## Management Endpoints / UI
|
||||
|
||||
- **Vector Stores**
|
||||
- Permission management for LiteLLM Keys, Teams, and Organizations - [PR](https://github.com/BerriAI/litellm/pull/11213)
|
||||
- UI display of vector store permissions - [PR](https://github.com/BerriAI/litellm/pull/11277)
|
||||
- Vector store access controls enforcement - [PR](https://github.com/BerriAI/litellm/pull/11281)
|
||||
- Object permissions fixes and QA improvements - [PR](https://github.com/BerriAI/litellm/pull/11291)
|
||||
- **Teams**
|
||||
- "All proxy models" display when no models selected - [PR](https://github.com/BerriAI/litellm/pull/11187)
|
||||
- Removed redundant teamInfo call, using existing teamsList - [PR](https://github.com/BerriAI/litellm/pull/11051)
|
||||
- Improved model tags display on Keys, Teams and Org pages - [PR](https://github.com/BerriAI/litellm/pull/11022)
|
||||
- **SSO/SCIM**
|
||||
- Bug fixes for showing SCIM token on UI - [PR](https://github.com/BerriAI/litellm/pull/11220)
|
||||
- **General UI**
|
||||
- Fix "UI Session Expired. Logging out" - [PR](https://github.com/BerriAI/litellm/pull/11279)
|
||||
- Support for forwarding /sso/key/generate to server root path URL - [PR](https://github.com/BerriAI/litellm/pull/11165)
|
||||
|
||||
|
||||
## Logging / Guardrails Integrations
|
||||
|
||||
#### Logging
|
||||
- **[Prometheus](../../docs/proxy/prometheus)**
|
||||
- End users will no longer be tracked by default on Prometheus. Tracking end_users on prometheus is now opt-in. [PR](https://github.com/BerriAI/litellm/pull/11192)
|
||||
- **[Langfuse](../../docs/proxy/logging#langfuse)**
|
||||
- Performance improvements: Fixed "Max langfuse clients reached" issue - [PR](https://github.com/BerriAI/litellm/pull/11285)
|
||||
- **[Helicone](../../docs/observability/helicone_integration)**
|
||||
- Base URL support - [PR](https://github.com/BerriAI/litellm/pull/11211)
|
||||
- **[Sentry](../../docs/proxy/logging#sentry)**
|
||||
- Added sentry sample rate configuration - [PR](https://github.com/BerriAI/litellm/pull/10283)
|
||||
|
||||
#### Guardrails
|
||||
- **[Bedrock Guardrails](../../docs/proxy/guardrails/bedrock)**
|
||||
- Streaming support for bedrock post guard - [PR](https://github.com/BerriAI/litellm/pull/11247)
|
||||
- Auth parameter persistence fixes - [PR](https://github.com/BerriAI/litellm/pull/11270)
|
||||
- **[Pangea Guardrails](../../docs/proxy/guardrails/pangea)**
|
||||
- Added Pangea provider to Guardrails hook - [PR](https://github.com/BerriAI/litellm/pull/10775)
|
||||
|
||||
|
||||
## Performance / Reliability Improvements
|
||||
- **aiohttp Transport**
|
||||
- Handling for aiohttp.ClientPayloadError - [PR](https://github.com/BerriAI/litellm/pull/11162)
|
||||
- SSL verification settings support - [PR](https://github.com/BerriAI/litellm/pull/11162)
|
||||
- Rollback to httpx==0.27.0 for stability - [PR](https://github.com/BerriAI/litellm/pull/11146)
|
||||
- **Request Limiting**
|
||||
- Sliding window logic for parallel request limiter v2 - [PR](https://github.com/BerriAI/litellm/pull/11283)
|
||||
|
||||
|
||||
## Bug Fixes
|
||||
|
||||
- **LLM API Fixes**
|
||||
- Added missing request_kwargs to get_available_deployment call - [PR](https://github.com/BerriAI/litellm/pull/11202)
|
||||
- Fixed calling Azure O-series models - [PR](https://github.com/BerriAI/litellm/pull/11212)
|
||||
- Support for dropping non-OpenAI params via additional_drop_params - [PR](https://github.com/BerriAI/litellm/pull/11246)
|
||||
- Fixed frequency_penalty to repeat_penalty parameter mapping - [PR](https://github.com/BerriAI/litellm/pull/11284)
|
||||
- Fix for embedding cache hits on string input - [PR](https://github.com/BerriAI/litellm/pull/11211)
|
||||
- **General**
|
||||
- OIDC provider improvements and audience bug fix - [PR](https://github.com/BerriAI/litellm/pull/10054)
|
||||
- Removed AzureCredentialType restriction on AZURE_CREDENTIAL - [PR](https://github.com/BerriAI/litellm/pull/11272)
|
||||
- Prevention of sensitive key leakage to Langfuse - [PR](https://github.com/BerriAI/litellm/pull/11165)
|
||||
- Fixed healthcheck test using curl when curl not in image - [PR](https://github.com/BerriAI/litellm/pull/9737)
|
||||
|
||||
## New Contributors
|
||||
* [@agajdosi](https://github.com/agajdosi) made their first contribution in [#9737](https://github.com/BerriAI/litellm/pull/9737)
|
||||
* [@ketangangal](https://github.com/ketangangal) made their first contribution in [#11161](https://github.com/BerriAI/litellm/pull/11161)
|
||||
* [@Aktsvigun](https://github.com/Aktsvigun) made their first contribution in [#11143](https://github.com/BerriAI/litellm/pull/11143)
|
||||
* [@ryanmeans](https://github.com/ryanmeans) made their first contribution in [#10775](https://github.com/BerriAI/litellm/pull/10775)
|
||||
* [@nikoizs](https://github.com/nikoizs) made their first contribution in [#10054](https://github.com/BerriAI/litellm/pull/10054)
|
||||
* [@Nitro963](https://github.com/Nitro963) made their first contribution in [#11202](https://github.com/BerriAI/litellm/pull/11202)
|
||||
* [@Jacobh2](https://github.com/Jacobh2) made their first contribution in [#11207](https://github.com/BerriAI/litellm/pull/11207)
|
||||
* [@regismesquita](https://github.com/regismesquita) made their first contribution in [#10729](https://github.com/BerriAI/litellm/pull/10729)
|
||||
* [@Vinnie-Singleton-NN](https://github.com/Vinnie-Singleton-NN) made their first contribution in [#10283](https://github.com/BerriAI/litellm/pull/10283)
|
||||
* [@trashhalo](https://github.com/trashhalo) made their first contribution in [#11219](https://github.com/BerriAI/litellm/pull/11219)
|
||||
* [@VigneshwarRajasekaran](https://github.com/VigneshwarRajasekaran) made their first contribution in [#11223](https://github.com/BerriAI/litellm/pull/11223)
|
||||
* [@AnilAren](https://github.com/AnilAren) made their first contribution in [#11233](https://github.com/BerriAI/litellm/pull/11233)
|
||||
* [@fadil4u](https://github.com/fadil4u) made their first contribution in [#11242](https://github.com/BerriAI/litellm/pull/11242)
|
||||
* [@whitfin](https://github.com/whitfin) made their first contribution in [#11279](https://github.com/BerriAI/litellm/pull/11279)
|
||||
* [@hcoona](https://github.com/hcoona) made their first contribution in [#11272](https://github.com/BerriAI/litellm/pull/11272)
|
||||
* [@keyute](https://github.com/keyute) made their first contribution in [#11173](https://github.com/BerriAI/litellm/pull/11173)
|
||||
* [@emmanuel-ferdman](https://github.com/emmanuel-ferdman) made their first contribution in [#11230](https://github.com/BerriAI/litellm/pull/11230)
|
||||
|
||||
## Demo Instance
|
||||
|
||||
Here's a Demo Instance to test changes:
|
||||
|
||||
- Instance: https://demo.litellm.ai/
|
||||
- Login Credentials:
|
||||
- Username: admin
|
||||
- Password: sk-1234
|
||||
|
||||
## [Git Diff](https://github.com/BerriAI/litellm/releases)
|
||||
|
|
@ -102,6 +102,7 @@ const sidebars = {
|
|||
items: [
|
||||
"proxy/ui",
|
||||
"proxy/admin_ui_sso",
|
||||
"proxy/custom_root_ui",
|
||||
"proxy/self_serve",
|
||||
"proxy/public_teams",
|
||||
"tutorials/scim_litellm",
|
||||
|
|
@ -328,6 +329,7 @@ const sidebars = {
|
|||
label: "Bedrock",
|
||||
items: [
|
||||
"providers/bedrock",
|
||||
"providers/bedrock_agents",
|
||||
"providers/bedrock_vector_store",
|
||||
]
|
||||
},
|
||||
|
|
@ -505,6 +507,7 @@ const sidebars = {
|
|||
items: [
|
||||
"tutorials/openweb_ui",
|
||||
"tutorials/openai_codex",
|
||||
"tutorials/anthropic_file_usage",
|
||||
"tutorials/msft_sso",
|
||||
"tutorials/prompt_caching",
|
||||
"tutorials/tag_management",
|
||||
|
|
|
|||
BIN
enterprise/dist/litellm_enterprise-0.1.7-py3-none-any.whl
vendored
Normal file
BIN
enterprise/dist/litellm_enterprise-0.1.7-py3-none-any.whl
vendored
Normal file
Binary file not shown.
BIN
enterprise/dist/litellm_enterprise-0.1.7.tar.gz
vendored
Normal file
BIN
enterprise/dist/litellm_enterprise-0.1.7.tar.gz
vendored
Normal file
Binary file not shown.
|
|
@ -1,17 +1,23 @@
|
|||
from litellm.proxy._types import SpendLogsPayload
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from typing import Optional, List, Union
|
||||
import json
|
||||
from litellm.types.utils import ModelResponse, Message
|
||||
from typing import TYPE_CHECKING, Any, List, Optional, Union, cast
|
||||
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from litellm.proxy._types import SpendLogsPayload
|
||||
from litellm.responses.utils import ResponsesAPIRequestUtils
|
||||
from litellm.types.llms.openai import (
|
||||
AllMessageValues,
|
||||
ChatCompletionResponseMessage,
|
||||
GenericChatCompletionMessage,
|
||||
ResponseInputParam,
|
||||
)
|
||||
from litellm.types.utils import ChatCompletionMessageToolCall
|
||||
from litellm.responses.utils import ResponsesAPIRequestUtils
|
||||
from litellm.responses.litellm_completion_transformation.transformation import ChatCompletionSession
|
||||
from litellm.types.utils import ChatCompletionMessageToolCall, Message, ModelResponse
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.responses.litellm_completion_transformation.transformation import (
|
||||
ChatCompletionSession,
|
||||
)
|
||||
else:
|
||||
ChatCompletionSession = Any
|
||||
|
||||
|
||||
class _ENTERPRISE_ResponsesSessionHandler:
|
||||
|
|
@ -22,9 +28,23 @@ class _ENTERPRISE_ResponsesSessionHandler:
|
|||
"""
|
||||
Return the chat completion message history for a previous response id
|
||||
"""
|
||||
from litellm.responses.litellm_completion_transformation.transformation import LiteLLMCompletionResponsesConfig
|
||||
all_spend_logs: List[SpendLogsPayload] = await _ENTERPRISE_ResponsesSessionHandler.get_all_spend_logs_for_previous_response_id(previous_response_id)
|
||||
|
||||
from litellm.responses.litellm_completion_transformation.transformation import (
|
||||
ChatCompletionSession,
|
||||
LiteLLMCompletionResponsesConfig,
|
||||
)
|
||||
|
||||
verbose_proxy_logger.debug(
|
||||
"inside get_chat_completion_message_history_for_previous_response_id"
|
||||
)
|
||||
all_spend_logs: List[
|
||||
SpendLogsPayload
|
||||
] = await _ENTERPRISE_ResponsesSessionHandler.get_all_spend_logs_for_previous_response_id(
|
||||
previous_response_id
|
||||
)
|
||||
verbose_proxy_logger.debug(
|
||||
"found %s spend logs for this response id", len(all_spend_logs)
|
||||
)
|
||||
|
||||
litellm_session_id: Optional[str] = None
|
||||
if len(all_spend_logs) > 0:
|
||||
litellm_session_id = all_spend_logs[0].get("session_id")
|
||||
|
|
@ -39,14 +59,16 @@ class _ENTERPRISE_ResponsesSessionHandler:
|
|||
]
|
||||
] = []
|
||||
for spend_log in all_spend_logs:
|
||||
proxy_server_request: Union[str, dict] = spend_log.get("proxy_server_request") or "{}"
|
||||
proxy_server_request: Union[str, dict] = (
|
||||
spend_log.get("proxy_server_request") or "{}"
|
||||
)
|
||||
proxy_server_request_dict: Optional[dict] = None
|
||||
response_input_param: Optional[Union[str, ResponseInputParam]] = None
|
||||
if isinstance(proxy_server_request, dict):
|
||||
proxy_server_request_dict = proxy_server_request
|
||||
else:
|
||||
proxy_server_request_dict = json.loads(proxy_server_request)
|
||||
|
||||
|
||||
############################################################
|
||||
# Add Input messages for this Spend Log
|
||||
############################################################
|
||||
|
|
@ -55,15 +77,17 @@ class _ENTERPRISE_ResponsesSessionHandler:
|
|||
if isinstance(_response_input_param, str):
|
||||
response_input_param = _response_input_param
|
||||
elif isinstance(_response_input_param, dict):
|
||||
response_input_param = ResponseInputParam(**_response_input_param)
|
||||
|
||||
response_input_param = cast(
|
||||
ResponseInputParam, _response_input_param
|
||||
)
|
||||
|
||||
if response_input_param:
|
||||
chat_completion_messages = LiteLLMCompletionResponsesConfig.transform_responses_api_input_to_messages(
|
||||
input=response_input_param,
|
||||
responses_api_request=proxy_server_request_dict or {}
|
||||
responses_api_request=proxy_server_request_dict or {},
|
||||
)
|
||||
chat_completion_message_history.extend(chat_completion_messages)
|
||||
|
||||
|
||||
############################################################
|
||||
# Add Output messages for this Spend Log
|
||||
############################################################
|
||||
|
|
@ -73,17 +97,22 @@ class _ENTERPRISE_ResponsesSessionHandler:
|
|||
model_response = ModelResponse(**_response_output)
|
||||
for choice in model_response.choices:
|
||||
if hasattr(choice, "message"):
|
||||
chat_completion_message_history.append(choice.message)
|
||||
|
||||
verbose_proxy_logger.debug("chat_completion_message_history %s", json.dumps(chat_completion_message_history, indent=4, default=str))
|
||||
chat_completion_message_history.append(
|
||||
getattr(choice, "message")
|
||||
)
|
||||
|
||||
verbose_proxy_logger.debug(
|
||||
"chat_completion_message_history %s",
|
||||
json.dumps(chat_completion_message_history, indent=4, default=str),
|
||||
)
|
||||
return ChatCompletionSession(
|
||||
messages=chat_completion_message_history,
|
||||
litellm_session_id=litellm_session_id
|
||||
litellm_session_id=litellm_session_id,
|
||||
)
|
||||
|
||||
@staticmethod
|
||||
async def get_all_spend_logs_for_previous_response_id(
|
||||
previous_response_id: str
|
||||
previous_response_id: str,
|
||||
) -> List[SpendLogsPayload]:
|
||||
"""
|
||||
Get all spend logs for a previous response id
|
||||
|
|
@ -94,8 +123,17 @@ class _ENTERPRISE_ResponsesSessionHandler:
|
|||
SELECT session_id FROM spend_logs WHERE response_id = previous_response_id, SELECT * FROM spend_logs WHERE session_id = session_id
|
||||
"""
|
||||
from litellm.proxy.proxy_server import prisma_client
|
||||
decoded_response_id = ResponsesAPIRequestUtils._decode_responses_api_response_id(previous_response_id)
|
||||
previous_response_id = decoded_response_id.get("response_id", previous_response_id)
|
||||
|
||||
verbose_proxy_logger.debug("decoding response id=%s", previous_response_id)
|
||||
|
||||
decoded_response_id = (
|
||||
ResponsesAPIRequestUtils._decode_responses_api_response_id(
|
||||
previous_response_id
|
||||
)
|
||||
)
|
||||
previous_response_id = decoded_response_id.get(
|
||||
"response_id", previous_response_id
|
||||
)
|
||||
if prisma_client is None:
|
||||
return []
|
||||
|
||||
|
|
@ -111,21 +149,12 @@ class _ENTERPRISE_ResponsesSessionHandler:
|
|||
ORDER BY "endTime" ASC;
|
||||
"""
|
||||
|
||||
spend_logs = await prisma_client.db.query_raw(
|
||||
query,
|
||||
previous_response_id
|
||||
)
|
||||
spend_logs = await prisma_client.db.query_raw(query, previous_response_id)
|
||||
|
||||
verbose_proxy_logger.debug(
|
||||
"Found the following spend logs for previous response id %s: %s",
|
||||
previous_response_id,
|
||||
json.dumps(spend_logs, indent=4, default=str)
|
||||
json.dumps(spend_logs, indent=4, default=str),
|
||||
)
|
||||
|
||||
|
||||
return spend_logs
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
|
@ -1,6 +1,6 @@
|
|||
[tool.poetry]
|
||||
name = "litellm-enterprise"
|
||||
version = "0.1.6"
|
||||
version = "0.1.7"
|
||||
description = "Package for LiteLLM Enterprise features"
|
||||
authors = ["BerriAI"]
|
||||
readme = "README.md"
|
||||
|
|
@ -22,7 +22,7 @@ requires = ["poetry-core"]
|
|||
build-backend = "poetry.core.masonry.api"
|
||||
|
||||
[tool.commitizen]
|
||||
version = "0.1.6"
|
||||
version = "0.1.7"
|
||||
version_files = [
|
||||
"pyproject.toml:version",
|
||||
"../requirements.txt:litellm-enterprise==",
|
||||
|
|
|
|||
|
|
@ -119,6 +119,7 @@ _custom_logger_compatible_callbacks_literal = Literal[
|
|||
"resend_email",
|
||||
"smtp_email",
|
||||
"deepeval",
|
||||
"s3_v2",
|
||||
]
|
||||
logged_real_time_event_types: Optional[Union[List[str], Literal["*"]]] = None
|
||||
_known_custom_logger_compatible_callbacks: List = list(
|
||||
|
|
@ -133,7 +134,7 @@ langsmith_batch_size: Optional[int] = None
|
|||
prometheus_initialize_budget_metrics: Optional[bool] = False
|
||||
require_auth_for_metrics_endpoint: Optional[bool] = False
|
||||
argilla_batch_size: Optional[int] = None
|
||||
datadog_use_v1: Optional[bool] = False # if you want to use v1 datadog logged payload
|
||||
datadog_use_v1: Optional[bool] = False # if you want to use v1 datadog logged payload.
|
||||
gcs_pub_sub_use_v1: Optional[
|
||||
bool
|
||||
] = False # if you want to use v1 gcs pubsub logged payload
|
||||
|
|
@ -190,6 +191,7 @@ maritalk_key: Optional[str] = None
|
|||
ai21_key: Optional[str] = None
|
||||
ollama_key: Optional[str] = None
|
||||
openrouter_key: Optional[str] = None
|
||||
datarobot_key: Optional[str] = None
|
||||
predibase_key: Optional[str] = None
|
||||
huggingface_key: Optional[str] = None
|
||||
vertex_project: Optional[str] = None
|
||||
|
|
@ -215,6 +217,7 @@ use_client: bool = False
|
|||
ssl_verify: Union[str, bool] = True
|
||||
ssl_certificate: Optional[str] = None
|
||||
disable_streaming_logging: bool = False
|
||||
disable_token_counter: bool = False
|
||||
disable_add_transform_inline_image_block: bool = False
|
||||
in_memory_llm_clients_cache: LLMClientCache = LLMClientCache()
|
||||
safe_memory_mode: bool = False
|
||||
|
|
@ -303,7 +306,8 @@ priority_reservation: Optional[Dict[str, float]] = None
|
|||
|
||||
|
||||
######## Networking Settings ########
|
||||
use_aiohttp_transport: bool = True
|
||||
use_aiohttp_transport: bool = True # Older variable, aiohttp is now the default. use disable_aiohttp_transport instead.
|
||||
disable_aiohttp_transport: bool = False # Set this to true to use httpx instead
|
||||
force_ipv4: bool = False # when True, litellm will force ipv4 for all LLM requests. Some users have seen httpx ConnectionError when using ipv6.
|
||||
module_level_aclient = AsyncHTTPHandler(
|
||||
timeout=request_timeout, client_alias="module level aclient"
|
||||
|
|
@ -401,6 +405,7 @@ mistral_chat_models: List = []
|
|||
text_completion_codestral_models: List = []
|
||||
anthropic_models: List = []
|
||||
openrouter_models: List = []
|
||||
datarobot_models: List = []
|
||||
vertex_language_models: List = []
|
||||
vertex_vision_models: List = []
|
||||
vertex_chat_models: List = []
|
||||
|
|
@ -508,6 +513,8 @@ def add_known_models():
|
|||
empower_models.append(key)
|
||||
elif value.get("litellm_provider") == "openrouter":
|
||||
openrouter_models.append(key)
|
||||
elif value.get("litellm_provider") == "datarobot":
|
||||
datarobot_models.append(key)
|
||||
elif value.get("litellm_provider") == "vertex_ai-text-models":
|
||||
vertex_text_models.append(key)
|
||||
elif value.get("litellm_provider") == "vertex_ai-code-text-models":
|
||||
|
|
@ -658,6 +665,7 @@ model_list = (
|
|||
+ anthropic_models
|
||||
+ replicate_models
|
||||
+ openrouter_models
|
||||
+ datarobot_models
|
||||
+ huggingface_models
|
||||
+ vertex_chat_models
|
||||
+ vertex_text_models
|
||||
|
|
@ -718,6 +726,7 @@ models_by_provider: dict = {
|
|||
"together_ai": together_ai_models,
|
||||
"baseten": baseten_models,
|
||||
"openrouter": openrouter_models,
|
||||
"datarobot": datarobot_models,
|
||||
"vertex_ai": vertex_chat_models
|
||||
+ vertex_text_models
|
||||
+ vertex_anthropic_models
|
||||
|
|
@ -868,6 +877,7 @@ from .llms.huggingface.embedding.transformation import HuggingFaceEmbeddingConfi
|
|||
from .llms.oobabooga.chat.transformation import OobaboogaConfig
|
||||
from .llms.maritalk import MaritalkConfig
|
||||
from .llms.openrouter.chat.transformation import OpenrouterConfig
|
||||
from .llms.datarobot.chat.transformation import DataRobotConfig
|
||||
from .llms.anthropic.chat.transformation import AnthropicConfig
|
||||
from .llms.anthropic.common_utils import AnthropicModelInfo
|
||||
from .llms.groq.stt.transformation import GroqSTTConfig
|
||||
|
|
@ -1136,3 +1146,6 @@ disable_hf_tokenizer_download: Optional[
|
|||
bool
|
||||
] = None # disable huggingface tokenizer download. Defaults to openai clk100
|
||||
global_disable_no_log_param: bool = False
|
||||
|
||||
### PASSTHROUGH ###
|
||||
from .passthrough import allm_passthrough_route, llm_passthrough_route
|
||||
|
|
|
|||
|
|
@ -84,6 +84,19 @@ class InMemoryCache(BaseCache):
|
|||
except Exception:
|
||||
return False
|
||||
|
||||
def _is_key_expired(self, key: str) -> bool:
|
||||
"""
|
||||
Check if a specific key is expired
|
||||
"""
|
||||
return key in self.ttl_dict and time.time() > self.ttl_dict[key]
|
||||
|
||||
def _remove_key(self, key: str) -> None:
|
||||
"""
|
||||
Remove a key from both cache_dict and ttl_dict
|
||||
"""
|
||||
self.cache_dict.pop(key, None)
|
||||
self.ttl_dict.pop(key, None)
|
||||
|
||||
def evict_cache(self):
|
||||
"""
|
||||
Eviction policy:
|
||||
|
|
@ -97,9 +110,8 @@ class InMemoryCache(BaseCache):
|
|||
|
||||
"""
|
||||
for key in list(self.ttl_dict.keys()):
|
||||
if time.time() > self.ttl_dict[key]:
|
||||
self.cache_dict.pop(key, None)
|
||||
self.ttl_dict.pop(key, None)
|
||||
if self._is_key_expired(key):
|
||||
self._remove_key(key)
|
||||
|
||||
# de-reference the removed item
|
||||
# https://www.geeksforgeeks.org/diagnosing-and-fixing-memory-leaks-in-python/
|
||||
|
|
@ -153,13 +165,21 @@ class InMemoryCache(BaseCache):
|
|||
self.set_cache(key, init_value, ttl=ttl)
|
||||
return value
|
||||
|
||||
def evict_element_if_expired(self, key: str) -> bool:
|
||||
"""
|
||||
Returns True if the element is expired and removed from the cache
|
||||
|
||||
Returns False if the element is not expired
|
||||
"""
|
||||
if self._is_key_expired(key):
|
||||
self._remove_key(key)
|
||||
return True
|
||||
return False
|
||||
|
||||
def get_cache(self, key, **kwargs):
|
||||
if key in self.cache_dict:
|
||||
if key in self.ttl_dict:
|
||||
if time.time() > self.ttl_dict[key]:
|
||||
self.cache_dict.pop(key, None)
|
||||
self.ttl_dict.pop(key, None)
|
||||
return None
|
||||
if self.evict_element_if_expired(key):
|
||||
return None
|
||||
original_cached_response = self.cache_dict[key]
|
||||
try:
|
||||
cached_response = json.loads(original_cached_response)
|
||||
|
|
@ -207,8 +227,7 @@ class InMemoryCache(BaseCache):
|
|||
pass
|
||||
|
||||
def delete_cache(self, key):
|
||||
self.cache_dict.pop(key, None)
|
||||
self.ttl_dict.pop(key, None)
|
||||
self._remove_key(key)
|
||||
|
||||
async def async_get_ttl(self, key: str) -> Optional[int]:
|
||||
"""
|
||||
|
|
|
|||
|
|
@ -4,6 +4,10 @@ from typing import List, Literal
|
|||
ROUTER_MAX_FALLBACKS = int(os.getenv("ROUTER_MAX_FALLBACKS", 5))
|
||||
DEFAULT_BATCH_SIZE = int(os.getenv("DEFAULT_BATCH_SIZE", 512))
|
||||
DEFAULT_FLUSH_INTERVAL_SECONDS = int(os.getenv("DEFAULT_FLUSH_INTERVAL_SECONDS", 5))
|
||||
DEFAULT_S3_FLUSH_INTERVAL_SECONDS = int(
|
||||
os.getenv("DEFAULT_S3_FLUSH_INTERVAL_SECONDS", 10)
|
||||
)
|
||||
DEFAULT_S3_BATCH_SIZE = int(os.getenv("DEFAULT_S3_BATCH_SIZE", 512))
|
||||
DEFAULT_MAX_RETRIES = int(os.getenv("DEFAULT_MAX_RETRIES", 2))
|
||||
DEFAULT_MAX_RECURSE_DEPTH = int(os.getenv("DEFAULT_MAX_RECURSE_DEPTH", 100))
|
||||
DEFAULT_MAX_RECURSE_DEPTH_SENSITIVE_DATA_MASKER = int(
|
||||
|
|
@ -154,7 +158,10 @@ FIREWORKS_AI_80_B = int(os.getenv("FIREWORKS_AI_80_B", 80))
|
|||
#### Logging callback constants ####
|
||||
REDACTED_BY_LITELM_STRING = "REDACTED_BY_LITELM"
|
||||
MAX_LANGFUSE_INITIALIZED_CLIENTS = int(
|
||||
os.getenv("MAX_LANGFUSE_INITIALIZED_CLIENTS", 20)
|
||||
os.getenv("MAX_LANGFUSE_INITIALIZED_CLIENTS", 50)
|
||||
)
|
||||
DD_TRACER_STREAMING_CHUNK_YIELD_RESOURCE = os.getenv(
|
||||
"DD_TRACER_STREAMING_CHUNK_YIELD_RESOURCE", "streaming.chunk.yield"
|
||||
)
|
||||
|
||||
############### LLM Provider Constants ###############
|
||||
|
|
@ -180,6 +187,7 @@ LITELLM_CHAT_PROVIDERS = [
|
|||
"replicate",
|
||||
"huggingface",
|
||||
"together_ai",
|
||||
"datarobot",
|
||||
"openrouter",
|
||||
"vertex_ai",
|
||||
"vertex_ai_beta",
|
||||
|
|
@ -587,6 +595,7 @@ BEDROCK_INVOKE_PROVIDERS_LITERAL = Literal[
|
|||
|
||||
open_ai_embedding_models: List = ["text-embedding-ada-002"]
|
||||
cohere_embedding_models: List = [
|
||||
"embed-v4.0",
|
||||
"embed-english-v3.0",
|
||||
"embed-english-light-v3.0",
|
||||
"embed-multilingual-v3.0",
|
||||
|
|
|
|||
438
litellm/integrations/s3_v2.py
Normal file
438
litellm/integrations/s3_v2.py
Normal file
|
|
@ -0,0 +1,438 @@
|
|||
"""
|
||||
s3 Bucket Logging Integration
|
||||
|
||||
async_log_success_event: Processes the event, stores it in memory for DEFAULT_S3_FLUSH_INTERVAL_SECONDS seconds or until DEFAULT_S3_BATCH_SIZE and then flushes to s3
|
||||
|
||||
NOTE 1: S3 does not provide a BATCH PUT API endpoint, so we create tasks to upload each element individually
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
from datetime import datetime
|
||||
from typing import List, Optional, cast
|
||||
|
||||
import litellm
|
||||
from litellm._logging import print_verbose, verbose_logger
|
||||
from litellm.constants import DEFAULT_S3_BATCH_SIZE, DEFAULT_S3_FLUSH_INTERVAL_SECONDS
|
||||
from litellm.integrations.s3 import get_s3_object_key
|
||||
from litellm.llms.bedrock.base_aws_llm import BaseAWSLLM
|
||||
from litellm.llms.custom_httpx.http_handler import (
|
||||
_get_httpx_client,
|
||||
get_async_httpx_client,
|
||||
httpxSpecialProvider,
|
||||
)
|
||||
from litellm.types.integrations.s3_v2 import s3BatchLoggingElement
|
||||
from litellm.types.utils import StandardLoggingPayload
|
||||
|
||||
from .custom_batch_logger import CustomBatchLogger
|
||||
|
||||
|
||||
class S3Logger(CustomBatchLogger, BaseAWSLLM):
|
||||
def __init__(
|
||||
self,
|
||||
s3_bucket_name: Optional[str] = None,
|
||||
s3_path: Optional[str] = None,
|
||||
s3_region_name: Optional[str] = None,
|
||||
s3_api_version: Optional[str] = None,
|
||||
s3_use_ssl: bool = True,
|
||||
s3_verify: Optional[bool] = None,
|
||||
s3_endpoint_url: Optional[str] = None,
|
||||
s3_aws_access_key_id: Optional[str] = None,
|
||||
s3_aws_secret_access_key: Optional[str] = None,
|
||||
s3_aws_session_token: Optional[str] = None,
|
||||
s3_aws_session_name: Optional[str] = None,
|
||||
s3_aws_profile_name: Optional[str] = None,
|
||||
s3_aws_role_name: Optional[str] = None,
|
||||
s3_aws_web_identity_token: Optional[str] = None,
|
||||
s3_aws_sts_endpoint: Optional[str] = None,
|
||||
s3_flush_interval: Optional[int] = DEFAULT_S3_FLUSH_INTERVAL_SECONDS,
|
||||
s3_batch_size: Optional[int] = DEFAULT_S3_BATCH_SIZE,
|
||||
s3_config=None,
|
||||
s3_use_team_prefix: bool = False,
|
||||
**kwargs,
|
||||
):
|
||||
try:
|
||||
verbose_logger.debug(
|
||||
f"in init s3 logger - s3_callback_params {litellm.s3_callback_params}"
|
||||
)
|
||||
|
||||
# IMPORTANT: We use a concurrent limit of 1 to upload to s3
|
||||
# Files should get uploaded BUT they should not impact latency of LLM calling logic
|
||||
self.async_httpx_client = get_async_httpx_client(
|
||||
llm_provider=httpxSpecialProvider.LoggingCallback,
|
||||
)
|
||||
|
||||
self._init_s3_params(
|
||||
s3_bucket_name=s3_bucket_name,
|
||||
s3_region_name=s3_region_name,
|
||||
s3_api_version=s3_api_version,
|
||||
s3_use_ssl=s3_use_ssl,
|
||||
s3_verify=s3_verify,
|
||||
s3_endpoint_url=s3_endpoint_url,
|
||||
s3_aws_access_key_id=s3_aws_access_key_id,
|
||||
s3_aws_secret_access_key=s3_aws_secret_access_key,
|
||||
s3_aws_session_token=s3_aws_session_token,
|
||||
s3_aws_session_name=s3_aws_session_name,
|
||||
s3_aws_profile_name=s3_aws_profile_name,
|
||||
s3_aws_role_name=s3_aws_role_name,
|
||||
s3_aws_web_identity_token=s3_aws_web_identity_token,
|
||||
s3_aws_sts_endpoint=s3_aws_sts_endpoint,
|
||||
s3_config=s3_config,
|
||||
s3_path=s3_path,
|
||||
s3_use_team_prefix=s3_use_team_prefix,
|
||||
)
|
||||
verbose_logger.debug(f"s3 logger using endpoint url {s3_endpoint_url}")
|
||||
|
||||
asyncio.create_task(self.periodic_flush())
|
||||
self.flush_lock = asyncio.Lock()
|
||||
|
||||
verbose_logger.debug(
|
||||
f"s3 flush interval: {s3_flush_interval}, s3 batch size: {s3_batch_size}"
|
||||
)
|
||||
# Call CustomLogger's __init__
|
||||
CustomBatchLogger.__init__(
|
||||
self,
|
||||
flush_lock=self.flush_lock,
|
||||
flush_interval=s3_flush_interval,
|
||||
batch_size=s3_batch_size,
|
||||
)
|
||||
self.log_queue: List[s3BatchLoggingElement] = []
|
||||
|
||||
# Call BaseAWSLLM's __init__
|
||||
BaseAWSLLM.__init__(self)
|
||||
|
||||
except Exception as e:
|
||||
print_verbose(f"Got exception on init s3 client {str(e)}")
|
||||
raise e
|
||||
|
||||
def _init_s3_params(
|
||||
self,
|
||||
s3_bucket_name: Optional[str] = None,
|
||||
s3_region_name: Optional[str] = None,
|
||||
s3_api_version: Optional[str] = None,
|
||||
s3_use_ssl: bool = True,
|
||||
s3_verify: Optional[bool] = None,
|
||||
s3_endpoint_url: Optional[str] = None,
|
||||
s3_aws_access_key_id: Optional[str] = None,
|
||||
s3_aws_secret_access_key: Optional[str] = None,
|
||||
s3_aws_session_token: Optional[str] = None,
|
||||
s3_aws_session_name: Optional[str] = None,
|
||||
s3_aws_profile_name: Optional[str] = None,
|
||||
s3_aws_role_name: Optional[str] = None,
|
||||
s3_aws_web_identity_token: Optional[str] = None,
|
||||
s3_aws_sts_endpoint: Optional[str] = None,
|
||||
s3_config=None,
|
||||
s3_path: Optional[str] = None,
|
||||
s3_use_team_prefix: bool = False,
|
||||
):
|
||||
"""
|
||||
Initialize the s3 params for this logging callback
|
||||
"""
|
||||
litellm.s3_callback_params = litellm.s3_callback_params or {}
|
||||
# read in .env variables - example os.environ/AWS_BUCKET_NAME
|
||||
for key, value in litellm.s3_callback_params.items():
|
||||
if isinstance(value, str) and value.startswith("os.environ/"):
|
||||
litellm.s3_callback_params[key] = litellm.get_secret(value)
|
||||
|
||||
self.s3_bucket_name = (
|
||||
litellm.s3_callback_params.get("s3_bucket_name") or s3_bucket_name
|
||||
)
|
||||
self.s3_region_name = (
|
||||
litellm.s3_callback_params.get("s3_region_name") or s3_region_name
|
||||
)
|
||||
self.s3_api_version = (
|
||||
litellm.s3_callback_params.get("s3_api_version") or s3_api_version
|
||||
)
|
||||
self.s3_use_ssl = (
|
||||
litellm.s3_callback_params.get("s3_use_ssl", True) or s3_use_ssl
|
||||
)
|
||||
self.s3_verify = litellm.s3_callback_params.get("s3_verify") or s3_verify
|
||||
self.s3_endpoint_url = (
|
||||
litellm.s3_callback_params.get("s3_endpoint_url") or s3_endpoint_url
|
||||
)
|
||||
self.s3_aws_access_key_id = (
|
||||
litellm.s3_callback_params.get("s3_aws_access_key_id")
|
||||
or s3_aws_access_key_id
|
||||
)
|
||||
|
||||
self.s3_aws_secret_access_key = (
|
||||
litellm.s3_callback_params.get("s3_aws_secret_access_key")
|
||||
or s3_aws_secret_access_key
|
||||
)
|
||||
|
||||
self.s3_aws_session_token = (
|
||||
litellm.s3_callback_params.get("s3_aws_session_token")
|
||||
or s3_aws_session_token
|
||||
)
|
||||
|
||||
self.s3_aws_session_name = (
|
||||
litellm.s3_callback_params.get("s3_aws_session_name") or s3_aws_session_name
|
||||
)
|
||||
|
||||
self.s3_aws_profile_name = (
|
||||
litellm.s3_callback_params.get("s3_aws_profile_name") or s3_aws_profile_name
|
||||
)
|
||||
|
||||
self.s3_aws_role_name = (
|
||||
litellm.s3_callback_params.get("s3_aws_role_name") or s3_aws_role_name
|
||||
)
|
||||
|
||||
self.s3_aws_web_identity_token = (
|
||||
litellm.s3_callback_params.get("s3_aws_web_identity_token")
|
||||
or s3_aws_web_identity_token
|
||||
)
|
||||
|
||||
self.s3_aws_sts_endpoint = (
|
||||
litellm.s3_callback_params.get("s3_aws_sts_endpoint") or s3_aws_sts_endpoint
|
||||
)
|
||||
|
||||
self.s3_config = litellm.s3_callback_params.get("s3_config") or s3_config
|
||||
self.s3_path = litellm.s3_callback_params.get("s3_path") or s3_path
|
||||
# done reading litellm.s3_callback_params
|
||||
self.s3_use_team_prefix = (
|
||||
bool(litellm.s3_callback_params.get("s3_use_team_prefix", False))
|
||||
or s3_use_team_prefix
|
||||
)
|
||||
|
||||
return
|
||||
|
||||
async def async_log_success_event(self, kwargs, response_obj, start_time, end_time):
|
||||
try:
|
||||
verbose_logger.debug(
|
||||
f"s3 Logging - Enters logging function for model {kwargs}"
|
||||
)
|
||||
|
||||
s3_batch_logging_element = self.create_s3_batch_logging_element(
|
||||
start_time=start_time,
|
||||
standard_logging_payload=kwargs.get("standard_logging_object", None),
|
||||
)
|
||||
|
||||
if s3_batch_logging_element is None:
|
||||
raise ValueError("s3_batch_logging_element is None")
|
||||
|
||||
verbose_logger.debug(
|
||||
"\ns3 Logger - Logging payload = %s", s3_batch_logging_element
|
||||
)
|
||||
|
||||
self.log_queue.append(s3_batch_logging_element)
|
||||
verbose_logger.debug(
|
||||
"s3 logging: queue length %s, batch size %s",
|
||||
len(self.log_queue),
|
||||
self.batch_size,
|
||||
)
|
||||
except Exception as e:
|
||||
verbose_logger.exception(f"s3 Layer Error - {str(e)}")
|
||||
pass
|
||||
|
||||
async def async_upload_data_to_s3(
|
||||
self, batch_logging_element: s3BatchLoggingElement
|
||||
):
|
||||
try:
|
||||
import hashlib
|
||||
|
||||
import requests
|
||||
from botocore.auth import SigV4Auth
|
||||
from botocore.awsrequest import AWSRequest
|
||||
except ImportError:
|
||||
raise ImportError("Missing boto3 to call bedrock. Run 'pip install boto3'.")
|
||||
try:
|
||||
from litellm.litellm_core_utils.asyncify import asyncify
|
||||
|
||||
asyncified_get_credentials = asyncify(self.get_credentials)
|
||||
credentials = await asyncified_get_credentials(
|
||||
aws_access_key_id=self.s3_aws_access_key_id,
|
||||
aws_secret_access_key=self.s3_aws_secret_access_key,
|
||||
aws_session_token=self.s3_aws_session_token,
|
||||
aws_region_name=self.s3_region_name,
|
||||
aws_session_name=self.s3_aws_session_name,
|
||||
aws_profile_name=self.s3_aws_profile_name,
|
||||
aws_role_name=self.s3_aws_role_name,
|
||||
aws_web_identity_token=self.s3_aws_web_identity_token,
|
||||
aws_sts_endpoint=self.s3_aws_sts_endpoint,
|
||||
)
|
||||
|
||||
verbose_logger.debug(
|
||||
f"s3_v2 logger - uploading data to s3 - {batch_logging_element.s3_object_key}"
|
||||
)
|
||||
|
||||
# Prepare the URL
|
||||
url = f"https://{self.s3_bucket_name}.s3.{self.s3_region_name}.amazonaws.com/{batch_logging_element.s3_object_key}"
|
||||
|
||||
if self.s3_endpoint_url:
|
||||
url = self.s3_endpoint_url + "/" + batch_logging_element.s3_object_key
|
||||
|
||||
# Convert JSON to string
|
||||
json_string = json.dumps(batch_logging_element.payload)
|
||||
|
||||
# Calculate SHA256 hash of the content
|
||||
content_hash = hashlib.sha256(json_string.encode("utf-8")).hexdigest()
|
||||
|
||||
# Prepare the request
|
||||
headers = {
|
||||
"Content-Type": "application/json",
|
||||
"x-amz-content-sha256": content_hash,
|
||||
"Content-Language": "en",
|
||||
"Content-Disposition": f'inline; filename="{batch_logging_element.s3_object_download_filename}"',
|
||||
"Cache-Control": "private, immutable, max-age=31536000, s-maxage=0",
|
||||
}
|
||||
req = requests.Request("PUT", url, data=json_string, headers=headers)
|
||||
prepped = req.prepare()
|
||||
|
||||
# Sign the request
|
||||
aws_request = AWSRequest(
|
||||
method=prepped.method,
|
||||
url=prepped.url,
|
||||
data=prepped.body,
|
||||
headers=prepped.headers,
|
||||
)
|
||||
SigV4Auth(credentials, "s3", self.s3_region_name).add_auth(aws_request)
|
||||
|
||||
# Prepare the signed headers
|
||||
signed_headers = dict(aws_request.headers.items())
|
||||
|
||||
# Make the request
|
||||
response = await self.async_httpx_client.put(
|
||||
url, data=json_string, headers=signed_headers
|
||||
)
|
||||
response.raise_for_status()
|
||||
except Exception as e:
|
||||
verbose_logger.exception(f"Error uploading to s3: {str(e)}")
|
||||
|
||||
async def async_send_batch(self):
|
||||
"""
|
||||
|
||||
Sends runs from self.log_queue
|
||||
|
||||
Returns: None
|
||||
|
||||
Raises: Does not raise an exception, will only verbose_logger.exception()
|
||||
"""
|
||||
verbose_logger.debug(f"s3_v2 logger - sending batch of {len(self.log_queue)}")
|
||||
if not self.log_queue:
|
||||
return
|
||||
|
||||
#########################################################
|
||||
# Flush the log queue to s3
|
||||
# the log queue can be bounded by DEFAULT_S3_BATCH_SIZE
|
||||
# see custom_batch_logger.py which triggers the flush
|
||||
#########################################################
|
||||
for payload in self.log_queue:
|
||||
asyncio.create_task(self.async_upload_data_to_s3(payload))
|
||||
|
||||
def create_s3_batch_logging_element(
|
||||
self,
|
||||
start_time: datetime,
|
||||
standard_logging_payload: Optional[StandardLoggingPayload],
|
||||
) -> Optional[s3BatchLoggingElement]:
|
||||
"""
|
||||
Helper function to create an s3BatchLoggingElement.
|
||||
|
||||
Args:
|
||||
start_time (datetime): The start time of the logging event.
|
||||
standard_logging_payload (Optional[StandardLoggingPayload]): The payload to be logged.
|
||||
s3_path (Optional[str]): The S3 path prefix.
|
||||
|
||||
Returns:
|
||||
Optional[s3BatchLoggingElement]: The created s3BatchLoggingElement, or None if payload is None.
|
||||
"""
|
||||
if standard_logging_payload is None:
|
||||
return None
|
||||
|
||||
team_alias = standard_logging_payload["metadata"].get("user_api_key_team_alias")
|
||||
|
||||
team_alias_prefix = ""
|
||||
if (
|
||||
litellm.enable_preview_features
|
||||
and self.s3_use_team_prefix
|
||||
and team_alias is not None
|
||||
):
|
||||
team_alias_prefix = f"{team_alias}/"
|
||||
|
||||
s3_file_name = (
|
||||
litellm.utils.get_logging_id(start_time, standard_logging_payload) or ""
|
||||
)
|
||||
s3_object_key = get_s3_object_key(
|
||||
s3_path=cast(Optional[str], self.s3_path) or "",
|
||||
team_alias_prefix=team_alias_prefix,
|
||||
start_time=start_time,
|
||||
s3_file_name=s3_file_name,
|
||||
)
|
||||
|
||||
s3_object_download_filename = (
|
||||
"time-"
|
||||
+ start_time.strftime("%Y-%m-%dT%H-%M-%S-%f")
|
||||
+ "_"
|
||||
+ standard_logging_payload["id"]
|
||||
+ ".json"
|
||||
)
|
||||
|
||||
s3_object_download_filename = f"time-{start_time.strftime('%Y-%m-%dT%H-%M-%S-%f')}_{standard_logging_payload['id']}.json"
|
||||
|
||||
return s3BatchLoggingElement(
|
||||
payload=dict(standard_logging_payload),
|
||||
s3_object_key=s3_object_key,
|
||||
s3_object_download_filename=s3_object_download_filename,
|
||||
)
|
||||
|
||||
def upload_data_to_s3(self, batch_logging_element: s3BatchLoggingElement):
|
||||
try:
|
||||
import hashlib
|
||||
|
||||
import requests
|
||||
from botocore.auth import SigV4Auth
|
||||
from botocore.awsrequest import AWSRequest
|
||||
from botocore.credentials import Credentials
|
||||
except ImportError:
|
||||
raise ImportError("Missing boto3 to call bedrock. Run 'pip install boto3'.")
|
||||
try:
|
||||
verbose_logger.debug(
|
||||
f"s3_v2 logger - uploading data to s3 - {batch_logging_element.s3_object_key}"
|
||||
)
|
||||
credentials: Credentials = self.get_credentials(
|
||||
aws_access_key_id=self.s3_aws_access_key_id,
|
||||
aws_secret_access_key=self.s3_aws_secret_access_key,
|
||||
aws_session_token=self.s3_aws_session_token,
|
||||
aws_region_name=self.s3_region_name,
|
||||
)
|
||||
|
||||
# Prepare the URL
|
||||
url = f"https://{self.s3_bucket_name}.s3.{self.s3_region_name}.amazonaws.com/{batch_logging_element.s3_object_key}"
|
||||
|
||||
if self.s3_endpoint_url:
|
||||
url = self.s3_endpoint_url + "/" + batch_logging_element.s3_object_key
|
||||
|
||||
# Convert JSON to string
|
||||
json_string = json.dumps(batch_logging_element.payload)
|
||||
|
||||
# Calculate SHA256 hash of the content
|
||||
content_hash = hashlib.sha256(json_string.encode("utf-8")).hexdigest()
|
||||
|
||||
# Prepare the request
|
||||
headers = {
|
||||
"Content-Type": "application/json",
|
||||
"x-amz-content-sha256": content_hash,
|
||||
"Content-Language": "en",
|
||||
"Content-Disposition": f'inline; filename="{batch_logging_element.s3_object_download_filename}"',
|
||||
"Cache-Control": "private, immutable, max-age=31536000, s-maxage=0",
|
||||
}
|
||||
req = requests.Request("PUT", url, data=json_string, headers=headers)
|
||||
prepped = req.prepare()
|
||||
|
||||
# Sign the request
|
||||
aws_request = AWSRequest(
|
||||
method=prepped.method,
|
||||
url=prepped.url,
|
||||
data=prepped.body,
|
||||
headers=prepped.headers,
|
||||
)
|
||||
SigV4Auth(credentials, "s3", self.s3_region_name).add_auth(aws_request)
|
||||
|
||||
# Prepare the signed headers
|
||||
signed_headers = dict(aws_request.headers.items())
|
||||
|
||||
httpx_client = _get_httpx_client()
|
||||
# Make the request
|
||||
response = httpx_client.put(url, data=json_string, headers=signed_headers)
|
||||
response.raise_for_status()
|
||||
except Exception as e:
|
||||
verbose_logger.exception(f"Error uploading to s3: {str(e)}")
|
||||
|
|
@ -34,7 +34,6 @@ from litellm.types.vector_stores import (
|
|||
VectorStoreSearchResponse,
|
||||
VectorStoreSearchResult,
|
||||
)
|
||||
from litellm.utils import load_credentials_from_list
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
|
||||
|
|
@ -258,22 +257,49 @@ class BedrockVectorStore(BaseVectorStore, BaseAWSLLM):
|
|||
from fastapi import HTTPException
|
||||
|
||||
non_default_params = non_default_params or {}
|
||||
load_credentials_from_list(kwargs=non_default_params)
|
||||
credentials_dict: Dict[str, Any] = {}
|
||||
if litellm.vector_store_registry is not None:
|
||||
credentials_dict = (
|
||||
litellm.vector_store_registry.get_credentials_for_vector_store(
|
||||
knowledge_base_id
|
||||
)
|
||||
)
|
||||
|
||||
credentials = self.get_credentials(
|
||||
aws_access_key_id=non_default_params.get("aws_access_key_id", None),
|
||||
aws_secret_access_key=non_default_params.get("aws_secret_access_key", None),
|
||||
aws_session_token=non_default_params.get("aws_session_token", None),
|
||||
aws_region_name=non_default_params.get("aws_region_name", None),
|
||||
aws_session_name=non_default_params.get("aws_session_name", None),
|
||||
aws_profile_name=non_default_params.get("aws_profile_name", None),
|
||||
aws_role_name=non_default_params.get("aws_role_name", None),
|
||||
aws_web_identity_token=non_default_params.get(
|
||||
"aws_web_identity_token", None
|
||||
aws_access_key_id=credentials_dict.get(
|
||||
"aws_access_key_id", non_default_params.get("aws_access_key_id", None)
|
||||
),
|
||||
aws_secret_access_key=credentials_dict.get(
|
||||
"aws_secret_access_key",
|
||||
non_default_params.get("aws_secret_access_key", None),
|
||||
),
|
||||
aws_session_token=credentials_dict.get(
|
||||
"aws_session_token", non_default_params.get("aws_session_token", None)
|
||||
),
|
||||
aws_region_name=credentials_dict.get(
|
||||
"aws_region_name", non_default_params.get("aws_region_name", None)
|
||||
),
|
||||
aws_session_name=credentials_dict.get(
|
||||
"aws_session_name", non_default_params.get("aws_session_name", None)
|
||||
),
|
||||
aws_profile_name=credentials_dict.get(
|
||||
"aws_profile_name", non_default_params.get("aws_profile_name", None)
|
||||
),
|
||||
aws_role_name=credentials_dict.get(
|
||||
"aws_role_name", non_default_params.get("aws_role_name", None)
|
||||
),
|
||||
aws_web_identity_token=credentials_dict.get(
|
||||
"aws_web_identity_token",
|
||||
non_default_params.get("aws_web_identity_token", None),
|
||||
),
|
||||
aws_sts_endpoint=credentials_dict.get(
|
||||
"aws_sts_endpoint", non_default_params.get("aws_sts_endpoint", None)
|
||||
),
|
||||
aws_sts_endpoint=non_default_params.get("aws_sts_endpoint", None),
|
||||
)
|
||||
aws_region_name = self._get_aws_region_name(
|
||||
optional_params=self.optional_params
|
||||
aws_region_name = self.get_aws_region_name_for_non_llm_api_calls(
|
||||
aws_region_name=credentials_dict.get(
|
||||
"aws_region_name", non_default_params.get("aws_region_name", None)
|
||||
),
|
||||
)
|
||||
|
||||
# Prepare request data
|
||||
|
|
|
|||
|
|
@ -514,6 +514,14 @@ def _get_openai_compatible_provider_info( # noqa: PLR0915
|
|||
) = litellm.LlamafileChatConfig()._get_openai_compatible_provider_info(
|
||||
api_base, api_key
|
||||
)
|
||||
elif custom_llm_provider == "datarobot":
|
||||
# DataRobot is OpenAI compatible.
|
||||
(
|
||||
api_base,
|
||||
dynamic_api_key
|
||||
) = litellm.DataRobotConfig()._get_openai_compatible_provider_info(
|
||||
api_base, api_key
|
||||
)
|
||||
elif custom_llm_provider == "lm_studio":
|
||||
# lm_studio is openai compatible, we just need to set this to custom_openai
|
||||
(
|
||||
|
|
|
|||
|
|
@ -135,6 +135,7 @@ from ..integrations.opik.opik import OpikLogger
|
|||
from ..integrations.prometheus import PrometheusLogger
|
||||
from ..integrations.prompt_layer import PromptLayerLogger
|
||||
from ..integrations.s3 import S3Logger
|
||||
from ..integrations.s3_v2 import S3Logger as S3V2Logger
|
||||
from ..integrations.supabase import Supabase
|
||||
from ..integrations.traceloop import TraceloopLogger
|
||||
from ..integrations.weights_biases import WeightsBiasesLogger
|
||||
|
|
@ -2699,7 +2700,9 @@ def set_callbacks(callback_list, function_id=None): # noqa: PLR0915
|
|||
sentry_sdk_instance.init(
|
||||
dsn=os.environ.get("SENTRY_DSN"),
|
||||
traces_sample_rate=float(sentry_trace_rate), # type: ignore
|
||||
sample_rate=float(sentry_sample_rate),
|
||||
sample_rate=float(
|
||||
sentry_sample_rate if sentry_sample_rate else 1.0
|
||||
),
|
||||
)
|
||||
capture_exception = sentry_sdk_instance.capture_exception
|
||||
add_breadcrumb = sentry_sdk_instance.add_breadcrumb
|
||||
|
|
@ -2867,6 +2870,14 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
_gcs_bucket_logger = GCSBucketLogger()
|
||||
_in_memory_loggers.append(_gcs_bucket_logger)
|
||||
return _gcs_bucket_logger # type: ignore
|
||||
elif logging_integration == "s3_v2":
|
||||
for callback in _in_memory_loggers:
|
||||
if isinstance(callback, S3V2Logger):
|
||||
return callback # type: ignore
|
||||
|
||||
_s3_v2_logger = S3V2Logger()
|
||||
_in_memory_loggers.append(_s3_v2_logger)
|
||||
return _s3_v2_logger # type: ignore
|
||||
elif logging_integration == "azure_storage":
|
||||
for callback in _in_memory_loggers:
|
||||
if isinstance(callback, AzureBlobStorageLogger):
|
||||
|
|
@ -2962,7 +2973,7 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
galileo_logger = GalileoObserve()
|
||||
_in_memory_loggers.append(galileo_logger)
|
||||
return galileo_logger # type: ignore
|
||||
|
||||
|
||||
elif logging_integration == "deepeval":
|
||||
for callback in _in_memory_loggers:
|
||||
if isinstance(callback, DeepEvalLogger):
|
||||
|
|
@ -2970,7 +2981,7 @@ def _init_custom_logger_compatible_class( # noqa: PLR0915
|
|||
deepeval_logger = DeepEvalLogger()
|
||||
_in_memory_loggers.append(deepeval_logger)
|
||||
return deepeval_logger # type: ignore
|
||||
|
||||
|
||||
elif logging_integration == "logfire":
|
||||
if "LOGFIRE_TOKEN" not in os.environ:
|
||||
raise ValueError("LOGFIRE_TOKEN not found in environment variables")
|
||||
|
|
@ -3172,6 +3183,10 @@ def get_custom_logger_compatible_class( # noqa: PLR0915
|
|||
for callback in _in_memory_loggers:
|
||||
if isinstance(callback, GCSBucketLogger):
|
||||
return callback
|
||||
elif logging_integration == "s3_v2":
|
||||
for callback in _in_memory_loggers:
|
||||
if isinstance(callback, S3V2Logger):
|
||||
return callback
|
||||
elif logging_integration == "azure_storage":
|
||||
for callback in _in_memory_loggers:
|
||||
if isinstance(callback, AzureBlobStorageLogger):
|
||||
|
|
|
|||
|
|
@ -532,6 +532,12 @@ def convert_to_model_response_object( # noqa: PLR0915
|
|||
if finish_reason is None:
|
||||
# gpt-4 vision can return 'finish_reason' or 'finish_details'
|
||||
finish_reason = choice.get("finish_details") or "stop"
|
||||
if (
|
||||
finish_reason == "stop"
|
||||
and message.tool_calls
|
||||
and len(message.tool_calls) > 0
|
||||
):
|
||||
finish_reason = "tool_calls"
|
||||
logprobs = choice.get("logprobs", None)
|
||||
enhancements = choice.get("enhancements", None)
|
||||
choice = Choices(
|
||||
|
|
|
|||
|
|
@ -72,13 +72,11 @@ def get_api_base(
|
|||
_optional_params.vertex_location is not None
|
||||
and _optional_params.vertex_project is not None
|
||||
):
|
||||
from litellm.llms.vertex_ai.vertex_ai_partner_models.main import (
|
||||
VertexPartnerProvider,
|
||||
create_vertex_url,
|
||||
)
|
||||
from litellm.llms.vertex_ai.vertex_llm_base import VertexBase
|
||||
from litellm.types.llms.vertex_ai import VertexPartnerProvider
|
||||
|
||||
if "claude" in model:
|
||||
_api_base = create_vertex_url(
|
||||
_api_base = VertexBase.create_vertex_url(
|
||||
vertex_location=_optional_params.vertex_location,
|
||||
vertex_project=_optional_params.vertex_project,
|
||||
model=model,
|
||||
|
|
|
|||
|
|
@ -6,7 +6,17 @@ import io
|
|||
import mimetypes
|
||||
import re
|
||||
from os import PathLike
|
||||
from typing import Any, Dict, List, Literal, Mapping, Optional, Union, cast
|
||||
from typing import (
|
||||
TYPE_CHECKING,
|
||||
Any,
|
||||
Dict,
|
||||
List,
|
||||
Literal,
|
||||
Mapping,
|
||||
Optional,
|
||||
Union,
|
||||
cast,
|
||||
)
|
||||
|
||||
from litellm.types.llms.openai import (
|
||||
AllMessageValues,
|
||||
|
|
@ -25,6 +35,9 @@ from litellm.types.utils import (
|
|||
StreamingChoices,
|
||||
)
|
||||
|
||||
if TYPE_CHECKING: # newer pattern to avoid importing pydantic objects on __init__.py
|
||||
from litellm.types.llms.openai import ChatCompletionImageObject
|
||||
|
||||
DEFAULT_USER_CONTINUE_MESSAGE = ChatCompletionUserMessage(
|
||||
content="Please continue.", role="user"
|
||||
)
|
||||
|
|
@ -33,6 +46,9 @@ DEFAULT_ASSISTANT_CONTINUE_MESSAGE = ChatCompletionAssistantMessage(
|
|||
content="Please continue.", role="assistant"
|
||||
)
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.litellm_core_utils.litellm_logging import Logging as LoggingClass
|
||||
|
||||
|
||||
def handle_any_messages_to_chat_completion_str_messages_conversion(
|
||||
messages: Any,
|
||||
|
|
@ -582,3 +598,93 @@ def is_function_call(optional_params: dict) -> bool:
|
|||
if "functions" in optional_params and optional_params.get("functions"):
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def get_file_ids_from_messages(messages: List[AllMessageValues]) -> List[str]:
|
||||
"""
|
||||
Gets file ids from messages
|
||||
"""
|
||||
file_ids = []
|
||||
for message in messages:
|
||||
if message.get("role") == "user":
|
||||
content = message.get("content")
|
||||
if content:
|
||||
if isinstance(content, str):
|
||||
continue
|
||||
for c in content:
|
||||
if c["type"] == "file":
|
||||
file_object = cast(ChatCompletionFileObject, c)
|
||||
file_object_file_field = file_object["file"]
|
||||
file_id = file_object_file_field.get("file_id")
|
||||
if file_id:
|
||||
file_ids.append(file_id)
|
||||
return file_ids
|
||||
|
||||
|
||||
|
||||
def check_is_function_call(logging_obj: "LoggingClass") -> bool:
|
||||
from litellm.litellm_core_utils.prompt_templates.common_utils import (
|
||||
is_function_call,
|
||||
)
|
||||
|
||||
if hasattr(logging_obj, "optional_params") and isinstance(
|
||||
logging_obj.optional_params, dict
|
||||
):
|
||||
if is_function_call(logging_obj.optional_params):
|
||||
return True
|
||||
|
||||
return False
|
||||
|
||||
def filter_value_from_dict(dictionary: dict, key: str, depth: int = 0) -> Any:
|
||||
"""
|
||||
Filters a value from a dictionary
|
||||
|
||||
Goes through the nested dict and removes the key if it exists
|
||||
"""
|
||||
from litellm.constants import DEFAULT_MAX_RECURSE_DEPTH
|
||||
|
||||
if depth > DEFAULT_MAX_RECURSE_DEPTH:
|
||||
return dictionary
|
||||
|
||||
# Create a copy of keys to avoid modifying dict during iteration
|
||||
keys = list(dictionary.keys())
|
||||
for k in keys:
|
||||
v = dictionary[k]
|
||||
if k == key:
|
||||
del dictionary[k]
|
||||
elif isinstance(v, dict):
|
||||
filter_value_from_dict(v, key, depth + 1)
|
||||
elif isinstance(v, list):
|
||||
for item in v:
|
||||
if isinstance(item, dict):
|
||||
filter_value_from_dict(item, key, depth + 1)
|
||||
return dictionary
|
||||
|
||||
|
||||
def migrate_file_to_image_url(
|
||||
message: "ChatCompletionFileObject",
|
||||
) -> "ChatCompletionImageObject":
|
||||
"""
|
||||
Migrate file to image_url
|
||||
"""
|
||||
from litellm.types.llms.openai import (
|
||||
ChatCompletionImageObject,
|
||||
ChatCompletionImageUrlObject,
|
||||
)
|
||||
|
||||
file_id = message["file"].get("file_id")
|
||||
file_data = message["file"].get("file_data")
|
||||
format = message["file"].get("format")
|
||||
if not file_id and not file_data:
|
||||
raise ValueError("file_id and file_data are both None")
|
||||
image_url_object = ChatCompletionImageObject(
|
||||
type="image_url",
|
||||
image_url=ChatCompletionImageUrlObject(
|
||||
url=cast(str, file_id or file_data),
|
||||
),
|
||||
)
|
||||
if format and isinstance(image_url_object["image_url"], dict):
|
||||
image_url_object["image_url"]["format"] = format
|
||||
return image_url_object
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -1385,6 +1385,84 @@ def _anthropic_content_element_factory(
|
|||
return _anthropic_content_element
|
||||
|
||||
|
||||
def select_anthropic_content_block_type_for_file(
|
||||
format: str,
|
||||
) -> Literal["document", "image", "container_upload"]:
|
||||
if format == "application/pdf" or format == "text/plain":
|
||||
return "document"
|
||||
elif format in ["image/jpeg", "image/png", "image/gif", "image/webp"]:
|
||||
return "image"
|
||||
else:
|
||||
return "container_upload"
|
||||
|
||||
|
||||
def anthropic_process_openai_file_message(
|
||||
message: ChatCompletionFileObject,
|
||||
) -> Union[
|
||||
AnthropicMessagesDocumentParam,
|
||||
AnthropicMessagesImageParam,
|
||||
AnthropicMessagesContainerUploadParam,
|
||||
]:
|
||||
file_message = cast(ChatCompletionFileObject, message)
|
||||
file_data = file_message["file"].get("file_data")
|
||||
file_id = file_message["file"].get("file_id")
|
||||
format = file_message["file"].get("format")
|
||||
if file_data:
|
||||
image_chunk = convert_to_anthropic_image_obj(
|
||||
openai_image_url=file_data,
|
||||
format=format,
|
||||
)
|
||||
anthropic_document_param = AnthropicMessagesDocumentParam(
|
||||
type="document",
|
||||
source=AnthropicContentParamSource(
|
||||
type="base64",
|
||||
media_type=image_chunk["media_type"],
|
||||
data=image_chunk["data"],
|
||||
),
|
||||
)
|
||||
return anthropic_document_param
|
||||
elif file_id:
|
||||
content_block_type = (
|
||||
select_anthropic_content_block_type_for_file(format)
|
||||
if format
|
||||
else "container_upload"
|
||||
)
|
||||
return_block_param: Optional[
|
||||
Union[
|
||||
AnthropicMessagesDocumentParam,
|
||||
AnthropicMessagesImageParam,
|
||||
AnthropicMessagesContainerUploadParam,
|
||||
]
|
||||
] = None
|
||||
if content_block_type == "document":
|
||||
return_block_param = AnthropicMessagesDocumentParam(
|
||||
type="document",
|
||||
source=AnthropicContentParamSourceFileId(
|
||||
type="file",
|
||||
file_id=file_id,
|
||||
),
|
||||
)
|
||||
elif content_block_type == "image":
|
||||
return_block_param = AnthropicMessagesImageParam(
|
||||
type="image",
|
||||
source=AnthropicContentParamSourceFileId(
|
||||
type="file",
|
||||
file_id=file_id,
|
||||
),
|
||||
)
|
||||
elif content_block_type == "container_upload":
|
||||
return_block_param = AnthropicMessagesContainerUploadParam(
|
||||
type="container_upload", file_id=file_id
|
||||
)
|
||||
|
||||
if return_block_param is None:
|
||||
raise Exception(f"Unable to parse anthropic file message: {message}")
|
||||
return return_block_param
|
||||
raise Exception(
|
||||
f"Either file_data or file_id must be present in the file message: {message}"
|
||||
)
|
||||
|
||||
|
||||
def anthropic_messages_pt( # noqa: PLR0915
|
||||
messages: List[AllMessageValues],
|
||||
model: str,
|
||||
|
|
@ -1489,24 +1567,11 @@ def anthropic_messages_pt( # noqa: PLR0915
|
|||
elif m.get("type", "") == "document":
|
||||
user_content.append(cast(AnthropicMessagesDocumentParam, m))
|
||||
elif m.get("type", "") == "file":
|
||||
file_message = cast(ChatCompletionFileObject, m)
|
||||
file_data = file_message["file"].get("file_data")
|
||||
if file_data:
|
||||
image_chunk = convert_to_anthropic_image_obj(
|
||||
openai_image_url=file_data,
|
||||
format=file_message["file"].get("format"),
|
||||
user_content.append(
|
||||
anthropic_process_openai_file_message(
|
||||
cast(ChatCompletionFileObject, m)
|
||||
)
|
||||
anthropic_document_param = (
|
||||
AnthropicMessagesDocumentParam(
|
||||
type="document",
|
||||
source=AnthropicContentParamSource(
|
||||
type="base64",
|
||||
media_type=image_chunk["media_type"],
|
||||
data=image_chunk["data"],
|
||||
),
|
||||
)
|
||||
)
|
||||
user_content.append(anthropic_document_param)
|
||||
)
|
||||
elif isinstance(user_message_types_block["content"], str):
|
||||
_anthropic_content_text_element: AnthropicMessagesTextParam = {
|
||||
"type": "text",
|
||||
|
|
|
|||
|
|
@ -1,10 +1,56 @@
|
|||
"""
|
||||
This is a cache for LangfuseLoggers.
|
||||
|
||||
Langfuse Python SDK initializes a thread for each client.
|
||||
|
||||
This ensures we do
|
||||
1. Proper cleanup of Langfuse initialized clients.
|
||||
2. Re-use created langfuse clients.
|
||||
"""
|
||||
import hashlib
|
||||
import json
|
||||
from typing import Any, Optional
|
||||
|
||||
import litellm
|
||||
from litellm.constants import _DEFAULT_TTL_FOR_HTTPX_CLIENTS
|
||||
|
||||
from ...caching import InMemoryCache
|
||||
|
||||
|
||||
class LangfuseInMemoryCache(InMemoryCache):
|
||||
"""
|
||||
Ensures we do proper cleanup of Langfuse initialized clients.
|
||||
|
||||
Langfuse Python SDK initializes a thread for each client, we need to call Langfuse.shutdown() to properly cleanup.
|
||||
|
||||
This ensures we do proper cleanup of Langfuse initialized clients.
|
||||
"""
|
||||
|
||||
def _remove_key(self, key: str) -> None:
|
||||
"""
|
||||
Override _remove_key in InMemoryCache to ensure we do proper cleanup of Langfuse initialized clients.
|
||||
|
||||
LangfuseLoggers consume threads when initalized, this shuts them down when they are expired
|
||||
|
||||
Relevant Issue: https://github.com/BerriAI/litellm/issues/11169
|
||||
"""
|
||||
from litellm.integrations.langfuse.langfuse import LangFuseLogger
|
||||
|
||||
if isinstance(self.cache_dict[key], LangFuseLogger):
|
||||
_created_langfuse_logger: LangFuseLogger = self.cache_dict[key]
|
||||
#########################################################
|
||||
# Clean up Langfuse initialized clients
|
||||
#########################################################
|
||||
litellm.initialized_langfuse_clients -= 1
|
||||
_created_langfuse_logger.Langfuse.flush()
|
||||
_created_langfuse_logger.Langfuse.shutdown()
|
||||
|
||||
#########################################################
|
||||
# Call parent class to remove key from cache
|
||||
#########################################################
|
||||
return super()._remove_key(key)
|
||||
|
||||
|
||||
class DynamicLoggingCache:
|
||||
"""
|
||||
Prevent memory leaks caused by initializing new logging clients on each request.
|
||||
|
|
@ -13,7 +59,7 @@ class DynamicLoggingCache:
|
|||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.cache = InMemoryCache()
|
||||
self.cache = LangfuseInMemoryCache(default_ttl=_DEFAULT_TTL_FOR_HTTPX_CLIENTS)
|
||||
|
||||
def get_cache_key(self, args: dict) -> str:
|
||||
args_str = json.dumps(args, sort_keys=True)
|
||||
|
|
|
|||
|
|
@ -1359,6 +1359,7 @@ class CustomStreamWrapper:
|
|||
print_verbose(f"self.sent_first_chunk: {self.sent_first_chunk}")
|
||||
|
||||
## CHECK FOR TOOL USE
|
||||
|
||||
if "tool_calls" in completion_obj and len(completion_obj["tool_calls"]) > 0:
|
||||
if self.is_function_call is True: # user passed in 'functions' param
|
||||
completion_obj["function_call"] = completion_obj["tool_calls"][0][
|
||||
|
|
@ -1633,7 +1634,8 @@ class CustomStreamWrapper:
|
|||
if is_async_iterable(self.completion_stream):
|
||||
async for chunk in self.completion_stream:
|
||||
if chunk == "None" or chunk is None:
|
||||
raise Exception
|
||||
continue # skip None chunks
|
||||
|
||||
elif (
|
||||
self.custom_llm_provider == "gemini"
|
||||
and hasattr(chunk, "parts")
|
||||
|
|
@ -1642,7 +1644,9 @@ class CustomStreamWrapper:
|
|||
continue
|
||||
# chunk_creator() does logging/stream chunk building. We need to let it know its being called in_async_func, so we don't double add chunks.
|
||||
# __anext__ also calls async_success_handler, which does logging
|
||||
print_verbose(f"PROCESSED ASYNC CHUNK PRE CHUNK CREATOR: {chunk}")
|
||||
verbose_logger.debug(
|
||||
f"PROCESSED ASYNC CHUNK PRE CHUNK CREATOR: {chunk}"
|
||||
)
|
||||
|
||||
processed_chunk: Optional[ModelResponseStream] = self.chunk_creator(
|
||||
chunk=chunk
|
||||
|
|
|
|||
|
|
@ -362,6 +362,15 @@ def token_counter(
|
|||
"""
|
||||
from litellm.utils import convert_list_message_to_dict
|
||||
|
||||
#########################################################
|
||||
# Flag to disable token counter
|
||||
# We've gotten reports of this consuming CPU cycles,
|
||||
# exposing this flag to allow users to disable
|
||||
# it to confirm if this is indeed the issue
|
||||
#########################################################
|
||||
if litellm.disable_token_counter is True:
|
||||
return 0
|
||||
|
||||
verbose_logger.debug(
|
||||
f"messages in token_counter: {messages}, text in token_counter: {text}"
|
||||
)
|
||||
|
|
|
|||
|
|
@ -18,7 +18,9 @@ from litellm.litellm_core_utils.prompt_templates.factory import anthropic_messag
|
|||
from litellm.llms.base_llm.base_utils import type_to_response_format_param
|
||||
from litellm.llms.base_llm.chat.transformation import BaseConfig, BaseLLMException
|
||||
from litellm.types.llms.anthropic import (
|
||||
AllAnthropicMessageValues,
|
||||
AllAnthropicToolsValues,
|
||||
AnthropicCodeExecutionTool,
|
||||
AnthropicComputerTool,
|
||||
AnthropicHostedTools,
|
||||
AnthropicInputSchema,
|
||||
|
|
@ -530,6 +532,40 @@ class AnthropicConfig(AnthropicModelInfo, BaseConfig):
|
|||
|
||||
return anthropic_system_message_list
|
||||
|
||||
def add_code_execution_tool(
|
||||
self,
|
||||
messages: List[AllAnthropicMessageValues],
|
||||
tools: List[Union[AllAnthropicToolsValues, Dict]],
|
||||
) -> List[Union[AllAnthropicToolsValues, Dict]]:
|
||||
"""if 'container_upload' in messages, add code_execution tool"""
|
||||
add_code_execution_tool = False
|
||||
for message in messages:
|
||||
message_content = message.get("content", None)
|
||||
if message_content and isinstance(message_content, list):
|
||||
for content in message_content:
|
||||
content_type = content.get("type", None)
|
||||
if content_type == "container_upload":
|
||||
add_code_execution_tool = True
|
||||
break
|
||||
|
||||
if add_code_execution_tool:
|
||||
## check if code_execution tool is already in tools
|
||||
for tool in tools:
|
||||
tool_type = tool.get("type", None)
|
||||
if (
|
||||
tool_type
|
||||
and isinstance(tool_type, str)
|
||||
and tool_type.startswith("code_execution")
|
||||
):
|
||||
return tools
|
||||
tools.append(
|
||||
AnthropicCodeExecutionTool(
|
||||
name="code_execution",
|
||||
type="code_execution_20250522",
|
||||
)
|
||||
)
|
||||
return tools
|
||||
|
||||
def transform_request(
|
||||
self,
|
||||
model: str,
|
||||
|
|
@ -579,6 +615,18 @@ class AnthropicConfig(AnthropicModelInfo, BaseConfig):
|
|||
message="{}\nReceived Messages={}".format(str(e), messages),
|
||||
) # don't use verbose_logger.exception, if exception is raised
|
||||
|
||||
## Add code_execution tool if container_upload is in messages
|
||||
_tools = (
|
||||
cast(
|
||||
Optional[List[Union[AllAnthropicToolsValues, Dict]]],
|
||||
optional_params.get("tools"),
|
||||
)
|
||||
or []
|
||||
)
|
||||
tools = self.add_code_execution_tool(messages=anthropic_messages, tools=_tools)
|
||||
if len(tools) > 1:
|
||||
optional_params["tools"] = tools
|
||||
|
||||
## Load Config
|
||||
config = litellm.AnthropicConfig.get_config()
|
||||
for k, v in config.items():
|
||||
|
|
|
|||
|
|
@ -7,6 +7,9 @@ from typing import Dict, List, Optional, Union
|
|||
import httpx
|
||||
|
||||
import litellm
|
||||
from litellm.litellm_core_utils.prompt_templates.common_utils import (
|
||||
get_file_ids_from_messages,
|
||||
)
|
||||
from litellm.llms.base_llm.base_utils import BaseLLMModelInfo
|
||||
from litellm.llms.base_llm.chat.transformation import BaseLLMException
|
||||
from litellm.secret_managers.main import get_secret_str
|
||||
|
|
@ -42,6 +45,13 @@ class AnthropicModelInfo(BaseLLMModelInfo):
|
|||
|
||||
return False
|
||||
|
||||
def is_file_id_used(self, messages: List[AllMessageValues]) -> bool:
|
||||
"""
|
||||
Return if {"source": {"type": "file", "file_id": ..}} in message content block
|
||||
"""
|
||||
file_ids = get_file_ids_from_messages(messages)
|
||||
return len(file_ids) > 0
|
||||
|
||||
def is_computer_tool_used(
|
||||
self, tools: Optional[List[AllAnthropicToolsValues]]
|
||||
) -> bool:
|
||||
|
|
@ -82,6 +92,7 @@ class AnthropicModelInfo(BaseLLMModelInfo):
|
|||
computer_tool_used: bool = False,
|
||||
prompt_caching_set: bool = False,
|
||||
pdf_used: bool = False,
|
||||
file_id_used: bool = False,
|
||||
is_vertex_request: bool = False,
|
||||
user_anthropic_beta_headers: Optional[List[str]] = None,
|
||||
) -> dict:
|
||||
|
|
@ -90,8 +101,11 @@ class AnthropicModelInfo(BaseLLMModelInfo):
|
|||
betas.add("prompt-caching-2024-07-31")
|
||||
if computer_tool_used:
|
||||
betas.add("computer-use-2024-10-22")
|
||||
if pdf_used:
|
||||
betas.add("pdfs-2024-09-25")
|
||||
# if pdf_used:
|
||||
# betas.add("pdfs-2024-09-25")
|
||||
if file_id_used:
|
||||
betas.add("files-api-2025-04-14")
|
||||
betas.add("code-execution-2025-05-22")
|
||||
headers = {
|
||||
"anthropic-version": anthropic_version or "2023-06-01",
|
||||
"x-api-key": api_key,
|
||||
|
|
@ -131,6 +145,7 @@ class AnthropicModelInfo(BaseLLMModelInfo):
|
|||
prompt_caching_set = self.is_cache_control_set(messages=messages)
|
||||
computer_tool_used = self.is_computer_tool_used(tools=tools)
|
||||
pdf_used = self.is_pdf_used(messages=messages)
|
||||
file_id_used = self.is_file_id_used(messages=messages)
|
||||
user_anthropic_beta_headers = self._get_user_anthropic_beta_headers(
|
||||
anthropic_beta_header=headers.get("anthropic-beta")
|
||||
)
|
||||
|
|
@ -139,6 +154,7 @@ class AnthropicModelInfo(BaseLLMModelInfo):
|
|||
prompt_caching_set=prompt_caching_set,
|
||||
pdf_used=pdf_used,
|
||||
api_key=api_key,
|
||||
file_id_used=file_id_used,
|
||||
is_vertex_request=optional_params.get("is_vertex_request", False),
|
||||
user_anthropic_beta_headers=user_anthropic_beta_headers,
|
||||
)
|
||||
|
|
|
|||
|
|
@ -140,7 +140,7 @@ def anthropic_messages_handler(
|
|||
)
|
||||
if anthropic_messages_provider_config is None:
|
||||
raise ValueError(
|
||||
f"Anthropic messages provider config not found for model: {model}"
|
||||
f"Anthropic messages provider config not found for model: {model}, custom_llm_provider: {custom_llm_provider}"
|
||||
)
|
||||
if custom_llm_provider is None:
|
||||
raise ValueError(
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
from typing import Any, AsyncIterator, Dict, List, Optional
|
||||
from typing import Any, AsyncIterator, Dict, List, Optional, Tuple
|
||||
|
||||
import httpx
|
||||
|
||||
|
|
@ -50,7 +50,7 @@ class AnthropicMessagesConfig(BaseAnthropicMessagesConfig):
|
|||
api_base = f"{api_base}/v1/messages"
|
||||
return api_base
|
||||
|
||||
def validate_environment(
|
||||
def validate_anthropic_messages_environment(
|
||||
self,
|
||||
headers: dict,
|
||||
model: str,
|
||||
|
|
@ -59,14 +59,14 @@ class AnthropicMessagesConfig(BaseAnthropicMessagesConfig):
|
|||
litellm_params: dict,
|
||||
api_key: Optional[str] = None,
|
||||
api_base: Optional[str] = None,
|
||||
) -> dict:
|
||||
if "x-api-key" not in headers:
|
||||
) -> Tuple[dict, Optional[str]]:
|
||||
if "x-api-key" not in headers and api_key:
|
||||
headers["x-api-key"] = api_key
|
||||
if "anthropic-version" not in headers:
|
||||
headers["anthropic-version"] = DEFAULT_ANTHROPIC_API_VERSION
|
||||
if "content-type" not in headers:
|
||||
headers["content-type"] = "application/json"
|
||||
return headers
|
||||
return headers, api_base
|
||||
|
||||
def transform_anthropic_messages_request(
|
||||
self,
|
||||
|
|
|
|||
|
|
@ -94,7 +94,7 @@ class AzureAudioTranscription(AzureChatCompletion):
|
|||
additional_args={"complete_input_dict": data},
|
||||
original_response=stringified_response,
|
||||
)
|
||||
hidden_params = {"model": "whisper-1", "custom_llm_provider": "azure"}
|
||||
hidden_params = {"model": model, "custom_llm_provider": "azure"}
|
||||
final_response: TranscriptionResponse = convert_to_model_response_object(response_object=stringified_response, model_response_object=model_response, hidden_params=hidden_params, response_type="audio_transcription") # type: ignore
|
||||
return final_response
|
||||
|
||||
|
|
@ -174,7 +174,7 @@ class AzureAudioTranscription(AzureChatCompletion):
|
|||
},
|
||||
original_response=stringified_response,
|
||||
)
|
||||
hidden_params = {"model": "whisper-1", "custom_llm_provider": "azure"}
|
||||
hidden_params = {"model": model, "custom_llm_provider": "azure"}
|
||||
response = convert_to_model_response_object(
|
||||
_response_headers=headers,
|
||||
response_object=stringified_response,
|
||||
|
|
|
|||
|
|
@ -18,7 +18,7 @@ else:
|
|||
|
||||
class BaseAnthropicMessagesConfig(ABC):
|
||||
@abstractmethod
|
||||
def validate_environment(
|
||||
def validate_anthropic_messages_environment( # use different name because return type is different from base config's validate_environment
|
||||
self,
|
||||
headers: dict,
|
||||
model: str,
|
||||
|
|
@ -27,13 +27,17 @@ class BaseAnthropicMessagesConfig(ABC):
|
|||
litellm_params: dict,
|
||||
api_key: Optional[str] = None,
|
||||
api_base: Optional[str] = None,
|
||||
) -> dict:
|
||||
) -> Tuple[dict, Optional[str]]:
|
||||
"""
|
||||
OPTIONAL
|
||||
|
||||
Validate the environment for the request
|
||||
|
||||
Returns:
|
||||
- headers: dict
|
||||
- api_base: Optional[str] - If the provider needs to update the api_base, return it here. Otherwise, return None.
|
||||
"""
|
||||
return headers
|
||||
return headers, api_base
|
||||
|
||||
@abstractmethod
|
||||
def get_complete_url(
|
||||
|
|
|
|||
|
|
@ -336,6 +336,36 @@ class BaseAWSLLM:
|
|||
|
||||
return aws_region_name
|
||||
|
||||
def get_aws_region_name_for_non_llm_api_calls(
|
||||
self,
|
||||
aws_region_name: Optional[str] = None,
|
||||
):
|
||||
"""
|
||||
Get the AWS region name for non-llm api calls.
|
||||
|
||||
LLM API calls check the model arn and end up using that as the region name.
|
||||
|
||||
For non-llm api calls eg. Guardrails, Vector Stores we just need to check the dynamic param or env vars.
|
||||
"""
|
||||
if aws_region_name is None:
|
||||
# check env #
|
||||
litellm_aws_region_name = get_secret("AWS_REGION_NAME", None)
|
||||
|
||||
if litellm_aws_region_name is not None and isinstance(
|
||||
litellm_aws_region_name, str
|
||||
):
|
||||
aws_region_name = litellm_aws_region_name
|
||||
|
||||
standard_aws_region_name = get_secret("AWS_REGION", None)
|
||||
if standard_aws_region_name is not None and isinstance(
|
||||
standard_aws_region_name, str
|
||||
):
|
||||
aws_region_name = standard_aws_region_name
|
||||
|
||||
if aws_region_name is None:
|
||||
aws_region_name = "us-west-2"
|
||||
return aws_region_name
|
||||
|
||||
@tracer.wrap()
|
||||
def _auth_with_web_identity_token(
|
||||
self,
|
||||
|
|
@ -527,6 +557,7 @@ class BaseAWSLLM:
|
|||
api_base: Optional[str],
|
||||
aws_bedrock_runtime_endpoint: Optional[str],
|
||||
aws_region_name: str,
|
||||
endpoint_type: Optional[Literal["runtime", "agent"]] = "runtime",
|
||||
) -> Tuple[str, str]:
|
||||
env_aws_bedrock_runtime_endpoint = get_secret("AWS_BEDROCK_RUNTIME_ENDPOINT")
|
||||
if api_base is not None:
|
||||
|
|
@ -540,7 +571,10 @@ class BaseAWSLLM:
|
|||
):
|
||||
endpoint_url = env_aws_bedrock_runtime_endpoint
|
||||
else:
|
||||
endpoint_url = f"https://bedrock-runtime.{aws_region_name}.amazonaws.com"
|
||||
endpoint_url = self._select_default_endpoint_url(
|
||||
endpoint_type=endpoint_type,
|
||||
aws_region_name=aws_region_name,
|
||||
)
|
||||
|
||||
# Determine proxy_endpoint_url
|
||||
if env_aws_bedrock_runtime_endpoint and isinstance(
|
||||
|
|
@ -556,6 +590,19 @@ class BaseAWSLLM:
|
|||
|
||||
return endpoint_url, proxy_endpoint_url
|
||||
|
||||
def _select_default_endpoint_url(
|
||||
self, endpoint_type: Optional[Literal["runtime", "agent"]], aws_region_name: str
|
||||
) -> str:
|
||||
"""
|
||||
Select the default endpoint url based on the endpoint type
|
||||
|
||||
Default endpoint url is https://bedrock-runtime.{aws_region_name}.amazonaws.com
|
||||
"""
|
||||
if endpoint_type == "agent":
|
||||
return f"https://bedrock-agent-runtime.{aws_region_name}.amazonaws.com"
|
||||
else:
|
||||
return f"https://bedrock-runtime.{aws_region_name}.amazonaws.com"
|
||||
|
||||
def _get_boto_credentials_from_optional_params(
|
||||
self, optional_params: dict, model: Optional[str] = None
|
||||
) -> Boto3CredentialsInfo:
|
||||
|
|
|
|||
527
litellm/llms/bedrock/chat/invoke_agent/transformation.py
Normal file
527
litellm/llms/bedrock/chat/invoke_agent/transformation.py
Normal file
|
|
@ -0,0 +1,527 @@
|
|||
"""
|
||||
Transformation for Bedrock Invoke Agent
|
||||
|
||||
https://docs.aws.amazon.com/bedrock/latest/APIReference/API_agent-runtime_InvokeAgent.html
|
||||
"""
|
||||
import base64
|
||||
import json
|
||||
import uuid
|
||||
from typing import TYPE_CHECKING, Any, Dict, List, Optional, Tuple, Union
|
||||
|
||||
import httpx
|
||||
|
||||
from litellm._logging import verbose_logger
|
||||
from litellm.litellm_core_utils.prompt_templates.common_utils import (
|
||||
convert_content_list_to_str,
|
||||
)
|
||||
from litellm.llms.base_llm.chat.transformation import BaseConfig, BaseLLMException
|
||||
from litellm.llms.bedrock.base_aws_llm import BaseAWSLLM
|
||||
from litellm.llms.bedrock.common_utils import BedrockError
|
||||
from litellm.types.llms.bedrock_invoke_agents import (
|
||||
InvokeAgentChunkPayload,
|
||||
InvokeAgentEvent,
|
||||
InvokeAgentEventHeaders,
|
||||
InvokeAgentEventList,
|
||||
InvokeAgentTrace,
|
||||
InvokeAgentTracePayload,
|
||||
InvokeAgentUsage,
|
||||
)
|
||||
from litellm.types.llms.openai import AllMessageValues
|
||||
from litellm.types.utils import Choices, Message, ModelResponse
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.litellm_core_utils.litellm_logging import Logging as _LiteLLMLoggingObj
|
||||
|
||||
LiteLLMLoggingObj = _LiteLLMLoggingObj
|
||||
else:
|
||||
LiteLLMLoggingObj = Any
|
||||
|
||||
|
||||
class AmazonInvokeAgentConfig(BaseConfig, BaseAWSLLM):
|
||||
def __init__(self, **kwargs):
|
||||
BaseConfig.__init__(self, **kwargs)
|
||||
BaseAWSLLM.__init__(self, **kwargs)
|
||||
|
||||
def get_supported_openai_params(self, model: str) -> List[str]:
|
||||
"""
|
||||
This is a base invoke agent model mapping. For Invoke Agent - define a bedrock provider specific config that extends this class.
|
||||
|
||||
Bedrock Invoke Agents has 0 OpenAI compatible params
|
||||
|
||||
As of May 29th, 2025 - they don't support streaming.
|
||||
"""
|
||||
return []
|
||||
|
||||
def map_openai_params(
|
||||
self,
|
||||
non_default_params: dict,
|
||||
optional_params: dict,
|
||||
model: str,
|
||||
drop_params: bool,
|
||||
) -> dict:
|
||||
"""
|
||||
This is a base invoke agent model mapping. For Invoke Agent - define a bedrock provider specific config that extends this class.
|
||||
"""
|
||||
return optional_params
|
||||
|
||||
def get_complete_url(
|
||||
self,
|
||||
api_base: Optional[str],
|
||||
api_key: Optional[str],
|
||||
model: str,
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
stream: Optional[bool] = None,
|
||||
) -> str:
|
||||
"""
|
||||
Get the complete url for the request
|
||||
"""
|
||||
### SET RUNTIME ENDPOINT ###
|
||||
aws_bedrock_runtime_endpoint = optional_params.get(
|
||||
"aws_bedrock_runtime_endpoint", None
|
||||
) # https://bedrock-runtime.{region_name}.amazonaws.com
|
||||
endpoint_url, _ = self.get_runtime_endpoint(
|
||||
api_base=api_base,
|
||||
aws_bedrock_runtime_endpoint=aws_bedrock_runtime_endpoint,
|
||||
aws_region_name=self._get_aws_region_name(
|
||||
optional_params=optional_params, model=model
|
||||
),
|
||||
endpoint_type="agent",
|
||||
)
|
||||
|
||||
agent_id, agent_alias_id = self._get_agent_id_and_alias_id(model)
|
||||
session_id = self._get_session_id(optional_params)
|
||||
|
||||
endpoint_url = f"{endpoint_url}/agents/{agent_id}/agentAliases/{agent_alias_id}/sessions/{session_id}/text"
|
||||
|
||||
return endpoint_url
|
||||
|
||||
def sign_request(
|
||||
self,
|
||||
headers: dict,
|
||||
optional_params: dict,
|
||||
request_data: dict,
|
||||
api_base: str,
|
||||
model: Optional[str] = None,
|
||||
stream: Optional[bool] = None,
|
||||
fake_stream: Optional[bool] = None,
|
||||
) -> Tuple[dict, Optional[bytes]]:
|
||||
return self._sign_request(
|
||||
service_name="bedrock",
|
||||
headers=headers,
|
||||
optional_params=optional_params,
|
||||
request_data=request_data,
|
||||
api_base=api_base,
|
||||
model=model,
|
||||
stream=stream,
|
||||
fake_stream=fake_stream,
|
||||
)
|
||||
|
||||
def _get_agent_id_and_alias_id(self, model: str) -> tuple[str, str]:
|
||||
"""
|
||||
model = "agent/L1RT58GYRW/MFPSBCXYTW"
|
||||
agent_id = "L1RT58GYRW"
|
||||
agent_alias_id = "MFPSBCXYTW"
|
||||
"""
|
||||
# Split the model string by '/' and extract components
|
||||
parts = model.split("/")
|
||||
if len(parts) != 3 or parts[0] != "agent":
|
||||
raise ValueError(
|
||||
"Invalid model format. Expected format: 'model=agent/AGENT_ID/ALIAS_ID'"
|
||||
)
|
||||
|
||||
return parts[1], parts[2] # Return (agent_id, agent_alias_id)
|
||||
|
||||
def _get_session_id(self, optional_params: dict) -> str:
|
||||
""" """
|
||||
return optional_params.get("sessionID", None) or str(uuid.uuid4())
|
||||
|
||||
def transform_request(
|
||||
self,
|
||||
model: str,
|
||||
messages: List[AllMessageValues],
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
headers: dict,
|
||||
) -> dict:
|
||||
# use the last message content as the query
|
||||
query: str = convert_content_list_to_str(messages[-1])
|
||||
return {
|
||||
"inputText": query,
|
||||
"enableTrace": True,
|
||||
**optional_params,
|
||||
}
|
||||
|
||||
def _parse_aws_event_stream(self, raw_content: bytes) -> InvokeAgentEventList:
|
||||
"""
|
||||
Parse AWS event stream format using boto3/botocore's built-in parser.
|
||||
This is the same approach used in the existing AWSEventStreamDecoder.
|
||||
"""
|
||||
try:
|
||||
from botocore.eventstream import EventStreamBuffer
|
||||
from botocore.parsers import EventStreamJSONParser
|
||||
except ImportError:
|
||||
raise ImportError("boto3/botocore is required for AWS event stream parsing")
|
||||
|
||||
events: InvokeAgentEventList = []
|
||||
parser = EventStreamJSONParser()
|
||||
event_stream_buffer = EventStreamBuffer()
|
||||
|
||||
# Add the entire response to the buffer
|
||||
event_stream_buffer.add_data(raw_content)
|
||||
|
||||
# Process all events in the buffer
|
||||
for event in event_stream_buffer:
|
||||
try:
|
||||
headers = self._extract_headers_from_event(event)
|
||||
|
||||
event_type = headers.get("event_type", "")
|
||||
|
||||
if event_type == "chunk":
|
||||
# Handle chunk events specially - they contain decoded content, not JSON
|
||||
message = self._parse_message_from_event(event, parser)
|
||||
parsed_event: InvokeAgentEvent = InvokeAgentEvent()
|
||||
if message:
|
||||
# For chunk events, create a payload with the decoded content
|
||||
parsed_event = {
|
||||
"headers": headers,
|
||||
"payload": {
|
||||
"bytes": base64.b64encode(
|
||||
message.encode("utf-8")
|
||||
).decode("utf-8")
|
||||
}, # Re-encode for consistency
|
||||
}
|
||||
events.append(parsed_event)
|
||||
|
||||
elif event_type == "trace":
|
||||
# Handle trace events normally - they contain JSON
|
||||
message = self._parse_message_from_event(event, parser)
|
||||
|
||||
if message:
|
||||
try:
|
||||
event_data = json.loads(message)
|
||||
parsed_event = {
|
||||
"headers": headers,
|
||||
"payload": event_data,
|
||||
}
|
||||
events.append(parsed_event)
|
||||
except json.JSONDecodeError as e:
|
||||
verbose_logger.warning(
|
||||
f"Failed to parse trace event JSON: {e}"
|
||||
)
|
||||
else:
|
||||
verbose_logger.debug(f"Unknown event type: {event_type}")
|
||||
|
||||
except Exception as e:
|
||||
verbose_logger.error(f"Error processing event: {e}")
|
||||
continue
|
||||
|
||||
return events
|
||||
|
||||
def _parse_message_from_event(self, event, parser) -> Optional[str]:
|
||||
"""Extract message content from an AWS event, adapted from AWSEventStreamDecoder."""
|
||||
try:
|
||||
response_dict = event.to_response_dict()
|
||||
verbose_logger.debug(f"Response dict: {response_dict}")
|
||||
|
||||
# Use the same response shape parsing as the existing decoder
|
||||
parsed_response = parser.parse(
|
||||
response_dict, self._get_response_stream_shape()
|
||||
)
|
||||
verbose_logger.debug(f"Parsed response: {parsed_response}")
|
||||
|
||||
if response_dict["status_code"] != 200:
|
||||
decoded_body = response_dict["body"].decode()
|
||||
if isinstance(decoded_body, dict):
|
||||
error_message = decoded_body.get("message")
|
||||
elif isinstance(decoded_body, str):
|
||||
error_message = decoded_body
|
||||
else:
|
||||
error_message = ""
|
||||
exception_status = response_dict["headers"].get(":exception-type")
|
||||
error_message = exception_status + " " + error_message
|
||||
raise BedrockError(
|
||||
status_code=response_dict["status_code"],
|
||||
message=(
|
||||
json.dumps(error_message)
|
||||
if isinstance(error_message, dict)
|
||||
else error_message
|
||||
),
|
||||
)
|
||||
|
||||
if "chunk" in parsed_response:
|
||||
chunk = parsed_response.get("chunk")
|
||||
if not chunk:
|
||||
return None
|
||||
return chunk.get("bytes").decode()
|
||||
else:
|
||||
chunk = response_dict.get("body")
|
||||
if not chunk:
|
||||
return None
|
||||
return chunk.decode()
|
||||
|
||||
except Exception as e:
|
||||
verbose_logger.debug(f"Error parsing message from event: {e}")
|
||||
return None
|
||||
|
||||
def _extract_headers_from_event(self, event) -> InvokeAgentEventHeaders:
|
||||
"""Extract headers from an AWS event for categorization."""
|
||||
try:
|
||||
response_dict = event.to_response_dict()
|
||||
headers = response_dict.get("headers", {})
|
||||
|
||||
# Extract the event-type and content-type headers that we care about
|
||||
return InvokeAgentEventHeaders(
|
||||
event_type=headers.get(":event-type", ""),
|
||||
content_type=headers.get(":content-type", ""),
|
||||
message_type=headers.get(":message-type", ""),
|
||||
)
|
||||
except Exception as e:
|
||||
verbose_logger.debug(f"Error extracting headers: {e}")
|
||||
return InvokeAgentEventHeaders(
|
||||
event_type="", content_type="", message_type=""
|
||||
)
|
||||
|
||||
def _get_response_stream_shape(self):
|
||||
"""Get the response stream shape for parsing, reusing existing logic."""
|
||||
try:
|
||||
# Try to reuse the cached shape from the existing decoder
|
||||
from litellm.llms.bedrock.chat.invoke_handler import (
|
||||
get_response_stream_shape,
|
||||
)
|
||||
|
||||
return get_response_stream_shape()
|
||||
except ImportError:
|
||||
# Fallback: create our own shape
|
||||
try:
|
||||
from botocore.loaders import Loader
|
||||
from botocore.model import ServiceModel
|
||||
|
||||
loader = Loader()
|
||||
bedrock_service_dict = loader.load_service_model(
|
||||
"bedrock-runtime", "service-2"
|
||||
)
|
||||
bedrock_service_model = ServiceModel(bedrock_service_dict)
|
||||
return bedrock_service_model.shape_for("ResponseStream")
|
||||
except Exception as e:
|
||||
verbose_logger.warning(f"Could not load response stream shape: {e}")
|
||||
return None
|
||||
|
||||
def _extract_response_content(self, events: InvokeAgentEventList) -> str:
|
||||
"""Extract the final response content from parsed events."""
|
||||
response_parts = []
|
||||
|
||||
for event in events:
|
||||
headers = event.get("headers", {})
|
||||
payload = event.get("payload")
|
||||
|
||||
event_type = headers.get(
|
||||
"event_type"
|
||||
) # Note: using event_type not event-type
|
||||
|
||||
if event_type == "chunk" and payload:
|
||||
# Extract base64 encoded content from chunk events
|
||||
chunk_payload: InvokeAgentChunkPayload = payload # type: ignore
|
||||
encoded_bytes = chunk_payload.get("bytes", "")
|
||||
if encoded_bytes:
|
||||
try:
|
||||
decoded_content = base64.b64decode(encoded_bytes).decode(
|
||||
"utf-8"
|
||||
)
|
||||
response_parts.append(decoded_content)
|
||||
except Exception as e:
|
||||
verbose_logger.warning(f"Failed to decode chunk content: {e}")
|
||||
|
||||
return "".join(response_parts)
|
||||
|
||||
def _extract_usage_info(self, events: InvokeAgentEventList) -> InvokeAgentUsage:
|
||||
"""Extract token usage information from trace events."""
|
||||
usage_info = InvokeAgentUsage(
|
||||
inputTokens=0,
|
||||
outputTokens=0,
|
||||
model=None,
|
||||
)
|
||||
|
||||
response_model: Optional[str] = None
|
||||
|
||||
for event in events:
|
||||
if not self._is_trace_event(event):
|
||||
continue
|
||||
|
||||
trace_data = self._get_trace_data(event)
|
||||
if not trace_data:
|
||||
continue
|
||||
|
||||
verbose_logger.debug(f"Trace event: {trace_data}")
|
||||
|
||||
# Extract usage from pre-processing trace
|
||||
self._extract_and_update_preprocessing_usage(
|
||||
trace_data=trace_data,
|
||||
usage_info=usage_info,
|
||||
)
|
||||
|
||||
# Extract model from orchestration trace
|
||||
if response_model is None:
|
||||
response_model = self._extract_orchestration_model(trace_data)
|
||||
|
||||
usage_info["model"] = response_model
|
||||
return usage_info
|
||||
|
||||
def _is_trace_event(self, event: InvokeAgentEvent) -> bool:
|
||||
"""Check if the event is a trace event."""
|
||||
headers = event.get("headers", {})
|
||||
event_type = headers.get("event_type")
|
||||
payload = event.get("payload")
|
||||
return event_type == "trace" and payload is not None
|
||||
|
||||
def _get_trace_data(self, event: InvokeAgentEvent) -> Optional[InvokeAgentTrace]:
|
||||
"""Extract trace data from a trace event."""
|
||||
payload = event.get("payload")
|
||||
if not payload:
|
||||
return None
|
||||
|
||||
trace_payload: InvokeAgentTracePayload = payload # type: ignore
|
||||
return trace_payload.get("trace", {})
|
||||
|
||||
def _extract_and_update_preprocessing_usage(
|
||||
self, trace_data: InvokeAgentTrace, usage_info: InvokeAgentUsage
|
||||
) -> None:
|
||||
"""Extract usage information from preprocessing trace."""
|
||||
pre_processing = trace_data.get("preProcessingTrace", {})
|
||||
if not pre_processing:
|
||||
return
|
||||
|
||||
model_output = pre_processing.get("modelInvocationOutput", {})
|
||||
if not model_output:
|
||||
return
|
||||
|
||||
metadata = model_output.get("metadata", {})
|
||||
if not metadata:
|
||||
return
|
||||
|
||||
usage: Optional[Union[InvokeAgentUsage, Dict]] = metadata.get("usage", {})
|
||||
if not usage:
|
||||
return
|
||||
|
||||
usage_info["inputTokens"] += usage.get("inputTokens", 0)
|
||||
usage_info["outputTokens"] += usage.get("outputTokens", 0)
|
||||
|
||||
def _extract_orchestration_model(
|
||||
self, trace_data: InvokeAgentTrace
|
||||
) -> Optional[str]:
|
||||
"""Extract model information from orchestration trace."""
|
||||
orchestration_trace = trace_data.get("orchestrationTrace", {})
|
||||
if not orchestration_trace:
|
||||
return None
|
||||
|
||||
model_invocation = orchestration_trace.get("modelInvocationInput", {})
|
||||
if not model_invocation:
|
||||
return None
|
||||
|
||||
return model_invocation.get("foundationModel")
|
||||
|
||||
def _build_model_response(
|
||||
self,
|
||||
content: str,
|
||||
model: str,
|
||||
usage_info: InvokeAgentUsage,
|
||||
model_response: ModelResponse,
|
||||
) -> ModelResponse:
|
||||
"""Build the final ModelResponse object."""
|
||||
|
||||
# Create the message content
|
||||
message = Message(content=content, role="assistant")
|
||||
|
||||
# Create choices
|
||||
choice = Choices(finish_reason="stop", index=0, message=message)
|
||||
|
||||
# Update model response
|
||||
model_response.choices = [choice]
|
||||
model_response.model = usage_info.get("model", model)
|
||||
|
||||
# Add usage information if available
|
||||
if usage_info:
|
||||
from litellm.types.utils import Usage
|
||||
|
||||
usage = Usage(
|
||||
prompt_tokens=usage_info.get("inputTokens", 0),
|
||||
completion_tokens=usage_info.get("outputTokens", 0),
|
||||
total_tokens=usage_info.get("inputTokens", 0)
|
||||
+ usage_info.get("outputTokens", 0),
|
||||
)
|
||||
setattr(model_response, "usage", usage)
|
||||
|
||||
return model_response
|
||||
|
||||
def transform_response(
|
||||
self,
|
||||
model: str,
|
||||
raw_response: httpx.Response,
|
||||
model_response: ModelResponse,
|
||||
logging_obj: LiteLLMLoggingObj,
|
||||
request_data: dict,
|
||||
messages: List[AllMessageValues],
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
encoding: Any,
|
||||
api_key: Optional[str] = None,
|
||||
json_mode: Optional[bool] = None,
|
||||
) -> ModelResponse:
|
||||
try:
|
||||
# Get the raw binary content
|
||||
raw_content = raw_response.content
|
||||
verbose_logger.debug(
|
||||
f"Processing {len(raw_content)} bytes of AWS event stream data"
|
||||
)
|
||||
|
||||
# Parse the AWS event stream format
|
||||
events = self._parse_aws_event_stream(raw_content)
|
||||
verbose_logger.debug(f"Parsed {len(events)} events from stream")
|
||||
|
||||
# Extract response content from chunk events
|
||||
content = self._extract_response_content(events)
|
||||
|
||||
# Extract usage information from trace events
|
||||
usage_info = self._extract_usage_info(events)
|
||||
|
||||
# Build and return the model response
|
||||
return self._build_model_response(
|
||||
content=content,
|
||||
model=model,
|
||||
usage_info=usage_info,
|
||||
model_response=model_response,
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
verbose_logger.error(
|
||||
f"Error processing Bedrock Invoke Agent response: {str(e)}"
|
||||
)
|
||||
raise BedrockError(
|
||||
message=f"Error processing response: {str(e)}",
|
||||
status_code=raw_response.status_code,
|
||||
)
|
||||
|
||||
def validate_environment(
|
||||
self,
|
||||
headers: dict,
|
||||
model: str,
|
||||
messages: List[AllMessageValues],
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
api_key: Optional[str] = None,
|
||||
api_base: Optional[str] = None,
|
||||
) -> dict:
|
||||
return headers
|
||||
|
||||
def get_error_class(
|
||||
self, error_message: str, status_code: int, headers: Union[dict, httpx.Headers]
|
||||
) -> BaseLLMException:
|
||||
return BedrockError(status_code=status_code, message=error_message)
|
||||
|
||||
def should_fake_stream(
|
||||
self,
|
||||
model: Optional[str],
|
||||
stream: Optional[bool],
|
||||
custom_llm_provider: Optional[str] = None,
|
||||
) -> bool:
|
||||
return True
|
||||
|
|
@ -402,7 +402,9 @@ class BedrockModelInfo(BaseLLMModelInfo):
|
|||
return ["us", "eu", "apac"]
|
||||
|
||||
@staticmethod
|
||||
def get_bedrock_route(model: str) -> Literal["converse", "invoke", "converse_like"]:
|
||||
def get_bedrock_route(
|
||||
model: str,
|
||||
) -> Literal["converse", "invoke", "converse_like", "agent"]:
|
||||
"""
|
||||
Get the bedrock route for the given model.
|
||||
"""
|
||||
|
|
@ -414,6 +416,8 @@ class BedrockModelInfo(BaseLLMModelInfo):
|
|||
return "converse_like"
|
||||
elif "converse/" in model:
|
||||
return "converse"
|
||||
elif "agent/" in model:
|
||||
return "agent"
|
||||
elif (
|
||||
base_model in litellm.bedrock_converse_models
|
||||
or alt_model in litellm.bedrock_converse_models
|
||||
|
|
|
|||
|
|
@ -38,6 +38,18 @@ class AmazonAnthropicClaude3MessagesConfig(
|
|||
BaseAnthropicMessagesConfig.__init__(self, **kwargs)
|
||||
AmazonInvokeConfig.__init__(self, **kwargs)
|
||||
|
||||
def validate_anthropic_messages_environment(
|
||||
self,
|
||||
headers: dict,
|
||||
model: str,
|
||||
messages: List[Any],
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
api_key: Optional[str] = None,
|
||||
api_base: Optional[str] = None,
|
||||
) -> Tuple[dict, Optional[str]]:
|
||||
return headers, api_base
|
||||
|
||||
def sign_request(
|
||||
self,
|
||||
headers: dict,
|
||||
|
|
@ -59,18 +71,6 @@ class AmazonAnthropicClaude3MessagesConfig(
|
|||
fake_stream=fake_stream,
|
||||
)
|
||||
|
||||
def validate_environment(
|
||||
self,
|
||||
headers: dict,
|
||||
model: str,
|
||||
messages: List[Any],
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
api_key: Optional[str] = None,
|
||||
api_base: Optional[str] = None,
|
||||
) -> dict:
|
||||
return headers
|
||||
|
||||
def get_complete_url(
|
||||
self,
|
||||
api_base: Optional[str],
|
||||
|
|
|
|||
|
|
@ -5,6 +5,7 @@ from typing import Callable, Dict, Union
|
|||
|
||||
import aiohttp
|
||||
import aiohttp.client_exceptions
|
||||
import aiohttp.http_exceptions
|
||||
import httpx
|
||||
from aiohttp.client import ClientResponse, ClientSession
|
||||
|
||||
|
|
@ -81,13 +82,20 @@ class AiohttpResponseStream(httpx.AsyncByteStream):
|
|||
self.CHUNK_SIZE
|
||||
):
|
||||
yield chunk
|
||||
except aiohttp.ClientPayloadError as e:
|
||||
except (
|
||||
aiohttp.ClientPayloadError,
|
||||
aiohttp.client_exceptions.ClientPayloadError,
|
||||
) as e:
|
||||
# Handle incomplete transfers more gracefully
|
||||
# Log the error but don't re-raise if we've already yielded some data
|
||||
verbose_logger.debug(f"Transfer incomplete, but continuing: {e}")
|
||||
# If the error is due to incomplete transfer encoding, we can still
|
||||
# return what we've received so far, similar to how httpx handles it
|
||||
return
|
||||
except aiohttp.http_exceptions.TransferEncodingError as e:
|
||||
# Handle transfer encoding errors gracefully
|
||||
verbose_logger.debug(f"Transfer encoding error, but continuing: {e}")
|
||||
return
|
||||
except Exception:
|
||||
# For other exceptions, use the normal mapping
|
||||
with map_aiohttp_exceptions():
|
||||
|
|
@ -203,7 +211,6 @@ class LiteLLMAiohttpTransport(AiohttpTransport):
|
|||
data=data,
|
||||
allow_redirects=False,
|
||||
auto_decompress=False,
|
||||
compress=False,
|
||||
timeout=ClientTimeout(
|
||||
sock_connect=timeout.get("connect"),
|
||||
sock_read=timeout.get("read"),
|
||||
|
|
|
|||
|
|
@ -505,20 +505,30 @@ class AsyncHTTPHandler:
|
|||
@staticmethod
|
||||
def _should_use_aiohttp_transport() -> bool:
|
||||
"""
|
||||
This is feature flagged for now and is opt in as we roll out to all users.
|
||||
AiohttpTransport is the default transport for litellm.
|
||||
|
||||
Controlled by either
|
||||
- litellm.use_aiohttp_transport or os.getenv("USE_AIOHTTP_TRANSPORT") = "True"
|
||||
Httpx can be used by the following
|
||||
- litellm.disable_aiohttp_transport = True
|
||||
- os.getenv("DISABLE_AIOHTTP_TRANSPORT") = "True"
|
||||
"""
|
||||
import os
|
||||
|
||||
from litellm.secret_managers.main import str_to_bool
|
||||
|
||||
#########################################################
|
||||
# Check if user disabled aiohttp transport
|
||||
########################################################
|
||||
if (
|
||||
str_to_bool(os.getenv("USE_AIOHTTP_TRANSPORT", "False"))
|
||||
or litellm.use_aiohttp_transport
|
||||
litellm.disable_aiohttp_transport is True
|
||||
or str_to_bool(os.getenv("DISABLE_AIOHTTP_TRANSPORT", "False")) is True
|
||||
):
|
||||
verbose_logger.debug("Using AiohttpTransport...")
|
||||
return True
|
||||
return False
|
||||
return False
|
||||
|
||||
#########################################################
|
||||
# Default: Use AiohttpTransport
|
||||
########################################################
|
||||
verbose_logger.debug("Using AiohttpTransport...")
|
||||
return True
|
||||
|
||||
@staticmethod
|
||||
def _create_aiohttp_transport(
|
||||
|
|
|
|||
|
|
@ -271,7 +271,6 @@ class BaseLLMHTTPHandler:
|
|||
):
|
||||
json_mode: bool = optional_params.pop("json_mode", False)
|
||||
extra_body: Optional[dict] = optional_params.pop("extra_body", None)
|
||||
fake_stream = fake_stream or optional_params.pop("fake_stream", False)
|
||||
|
||||
provider_config = (
|
||||
provider_config
|
||||
|
|
@ -284,6 +283,14 @@ class BaseLLMHTTPHandler:
|
|||
f"Provider config not found for model: {model} and provider: {custom_llm_provider}"
|
||||
)
|
||||
|
||||
fake_stream = (
|
||||
fake_stream
|
||||
or optional_params.pop("fake_stream", False)
|
||||
or provider_config.should_fake_stream(
|
||||
model=model, custom_llm_provider=custom_llm_provider, stream=stream
|
||||
)
|
||||
)
|
||||
|
||||
# get config from model, custom llm provider
|
||||
headers = provider_config.validate_environment(
|
||||
api_key=api_key,
|
||||
|
|
@ -1090,7 +1097,10 @@ class BaseLLMHTTPHandler:
|
|||
if provider_specific_header
|
||||
else {}
|
||||
)
|
||||
headers = anthropic_messages_provider_config.validate_environment(
|
||||
(
|
||||
headers,
|
||||
api_base,
|
||||
) = anthropic_messages_provider_config.validate_anthropic_messages_environment(
|
||||
headers=extra_headers or {},
|
||||
model=model,
|
||||
messages=messages,
|
||||
|
|
|
|||
80
litellm/llms/datarobot/chat/transformation.py
Normal file
80
litellm/llms/datarobot/chat/transformation.py
Normal file
|
|
@ -0,0 +1,80 @@
|
|||
"""
|
||||
Support for OpenAI's `/v1/chat/completions` endpoint.
|
||||
|
||||
Calls done in OpenAI/openai.py as DataRobot is openai-compatible.
|
||||
"""
|
||||
|
||||
from typing import Optional, Tuple
|
||||
from litellm.secret_managers.main import get_secret_str
|
||||
from ...openai_like.chat.transformation import OpenAILikeChatConfig
|
||||
|
||||
|
||||
class DataRobotConfig(OpenAILikeChatConfig):
|
||||
@staticmethod
|
||||
def _resolve_api_key(api_key: Optional[str] = None) -> str:
|
||||
"""Attempt to ensure that the API key is set, preferring the user-provided key
|
||||
over the secret manager key (``DATAROBOT_API_TOKEN``).
|
||||
|
||||
If both are None, a fake API key is returned for testing.
|
||||
"""
|
||||
return api_key or get_secret_str("DATAROBOT_API_TOKEN") or "fake-api-key"
|
||||
|
||||
@staticmethod
|
||||
def _resolve_api_base(api_base: Optional[str] = None) -> Optional[str]:
|
||||
"""Attempt to ensure that the API base is set, preferring the user-provided key
|
||||
over the secret manager key (``DATAROBOT_ENDPOINT``).
|
||||
|
||||
If both are None, a default Llamafile server URL is returned.
|
||||
See: https://github.com/Mozilla-Ocho/llamafile/blob/bd1bbe9aabb1ee12dbdcafa8936db443c571eb9d/README.md#L61
|
||||
"""
|
||||
api_base = api_base or get_secret_str("DATAROBOT_ENDPOINT")
|
||||
|
||||
if api_base is None:
|
||||
api_base = "https://app.datarobot.com"
|
||||
|
||||
# If the api_base is a deployment URL, we do not append the chat completions path
|
||||
if "api/v2/deployments" not in api_base:
|
||||
# If the api_base is not a deployment URL, we need to append the chat completions path
|
||||
if "api/v2/genai/llmgw/chat/completions" not in api_base:
|
||||
api_base += "/api/v2/genai/llmgw/chat/completions"
|
||||
|
||||
# Ensure the url ends with a trailing slash
|
||||
if not api_base.endswith("/"):
|
||||
api_base += "/"
|
||||
|
||||
return api_base # type: ignore
|
||||
|
||||
def _get_openai_compatible_provider_info(
|
||||
self,
|
||||
api_base: Optional[str],
|
||||
api_key: Optional[str]
|
||||
) -> Tuple[Optional[str], Optional[str]]:
|
||||
"""Attempts to ensure that the API base and key are set, preferring user-provided values,
|
||||
before falling back to secret manager values (``DATAROBOT_ENDPOINT`` and ``DATAROBOT_API_TOKEN``
|
||||
respectively).
|
||||
|
||||
If an API key cannot be resolved via either method, a fake key is returned.
|
||||
"""
|
||||
api_base = DataRobotConfig._resolve_api_base(api_base)
|
||||
dynamic_api_key = DataRobotConfig._resolve_api_key(api_key)
|
||||
|
||||
return api_base, dynamic_api_key
|
||||
|
||||
def get_complete_url(
|
||||
self,
|
||||
api_base: Optional[str],
|
||||
api_key: Optional[str],
|
||||
model: str,
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
stream: Optional[bool] = None,
|
||||
) -> str:
|
||||
"""
|
||||
Get the complete URL for the API call. Datarobot's API base is set to
|
||||
the complete value, so it does not need to be updated to additionally add
|
||||
chat completions.
|
||||
|
||||
Returns:
|
||||
str: The complete URL for the API call.
|
||||
"""
|
||||
return str(api_base) # type: ignore
|
||||
|
|
@ -186,11 +186,24 @@ class FireworksAIConfig(OpenAIGPTConfig):
|
|||
"""
|
||||
Add 'transform=inline' to the url of the image_url
|
||||
"""
|
||||
from litellm.litellm_core_utils.prompt_templates.common_utils import (
|
||||
filter_value_from_dict,
|
||||
migrate_file_to_image_url,
|
||||
)
|
||||
|
||||
disable_add_transform_inline_image_block = cast(
|
||||
Optional[bool],
|
||||
litellm_params.get("disable_add_transform_inline_image_block")
|
||||
or litellm.disable_add_transform_inline_image_block,
|
||||
)
|
||||
## For any 'file' message type with pdf content, move to 'image_url' message type
|
||||
for message in messages:
|
||||
if message["role"] == "user":
|
||||
_message_content = message.get("content")
|
||||
if _message_content is not None and isinstance(_message_content, list):
|
||||
for idx, content in enumerate(_message_content):
|
||||
if content["type"] == "file":
|
||||
_message_content[idx] = migrate_file_to_image_url(content)
|
||||
for message in messages:
|
||||
if message["role"] == "user":
|
||||
_message_content = message.get("content")
|
||||
|
|
@ -202,6 +215,8 @@ class FireworksAIConfig(OpenAIGPTConfig):
|
|||
model=model,
|
||||
disable_add_transform_inline_image_block=disable_add_transform_inline_image_block,
|
||||
)
|
||||
filter_value_from_dict(cast(dict, message), "cache_control")
|
||||
|
||||
return messages
|
||||
|
||||
def get_provider_info(self, model: str) -> ProviderSpecificModelInfo:
|
||||
|
|
|
|||
|
|
@ -6,7 +6,7 @@ from litellm.litellm_core_utils.prompt_templates.factory import (
|
|||
convert_to_anthropic_image_obj,
|
||||
)
|
||||
from litellm.types.llms.openai import AllMessageValues
|
||||
from litellm.types.llms.vertex_ai import ContentType, PartType
|
||||
from litellm.types.llms.vertex_ai import ContentType, PartType, SpeechConfig, VoiceConfig, PrebuiltVoiceConfig
|
||||
from litellm.utils import supports_reasoning
|
||||
|
||||
from ...vertex_ai.gemini.transformation import _gemini_convert_messages_with_history
|
||||
|
|
@ -67,6 +67,9 @@ class GoogleAIStudioGeminiConfig(VertexGeminiConfig):
|
|||
def get_config(cls):
|
||||
return super().get_config()
|
||||
|
||||
def is_model_gemini_audio_model(self, model: str) -> bool:
|
||||
return "tts" in model
|
||||
|
||||
def get_supported_openai_params(self, model: str) -> List[str]:
|
||||
supported_params = [
|
||||
"temperature",
|
||||
|
|
@ -84,10 +87,13 @@ class GoogleAIStudioGeminiConfig(VertexGeminiConfig):
|
|||
"frequency_penalty",
|
||||
"modalities",
|
||||
"parallel_tool_calls",
|
||||
"web_search_options",
|
||||
]
|
||||
if supports_reasoning(model):
|
||||
supported_params.append("reasoning_effort")
|
||||
supported_params.append("thinking")
|
||||
if self.is_model_gemini_audio_model(model):
|
||||
supported_params.append("audio")
|
||||
return supported_params
|
||||
|
||||
def map_openai_params(
|
||||
|
|
@ -97,6 +103,40 @@ class GoogleAIStudioGeminiConfig(VertexGeminiConfig):
|
|||
model: str,
|
||||
drop_params: bool,
|
||||
) -> Dict:
|
||||
# Handle audio parameter for TTS models
|
||||
if self.is_model_gemini_audio_model(model):
|
||||
for param, value in non_default_params.items():
|
||||
if param == "audio" and isinstance(value, dict):
|
||||
# Validate audio format - Gemini TTS only supports pcm16
|
||||
audio_format = value.get("format")
|
||||
if audio_format is not None and audio_format != "pcm16":
|
||||
raise ValueError(
|
||||
f"Unsupported audio format for Gemini TTS models: {audio_format}. "
|
||||
f"Gemini TTS models only support 'pcm16' format as they return audio data in L16 PCM format. "
|
||||
f"Please set audio format to 'pcm16'."
|
||||
)
|
||||
|
||||
# Map OpenAI audio parameter to Gemini speech config
|
||||
speech_config: SpeechConfig = {}
|
||||
|
||||
if "voice" in value:
|
||||
prebuilt_voice_config: PrebuiltVoiceConfig = {
|
||||
"voiceName": value["voice"]
|
||||
}
|
||||
voice_config: VoiceConfig = {
|
||||
"prebuiltVoiceConfig": prebuilt_voice_config
|
||||
}
|
||||
speech_config["voiceConfig"] = voice_config
|
||||
|
||||
if speech_config:
|
||||
optional_params["speechConfig"] = speech_config
|
||||
|
||||
# Ensure audio modality is set
|
||||
if "responseModalities" not in optional_params:
|
||||
optional_params["responseModalities"] = ["AUDIO"]
|
||||
elif "AUDIO" not in optional_params["responseModalities"]:
|
||||
optional_params["responseModalities"].append("AUDIO")
|
||||
|
||||
if litellm.vertex_ai_safety_settings is not None:
|
||||
optional_params["safety_settings"] = litellm.vertex_ai_safety_settings
|
||||
return super().map_openai_params(
|
||||
|
|
|
|||
|
|
@ -173,7 +173,7 @@ class OllamaConfig(BaseConfig):
|
|||
if param == "top_p":
|
||||
optional_params["top_p"] = value
|
||||
if param == "frequency_penalty":
|
||||
optional_params["repeat_penalty"] = value
|
||||
optional_params["frequency_penalty"] = value
|
||||
if param == "stop":
|
||||
optional_params["stop"] = value
|
||||
if param == "response_format" and isinstance(value, dict):
|
||||
|
|
|
|||
|
|
@ -155,7 +155,7 @@ class OpenAIAudioTranscription(OpenAIChatCompletion):
|
|||
additional_args={"complete_input_dict": data},
|
||||
original_response=stringified_response,
|
||||
)
|
||||
hidden_params = {"model": "whisper-1", "custom_llm_provider": "openai"}
|
||||
hidden_params = {"model": model, "custom_llm_provider": "openai"}
|
||||
final_response: TranscriptionResponse = convert_to_model_response_object(response_object=stringified_response, model_response_object=model_response, hidden_params=hidden_params, response_type="audio_transcription") # type: ignore
|
||||
return final_response
|
||||
|
||||
|
|
@ -210,7 +210,9 @@ class OpenAIAudioTranscription(OpenAIChatCompletion):
|
|||
additional_args={"complete_input_dict": data},
|
||||
original_response=stringified_response,
|
||||
)
|
||||
hidden_params = {"model": "whisper-1", "custom_llm_provider": "openai"}
|
||||
# Extract the actual model from data instead of hardcoding "whisper-1"
|
||||
actual_model = data.get("model", "whisper-1")
|
||||
hidden_params = {"model": actual_model, "custom_llm_provider": "openai"}
|
||||
return convert_to_model_response_object(response_object=stringified_response, model_response_object=model_response, hidden_params=hidden_params, response_type="audio_transcription") # type: ignore
|
||||
except Exception as e:
|
||||
## LOGGING
|
||||
|
|
|
|||
|
|
@ -43,7 +43,7 @@ class VertexAIBatchPrediction(VertexLLM):
|
|||
custom_llm_provider="vertex_ai",
|
||||
)
|
||||
|
||||
default_api_base = self.create_vertex_url(
|
||||
default_api_base = self.create_vertex_batch_url(
|
||||
vertex_location=vertex_location or "us-central1",
|
||||
vertex_project=vertex_project or project_id,
|
||||
)
|
||||
|
|
@ -117,7 +117,7 @@ class VertexAIBatchPrediction(VertexLLM):
|
|||
)
|
||||
return vertex_batch_response
|
||||
|
||||
def create_vertex_url(
|
||||
def create_vertex_batch_url(
|
||||
self,
|
||||
vertex_location: str,
|
||||
vertex_project: str,
|
||||
|
|
@ -145,7 +145,7 @@ class VertexAIBatchPrediction(VertexLLM):
|
|||
custom_llm_provider="vertex_ai",
|
||||
)
|
||||
|
||||
default_api_base = self.create_vertex_url(
|
||||
default_api_base = self.create_vertex_batch_url(
|
||||
vertex_location=vertex_location or "us-central1",
|
||||
vertex_project=vertex_project or project_id,
|
||||
)
|
||||
|
|
|
|||
|
|
@ -2,6 +2,7 @@
|
|||
## httpx client for vertex ai calls
|
||||
## Initial implementation - covers gemini + image gen calls
|
||||
import json
|
||||
import time
|
||||
import uuid
|
||||
from copy import deepcopy
|
||||
from functools import partial
|
||||
|
|
@ -61,6 +62,7 @@ from litellm.types.llms.vertex_ai import (
|
|||
UsageMetadata,
|
||||
)
|
||||
from litellm.types.utils import (
|
||||
ChatCompletionAudioResponse,
|
||||
ChatCompletionTokenLogprob,
|
||||
ChoiceLogprobs,
|
||||
CompletionTokensDetailsWrapper,
|
||||
|
|
@ -69,7 +71,7 @@ from litellm.types.utils import (
|
|||
TopLogprob,
|
||||
Usage,
|
||||
)
|
||||
from litellm.utils import CustomStreamWrapper, ModelResponse, supports_reasoning
|
||||
from litellm.utils import CustomStreamWrapper, ModelResponse, is_base64_encoded, supports_reasoning
|
||||
|
||||
from ....utils import _remove_additional_properties, _remove_strict_from_schema
|
||||
from ..common_utils import VertexAIError, _build_vertex_schema
|
||||
|
|
@ -220,6 +222,7 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
"top_logprobs",
|
||||
"modalities",
|
||||
"parallel_tool_calls",
|
||||
"web_search_options",
|
||||
]
|
||||
if supports_reasoning(model):
|
||||
supported_params.append("reasoning_effort")
|
||||
|
|
@ -251,6 +254,14 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
status_code=400,
|
||||
)
|
||||
|
||||
def _map_web_search_options(self, value: dict) -> Tools:
|
||||
"""
|
||||
Base Case: empty dict
|
||||
|
||||
Google doesn't support user_location or search_context_size params
|
||||
"""
|
||||
return Tools(googleSearch={})
|
||||
|
||||
def _map_function(self, value: List[dict]) -> List[Tools]:
|
||||
gtool_func_declarations = []
|
||||
googleSearch: Optional[dict] = None
|
||||
|
|
@ -445,6 +456,19 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
response_modalities.append("MODALITY_UNSPECIFIED")
|
||||
return response_modalities
|
||||
|
||||
def validate_parallel_tool_calls(self, value: bool, non_default_params: dict):
|
||||
tools = non_default_params.get("tools", non_default_params.get("functions"))
|
||||
num_function_declarations = len(tools) if isinstance(tools, list) else 0
|
||||
if num_function_declarations > 1:
|
||||
raise litellm.utils.UnsupportedParamsError(
|
||||
message=(
|
||||
"`parallel_tool_calls=False` is not supported by Gemini when multiple tools are "
|
||||
"provided. Specify a single tool, or set "
|
||||
"`parallel_tool_calls=True`. If you want to drop this param, set `litellm.drop_params = True` or pass in `(.., drop_params=True)` in the requst - https://docs.litellm.ai/docs/completion/drop_params"
|
||||
),
|
||||
status_code=400,
|
||||
)
|
||||
|
||||
def map_openai_params(
|
||||
self,
|
||||
non_default_params: Dict,
|
||||
|
|
@ -487,7 +511,9 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
and isinstance(value, list)
|
||||
and value
|
||||
):
|
||||
optional_params["tools"] = self._map_function(value=value)
|
||||
optional_params = self._add_tools_to_optional_params(
|
||||
optional_params, self._map_function(value=value)
|
||||
)
|
||||
elif param == "tool_choice" and (
|
||||
isinstance(value, str) or isinstance(value, dict)
|
||||
):
|
||||
|
|
@ -500,21 +526,7 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
if value is False and not (
|
||||
drop_params or litellm.drop_params
|
||||
): # if drop params is True, then we should just ignore this
|
||||
tools = non_default_params.get(
|
||||
"tools", non_default_params.get("functions")
|
||||
)
|
||||
num_function_declarations = (
|
||||
len(tools) if isinstance(tools, list) else 0
|
||||
)
|
||||
if num_function_declarations > 1:
|
||||
raise litellm.utils.UnsupportedParamsError(
|
||||
message=(
|
||||
"`parallel_tool_calls=False` is not supported when multiple tools are "
|
||||
"provided for Gemini. Specify a single tool, or set "
|
||||
"`parallel_tool_calls=True`. If you want to drop this param, set `litellm.drop_params = True` or pass in `(.., drop_params=True)` in the requst - https://docs.litellm.ai/docs/completion/drop_params"
|
||||
),
|
||||
status_code=400,
|
||||
)
|
||||
self.validate_parallel_tool_calls(value, non_default_params)
|
||||
else:
|
||||
optional_params["parallel_tool_calls"] = value
|
||||
elif param == "seed":
|
||||
|
|
@ -532,7 +544,11 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
elif param == "modalities" and isinstance(value, list):
|
||||
response_modalities = self.map_response_modalities(value)
|
||||
optional_params["responseModalities"] = response_modalities
|
||||
|
||||
elif param == "web_search_options" and value and isinstance(value, dict):
|
||||
_tools = self._map_web_search_options(value)
|
||||
optional_params = self._add_tools_to_optional_params(
|
||||
optional_params, [_tools]
|
||||
)
|
||||
if litellm.vertex_ai_safety_settings is not None:
|
||||
optional_params["safety_settings"] = litellm.vertex_ai_safety_settings
|
||||
return optional_params
|
||||
|
|
@ -662,14 +678,30 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
) -> Tuple[Optional[str], Optional[str]]:
|
||||
content_str: Optional[str] = None
|
||||
reasoning_content_str: Optional[str] = None
|
||||
|
||||
for part in parts:
|
||||
_content_str = ""
|
||||
if "text" in part:
|
||||
_content_str += part["text"]
|
||||
elif "inlineData" in part: # base64 encoded image
|
||||
_content_str += "data:{};base64,{}".format(
|
||||
part["inlineData"]["mimeType"], part["inlineData"]["data"]
|
||||
)
|
||||
text_content = part["text"]
|
||||
# Check if text content is audio data URI - if so, exclude from text content
|
||||
if text_content.startswith("data:audio") and ";base64," in text_content:
|
||||
try:
|
||||
if is_base64_encoded(text_content):
|
||||
media_type, _ = text_content.split("data:")[1].split(";base64,")
|
||||
if media_type.startswith("audio/"):
|
||||
continue
|
||||
except (ValueError, IndexError):
|
||||
# If parsing fails, treat as regular text
|
||||
pass
|
||||
_content_str += text_content
|
||||
elif "inlineData" in part:
|
||||
mime_type = part["inlineData"]["mimeType"]
|
||||
data = part["inlineData"]["data"]
|
||||
# Check if inline data is audio - if so, exclude from text content
|
||||
if mime_type.startswith("audio/"):
|
||||
continue
|
||||
_content_str += "data:{};base64,{}".format(mime_type, data)
|
||||
|
||||
if len(_content_str) > 0:
|
||||
if part.get("thought") is True:
|
||||
if reasoning_content_str is None:
|
||||
|
|
@ -682,6 +714,47 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
|
||||
return content_str, reasoning_content_str
|
||||
|
||||
def _extract_audio_response_from_parts(
|
||||
self, parts: List[HttpxPartType]
|
||||
) -> Optional[ChatCompletionAudioResponse]:
|
||||
"""Extract audio response from parts if present"""
|
||||
for part in parts:
|
||||
if "text" in part:
|
||||
text_content = part["text"]
|
||||
# Check if text content contains audio data URI
|
||||
if text_content.startswith("data:audio") and ";base64," in text_content:
|
||||
try:
|
||||
if is_base64_encoded(text_content):
|
||||
media_type, audio_data = text_content.split("data:")[1].split(";base64,")
|
||||
|
||||
if media_type.startswith("audio/"):
|
||||
expires_at = int(time.time()) + (24 * 60 * 60)
|
||||
transcript = "" # Gemini doesn't provide transcript
|
||||
|
||||
return ChatCompletionAudioResponse(
|
||||
data=audio_data,
|
||||
expires_at=expires_at,
|
||||
transcript=transcript
|
||||
)
|
||||
except (ValueError, IndexError):
|
||||
pass
|
||||
|
||||
elif "inlineData" in part:
|
||||
mime_type = part["inlineData"]["mimeType"]
|
||||
data = part["inlineData"]["data"]
|
||||
|
||||
if mime_type.startswith("audio/"):
|
||||
expires_at = int(time.time()) + (24 * 60 * 60)
|
||||
transcript = "" # Gemini doesn't provide transcript
|
||||
|
||||
return ChatCompletionAudioResponse(
|
||||
data=data,
|
||||
expires_at=expires_at,
|
||||
transcript=transcript
|
||||
)
|
||||
|
||||
return None
|
||||
|
||||
def _transform_parts(
|
||||
self,
|
||||
parts: List[HttpxPartType],
|
||||
|
|
@ -967,8 +1040,17 @@ class VertexGeminiConfig(VertexAIBaseConfig, BaseConfig):
|
|||
) = VertexGeminiConfig().get_assistant_content_message(
|
||||
parts=candidate["content"]["parts"]
|
||||
)
|
||||
if content is not None:
|
||||
|
||||
audio_response = VertexGeminiConfig()._extract_audio_response_from_parts(
|
||||
parts=candidate["content"]["parts"]
|
||||
)
|
||||
|
||||
if audio_response is not None:
|
||||
cast(Dict[str, Any], chat_completion_message)["audio"] = audio_response
|
||||
chat_completion_message["content"] = None # OpenAI spec
|
||||
elif content is not None:
|
||||
chat_completion_message["content"] = content
|
||||
|
||||
if reasoning_content is not None:
|
||||
chat_completion_message["reasoning_content"] = reasoning_content
|
||||
|
||||
|
|
@ -1185,7 +1267,9 @@ async def make_call(
|
|||
)
|
||||
|
||||
completion_stream = ModelResponseIterator(
|
||||
streaming_response=response.aiter_lines(), sync_stream=False
|
||||
streaming_response=response.aiter_lines(),
|
||||
sync_stream=False,
|
||||
logging_obj=logging_obj,
|
||||
)
|
||||
# LOGGING
|
||||
logging_obj.post_call(
|
||||
|
|
@ -1223,7 +1307,9 @@ def make_sync_call(
|
|||
)
|
||||
|
||||
completion_stream = ModelResponseIterator(
|
||||
streaming_response=response.iter_lines(), sync_stream=True
|
||||
streaming_response=response.iter_lines(),
|
||||
sync_stream=True,
|
||||
logging_obj=logging_obj,
|
||||
)
|
||||
|
||||
# LOGGING
|
||||
|
|
@ -1644,11 +1730,19 @@ class VertexLLM(VertexBase):
|
|||
|
||||
|
||||
class ModelResponseIterator:
|
||||
def __init__(self, streaming_response, sync_stream: bool):
|
||||
def __init__(
|
||||
self, streaming_response, sync_stream: bool, logging_obj: LoggingClass
|
||||
):
|
||||
from litellm.litellm_core_utils.prompt_templates.common_utils import (
|
||||
check_is_function_call,
|
||||
)
|
||||
|
||||
self.streaming_response = streaming_response
|
||||
self.chunk_type: Literal["valid_json", "accumulated_json"] = "valid_json"
|
||||
self.accumulated_json = ""
|
||||
self.sent_first_chunk = False
|
||||
self.logging_obj = logging_obj
|
||||
self.is_function_call = check_is_function_call(logging_obj)
|
||||
|
||||
def chunk_parser(self, chunk: dict) -> GenericStreamingChunk:
|
||||
try:
|
||||
|
|
@ -1712,11 +1806,23 @@ class ModelResponseIterator:
|
|||
},
|
||||
)
|
||||
|
||||
returned_chunk = GenericStreamingChunk(
|
||||
text=text,
|
||||
tool_use=tool_use,
|
||||
is_finished=False,
|
||||
finish_reason=finish_reason,
|
||||
args: Dict[str, Any] = {
|
||||
"content": text or None,
|
||||
"reasoning_content": reasoning_content,
|
||||
}
|
||||
if self.is_function_call and tool_use is not None:
|
||||
args["function_call"] = tool_use["function"]
|
||||
elif tool_use is not None:
|
||||
args["tool_calls"] = [tool_use]
|
||||
|
||||
returned_chunk = ModelResponseStream(
|
||||
choices=[
|
||||
StreamingChoices(
|
||||
index=0,
|
||||
delta=Delta(**args),
|
||||
finish_reason=finish_reason,
|
||||
)
|
||||
],
|
||||
usage=usage,
|
||||
index=0,
|
||||
)
|
||||
|
|
@ -1729,7 +1835,7 @@ class ModelResponseIterator:
|
|||
self.response_iterator = self.streaming_response
|
||||
return self
|
||||
|
||||
def handle_valid_json_chunk(self, chunk: str) -> GenericStreamingChunk:
|
||||
def handle_valid_json_chunk(self, chunk: str) -> Optional[ModelResponseStream]:
|
||||
chunk = chunk.strip()
|
||||
try:
|
||||
json_chunk = json.loads(chunk)
|
||||
|
|
@ -1747,7 +1853,9 @@ class ModelResponseIterator:
|
|||
|
||||
return self.chunk_parser(chunk=json_chunk)
|
||||
|
||||
def handle_accumulated_json_chunk(self, chunk: str) -> GenericStreamingChunk:
|
||||
def handle_accumulated_json_chunk(
|
||||
self, chunk: str
|
||||
) -> Optional[ModelResponseStream]:
|
||||
chunk = litellm.CustomStreamWrapper._strip_sse_data_from_chunk(chunk) or ""
|
||||
message = chunk.replace("\n\n", "")
|
||||
|
||||
|
|
@ -1761,16 +1869,9 @@ class ModelResponseIterator:
|
|||
return self.chunk_parser(chunk=_data)
|
||||
except json.JSONDecodeError:
|
||||
# If it's not valid JSON yet, continue to the next event
|
||||
return GenericStreamingChunk(
|
||||
text="",
|
||||
is_finished=False,
|
||||
finish_reason="",
|
||||
usage=None,
|
||||
index=0,
|
||||
tool_use=None,
|
||||
)
|
||||
return None
|
||||
|
||||
def _common_chunk_parsing_logic(self, chunk: str) -> GenericStreamingChunk:
|
||||
def _common_chunk_parsing_logic(self, chunk: str) -> Optional[ModelResponseStream]:
|
||||
try:
|
||||
chunk = litellm.CustomStreamWrapper._strip_sse_data_from_chunk(chunk) or ""
|
||||
if len(chunk) > 0:
|
||||
|
|
@ -1783,15 +1884,7 @@ class ModelResponseIterator:
|
|||
return self.handle_valid_json_chunk(chunk=chunk)
|
||||
elif self.chunk_type == "accumulated_json":
|
||||
return self.handle_accumulated_json_chunk(chunk=chunk)
|
||||
|
||||
return GenericStreamingChunk(
|
||||
text="",
|
||||
is_finished=False,
|
||||
finish_reason="",
|
||||
usage=None,
|
||||
index=0,
|
||||
tool_use=None,
|
||||
)
|
||||
return None
|
||||
except Exception:
|
||||
raise
|
||||
|
||||
|
|
|
|||
|
|
@ -0,0 +1,96 @@
|
|||
from typing import Any, Dict, List, Optional, Tuple
|
||||
|
||||
import litellm
|
||||
from litellm.llms.anthropic.experimental_pass_through.messages.transformation import (
|
||||
AnthropicMessagesConfig,
|
||||
)
|
||||
from litellm.secret_managers.main import get_secret_str
|
||||
from litellm.types.llms.vertex_ai import VertexPartnerProvider
|
||||
from litellm.types.router import GenericLiteLLMParams
|
||||
|
||||
from ....vertex_llm_base import VertexBase
|
||||
|
||||
|
||||
class VertexAIPartnerModelsAnthropicMessagesConfig(AnthropicMessagesConfig, VertexBase):
|
||||
def validate_anthropic_messages_environment(
|
||||
self,
|
||||
headers: dict,
|
||||
model: str,
|
||||
messages: List[Any],
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
api_key: Optional[str] = None,
|
||||
api_base: Optional[str] = None,
|
||||
) -> Tuple[dict, Optional[str]]:
|
||||
"""
|
||||
OPTIONAL
|
||||
|
||||
Validate the environment for the request
|
||||
"""
|
||||
if "Authorization" not in headers:
|
||||
vertex_ai_project = (
|
||||
optional_params.pop("vertex_project", None)
|
||||
or optional_params.pop("vertex_ai_project", None)
|
||||
or litellm.vertex_project
|
||||
or get_secret_str("VERTEXAI_PROJECT")
|
||||
)
|
||||
vertex_credentials = (
|
||||
optional_params.pop("vertex_credentials", None)
|
||||
or optional_params.pop("vertex_ai_credentials", None)
|
||||
or get_secret_str("VERTEXAI_CREDENTIALS")
|
||||
)
|
||||
|
||||
access_token, project_id = self._ensure_access_token(
|
||||
credentials=vertex_credentials,
|
||||
project_id=vertex_ai_project,
|
||||
custom_llm_provider="vertex_ai",
|
||||
)
|
||||
|
||||
headers["Authorization"] = f"Bearer {access_token}"
|
||||
|
||||
api_base = self.get_complete_vertex_url(
|
||||
custom_api_base=api_base,
|
||||
vertex_location=optional_params.pop("vertex_location", None),
|
||||
vertex_project=vertex_ai_project,
|
||||
project_id=project_id,
|
||||
partner=VertexPartnerProvider.claude,
|
||||
stream=optional_params.get("stream", False),
|
||||
model=model,
|
||||
)
|
||||
|
||||
headers["content-type"] = "application/json"
|
||||
return headers, api_base
|
||||
|
||||
def get_complete_url(
|
||||
self,
|
||||
api_base: Optional[str],
|
||||
api_key: Optional[str],
|
||||
model: str,
|
||||
optional_params: dict,
|
||||
litellm_params: dict,
|
||||
stream: Optional[bool] = None,
|
||||
) -> str:
|
||||
if api_base is None:
|
||||
raise ValueError(
|
||||
"api_base is required. Unable to determine the correct api_base for the request."
|
||||
)
|
||||
return api_base # no transformation is needed - handled in validate_environment
|
||||
|
||||
def transform_anthropic_messages_request(
|
||||
self,
|
||||
model: str,
|
||||
messages: List[Dict],
|
||||
anthropic_messages_optional_request_params: Dict,
|
||||
litellm_params: GenericLiteLLMParams,
|
||||
headers: dict,
|
||||
) -> Dict:
|
||||
anthropic_messages_request = super().transform_anthropic_messages_request(
|
||||
model=model,
|
||||
messages=messages,
|
||||
anthropic_messages_optional_request_params=anthropic_messages_optional_request_params,
|
||||
litellm_params=litellm_params,
|
||||
headers=headers,
|
||||
)
|
||||
|
||||
anthropic_messages_request["anthropic_version"] = "vertex-2023-10-16"
|
||||
return anthropic_messages_request
|
||||
|
|
@ -1,12 +1,12 @@
|
|||
# What is this?
|
||||
## API Handler for calling Vertex AI Partner Models
|
||||
from enum import Enum
|
||||
from typing import Callable, Optional, Union
|
||||
|
||||
import httpx # type: ignore
|
||||
|
||||
import litellm
|
||||
from litellm import LlmProviders
|
||||
from litellm.types.llms.vertex_ai import VertexPartnerProvider
|
||||
from litellm.utils import ModelResponse
|
||||
|
||||
from ...custom_httpx.llm_http_handler import BaseLLMHTTPHandler
|
||||
|
|
@ -15,13 +15,6 @@ from ..vertex_llm_base import VertexBase
|
|||
base_llm_http_handler = BaseLLMHTTPHandler()
|
||||
|
||||
|
||||
class VertexPartnerProvider(str, Enum):
|
||||
mistralai = "mistralai"
|
||||
llama = "llama"
|
||||
ai21 = "ai21"
|
||||
claude = "claude"
|
||||
|
||||
|
||||
class VertexAIError(Exception):
|
||||
def __init__(self, status_code, message):
|
||||
self.status_code = status_code
|
||||
|
|
@ -35,78 +28,10 @@ class VertexAIError(Exception):
|
|||
) # Call the base class constructor with the parameters it needs
|
||||
|
||||
|
||||
def create_vertex_url(
|
||||
vertex_location: str,
|
||||
vertex_project: str,
|
||||
partner: VertexPartnerProvider,
|
||||
stream: Optional[bool],
|
||||
model: str,
|
||||
api_base: Optional[str] = None,
|
||||
) -> str:
|
||||
"""Return the base url for the vertex partner models"""
|
||||
|
||||
api_base = api_base or f"https://{vertex_location}-aiplatform.googleapis.com"
|
||||
if partner == VertexPartnerProvider.llama:
|
||||
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/endpoints/openapi/chat/completions"
|
||||
elif partner == VertexPartnerProvider.mistralai:
|
||||
if stream:
|
||||
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/mistralai/models/{model}:streamRawPredict"
|
||||
else:
|
||||
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/mistralai/models/{model}:rawPredict"
|
||||
elif partner == VertexPartnerProvider.ai21:
|
||||
if stream:
|
||||
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/publishers/ai21/models/{model}:streamRawPredict"
|
||||
else:
|
||||
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/publishers/ai21/models/{model}:rawPredict"
|
||||
elif partner == VertexPartnerProvider.claude:
|
||||
if stream:
|
||||
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/anthropic/models/{model}:streamRawPredict"
|
||||
else:
|
||||
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/anthropic/models/{model}:rawPredict"
|
||||
|
||||
|
||||
class VertexAIPartnerModels(VertexBase):
|
||||
def __init__(self) -> None:
|
||||
pass
|
||||
|
||||
def get_complete_url(
|
||||
self,
|
||||
custom_api_base: Optional[str],
|
||||
vertex_location: Optional[str],
|
||||
vertex_project: Optional[str],
|
||||
project_id: str,
|
||||
partner: VertexPartnerProvider,
|
||||
stream: Optional[bool],
|
||||
model: str,
|
||||
) -> str:
|
||||
api_base = self.get_api_base(
|
||||
api_base=custom_api_base, vertex_location=vertex_location
|
||||
)
|
||||
default_api_base = create_vertex_url(
|
||||
vertex_location=vertex_location or "us-central1",
|
||||
vertex_project=vertex_project or project_id,
|
||||
partner=partner, # type: ignore
|
||||
stream=stream,
|
||||
model=model,
|
||||
api_base=api_base,
|
||||
)
|
||||
|
||||
if len(default_api_base.split(":")) > 1:
|
||||
endpoint = default_api_base.split(":")[-1]
|
||||
else:
|
||||
endpoint = ""
|
||||
|
||||
_, api_base = self._check_custom_proxy(
|
||||
api_base=custom_api_base,
|
||||
custom_llm_provider="vertex_ai",
|
||||
gemini_api_key=None,
|
||||
endpoint=endpoint,
|
||||
stream=stream,
|
||||
auth_header=None,
|
||||
url=default_api_base,
|
||||
)
|
||||
return api_base
|
||||
|
||||
def completion(
|
||||
self,
|
||||
model: str,
|
||||
|
|
@ -181,7 +106,7 @@ class VertexAIPartnerModels(VertexBase):
|
|||
else:
|
||||
raise ValueError(f"Unknown partner model: {model}")
|
||||
|
||||
api_base = self.get_complete_url(
|
||||
api_base = self.get_complete_vertex_url(
|
||||
custom_api_base=api_base,
|
||||
vertex_location=vertex_location,
|
||||
vertex_project=vertex_project,
|
||||
|
|
|
|||
|
|
@ -11,7 +11,7 @@ from typing import TYPE_CHECKING, Any, Dict, Literal, Optional, Tuple
|
|||
from litellm._logging import verbose_logger
|
||||
from litellm.litellm_core_utils.asyncify import asyncify
|
||||
from litellm.llms.custom_httpx.http_handler import AsyncHTTPHandler
|
||||
from litellm.types.llms.vertex_ai import VERTEX_CREDENTIALS_TYPES
|
||||
from litellm.types.llms.vertex_ai import VERTEX_CREDENTIALS_TYPES, VertexPartnerProvider
|
||||
|
||||
from .common_utils import _get_gemini_url, _get_vertex_url, all_gemini_url_modes
|
||||
|
||||
|
|
@ -150,6 +150,74 @@ class VertexBase:
|
|||
else:
|
||||
return f"https://{self.get_default_vertex_location()}-aiplatform.googleapis.com"
|
||||
|
||||
@staticmethod
|
||||
def create_vertex_url(
|
||||
vertex_location: str,
|
||||
vertex_project: str,
|
||||
partner: VertexPartnerProvider,
|
||||
stream: Optional[bool],
|
||||
model: str,
|
||||
api_base: Optional[str] = None,
|
||||
) -> str:
|
||||
"""Return the base url for the vertex partner models"""
|
||||
|
||||
api_base = api_base or f"https://{vertex_location}-aiplatform.googleapis.com"
|
||||
if partner == VertexPartnerProvider.llama:
|
||||
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/endpoints/openapi/chat/completions"
|
||||
elif partner == VertexPartnerProvider.mistralai:
|
||||
if stream:
|
||||
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/mistralai/models/{model}:streamRawPredict"
|
||||
else:
|
||||
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/mistralai/models/{model}:rawPredict"
|
||||
elif partner == VertexPartnerProvider.ai21:
|
||||
if stream:
|
||||
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/publishers/ai21/models/{model}:streamRawPredict"
|
||||
else:
|
||||
return f"{api_base}/v1beta1/projects/{vertex_project}/locations/{vertex_location}/publishers/ai21/models/{model}:rawPredict"
|
||||
elif partner == VertexPartnerProvider.claude:
|
||||
if stream:
|
||||
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/anthropic/models/{model}:streamRawPredict"
|
||||
else:
|
||||
return f"{api_base}/v1/projects/{vertex_project}/locations/{vertex_location}/publishers/anthropic/models/{model}:rawPredict"
|
||||
|
||||
def get_complete_vertex_url(
|
||||
self,
|
||||
custom_api_base: Optional[str],
|
||||
vertex_location: Optional[str],
|
||||
vertex_project: Optional[str],
|
||||
project_id: str,
|
||||
partner: VertexPartnerProvider,
|
||||
stream: Optional[bool],
|
||||
model: str,
|
||||
) -> str:
|
||||
api_base = self.get_api_base(
|
||||
api_base=custom_api_base, vertex_location=vertex_location
|
||||
)
|
||||
default_api_base = VertexBase.create_vertex_url(
|
||||
vertex_location=vertex_location or "us-central1",
|
||||
vertex_project=vertex_project or project_id,
|
||||
partner=partner,
|
||||
stream=stream,
|
||||
model=model,
|
||||
api_base=api_base,
|
||||
)
|
||||
|
||||
if len(default_api_base.split(":")) > 1:
|
||||
endpoint = default_api_base.split(":")[-1]
|
||||
else:
|
||||
endpoint = ""
|
||||
|
||||
_, api_base = self._check_custom_proxy(
|
||||
api_base=custom_api_base,
|
||||
custom_llm_provider="vertex_ai",
|
||||
gemini_api_key=None,
|
||||
endpoint=endpoint,
|
||||
stream=stream,
|
||||
auth_header=None,
|
||||
url=default_api_base,
|
||||
)
|
||||
return api_base
|
||||
|
||||
def refresh_auth(self, credentials: Any) -> None:
|
||||
from google.auth.transport.requests import (
|
||||
Request, # type: ignore[import-untyped]
|
||||
|
|
|
|||
|
|
@ -3,6 +3,7 @@ from typing import List, Optional, Tuple
|
|||
import litellm
|
||||
from litellm._logging import verbose_logger
|
||||
from litellm.litellm_core_utils.prompt_templates.common_utils import (
|
||||
filter_value_from_dict,
|
||||
strip_name_from_messages,
|
||||
)
|
||||
from litellm.secret_managers.main import get_secret_str
|
||||
|
|
@ -44,6 +45,7 @@ class XAIChatConfig(OpenAIGPTConfig):
|
|||
"top_logprobs",
|
||||
"top_p",
|
||||
"user",
|
||||
"web_search_options",
|
||||
]
|
||||
try:
|
||||
if litellm.supports_reasoning(
|
||||
|
|
@ -66,6 +68,14 @@ class XAIChatConfig(OpenAIGPTConfig):
|
|||
for param, value in non_default_params.items():
|
||||
if param == "max_completion_tokens":
|
||||
optional_params["max_tokens"] = value
|
||||
elif param == "tools" and value is not None:
|
||||
tools = []
|
||||
for tool in value:
|
||||
tool = filter_value_from_dict(tool, "strict")
|
||||
if tool is not None:
|
||||
tools.append(tool)
|
||||
if len(tools) > 0:
|
||||
optional_params["tools"] = tools
|
||||
elif param in supported_openai_params:
|
||||
if value is not None:
|
||||
optional_params[param] = value
|
||||
|
|
|
|||
|
|
@ -6,9 +6,21 @@ import litellm
|
|||
from litellm.llms.base_llm.base_utils import BaseLLMModelInfo
|
||||
from litellm.secret_managers.main import get_secret_str
|
||||
from litellm.types.llms.openai import AllMessageValues
|
||||
from litellm.types.utils import ProviderSpecificModelInfo
|
||||
|
||||
|
||||
class XAIModelInfo(BaseLLMModelInfo):
|
||||
def get_provider_info(
|
||||
self,
|
||||
model: str,
|
||||
) -> Optional[ProviderSpecificModelInfo]:
|
||||
"""
|
||||
Default values all models of this provider support.
|
||||
"""
|
||||
return {
|
||||
"supports_web_search": True,
|
||||
}
|
||||
|
||||
def validate_environment(
|
||||
self,
|
||||
headers: dict,
|
||||
|
|
|
|||
|
|
@ -59,6 +59,7 @@ from litellm.constants import (
|
|||
from litellm.exceptions import LiteLLMUnknownProvider
|
||||
from litellm.integrations.custom_logger import CustomLogger
|
||||
from litellm.litellm_core_utils.audio_utils.utils import get_audio_file_for_health_check
|
||||
from litellm.litellm_core_utils.dd_tracing import tracer
|
||||
from litellm.litellm_core_utils.health_check_utils import (
|
||||
_create_health_check_response,
|
||||
_filter_model_params,
|
||||
|
|
@ -314,6 +315,7 @@ class AsyncCompletions:
|
|||
return response
|
||||
|
||||
|
||||
@tracer.wrap()
|
||||
@client
|
||||
async def acompletion(
|
||||
model: str,
|
||||
|
|
@ -810,6 +812,7 @@ def mock_completion(
|
|||
raise Exception("Mock completion response failed - {}".format(e))
|
||||
|
||||
|
||||
@tracer.wrap()
|
||||
@client
|
||||
def completion( # type: ignore # noqa: PLR0915
|
||||
model: str,
|
||||
|
|
@ -2335,6 +2338,26 @@ def completion( # type: ignore # noqa: PLR0915
|
|||
original_response=response,
|
||||
additional_args={"headers": headers},
|
||||
)
|
||||
|
||||
elif custom_llm_provider == "datarobot":
|
||||
response = base_llm_http_handler.completion(
|
||||
model=model,
|
||||
messages=messages,
|
||||
headers=headers,
|
||||
model_response=model_response,
|
||||
api_key=api_key,
|
||||
api_base=api_base,
|
||||
acompletion=acompletion,
|
||||
logging_obj=logging,
|
||||
optional_params=optional_params,
|
||||
litellm_params=litellm_params,
|
||||
timeout=timeout, # type: ignore
|
||||
client=client,
|
||||
custom_llm_provider=custom_llm_provider,
|
||||
encoding=encoding,
|
||||
stream=stream,
|
||||
provider_config=provider_config,
|
||||
)
|
||||
elif custom_llm_provider == "openrouter":
|
||||
api_base = (
|
||||
api_base
|
||||
|
|
|
|||
|
|
@ -96,13 +96,7 @@
|
|||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_native_streaming": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.03,
|
||||
"search_context_size_medium": 0.035,
|
||||
"search_context_size_high": 0.05
|
||||
}
|
||||
"supports_native_streaming": true
|
||||
},
|
||||
"gpt-4.1-2025-04-14": {
|
||||
"max_tokens": 32768,
|
||||
|
|
@ -135,13 +129,7 @@
|
|||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_native_streaming": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.03,
|
||||
"search_context_size_medium": 0.035,
|
||||
"search_context_size_high": 0.05
|
||||
}
|
||||
"supports_native_streaming": true
|
||||
},
|
||||
"gpt-4.1-mini": {
|
||||
"max_tokens": 32768,
|
||||
|
|
@ -174,13 +162,7 @@
|
|||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_native_streaming": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.025,
|
||||
"search_context_size_medium": 0.0275,
|
||||
"search_context_size_high": 0.03
|
||||
}
|
||||
"supports_native_streaming": true
|
||||
},
|
||||
"gpt-4.1-mini-2025-04-14": {
|
||||
"max_tokens": 32768,
|
||||
|
|
@ -213,13 +195,7 @@
|
|||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_native_streaming": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.025,
|
||||
"search_context_size_medium": 0.0275,
|
||||
"search_context_size_high": 0.03
|
||||
}
|
||||
"supports_native_streaming": true
|
||||
},
|
||||
"gpt-4.1-nano": {
|
||||
"max_tokens": 32768,
|
||||
|
|
@ -305,13 +281,7 @@
|
|||
"supports_vision": true,
|
||||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.03,
|
||||
"search_context_size_medium": 0.035,
|
||||
"search_context_size_high": 0.05
|
||||
}
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"watsonx/ibm/granite-3-8b-instruct": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -349,13 +319,7 @@
|
|||
"supports_vision": true,
|
||||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.03,
|
||||
"search_context_size_medium": 0.035,
|
||||
"search_context_size_high": 0.05
|
||||
}
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"gpt-4o-search-preview": {
|
||||
"max_tokens": 16384,
|
||||
|
|
@ -527,13 +491,7 @@
|
|||
"supports_vision": true,
|
||||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.025,
|
||||
"search_context_size_medium": 0.0275,
|
||||
"search_context_size_high": 0.03
|
||||
}
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"gpt-4o-mini-search-preview-2025-03-11": {
|
||||
"max_tokens": 16384,
|
||||
|
|
@ -553,13 +511,7 @@
|
|||
"supports_vision": true,
|
||||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.025,
|
||||
"search_context_size_medium": 0.0275,
|
||||
"search_context_size_high": 0.03
|
||||
}
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"gpt-4o-mini-search-preview": {
|
||||
"max_tokens": 16384,
|
||||
|
|
@ -954,13 +906,7 @@
|
|||
"supports_vision": true,
|
||||
"supports_prompt_caching": true,
|
||||
"supports_system_messages": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true,
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.03,
|
||||
"search_context_size_medium": 0.035,
|
||||
"search_context_size_high": 0.05
|
||||
}
|
||||
"supports_tool_choice": true
|
||||
},
|
||||
"gpt-4o-2024-11-20": {
|
||||
"max_tokens": 16384,
|
||||
|
|
@ -4217,7 +4163,8 @@
|
|||
"mode": "chat",
|
||||
"supports_function_calling": true,
|
||||
"supports_vision": true,
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-2-vision-1212": {
|
||||
"max_tokens": 32768,
|
||||
|
|
@ -4230,7 +4177,8 @@
|
|||
"mode": "chat",
|
||||
"supports_function_calling": true,
|
||||
"supports_vision": true,
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-2-vision-latest": {
|
||||
"max_tokens": 32768,
|
||||
|
|
@ -4243,7 +4191,8 @@
|
|||
"mode": "chat",
|
||||
"supports_function_calling": true,
|
||||
"supports_vision": true,
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-2-vision": {
|
||||
"max_tokens": 32768,
|
||||
|
|
@ -4256,7 +4205,8 @@
|
|||
"mode": "chat",
|
||||
"supports_function_calling": true,
|
||||
"supports_vision": true,
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-3": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4269,7 +4219,8 @@
|
|||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_response_schema": false,
|
||||
"source": "https://x.ai/api#pricing"
|
||||
"source": "https://x.ai/api#pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-3-beta": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4282,7 +4233,8 @@
|
|||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_response_schema": false,
|
||||
"source": "https://x.ai/api#pricing"
|
||||
"source": "https://x.ai/api#pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-3-fast-beta": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4295,7 +4247,8 @@
|
|||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_response_schema": false,
|
||||
"source": "https://x.ai/api#pricing"
|
||||
"source": "https://x.ai/api#pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-3-fast-latest": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4308,7 +4261,8 @@
|
|||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_response_schema": false,
|
||||
"source": "https://x.ai/api#pricing"
|
||||
"source": "https://x.ai/api#pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-3-mini-beta": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4322,7 +4276,8 @@
|
|||
"supports_tool_choice": true,
|
||||
"supports_reasoning": true,
|
||||
"supports_response_schema": false,
|
||||
"source": "https://x.ai/api#pricing"
|
||||
"source": "https://x.ai/api#pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-3-mini-fast-beta": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4336,7 +4291,8 @@
|
|||
"supports_tool_choice": true,
|
||||
"supports_reasoning": true,
|
||||
"supports_response_schema": false,
|
||||
"source": "https://x.ai/api#pricing"
|
||||
"source": "https://x.ai/api#pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-3-mini-fast-latest": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4350,7 +4306,8 @@
|
|||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true,
|
||||
"supports_response_schema": false,
|
||||
"source": "https://x.ai/api#pricing"
|
||||
"source": "https://x.ai/api#pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-vision-beta": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -4363,7 +4320,8 @@
|
|||
"mode": "chat",
|
||||
"supports_function_calling": true,
|
||||
"supports_vision": true,
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-2-1212": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4374,7 +4332,8 @@
|
|||
"litellm_provider": "xai",
|
||||
"mode": "chat",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-2": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4385,7 +4344,8 @@
|
|||
"litellm_provider": "xai",
|
||||
"mode": "chat",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"xai/grok-2-latest": {
|
||||
"max_tokens": 131072,
|
||||
|
|
@ -4396,7 +4356,8 @@
|
|||
"litellm_provider": "xai",
|
||||
"mode": "chat",
|
||||
"supports_function_calling": true,
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"deepseek/deepseek-coder": {
|
||||
"max_tokens": 4096,
|
||||
|
|
@ -6112,7 +6073,8 @@
|
|||
"text"
|
||||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-pro-exp-02-05": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6152,7 +6114,8 @@
|
|||
"text"
|
||||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-flash-exp": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6197,7 +6160,8 @@
|
|||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
|
||||
"supports_tool_choice": true,
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-flash-001": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6232,7 +6196,8 @@
|
|||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
|
||||
"deprecation_date": "2026-02-05",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-flash-thinking-exp": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6277,7 +6242,8 @@
|
|||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
|
||||
"supports_tool_choice": true,
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-flash-thinking-exp-01-21": {
|
||||
"max_tokens": 65536,
|
||||
|
|
@ -6322,7 +6288,8 @@
|
|||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
|
||||
"supports_tool_choice": true,
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.5-pro-exp-03-25": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6363,7 +6330,8 @@
|
|||
"supported_output_modalities": [
|
||||
"text"
|
||||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing"
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.5-flash-preview-tts": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6400,7 +6368,8 @@
|
|||
"supported_output_modalities": [
|
||||
"audio"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview"
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.5-flash-preview-05-20": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6440,7 +6409,8 @@
|
|||
"supported_output_modalities": [
|
||||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview"
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.5-flash-preview-04-17": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6480,7 +6450,8 @@
|
|||
"supported_output_modalities": [
|
||||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview"
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.5-flash-preview-05-20": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6520,7 +6491,8 @@
|
|||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.5-flash-preview-04-17": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6560,7 +6532,8 @@
|
|||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-flash": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6595,7 +6568,8 @@
|
|||
],
|
||||
"supports_tool_choice": true,
|
||||
"source": "https://ai.google.dev/pricing#2_0flash",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-flash-lite": {
|
||||
"max_input_tokens": 1048576,
|
||||
|
|
@ -6627,7 +6601,8 @@
|
|||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
|
||||
"supports_tool_choice": true,
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-flash-lite-001": {
|
||||
"max_input_tokens": 1048576,
|
||||
|
|
@ -6660,7 +6635,8 @@
|
|||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
|
||||
"supports_tool_choice": true,
|
||||
"deprecation_date": "2026-02-25",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.5-pro-preview-05-06": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6701,7 +6677,8 @@
|
|||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.5-pro-preview-03-25": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6742,7 +6719,8 @@
|
|||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/models#gemini-2.5-flash-preview",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.0-flash-preview-image-generation": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6777,7 +6755,8 @@
|
|||
],
|
||||
"supports_tool_choice": true,
|
||||
"source": "https://ai.google.dev/pricing#2_0flash",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini-2.5-pro-preview-tts": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -6809,7 +6788,8 @@
|
|||
"audio"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview",
|
||||
"supports_parallel_function_calling": true
|
||||
"supports_parallel_function_calling": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-pro-exp-02-05": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6847,7 +6827,8 @@
|
|||
"supports_pdf_input": true,
|
||||
"supports_response_schema": true,
|
||||
"supports_tool_choice": true,
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing"
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/pricing",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-flash-preview-image-generation": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6883,7 +6864,8 @@
|
|||
"image"
|
||||
],
|
||||
"supports_tool_choice": true,
|
||||
"source": "https://ai.google.dev/pricing#2_0flash"
|
||||
"source": "https://ai.google.dev/pricing#2_0flash",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-flash": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6919,7 +6901,8 @@
|
|||
"image"
|
||||
],
|
||||
"supports_tool_choice": true,
|
||||
"source": "https://ai.google.dev/pricing#2_0flash"
|
||||
"source": "https://ai.google.dev/pricing#2_0flash",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-flash-lite": {
|
||||
"max_input_tokens": 1048576,
|
||||
|
|
@ -6952,7 +6935,8 @@
|
|||
"supported_output_modalities": [
|
||||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.0-flash-lite"
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.0-flash-lite",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-flash-001": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -6987,7 +6971,8 @@
|
|||
"text",
|
||||
"image"
|
||||
],
|
||||
"source": "https://ai.google.dev/pricing#2_0flash"
|
||||
"source": "https://ai.google.dev/pricing#2_0flash",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.5-pro-preview-tts": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -7020,7 +7005,8 @@
|
|||
"supported_output_modalities": [
|
||||
"audio"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview"
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.5-pro-preview-05-06": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -7056,7 +7042,8 @@
|
|||
"supported_output_modalities": [
|
||||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview"
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.5-pro-preview-03-25": {
|
||||
"max_tokens": 65535,
|
||||
|
|
@ -7092,7 +7079,8 @@
|
|||
"supported_output_modalities": [
|
||||
"text"
|
||||
],
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview"
|
||||
"source": "https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro-preview",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-flash-exp": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -7138,7 +7126,8 @@
|
|||
"image"
|
||||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-flash-lite-preview-02-05": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -7172,7 +7161,8 @@
|
|||
"supported_output_modalities": [
|
||||
"text"
|
||||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash-lite"
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash-lite",
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-flash-thinking-exp": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -7218,7 +7208,8 @@
|
|||
"image"
|
||||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemini-2.0-flash-thinking-exp-01-21": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -7264,7 +7255,8 @@
|
|||
"image"
|
||||
],
|
||||
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models#gemini-2.0-flash",
|
||||
"supports_tool_choice": true
|
||||
"supports_tool_choice": true,
|
||||
"supports_web_search": true
|
||||
},
|
||||
"gemini/gemma-3-27b-it": {
|
||||
"max_tokens": 8192,
|
||||
|
|
@ -8758,6 +8750,20 @@
|
|||
"notes": "'supports_image_input' is a deprecated field. Use 'supports_embedding_image_input' instead."
|
||||
}
|
||||
},
|
||||
"embed-v4.0": {
|
||||
"max_tokens": 1024,
|
||||
"max_input_tokens": 1024,
|
||||
"input_cost_per_token": 1.2e-07,
|
||||
"input_cost_per_image": 4.7e-07,
|
||||
"output_cost_per_token": 0.0,
|
||||
"litellm_provider": "cohere",
|
||||
"mode": "embedding",
|
||||
"supports_image_input": true,
|
||||
"supports_embedding_image_input": true,
|
||||
"metadata": {
|
||||
"notes": "'supports_image_input' is a deprecated field. Use 'supports_embedding_image_input' instead."
|
||||
}
|
||||
},
|
||||
"replicate/meta/llama-2-13b": {
|
||||
"max_tokens": 4096,
|
||||
"max_input_tokens": 4096,
|
||||
|
|
|
|||
118
litellm/passthrough/README.md
Normal file
118
litellm/passthrough/README.md
Normal file
|
|
@ -0,0 +1,118 @@
|
|||
This makes it easier to pass through requests to the LLM APIs.
|
||||
|
||||
E.g. Route to VLLM's `/classify` endpoint:
|
||||
|
||||
|
||||
## SDK (Basic)
|
||||
|
||||
```python
|
||||
import litellm
|
||||
|
||||
|
||||
response = litellm.llm_passthrough_route(
|
||||
model="hosted_vllm/papluca/xlm-roberta-base-language-detection",
|
||||
method="POST",
|
||||
endpoint="classify",
|
||||
api_base="http://localhost:8090",
|
||||
api_key=None,
|
||||
json={
|
||||
"model": "swapped-for-litellm-model",
|
||||
"input": "Hello, world!",
|
||||
}
|
||||
)
|
||||
|
||||
print(response)
|
||||
```
|
||||
|
||||
## SDK (Router)
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from litellm import Router
|
||||
|
||||
router = Router(
|
||||
model_list=[
|
||||
{
|
||||
"model_name": "roberta-base-language-detection",
|
||||
"litellm_params": {
|
||||
"model": "hosted_vllm/papluca/xlm-roberta-base-language-detection",
|
||||
"api_base": "http://localhost:8090",
|
||||
}
|
||||
}
|
||||
]
|
||||
)
|
||||
|
||||
request_data = {
|
||||
"model": "roberta-base-language-detection",
|
||||
"method": "POST",
|
||||
"endpoint": "classify",
|
||||
"api_base": "http://localhost:8090",
|
||||
"api_key": None,
|
||||
"json": {
|
||||
"model": "roberta-base-language-detection",
|
||||
"input": "Hello, world!",
|
||||
}
|
||||
}
|
||||
|
||||
async def main():
|
||||
response = await router.allm_passthrough_route(**request_data)
|
||||
print(response)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## PROXY
|
||||
|
||||
1. Setup config.yaml
|
||||
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: roberta-base-language-detection
|
||||
litellm_params:
|
||||
model: hosted_vllm/papluca/xlm-roberta-base-language-detection
|
||||
api_base: http://localhost:8090
|
||||
```
|
||||
|
||||
2. Run the proxy
|
||||
|
||||
```bash
|
||||
litellm proxy --config config.yaml
|
||||
|
||||
# RUNNING on http://localhost:4000
|
||||
```
|
||||
|
||||
3. Use the proxy
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:4000/vllm/classify \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer <your-api-key>" \
|
||||
-d '{"model": "roberta-base-language-detection", "input": "Hello, world!"}' \
|
||||
```
|
||||
|
||||
# How to add a provider for passthrough
|
||||
|
||||
See [VLLMModelInfo](https://github.com/BerriAI/litellm/blob/main/litellm/llms/vllm/common_utils.py) for an example.
|
||||
|
||||
1. Inherit from BaseModelInfo
|
||||
|
||||
```python
|
||||
from litellm.llms.base_llm.base_utils import BaseLLMModelInfo
|
||||
|
||||
class VLLMModelInfo(BaseLLMModelInfo):
|
||||
pass
|
||||
```
|
||||
|
||||
2. Register the provider in the ProviderConfigManager.get_provider_model_info
|
||||
|
||||
```python
|
||||
from litellm.utils import ProviderConfigManager
|
||||
from litellm.types.utils import LlmProviders
|
||||
|
||||
provider_config = ProviderConfigManager.get_provider_model_info(
|
||||
model="my-test-model", provider=LlmProviders.VLLM
|
||||
)
|
||||
|
||||
print(provider_config)
|
||||
```
|
||||
8
litellm/passthrough/__init__.py
Normal file
8
litellm/passthrough/__init__.py
Normal file
|
|
@ -0,0 +1,8 @@
|
|||
from .main import allm_passthrough_route, llm_passthrough_route
|
||||
from .utils import BasePassthroughUtils
|
||||
|
||||
__all__ = [
|
||||
"allm_passthrough_route",
|
||||
"llm_passthrough_route",
|
||||
"BasePassthroughUtils",
|
||||
]
|
||||
193
litellm/passthrough/main.py
Normal file
193
litellm/passthrough/main.py
Normal file
|
|
@ -0,0 +1,193 @@
|
|||
"""
|
||||
This module is used to pass through requests to the LLM APIs.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import contextvars
|
||||
from functools import partial
|
||||
from typing import Any, Coroutine, Optional, Union
|
||||
from urllib.parse import urlencode
|
||||
|
||||
import httpx
|
||||
from httpx._types import CookieTypes, QueryParamTypes, RequestFiles
|
||||
|
||||
import litellm
|
||||
from litellm.litellm_core_utils.get_llm_provider_logic import get_llm_provider
|
||||
from litellm.llms.custom_httpx.http_handler import AsyncHTTPHandler, HTTPHandler
|
||||
from litellm.utils import client
|
||||
|
||||
from .utils import BasePassthroughUtils
|
||||
|
||||
|
||||
@client
|
||||
async def allm_passthrough_route(
|
||||
*,
|
||||
method: str,
|
||||
endpoint: str,
|
||||
custom_llm_provider: Optional[str] = None,
|
||||
api_base: Optional[str] = None,
|
||||
api_key: Optional[str] = None,
|
||||
request_query_params: Optional[dict] = None,
|
||||
request_headers: Optional[dict] = None,
|
||||
stream: bool = False,
|
||||
content: Optional[Any] = None,
|
||||
data: Optional[dict] = None,
|
||||
files: Optional[RequestFiles] = None,
|
||||
json: Optional[Any] = None,
|
||||
params: Optional[QueryParamTypes] = None,
|
||||
cookies: Optional[CookieTypes] = None,
|
||||
client: Optional[Union[HTTPHandler, AsyncHTTPHandler]] = None,
|
||||
**kwargs,
|
||||
) -> Union[httpx.Response, Coroutine[Any, Any, httpx.Response]]:
|
||||
"""
|
||||
Async: Reranks a list of documents based on their relevance to the query
|
||||
"""
|
||||
try:
|
||||
loop = asyncio.get_event_loop()
|
||||
kwargs["allm_passthrough_route"] = True
|
||||
|
||||
func = partial(
|
||||
llm_passthrough_route,
|
||||
method=method,
|
||||
endpoint=endpoint,
|
||||
custom_llm_provider=custom_llm_provider,
|
||||
api_base=api_base,
|
||||
api_key=api_key,
|
||||
request_query_params=request_query_params,
|
||||
request_headers=request_headers,
|
||||
stream=stream,
|
||||
content=content,
|
||||
data=data,
|
||||
files=files,
|
||||
json=json,
|
||||
params=params,
|
||||
cookies=cookies,
|
||||
client=client,
|
||||
**kwargs,
|
||||
)
|
||||
|
||||
ctx = contextvars.copy_context()
|
||||
func_with_context = partial(ctx.run, func)
|
||||
init_response = await loop.run_in_executor(None, func_with_context)
|
||||
|
||||
if asyncio.iscoroutine(init_response):
|
||||
response = await init_response
|
||||
else:
|
||||
response = init_response
|
||||
return response
|
||||
except Exception as e:
|
||||
raise e
|
||||
|
||||
|
||||
@client
|
||||
def llm_passthrough_route(
|
||||
*,
|
||||
method: str,
|
||||
endpoint: str,
|
||||
model: str,
|
||||
custom_llm_provider: Optional[str] = None,
|
||||
api_base: Optional[str] = None,
|
||||
api_key: Optional[str] = None,
|
||||
request_query_params: Optional[dict] = None,
|
||||
request_headers: Optional[dict] = None,
|
||||
allm_passthrough_route: bool = False,
|
||||
stream: bool = False,
|
||||
content: Optional[Any] = None,
|
||||
data: Optional[dict] = None,
|
||||
files: Optional[RequestFiles] = None,
|
||||
json: Optional[Any] = None,
|
||||
params: Optional[QueryParamTypes] = None,
|
||||
cookies: Optional[CookieTypes] = None,
|
||||
client: Optional[Union[HTTPHandler, AsyncHTTPHandler]] = None,
|
||||
**kwargs,
|
||||
) -> Union[httpx.Response, Coroutine[Any, Any, httpx.Response]]:
|
||||
"""
|
||||
Pass through requests to the LLM APIs.
|
||||
|
||||
Step 1. Build the request
|
||||
Step 2. Send the request
|
||||
Step 3. Return the response
|
||||
|
||||
[TODO] Refactor this into a provider-config pattern, once we expand this to non-vllm providers.
|
||||
"""
|
||||
if client is None:
|
||||
if allm_passthrough_route:
|
||||
client = litellm.module_level_aclient
|
||||
else:
|
||||
client = litellm.module_level_client
|
||||
|
||||
model, custom_llm_provider, api_key, api_base = get_llm_provider(
|
||||
model=model,
|
||||
custom_llm_provider=custom_llm_provider,
|
||||
api_base=api_base,
|
||||
api_key=api_key,
|
||||
)
|
||||
|
||||
from litellm.types.utils import LlmProviders
|
||||
from litellm.utils import ProviderConfigManager
|
||||
|
||||
provider_config = ProviderConfigManager.get_provider_model_info(
|
||||
provider=LlmProviders(custom_llm_provider),
|
||||
model=model,
|
||||
)
|
||||
if provider_config is None:
|
||||
raise Exception(f"Provider {custom_llm_provider} not found")
|
||||
|
||||
base_target_url = provider_config.get_api_base(api_base)
|
||||
|
||||
if base_target_url is None:
|
||||
raise Exception(f"Provider {custom_llm_provider} api base not found")
|
||||
|
||||
encoded_endpoint = httpx.URL(endpoint).path
|
||||
|
||||
# Ensure endpoint starts with '/' for proper URL construction
|
||||
if not encoded_endpoint.startswith("/"):
|
||||
encoded_endpoint = "/" + encoded_endpoint
|
||||
|
||||
# Construct the full target URL using httpx
|
||||
base_url = httpx.URL(base_target_url)
|
||||
updated_url = base_url.copy_with(path=encoded_endpoint)
|
||||
|
||||
if request_query_params:
|
||||
# Create a new URL with the merged query params
|
||||
updated_url = updated_url.copy_with(
|
||||
query=urlencode(request_query_params).encode("ascii")
|
||||
)
|
||||
|
||||
# Add or update query parameters
|
||||
provider_api_key = provider_config.get_api_key(api_key)
|
||||
|
||||
auth_headers = provider_config.validate_environment(
|
||||
headers={},
|
||||
model=model,
|
||||
messages=[],
|
||||
optional_params={},
|
||||
litellm_params={},
|
||||
api_key=provider_api_key,
|
||||
api_base=base_target_url,
|
||||
)
|
||||
|
||||
headers = BasePassthroughUtils.forward_headers_from_request(
|
||||
request_headers=request_headers or {},
|
||||
headers=auth_headers,
|
||||
forward_headers=False,
|
||||
)
|
||||
|
||||
## SWAP MODEL IN JSON BODY
|
||||
if json and isinstance(json, dict) and "model" in json:
|
||||
json["model"] = model
|
||||
|
||||
request = client.client.build_request(
|
||||
method=method,
|
||||
url=updated_url,
|
||||
content=content,
|
||||
data=data,
|
||||
files=files,
|
||||
json=json,
|
||||
params=params,
|
||||
headers=headers,
|
||||
cookies=cookies,
|
||||
)
|
||||
|
||||
response = client.client.send(request=request, stream=stream)
|
||||
return response
|
||||
39
litellm/passthrough/utils.py
Normal file
39
litellm/passthrough/utils.py
Normal file
|
|
@ -0,0 +1,39 @@
|
|||
from typing import Dict, List, Optional, Union
|
||||
from urllib.parse import parse_qs
|
||||
|
||||
import httpx
|
||||
|
||||
|
||||
class BasePassthroughUtils:
|
||||
@staticmethod
|
||||
def get_merged_query_parameters(
|
||||
existing_url: httpx.URL, request_query_params: Dict[str, Union[str, list]]
|
||||
) -> Dict[str, Union[str, List[str]]]:
|
||||
# Get the existing query params from the target URL
|
||||
existing_query_string = existing_url.query.decode("utf-8")
|
||||
existing_query_params = parse_qs(existing_query_string)
|
||||
|
||||
# parse_qs returns a dict where each value is a list, so let's flatten it
|
||||
updated_existing_query_params = {
|
||||
k: v[0] if len(v) == 1 else v for k, v in existing_query_params.items()
|
||||
}
|
||||
# Merge the query params, giving priority to the existing ones
|
||||
return {**request_query_params, **updated_existing_query_params}
|
||||
|
||||
@staticmethod
|
||||
def forward_headers_from_request(
|
||||
request_headers: dict,
|
||||
headers: dict,
|
||||
forward_headers: Optional[bool] = False,
|
||||
):
|
||||
"""
|
||||
Helper to forward headers from original request
|
||||
"""
|
||||
if forward_headers is True:
|
||||
# Header We Should NOT forward
|
||||
request_headers.pop("content-length", None)
|
||||
request_headers.pop("host", None)
|
||||
|
||||
# Combine request headers with custom headers
|
||||
headers = {**request_headers, **headers}
|
||||
return headers
|
||||
|
|
@ -73,13 +73,13 @@ async def get_mcp_servers_by_verificationtoken(
|
|||
)
|
||||
)
|
||||
|
||||
mcp_servers = []
|
||||
mcp_servers: Optional[List[str]] = []
|
||||
if (
|
||||
verification_token_record is not None
|
||||
and verification_token_record.object_permission is not None
|
||||
):
|
||||
mcp_servers = verification_token_record.object_permission.mcp_servers
|
||||
return mcp_servers
|
||||
return mcp_servers or []
|
||||
|
||||
|
||||
async def get_mcp_servers_by_team(
|
||||
|
|
@ -99,10 +99,10 @@ async def get_mcp_servers_by_team(
|
|||
)
|
||||
)
|
||||
|
||||
mcp_servers = []
|
||||
mcp_servers: Optional[List[str]] = []
|
||||
if team_record is not None and team_record.object_permission is not None:
|
||||
mcp_servers = team_record.object_permission.mcp_servers
|
||||
return mcp_servers
|
||||
return mcp_servers or []
|
||||
|
||||
|
||||
async def get_all_mcp_servers_for_user(
|
||||
|
|
|
|||
|
|
@ -72,8 +72,9 @@ class MCPServerManager:
|
|||
mcp_info = MCPInfo(**_mcp_info)
|
||||
mcp_info["server_name"] = server_name
|
||||
mcp_info["description"] = server_config.get("description", None)
|
||||
server_id = str(uuid.uuid4())
|
||||
new_server = MCPServer(
|
||||
server_id=str(uuid.uuid4()),
|
||||
server_id=server_id,
|
||||
name=server_name,
|
||||
url=server_config["url"],
|
||||
# TODO: utility fn the default values
|
||||
|
|
@ -82,7 +83,6 @@ class MCPServerManager:
|
|||
auth_type=server_config.get("auth_type", None),
|
||||
mcp_info=mcp_info,
|
||||
)
|
||||
server_id = str(uuid.uuid4())
|
||||
self.config_mcp_servers[server_id] = new_server
|
||||
verbose_logger.debug(
|
||||
f"Loaded MCP Servers: {json.dumps(self.config_mcp_servers, indent=4, default=str)}"
|
||||
|
|
|
|||
12
litellm/proxy/_experimental/mcp_server/utils.py
Normal file
12
litellm/proxy/_experimental/mcp_server/utils.py
Normal file
|
|
@ -0,0 +1,12 @@
|
|||
import importlib
|
||||
|
||||
|
||||
def is_mcp_available() -> bool:
|
||||
"""
|
||||
Returns True if the MCP module is available, False otherwise
|
||||
"""
|
||||
try:
|
||||
importlib.import_module("mcp")
|
||||
return True
|
||||
except ImportError:
|
||||
return False
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1 @@
|
|||
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[185],{96443:function(e,n,t){Promise.resolve().then(t.t.bind(t,39974,23)),Promise.resolve().then(t.t.bind(t,2778,23))},2778:function(){},39974:function(e){e.exports={style:{fontFamily:"'__Inter_3373e4', '__Inter_Fallback_3373e4'",fontStyle:"normal"},className:"__className_3373e4"}}},function(e){e.O(0,[919,986,971,117,744],function(){return e(e.s=96443)}),_N_E=e.O()}]);
|
||||
|
|
@ -1 +0,0 @@
|
|||
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[185],{6580:function(n,e,t){Promise.resolve().then(t.t.bind(t,39974,23)),Promise.resolve().then(t.t.bind(t,2778,23))},2778:function(){},39974:function(n){n.exports={style:{fontFamily:"'__Inter_cf7686', '__Inter_Fallback_cf7686'",fontStyle:"normal"},className:"__className_cf7686"}}},function(n){n.O(0,[919,191,971,117,744],function(){return n(n.s=6580)}),_N_E=n.O()}]);
|
||||
|
|
@ -1 +1 @@
|
|||
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[418],{11790:function(e,n,u){Promise.resolve().then(u.bind(u,52829))},52829:function(e,n,u){"use strict";u.r(n),u.d(n,{default:function(){return f}});var t=u(57437),s=u(2265),r=u(99376),c=u(92699);function f(){let e=(0,r.useSearchParams)().get("key"),[n,u]=(0,s.useState)(null);return(0,s.useEffect)(()=>{e&&u(e)},[e]),(0,t.jsx)(c.Z,{accessToken:n,publicPage:!0,premiumUser:!1})}}},function(e){e.O(0,[402,313,250,699,971,117,744],function(){return e(e.s=11790)}),_N_E=e.O()}]);
|
||||
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[418],{21024:function(e,n,u){Promise.resolve().then(u.bind(u,52829))},52829:function(e,n,u){"use strict";u.r(n),u.d(n,{default:function(){return f}});var t=u(57437),s=u(2265),r=u(99376),c=u(92699);function f(){let e=(0,r.useSearchParams)().get("key"),[n,u]=(0,s.useState)(null);return(0,s.useEffect)(()=>{e&&u(e)},[e]),(0,t.jsx)(c.Z,{accessToken:n,publicPage:!0,premiumUser:!1})}}},function(e){e.O(0,[402,313,250,699,971,117,744],function(){return e(e.s=21024)}),_N_E=e.O()}]);
|
||||
|
|
@ -0,0 +1 @@
|
|||
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[461],{8672:function(e,t,n){Promise.resolve().then(n.bind(n,12011))},12011:function(e,t,n){"use strict";n.r(t),n.d(t,{default:function(){return S}});var o=n(57437),s=n(2265),a=n(99376),c=n(20831),i=n(94789),l=n(12514),r=n(49804),u=n(67101),d=n(84264),m=n(49566),h=n(96761),x=n(84566),f=n(19250),p=n(14474),g=n(13634),k=n(73002),j=n(3914);function S(){let[e]=g.Z.useForm(),t=(0,a.useSearchParams)();(0,j.e)("token");let n=t.get("invitation_id"),[S,w]=(0,s.useState)(null),[Z,_]=(0,s.useState)(""),[b,N]=(0,s.useState)(""),[y,T]=(0,s.useState)(null),[E,v]=(0,s.useState)(""),[C,U]=(0,s.useState)(""),[F,J]=(0,s.useState)(!0);return(0,s.useEffect)(()=>{(0,f.MO)().then(e=>{console.log("ui config in onboarding.tsx:",e),J(!1)})},[]),(0,s.useEffect)(()=>{n&&!F&&(0,f.W_)(n).then(e=>{let t=e.login_url;console.log("login_url:",t),v(t);let n=e.token,o=(0,p.o)(n);U(n),console.log("decoded:",o),w(o.key),console.log("decoded user email:",o.user_email),N(o.user_email),T(o.user_id)})},[n,F]),(0,o.jsx)("div",{className:"mx-auto w-full max-w-md mt-10",children:(0,o.jsxs)(l.Z,{children:[(0,o.jsx)(h.Z,{className:"text-sm mb-5 text-center",children:"\uD83D\uDE85 LiteLLM"}),(0,o.jsx)(h.Z,{className:"text-xl",children:"Sign up"}),(0,o.jsx)(d.Z,{children:"Claim your user account to login to Admin UI."}),(0,o.jsx)(i.Z,{className:"mt-4",title:"SSO",icon:x.GH$,color:"sky",children:(0,o.jsxs)(u.Z,{numItems:2,className:"flex justify-between items-center",children:[(0,o.jsx)(r.Z,{children:"SSO is under the Enterprise Tier."}),(0,o.jsx)(r.Z,{children:(0,o.jsx)(c.Z,{variant:"primary",className:"mb-2",children:(0,o.jsx)("a",{href:"https://forms.gle/W3U4PZpJGFHWtHyA9",target:"_blank",children:"Get Free Trial"})})})]})}),(0,o.jsxs)(g.Z,{className:"mt-10 mb-5 mx-auto",layout:"vertical",onFinish:e=>{console.log("in handle submit. accessToken:",S,"token:",C,"formValues:",e),S&&C&&(e.user_email=b,y&&n&&(0,f.m_)(S,n,y,e.password).then(e=>{let t="/ui/";t+="?login=success",document.cookie="token="+C,console.log("redirecting to:",t);let n=(0,f.zX)();console.log("proxyBaseUrl:",n),n?window.location.href=n+t:window.location.href=t}))},children:[(0,o.jsxs)(o.Fragment,{children:[(0,o.jsx)(g.Z.Item,{label:"Email Address",name:"user_email",children:(0,o.jsx)(m.Z,{type:"email",disabled:!0,value:b,defaultValue:b,className:"max-w-md"})}),(0,o.jsx)(g.Z.Item,{label:"Password",name:"password",rules:[{required:!0,message:"password required to sign up"}],help:"Create a password for your account",children:(0,o.jsx)(m.Z,{placeholder:"",type:"password",className:"max-w-md"})})]}),(0,o.jsx)("div",{className:"mt-10",children:(0,o.jsx)(k.ZP,{htmlType:"submit",children:"Sign Up"})})]})]})})}},3914:function(e,t,n){"use strict";function o(){let e=window.location.hostname,t=["Lax","Strict","None"];["/","/ui"].forEach(n=>{document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,";"),document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; domain=").concat(e,";"),t.forEach(t=>{let o="None"===t?" Secure;":"";document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; SameSite=").concat(t,";").concat(o),document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; domain=").concat(e,"; SameSite=").concat(t,";").concat(o)})}),console.log("After clearing cookies:",document.cookie)}function s(e){let t=document.cookie.split("; ").find(t=>t.startsWith(e+"="));return t?t.split("=")[1]:null}n.d(t,{b:function(){return o},e:function(){return s}})}},function(e){e.O(0,[665,402,899,250,971,117,744],function(){return e(e.s=8672)}),_N_E=e.O()}]);
|
||||
|
|
@ -1 +0,0 @@
|
|||
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[461],{32922:function(e,t,n){Promise.resolve().then(n.bind(n,12011))},12011:function(e,t,n){"use strict";n.r(t),n.d(t,{default:function(){return S}});var s=n(57437),o=n(2265),a=n(99376),c=n(20831),i=n(94789),l=n(12514),r=n(49804),u=n(67101),m=n(84264),d=n(49566),h=n(96761),x=n(84566),p=n(19250),f=n(14474),k=n(13634),g=n(73002),j=n(3914);function S(){let[e]=k.Z.useForm(),t=(0,a.useSearchParams)();(0,j.e)("token");let n=t.get("invitation_id"),[S,w]=(0,o.useState)(null),[Z,_]=(0,o.useState)(""),[N,b]=(0,o.useState)(""),[T,y]=(0,o.useState)(null),[E,v]=(0,o.useState)(""),[C,U]=(0,o.useState)("");return(0,o.useEffect)(()=>{n&&(0,p.W_)(n).then(e=>{let t=e.login_url;console.log("login_url:",t),v(t);let n=e.token,s=(0,f.o)(n);U(n),console.log("decoded:",s),w(s.key),console.log("decoded user email:",s.user_email),b(s.user_email),y(s.user_id)})},[n]),(0,s.jsx)("div",{className:"mx-auto w-full max-w-md mt-10",children:(0,s.jsxs)(l.Z,{children:[(0,s.jsx)(h.Z,{className:"text-sm mb-5 text-center",children:"\uD83D\uDE85 LiteLLM"}),(0,s.jsx)(h.Z,{className:"text-xl",children:"Sign up"}),(0,s.jsx)(m.Z,{children:"Claim your user account to login to Admin UI."}),(0,s.jsx)(i.Z,{className:"mt-4",title:"SSO",icon:x.GH$,color:"sky",children:(0,s.jsxs)(u.Z,{numItems:2,className:"flex justify-between items-center",children:[(0,s.jsx)(r.Z,{children:"SSO is under the Enterprise Tier."}),(0,s.jsx)(r.Z,{children:(0,s.jsx)(c.Z,{variant:"primary",className:"mb-2",children:(0,s.jsx)("a",{href:"https://forms.gle/W3U4PZpJGFHWtHyA9",target:"_blank",children:"Get Free Trial"})})})]})}),(0,s.jsxs)(k.Z,{className:"mt-10 mb-5 mx-auto",layout:"vertical",onFinish:e=>{console.log("in handle submit. accessToken:",S,"token:",C,"formValues:",e),S&&C&&(e.user_email=N,T&&n&&(0,p.m_)(S,n,T,e.password).then(e=>{let t="/ui/";t+="?login=success",document.cookie="token="+C,console.log("redirecting to:",t),window.location.href=t}))},children:[(0,s.jsxs)(s.Fragment,{children:[(0,s.jsx)(k.Z.Item,{label:"Email Address",name:"user_email",children:(0,s.jsx)(d.Z,{type:"email",disabled:!0,value:N,defaultValue:N,className:"max-w-md"})}),(0,s.jsx)(k.Z.Item,{label:"Password",name:"password",rules:[{required:!0,message:"password required to sign up"}],help:"Create a password for your account",children:(0,s.jsx)(d.Z,{placeholder:"",type:"password",className:"max-w-md"})})]}),(0,s.jsx)("div",{className:"mt-10",children:(0,s.jsx)(g.ZP,{htmlType:"submit",children:"Sign Up"})})]})]})})}},3914:function(e,t,n){"use strict";function s(){let e=window.location.hostname,t=["Lax","Strict","None"];["/","/ui"].forEach(n=>{document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,";"),document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; domain=").concat(e,";"),t.forEach(t=>{let s="None"===t?" Secure;":"";document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; SameSite=").concat(t,";").concat(s),document.cookie="token=; expires=Thu, 01 Jan 1970 00:00:00 UTC; path=".concat(n,"; domain=").concat(e,"; SameSite=").concat(t,";").concat(s)})}),console.log("After clearing cookies:",document.cookie)}function o(e){let t=document.cookie.split("; ").find(t=>t.startsWith(e+"="));return t?t.split("=")[1]:null}n.d(t,{b:function(){return s},e:function(){return o}})}},function(e){e.O(0,[665,402,899,250,971,117,744],function(){return e(e.s=32922)}),_N_E=e.O()}]);
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -1 +1 @@
|
|||
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[744],{20169:function(e,n,t){Promise.resolve().then(t.t.bind(t,12846,23)),Promise.resolve().then(t.t.bind(t,19107,23)),Promise.resolve().then(t.t.bind(t,61060,23)),Promise.resolve().then(t.t.bind(t,4707,23)),Promise.resolve().then(t.t.bind(t,80,23)),Promise.resolve().then(t.t.bind(t,36423,23))}},function(e){var n=function(n){return e(e.s=n)};e.O(0,[971,117],function(){return n(54278),n(20169)}),_N_E=e.O()}]);
|
||||
(self.webpackChunk_N_E=self.webpackChunk_N_E||[]).push([[744],{10264:function(e,n,t){Promise.resolve().then(t.t.bind(t,12846,23)),Promise.resolve().then(t.t.bind(t,19107,23)),Promise.resolve().then(t.t.bind(t,61060,23)),Promise.resolve().then(t.t.bind(t,4707,23)),Promise.resolve().then(t.t.bind(t,80,23)),Promise.resolve().then(t.t.bind(t,36423,23))}},function(e){var n=function(n){return e(e.s=n)};e.O(0,[971,117],function(){return n(54278),n(10264)}),_N_E=e.O()}]);
|
||||
File diff suppressed because one or more lines are too long
|
|
@ -1 +1 @@
|
|||
!function(){"use strict";var e,t,n,r,o,u,i,c,f,a={},l={};function d(e){var t=l[e];if(void 0!==t)return t.exports;var n=l[e]={id:e,loaded:!1,exports:{}},r=!0;try{a[e].call(n.exports,n,n.exports,d),r=!1}finally{r&&delete l[e]}return n.loaded=!0,n.exports}d.m=a,e=[],d.O=function(t,n,r,o){if(n){o=o||0;for(var u=e.length;u>0&&e[u-1][2]>o;u--)e[u]=e[u-1];e[u]=[n,r,o];return}for(var i=1/0,u=0;u<e.length;u++){for(var n=e[u][0],r=e[u][1],o=e[u][2],c=!0,f=0;f<n.length;f++)i>=o&&Object.keys(d.O).every(function(e){return d.O[e](n[f])})?n.splice(f--,1):(c=!1,o<i&&(i=o));if(c){e.splice(u--,1);var a=r();void 0!==a&&(t=a)}}return t},d.n=function(e){var t=e&&e.__esModule?function(){return e.default}:function(){return e};return d.d(t,{a:t}),t},n=Object.getPrototypeOf?function(e){return Object.getPrototypeOf(e)}:function(e){return e.__proto__},d.t=function(e,r){if(1&r&&(e=this(e)),8&r||"object"==typeof e&&e&&(4&r&&e.__esModule||16&r&&"function"==typeof e.then))return e;var o=Object.create(null);d.r(o);var u={};t=t||[null,n({}),n([]),n(n)];for(var i=2&r&&e;"object"==typeof i&&!~t.indexOf(i);i=n(i))Object.getOwnPropertyNames(i).forEach(function(t){u[t]=function(){return e[t]}});return u.default=function(){return e},d.d(o,u),o},d.d=function(e,t){for(var n in t)d.o(t,n)&&!d.o(e,n)&&Object.defineProperty(e,n,{enumerable:!0,get:t[n]})},d.f={},d.e=function(e){return Promise.all(Object.keys(d.f).reduce(function(t,n){return d.f[n](e,t),t},[]))},d.u=function(e){},d.miniCssF=function(e){},d.g=function(){if("object"==typeof globalThis)return globalThis;try{return this||Function("return this")()}catch(e){if("object"==typeof window)return window}}(),d.o=function(e,t){return Object.prototype.hasOwnProperty.call(e,t)},r={},o="_N_E:",d.l=function(e,t,n,u){if(r[e]){r[e].push(t);return}if(void 0!==n)for(var i,c,f=document.getElementsByTagName("script"),a=0;a<f.length;a++){var l=f[a];if(l.getAttribute("src")==e||l.getAttribute("data-webpack")==o+n){i=l;break}}i||(c=!0,(i=document.createElement("script")).charset="utf-8",i.timeout=120,d.nc&&i.setAttribute("nonce",d.nc),i.setAttribute("data-webpack",o+n),i.src=d.tu(e)),r[e]=[t];var s=function(t,n){i.onerror=i.onload=null,clearTimeout(p);var o=r[e];if(delete r[e],i.parentNode&&i.parentNode.removeChild(i),o&&o.forEach(function(e){return e(n)}),t)return t(n)},p=setTimeout(s.bind(null,void 0,{type:"timeout",target:i}),12e4);i.onerror=s.bind(null,i.onerror),i.onload=s.bind(null,i.onload),c&&document.head.appendChild(i)},d.r=function(e){"undefined"!=typeof Symbol&&Symbol.toStringTag&&Object.defineProperty(e,Symbol.toStringTag,{value:"Module"}),Object.defineProperty(e,"__esModule",{value:!0})},d.nmd=function(e){return e.paths=[],e.children||(e.children=[]),e},d.tt=function(){return void 0===u&&(u={createScriptURL:function(e){return e}},"undefined"!=typeof trustedTypes&&trustedTypes.createPolicy&&(u=trustedTypes.createPolicy("nextjs#bundler",u))),u},d.tu=function(e){return d.tt().createScriptURL(e)},d.p="/ui/_next/",i={272:0,919:0,191:0},d.f.j=function(e,t){var n=d.o(i,e)?i[e]:void 0;if(0!==n){if(n)t.push(n[2]);else if(/^(191|272|919)$/.test(e))i[e]=0;else{var r=new Promise(function(t,r){n=i[e]=[t,r]});t.push(n[2]=r);var o=d.p+d.u(e),u=Error();d.l(o,function(t){if(d.o(i,e)&&(0!==(n=i[e])&&(i[e]=void 0),n)){var r=t&&("load"===t.type?"missing":t.type),o=t&&t.target&&t.target.src;u.message="Loading chunk "+e+" failed.\n("+r+": "+o+")",u.name="ChunkLoadError",u.type=r,u.request=o,n[1](u)}},"chunk-"+e,e)}}},d.O.j=function(e){return 0===i[e]},c=function(e,t){var n,r,o=t[0],u=t[1],c=t[2],f=0;if(o.some(function(e){return 0!==i[e]})){for(n in u)d.o(u,n)&&(d.m[n]=u[n]);if(c)var a=c(d)}for(e&&e(t);f<o.length;f++)r=o[f],d.o(i,r)&&i[r]&&i[r][0](),i[r]=0;return d.O(a)},(f=self.webpackChunk_N_E=self.webpackChunk_N_E||[]).forEach(c.bind(null,0)),f.push=c.bind(null,f.push.bind(f))}();
|
||||
!function(){"use strict";var e,t,n,r,o,u,i,c,f,a={},l={};function d(e){var t=l[e];if(void 0!==t)return t.exports;var n=l[e]={id:e,loaded:!1,exports:{}},r=!0;try{a[e].call(n.exports,n,n.exports,d),r=!1}finally{r&&delete l[e]}return n.loaded=!0,n.exports}d.m=a,e=[],d.O=function(t,n,r,o){if(n){o=o||0;for(var u=e.length;u>0&&e[u-1][2]>o;u--)e[u]=e[u-1];e[u]=[n,r,o];return}for(var i=1/0,u=0;u<e.length;u++){for(var n=e[u][0],r=e[u][1],o=e[u][2],c=!0,f=0;f<n.length;f++)i>=o&&Object.keys(d.O).every(function(e){return d.O[e](n[f])})?n.splice(f--,1):(c=!1,o<i&&(i=o));if(c){e.splice(u--,1);var a=r();void 0!==a&&(t=a)}}return t},d.n=function(e){var t=e&&e.__esModule?function(){return e.default}:function(){return e};return d.d(t,{a:t}),t},n=Object.getPrototypeOf?function(e){return Object.getPrototypeOf(e)}:function(e){return e.__proto__},d.t=function(e,r){if(1&r&&(e=this(e)),8&r||"object"==typeof e&&e&&(4&r&&e.__esModule||16&r&&"function"==typeof e.then))return e;var o=Object.create(null);d.r(o);var u={};t=t||[null,n({}),n([]),n(n)];for(var i=2&r&&e;"object"==typeof i&&!~t.indexOf(i);i=n(i))Object.getOwnPropertyNames(i).forEach(function(t){u[t]=function(){return e[t]}});return u.default=function(){return e},d.d(o,u),o},d.d=function(e,t){for(var n in t)d.o(t,n)&&!d.o(e,n)&&Object.defineProperty(e,n,{enumerable:!0,get:t[n]})},d.f={},d.e=function(e){return Promise.all(Object.keys(d.f).reduce(function(t,n){return d.f[n](e,t),t},[]))},d.u=function(e){},d.miniCssF=function(e){},d.g=function(){if("object"==typeof globalThis)return globalThis;try{return this||Function("return this")()}catch(e){if("object"==typeof window)return window}}(),d.o=function(e,t){return Object.prototype.hasOwnProperty.call(e,t)},r={},o="_N_E:",d.l=function(e,t,n,u){if(r[e]){r[e].push(t);return}if(void 0!==n)for(var i,c,f=document.getElementsByTagName("script"),a=0;a<f.length;a++){var l=f[a];if(l.getAttribute("src")==e||l.getAttribute("data-webpack")==o+n){i=l;break}}i||(c=!0,(i=document.createElement("script")).charset="utf-8",i.timeout=120,d.nc&&i.setAttribute("nonce",d.nc),i.setAttribute("data-webpack",o+n),i.src=d.tu(e)),r[e]=[t];var s=function(t,n){i.onerror=i.onload=null,clearTimeout(p);var o=r[e];if(delete r[e],i.parentNode&&i.parentNode.removeChild(i),o&&o.forEach(function(e){return e(n)}),t)return t(n)},p=setTimeout(s.bind(null,void 0,{type:"timeout",target:i}),12e4);i.onerror=s.bind(null,i.onerror),i.onload=s.bind(null,i.onload),c&&document.head.appendChild(i)},d.r=function(e){"undefined"!=typeof Symbol&&Symbol.toStringTag&&Object.defineProperty(e,Symbol.toStringTag,{value:"Module"}),Object.defineProperty(e,"__esModule",{value:!0})},d.nmd=function(e){return e.paths=[],e.children||(e.children=[]),e},d.tt=function(){return void 0===u&&(u={createScriptURL:function(e){return e}},"undefined"!=typeof trustedTypes&&trustedTypes.createPolicy&&(u=trustedTypes.createPolicy("nextjs#bundler",u))),u},d.tu=function(e){return d.tt().createScriptURL(e)},d.p="/litellm/_next/",i={272:0,919:0,986:0},d.f.j=function(e,t){var n=d.o(i,e)?i[e]:void 0;if(0!==n){if(n)t.push(n[2]);else if(/^(272|919|986)$/.test(e))i[e]=0;else{var r=new Promise(function(t,r){n=i[e]=[t,r]});t.push(n[2]=r);var o=d.p+d.u(e),u=Error();d.l(o,function(t){if(d.o(i,e)&&(0!==(n=i[e])&&(i[e]=void 0),n)){var r=t&&("load"===t.type?"missing":t.type),o=t&&t.target&&t.target.src;u.message="Loading chunk "+e+" failed.\n("+r+": "+o+")",u.name="ChunkLoadError",u.type=r,u.request=o,n[1](u)}},"chunk-"+e,e)}}},d.O.j=function(e){return 0===i[e]},c=function(e,t){var n,r,o=t[0],u=t[1],c=t[2],f=0;if(o.some(function(e){return 0!==i[e]})){for(n in u)d.o(u,n)&&(d.m[n]=u[n]);if(c)var a=c(d)}for(e&&e(t);f<o.length;f++)r=o[f],d.o(i,r)&&i[r]&&i[r][0](),i[r]=0;return d.O(a)},(f=self.webpackChunk_N_E=self.webpackChunk_N_E||[]).forEach(c.bind(null,0)),f.push=c.bind(null,f.push.bind(f))}();
|
||||
File diff suppressed because one or more lines are too long
|
|
@ -1 +0,0 @@
|
|||
@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/55c55f0601d81cf3-s.woff2) format("woff2");unicode-range:u+0460-052f,u+1c80-1c8a,u+20b4,u+2de0-2dff,u+a640-a69f,u+fe2e-fe2f}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/26a46d62cd723877-s.woff2) format("woff2");unicode-range:u+0301,u+0400-045f,u+0490-0491,u+04b0-04b1,u+2116}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/97e0cb1ae144a2a9-s.woff2) format("woff2");unicode-range:u+1f??}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/581909926a08bbc8-s.woff2) format("woff2");unicode-range:u+0370-0377,u+037a-037f,u+0384-038a,u+038c,u+038e-03a1,u+03a3-03ff}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/df0a9ae256c0569c-s.woff2) format("woff2");unicode-range:u+0102-0103,u+0110-0111,u+0128-0129,u+0168-0169,u+01a0-01a1,u+01af-01b0,u+0300-0301,u+0303-0304,u+0308-0309,u+0323,u+0329,u+1ea0-1ef9,u+20ab}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/6d93bde91c0c2823-s.woff2) format("woff2");unicode-range:u+0100-02ba,u+02bd-02c5,u+02c7-02cc,u+02ce-02d7,u+02dd-02ff,u+0304,u+0308,u+0329,u+1d00-1dbf,u+1e00-1e9f,u+1ef2-1eff,u+2020,u+20a0-20ab,u+20ad-20c0,u+2113,u+2c60-2c7f,u+a720-a7ff}@font-face{font-family:__Inter_cf7686;font-style:normal;font-weight:100 900;font-display:swap;src:url(/ui/_next/static/media/a34f9d1faa5f3315-s.p.woff2) format("woff2");unicode-range:u+00??,u+0131,u+0152-0153,u+02bb-02bc,u+02c6,u+02da,u+02dc,u+0304,u+0308,u+0329,u+2000-206f,u+20ac,u+2122,u+2191,u+2193,u+2212,u+2215,u+feff,u+fffd}@font-face{font-family:__Inter_Fallback_cf7686;src:local("Arial");ascent-override:90.49%;descent-override:22.56%;line-gap-override:0.00%;size-adjust:107.06%}.__className_cf7686{font-family:__Inter_cf7686,__Inter_Fallback_cf7686;font-style:normal}
|
||||
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1 @@
|
|||
@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/55c55f0601d81cf3-s.woff2) format("woff2");unicode-range:u+0460-052f,u+1c80-1c8a,u+20b4,u+2de0-2dff,u+a640-a69f,u+fe2e-fe2f}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/26a46d62cd723877-s.woff2) format("woff2");unicode-range:u+0301,u+0400-045f,u+0490-0491,u+04b0-04b1,u+2116}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/97e0cb1ae144a2a9-s.woff2) format("woff2");unicode-range:u+1f??}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/581909926a08bbc8-s.woff2) format("woff2");unicode-range:u+0370-0377,u+037a-037f,u+0384-038a,u+038c,u+038e-03a1,u+03a3-03ff}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/df0a9ae256c0569c-s.woff2) format("woff2");unicode-range:u+0102-0103,u+0110-0111,u+0128-0129,u+0168-0169,u+01a0-01a1,u+01af-01b0,u+0300-0301,u+0303-0304,u+0308-0309,u+0323,u+0329,u+1ea0-1ef9,u+20ab}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/8e9860b6e62d6359-s.woff2) format("woff2");unicode-range:u+0100-02ba,u+02bd-02c5,u+02c7-02cc,u+02ce-02d7,u+02dd-02ff,u+0304,u+0308,u+0329,u+1d00-1dbf,u+1e00-1e9f,u+1ef2-1eff,u+2020,u+20a0-20ab,u+20ad-20c0,u+2113,u+2c60-2c7f,u+a720-a7ff}@font-face{font-family:__Inter_3373e4;font-style:normal;font-weight:100 900;font-display:swap;src:url(/litellm/_next/static/media/e4af272ccee01ff0-s.p.woff2) format("woff2");unicode-range:u+00??,u+0131,u+0152-0153,u+02bb-02bc,u+02c6,u+02da,u+02dc,u+0304,u+0308,u+0329,u+2000-206f,u+20ac,u+2122,u+2191,u+2193,u+2212,u+2215,u+feff,u+fffd}@font-face{font-family:__Inter_Fallback_3373e4;src:local("Arial");ascent-override:90.49%;descent-override:22.56%;line-gap-override:0.00%;size-adjust:107.06%}.__className_3373e4{font-family:__Inter_3373e4,__Inter_Fallback_3373e4;font-style:normal}
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Reference in a new issue