diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md
index f0cce63b5dd..766f71cf2c5 100644
--- a/ARCHITECTURE.md
+++ b/ARCHITECTURE.md
@@ -1,629 +1,220 @@
# LiteLLM Architecture
-## 1. System Overview
+This document explains the internal architecture of LiteLLM for contributors. It describes the
+major components, how they interact, and the key design decisions that shape the library.
-LiteLLM is a unified interface for 100+ LLM providers. The system consists of two main components:
-the **Core Library** for direct LLM interactions and the **Proxy Server** (LLM Gateway) for
-production deployments with authentication, rate limiting, and observability.
+## Overview
+
+LiteLLM provides a unified interface for 100+ LLM providers. The system has two main components:
+
+1. **Core Library (`litellm/`)** - Direct Python SDK for LLM calls
+2. **Proxy Server (`litellm/proxy/`)** - Production LLM Gateway with auth, rate limiting, and observability
+
+## Request Flow
+
+The data flow for a completion request is as follows:
```mermaid
graph TD
- subgraph Client Application
- A[User/Application] --> B["litellm.completion()
litellm/main.py"]
+ subgraph "User Code"
+ Client["litellm.completion()"]
end
- subgraph "LiteLLM Core (litellm/)"
- B --> C["get_llm_provider()
litellm/utils.py"]
- C --> D["Provider Handler
litellm/llms/{provider}/"]
- D --> E["BaseConfig.transform_request()
litellm/llms/base_llm/"]
- E --> F["HTTPHandler
litellm/llms/custom_httpx/http_handler.py"]
+ subgraph "litellm/main.py"
+ GetProvider["get_llm_provider()"]
+ ProviderSwitch{"Provider
Switch"}
end
- subgraph LLM Providers
- F --> G[OpenAI API]
- F --> H[Anthropic API]
- F --> I[Azure API]
- F --> J[Bedrock API]
- F --> K[100+ Provider APIs]
+ subgraph "litellm/llms/custom_httpx/llm_http_handler.py"
+ BaseLLMHTTPHandler["BaseLLMHTTPHandler"]
end
+
+ subgraph "litellm/llms/{provider}/chat/transformation.py"
+ TransformRequest["ProviderConfig.transform_request()"]
+ TransformResponse["ProviderConfig.transform_response()"]
+ end
+
+ subgraph "litellm/llms/custom_httpx/http_handler.py"
+ HTTPHandler["HTTPHandler / AsyncHTTPHandler"]
+ end
+
+ subgraph "External"
+ ProviderAPI["LLM Provider API"]
+ end
+
+ Client --> GetProvider
+ GetProvider --> ProviderSwitch
+ ProviderSwitch --> BaseLLMHTTPHandler
+ BaseLLMHTTPHandler --> TransformRequest
+ TransformRequest --> HTTPHandler
+ HTTPHandler --> ProviderAPI
+ ProviderAPI --> HTTPHandler
+ HTTPHandler --> TransformResponse
+ TransformResponse --> BaseLLMHTTPHandler
+ BaseLLMHTTPHandler --> Client
```
-**Core Library (`litellm/`):**
-- **Purpose:** Provides a unified `completion()` interface that translates OpenAI-format requests
- to provider-specific formats and normalizes responses back to OpenAI format.
-- **Mechanism:** Uses transformation classes to convert inputs/outputs, handles streaming,
- function calling, and error mapping across all providers.
-- **Use Case:** Direct integration into Python applications for LLM calls.
+## Provider Configuration
-**Proxy Server (`litellm/proxy/`):**
-- **Purpose:** Production-ready LLM Gateway with authentication, rate limiting, load balancing,
- spend tracking, and admin UI.
-- **Mechanism:** FastAPI server that wraps the core library with enterprise features.
-- **Use Case:** Centralized LLM access for organizations with multiple teams and applications.
+There is an architectural separation between the HTTP handling logic and provider-specific
+transformations. The `BaseConfig` class is the base for all provider transformations. The
+`BaseLLMHTTPHandler` in `llms/custom_httpx/llm_http_handler.py` is the central handler that
+orchestrates all provider calls.
-## 2. Request Flow
+Each provider implements a `Config` class that inherits from `BaseConfig` and defines:
+- `transform_request()` - Convert OpenAI format to provider format
+- `transform_response()` - Convert provider response to OpenAI format
+- `get_supported_openai_params()` - List of supported parameters
+- `get_complete_url()` - Build the provider API endpoint
-Every request flows through a standard chain of handlers. The transformation layer converts
-inputs to provider-specific formats only after routing decisions are made.
-
-```mermaid
-sequenceDiagram
- participant Client
- participant ProxyServer as proxy/proxy_server.py
chat_completion()
- participant Auth as proxy/auth/
user_api_key_auth.py
- participant PreCall as proxy/litellm_pre_call_utils.py
add_litellm_data_to_request()
- participant Router as router.py
Router.acompletion()
- participant Main as main.py
completion()
- participant Transform as llms/base_llm/chat/
transformation.py
- participant HTTP as llms/custom_httpx/
http_handler.py
- participant Provider as LLM Provider API
-
- Client->>ProxyServer: POST /v1/chat/completions
- ProxyServer->>Auth: user_api_key_auth()
- Auth-->>ProxyServer: UserAPIKeyAuth
- ProxyServer->>PreCall: add_litellm_data_to_request()
- PreCall-->>ProxyServer: Enhanced Request Data
- ProxyServer->>Router: route_request() -> acompletion()
- Router->>Main: litellm.acompletion()
- Main->>Transform: BaseConfig.transform_request()
- Transform->>HTTP: AsyncHTTPHandler.post()
- HTTP->>Provider: Provider-specific HTTP Request
- Provider-->>HTTP: Provider Response
- HTTP-->>Transform: Raw Response
- Transform-->>Main: BaseConfig.transform_response()
- Main-->>Router: ModelResponse
- Router-->>ProxyServer: ModelResponse
- ProxyServer-->>Client: OpenAI-format JSON Response
-```
-
-### Request Processing Stages
-
-1. **Authentication (`proxy/auth/user_api_key_auth.py`):** Validates API keys, JWT tokens, or OAuth2 credentials.
- Extracts user, team, and organization context for downstream processing.
-
-2. **Pre-call Processing (`proxy/litellm_pre_call_utils.py`):** Adds metadata, applies guardrails via
- `proxy/hooks/`, and prepares request data.
-
-3. **Routing (`proxy/route_llm_request.py` -> `router.py`):** The `route_request()` function selects
- the appropriate model deployment based on load balancing strategy, cooldowns, and rate limits.
-
-4. **Provider Resolution (`litellm/utils.py`):** The `get_llm_provider()` function determines which
- provider handler to use based on the model name.
-
-5. **Transformation (`llms/base_llm/chat/transformation.py`):** The provider's `BaseConfig` subclass
- converts OpenAI-format requests to provider-specific formats.
-
-6. **HTTP Request (`llms/custom_httpx/http_handler.py`):** `AsyncHTTPHandler` or `HTTPHandler` makes
- the actual HTTP request to the LLM provider.
-
-7. **Response Processing (`llms/{provider}/chat/transformation.py`):** Provider's `transform_response()`
- normalizes the response back to OpenAI format.
-
-8. **Post-call Hooks (`integrations/custom_logger.py`):** Logs to observability platforms, updates
- spend tracking via callbacks registered in `litellm.callbacks`.
-
-## 3. Core Library Architecture
-
-### Main Entry Points
-
-The core library exposes several main functions in `litellm/main.py`:
-
-| Function | File Location | Purpose |
-|----------|---------------|---------|
-| `completion()` | `litellm/main.py:992` | Chat completions (sync) |
-| `acompletion()` | `litellm/main.py:369` | Chat completions (async) |
-| `embedding()` | `litellm/main.py:4352` | Text embeddings |
-| `text_completion()` | `litellm/main.py:5445` | Legacy text completions |
-| `image_generation()` | `litellm/main.py` | Image generation |
-| `transcription()` | `litellm/main.py` | Audio transcription |
-| `speech()` | `litellm/main.py` | Text-to-speech |
-
-### Provider Resolution
-
-When `completion()` is called, the provider is determined by `get_llm_provider()` in `litellm/utils.py`.
-The function parses the model string (e.g., `anthropic/claude-3-opus`) and returns:
-- `model` - The model name without provider prefix
-- `custom_llm_provider` - The provider identifier (e.g., "anthropic", "openai", "bedrock")
-- `api_key` - Resolved API key
-- `api_base` - Provider endpoint URL
-
-### Provider Implementation Pattern
-
-Each provider follows a consistent implementation pattern:
-
-```mermaid
-graph TD
- subgraph "Provider Implementation (litellm/llms/anthropic/)"
- A["BaseConfig
llms/base_llm/chat/transformation.py"] --> B["AnthropicConfig
llms/anthropic/chat/transformation.py"]
- B --> C["transform_request()"]
- B --> D["transform_response()"]
- B --> E["get_supported_openai_params()"]
- end
-
- subgraph "Provider Directory Structure"
- F["litellm/llms/anthropic/"]
- F --> G["chat/transformation.py
(AnthropicConfig)"]
- F --> H["chat/handler.py
(AnthropicChatCompletion)"]
- F --> I["common_utils.py
(AnthropicModelInfo)"]
- end
-```
-
-**Key Files per Provider (`litellm/llms/{provider}/`):**
-- `chat/transformation.py` - Request/response transformation inheriting from `BaseConfig`
-- `chat/handler.py` - HTTP request handling and streaming logic
-- `common_utils.py` - Shared utilities, model info, and constants
-
-### Transformation Layer
-
-The transformation layer (`litellm/llms/base_llm/`) provides base classes for all API types:
-
-| Base Class | File | Purpose |
-|------------|------|---------|
-| `BaseConfig` | `llms/base_llm/chat/transformation.py` | Chat completions transformation |
-| `BaseEmbeddingConfig` | `llms/base_llm/embedding/transformation.py` | Embedding transformation |
-| `BaseImageGenerationConfig` | `llms/base_llm/image_generation/transformation.py` | Image generation transformation |
-| `BaseAudioTranscriptionConfig` | `llms/base_llm/audio_transcription/transformation.py` | Audio transcription transformation |
-| `BaseBatchesConfig` | `llms/base_llm/batches/transformation.py` | Batch API transformation |
-
-Each provider implements these base classes to handle format conversion. Example from
-`litellm/llms/anthropic/chat/transformation.py`:
+Example from `litellm/llms/bedrock/chat/agentcore/transformation.py`:
```python
-class AnthropicConfig(BaseConfig):
- def transform_request(
- self,
- model: str,
- messages: List[AllMessageValues],
- optional_params: dict,
- litellm_params: dict,
- headers: dict,
- ) -> dict:
- # Convert OpenAI format to Anthropic format
- return {"model": model, "messages": transformed_messages, ...}
+class AmazonAgentCoreConfig(BaseConfig, BaseAWSLLM):
+ def transform_request(self, model, messages, optional_params, litellm_params, headers):
+ # Convert OpenAI messages to AgentCore format
+ return {"messages": transformed_messages, ...}
- def transform_response(
- self,
- model: str,
- raw_response: httpx.Response,
- model_response: ModelResponse,
- logging_obj: LiteLLMLoggingObj,
- ...
- ) -> ModelResponse:
- # Convert Anthropic response to OpenAI format
+ def transform_response(self, model, raw_response, model_response, logging_obj, ...):
+ # Convert AgentCore response to OpenAI ModelResponse
return ModelResponse(choices=[...], usage=Usage(...))
```
-## 4. Router System
+## HTTP Handler
-The Router (`litellm/router.py`) manages multiple model deployments with load balancing,
-fallbacks, and health monitoring.
+The `BaseLLMHTTPHandler` (`llms/custom_httpx/llm_http_handler.py`) is the central orchestrator
+for all provider calls. It:
-```mermaid
-graph TD
- subgraph "Router Configuration (router.py)"
- A["Model Group: gpt-4
Router.model_list"] --> B["Deployment 1: Azure
litellm_params.model=azure/gpt-4"]
- A --> C["Deployment 2: OpenAI
litellm_params.model=gpt-4"]
- A --> D["Deployment 3: Bedrock
litellm_params.model=bedrock/anthropic.claude-3"]
- end
+1. Receives the provider's `Config` class
+2. Calls `transform_request()` to build the provider-specific request
+3. Makes the HTTP call via `HTTPHandler` or `AsyncHTTPHandler`
+4. Calls `transform_response()` to normalize the response
+5. Handles streaming, retries, and error mapping
- subgraph "Routing Decision (router_strategy/)"
- E["Router.acompletion()"] --> F{"routing_strategy
router.py:200"}
- F -->|"simple-shuffle"| G["simple_shuffle.py"]
- F -->|"least-busy"| H["least_busy.py"]
- F -->|"lowest-latency"| I["lowest_latency.py"]
- F -->|"lowest-cost"| J["lowest_cost.py"]
- F -->|"lowest-tpm-rpm"| K["lowest_tpm_rpm.py"]
- end
+This design means adding a new provider only requires implementing a `Config` class - no changes
+to the HTTP handling logic.
- subgraph "Health Management (router_utils/)"
- L["cooldown_cache.py
CooldownCache"] --> M["Failed Deployments"]
- N["cooldown_handlers.py"] --> O["_set_cooldown_deployments()"]
- end
-```
+## Router
-### Routing Strategies (`litellm/router_strategy/`)
+The `Router` (`litellm/router.py`) manages multiple model deployments with load balancing,
+fallbacks, and health monitoring. It wraps the core `completion()` function.
-| Strategy | File | Description |
-|----------|------|-------------|
-| `simple-shuffle` | `simple_shuffle.py` | Random selection across healthy deployments |
-| `least-busy` | `least_busy.py` | Routes to deployment with lowest active requests |
-| `lowest-latency` | `lowest_latency.py` | Routes to historically fastest deployment |
-| `lowest-cost` | `lowest_cost.py` | Routes to cheapest available deployment |
-| `lowest-tpm-rpm` | `lowest_tpm_rpm.py` | Routes based on token/request capacity |
-| `tag-based` | `tag_based_routing.py` | Routes based on request metadata tags |
+Key concepts:
+- **Model Group**: A logical name (e.g., "gpt-4") that maps to multiple deployments
+- **Deployment**: A specific provider endpoint with credentials
+- **Routing Strategy**: Algorithm for selecting a deployment (e.g., `lowest-latency`, `simple-shuffle`)
+- **Cooldown**: Temporarily removing failed deployments from rotation
-### Fallback and Retry Logic (`litellm/router_utils/`)
+The routing strategies are implemented in `litellm/router_strategy/`:
+- `simple_shuffle.py` - Random selection
+- `lowest_latency.py` - Route to fastest deployment
+- `lowest_cost.py` - Route to cheapest deployment
+- `lowest_tpm_rpm.py` - Route based on available capacity
-| File | Purpose |
-|------|---------|
-| `fallback_event_handlers.py` | `run_async_fallback()`, `get_fallback_model_group()` |
-| `cooldown_handlers.py` | `_set_cooldown_deployments()`, `_async_get_cooldown_deployments()` |
-| `cooldown_cache.py` | `CooldownCache` class for tracking failed deployments |
-| `handle_error.py` | `send_llm_exception_alert()`, exception handling |
-| `get_retry_from_policy.py` | `get_num_retries_from_retry_policy()` |
-
-## 5. Proxy Server Architecture
+## Proxy Server
The Proxy Server (`litellm/proxy/proxy_server.py`) is a FastAPI application that wraps the
-core library with enterprise features.
+core library with enterprise features:
```mermaid
graph TD
- subgraph "Proxy Server (proxy/proxy_server.py)"
- A["FastAPI app"] --> B["chat_completion()
Line 5149"]
- A --> C["embeddings()
proxy_server.py"]
- A --> D["image_generation()
proxy_server.py"]
+ subgraph "Incoming Request"
+ Client["POST /v1/chat/completions"]
end
- subgraph "Authentication (proxy/auth/)"
- E["user_api_key_auth.py
user_api_key_auth()"] --> F["auth_checks.py"]
- G["handle_jwt.py
JWTHandler"] --> F
- H["oauth2_check.py"] --> F
+ subgraph "proxy/proxy_server.py"
+ Endpoint["chat_completion()"]
end
- subgraph "Request Processing"
- I["common_request_processing.py
ProxyBaseLLMRequestProcessing"] --> J["route_llm_request.py
route_request()"]
- J --> K["router.py
Router.acompletion()"]
+ subgraph "proxy/auth/"
+ Auth["user_api_key_auth()"]
end
- subgraph "Data Layer (proxy/db/)"
- L["prisma_client.py
PrismaClient"] --> M["PostgreSQL/SQLite"]
- N["caching/redis_cache.py"] --> O["Redis"]
+ subgraph "proxy/"
+ PreCall["litellm_pre_call_utils.py"]
+ RouteRequest["route_llm_request.py"]
end
+
+ subgraph "litellm/"
+ Router["router.py"]
+ Main["main.py"]
+ end
+
+ Client --> Endpoint
+ Endpoint --> Auth
+ Auth --> PreCall
+ PreCall --> RouteRequest
+ RouteRequest --> Router
+ Router --> Main
+ Main --> Client
```
-### Endpoint Categories
+Key subsystems:
+- **Authentication** (`proxy/auth/`) - API key, JWT, OAuth2 validation
+- **Management APIs** (`proxy/management_endpoints/`) - Key, team, model, budget management
+- **Guardrails** (`proxy/guardrails/`) - Content filtering and safety checks
+- **Database** (`proxy/db/`) - Prisma ORM for PostgreSQL/SQLite
-**OpenAI-compatible Endpoints (defined in `proxy/proxy_server.py`):**
+## Caching
-| Endpoint | Function | Line |
-|----------|----------|------|
-| `POST /v1/chat/completions` | `chat_completion()` | ~5149 |
-| `POST /v1/completions` | `completion()` | proxy_server.py |
-| `POST /v1/embeddings` | `embeddings()` | proxy_server.py |
-| `POST /v1/images/generations` | `image_generation()` | image_endpoints/ |
-| `POST /v1/audio/transcriptions` | `audio_transcriptions()` | proxy_server.py |
-| `POST /v1/audio/speech` | `audio_speech()` | proxy_server.py |
+LiteLLM provides multiple caching backends in `litellm/caching/`:
-**Management Endpoints (`proxy/management_endpoints/`):**
+| Backend | Use Case |
+|---------|----------|
+| `InMemoryCache` | Single-instance deployments |
+| `RedisCache` | Multi-instance with shared state |
+| `DualCache` | Fast local + persistent remote |
+| `S3Cache` | Long-term response storage |
-| File | Endpoints |
-|------|-----------|
-| `key_management_endpoints.py` | `/key/generate`, `/key/delete`, `/key/info` |
-| `team_endpoints.py` | `/team/new`, `/team/update`, `/team/delete` |
-| `internal_user_endpoints.py` | `/user/new`, `/user/update`, `/user/delete` |
-| `model_management_endpoints.py` | `/model/new`, `/model/delete`, `/model/info` |
-| `budget_management_endpoints.py` | `/budget/new`, `/budget/info` |
-| `organization_endpoints.py` | `/organization/new`, `/organization/update` |
+## Callbacks and Observability
-**Pass-through Endpoints (`proxy/pass_through_endpoints/`):**
-
-| File | Purpose |
-|------|---------|
-| `llm_passthrough_endpoints.py` | Provider-specific API forwarding |
-| `pass_through_endpoints.py` | Custom pass-through route initialization |
-
-### Authentication System (`proxy/auth/`)
-
-| File | Purpose |
-|------|---------|
-| `user_api_key_auth.py` | Main `user_api_key_auth()` dependency for FastAPI routes |
-| `auth_checks.py` | `get_team_object()`, permission and budget validation |
-| `handle_jwt.py` | `JWTHandler` class for JWT token processing |
-| `oauth2_check.py` | OAuth2 flow handling |
-| `model_checks.py` | `get_key_models()`, `get_team_models()` for access validation |
-| `route_checks.py` | Endpoint permission checks |
-
-### Database Schema (`proxy/schema.prisma`)
-
-The proxy uses Prisma ORM with the following key entities:
-
-| Table | Purpose |
-|-------|---------|
-| `LiteLLM_UserTable` | User accounts and settings |
-| `LiteLLM_TeamTable` | Team definitions and membership |
-| `LiteLLM_OrganizationTable` | Organization hierarchy |
-| `LiteLLM_VerificationToken` | API keys (hashed) |
-| `LiteLLM_SpendLogs` | Usage and spend tracking |
-| `LiteLLM_ModelTable` | Model configurations |
-| `LiteLLM_BudgetTable` | Budget definitions |
-
-## 6. Caching System
-
-LiteLLM provides multiple caching backends (`litellm/caching/`):
-
-```mermaid
-graph TD
- subgraph "Cache Backends (litellm/caching/)"
- A["in_memory_cache.py
InMemoryCache"] --> B["Local LRU Cache"]
- C["redis_cache.py
RedisCache"] --> D["Redis Server"]
- E["redis_cluster_cache.py
RedisClusterCache"] --> F["Redis Cluster"]
- G["s3_cache.py
S3Cache"] --> H["S3 Bucket"]
- I["disk_cache.py
DiskCache"] --> J["Local Filesystem"]
- end
-
- subgraph "Cache Strategy"
- K["dual_cache.py
DualCache"] --> L["In-Memory + Redis"]
- M["redis_semantic_cache.py
RedisSemanticCache"] --> N["Vector Similarity"]
- end
-```
-
-| Cache Type | File | Use Case |
-|------------|------|----------|
-| `InMemoryCache` | `in_memory_cache.py` | Single-instance deployments |
-| `RedisCache` | `redis_cache.py` | Multi-instance with shared state |
-| `RedisClusterCache` | `redis_cluster_cache.py` | Redis Cluster deployments |
-| `DualCache` | `dual_cache.py` | Fast local + persistent remote |
-| `S3Cache` | `s3_cache.py` | Long-term response storage |
-| `RedisSemanticCache` | `redis_semantic_cache.py` | Similar query deduplication |
-| `DiskCache` | `disk_cache.py` | Local filesystem caching |
-
-## 7. Integrations and Observability
-
-### Callback System (`litellm/integrations/`)
-
-LiteLLM supports 30+ observability integrations through a callback system:
-
-```mermaid
-graph LR
- subgraph "LiteLLM (litellm/main.py)"
- A["completion()"] --> B["litellm.callbacks
List[CustomLogger]"]
- end
-
- subgraph "Observability (integrations/)"
- B --> C["langfuse/
langfuse.py"]
- B --> D["datadog/
datadog.py"]
- B --> E["prometheus.py"]
- B --> F["opentelemetry.py"]
- B --> G["weights_biases.py"]
- B --> H["mlflow.py"]
- end
-
- subgraph "Alerting (integrations/)"
- B --> I["SlackAlerting/
slack_alerting.py"]
- B --> J["email_alerting.py"]
- end
-```
-
-**Key Integration Files (`litellm/integrations/`):**
-
-| Category | Files |
-|----------|-------|
-| **Tracing** | `langfuse/langfuse.py`, `datadog/datadog.py`, `opentelemetry.py`, `arize/arize.py` |
-| **Metrics** | `prometheus.py`, `cloudzero/cloudzero.py`, `openmeter.py` |
-| **Logging** | `s3.py`, `gcs_bucket/gcs_bucket.py`, `dynamodb.py` |
-| **Alerting** | `SlackAlerting/slack_alerting.py`, `email_alerting.py` |
-
-### Custom Callbacks
-
-Implement `CustomLogger` (from `integrations/custom_logger.py`) for custom integrations:
+The callback system (`litellm/integrations/`) enables logging to 30+ observability platforms.
+Callbacks implement the `CustomLogger` interface from `integrations/custom_logger.py`:
```python
-from litellm.integrations.custom_logger import CustomLogger
-
-class MyCallback(CustomLogger):
- def log_success_event(self, kwargs, response_obj, start_time, end_time):
- # Log successful completion
- pass
-
- def log_failure_event(self, kwargs, response_obj, start_time, end_time):
- # Log failed completion
- pass
-
- async def async_log_success_event(self, kwargs, response_obj, start_time, end_time):
- # Async version for non-blocking logging
- pass
+class CustomLogger:
+ def log_success_event(self, kwargs, response_obj, start_time, end_time): ...
+ def log_failure_event(self, kwargs, response_obj, start_time, end_time): ...
+ async def async_log_success_event(self, kwargs, response_obj, start_time, end_time): ...
```
Register callbacks via `litellm.callbacks.append(MyCallback())` or in proxy config YAML.
-## 8. Guardrails System
+## Adding a New Provider
-The guardrails system (`litellm/proxy/guardrails/`) provides content filtering and safety checks:
+1. Create `litellm/llms/{provider}/chat/transformation.py`
+2. Implement a `Config` class inheriting from `BaseConfig`
+3. Implement `transform_request()`, `transform_response()`, `get_complete_url()`
+4. Add provider routing in `litellm/main.py`
+5. Add tests in `tests/llm_translation/`
+6. Update `model_prices_and_context_window.json`
-```mermaid
-graph TD
- subgraph "Pre-call Guardrails (proxy/hooks/)"
- A["prompt_injection_detection.py"] --> B["Content Filtering"]
- B --> C["guardrails/guardrail_hooks/"]
- end
+See `litellm/llms/bedrock/chat/agentcore/transformation.py` for a complete example.
- subgraph "Guardrail Providers (guardrails/guardrail_hooks/)"
- D["lakera_ai.py"]
- E["bedrock_guardrails.py"]
- F["azure/prompt_shield.py"]
- G["presidio.py"]
- H["custom_guardrail.py"]
- end
-
- subgraph "Initialization"
- I["init_guardrails.py
init_guardrails_v2()"] --> J["guardrail_registry.py"]
- end
-```
-
-**Guardrail Providers (`proxy/guardrails/guardrail_hooks/`):**
-
-| Provider | File |
-|----------|------|
-| Lakera AI | `lakera_ai.py`, `lakera_ai_v2.py` |
-| Bedrock Guardrails | `bedrock_guardrails.py` |
-| Azure Content Safety | `azure/prompt_shield.py`, `azure/text_moderation.py` |
-| OpenAI Moderation | `openai/moderations.py` |
-| Presidio (PII) | `presidio.py` |
-| Aporia AI | `aporia_ai/aporia_ai.py` |
-| Custom | `custom_guardrail.py` |
-
-**Key Files:**
-- `init_guardrails.py` - `init_guardrails_v2()` initializes guardrails from config
-- `guardrail_registry.py` - Registry of available guardrail implementations
-- `guardrail_helpers.py` - Shared helper functions
-
-## 9. Type System
-
-LiteLLM uses Pydantic models for type safety (`litellm/types/`):
-
-| File | Purpose |
-|------|---------|
-| `types/utils.py` | Core response types (`ModelResponse`, `Usage`, `EmbeddingResponse`) |
-| `types/router.py` | Router configuration (`Deployment`, `LiteLLM_Params`, `RetryPolicy`) |
-| `types/llms/openai.py` | OpenAI-specific types (`ChatCompletionRequest`, `AllMessageValues`) |
-| `types/llms/anthropic.py` | Anthropic-specific types |
-| `types/integrations/*.py` | Integration configuration types |
-| `types/guardrails.py` | Guardrail configuration types |
-
-### Key Types (`litellm/types/utils.py`)
-
-```python
-# Core response type - returned by all completion calls
-class ModelResponse(BaseModel):
- id: str
- choices: List[Choices]
- created: int
- model: str
- usage: Usage
-
-# Usage tracking
-class Usage(BaseModel):
- prompt_tokens: int
- completion_tokens: int
- total_tokens: int
-```
-
-### Router Types (`litellm/types/router.py`)
-
-```python
-# Model deployment configuration
-class Deployment(BaseModel):
- model_name: str # User-facing model name
- litellm_params: LiteLLM_Params # Provider-specific params
- model_info: Optional[ModelInfo] # Pricing, context window info
-
-class LiteLLM_Params(BaseModel):
- model: str # Provider model string (e.g., "azure/gpt-4")
- api_key: Optional[str]
- api_base: Optional[str]
- # ... additional provider params
-```
-
-## 10. Directory Structure Reference
+## Directory Structure
```
litellm/
-├── main.py # completion(), acompletion(), embedding() - core entry points
-├── router.py # Router class - load balancing, fallbacks, health checks
-├── utils.py # get_llm_provider(), helper functions, response types
-├── exceptions.py # LiteLLM exception classes
-├── cost_calculator.py # completion_cost(), token counting
-├── _logging.py # Logging configuration
+├── main.py # Entry points: completion(), embedding(), etc.
+├── router.py # Load balancing and fallbacks
+├── utils.py # get_llm_provider(), helpers
│
-├── llms/ # Provider implementations (100+ providers)
-│ ├── base_llm/ # Base classes all providers inherit from
-│ │ ├── chat/transformation.py # BaseConfig class
-│ │ ├── embedding/transformation.py # BaseEmbeddingConfig
-│ │ └── ...
-│ ├── openai/
-│ │ ├── chat/transformation.py # OpenAIConfig
-│ │ ├── chat/handler.py # OpenAIChatCompletion
-│ │ └── openai.py
-│ ├── anthropic/
-│ │ ├── chat/transformation.py # AnthropicConfig
-│ │ └── chat/handler.py # AnthropicChatCompletion
-│ ├── azure/ # Azure OpenAI
-│ ├── bedrock/ # AWS Bedrock
-│ ├── vertex_ai/ # Google Vertex AI
-│ └── custom_httpx/
-│ └── http_handler.py # HTTPHandler, AsyncHTTPHandler
+├── llms/
+│ ├── base_llm/ # Base transformation classes
+│ │ └── chat/transformation.py # BaseConfig
+│ ├── custom_httpx/
+│ │ ├── llm_http_handler.py # BaseLLMHTTPHandler (central orchestrator)
+│ │ └── http_handler.py # HTTPHandler, AsyncHTTPHandler
+│ └── {provider}/
+│ └── chat/transformation.py # ProviderConfig
│
-├── proxy/ # Proxy server (LLM Gateway)
-│ ├── proxy_server.py # FastAPI app, chat_completion(), embeddings()
-│ ├── route_llm_request.py # route_request() - routes to router
-│ ├── common_request_processing.py # ProxyBaseLLMRequestProcessing
-│ ├── litellm_pre_call_utils.py # add_litellm_data_to_request()
-│ ├── auth/
-│ │ ├── user_api_key_auth.py # user_api_key_auth() dependency
-│ │ ├── auth_checks.py # Permission validation
-│ │ └── handle_jwt.py # JWTHandler
-│ ├── management_endpoints/
-│ │ ├── key_management_endpoints.py
-│ │ ├── team_endpoints.py
-│ │ └── model_management_endpoints.py
-│ ├── guardrails/
-│ │ ├── init_guardrails.py
-│ │ └── guardrail_hooks/ # Provider implementations
-│ ├── hooks/ # Pre/post call hooks
-│ ├── db/
-│ │ └── prisma_client.py # PrismaClient
-│ └── schema.prisma # Database schema
-│
-├── router_utils/ # Router helper modules
-│ ├── cooldown_handlers.py # Deployment cooldown logic
-│ ├── fallback_event_handlers.py # Fallback handling
-│ └── handle_error.py # Error handling utilities
-│
-├── router_strategy/ # Load balancing strategies
-│ ├── simple_shuffle.py
-│ ├── lowest_latency.py
-│ ├── lowest_cost.py
-│ └── tag_based_routing.py
-│
-├── caching/ # Cache implementations
-│ ├── redis_cache.py
-│ ├── in_memory_cache.py
-│ ├── dual_cache.py
-│ └── caching_handler.py
+├── proxy/
+│ ├── proxy_server.py # FastAPI application
+│ ├── auth/ # Authentication
+│ ├── management_endpoints/ # Admin APIs
+│ └── guardrails/ # Content filtering
│
+├── caching/ # Cache backends
├── integrations/ # Observability callbacks
-│ ├── custom_logger.py # CustomLogger base class
-│ ├── langfuse/langfuse.py
-│ ├── datadog/datadog.py
-│ ├── prometheus.py
-│ └── SlackAlerting/slack_alerting.py
-│
-├── types/ # Pydantic type definitions
-│ ├── utils.py # ModelResponse, Usage, etc.
-│ ├── router.py # Deployment, LiteLLM_Params
-│ └── llms/ # Provider-specific types
-│
-└── litellm_core_utils/ # Internal utilities
- ├── litellm_logging.py # Logging class
- ├── streaming_handler.py # Stream processing
- └── exception_mapping_utils.py # Exception mapping
+└── types/ # Pydantic models
```
-
-## 11. Contributing Guidelines
-
-### Adding a New Provider
-
-1. Create directory: `litellm/llms/{provider}/`
-2. Create `chat/transformation.py` with class inheriting from `BaseConfig` (`llms/base_llm/chat/transformation.py`)
-3. Implement `transform_request()` and `transform_response()` methods
-4. Add provider routing in `litellm/main.py` (search for `custom_llm_provider ==`)
-5. Add tests in `tests/llm_translation/test_{provider}.py`
-6. Update `model_prices_and_context_window.json` with model pricing
-
-### Adding a New Integration
-
-1. Create file in `litellm/integrations/{integration}.py`
-2. Implement class inheriting from `CustomLogger` (`integrations/custom_logger.py`)
-3. Implement `log_success_event()`, `log_failure_event()`, and async variants
-4. Register callback name in `litellm/__init__.py` (add to `_known_custom_logger_compatible_callbacks`)
-5. Add configuration types in `litellm/types/integrations/`
-6. Add tests in `tests/`
-
-### Adding a New Guardrail
-
-1. Create directory in `litellm/proxy/guardrails/guardrail_hooks/{guardrail}/`
-2. Implement guardrail class with `async_pre_call_hook()` and `async_post_call_hook()` methods
-3. Register in `proxy/guardrails/guardrail_registry.py`
-4. Add configuration schema in `litellm/types/guardrails.py`
-5. Add tests in `tests/proxy_unit_tests/`
-
-## 12. Security Considerations
-
-| Feature | Implementation |
-|---------|----------------|
-| API Key Storage | Keys hashed via `proxy/auth/auth_utils.py` before database storage |
-| Secret Management | `litellm/secret_managers/` - AWS Secrets Manager, Azure Key Vault, HashiCorp Vault |
-| Input Validation | Request validation in `proxy/litellm_pre_call_utils.py` |
-| Rate Limiting | `proxy/hooks/` - per-key, per-user, per-team limits |
-| Audit Logging | `proxy/spend_tracking/` - all requests logged with user context |