litellm/deploy/aws/ARCHITECTURE.md
Julio Quinteros Pro 445c67cfec Add AWS ECS deployment template matching benchmark specifications
This commit adds a complete 1-click deployment solution for LiteLLM on AWS ECS,
configured to match the benchmark specifications from https://docs.litellm.ai/docs/benchmarks

## What's Added

### Infrastructure (1 file)
- cloudformation-ecs.yaml: AWS CloudFormation template for ECS deployment
  - 4 ECS Fargate tasks (4 vCPU, 8 GB RAM each)
  - 4 workers per task (16 total workers)
  - RDS PostgreSQL database (db.t3.medium)
  - Application Load Balancer
  - VPC with public/private subnets across 2 AZs
  - Security groups, NAT Gateway, monitoring

### Deployment Tools (3 files)
- deploy.sh: Automated deployment script with interactive prompts
- test-deployment.sh: Deployment validation and health check script
- cost-calculator.sh: Interactive cost estimation tool

### Documentation (6 files)
- 00-START-HERE.md: Quick start guide and overview
- QUICKSTART.md: 5-minute deployment guide
- README.md: Complete deployment documentation
- ARCHITECTURE.md: Detailed architecture deep-dive with diagrams
- INDEX.md: Master index of all files
- .summary.md: Internal summary document

### Testing & Configuration (2 files)
- locustfile.py: Load testing script to replicate benchmark tests
- example-config.yaml: LiteLLM configuration example

## Configuration

- 4 instances with 4 vCPU and 8 GB RAM each
- 4 workers per instance
- Expected performance:
  - Median latency: ~100 ms
  - P95 latency: ~150 ms
  - Throughput: ~1,170 RPS
  - LiteLLM overhead: ~2 ms

## Usage

```bash
cd deploy/aws
./deploy.sh
```

## Monthly Cost

~$440-460 (pay-as-you-go) or ~$270-370 (with reserved capacity)

## Features

-  CloudFormation template validated with AWS
-  Production-ready with high availability
-  Secure by default (private subnets, security groups, encrypted secrets)
-  Well-documented with comprehensive guides
-  Includes validation and load testing tools
-  Cost-optimized configuration

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-16 13:03:05 -03:00

18 KiB
Raw Permalink Blame History

AWS Deployment Architecture

This document describes the architecture of the LiteLLM AWS deployment configured for benchmark performance.

Architecture Diagram

                                    ┌─────────────────┐
                                    │   Internet      │
                                    └────────┬────────┘
                                             │
                                             │ HTTPS/HTTP
                                             │
                   ┌─────────────────────────▼───────────────────────────┐
                   │         Application Load Balancer (ALB)             │
                   │                                                      │
                   │  - Internet-facing                                   │
                   │  - HTTP/HTTPS listeners                              │
                   │  - Health checks: /health/readiness                  │
                   └──────────────────┬───────────────────────────────────┘
                                      │
                      ┌───────────────┼───────────────┐
                      │               │               │
        ┌─────────────▼──┐  ┌────────▼────┐  ┌──────▼───────────┐
        │  ECS Task 1    │  │ ECS Task 2  │  │  ECS Task 3-4    │
        │  (Fargate)     │  │ (Fargate)   │  │  (Fargate)       │
        │                │  │             │  │                  │
        │  - 4 vCPU      │  │ - 4 vCPU    │  │  - 4 vCPU        │
        │  - 8 GB RAM    │  │ - 8 GB RAM  │  │  - 8 GB RAM      │
        │  - 4 workers   │  │ - 4 workers │  │  - 4 workers     │
        │                │  │             │  │                  │
        │  LiteLLM       │  │  LiteLLM    │  │  LiteLLM         │
        │  Port: 4000    │  │  Port: 4000 │  │  Port: 4000      │
        └────────┬───────┘  └──────┬──────┘  └────────┬─────────┘
                 │                 │                   │
                 └─────────────────┼───────────────────┘
                                   │
                                   │ PostgreSQL Protocol
                                   │ Port: 5432
                                   │
                         ┌─────────▼──────────┐
                         │  RDS PostgreSQL    │
                         │                    │
                         │  - db.t3.medium    │
                         │  - 2 vCPU          │
                         │  - 4 GB RAM        │
                         │  - 100 GB Storage  │
                         │  - Multi-AZ        │
                         │  - Auto backup     │
                         └────────────────────┘

Network Architecture

┌─────────────────────────────────────────────────────────────────┐
│                          VPC (10.0.0.0/16)                      │
│                                                                 │
│  ┌───────────────────────────────────────────────────────────┐ │
│  │              Public Subnets (2 AZs)                       │ │
│  │                                                            │ │
│  │  ┌─────────────────────┐    ┌─────────────────────┐     │ │
│  │  │  Public Subnet 1    │    │  Public Subnet 2    │     │ │
│  │  │  (10.0.1.0/24)      │    │  (10.0.2.0/24)      │     │ │
│  │  │                     │    │                     │     │ │
│  │  │  - ALB              │    │  - ALB              │     │ │
│  │  │  - NAT Gateway      │    │                     │     │ │
│  │  │  - Internet Gateway │    │                     │     │ │
│  │  └─────────────────────┘    └─────────────────────┘     │ │
│  └───────────────────────────────────────────────────────────┘ │
│                                                                 │
│  ┌───────────────────────────────────────────────────────────┐ │
│  │              Private Subnets (2 AZs)                      │ │
│  │                                                            │ │
│  │  ┌─────────────────────┐    ┌─────────────────────┐     │ │
│  │  │  Private Subnet 1   │    │  Private Subnet 2   │     │ │
│  │  │  (10.0.11.0/24)     │    │  (10.0.12.0/24)     │     │ │
│  │  │                     │    │                     │     │ │
│  │  │  - ECS Tasks        │    │  - ECS Tasks        │     │ │
│  │  │  - RDS Primary      │    │  - RDS Standby      │     │ │
│  │  └─────────────────────┘    └─────────────────────┘     │ │
│  └───────────────────────────────────────────────────────────┘ │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘

Security Groups

┌──────────────────────────────────────────────────────────────┐
│                      Security Groups                          │
└──────────────────────────────────────────────────────────────┘

┌─────────────────┐       ┌─────────────────┐      ┌──────────────────┐
│  ALB Security   │       │  ECS Security   │      │  RDS Security    │
│     Group       │       │     Group       │      │     Group        │
│                 │       │                 │      │                  │
│  Inbound:       │       │  Inbound:       │      │  Inbound:        │
│  - 80 (HTTP)    │──────▶│  - 4000 (HTTP)  │─────▶│  - 5432 (PG)     │
│  - 443 (HTTPS)  │  from │    from ALB SG  │ from │    from ECS SG   │
│    from 0.0.0.0 │  ALB  │                 │ ECS  │                  │
│                 │       │  Outbound:      │      │  Outbound:       │
│  Outbound:      │       │  - All          │      │  - All           │
│  - All          │       │                 │      │                  │
└─────────────────┘       └─────────────────┘      └──────────────────┘

Components

1. Application Load Balancer (ALB)

Purpose: Distributes incoming traffic across ECS tasks

Configuration:

  • Type: Application Load Balancer
  • Scheme: Internet-facing
  • Subnets: Public subnets in 2 availability zones
  • Listeners: HTTP (port 80), optionally HTTPS (port 443)
  • Health check: /health/readiness
  • Health check interval: 30 seconds
  • Healthy threshold: 2 consecutive successes
  • Unhealthy threshold: 3 consecutive failures

Benefits:

  • Automatic SSL termination (with HTTPS)
  • Health monitoring and automatic failover
  • Connection draining during deployments
  • Path-based routing (if needed)

2. ECS Fargate Tasks

Purpose: Run LiteLLM proxy containers

Configuration:

  • Launch type: Fargate
  • Task count: 4 (configurable)
  • CPU: 4 vCPU (4096 units) per task
  • Memory: 8 GB (8192 MB) per task
  • Workers: 4 per task
  • Total capacity: 16 vCPU, 32 GB RAM, 16 workers

Container Configuration:

  • Image: ghcr.io/berriai/litellm-database:main-latest
  • Port: 4000
  • Command: --port 4000 --num_workers 4
  • Health check: HTTP GET /health/liveliness
  • Environment variables:
    • DATABASE_URL: PostgreSQL connection string
    • STORE_MODEL_IN_DB: True
    • PROXY_MASTER_KEY: From Secrets Manager

Benefits:

  • Serverless containers (no EC2 management)
  • Automatic scaling capability
  • High availability across AZs
  • Isolated execution environment

3. RDS PostgreSQL

Purpose: Persistent storage for LiteLLM configuration and logs

Configuration:

  • Engine: PostgreSQL 16.3
  • Instance class: db.t3.medium (2 vCPU, 4 GB RAM)
  • Storage: 100 GB GP3 SSD
  • Multi-AZ: No (can be enabled for HA)
  • Backup retention: 7 days
  • Automated backups: Yes

Database Schema:

  • Managed by Prisma ORM
  • Tables: models, users, teams, keys, logs, etc.
  • Automatic migrations on deployment

Recommended Settings:

-- For 1-2K RPS workload
max_connections = 200
shared_buffers = 1GB
effective_cache_size = 3GB
maintenance_work_mem = 256MB
work_mem = 5MB

Benefits:

  • Automatic backups and point-in-time recovery
  • Automatic software patching
  • Monitoring via CloudWatch
  • Easy scaling (vertical and storage)

4. VPC and Networking

Configuration:

  • VPC CIDR: 10.0.0.0/16
  • Public Subnets: 10.0.1.0/24, 10.0.2.0/24
  • Private Subnets: 10.0.11.0/24, 10.0.12.0/24
  • NAT Gateway: 1 (in Public Subnet 1)
  • Internet Gateway: 1

Routing:

  • Public subnets → Internet Gateway
  • Private subnets → NAT Gateway → Internet Gateway

Benefits:

  • ECS tasks in private subnets for security
  • Database isolated from internet
  • Controlled outbound access via NAT Gateway
  • High availability across 2 AZs

5. Secrets Management

Configuration:

  • AWS Secrets Manager for sensitive data
  • Secrets:
    • Database password
    • LiteLLM master key
    • API keys (stored separately)

Benefits:

  • Encrypted at rest
  • Automatic rotation support
  • Audit logging via CloudTrail
  • Fine-grained IAM access control

6. Logging and Monitoring

CloudWatch Logs:

  • Log group: /ecs/[stack-name]-litellm
  • Retention: 7 days (configurable)
  • Logs from all ECS tasks

CloudWatch Metrics:

  • ECS: CPU, Memory, Task Count
  • ALB: Request Count, Latency, Target Health
  • RDS: Connections, CPU, Storage

Custom Metrics:

  • LiteLLM reports overhead in x-litellm-overhead-duration-ms header
  • Can be extracted and sent to CloudWatch

Data Flow

Request Flow

  1. Client Request

    Client → ALB (port 80/443)
    
  2. Load Balancing

    ALB → Target Group → Healthy ECS Tasks
    
    • ALB selects a healthy task using round-robin
    • Sticky sessions not enabled (stateless)
  3. LiteLLM Processing

    ECS Task → LiteLLM Proxy (4 workers)
    
    • Request handled by one of 4 workers
    • Worker selection by internal load balancing (Uvicorn)
  4. Database Operations

    LiteLLM → RDS PostgreSQL
    
    • Validate API key
    • Log request
    • Retrieve model configuration
  5. External LLM Call

    LiteLLM → External LLM Provider (OpenAI, Anthropic, etc.)
    
    • Transform request to provider format
    • Forward request via NAT Gateway
    • Receive and transform response
  6. Response Flow

    LiteLLM → ALB → Client
    
    • Response sent back through ALB
    • Overhead metrics in headers

Database Connection Pooling

┌──────────────────────────────────────────────┐
│  4 ECS Tasks × 4 Workers = 16 Workers       │
│                                              │
│  Each Worker → Connection Pool               │
│  Pool size: ~10 connections per worker       │
│  Total connections: ~160                     │
│                                              │
│  RDS max_connections: 200                    │
│  Available headroom: 40 connections          │
└──────────────────────────────────────────────┘

High Availability

Availability Zones

  • Resources deployed across 2 AZs
  • ECS tasks distributed automatically
  • RDS can be configured for Multi-AZ
  • ALB spans both AZs

Failure Scenarios

Single ECS Task Failure:

  • ALB marks task unhealthy
  • Traffic routed to other tasks
  • ECS starts replacement task
  • Impact: 25% capacity reduction (temporary)

Availability Zone Failure:

  • ALB routes all traffic to healthy AZ
  • ECS maintains tasks in remaining AZ
  • Impact: 50% capacity reduction (until AZ recovers)

Database Failure:

  • With Multi-AZ: Automatic failover to standby (~60-120s)
  • Without Multi-AZ: Manual restore from backup

Recovery Time Objectives

Scenario RTO RPO
Single task failure < 2 minutes None (stateless)
AZ failure < 1 minute None (stateless)
Database failure (Multi-AZ) < 2 minutes ~0 (sync replication)
Database failure (Single-AZ) 30-60 minutes ~5 minutes (backup)
Complete region failure Hours Depends on backup strategy

Scaling

Horizontal Scaling (Task Count)

Manual Scaling:

aws ecs update-service \
  --cluster [cluster-name] \
  --service [service-name] \
  --desired-count 8

Auto Scaling (CPU-based):

  • Scale out: When average CPU > 70%
  • Scale in: When average CPU < 30%
  • Min tasks: 2
  • Max tasks: 10

Expected Performance by Scale:

Tasks Workers Expected RPS Median Latency
2 8 ~1,035 ~200ms
4 16 ~1,170 ~100ms
8 32 ~2,000+ ~50-75ms

Vertical Scaling (Task Size)

Upgrade to 8 vCPU, 16 GB:

TaskCPU: 8192
TaskMemory: 16384

Benefits:

  • More workers per task (8-16 workers)
  • Better performance per task
  • Fewer tasks needed for same throughput

Database Scaling

Vertical Scaling:

  • Upgrade to db.r6g.large (2 vCPU → 8 vCPU)
  • Minimal downtime (~1-2 minutes)

Read Replicas:

  • Offload read queries
  • Reduce primary load
  • Not needed for typical LiteLLM workload

Cost Optimization

Reserved Capacity

ECS Fargate Savings Plans:

  • 1-year: ~20-30% savings
  • 3-year: ~40-50% savings
  • Applies to Fargate compute usage

RDS Reserved Instances:

  • 1-year: ~30% savings
  • 3-year: ~60% savings
  • Partial or full upfront payment

Right-Sizing

Monitor and adjust:

  • Use CloudWatch to track actual CPU/Memory usage
  • Scale down if consistently under 50% utilization
  • Scale up if consistently over 80% utilization

Alternative Configurations

Lower Cost (Dev/Test):

  • 2 tasks × 2 vCPU × 4 GB
  • db.t3.micro
  • Single AZ
  • Cost: ~$150-200/month

Production (HA + Performance):

  • 8 tasks × 4 vCPU × 8 GB
  • db.r6g.large (Multi-AZ)
  • Redis cluster
  • Cost: ~$1,200-1,500/month

Security Best Practices

Network Security

  • ECS tasks in private subnets
  • Database not publicly accessible
  • Security groups with principle of least privilege
  • NAT Gateway for controlled outbound access
  • ⚠️ Consider VPC endpoints for AWS services (S3, Secrets Manager)

Authentication & Authorization

  • Master key stored in Secrets Manager
  • IAM roles for task execution
  • IAM roles for task operations
  • ⚠️ Implement key rotation policy
  • ⚠️ Use IAM-based database authentication

Data Protection

  • RDS encryption at rest
  • Secrets Manager encryption
  • HTTPS termination at ALB (with certificate)
  • ⚠️ Enable CloudTrail for audit logging
  • ⚠️ Enable VPC Flow Logs

Compliance

  • Enable CloudWatch Logs encryption
  • Configure S3 for long-term log archival
  • Implement backup retention policies
  • Regular security assessments

Monitoring and Alerting

Key Metrics to Monitor

Application Performance:

  • Request latency (P50, P95, P99)
  • Request rate (RPS)
  • Error rate (4xx, 5xx)
  • LiteLLM overhead (custom metric)

Infrastructure Health:

  • ECS task count and health
  • CPU/Memory utilization
  • Database connections
  • Target health

Cost Metrics:

  • Fargate compute hours
  • Data transfer costs
  • RDS instance hours
  • NAT Gateway data transfer
Alarms:
  - High 5xx rate (> 1%)
  - High latency (P95 > 500ms)
  - Low healthy target count (< 2)
  - High database CPU (> 80%)
  - High database connections (> 180)
  - Task stopped unexpectedly

References