mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-22 00:31:44 +00:00
The model_max_budget limiter tracks spend in one code path (async_log_success_event) and enforces budget limits in another (is_key_within_model_budget via user_api_key_auth). These two paths used different model name formats to build cache keys: - Tracking used standard_logging_payload["model"], which is the deployment-level model name (e.g. "vertex_ai/claude-opus-4-6@default") - Enforcement used request_data["model"], which is the model group alias (e.g. "claude-opus-4-6") Because the cache keys never matched, the enforcement path always read None for current spend, silently allowing all requests through even after the budget was exceeded. This affected any provider that decorates model names with provider prefixes or version suffixes (Vertex AI, Bedrock, etc.). Fix: use model_group (the user-facing alias) from StandardLoggingPayload for spend tracking, falling back to model when model_group is None. This aligns the tracking cache key with the enforcement cache key. Fixes the same root cause reported in #15223 and #10052. Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| litellm_skills | ||
| mcp_semantic_filter | ||
| __init__.py | ||
| azure_content_safety.py | ||
| batch_rate_limiter.py | ||
| batch_redis_get.py | ||
| cache_control_check.py | ||
| dynamic_rate_limiter.py | ||
| dynamic_rate_limiter_v3.py | ||
| example_presidio_ad_hoc_recognizer.json | ||
| key_management_event_hooks.py | ||
| max_budget_limiter.py | ||
| max_budget_per_session_limiter.py | ||
| max_iterations_limiter.py | ||
| model_max_budget_limiter.py | ||
| parallel_request_limiter.py | ||
| parallel_request_limiter_v3.py | ||
| prompt_injection_detection.py | ||
| proxy_track_cost_callback.py | ||
| rate_limiter_utils.py | ||
| README.dynamic_rate_limiter_v3.md | ||
| responses_id_security.py | ||
| user_management_event_hooks.py | ||
Dynamic Rate Limiter v3 - Saturation-Aware Priority-Based Rate Limiting
Overview
The v3 dynamic rate limiter implements saturation-aware rate limiting with priority-based allocation. It balances resource efficiency (allowing unused capacity to be borrowed) with fairness guarantees (enforcing priorities during high load).
Key Behavior:
- When system is under 80% capacity: Generous mode - allows priority borrowing
- When system is at/above 80% capacity: Strict mode - enforces normalized priority limits
How It Works
Flow Diagram
┌─────────────────────────────────────────────────────────────┐
│ Incoming Request │
└────────────────────────┬────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ 1. Check Model Saturation │
│ - Query v3 limiter's Redis counters │
│ - Calculate: current_usage / capacity │
│ - Returns: 0.0 (empty) to 1.0+ (saturated) │
└────────────────────────┬────────────────────────────────────┘
│
▼
┌────────┴────────┐
│ Saturation? │
└────────┬────────┘
│
┌───────────────┴───────────────┐
│ │
▼ ▼
< 80% (Generous) >= 80% (Strict)
│ │
▼ ▼
┌─────────────────────┐ ┌─────────────────────┐
│ Generous Mode │ │ Strict Mode │
│ │ │ │
│ - Enforce model- │ │ - Normalize │
│ wide capacity │ │ priority weights │
│ - No priority │ │ (if over 1.0) │
│ restrictions │ │ │
│ - Allows borrowing │ │ - Create priority- │
│ │ │ specific │
│ - First-come- │ │ descriptors │
│ first-served │ │ │
│ until capacity │ │ - Enforce strict │
│ │ │ limits per │
│ │ │ priority │
└──────────┬──────────┘ └──────────┬──────────┘
│ │
│ ▼
│ ┌──────────────────────┐
│ │ Track model usage │
│ │ for future │
│ │ saturation checks │
│ └──────────┬───────────┘
│ │
└───────────────┬───────────────┘
│
▼
┌──────────────┐
│ v3 Limiter │
│ Check │
└──────┬───────┘
│
┌───────────────┴───────────────┐
│ │
▼ ▼
OVER_LIMIT OK
│ │
▼ ▼
Return 429 Error Allow Request
Configuration
Priority Reservation
Set priority weights in your proxy configuration:
litellm.priority_reservation = {
"premium": 0.75, # 75% of capacity
"standard": 0.25 # 25% of capacity
}
Priority Reservation Settings
Configure saturation-aware behavior:
litellm.priority_reservation_settings = PriorityReservationSettings(
default_priority=0.5, # Default weight for users without explicit priority
saturation_threshold=0.80, # 80% - threshold for strict mode enforcement
tracking_multiplier=10 # 10x - multiplier for non-blocking tracking in strict mode
)
Settings:
default_priority(default: 0.5) - Priority weight for users without explicit priority metadatasaturation_threshold(default: 0.80) - Saturation level (0.0-1.0) at which strict priority enforcement beginstracking_multiplier(default: 10) - Multiplier for model-wide tracking limits in strict mode
User Priority Assignment
Set priority in user metadata:
user_api_key_dict.metadata = {"priority": "premium"}
Priority Weight Normalization
If priorities sum to > 1.0, they are automatically normalized:
Input: {key_a: 0.60, key_b: 0.80} = 1.40 total
Output: {key_a: 0.43, key_b: 0.57} = 1.00 total
This ensures total allocation never exceeds model capacity.
Implementation Details
Saturation Detection
- Queries v3 limiter's Redis counters for model-wide usage
- Checks both RPM and TPM, returns higher saturation value
- Non-blocking reads (doesn't increment counters)
Mode Selection
Generous Mode (< 80% saturation):
- Creates single model-wide descriptor
- Enforces total capacity only
- Allows any priority to use available capacity
- Prevents over-subscription via model-wide limit
Strict Mode (>= 80% saturation):
- Creates priority-specific descriptors with normalized weights
- Each priority gets its reserved allocation
- Tracks model-wide usage separately (non-blocking, 10x multiplier)
- Ensures fairness under load
Test scenarios covered:
- No rate limiting when under capacity
- Priority queue behavior during saturation
- Spillover capacity for default keys
- Over-allocated priorities with normalization
- Default priority value handling
_PROXY_DynamicRateLimitHandlerV3
Main handler class inheriting from CustomLogger.
Key Methods:
async_pre_call_hook()- Main entry point, routes to generous/strict mode_check_model_saturation()- Queries Redis for current usage_handle_generous_mode()- Enforces model-wide capacity only_handle_strict_mode()- Enforces normalized priority limits_normalize_priority_weights()- Handles over-allocation_create_priority_based_descriptors()- Creates rate limit descriptors