litellm/litellm/proxy/hooks
Darien Kindlund 17e145a083
fix(proxy): use model_group for model_max_budget spend tracking cache key (#25549)
The model_max_budget limiter tracks spend in one code path
(async_log_success_event) and enforces budget limits in another
(is_key_within_model_budget via user_api_key_auth). These two paths
used different model name formats to build cache keys:

- Tracking used standard_logging_payload["model"], which is the
  deployment-level model name (e.g. "vertex_ai/claude-opus-4-6@default")
- Enforcement used request_data["model"], which is the model group
  alias (e.g. "claude-opus-4-6")

Because the cache keys never matched, the enforcement path always read
None for current spend, silently allowing all requests through even
after the budget was exceeded. This affected any provider that decorates
model names with provider prefixes or version suffixes (Vertex AI,
Bedrock, etc.).

Fix: use model_group (the user-facing alias) from StandardLoggingPayload
for spend tracking, falling back to model when model_group is None.
This aligns the tracking cache key with the enforcement cache key.

Fixes the same root cause reported in #15223 and #10052.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 19:37:58 -07:00
..
litellm_skills style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
mcp_semantic_filter fix(mypy): fix presidio, panw, perplexity, and mcp hook type issues 2026-03-13 00:01:26 +00:00
__init__.py Agents - add max budget + tpm/rpm limiting per agent AND per agent session (#22849) 2026-03-07 19:12:42 -08:00
azure_content_safety.py (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
batch_rate_limiter.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
batch_redis_get.py (code quality) run ruff rule to ban unused imports (#7313) 2024-12-19 12:33:42 -08:00
cache_control_check.py (minor latency fixes / proxy) - use verbose_proxy_logger.debug() instead of litellm.print_verbose (#7664) 2025-01-09 21:06:09 -08:00
dynamic_rate_limiter.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
dynamic_rate_limiter_v3.py [Feat] New LiteLLM Policy engine - create policies to manage guardrails, conditions - permissions per Key, Team (#19612) 2026-01-22 19:49:53 -08:00
example_presidio_ad_hoc_recognizer.json fix(presidio_pii_masking.py): enable user to pass ad hoc recognizer for pii masking 2024-02-20 16:01:15 -08:00
key_management_event_hooks.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
max_budget_limiter.py (minor latency fixes / proxy) - use verbose_proxy_logger.debug() instead of litellm.print_verbose (#7664) 2025-01-09 21:06:09 -08:00
max_budget_per_session_limiter.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
max_iterations_limiter.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
model_max_budget_limiter.py fix(proxy): use model_group for model_max_budget spend tracking cache key (#25549) 2026-04-11 19:37:58 -07:00
parallel_request_limiter.py Litellm ishaan march30 (#24887) (#25151) 2026-04-04 14:44:07 -07:00
parallel_request_limiter_v3.py feat: add proxy-wide default tpm/rpm limits per deployment 2026-03-19 01:30:18 -04:00
prompt_injection_detection.py fix: prompt injection not working (#16701) 2025-11-17 20:04:57 -08:00
proxy_track_cost_callback.py Litellm ishaan april1 try2 (#25110) 2026-04-03 14:57:44 -07:00
rate_limiter_utils.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
README.dynamic_rate_limiter_v3.md [Feat] Fixes to dynamic rate limiter v3 - add saturatation detection (#15119) 2025-10-01 18:35:34 -07:00
responses_id_security.py style: run black formatter on entire codebase 2026-03-11 17:07:57 -03:00
user_management_event_hooks.py Fix User Invite & Key Generation Email Notification Logic (#18524) 2026-01-06 01:35:52 +05:30

Dynamic Rate Limiter v3 - Saturation-Aware Priority-Based Rate Limiting

Overview

The v3 dynamic rate limiter implements saturation-aware rate limiting with priority-based allocation. It balances resource efficiency (allowing unused capacity to be borrowed) with fairness guarantees (enforcing priorities during high load).

Key Behavior:

  • When system is under 80% capacity: Generous mode - allows priority borrowing
  • When system is at/above 80% capacity: Strict mode - enforces normalized priority limits

How It Works

Flow Diagram

┌─────────────────────────────────────────────────────────────┐
│                    Incoming Request                          │
└────────────────────────┬────────────────────────────────────┘
                         │
                         ▼
┌─────────────────────────────────────────────────────────────┐
│  1. Check Model Saturation                                   │
│     - Query v3 limiter's Redis counters                      │
│     - Calculate: current_usage / capacity                    │
│     - Returns: 0.0 (empty) to 1.0+ (saturated)              │
└────────────────────────┬────────────────────────────────────┘
                         │
                         ▼
                ┌────────┴────────┐
                │  Saturation?    │
                └────────┬────────┘
                         │
         ┌───────────────┴───────────────┐
         │                               │
         ▼                               ▼
   < 80% (Generous)                >= 80% (Strict)
         │                               │
         ▼                               ▼
┌─────────────────────┐         ┌─────────────────────┐
│  Generous Mode      │         │  Strict Mode        │
│                     │         │                     │
│  - Enforce model-   │         │  - Normalize        │
│    wide capacity    │         │    priority weights │
│  - No priority      │         │    (if over 1.0)    │
│    restrictions     │         │                     │
│  - Allows borrowing │         │  - Create priority- │
│                     │         │    specific         │
│  - First-come-      │         │    descriptors      │
│    first-served     │         │                     │
│    until capacity   │         │  - Enforce strict   │
│                     │         │    limits per       │
│                     │         │    priority         │
└──────────┬──────────┘         └──────────┬──────────┘
           │                               │
           │                               ▼
           │                    ┌──────────────────────┐
           │                    │  Track model usage   │
           │                    │  for future          │
           │                    │  saturation checks   │
           │                    └──────────┬───────────┘
           │                               │
           └───────────────┬───────────────┘
                           │
                           ▼
                    ┌──────────────┐
                    │  v3 Limiter  │
                    │  Check       │
                    └──────┬───────┘
                           │
           ┌───────────────┴───────────────┐
           │                               │
           ▼                               ▼
     OVER_LIMIT                        OK
           │                               │
           ▼                               ▼
   Return 429 Error              Allow Request

Configuration

Priority Reservation

Set priority weights in your proxy configuration:

litellm.priority_reservation = {
    "premium": 0.75,    # 75% of capacity
    "standard": 0.25    # 25% of capacity
}

Priority Reservation Settings

Configure saturation-aware behavior:

litellm.priority_reservation_settings = PriorityReservationSettings(
    default_priority=0.5,           # Default weight for users without explicit priority
    saturation_threshold=0.80,      # 80% - threshold for strict mode enforcement
    tracking_multiplier=10          # 10x - multiplier for non-blocking tracking in strict mode
)

Settings:

  • default_priority (default: 0.5) - Priority weight for users without explicit priority metadata
  • saturation_threshold (default: 0.80) - Saturation level (0.0-1.0) at which strict priority enforcement begins
  • tracking_multiplier (default: 10) - Multiplier for model-wide tracking limits in strict mode

User Priority Assignment

Set priority in user metadata:

user_api_key_dict.metadata = {"priority": "premium"}

Priority Weight Normalization

If priorities sum to > 1.0, they are automatically normalized:

Input:  {key_a: 0.60, key_b: 0.80} = 1.40 total
Output: {key_a: 0.43, key_b: 0.57} = 1.00 total

This ensures total allocation never exceeds model capacity.

Implementation Details

Saturation Detection

  • Queries v3 limiter's Redis counters for model-wide usage
  • Checks both RPM and TPM, returns higher saturation value
  • Non-blocking reads (doesn't increment counters)

Mode Selection

Generous Mode (< 80% saturation):

  • Creates single model-wide descriptor
  • Enforces total capacity only
  • Allows any priority to use available capacity
  • Prevents over-subscription via model-wide limit

Strict Mode (>= 80% saturation):

  • Creates priority-specific descriptors with normalized weights
  • Each priority gets its reserved allocation
  • Tracks model-wide usage separately (non-blocking, 10x multiplier)
  • Ensures fairness under load

Test scenarios covered:

  1. No rate limiting when under capacity
  2. Priority queue behavior during saturation
  3. Spillover capacity for default keys
  4. Over-allocated priorities with normalization
  5. Default priority value handling

_PROXY_DynamicRateLimitHandlerV3

Main handler class inheriting from CustomLogger.

Key Methods:

  • async_pre_call_hook() - Main entry point, routes to generous/strict mode
  • _check_model_saturation() - Queries Redis for current usage
  • _handle_generous_mode() - Enforces model-wide capacity only
  • _handle_strict_mode() - Enforces normalized priority limits
  • _normalize_priority_weights() - Handles over-allocation
  • _create_priority_based_descriptors() - Creates rate limit descriptors