litellm/litellm/proxy/common_utils/performance_utils.md
stuxf a6c30b30bf
build: migrate packaging, CI, and Docker from Poetry to uv (#25007)
* build: migrate packaging metadata to uv

* ci: move automation and local tooling to uv

* docker: migrate image builds and runtime setup to uv

* docs: update install and deployment guidance for uv

* chore: align auxiliary scripts and tests with uv

* test: harden test_litellm isolation

* fix: keep release and health check images self-contained

* build: pin uv tooling and health check deps

* test: isolate bedrock image request formatting from suite state

* test: cover sandbox executor requirements flow

* ci: fix circleci no-op command steps

* ci: fix circleci publish workflow parsing

* fix: stabilize remaining uv migration CI checks

* ci: increase matrix test timeout headroom

* fix: restore published docker and license coverage

* fix: restore proxy runtime build parity

* fix: restore proxy extras parity and venv migrations

* ci: persist uv path across circleci steps

* fix: keep psycopg binary in default test env

* docker: preserve prisma cache across stages

* test: run local proxy checks through uv python

* build: restore runtime deps moved into ci

* build: refresh uv lock after upstream merge

* fix: restore module import in test_check_migration after merge

The conflict resolution imported only the function but the test body
references check_migration as a module throughout.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: revert dependency promotions, remove nodejs-wheel-binaries, fix Docker layer caching

- Move google-generativeai, Pillow, tenacity back to ci group (they are
  lazily imported and bloat the base SDK install needlessly)
- Remove nodejs-wheel-binaries from extra_proxy and proxy-dev (redundant
  in Docker where system Node.js is already installed via apk)
- Remove all nodejs-wheel node replacement and venv npm patching blocks
  from Dockerfiles since the wheel is no longer installed
- Add --no-default-groups to CodSpeed benchmark workflow so the benchmark
  environment matches the old minimal pip install footprint
- Apply standard uv two-phase Docker pattern: copy metadata first, install
  deps (cached layer), then copy source and install project
- Replace CircleCI enterprise no-op with proper uv sync command

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: regenerate uv.lock after removing nodejs-wheel-binaries

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): use cache/restore instead of cache to prevent cache poisoning

The old workflow used actions/cache/restore (read-only). The uv migration
changed it to actions/cache (read-write), which zizmor flags as a cache
poisoning risk. Restore the safer read-only variant.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): disable setup-uv built-in cache to silence cache-poisoning alert

The setup-uv action enables caching by default, which zizmor flags as a
cache poisoning risk. Disable it since we already use a read-only
cache/restore step.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): disable setup-uv cache in publish workflow

Silences zizmor cache-poisoning alert. Publishing workflow runs
infrequently on protected branches so caching adds no real benefit.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): remove duplicate verbose_logger mock in test_check_migration

The logger was patched twice — first via mocker.patch() then via
mocker.patch.object(autospec=True). The second call fails because
autospec cannot inspect an already-mocked attribute. Remove the
redundant first patch.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): free disk space before Docker build in test-server-root-path

The Dockerfile.non_root build ran out of disk on the CI runner. Remove
Android SDK, .NET, Boost, and GHC toolchains (~12GB) to free space.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-09 11:46:23 -07:00

6.8 KiB

Performance Utilities Documentation

This module provides performance monitoring and profiling functionality for LiteLLM proxy server using cProfile and line_profiler.

Table of Contents

Line Profiler Usage

Example 1: Wrapping a function directly

This is how it's used in litellm/utils.py to profile wrapper_async:

from litellm.proxy.common_utils.performance_utils import (
    register_shutdown_handler,
    wrap_function_directly,
)

def client(original_function):
    @wraps(original_function)
    async def wrapper_async(*args, **kwargs):
        # ... function implementation ...
        pass
    
    # Wrap the function with line_profiler
    wrapper_async = wrap_function_directly(wrapper_async)
    
    # Register shutdown handler to collect stats on server shutdown
    register_shutdown_handler(output_file="wrapper_async_line_profile.lprof")
    
    return wrapper_async

Example 2: Wrapping a module function dynamically

import my_module
from litellm.proxy.common_utils.performance_utils import (
    wrap_function_with_line_profiler,
    register_shutdown_handler,
)

# Wrap a function in a module
wrap_function_with_line_profiler(my_module, "expensive_function")

# Register shutdown handler
register_shutdown_handler(output_file="my_profile.lprof")

# Now all calls to my_module.expensive_function will be profiled
my_module.expensive_function()

Example 3: Manual stats collection

from litellm.proxy.common_utils.performance_utils import (
    wrap_function_directly,
    collect_line_profiler_stats,
)

def my_function():
    # ... implementation ...
    pass

# Wrap the function
my_function = wrap_function_directly(my_function)

# Run your code
my_function()

# Collect stats manually (instead of waiting for shutdown)
collect_line_profiler_stats(output_file="manual_profile.lprof")

Example 4: Analyzing the profile output

After running your code, analyze the .lprof file:

# View the profile
python -m line_profiler wrapper_async_line_profile.lprof

# Save to text file
python -m line_profiler wrapper_async_line_profile.lprof > profile_report.txt

The output shows:

  • Line #: Line number in the source file
  • Hits: Number of times the line was executed
  • Time: Total time spent on that line (in microseconds)
  • Per Hit: Average time per execution
  • % Time: Percentage of total function time
  • Line Contents: The actual source code

Example output:

Timer unit: 1e-06 s

Total time: 3.73697 s
File: litellm/utils.py
Function: client.<locals>.wrapper_async at line 1657

Line #      Hits         Time  Per Hit   % Time  Line Contents
==============================================================
  1657                                               @wraps(original_function)
  1658                                               async def wrapper_async(*args, **kwargs):
  1659      2005       7577.1      3.8      0.2          print_args_passed_to_litellm(...)
  1763      2005    1351909.0    674.3    36.2          result = await original_function(*args, **kwargs)
  1846      4010    1543688.1    385.0    41.3          update_response_metadata(...)

Example 5: Using in a decorator pattern

from litellm.proxy.common_utils.performance_utils import (
    wrap_function_directly,
    register_shutdown_handler,
)

def profile_decorator(func):
    # Wrap the function
    profiled_func = wrap_function_directly(func)
    
    # Register shutdown handler (only once)
    if not hasattr(profile_decorator, '_registered'):
        register_shutdown_handler(output_file="decorated_functions.lprof")
        profile_decorator._registered = True
    
    return profiled_func

@profile_decorator
async def my_async_function():
    # This function will be profiled
    pass

cProfile Usage

Example: Using the profile_endpoint decorator

from litellm.proxy.common_utils.performance_utils import profile_endpoint

@profile_endpoint(sampling_rate=0.1)  # Profile 10% of requests
async def my_endpoint():
    # ... implementation ...
    pass

The sampling_rate parameter controls what percentage of requests are profiled:

  • 1.0: Profile all requests (100%)
  • 0.1: Profile 1 in 10 requests (10%)
  • 0.0: Profile no requests (0%)

Installation

line_profiler must be installed to use the line profiling functionality:

uv add --dev line-profiler

On Windows with Python 3.14+, you may need to install Microsoft Visual C++ Build Tools to compile line_profiler from source.

Notes

  • The profiler aggregates stats by source code location, so multiple instances of the same function (e.g., closures) will be profiled together
  • Stats are automatically collected on server shutdown via atexit handler when using register_shutdown_handler()
  • You can also manually collect stats using collect_line_profiler_stats()
  • The line profiler will fail with an ImportError if line_profiler is not installed (as configured in litellm/utils.py)

API Reference

wrap_function_directly(func: Callable) -> Callable

Wrap a function directly with line_profiler. This is the recommended way to profile functions, especially closures or functions created dynamically.

Raises:

  • ImportError: If line_profiler is not available
  • RuntimeError: If line_profiler cannot be enabled or function cannot be wrapped

wrap_function_with_line_profiler(module: Any, function_name: str) -> bool

Dynamically wrap a function in a module with line_profiler.

Returns: True if wrapping was successful, False otherwise

collect_line_profiler_stats(output_file: Optional[str] = None) -> None

Collect and save line_profiler statistics. If output_file is provided, saves to file. Otherwise, prints to stdout.

register_shutdown_handler(output_file: Optional[str] = None) -> None

Register an atexit handler that will automatically save profiling statistics when the Python process exits. Safe to call multiple times (only registers once).

Default output file: line_profile_stats.lprof if not specified

profile_endpoint(sampling_rate: float = 1.0)

Decorator to sample endpoint hits and save to a profile file using cProfile.

Args:

  • sampling_rate: Rate of requests to profile (0.0 to 1.0)