- Cache StandardLoggingMetadata.__annotations__.keys() as module-level frozenset
- Use set intersection to iterate only keys present in both metadata and supported keys
- Single lookup for user_api_key instead of 3 separate .get() calls
Results:
- get_standard_logging_metadata: 1.55s → 1.41s (9.2% faster)
The get_llm_provider call was used to determine if the target model
is a Gemini model for thought signature removal. This call is redundant
because _is_gemini_model already has a fallback that checks if "gemini"
is in the model name, which covers all cases where get_llm_provider
would return "gemini" as the provider.
Profiling shows this reduces function_setup time by ~37% (8.18s -> 5.19s
across 6000 requests), saving ~500µs per request.
- Add early return when all callbacks are None (common case)
- Add per-callback conditionals to only process callbacks that are set
- Reorder processing: success before async_success, failure before async_failure
(required because success/failure processing adds items to async lists)
Profiling shows ~79% reduction in process_dynamic_callbacks time (11.5% -> 0.8%)
Guard verbose_logger.debug() f-strings with isEnabledFor(logging.DEBUG)
checks in the router and cost calculation hot paths. Python evaluates
f-string arguments before the logging framework checks the log level,
causing expensive formatting on every request even with debug logging
disabled.
Changes:
- Remove redundant litellm_params.copy() in _completion/_acompletion
- Guard 5 debug logs in router.py (+ remove 1 duplicate log)
- Guard 6 debug logs in cost_calculator.py and utils.py
- get_model_info(): formatted 50+ field dict every call
- _apply_cost_margin(): called list(dict.keys()) every request
Profiled improvement: completion_cost 769µs → 637µs/call (-17.2%)
Tests that the fix using `getattr(type(self), method_name) is not getattr(CustomLogger, method_name)` correctly walks the full Method Resolution Order, unlike the buggy `method_name in type(self).__dict__` which only checks the immediate class.
Check if redact_standard_logging_payload_from_model_call_details was
overridden in a subclass before skipping redundant redaction. This
ensures custom callbacks with additional redaction logic still execute.
- Add global_redaction_applied flag to skip per-callback redaction
- Add should_redact param to avoid calling should_redact_message_logging twice
- Add tests for both flag=True (skip) and flag=False (proceed) cases
When turn_off_message_logging is enabled globally, skip per-callback
redaction functions that would re-process already-redacted data.
17% faster async_success_handler when global redaction is ON.
- Add _OPTIONAL_KWARGS_KEYS frozenset for O(1) lookups
- Replace 28 unconditional kwargs.get() calls with sparse extraction
- Only add kwargs keys that are actually present in the dict
- Simplify _get_base_model_from_litellm_call_metadata by removing redundant None checks
This reduces get_litellm_params() time by ~31% (743ms → 509ms across 6000 calls)
and Logging.__init__ total time by ~24% (1.61s → 1.23s).
Add fast path to check optional_params.get("api_base") directly before
constructing a full LiteLLM_Params Pydantic model. When api_base is present
(the common case via router), return it immediately — avoiding ~29µs of
Pydantic validation overhead per request.
Profiled: get_api_base 30.5µs/call → 1.2µs/call (-96%)
Revert the guard that skipped built-in tool cost when standard_built_in_tools_params
was falsy. get_cost_for_built_in_tools can return non-zero from response/usage alone
(e.g. web search) even when params is None/empty, so skipping the call caused
under-counting. Keep discount/margin/logging_obj guards (crash fix + optimization).
Skip function calls to get_cost_for_built_in_tools, _apply_cost_discount,
_apply_cost_margin, and _store_cost_breakdown_in_logging_obj when their
respective features are not configured. Reduces completion_cost() time
by ~20% (4.39s → 3.53s over 6K requests) for the common case where
built-in tools, discounts, margins, and logging object are not active.
The set of valid LLM API parameter names for logging was being rebuilt
on every request from 8 OpenAI SDK type annotations + set operations.
Since these are static TypedDict annotations that never change at
runtime, compute once at import time and store as a class-level
frozenset.
Line profiler: get_standard_logging_model_parameters() dropped from
774ms to 77ms across 12K calls (90% reduction, ~25µs/req saved).
- Pre-compute CallTypes enum values as module-level frozenset and dict map,
replacing per-request list comprehension (133µs → 0.4µs/call)
- Guard debug f-string with _is_debugging_on() to skip evaluation when off
- Cache update_response_metadata getattr lookup once per call instead of twice