Guard verbose_logger.debug() f-strings with isEnabledFor(logging.DEBUG)
checks in the router and cost calculation hot paths. Python evaluates
f-string arguments before the logging framework checks the log level,
causing expensive formatting on every request even with debug logging
disabled.
Changes:
- Remove redundant litellm_params.copy() in _completion/_acompletion
- Guard 5 debug logs in router.py (+ remove 1 duplicate log)
- Guard 6 debug logs in cost_calculator.py and utils.py
- get_model_info(): formatted 50+ field dict every call
- _apply_cost_margin(): called list(dict.keys()) every request
Profiled improvement: completion_cost 769µs → 637µs/call (-17.2%)
Tests that the fix using `getattr(type(self), method_name) is not getattr(CustomLogger, method_name)` correctly walks the full Method Resolution Order, unlike the buggy `method_name in type(self).__dict__` which only checks the immediate class.
Check if redact_standard_logging_payload_from_model_call_details was
overridden in a subclass before skipping redundant redaction. This
ensures custom callbacks with additional redaction logic still execute.
- Add global_redaction_applied flag to skip per-callback redaction
- Add should_redact param to avoid calling should_redact_message_logging twice
- Add tests for both flag=True (skip) and flag=False (proceed) cases
When turn_off_message_logging is enabled globally, skip per-callback
redaction functions that would re-process already-redacted data.
17% faster async_success_handler when global redaction is ON.
- Add _OPTIONAL_KWARGS_KEYS frozenset for O(1) lookups
- Replace 28 unconditional kwargs.get() calls with sparse extraction
- Only add kwargs keys that are actually present in the dict
- Simplify _get_base_model_from_litellm_call_metadata by removing redundant None checks
This reduces get_litellm_params() time by ~31% (743ms → 509ms across 6000 calls)
and Logging.__init__ total time by ~24% (1.61s → 1.23s).
Add fast path to check optional_params.get("api_base") directly before
constructing a full LiteLLM_Params Pydantic model. When api_base is present
(the common case via router), return it immediately — avoiding ~29µs of
Pydantic validation overhead per request.
Profiled: get_api_base 30.5µs/call → 1.2µs/call (-96%)
Revert the guard that skipped built-in tool cost when standard_built_in_tools_params
was falsy. get_cost_for_built_in_tools can return non-zero from response/usage alone
(e.g. web search) even when params is None/empty, so skipping the call caused
under-counting. Keep discount/margin/logging_obj guards (crash fix + optimization).
Skip function calls to get_cost_for_built_in_tools, _apply_cost_discount,
_apply_cost_margin, and _store_cost_breakdown_in_logging_obj when their
respective features are not configured. Reduces completion_cost() time
by ~20% (4.39s → 3.53s over 6K requests) for the common case where
built-in tools, discounts, margins, and logging object are not active.
The set of valid LLM API parameter names for logging was being rebuilt
on every request from 8 OpenAI SDK type annotations + set operations.
Since these are static TypedDict annotations that never change at
runtime, compute once at import time and store as a class-level
frozenset.
Line profiler: get_standard_logging_model_parameters() dropped from
774ms to 77ms across 12K calls (90% reduction, ~25µs/req saved).
- Pre-compute CallTypes enum values as module-level frozenset and dict map,
replacing per-request list comprehension (133µs → 0.4µs/call)
- Guard debug f-string with _is_debugging_on() to skip evaluation when off
- Cache update_response_metadata getattr lookup once per call instead of twice