- Add _OPTIONAL_KWARGS_KEYS frozenset for O(1) lookups
- Replace 28 unconditional kwargs.get() calls with sparse extraction
- Only add kwargs keys that are actually present in the dict
- Simplify _get_base_model_from_litellm_call_metadata by removing redundant None checks
This reduces get_litellm_params() time by ~31% (743ms → 509ms across 6000 calls)
and Logging.__init__ total time by ~24% (1.61s → 1.23s).
Add fast path to check optional_params.get("api_base") directly before
constructing a full LiteLLM_Params Pydantic model. When api_base is present
(the common case via router), return it immediately — avoiding ~29µs of
Pydantic validation overhead per request.
Profiled: get_api_base 30.5µs/call → 1.2µs/call (-96%)
Revert the guard that skipped built-in tool cost when standard_built_in_tools_params
was falsy. get_cost_for_built_in_tools can return non-zero from response/usage alone
(e.g. web search) even when params is None/empty, so skipping the call caused
under-counting. Keep discount/margin/logging_obj guards (crash fix + optimization).
Skip function calls to get_cost_for_built_in_tools, _apply_cost_discount,
_apply_cost_margin, and _store_cost_breakdown_in_logging_obj when their
respective features are not configured. Reduces completion_cost() time
by ~20% (4.39s → 3.53s over 6K requests) for the common case where
built-in tools, discounts, margins, and logging object are not active.
The set of valid LLM API parameter names for logging was being rebuilt
on every request from 8 OpenAI SDK type annotations + set operations.
Since these are static TypedDict annotations that never change at
runtime, compute once at import time and store as a class-level
frozenset.
Line profiler: get_standard_logging_model_parameters() dropped from
774ms to 77ms across 12K calls (90% reduction, ~25µs/req saved).
- Pre-compute CallTypes enum values as module-level frozenset and dict map,
replacing per-request list comprehension (133µs → 0.4µs/call)
- Guard debug f-string with _is_debugging_on() to skip evaluation when off
- Cache update_response_metadata getattr lookup once per call instead of twice