Commit graph

30515 commits

Author SHA1 Message Date
Ryan Crabbe
f51870811d perf: optimize initialize_standard_built_in_tools_params with early exit and single-pass extraction
- Add early exit when kwargs lacks web_search_options and tools keys
- Add get_built_in_tools_from_kwargs() for single-pass tool extraction
- Avoids calling _get_web_search_options() and _get_file_search_tool_call()
  separately (which each iterated tools list)

Performance improvement:
- No-tools workloads: 83.9% faster (117.9ms → 19.0ms per 6000 calls)
- With-tools workloads: 14.7% faster (240.9ms → 205.5ms per 6000 calls)
2026-01-31 15:27:29 -08:00
Ryan Crabbe
eb264c802c perf: guard debug log f-strings and remove redundant dict copy in hot path
Guard verbose_logger.debug() f-strings with isEnabledFor(logging.DEBUG)
checks in the router and cost calculation hot paths. Python evaluates
f-string arguments before the logging framework checks the log level,
causing expensive formatting on every request even with debug logging
disabled.

Changes:
- Remove redundant litellm_params.copy() in _completion/_acompletion
- Guard 5 debug logs in router.py (+ remove 1 duplicate log)
- Guard 6 debug logs in cost_calculator.py and utils.py
  - get_model_info(): formatted 50+ field dict every call
  - _apply_cost_margin(): called list(dict.keys()) every request

Profiled improvement: completion_cost 769µs → 637µs/call (-17.2%)
2026-01-31 15:27:29 -08:00
Ryan Crabbe
33440858aa Add unit test for MRO-aware method override detection
Tests that the fix using `getattr(type(self), method_name) is not getattr(CustomLogger, method_name)` correctly walks the full Method Resolution Order, unlike the buggy `method_name in type(self).__dict__` which only checks the immediate class.
2026-01-31 15:27:29 -08:00
Ryan Crabbe
31eaa70111 fix: Check full MRO for method override detection 2026-01-31 15:27:29 -08:00
Ryan Crabbe
8788882b44 fix: Add safeguard to allow overridden redaction methods to run
Check if redact_standard_logging_payload_from_model_call_details was
overridden in a subclass before skipping redundant redaction. This
ensures custom callbacks with additional redaction logic still execute.
2026-01-31 15:27:28 -08:00
Ryan Crabbe
49c6afe64f perf: skip redundant redaction + avoid double check
- Add global_redaction_applied flag to skip per-callback redaction
- Add should_redact param to avoid calling should_redact_message_logging twice
- Add tests for both flag=True (skip) and flag=False (proceed) cases
2026-01-31 15:27:28 -08:00
Ryan Crabbe
e3b4ef7cad perf: skip redundant redaction when global redaction enabled
When turn_off_message_logging is enabled globally, skip per-callback
redaction functions that would re-process already-redacted data.

17% faster async_success_handler when global redaction is ON.
2026-01-31 15:27:28 -08:00
Ryan Crabbe
8216ec4389 perf: add LRU cache to normalize_request_route
Add @lru_cache(maxsize=256) to eliminate redundant regex work for
repeated routes. Reduces time from 1.04s to ~0s for 6,006 calls.
2026-01-31 15:27:28 -08:00
Ryan Crabbe
d1f428de32 test: add unit tests for get_litellm_params sparse kwargs extraction 2026-01-31 15:27:28 -08:00
Ryan Crabbe
ccc9927407 perf: Optimize get_litellm_params with sparse kwargs extraction
- Add _OPTIONAL_KWARGS_KEYS frozenset for O(1) lookups
- Replace 28 unconditional kwargs.get() calls with sparse extraction
- Only add kwargs keys that are actually present in the dict
- Simplify _get_base_model_from_litellm_call_metadata by removing redundant None checks

This reduces get_litellm_params() time by ~31% (743ms → 509ms across 6000 calls)
and Logging.__init__ total time by ~24% (1.61s → 1.23s).
2026-01-31 15:27:28 -08:00
Ryan Crabbe
227fb6e6c8 perf: skip Pydantic model construction in get_api_base when api_base is in dict
Add fast path to check optional_params.get("api_base") directly before
constructing a full LiteLLM_Params Pydantic model. When api_base is present
(the common case via router), return it immediately — avoiding ~29µs of
Pydantic validation overhead per request.

Profiled: get_api_base 30.5µs/call → 1.2µs/call (-96%)
2026-01-31 15:27:28 -08:00
Alexsander Hamir
535f721849 fix(cost): always call get_cost_for_built_in_tools to avoid under-counting
Revert the guard that skipped built-in tool cost when standard_built_in_tools_params
was falsy. get_cost_for_built_in_tools can return non-zero from response/usage alone
(e.g. web search) even when params is None/empty, so skipping the call caused
under-counting. Keep discount/margin/logging_obj guards (crash fix + optimization).
2026-01-31 15:27:28 -08:00
Ryan Crabbe
b8380f237d perf: add early-exit guards in completion_cost for unused features
Skip function calls to get_cost_for_built_in_tools, _apply_cost_discount,
_apply_cost_margin, and _store_cost_breakdown_in_logging_obj when their
respective features are not configured. Reduces completion_cost() time
by ~20% (4.39s → 3.53s over 6K requests) for the common case where
built-in tools, discounts, margins, and logging object are not active.
2026-01-31 15:27:28 -08:00
Ryan Crabbe
dc491e3ab1 test: add tests for cached ModelParamHelper logging args
Verify cached frozenset matches dynamic computation and that
prompt content keys (messages, prompt, input) are excluded from
logged model parameters.
2026-01-31 15:27:28 -08:00
Ryan Crabbe
a06359d40a perf: cache _get_relevant_args_to_use_for_logging() as module-level frozenset
The set of valid LLM API parameter names for logging was being rebuilt
on every request from 8 OpenAI SDK type annotations + set operations.
Since these are static TypedDict annotations that never change at
runtime, compute once at import time and store as a class-level
frozenset.

Line profiler: get_standard_logging_model_parameters() dropped from
774ms to 77ms across 12K calls (90% reduction, ~25µs/req saved).
2026-01-31 15:27:28 -08:00
Ryan Crabbe
cbc366f0d7 perf: optimize wrapper_async hot path with CallTypes caching and reduced lookups
- Pre-compute CallTypes enum values as module-level frozenset and dict map,
  replacing per-request list comprehension (133µs → 0.4µs/call)
- Guard debug f-string with _is_debugging_on() to skip evaluation when off
- Cache update_response_metadata getattr lookup once per call instead of twice
2026-01-31 15:27:27 -08:00
Ishaan Jaffer
534fa9f4c0 docs fix 2026-01-17 17:26:58 -08:00
Ishaan Jaffer
26497b415b docs fix 2026-01-17 17:21:31 -08:00
Ishaan Jaffer
c6998823c0 docs fix 2026-01-17 17:17:34 -08:00
Ishaan Jaffer
c158f83cff docs fix 2026-01-17 17:17:13 -08:00
Ishaan Jaffer
4610d1d43c docs fix 2026-01-17 17:16:22 -08:00
Ishaan Jaffer
7d24bbed42 qa fixes 2026-01-17 17:14:51 -08:00
Ishaan Jaffer
e15526a60e fix 2026-01-17 17:13:22 -08:00
Ishaan Jaffer
60dd04ac95 test_aiohttp_openai 2026-01-17 17:05:00 -08:00
Ishaan Jaffer
c30b17aa9b docs fix 2026-01-17 17:03:48 -08:00
Ishaan Jaffer
0a84120be5 v1.81.0 2026-01-17 16:43:00 -08:00
Ishaan Jaffer
db7de13818 test_deepseek_mock_completion 2026-01-17 16:36:42 -08:00
Ishaan Jaffer
5812654bdd test_router_fallbacks_with_custom_model_costs 2026-01-17 16:34:46 -08:00
Ishaan Jaff
1417b002a3
[Feat] Claude Code x LiteLLM WebSearch - QA Fixes to work with Claude Code (#19294)
* fix websearch_interception_converted_stream

* test_websearch_interception_no_tool_call_streaming

* FakeAnthropicMessagesStreamIterator

* LITELLM_WEB_SEARCH_TOOL_NAME

* fixes tools def for litellm web search

* fixes FakeAnthropicMessagesStreamIterator

* test_litellm_standard_websearch_tool

* use new hook for modfying before any transfroms from litellm

* init WebSearchInterceptionLogger + ARCHITECTURE

* fix config.yaml

* init doc for claude code web search

* docs fix

* doc fix

* fix mypy linting
2026-01-17 16:30:31 -08:00
yuneng-jiang
42c0136bf4
Merge pull request #19293 from BerriAI/litellm_ui_build_fix_22
[Infra] Fix UI Build
2026-01-17 15:58:24 -08:00
yuneng-jiang
91eb047761 testing adding entire out 2026-01-17 15:49:20 -08:00
yuneng-jiang
c1f194cde9 fix build attempt 2026-01-17 15:36:53 -08:00
YutaSaito
d28bf983eb
Merge pull request #19272 from Harshit28j/feature/panw-custom-violation-msg
feat(panw_prisma_airs): add custom violation message support
2026-01-18 06:55:39 +09:00
yuneng-jiang
953e2736d4
Merge pull request #19291 from BerriAI/deleted_keys_docs_2
[Docs] Deleted Keys and Teams Docs
2026-01-17 13:20:48 -08:00
yuneng-jiang
19a69a89f0 deleted keys and teams docs 2026-01-17 13:19:48 -08:00
Ishaan Jaffer
e238cb2ca0 docs clean up 2026-01-17 12:30:26 -08:00
Ishaan Jaffer
58570e5b13 fix doc 2026-01-17 12:28:04 -08:00
Ishaan Jaffer
1115e6b8c7 docs fix 2026-01-17 12:20:09 -08:00
Ishaan Jaffer
eb26ebc926 docs ui usage 2026-01-17 12:15:37 -08:00
yuneng-jiang
052aa4f7ed
Merge pull request #19287 from BerriAI/ui_build_jan17_1
[Infra] Fixing UI Build
2026-01-17 12:09:41 -08:00
yuneng-jiang
715fa8fa5f fixing ui build 2026-01-17 12:08:20 -08:00
yuneng-jiang
ff73d8f640
Merge pull request #19286 from BerriAI/8015_docs_1
[Docs] Updating Docs for v1.80.15-stable.1
2026-01-17 11:55:10 -08:00
yuneng-jiang
b964e2a1a0 updating docker pull cmd 2026-01-17 11:54:04 -08:00
yuneng-jiang
32626954c3
Merge pull request #19283 from BerriAI/ui_build_jan17
[Infra] Re-Building UI
2026-01-17 11:13:29 -08:00
yuneng-jiang
d160ac0319 rebuilding ui 2026-01-17 11:11:25 -08:00
yuneng-jiang
3619fdce31
Merge pull request #19282 from BerriAI/yj_ui_test_branch
[Fix] UI - Deleted Teams Endpoint Fix
2026-01-17 11:07:31 -08:00
yuneng-jiang
374662c60f deleted teams endpoint fix 2026-01-17 11:00:32 -08:00
Ishaan Jaffer
e9e323bf32 png fixes 2026-01-17 10:28:41 -08:00
Ishaan Jaffer
f0569bc102 docs fix 2026-01-17 09:40:08 -08:00
yuneng-jiang
0d3a9b1c87
Merge pull request #19279 from BerriAI/ui_build_branch
[Infra] Building UI
2026-01-17 09:20:28 -08:00