Commit graph

30508 commits

Author SHA1 Message Date
Ryan Crabbe
8216ec4389 perf: add LRU cache to normalize_request_route
Add @lru_cache(maxsize=256) to eliminate redundant regex work for
repeated routes. Reduces time from 1.04s to ~0s for 6,006 calls.
2026-01-31 15:27:28 -08:00
Ryan Crabbe
d1f428de32 test: add unit tests for get_litellm_params sparse kwargs extraction 2026-01-31 15:27:28 -08:00
Ryan Crabbe
ccc9927407 perf: Optimize get_litellm_params with sparse kwargs extraction
- Add _OPTIONAL_KWARGS_KEYS frozenset for O(1) lookups
- Replace 28 unconditional kwargs.get() calls with sparse extraction
- Only add kwargs keys that are actually present in the dict
- Simplify _get_base_model_from_litellm_call_metadata by removing redundant None checks

This reduces get_litellm_params() time by ~31% (743ms → 509ms across 6000 calls)
and Logging.__init__ total time by ~24% (1.61s → 1.23s).
2026-01-31 15:27:28 -08:00
Ryan Crabbe
227fb6e6c8 perf: skip Pydantic model construction in get_api_base when api_base is in dict
Add fast path to check optional_params.get("api_base") directly before
constructing a full LiteLLM_Params Pydantic model. When api_base is present
(the common case via router), return it immediately — avoiding ~29µs of
Pydantic validation overhead per request.

Profiled: get_api_base 30.5µs/call → 1.2µs/call (-96%)
2026-01-31 15:27:28 -08:00
Alexsander Hamir
535f721849 fix(cost): always call get_cost_for_built_in_tools to avoid under-counting
Revert the guard that skipped built-in tool cost when standard_built_in_tools_params
was falsy. get_cost_for_built_in_tools can return non-zero from response/usage alone
(e.g. web search) even when params is None/empty, so skipping the call caused
under-counting. Keep discount/margin/logging_obj guards (crash fix + optimization).
2026-01-31 15:27:28 -08:00
Ryan Crabbe
b8380f237d perf: add early-exit guards in completion_cost for unused features
Skip function calls to get_cost_for_built_in_tools, _apply_cost_discount,
_apply_cost_margin, and _store_cost_breakdown_in_logging_obj when their
respective features are not configured. Reduces completion_cost() time
by ~20% (4.39s → 3.53s over 6K requests) for the common case where
built-in tools, discounts, margins, and logging object are not active.
2026-01-31 15:27:28 -08:00
Ryan Crabbe
dc491e3ab1 test: add tests for cached ModelParamHelper logging args
Verify cached frozenset matches dynamic computation and that
prompt content keys (messages, prompt, input) are excluded from
logged model parameters.
2026-01-31 15:27:28 -08:00
Ryan Crabbe
a06359d40a perf: cache _get_relevant_args_to_use_for_logging() as module-level frozenset
The set of valid LLM API parameter names for logging was being rebuilt
on every request from 8 OpenAI SDK type annotations + set operations.
Since these are static TypedDict annotations that never change at
runtime, compute once at import time and store as a class-level
frozenset.

Line profiler: get_standard_logging_model_parameters() dropped from
774ms to 77ms across 12K calls (90% reduction, ~25µs/req saved).
2026-01-31 15:27:28 -08:00
Ryan Crabbe
cbc366f0d7 perf: optimize wrapper_async hot path with CallTypes caching and reduced lookups
- Pre-compute CallTypes enum values as module-level frozenset and dict map,
  replacing per-request list comprehension (133µs → 0.4µs/call)
- Guard debug f-string with _is_debugging_on() to skip evaluation when off
- Cache update_response_metadata getattr lookup once per call instead of twice
2026-01-31 15:27:27 -08:00
Ishaan Jaffer
534fa9f4c0 docs fix 2026-01-17 17:26:58 -08:00
Ishaan Jaffer
26497b415b docs fix 2026-01-17 17:21:31 -08:00
Ishaan Jaffer
c6998823c0 docs fix 2026-01-17 17:17:34 -08:00
Ishaan Jaffer
c158f83cff docs fix 2026-01-17 17:17:13 -08:00
Ishaan Jaffer
4610d1d43c docs fix 2026-01-17 17:16:22 -08:00
Ishaan Jaffer
7d24bbed42 qa fixes 2026-01-17 17:14:51 -08:00
Ishaan Jaffer
e15526a60e fix 2026-01-17 17:13:22 -08:00
Ishaan Jaffer
60dd04ac95 test_aiohttp_openai 2026-01-17 17:05:00 -08:00
Ishaan Jaffer
c30b17aa9b docs fix 2026-01-17 17:03:48 -08:00
Ishaan Jaffer
0a84120be5 v1.81.0 2026-01-17 16:43:00 -08:00
Ishaan Jaffer
db7de13818 test_deepseek_mock_completion 2026-01-17 16:36:42 -08:00
Ishaan Jaffer
5812654bdd test_router_fallbacks_with_custom_model_costs 2026-01-17 16:34:46 -08:00
Ishaan Jaff
1417b002a3
[Feat] Claude Code x LiteLLM WebSearch - QA Fixes to work with Claude Code (#19294)
* fix websearch_interception_converted_stream

* test_websearch_interception_no_tool_call_streaming

* FakeAnthropicMessagesStreamIterator

* LITELLM_WEB_SEARCH_TOOL_NAME

* fixes tools def for litellm web search

* fixes FakeAnthropicMessagesStreamIterator

* test_litellm_standard_websearch_tool

* use new hook for modfying before any transfroms from litellm

* init WebSearchInterceptionLogger + ARCHITECTURE

* fix config.yaml

* init doc for claude code web search

* docs fix

* doc fix

* fix mypy linting
2026-01-17 16:30:31 -08:00
yuneng-jiang
42c0136bf4
Merge pull request #19293 from BerriAI/litellm_ui_build_fix_22
[Infra] Fix UI Build
2026-01-17 15:58:24 -08:00
yuneng-jiang
91eb047761 testing adding entire out 2026-01-17 15:49:20 -08:00
yuneng-jiang
c1f194cde9 fix build attempt 2026-01-17 15:36:53 -08:00
YutaSaito
d28bf983eb
Merge pull request #19272 from Harshit28j/feature/panw-custom-violation-msg
feat(panw_prisma_airs): add custom violation message support
2026-01-18 06:55:39 +09:00
yuneng-jiang
953e2736d4
Merge pull request #19291 from BerriAI/deleted_keys_docs_2
[Docs] Deleted Keys and Teams Docs
2026-01-17 13:20:48 -08:00
yuneng-jiang
19a69a89f0 deleted keys and teams docs 2026-01-17 13:19:48 -08:00
Ishaan Jaffer
e238cb2ca0 docs clean up 2026-01-17 12:30:26 -08:00
Ishaan Jaffer
58570e5b13 fix doc 2026-01-17 12:28:04 -08:00
Ishaan Jaffer
1115e6b8c7 docs fix 2026-01-17 12:20:09 -08:00
Ishaan Jaffer
eb26ebc926 docs ui usage 2026-01-17 12:15:37 -08:00
yuneng-jiang
052aa4f7ed
Merge pull request #19287 from BerriAI/ui_build_jan17_1
[Infra] Fixing UI Build
2026-01-17 12:09:41 -08:00
yuneng-jiang
715fa8fa5f fixing ui build 2026-01-17 12:08:20 -08:00
yuneng-jiang
ff73d8f640
Merge pull request #19286 from BerriAI/8015_docs_1
[Docs] Updating Docs for v1.80.15-stable.1
2026-01-17 11:55:10 -08:00
yuneng-jiang
b964e2a1a0 updating docker pull cmd 2026-01-17 11:54:04 -08:00
yuneng-jiang
32626954c3
Merge pull request #19283 from BerriAI/ui_build_jan17
[Infra] Re-Building UI
2026-01-17 11:13:29 -08:00
yuneng-jiang
d160ac0319 rebuilding ui 2026-01-17 11:11:25 -08:00
yuneng-jiang
3619fdce31
Merge pull request #19282 from BerriAI/yj_ui_test_branch
[Fix] UI - Deleted Teams Endpoint Fix
2026-01-17 11:07:31 -08:00
yuneng-jiang
374662c60f deleted teams endpoint fix 2026-01-17 11:00:32 -08:00
Ishaan Jaffer
e9e323bf32 png fixes 2026-01-17 10:28:41 -08:00
Ishaan Jaffer
f0569bc102 docs fix 2026-01-17 09:40:08 -08:00
yuneng-jiang
0d3a9b1c87
Merge pull request #19279 from BerriAI/ui_build_branch
[Infra] Building UI
2026-01-17 09:20:28 -08:00
yuneng-jiang
581ba2def3 building ui 2026-01-17 09:19:05 -08:00
yuneng-jiang
23a06a04a3
Merge pull request #19278 from BerriAI/litellm_ui_update_new_badges
[Refactor] UI - Adjusting New Badges
2026-01-17 09:17:15 -08:00
yuneng-jiang
1301896e03 Adjusting new badges 2026-01-17 09:08:41 -08:00
Harshit Jain
0683f29671
feat(panw_prisma_airs): add custom violation message support 2026-01-17 17:52:25 +05:30
yuneng-jiang
362081c2b9
Merge pull request #19258 from BerriAI/litellm_ui_model_hub_health_1
[Feature] UI - Public Model Hub: Health Checks
2026-01-16 22:28:45 -08:00
yuneng-jiang
480d9356af
Merge pull request #19268 from BerriAI/litellm_ui_deleted_keys_teams_table
[Feature] UI - Logs: Deleted Keys and Teams Table
2026-01-16 22:28:24 -08:00
yuneng-jiang
d48e41bd94 fixing tests 2026-01-16 22:20:10 -08:00