From aad7199906741b19be8700773d0d2312270ef325 Mon Sep 17 00:00:00 2001 From: Alexsander Hamir Date: Sat, 17 Jan 2026 17:57:52 -0800 Subject: [PATCH] docs: add performance improvement section to v1.81.0 release notes - Add 'Performance - 25% CPU Usage Reduction' section - Document removal of premature model.dump() calls from hot path - Add to Key Highlights section --- docs/my-website/release_notes/v1.81.0/index.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/docs/my-website/release_notes/v1.81.0/index.md b/docs/my-website/release_notes/v1.81.0/index.md index 071422f96be..afd65eaade7 100644 --- a/docs/my-website/release_notes/v1.81.0/index.md +++ b/docs/my-website/release_notes/v1.81.0/index.md @@ -47,6 +47,7 @@ pip install litellm==1.81.0 - **Claude Code** - Support for using web search across Bedrock, Vertex AI, and all LiteLLM providers - **Major Change** - [50MB limit on image URL downloads](#major-change---chatcompletions-image-url-download-size-limit) to improve reliability +- **Performance** - [25% CPU Usage Reduction](#performance---25-cpu-usage-reduction) by preventing unbounded queue growth in GCS Bucket logging --- @@ -140,6 +141,13 @@ This feature improves reliability by: --- +## Performance - 25% CPU Usage Reduction + + +LiteLLM now reduces CPU usage by removing premature `model.dump()` calls from the hot path in request processing. Previously, Pydantic model serialization was performed earlier and more frequently than necessary, causing unnecessary CPU overhead on every request. By deferring serialization until it is actually needed, LiteLLM reduces CPU usage and improves request throughput under high load. + +--- + ## New Models / Updated Models #### New Model Support