docs: add performance improvement section to v1.81.0 release notes

- Add 'Performance - 25% CPU Usage Reduction' section
- Document removal of premature model.dump() calls from hot path
- Add to Key Highlights section
This commit is contained in:
Alexsander Hamir 2026-01-17 17:57:52 -08:00
parent 7eecf81cdc
commit aad7199906

View file

@ -47,6 +47,7 @@ pip install litellm==1.81.0
- **Claude Code** - Support for using web search across Bedrock, Vertex AI, and all LiteLLM providers
- **Major Change** - [50MB limit on image URL downloads](#major-change---chatcompletions-image-url-download-size-limit) to improve reliability
- **Performance** - [25% CPU Usage Reduction](#performance---25-cpu-usage-reduction) by preventing unbounded queue growth in GCS Bucket logging
---
@ -140,6 +141,13 @@ This feature improves reliability by:
---
## Performance - 25% CPU Usage Reduction
LiteLLM now reduces CPU usage by removing premature `model.dump()` calls from the hot path in request processing. Previously, Pydantic model serialization was performed earlier and more frequently than necessary, causing unnecessary CPU overhead on every request. By deferring serialization until it is actually needed, LiteLLM reduces CPU usage and improves request throughput under high load.
---
## New Models / Updated Models
#### New Model Support