From 914ab0080594299591d64fecefb7a23c3ac9931f Mon Sep 17 00:00:00 2001 From: Krrish Dholakia Date: Sat, 3 May 2025 22:04:22 -0700 Subject: [PATCH] docs(index.md): add key highlights to docs --- .../release_notes/v1.68.0-stable/index.md | 29 ++++++++++++++++++- 1 file changed, 28 insertions(+), 1 deletion(-) diff --git a/docs/my-website/release_notes/v1.68.0-stable/index.md b/docs/my-website/release_notes/v1.68.0-stable/index.md index 0b2e7793327..eb30853e9b9 100644 --- a/docs/my-website/release_notes/v1.68.0-stable/index.md +++ b/docs/my-website/release_notes/v1.68.0-stable/index.md @@ -41,7 +41,16 @@ pip install litellm==1.68.0.post1 -## Bedrock Vector Stores +## Key Highlights + +LiteLLM v1.68.0-stable will be live soon. Here are the key highlights of this release: + +- **Bedrock Knowledge Base**: You can now call query your Bedrock Knowledge Base with all LiteLLM models via `/chat/completion` or `/responses` API. +- **Rate Limits**: This release brings accurate rate limiting across multiple instances, reducing spillover to at most 10 additional requests in high traffic. +- **Meta Llama API**: Added support for Meta Llama API [Get Started](https://docs.litellm.ai/docs/providers/meta_llama) +- **LlamaFile**: Added support for LlamaFile [Get Started](https://docs.litellm.ai/docs/providers/llamafile) + +## Bedrock Knowledge Base (Vector Store)
@@ -57,6 +66,24 @@ For the next release we plan on allowing you to set key, user, team, org permiss [Read more here](https://docs.litellm.ai/docs/completion/knowledgebase) +## Rate Limiting + +This release brings accurate multi-instance rate limiting across keys/users/teams. Outlining key engineering changes below: + +- **Change**: Instances now increment cache value instead of setting it. To avoid calling Redis on each request, this is synced every 0.01s. +- **Accuracy**: In testing, we saw a maximum spill over from expected of 10 requests, in high traffic (100 RPS, 3 instances), vs. current 189 request spillover +- **Performance**: Our load tests show this to reduce median response time by 100ms in high traffic  + +This is currently behind a feature flag, and we plan to have this be the default by next week. To enable this today, just add this environment variable: + +``` +export LITELLM_RATE_LIMIT_ACCURACY=true +``` + +[Read more here](../../docs/proxy/users#beta-multi-instance-rate-limiting) + + + ## New Models / Updated Models - **Gemini ([VertexAI](https://docs.litellm.ai/docs/providers/vertex#usage-with-litellm-proxy-server) + [Google AI Studio](https://docs.litellm.ai/docs/providers/gemini))** - Handle more json schema - openapi schema conversion edge cases [PR](https://github.com/BerriAI/litellm/pull/10351)