add docs for spend logs (#10913)

* add docs for spend logs

* change docs

* change docs

* change docs and change default run loops.

There was a calc error previously, it was not a million when run 100 times, but 100k, now default run would delete 500k records in 50s (running 500 loops of function before exiting)

* add docs changes to ui logs to metion deletion
This commit is contained in:
Jugal D. Bhatt 2025-05-17 18:00:28 -05:00 • committed by GitHub
parent 9846abb5c0
commit b342aa9253
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
6 changed files with 127 additions and 2 deletions

View file

@ -0,0 +1,98 @@
---
id: spend_logs_deletion
title: Spend Logs Deletion
---
# Spend Log Cleanup
LiteLLM stores a log for every request. Over time, these logs can grow large and slow down your database. The Spend Log Cleanup feature helps manage database size by deleting old logs automatically.
---
## Usage
### Requirements
- **Postgres** (for log storage)
- **Redis** *(optional)* — required only if you're running multiple proxy instances and want to enable distributed locking
### Setup
Add this to your `proxy_config.yaml` under `general_settings`:
```yaml title="proxy_config.yaml"
general_settings:
maximum_spend_logs_retention_period: "7d" # Keep logs for 7 days
# Optional: set how frequently cleanup should run
maximum_spend_logs_retention_interval: "1d" # Run cleanup every day
litellm_settings:
cache: true
cache_params:
type: redis
```
### Configuration Options
#### `maximum_spend_logs_retention_period` (required)
How long logs should be kept before deletion. Supported formats:
- `"7d"` – 7 days
- `"24h"` – 24 hours
- `"60m"` – 60 minutes
- `"3600s"` – 3600 seconds
#### `maximum_spend_logs_retention_interval` (optional)
How often the cleanup job should run. Uses the same format as above. If not set, cleanup will run every 24 hours if and only if `maximum_spend_logs_retention_period` is set.
---
## How it works
### Step 1. Lock Acquisition (Optional with Redis)
If Redis is enabled, LiteLLM uses it to make sure only one instance runs the cleanup at a time.
- If the lock is acquired:
- This instance proceeds with cleanup
- Others skip it
- If no lock is present:
- Cleanup still runs (useful for single-node setups)
![Working of spend log deletions](../../img/spend_log_deletion_working.png)
*Working of spend log deletions*
---
### Step 2. Batch Deletion
Once cleanup starts:
- It calculates the cutoff date using the configured retention period
- Deletes logs older than the cutoff in **batches of 1000**
- Adds a short delay between batches to avoid overloading the database
### Default settings:
- **Batch size**: 1000 logs
- **Max batches per run**: 500
- **Max deletions per run**: 500,000 logs
You can change the number of batches using an environment variable:
```bash
SPEND_LOG_RUN_LOOPS=200
```
This would allow up to 200,000 logs to be deleted in one run.
![Batch deletion of old logs](../../img/spend_log_deletion_multi_pod.jpg)
*Batch deletion of old logs*
---
## Summary
Spend Log Cleanup helps keep your database fast by regularly deleting old logs. It’s safe, customizable, and works well for both single-node and multi-node deployments.

View file

@ -52,3 +52,30 @@ If you do not want to store spend logs in DB, you can opt out with this setting
general_settings:
disable_spend_logs: True # Disable writing spend logs to DB
```
## Automatically Deleting Old Spend Logs
If you're storing spend logs, it might be a good idea to delete them regularly to keep the database fast.
LiteLLM lets you configure this in your `proxy_config.yaml`:
```yaml
general_settings:
maximum_spend_logs_retention_period: "7d" # Delete logs older than 7 days
# Optional: how often to run cleanup
maximum_spend_logs_retention_interval: "1d" # Run once per day
```
You can control how many logs are deleted per run using this environment variable:
`SPEND_LOG_RUN_LOOPS=200 # Deletes up to 200,000 logs in one run (batch size = 1000)`
For detailed architecture and how it works, see [Spend Logs Deletion](../proxy/spend_logs_deletion).

Binary file not shown.

After

Width:  |  Height:  |  Size: 189 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 151 KiB

View file

@ -54,7 +54,7 @@ const sidebars = {
{
type: "category",
label: "Architecture",
items: ["proxy/architecture", "proxy/db_info", "proxy/db_deadlocks", "router_architecture", "proxy/user_management_heirarchy", "proxy/jwt_auth_arch", "proxy/image_handling"],
items: ["proxy/architecture", "proxy/db_info", "proxy/db_deadlocks", "router_architecture", "proxy/user_management_heirarchy", "proxy/jwt_auth_arch", "proxy/image_handling", "proxy/spend_logs_deletion"],
},
{
type: "link",

View file

@ -618,7 +618,7 @@ LITELLM_PROXY_ADMIN_NAME = "default_user_id"
DB_SPEND_UPDATE_JOB_NAME = "db_spend_update_job"
PROMETHEUS_EMIT_BUDGET_METRICS_JOB_NAME = "prometheus_emit_budget_metrics"
SPEND_LOG_CLEANUP_JOB_NAME = "spend_log_cleanup"
SPEND_LOG_RUN_LOOPS = int(os.getenv("SPEND_LOG_RUN_LOOPS", 100))
SPEND_LOG_RUN_LOOPS = int(os.getenv("SPEND_LOG_RUN_LOOPS", 500))
DEFAULT_CRON_JOB_LOCK_TTL_SECONDS = int(
os.getenv("DEFAULT_CRON_JOB_LOCK_TTL_SECONDS", 60)
) # 1 minute