Commit graph

33599 commits

Author SHA1 Message Date
Yuneng Jiang
5aa115611c
chore: fixes 2026-04-04 23:03:03 -07:00
Cursor Agent
d919b1048e fix: reduce memory retention after traffic spikes
Two-pronged fix for the reported memory leak where RSS grows from ~1.9 GiB
to ~3.5 GiB under load (1000-1500 concurrent users) and never returns to
baseline:

1. Logging._cleanup_after_logging(): Clears large payload fields (input,
   original_response, raw_request_typed_dict) and streaming chunk lists from
   the Logging object after all success/failure callbacks have consumed them.
   This prevents per-request data (which can be kilobytes per request) from
   being held in memory by async task references.

2. _periodic_memory_cleanup(): A scheduled background job (every 60s by
   default, configurable via LITELLM_MEMORY_CLEANUP_INTERVAL) that runs
   gc.collect() + malloc_trim(0) on Linux. Python's pymalloc allocator does
   not return freed memory arenas to the OS by default; malloc_trim forces
   glibc to release unused heap pages, allowing RSS to shrink after traffic
   spikes subside.

Together these ensure that:
- Per-request memory is released promptly after logging completes
- The OS reclaims freed heap memory periodically, so RSS drops during idle
  periods instead of staying at the peak level indefinitely

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-02-25 23:05:12 +00:00
Ishaan Jaff
cc85fe5921
Proxy request tags docs (#22129)
* docs: document x-litellm-tags header and request body tags parameter

- Add documentation for x-litellm-tags header (comma-separated or array)
- Add documentation for tags in request body
- Clarify that dynamic tags override config tags

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: consolidate tag documentation and improve cross-references

- Make request_tags.md the single source of truth for all tag options
- Add cross-reference from cost_tracking.md to request_tags.md
- Document both direct tags and metadata.tags formats
- Add key/team tag setup and custom header tracking to request_tags.md
- Reduce duplication and make navigation clearer

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: use generic examples instead of specific company names

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: clarify x-litellm-tags header format is comma-separated string

HTTP headers are always strings, not arrays. Remove misleading
array format documentation for the header parameter.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* Update docs/my-website/docs/proxy/request_tags.md

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-25 14:29:00 -08:00
Julio Quinteros Pro
54e4af84d6
fix(auth): remove hardcoded base64 string flagged by secret scanner (#22125)
Replace literal Base64-encoded example value in docstring with a
placeholder to prevent GitGuardian/ggshield from flagging it as a
leaked Basic Authentication credential in container scans.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 14:18:36 -08:00
ryan-crabbe
b9dd36c14b
Merge pull request #22028 from BerriAI/litellm_batch_update_database_tasks
perf(proxy): batch 11 create_task() calls into 1 in update_database()
2026-02-25 12:16:58 -08:00
Ryan Crabbe
1c1b9cda42 merge: resolve conflict with origin/main in test_db_spend_update_writer.py
Keep both batch update tests (from this branch) and pipeline test (from main).
2026-02-25 12:15:53 -08:00
Ryan Crabbe
e1e05d27f8 Merge remote-tracking branch 'origin/main' into litellm_batch_update_database_tasks 2026-02-25 12:12:10 -08:00
ryan-crabbe
9857643eb4
Merge pull request #22044 from ryan-crabbe/litellm_redis_pipeline_spend_updates
Litellm redis pipeline spend updates
2026-02-25 12:11:09 -08:00
github-actions[bot]
6e10428e00
chore: regenerate poetry.lock to match pyproject.toml (#22120)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-02-25 20:10:37 +00:00
yuneng-jiang
243a207b1e
Merge pull request #22066 from BerriAI/litellm_spend_log_duration
[Feature] Add request_duration_ms to SpendLogs
2026-02-25 12:09:31 -08:00
yuneng-jiang
db05a5f025 adding build artifacts 2026-02-25 12:08:41 -08:00
yuneng-jiang
1132176289 bump: version 0.4.47 → 0.4.48 2026-02-25 12:08:15 -08:00
yuneng-jiang
9d6f02e8b7 Merge remote-tracking branch 'origin' into litellm_spend_log_duration 2026-02-25 12:06:19 -08:00
Krish Dholakia
12c4876891
Agents - assign tools (#22064)
* feat(proxy): add max_iterations limiter for agent session loops (#22058)

Adds a new proxy hook that enforces a per-session cap on the number of
LLM calls an agentic loop can make. Callers send a session_id with each
request, and the hook counts calls per session, returning 429 when the
configured max_iterations limit is exceeded.

- Uses Redis Lua script for atomic increment (multi-instance safe)
- Falls back to in-memory cache when Redis unavailable
- Follows parallel_request_limiter_v3 pattern
- Configurable via key metadata: {"max_iterations": 25}
- Session counters auto-expire via TTL (default 1hr)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add new code execution dataset

* feat(agent_endpoints/): allow giving agents keys

* fix: ui fixes

* feat: allow assigning mcp servers to agents

* fix: eliminate duplicate DB queries in MCP agent auth and N+1 in agent listing (#22110)

- Extract _get_agent_object_permission helper so _get_allowed_mcp_servers_for_agent
  and _get_agent_tool_permissions_for_server share a single DB fetch instead of
  each independently querying the same agent row (was 1+N queries per MCP request)
- Use include={"object_permission": True} on find_many in get_all_agents_from_db
  to eagerly load permissions in one query instead of N+1
- Use include={"object_permission": True} on create/update/find_unique in all
  agent CRUD operations, removing attach_object_permission_to_dict follow-up calls

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 11:44:30 -08:00
yuneng-jiang
a4baa022c5
Merge pull request #22116 from BerriAI/ui_build_fix_feb25
[Infra] Ui Build and Test Fixes
2026-02-25 11:31:45 -08:00
yuneng-jiang
30e8151288 fixing build and tests 2026-02-25 11:30:09 -08:00
yuneng-jiang
edd00c025c
Merge pull request #22112 from BerriAI/litellm_credential_tag_docs
[Docs] Credential Usage Tracking
2026-02-25 10:55:52 -08:00
yuneng-jiang
7daeaf8106 [Docs] Add Credential Usage Tracking documentation
Add new document explaining automatic credential usage tracking and tagging. When models use reusable credentials, LiteLLM automatically injects a Credential: <name> tag on requests, enabling credential-level spend tracking on the Usage page with no additional configuration.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-25 10:52:24 -08:00
yuneng-jiang
b9d3fdf93e Merge remote-tracking branch 'origin' into ui_build_fix_feb25 2026-02-25 10:51:50 -08:00
yuneng-jiang
cd41b49061 fixing build 2026-02-25 10:51:42 -08:00
yuneng-jiang
243771f99a
Merge pull request #22108 from BerriAI/litellm_pricing_calc_tests
[Test] UI - Pricing Calculator: Add comprehensive unit tests
2026-02-25 10:51:12 -08:00
Krish Dholakia
5662228e20
feat(ui): add user filtering to usage page (#22059)
* feat(ui): add user filtering to usage page

Adds "User Usage" as a new view option in the usage page dropdown,
allowing admins to view and filter usage data by individual users
via the existing /user/daily/activity backend endpoint.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(ui/): working usage filtering

* fix(ui): use single-select for user filter and add tests

The user entity type's backend endpoint only accepts a single user_id,
so the filter now uses single-select mode instead of multi-select.
Added tests for the new user entity type in EntityUsage and
UsageViewSelect. Updated CLAUDE.md and AGENTS.md with guidance on
UI/backend contract consistency and test coverage for new entity types.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* revert: remove unintended package-lock.json changes

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* revert: restore package-lock.json to merge base state

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 10:37:45 -08:00
yuneng-jiang
a38993b871
Merge pull request #22107 from BerriAI/litellm_fix_ui_duplicate_import_v2
fix(ui): remove duplicate antd import in ToolPolicies
2026-02-25 10:25:56 -08:00
yuneng-jiang
6ae7e84a0b [Test] UI - Pricing Calculator: Add comprehensive unit tests
Added unit tests for all pricing calculator components with 64 passing tests across 5 test files:
- multi_export_utils.test.ts (16 tests for PDF/CSV export functions)
- use_multi_cost_estimate.test.ts (15 tests for cost estimation hook)
- multi_export_dropdown.test.tsx (8 tests for export dropdown component)
- multi_cost_results.test.tsx (15 tests for results display and UI states)
- index.test.tsx (10 tests for main calculator component)

Also fixed missing page description for tool-policies page in page_metadata.ts.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-25 10:25:25 -08:00
Shin
115b9a2d29 fix(ui): remove duplicate antd import in ToolPolicies
Duplicate import was causing UI build to fail:
- Line 4: import { Select, Switch, Tooltip } from 'antd'
- Line 5: import { Select, Tooltip } from 'antd' (duplicate)

Removed the duplicate line 5.
2026-02-25 18:09:17 +00:00
Sameer Kankute
47c24ef8ae
Merge pull request #22094 from BerriAI/litellm_bump_openai_25_02
[chore] bump OpenAI package version
2026-02-25 18:46:10 +05:30
Sameer Kankute
05165779ac
Merge pull request #22081 from BerriAI/litellm_fix_bedrock_batch_health_check
Fix: Test connect failing for bedrock batches mode
2026-02-25 18:46:01 +05:30
Sameer Kankute
ebe35aad9a
Merge pull request #22080 from BerriAI/litellm_vllm_response_fix
Fix None (TypeError: 'NoneType' object is not a mapping)
2026-02-25 18:45:54 +05:30
Sameer Kankute
ec3ae25a3a
Merge pull request #22070 from BerriAI/litellm_forward_auth_headers
[Feat]Add forward auth headers of provider
2026-02-25 18:45:38 +05:30
Harshit Jain
fbffa9a215
Merge pull request #22095 from Harshit28j/litellm_fix_RCE_issues
fix Unauthenticated RCE and Sandbox Escape in Custom Code Guardrail
2026-02-25 18:23:04 +05:30
Harshit28j
c3e32b450e Address CR feedback: fix auth dependency duplication, correct logger wording, clean imports 2026-02-25 18:02:12 +05:30
Harshit28j
e50b4486d0 fix Unauthenticated RCE and Sandbox Escape in Custom Code Guardrail 2026-02-25 17:57:32 +05:30
Sameer Kankute
ec9ec6b771 remove comment 2026-02-25 17:40:46 +05:30
Sameer Kankute
955b85015f Bump openai package 2026-02-25 17:40:41 +05:30
Harshit Jain
d2aeb3e513
Merge pull request #22084 from Harshit28j/litellm_presidio-non-json-response-handling
fix(guardrails): prevent presidio crash on non-json responses
2026-02-25 16:37:39 +05:30
Harshit28j
6a8052295b fix(guardrails): prevent presidio crash on non-json responses 2026-02-25 16:11:24 +05:30
Harshit Jain
66c49dbb9c
Merge pull request #22077 from Harshit28j/litellm_gemini_trace_id_missingv2
Litellm gemini trace id missingv2
2026-02-25 15:09:42 +05:30
Sameer Kankute
61f9222156 Fix: Test connect failing for bedrock batches mode 2026-02-25 15:03:27 +05:30
Sameer Kankute
45c87a714e Fix None (TypeError: 'NoneType' object is not a mapping) 2026-02-25 14:39:12 +05:30
Harshit Jain
23e84eb789 fix: Metadata / Trace ID Missing in S3 Streaming Callbacks 2026-02-25 14:16:42 +05:30
Harshit Jain
897b472011 fix: Metadata / Trace ID Missing in S3 Streaming Callbacks 2026-02-25 14:16:21 +05:30
Sameer Kankute
83afa310d6 Fix clean header logger 2026-02-25 12:31:35 +05:30
Sameer Kankute
0e806c83c1 Fix docs 2026-02-25 12:13:56 +05:30
Sameer Kankute
a43d6139c7 Fix docs 2026-02-25 12:13:07 +05:30
Sameer Kankute
1f8f66de69 add docs for Authentication Headers forwarding 2026-02-25 12:10:12 +05:30
Sameer Kankute
886f9d6a70 Add support for forwarding provider's auth headers 2026-02-25 12:08:25 +05:30
yuneng-jiang
b78a30f773 [Feature] Add request_duration_ms to SpendLogs
Add a `request_duration_ms` column to `LiteLLM_SpendLogs` to track request
duration. New rows are computed at write time. Legacy rows use a COALESCE
fallback in the `/spend/logs/ui` query to compute duration on the fly from
`endTime - startTime`. The field is also sortable in the UI endpoint.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-02-24 21:04:53 -08:00
Krish Dholakia
00ab4d2067
feat: add new code execution dataset (#22065) 2026-02-24 20:58:23 -08:00
yuneng-jiang
ec7dc44e73
Merge pull request #22049 from BerriAI/litellm_spend_log_key_info
[Fix] Enrich Failure Spend Logs With Key/Team Metadata
2026-02-24 20:56:51 -08:00
shin-bot-litellm
b2c0bec5c5
fix(test): Update status enum values to match Google Interactions OpenAPI spec (#22061)
Google's Interactions API spec changed the status enum:
- Values are now lowercase (was uppercase)
- 'UNSPECIFIED' value was removed

Updated test to match the current spec from:
https://ai.google.dev/static/api/interactions.openapi.json
2026-02-24 20:26:11 -08:00