Commit graph

33599 commits

Author SHA1 Message Date
Chesars
c4458c09fe fix(count_tokens): include system and tools in token counting API requests
The /v1/messages/count_tokens proxy endpoint was only passing `messages`
to provider token counting APIs, discarding `system` and `tools`. This
caused clients like Claude Code to receive artificially low token counts
(e.g. 10 instead of 531), preventing proper context window management
and leading to context overflow errors.

Pass system and tools through the full chain:
- TokenCountRequest → proxy_server → provider counters → API handlers
- Bedrock: transform tools to toolConfig format, system to text blocks
- Anthropic/Azure AI: pass through directly (same API format)
2026-02-27 15:39:35 -03:00
Ishaan Jaff
adba088df2
Realtime API: spend log storage, playground UI, tools logging, and guardrail support (#22105)
Backend - Spend Log Storage for Realtime Calls:
- Collect user voice transcripts and text input during WebSocket sessions
- Store collected messages in spend logs when store_prompts_in_spend_logs enabled
- Capture tool definitions from session.update and tool calls from response.done
- Enrich proxy_server_request with tools and response with tool_calls for UI

Backend - WebSocket Auth:
- Support browser-based auth via Sec-WebSocket-Protocol subprotocol
- Echo back subprotocol on WebSocket accept

UI - Realtime Playground:
- New RealtimePlayground component with WebSocket voice+text chat
- Mic recording (PCM16 24kHz), server VAD, audio playback, text input
- Handle binary WebSocket frames (Blob/ArrayBuffer decoding)
- Add /v1/realtime endpoint option to playground endpoint selector

UI - Tools Section for Realtime Logs:
- Extract tool calls from realtime response format (response.tool_calls
  and response.results[].response.output[].type=function_call)

Tests:
- 15 new backend tests for realtime streaming and spend log storage
- 4 new UI tests for realtime tool call extraction

Fixes pre-existing build errors:
- ToolPolicies.tsx: duplicate import, antd styles type
- create_key_button.tsx: missing message import

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-02-25 14:55:27 -08:00
Ishaan Jaff
cc85fe5921
Proxy request tags docs (#22129)
* docs: document x-litellm-tags header and request body tags parameter

- Add documentation for x-litellm-tags header (comma-separated or array)
- Add documentation for tags in request body
- Clarify that dynamic tags override config tags

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: consolidate tag documentation and improve cross-references

- Make request_tags.md the single source of truth for all tag options
- Add cross-reference from cost_tracking.md to request_tags.md
- Document both direct tags and metadata.tags formats
- Add key/team tag setup and custom header tracking to request_tags.md
- Reduce duplication and make navigation clearer

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: use generic examples instead of specific company names

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* docs: clarify x-litellm-tags header format is comma-separated string

HTTP headers are always strings, not arrays. Remove misleading
array format documentation for the header parameter.

Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>

* Update docs/my-website/docs/proxy/request_tags.md

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-25 14:29:00 -08:00
Julio Quinteros Pro
54e4af84d6
fix(auth): remove hardcoded base64 string flagged by secret scanner (#22125)
Replace literal Base64-encoded example value in docstring with a
placeholder to prevent GitGuardian/ggshield from flagging it as a
leaked Basic Authentication credential in container scans.

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 14:18:36 -08:00
ryan-crabbe
b9dd36c14b
Merge pull request #22028 from BerriAI/litellm_batch_update_database_tasks
perf(proxy): batch 11 create_task() calls into 1 in update_database()
2026-02-25 12:16:58 -08:00
Ryan Crabbe
1c1b9cda42 merge: resolve conflict with origin/main in test_db_spend_update_writer.py
Keep both batch update tests (from this branch) and pipeline test (from main).
2026-02-25 12:15:53 -08:00
Ryan Crabbe
e1e05d27f8 Merge remote-tracking branch 'origin/main' into litellm_batch_update_database_tasks 2026-02-25 12:12:10 -08:00
ryan-crabbe
9857643eb4
Merge pull request #22044 from ryan-crabbe/litellm_redis_pipeline_spend_updates
Litellm redis pipeline spend updates
2026-02-25 12:11:09 -08:00
github-actions[bot]
6e10428e00
chore: regenerate poetry.lock to match pyproject.toml (#22120)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-02-25 20:10:37 +00:00
yuneng-jiang
243a207b1e
Merge pull request #22066 from BerriAI/litellm_spend_log_duration
[Feature] Add request_duration_ms to SpendLogs
2026-02-25 12:09:31 -08:00
yuneng-jiang
db05a5f025 adding build artifacts 2026-02-25 12:08:41 -08:00
yuneng-jiang
1132176289 bump: version 0.4.47 → 0.4.48 2026-02-25 12:08:15 -08:00
yuneng-jiang
9d6f02e8b7 Merge remote-tracking branch 'origin' into litellm_spend_log_duration 2026-02-25 12:06:19 -08:00
Krish Dholakia
12c4876891
Agents - assign tools (#22064)
* feat(proxy): add max_iterations limiter for agent session loops (#22058)

Adds a new proxy hook that enforces a per-session cap on the number of
LLM calls an agentic loop can make. Callers send a session_id with each
request, and the hook counts calls per session, returning 429 when the
configured max_iterations limit is exceeded.

- Uses Redis Lua script for atomic increment (multi-instance safe)
- Falls back to in-memory cache when Redis unavailable
- Follows parallel_request_limiter_v3 pattern
- Configurable via key metadata: {"max_iterations": 25}
- Session counters auto-expire via TTL (default 1hr)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add new code execution dataset

* feat(agent_endpoints/): allow giving agents keys

* fix: ui fixes

* feat: allow assigning mcp servers to agents

* fix: eliminate duplicate DB queries in MCP agent auth and N+1 in agent listing (#22110)

- Extract _get_agent_object_permission helper so _get_allowed_mcp_servers_for_agent
  and _get_agent_tool_permissions_for_server share a single DB fetch instead of
  each independently querying the same agent row (was 1+N queries per MCP request)
- Use include={"object_permission": True} on find_many in get_all_agents_from_db
  to eagerly load permissions in one query instead of N+1
- Use include={"object_permission": True} on create/update/find_unique in all
  agent CRUD operations, removing attach_object_permission_to_dict follow-up calls

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 11:44:30 -08:00
yuneng-jiang
a4baa022c5
Merge pull request #22116 from BerriAI/ui_build_fix_feb25
[Infra] Ui Build and Test Fixes
2026-02-25 11:31:45 -08:00
yuneng-jiang
30e8151288 fixing build and tests 2026-02-25 11:30:09 -08:00
yuneng-jiang
edd00c025c
Merge pull request #22112 from BerriAI/litellm_credential_tag_docs
[Docs] Credential Usage Tracking
2026-02-25 10:55:52 -08:00
yuneng-jiang
7daeaf8106 [Docs] Add Credential Usage Tracking documentation
Add new document explaining automatic credential usage tracking and tagging. When models use reusable credentials, LiteLLM automatically injects a Credential: <name> tag on requests, enabling credential-level spend tracking on the Usage page with no additional configuration.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-25 10:52:24 -08:00
yuneng-jiang
b9d3fdf93e Merge remote-tracking branch 'origin' into ui_build_fix_feb25 2026-02-25 10:51:50 -08:00
yuneng-jiang
cd41b49061 fixing build 2026-02-25 10:51:42 -08:00
yuneng-jiang
243771f99a
Merge pull request #22108 from BerriAI/litellm_pricing_calc_tests
[Test] UI - Pricing Calculator: Add comprehensive unit tests
2026-02-25 10:51:12 -08:00
Krish Dholakia
5662228e20
feat(ui): add user filtering to usage page (#22059)
* feat(ui): add user filtering to usage page

Adds "User Usage" as a new view option in the usage page dropdown,
allowing admins to view and filter usage data by individual users
via the existing /user/daily/activity backend endpoint.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(ui/): working usage filtering

* fix(ui): use single-select for user filter and add tests

The user entity type's backend endpoint only accepts a single user_id,
so the filter now uses single-select mode instead of multi-select.
Added tests for the new user entity type in EntityUsage and
UsageViewSelect. Updated CLAUDE.md and AGENTS.md with guidance on
UI/backend contract consistency and test coverage for new entity types.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* revert: remove unintended package-lock.json changes

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* revert: restore package-lock.json to merge base state

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 10:37:45 -08:00
yuneng-jiang
a38993b871
Merge pull request #22107 from BerriAI/litellm_fix_ui_duplicate_import_v2
fix(ui): remove duplicate antd import in ToolPolicies
2026-02-25 10:25:56 -08:00
yuneng-jiang
6ae7e84a0b [Test] UI - Pricing Calculator: Add comprehensive unit tests
Added unit tests for all pricing calculator components with 64 passing tests across 5 test files:
- multi_export_utils.test.ts (16 tests for PDF/CSV export functions)
- use_multi_cost_estimate.test.ts (15 tests for cost estimation hook)
- multi_export_dropdown.test.tsx (8 tests for export dropdown component)
- multi_cost_results.test.tsx (15 tests for results display and UI states)
- index.test.tsx (10 tests for main calculator component)

Also fixed missing page description for tool-policies page in page_metadata.ts.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-25 10:25:25 -08:00
Shin
115b9a2d29 fix(ui): remove duplicate antd import in ToolPolicies
Duplicate import was causing UI build to fail:
- Line 4: import { Select, Switch, Tooltip } from 'antd'
- Line 5: import { Select, Tooltip } from 'antd' (duplicate)

Removed the duplicate line 5.
2026-02-25 18:09:17 +00:00
Sameer Kankute
47c24ef8ae
Merge pull request #22094 from BerriAI/litellm_bump_openai_25_02
[chore] bump OpenAI package version
2026-02-25 18:46:10 +05:30
Sameer Kankute
05165779ac
Merge pull request #22081 from BerriAI/litellm_fix_bedrock_batch_health_check
Fix: Test connect failing for bedrock batches mode
2026-02-25 18:46:01 +05:30
Sameer Kankute
ebe35aad9a
Merge pull request #22080 from BerriAI/litellm_vllm_response_fix
Fix None (TypeError: 'NoneType' object is not a mapping)
2026-02-25 18:45:54 +05:30
Sameer Kankute
ec3ae25a3a
Merge pull request #22070 from BerriAI/litellm_forward_auth_headers
[Feat]Add forward auth headers of provider
2026-02-25 18:45:38 +05:30
Harshit Jain
fbffa9a215
Merge pull request #22095 from Harshit28j/litellm_fix_RCE_issues
fix Unauthenticated RCE and Sandbox Escape in Custom Code Guardrail
2026-02-25 18:23:04 +05:30
Harshit28j
c3e32b450e Address CR feedback: fix auth dependency duplication, correct logger wording, clean imports 2026-02-25 18:02:12 +05:30
Harshit28j
e50b4486d0 fix Unauthenticated RCE and Sandbox Escape in Custom Code Guardrail 2026-02-25 17:57:32 +05:30
Sameer Kankute
ec9ec6b771 remove comment 2026-02-25 17:40:46 +05:30
Sameer Kankute
955b85015f Bump openai package 2026-02-25 17:40:41 +05:30
Harshit Jain
d2aeb3e513
Merge pull request #22084 from Harshit28j/litellm_presidio-non-json-response-handling
fix(guardrails): prevent presidio crash on non-json responses
2026-02-25 16:37:39 +05:30
Harshit28j
6a8052295b fix(guardrails): prevent presidio crash on non-json responses 2026-02-25 16:11:24 +05:30
Harshit Jain
66c49dbb9c
Merge pull request #22077 from Harshit28j/litellm_gemini_trace_id_missingv2
Litellm gemini trace id missingv2
2026-02-25 15:09:42 +05:30
Sameer Kankute
61f9222156 Fix: Test connect failing for bedrock batches mode 2026-02-25 15:03:27 +05:30
Sameer Kankute
45c87a714e Fix None (TypeError: 'NoneType' object is not a mapping) 2026-02-25 14:39:12 +05:30
Harshit Jain
23e84eb789 fix: Metadata / Trace ID Missing in S3 Streaming Callbacks 2026-02-25 14:16:42 +05:30
Harshit Jain
897b472011 fix: Metadata / Trace ID Missing in S3 Streaming Callbacks 2026-02-25 14:16:21 +05:30
Sameer Kankute
83afa310d6 Fix clean header logger 2026-02-25 12:31:35 +05:30
Sameer Kankute
0e806c83c1 Fix docs 2026-02-25 12:13:56 +05:30
Sameer Kankute
a43d6139c7 Fix docs 2026-02-25 12:13:07 +05:30
Sameer Kankute
1f8f66de69 add docs for Authentication Headers forwarding 2026-02-25 12:10:12 +05:30
Sameer Kankute
886f9d6a70 Add support for forwarding provider's auth headers 2026-02-25 12:08:25 +05:30
yuneng-jiang
b78a30f773 [Feature] Add request_duration_ms to SpendLogs
Add a `request_duration_ms` column to `LiteLLM_SpendLogs` to track request
duration. New rows are computed at write time. Legacy rows use a COALESCE
fallback in the `/spend/logs/ui` query to compute duration on the fly from
`endTime - startTime`. The field is also sortable in the UI endpoint.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-02-24 21:04:53 -08:00
Krish Dholakia
00ab4d2067
feat: add new code execution dataset (#22065) 2026-02-24 20:58:23 -08:00
yuneng-jiang
ec7dc44e73
Merge pull request #22049 from BerriAI/litellm_spend_log_key_info
[Fix] Enrich Failure Spend Logs With Key/Team Metadata
2026-02-24 20:56:51 -08:00
shin-bot-litellm
b2c0bec5c5
fix(test): Update status enum values to match Google Interactions OpenAPI spec (#22061)
Google's Interactions API spec changed the status enum:
- Values are now lowercase (was uppercase)
- 'UNSPECIFIED' value was removed

Updated test to match the current spec from:
https://ai.google.dev/static/api/interactions.openapi.json
2026-02-24 20:26:11 -08:00