Commit graph

42355 commits

Author SHA1 Message Date
yuneng-jiang
c3c79b9d80 [Fix] Virtual Keys pagination displays stale totals when filtering
When applying a key_alias filter on the Virtual Keys page, the pagination display (total count, page count) showed stale unfiltered values. The bug was that the table component tracked total_count only from the unfiltered useKeys hook, not from the filtered API response returned by useFilterLogic.

Fixes: When selecting a key alias that matches 1 key, the UI previously showed "Showing 1-50 of 509 results" and "Page 1 of 11" instead of the correct "Showing 1-1 of 1 results" and "Page 1 of 1".

Changes:
- Added filteredTotalCount state to useFilterLogic to track the total_count from filtered API responses
- Updated VirtualKeysTable to use filteredTotalCount (when set) instead of always using the unfiltered total
- Added comprehensive tests to prevent regression of pagination display logic

Type: 🐛 Bug Fix, ✅ Test

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-26 12:28:43 -08:00
Chesars
5aff0e4da6 refactor: move _is_claude_4_6_model to AnthropicModelInfo, use explicit effort_map
- Move _is_claude_4_6_model from AnthropicConfig to AnthropicModelInfo
  to eliminate duplicated logic in is_effort_used
- Use explicit effort_map dict instead of passing unknown values through
  to output_config
2026-02-26 17:25:23 -03:00
Chesars
a8c95392fb fix(anthropic): map reasoning_effort to output_config for Claude 4.6 models
Claude 4.6 models use output_config as a stable API feature. This commit:
- Maps reasoning_effort to output_config for 4.6 models (minimal → low)
- Restricts effort="max" to Opus 4.6 only
- Skips beta header injection for 4.6 models
- Updates docs for Claude 4.6 effort support
2026-02-26 17:18:32 -03:00
Ryan Crabbe
0ee8cb5f02 refactor: remove shutdown cleanup, rely solely on startup wipe
Phase 2 (per-worker mark_process_dead on shutdown) only ever fired
when all workers shut down together, making it redundant — Phase 1
wipes everything on next startup anyway. This aligns with the
prometheus_client docs: just wipe the directory between runs.
2026-02-26 12:17:31 -08:00
yuneng-jiang
719b7fd013
Merge pull request #22157 from BerriAI/litellm_paginated_key_alias
[Feature] UI - Paginated Key Alias Select
2026-02-26 12:09:10 -08:00
yuneng-jiang
4b75a89673 Merge remote-tracking branch 'origin' into litellm_paginated_key_alias 2026-02-26 12:06:28 -08:00
Cesar Garcia
ec6a55c6db
Merge pull request #22208 from Chesars/fix/transcription-duration-hidden-params
fix(transcription): move duration to _hidden_params to match OpenAI response spec
2026-02-26 16:51:24 -03:00
Ishaan Jaff
cbdaaaeba4
[Feat] Add control for setting upperbound on chunk processing time (#22209)
* add LITELLM_MAX_STREAMING_DURATION_SECONDS

* add add LITELLM_MAX_STREAMING_DURATION_SECONDS

* fix: address Greptile review - rename constant, add sync check, add tests

- Rename MAX_STREAMING_CHUNK_DURATION_S → MAX_STREAMING_DURATION_S (misleading "CHUNK")
- Add _check_max_streaming_duration to SyncResponsesAPIStreamingIterator.__next__
- Add 8 unit tests covering both CustomStreamWrapper and ResponsesAPI paths
- Fix pre-existing pyright errors in streaming_handler.py

Made-with: Cursor

* add add LITELLM_MAX_STREAMING_DURATION_SECONDS
2026-02-26 11:39:09 -08:00
Cesar Garcia
b5a0a72acb
Merge pull request #21597 from Chesars/fix/gemini-json-schema-keep-refs
fix(gemini): preserve $ref in JSON Schema for Gemini 2.0+
2026-02-26 16:37:26 -03:00
Harshit28j
d3d11fb06e feat: add tags in project 2026-02-27 01:07:18 +05:30
Chesars
121669090d test: use real completion_cost() instead of duplicating inline logic 2026-02-26 16:24:48 -03:00
Chesars
5f957add18 fix(azure): apply same duration hidden_params fix to Azure transcription handler 2026-02-26 16:16:41 -03:00
Chesars
b6784e7d8c fix(types): annotate _FINISH_REASON_MAP and map_finish_reason with OpenAIChatCompletionFinishReason
Fixes mypy errors where dict[str, str] was incompatible with the
expected Literal type in get_finish_reason_mapping() and
_check_finish_reason() return types.
2026-02-26 16:13:23 -03:00
Chesars
7dd4f17021 fix(transcription): store duration in _hidden_params to avoid OpenAI SDK deserialization issues
LiteLLM was adding a `duration` field to audio transcription responses
for internal cost tracking. The OpenAI Python SDK uses "best match
deserialization" to determine the response type from present fields —
seeing `duration` caused it to incorrectly match plain Transcription
responses as TranscriptionVerbose/TranscriptionDiarized types.

Move the internally-calculated duration to `_hidden_params` so it
remains available for cost calculation without polluting the response
body. Provider-returned duration (e.g. from verbose_json format) is
still preserved in the response as expected.
2026-02-26 16:06:08 -03:00
Chesars
3196d40a04 fix(vertex): delegate Gemini finish reason mapping to centralized _FINISH_REASON_MAP
Addresses Greptile review feedback on PR #22138 — removes duplicated
Gemini finish reason dict in VertexGeminiConfig and delegates to the
shared map_finish_reason() to prevent the two mappings from drifting
apart.
2026-02-26 15:48:05 -03:00
yuneng-jiang
947586b62d Merge remote-tracking branch 'origin' into litellm_paginated_key_alias 2026-02-26 10:28:35 -08:00
yuneng-jiang
50bf2da05e
Merge pull request #22137 from BerriAI/litellm_key_info_crash_fix
[Fix] /key/aliases: Add pagination and search to prevent OOMs
2026-02-26 10:27:10 -08:00
Harshit Jain
4d2fab49a7
Merge pull request #22164 from Harshit28j/litellm_custom_auth_budget_fix
fix: custom auth budget issue
2026-02-26 23:38:38 +05:30
yuneng-jiang
857a324f7a Merge remote-tracking branch 'origin' into litellm_key_info_crash_fix 2026-02-26 10:04:54 -08:00
Ryan Crabbe
acfb5ade97 refactor: remove periodic dead PID cleanup, trim redundant tests
Startup wipe + graceful shutdown cleanup are sufficient. Remove
hourly mark_dead_pids scan, its helpers, and redundant test cases.
2026-02-26 10:03:51 -08:00
michelligabriele
ae13a40c01
test(mcp): add e2e test for stateless StreamableHTTP behavior (#22033)
Adds TestProxyMcpStatelessBehavior to test_proxy_mcp_e2e.py with a test
that verifies two independent MCP clients can connect, initialize, and
call tools without sharing session state. This catches the regression
from PR #19809 where stateless=False broke clients that don't manage
mcp-session-id headers.

Regression test for #20242
2026-02-26 09:30:03 -08:00
Sameer Kankute
1790a6bf82
Merge pull request #21604 from michelligabriele/fix/websearch-thinking-blocks
fix(websearch_interception): preserve thinking blocks in agentic loop follow-up messages
2026-02-26 22:26:38 +05:30
Julio Quinteros Pro
21d958e46d fix(mypy): add union-attr suppression to all backend_ws method calls
Address Greptile review: apply type: ignore[union-attr] consistently
on all backend_ws.send(), .recv(), and .close() calls, not just the
three that CI flagged.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 13:32:13 -03:00
Julio Quinteros Pro
494d783a80 fix(mypy): suppress attr-defined errors on websocket send/close calls
CI MyPy resolves CLIENT_CONNECTION_CLASS as Optional[ClientConnection]
and flags .send() and .close() as attr-defined errors. These methods
exist at runtime on the websocket connection object. Add type: ignore
comments to unblock the linting CI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 13:22:56 -03:00
Julio Quinteros Pro
4419c7a5f9
Merge pull request #22198 from BerriAI/fix/mcp-server-test-mock-mismatch
fix(tests): update MCP server test mocks to match production API
2026-02-26 13:16:40 -03:00
Julio Quinteros Pro
ace49b18d3 fix(tests): update MCP server test mocks to match production API
The tests were mocking `filter_server_ids_by_ip` but the production
code in server.py now calls `filter_server_ids_by_ip_with_info` which
returns a (server_ids, blocked_count) tuple. Update all 8 mock sites
to use the correct method name and return signature.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 13:11:57 -03:00
Julio Quinteros Pro
1c376afc85 fix(ci): use secrets context in ggshield step condition
Step-level env is not visible to the if condition — reference
secrets directly so ggshield actually runs when the key is configured.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 12:51:28 -03:00
Julio Quinteros Pro
05c3a95da8 fix(ci): add permissions block to secret-scan job
Address github-advanced-security bot review comment by setting explicit
minimal permissions (contents: read) for the GITHUB_TOKEN.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 12:48:43 -03:00
Julio Quinteros Pro
2fce35a162 test(ci): add secret scan test and CI job to prevent hardcoded credentials
- Add unit test that scans Python source for Base64 Basic Auth patterns
  that would be flagged by secret scanners like GitGuardian/ggshield
- Add secret-scan job to the linting CI workflow that runs the test on
  every PR and optionally runs ggshield if GITGUARDIAN_API_KEY is set

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 12:46:42 -03:00
Harshit Jain
43054a2390
fix: langfuse trace leak key on model params 2026-02-26 19:03:49 +05:30
Sameer Kankute
6600c86dbd
Merge pull request #22170 from BerriAI/litellm_fix_video_veo_vertex
Fix: Passing of image and parameters in videos api
2026-02-26 18:55:25 +05:30
Sameer Kankute
ec31845469
Merge pull request #22092 from BerriAI/litellm_fix_tts_vertex_ai
Add audio as supported openai param
2026-02-26 18:51:53 +05:30
Sameer Kankute
61d2f28545 Fix based on review 2026-02-26 18:49:25 +05:30
Sameer Kankute
16d6c279da
Merge pull request #22180 from BerriAI/litellm_fix_vllm_test
Add JSON exact match test for vLLM embeddings
2026-02-26 18:43:52 +05:30
Sameer Kankute
dc1e97d345
Merge pull request #22144 from BerriAI/litellm_allowed_openai_params_embeddings
fix(embeddings): allow dimensions param passthrough via allowed_openai_params for non-text-embedding-3 OpenAI models
2026-02-26 18:41:47 +05:30
Sameer Kankute
2772c88864
Merge pull request #22142 from BerriAI/litellm_fix_mcp_server_ip
Return Clear error message why no tools are available / IP Filtering occured
2026-02-26 18:41:22 +05:30
Sameer Kankute
e1df85ebea
Merge pull request #22166 from BerriAI/litellm_oss_staging_02_26_2026
Litellm oss staging 02 26 2026
2026-02-26 18:41:04 +05:30
Sameer Kankute
81455dbd57
Merge pull request #22187 from BerriAI/revert-22099-fix/improve-auth-exception-logging
Revert "fix(proxy): improve auth exception logging levels and add structured context"
2026-02-26 18:39:04 +05:30
Sameer Kankute
95b8fb823b
Revert "fix(proxy): improve auth exception logging levels and add structured …"
This reverts commit efeaf650aa.
2026-02-26 18:38:52 +05:30
Sameer Kankute
87b4fed967
Merge pull request #22186 from BerriAI/main
merge main
2026-02-26 18:21:38 +05:30
Sameer Kankute
27f9903765
Merge pull request #22184 from BerriAI/litellm_bump_litellm_26_02
Bump litellm version to 1.81.16
2026-02-26 18:19:10 +05:30
Sameer Kankute
678200ee48 Bump litellm version to 1.81.16 2026-02-26 18:18:03 +05:30
Sameer Kankute
6a68e3bba3 Add tests for messages to responses transformation: 2026-02-26 18:14:06 +05:30
shivam
ffb438f3f2 fix: clarify EXPERIMENTAL_UI_LOGIN ignores LITELLM_UI_SESSION_DURATION, add regression test 2026-02-26 04:41:07 -08:00
Sameer Kankute
7adaf49db7 Add tranlation of context_management 2026-02-26 18:05:20 +05:30
shivam
44557261a3 greptile issue 2026-02-26 04:29:42 -08:00
Shivam Rawat
d276d273c6
Update litellm/constants.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-26 04:21:35 -08:00
shivam
6974b6de68 resolved greptile comment 2026-02-26 04:13:52 -08:00
shivam
f4f834489c added variable for invitation link 2026-02-26 04:04:20 -08:00
shivam
a7f5163976 added the env flag 2026-02-26 03:49:54 -08:00