Commit graph

42355 commits

Author SHA1 Message Date
Emerson Gomes
72682f4bd4
fix(proxy): avoid in-place mutation in SpendUpdateQueue aggregation (#20876)
* fix(proxy): prevent spend queue aggregation from mutating input updates

* test(proxy): avoid order-dependent spend queue aggregation assertion
2026-02-10 22:36:03 -08:00
pb
713d3022ae
fix(scheduler): remove orphan entries from queue - causing memory leak. (#20866)
* fix(scheduler): remove timed-out requests from queue to prevent memory leak

Fixes #20059

* fix(scheduler): use actual model param instead of hardcoded gpt-3.5-turbo in schedule_acompletion

* trigger CLA recheck

---------

Co-authored-by: Piyush Bhawsar <piyush100x@Piyushs-MacBook-Pro-3.local>
2026-02-10 22:34:52 -08:00
shin-bot-litellm
8c13001eb1
fix(mcp): use anyio.fail_after instead of asyncio.wait_for for StreamableHTTP backends (#20891)
Replace asyncio.wait_for() with anyio.fail_after() in _fetch_tools_with_timeout()
to fix conflict with MCP SDK's anyio TaskGroup that causes 0 tools to be returned
for external StreamableHTTP MCP backends.

Root cause: asyncio.wait_for() wrapping anyio-managed code causes inconsistent
CancelledError propagation, resulting in false cancellations even when the
operation hasn't timed out.

Fixes #20715
2026-02-10 22:20:37 -08:00
shin-bot-litellm
c7a22921dc
feat: add standard_logging_payload_excluded_fields config option (#20831)
Adds a new config option to exclude specific fields from StandardLoggingPayload
before any callback receives it. This provides a general approach to control
what data is logged across ALL integrations (S3, GCS, Datadog, etc.).

## Changes

1. **litellm/__init__.py**: Added new global setting
   `standard_logging_payload_excluded_fields: Optional[List[str]] = None`

2. **litellm/integrations/custom_logger.py**: Modified
   `redact_standard_logging_payload_from_model_call_details()` to:
   - Remove specified fields entirely from the StandardLoggingPayload
   - Works alongside existing `turn_off_message_logging` feature
   - Excluded fields take precedence (removed rather than redacted)

3. **tests/**: Added comprehensive test suite with 17 tests covering:
   - Single/multiple field exclusion
   - Interaction with turn_off_message_logging
   - Original payload immutability
   - Config loading via setattr (proxy pattern)
   - Edge cases (empty list, non-existent fields, None standard_logging_object)

## Usage

```yaml
litellm_settings:
  success_callback: ["s3"]
  standard_logging_payload_excluded_fields: ["response", "messages"]
```

This removes the `response` and `messages` fields from logs before any
callback processes them, reducing log size and improving privacy compliance.

## Available Fields

The fields match StandardLoggingPayload TypedDict keys including:
- messages, response (large payload fields)
- metadata, hidden_params, model_parameters
- error_str, error_information
- And all other StandardLoggingPayload fields

Closes the need for per-integration flags like `s3_log_response`.
2026-02-10 22:16:41 -08:00
shin-bot-litellm
d9c69ae9e5
docs: add Greptile review requirement to PR template (#20762) 2026-02-10 22:08:04 -08:00
The Mavik
6b4db0caeb
fix: stop leaking Python tracebacks in streaming SSE error responses (#20850)
When a guardrail (e.g. Zscaler AI Guard) or other exception is raised
during streaming, the error handler in `async_data_generator` includes
the full Python traceback in the SSE response sent to clients:

    error_msg = f"{str(e)}\n\n{traceback.format_exc()}"

This leaks internal server details (file paths, line numbers, call
stacks) to end users. The traceback is already logged server-side via
`verbose_proxy_logger.exception()`, so including it in the client
response is unnecessary.

Change to only include the exception message (`str(e)`) in the SSE
error payload, consistent with how `StreamingCallbackError` is already
handled.

Fixes #20610
2026-02-10 22:04:33 -08:00
BlueT - Matthew Lien - 練喆明
c0de6c5c6c
[Fix] handle metadata=None in SDK path retry/error logic (utils.py) (#20873)
* [Fix] handle metadata=None in SDK path retry/error logic (utils.py)

Fixes #20871

Same class of bug as #9717 (fixed by #9764 for the proxy path).
The SDK path in utils.py has the same fragile pattern at 7 locations.

Replace `kwargs.get("metadata", {})` with `(kwargs.get("metadata") or {})`
to handle the case where metadata key exists with value None (e.g. from
Azure OpenAI streaming responses).

This is consistent with the existing correct pattern at line 602:
`metadata = kwargs.get("metadata") or {}`

Adds TestMetadataNoneHandling with 6 unit tests in test_utils.py.

* fix: remove duplicate PerplexityResponsesConfig key in lazy imports registry

Removes duplicate dictionary key added in commit be0ebb15 (PR #20860).
The entry at line 1042 is identical to the existing entry at line 906.
This causes ruff F601 lint failure on all PRs targeting main.
2026-02-10 22:03:33 -08:00
yuneng-jiang
d050404393
Merge pull request #20910 from BerriAI/litellm_ui_hide_usage_modal
[Feature] UI - Navbar: Option to hide Usage Popup
2026-02-10 20:07:40 -08:00
yuneng-jiang
dc8934cf96
Merge pull request #20908 from BerriAI/litellm_ui_login_sso_redir
[Feature] UI - Login: New Login With SSO Button
2026-02-10 20:07:24 -08:00
Sameer Kankute
bb53e9dd2e
Merge pull request #20548 from kelvin-tran/kt/anthropic-opus-4-6-structured-outputs
feat: enable support for non-tool structured outputs on Anthropic Claude Opus 4.5 and 4.6 (use `output_format` param)
2026-02-11 09:22:57 +05:30
Ishaan Jaffer
f29165561f fix linting 2026-02-10 19:28:04 -08:00
ken
2913db783e
feat: add dashscope/qwen3-max model with tiered pricing (#20919)
Add support for Alibaba Cloud's Qwen3-Max model with:
- 258K input tokens, 65K output tokens
- Tiered pricing based on context window usage (0-32K, 32K-128K, 128K-252K)
- Function calling and tool choice support
- Reasoning capabilities enabled

Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-10 18:53:32 -08:00
Ishaan Jaff
3407006120
[Docs] Add docs guide for using policies (#20914)
* init schema with TAGS

* ui: add policy test

* resolvePoliciesCall

* add_policy_sources_to_metadata + headers

* types Policy

* preview Impact

* def _describe_match_reason(

* match based on TAGs

* TestTagBasedAttachments

* test fixes

* add policy_resolve_router

* add_guardrails_from_policy_engine

* TestMatchAttribution

* refactor

* fix

* fix: address Greptile review feedback on policy resolve endpoints

- Track unnamed keys/teams as separate counts instead of inflating
  affected_keys_count with duplicate "(unnamed key)" placeholders.
  Added unnamed_keys_count and unnamed_teams_count to response.
- Push alias pattern matching to DB via _build_alias_where() which
  converts exact patterns to Prisma "in" and suffix wildcards to
  "startsWith" filters.
- Gate sync_policies_from_db/sync_attachments_from_db behind
  force_sync query param (default false) to avoid 2 DB round-trips
  on every /policies/resolve request.
- Remove worktree-only conftest.py that cleared sys.modules at import
  time — no longer needed since code moved to main repo.
- Rename MAX_ESTIMATE_IMPACT_ROWS → MAX_POLICY_ESTIMATE_IMPACT_ROWS.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: eliminate duplicate DB queries and fix header delimiter ambiguity

- Fetch teams table once in estimate_attachment_impact and reuse for
  both tag-based and alias-based lookups (was querying teams twice when
  both tag_patterns and team_patterns were provided).
- Convert tag/team filter functions from async DB queries to sync
  filters that operate on pre-fetched data (_filter_keys_by_tags,
  _filter_teams_by_tags).
- Fix comma ambiguity in x-litellm-policy-sources header: use '; '
  as entry delimiter since matched_via values can contain commas.
- Use '+' as the within-value separator in matched_via reason strings
  (e.g. "tag:healthcare+team:health-team") to avoid conflict with
  header delimiters.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs v1 guide with UI imgs

* docs fix

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-10 18:52:31 -08:00
Alejandro Tapia
aad90ede43 fallback-display: updated fallback display table to use arrows and card structure for better visibility. 2026-02-10 18:48:00 -08:00
Harshit Jain
bc0622692d
fix: type error & better error handling (#20689) 2026-02-10 18:33:57 -08:00
Ishaan Jaff
f83620157e
[Feat] Policies - Allow connecting Policies to Tags, Simulating Policies, Viewing how many keys, teams it applies on (#20904)
* init schema with TAGS

* ui: add policy test

* resolvePoliciesCall

* add_policy_sources_to_metadata + headers

* types Policy

* preview Impact

* def _describe_match_reason(

* match based on TAGs

* TestTagBasedAttachments

* test fixes

* add policy_resolve_router

* add_guardrails_from_policy_engine

* TestMatchAttribution

* refactor

* fix

* fix: address Greptile review feedback on policy resolve endpoints

- Track unnamed keys/teams as separate counts instead of inflating
  affected_keys_count with duplicate "(unnamed key)" placeholders.
  Added unnamed_keys_count and unnamed_teams_count to response.
- Push alias pattern matching to DB via _build_alias_where() which
  converts exact patterns to Prisma "in" and suffix wildcards to
  "startsWith" filters.
- Gate sync_policies_from_db/sync_attachments_from_db behind
  force_sync query param (default false) to avoid 2 DB round-trips
  on every /policies/resolve request.
- Remove worktree-only conftest.py that cleared sys.modules at import
  time — no longer needed since code moved to main repo.
- Rename MAX_ESTIMATE_IMPACT_ROWS → MAX_POLICY_ESTIMATE_IMPACT_ROWS.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: eliminate duplicate DB queries and fix header delimiter ambiguity

- Fetch teams table once in estimate_attachment_impact and reuse for
  both tag-based and alias-based lookups (was querying teams twice when
  both tag_patterns and team_patterns were provided).
- Convert tag/team filter functions from async DB queries to sync
  filters that operate on pre-fetched data (_filter_keys_by_tags,
  _filter_teams_by_tags).
- Fix comma ambiguity in x-litellm-policy-sources header: use '; '
  as entry delimiter since matched_via values can contain commas.
- Use '+' as the within-value separator in matched_via reason strings
  (e.g. "tag:healthcare+team:health-team") to avoid conflict with
  header delimiters.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* Update litellm/proxy/policy_engine/policy_resolve_endpoints.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-10 17:50:37 -08:00
Alexsander Hamir
b7993b14cf
Add semgrep & Fix OOMs (#20912) 2026-02-10 17:50:14 -08:00
yuneng-jiang
9a418443d1 Add banner notifying of breaking change 2026-02-10 17:37:06 -08:00
shin-bot-litellm
26d561081b
fix(cloudzero): update CBF field mappings per LIT-1907 (#20906)
* fix(cloudzero): update CBF field mappings per LIT-1907

Phase 1 field updates for CloudZero integration:

ADD/UPDATE:
- resource/account: Send concat(api_key_alias, '|', api_key_prefix)
- resource/service: Send model_group instead of service_type
- resource/usage_family: Send provider instead of hardcoded 'llm-usage'
- action/operation: NEW - Send team_id
- resource/id: Send model name instead of CZRN
- resource/tag:organization_alias: Add if exists
- resource/tag:project_alias: Add if exists
- resource/tag:user_alias: Add if exists

REMOVE:
- resource/tag:total_tokens: Removed
- resource/tag:team_id: Removed (team_id now in action/operation)

Fixes LIT-1907

* Update litellm/integrations/cloudzero/transform.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* fix: define api_key_alias variable, update CBFRecord docstring

- Fix F821 lint error: api_key_alias was used but not defined
- Update CBFRecord docstring to reflect LIT-1907 field mappings
- Remove unused Optional import

---------

Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-10 17:33:16 -08:00
yuneng-jiang
97957ae9a3 option to hide usage indicator 2026-02-10 17:21:41 -08:00
Julio Quinteros Pro
80a3e072be fix(ui): remove duplicate URL in tagsSpendLogsCall query string
The template literal in the tags query parameter concatenation included
`${url}` inside a `+=` assignment, causing the full URL to be doubled.

Supersedes #8793 (original fix by @Mte90, now stale).

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-10 22:13:00 -03:00
yuneng-jiang
e86d7f59c6 new login with sso button in login page 2026-02-10 17:04:52 -08:00
yuneng-jiang
c68383068a
Merge pull request #20803 from BerriAI/litellm_ui_e2e_02
[Infra] UI - E2E Tests: Key Delete, Regenerate, and Update TPM/RPM Limits
2026-02-10 17:02:37 -08:00
yuneng-jiang
a8b4b4ccba
Merge pull request #20894 from BerriAI/litellm_ui_chart_fix_02
[Fix] UI - Usage: Request Chart stack variant
2026-02-10 16:32:04 -08:00
yuneng-jiang
2730e91356
Merge pull request #20898 from BerriAI/litellm_config_pt_endpoints
[Feature] Include Config Defined Pass Through Endpoints
2026-02-10 16:31:36 -08:00
yuneng-jiang
df37bc1900
Merge pull request #20796 from BerriAI/litellm_guardrail_list_sec
[Fix] /v2/guardrails/list Returns Sensitive Values
2026-02-10 16:31:16 -08:00
yuneng-jiang
cfd261d679 Split e2e ui testing for browser 2026-02-10 16:30:31 -08:00
Alexsander Hamir
ebce0e5f8c
[Release - 02/10/2026] v1.81.10-nightly 2026-02-10 16:26:30 -08:00
michelligabriele
8507df483c
fix(router): propagate model-level tags from config to SpendLogs (#20769) 2026-02-10 15:52:52 -08:00
yuneng-jiang
39bf5b780b addressing comments 2026-02-10 15:29:07 -08:00
Ishaan Jaffer
f311fba194 fix 2026-02-10 15:24:46 -08:00
Ishaan Jaff
f8619e2000
[Stability] Investigate + fix issue where model cost map became poorly formatted (#20895)
* init: GetModelCostMap

* fix

* docs

* docs fix

* docs fixes

* docs fix

* test model cost map resilience

* MODEL_COST_MAP_MIN_MODEL_COUNT

* validate_model_cost_map

* test_should_have_minimum_models_in_backup

* docs fix

* docs fix

* fix

* dos fix

* docs fix

* docs fix

* docs fix

* docs fix

* validate_model_cost_map

* fix

* cleanup
2026-02-10 15:17:01 -08:00
yuneng-jiang
e002d6afe8 addressing comments 2026-02-10 15:16:18 -08:00
Krish Dholakia
10d891a365
Guardrails - add logging to all unified_guardrails + link to custom code guardrail templates (#20900)
* feat(guardrail_hooks/): add guardrail logging to all unified guardrails

ensures unified guardrails use the 'log_guardrail_information' decorator for logging

* fix(custom_guardrail.py): don't log inputs on guardrail response - just emit state

* refactor: don't double log bedrock guardrail information

* feat: add in-product nudges for contributing + trying community custom code guardrails

allows users to contribute / share custom code guardrails
2026-02-10 15:13:54 -08:00
yuneng-jiang
fc0563fab3 get pass through include config defined pass through 2026-02-10 14:55:37 -08:00
Emerson Gomes
a6f90586ac
feat(model-db): add azure_ai/kimi-k2.5 pricing entry (#20896) 2026-02-10 14:46:40 -08:00
yuneng-jiang
79b24c8f25 remove stack from charts 2026-02-10 14:11:55 -08:00
yuneng-jiang
9f8878ee17
Merge pull request #20893 from BerriAI/pypi_fix_feb10
[Infra] CI/CD - Fix PyPI CI Step
2026-02-10 14:01:05 -08:00
yuneng-jiang
7d2c874434 Fixing ci pypi build 2026-02-10 13:59:48 -08:00
yuneng-jiang
ea38630e7c bump: version 0.4.33 → 0.4.34 2026-02-10 13:58:57 -08:00
yuneng-jiang
b7107ab803
Merge pull request #20892 from BerriAI/litellm_ui_spend_logs_model
[Feature] UI - Spend Logs: Paginated Searchable Model Select
2026-02-10 13:51:53 -08:00
yuneng-jiang
3fe1c1ba24 fixing build 2026-02-10 12:53:51 -08:00
yuneng-jiang
7fd8c0e160 Searchable Paginated Model Select For Spend Logs 2026-02-10 12:44:38 -08:00
michelligabriele
3bbc25a3f0
fix(aiohttp): respect ssl_verify with shared sessions (#20349)
* fix(aiohttp): respect ssl_verify with shared sessions

* fix(aiohttp): resolve mypy error for ssl parameter type

Pass ssl kwarg conditionally to aiohttp request() only when explicitly
configured, since None is not a valid value for the ssl parameter
(expected SSLContext | bool | Fingerprint).
2026-02-10 10:17:35 -08:00
Ryan Crabbe
19122ed271 perf: optimize model_dump_with_preserved_fields (15.9x faster)
Replace 3 full tree traversals with 1 Pydantic dump + 3 dict lookups
per choice. Deletes _path_matches_pattern, _build_preserved_paths, and
_remove_none_except_preserved helpers (~140 lines removed). Adds 13
regression tests including full output structure snapshots.
2026-02-10 09:49:44 -08:00
michelligabriele
1afe3032fd
fix(otel): auto-infer otlp_http exporter when endpoint is configured (#20438)
When OpenTelemetry is configured via the UI, only OTEL_ENDPOINT and
OTEL_HEADERS are set, but OTEL_EXPORTER is not specified. This caused
the exporter to default to "console", meaning traces were printed to
stdout instead of being sent to the configured endpoint.

This fix adds logic in OpenTelemetryConfig.__post_init__ to automatically
infer "otlp_http" as the exporter when an endpoint is specified but the
exporter is still the default "console".

Fixes issue reported by Elastic team where traces weren't being sent
to their OTEL endpoint when configured through the LiteLLM UI.
2026-02-10 09:33:16 -08:00
Sameer Kankute
0f01802dde
Merge pull request #20845 from BerriAI/litellm_gemini_image_handling
Handle image in assitant message for gemini
2026-02-10 18:24:09 +05:30
Sameer Kankute
3de892b8ca
Merge pull request #20860 from BerriAI/litellm_perplexity_research_api_support
[Feat] Perplexity research api support
2026-02-10 18:22:30 +05:30
Sameer Kankute
cdab87dec0
Merge pull request #20838 from BerriAI/litellm_managed_error_file
Add support managed error file
2026-02-10 18:20:26 +05:30
Sameer Kankute
f6228fda3e Fix mypy issues 2026-02-10 18:18:41 +05:30