Commit graph

33248 commits

Author SHA1 Message Date
yuneng-jiang
2864ce73da
Merge pull request #21022 from BerriAI/litellm_unified_ag
[Feature] Access Groups
2026-02-12 15:34:38 -08:00
Milan
de42f733df fix: Update alias tooltip - remove outdated space replacement text
Since spaces are now blocked in server names, the tooltip text about
'spaces replaced by underscores' is no longer accurate.
2026-02-13 01:20:57 +02:00
Harshit Jain
c56bbb9067
Update tests/test_litellm/proxy/management_endpoints/test_ui_sso.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-13 04:48:36 +05:30
Milan
b769fa08d2 chore: Remove accidentally committed image file 2026-02-13 01:18:19 +02:00
Milan
bab5173500 fix: Make validation message generic and restore alias tooltip text
- Change error message to be generic (works for both server_name and alias)
- Restore 'Defaults to server name with spaces replaced' text in alias tooltip
2026-02-13 01:17:01 +02:00
Milan
1d67476ed0 refactor: Simplify validateMCPServerName to match original ternary style 2026-02-13 01:14:37 +02:00
Ishaan Jaff
5f40f93846
fix: MCP - inject NPM_CONFIG_CACHE into STDIO MCP subprocess env (#21069)
* fix: inject NPM_CONFIG_CACHE into STDIO MCP subprocess env for Docker

npm/npx needs a writable cache directory. In containers the default
(~/.npm) may not exist or be read-only, causing STDIO MCP servers
launched via npx to fail with ENOENT. Inject NPM_CONFIG_CACHE=/tmp/.npm_mcp_cache
into the subprocess env when not already set.

* test: add unit test for NPM_CONFIG_CACHE injection in STDIO MCP

Verifies that NPM_CONFIG_CACHE is auto-injected when not set, and
preserved when explicitly provided. Also moves the import to module
level per code style rules.

* Update litellm/constants.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Apply suggestion from @greptile-apps[bot]

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-12 15:11:37 -08:00
Milan
8fa2734830 fix(ui): Block spaces and hyphens in MCP server names and aliases
- Update validateMCPServerName to reject both spaces and hyphens
- Apply shared validation to alias field in create form (was inline)
- Update tooltips to mention space restriction
- Ensures consistency across create/edit forms for server_name and alias fields
2026-02-13 01:11:06 +02:00
Harshit Jain
a2b4728e74
Update tests/test_litellm/proxy/management_endpoints/test_ui_sso.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-13 04:40:29 +05:30
yuneng-jiang
a37623945d migration and build 2026-02-12 14:34:10 -08:00
Harshit Jain
e0846389e9
add modify test to perform async run 2026-02-13 04:04:03 +05:30
yuneng-jiang
ed59c7c84d bump: version 0.4.35 → 0.4.36 2026-02-12 14:33:38 -08:00
yuneng-jiang
a45028f623 Merge remote-tracking branch 'origin' into litellm_unified_ag 2026-02-12 14:32:52 -08:00
Ryan Crabbe
2065e5b88b perf: cache model_fields.keys() as frozensets in convert_to_model_response_object (15% faster)
Replace per-call .model_fields.keys() allocations and linear-scan membership
checks with module-level frozenset constants and dict.keys() set difference.
Defer locals() from hot path to except block. 617µs → 524µs/call.
2026-02-12 14:25:09 -08:00
Harshit Jain
847402b68d
Merge branch 'fix/sso_PKCE_deployments' of https://github.com/Harshit28j/litellm into fix/sso_PKCE_deployments 2026-02-13 03:50:16 +05:30
Harshit Jain
eb249b2f06
fix: add await in tests 2026-02-13 03:46:31 +05:30
Harshit Jain
bc5543cfdc
Update tests/test_litellm/proxy/management_endpoints/test_ui_sso.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-13 03:39:57 +05:30
yuneng-jiang
ea8c89ea3f rename file and add tests 2026-02-12 14:05:51 -08:00
Harshit Jain
1792b3c8e5
fix: add async call to avoid server pauses 2026-02-13 03:28:51 +05:30
yuneng-jiang
c34cdb29cd remove double auth and add alias 2026-02-12 13:40:58 -08:00
Ishaan Jaff
736daf0a7d
[Feat] Adds Shell tool support for the OpenAI Responses API (#21063)
* test_responses_api_context_management_server_side_compaction

* Server-side compaction

* docs fix

* test_responses_api_shell_tool

* add SHELL tool

* test_responses_api_shell_tool

* add SHELL_CALL_IN_PROGRESS

* add SHELL_CALL_IN_PROGRESS events

* TestOpenAIResponsesAPITest

* transform_streaming_response

* test_responses_api_shell_tool_streaming_sees_shell_output

* test_responses_api_shell_tool_streaming_sees_shell_output

* test_responses_api_shell_tool

* docs fix
2026-02-12 13:04:29 -08:00
yuneng-jiang
e6df587bfb adding tests and fixing prisma lookup table 2026-02-12 12:48:05 -08:00
yuneng-jiang
fbfaa6c8af rename unified access group to access group 2026-02-12 12:30:22 -08:00
yuneng-jiang
5a0db00d8a
Merge pull request #21061 from BerriAI/migration_yj_feb12
[Infra] Add Mmigration for Tags Adjustment on Policy Table
2026-02-12 10:36:10 -08:00
yuneng-jiang
5147515d78 add migration + build files 2026-02-12 10:34:59 -08:00
yuneng-jiang
5152bf4f4e bump: version 0.4.34 → 0.4.35 2026-02-12 10:34:17 -08:00
yuneng-jiang
df15456bcc
Merge pull request #20598 from muraliavarma/fix/team-update-empty-premium-fields-403
fix(proxy): skip premium check for empty metadata fields on team/key update
2026-02-12 10:10:27 -08:00
Ishaan Jaff
89565c97cc
[Feat] AI Gateway - Add Tracing for MCP Calls running through AI Gateway (#21018)
* commit new expansion

* fix MCP

* fix: LiteLLMProxyRequestSetup

* _process_mcp_tools_without_openai_transform

* UI fixes

* UI refactor view logs/sessions

* index

* _add_mcp_tool_metadata_to_final_chunk

* add badges

* add getEventDisplayName

* ui fixes

* backend fix

* fix

* UI fix

* UI fix

* fix row

* fix: address Greptile review feedback on PR #21018 (#21057)

- Fix session time range calculation: use Math.min/Math.max across all
  entries instead of relying on array order (sessionLogs is sorted by
  type, not time).

Other Greptile comments were already addressed in the branch:
- LogDetailContent.tsx exists
- Clipboard call already wrapped in try/catch
- Dedup already uses O(1) Map lookup
- model_dump() serialization is documented
- GROUP BY performance comment already present

---------

Co-authored-by: shin-bot-litellm <shin-bot-litellm@berri.ai>
2026-02-12 10:02:03 -08:00
Ishaan Jaff
3d9b145b04
[Feat] Adds support for server-side compaction on the OpenAI Responses API context_management (#21058)
* test_responses_api_context_management_server_side_compaction

* Server-side compaction

* docs fix

* test_responses_api_shell_tool
2026-02-12 10:00:30 -08:00
Krrish Dholakia
f5382ebac9 docs: fix docs 2026-02-12 08:45:57 -08:00
Sameer Kankute
556bcd7203
Merge pull request #21055 from BerriAI/litellm_day_0_MiniMax-M2.1
fix docs
2026-02-12 22:04:23 +05:30
Sameer Kankute
9f15eca6b6 fix docs 2026-02-12 22:03:21 +05:30
Sameer Kankute
4e62386c65
Merge pull request #21054 from BerriAI/litellm_day_0_MiniMax-M2.1
Add support for MiniMax-M2.1 and MiniMax-M2.1-lightining
2026-02-12 21:51:46 +05:30
Sameer Kankute
7a3b227aeb
Apply suggestion from @greptile-apps[bot]
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-02-12 21:50:55 +05:30
Sameer Kankute
9b32c516ad Add support for MiniMax-M2.1 and MiniMax-M2.1-lightining 2026-02-12 21:45:49 +05:30
Sameer Kankute
a7179797f7
Merge pull request #20930 from BerriAI/litellm_oss_staging_02_11_2026
oss staging 02 / 11/ 2026
2026-02-12 21:28:59 +05:30
Sameer Kankute
5471b14c56
Merge pull request #21051 from BerriAI/revert-20569-fix/20557-gemini-multiturn-tool-calling
Revert "Fix #20557: Fix Gemini multi-turn tool calling message formatting"
2026-02-12 21:17:04 +05:30
Sameer Kankute
4931bac3da
Revert "Fix #20557: Fix Gemini multi-turn tool calling message formatting (#2…"
This reverts commit 49078b3c6b.
2026-02-12 21:16:33 +05:30
Sameer Kankute
e68b970953
Merge branch 'main' into litellm_oss_staging_02_11_2026 2026-02-12 21:14:08 +05:30
Sameer Kankute
3d7126b11a
Merge pull request #20587 from BerriAI/litellm_oss_staging_02_06_2026
Litellm oss staging 02 06 2026
2026-02-12 20:07:35 +05:30
Sameer Kankute
59d6ab8a00
Merge branch 'main' into litellm_oss_staging_02_11_2026 2026-02-12 20:04:46 +05:30
Cesar Garcia
622983cf89 fix(helm): add OCI annotations so GHCR shows helm pull instead of docker pull (#20617)
The Helm chart on GHCR displays a `docker pull` command instead of
the correct `helm pull oci://` command. This is because the OCI artifact
is missing the `org.opencontainers.image.source` annotation that GHCR
uses to identify and properly display Helm charts.

Changes:
- Add OCI annotations to Chart.yaml (source + url) which Helm 3.10+
  propagates to the OCI manifest on push
- Install explicit Helm v3.20.0 via azure/setup-helm@v4 for reproducible
  builds and proper OCI annotation support
- Remove deprecated HELM_EXPERIMENTAL_OCI env var (OCI is GA since Helm 3.8)
2026-02-12 19:58:16 +05:30
Cesar Garcia
df38de5683 docs(web_search): add gpt-5-search-api usage examples for SDK and AI Gateway (#20616)
- Document two OpenAI web search approaches: search models (/chat/completions) vs web_search_preview tool (/responses)
- Add gpt-5-search-api examples across all sections in web_search.md
- Update /responses examples to use gpt-5 with web_search_preview tool
- Add OpenAI Web Search Models section to providers/openai.md
- Add web search example to providers/openai/responses_api.md
2026-02-12 19:58:12 +05:30
shin-bot-litellm
7b6b97cc10 fix: mask API keys in error responses for invalid/malformed keys (#20289)
Fixes AT&T customer issue where API keys are returned in plain text
in error responses.

Changes:
1. user_api_key_auth.py: Mask the API key in the AssertionError when a
   key doesn't start with 'sk-' (e.g. key with leading space). Shows
   first 4 + last 4 chars with **** in between instead of the full key.

2. key_management_endpoints.py: Same masking for the key format
   validation error when creating keys with invalid prefix.

3. presidio.py: Sanitize exceptions from Presidio analyze/anonymize
   calls to prevent leaking original request text (which may contain
   API keys) in error responses. Error messages now show only the
   exception type, not the full payload.
2026-02-12 19:58:05 +05:30
shin-bot-litellm
7ee36c2a3a fix(http_handler): bypass cache when shared_session is provided for aiohttp tracing (#20630)
* Add http support to custom code guardrails + Unified guardrails for MCP + Agent guardrail support (#20619)

* fix: fix styling

* fix(custom_code_guardrail.py): add http support for custom code guardrails

allows users to call external guardrails on litellm with minimal code changes (no custom handlers)

Test guardrail integrations more easily

* feat(a2a/): add guardrails for agent interactions

allows the same guardrails for llm's to be applied to agents as well

* fix(a2a/): support passing guardrails to a2a from the UI

* style(code-editor): allow editing custom code guardrails on ui + add examples of pre/post calls for custom code guardrails

* feat(mcp/): support custom code guardrails for mcp calls

allows custom code guardrails to work on mcp input

* feat(chatui.tsx): support guardrails on mcp tool calls on playground

* fix(mypy): resolve missing return statements and type casting issues (#20618)

* fix(mypy): resolve missing return statements and type casting issues

* fix(pangea): use elif to prevent UnboundLocalError and handle None messages

Address Greptile review feedback:
- Make branches mutually exclusive using elif to prevent input_messages from being overwritten
- Handle case where data.get('messages') returns None to avoid passing invalid payload to Pangea API

---------

Co-authored-by: Shin <shin@openclaw.ai>

* [Feat] MCP Gateway - Allow setting MCP Servers as Private/Public available on Internet (#20607)

* update MCPAuthenticatedUser

* add available_on_public_internet for MCPs

* update claude.md

* init IPAddressUtils

* init available_on_public_internet

* add on REST endpoints

* filter with IP

* TestIsInternalIp

* _extract_mcp_headers_from_request

* init get_mcp_client_ip

* _get_general_settings

* allowed_server_ids

* address PR comments

* get_mcp_server_by_name fix

* fix server

* fix review comments

* get_public_mcp_servers

* address _get_allowed_mcp_servers

* fixing user_id

* [Feat] IP-Based Access Control for MCP Servers (#20620)

* update MCPAuthenticatedUser

* add available_on_public_internet for MCPs

* update claude.md

* init IPAddressUtils

* init available_on_public_internet

* add on REST endpoints

* filter with IP

* TestIsInternalIp

* _extract_mcp_headers_from_request

* init get_mcp_client_ip

* _get_general_settings

* allowed_server_ids

* address PR comments

* get_mcp_server_by_name fix

* fix server

* fix review comments

* get_public_mcp_servers

* address _get_allowed_mcp_servers

* test fix

* fix linting

* inint ui types

* add ui for managing MCP private/public

* add ui

* fixes

* add to schema

* add types

* fix endpoint

* add endpoint

* update manager

* test mcp

* dont use external party for ip address

* Add OpenAI/Azure release test suite with HTTP client lifecycle regression detection (#20622)

* docs (#20626)

* docs

* fix(mypy): resolve type checking errors in 5 files (#20627)

- a2a_protocol/exception_mapping_utils.py: Fix type ignore comment for None assignment
- caching/redis_cache.py: Add type ignore for async ping return type
- caching/redis_cluster_cache.py: Add type ignore for async ping return type
- llms/deprecated_providers/palm.py: Add type ignore for palm.generate_text
- proxy/auth/handle_jwt.py: Add type ignore for jwt.decode options argument

All changes add appropriate type: ignore comments to handle library typing inconsistencies.

* fix(test): update deprecated gemini embedding model (#20621)

Replace text-embedding-004 with gemini-embedding-001.

The old model was deprecated and returns 404:
'models/text-embedding-004 is not found for API version v1beta'

Co-authored-by: Shin <shin@openclaw.ai>

* ui new buil

* fix(http_handler): bypass cache when shared_session is provided for aiohttp tracing

When users pass a shared_session with trace_configs to acompletion(),
the get_async_httpx_client() function was ignoring it and returning
a cached client without the user's tracing configuration.

This fix bypasses the cache when shared_session is provided, ensuring
the user's ClientSession (with its trace_configs, connector settings, etc.)
is actually used for the request.

Fixes #20174

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
Co-authored-by: Shin <shin@openclaw.ai>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: yuneng-jiang <yuneng.jiang@gmail.com>
Co-authored-by: Alexsander Hamir <alexsanderhamirgomesbaptista@gmail.com>
Co-authored-by: shin-bot-litellm <shin-bot-litellm@users.noreply.github.com>
2026-02-12 19:57:57 +05:30
Varun Chawla
e587370f67 fix(proxy): add regression tests for #20441 - <script> tags in messages (#20573)
* fix: empty guardrails/policies arrays should not trigger enterprise license check (#20304)

The UI sends empty arrays for enterprise-only fields (guardrails, policies,
logging) even when the user has not configured these features. The backend
`is not None` check treated `[]` as a truthy intent to use the feature,
falsely requiring an enterprise license for basic team operations.

Backend: Add `and updated_kv[field] != [] and updated_kv[field] != {}`
guards in `_update_metadata_fields` so empty collections are skipped.

UI: Conditionally omit guardrails, logging, and policies from the
payload when empty instead of defaulting to `[]`.

Fixes #20304

* fix: allow clearing fields with empty collections while skipping enterprise check

Address PR review feedback:

1. Move the empty-collection guard into _update_metadata_field (singular)
   so that empty lists/dicts skip only the premium license check but still
   get written into metadata. This lets users intentionally clear a
   previously-set field (e.g. guardrails: []) without being blocked, while
   the UI's default empty arrays still don't trigger a false enterprise
   error.

2. Remove sys.path hack from test file; use standard imports that work
   with pytest discovery.

3. Add tests verifying that empty collections are moved into metadata
   (field clearing works) even though they bypass the premium check.

Fixes #20304

* fix(proxy): add regression tests for #20441 - ensure <script> tags in LLM messages are not blocked

The 403 Forbidden error when sending messages containing `<script>` is caused
by external WAF/reverse proxy infrastructure (confirmed by the standard nginx
HTML 403 response format), not by LiteLLM's own content filtering. However,
these regression tests ensure that:

1. The content filter guardrail's built-in patterns do not match HTML tags
2. Messages containing <script> and other HTML tags pass through the content
   filter unchanged when no explicit HTML-blocking rules are configured
3. The HTTP request body parser correctly handles JSON payloads containing
   HTML content without modification

These tests guard against accidentally introducing HTML/XSS filtering that
would break legitimate LLM API usage (e.g., discussing HTML/JavaScript code).

Closes #20441

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-12 19:57:40 +05:30
Piotr Grabowski
8a5feb18e3 fix(openrouter): fix crash of gpt-5.2-codex by using mode "chat" (#20577)
Commit 1cdda28b6 changed "openrouter/openai/gpt-5.2-codex" to mode "responses",
but this broke GPT-5.2-Codex with OpenRouter:

```
response = await litellm.acompletion(
            model="openrouter/openai/gpt-5.2-codex",
            messages=[{"role": "user", "content": "Hello"}],
            api_key=os.environ.get("OPENROUTER_API_KEY"),
)
```
crashes with: `OpenrouterException - argument of type 'NoneType' is not iterable`

Responses API is in beta in OpenRouter and no other OpenRouter models use "responses"
mode. The commit that changed this probably did it by mistake.

Therefore change the mode to "chat" and fix the crash.
2026-02-12 19:56:03 +05:30
Varun Chawla
6fc335030a fix(responses): handle Pydantic ValidationError when provider omits required fields in streaming events (#20580)
When an OpenAI-compatible upstream provider emits minimal streaming event
payloads that omit required fields (e.g. created_at, output, output_index,
content_index), Pydantic raises a ValidationError crashing the SSE stream
and returning HTTP 500.

Fall back to model_construct() on ValidationError, consistent with the
existing pattern in transform_response_api_response for non-streaming.

Fixes https://github.com/BerriAI/litellm/issues/20570

Signed-off-by: Varun Chawla <varun_6april@hotmail.com>
2026-02-12 19:55:58 +05:30
jquinter
e78d17cd0a Fix/mcp health check cancelled error (#19851)
* Fix MCP health check CancelledError handling for parallel test execution

Add asyncio.CancelledError handler in health_check_server() and missing
@pytest.mark.asyncio decorator on test_mcp_server_manager_config_integration_with_database.

In Python 3.8+, CancelledError inherits from BaseException, not Exception,
so it bypassed the generic exception handler when pytest-xdist cancels
running tasks after a failure.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

* Regenerate poetry.lock to resolve merge conflict markers

The lock file had unresolved conflict markers from a previous merge,
causing poetry to fail with "Invalid statement (at line 8534, column 1)".

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-02-12 19:40:08 +05:30
jquinter
2875fe8e49 ci: add matrix-based parallel test workflow (#19942)
Split tests/test_litellm into 10 parallel CI jobs using GitHub Actions
matrix strategy to reduce PR feedback time from ~25 min to ~8-10 min.

Changes:
- Add new test-litellm-matrix.yml workflow with 10 matrix jobs:
  - llms (~225 files, 4 workers)
  - proxy-guardrails (~51 files, 4 workers)
  - proxy-core (~52 files, 4 workers)
  - proxy-misc (~77 files, 4 workers)
  - integrations (~60 files, 4 workers)
  - core-utils (~32 files, 2 workers)
  - other (~69 files, 4 workers) - includes all previously uncovered dirs
  - root (~34 files, 4 workers)
  - proxy-unit-a (~20 files, 2 workers)
  - proxy-unit-b (~28 files, 2 workers)

- Deprecate test-litellm.yml (moved to workflow_dispatch for manual use)

- Add matching Makefile targets for local testing:
  - make test-unit-llms
  - make test-unit-proxy-guardrails
  - make test-unit-proxy-core
  - make test-unit-proxy-misc
  - make test-unit-integrations
  - make test-unit-core-utils
  - make test-unit-other
  - make test-unit-root
  - make test-proxy-unit-a
  - make test-proxy-unit-b

Benefits:
- ~3x faster wall-clock time through parallelization
- Dependency caching for faster subsequent runs
- Concurrency control to cancel stale runs
- Better failure isolation per test group

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-02-12 19:39:05 +05:30