The Helm chart on GHCR displays a `docker pull` command instead of
the correct `helm pull oci://` command. This is because the OCI artifact
is missing the `org.opencontainers.image.source` annotation that GHCR
uses to identify and properly display Helm charts.
Changes:
- Add OCI annotations to Chart.yaml (source + url) which Helm 3.10+
propagates to the OCI manifest on push
- Install explicit Helm v3.20.0 via azure/setup-helm@v4 for reproducible
builds and proper OCI annotation support
- Remove deprecated HELM_EXPERIMENTAL_OCI env var (OCI is GA since Helm 3.8)
- Document two OpenAI web search approaches: search models (/chat/completions) vs web_search_preview tool (/responses)
- Add gpt-5-search-api examples across all sections in web_search.md
- Update /responses examples to use gpt-5 with web_search_preview tool
- Add OpenAI Web Search Models section to providers/openai.md
- Add web search example to providers/openai/responses_api.md
Fixes AT&T customer issue where API keys are returned in plain text
in error responses.
Changes:
1. user_api_key_auth.py: Mask the API key in the AssertionError when a
key doesn't start with 'sk-' (e.g. key with leading space). Shows
first 4 + last 4 chars with **** in between instead of the full key.
2. key_management_endpoints.py: Same masking for the key format
validation error when creating keys with invalid prefix.
3. presidio.py: Sanitize exceptions from Presidio analyze/anonymize
calls to prevent leaking original request text (which may contain
API keys) in error responses. Error messages now show only the
exception type, not the full payload.
* Add http support to custom code guardrails + Unified guardrails for MCP + Agent guardrail support (#20619)
* fix: fix styling
* fix(custom_code_guardrail.py): add http support for custom code guardrails
allows users to call external guardrails on litellm with minimal code changes (no custom handlers)
Test guardrail integrations more easily
* feat(a2a/): add guardrails for agent interactions
allows the same guardrails for llm's to be applied to agents as well
* fix(a2a/): support passing guardrails to a2a from the UI
* style(code-editor): allow editing custom code guardrails on ui + add examples of pre/post calls for custom code guardrails
* feat(mcp/): support custom code guardrails for mcp calls
allows custom code guardrails to work on mcp input
* feat(chatui.tsx): support guardrails on mcp tool calls on playground
* fix(mypy): resolve missing return statements and type casting issues (#20618)
* fix(mypy): resolve missing return statements and type casting issues
* fix(pangea): use elif to prevent UnboundLocalError and handle None messages
Address Greptile review feedback:
- Make branches mutually exclusive using elif to prevent input_messages from being overwritten
- Handle case where data.get('messages') returns None to avoid passing invalid payload to Pangea API
---------
Co-authored-by: Shin <shin@openclaw.ai>
* [Feat] MCP Gateway - Allow setting MCP Servers as Private/Public available on Internet (#20607)
* update MCPAuthenticatedUser
* add available_on_public_internet for MCPs
* update claude.md
* init IPAddressUtils
* init available_on_public_internet
* add on REST endpoints
* filter with IP
* TestIsInternalIp
* _extract_mcp_headers_from_request
* init get_mcp_client_ip
* _get_general_settings
* allowed_server_ids
* address PR comments
* get_mcp_server_by_name fix
* fix server
* fix review comments
* get_public_mcp_servers
* address _get_allowed_mcp_servers
* fixing user_id
* [Feat] IP-Based Access Control for MCP Servers (#20620)
* update MCPAuthenticatedUser
* add available_on_public_internet for MCPs
* update claude.md
* init IPAddressUtils
* init available_on_public_internet
* add on REST endpoints
* filter with IP
* TestIsInternalIp
* _extract_mcp_headers_from_request
* init get_mcp_client_ip
* _get_general_settings
* allowed_server_ids
* address PR comments
* get_mcp_server_by_name fix
* fix server
* fix review comments
* get_public_mcp_servers
* address _get_allowed_mcp_servers
* test fix
* fix linting
* inint ui types
* add ui for managing MCP private/public
* add ui
* fixes
* add to schema
* add types
* fix endpoint
* add endpoint
* update manager
* test mcp
* dont use external party for ip address
* Add OpenAI/Azure release test suite with HTTP client lifecycle regression detection (#20622)
* docs (#20626)
* docs
* fix(mypy): resolve type checking errors in 5 files (#20627)
- a2a_protocol/exception_mapping_utils.py: Fix type ignore comment for None assignment
- caching/redis_cache.py: Add type ignore for async ping return type
- caching/redis_cluster_cache.py: Add type ignore for async ping return type
- llms/deprecated_providers/palm.py: Add type ignore for palm.generate_text
- proxy/auth/handle_jwt.py: Add type ignore for jwt.decode options argument
All changes add appropriate type: ignore comments to handle library typing inconsistencies.
* fix(test): update deprecated gemini embedding model (#20621)
Replace text-embedding-004 with gemini-embedding-001.
The old model was deprecated and returns 404:
'models/text-embedding-004 is not found for API version v1beta'
Co-authored-by: Shin <shin@openclaw.ai>
* ui new buil
* fix(http_handler): bypass cache when shared_session is provided for aiohttp tracing
When users pass a shared_session with trace_configs to acompletion(),
the get_async_httpx_client() function was ignoring it and returning
a cached client without the user's tracing configuration.
This fix bypasses the cache when shared_session is provided, ensuring
the user's ClientSession (with its trace_configs, connector settings, etc.)
is actually used for the request.
Fixes#20174
---------
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
Co-authored-by: Shin <shin@openclaw.ai>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: yuneng-jiang <yuneng.jiang@gmail.com>
Co-authored-by: Alexsander Hamir <alexsanderhamirgomesbaptista@gmail.com>
Co-authored-by: shin-bot-litellm <shin-bot-litellm@users.noreply.github.com>
* fix: empty guardrails/policies arrays should not trigger enterprise license check (#20304)
The UI sends empty arrays for enterprise-only fields (guardrails, policies,
logging) even when the user has not configured these features. The backend
`is not None` check treated `[]` as a truthy intent to use the feature,
falsely requiring an enterprise license for basic team operations.
Backend: Add `and updated_kv[field] != [] and updated_kv[field] != {}`
guards in `_update_metadata_fields` so empty collections are skipped.
UI: Conditionally omit guardrails, logging, and policies from the
payload when empty instead of defaulting to `[]`.
Fixes#20304
* fix: allow clearing fields with empty collections while skipping enterprise check
Address PR review feedback:
1. Move the empty-collection guard into _update_metadata_field (singular)
so that empty lists/dicts skip only the premium license check but still
get written into metadata. This lets users intentionally clear a
previously-set field (e.g. guardrails: []) without being blocked, while
the UI's default empty arrays still don't trigger a false enterprise
error.
2. Remove sys.path hack from test file; use standard imports that work
with pytest discovery.
3. Add tests verifying that empty collections are moved into metadata
(field clearing works) even though they bypass the premium check.
Fixes#20304
* fix(proxy): add regression tests for #20441 - ensure <script> tags in LLM messages are not blocked
The 403 Forbidden error when sending messages containing `<script>` is caused
by external WAF/reverse proxy infrastructure (confirmed by the standard nginx
HTML 403 response format), not by LiteLLM's own content filtering. However,
these regression tests ensure that:
1. The content filter guardrail's built-in patterns do not match HTML tags
2. Messages containing <script> and other HTML tags pass through the content
filter unchanged when no explicit HTML-blocking rules are configured
3. The HTTP request body parser correctly handles JSON payloads containing
HTML content without modification
These tests guard against accidentally introducing HTML/XSS filtering that
would break legitimate LLM API usage (e.g., discussing HTML/JavaScript code).
Closes#20441
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Commit 1cdda28b6 changed "openrouter/openai/gpt-5.2-codex" to mode "responses",
but this broke GPT-5.2-Codex with OpenRouter:
```
response = await litellm.acompletion(
model="openrouter/openai/gpt-5.2-codex",
messages=[{"role": "user", "content": "Hello"}],
api_key=os.environ.get("OPENROUTER_API_KEY"),
)
```
crashes with: `OpenrouterException - argument of type 'NoneType' is not iterable`
Responses API is in beta in OpenRouter and no other OpenRouter models use "responses"
mode. The commit that changed this probably did it by mistake.
Therefore change the mode to "chat" and fix the crash.
When an OpenAI-compatible upstream provider emits minimal streaming event
payloads that omit required fields (e.g. created_at, output, output_index,
content_index), Pydantic raises a ValidationError crashing the SSE stream
and returning HTTP 500.
Fall back to model_construct() on ValidationError, consistent with the
existing pattern in transform_response_api_response for non-streaming.
Fixes https://github.com/BerriAI/litellm/issues/20570
Signed-off-by: Varun Chawla <varun_6april@hotmail.com>
* Fix MCP health check CancelledError handling for parallel test execution
Add asyncio.CancelledError handler in health_check_server() and missing
@pytest.mark.asyncio decorator on test_mcp_server_manager_config_integration_with_database.
In Python 3.8+, CancelledError inherits from BaseException, not Exception,
so it bypassed the generic exception handler when pytest-xdist cancels
running tasks after a failure.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Regenerate poetry.lock to resolve merge conflict markers
The lock file had unresolved conflict markers from a previous merge,
causing poetry to fail with "Invalid statement (at line 8534, column 1)".
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Split tests/test_litellm into 10 parallel CI jobs using GitHub Actions
matrix strategy to reduce PR feedback time from ~25 min to ~8-10 min.
Changes:
- Add new test-litellm-matrix.yml workflow with 10 matrix jobs:
- llms (~225 files, 4 workers)
- proxy-guardrails (~51 files, 4 workers)
- proxy-core (~52 files, 4 workers)
- proxy-misc (~77 files, 4 workers)
- integrations (~60 files, 4 workers)
- core-utils (~32 files, 2 workers)
- other (~69 files, 4 workers) - includes all previously uncovered dirs
- root (~34 files, 4 workers)
- proxy-unit-a (~20 files, 2 workers)
- proxy-unit-b (~28 files, 2 workers)
- Deprecate test-litellm.yml (moved to workflow_dispatch for manual use)
- Add matching Makefile targets for local testing:
- make test-unit-llms
- make test-unit-proxy-guardrails
- make test-unit-proxy-core
- make test-unit-proxy-misc
- make test-unit-integrations
- make test-unit-core-utils
- make test-unit-other
- make test-unit-root
- make test-proxy-unit-a
- make test-proxy-unit-b
Benefits:
- ~3x faster wall-clock time through parallelization
- Dependency caching for faster subsequent runs
- Concurrency control to cancel stale runs
- Better failure isolation per test group
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
- Fix skip condition to detect claude models (was only checking for
"anthropic" in model name, missing "claude-haiku-4-5")
- Add missing skip for OpenAI tests when OPENAI_API_KEY is not set
- Fix TypeError in utils.py when metadata is explicitly None instead
of missing (use `or {}` fallback)
HuggingFace Text Embeddings Inference (TEI) returns embeddings as raw
arrays [[0.1, 0.2, ...]] instead of wrapped format {"embedding": [...]}.
This change handles both formats:
- Raw array: [[...]] (TEI, some HF models)
- Wrapped: {"embedding": [[...]]} (standard HF format)
Fixes SagemakerError: "HF response missing 'embedding' field" when using
TEI containers on SageMaker.
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
* fix(oldteams.tsx): show policies when creating
* fix(proxy/_types.py): ensure mcp rest endpoints can be called by virtual key
ensures UI works with virtual key testing mcp endpoints
* refactor: migrate get object permissions table logic to happen in user api key auth - allows functions to trust user api key object they receive has what they need
* fix(rest_endpoints.py): filter for allowed tools based on what key has access to
* fix(mcp_server_manager.py): ensure only allowed MCP's are returned to the user, via rest endpoints
- Updated benchmarks.md with a section on setting up fake OpenAI endpoints
- Updated load_test.md to mention the self-hosted option
- Updated load_test_advanced.md with a tip box about the example repo
Reference: https://github.com/BerriAI/example_openai_endpoint
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
- Updated create_mcp_server.tsx to show 'Streamable HTTP (Recommended)' label
- Updated mcp_server_edit.tsx to show 'Streamable HTTP (Recommended)' label
- Both Add New MCP Server and Edit MCP Server pages now display the updated label
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix:fix: prompt_cache_key OAI + Azure OpenAI
* test_prompt_cache_key_supported
* test_azure_openai_with_prompt_cache_key
* fix: remove unnecessary async from test_azure_openai_with_prompt_cache_key
Addresses Greptile feedback: litellm.completion() is synchronous, so
async def is unnecessary and would silently pass without running.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: remove unused filter_and_transform_beta_headers imports
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* test_azure_openai_with_prompt_cache_key
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* fix: _should_use_api_key_header
* test_azure_ai_validate_environment_with_api_key
* fix: remove unused top-level RouteChecks import
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* docs: add missing env keys to config_settings reference
Add MODEL_COST_MAP_MIN_MODEL_COUNT, MODEL_COST_MAP_MAX_SHRINK_RATIO,
and MAX_POLICY_ESTIMATE_IMPACT_ROWS to the environment variables
reference table.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
- Add missing 'role' field to tool result ContentType objects
- Tool result messages must have role='user' to match Gemini API requirements
- Fixes 'contents[2].parts[0].data: required oneof field data' error on second turn
- Prevents uninitialized data field in tool result messages
- Add verbose_logger to imports when LITELLM_LOG=DEBUG is set
- Set verbose_logger to DEBUG level alongside verbose_proxy_logger and verbose_router_logger
- Fixes issue where callback integrations (Langsmith, Langfuse, etc.) don't show debug logs with LITELLM_LOG=DEBUG
- Makes LITELLM_LOG=DEBUG behavior consistent with --detailed_debug flag
PermissionDeniedError (403) is defined in litellm/exceptions.py but was
never added to the import block in litellm/__init__.py. This makes it
the only standard HTTP error exception not accessible as
litellm.PermissionDeniedError, forcing users to import from
litellm.exceptions directly.
Fixes#20959