Unit Tests: Proxy DB Operations / proxy-db (auth-checks, tests/proxy_unit_tests/test_auth_checks.py tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-db (remaining, tests/proxy_unit_tests --ignore=tests/proxy_unit_tests/test_key_generate_prisma.py --ignore=tests/proxy_unit_tests/test_auth_checks.py --ignore=tests/proxy_unit_tests/test_user_api_key_auth.py, 20, 8) (push) Has been cancelled
Added two new FAQ entries to the Enterprise docs page:
- How to set up your Enterprise License (LITELLM_LICENSE) via .env, Docker, or docker-compose
- How to verify the license is active by checking for 'Enterprise Edition' in the Swagger UI
* perf: cache _get_relevant_args_to_use_for_logging() as module-level frozenset
The set of valid LLM API parameter names for logging was being rebuilt
on every request from 8 OpenAI SDK type annotations + set operations.
Since these are static TypedDict annotations that never change at
runtime, compute once at import time and store as a class-level
frozenset.
Line profiler: get_standard_logging_model_parameters() dropped from
774ms to 77ms across 12K calls (90% reduction, ~25µs/req saved).
* test: add tests for cached ModelParamHelper logging args
Verify cached frozenset matches dynamic computation and that
prompt content keys (messages, prompt, input) are excluded from
logged model parameters.
- Cache CallTypes enum values as module-level dict to avoid repeated list
comprehension and enum construction on every call
- Hoist update_response_metadata getattr lookup to top of function
- Guard verbose print_verbose call behind _is_debugging_on() check
Pass-through endpoints (like vLLM classify) were not setting
standard_logging_object because _get_assembled_streaming_response
returns None for non-ModelResponse results.
This caused model_max_budget_limiter.async_log_success_event to raise
ValueError('standard_logging_payload is required').
The fix adds an elif branch in async_success_handler that mirrors the
non-pass-through code path.
Co-authored-by: openhands <openhands@all-hands.dev>
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
Add x-api-key header to CountTokens handler to match chat completion
authentication. Azure AI Anthropic requires this header per Microsoft's
native API format.
Bedrock rejects requests when toolResult or toolUse blocks within a
single message contain duplicate IDs. The Converse message transformer
merges consecutive tool/assistant messages without checking for
duplicate toolUseId values, causing BedrockException errors.
Add _deduplicate_bedrock_content_blocks() — a generalized helper that
removes duplicate blocks by ID, logs a warning for each dropped
duplicate via verbose_logger, and preserves non-tool blocks (e.g.
cachePoint). Apply it at all four merge sites (sync/async × toolResult/
toolUse).
The Anthropic /messages path was fixed in PR #19324; this applies the
equivalent fix to the Bedrock Converse path.
Fixes#20048
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
* Fix Nova grounding web_search_options={} not applying systemTool
Two bugs prevented web_search_options={} from working for Nova grounding:
1. Empty dict falsy check: The condition `value and isinstance(value, dict)`
short-circuits to False when value is {} (empty dict is falsy in Python).
Changed to `isinstance(value, dict)` to match Anthropic's implementation.
2. Pre-formatted tools mangled by _bedrock_tools_pt: The systemTool
(already in Bedrock format) was added to optional_params["tools"], but
_process_tools_and_beta passed all tools through _bedrock_tools_pt which
expects OpenAI-format tools. This corrupted the systemTool into an empty
toolSpec. Fixed by separating systemTool blocks before transformation
and appending them after.
Fixes follow-up to #19598
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Fix python-multipart Python version constraint for Poetry lock
python-multipart ^0.0.22 requires Python >=3.10 but the project supports
>=3.9. Add python = ">=3.10" marker so Poetry can resolve dependencies
for Python 3.9.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Previously, when Model Armor guardrail blocked a request/response,
the `applied_guardrails` field was not populated in the logs because
`add_guardrail_to_applied_guardrails_header()` was called after the
HTTPException was raised.
This fix moves the `add_guardrail_to_applied_guardrails_header()` call
to before the blocking check in all hooks:
- async_pre_call_hook (pre_call mode)
- async_moderation_hook (during_call mode)
- async_post_call_success_hook (post_call mode)
- async_post_call_streaming_iterator_hook (streaming)
This ensures that even when a guardrail blocks content, the guardrail
name is properly recorded in the logs for observability.
Added regression tests to verify applied_guardrails is populated when
content is blocked.
Co-authored-by: Cursor <cursoragent@cursor.com>
When using LiteLLM's Anthropic /v1/messages endpoint to route requests to
OpenAI models, requests fail if any tool name exceeds OpenAI's 64-character
limit. Anthropic API has no such limit, causing compatibility issues.
Changes:
- Add truncate_tool_name() function using {55-char-prefix}_{8-char-hash} format
- Modify translate_anthropic_tools_to_openai() to truncate and return mapping
- Modify translate_anthropic_tool_choice_to_openai() to truncate tool name
- Restore original tool names in responses using the mapping
- Support tool name restoration in streaming responses
- Add backwards-compatible API (existing methods still work)
The fix only applies when routing Anthropic requests to OpenAI models.
Native Anthropic/Claude requests pass through unchanged.
* fix: models loadbalancing billing issue by filter (#18891)
* fix: models loadbalancing billing issue by filter
* fix: separate key and team access groups in metadata
* fix: lint issues
Fixes#19788
- Add `supported_regions: ["global"]` to Qwen MaaS models in model_prices_and_context_window.json
- Update `get_supported_regions()` to read directly from `model_cost` dict
- Update `get_complete_vertex_url()` to use `get_vertex_region()` for global-only models
- Update `create_vertex_url()` to generate correct URL for global location (without region prefix)
- Add tests for Qwen global endpoint support
Per review feedback, thought_signature should not be a root-level
param on ImageObject as it's not OpenAI compatible. Moved to
provider_specific_fields dict to match the pattern used in chat
completions (Message, Delta, Choices, etc).
Fixes#17184 - Gemini 3 Pro image preview model returns a thoughtSignature
field required for interactive image editing. This change:
- Adds thought_signature field to ImageObject class
- Updates Gemini and Vertex AI transformations to extract thoughtSignature
- Adds test for thought_signature in response transformation