DualCache.async_set_cache and async_set_cache_pipeline were missing the
default_in_memory_ttl injection that the sync set_cache method has. This
caused InMemoryCache to fall back to its own default_ttl (600s) instead
of using DualCache's configured default_in_memory_ttl (typically 60s).
This is particularly impactful for end-user budget enforcement in the
proxy, where cached spend values could remain stale for 10 minutes
instead of 1 minute, allowing users to exceed their budgets.
The response.completed handler in the completion→responses streaming
bridge was discarding the usage object, causing prompt_tokens_details
(and cached_tokens) to always be None when streaming with models that
use the Responses API (e.g. gpt-5.2-codex, gpt-5.3-codex).
Extract usage from the response.completed event and translate it via
the existing _transform_response_api_usage_to_chat_usage helper.
Fixes#22192
All 55 deepinfra models that had `supports_tool_choice: true` were
missing the `supports_function_calling` flag, causing
`litellm.supports_function_calling()` to incorrectly return False.
Fixes#22619
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* Add CrowdStrike AIDR guardrail hook
* fixup! use apply_guardrail event hook
* fixup! update imports
* fix(guardrails): include AI response in CrowdStrike AIDR output events
Issue:
_build_guard_input_for_response() was:
- Sending only the original user input (messages).
- Not sending the AI provider response.
This fix will:
- Extract response.choices from the ModelResponse object and include them in guard_input payload.
- Thus, ensure AIDR output rules receive the AI-generated content for analysis.
- Fix and update tests.
* fix(guardrails): prevent duplicate input events in CrowdStrike AIDR guardrail
Issue:
The CrowdStrike AIDR guardrail was running on during_call hooks wihtout event_hook configured.
This fix will:
- Set event_hook to ["pre_call", "post_call"] (AIDR admins will control what policy is applied)
This change will:
- Require default_on parameter
- Prevent duplicate API calls to AIDR for the same input
- Avoid unchecked AI provider API calls on during_call hook
* docs: add CrowdStrike AIDR to the list of Guardrails under Integrations
* docs: update CrowdStrike AIDR documentation page
---------
Co-authored-by: Konstantin Lapine <konstantin.lapine@crowdstrike.com>
Reorder elif branches so is_vertex_ai is checked before "gemini" in model.
Previously, Vertex AI Gemini models (e.g. vertex_ai/gemini-2.5-flash) matched
the "gemini" substring check first and were logged with the Google AI Studio
URL instead of the Vertex AI URL.
mistral-small-latest now points to Small 3.2 (since June 2025).
Updated pricing from $0.10/$0.30 to $0.06/$0.18 per 1M tokens,
context from 32k to 131k, and added vision support to match
mistral-small-3-2-2506.
Add 9 new Mistral models (mistral-large-2512, mistral-medium-3-1-2508,
mistral-small-3-2-2506, ministral-3-3b/8b/14b-2512, saba-2502,
magistral-medium/small-1-2-2509) and update mistral-large-latest,
mistral-large-3, and mistral-medium-latest with correct pricing and
context windows.
Fixes#22585
Fixes#22591 - These models were missing from the pricing JSON, causing
$0 cost tracking when routed via the dashscope/* wildcard.
Pricing sourced from official Alibaba Cloud Model Studio docs (international tier).
Fixes NameError at runtime when ChatGPTToolCallNormalizer is
instantiated. The imports were missed when type hints were changed
from Python 3.10+ syntax (dict[], str | None) to typing module
syntax (Dict[], Optional[str]).
The previous check in _get_openai_compatible_provider_info() ran after
the model name was already split, so it never caught the second
get_llm_provider() call from the anthropic_messages bridge.
Moved the check to get_llm_provider() before the provider-list
stripping, using a pattern-based approach (custom_llm_provider ==
"openrouter" and model.startswith("openrouter/")) instead of a
hardcoded set. This covers all current and future native OpenRouter
models.
Updated tests to verify the bridge double-call scenario with
custom_llm_provider passed through.
Empty JSON schemas `{}` mean "any JSON value is valid" per the spec,
but _build_vertex_schema was coercing them to `{"type": "object"}` in
three places (process_items, convert_anyof_null_to_nullable, add_object_type),
breaking Pydantic JsonValue fields on Gemini.
Adds _is_any_type_schema() to detect unconstrained schemas and skip
the type coercion, preserving Gemini's TYPE_UNSPECIFIED semantics.
Fixes#22391
Bedrock Claude models were missing cache_read_input_token_cost and
cache_creation_input_token_cost fields, causing cache tokens to be
billed at the full input rate instead of the discounted cache rate.
Added pricing using Bedrock's documented multipliers (0.1x for cache
read, 1.25x for cache write) consistent with all existing entries.
The proxy has two separate failure paths:
1. async_failure_handler → Langfuse callback (uses model_call_details with
standard_logging_object containing the correct trace_id)
2. post_call_failure_hook → _ProxyDBLogger → spend log (uses request_data
which did NOT have standard_logging_object, so session_id fell to
random uuid4())
These two paths used different data dicts, so the DB session_id was a
random UUID unrelated to the Langfuse trace_id. Users could not search
by the Session ID from LiteLLM logs in Langfuse for failed requests.
Fix: In _ProxyDBLogger.async_post_call_failure_hook, propagate
standard_logging_object and litellm_trace_id from the litellm_logging_obj
(already present in request_data) before writing the spend log.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Logging.__init__ stored the raw litellm_trace_id parameter (None when not
explicitly provided) in model_call_details, while self.litellm_trace_id
always held a valid UUID. When get_standard_logging_object_payload() failed,
both the DB and Langfuse fell back to kwargs["litellm_trace_id"] which was
None, causing each to generate different random UUIDs. Now model_call_details
stores self.litellm_trace_id (always valid), so all fallback paths use the
same ID.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When standard_logging_object is None (failure case), Langfuse was falling
back to litellm_call_id while the DB used litellm_trace_id as session_id.
This caused the Session ID in LiteLLM logs to not match the trace in
Langfuse. Now Langfuse checks litellm_trace_id first, matching the DB.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The proxy uses async_failure_handler → LangfusePromptManagement.async_log_failure_event(),
which silently returned when standard_logging_object was None. This meant failed LLM calls
never created traces in Langfuse. Remove the early return and fall back to extracting the
error message from kwargs["exception"] when standard_logging_object is unavailable.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>