When create_batch routes via x-litellm-model header, the response batch_id
was returned raw without model routing info. This meant retrieve_batch could
not determine which provider/credentials to use, defaulting to "openai"
instead of the correct provider (e.g., VLLM).
Now encodes batch_id, output_file_id, and error_file_id with model info
(same pattern as the model-embedded file_id flow in Scenario 1), so
retrieve_batch can decode and route back to the correct provider.
Fixes#22646
Adds pricing for DashScope models that were missing from the cost map,
causing $0 spend tracking in the proxy dashboard:
- dashscope/qwen3-max-2026-01-23 (tiered, same as qwen3-max)
- dashscope/qwen3-next-80b-a3b-instruct ($0.15/$1.20 per 1M)
- dashscope/qwen3-next-80b-a3b-thinking ($0.15/$1.20 per 1M)
- dashscope/qwen3-vl-235b-a22b-instruct ($0.40/$1.60 per 1M)
- dashscope/qwen3-vl-235b-a22b-thinking ($0.40/$4.00 per 1M)
- dashscope/qwen3-vl-32b-instruct ($0.16/$0.64 per 1M)
- dashscope/qwen3-vl-32b-thinking ($0.16/$2.87 per 1M)
Fixes#22609
Adds pricing for OpenRouter models that were routing correctly but
returning $0 for spend tracking due to missing cost map entries:
- openrouter/anthropic/claude-sonnet-4.6 ($3.00/$15.00 per 1M tokens)
- openrouter/google/gemini-3.1-pro-preview ($2.00/$12.00 per 1M tokens)
- openrouter/openai/gpt-5.1-codex-max ($1.25/$10.00 per 1M tokens)
- openrouter/qwen/qwen3-coder-plus ($1.00/$5.00 per 1M tokens)
- openrouter/z-ai/glm-5 ($0.80/$2.56 per 1M tokens)
* feat(guardrails): team-based guardrail registration and approval workflow
Add team-based guardrail submission system where teams can register
Generic Guardrail API guardrails for admin review. Includes:
- POST /guardrails/register endpoint for team-scoped submissions
- Admin review endpoints (list/get/approve/reject submissions)
- Team Guardrails tab in the UI dashboard
- extra_headers support for forwarding client headers to guardrail APIs
- Prisma schema migration for status, submitted_at, reviewed_at fields
- Documentation for team-based guardrails and static/dynamic headers
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(guardrails): address review feedback - SSRF, silent failure, redundant query
- Validate api_base URL scheme (http/https only) and hostname in
register_guardrail to prevent SSRF via team submissions
- Return warning field in approve response when in-memory initialization
fails so admins know the guardrail won't work until next sync cycle
- Eliminate redundant DB query in list_guardrail_submissions by fetching
all team guardrails once and deriving both filtered list and summary
counts from the single result set
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(guardrails): add pending_review status guard to reject endpoint
Prevent rejecting already-active or already-rejected guardrails, which
would create a DB/memory inconsistency (active in memory but rejected
in DB). Now mirrors the approve endpoint's status check.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
OpenRouter supports the Responses API at /api/v1/responses with
encrypted_content for multi-turn stateless reasoning workflows.
Without native registration, requests fall through to the chat
completion bridge, which uses a different format (reasoning_details)
and drops encrypted_content entirely.
This adds OpenRouterResponsesAPIConfig to route requests directly to
OpenRouter's Responses API endpoint, preserving encrypted_content.
Fixes https://github.com/BerriAI/litellm/issues/22189
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
The _resolve_jwks_url method checks response.status_code != 200, but
MagicMock returns a MagicMock object for status_code which is always
truthy (!= 200). Explicitly set mock_response.status_code = 200 so the
tests exercise the intended code path.
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
mapped passthrough routes (vertex_ai, bedrock, etc) were compared
against the raw request path without prepending SERVER_ROOT_PATH.
db-registered routes already used _build_full_path_with_root for this
but the mapped routes branch was missed.
fixes#22272
DualCache.async_set_cache and async_set_cache_pipeline were missing the
default_in_memory_ttl injection that the sync set_cache method has. This
caused InMemoryCache to fall back to its own default_ttl (600s) instead
of using DualCache's configured default_in_memory_ttl (typically 60s).
This is particularly impactful for end-user budget enforcement in the
proxy, where cached spend values could remain stale for 10 minutes
instead of 1 minute, allowing users to exceed their budgets.
The response.completed handler in the completion→responses streaming
bridge was discarding the usage object, causing prompt_tokens_details
(and cached_tokens) to always be None when streaming with models that
use the Responses API (e.g. gpt-5.2-codex, gpt-5.3-codex).
Extract usage from the response.completed event and translate it via
the existing _transform_response_api_usage_to_chat_usage helper.
Fixes#22192
All 55 deepinfra models that had `supports_tool_choice: true` were
missing the `supports_function_calling` flag, causing
`litellm.supports_function_calling()` to incorrectly return False.
Fixes#22619
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* Add CrowdStrike AIDR guardrail hook
* fixup! use apply_guardrail event hook
* fixup! update imports
* fix(guardrails): include AI response in CrowdStrike AIDR output events
Issue:
_build_guard_input_for_response() was:
- Sending only the original user input (messages).
- Not sending the AI provider response.
This fix will:
- Extract response.choices from the ModelResponse object and include them in guard_input payload.
- Thus, ensure AIDR output rules receive the AI-generated content for analysis.
- Fix and update tests.
* fix(guardrails): prevent duplicate input events in CrowdStrike AIDR guardrail
Issue:
The CrowdStrike AIDR guardrail was running on during_call hooks wihtout event_hook configured.
This fix will:
- Set event_hook to ["pre_call", "post_call"] (AIDR admins will control what policy is applied)
This change will:
- Require default_on parameter
- Prevent duplicate API calls to AIDR for the same input
- Avoid unchecked AI provider API calls on during_call hook
* docs: add CrowdStrike AIDR to the list of Guardrails under Integrations
* docs: update CrowdStrike AIDR documentation page
---------
Co-authored-by: Konstantin Lapine <konstantin.lapine@crowdstrike.com>
Reorder elif branches so is_vertex_ai is checked before "gemini" in model.
Previously, Vertex AI Gemini models (e.g. vertex_ai/gemini-2.5-flash) matched
the "gemini" substring check first and were logged with the Google AI Studio
URL instead of the Vertex AI URL.