- Change chunk["id"] to chunk.get("id") for compatibility with MiniMax
- ModelResponseStream auto-generates id when None is passed
- Add regression test test_chunk_parser_without_id_field
Replace bare _get_deployment_default_tpm/rpm_limit calls in the
async_log_success_event condition with get_key_model_tpm/rpm_limit
(model_name=model_group). The higher-level getters short-circuit on
key/team metadata hits before ever reaching the router, so requests
that don't use deployment defaults incur no extra router lookup. Remove
the now-unused bare helper imports.
Also fix invalid `int = None` type hints in test helper signatures
to `Optional[int] = None`.
Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
async_log_success_event only updated the per-model cache counter when
model_rpm_limit / model_tpm_limit were present in key metadata or
model_max_budget was set. For the new deployment-default path
(default_api_key_tpm_limit / default_api_key_rpm_limit), none of those
conditions held, so current_tpm stayed at zero and tpm enforcement was
never applied across multiple requests.
Extend the guard condition to also trigger when the model group has a
deployment-default tpm or rpm limit, and import the two helpers at
module level.
Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
Compute get_key_model_tpm/rpm_limit once before the guard condition
instead of calling each function twice (once to check non-None, once to
retrieve). Removes 2 extra llm_router.get_model_list() calls per request
when deployment defaults are active.
Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
- Use min() across all matching deployments instead of first-wins when
resolving default_api_key_tpm/rpm_limit for a model group, so
load-balanced setups with different per-deployment limits always apply
the most conservative value
- Replace the global SensitiveDataMasker non_sensitive_overrides change
with a targeted excluded_keys set at the remove_sensitive_info_from_deployment
call site, avoiding unintended suppression of other fields
- Update the v1 parallel request limiter to pass model_name to
get_key_model_tpm/rpm_limit so deployment defaults apply there too
- Add 4 tests covering multi-deployment min semantics
Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
Helicone (PRs #19288, #22603) and Langfuse (#22390) were present in the
v1.82.0-stable...v1.82.3-stable diff but omitted from the AI Integrations
logging section. Also updates the AI Integrations diff summary count from 2 to 4.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds `default_api_key_tpm_limit` and `default_api_key_rpm_limit` to
`GenericLiteLLMParams` so operators can set per-deployment rate limit
defaults in config.yaml. When a key has no model-specific tpm/rpm limit
configured, the proxy falls back to these deployment defaults (Case 2 in
spec). Key-level limits always take priority (Case 1).
- Extends `get_key_model_tpm_limit` / `get_key_model_rpm_limit` with a
`model_name` param and a priority-4 deployment-default fallback
- Passes `model_name=requested_model` in the parallel request limiter so
the fallback is triggered at enforcement time
- Adds `"limit"` to `SensitiveDataMasker` non-sensitive overrides so
`*_limit` fields are not masked in `/model/info` responses
- Adds 17 unit tests covering both spec cases and the `/model/info` path
Co-Authored-By: Claude (claude-sonnet-4-6) <noreply@anthropic.com>
- Add 'Contributing to Guardrails' category with links to:
- Generic Guardrail API (integrate without PR)
- Adding a New Guardrail Integration tutorial
- Adding Guardrail Support to Endpoints
- Add 'Team Bring-Your-Own Guardrails' link for team BYOG workflow
These docs existed but were only accessible from the 'LiteLLM AI Gateway'
sidebar. Now they're also accessible when browsing the 'Guardrail Providers'
section.
Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
- Only remove wildcard path from openai_routes when the route entry has
type="subpath", avoiding accidental removal when two endpoints share
the same base path but differ in include_subpath
- Clean up _registered_pass_through_routes in the test finally block to
prevent stale entries from polluting subsequent tests on failure
- Add dedup guard for base path registration (prevents unbounded list
growth on config reload)
- Clean up base path and wildcard path from openai_routes when an
endpoint is removed via remove_endpoint_routes
- Rewrite test to exercise initialize_pass_through_endpoints directly,
covering registration, dedup on reload, and cleanup on removal
When a pass-through endpoint has both auth=true and include_subpath=true,
non-admin users got 401 errors on subpath requests because only the base
path was registered in openai_routes. Now the wildcard path is also
registered so the auth check recognizes subpath requests as LLM API routes.
Also fixes pre-existing pyright error where logging_obj was possibly
unbound in the except block.
Add support for {"location": "tool_config"} in cache_control_injection_points,
which appends a cachePoint block to the Bedrock Converse toolConfig.tools array.
This enables prompt caching of tool definitions on Bedrock Claude models.
Also update the cache control hook to pass through non-message injection points
to provider-specific handling instead of silently dropping them.
Fixes#21969
OpenAI strict mode requires both additionalProperties:false AND all
property keys in required. Without required, OpenAI rejects the schema
even with additionalProperties:false set.
Enables Gemini 3+ models to combine built-in tools (Google Search, etc.)
with custom functions via `include_server_side_tool_invocations=True`.
Server-side invocations are surfaced in provider_specific_fields and
automatically re-injected on subsequent turns for multi-turn coherence.
Closes#24047
When translating Anthropic output_format to OpenAI response_format,
the adapter sets strict: true but didn't add additionalProperties: false,
which OpenAI requires at every object nesting level. This caused
BadRequestError for structured output requests routed to OpenAI models.
Fixes#20997
The check `content.get("thinking", None) is not None` incorrectly
drops thinking blocks when the `thinking` key is explicitly null or
absent. Changed to `content.get("type") == "thinking"` to match
the fix already applied in the experimental pass-through path (PR #15501).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add ExportOutlined icon next to nav items that link to external pages,
making it clear to users when a link opens in a new tab.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ChatUI stores endpointType as string but the narrowed prop expects
EndpointType — add explicit cast at the call site.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Deduplicate: update_key_fn now delegates to _get_and_validate_existing_key()
instead of inlining its own copy of the lookup logic
- Use _hash_token_if_needed (already imported at module level) instead of
inline `from proxy_server import hash_token` + manual conditional
- Fix stale docstring: _get_and_validate_existing_key raises ProxyException,
not HTTPException
- Add unit test: test_update_key_nonexistent_key_returns_404
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Matches the cast used in ChatUI.tsx — the react-syntax-highlighter
type definitions don't accept CSSProperties directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The /key/update endpoint's get_data() call raises a 401 when the body
`key` field doesn't exist in the DB, because get_data() treats the
token as an auth credential. This caused the auth layer to resolve the
body key instead of the Authorization header bearer token.
Replace prisma_client.get_data() with direct Prisma find_unique() in
both _get_and_validate_existing_key() and update_key_fn(), matching
the pattern used in the /key/block and /key/unblock fix (PR #23977).
Also fix the incorrect "Team not found" error message in update_key_fn.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Narrow endpointType prop from string to EndpointType enum
- Add missing test for MCP events on CHAT endpoint
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Extract the chat message bubble rendering (~165 lines) into a dedicated
ChatMessageBubble component with 15 Vitest tests covering all display branches.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Accept theirs for UI test file conflicts (not related to our changes).
Our key_management_endpoints.py merged cleanly.
Co-authored-by: yuneng-jiang <yuneng-jiang@users.noreply.github.com>
Incorporate new _check_key_admin_access() calls from the base branch
into block_key/unblock_key alongside our existence-check fix.
Update test mocks: replace references to removed get_key_object and
_cache_key_object with _delete_cache_key_object in both the shared
_setup_block_unblock_mocks helper and individual test functions.
Co-authored-by: yuneng-jiang <yuneng-jiang@users.noreply.github.com>