The _resolve_model_for_cost_lookup function was only checking
litellm_params.model when resolving model names from the router.
For Azure custom deployment names (e.g. azure/openai/gpt-5.3-codex),
this deployment name doesn't exist in the model cost map, so cost
returned /bin/zsh.
Now checks model_info.base_model and litellm_params.base_model first,
falling back to litellm_params.model only if no base_model is set.
This matches how the router resolves base_model everywhere else.
- Move Reset Spend button after Regenerate Key in header
- Make modal OK button danger style with text "Reset"
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Use TransactionOutlined icon and danger style to match Delete Key button
- Rewrite modal description: remove assumption about key being blocked,
clarify spend history is preserved in logs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The _safe_get_request_headers caching (commit e7175a52) uses
request.state._cached_headers. With Mock(spec=Request), getattr on
state returns a Mock (truthy), causing RedactedDict to receive a Mock
instead of a dict. Using a real starlette State object fixes this.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Address Greptile review: test_resolve_jwks_url_resolves_oidc_discovery_document
also used the inconsistent patch.object pattern.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Vertex AI / Gemini uses Pydantic's model_json_schema() which omits
additionalProperties: False (Gemini rejects it). The test expected
the same schema for all providers.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The patch.object with new_callable=AsyncMock can behave inconsistently
across Python versions, causing mock_response.status_code to return a
MagicMock instead of the assigned value. Direct assignment is simpler
and more reliable.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The CompletionTokensDetailsWrapper type now includes video_tokens field,
but this test's expected dict was not updated to include it.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The Gemini API requires role="user" on function_response content blocks
(added in commit 273cf12afa), but these tests were never updated to match.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Adds a "Reset Spend" button to the key detail view so proxy admins and team
admins can immediately reset a key's spend to $0, unblocking keys that have
hit their budget limit without waiting for the next scheduled budget reset.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Validate anchor is "created_at" in enforced_file_expires_after (matching
user-provided path). Add key existence validation to batch endpoint for
enforced_batch_output_expires_after.
Cleanup per review: this class attribute is no longer used after the
__next__ refactor to queue-based approach.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Two independent fixes for pre-existing test failures on main:
1. Anthropic streaming: The sync __next__ method used a simple
holding_chunk pattern that lost chunks when multiple events needed
to be returned. Refactored to use the same chunk_queue approach as
the async __anext__ method. Also fixed tests that used ModelResponse
(which defaults finish_reason to 'stop') instead of ModelResponseStream.
2. Azure GPT-5.1 logprobs: The base OpenAI class includes logprobs for
gpt-5.1+ models, but Azure hasn't verified support for gpt-5.1.
Added explicit removal of logprobs/top_logprobs for gpt-5.1 (non-5.2)
models in the Azure config.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace **kwargs with explicit tools and system parameters to match
the BaseTokenCounter.count_tokens abstract method signature.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Suppress noisy error log fired every cron tick when spend log cleanup
is simply not configured. _should_delete_spend_logs already logs the
specific reason at the right level (info for None, warning for
invalid value), so the redundant blanket error log in
cleanup_old_spend_logs is removed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Only release distributed lock in finally if it was actually acquired;
prevents spurious Redis release_lock calls on early returns
- Treat bare integer maximum_spend_logs_retention_period as days (e.g. 3 → "3d")
instead of silently failing with a ValueError
- Elevate "Skipping cleanup" log from info to error so misconfigured
retention settings are visible without verbose logging
- Add tests for all three fixes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add created_at field to MCPServer type (was missing)
- Map created_at from LiteLLM_MCPServerTable in build_mcp_server_from_table()
- Use server.created_at and server.updated_at instead of datetime.now() in _build_mcp_server_table() and health check table builder
- Add regression tests to verify timestamps are preserved through round-trip conversions
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>