Adds cleanup utilities for PROMETHEUS_MULTIPROC_DIR to prevent
unbounded RAM/disk growth from stale .db files in multi-worker setups.
Three-part lifecycle aligned with upstream prometheus_client docs:
1. Startup: wipe entire directory before workers fork (clean slate)
2. Shutdown: mark_process_dead() for own PID (removes gauge_live* only)
3. Periodic (hourly): scan for dead PIDs and call mark_process_dead()
Counter/histogram files are never individually deleted at runtime to
avoid partial counter resets that cause false spikes in rate()/increase().
Also auto-creates PROMETHEUS_MULTIPROC_DIR when prometheus callback is
configured with multiple workers and the env var is not already set.
Address Greptile review: apply type: ignore[union-attr] consistently
on all backend_ws.send(), .recv(), and .close() calls, not just the
three that CI flagged.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
CI MyPy resolves CLIENT_CONNECTION_CLASS as Optional[ClientConnection]
and flags .send() and .close() as attr-defined errors. These methods
exist at runtime on the websocket connection object. Add type: ignore
comments to unblock the linting CI.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace Prisma ORM count/find_many calls with two query_raw calls that
only project the key_alias column. The Prisma client wrapper does not
support SELECT projection via find_many, so raw SQL is used to keep
memory usage proportional to the page size rather than total key count.
Update tests to mock query_raw instead of count/find_many.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
LiteLLM_VerificationTokenActions.find_many() does not support the
select keyword argument. Remove it and drop the corresponding test.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
Add select={"key_alias": True} to the find_many call so only the alias
column is fetched from the database instead of full token rows. Add
five unit tests in test_key_management_endpoints.py covering response
shape, pagination skip/take computation, search filter injection,
absence of contains filter when no search term is given, and the
select-only-alias optimization.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
The /key/aliases endpoint previously fetched all key aliases from the database without limit, causing OOM crashes with large key sets. Added page, size, and search query parameters with database-level filtering to enable paginated and searchable key alias retrieval. Updated the response to include pagination metadata (total_count, current_page, total_pages, size) matching the /v2/model/info pattern.
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
Adds TestProxyMcpStatelessBehavior to test_proxy_mcp_e2e.py with a test
that verifies two independent MCP clients can connect, initialize, and
call tools without sharing session state. This catches the regression
from PR #19809 where stateless=False broke clients that don't manage
mcp-session-id headers.
Regression test for #20242
When extended thinking is enabled, the websearch interception agentic loop
builds a follow-up assistant message with only tool_use blocks. Anthropic's
API requires assistant messages to start with thinking/redacted_thinking
blocks when thinking is enabled, causing a 400 Bad Request.
Extract thinking blocks from the model's initial response, thread them
through the agentic loop, and prepend them to the follow-up assistant
message — matching the pattern used by anthropic_messages_pt in factory.py.
Fixes the error: "Expected 'thinking' or 'redacted_thinking', but found
'tool_use'"
The tests were mocking `filter_server_ids_by_ip` but the production
code in server.py now calls `filter_server_ids_by_ip_with_info` which
returns a (server_ids, blocked_count) tuple. Update all 8 mock sites
to use the correct method name and return signature.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When calling non-text-embedding-3 models routed through the openai provider
(e.g. nvidia/llama-3.2-nv-embedqa-1b-v2), passing `dimensions` previously
raised an UnsupportedParamsError unconditionally. This fix threads
`allowed_openai_params` through the embedding call stack so that providers
can opt-in to passing `dimensions` by including it in the list.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The old test assumed ArizePhoenixLogger reused the global TracerProvider.
With the nested traces fix, Phoenix now creates its own dedicated provider
and produces litellm_proxy_request + litellm_request + raw_gen_ai_request
spans independently.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
- ArizePhoenixLogger now creates spans on its own dedicated TracerProvider
instead of trying to reuse parent spans from the global otel TracerProvider
(which were invisible in Phoenix since they go to a different exporter)
- Auto-initialize ArizePhoenixLogger when otel callback is configured and
Phoenix env vars (PHOENIX_API_KEY, PHOENIX_COLLECTOR_*) are detected
- Use exact type check in get_custom_logger_compatible_class to prevent
ArizePhoenixLogger (subclass) from being returned when looking up otel
- Fix tool_permission guardrail to check non-function tools like
code_interpreter (previously skipped with `type != "function"`)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* Adjust input and output cost per token for mistral-small-2503
Cost per million for mistral-small-2503 is not correct.
In Azure Documentation:
Pay-as-you-go (per 1,000 tokens)
$0.0001
Model input
$0.0003
Model output
* Update input and output cost per token for model