mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-15 23:31:29 +00:00
Some checks are pending
GitHub Actions Security Analysis / zizmor (push) Waiting to run
* feat(router): add separate ITPM/OTPM deployment rate limits Support input/output tokens per minute on deployments via enforce_model_rate_limits, with reservation, reconciliation, refund on failure, and rate-limit headers. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(router): keep ITPM/OTPM diff minimal in router.py Drop unrelated Black reformatting from router.py and types/router.py so the PR only contains functional ITPM/OTPM changes. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(router): make ITPM/OTPM limits separate and atomic Address Greptile review on separate ITPM/OTPM deployment rate limits. - OTPM is now reserved atomically pre-call with rollback, matching the ITPM path, so concurrent requests can no longer overshoot the configured output limit before reconciliation - ITPM counts input tokens only; it no longer accumulates completion tokens, so the input-token limit and x-ratelimit-limit-input-tokens header describe input usage as their names imply - _read_reservation_from_kwargs only falls back to litellm_params.metadata when the top-level metadata channel is absent, so production requests carrying a litellm_params.metadata dict still reconcile and refund their reservation Adds regression tests for OTPM atomicity under concurrency, input-only ITPM enforcement, and reservation lookup when litellm_params.metadata is present. * fix(router): subtract input tokens only from remaining-input-tokens header The in-flight replay for x-ratelimit-remaining-input-tokens subtracted total tokens (input + output) instead of input tokens only, so clients saw remaining input quota understated by the completion token count on every response. Now consistent with the input-only ITPM counter. * fix(router): make itpm/otpm vs tpm/rpm precedence explicit When a deployment configures itpm/otpm alongside tpm/rpm, the io-token path takes over and the tpm/rpm limits are not enforced. Log a warning the first time such a conflicting deployment is seen so the supersession is not silent, and document the mutual exclusivity. Post-call reconciliation now only trues up a counter that was actually reserved against, so the itpm/otpm keys are no longer incremented for deployments that never configured that limit. * fix(router): track actual io-token usage on the reservation-minute key Post-call reconciliation now keys off the exact cache key stashed at pre-call time rather than one recomputed from the response-time minute. This fixes two issues: a request whose pre-call estimate was 0 now still writes its actual billable input to the ITPM counter (previously it was skipped, leaving the limit unenforceable for that request), and a call that finishes in a later minute reconciles against the minute it reserved against instead of pushing a negative delta into the next minute. Counters are only touched when their limit is configured. * fix(router): run io-token reconciliation before the model_id guard async_log_success_event gated IO reconciliation behind the model_id guard that only the TPM tracking path needs. Since reconciliation works entirely from the cache keys stashed in kwargs, a success event whose standard_logging_object lacks model_id would skip reconciliation and leave the reservation on the counter until the TTL expired, wasting quota. Route the IO path first. * fix(router): don't replay in-flight delta for itpm/otpm headers For ITPM/OTPM model groups the counter is incremented at reservation time (pre-call), so the remaining values returned by get_remaining_model_group_usage already account for the current request. Replaying the in-flight delta on top double-counted it and understated x-ratelimit-remaining-input/output-tokens by up to max_tokens on every response. Skip the delta for io-token groups; the legacy TPM/RPM replay path is unchanged. * fix(router): clear io-token reservation after reconcile/refund async_io_token_refund_failure and async_io_token_reconcile_success now clear the stashed reservation keys from the request metadata once done. Otherwise, on a model group mixing IO-limited and non-IO deployments, a failed IO call that retries on a non-IO fallback left the stale sentinel in the shared request metadata; the fallback's success handler would divert into IO reconciliation against the already-refunded key, driving the ITPM counter negative and skipping the non-IO deployment's TPM tracking. * fix(router): tidy reservation channel lookup and header guard Consolidate the reservation channel lookup into a single ordered helper shared by read and clear, so top-level metadata always wins over litellm_params metadata without the tangled per-iteration fallback. Also stop gating the router rate-limit header block on the presence of x-ratelimit-remaining-input/output-tokens. That block only emits those headers for ITPM/OTPM groups; for a non-IO group backed by a provider that natively returns input/output token headers, the extra conditions suppressed the router's own remaining-tokens/requests headers. * fix(router): strip client-supplied io-token reservation keys The reservation sentinels (_litellm_itpm_reserved, _litellm_itpm_cache_key, and the otpm equivalents) are server-only, but metadata is caller-controlled on proxy requests. An authenticated caller could forge these fields with an arbitrary cache key so the post-call reconcile/refund path would decrement any deployment's ITPM/OTPM counter and let it exceed the configured limit. Strip the reserved keys from the request metadata in set_io_token_rate_limit_request_kwargs, which runs before the router stashes its own reservation, so only a genuine server-side reservation is ever read post-call. * fix(router): track TPM routing load for io-limited deployments deployment_callback_on_success early-returned for any deployment with itpm/otpm set, so its total-token usage never landed in the router's TPM routing counter. TPM-aware routing strategies then saw 0 load for IO deployments and over-routed to them in mixed model groups. Only skip tracking when neither tpm/rpm nor itpm/otpm are configured; itpm/otpm enforcement still runs separately in ModelRateLimitingCheck, so the routing counter and the enforcement counters stay independent. * fix(router): expose standard tpm/rpm headers for io-limited groups get_remaining_model_group_usage returned early for ITPM/OTPM groups, so a group that also set tpm/rpm never emitted x-ratelimit-remaining-tokens / -requests; clients and prometheus gauges reading those saw no data. Build both header sets instead of returning early. Also simplify the in-flight header replay: only the tpm/rpm counters are incremented post-response, so the delta now adjusts just those. The itpm/otpm counters are incremented at reservation time (pre-call), so the input/output token headers already reflect the request and are left untouched - which removes the need for the separate io-group special case. * fix(router): roll back ITPM on any OTPM reservation error; dedup warning per instance Two follow-ups from review. The pre-call OTPM reservation only rolled back the ITPM reservation on a RateLimitError, so a transient cache error while reserving OTPM left the ITPM counter inflated until the TTL expired; catch any exception, release the ITPM reservation, then re-raise. Replace the module-level lru_cache warn-once (caching a logging side effect, which never re-warns in a long-lived process) with an instance-scoped set of already-warned deployment ids on ModelRateLimitingCheck. * fix(router): always clear reservation stash on reconcile; don't collapse id-less warning dedup Clear the reservation in a finally block so a mid-reconciliation cache error still removes the stash and a duplicate success event can't re-process it. Dedup the itpm/otpm-vs-tpm/rpm conflict warning per real deployment id; a deployment with no id no longer collapses every id-less deployment onto the str(None) key (which would suppress all but the first warning). * fix(router): skip io reservation when deployment can't be keyed _get_cache_keys returned a shared 'global_router:None:None:...' key when a deployment was missing model_info.id or litellm_params.model, so misconfigured deployments could share one rate-limit bucket. Return None in that case and skip io reservation for the request. * fix(router): honor explicit max_tokens=0 in io reservation _resolve_max_tokens used 'max_tokens or max_completion_tokens', so an explicit max_tokens=0 fell through to the model default. Only fall back to max_completion_tokens when max_tokens is absent. * fix(ci): satisfy lint budget, router coverage, and dashboard schema sync - Modernize the new itpm/otpm module's type hints to PEP 585 lowercase generics (Dict/Tuple/List -> dict/tuple/list) to clear the added UP006 violations; ratchet ruff-strict-budget.json's UP006 ceiling down to match. - Replace three try/except Exception blocks that must stay broad by design (token_counter and litellm.get_model_info raise untyped exceptions, and an io-token refund failure must never break the logging pipeline) with contextlib.suppress(Exception), matching the codebase's existing resolution for this exact BLE001 pattern. - Add direct unit tests for get_model_group_io_token_usage (multi-deployment aggregation and the empty-model-list case) in test_router_helper_utils.py, satisfying the router function-coverage check. - Regenerate the dashboard's schema.d.ts so the new itpm/otpm fields on GenericLiteLLMParams and ModelGroupInfo are reflected in the OpenAPI types. * fix: enforce io token rate limits consistently * fix: honor zero max tokens in otpm reservation * fix(lint): fix UP007 violation and resync ruff-strict-budget.json to base Convert Union[_Span, Any] to _Span | Any (safe on this repo's Python >=3.10 floor) to clear the new UP007 violation from the TYPE_CHECKING-gated Span alias. The previously committed ruff-strict-budget.json ratcheted UP006 down from a stale base; litellm_internal_staging has since tightened that same ceiling further on its own. Reset the file to the current base's committed values and re-ratchet from there so the budget only ever moves down relative to the actual merge-base, never against a stale snapshot. * fix(router): attach ITPM/OTPM headers on dict responses and harden reservation Strip itpm/otpm from provider kwargs, ensure messages are available for ITPM estimation, honor max_output_tokens on /v1/responses, and propagate rate-limit headers through /v1/messages dict responses via _hidden_params. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(router): attach ITPM/OTPM headers to streaming /v1/messages responses Wrap bare async iterators in HiddenParamsAsyncIteratorWrapper so set_response_headers can attach rate-limit headers to streaming Anthropic messages responses that lack a _hidden_params slot. Co-authored-by: Cursor <cursoragent@cursor.com> * style: ruff format add_retry_fallback_headers.py Fix CI ruff format check failure on get_hidden_params_dict call site. Co-authored-by: Cursor <cursoragent@cursor.com> * refactor(router): extract set_response_headers helpers to fix C901 budget Move header-attachment logic into add_retry_fallback_headers helpers so set_response_headers stays under the strict complexity ceiling. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: keep IO token reservation when response usage is missing Missing usage was reconciled as zero and fully refunded the pre-call reservation, allowing limit bypass on repeated successful calls. Only adjust counters when usage is resolved from the response or standard logging fields; otherwise keep the reservation until TTL expires. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: enforce RPM/TPM alongside IO-token limits on mixed deployments Deployments with both itpm/otpm and tpm/rpm previously returned after the IO reservation and skipped RPM/TPM checks. Run both paths and refund the IO reservation only when RPM/TPM rejects after a successful reservation. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: track TPM usage on success for mixed IO+TPM deployments The early return after IO-token reconciliation in log_success_event and async_log_success_event skipped the TPM counter increment, so the tpm_key the pre-call check reads was never written and tpm_limit was never actually enforced on deployments that also configure itpm/otpm. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: treat total-only usage as unresolved in IO-token reconcile usage/standard_logging_object entries carrying only total_tokens (no prompt/completion or input/output breakdown) were treated as resolved usage, resolving to (0, 0) and refunding the full reservation. Both _usage_is_present and the standard_logging_object fallback now require an actual input/output breakdown before reconciling, keeping the reservation otherwise. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: reserve minimal token when input/output estimation fails _reservation_value(0, limit) reserved the entire limit whenever token estimation failed (empty/unsupported input, tokenizer error), letting one such request claim the whole bucket and 429 every concurrent request to the deployment until it completed. Reserve 1 token instead so estimation failures no longer serialize traffic. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: refund IO reservation synchronously before retry deployment pick On retry, set_io_token_rate_limit_request_kwargs clears reservation sentinels from the shared kwargs dict before a background failure handler can refund them, stranding the counter until TTL. Refund and clear any stale reservation in _update_kwargs_with_deployment before stripping sentinels for the next attempt. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(io_token_rate_limit_check): use model-specific tokenizer for ITPM estimate; document sync-refund Redis ceiling Pass the deployment litellm_params.model to token_counter so it uses the model's native tokenizer instead of the generic fallback, narrowing the reservation over/under-estimate window between pre-call and post-call reconcile. Add a ponytail: comment to refund_stale_reservation_before_retry explaining the known ceiling: the synchronous DualCache.increment_cache issues a blocking Redis INCR when a Redis backend is configured. This only fires on streaming mid-stream retries (non-streaming failures await their failure handler before the retry picks a new deployment, leaving no sentinels to refund). Upgrade path: make _update_kwargs_with_deployment async. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
2835 lines
97 KiB
Python
2835 lines
97 KiB
Python
import sys
|
|
import os
|
|
import traceback
|
|
from dotenv import load_dotenv
|
|
from fastapi import Request
|
|
from datetime import datetime
|
|
|
|
sys.path.insert(
|
|
0, os.path.abspath("../..")
|
|
) # Adds the parent directory to the system path
|
|
from litellm import Router
|
|
import pytest
|
|
import litellm
|
|
from unittest.mock import patch, MagicMock, AsyncMock
|
|
from create_mock_standard_logging_payload import create_standard_logging_payload
|
|
from litellm.types.utils import StandardLoggingPayload
|
|
from litellm.types.router import Deployment, LiteLLM_Params, ModelInfo
|
|
|
|
|
|
@pytest.fixture
|
|
def model_list():
|
|
return [
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {
|
|
"model": "gpt-5-mini",
|
|
"api_key": os.getenv("OPENAI_API_KEY"),
|
|
"tpm": 1000, # Add TPM limit so async method doesn't return early
|
|
"rpm": 100, # Add RPM limit so async method doesn't return early
|
|
},
|
|
"model_info": {
|
|
"access_groups": ["group1", "group2"],
|
|
},
|
|
},
|
|
{
|
|
"model_name": "gpt-5.5",
|
|
"litellm_params": {
|
|
"model": "gpt-5.5",
|
|
"api_key": os.getenv("OPENAI_API_KEY"),
|
|
},
|
|
},
|
|
{
|
|
"model_name": "gpt-image-1",
|
|
"litellm_params": {
|
|
"model": "gpt-image-1",
|
|
"api_key": os.getenv("OPENAI_API_KEY"),
|
|
},
|
|
},
|
|
{
|
|
"model_name": "*",
|
|
"litellm_params": {
|
|
"model": "openai/*",
|
|
"api_key": os.getenv("OPENAI_API_KEY"),
|
|
},
|
|
},
|
|
{
|
|
"model_name": "claude-*",
|
|
"litellm_params": {
|
|
"model": "anthropic/*",
|
|
"api_key": os.getenv("ANTHROPIC_API_KEY"),
|
|
},
|
|
},
|
|
]
|
|
|
|
|
|
def test_validate_fallbacks(model_list):
|
|
router = Router(model_list=model_list, fallbacks=[{"gpt-5.5": "gpt-5-mini"}])
|
|
router.validate_fallbacks(fallback_param=[{"gpt-5.5": "gpt-5-mini"}])
|
|
|
|
|
|
def test_routing_strategy_init(model_list):
|
|
"""Test if all routing strategies are initialized correctly"""
|
|
from litellm.types.router import RoutingStrategy
|
|
|
|
router = Router(model_list=model_list)
|
|
for strategy in RoutingStrategy:
|
|
router.routing_strategy_init(
|
|
routing_strategy=strategy, routing_strategy_args={}
|
|
)
|
|
|
|
|
|
def test_routing_strategy_init_invalid_strategy(model_list):
|
|
"""Test that invalid routing_strategy raises ValueError with helpful message.
|
|
|
|
See: https://github.com/BerriAI/litellm/issues/11330
|
|
Invalid strategies like 'simple' (without '-shuffle') should fail fast
|
|
with a clear error, not silently cause 'No deployments available' errors.
|
|
"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Test common mistake: "simple" instead of "simple-shuffle"
|
|
with pytest.raises(ValueError) as exc_info:
|
|
router.routing_strategy_init(
|
|
routing_strategy="simple", routing_strategy_args={}
|
|
)
|
|
|
|
# Verify error message is helpful
|
|
error_msg = str(exc_info.value)
|
|
assert "Invalid routing_strategy" in error_msg
|
|
assert "simple" in error_msg
|
|
assert "simple-shuffle" in error_msg # Suggests the correct option
|
|
# Verify error message tells user WHERE to fix it
|
|
assert "config.yaml" in error_msg
|
|
assert "router_settings.routing_strategy" in error_msg
|
|
assert "Router SDK" in error_msg
|
|
|
|
# Test completely invalid strategy
|
|
with pytest.raises(ValueError) as exc_info:
|
|
router.routing_strategy_init(
|
|
routing_strategy="not-a-real-strategy", routing_strategy_args={}
|
|
)
|
|
assert "Invalid routing_strategy" in str(exc_info.value)
|
|
|
|
|
|
def test_routing_strategy_init_valid_string_strategies(model_list):
|
|
"""Test that all valid string routing strategies work without error.
|
|
|
|
Valid strategies are derived from RoutingStrategy enum values plus 'simple-shuffle'.
|
|
"""
|
|
from litellm.types.router import RoutingStrategy
|
|
|
|
router = Router(model_list=model_list)
|
|
|
|
# All strategies from enum + simple-shuffle (default, not in enum)
|
|
valid_strategies = ["simple-shuffle"] + [s.value for s in RoutingStrategy]
|
|
|
|
for strategy in valid_strategies:
|
|
# Should not raise
|
|
router.routing_strategy_init(
|
|
routing_strategy=strategy, routing_strategy_args={}
|
|
)
|
|
|
|
|
|
def test_routing_strategy_init_valid_enum_strategies(model_list):
|
|
"""Test that RoutingStrategy enum values work without error."""
|
|
from litellm.types.router import RoutingStrategy
|
|
|
|
router = Router(model_list=model_list)
|
|
|
|
for strategy in RoutingStrategy:
|
|
# Should not raise when passing enum directly
|
|
router.routing_strategy_init(
|
|
routing_strategy=strategy, routing_strategy_args={}
|
|
)
|
|
|
|
|
|
def test_print_deployment(model_list):
|
|
"""Test if the api key is masked correctly"""
|
|
|
|
router = Router(model_list=model_list)
|
|
deployment = {
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {
|
|
"model": "gpt-5-mini",
|
|
"api_key": os.getenv("OPENAI_API_KEY"),
|
|
},
|
|
}
|
|
printed_deployment = router.print_deployment(deployment)
|
|
assert 10 * "*" in printed_deployment["litellm_params"]["api_key"]
|
|
|
|
|
|
def test_print_deployment_with_redact_enabled(model_list):
|
|
"""Test if sensitive credentials are masked when redact_user_api_key_info is enabled"""
|
|
import litellm
|
|
|
|
router = Router(model_list=model_list)
|
|
deployment = {
|
|
"model_name": "bedrock-claude",
|
|
"litellm_params": {
|
|
"model": "bedrock/anthropic.claude-v2",
|
|
"aws_access_key_id": "AKIAIOSFODNN7EXAMPLE",
|
|
"aws_secret_access_key": "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
|
|
"aws_region_name": "us-west-2",
|
|
},
|
|
}
|
|
|
|
original_setting = litellm.redact_user_api_key_info
|
|
try:
|
|
litellm.redact_user_api_key_info = True
|
|
printed_deployment = router.print_deployment(deployment)
|
|
|
|
assert "*" in printed_deployment["litellm_params"]["aws_access_key_id"]
|
|
assert "*" in printed_deployment["litellm_params"]["aws_secret_access_key"]
|
|
assert "us-west-2" == printed_deployment["litellm_params"]["aws_region_name"]
|
|
finally:
|
|
litellm.redact_user_api_key_info = original_setting
|
|
|
|
|
|
def test_completion(model_list):
|
|
"""Test if the completion function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
response = router._completion(
|
|
model="gpt-5-mini",
|
|
messages=[{"role": "user", "content": "Hello, how are you?"}],
|
|
mock_response="I'm fine, thank you!",
|
|
)
|
|
assert response["choices"][0]["message"]["content"] == "I'm fine, thank you!"
|
|
|
|
|
|
@pytest.mark.parametrize("sync_mode", [True, False])
|
|
@pytest.mark.flaky(retries=6, delay=1)
|
|
@pytest.mark.asyncio
|
|
async def test_image_generation(model_list, sync_mode):
|
|
"""Test if the underlying '_image_generation' function is working correctly"""
|
|
from litellm.types.utils import ImageResponse
|
|
|
|
router = Router(model_list=model_list)
|
|
if sync_mode:
|
|
response = router._image_generation(
|
|
model="gpt-image-1",
|
|
prompt="A cute baby sea otter",
|
|
)
|
|
else:
|
|
response = await router._aimage_generation(
|
|
model="gpt-image-1",
|
|
prompt="A cute baby sea otter",
|
|
)
|
|
|
|
ImageResponse.model_validate(response)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_router_acompletion_util(model_list):
|
|
"""Test if the underlying '_acompletion' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
response = await router._acompletion(
|
|
model="gpt-5-mini",
|
|
messages=[{"role": "user", "content": "Hello, how are you?"}],
|
|
mock_response="I'm fine, thank you!",
|
|
)
|
|
assert response["choices"][0]["message"]["content"] == "I'm fine, thank you!"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_router_abatch_completion_one_model_multiple_requests_util(model_list):
|
|
"""Test if the 'abatch_completion_one_model_multiple_requests' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
response = await router.abatch_completion_one_model_multiple_requests(
|
|
model="gpt-5-mini",
|
|
messages=[
|
|
[{"role": "user", "content": "Hello, how are you?"}],
|
|
[{"role": "user", "content": "Hello, how are you?"}],
|
|
],
|
|
mock_response="I'm fine, thank you!",
|
|
)
|
|
print(response)
|
|
assert response[0]["choices"][0]["message"]["content"] == "I'm fine, thank you!"
|
|
assert response[1]["choices"][0]["message"]["content"] == "I'm fine, thank you!"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_router_schedule_acompletion(model_list):
|
|
"""Test if the 'schedule_acompletion' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
response = await router.schedule_acompletion(
|
|
model="gpt-5-mini",
|
|
messages=[{"role": "user", "content": "Hello, how are you?"}],
|
|
mock_response="I'm fine, thank you!",
|
|
priority=1,
|
|
)
|
|
assert response["choices"][0]["message"]["content"] == "I'm fine, thank you!"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_router_schedule_atext_completion(model_list):
|
|
"""Test if the 'schedule_atext_completion' function is working correctly"""
|
|
from litellm.types.utils import TextCompletionResponse
|
|
|
|
router = Router(model_list=model_list)
|
|
with patch.object(
|
|
router, "_atext_completion", AsyncMock()
|
|
) as mock_atext_completion:
|
|
mock_atext_completion.return_value = TextCompletionResponse()
|
|
response = await router.atext_completion(
|
|
model="gpt-5-mini",
|
|
prompt="Hello, how are you?",
|
|
priority=1,
|
|
)
|
|
mock_atext_completion.assert_awaited_once()
|
|
assert "priority" not in mock_atext_completion.call_args.kwargs
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_router_schedule_factory(model_list):
|
|
"""Test if the 'schedule_atext_completion' function is working correctly"""
|
|
from litellm.types.utils import TextCompletionResponse
|
|
|
|
router = Router(model_list=model_list)
|
|
with patch.object(
|
|
router, "_atext_completion", AsyncMock()
|
|
) as mock_atext_completion:
|
|
mock_atext_completion.return_value = TextCompletionResponse()
|
|
response = await router._schedule_factory(
|
|
model="gpt-5-mini",
|
|
args=(
|
|
"gpt-5-mini",
|
|
"Hello, how are you?",
|
|
),
|
|
priority=1,
|
|
kwargs={},
|
|
original_function=router.atext_completion,
|
|
)
|
|
mock_atext_completion.assert_awaited_once()
|
|
assert "priority" not in mock_atext_completion.call_args.kwargs
|
|
|
|
|
|
@pytest.mark.parametrize("sync_mode", [True, False])
|
|
@pytest.mark.asyncio
|
|
async def test_router_function_with_fallbacks(model_list, sync_mode):
|
|
"""Test if the router 'async_function_with_fallbacks' + 'function_with_fallbacks' are working correctly"""
|
|
router = Router(model_list=model_list)
|
|
data = {
|
|
"model": "gpt-5-mini",
|
|
"messages": [{"role": "user", "content": "Hello, how are you?"}],
|
|
"mock_response": "I'm fine, thank you!",
|
|
"num_retries": 0,
|
|
}
|
|
if sync_mode:
|
|
response = router.function_with_fallbacks(
|
|
original_function=router._completion,
|
|
**data,
|
|
)
|
|
else:
|
|
response = await router.async_function_with_fallbacks(
|
|
original_function=router._acompletion,
|
|
**data,
|
|
)
|
|
assert response.choices[0].message.content == "I'm fine, thank you!"
|
|
|
|
|
|
@pytest.mark.parametrize("sync_mode", [True, False])
|
|
@pytest.mark.asyncio
|
|
async def test_router_function_with_retries(model_list, sync_mode):
|
|
"""Test if the router 'async_function_with_retries' + 'function_with_retries' are working correctly"""
|
|
router = Router(model_list=model_list)
|
|
data = {
|
|
"model": "gpt-5-mini",
|
|
"messages": [{"role": "user", "content": "Hello, how are you?"}],
|
|
"mock_response": "I'm fine, thank you!",
|
|
"num_retries": 0,
|
|
}
|
|
response = await router.async_function_with_retries(
|
|
original_function=router._acompletion,
|
|
**data,
|
|
)
|
|
|
|
assert response.choices[0].message.content == "I'm fine, thank you!"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_router_make_call(model_list):
|
|
"""Test if the router 'make_call' function is working correctly"""
|
|
|
|
## ACOMPLETION
|
|
router = Router(model_list=model_list)
|
|
response = await router.make_call(
|
|
original_function=router._acompletion,
|
|
model="gpt-5-mini",
|
|
messages=[{"role": "user", "content": "Hello, how are you?"}],
|
|
mock_response="I'm fine, thank you!",
|
|
)
|
|
assert response.choices[0].message.content == "I'm fine, thank you!"
|
|
|
|
## ATEXT_COMPLETION
|
|
response = await router.make_call(
|
|
original_function=router._atext_completion,
|
|
model="gpt-5-mini",
|
|
prompt="Hello, how are you?",
|
|
mock_response="I'm fine, thank you!",
|
|
)
|
|
assert response.choices[0].text == "I'm fine, thank you!"
|
|
|
|
## AEMBEDDING
|
|
response = await router.make_call(
|
|
original_function=router._aembedding,
|
|
model="gpt-5-mini",
|
|
input="Hello, how are you?",
|
|
mock_response=[0.1, 0.2, 0.3],
|
|
)
|
|
assert response.data[0].embedding == [0.1, 0.2, 0.3]
|
|
|
|
## AIMAGE_GENERATION
|
|
response = await router.make_call(
|
|
original_function=router._aimage_generation,
|
|
model="gpt-image-1",
|
|
prompt="A cute baby sea otter",
|
|
mock_response="https://example.com/image.png",
|
|
)
|
|
assert response.data[0].url == "https://example.com/image.png"
|
|
|
|
|
|
def test_update_kwargs_with_deployment(model_list):
|
|
"""Test if the '_update_kwargs_with_deployment' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
kwargs: dict = {"metadata": {}}
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
router._update_kwargs_with_deployment(
|
|
deployment=deployment,
|
|
kwargs=kwargs,
|
|
)
|
|
set_fields = ["deployment", "api_base", "model_info"]
|
|
assert all(field in kwargs["metadata"] for field in set_fields)
|
|
|
|
|
|
def test_update_kwargs_with_default_litellm_params(model_list):
|
|
"""Test if the '_update_kwargs_with_default_litellm_params' function is working correctly"""
|
|
router = Router(
|
|
model_list=model_list,
|
|
default_litellm_params={"api_key": "test", "metadata": {"key": "value"}},
|
|
)
|
|
kwargs: dict = {"metadata": {"key2": "value2"}}
|
|
router._update_kwargs_with_default_litellm_params(kwargs=kwargs)
|
|
assert kwargs["api_key"] == "test"
|
|
assert kwargs["metadata"]["key"] == "value"
|
|
assert kwargs["metadata"]["key2"] == "value2"
|
|
|
|
|
|
def test_get_timeout(model_list):
|
|
"""Test if the '_get_timeout' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
timeout = router._get_timeout(kwargs={}, data={"timeout": 100})
|
|
assert timeout == 100
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"fallback_kwarg, expected_error",
|
|
[
|
|
("mock_testing_fallbacks", litellm.InternalServerError),
|
|
("mock_testing_context_fallbacks", litellm.ContextWindowExceededError),
|
|
("mock_testing_content_policy_fallbacks", litellm.ContentPolicyViolationError),
|
|
],
|
|
)
|
|
def test_handle_mock_testing_fallbacks(model_list, fallback_kwarg, expected_error):
|
|
"""Test if the '_handle_mock_testing_fallbacks' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
with pytest.raises(expected_error):
|
|
data = {
|
|
fallback_kwarg: True,
|
|
}
|
|
router._handle_mock_testing_fallbacks(
|
|
kwargs=data,
|
|
)
|
|
|
|
|
|
def test_handle_mock_testing_rate_limit_error(model_list):
|
|
"""Test if the '_handle_mock_testing_rate_limit_error' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
with pytest.raises(litellm.RateLimitError):
|
|
data = {
|
|
"mock_testing_rate_limit_error": True,
|
|
}
|
|
router._handle_mock_testing_rate_limit_error(
|
|
kwargs=data,
|
|
)
|
|
|
|
|
|
def test_get_fallback_model_group_from_fallbacks(model_list):
|
|
"""Test if the '_get_fallback_model_group_from_fallbacks' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
fallback_model_group_name = router._get_fallback_model_group_from_fallbacks(
|
|
model_group="gpt-5.5",
|
|
fallbacks=[{"gpt-5.5": "gpt-5-mini"}],
|
|
)
|
|
assert fallback_model_group_name == "gpt-5-mini"
|
|
|
|
|
|
@pytest.mark.parametrize("sync_mode", [True, False])
|
|
@pytest.mark.asyncio
|
|
async def test_deployment_callback_on_success(sync_mode):
|
|
"""Test if the '_deployment_callback_on_success' function is working correctly"""
|
|
import time
|
|
|
|
model_list = [
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {
|
|
"model": "gpt-5-mini",
|
|
"api_key": os.getenv("OPENAI_API_KEY"),
|
|
"rpm": 100,
|
|
},
|
|
"model_info": {"id": "100"},
|
|
}
|
|
]
|
|
router = Router(model_list=model_list)
|
|
# Get the actual deployment ID that was generated
|
|
gpt_deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
deployment_id = gpt_deployment["model_info"]["id"]
|
|
|
|
standard_logging_payload = create_standard_logging_payload()
|
|
standard_logging_payload["total_tokens"] = 100
|
|
standard_logging_payload["model_id"] = "100"
|
|
kwargs = {
|
|
"litellm_params": {
|
|
"metadata": {
|
|
"model_group": "gpt-5-mini",
|
|
},
|
|
"model_info": {"id": deployment_id},
|
|
},
|
|
"standard_logging_object": standard_logging_payload,
|
|
}
|
|
response = litellm.ModelResponse(
|
|
model="gpt-5-mini",
|
|
usage={"total_tokens": 100},
|
|
)
|
|
if sync_mode:
|
|
tpm_key = router.sync_deployment_callback_on_success(
|
|
kwargs=kwargs,
|
|
completion_response=response,
|
|
start_time=time.time(),
|
|
end_time=time.time(),
|
|
)
|
|
else:
|
|
tpm_key = await router.deployment_callback_on_success(
|
|
kwargs=kwargs,
|
|
completion_response=response,
|
|
start_time=time.time(),
|
|
end_time=time.time(),
|
|
)
|
|
assert tpm_key is not None
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_deployment_callback_on_success_tracks_tpm_for_io_deployment():
|
|
"""
|
|
An IO-limited deployment (itpm/otpm, no tpm/rpm) must still record TPM usage
|
|
in the router's routing counter so TPM-aware routing strategies see its real
|
|
load in mixed model groups; its itpm/otpm enforcement runs separately.
|
|
"""
|
|
import time
|
|
|
|
model_list = [
|
|
{
|
|
"model_name": "opus",
|
|
"litellm_params": {
|
|
"model": "openai/gpt-4o-mini",
|
|
"api_key": "sk-fake",
|
|
"itpm": 1000,
|
|
},
|
|
"model_info": {"id": "io-100"},
|
|
}
|
|
]
|
|
router = Router(model_list=model_list)
|
|
|
|
standard_logging_payload = create_standard_logging_payload()
|
|
standard_logging_payload["total_tokens"] = 100
|
|
standard_logging_payload["model_id"] = "io-100"
|
|
kwargs = {
|
|
"litellm_params": {
|
|
"metadata": {
|
|
"deployment": "openai/gpt-4o-mini",
|
|
"model_group": "opus",
|
|
},
|
|
"model_info": {"id": "io-100"},
|
|
},
|
|
"standard_logging_object": standard_logging_payload,
|
|
}
|
|
response = litellm.ModelResponse(model="openai/gpt-4o-mini", usage={"total_tokens": 100})
|
|
|
|
tpm_key = await router.deployment_callback_on_success(
|
|
kwargs=kwargs,
|
|
completion_response=response,
|
|
start_time=time.time(),
|
|
end_time=time.time(),
|
|
)
|
|
|
|
# The IO deployment is no longer skipped: its TPM routing counter is tracked.
|
|
assert tpm_key is not None
|
|
assert await router.cache.async_get_cache(key=tpm_key) == 100
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_deployment_callback_on_failure(model_list):
|
|
"""Test if the '_deployment_callback_on_failure' function is working correctly"""
|
|
import time
|
|
|
|
router = Router(model_list=model_list)
|
|
kwargs = {
|
|
"litellm_params": {
|
|
"metadata": {
|
|
"model_group": "gpt-5-mini",
|
|
},
|
|
"model_info": {"id": 100},
|
|
},
|
|
}
|
|
result = router.deployment_callback_on_failure(
|
|
kwargs=kwargs,
|
|
completion_response=None,
|
|
start_time=time.time(),
|
|
end_time=time.time(),
|
|
)
|
|
assert isinstance(result, bool)
|
|
assert result is False
|
|
|
|
model_response = router.completion(
|
|
model="gpt-5-mini",
|
|
messages=[{"role": "user", "content": "Hello, how are you?"}],
|
|
mock_response="I'm fine, thank you!",
|
|
)
|
|
result = await router.async_deployment_callback_on_failure(
|
|
kwargs=kwargs,
|
|
completion_response=model_response,
|
|
start_time=time.time(),
|
|
end_time=time.time(),
|
|
)
|
|
|
|
|
|
def test_deployment_callback_respects_cooldown_time(model_list):
|
|
"""Ensure per-model cooldown_time is honored even when exception headers are present."""
|
|
import httpx
|
|
import time
|
|
from unittest.mock import patch
|
|
|
|
router = Router(model_list=model_list)
|
|
|
|
class FakeException(Exception):
|
|
def __init__(self):
|
|
self.status_code = 429
|
|
self.headers = httpx.Headers({"x-test": "1"})
|
|
|
|
kwargs = {
|
|
"exception": FakeException(),
|
|
"litellm_params": {
|
|
"metadata": {"model_group": "gpt-5-mini"},
|
|
"model_info": {"id": 100},
|
|
"cooldown_time": 0,
|
|
},
|
|
}
|
|
|
|
with patch("litellm.router._set_cooldown_deployments") as mock_set:
|
|
router.deployment_callback_on_failure(
|
|
kwargs=kwargs,
|
|
completion_response=None,
|
|
start_time=time.time(),
|
|
end_time=time.time(),
|
|
)
|
|
|
|
mock_set.assert_called_once()
|
|
assert mock_set.call_args.kwargs["time_to_cooldown"] == 0
|
|
|
|
|
|
def test_log_retry(model_list):
|
|
"""Test if the '_log_retry' function is working correctly"""
|
|
import time
|
|
|
|
router = Router(model_list=model_list)
|
|
new_kwargs = router.log_retry(
|
|
kwargs={"metadata": {}},
|
|
e=Exception(),
|
|
)
|
|
assert "metadata" in new_kwargs
|
|
assert "previous_models" in new_kwargs["metadata"]
|
|
|
|
|
|
def test_update_usage(model_list):
|
|
"""Test if the '_update_usage' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
deployment_id = deployment["model_info"]["id"]
|
|
request_count = router._update_usage(
|
|
deployment_id=deployment_id, parent_otel_span=None
|
|
)
|
|
assert request_count == 1
|
|
|
|
request_count = router._update_usage(
|
|
deployment_id=deployment_id, parent_otel_span=None
|
|
)
|
|
|
|
assert request_count == 2
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"finish_reason, expected_fallback", [("content_filter", True), ("stop", False)]
|
|
)
|
|
@pytest.mark.parametrize("fallback_type", ["model-specific", "default"])
|
|
def test_should_raise_content_policy_error(
|
|
model_list, finish_reason, expected_fallback, fallback_type
|
|
):
|
|
"""Test if the '_should_raise_content_policy_error' function is working correctly"""
|
|
router = Router(
|
|
model_list=model_list,
|
|
default_fallbacks=["gpt-5.5"] if fallback_type == "default" else None,
|
|
)
|
|
|
|
assert (
|
|
router._should_raise_content_policy_error(
|
|
model="gpt-5-mini",
|
|
response=litellm.ModelResponse(
|
|
model="gpt-5-mini",
|
|
choices=[
|
|
{
|
|
"finish_reason": finish_reason,
|
|
"message": {"content": "I'm fine, thank you!"},
|
|
}
|
|
],
|
|
usage={"total_tokens": 100},
|
|
),
|
|
kwargs={
|
|
"content_policy_fallbacks": (
|
|
[{"gpt-5-mini": "gpt-5.5"}]
|
|
if fallback_type == "model-specific"
|
|
else None
|
|
)
|
|
},
|
|
)
|
|
is expected_fallback
|
|
)
|
|
|
|
|
|
def test_get_healthy_deployments(model_list):
|
|
"""Test if the '_get_healthy_deployments' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployments = router._get_healthy_deployments(
|
|
model="gpt-5-mini", parent_otel_span=None
|
|
)
|
|
assert len(deployments) > 0
|
|
|
|
|
|
@pytest.mark.parametrize("sync_mode", [True, False])
|
|
@pytest.mark.asyncio
|
|
async def test_routing_strategy_pre_call_checks(model_list, sync_mode):
|
|
"""Test if the '_routing_strategy_pre_call_checks' function is working correctly"""
|
|
from litellm.integrations.custom_logger import CustomLogger
|
|
from litellm.litellm_core_utils.litellm_logging import Logging
|
|
|
|
callback = CustomLogger()
|
|
litellm.callbacks = [callback]
|
|
|
|
router = Router(model_list=model_list)
|
|
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
|
|
litellm_logging_obj = Logging(
|
|
model="gpt-5-mini",
|
|
messages=[{"role": "user", "content": "hi"}],
|
|
stream=False,
|
|
call_type="acompletion",
|
|
litellm_call_id="1234",
|
|
start_time=datetime.now(),
|
|
function_id="1234",
|
|
)
|
|
if sync_mode:
|
|
router.routing_strategy_pre_call_checks(deployment)
|
|
else:
|
|
## NO EXCEPTION
|
|
await router.async_routing_strategy_pre_call_checks(
|
|
deployment, litellm_logging_obj
|
|
)
|
|
|
|
## WITH EXCEPTION - rate limit error
|
|
with patch.object(
|
|
callback,
|
|
"async_pre_call_check",
|
|
AsyncMock(
|
|
side_effect=litellm.RateLimitError(
|
|
message="Rate limit error",
|
|
llm_provider="openai",
|
|
model="gpt-5-mini",
|
|
)
|
|
),
|
|
):
|
|
try:
|
|
await router.async_routing_strategy_pre_call_checks(
|
|
deployment, litellm_logging_obj
|
|
)
|
|
pytest.fail("Exception was not raised")
|
|
except Exception as e:
|
|
assert isinstance(e, litellm.RateLimitError)
|
|
|
|
## WITH EXCEPTION - generic error
|
|
with patch.object(
|
|
callback, "async_pre_call_check", AsyncMock(side_effect=Exception("Error"))
|
|
):
|
|
try:
|
|
await router.async_routing_strategy_pre_call_checks(
|
|
deployment, litellm_logging_obj
|
|
)
|
|
pytest.fail("Exception was not raised")
|
|
except Exception as e:
|
|
assert isinstance(e, Exception)
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"set_supported_environments, supported_environments, is_supported",
|
|
[(True, ["staging"], True), (False, None, True), (True, ["development"], False)],
|
|
)
|
|
def test_create_deployment(
|
|
model_list, set_supported_environments, supported_environments, is_supported
|
|
):
|
|
"""Test if the '_create_deployment' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
|
|
if set_supported_environments:
|
|
os.environ["LITELLM_ENVIRONMENT"] = "staging"
|
|
deployment = router._create_deployment(
|
|
deployment_info={},
|
|
_model_name="gpt-5-mini",
|
|
_litellm_params={
|
|
"model": "gpt-5-mini",
|
|
"api_key": "test",
|
|
"custom_llm_provider": "openai",
|
|
},
|
|
_model_info={
|
|
"id": 100,
|
|
"supported_environments": supported_environments,
|
|
},
|
|
)
|
|
if is_supported:
|
|
assert deployment is not None
|
|
else:
|
|
assert deployment is None
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"set_supported_environments, supported_environments, is_supported",
|
|
[(True, ["staging"], True), (False, None, True), (True, ["development"], False)],
|
|
)
|
|
def test_deployment_is_active_for_environment(
|
|
model_list, set_supported_environments, supported_environments, is_supported
|
|
):
|
|
"""Test if the '_deployment_is_active_for_environment' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
if set_supported_environments:
|
|
os.environ["LITELLM_ENVIRONMENT"] = "staging"
|
|
deployment["model_info"]["supported_environments"] = supported_environments
|
|
if is_supported:
|
|
assert (
|
|
router.deployment_is_active_for_environment(deployment=deployment) is True
|
|
)
|
|
else:
|
|
assert (
|
|
router.deployment_is_active_for_environment(deployment=deployment) is False
|
|
)
|
|
|
|
|
|
def test_set_model_list(model_list):
|
|
"""Test if the '_set_model_list' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
router.set_model_list(model_list=model_list)
|
|
assert len(router.model_list) == len(model_list)
|
|
|
|
|
|
def test_add_deployment(model_list):
|
|
"""Test if the '_add_deployment' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
deployment["model_info"]["id"] = "100"
|
|
## Test 1: call user facing function
|
|
router.add_deployment(deployment=deployment)
|
|
|
|
## Test 2: call internal function
|
|
router._add_deployment(deployment=deployment)
|
|
assert len(router.model_list) == len(model_list) + 1
|
|
|
|
|
|
def test_upsert_deployment(model_list):
|
|
"""Test if the 'upsert_deployment' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
print("model list", len(router.model_list))
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
deployment.litellm_params.model = "gpt-5.5"
|
|
router.upsert_deployment(deployment=deployment)
|
|
assert len(router.model_list) == len(model_list)
|
|
|
|
|
|
def test_delete_deployment(model_list):
|
|
"""Test if the 'delete_deployment' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
router.delete_deployment(id=deployment["model_info"]["id"])
|
|
assert len(router.model_list) == len(model_list) - 1
|
|
|
|
|
|
def test_get_model_info(model_list):
|
|
"""Test if the 'get_model_info' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
model_info = router.get_model_info(id=deployment["model_info"]["id"])
|
|
assert model_info is not None
|
|
|
|
|
|
def test_get_model_group(model_list):
|
|
"""Test if the 'get_model_group' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployment = router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
)
|
|
model_group = router.get_model_group(id=deployment["model_info"]["id"])
|
|
assert model_group is not None
|
|
assert model_group[0]["model_name"] == "gpt-5-mini"
|
|
|
|
|
|
@pytest.mark.parametrize("user_facing_model_group_name", ["gpt-5-mini", "gpt-5.5"])
|
|
def test_set_model_group_info(model_list, user_facing_model_group_name):
|
|
"""Test if the 'set_model_group_info' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
resp = router._set_model_group_info(
|
|
model_group="gpt-5-mini",
|
|
user_facing_model_group_name=user_facing_model_group_name,
|
|
)
|
|
assert resp is not None
|
|
assert resp.model_group == user_facing_model_group_name
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_set_response_headers(model_list):
|
|
"""Test if the 'set_response_headers' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
resp = await router.set_response_headers(response=None, model_group=None)
|
|
assert resp is None
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_set_response_headers_subtracts_in_flight_delta(model_list):
|
|
"""
|
|
LIT-2719: router-derived `x-ratelimit-remaining-*` headers must be
|
|
post-decrement (match OpenAI/Anthropic vendor semantics) so the proxy's
|
|
HTTP response headers and the prometheus gauges that read them stay
|
|
comparable across providers.
|
|
|
|
Router's TPM/RPM counter is incremented post-response by
|
|
`deployment_callback_on_success`, so `get_remaining_model_group_usage`
|
|
sees pre-decrement values. `set_response_headers` must replay the
|
|
in-flight increment before writing the headers.
|
|
"""
|
|
from pydantic import BaseModel
|
|
|
|
class _Usage(BaseModel):
|
|
total_tokens: int = 42
|
|
|
|
class _Resp(BaseModel):
|
|
usage: _Usage = _Usage()
|
|
_hidden_params: dict = {}
|
|
|
|
router = Router(model_list=model_list)
|
|
router.get_remaining_model_group_usage = AsyncMock(
|
|
return_value={
|
|
"x-ratelimit-remaining-tokens": 1000,
|
|
"x-ratelimit-limit-tokens": 1000,
|
|
"x-ratelimit-remaining-requests": 100,
|
|
"x-ratelimit-limit-requests": 100,
|
|
}
|
|
)
|
|
|
|
resp = _Resp()
|
|
resp._hidden_params = {}
|
|
await router.set_response_headers(response=resp, model_group="gpt-3.5-turbo")
|
|
|
|
headers = resp._hidden_params["additional_headers"]
|
|
assert headers["x-ratelimit-remaining-tokens"] == 958
|
|
assert headers["x-ratelimit-remaining-requests"] == 99
|
|
# Limit headers pass through unmodified.
|
|
assert headers["x-ratelimit-limit-tokens"] == 1000
|
|
assert headers["x-ratelimit-limit-requests"] == 100
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_set_response_headers_in_flight_delta_only_adjusts_tpm_rpm(model_list):
|
|
"""
|
|
The in-flight replay applies only to the post-incremented TPM/RPM counters
|
|
(`x-ratelimit-remaining-tokens` / `-requests`). The ITPM/OTPM counters are
|
|
incremented at reservation time (pre-call), so the input/output token
|
|
headers already reflect this request and must pass through untouched.
|
|
"""
|
|
from pydantic import BaseModel
|
|
|
|
class _Usage(BaseModel):
|
|
total_tokens: int = 30
|
|
prompt_tokens: int = 20
|
|
completion_tokens: int = 10
|
|
|
|
class _Resp(BaseModel):
|
|
usage: _Usage = _Usage()
|
|
_hidden_params: dict = {}
|
|
|
|
router = Router(model_list=model_list)
|
|
router.get_remaining_model_group_usage = AsyncMock(
|
|
return_value={
|
|
"x-ratelimit-remaining-tokens": 1000,
|
|
"x-ratelimit-remaining-requests": 100,
|
|
"x-ratelimit-remaining-input-tokens": 1000,
|
|
"x-ratelimit-remaining-output-tokens": 500,
|
|
}
|
|
)
|
|
|
|
resp = _Resp()
|
|
resp._hidden_params = {}
|
|
await router.set_response_headers(response=resp, model_group="gpt-3.5-turbo")
|
|
|
|
headers = resp._hidden_params["additional_headers"]
|
|
# TPM/RPM headers replay the in-flight increment...
|
|
assert headers["x-ratelimit-remaining-tokens"] == 970
|
|
assert headers["x-ratelimit-remaining-requests"] == 99
|
|
# ...but the reservation-based input/output headers pass through unchanged.
|
|
assert headers["x-ratelimit-remaining-input-tokens"] == 1000
|
|
assert headers["x-ratelimit-remaining-output-tokens"] == 500
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_get_model_group_io_token_usage_sums_across_deployments():
|
|
"""
|
|
get_model_group_io_token_usage must sum ITPM/OTPM across every deployment
|
|
in the model group (not just the first), reading the same per-deployment
|
|
cache keys the pre-call reservation writes to.
|
|
"""
|
|
from litellm.types.router import RouterCacheEnum
|
|
from litellm.utils import get_utc_datetime
|
|
|
|
router = Router(
|
|
model_list=[
|
|
{
|
|
"model_name": "opus",
|
|
"litellm_params": {
|
|
"model": "openai/gpt-4o-mini",
|
|
"itpm": 1000,
|
|
"otpm": 500,
|
|
},
|
|
"model_info": {"id": "io-usage-dep-1"},
|
|
},
|
|
{
|
|
"model_name": "opus",
|
|
"litellm_params": {
|
|
"model": "openai/gpt-4o",
|
|
"itpm": 1000,
|
|
"otpm": 500,
|
|
},
|
|
"model_info": {"id": "io-usage-dep-2"},
|
|
},
|
|
]
|
|
)
|
|
|
|
minute = get_utc_datetime().strftime("%H-%M")
|
|
keys_and_values = [
|
|
(
|
|
RouterCacheEnum.ITPM.value.format(
|
|
id="io-usage-dep-1", model="openai/gpt-4o-mini", current_minute=minute
|
|
),
|
|
30,
|
|
),
|
|
(
|
|
RouterCacheEnum.OTPM.value.format(
|
|
id="io-usage-dep-1", model="openai/gpt-4o-mini", current_minute=minute
|
|
),
|
|
10,
|
|
),
|
|
(
|
|
RouterCacheEnum.ITPM.value.format(
|
|
id="io-usage-dep-2", model="openai/gpt-4o", current_minute=minute
|
|
),
|
|
70,
|
|
),
|
|
(
|
|
RouterCacheEnum.OTPM.value.format(
|
|
id="io-usage-dep-2", model="openai/gpt-4o", current_minute=minute
|
|
),
|
|
20,
|
|
),
|
|
]
|
|
for key, value in keys_and_values:
|
|
await router.cache.async_increment_cache(key=key, value=value, ttl=60)
|
|
|
|
current_itpm, current_otpm = await router.get_model_group_io_token_usage("opus")
|
|
|
|
assert current_itpm == 100
|
|
assert current_otpm == 30
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_get_model_group_io_token_usage_no_deployments_returns_none():
|
|
router = Router(model_list=[])
|
|
current_itpm, current_otpm = await router.get_model_group_io_token_usage(
|
|
"nonexistent-group"
|
|
)
|
|
assert current_itpm is None
|
|
assert current_otpm is None
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_get_remaining_model_group_usage_merges_io_and_tpm_headers(model_list):
|
|
"""
|
|
A model group with both itpm/otpm and tpm/rpm limits must expose the
|
|
standard remaining-tokens/requests headers alongside the input/output token
|
|
headers, so clients and prometheus gauges relying on either still get data.
|
|
"""
|
|
from unittest.mock import Mock
|
|
|
|
from litellm.types.router import ModelGroupInfo
|
|
|
|
router = Router(model_list=model_list)
|
|
router._cached_get_model_group_info = Mock(
|
|
return_value=ModelGroupInfo(
|
|
model_group="gpt-3.5-turbo",
|
|
providers=["openai"],
|
|
itpm=2000,
|
|
otpm=1000,
|
|
tpm=5000,
|
|
rpm=50,
|
|
)
|
|
)
|
|
router.get_model_group_io_token_usage = AsyncMock(return_value=(100, 40))
|
|
router.get_model_group_usage = AsyncMock(return_value=(500, 5))
|
|
|
|
headers = await router.get_remaining_model_group_usage("gpt-3.5-turbo")
|
|
|
|
assert headers["x-ratelimit-remaining-input-tokens"] == 1900
|
|
assert headers["x-ratelimit-remaining-output-tokens"] == 960
|
|
assert headers["x-ratelimit-remaining-tokens"] == 4500
|
|
assert headers["x-ratelimit-remaining-requests"] == 45
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_set_response_headers_native_input_token_header_does_not_suppress_router_headers(model_list):
|
|
"""
|
|
A provider that natively returns `x-ratelimit-remaining-input-tokens` must
|
|
not suppress the router's own remaining-tokens/requests headers for a
|
|
non-IO model group.
|
|
"""
|
|
from pydantic import BaseModel
|
|
|
|
class _Usage(BaseModel):
|
|
total_tokens: int = 42
|
|
|
|
class _Resp(BaseModel):
|
|
usage: _Usage = _Usage()
|
|
_hidden_params: dict = {}
|
|
|
|
router = Router(model_list=model_list)
|
|
router.get_remaining_model_group_usage = AsyncMock(
|
|
return_value={
|
|
"x-ratelimit-remaining-tokens": 1000,
|
|
"x-ratelimit-remaining-requests": 100,
|
|
}
|
|
)
|
|
|
|
resp = _Resp()
|
|
resp._hidden_params = {"additional_headers": {"x-ratelimit-remaining-input-tokens": 5}}
|
|
await router.set_response_headers(response=resp, model_group="gpt-3.5-turbo")
|
|
|
|
headers = resp._hidden_params["additional_headers"]
|
|
assert headers["x-ratelimit-remaining-tokens"] == 958
|
|
assert headers["x-ratelimit-remaining-requests"] == 99
|
|
# the provider's native header is left untouched
|
|
assert headers["x-ratelimit-remaining-input-tokens"] == 5
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_set_response_headers_native_token_header_does_not_suppress_io_headers(model_list):
|
|
from pydantic import BaseModel
|
|
|
|
class _Usage(BaseModel):
|
|
total_tokens: int = 42
|
|
|
|
class _Resp(BaseModel):
|
|
usage: _Usage = _Usage()
|
|
_hidden_params: dict = {}
|
|
|
|
router = Router(model_list=model_list)
|
|
router.get_remaining_model_group_usage = AsyncMock(
|
|
return_value={
|
|
"x-ratelimit-remaining-tokens": 1000,
|
|
"x-ratelimit-remaining-requests": 100,
|
|
"x-ratelimit-remaining-input-tokens": 900,
|
|
"x-ratelimit-remaining-output-tokens": 450,
|
|
}
|
|
)
|
|
|
|
resp = _Resp()
|
|
resp._hidden_params = {"additional_headers": {"x-ratelimit-remaining-tokens": 5}}
|
|
await router.set_response_headers(response=resp, model_group="gpt-3.5-turbo")
|
|
|
|
headers = resp._hidden_params["additional_headers"]
|
|
assert headers["x-ratelimit-remaining-tokens"] == 5
|
|
assert headers["x-ratelimit-remaining-requests"] == 99
|
|
assert headers["x-ratelimit-remaining-input-tokens"] == 900
|
|
assert headers["x-ratelimit-remaining-output-tokens"] == 450
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_set_response_headers_handles_missing_usage(model_list):
|
|
"""
|
|
Streaming chunks and some response shapes may lack a `usage` attribute or
|
|
populated `total_tokens`. The in-flight subtraction must default to 0
|
|
tokens (still subtract 1 from requests) and never raise.
|
|
"""
|
|
from pydantic import BaseModel
|
|
|
|
class _Resp(BaseModel):
|
|
_hidden_params: dict = {}
|
|
|
|
router = Router(model_list=model_list)
|
|
router.get_remaining_model_group_usage = AsyncMock(
|
|
return_value={
|
|
"x-ratelimit-remaining-tokens": 1000,
|
|
"x-ratelimit-remaining-requests": 100,
|
|
}
|
|
)
|
|
|
|
resp = _Resp()
|
|
resp._hidden_params = {}
|
|
await router.set_response_headers(response=resp, model_group="gpt-3.5-turbo")
|
|
|
|
headers = resp._hidden_params["additional_headers"]
|
|
assert headers["x-ratelimit-remaining-tokens"] == 1000
|
|
assert headers["x-ratelimit-remaining-requests"] == 99
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_set_response_headers_dict_anthropic_messages_response(model_list):
|
|
"""Anthropic /v1/messages returns a dict; IO rate-limit headers must attach."""
|
|
router = Router(model_list=model_list)
|
|
router.get_remaining_model_group_usage = AsyncMock(
|
|
return_value={
|
|
"x-ratelimit-limit-input-tokens": 25,
|
|
"x-ratelimit-remaining-input-tokens": 20,
|
|
"x-ratelimit-limit-output-tokens": 100,
|
|
"x-ratelimit-remaining-output-tokens": 95,
|
|
}
|
|
)
|
|
|
|
resp = {
|
|
"id": "msg_123",
|
|
"type": "message",
|
|
"role": "assistant",
|
|
"content": [{"type": "text", "text": "hi"}],
|
|
"usage": {"input_tokens": 5, "output_tokens": 1},
|
|
}
|
|
await router.set_response_headers(response=resp, model_group="io-itpm-strict")
|
|
|
|
assert "_hidden_params" in resp
|
|
headers = resp["_hidden_params"]["additional_headers"]
|
|
assert headers["x-litellm-model-group"] == "io-itpm-strict"
|
|
assert headers["x-ratelimit-limit-input-tokens"] == 25
|
|
assert headers["x-ratelimit-remaining-input-tokens"] == 20
|
|
assert headers["x-ratelimit-remaining-output-tokens"] == 95
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_set_response_headers_wraps_bare_async_generator(model_list):
|
|
"""
|
|
Streaming responses that never go through Router.make_call's usual
|
|
object-based wrappers (e.g. the Anthropic /v1/messages -> Responses API
|
|
bridge, which yields a raw async generator with no `_hidden_params` slot)
|
|
must still get IO rate-limit headers attached via a thin wrapper.
|
|
"""
|
|
|
|
async def _raw_generator():
|
|
yield {"type": "message_start"}
|
|
yield {"type": "message_stop"}
|
|
|
|
router = Router(model_list=model_list)
|
|
router.get_remaining_model_group_usage = AsyncMock(
|
|
return_value={
|
|
"x-ratelimit-limit-input-tokens": 25,
|
|
"x-ratelimit-remaining-input-tokens": 20,
|
|
}
|
|
)
|
|
|
|
wrapped = await router.set_response_headers(response=_raw_generator(), model_group="io-itpm-strict")
|
|
|
|
assert hasattr(wrapped, "_hidden_params")
|
|
headers = wrapped._hidden_params["additional_headers"]
|
|
assert headers["x-litellm-model-group"] == "io-itpm-strict"
|
|
assert headers["x-ratelimit-limit-input-tokens"] == 25
|
|
assert headers["x-ratelimit-remaining-input-tokens"] == 20
|
|
|
|
from collections.abc import AsyncIterator
|
|
|
|
assert isinstance(wrapped, AsyncIterator)
|
|
chunks = [chunk async for chunk in wrapped]
|
|
assert chunks == [{"type": "message_start"}, {"type": "message_stop"}]
|
|
|
|
|
|
def test_get_all_deployments(model_list):
|
|
"""Test if the 'get_all_deployments' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployments = router._get_all_deployments(
|
|
model_name="gpt-5-mini", model_alias="gpt-5-mini"
|
|
)
|
|
assert len(deployments) > 0
|
|
|
|
|
|
def test_get_model_access_groups(model_list):
|
|
"""Test if the 'get_model_access_groups' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
access_groups = router.get_model_access_groups()
|
|
assert len(access_groups) == 2
|
|
|
|
|
|
def test_update_settings(model_list):
|
|
"""Test if the 'update_settings' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
pre_update_allowed_fails = router.allowed_fails
|
|
router.update_settings(**{"allowed_fails": 20})
|
|
assert router.allowed_fails != pre_update_allowed_fails
|
|
assert router.allowed_fails == 20
|
|
|
|
|
|
def test_common_checks_available_deployment(model_list):
|
|
"""Test if the 'common_checks_available_deployment' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
_, available_deployments = router._common_checks_available_deployment(
|
|
model="gpt-5-mini",
|
|
messages=[{"role": "user", "content": "hi"}],
|
|
input="hi",
|
|
specific_deployment=False,
|
|
)
|
|
|
|
assert len(available_deployments) > 0
|
|
|
|
|
|
def test_filter_cooldown_deployments(model_list):
|
|
"""Test if the 'filter_cooldown_deployments' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployments = router._filter_cooldown_deployments(
|
|
healthy_deployments=router._get_all_deployments(model_name="gpt-5-mini"), # type: ignore
|
|
cooldown_deployments=[],
|
|
)
|
|
assert len(deployments) == len(router._get_all_deployments(model_name="gpt-5-mini"))
|
|
|
|
|
|
def test_track_deployment_metrics(model_list):
|
|
"""Test if the 'track_deployment_metrics' function is working correctly"""
|
|
from litellm.types.utils import ModelResponse
|
|
|
|
router = Router(model_list=model_list)
|
|
router._track_deployment_metrics(
|
|
deployment=router.get_deployment_by_model_group_name(
|
|
model_group_name="gpt-5-mini"
|
|
),
|
|
response=ModelResponse(
|
|
model="gpt-5-mini",
|
|
usage={"total_tokens": 100},
|
|
),
|
|
parent_otel_span=None,
|
|
)
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"exception_type, exception_name, num_retries",
|
|
[
|
|
(litellm.exceptions.BadRequestError, "BadRequestError", 3),
|
|
(litellm.exceptions.AuthenticationError, "AuthenticationError", 4),
|
|
(litellm.exceptions.RateLimitError, "RateLimitError", 6),
|
|
(
|
|
litellm.exceptions.ContentPolicyViolationError,
|
|
"ContentPolicyViolationError",
|
|
7,
|
|
),
|
|
],
|
|
)
|
|
def test_get_num_retries_from_retry_policy(
|
|
model_list, exception_type, exception_name, num_retries
|
|
):
|
|
"""Test if the 'get_num_retries_from_retry_policy' function is working correctly"""
|
|
from litellm.router import RetryPolicy
|
|
|
|
data = {exception_name + "Retries": num_retries}
|
|
print("data", data)
|
|
router = Router(
|
|
model_list=model_list,
|
|
retry_policy=RetryPolicy(**data),
|
|
)
|
|
print("exception_type", exception_type)
|
|
calc_num_retries = router.get_num_retries_from_retry_policy(
|
|
exception=exception_type(
|
|
message="test", llm_provider="openai", model="gpt-5-mini"
|
|
)
|
|
)
|
|
assert calc_num_retries == num_retries
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"exception_type, exception_name, allowed_fails",
|
|
[
|
|
(litellm.exceptions.BadRequestError, "BadRequestError", 3),
|
|
(litellm.exceptions.AuthenticationError, "AuthenticationError", 4),
|
|
(litellm.exceptions.RateLimitError, "RateLimitError", 6),
|
|
(
|
|
litellm.exceptions.ContentPolicyViolationError,
|
|
"ContentPolicyViolationError",
|
|
7,
|
|
),
|
|
],
|
|
)
|
|
def test_get_allowed_fails_from_policy(
|
|
model_list, exception_type, exception_name, allowed_fails
|
|
):
|
|
"""Test if the 'get_allowed_fails_from_policy' function is working correctly"""
|
|
from litellm.types.router import AllowedFailsPolicy
|
|
|
|
data = {exception_name + "AllowedFails": allowed_fails}
|
|
router = Router(
|
|
model_list=model_list, allowed_fails_policy=AllowedFailsPolicy(**data)
|
|
)
|
|
calc_allowed_fails = router.get_allowed_fails_from_policy(
|
|
exception=exception_type(
|
|
message="test", llm_provider="openai", model="gpt-5-mini"
|
|
)
|
|
)
|
|
assert calc_allowed_fails == allowed_fails
|
|
|
|
|
|
def test_initialize_alerting(model_list):
|
|
"""Test if the 'initialize_alerting' function is working correctly"""
|
|
from litellm.types.router import AlertingConfig
|
|
from litellm.integrations.SlackAlerting.slack_alerting import SlackAlerting
|
|
|
|
router = Router(
|
|
model_list=model_list, alerting_config=AlertingConfig(webhook_url="test")
|
|
)
|
|
router._initialize_alerting()
|
|
|
|
callback_added = False
|
|
for callback in litellm.callbacks:
|
|
if isinstance(callback, SlackAlerting):
|
|
callback_added = True
|
|
assert callback_added is True
|
|
|
|
|
|
def test_flush_cache(model_list):
|
|
"""Test if the 'flush_cache' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
router.cache.set_cache("test", "test")
|
|
assert router.cache.get_cache("test") == "test"
|
|
router.flush_cache()
|
|
assert router.cache.get_cache("test") is None
|
|
|
|
|
|
def test_discard(model_list):
|
|
"""
|
|
Test that discard properly removes a Router from the callback lists
|
|
"""
|
|
litellm.callbacks = []
|
|
litellm.success_callback = []
|
|
litellm._async_success_callback = []
|
|
litellm.failure_callback = []
|
|
litellm._async_failure_callback = []
|
|
litellm.input_callback = []
|
|
litellm.service_callback = []
|
|
|
|
router = Router(model_list=model_list)
|
|
router.discard()
|
|
|
|
# Verify all callback lists are empty
|
|
assert len(litellm.callbacks) == 0
|
|
assert len(litellm.success_callback) == 0
|
|
assert len(litellm.failure_callback) == 0
|
|
assert len(litellm._async_success_callback) == 0
|
|
assert len(litellm._async_failure_callback) == 0
|
|
assert len(litellm.input_callback) == 0
|
|
assert len(litellm.service_callback) == 0
|
|
|
|
|
|
def test_initialize_assistants_endpoint(model_list):
|
|
"""Test if the 'initialize_assistants_endpoint' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
router.initialize_assistants_endpoint()
|
|
assert router.acreate_assistants is not None
|
|
assert router.adelete_assistant is not None
|
|
assert router.aget_assistants is not None
|
|
assert router.acreate_thread is not None
|
|
assert router.aget_thread is not None
|
|
assert router.arun_thread is not None
|
|
assert router.aget_messages is not None
|
|
assert router.a_add_message is not None
|
|
|
|
|
|
def test_pass_through_assistants_endpoint_factory(model_list):
|
|
"""Test if the 'pass_through_assistants_endpoint_factory' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
router._pass_through_assistants_endpoint_factory(
|
|
original_function=litellm.acreate_assistants,
|
|
custom_llm_provider="openai",
|
|
client=None,
|
|
**{},
|
|
)
|
|
|
|
|
|
def test_factory_function(model_list):
|
|
"""Test if the 'factory_function' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
router.factory_function(litellm.acreate_assistants)
|
|
|
|
|
|
def test_get_model_from_alias(model_list):
|
|
"""Test if the 'get_model_from_alias' function is working correctly"""
|
|
router = Router(
|
|
model_list=model_list,
|
|
model_group_alias={"gpt-5.5": "gpt-5-mini"},
|
|
)
|
|
model = router._get_model_from_alias(model="gpt-5.5")
|
|
assert model == "gpt-5-mini"
|
|
|
|
|
|
def test_get_deployment_by_litellm_model(model_list):
|
|
"""Test if the 'get_deployment_by_litellm_model' function is working correctly"""
|
|
router = Router(model_list=model_list)
|
|
deployment = router._get_deployment_by_litellm_model(model="gpt-5-mini")
|
|
assert deployment is not None
|
|
|
|
|
|
def test_get_pattern(model_list):
|
|
router = Router(model_list=model_list)
|
|
pattern = router.pattern_router.get_pattern(model="claude-3")
|
|
assert pattern is not None
|
|
|
|
|
|
def test_deployments_by_pattern(model_list):
|
|
router = Router(model_list=model_list)
|
|
deployments = router.pattern_router.get_deployments_by_pattern(model="claude-3")
|
|
assert deployments is not None
|
|
|
|
|
|
def test_replace_model_in_jsonl(model_list):
|
|
router = Router(model_list=model_list)
|
|
deployments = router.pattern_router.get_deployments_by_pattern(model="claude-3")
|
|
assert deployments is not None
|
|
|
|
|
|
# def test_pattern_match_deployments(model_list):
|
|
# from litellm.router_utils.pattern_match_deployments import PatternMatchRouter
|
|
# import re
|
|
|
|
# patter_router = PatternMatchRouter()
|
|
|
|
# request = "fo::hi::static::hello"
|
|
# model_name = "fo::*:static::*"
|
|
|
|
# model_name_regex = patter_router._pattern_to_regex(model_name)
|
|
|
|
# # Match against the request
|
|
# match = re.match(model_name_regex, request)
|
|
|
|
# print(f"match: {match}")
|
|
# print(f"match.end: {match.end()}")
|
|
# if match is None:
|
|
# raise ValueError("Match not found")
|
|
# updated_model = patter_router.set_deployment_model_name(
|
|
# matched_pattern=match, litellm_deployment_litellm_model="openai/*"
|
|
# )
|
|
# assert updated_model == "openai/fo::hi:static::hello"
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"user_request_model, model_name, litellm_model, expected_model",
|
|
[
|
|
("llmengine/foo", "llmengine/*", "openai/foo", "openai/foo"),
|
|
("llmengine/foo", "llmengine/*", "openai/*", "openai/foo"),
|
|
(
|
|
"fo::hi::static::hello",
|
|
"fo::*::static::*",
|
|
"openai/fo::*:static::*",
|
|
"openai/fo::hi:static::hello",
|
|
),
|
|
(
|
|
"fo::hi::static::hello",
|
|
"fo::*::static::*",
|
|
"openai/gpt-5-mini",
|
|
"openai/gpt-5-mini",
|
|
),
|
|
(
|
|
"bedrock/meta.llama3-70b",
|
|
"*meta.llama3*",
|
|
"bedrock/meta.llama3-*",
|
|
"bedrock/meta.llama3-70b",
|
|
),
|
|
(
|
|
"meta.llama3-70b",
|
|
"*meta.llama3*",
|
|
"bedrock/meta.llama3-*",
|
|
"meta.llama3-70b",
|
|
),
|
|
],
|
|
)
|
|
def test_pattern_match_deployment_set_model_name(
|
|
user_request_model, model_name, litellm_model, expected_model
|
|
):
|
|
from re import Match
|
|
from litellm.router_utils.pattern_match_deployments import PatternMatchRouter
|
|
|
|
pattern_router = PatternMatchRouter()
|
|
|
|
import re
|
|
|
|
# Convert model_name into a proper regex
|
|
model_name_regex = pattern_router._pattern_to_regex(model_name)
|
|
|
|
# Match against the request
|
|
match = re.match(model_name_regex, user_request_model)
|
|
|
|
if match is None:
|
|
raise ValueError("Match not found")
|
|
|
|
# Call the set_deployment_model_name function
|
|
updated_model = pattern_router.set_deployment_model_name(match, litellm_model)
|
|
|
|
print(updated_model) # Expected output: "openai/fo::hi:static::hello"
|
|
assert updated_model == expected_model
|
|
|
|
updated_models = pattern_router._return_pattern_matched_deployments(
|
|
match,
|
|
deployments=[
|
|
{
|
|
"model_name": model_name,
|
|
"litellm_params": {"model": litellm_model},
|
|
}
|
|
],
|
|
)
|
|
|
|
for model in updated_models:
|
|
assert model["litellm_params"]["model"] == expected_model
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_pass_through_moderation_endpoint_factory(model_list):
|
|
router = Router(model_list=model_list)
|
|
response = await router._pass_through_moderation_endpoint_factory(
|
|
original_function=litellm.amoderation,
|
|
input="this is valid good text",
|
|
model=None,
|
|
)
|
|
assert response is not None
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"has_default_fallbacks, expected_result",
|
|
[(True, True), (False, False)],
|
|
)
|
|
def test_has_default_fallbacks(model_list, has_default_fallbacks, expected_result):
|
|
router = Router(
|
|
model_list=model_list,
|
|
default_fallbacks=(
|
|
["my-default-fallback-model"] if has_default_fallbacks else None
|
|
),
|
|
)
|
|
assert router._has_default_fallbacks() is expected_result
|
|
|
|
|
|
def test_add_optional_pre_call_checks(model_list):
|
|
router = Router(model_list=model_list)
|
|
|
|
router.add_optional_pre_call_checks(["prompt_caching"])
|
|
assert len(litellm.callbacks) > 0
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_async_callback_filter_deployments(model_list):
|
|
from litellm.router_strategy.budget_limiter import RouterBudgetLimiting
|
|
|
|
router = Router(model_list=model_list)
|
|
|
|
healthy_deployments = router.get_model_list(model_name="gpt-5-mini")
|
|
|
|
new_healthy_deployments = await router.async_callback_filter_deployments(
|
|
model="gpt-5-mini",
|
|
healthy_deployments=healthy_deployments,
|
|
messages=[],
|
|
parent_otel_span=None,
|
|
)
|
|
|
|
assert len(new_healthy_deployments) == len(healthy_deployments)
|
|
|
|
|
|
def test_cached_get_model_group_info(model_list):
|
|
"""Test if the '_cached_get_model_group_info' function is working correctly with LRU cache"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# First call - should hit the actual function
|
|
result1 = router._cached_get_model_group_info("gpt-5-mini")
|
|
|
|
# Second call with same argument - should hit the cache
|
|
result2 = router._cached_get_model_group_info("gpt-5-mini")
|
|
|
|
# Verify results are the same
|
|
assert result1 == result2
|
|
|
|
# Verify the cache info shows hits
|
|
cache_info = router._cached_get_model_group_info.cache_info()
|
|
assert cache_info.hits > 0 # Should have at least one cache hit
|
|
|
|
|
|
def test_init_responses_api_endpoints(model_list):
|
|
"""Test if the '_init_responses_api_endpoints' function is working correctly"""
|
|
from typing import Callable
|
|
|
|
router = Router(model_list=model_list)
|
|
|
|
assert router.aget_responses is not None
|
|
assert isinstance(router.aget_responses, Callable)
|
|
assert router.adelete_responses is not None
|
|
assert isinstance(router.adelete_responses, Callable)
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"mock_testing_fallbacks, mock_testing_context_fallbacks, mock_testing_content_policy_fallbacks, expected_fallbacks, expected_context, expected_content_policy",
|
|
[
|
|
# Test string to bool conversion
|
|
("true", "false", "True", True, False, True),
|
|
("TRUE", "FALSE", "False", True, False, False),
|
|
("false", "true", "false", False, True, False),
|
|
# Test actual boolean values (should pass through unchanged)
|
|
(True, False, True, True, False, True),
|
|
(False, True, False, False, True, False),
|
|
# Test None values
|
|
(None, None, None, None, None, None),
|
|
# Test mixed types
|
|
("true", False, None, True, False, None),
|
|
],
|
|
)
|
|
def test_mock_router_testing_params_str_to_bool_conversion(
|
|
mock_testing_fallbacks,
|
|
mock_testing_context_fallbacks,
|
|
mock_testing_content_policy_fallbacks,
|
|
expected_fallbacks,
|
|
expected_context,
|
|
expected_content_policy,
|
|
):
|
|
"""Test if MockRouterTestingParams.from_kwargs correctly converts string values to booleans using str_to_bool"""
|
|
from litellm.types.router import MockRouterTestingParams
|
|
|
|
kwargs = {
|
|
"mock_testing_fallbacks": mock_testing_fallbacks,
|
|
"mock_testing_context_fallbacks": mock_testing_context_fallbacks,
|
|
"mock_testing_content_policy_fallbacks": mock_testing_content_policy_fallbacks,
|
|
"other_param": "should_remain", # This should not be affected
|
|
}
|
|
|
|
# Make a copy to verify kwargs are properly popped
|
|
original_kwargs = kwargs.copy()
|
|
|
|
mock_params = MockRouterTestingParams.from_kwargs(kwargs)
|
|
|
|
# Verify the converted values
|
|
assert mock_params.mock_testing_fallbacks == expected_fallbacks
|
|
assert mock_params.mock_testing_context_fallbacks == expected_context
|
|
assert mock_params.mock_testing_content_policy_fallbacks == expected_content_policy
|
|
|
|
# Verify that the mock testing params were popped from kwargs
|
|
assert "mock_testing_fallbacks" not in kwargs
|
|
assert "mock_testing_context_fallbacks" not in kwargs
|
|
assert "mock_testing_content_policy_fallbacks" not in kwargs
|
|
|
|
# Verify other params remain unchanged
|
|
assert kwargs["other_param"] == "should_remain"
|
|
|
|
|
|
def test_is_auto_router_deployment(model_list):
|
|
"""Test if the '_is_auto_router_deployment' function correctly identifies auto-router deployments"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Test case 1: Model starts with "auto_router/" - should return True
|
|
litellm_params_auto = LiteLLM_Params(model="auto_router/my-auto-router")
|
|
assert router._is_auto_router_deployment(litellm_params_auto) is True
|
|
|
|
# Test case 2: Model doesn't start with "auto_router/" - should return False
|
|
litellm_params_regular = LiteLLM_Params(model="gpt-5-mini")
|
|
assert router._is_auto_router_deployment(litellm_params_regular) is False
|
|
|
|
# Test case 3: Model is empty string - should return False
|
|
litellm_params_empty = LiteLLM_Params(model="")
|
|
assert router._is_auto_router_deployment(litellm_params_empty) is False
|
|
|
|
# Test case 4: Model contains "auto_router/" but doesn't start with it - should return False
|
|
litellm_params_contains = LiteLLM_Params(model="prefix_auto_router/something")
|
|
assert router._is_auto_router_deployment(litellm_params_contains) is False
|
|
|
|
|
|
@patch("litellm.router_strategy.auto_router.auto_router.AutoRouter")
|
|
def test_init_auto_router_deployment_success(mock_auto_router, model_list):
|
|
"""Test if the 'init_auto_router_deployment' function successfully initializes auto-router when all params provided"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Create a mock AutoRouter instance
|
|
mock_auto_router_instance = MagicMock()
|
|
mock_auto_router.return_value = mock_auto_router_instance
|
|
|
|
# Test case: All required parameters provided
|
|
litellm_params = LiteLLM_Params(
|
|
model="auto_router/test",
|
|
auto_router_config_path="/path/to/config",
|
|
auto_router_default_model="gpt-5-mini",
|
|
auto_router_embedding_model="text-embedding-3-small",
|
|
)
|
|
deployment = Deployment(
|
|
model_name="test-auto-router",
|
|
litellm_params=litellm_params,
|
|
model_info={"id": "test-id"},
|
|
)
|
|
|
|
# Should not raise any exception
|
|
router.init_auto_router_deployment(deployment)
|
|
|
|
# Verify AutoRouter was called with correct parameters
|
|
mock_auto_router.assert_called_once_with(
|
|
model_name="test-auto-router",
|
|
auto_router_config_path="/path/to/config",
|
|
auto_router_config=None,
|
|
default_model="gpt-5-mini",
|
|
embedding_model="text-embedding-3-small",
|
|
litellm_router_instance=router,
|
|
)
|
|
|
|
# Verify the auto-router was added to the router's auto_routers dict
|
|
assert "test-auto-router" in router.auto_routers
|
|
assert router.auto_routers["test-auto-router"] == mock_auto_router_instance
|
|
|
|
|
|
@patch("litellm.router_strategy.auto_router.auto_router.AutoRouter")
|
|
def test_init_auto_router_deployment_duplicate_model_name(mock_auto_router, model_list):
|
|
"""Test if the 'init_auto_router_deployment' function raises ValueError when model_name already exists"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Create a mock AutoRouter instance
|
|
mock_auto_router_instance = MagicMock()
|
|
mock_auto_router.return_value = mock_auto_router_instance
|
|
|
|
# Add an existing auto-router
|
|
router.auto_routers["test-auto-router"] = mock_auto_router_instance
|
|
|
|
# Try to add another auto-router with the same name
|
|
litellm_params = LiteLLM_Params(
|
|
model="auto_router/test",
|
|
auto_router_config_path="/path/to/config",
|
|
auto_router_default_model="gpt-5-mini",
|
|
auto_router_embedding_model="text-embedding-3-small",
|
|
)
|
|
deployment = Deployment(
|
|
model_name="test-auto-router",
|
|
litellm_params=litellm_params,
|
|
model_info={"id": "test-id"},
|
|
)
|
|
|
|
with pytest.raises(
|
|
ValueError, match="Auto-router deployment test-auto-router already exists"
|
|
):
|
|
router.init_auto_router_deployment(deployment)
|
|
|
|
|
|
def test_generate_model_id_with_deployment_model_name(model_list):
|
|
"""Test that _generate_model_id works correctly with deployment model_name and handles None values properly"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Test case 1: Normal case with valid model_group and litellm_params
|
|
model_group = "gpt-4.1"
|
|
litellm_params = {
|
|
"model": "gpt-4.1",
|
|
"api_key": "test_key",
|
|
"api_base": "https://api.openai.com/v1",
|
|
}
|
|
|
|
try:
|
|
result = router._generate_model_id(
|
|
model_group=model_group, litellm_params=litellm_params
|
|
)
|
|
assert isinstance(result, str)
|
|
assert len(result) > 0
|
|
print(f"✓ Success with valid model_group: {result}")
|
|
except Exception as e:
|
|
pytest.fail(f"Failed with valid model_group: {e}")
|
|
|
|
# Test case 2: Edge case with None model_group (this should fail as expected - our fix prevents this from happening)
|
|
try:
|
|
result = router._generate_model_id(
|
|
model_group=None, litellm_params=litellm_params
|
|
)
|
|
pytest.fail(
|
|
"Expected TypeError when model_group is None - this confirms our fix is needed"
|
|
)
|
|
except TypeError as e:
|
|
# After optimization, error message changed but still fails appropriately on None
|
|
assert "unsupported operand type(s) for +=" in str(
|
|
e
|
|
) or "expected str instance, NoneType found" in str(e)
|
|
print(f"✓ Correctly failed with None model_group (as expected): {e}")
|
|
except Exception as e:
|
|
pytest.fail(f"Unexpected error with None model_group: {e}")
|
|
|
|
# Test case 3: Edge case with None key in litellm_params
|
|
litellm_params_with_none_key = {
|
|
"model": "gpt-4.1",
|
|
"api_key": "test_key",
|
|
None: "should_be_skipped", # This should be handled gracefully
|
|
}
|
|
|
|
try:
|
|
result = router._generate_model_id(
|
|
model_group=model_group, litellm_params=litellm_params_with_none_key
|
|
)
|
|
assert isinstance(result, str)
|
|
assert len(result) > 0
|
|
print(f"✓ Success with None key in litellm_params: {result}")
|
|
except Exception as e:
|
|
pytest.fail(f"Failed with None key in litellm_params: {e}")
|
|
|
|
# Test case 4: Edge case with empty litellm_params
|
|
try:
|
|
result = router._generate_model_id(model_group=model_group, litellm_params={})
|
|
assert isinstance(result, str)
|
|
assert len(result) > 0
|
|
print(f"✓ Success with empty litellm_params: {result}")
|
|
except Exception as e:
|
|
pytest.fail(f"Failed with empty litellm_params: {e}")
|
|
|
|
# Test case 5: Verify that the same inputs produce the same result (deterministic)
|
|
result1 = router._generate_model_id(
|
|
model_group=model_group, litellm_params=litellm_params
|
|
)
|
|
result2 = router._generate_model_id(
|
|
model_group=model_group, litellm_params=litellm_params
|
|
)
|
|
assert result1 == result2, "Model ID generation should be deterministic"
|
|
|
|
print("✓ All _generate_model_id tests passed!")
|
|
|
|
|
|
def test_handle_clientside_credential_with_deployment_model_name(model_list):
|
|
"""Test that _handle_clientside_credential uses deployment model_name correctly"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Mock deployment with model_name
|
|
deployment = {
|
|
"model_name": "gpt-4.1",
|
|
"litellm_params": {"model": "gpt-4.1", "api_key": "test_key"},
|
|
}
|
|
|
|
# Mock kwargs with empty metadata (simulating the original issue)
|
|
kwargs = {
|
|
"metadata": {}, # Empty metadata, no model_group
|
|
"litellm_params": {
|
|
"api_key": "client_side_key",
|
|
"api_base": "https://api.openai.com/v1",
|
|
},
|
|
}
|
|
|
|
# Mock dynamic_litellm_params that would be returned by get_dynamic_litellm_params
|
|
dynamic_litellm_params = {
|
|
"api_key": "client_side_key",
|
|
"api_base": "https://api.openai.com/v1",
|
|
}
|
|
|
|
# Test that the method doesn't fail when metadata is empty
|
|
try:
|
|
# This would normally call _generate_model_id internally
|
|
# We're testing that the fix prevents the TypeError
|
|
model_group = deployment["model_name"] # This is what our fix does
|
|
assert model_group == "gpt-4.1"
|
|
|
|
# Verify that _generate_model_id works with this model_group
|
|
result = router._generate_model_id(
|
|
model_group=model_group, litellm_params=dynamic_litellm_params
|
|
)
|
|
assert isinstance(result, str)
|
|
assert len(result) > 0
|
|
|
|
print(f"✓ Success with deployment model_name: {result}")
|
|
except Exception as e:
|
|
pytest.fail(f"Failed with deployment model_name: {e}")
|
|
|
|
print("✓ _handle_clientside_credential test passed!")
|
|
|
|
|
|
def test_sync_generic_api_call_preserves_requested_model_group_in_logs():
|
|
router = Router(
|
|
model_list=[
|
|
{
|
|
"model_name": "claude-sonnet-4-6",
|
|
"litellm_params": {
|
|
"model": "bedrock/global.anthropic.claude-sonnet-4-6",
|
|
"aws_access_key_id": "test-access-key",
|
|
"aws_secret_access_key": "test-secret-key",
|
|
"aws_region_name": "us-west-2",
|
|
},
|
|
}
|
|
]
|
|
)
|
|
|
|
try:
|
|
captured_kwargs = {}
|
|
|
|
def mock_original_function(**kwargs):
|
|
captured_kwargs.update(kwargs)
|
|
return {"status": "ok"}
|
|
|
|
response = router._generic_api_call_with_fallbacks(
|
|
model="claude-sonnet-4-6",
|
|
original_function=mock_original_function,
|
|
)
|
|
|
|
assert response == {"status": "ok"}
|
|
assert captured_kwargs["model"] == "bedrock/global.anthropic.claude-sonnet-4-6"
|
|
assert captured_kwargs["litellm_metadata"]["model_group"] == "claude-sonnet-4-6"
|
|
assert (
|
|
captured_kwargs["litellm_metadata"]["deployment"]
|
|
== "bedrock/global.anthropic.claude-sonnet-4-6"
|
|
)
|
|
finally:
|
|
router.discard()
|
|
|
|
|
|
def test_sync_generic_api_call_uses_request_kwargs_for_deployment_selection():
|
|
router = Router(
|
|
model_list=[
|
|
{
|
|
"model_name": "regional-model",
|
|
"litellm_params": {
|
|
"model": "anthropic/us-model",
|
|
"api_key": "test-api-key",
|
|
"region_name": "us",
|
|
},
|
|
},
|
|
{
|
|
"model_name": "regional-model",
|
|
"litellm_params": {
|
|
"model": "anthropic/eu-model",
|
|
"api_key": "test-api-key",
|
|
"region_name": "eu",
|
|
},
|
|
},
|
|
],
|
|
enable_pre_call_checks=True,
|
|
)
|
|
|
|
try:
|
|
captured_kwargs = {}
|
|
|
|
def mock_original_function(**kwargs):
|
|
captured_kwargs.update(kwargs)
|
|
return {"status": "ok"}
|
|
|
|
response = router._generic_api_call_with_fallbacks(
|
|
model="regional-model",
|
|
original_function=mock_original_function,
|
|
messages=[{"role": "user", "content": "Hello from Europe"}],
|
|
allowed_model_region="eu",
|
|
)
|
|
|
|
assert response == {"status": "ok"}
|
|
assert captured_kwargs["model"] == "anthropic/eu-model"
|
|
finally:
|
|
router.discard()
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"function_name, expected_metadata_key",
|
|
[
|
|
("acompletion", "metadata"),
|
|
("_ageneric_api_call_with_fallbacks", "litellm_metadata"),
|
|
("batch", "litellm_metadata"),
|
|
("completion", "metadata"),
|
|
("acreate_file", "litellm_metadata"),
|
|
("aget_file", "litellm_metadata"),
|
|
],
|
|
)
|
|
def test_handle_clientside_credential_metadata_loading(
|
|
model_list, function_name, expected_metadata_key
|
|
):
|
|
"""Test that _handle_clientside_credential correctly loads metadata based on function name"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Mock deployment
|
|
deployment = {
|
|
"model_name": "gpt-4.1",
|
|
"litellm_params": {"model": "gpt-4.1", "api_key": "test_key"},
|
|
"model_info": {"id": "original-id-123"},
|
|
}
|
|
|
|
# Mock kwargs with clientside credentials and metadata
|
|
kwargs = {
|
|
"api_key": "client_side_key",
|
|
"api_base": "https://api.openai.com/v1",
|
|
expected_metadata_key: {"model_group": "gpt-4.1", "custom_field": "test_value"},
|
|
}
|
|
|
|
# Call the function
|
|
result_deployment = router._handle_clientside_credential(
|
|
deployment=deployment, kwargs=kwargs, function_name=function_name
|
|
)
|
|
|
|
# Verify the result is a Deployment object
|
|
assert isinstance(result_deployment, Deployment)
|
|
|
|
# Verify the deployment has the correct model_name (should be the model_group from metadata)
|
|
assert result_deployment.model_name == "gpt-4.1"
|
|
|
|
# Verify the litellm_params contain the clientside credentials
|
|
assert result_deployment.litellm_params.api_key == "client_side_key"
|
|
assert result_deployment.litellm_params.api_base == "https://api.openai.com/v1"
|
|
|
|
# Verify the model_info has been updated with a new ID
|
|
assert result_deployment.model_info.id != "original-id-123"
|
|
assert result_deployment.model_info.original_model_id == "original-id-123"
|
|
|
|
# Verify the deployment was added to the router
|
|
assert len(router.model_list) == len(model_list) + 1
|
|
|
|
# Test that the function correctly uses the right metadata key
|
|
# For acompletion, it should use "metadata"
|
|
# For _ageneric_api_call_with_fallbacks/batch, it should use "litellm_metadata"
|
|
if function_name == "acompletion":
|
|
assert "metadata" in kwargs
|
|
assert "litellm_metadata" not in kwargs
|
|
elif function_name in [
|
|
"_ageneric_api_call_with_fallbacks",
|
|
"batch",
|
|
"acreate_file",
|
|
"aget_file",
|
|
]:
|
|
assert "litellm_metadata" in kwargs
|
|
# Note: acompletion would not have litellm_metadata, but other functions might have both
|
|
|
|
print(
|
|
f"✓ Success with function_name '{function_name}' using '{expected_metadata_key}' metadata key"
|
|
)
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"function_name, metadata_key",
|
|
[
|
|
("acompletion", "metadata"),
|
|
("_ageneric_api_call_with_fallbacks", "litellm_metadata"),
|
|
],
|
|
)
|
|
def test_handle_clientside_credential_metadata_variable_name(
|
|
model_list, function_name, metadata_key
|
|
):
|
|
"""Test that _handle_clientside_credential uses the correct metadata variable name based on function name"""
|
|
from litellm.router_utils.batch_utils import _get_router_metadata_variable_name
|
|
|
|
router = Router(model_list=model_list)
|
|
|
|
# Verify the metadata variable name is correct for each function
|
|
expected_metadata_key = _get_router_metadata_variable_name(
|
|
function_name=function_name
|
|
)
|
|
assert expected_metadata_key == metadata_key
|
|
|
|
# Mock deployment
|
|
deployment = {
|
|
"model_name": "gpt-4.1",
|
|
"litellm_params": {"model": "gpt-4.1", "api_key": "test_key"},
|
|
"model_info": {"id": "original-id-456"},
|
|
}
|
|
|
|
# Mock kwargs with clientside credentials and the correct metadata key
|
|
kwargs = {
|
|
"api_key": "client_side_key",
|
|
"api_base": "https://api.openai.com/v1",
|
|
metadata_key: {"model_group": "gpt-4.1", "test_field": "test_value"},
|
|
}
|
|
|
|
# Call the function
|
|
result_deployment = router._handle_clientside_credential(
|
|
deployment=deployment, kwargs=kwargs, function_name=function_name
|
|
)
|
|
|
|
# Verify the function correctly extracted model_group from the right metadata key
|
|
assert result_deployment.model_name == "gpt-4.1"
|
|
|
|
# Verify the deployment was created with the correct metadata
|
|
assert result_deployment.litellm_params.api_key == "client_side_key"
|
|
assert result_deployment.litellm_params.api_base == "https://api.openai.com/v1"
|
|
|
|
print(
|
|
f"✓ Success with function_name '{function_name}' correctly using '{metadata_key}' for metadata"
|
|
)
|
|
|
|
|
|
def test_handle_clientside_credential_no_metadata(model_list):
|
|
"""Test that _handle_clientside_credential handles cases where no metadata is provided"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Mock deployment
|
|
deployment = {
|
|
"model_name": "gpt-4.1",
|
|
"litellm_params": {"model": "gpt-4.1", "api_key": "test_key"},
|
|
"model_info": {"id": "original-id-789"},
|
|
}
|
|
|
|
# Mock kwargs with clientside credentials but NO metadata
|
|
kwargs = {
|
|
"api_key": "client_side_key",
|
|
"api_base": "https://api.openai.com/v1",
|
|
# No metadata key at all
|
|
}
|
|
|
|
# This should fail because there's no model_group in metadata
|
|
# The function expects to find model_group in the metadata
|
|
try:
|
|
result_deployment = router._handle_clientside_credential(
|
|
deployment=deployment, kwargs=kwargs, function_name="acompletion"
|
|
)
|
|
# If we get here, the function should have used deployment.model_name as fallback
|
|
assert result_deployment.model_name == "gpt-4.1"
|
|
print("✓ Success with no metadata - used deployment.model_name as fallback")
|
|
except Exception as e:
|
|
# This is expected behavior - the function needs model_group to generate model_id
|
|
print(f"✓ Correctly handled no metadata case: {e}")
|
|
|
|
# Test with empty metadata
|
|
kwargs_with_empty_metadata = {
|
|
"api_key": "client_side_key",
|
|
"api_base": "https://api.openai.com/v1",
|
|
"metadata": {}, # Empty metadata
|
|
}
|
|
|
|
try:
|
|
result_deployment = router._handle_clientside_credential(
|
|
deployment=deployment,
|
|
kwargs=kwargs_with_empty_metadata,
|
|
function_name="acompletion",
|
|
)
|
|
# Should fail because empty metadata has no model_group
|
|
pytest.fail("Expected failure with empty metadata")
|
|
except Exception as e:
|
|
print(f"✓ Correctly handled empty metadata case: {e}")
|
|
|
|
|
|
def test_handle_clientside_credential_with_responses_function(model_list):
|
|
"""Test that _handle_clientside_credential works correctly with responses function name"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Mock deployment
|
|
deployment = {
|
|
"model_name": "gpt-4.1",
|
|
"litellm_params": {"model": "gpt-4.1", "api_key": "test_key"},
|
|
"model_info": {"id": "original-id-responses"},
|
|
}
|
|
|
|
# Mock kwargs with clientside credentials and litellm_metadata (for responses function)
|
|
kwargs = {
|
|
"api_key": "client_side_key",
|
|
"api_base": "https://api.openai.com/v1",
|
|
"litellm_metadata": {
|
|
"model_group": "gpt-4.1",
|
|
"responses_field": "responses_value",
|
|
},
|
|
}
|
|
|
|
# Call the function with _ageneric_api_call_with_fallbacks function name (which handles responses)
|
|
result_deployment = router._handle_clientside_credential(
|
|
deployment=deployment,
|
|
kwargs=kwargs,
|
|
function_name="_ageneric_api_call_with_fallbacks",
|
|
)
|
|
|
|
# Verify the result
|
|
assert isinstance(result_deployment, Deployment)
|
|
assert result_deployment.model_name == "gpt-4.1"
|
|
assert result_deployment.litellm_params.api_key == "client_side_key"
|
|
assert result_deployment.litellm_params.api_base == "https://api.openai.com/v1"
|
|
assert result_deployment.model_info.id != "original-id-responses"
|
|
assert result_deployment.model_info.original_model_id == "original-id-responses"
|
|
|
|
# Verify the deployment was added to the router
|
|
assert len(router.model_list) == len(model_list) + 1
|
|
|
|
print(
|
|
"✓ Success with _ageneric_api_call_with_fallbacks function name and litellm_metadata"
|
|
)
|
|
|
|
|
|
def test_get_metadata_variable_name_from_kwargs(model_list):
|
|
"""
|
|
Test _get_metadata_variable_name_from_kwargs method returns correct metadata variable name based on kwargs content.
|
|
"""
|
|
router = Router(model_list=model_list)
|
|
|
|
# Test case 1: kwargs contains litellm_metadata - should return "litellm_metadata"
|
|
kwargs_with_litellm_metadata = {
|
|
"litellm_metadata": {"user": "test"},
|
|
"metadata": {"other": "data"},
|
|
}
|
|
result = router._get_metadata_variable_name_from_kwargs(
|
|
kwargs_with_litellm_metadata
|
|
)
|
|
assert result == "litellm_metadata"
|
|
|
|
# Test case 2: kwargs only contains metadata - should return "metadata"
|
|
kwargs_with_metadata_only = {"metadata": {"user": "test"}}
|
|
result = router._get_metadata_variable_name_from_kwargs(kwargs_with_metadata_only)
|
|
assert result == "metadata"
|
|
|
|
# Test case 3: kwargs contains neither - should return "metadata" (default)
|
|
kwargs_empty = {}
|
|
result = router._get_metadata_variable_name_from_kwargs(kwargs_empty)
|
|
assert result == "metadata"
|
|
|
|
# Test case 4: kwargs contains other keys but no metadata keys - should return "metadata"
|
|
kwargs_other = {
|
|
"model": "gpt-5.5",
|
|
"messages": [{"role": "user", "content": "hello"}],
|
|
}
|
|
result = router._get_metadata_variable_name_from_kwargs(kwargs_other)
|
|
assert result == "metadata"
|
|
|
|
|
|
@pytest.fixture
|
|
def search_tools():
|
|
"""Fixture for search tools configuration"""
|
|
return [
|
|
{
|
|
"search_tool_name": "test-search-tool",
|
|
"litellm_params": {
|
|
"search_provider": "perplexity",
|
|
"api_key": "test-api-key",
|
|
"api_base": "https://api.perplexity.ai",
|
|
},
|
|
},
|
|
{
|
|
"search_tool_name": "test-search-tool",
|
|
"litellm_params": {
|
|
"search_provider": "perplexity",
|
|
"api_key": "test-api-key-2",
|
|
"api_base": "https://api.perplexity.ai",
|
|
},
|
|
},
|
|
]
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_asearch_with_fallbacks(search_tools):
|
|
"""
|
|
Test _asearch_with_fallbacks method of Router.
|
|
|
|
Tests that the _asearch_with_fallbacks method correctly:
|
|
- Accepts search parameters
|
|
- Calls async_function_with_fallbacks with correct configuration
|
|
- Returns SearchResponse
|
|
"""
|
|
from litellm.llms.base_llm.search.transformation import SearchResponse, SearchResult
|
|
|
|
router = Router(search_tools=search_tools)
|
|
|
|
# Create a mock search response
|
|
mock_response = SearchResponse(
|
|
object="search",
|
|
results=[
|
|
SearchResult(
|
|
title="Test Result",
|
|
url="https://example.com",
|
|
snippet="Test snippet content",
|
|
)
|
|
],
|
|
)
|
|
|
|
# Mock the async_function_with_fallbacks to return our mock response
|
|
with patch.object(
|
|
router, "async_function_with_fallbacks", new_callable=AsyncMock
|
|
) as mock_fallbacks:
|
|
mock_fallbacks.return_value = mock_response
|
|
|
|
# Mock original function
|
|
async def mock_asearch(**kwargs):
|
|
return mock_response
|
|
|
|
# Call _asearch_with_fallbacks
|
|
response = await router._asearch_with_fallbacks(
|
|
original_function=mock_asearch,
|
|
search_tool_name="test-search-tool",
|
|
query="test query",
|
|
max_results=5,
|
|
)
|
|
|
|
# Verify async_function_with_fallbacks was called
|
|
assert mock_fallbacks.called
|
|
|
|
# Verify the response
|
|
assert isinstance(response, SearchResponse)
|
|
assert response.object == "search"
|
|
assert len(response.results) == 1
|
|
assert response.results[0].title == "Test Result"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_asearch_with_fallbacks_helper(search_tools):
|
|
"""
|
|
Test _asearch_with_fallbacks_helper method of Router.
|
|
|
|
Tests that the _asearch_with_fallbacks_helper method correctly:
|
|
- Selects a search tool from available options
|
|
- Calls the original search function with correct provider parameters
|
|
- Returns SearchResponse
|
|
"""
|
|
from litellm.llms.base_llm.search.transformation import SearchResponse, SearchResult
|
|
|
|
router = Router(search_tools=search_tools)
|
|
|
|
# Create a mock search response
|
|
mock_response = SearchResponse(
|
|
object="search",
|
|
results=[
|
|
SearchResult(
|
|
title="Helper Test Result",
|
|
url="https://example.com/helper",
|
|
snippet="Helper test snippet",
|
|
)
|
|
],
|
|
)
|
|
|
|
# Mock the original generic function
|
|
async def mock_original_function(**kwargs):
|
|
# Verify correct parameters are passed
|
|
assert "search_provider" in kwargs
|
|
assert kwargs["search_provider"] == "perplexity"
|
|
assert "api_key" in kwargs
|
|
assert kwargs["query"] == "helper test query"
|
|
return mock_response
|
|
|
|
# Call _asearch_with_fallbacks_helper
|
|
response = await router._asearch_with_fallbacks_helper(
|
|
model="test-search-tool",
|
|
original_generic_function=mock_original_function,
|
|
query="helper test query",
|
|
max_results=3,
|
|
)
|
|
|
|
# Verify the response
|
|
assert isinstance(response, SearchResponse)
|
|
assert response.object == "search"
|
|
assert len(response.results) == 1
|
|
assert response.results[0].title == "Helper Test Result"
|
|
assert response.results[0].url == "https://example.com/helper"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_asearch_with_fallbacks_helper_missing_search_tool():
|
|
"""
|
|
Test _asearch_with_fallbacks_helper raises error when search tool not found.
|
|
|
|
Tests that the helper method raises a ValueError when the requested
|
|
search tool name doesn't exist in the router's search_tools configuration.
|
|
"""
|
|
# Create router with no search tools
|
|
router = Router(model_list=[])
|
|
|
|
async def mock_original_function(**kwargs):
|
|
return None
|
|
|
|
# Should raise ValueError for missing search tool
|
|
with pytest.raises(ValueError, match="Search tool 'nonexistent-tool' not found"):
|
|
await router._asearch_with_fallbacks_helper(
|
|
model="nonexistent-tool",
|
|
original_generic_function=mock_original_function,
|
|
query="test query",
|
|
)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_asearch_with_fallbacks_helper_missing_search_provider():
|
|
"""
|
|
Test _asearch_with_fallbacks_helper raises error when search_provider not configured.
|
|
|
|
Tests that the helper method raises a ValueError when a search tool
|
|
is found but doesn't have search_provider in its litellm_params.
|
|
"""
|
|
# Create router with misconfigured search tool (missing search_provider)
|
|
search_tools_bad = [
|
|
{
|
|
"search_tool_name": "bad-tool",
|
|
"litellm_params": {
|
|
"api_key": "test-key"
|
|
# Missing search_provider
|
|
},
|
|
}
|
|
]
|
|
|
|
router = Router(search_tools=search_tools_bad)
|
|
|
|
async def mock_original_function(**kwargs):
|
|
return None
|
|
|
|
# Should raise ValueError for missing search_provider
|
|
with pytest.raises(ValueError, match="search_provider not found in litellm_params"):
|
|
await router._asearch_with_fallbacks_helper(
|
|
model="bad-tool",
|
|
original_generic_function=mock_original_function,
|
|
query="test query",
|
|
)
|
|
|
|
|
|
def test_get_first_default_fallback():
|
|
"""Test _get_first_default_fallback method"""
|
|
# Test with default fallback ("*")
|
|
model_list = [
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {"model": "gpt-5-mini", "api_key": "fake-key"},
|
|
}
|
|
]
|
|
|
|
router = Router(model_list=model_list, fallbacks=[{"*": ["gpt-5-mini"]}])
|
|
|
|
result = router._get_first_default_fallback()
|
|
assert result == "gpt-5-mini"
|
|
|
|
# Test with no fallbacks
|
|
router_no_fallbacks = Router(model_list=model_list)
|
|
result = router_no_fallbacks._get_first_default_fallback()
|
|
assert result is None
|
|
|
|
# Test with fallbacks but no default
|
|
router_no_default = Router(
|
|
model_list=model_list, fallbacks=[{"gpt-5.5": ["gpt-5-mini"]}]
|
|
)
|
|
result = router_no_default._get_first_default_fallback()
|
|
assert result is None
|
|
|
|
# Test with empty default list
|
|
router_empty_list = Router(model_list=model_list, fallbacks=[{"*": []}])
|
|
result = router_empty_list._get_first_default_fallback()
|
|
assert result is None
|
|
|
|
|
|
def test_resolve_model_name_from_model_id():
|
|
"""Test resolve_model_name_from_model_id function with various scenarios"""
|
|
|
|
# Test case 1: model_id is None
|
|
router = Router(model_list=[])
|
|
result = router.resolve_model_name_from_model_id(None)
|
|
assert result is None
|
|
|
|
# Test case 2: model_id directly matches a model_name
|
|
model_list = [
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {
|
|
"model": "gpt-5-mini",
|
|
"api_key": "test-key",
|
|
},
|
|
},
|
|
]
|
|
router = Router(model_list=model_list)
|
|
result = router.resolve_model_name_from_model_id("gpt-5-mini")
|
|
assert result == "gpt-5-mini"
|
|
|
|
# Test case 3: model_id matches litellm_params.model exactly
|
|
model_list = [
|
|
{
|
|
"model_name": "vertex-ai-sora-2",
|
|
"litellm_params": {
|
|
"model": "vertex_ai/veo-2.0-generate-001",
|
|
"api_key": "test-key",
|
|
},
|
|
},
|
|
]
|
|
router = Router(model_list=model_list)
|
|
result = router.resolve_model_name_from_model_id("vertex_ai/veo-2.0-generate-001")
|
|
assert result == "vertex-ai-sora-2"
|
|
|
|
# Test case 4: model_id matches when actual_model ends with /model_id
|
|
model_list = [
|
|
{
|
|
"model_name": "vertex-ai-sora-2",
|
|
"litellm_params": {
|
|
"model": "vertex_ai/veo-2.0-generate-001",
|
|
"api_key": "test-key",
|
|
},
|
|
},
|
|
]
|
|
router = Router(model_list=model_list)
|
|
result = router.resolve_model_name_from_model_id("veo-2.0-generate-001")
|
|
assert result == "vertex-ai-sora-2"
|
|
|
|
# Test case 5: model_id matches when actual_model ends with :model_id
|
|
# Note: We use a valid model format for router initialization, but test the function
|
|
# with a model_id that would match the pattern vertex_ai:model_id
|
|
# Since the router validates models on init, we'll test this by manually setting up
|
|
# the model_list after initialization or using a valid format
|
|
model_list = [
|
|
{
|
|
"model_name": "vertex-ai-sora-2",
|
|
"litellm_params": {
|
|
"model": "vertex_ai/veo-2.0-generate-001",
|
|
"api_key": "test-key",
|
|
},
|
|
},
|
|
]
|
|
router = Router(model_list=model_list)
|
|
# Test that the function can handle model_id that would match if the format was vertex_ai:model_id
|
|
# We'll test with a model_id that matches the end of the actual_model
|
|
result = router.resolve_model_name_from_model_id("veo-2.0-generate-001")
|
|
assert result == "vertex-ai-sora-2"
|
|
|
|
# Test case 6: model_id doesn't match anything
|
|
model_list = [
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {
|
|
"model": "gpt-5-mini",
|
|
"api_key": "test-key",
|
|
},
|
|
},
|
|
]
|
|
router = Router(model_list=model_list)
|
|
result = router.resolve_model_name_from_model_id("non-existent-model")
|
|
assert result is None
|
|
|
|
# Test case 7: Empty model_list
|
|
router = Router(model_list=[])
|
|
result = router.resolve_model_name_from_model_id("some-model")
|
|
assert result is None
|
|
|
|
# Test case 8: Multiple models, find the correct one
|
|
model_list = [
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {
|
|
"model": "gpt-5-mini",
|
|
"api_key": "test-key",
|
|
},
|
|
},
|
|
{
|
|
"model_name": "vertex-ai-sora-2",
|
|
"litellm_params": {
|
|
"model": "vertex_ai/veo-2.0-generate-001",
|
|
"api_key": "test-key",
|
|
},
|
|
},
|
|
]
|
|
router = Router(model_list=model_list)
|
|
result = router.resolve_model_name_from_model_id("veo-2.0-generate-001")
|
|
assert result == "vertex-ai-sora-2"
|
|
|
|
# Test case 9: model_id matches deployment ID (has_model_id check)
|
|
# This tests the has_model_id path in Strategy 1
|
|
model_list = [
|
|
{
|
|
"model_name": "gpt-5-mini",
|
|
"litellm_params": {
|
|
"model": "gpt-5-mini",
|
|
"api_key": "test-key",
|
|
},
|
|
},
|
|
]
|
|
router = Router(model_list=model_list)
|
|
|
|
result = router.resolve_model_name_from_model_id("gpt-5-mini")
|
|
assert result == "gpt-5-mini"
|
|
|
|
|
|
def test_get_valid_args():
|
|
"""Test get_valid_args static method returns valid Router.__init__ arguments"""
|
|
# Call the static method
|
|
valid_args = Router.get_valid_args()
|
|
|
|
# Verify it returns a list
|
|
assert isinstance(valid_args, list)
|
|
assert len(valid_args) > 0
|
|
|
|
# Verify it contains expected Router.__init__ arguments
|
|
expected_args = [
|
|
"model_list",
|
|
"routing_strategy",
|
|
"cache_responses",
|
|
"num_retries",
|
|
"timeout",
|
|
"fallbacks",
|
|
]
|
|
for arg in expected_args:
|
|
assert arg in valid_args, f"Expected argument '{arg}' not found in valid_args"
|
|
|
|
# Verify "self" is not in the list (since it's removed)
|
|
assert "self" not in valid_args
|
|
|
|
# Verify it contains keyword-only arguments too
|
|
# These are common Router.__init__ parameters
|
|
assert "assistants_config" in valid_args or "search_tools" in valid_args
|
|
|
|
|
|
def test_get_router_model_info_with_deployment_object():
|
|
"""Test get_router_model_info accepts Deployment object directly and reuses LiteLLM_Params"""
|
|
router = Router(
|
|
model_list=[
|
|
{
|
|
"model_name": "gpt-5.5",
|
|
"litellm_params": {"model": "gpt-5.5", "api_key": "test-key"},
|
|
"model_info": {"id": "test-id"},
|
|
}
|
|
]
|
|
)
|
|
|
|
# Get the Deployment object (not dict)
|
|
deployment = router.get_deployment(model_id="test-id")
|
|
assert deployment is not None
|
|
assert isinstance(deployment, Deployment)
|
|
assert isinstance(deployment.litellm_params, LiteLLM_Params)
|
|
|
|
# Pass Deployment directly (not .model_dump()) - this exercises the isinstance check
|
|
# that reuses the existing LiteLLM_Params instead of reconstructing it
|
|
model_info = router.get_router_model_info(
|
|
deployment=deployment,
|
|
received_model_name="gpt-5.5",
|
|
)
|
|
|
|
# Verify we got valid model info back
|
|
assert model_info is not None
|
|
assert isinstance(model_info, dict)
|
|
|
|
|
|
def test_deployment_has_budget_limits():
|
|
router = Router(model_list=[])
|
|
|
|
with_budget = Deployment(
|
|
model_name="budgeted-model",
|
|
litellm_params=LiteLLM_Params(
|
|
model="openai/gpt-4o-mini",
|
|
max_budget=0.001,
|
|
budget_duration="1d",
|
|
),
|
|
model_info=ModelInfo(id="budget-deployment-id"),
|
|
)
|
|
without_budget = Deployment(
|
|
model_name="unbudgeted-model",
|
|
litellm_params=LiteLLM_Params(model="openai/gpt-4o-mini"),
|
|
model_info=ModelInfo(id="no-budget-deployment-id"),
|
|
)
|
|
|
|
assert router._deployment_has_budget_limits(deployment=with_budget) is True
|
|
assert router._deployment_has_budget_limits(deployment=without_budget) is False
|
|
|
|
|
|
def test_sync_deployment_budget_config(monkeypatch):
|
|
import asyncio
|
|
|
|
monkeypatch.setattr(asyncio, "create_task", lambda coro: None)
|
|
|
|
router = Router(model_list=[], optional_pre_call_checks=[])
|
|
deployment = Deployment(
|
|
model_name="dynamic-budget-model",
|
|
litellm_params=LiteLLM_Params(
|
|
model="openai/gpt-4o-mini",
|
|
api_key="fake-key",
|
|
max_budget=0.000000000001,
|
|
budget_duration="1d",
|
|
),
|
|
model_info=ModelInfo(id="runtime-budget-deployment"),
|
|
)
|
|
|
|
router._sync_deployment_budget_config(deployment=deployment)
|
|
|
|
budget_limiter = router._get_router_deployment_budget_limiter()
|
|
assert budget_limiter is not None
|
|
config = budget_limiter._get_budget_config_for_deployment(
|
|
"runtime-budget-deployment"
|
|
)
|
|
assert config is not None
|
|
assert config.max_budget == 0.000000000001
|
|
|
|
|
|
def test_sync_deployment_budget_config_clears_removed_limits(monkeypatch):
|
|
import asyncio
|
|
|
|
monkeypatch.setattr(asyncio, "create_task", lambda coro: None)
|
|
|
|
router = Router(model_list=[], optional_pre_call_checks=[])
|
|
model_id = "runtime-budget-deployment"
|
|
budgeted = Deployment(
|
|
model_name="dynamic-budget-model",
|
|
litellm_params=LiteLLM_Params(
|
|
model="openai/gpt-4o-mini",
|
|
api_key="fake-key",
|
|
max_budget=0.000000000001,
|
|
budget_duration="1d",
|
|
),
|
|
model_info=ModelInfo(id=model_id),
|
|
)
|
|
unbudgeted = Deployment(
|
|
model_name="dynamic-budget-model",
|
|
litellm_params=LiteLLM_Params(
|
|
model="openai/gpt-4o-mini",
|
|
api_key="fake-key",
|
|
),
|
|
model_info=ModelInfo(id=model_id),
|
|
)
|
|
|
|
router._sync_deployment_budget_config(deployment=budgeted)
|
|
budget_limiter = router._get_router_deployment_budget_limiter()
|
|
assert budget_limiter is not None
|
|
assert budget_limiter._get_budget_config_for_deployment(model_id) is not None
|
|
|
|
router._sync_deployment_budget_config(deployment=unbudgeted)
|
|
assert budget_limiter._get_budget_config_for_deployment(model_id) is None
|
|
|
|
|
|
def test_upsert_deployment_clears_stale_budget_config(monkeypatch):
|
|
import asyncio
|
|
|
|
monkeypatch.setattr(asyncio, "create_task", lambda coro: None)
|
|
|
|
router = Router(model_list=[], optional_pre_call_checks=[])
|
|
model_id = "upsert-budget-deployment"
|
|
budgeted = Deployment(
|
|
model_name="dynamic-budget-model",
|
|
litellm_params=LiteLLM_Params(
|
|
model="openai/gpt-4o-mini",
|
|
api_key="fake-key",
|
|
max_budget=0.000000000001,
|
|
budget_duration="1d",
|
|
),
|
|
model_info=ModelInfo(id=model_id),
|
|
)
|
|
unbudgeted = Deployment(
|
|
model_name="dynamic-budget-model",
|
|
litellm_params=LiteLLM_Params(
|
|
model="openai/gpt-4o-mini",
|
|
api_key="fake-key",
|
|
),
|
|
model_info=ModelInfo(id=model_id),
|
|
)
|
|
|
|
router.upsert_deployment(deployment=budgeted)
|
|
budget_limiter = router._get_router_deployment_budget_limiter()
|
|
assert budget_limiter is not None
|
|
assert budget_limiter._get_budget_config_for_deployment(model_id) is not None
|
|
|
|
router.upsert_deployment(deployment=unbudgeted)
|
|
assert budget_limiter._get_budget_config_for_deployment(model_id) is None
|