mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-10 03:28:53 +00:00
Merge pull request #21917 from BerriAI/litellm_fix_model_cost_map_wildcard
Fix: Anthropic model wildcard access issue
This commit is contained in:
commit
37d45139f2
7 changed files with 537 additions and 10 deletions
|
|
@ -0,0 +1,147 @@
|
|||
---
|
||||
slug: anthropic-wildcard-model-access-incident
|
||||
title: "Incident Report: Wildcard Blocking New Models After Cost Map Reload"
|
||||
date: 2026-02-23T10:00:00
|
||||
authors:
|
||||
- name: Sameer Kankute
|
||||
title: SWE @ LiteLLM (LLM Translation)
|
||||
url: https://www.linkedin.com/in/sameer-kankute/
|
||||
image_url: https://pbs.twimg.com/profile_images/2001352686994907136/ONgNuSk5_400x400.jpg
|
||||
- name: Krrish Dholakia
|
||||
title: "CEO, LiteLLM"
|
||||
url: https://www.linkedin.com/in/krish-d/
|
||||
image_url: https://pbs.twimg.com/profile_images/1298587542745358340/DZv3Oj-h_400x400.jpg
|
||||
- name: Ishaan Jaff
|
||||
title: "CTO, LiteLLM"
|
||||
url: https://www.linkedin.com/in/reffajnaahsi/
|
||||
image_url: https://pbs.twimg.com/profile_images/1613813310264340481/lz54oEiB_400x400.jpg
|
||||
tags: [incident-report, proxy, auth, model-access]
|
||||
hide_table_of_contents: false
|
||||
---
|
||||
|
||||
**Date:** Feb 23, 2026
|
||||
**Duration:** ~3 hours
|
||||
**Severity:** High (for users with provider wildcard access rules)
|
||||
**Status:** Resolved
|
||||
|
||||
## Summary
|
||||
|
||||
When a new Anthropic model (e.g. `claude-sonnet-4-6`) was added to the LiteLLM model cost map and a cost map reload was triggered, requests to the new model were rejected with:
|
||||
|
||||
```
|
||||
key not allowed to access model. This key can only access models=['anthropic/*']. Tried to access claude-sonnet-4-6.
|
||||
```
|
||||
|
||||
The reload updated `litellm.model_cost` correctly but never re-ran `add_known_models()`, so `litellm.anthropic_models` (the in-memory set used by the wildcard resolver) remained stale. The new model was invisible to the `anthropic/*` wildcard even though the cost map knew about it.
|
||||
|
||||
- **LLM calls:** All requests to newly-added Anthropic models were blocked with a 401.
|
||||
- **Existing models:** Unaffected — only models missing from the stale provider set were impacted.
|
||||
- **Other providers:** Same bug class existed for any provider wildcard (e.g. `openai/*`, `gemini/*`).
|
||||
|
||||
{/* truncate */}
|
||||
|
||||
---
|
||||
|
||||
## Background
|
||||
|
||||
LiteLLM supports provider-level wildcard access rules. When an admin configures a key or team with `models=['anthropic/*']`, any model whose provider resolves to `anthropic` should be allowed. The resolution happens in `_model_custom_llm_provider_matches_wildcard_pattern`:
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
A["1. Request arrives for claude-sonnet-4-6"] --> B["2. Auth check: can this key call this model?
|
||||
proxy/auth/auth_checks.py"]
|
||||
B --> C["3. Key has models=['anthropic/*']
|
||||
→ wildcard match attempted"]
|
||||
C --> D["4. get_llm_provider('claude-sonnet-4-6')
|
||||
checks litellm.anthropic_models set"]
|
||||
D -->|"model IN set"| E["5a. ✅ Provider = 'anthropic'
|
||||
→ 'anthropic/claude-sonnet-4-6' matches 'anthropic/*'"]
|
||||
D -->|"model NOT IN set"| F["5b. ❌ Provider unknown
|
||||
→ exception raised → wildcard returns False"]
|
||||
E --> G["6. Request allowed"]
|
||||
F --> H["6. 401: key not allowed to access model"]
|
||||
|
||||
style E fill:#d4edda,stroke:#28a745
|
||||
style F fill:#f8d7da,stroke:#dc3545
|
||||
style H fill:#f8d7da,stroke:#dc3545
|
||||
style D fill:#fff3cd,stroke:#ffc107
|
||||
```
|
||||
|
||||
`litellm.anthropic_models` is a Python `set` populated at import time by `add_known_models()`. It is the source `get_llm_provider()` consults to map a bare model name like `claude-sonnet-4-6` to the provider string `"anthropic"`.
|
||||
|
||||
---
|
||||
|
||||
## Root Cause
|
||||
|
||||
`add_known_models()` is called **once** at module import time. Both reload paths in `proxy_server.py` updated `litellm.model_cost` with the fresh map but never called `add_known_models()` again:
|
||||
|
||||
```python
|
||||
# Before the fix — both reload paths looked like this:
|
||||
new_model_cost_map = get_model_cost_map(url=model_cost_map_url)
|
||||
litellm.model_cost = new_model_cost_map # ✅ cost map updated
|
||||
_invalidate_model_cost_lowercase_map() # ✅ cache cleared
|
||||
# ❌ add_known_models() never called
|
||||
# → litellm.anthropic_models still has the old set
|
||||
# → new model not in the set
|
||||
# → get_llm_provider() raises for the new model
|
||||
# → wildcard match returns False
|
||||
# → 401 for every request to the new model
|
||||
```
|
||||
|
||||
The gap existed in two places:
|
||||
1. `_check_and_reload_model_cost_map` — the periodic automatic reload (every 10 s)
|
||||
2. The `/reload/model_cost_map` admin endpoint — the manual reload
|
||||
|
||||
**Timeline:**
|
||||
|
||||
1. New model (`claude-sonnet-4-6`) added to `model_prices_and_context_window.json`
|
||||
2. Admin triggers cost map reload via UI → `litellm.model_cost` updated
|
||||
3. Users with `anthropic/*` wildcard keys attempt requests to `claude-sonnet-4-6`
|
||||
4. `get_llm_provider('claude-sonnet-4-6')` raises → wildcard returns False → 401
|
||||
5. Admin reloads cost map again — same result (root cause not addressed)
|
||||
6. ~3 hours of investigation → root cause identified → fix deployed
|
||||
|
||||
---
|
||||
|
||||
## The Fix
|
||||
|
||||
After each reload, `add_known_models()` is called with the freshly fetched map passed explicitly. Passing the map directly (rather than relying on the module-level reference) removes any ambiguity about which dict is iterated:
|
||||
|
||||
```python
|
||||
# After the fix — both reload paths now do:
|
||||
new_model_cost_map = get_model_cost_map(url=model_cost_map_url)
|
||||
litellm.model_cost = new_model_cost_map
|
||||
_invalidate_model_cost_lowercase_map()
|
||||
litellm.add_known_models(model_cost_map=new_model_cost_map) # ✅ sets repopulated
|
||||
```
|
||||
|
||||
`add_known_models()` was also updated to accept an optional explicit map so callers cannot accidentally iterate a stale module-level reference:
|
||||
|
||||
```python
|
||||
# Before
|
||||
def add_known_models():
|
||||
for key, value in model_cost.items(): # reads module global — ambiguous after reload
|
||||
...
|
||||
|
||||
# After
|
||||
def add_known_models(model_cost_map: Optional[Dict] = None):
|
||||
_map = model_cost_map if model_cost_map is not None else model_cost
|
||||
for key, value in _map.items(): # always iterates the map you just fetched
|
||||
...
|
||||
```
|
||||
|
||||
After the fix, the provider sets (`anthropic_models`, `open_ai_chat_completion_models`, etc.) are always consistent with `litellm.model_cost` immediately after every reload. New models become accessible via wildcard rules without any proxy restart.
|
||||
|
||||
---
|
||||
|
||||
## Remediation
|
||||
|
||||
| # | Action | Status | Code |
|
||||
|---|---|---|---|
|
||||
| 1 | Call `add_known_models(model_cost_map=...)` in the periodic reload path | ✅ Done | [`proxy_server.py#L4393`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py#L4393) |
|
||||
| 2 | Call `add_known_models(model_cost_map=...)` in the `/reload/model_cost_map` endpoint | ✅ Done | [`proxy_server.py#L11904`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py#L11904) |
|
||||
| 3 | Update `add_known_models()` to accept an explicit map parameter | ✅ Done | [`__init__.py#L617`](https://github.com/BerriAI/litellm/blob/main/litellm/__init__.py#L617) |
|
||||
| 4 | Regression test: `add_known_models(model_cost_map=...)` populates provider sets | ✅ Done | [`test_auth_checks.py`](https://github.com/BerriAI/litellm/blob/main/tests/proxy_unit_tests/test_auth_checks.py) |
|
||||
| 5 | Regression test: `anthropic/*` wildcard grants/denies access correctly after reload | ✅ Done | [`test_auth_checks.py`](https://github.com/BerriAI/litellm/blob/main/tests/proxy_unit_tests/test_auth_checks.py) |
|
||||
|
||||
---
|
||||
|
|
@ -614,8 +614,9 @@ def is_openai_finetune_model(key: str) -> bool:
|
|||
return key.startswith("ft:") and not key.count(":") > 1
|
||||
|
||||
|
||||
def add_known_models():
|
||||
for key, value in model_cost.items():
|
||||
def add_known_models(model_cost_map: Optional[Dict] = None):
|
||||
_map = model_cost_map if model_cost_map is not None else model_cost
|
||||
for key, value in _map.items():
|
||||
if value.get("litellm_provider") == "openai" and not is_openai_finetune_model(
|
||||
key
|
||||
):
|
||||
|
|
|
|||
|
|
@ -11,6 +11,7 @@ export LITELLM_LOCAL_MODEL_COST_MAP=True
|
|||
import json
|
||||
import os
|
||||
from importlib.resources import files
|
||||
from typing import Optional
|
||||
|
||||
import httpx
|
||||
|
||||
|
|
@ -151,6 +152,37 @@ class GetModelCostMap:
|
|||
return response.json()
|
||||
|
||||
|
||||
class ModelCostMapSourceInfo:
|
||||
"""Tracks the source of the currently loaded model cost map."""
|
||||
|
||||
source: str = "local" # "local" or "remote"
|
||||
url: Optional[str] = None
|
||||
is_env_forced: bool = False
|
||||
fallback_reason: Optional[str] = None
|
||||
|
||||
|
||||
# Module-level singleton tracking the source of the current cost map
|
||||
_cost_map_source_info = ModelCostMapSourceInfo()
|
||||
|
||||
|
||||
def get_model_cost_map_source_info() -> dict:
|
||||
"""
|
||||
Return metadata about where the current model cost map was loaded from.
|
||||
|
||||
Returns a dict with:
|
||||
- source: "local" or "remote"
|
||||
- url: the remote URL attempted (or None for local-only)
|
||||
- is_env_forced: True if LITELLM_LOCAL_MODEL_COST_MAP=True forced local usage
|
||||
- fallback_reason: human-readable reason if remote failed and local was used
|
||||
"""
|
||||
return {
|
||||
"source": _cost_map_source_info.source,
|
||||
"url": _cost_map_source_info.url,
|
||||
"is_env_forced": _cost_map_source_info.is_env_forced,
|
||||
"fallback_reason": _cost_map_source_info.fallback_reason,
|
||||
}
|
||||
|
||||
|
||||
def get_model_cost_map(url: str) -> dict:
|
||||
"""
|
||||
Public entry point — returns the model cost map dict.
|
||||
|
|
@ -166,8 +198,15 @@ def get_model_cost_map(url: str) -> dict:
|
|||
# Note: can't use get_secret_bool here — this runs during litellm.__init__
|
||||
# before litellm._key_management_settings is set.
|
||||
if os.getenv("LITELLM_LOCAL_MODEL_COST_MAP", "").lower() == "true":
|
||||
_cost_map_source_info.source = "local"
|
||||
_cost_map_source_info.url = None
|
||||
_cost_map_source_info.is_env_forced = True
|
||||
_cost_map_source_info.fallback_reason = None
|
||||
return GetModelCostMap.load_local_model_cost_map()
|
||||
|
||||
_cost_map_source_info.url = url
|
||||
_cost_map_source_info.is_env_forced = False
|
||||
|
||||
try:
|
||||
content = GetModelCostMap.fetch_remote_model_cost_map(url)
|
||||
except Exception as e:
|
||||
|
|
@ -177,6 +216,8 @@ def get_model_cost_map(url: str) -> dict:
|
|||
url,
|
||||
str(e),
|
||||
)
|
||||
_cost_map_source_info.source = "local"
|
||||
_cost_map_source_info.fallback_reason = f"Remote fetch failed: {str(e)}"
|
||||
return GetModelCostMap.load_local_model_cost_map()
|
||||
|
||||
# Validate using cached count (cheap int comparison, no file I/O)
|
||||
|
|
@ -189,6 +230,10 @@ def get_model_cost_map(url: str) -> dict:
|
|||
"Using local backup instead. url=%s",
|
||||
url,
|
||||
)
|
||||
_cost_map_source_info.source = "local"
|
||||
_cost_map_source_info.fallback_reason = "Remote data failed integrity validation"
|
||||
return GetModelCostMap.load_local_model_cost_map()
|
||||
|
||||
_cost_map_source_info.source = "remote"
|
||||
_cost_map_source_info.fallback_reason = None
|
||||
return content
|
||||
|
|
|
|||
|
|
@ -1,4 +1,3 @@
|
|||
import anyio
|
||||
import asyncio
|
||||
import copy
|
||||
import enum
|
||||
|
|
@ -31,6 +30,7 @@ from typing import (
|
|||
get_type_hints,
|
||||
)
|
||||
|
||||
import anyio
|
||||
from pydantic import BaseModel, Json
|
||||
|
||||
from litellm._uuid import uuid
|
||||
|
|
@ -3598,9 +3598,6 @@ class ProxyConfig:
|
|||
parsed = value
|
||||
elif isinstance(value, str):
|
||||
import json
|
||||
|
||||
import yaml
|
||||
|
||||
try:
|
||||
parsed = yaml.safe_load(value)
|
||||
except (yaml.YAMLError, json.JSONDecodeError):
|
||||
|
|
@ -4388,6 +4385,9 @@ class ProxyConfig:
|
|||
litellm.model_cost = new_model_cost_map
|
||||
# Invalidate case-insensitive lookup map since model_cost was replaced
|
||||
_invalidate_model_cost_lowercase_map()
|
||||
# Repopulate provider model sets (e.g. litellm.anthropic_models) so that
|
||||
# wildcard patterns like "anthropic/*" include any newly added models.
|
||||
litellm.add_known_models(model_cost_map=new_model_cost_map)
|
||||
|
||||
# Update pod's in-memory last reload time
|
||||
last_model_cost_map_reload = current_time.isoformat()
|
||||
|
|
@ -11890,6 +11890,9 @@ async def reload_model_cost_map(
|
|||
litellm.model_cost = new_model_cost_map
|
||||
# Invalidate case-insensitive lookup map since model_cost was replaced
|
||||
_invalidate_model_cost_lowercase_map()
|
||||
# Repopulate provider model sets (e.g. litellm.anthropic_models) so that
|
||||
# wildcard patterns like "anthropic/*" include any newly added models.
|
||||
litellm.add_known_models(model_cost_map=new_model_cost_map)
|
||||
|
||||
# Update pod's in-memory last reload time
|
||||
global last_model_cost_map_reload
|
||||
|
|
@ -12144,6 +12147,55 @@ async def get_model_cost_map_reload_status(
|
|||
)
|
||||
|
||||
|
||||
@router.get(
|
||||
"/model/cost_map/source",
|
||||
tags=["model management"],
|
||||
dependencies=[Depends(user_api_key_auth)],
|
||||
include_in_schema=False,
|
||||
)
|
||||
async def get_model_cost_map_source(
|
||||
user_api_key_dict: UserAPIKeyAuth = Depends(user_api_key_auth),
|
||||
):
|
||||
"""
|
||||
ADMIN ONLY / MASTER KEY Only Endpoint
|
||||
|
||||
Returns information about where the current model cost/pricing data was loaded from.
|
||||
|
||||
Response fields:
|
||||
- source: "local" (bundled backup) or "remote" (fetched from URL)
|
||||
- url: the remote URL that was attempted (null when env-forced local)
|
||||
- is_env_forced: true if LITELLM_LOCAL_MODEL_COST_MAP=True forced local usage
|
||||
- fallback_reason: human-readable reason why remote failed (null on success)
|
||||
- model_count: number of models in the currently loaded cost map
|
||||
"""
|
||||
if user_api_key_dict.user_role != LitellmUserRoles.PROXY_ADMIN:
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail=f"Access denied. Admin role required. Current role: {user_api_key_dict.user_role}",
|
||||
)
|
||||
|
||||
try:
|
||||
from litellm.litellm_core_utils.get_model_cost_map import (
|
||||
get_model_cost_map_source_info,
|
||||
)
|
||||
|
||||
source_info = get_model_cost_map_source_info()
|
||||
model_count = len(litellm.model_cost) if litellm.model_cost else 0
|
||||
|
||||
return {
|
||||
**source_info,
|
||||
"model_count": model_count,
|
||||
}
|
||||
except Exception as e:
|
||||
verbose_proxy_logger.exception(
|
||||
f"Failed to get model cost map source info: {str(e)}"
|
||||
)
|
||||
raise HTTPException(
|
||||
status_code=500,
|
||||
detail=f"Failed to get model cost map source info: {str(e)}",
|
||||
)
|
||||
|
||||
|
||||
#### ANTHROPIC BETA HEADERS RELOAD ENDPOINTS ####
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -261,6 +261,132 @@ async def test_can_key_call_model_wildcard_access(key_models, model, expect_to_w
|
|||
print(e)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"key_models, model, expect_to_work",
|
||||
[
|
||||
# After a cost-map reload, add_known_models() updates anthropic_models so
|
||||
# the anthropic/* wildcard can match a newly-added Anthropic model.
|
||||
(["anthropic/*"], "claude-brand-new-model-reload-test", True),
|
||||
# Wrong provider wildcard must still be denied even after reload.
|
||||
(["openai/*"], "claude-brand-new-model-reload-test", False),
|
||||
],
|
||||
)
|
||||
@pytest.mark.asyncio
|
||||
async def test_wildcard_access_after_cost_map_reload(key_models, model, expect_to_work):
|
||||
"""
|
||||
Regression test: after a cost-map hot-reload, calling
|
||||
add_known_models(model_cost_map=new_map) must update litellm.anthropic_models
|
||||
so that the anthropic/* wildcard correctly grants (or denies) access to
|
||||
newly-added models.
|
||||
|
||||
Root cause: both reload paths in proxy_server.py only updated
|
||||
litellm.model_cost but never re-ran add_known_models(), so the provider sets
|
||||
stayed stale and wildcard matching failed for new models.
|
||||
|
||||
Fix: each reload now calls litellm.add_known_models(model_cost_map=new_map)
|
||||
with the fetched map passed explicitly to avoid any reference ambiguity.
|
||||
"""
|
||||
from litellm.proxy.auth.auth_checks import can_key_call_model
|
||||
|
||||
# Build a new cost map that includes the brand-new model — exactly what
|
||||
# proxy_server.py receives from get_model_cost_map() during a reload.
|
||||
new_cost_map = dict(litellm.model_cost)
|
||||
new_cost_map[model] = {
|
||||
"litellm_provider": "anthropic",
|
||||
"max_tokens": 8192,
|
||||
"input_cost_per_token": 0.000003,
|
||||
"output_cost_per_token": 0.000015,
|
||||
}
|
||||
|
||||
original_model_cost = litellm.model_cost
|
||||
litellm.model_cost = new_cost_map
|
||||
|
||||
# Confirm the model is NOT yet in the provider set before reload propagation.
|
||||
assert model not in litellm.anthropic_models
|
||||
|
||||
# Simulate what proxy_server.py now does after every reload.
|
||||
litellm.add_known_models(model_cost_map=new_cost_map)
|
||||
|
||||
# After add_known_models(), the model must be in the set.
|
||||
assert model in litellm.anthropic_models
|
||||
|
||||
llm_model_list = [
|
||||
{
|
||||
"model_name": "anthropic/*",
|
||||
"litellm_params": {"model": "anthropic/*", "api_key": "test-api-key"},
|
||||
"model_info": {"id": "test-id-anthropic-wildcard", "db_model": False},
|
||||
},
|
||||
{
|
||||
"model_name": "openai/*",
|
||||
"litellm_params": {"model": "openai/*", "api_key": "test-api-key"},
|
||||
"model_info": {"id": "test-id-openai-wildcard", "db_model": False},
|
||||
},
|
||||
]
|
||||
router = litellm.Router(model_list=llm_model_list)
|
||||
user_api_key_object = UserAPIKeyAuth(models=key_models)
|
||||
|
||||
try:
|
||||
if expect_to_work:
|
||||
await can_key_call_model(
|
||||
model=model,
|
||||
llm_model_list=llm_model_list,
|
||||
valid_token=user_api_key_object,
|
||||
llm_router=router,
|
||||
)
|
||||
else:
|
||||
with pytest.raises(Exception):
|
||||
await can_key_call_model(
|
||||
model=model,
|
||||
llm_model_list=llm_model_list,
|
||||
valid_token=user_api_key_object,
|
||||
llm_router=router,
|
||||
)
|
||||
finally:
|
||||
litellm.model_cost = original_model_cost
|
||||
litellm.anthropic_models.discard(model)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_add_known_models_explicit_map_updates_provider_sets():
|
||||
"""
|
||||
Regression test: after a cost-map hot-reload, calling
|
||||
add_known_models(model_cost_map=new_map) with the new map passed explicitly
|
||||
must add any new provider models to the correct provider sets so that
|
||||
wildcard access checks (anthropic/*, openai/*, …) work immediately.
|
||||
|
||||
This covers the proxy_server.py fix where both reload paths now call
|
||||
litellm.add_known_models(model_cost_map=new_model_cost_map) instead of
|
||||
relying on the module-level model_cost being up to date.
|
||||
"""
|
||||
fake_new_model = "claude-brand-new-explicit-map-test"
|
||||
|
||||
# Baseline: the model must not be in the sets before we do anything.
|
||||
assert fake_new_model not in litellm.anthropic_models
|
||||
|
||||
new_cost_map = dict(litellm.model_cost)
|
||||
new_cost_map[fake_new_model] = {
|
||||
"litellm_provider": "anthropic",
|
||||
"max_tokens": 8192,
|
||||
"input_cost_per_token": 0.000003,
|
||||
"output_cost_per_token": 0.000015,
|
||||
}
|
||||
|
||||
# Simulate what proxy_server.py does on reload.
|
||||
original_model_cost = litellm.model_cost
|
||||
litellm.model_cost = new_cost_map
|
||||
litellm.add_known_models(model_cost_map=new_cost_map)
|
||||
|
||||
try:
|
||||
assert fake_new_model in litellm.anthropic_models, (
|
||||
"add_known_models(model_cost_map=...) did not add the new model to "
|
||||
"litellm.anthropic_models — wildcard access checks would fail."
|
||||
)
|
||||
finally:
|
||||
# Clean up: restore original state.
|
||||
litellm.model_cost = original_model_cost
|
||||
litellm.anthropic_models.discard(fake_new_model)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_is_valid_fallback_model():
|
||||
from litellm.proxy.auth.auth_checks import is_valid_fallback_model
|
||||
|
|
|
|||
|
|
@ -466,6 +466,33 @@ export const cancelModelCostMapReload = async (accessToken: string) => {
|
|||
}
|
||||
};
|
||||
|
||||
export const getModelCostMapSource = async (accessToken: string) => {
|
||||
try {
|
||||
const url = proxyBaseUrl
|
||||
? `${proxyBaseUrl}/model/cost_map/source`
|
||||
: `/model/cost_map/source`;
|
||||
const response = await fetch(url, {
|
||||
method: "GET",
|
||||
headers: {
|
||||
[globalLitellmHeaderName]: `Bearer ${accessToken}`,
|
||||
"Content-Type": "application/json",
|
||||
},
|
||||
});
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text();
|
||||
throw new Error(`HTTP ${response.status}: ${errorText}`);
|
||||
}
|
||||
|
||||
const jsonData = await response.json();
|
||||
console.log("Model cost map source info:", jsonData);
|
||||
return jsonData;
|
||||
} catch (error) {
|
||||
console.error("Failed to get model cost map source info:", error);
|
||||
throw error;
|
||||
}
|
||||
};
|
||||
|
||||
export const getModelCostMapReloadStatus = async (accessToken: string) => {
|
||||
try {
|
||||
const url = proxyBaseUrl
|
||||
|
|
|
|||
|
|
@ -1,11 +1,12 @@
|
|||
import React, { useState, useEffect } from "react";
|
||||
import { Button, Popconfirm, Modal, InputNumber, Space, Typography, Tag, Card } from "antd";
|
||||
import { ReloadOutlined, ClockCircleOutlined, StopOutlined } from "@ant-design/icons";
|
||||
import { Button, Popconfirm, Modal, InputNumber, Space, Typography, Tag, Card, Tooltip, Divider } from "antd";
|
||||
import { ReloadOutlined, ClockCircleOutlined, StopOutlined, CloudOutlined, DatabaseOutlined, InfoCircleOutlined, WarningOutlined } from "@ant-design/icons";
|
||||
import {
|
||||
reloadModelCostMap,
|
||||
scheduleModelCostMapReload,
|
||||
cancelModelCostMapReload,
|
||||
getModelCostMapReloadStatus,
|
||||
getModelCostMapSource,
|
||||
} from "./networking";
|
||||
import NotificationsManager from "./molecules/notifications_manager";
|
||||
|
||||
|
|
@ -18,6 +19,14 @@ interface ReloadStatus {
|
|||
next_run: string | null;
|
||||
}
|
||||
|
||||
interface CostMapSourceInfo {
|
||||
source: "local" | "remote";
|
||||
url: string | null;
|
||||
is_env_forced: boolean;
|
||||
fallback_reason: string | null;
|
||||
model_count: number;
|
||||
}
|
||||
|
||||
interface PriceDataReloadProps {
|
||||
accessToken: string;
|
||||
onReloadSuccess?: () => void;
|
||||
|
|
@ -44,14 +53,18 @@ const PriceDataReload: React.FC<PriceDataReloadProps> = ({
|
|||
const [hours, setHours] = useState<number>(6);
|
||||
const [reloadStatus, setReloadStatus] = useState<ReloadStatus | null>(null);
|
||||
const [loadingStatus, setLoadingStatus] = useState(false);
|
||||
const [sourceInfo, setSourceInfo] = useState<CostMapSourceInfo | null>(null);
|
||||
const [loadingSource, setLoadingSource] = useState(false);
|
||||
|
||||
// Fetch status on component mount and periodically
|
||||
useEffect(() => {
|
||||
fetchReloadStatus();
|
||||
fetchSourceInfo();
|
||||
|
||||
// Refresh status every 30 seconds to keep it up to date
|
||||
const interval = setInterval(() => {
|
||||
fetchReloadStatus();
|
||||
fetchSourceInfo();
|
||||
}, 30000);
|
||||
|
||||
return () => clearInterval(interval);
|
||||
|
|
@ -80,6 +93,20 @@ const PriceDataReload: React.FC<PriceDataReloadProps> = ({
|
|||
}
|
||||
};
|
||||
|
||||
const fetchSourceInfo = async () => {
|
||||
if (!accessToken) return;
|
||||
|
||||
setLoadingSource(true);
|
||||
try {
|
||||
const info = await getModelCostMapSource(accessToken);
|
||||
setSourceInfo(info);
|
||||
} catch (error) {
|
||||
console.error("Failed to fetch cost map source info:", error);
|
||||
} finally {
|
||||
setLoadingSource(false);
|
||||
}
|
||||
};
|
||||
|
||||
const handleHardRefresh = async () => {
|
||||
if (!accessToken) {
|
||||
NotificationsManager.fromBackend("No access token available");
|
||||
|
|
@ -93,8 +120,9 @@ const PriceDataReload: React.FC<PriceDataReloadProps> = ({
|
|||
if (response.status === "success") {
|
||||
NotificationsManager.success(`Price data reloaded successfully! ${response.models_count || 0} models updated.`);
|
||||
onReloadSuccess?.();
|
||||
// Refresh status after successful reload
|
||||
// Refresh status and source info after successful reload
|
||||
await fetchReloadStatus();
|
||||
await fetchSourceInfo();
|
||||
} else {
|
||||
NotificationsManager.fromBackend("Failed to reload price data");
|
||||
}
|
||||
|
|
@ -284,7 +312,108 @@ const PriceDataReload: React.FC<PriceDataReloadProps> = ({
|
|||
)}
|
||||
</Space>
|
||||
|
||||
{/* Status Card */}
|
||||
{/* Cost Map Source Info Card */}
|
||||
{sourceInfo && (
|
||||
<Card
|
||||
size="small"
|
||||
style={{
|
||||
backgroundColor: sourceInfo.source === "remote" ? "#f0f7ff" : "#fff8f0",
|
||||
border: `1px solid ${sourceInfo.source === "remote" ? "#bae0ff" : "#ffd591"}`,
|
||||
borderRadius: 8,
|
||||
marginBottom: 12,
|
||||
}}
|
||||
>
|
||||
<Space direction="vertical" size="small" style={{ width: "100%" }}>
|
||||
{/* Header row */}
|
||||
<div style={{ display: "flex", alignItems: "center", gap: 8 }}>
|
||||
{sourceInfo.source === "remote" ? (
|
||||
<CloudOutlined style={{ color: "#1677ff", fontSize: 16 }} />
|
||||
) : (
|
||||
<DatabaseOutlined style={{ color: "#fa8c16", fontSize: 16 }} />
|
||||
)}
|
||||
<Text strong style={{ fontSize: "13px" }}>
|
||||
Pricing Data Source
|
||||
</Text>
|
||||
<Tag
|
||||
color={sourceInfo.source === "remote" ? "blue" : "orange"}
|
||||
style={{ marginLeft: "auto", fontWeight: 600, textTransform: "uppercase", fontSize: "11px" }}
|
||||
>
|
||||
{sourceInfo.source === "remote" ? "Remote" : "Local"}
|
||||
</Tag>
|
||||
</div>
|
||||
|
||||
<Divider style={{ margin: "6px 0" }} />
|
||||
|
||||
{/* Model count */}
|
||||
<div style={{ display: "flex", justifyContent: "space-between", alignItems: "center" }}>
|
||||
<Text type="secondary" style={{ fontSize: "12px" }}>
|
||||
Models loaded:
|
||||
</Text>
|
||||
<Text strong style={{ fontSize: "12px" }}>
|
||||
{sourceInfo.model_count.toLocaleString()}
|
||||
</Text>
|
||||
</div>
|
||||
|
||||
{/* URL (when remote or attempted) */}
|
||||
{sourceInfo.url && (
|
||||
<div style={{ display: "flex", justifyContent: "space-between", alignItems: "flex-start", gap: 8 }}>
|
||||
<Text type="secondary" style={{ fontSize: "12px", whiteSpace: "nowrap" }}>
|
||||
{sourceInfo.source === "remote" ? "Loaded from:" : "Attempted URL:"}
|
||||
</Text>
|
||||
<Tooltip title={sourceInfo.url}>
|
||||
<Text
|
||||
style={{
|
||||
fontSize: "11px",
|
||||
maxWidth: 240,
|
||||
overflow: "hidden",
|
||||
textOverflow: "ellipsis",
|
||||
whiteSpace: "nowrap",
|
||||
display: "block",
|
||||
color: "#1677ff",
|
||||
cursor: "default",
|
||||
}}
|
||||
>
|
||||
{sourceInfo.url}
|
||||
</Text>
|
||||
</Tooltip>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{/* Env forced notice */}
|
||||
{sourceInfo.is_env_forced && (
|
||||
<div style={{ display: "flex", alignItems: "center", gap: 6, marginTop: 2 }}>
|
||||
<InfoCircleOutlined style={{ color: "#fa8c16", fontSize: 12 }} />
|
||||
<Text type="secondary" style={{ fontSize: "11px" }}>
|
||||
Local mode forced via <code>LITELLM_LOCAL_MODEL_COST_MAP=True</code>
|
||||
</Text>
|
||||
</div>
|
||||
)}
|
||||
|
||||
{/* Fallback reason */}
|
||||
{sourceInfo.fallback_reason && (
|
||||
<div
|
||||
style={{
|
||||
display: "flex",
|
||||
alignItems: "flex-start",
|
||||
gap: 6,
|
||||
backgroundColor: "#fff7e6",
|
||||
border: "1px solid #ffd591",
|
||||
borderRadius: 4,
|
||||
padding: "4px 8px",
|
||||
marginTop: 2,
|
||||
}}
|
||||
>
|
||||
<WarningOutlined style={{ color: "#fa8c16", fontSize: 12, marginTop: 2 }} />
|
||||
<Text style={{ fontSize: "11px", color: "#614700" }}>
|
||||
Fell back to local: {sourceInfo.fallback_reason}
|
||||
</Text>
|
||||
</div>
|
||||
)}
|
||||
</Space>
|
||||
</Card>
|
||||
)}
|
||||
|
||||
{/* Reload Schedule Status Card */}
|
||||
{reloadStatus && (
|
||||
<Card
|
||||
size="small"
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue