test(pricing): pin the realtime mode assertion to the bundled cost map (#33806)

test_get_model_info_reports_realtime_mode resolved gpt-realtime-mini through
litellm.get_model_info, which reads the cost map litellm fetches at import from
raw.githubusercontent.com/BerriAI/litellm/main. The mode=realtime retag from
#33728 is in this repo's json and its bundled backup but has not reached main
yet, so the test failed whenever the fetch succeeded and passed whenever the
runner was rate limited and litellm fell back to the backup, flapping the
Unit Tests: MCP, Secrets, Containers & Misc job on unrelated PRs

Resolve the lookup against the bundled backup instead, the way
tests/test_litellm/test_cost_calculator.py already does: force
LITELLM_LOCAL_MODEL_COST_MAP, rebind litellm.model_cost, and clear the
get_model_info lru cache before asserting so a remote-backed entry cached
earlier in the same worker cannot leak through, then clear it again afterwards
so no locally-backed entry outlives the test
This commit is contained in:
yuneng-jiang 2026-07-17 19:52:55 -07:00 • committed by GitHub
parent 40e914cfa7
commit c8b36dc1d4
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -68,8 +68,16 @@ def test_realtime_only_gpt_4o_models_are_mode_realtime(model):
assert _load_cost_map()[model]["mode"] == "realtime"
def test_get_model_info_reports_realtime_mode():
assert litellm.get_model_info("gpt-realtime-mini")["mode"] == "realtime"
def test_get_model_info_reports_realtime_mode(monkeypatch):
"""get_model_info must resolve the retag against the bundled cost map, not the
hosted map fetched from main, which lags this repo until the next promotion."""
monkeypatch.setenv("LITELLM_LOCAL_MODEL_COST_MAP", "True")
monkeypatch.setattr(litellm, "model_cost", litellm.get_model_cost_map(url=""))
litellm.get_model_info.cache_clear()
try:
assert litellm.get_model_info("gpt-realtime-mini")["mode"] == "realtime"
finally:
litellm.get_model_info.cache_clear()
def test_backup_matches_main_for_realtime_models():