This commit is contained in:
Ishaan Jaffer 2026-02-10 14:12:40 -08:00
parent 23b9d7181c
commit aca901096d

View file

@ -1,62 +1,85 @@
---
slug: model-cost-map-incident
title: "Incident Report: Broken Model Cost Map on main"
title: "Incident Report: Invalid model cost map on main"
date: 2026-02-10T10:00:00
authors:
- name: Ishaan Jaffer
title: "CTO, LiteLLM"
url: https://www.linkedin.com/in/ishaanjaffer/
image_url: https://pbs.twimg.com/profile_images/1298587542745358340/DZv3Oj-h_400x400.jpg
tags: [incident-report, stability, model-cost-map]
tags: [incident-report, stability]
hide_table_of_contents: false
---
# Incident Report: Broken Model Cost Map on `main`
# Incident Report: Invalid model cost map on `main`
## What happened?
**Date:** January 27, 2026
**Duration:** ~20 minutes
**Severity:** Low
**Status:** Resolved
A contributor PR with changes to the model cost map had a poorly formatted JSON entry (extra `{` bracket). When this was merged into `main` ([commit `562f0a0`](https://github.com/BerriAI/litellm/commit/562f0a028251750e3d75386bee0e630d9796d0df)) the remote `model_prices_and_context_window.json` became invalid JSON. Since `litellm` fetches this file from GitHub `main` at import time, every installation silently fell back to its local backup copy. Customers on older versions had backups missing newer models.
## Summary
**Impact:**
A malformed JSON entry in `model_prices_and_context_window.json` was merged to `main` ([`562f0a0`](https://github.com/BerriAI/litellm/commit/562f0a028251750e3d75386bee0e630d9796d0df)). This caused LiteLLM to silently fall back to a stale local copy of the model cost map. Users on older package versions lost cost tracking for **newer models only** (e.g. `azure/gpt-5.2`). No LLM calls were blocked.
- **SDK calls (`litellm.completion`)** -- Worked. The SDK catches model map errors internally and proceeds to the API.
- **AI Gateway calls (proxy routing)** -- Worked. The proxy routes based on its config `model_list`, not the cost map.
- **Cost tracking, `get_model_info` on SDK and proxy** -- Impacted for users relying on the model cost map. `get_model_info()` raised `"This model isn't mapped yet"` for models missing from the stale backup. The incident lasted ~20 minutes until we fixed it.
## Impact
- **LLM calls (`litellm.completion`, proxy routing):** No impact.
- **Cost tracking for newer models:** Impacted. Models not present in the local backup (e.g. `gpt-5.2`) returned `"This model isn't mapped yet"` during cost lookups. Older models already in the backup were unaffected.
{/* truncate */}
---
## What caused this?
## How the model cost map fits into a request
1. At import time, `litellm` fetches `model_prices_and_context_window.json` from GitHub `main`
2. If the fetch fails (network error, invalid JSON, etc.), it **silently** falls back to a local backup bundled with the installed package
3. The bad commit broke the JSON on `main` -- every litellm installation hit the fallback
4. Customers on older versions (e.g. v1.80.5) had backups missing 661+ newer models including `azure/gpt-5.2`
5. Any call to `get_model_info("azure/gpt-5.2")` then raised `"This model isn't mapped yet"`
The model cost map is **not** in the request path. It is only used **after** the LLM response comes back, inside a try/catch. A missing entry never blocks a call.
**Root cause:** No CI validation on the JSON file, and the fallback was completely silent -- no log, no warning.
```mermaid
flowchart LR
A["SDK / Proxy receives request"] --> B["Route to provider (Azure, OpenAI, ...)"]
B --> C["LLM returns response"]
C --> D{"Post-call: look up model in cost map"}
D -->|found| E["Calculate cost, log spend"]
D -->|not found| F["Log warning, return response with cost=0"]
```
When the cost map lookup fails, the response is still returned to the caller. The only impact is that spend tracking reports `cost=0` for that request.
---
## Shippable improvements
## Root cause
| # | Improvement | Status | Details |
|---|---|---|---|
| 1 | **CI validation for model cost map JSON** | Shipped | [PR #20605](https://github.com/BerriAI/litellm/pull/20605) -- Validates JSON schema + structure on every PR |
| 2 | **Warning logging on fallback** | Shipped | `get_model_cost_map()` now logs a `WARNING` when falling back to backup instead of silently swallowing the error |
| 3 | **Fetched JSON integrity validation** | Shipped | `GetModelCostMap.validate_model_cost_map()` checks fetched map is a dict, has minimum model count, and hasn't shrunk >50% vs backup |
| 4 | **CI/CD resilience tests** | Shipped | `tests/llm_translation/test_model_cost_map_resilience.py` -- 13 tests for empty map, invalid JSON, network errors, shrinkage, and `litellm.completion()` resilience |
| 5 | Keep backup file in sync on every release | Planned | Update backup as part of release so fallback data is never more than 1 release behind |
| 6 | `LITELLM_LOCAL_MODEL_COST_MAP=True` as default for production | Planned | Eliminates runtime GitHub dependency entirely |
| 7 | Health check endpoint for external deps | Planned | Proxy endpoint reporting status of all external fetches |
LiteLLM fetches the model cost map from GitHub `main` at import time. If the fetch fails, it falls back to a local backup bundled with the package. The fallback was silent -- no warning was logged.
A contributor PR introduced an extra `{` bracket, producing invalid JSON. The remote fetch failed with `JSONDecodeError`, triggering the silent fallback. Users on older package versions had backup files missing newer models.
## Timeline
1. Malformed JSON merged to `main`
2. LiteLLM installations fall back to local backup on next import
3. Users report `"This model isn't mapped yet"` for newer models
4. Bad commit identified and reverted (~20 minutes)
---
## Other upstream dependencies in the codebase
## Remediation
| Dependency | Impact | Fallback | Silent? |
|---|---|---|---|
| **Model cost map** (GitHub `main`) | Critical -- cost tracking breaks | Local backup file | Now logs warning |
| **JWT public keys** (`JWT_PUBLIC_KEY_URL`) | Critical -- auth breaks | None (raises exception) | No |
| **OIDC UserInfo** (`oidc_userinfo_endpoint`) | Critical -- auth breaks | None (raises exception) | No |
| **HuggingFace provider mapping** (`huggingface.co/api`) | Medium -- HF calls fail | Raises `HuggingFaceError` | No |
| **Ollama model tags** (localhost) | Low | Static model list | Warning logged |
| **Together AI model info** (`api.together.xyz`) | Low | Returns `None` | Silent |
| # | Action | Status |
|---|---|---|
| 1 | CI validation on `model_prices_and_context_window.json` | Shipped ([PR #20605](https://github.com/BerriAI/litellm/pull/20605)) |
| 2 | Warning log on fallback to local backup | Shipped |
| 3 | Integrity validation of fetched map (min model count, shrinkage check) | Shipped |
| 4 | Resilience test suite for bad/missing cost maps | Shipped |
| 5 | Sync backup file on every release | Planned |
| 6 | Default to local-only cost map in production | Planned |
## Other upstream dependencies
| Dependency | Impact if unavailable | Fallback |
|---|---|---|
| Model cost map (GitHub) | Cost tracking for newer models | Local backup (now with warning) |
| JWT public keys | Auth fails | None |
| OIDC UserInfo | Auth fails | None |
| HuggingFace model API | HF provider calls fail | None |
| Ollama tags (localhost) | Ollama model list stale | Static list |