mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-14 23:21:35 +00:00
docs fix
This commit is contained in:
parent
23b9d7181c
commit
aca901096d
1 changed files with 58 additions and 35 deletions
|
|
@ -1,62 +1,85 @@
|
|||
---
|
||||
slug: model-cost-map-incident
|
||||
title: "Incident Report: Broken Model Cost Map on main"
|
||||
title: "Incident Report: Invalid model cost map on main"
|
||||
date: 2026-02-10T10:00:00
|
||||
authors:
|
||||
- name: Ishaan Jaffer
|
||||
title: "CTO, LiteLLM"
|
||||
url: https://www.linkedin.com/in/ishaanjaffer/
|
||||
image_url: https://pbs.twimg.com/profile_images/1298587542745358340/DZv3Oj-h_400x400.jpg
|
||||
tags: [incident-report, stability, model-cost-map]
|
||||
tags: [incident-report, stability]
|
||||
hide_table_of_contents: false
|
||||
---
|
||||
|
||||
# Incident Report: Broken Model Cost Map on `main`
|
||||
# Incident Report: Invalid model cost map on `main`
|
||||
|
||||
## What happened?
|
||||
**Date:** January 27, 2026
|
||||
**Duration:** ~20 minutes
|
||||
**Severity:** Low
|
||||
**Status:** Resolved
|
||||
|
||||
A contributor PR with changes to the model cost map had a poorly formatted JSON entry (extra `{` bracket). When this was merged into `main` ([commit `562f0a0`](https://github.com/BerriAI/litellm/commit/562f0a028251750e3d75386bee0e630d9796d0df)) the remote `model_prices_and_context_window.json` became invalid JSON. Since `litellm` fetches this file from GitHub `main` at import time, every installation silently fell back to its local backup copy. Customers on older versions had backups missing newer models.
|
||||
## Summary
|
||||
|
||||
**Impact:**
|
||||
A malformed JSON entry in `model_prices_and_context_window.json` was merged to `main` ([`562f0a0`](https://github.com/BerriAI/litellm/commit/562f0a028251750e3d75386bee0e630d9796d0df)). This caused LiteLLM to silently fall back to a stale local copy of the model cost map. Users on older package versions lost cost tracking for **newer models only** (e.g. `azure/gpt-5.2`). No LLM calls were blocked.
|
||||
|
||||
- **SDK calls (`litellm.completion`)** -- Worked. The SDK catches model map errors internally and proceeds to the API.
|
||||
- **AI Gateway calls (proxy routing)** -- Worked. The proxy routes based on its config `model_list`, not the cost map.
|
||||
- **Cost tracking, `get_model_info` on SDK and proxy** -- Impacted for users relying on the model cost map. `get_model_info()` raised `"This model isn't mapped yet"` for models missing from the stale backup. The incident lasted ~20 minutes until we fixed it.
|
||||
## Impact
|
||||
|
||||
- **LLM calls (`litellm.completion`, proxy routing):** No impact.
|
||||
- **Cost tracking for newer models:** Impacted. Models not present in the local backup (e.g. `gpt-5.2`) returned `"This model isn't mapped yet"` during cost lookups. Older models already in the backup were unaffected.
|
||||
|
||||
{/* truncate */}
|
||||
|
||||
---
|
||||
|
||||
## What caused this?
|
||||
## How the model cost map fits into a request
|
||||
|
||||
1. At import time, `litellm` fetches `model_prices_and_context_window.json` from GitHub `main`
|
||||
2. If the fetch fails (network error, invalid JSON, etc.), it **silently** falls back to a local backup bundled with the installed package
|
||||
3. The bad commit broke the JSON on `main` -- every litellm installation hit the fallback
|
||||
4. Customers on older versions (e.g. v1.80.5) had backups missing 661+ newer models including `azure/gpt-5.2`
|
||||
5. Any call to `get_model_info("azure/gpt-5.2")` then raised `"This model isn't mapped yet"`
|
||||
The model cost map is **not** in the request path. It is only used **after** the LLM response comes back, inside a try/catch. A missing entry never blocks a call.
|
||||
|
||||
**Root cause:** No CI validation on the JSON file, and the fallback was completely silent -- no log, no warning.
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A["SDK / Proxy receives request"] --> B["Route to provider (Azure, OpenAI, ...)"]
|
||||
B --> C["LLM returns response"]
|
||||
C --> D{"Post-call: look up model in cost map"}
|
||||
D -->|found| E["Calculate cost, log spend"]
|
||||
D -->|not found| F["Log warning, return response with cost=0"]
|
||||
```
|
||||
|
||||
When the cost map lookup fails, the response is still returned to the caller. The only impact is that spend tracking reports `cost=0` for that request.
|
||||
|
||||
---
|
||||
|
||||
## Shippable improvements
|
||||
## Root cause
|
||||
|
||||
| # | Improvement | Status | Details |
|
||||
|---|---|---|---|
|
||||
| 1 | **CI validation for model cost map JSON** | Shipped | [PR #20605](https://github.com/BerriAI/litellm/pull/20605) -- Validates JSON schema + structure on every PR |
|
||||
| 2 | **Warning logging on fallback** | Shipped | `get_model_cost_map()` now logs a `WARNING` when falling back to backup instead of silently swallowing the error |
|
||||
| 3 | **Fetched JSON integrity validation** | Shipped | `GetModelCostMap.validate_model_cost_map()` checks fetched map is a dict, has minimum model count, and hasn't shrunk >50% vs backup |
|
||||
| 4 | **CI/CD resilience tests** | Shipped | `tests/llm_translation/test_model_cost_map_resilience.py` -- 13 tests for empty map, invalid JSON, network errors, shrinkage, and `litellm.completion()` resilience |
|
||||
| 5 | Keep backup file in sync on every release | Planned | Update backup as part of release so fallback data is never more than 1 release behind |
|
||||
| 6 | `LITELLM_LOCAL_MODEL_COST_MAP=True` as default for production | Planned | Eliminates runtime GitHub dependency entirely |
|
||||
| 7 | Health check endpoint for external deps | Planned | Proxy endpoint reporting status of all external fetches |
|
||||
LiteLLM fetches the model cost map from GitHub `main` at import time. If the fetch fails, it falls back to a local backup bundled with the package. The fallback was silent -- no warning was logged.
|
||||
|
||||
A contributor PR introduced an extra `{` bracket, producing invalid JSON. The remote fetch failed with `JSONDecodeError`, triggering the silent fallback. Users on older package versions had backup files missing newer models.
|
||||
|
||||
## Timeline
|
||||
|
||||
1. Malformed JSON merged to `main`
|
||||
2. LiteLLM installations fall back to local backup on next import
|
||||
3. Users report `"This model isn't mapped yet"` for newer models
|
||||
4. Bad commit identified and reverted (~20 minutes)
|
||||
|
||||
---
|
||||
|
||||
## Other upstream dependencies in the codebase
|
||||
## Remediation
|
||||
|
||||
| Dependency | Impact | Fallback | Silent? |
|
||||
|---|---|---|---|
|
||||
| **Model cost map** (GitHub `main`) | Critical -- cost tracking breaks | Local backup file | Now logs warning |
|
||||
| **JWT public keys** (`JWT_PUBLIC_KEY_URL`) | Critical -- auth breaks | None (raises exception) | No |
|
||||
| **OIDC UserInfo** (`oidc_userinfo_endpoint`) | Critical -- auth breaks | None (raises exception) | No |
|
||||
| **HuggingFace provider mapping** (`huggingface.co/api`) | Medium -- HF calls fail | Raises `HuggingFaceError` | No |
|
||||
| **Ollama model tags** (localhost) | Low | Static model list | Warning logged |
|
||||
| **Together AI model info** (`api.together.xyz`) | Low | Returns `None` | Silent |
|
||||
| # | Action | Status |
|
||||
|---|---|---|
|
||||
| 1 | CI validation on `model_prices_and_context_window.json` | Shipped ([PR #20605](https://github.com/BerriAI/litellm/pull/20605)) |
|
||||
| 2 | Warning log on fallback to local backup | Shipped |
|
||||
| 3 | Integrity validation of fetched map (min model count, shrinkage check) | Shipped |
|
||||
| 4 | Resilience test suite for bad/missing cost maps | Shipped |
|
||||
| 5 | Sync backup file on every release | Planned |
|
||||
| 6 | Default to local-only cost map in production | Planned |
|
||||
|
||||
## Other upstream dependencies
|
||||
|
||||
| Dependency | Impact if unavailable | Fallback |
|
||||
|---|---|---|
|
||||
| Model cost map (GitHub) | Cost tracking for newer models | Local backup (now with warning) |
|
||||
| JWT public keys | Auth fails | None |
|
||||
| OIDC UserInfo | Auth fails | None |
|
||||
| HuggingFace model API | HF provider calls fail | None |
|
||||
| Ollama tags (localhost) | Ollama model list stale | Static list |
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue