From a0d1d22bcfa53f52de5f76ebdafbcde9cfad35c5 Mon Sep 17 00:00:00 2001 From: Chesars Date: Sun, 22 Mar 2026 18:25:01 -0300 Subject: [PATCH] docs: add Web Search Cost Tracking section Document how each provider bills for web search, the search_context_cost_per_query field in model_prices JSON, how to override pricing via proxy config, and how LiteLLM extracts web_search_requests from each provider's response. --- docs/my-website/docs/completion/web_search.md | 75 +++++++++++++++++++ 1 file changed, 75 insertions(+) diff --git a/docs/my-website/docs/completion/web_search.md b/docs/my-website/docs/completion/web_search.md index 1f5ba2dee4e..375e9a6375b 100644 --- a/docs/my-website/docs/completion/web_search.md +++ b/docs/my-website/docs/completion/web_search.md @@ -596,3 +596,78 @@ Expected Response + +## Web Search Cost Tracking + +LiteLLM tracks web search costs automatically based on provider-specific billing models. The cost is added on top of the standard token-based pricing. + +### How providers charge for web search + +| Provider | Billing Unit | How it works | +|----------|-------------|--------------| +| **Gemini 3.x** (3-flash, 3-pro, 3.1-*) | Per search query | Each internal search query is billed individually. One prompt may trigger multiple queries. | +| **Gemini 2.x** (2.0-flash, 2.5-flash, 2.5-pro) | Per grounded prompt | Flat fee per API call that uses grounding, regardless of how many queries are executed internally. | +| **OpenAI** (gpt-4o-search, gpt-5-search) | Per search context size | Cost varies by `search_context_size` (`low`, `medium`, `high`). | +| **Anthropic** (Claude with web search) | Per search request | Fixed cost per web search tool invocation. | +| **Perplexity** (sonar, sonar-pro) | Per search context size | Cost varies by `search_context_size`. | + +### Pricing configuration + +Web search costs are defined in `model_prices_and_context_window.json` using the `search_context_cost_per_query` field: + +```json +{ + "gemini/gemini-3-flash-preview": { + "search_context_cost_per_query": { + "search_context_size_low": 0.014, + "search_context_size_medium": 0.014, + "search_context_size_high": 0.014 + } + }, + "gemini/gemini-2.5-flash": { + "search_context_cost_per_query": { + "search_context_size_low": 0.035, + "search_context_size_medium": 0.035, + "search_context_size_high": 0.035 + } + } +} +``` + +You can override these costs in your proxy config using `model_info`: + +```yaml +model_list: + - model_name: gemini-3-flash + litellm_params: + model: gemini/gemini-3-flash-preview + model_info: + search_context_cost_per_query: + search_context_size_low: 0.014 + search_context_size_medium: 0.014 + search_context_size_high: 0.014 +``` + +### How LiteLLM tracks search usage + +The number of web search requests is stored in `usage.prompt_tokens_details.web_search_requests`. LiteLLM extracts this from each provider's response: + +- **Gemini**: Extracted from `groundingMetadata.webSearchQueries` in the response. For Gemini 2.x, clamped to 1 (per-prompt billing). +- **OpenAI**: Reported directly in the usage metadata. +- **Anthropic**: Reported via `server_tool_use.web_search_requests`. +- **xAI**: Mapped from `num_sources_used` in the response. + +```python +response = litellm.completion( + model="gemini/gemini-3-flash-preview", + messages=[{"role": "user", "content": "Latest tech news?"}], + web_search_options={"search_context_size": "medium"}, +) + +# Check web search usage +print(response.usage.prompt_tokens_details.web_search_requests) # e.g., 3 + +# Get total cost (includes token cost + web search cost) +cost = litellm.completion_cost(completion_response=response) +print(f"Total cost: ${cost}") +```