mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-12 23:01:41 +00:00
docs: add Web Search Cost Tracking section
Document how each provider bills for web search, the search_context_cost_per_query field in model_prices JSON, how to override pricing via proxy config, and how LiteLLM extracts web_search_requests from each provider's response.
This commit is contained in:
parent
4c99f3ddd8
commit
a0d1d22bcf
1 changed files with 75 additions and 0 deletions
|
|
@ -596,3 +596,78 @@ Expected Response
|
|||
|
||||
</TabItem>
|
||||
</Tabs>
|
||||
|
||||
## Web Search Cost Tracking
|
||||
|
||||
LiteLLM tracks web search costs automatically based on provider-specific billing models. The cost is added on top of the standard token-based pricing.
|
||||
|
||||
### How providers charge for web search
|
||||
|
||||
| Provider | Billing Unit | How it works |
|
||||
|----------|-------------|--------------|
|
||||
| **Gemini 3.x** (3-flash, 3-pro, 3.1-*) | Per search query | Each internal search query is billed individually. One prompt may trigger multiple queries. |
|
||||
| **Gemini 2.x** (2.0-flash, 2.5-flash, 2.5-pro) | Per grounded prompt | Flat fee per API call that uses grounding, regardless of how many queries are executed internally. |
|
||||
| **OpenAI** (gpt-4o-search, gpt-5-search) | Per search context size | Cost varies by `search_context_size` (`low`, `medium`, `high`). |
|
||||
| **Anthropic** (Claude with web search) | Per search request | Fixed cost per web search tool invocation. |
|
||||
| **Perplexity** (sonar, sonar-pro) | Per search context size | Cost varies by `search_context_size`. |
|
||||
|
||||
### Pricing configuration
|
||||
|
||||
Web search costs are defined in `model_prices_and_context_window.json` using the `search_context_cost_per_query` field:
|
||||
|
||||
```json
|
||||
{
|
||||
"gemini/gemini-3-flash-preview": {
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.014,
|
||||
"search_context_size_medium": 0.014,
|
||||
"search_context_size_high": 0.014
|
||||
}
|
||||
},
|
||||
"gemini/gemini-2.5-flash": {
|
||||
"search_context_cost_per_query": {
|
||||
"search_context_size_low": 0.035,
|
||||
"search_context_size_medium": 0.035,
|
||||
"search_context_size_high": 0.035
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
You can override these costs in your proxy config using `model_info`:
|
||||
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: gemini-3-flash
|
||||
litellm_params:
|
||||
model: gemini/gemini-3-flash-preview
|
||||
model_info:
|
||||
search_context_cost_per_query:
|
||||
search_context_size_low: 0.014
|
||||
search_context_size_medium: 0.014
|
||||
search_context_size_high: 0.014
|
||||
```
|
||||
|
||||
### How LiteLLM tracks search usage
|
||||
|
||||
The number of web search requests is stored in `usage.prompt_tokens_details.web_search_requests`. LiteLLM extracts this from each provider's response:
|
||||
|
||||
- **Gemini**: Extracted from `groundingMetadata.webSearchQueries` in the response. For Gemini 2.x, clamped to 1 (per-prompt billing).
|
||||
- **OpenAI**: Reported directly in the usage metadata.
|
||||
- **Anthropic**: Reported via `server_tool_use.web_search_requests`.
|
||||
- **xAI**: Mapped from `num_sources_used` in the response.
|
||||
|
||||
```python
|
||||
response = litellm.completion(
|
||||
model="gemini/gemini-3-flash-preview",
|
||||
messages=[{"role": "user", "content": "Latest tech news?"}],
|
||||
web_search_options={"search_context_size": "medium"},
|
||||
)
|
||||
|
||||
# Check web search usage
|
||||
print(response.usage.prompt_tokens_details.web_search_requests) # e.g., 3
|
||||
|
||||
# Get total cost (includes token cost + web search cost)
|
||||
cost = litellm.completion_cost(completion_response=response)
|
||||
print(f"Total cost: ${cost}")
|
||||
```
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue