litellm/litellm
Cesar Garcia 0e601d0bfe
Fix: Map Gemini cached_tokens to Langfuse cache_read_input_tokens (#18614)
* Fix: Map Gemini cached_tokens to Langfuse cache_read_input_tokens

Fixes #18520

## Problem
Langfuse integration was not capturing cached tokens from Gemini models.
Gemini returns cached tokens in `usage.prompt_tokens_details.cached_tokens`,
but Langfuse only read from top-level `usage.cache_read_input_tokens`
(which only Anthropic populates).

## Solution
Updated langfuse.py to check both locations:
1. First check top-level cache_read_input_tokens (for Anthropic)
2. Then check prompt_tokens_details.cached_tokens (for Gemini, OpenAI, others)

This ensures all providers' cached tokens are properly reported to Langfuse.

## Changes
- Modified litellm/integrations/langfuse/langfuse.py (lines 742-761)
- Added 3 unit tests in tests/test_litellm/integrations/langfuse/test_gemini_cached_tokens.py
- All existing Langfuse tests still pass (11/11)

## Testing
- test_cached_tokens_extraction: Verifies Gemini cached_tokens extraction
- test_cached_tokens_not_present: Backward compatibility (no cached_tokens)
- test_cached_tokens_is_zero: Edge case when cached_tokens = 0

* Refactor: Extract cache token logic into helper function

Address review feedback from @officer47p

- Created _extract_cache_read_input_tokens() helper function
- Reduces code bloat in _log_langfuse_v2 method
- Improves testability and reusability
- All tests still passing (11/11)
2026-01-06 01:34:09 +05:30
..
a2a_protocol [Refactor] litellm/init.py: lazy load LLMClientCache (#18008) 2025-12-16 05:44:06 -08:00
anthropic_interface [Feat] Unified Skills API - works across Anthropic, Vertex, Azure, Bedrock (#18232) 2025-12-19 18:55:59 +05:30
assistants Contributor PR - Support OPENAI_BASE_URL in addition to OPENAI_API_BASE (#9995) (#10423) 2025-04-29 21:27:37 -07:00
batch_completion
batches Revert batch utils with original logic 2025-12-10 17:01:21 +05:30
caching 3[Fix] CI/CD - logging_testing (#18204) 2025-12-18 10:52:24 -08:00
completion_extras fix(responses-api): use list format with input_text for tool results (#18257) 2025-12-20 13:46:14 +05:30
containers [Feat] Containers API - add new container API file management + UI Interface (#17745) 2025-12-09 17:33:26 -08:00
endpoints/speech/speech_to_completion_bridge fix: fix import errors 2025-09-14 09:32:21 -07:00
experimental_mcp_client [Feat] mcp resources support (#16800) 2025-11-20 14:53:44 -08:00
files Add support for expires after param 2025-12-12 10:01:18 +05:30
fine_tuning Litellm managed file updates combined (#11040) 2025-05-22 17:20:41 -07:00
google_genai google genai adapter inline data support (#18477) 2026-01-04 00:43:22 +05:30
images Add support for image generation via azure ad token 2025-12-24 15:15:20 +05:30
integrations Fix: Map Gemini cached_tokens to Langfuse cache_read_input_tokens (#18614) 2026-01-06 01:34:09 +05:30
interactions [Feat] Interactions API - allow using all litellm providers (interactions -> responses api bridge) (#18373) 2025-12-23 22:30:22 +05:30
litellm_core_utils fix: extract pure base64 data from data URLs for Ollama (#18465) 2026-01-04 00:47:38 +05:30
llms feat(zai): Add GLM-4.7 model with reasoning support (#18476) 2026-01-04 00:44:19 +05:30
ocr [Feat] /ocr - Add VertexAI OCR provider support + cost tracking (#16216) 2025-11-03 15:56:49 -08:00
passthrough fix bedrock passthrough auth issue (#16879) 2025-11-24 18:44:59 -08:00
proxy feat(mcp): parallelize tool fetching from multiple MCP servers (#18627) 2026-01-05 16:54:24 +05:30
rag code QA fixes 2025-12-20 13:56:04 +05:30
realtime_api refactor: Add lazy loading for get_llm_provider (#18591) 2026-01-02 13:18:57 -08:00
rerank_api Add tests for header forwarding 2025-12-12 17:54:17 +05:30
responses fix(cost_calculator): correct gpt-image-1 cost calculation using token-based pricing (#17906) 2026-01-02 23:08:52 +05:30
router_strategy 3[Fix] CI/CD - logging_testing (#18204) 2025-12-18 10:52:24 -08:00
router_utils refactor: Add lazy loading for get_llm_provider (#18591) 2026-01-02 13:18:57 -08:00
search [Bug Fix] Exa Search API - ensure request params are sent to Exa AI (#15855) 2025-10-23 11:56:30 -07:00
secret_managers feat: allow per-team Vault overrides when storing keys 2025-12-18 06:21:53 +09:00
skills [Feat] Unified Skills API - works across Anthropic, Vertex, Azure, Bedrock (#18232) 2025-12-19 18:55:59 +05:30
types added the option of adding langsmith tenant id in the env (#18623) 2026-01-06 01:19:27 +05:30
vector_store_files Vector store files Stable Release (#16643) 2025-11-15 13:00:33 -08:00
vector_stores Fix vector store configuration synchronization failure 2025-12-05 11:46:14 +05:30
videos fix: respect videos content db creds 2025-12-10 23:00:01 +05:30
__init__.py feat: lazy load heavy imports to reduce memory usage at import time (#18592) 2026-01-02 13:51:28 -08:00
_lazy_imports.py refactor: Add lazy loading for get_llm_provider (#18591) 2026-01-02 13:18:57 -08:00
_lazy_imports_registry.py refactor: lazy load get_llm_provider and remove_index_from_tool_calls (#18608) 2026-01-03 12:12:24 -08:00
_logging.py [Feat] Add support for returning images with gemini/gemini-2.5-flash-image-preview with /chat/completions (#13983) 2025-08-27 16:16:19 -07:00
_redis.py fix: Apply max_connections configuration to Redis async client (#15797) 2025-10-22 09:19:08 -07:00
_service_logger.py [⚡️ Python SDK import] - reduce python sdk import time by .3s (#12140) 2025-06-28 14:57:10 -07:00
_uuid.py Fix: revert fastuuid optional dependency, always use fastuuid in .__uid helper (#14941) 2025-09-26 09:14:20 -07:00
_version.py
budget_manager.py
constants.py [Feat] AI Gateway - Add support for Platform Fee / Margins (#18427) 2025-12-25 11:07:27 +05:30
cost.json
cost_calculator.py [Feat] AI Gateway - Add support for Platform Fee / Margins (#18427) 2025-12-25 11:07:27 +05:30
exceptions.py Adds support for returning Azure Content Policy error information when exceptions from Azure OpenAI occur (#16231) 2025-11-08 16:04:36 -08:00
main.py Remove double imports 2025-12-24 12:22:19 +05:30
model_prices_and_context_window_backup.json feat(zai): Add GLM-4.7 model with reasoning support (#18476) 2026-01-04 00:44:19 +05:30
mypy.ini fix mypy 2025-09-27 12:21:32 -07:00
py.typed
router.py fix(router): Validate routing_strategy at startup to fail fast with helpful error. (#18624) 2026-01-06 01:22:09 +05:30
scheduler.py
timeout.py
utils.py perf(utils): lazy load 15+ unused imports (#18616) 2026-01-03 16:16:17 -08:00