mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-13 23:11:40 +00:00
Streaming and pass-through requests could be logged with $0 cost or dropped from
SpendLogs entirely while the upstream provider still billed every token. This
closes the leak paths not already covered by #30160, #30787 and #30788.
- Catch a stream_chunk_builder raise in the core CustomStreamWrapper (sync and
async). Large agentic tool-use / thinking streams can make assembly re-raise
as APIError from inside the except-StopIteration handler, where the sibling
except does not catch it, so it escaped __next__/__anext__ and dropped the
request; recover best-effort usage from the raw chunks instead
- Add a usage-only fallback for Anthropic streaming pass-through: when
stream_chunk_builder returns None or raises, rebuild usage from the
message_start / message_delta SSE events via AnthropicConfig.calculate_usage so
cache, web-search and geo tokens are priced instead of left at $0
- Decode buffered pass-through bytes with errors="replace" so a stream cut
mid-multibyte-sequence still logs the usage events already received
- Record response_cost into model_call_details on the pass-through success path
(it is read from there, not from kwargs), matching the gemini/cohere/openai
handlers
- Name the key (alias + masked key) in the virtual-key BudgetExceededError so
operators don't have to reverse-map spend back to a key
(cherry picked from commit
|
||
|---|---|---|
| .. | ||
| llm_provider_handlers | ||
| test_carry_guardrail_logging_info.py | ||
| test_llm_pass_through_endpoints.py | ||
| test_method_specific_routing.py | ||
| test_pass_through_endpoints.py | ||
| test_passthrough_auth_default.py | ||
| test_passthrough_endpoints_common_utils.py | ||
| test_passthrough_guardrail_block_otel_span.py | ||
| test_passthrough_guardrails.py | ||
| test_passthrough_guardrails_field_targeting.py | ||
| test_passthrough_post_call_guardrails.py | ||
| test_streaming_handler_interrupt.py | ||
| test_vertex_ai_batch_passthrough.py | ||
| test_vertex_passthrough_load_balancing.py | ||
| test_watsonx_proxy_route.py | ||