fix(logging): stop billing and logging response reads as LLM calls (#36890)

* fix(logging): stop billing and logging response reads as LLM calls

Retrieving, deleting or cancelling a stored response, and vector store management calls, run through the same logging lifecycle as inference. A retrieved response replays the usage of the call that created it, so every read priced it again and wrote a second spend log row for the same tokens. Non-inference calls now cost 0, report no usage, log no placeholder chat message, and get a litellm.responses_management operation name instead of reading as chat.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep billing background response jobs after the poll

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): use an empty list for read-call messages

A tuple matches no branch in the loggers that walk this value, so lunary's
parse_messages falls through to clean_message and raises AttributeError on the
success hook. An empty list reads as no messages everywhere: it satisfies the
isinstance(list) checks in newrelic, mlflow and datadog, iterates zero times in
traceloop and helicone, and is what StandardLoggingPayload.messages is typed to
hold. None would be type-legal too but is not iterable, so it trades one crash
for another in mlflow and traceloop.

* fix(otel): stop the legacy emitter reporting replayed tokens on response reads

The zeroing so far lands in the standard logging payload, which the legacy
OpenTelemetry emitter does not read for usage: it takes prompt, completion and
total tokens straight off the response object, so a retrieval span still carried
the token counts of the call that produced the response, and the token usage
histogram still recorded them. That emitter is the default, so the spend row said
zero while the trace said otherwise. The background cost poller keeps its counts,
the same exemption the pricing path already makes.

* fix(logging): keep billing a background response when its retrieval is read

A response created with background=true comes back queued and carries no usage, so
its create bills nothing. The retrieval that first sees the finished job is the only
place that job's tokens are ever visible, and pricing every read at zero therefore
loses the spend outright rather than deduplicating it. On a proxy without the
enterprise cost poller a background job ended up costing $0 end to end.

is_unbilled_non_inference_call now takes the response it is deciding about and treats
a background response the same way it already treats the poller's own read, which is
the same exemption seen from the other side. The legacy OpenTelemetry emitter's time
per output token metric picks up the read gate it was missing, so it stops dividing a
read's latency by the replayed completion token count.

* test(proxy): pass the read response to the non-inference predicate

The poller test called is_unbilled_non_inference_call with the pre-background signature, so it broke when the predicate gained the response it classifies. It now hands the predicate a foreground read, and asserts that the same read is free without the origin stamp, so the stamp is what the test proves.

* fix(otel): stop the v2 metrics recorder reporting replayed tokens on response reads

The v2 span builder sources usage from the standard logging payload, so the
earlier fix already zeroes it there. The metrics recorder reads response_obj
directly, so a responses-management read still recorded the original
generation's tokens into gen_ai.client.token.usage and divided generation time
by them for gen_ai.server.time_per_output_token.

The read still records operation and response duration, under the
litellm.responses_management operation, so it stays observable.

* fix(proxy): keep the response-cost headers on calls priced at zero

Pricing responses reads and vector-store management routes at zero dropped the whole
x-litellm-response-cost family off those replies. The header build reads a falsy zero as
a cost this response never recorded and filters it out, and a call that returns before
pricing stores no cost breakdown for the component headers to read, so a client parsing
the cost off a read got a KeyError where it had previously been handed a number.

Those calls now advertise the family at zero. Retrieving a background response, and the
cost poller's read of one, still report their real cost.

The params-taking form of the predicate moves from opentelemetry into
internal_call_metadata so the proxy header build and the OTEL recorders share one copy.

* fix(proxy): report a zero cost split only under a zero cost total

The component headers were filled from call-type membership alone, while the
total they sit beside keeps its real value when the read priced normally, so a
breakdown that had not landed by the time headers were built could advertise a
real total next to an all-zero split. The split is now reported as zero only
when the total agrees with it, and is otherwise left absent.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
This commit is contained in:
devin-ai-integration[bot] 2026-08-26 18:34:17 -07:00 • committed by GitHub
parent 77765fd302
commit 4bf40c4e8d
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
18 changed files with 818 additions and 15 deletions

View file

@ -1,6 +1,8 @@
"""
Polls LiteLLM_ManagedObjectTable to check if the response is complete.
Cost tracking is handled automatically by the get-responses call.
Cost tracking is handled by the get-responses call, which prices normally only because the
poll stamps itself with BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN; user-facing reads of the
same route are non-inference and free.
"""
from datetime import datetime, timedelta, timezone
@ -9,12 +11,14 @@ from typing import TYPE_CHECKING, Dict, Optional, cast
import litellm
from litellm._logging import verbose_proxy_logger
from litellm.constants import (
INTERNAL_CALL_ORIGIN_METADATA_KEY,
MANAGED_OBJECT_STALENESS_CUTOFF_DAYS,
MAX_OBJECTS_PER_POLL_CYCLE,
STALE_OBJECT_CLEANUP_BATCH_SIZE,
)
from litellm.responses.utils import ResponsesAPIRequestUtils
from litellm.types.llms.openai import ResponsesAPIResponse
from litellm.types.utils import BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN
if TYPE_CHECKING:
from litellm.proxy.utils import PrismaClient, ProxyLogging
@ -113,7 +117,8 @@ class CheckResponsesCost:
Check if background responses are complete and track their cost.
- Get all status="queued" or "in_progress" and file_purpose="response" jobs
- Query the provider to check if response is complete
- Cost is automatically tracked by the get-responses call
- Cost is tracked by the get-responses call, billed because the poll is stamped
with BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN
- Mark responses in a terminal state as complete in the database
"""
try:
@ -153,6 +158,7 @@ class CheckResponsesCost:
# Prepare metadata with model information for cost tracking
litellm_metadata = {
"user_api_key_user_id": job.created_by or "default-user-id",
INTERNAL_CALL_ORIGIN_METADATA_KEY: BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN,
}
# Add model information if available

View file

@ -1812,6 +1812,43 @@ BROWSER_SECURITY_HEADERS: Final[frozenset[str]] = frozenset(
UNSAFE_PROXY_RESPONSE_HEADERS: Final[frozenset[str]] = HTTP_FRAMING_HEADERS | BROWSER_SECURITY_HEADERS
# A retrieved response replays the usage of the call that created it, so pricing these
# read/management routes like inference bills the same tokens twice.
NON_INFERENCE_CALL_TYPES: Final[frozenset[str]] = frozenset(
{
"get_responses",
"aget_responses",
"delete_responses",
"adelete_responses",
"cancel_responses",
"acancel_responses",
"list_input_items",
"alist_input_items",
"vector_store_create",
"avector_store_create",
"vector_store_retrieve",
"avector_store_retrieve",
"vector_store_list",
"avector_store_list",
"vector_store_update",
"avector_store_update",
"vector_store_delete",
"avector_store_delete",
"vector_store_file_create",
"avector_store_file_create",
"vector_store_file_list",
"avector_store_file_list",
"vector_store_file_retrieve",
"avector_store_file_retrieve",
"vector_store_file_content",
"avector_store_file_content",
"vector_store_file_update",
"avector_store_file_update",
"vector_store_file_delete",
"avector_store_file_delete",
}
)
# PTU reservation rollup writes rows to LiteLLM_DailyTeamSpend with this
# sentinel api_key so PTU flat cost stays distinguishable from real per-request
# spend under the table's composite unique constraint.

View file

@ -22,6 +22,7 @@ from litellm.integrations.opentelemetry_utils.gen_ai_semconv import (
)
from litellm.integrations.otel.model.db_endpoint import db_span_attributes
from litellm.integrations.otel.model.semconv import Metric
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call_from_params
from litellm.litellm_core_utils.safe_json_dumps import safe_dumps
from litellm.litellm_core_utils.secret_redaction import redact_string
from litellm.litellm_core_utils.service_tier_utils import (
@ -1643,7 +1644,12 @@ class OpenTelemetry(OTELGenAISemconvMixin, CustomLogger):
if self._operation_duration_histogram:
self._operation_duration_histogram.record(duration_s, attributes=common_attrs)
if response_obj and (usage := response_obj.get("usage")) and self._token_usage_histogram:
if (
self._token_usage_histogram
and response_obj
and not is_unbilled_non_inference_call_from_params(kwargs.get("call_type"), params, response_obj)
and (usage := response_obj.get("usage"))
):
in_attrs: Final = {**common_attrs, TOKEN_TYPE_ATTRIBUTE: "input"}
out_attrs: Final = {**common_attrs, TOKEN_TYPE_ATTRIBUTE: "output"}
self._token_usage_histogram.record(usage.get("prompt_tokens", 0), attributes=in_attrs)
@ -1719,6 +1725,11 @@ class OpenTelemetry(OTELGenAISemconvMixin, CustomLogger):
if not self._time_per_output_token_histogram:
return
if is_unbilled_non_inference_call_from_params(
kwargs.get("call_type"), kwargs.get("litellm_params"), response_obj
):
return
# Get completion tokens from response_obj
completion_tokens = None
if response_obj and (usage := response_obj.get("usage")):
@ -2488,7 +2499,14 @@ class OpenTelemetry(OTELGenAISemconvMixin, CustomLogger):
self._set_service_tier_attributes(span=span, standard_logging_payload=standard_logging_payload)
usage: Final = response_obj and response_obj.get("usage")
usage: Final = (
response_obj.get("usage")
if response_obj
and not is_unbilled_non_inference_call_from_params(
kwargs.get("call_type"), litellm_params, response_obj
)
else None
)
if usage:
self.safe_set_attribute(
span=span,

View file

@ -32,6 +32,7 @@ class GenAIOperation(str, Enum):
EXECUTE_TOOL = "execute_tool" # MCP tool-call spans
LITELLM_VECTOR_STORE_MANAGEMENT = "litellm.vector_store_management"
LITELLM_VECTOR_STORE_FILE_MANAGEMENT = "litellm.vector_store_file_management"
LITELLM_RESPONSES_MANAGEMENT = "litellm.responses_management"
LITELLM_MODERATION = "litellm.moderation"
@ -383,6 +384,14 @@ _OPERATION_BY_CALL_TYPE: Final[dict[str, GenAIOperation]] = {
"aembedding": GenAIOperation.EMBEDDINGS,
"responses": GenAIOperation.CHAT,
"aresponses": GenAIOperation.CHAT,
"get_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
"aget_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
"delete_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
"adelete_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
"cancel_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
"acancel_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
"list_input_items": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
"alist_input_items": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
"image_generation": GenAIOperation.GENERATE_CONTENT,
"aimage_generation": GenAIOperation.GENERATE_CONTENT,
"moderation": GenAIOperation.LITELLM_MODERATION,

View file

@ -32,6 +32,7 @@ from litellm.integrations.otel.model.semconv import (
resolve_provider,
)
from litellm.integrations.otel.model.utils import to_seconds
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call_from_params
from litellm.litellm_core_utils.safe_json_dumps import safe_dumps
@ -198,16 +199,21 @@ class GenAIMetricRecorder:
) -> None:
common_attrs: Final = self._filter_attributes(self._bounded_attributes(kwargs))
duration_s: Final = (end_time - start_time).total_seconds()
usage_is_replayed: Final = is_unbilled_non_inference_call_from_params(
kwargs.get("call_type"), kwargs.get("litellm_params"), response_obj
)
self._metrics.operation_duration.record(duration_s, attributes=common_attrs)
self._record_token_usage(response_obj, common_attrs)
if not usage_is_replayed:
self._record_token_usage(response_obj, common_attrs)
cost: Final = kwargs.get("response_cost")
if cost:
self._metrics.token_cost.record(cost, attributes=common_attrs)
self._record_time_to_first_token(kwargs, common_attrs)
self._record_time_per_output_token(kwargs, response_obj, end_time, duration_s, common_attrs)
if not usage_is_replayed:
self._record_time_per_output_token(kwargs, response_obj, end_time, duration_s, common_attrs)
self._record_response_duration(kwargs, end_time, common_attrs)
def record_failure(

View file

@ -20,8 +20,8 @@ from __future__ import annotations
from collections.abc import Mapping
from typing import Final
from litellm.constants import INTERNAL_CALL_ORIGIN_METADATA_KEY
from litellm.types.utils import InternalCallOrigin
from litellm.constants import INTERNAL_CALL_ORIGIN_METADATA_KEY, NON_INFERENCE_CALL_TYPES
from litellm.types.utils import BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN, InternalCallOrigin
BUDGET_RESERVATION_METADATA_KEYS: Final = frozenset({"user_api_key_budget_reservation"})
@ -45,6 +45,60 @@ budget-checked like the request that spawned it. Everything else on the parent's
be a lie on a sub-call that runs after it returned."""
def is_background_response(response: object) -> bool:
"""Whether a retrieved object is a response created with ``background=true``.
Such a create returns ``status="queued"`` and no usage at all, so nothing has billed the
job by the time anyone reads it back. Accepts the response as a mapping or a model,
because the callers hold it in both shapes.
"""
if isinstance(response, Mapping):
return response.get("background") is True
return getattr(response, "background", None) is True
def is_unbilled_non_inference_call(
call_type: str | None,
metadata: Mapping[str, object] | None,
response: object,
) -> bool:
"""A read/management route priced at zero, because the usage it reports belongs to the
call that created the object it just read.
Retrieving a background response is the exception, and the enterprise cost poller's read
is the same exception seen from the other side: that job's create billed nothing, so its
retrieval is the only place the spend is ever visible. Pricing those at zero would lose
the spend rather than deduplicate it.
"""
if call_type not in NON_INFERENCE_CALL_TYPES:
return False
if is_background_response(response):
return False
if metadata is None:
return True
return metadata.get(INTERNAL_CALL_ORIGIN_METADATA_KEY) != BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN
def is_unbilled_non_inference_call_from_params(
call_type: str | None,
litellm_params: Mapping[str, object] | None,
response: object,
) -> bool:
""":func:`is_unbilled_non_inference_call` for callers holding raw ``litellm_params``.
The call-type membership test runs first so that inference traffic, which is every
request in a normal workload, never pays for the metadata merge behind it.
"""
if call_type not in NON_INFERENCE_CALL_TYPES:
return False
from litellm.litellm_core_utils.litellm_logging import StandardLoggingPayloadSetup
metadata: Final = (
StandardLoggingPayloadSetup.merge_litellm_metadata(litellm_params) if litellm_params is not None else None
)
return is_unbilled_non_inference_call(call_type, metadata, response)
def sanitize_user_api_key_auth(auth: object) -> object:
"""Copy of the auth object with its budget reservation removed; the cost callback
falls back to reading the reservation from inside the auth object."""

View file

@ -64,6 +64,7 @@ from litellm.integrations.mlflow import MlflowLogger
from litellm.integrations.sqs import SQSLogger
from litellm.litellm_core_utils.core_helpers import is_expected_client_error, reconstruct_model_name
from litellm.litellm_core_utils.get_litellm_params import get_litellm_params
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call
from litellm.litellm_core_utils.llm_cost_calc.guardrail_cost import (
cost_breakdown_with_guardrail,
guardrail_information_cost,
@ -1586,6 +1587,11 @@ class Logging(LiteLLMLoggingBaseClass):
if cache_hit is True:
return 0.0
if is_unbilled_non_inference_call(
self.call_type, StandardLoggingPayloadSetup.merge_litellm_metadata(self.litellm_params), result
):
return 0.0
transformed_result: Final = self._generate_content_result_as_model_response(result)
if transformed_result is not None:
result = transformed_result
@ -5057,7 +5063,7 @@ class StandardLoggingPayloadSetup:
return messages
@staticmethod
def merge_litellm_metadata(litellm_params: dict) -> dict:
def merge_litellm_metadata(litellm_params: Mapping[str, object]) -> dict:
"""
Merge both litellm_metadata and metadata from litellm_params.
@ -5819,7 +5825,7 @@ def get_standard_logging_object_payload(
cache_hit: Final = kwargs.get("cache_hit", False)
# Extract usage as a plain dict, avoiding Pydantic round-trip
raw_usage_dict: Final = StandardLoggingPayloadSetup.get_usage_as_dict(
response_obj=response_obj,
response_obj=None if is_unbilled_non_inference_call(call_type, metadata, response_obj) else response_obj,
combined_usage_object=cast(Usage | None, kwargs.get("combined_usage_object")),
)
usage_dict: Final = (

View file

@ -26,6 +26,7 @@ from litellm.constants import (
LITELLM_DETAILED_TIMING,
LITELLM_HTTP_STATUS_CLIENT_DISCONNECTED,
MAX_PAYLOAD_SIZE_FOR_DEBUG_LOG,
NON_INFERENCE_CALL_TYPES,
RETURN_RAW_MODEL_NAME_METADATA_KEY,
STREAM_SSE_DATA_PREFIX,
STREAM_SSE_KEEPALIVE_PING_BYTES,
@ -37,6 +38,7 @@ from litellm.litellm_core_utils.dd_tracing import NullTracer, tracer
from litellm.litellm_core_utils.get_supported_openai_params import (
get_supported_openai_params,
)
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call_from_params
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
from litellm.litellm_core_utils.llm_cost_calc.guardrail_cost import guardrail_information_cost
from litellm.litellm_core_utils.llm_response_utils.get_headers import (
@ -1300,15 +1302,51 @@ def _uncached_input_cost(
return input_cost - (cache_read_cost or 0.0) - (cache_creation_cost or 0.0)
_ZERO_COST_BREAKDOWN: Final = CostBreakdownHeaderValues(
original_cost=0.0,
discount_amount=0.0,
margin_total_amount=0.0,
margin_percent=0.0,
input_cost=0.0,
output_cost=0.0,
tool_usage_cost=0.0,
)
"""The component split a call priced at zero advertises, so a client reading the cost headers off a
read or management route still finds the whole family rather than a partially populated one."""
def _totals_to_zero(response_cost: float | str | None) -> bool:
"""Whether the total these headers carry is zero, counting a total no route ever priced as one.
A component split is only reported as zero alongside a total that agrees with it, so a read
that did price normally never advertises a real total beside an all-zero split.
"""
if response_cost is None or response_cost == "":
return True
try:
return float(response_cost) == 0.0
except (TypeError, ValueError):
return False
def _get_cost_breakdown_from_logging_obj(
litellm_logging_obj: LiteLLMLoggingObj | None,
response_cost: float | str | None = None,
) -> CostBreakdownHeaderValues:
"""Extract discount, margin, and per-component cost information from logging object's cost breakdown."""
"""Extract discount, margin, and per-component cost information from logging object's cost breakdown.
A non-inference call that priced at zero never records a breakdown, so its components are
reported as zero here. Any such call that did price normally (retrieving a background response,
and the cost poller's read of one) reports the breakdown it stored, or nothing at all when the
breakdown has not landed yet.
"""
if not litellm_logging_obj or not hasattr(litellm_logging_obj, "cost_breakdown"):
return CostBreakdownHeaderValues()
cost_breakdown: Final = litellm_logging_obj.cost_breakdown
if not cost_breakdown:
if litellm_logging_obj.call_type in NON_INFERENCE_CALL_TYPES and _totals_to_zero(response_cost):
return _ZERO_COST_BREAKDOWN
return CostBreakdownHeaderValues()
return CostBreakdownHeaderValues(
@ -1457,7 +1495,9 @@ class ProxyBaseLLMRequestProcessing:
exclude_values: Final = {"", None, "None"}
hidden_params = hidden_params or {}
cost_breakdown: Final = _get_cost_breakdown_from_logging_obj(litellm_logging_obj=litellm_logging_obj)
cost_breakdown: Final = _get_cost_breakdown_from_logging_obj(
litellm_logging_obj=litellm_logging_obj, response_cost=response_cost
)
# Calculate updated spend for header (include current response_cost)
current_spend: Final = user_api_key_dict.spend or 0.0
@ -2537,11 +2577,16 @@ class ProxyBaseLLMRequestProcessing:
additional_headers = hidden_params.get("additional_headers", {}) or {}
recover_response_cost: Final = not response_cost and hidden_params.get("response_cost") is None
llm_cost_for_headers: Final = (
computed_cost_for_headers: Final = (
self._response_cost_from_logging_obj(response=response, logging_obj=logging_obj) or ""
if recover_response_cost
else response_cost
)
llm_cost_for_headers: Final = (
0.0
if is_unbilled_non_inference_call_from_params(logging_obj.call_type, logging_obj.litellm_params, response)
else computed_cost_for_headers
)
_, request_metadata_bucket = get_or_create_metadata_bucket(self.data)
guardrail_cost_for_headers: Final = guardrail_information_cost(
request_metadata_bucket.get("standard_logging_guardrail_information")

View file

@ -22,6 +22,7 @@ from litellm.litellm_core_utils.core_helpers import (
get_litellm_metadata_from_kwargs,
reconstruct_model_name,
)
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call
from litellm.litellm_core_utils.litellm_logging import is_valid_sha256_hash
from litellm.litellm_core_utils.safe_json_dumps import safe_dumps, strip_null_bytes
from litellm.proxy._types import SpendLogsMetadata, SpendLogsPayload
@ -277,7 +278,7 @@ def get_logging_payload(kwargs, response_obj, start_time, end_time) -> SpendLogs
usage: dict = {}
if call_type in ["ocr", "aocr"]:
usage = _extract_usage_for_ocr_call(response_obj, response_obj_dict)
else:
elif not is_unbilled_non_inference_call(call_type, metadata, response_obj_dict):
# Use response_obj_dict instead of response_obj to avoid calling .get() on Pydantic models
_usage: Final = response_obj_dict.get("usage", None) or {}
if isinstance(_usage, litellm.Usage):

View file

@ -2834,13 +2834,19 @@ RoutingDecisionCause = Literal[
]
InternalCallOrigin = Literal["autorouter_classifier", "shadow_eval_router", "shadow_eval_judge"]
InternalCallOrigin = Literal[
"autorouter_classifier",
"shadow_eval_router",
"shadow_eval_judge",
"background_response_cost_poll",
]
"""Which internal litellm feature originated a billed sub-call, so a spend log row
records that it is not traffic the caller sent."""
AUTOROUTER_CLASSIFIER_CALL_ORIGIN: Final[InternalCallOrigin] = "autorouter_classifier"
SHADOW_EVAL_ROUTER_CALL_ORIGIN: Final[InternalCallOrigin] = "shadow_eval_router"
SHADOW_EVAL_JUDGE_CALL_ORIGIN: Final[InternalCallOrigin] = "shadow_eval_judge"
BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN: Final[InternalCallOrigin] = "background_response_cost_poll"
class StandardLoggingRoutingDecision(TypedDict, total=False):

View file

@ -74,6 +74,7 @@ from litellm.constants import (
MAX_RETRY_DELAY,
MAX_TOKEN_TRIMMING_ATTEMPTS,
MINIMUM_PROMPT_CACHE_TOKEN_COUNT_OVERRIDE,
NON_INFERENCE_CALL_TYPES,
OPENAI_EMBEDDING_PARAMS,
TOOL_CHOICE_OBJECT_TOKEN_COUNT,
)
@ -1109,6 +1110,8 @@ def function_setup(
except Exception as e:
verbose_logger.debug("Error extracting messages from Google contents: %s", e)
messages = "default-message-value"
elif call_type in NON_INFERENCE_CALL_TYPES:
messages = [] # mutable-ok: loggers require a list here and Logging copies it
else:
messages = "default-message-value"
stream = False

View file

@ -753,3 +753,41 @@ class TestCheckResponsesCost:
call_kwargs = mock_aget.call_args[1]
assert "model" not in call_kwargs.get("litellm_metadata", {})
assert "model_group" not in call_kwargs.get("litellm_metadata", {})
@pytest.mark.asyncio
async def test_poll_stamps_internal_call_origin_so_the_read_is_billed(
self, check_responses_cost_instance, mock_prisma_client
):
"""A background create returns queued with no usage, so this poll's retrieval is the only
place the job's spend is ever seen. Without the origin stamp it is priced at zero like a
user-facing read (LIT-5602) and the job is never billed."""
from litellm.constants import INTERNAL_CALL_ORIGIN_METADATA_KEY
from litellm.litellm_core_utils.internal_call_metadata import (
is_unbilled_non_inference_call,
)
mock_job = MagicMock()
mock_job.unified_object_id = "resp_test_billed"
mock_job.created_by = "test-user"
mock_job.id = "job-billed"
mock_job.file_object = {"model": "gpt-5", "id": "resp_test_billed"}
mock_prisma_client.db.litellm_managedobjecttable.find_many = AsyncMock(
return_value=[mock_job]
)
mock_prisma_client.db.litellm_managedobjecttable.update_many = AsyncMock(
return_value=0
)
mock_response = MagicMock()
mock_response.status = "completed"
with patch("litellm.aget_responses", new_callable=AsyncMock) as mock_aget:
mock_aget.return_value = mock_response
await check_responses_cost_instance.check_responses_cost()
metadata = mock_aget.call_args[1]["litellm_metadata"]
foreground_read = {"background": False}
assert metadata[INTERNAL_CALL_ORIGIN_METADATA_KEY] == "background_response_cost_poll"
assert is_unbilled_non_inference_call("aget_responses", metadata, foreground_read) is False
assert is_unbilled_non_inference_call("aget_responses", None, foreground_read) is True

View file

@ -201,6 +201,36 @@ def test_time_to_first_token_is_streaming_only():
assert names == set(ALL_METRICS) - {TIME_TO_FIRST_TOKEN}
def test_response_read_does_not_replay_the_generation_usage():
"""A responses-management read returns the ORIGINAL generation's usage on the
object it fetches. Recording it would add those tokens again on every poll, so
the two usage-derived instruments are skipped while the duration ones, which
describe the read itself, still fire."""
metrics = _drive_success(InMemoryMetricReader(), call_type="aget_responses")
assert TOKEN_USAGE not in metrics
assert TIME_PER_OUTPUT_TOKEN not in metrics
assert OPERATION_DURATION in metrics
assert RESPONSE_DURATION in metrics
def test_background_response_read_still_records_usage():
"""A background=true create returns no usage, so its completed read is the only
place the generation's tokens are ever seen. Skipping it would lose them
entirely rather than deduplicate them."""
reader = InMemoryMetricReader()
logger = _logger(reader, enable_metrics=True)
kwargs, response_obj, start, end = _build_call(call_type="aget_responses")
response_obj["background"] = True
asyncio.run(logger.async_log_success_event(kwargs, response_obj, start, end))
metrics = _metrics_by_name(reader)
by_type = {dp.attributes[TOKEN_TYPE]: dp for dp in metrics[TOKEN_USAGE]}
assert by_type["input"].sum == PROMPT_TOKENS
assert by_type["output"].sum == COMPLETION_TOKENS
assert TIME_PER_OUTPUT_TOKEN in metrics
def test_metrics_disabled_records_nothing():
"""enable_metrics=False: the recorder is never built, so the injected reader
sees no gen_ai.client.* series even though the success hook runs."""

View file

@ -268,6 +268,27 @@ def test_vector_store_file_management_is_not_chat(call_type):
assert resolve_operation(call_type).value == "litellm.vector_store_file_management"
@pytest.mark.parametrize(
"call_type",
[
f"{prefix}{operation}"
for operation in ("get_responses", "delete_responses", "cancel_responses", "list_input_items")
for prefix in ("", "a")
],
)
def test_responses_management_is_not_chat(call_type):
"""Fetching, deleting or cancelling a stored response runs no inference, so it must not
read as a chat completion: the retrieved object replays the original call's tokens and
would inflate the chat series on every read. Regression test for LIT-5602."""
assert resolve_operation(call_type) is GenAIOperation.LITELLM_RESPONSES_MANAGEMENT
assert resolve_operation(call_type).value == "litellm.responses_management"
def test_creating_a_response_is_still_chat():
"""Guards the test above: ``/v1/responses`` itself is a chat completion."""
assert resolve_operation("aresponses") is GenAIOperation.CHAT
_NON_CHAT_ROUTES: Final = (
("image_generation", GenAIOperation.GENERATE_CONTENT, GenAIOutputType.IMAGE),
("speech", GenAIOperation.GENERATE_CONTENT, GenAIOutputType.SPEECH),

View file

@ -6345,3 +6345,95 @@ class TestOpenTelemetryDatabaseSemconvAttributes(unittest.TestCase):
span = self._service_span(ServiceTypes.DB, "get_data", None)
self.assertEqual(span.attributes["db.system.name"], "postgresql")
self.assertNotIn("server.address", span.attributes)
class TestOpenTelemetryNonInferenceUsage(unittest.TestCase):
"""Reading a stored response replays the usage of the call that created it, so emitting those
token counts again on the read's span reports the same tokens a second time. Regression tests
for LIT-5602, covering the legacy emitter that runs by default."""
USAGE = {"prompt_tokens": 4000, "completion_tokens": 2000, "total_tokens": 6000}
TOKEN_KEYS = frozenset({"gen_ai.usage.input_tokens", "gen_ai.usage.output_tokens", "gen_ai.usage.total_tokens"})
BACKGROUND_POLL = {"internal_call_origin": "background_response_cost_poll"}
RESPONSE_OBJ = {"id": "resp_lit5602", "model": "gpt-4o", "usage": USAGE}
BACKGROUND_RESPONSE_OBJ = {**RESPONSE_OBJ, "background": True}
def _kwargs(self, call_type, litellm_metadata=None):
return {
"model": "gpt-4o",
"call_type": call_type,
"optional_params": {},
"litellm_params": {
"custom_llm_provider": "openai",
"litellm_metadata": litellm_metadata or {},
},
"standard_logging_object": {"id": "lit5602", "call_type": call_type, "metadata": {}},
}
def _token_attributes_on_span(self, call_type, litellm_metadata=None, response_obj=None):
otel = OpenTelemetry()
mock_span = MagicMock()
otel.set_attributes(
span=mock_span,
kwargs=self._kwargs(call_type, litellm_metadata),
response_obj=response_obj or dict(self.RESPONSE_OBJ),
)
return {call[0][0] for call in mock_span.set_attribute.call_args_list if call[0][0] in self.TOKEN_KEYS}
def _token_histogram_calls(self, call_type, litellm_metadata=None, response_obj=None):
otel = OpenTelemetry()
otel._operation_duration_histogram = MagicMock()
otel._token_usage_histogram = MagicMock()
otel._cost_histogram = None
now = datetime.now()
otel._record_metrics(
self._kwargs(call_type, litellm_metadata), response_obj or dict(self.RESPONSE_OBJ), now, now
)
return otel._token_usage_histogram.record.call_count
def _time_per_output_token_calls(self, call_type, litellm_metadata=None, response_obj=None):
otel = OpenTelemetry()
otel._time_per_output_token_histogram = MagicMock()
now = datetime.now()
otel._record_time_per_output_token_metric(
self._kwargs(call_type, litellm_metadata), response_obj or dict(self.RESPONSE_OBJ), now, 1.0, {}
)
return otel._time_per_output_token_histogram.record.call_count
def test_inference_call_still_reports_its_tokens_on_the_span(self):
self.assertEqual(self._token_attributes_on_span("acompletion"), set(self.TOKEN_KEYS))
def test_response_read_does_not_report_the_retrieved_tokens_on_the_span(self):
self.assertEqual(self._token_attributes_on_span("aget_responses"), set())
def test_background_cost_poll_read_still_reports_its_tokens_on_the_span(self):
self.assertEqual(self._token_attributes_on_span("aget_responses", self.BACKGROUND_POLL), set(self.TOKEN_KEYS))
def test_inference_call_still_records_the_token_usage_histogram(self):
self.assertEqual(self._token_histogram_calls("acompletion"), 2)
def test_response_read_does_not_record_the_token_usage_histogram(self):
self.assertEqual(self._token_histogram_calls("aget_responses"), 0)
def test_background_cost_poll_read_still_records_the_token_usage_histogram(self):
self.assertEqual(self._token_histogram_calls("aget_responses", self.BACKGROUND_POLL), 2)
def test_background_response_read_still_reports_its_tokens_on_the_span(self):
self.assertEqual(
self._token_attributes_on_span("aget_responses", response_obj=self.BACKGROUND_RESPONSE_OBJ),
set(self.TOKEN_KEYS),
)
def test_background_response_read_still_records_the_token_usage_histogram(self):
self.assertEqual(self._token_histogram_calls("aget_responses", response_obj=self.BACKGROUND_RESPONSE_OBJ), 2)
def test_inference_call_still_records_time_per_output_token(self):
self.assertEqual(self._time_per_output_token_calls("acompletion"), 1)
def test_response_read_does_not_divide_its_latency_by_the_retrieved_token_count(self):
self.assertEqual(self._time_per_output_token_calls("aget_responses"), 0)
def test_background_response_read_still_records_time_per_output_token(self):
self.assertEqual(
self._time_per_output_token_calls("aget_responses", response_obj=self.BACKGROUND_RESPONSE_OBJ), 1
)

View file

@ -5225,6 +5225,197 @@ async def test_restore_correlation_context_works_across_asyncio_task_boundary():
session_id_var.set("")
class TestNonInferenceCallTypesAreNotBilled:
"""A retrieved response replays the usage of the call that created it, so pricing a read
of it double bills the same tokens. Regression tests for LIT-5602."""
RETRIEVED_RESPONSE_USAGE = {"input_tokens": 4000, "output_tokens": 2000, "total_tokens": 6000}
BACKGROUND_POLL_METADATA = {"internal_call_origin": "background_response_cost_poll"}
def _logging_obj(self, call_type: str, litellm_metadata: dict | None = None):
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
obj = LiteLLMLoggingObj(
model="gpt-4o",
messages=[],
stream=False,
call_type=call_type,
start_time=time.time(),
litellm_call_id=f"lit5602-{call_type}",
function_id="fn-lit5602",
)
obj.update_environment_variables(
model="gpt-4o",
user="",
optional_params={},
litellm_params={
"api_base": "",
"custom_llm_provider": "openai",
"litellm_metadata": litellm_metadata or {},
},
)
return obj
def _retrieved_response(self, background: bool | None = None):
from litellm.types.llms.openai import ResponsesAPIResponse
return ResponsesAPIResponse(
id="resp_lit5602",
created_at=1234567890,
model="gpt-4o",
output=[],
usage=self.RETRIEVED_RESPONSE_USAGE,
background=background,
)
def test_creating_a_response_is_still_priced(self):
"""Guards the tests below: the same response object must cost money on the create path."""
cost = self._logging_obj("aresponses")._response_cost_calculator(result=self._retrieved_response())
assert cost is not None and cost > 0
@pytest.mark.parametrize(
"call_type",
[
"aget_responses",
"adelete_responses",
"acancel_responses",
"alist_input_items",
"avector_store_delete",
"avector_store_file_content",
"avector_store_file_delete",
],
)
def test_read_and_management_calls_cost_nothing(self, call_type):
cost = self._logging_obj(call_type)._response_cost_calculator(result=self._retrieved_response())
assert cost == 0.0
def test_retrieved_usage_is_not_re_reported_in_standard_logging_payload(self):
from litellm.litellm_core_utils.litellm_logging import (
get_standard_logging_object_payload,
)
from datetime import datetime
logging_obj = self._logging_obj("aget_responses")
now = datetime.now()
payload = get_standard_logging_object_payload(
kwargs={
"litellm_call_id": "lit5602-payload",
"model": "gpt-4o",
"call_type": "aget_responses",
"litellm_params": {},
},
init_response_obj=self._retrieved_response(),
start_time=now,
end_time=now,
logging_obj=logging_obj,
status="success",
)
assert payload is not None
assert payload["prompt_tokens"] == 0
assert payload["completion_tokens"] == 0
assert payload["total_tokens"] == 0
assert payload["response_cost"] == 0.0
def test_background_cost_poll_read_is_still_priced(self):
"""A background create returns queued with no usage, so the poller's read carries the job's
only billable usage. Zeroing it there means background jobs are never billed."""
cost = self._logging_obj(
"aget_responses", litellm_metadata=self.BACKGROUND_POLL_METADATA
)._response_cost_calculator(result=self._retrieved_response())
assert cost is not None and cost > 0
def test_background_cost_poll_reports_usage_in_standard_logging_payload(self):
from datetime import datetime
from litellm.litellm_core_utils.litellm_logging import (
get_standard_logging_object_payload,
)
now = datetime.now()
payload = get_standard_logging_object_payload(
kwargs={
"litellm_call_id": "lit5602-poll-payload",
"model": "gpt-4o",
"call_type": "aget_responses",
"litellm_params": {"litellm_metadata": self.BACKGROUND_POLL_METADATA},
},
init_response_obj=self._retrieved_response(),
start_time=now,
end_time=now,
logging_obj=self._logging_obj(
"aget_responses", litellm_metadata=self.BACKGROUND_POLL_METADATA
),
status="success",
)
assert payload is not None
assert payload["total_tokens"] == 6000
def test_reading_a_background_response_is_still_priced(self):
"""A background create answers queued with no usage at all, so whoever reads the finished
job is the first and only caller to see its tokens. Zeroing that read bills the job nothing."""
cost = self._logging_obj("aget_responses")._response_cost_calculator(
result=self._retrieved_response(background=True)
)
assert cost is not None and cost > 0
def test_reading_a_background_response_reports_usage_in_standard_logging_payload(self):
from datetime import datetime
from litellm.litellm_core_utils.litellm_logging import (
get_standard_logging_object_payload,
)
now = datetime.now()
payload = get_standard_logging_object_payload(
kwargs={
"litellm_call_id": "lit5602-background-payload",
"model": "gpt-4o",
"call_type": "aget_responses",
"litellm_params": {},
},
init_response_obj=self._retrieved_response(background=True),
start_time=now,
end_time=now,
logging_obj=self._logging_obj("aget_responses"),
status="success",
)
assert payload is not None
assert payload["total_tokens"] == 6000
def test_reading_a_foreground_response_is_still_free(self):
"""Guards the test above against a blanket exemption: an explicit background=false read was
already billed by its create and must stay at zero."""
cost = self._logging_obj("aget_responses")._response_cost_calculator(
result=self._retrieved_response(background=False)
)
assert cost == 0.0
def _read_call_messages(self):
logging_obj, _ = litellm.utils.function_setup(
original_function="aget_responses",
rules_obj=litellm.utils.Rules(),
start_time=time.time(),
**{"litellm_call_id": "lit5602-setup", "response_id": "resp_lit5602"},
)
return logging_obj.model_call_details["messages"]
def test_read_calls_do_not_log_a_placeholder_chat_message(self):
assert self._read_call_messages() == []
def test_read_call_messages_survive_a_logger_that_walks_them(self):
"""Loggers reach into this value expecting a chat history and branch on it being a list.
An empty list reads as no messages; a tuple matches no branch and crashes the success hook,
and None is not iterable where other loggers walk it."""
from litellm.integrations.lunary import parse_messages
assert parse_messages(self._read_call_messages()) == []
def _build_success_payload(logging_obj, kwargs):
import datetime

View file

@ -3274,6 +3274,82 @@ def test_user_traffic_carries_no_internal_call_origin():
assert metadata["internal_call_origin"] is None
def _spend_log_for_call_type(
call_type: str, internal_call_origin: str | None = None, background: bool | None = None
) -> dict:
from litellm.types.llms.openai import ResponsesAPIResponse
return cast(
dict,
get_logging_payload(
kwargs={
"model": "gpt-4o",
"call_type": call_type,
"response_cost": 0.0,
"litellm_params": {
"metadata": {
"user_api_key": "test-key",
"internal_call_origin": internal_call_origin,
}
},
},
response_obj=ResponsesAPIResponse(
id="resp_lit5602",
created_at=1234567890,
model="gpt-4o",
output=[],
usage={"input_tokens": 4000, "output_tokens": 2000, "total_tokens": 6000},
background=background,
),
start_time=datetime.datetime.now(timezone.utc),
end_time=datetime.datetime.now(timezone.utc),
),
)
def test_spend_log_for_response_retrieval_does_not_replay_the_created_responses_tokens():
"""A retrieved response carries the usage of the call that created it, so counting it again
bills the same tokens twice. Regression test for LIT-5602."""
payload = _spend_log_for_call_type("aget_responses")
assert payload["prompt_tokens"] == 0
assert payload["completion_tokens"] == 0
assert payload["total_tokens"] == 0
assert payload["spend"] == 0.0
def test_spend_log_for_background_response_cost_poll_counts_tokens():
"""The poller's read is where a background job's usage first shows up, so dropping it there
leaves the job unbilled forever."""
payload = _spend_log_for_call_type("aget_responses", internal_call_origin="background_response_cost_poll")
assert payload["total_tokens"] == 6000
def test_spend_log_for_background_response_retrieval_counts_tokens():
"""A background create answers queued carrying no usage, so its retrieval is the first and only
place the job's tokens are ever visible. Zeroing that read bills the whole job nothing on any
proxy that is not running the enterprise cost poller."""
payload = _spend_log_for_call_type("aget_responses", background=True)
assert payload["total_tokens"] == 6000
def test_spend_log_for_foreground_response_retrieval_still_counts_nothing():
"""Guards the test above against a blanket exemption: an explicit background=false read was
already billed by its create and must stay at zero."""
payload = _spend_log_for_call_type("aget_responses", background=False)
assert payload["total_tokens"] == 0
def test_spend_log_for_response_creation_still_counts_tokens():
"""Guards the test above: the same response object must still be counted on the create path."""
payload = _spend_log_for_call_type("aresponses")
assert payload["total_tokens"] == 6000
REDACTED_RESPONSE_PLACEHOLDER: Final = {"text": "redacted-by-litellm"}
CONSTANT_ID_FROM_HASHED_PLACEHOLDER: Final = "00fcbef15a3b0097e14b0ca016ed30a0"

View file

@ -26,6 +26,7 @@ from litellm.proxy.common_request_processing import (
_ClientDisconnectedBeforeFirstChunk,
_extract_error_from_sse_chunk,
_get_cost_breakdown_from_logging_obj,
CostBreakdownHeaderValues,
_has_attribute_error_in_chain,
_is_azure_model_router_request,
open_sse_before_first_byte,
@ -5018,6 +5019,169 @@ class TestResponseCostHeaderForTypedDictResponses:
assert fastapi_response.headers["x-litellm-response-cost"] == "0.00123"
class TestCostHeadersForCallsPricedAtZero:
"""
Regression for LIT-5602. Pricing responses reads and vector-store management routes at
zero dropped the entire x-litellm-response-cost family off those replies: the header
build reads a falsy zero as "this response never recorded a cost" and filters it out,
and a call that returns before pricing stores no cost breakdown for the component
headers to read. A client parsing the cost off a read got a KeyError where it had
previously been handed a number. Those calls now advertise the whole family at zero.
"""
@staticmethod
def _responses_read(*, background=False):
from litellm.types.llms.openai import ResponsesAPIResponse
return ResponsesAPIResponse(
id="resp_lit5602",
created_at=0,
model="gpt-4.1-mini",
object="response",
output=[],
status="completed",
background=background,
usage={"input_tokens": 10, "output_tokens": 5, "total_tokens": 15},
)
@staticmethod
def _logging_obj(*, call_type, recovered_cost=0.0):
logging_obj = MagicMock()
logging_obj.litellm_call_id = "call-lit5602"
logging_obj.call_type = call_type
logging_obj.litellm_params = {}
logging_obj.cost_breakdown = None
logging_obj.model_call_details = {"response_cost": recovered_cost}
logging_obj._response_cost_calculator = MagicMock(return_value=recovered_cost)
logging_obj._enqueue_deferred_logging = None
logging_obj._on_deferred_stream_complete = None
return logging_obj
async def _drive(self, *, monkeypatch, response, logging_obj, route_type):
import litellm.proxy.common_request_processing as crp
from litellm.proxy._types import UserAPIKeyAuth as RealUserAPIKeyAuth
async def fake_route_request(**kwargs):
async def _llm_call():
return response
return _llm_call()
monkeypatch.setattr(crp, "route_request", fake_route_request)
async def fake_post_call_success_hook(data, user_api_key_dict, response):
return response
proxy_logging_obj = MagicMock(spec=ProxyLogging)
proxy_logging_obj.during_call_hook = AsyncMock(return_value=None)
proxy_logging_obj.update_request_status = AsyncMock(return_value=None)
proxy_logging_obj.post_call_response_headers_hook = AsyncMock(return_value={})
proxy_logging_obj.post_call_success_hook = fake_post_call_success_hook
fastapi_response = Response()
processing_obj = ProxyBaseLLMRequestProcessing(data={"litellm_logging_obj": logging_obj})
with patch.object(
ProxyBaseLLMRequestProcessing, "_has_post_call_guardrails", return_value=False
):
await processing_obj.base_process_llm_request(
request=MagicMock(spec=Request, headers={}),
fastapi_response=fastapi_response,
user_api_key_dict=RealUserAPIKeyAuth(api_key="sk-test"),
route_type=route_type,
proxy_logging_obj=proxy_logging_obj,
general_settings={},
proxy_config=MagicMock(spec=ProxyConfig),
select_data_generator=None,
llm_router=None,
skip_pre_call_logic=True,
)
return fastapi_response
@pytest.mark.asyncio
async def test_responses_read_emits_the_cost_header_family_at_zero(self, monkeypatch):
fastapi_response = await self._drive(
monkeypatch=monkeypatch,
response=self._responses_read(),
logging_obj=self._logging_obj(call_type="aget_responses"),
route_type="aget_responses",
)
assert fastapi_response.headers["x-litellm-response-cost"] == "0.0"
for component in (
"original",
"discount-amount",
"margin-amount",
"margin-percent",
"input",
"output",
"tool-usage",
):
assert fastapi_response.headers[f"x-litellm-response-cost-{component}"] == "0.0"
@pytest.mark.asyncio
async def test_reading_a_background_response_keeps_its_real_cost(self, monkeypatch):
fastapi_response = await self._drive(
monkeypatch=monkeypatch,
response=self._responses_read(background=True),
logging_obj=self._logging_obj(call_type="aget_responses", recovered_cost=0.00042),
route_type="aget_responses",
)
assert float(fastapi_response.headers["x-litellm-response-cost"]) == pytest.approx(0.00042)
@pytest.mark.asyncio
async def test_an_inference_call_without_a_recorded_cost_still_omits_the_header(self, monkeypatch):
"""A chat completion has no zero-priced route, so a falsy cost there means the cost was
never recorded and the header stays absent rather than advertising a made-up zero."""
fastapi_response = await self._drive(
monkeypatch=monkeypatch,
response=SimpleNamespace(_hidden_params={}),
logging_obj=self._logging_obj(call_type="acompletion"),
route_type="acompletion",
)
assert "x-litellm-response-cost" not in fastapi_response.headers
def test_cost_breakdown_reports_zero_components_for_a_call_priced_at_zero(self):
breakdown = _get_cost_breakdown_from_logging_obj(
litellm_logging_obj=self._logging_obj(call_type="aget_responses")
)
assert breakdown.original_cost == 0.0
assert breakdown.input_cost == 0.0
assert breakdown.output_cost == 0.0
assert breakdown.tool_usage_cost == 0.0
def test_cost_breakdown_stays_empty_for_an_inference_call(self):
breakdown = _get_cost_breakdown_from_logging_obj(
litellm_logging_obj=self._logging_obj(call_type="acompletion")
)
assert breakdown == CostBreakdownHeaderValues()
def test_cost_breakdown_never_zeroes_the_split_under_a_real_total(self):
"""Reading a background response prices normally, so a breakdown that has not landed by the
time headers are built is reported as absent rather than as a zero split contradicting the
real total alongside it."""
breakdown = _get_cost_breakdown_from_logging_obj(
litellm_logging_obj=self._logging_obj(call_type="aget_responses"),
response_cost=1.96e-05,
)
assert breakdown == CostBreakdownHeaderValues()
def test_cost_breakdown_reports_zero_components_under_a_zero_total(self):
breakdown = _get_cost_breakdown_from_logging_obj(
litellm_logging_obj=self._logging_obj(call_type="aget_responses"),
response_cost=0.0,
)
assert breakdown.original_cost == 0.0
assert breakdown.input_cost == 0.0
assert breakdown.output_cost == 0.0
class TestPreCallWithFallbacksOnLocalRateLimit:
@pytest.mark.asyncio