mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-09 03:18:44 +00:00
fix(logging): stop billing and logging response reads as LLM calls (#36890)
* fix(logging): stop billing and logging response reads as LLM calls Retrieving, deleting or cancelling a stored response, and vector store management calls, run through the same logging lifecycle as inference. A retrieved response replays the usage of the call that created it, so every read priced it again and wrote a second spend log row for the same tokens. Non-inference calls now cost 0, report no usage, log no placeholder chat message, and get a litellm.responses_management operation name instead of reading as chat. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(responses): keep billing background response jobs after the poll Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(logging): use an empty list for read-call messages A tuple matches no branch in the loggers that walk this value, so lunary's parse_messages falls through to clean_message and raises AttributeError on the success hook. An empty list reads as no messages everywhere: it satisfies the isinstance(list) checks in newrelic, mlflow and datadog, iterates zero times in traceloop and helicone, and is what StandardLoggingPayload.messages is typed to hold. None would be type-legal too but is not iterable, so it trades one crash for another in mlflow and traceloop. * fix(otel): stop the legacy emitter reporting replayed tokens on response reads The zeroing so far lands in the standard logging payload, which the legacy OpenTelemetry emitter does not read for usage: it takes prompt, completion and total tokens straight off the response object, so a retrieval span still carried the token counts of the call that produced the response, and the token usage histogram still recorded them. That emitter is the default, so the spend row said zero while the trace said otherwise. The background cost poller keeps its counts, the same exemption the pricing path already makes. * fix(logging): keep billing a background response when its retrieval is read A response created with background=true comes back queued and carries no usage, so its create bills nothing. The retrieval that first sees the finished job is the only place that job's tokens are ever visible, and pricing every read at zero therefore loses the spend outright rather than deduplicating it. On a proxy without the enterprise cost poller a background job ended up costing $0 end to end. is_unbilled_non_inference_call now takes the response it is deciding about and treats a background response the same way it already treats the poller's own read, which is the same exemption seen from the other side. The legacy OpenTelemetry emitter's time per output token metric picks up the read gate it was missing, so it stops dividing a read's latency by the replayed completion token count. * test(proxy): pass the read response to the non-inference predicate The poller test called is_unbilled_non_inference_call with the pre-background signature, so it broke when the predicate gained the response it classifies. It now hands the predicate a foreground read, and asserts that the same read is free without the origin stamp, so the stamp is what the test proves. * fix(otel): stop the v2 metrics recorder reporting replayed tokens on response reads The v2 span builder sources usage from the standard logging payload, so the earlier fix already zeroes it there. The metrics recorder reads response_obj directly, so a responses-management read still recorded the original generation's tokens into gen_ai.client.token.usage and divided generation time by them for gen_ai.server.time_per_output_token. The read still records operation and response duration, under the litellm.responses_management operation, so it stays observable. * fix(proxy): keep the response-cost headers on calls priced at zero Pricing responses reads and vector-store management routes at zero dropped the whole x-litellm-response-cost family off those replies. The header build reads a falsy zero as a cost this response never recorded and filters it out, and a call that returns before pricing stores no cost breakdown for the component headers to read, so a client parsing the cost off a read got a KeyError where it had previously been handed a number. Those calls now advertise the family at zero. Retrieving a background response, and the cost poller's read of one, still report their real cost. The params-taking form of the predicate moves from opentelemetry into internal_call_metadata so the proxy header build and the OTEL recorders share one copy. * fix(proxy): report a zero cost split only under a zero cost total The component headers were filled from call-type membership alone, while the total they sit beside keeps its real value when the read priced normally, so a breakdown that had not landed by the time headers were built could advertise a real total next to an all-zero split. The split is now reported as zero only when the total agrees with it, and is otherwise left absent. --------- Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
This commit is contained in:
parent
77765fd302
commit
4bf40c4e8d
18 changed files with 818 additions and 15 deletions
|
|
@ -1,6 +1,8 @@
|
|||
"""
|
||||
Polls LiteLLM_ManagedObjectTable to check if the response is complete.
|
||||
Cost tracking is handled automatically by the get-responses call.
|
||||
Cost tracking is handled by the get-responses call, which prices normally only because the
|
||||
poll stamps itself with BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN; user-facing reads of the
|
||||
same route are non-inference and free.
|
||||
"""
|
||||
|
||||
from datetime import datetime, timedelta, timezone
|
||||
|
|
@ -9,12 +11,14 @@ from typing import TYPE_CHECKING, Dict, Optional, cast
|
|||
import litellm
|
||||
from litellm._logging import verbose_proxy_logger
|
||||
from litellm.constants import (
|
||||
INTERNAL_CALL_ORIGIN_METADATA_KEY,
|
||||
MANAGED_OBJECT_STALENESS_CUTOFF_DAYS,
|
||||
MAX_OBJECTS_PER_POLL_CYCLE,
|
||||
STALE_OBJECT_CLEANUP_BATCH_SIZE,
|
||||
)
|
||||
from litellm.responses.utils import ResponsesAPIRequestUtils
|
||||
from litellm.types.llms.openai import ResponsesAPIResponse
|
||||
from litellm.types.utils import BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from litellm.proxy.utils import PrismaClient, ProxyLogging
|
||||
|
|
@ -113,7 +117,8 @@ class CheckResponsesCost:
|
|||
Check if background responses are complete and track their cost.
|
||||
- Get all status="queued" or "in_progress" and file_purpose="response" jobs
|
||||
- Query the provider to check if response is complete
|
||||
- Cost is automatically tracked by the get-responses call
|
||||
- Cost is tracked by the get-responses call, billed because the poll is stamped
|
||||
with BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN
|
||||
- Mark responses in a terminal state as complete in the database
|
||||
"""
|
||||
try:
|
||||
|
|
@ -153,6 +158,7 @@ class CheckResponsesCost:
|
|||
# Prepare metadata with model information for cost tracking
|
||||
litellm_metadata = {
|
||||
"user_api_key_user_id": job.created_by or "default-user-id",
|
||||
INTERNAL_CALL_ORIGIN_METADATA_KEY: BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN,
|
||||
}
|
||||
|
||||
# Add model information if available
|
||||
|
|
|
|||
|
|
@ -1812,6 +1812,43 @@ BROWSER_SECURITY_HEADERS: Final[frozenset[str]] = frozenset(
|
|||
|
||||
UNSAFE_PROXY_RESPONSE_HEADERS: Final[frozenset[str]] = HTTP_FRAMING_HEADERS | BROWSER_SECURITY_HEADERS
|
||||
|
||||
# A retrieved response replays the usage of the call that created it, so pricing these
|
||||
# read/management routes like inference bills the same tokens twice.
|
||||
NON_INFERENCE_CALL_TYPES: Final[frozenset[str]] = frozenset(
|
||||
{
|
||||
"get_responses",
|
||||
"aget_responses",
|
||||
"delete_responses",
|
||||
"adelete_responses",
|
||||
"cancel_responses",
|
||||
"acancel_responses",
|
||||
"list_input_items",
|
||||
"alist_input_items",
|
||||
"vector_store_create",
|
||||
"avector_store_create",
|
||||
"vector_store_retrieve",
|
||||
"avector_store_retrieve",
|
||||
"vector_store_list",
|
||||
"avector_store_list",
|
||||
"vector_store_update",
|
||||
"avector_store_update",
|
||||
"vector_store_delete",
|
||||
"avector_store_delete",
|
||||
"vector_store_file_create",
|
||||
"avector_store_file_create",
|
||||
"vector_store_file_list",
|
||||
"avector_store_file_list",
|
||||
"vector_store_file_retrieve",
|
||||
"avector_store_file_retrieve",
|
||||
"vector_store_file_content",
|
||||
"avector_store_file_content",
|
||||
"vector_store_file_update",
|
||||
"avector_store_file_update",
|
||||
"vector_store_file_delete",
|
||||
"avector_store_file_delete",
|
||||
}
|
||||
)
|
||||
|
||||
# PTU reservation rollup writes rows to LiteLLM_DailyTeamSpend with this
|
||||
# sentinel api_key so PTU flat cost stays distinguishable from real per-request
|
||||
# spend under the table's composite unique constraint.
|
||||
|
|
|
|||
|
|
@ -22,6 +22,7 @@ from litellm.integrations.opentelemetry_utils.gen_ai_semconv import (
|
|||
)
|
||||
from litellm.integrations.otel.model.db_endpoint import db_span_attributes
|
||||
from litellm.integrations.otel.model.semconv import Metric
|
||||
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call_from_params
|
||||
from litellm.litellm_core_utils.safe_json_dumps import safe_dumps
|
||||
from litellm.litellm_core_utils.secret_redaction import redact_string
|
||||
from litellm.litellm_core_utils.service_tier_utils import (
|
||||
|
|
@ -1643,7 +1644,12 @@ class OpenTelemetry(OTELGenAISemconvMixin, CustomLogger):
|
|||
|
||||
if self._operation_duration_histogram:
|
||||
self._operation_duration_histogram.record(duration_s, attributes=common_attrs)
|
||||
if response_obj and (usage := response_obj.get("usage")) and self._token_usage_histogram:
|
||||
if (
|
||||
self._token_usage_histogram
|
||||
and response_obj
|
||||
and not is_unbilled_non_inference_call_from_params(kwargs.get("call_type"), params, response_obj)
|
||||
and (usage := response_obj.get("usage"))
|
||||
):
|
||||
in_attrs: Final = {**common_attrs, TOKEN_TYPE_ATTRIBUTE: "input"}
|
||||
out_attrs: Final = {**common_attrs, TOKEN_TYPE_ATTRIBUTE: "output"}
|
||||
self._token_usage_histogram.record(usage.get("prompt_tokens", 0), attributes=in_attrs)
|
||||
|
|
@ -1719,6 +1725,11 @@ class OpenTelemetry(OTELGenAISemconvMixin, CustomLogger):
|
|||
if not self._time_per_output_token_histogram:
|
||||
return
|
||||
|
||||
if is_unbilled_non_inference_call_from_params(
|
||||
kwargs.get("call_type"), kwargs.get("litellm_params"), response_obj
|
||||
):
|
||||
return
|
||||
|
||||
# Get completion tokens from response_obj
|
||||
completion_tokens = None
|
||||
if response_obj and (usage := response_obj.get("usage")):
|
||||
|
|
@ -2488,7 +2499,14 @@ class OpenTelemetry(OTELGenAISemconvMixin, CustomLogger):
|
|||
|
||||
self._set_service_tier_attributes(span=span, standard_logging_payload=standard_logging_payload)
|
||||
|
||||
usage: Final = response_obj and response_obj.get("usage")
|
||||
usage: Final = (
|
||||
response_obj.get("usage")
|
||||
if response_obj
|
||||
and not is_unbilled_non_inference_call_from_params(
|
||||
kwargs.get("call_type"), litellm_params, response_obj
|
||||
)
|
||||
else None
|
||||
)
|
||||
if usage:
|
||||
self.safe_set_attribute(
|
||||
span=span,
|
||||
|
|
|
|||
|
|
@ -32,6 +32,7 @@ class GenAIOperation(str, Enum):
|
|||
EXECUTE_TOOL = "execute_tool" # MCP tool-call spans
|
||||
LITELLM_VECTOR_STORE_MANAGEMENT = "litellm.vector_store_management"
|
||||
LITELLM_VECTOR_STORE_FILE_MANAGEMENT = "litellm.vector_store_file_management"
|
||||
LITELLM_RESPONSES_MANAGEMENT = "litellm.responses_management"
|
||||
LITELLM_MODERATION = "litellm.moderation"
|
||||
|
||||
|
||||
|
|
@ -383,6 +384,14 @@ _OPERATION_BY_CALL_TYPE: Final[dict[str, GenAIOperation]] = {
|
|||
"aembedding": GenAIOperation.EMBEDDINGS,
|
||||
"responses": GenAIOperation.CHAT,
|
||||
"aresponses": GenAIOperation.CHAT,
|
||||
"get_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
|
||||
"aget_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
|
||||
"delete_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
|
||||
"adelete_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
|
||||
"cancel_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
|
||||
"acancel_responses": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
|
||||
"list_input_items": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
|
||||
"alist_input_items": GenAIOperation.LITELLM_RESPONSES_MANAGEMENT,
|
||||
"image_generation": GenAIOperation.GENERATE_CONTENT,
|
||||
"aimage_generation": GenAIOperation.GENERATE_CONTENT,
|
||||
"moderation": GenAIOperation.LITELLM_MODERATION,
|
||||
|
|
|
|||
|
|
@ -32,6 +32,7 @@ from litellm.integrations.otel.model.semconv import (
|
|||
resolve_provider,
|
||||
)
|
||||
from litellm.integrations.otel.model.utils import to_seconds
|
||||
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call_from_params
|
||||
from litellm.litellm_core_utils.safe_json_dumps import safe_dumps
|
||||
|
||||
|
||||
|
|
@ -198,16 +199,21 @@ class GenAIMetricRecorder:
|
|||
) -> None:
|
||||
common_attrs: Final = self._filter_attributes(self._bounded_attributes(kwargs))
|
||||
duration_s: Final = (end_time - start_time).total_seconds()
|
||||
usage_is_replayed: Final = is_unbilled_non_inference_call_from_params(
|
||||
kwargs.get("call_type"), kwargs.get("litellm_params"), response_obj
|
||||
)
|
||||
|
||||
self._metrics.operation_duration.record(duration_s, attributes=common_attrs)
|
||||
self._record_token_usage(response_obj, common_attrs)
|
||||
if not usage_is_replayed:
|
||||
self._record_token_usage(response_obj, common_attrs)
|
||||
|
||||
cost: Final = kwargs.get("response_cost")
|
||||
if cost:
|
||||
self._metrics.token_cost.record(cost, attributes=common_attrs)
|
||||
|
||||
self._record_time_to_first_token(kwargs, common_attrs)
|
||||
self._record_time_per_output_token(kwargs, response_obj, end_time, duration_s, common_attrs)
|
||||
if not usage_is_replayed:
|
||||
self._record_time_per_output_token(kwargs, response_obj, end_time, duration_s, common_attrs)
|
||||
self._record_response_duration(kwargs, end_time, common_attrs)
|
||||
|
||||
def record_failure(
|
||||
|
|
|
|||
|
|
@ -20,8 +20,8 @@ from __future__ import annotations
|
|||
from collections.abc import Mapping
|
||||
from typing import Final
|
||||
|
||||
from litellm.constants import INTERNAL_CALL_ORIGIN_METADATA_KEY
|
||||
from litellm.types.utils import InternalCallOrigin
|
||||
from litellm.constants import INTERNAL_CALL_ORIGIN_METADATA_KEY, NON_INFERENCE_CALL_TYPES
|
||||
from litellm.types.utils import BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN, InternalCallOrigin
|
||||
|
||||
BUDGET_RESERVATION_METADATA_KEYS: Final = frozenset({"user_api_key_budget_reservation"})
|
||||
|
||||
|
|
@ -45,6 +45,60 @@ budget-checked like the request that spawned it. Everything else on the parent's
|
|||
be a lie on a sub-call that runs after it returned."""
|
||||
|
||||
|
||||
def is_background_response(response: object) -> bool:
|
||||
"""Whether a retrieved object is a response created with ``background=true``.
|
||||
|
||||
Such a create returns ``status="queued"`` and no usage at all, so nothing has billed the
|
||||
job by the time anyone reads it back. Accepts the response as a mapping or a model,
|
||||
because the callers hold it in both shapes.
|
||||
"""
|
||||
if isinstance(response, Mapping):
|
||||
return response.get("background") is True
|
||||
return getattr(response, "background", None) is True
|
||||
|
||||
|
||||
def is_unbilled_non_inference_call(
|
||||
call_type: str | None,
|
||||
metadata: Mapping[str, object] | None,
|
||||
response: object,
|
||||
) -> bool:
|
||||
"""A read/management route priced at zero, because the usage it reports belongs to the
|
||||
call that created the object it just read.
|
||||
|
||||
Retrieving a background response is the exception, and the enterprise cost poller's read
|
||||
is the same exception seen from the other side: that job's create billed nothing, so its
|
||||
retrieval is the only place the spend is ever visible. Pricing those at zero would lose
|
||||
the spend rather than deduplicate it.
|
||||
"""
|
||||
if call_type not in NON_INFERENCE_CALL_TYPES:
|
||||
return False
|
||||
if is_background_response(response):
|
||||
return False
|
||||
if metadata is None:
|
||||
return True
|
||||
return metadata.get(INTERNAL_CALL_ORIGIN_METADATA_KEY) != BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN
|
||||
|
||||
|
||||
def is_unbilled_non_inference_call_from_params(
|
||||
call_type: str | None,
|
||||
litellm_params: Mapping[str, object] | None,
|
||||
response: object,
|
||||
) -> bool:
|
||||
""":func:`is_unbilled_non_inference_call` for callers holding raw ``litellm_params``.
|
||||
|
||||
The call-type membership test runs first so that inference traffic, which is every
|
||||
request in a normal workload, never pays for the metadata merge behind it.
|
||||
"""
|
||||
if call_type not in NON_INFERENCE_CALL_TYPES:
|
||||
return False
|
||||
from litellm.litellm_core_utils.litellm_logging import StandardLoggingPayloadSetup
|
||||
|
||||
metadata: Final = (
|
||||
StandardLoggingPayloadSetup.merge_litellm_metadata(litellm_params) if litellm_params is not None else None
|
||||
)
|
||||
return is_unbilled_non_inference_call(call_type, metadata, response)
|
||||
|
||||
|
||||
def sanitize_user_api_key_auth(auth: object) -> object:
|
||||
"""Copy of the auth object with its budget reservation removed; the cost callback
|
||||
falls back to reading the reservation from inside the auth object."""
|
||||
|
|
|
|||
|
|
@ -64,6 +64,7 @@ from litellm.integrations.mlflow import MlflowLogger
|
|||
from litellm.integrations.sqs import SQSLogger
|
||||
from litellm.litellm_core_utils.core_helpers import is_expected_client_error, reconstruct_model_name
|
||||
from litellm.litellm_core_utils.get_litellm_params import get_litellm_params
|
||||
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call
|
||||
from litellm.litellm_core_utils.llm_cost_calc.guardrail_cost import (
|
||||
cost_breakdown_with_guardrail,
|
||||
guardrail_information_cost,
|
||||
|
|
@ -1586,6 +1587,11 @@ class Logging(LiteLLMLoggingBaseClass):
|
|||
if cache_hit is True:
|
||||
return 0.0
|
||||
|
||||
if is_unbilled_non_inference_call(
|
||||
self.call_type, StandardLoggingPayloadSetup.merge_litellm_metadata(self.litellm_params), result
|
||||
):
|
||||
return 0.0
|
||||
|
||||
transformed_result: Final = self._generate_content_result_as_model_response(result)
|
||||
if transformed_result is not None:
|
||||
result = transformed_result
|
||||
|
|
@ -5057,7 +5063,7 @@ class StandardLoggingPayloadSetup:
|
|||
return messages
|
||||
|
||||
@staticmethod
|
||||
def merge_litellm_metadata(litellm_params: dict) -> dict:
|
||||
def merge_litellm_metadata(litellm_params: Mapping[str, object]) -> dict:
|
||||
"""
|
||||
Merge both litellm_metadata and metadata from litellm_params.
|
||||
|
||||
|
|
@ -5819,7 +5825,7 @@ def get_standard_logging_object_payload(
|
|||
cache_hit: Final = kwargs.get("cache_hit", False)
|
||||
# Extract usage as a plain dict, avoiding Pydantic round-trip
|
||||
raw_usage_dict: Final = StandardLoggingPayloadSetup.get_usage_as_dict(
|
||||
response_obj=response_obj,
|
||||
response_obj=None if is_unbilled_non_inference_call(call_type, metadata, response_obj) else response_obj,
|
||||
combined_usage_object=cast(Usage | None, kwargs.get("combined_usage_object")),
|
||||
)
|
||||
usage_dict: Final = (
|
||||
|
|
|
|||
|
|
@ -26,6 +26,7 @@ from litellm.constants import (
|
|||
LITELLM_DETAILED_TIMING,
|
||||
LITELLM_HTTP_STATUS_CLIENT_DISCONNECTED,
|
||||
MAX_PAYLOAD_SIZE_FOR_DEBUG_LOG,
|
||||
NON_INFERENCE_CALL_TYPES,
|
||||
RETURN_RAW_MODEL_NAME_METADATA_KEY,
|
||||
STREAM_SSE_DATA_PREFIX,
|
||||
STREAM_SSE_KEEPALIVE_PING_BYTES,
|
||||
|
|
@ -37,6 +38,7 @@ from litellm.litellm_core_utils.dd_tracing import NullTracer, tracer
|
|||
from litellm.litellm_core_utils.get_supported_openai_params import (
|
||||
get_supported_openai_params,
|
||||
)
|
||||
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call_from_params
|
||||
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
|
||||
from litellm.litellm_core_utils.llm_cost_calc.guardrail_cost import guardrail_information_cost
|
||||
from litellm.litellm_core_utils.llm_response_utils.get_headers import (
|
||||
|
|
@ -1300,15 +1302,51 @@ def _uncached_input_cost(
|
|||
return input_cost - (cache_read_cost or 0.0) - (cache_creation_cost or 0.0)
|
||||
|
||||
|
||||
_ZERO_COST_BREAKDOWN: Final = CostBreakdownHeaderValues(
|
||||
original_cost=0.0,
|
||||
discount_amount=0.0,
|
||||
margin_total_amount=0.0,
|
||||
margin_percent=0.0,
|
||||
input_cost=0.0,
|
||||
output_cost=0.0,
|
||||
tool_usage_cost=0.0,
|
||||
)
|
||||
"""The component split a call priced at zero advertises, so a client reading the cost headers off a
|
||||
read or management route still finds the whole family rather than a partially populated one."""
|
||||
|
||||
|
||||
def _totals_to_zero(response_cost: float | str | None) -> bool:
|
||||
"""Whether the total these headers carry is zero, counting a total no route ever priced as one.
|
||||
|
||||
A component split is only reported as zero alongside a total that agrees with it, so a read
|
||||
that did price normally never advertises a real total beside an all-zero split.
|
||||
"""
|
||||
if response_cost is None or response_cost == "":
|
||||
return True
|
||||
try:
|
||||
return float(response_cost) == 0.0
|
||||
except (TypeError, ValueError):
|
||||
return False
|
||||
|
||||
|
||||
def _get_cost_breakdown_from_logging_obj(
|
||||
litellm_logging_obj: LiteLLMLoggingObj | None,
|
||||
response_cost: float | str | None = None,
|
||||
) -> CostBreakdownHeaderValues:
|
||||
"""Extract discount, margin, and per-component cost information from logging object's cost breakdown."""
|
||||
"""Extract discount, margin, and per-component cost information from logging object's cost breakdown.
|
||||
|
||||
A non-inference call that priced at zero never records a breakdown, so its components are
|
||||
reported as zero here. Any such call that did price normally (retrieving a background response,
|
||||
and the cost poller's read of one) reports the breakdown it stored, or nothing at all when the
|
||||
breakdown has not landed yet.
|
||||
"""
|
||||
if not litellm_logging_obj or not hasattr(litellm_logging_obj, "cost_breakdown"):
|
||||
return CostBreakdownHeaderValues()
|
||||
|
||||
cost_breakdown: Final = litellm_logging_obj.cost_breakdown
|
||||
if not cost_breakdown:
|
||||
if litellm_logging_obj.call_type in NON_INFERENCE_CALL_TYPES and _totals_to_zero(response_cost):
|
||||
return _ZERO_COST_BREAKDOWN
|
||||
return CostBreakdownHeaderValues()
|
||||
|
||||
return CostBreakdownHeaderValues(
|
||||
|
|
@ -1457,7 +1495,9 @@ class ProxyBaseLLMRequestProcessing:
|
|||
exclude_values: Final = {"", None, "None"}
|
||||
hidden_params = hidden_params or {}
|
||||
|
||||
cost_breakdown: Final = _get_cost_breakdown_from_logging_obj(litellm_logging_obj=litellm_logging_obj)
|
||||
cost_breakdown: Final = _get_cost_breakdown_from_logging_obj(
|
||||
litellm_logging_obj=litellm_logging_obj, response_cost=response_cost
|
||||
)
|
||||
|
||||
# Calculate updated spend for header (include current response_cost)
|
||||
current_spend: Final = user_api_key_dict.spend or 0.0
|
||||
|
|
@ -2537,11 +2577,16 @@ class ProxyBaseLLMRequestProcessing:
|
|||
additional_headers = hidden_params.get("additional_headers", {}) or {}
|
||||
|
||||
recover_response_cost: Final = not response_cost and hidden_params.get("response_cost") is None
|
||||
llm_cost_for_headers: Final = (
|
||||
computed_cost_for_headers: Final = (
|
||||
self._response_cost_from_logging_obj(response=response, logging_obj=logging_obj) or ""
|
||||
if recover_response_cost
|
||||
else response_cost
|
||||
)
|
||||
llm_cost_for_headers: Final = (
|
||||
0.0
|
||||
if is_unbilled_non_inference_call_from_params(logging_obj.call_type, logging_obj.litellm_params, response)
|
||||
else computed_cost_for_headers
|
||||
)
|
||||
_, request_metadata_bucket = get_or_create_metadata_bucket(self.data)
|
||||
guardrail_cost_for_headers: Final = guardrail_information_cost(
|
||||
request_metadata_bucket.get("standard_logging_guardrail_information")
|
||||
|
|
|
|||
|
|
@ -22,6 +22,7 @@ from litellm.litellm_core_utils.core_helpers import (
|
|||
get_litellm_metadata_from_kwargs,
|
||||
reconstruct_model_name,
|
||||
)
|
||||
from litellm.litellm_core_utils.internal_call_metadata import is_unbilled_non_inference_call
|
||||
from litellm.litellm_core_utils.litellm_logging import is_valid_sha256_hash
|
||||
from litellm.litellm_core_utils.safe_json_dumps import safe_dumps, strip_null_bytes
|
||||
from litellm.proxy._types import SpendLogsMetadata, SpendLogsPayload
|
||||
|
|
@ -277,7 +278,7 @@ def get_logging_payload(kwargs, response_obj, start_time, end_time) -> SpendLogs
|
|||
usage: dict = {}
|
||||
if call_type in ["ocr", "aocr"]:
|
||||
usage = _extract_usage_for_ocr_call(response_obj, response_obj_dict)
|
||||
else:
|
||||
elif not is_unbilled_non_inference_call(call_type, metadata, response_obj_dict):
|
||||
# Use response_obj_dict instead of response_obj to avoid calling .get() on Pydantic models
|
||||
_usage: Final = response_obj_dict.get("usage", None) or {}
|
||||
if isinstance(_usage, litellm.Usage):
|
||||
|
|
|
|||
|
|
@ -2834,13 +2834,19 @@ RoutingDecisionCause = Literal[
|
|||
]
|
||||
|
||||
|
||||
InternalCallOrigin = Literal["autorouter_classifier", "shadow_eval_router", "shadow_eval_judge"]
|
||||
InternalCallOrigin = Literal[
|
||||
"autorouter_classifier",
|
||||
"shadow_eval_router",
|
||||
"shadow_eval_judge",
|
||||
"background_response_cost_poll",
|
||||
]
|
||||
"""Which internal litellm feature originated a billed sub-call, so a spend log row
|
||||
records that it is not traffic the caller sent."""
|
||||
|
||||
AUTOROUTER_CLASSIFIER_CALL_ORIGIN: Final[InternalCallOrigin] = "autorouter_classifier"
|
||||
SHADOW_EVAL_ROUTER_CALL_ORIGIN: Final[InternalCallOrigin] = "shadow_eval_router"
|
||||
SHADOW_EVAL_JUDGE_CALL_ORIGIN: Final[InternalCallOrigin] = "shadow_eval_judge"
|
||||
BACKGROUND_RESPONSE_COST_POLL_CALL_ORIGIN: Final[InternalCallOrigin] = "background_response_cost_poll"
|
||||
|
||||
|
||||
class StandardLoggingRoutingDecision(TypedDict, total=False):
|
||||
|
|
|
|||
|
|
@ -74,6 +74,7 @@ from litellm.constants import (
|
|||
MAX_RETRY_DELAY,
|
||||
MAX_TOKEN_TRIMMING_ATTEMPTS,
|
||||
MINIMUM_PROMPT_CACHE_TOKEN_COUNT_OVERRIDE,
|
||||
NON_INFERENCE_CALL_TYPES,
|
||||
OPENAI_EMBEDDING_PARAMS,
|
||||
TOOL_CHOICE_OBJECT_TOKEN_COUNT,
|
||||
)
|
||||
|
|
@ -1109,6 +1110,8 @@ def function_setup(
|
|||
except Exception as e:
|
||||
verbose_logger.debug("Error extracting messages from Google contents: %s", e)
|
||||
messages = "default-message-value"
|
||||
elif call_type in NON_INFERENCE_CALL_TYPES:
|
||||
messages = [] # mutable-ok: loggers require a list here and Logging copies it
|
||||
else:
|
||||
messages = "default-message-value"
|
||||
stream = False
|
||||
|
|
|
|||
|
|
@ -753,3 +753,41 @@ class TestCheckResponsesCost:
|
|||
call_kwargs = mock_aget.call_args[1]
|
||||
assert "model" not in call_kwargs.get("litellm_metadata", {})
|
||||
assert "model_group" not in call_kwargs.get("litellm_metadata", {})
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_poll_stamps_internal_call_origin_so_the_read_is_billed(
|
||||
self, check_responses_cost_instance, mock_prisma_client
|
||||
):
|
||||
"""A background create returns queued with no usage, so this poll's retrieval is the only
|
||||
place the job's spend is ever seen. Without the origin stamp it is priced at zero like a
|
||||
user-facing read (LIT-5602) and the job is never billed."""
|
||||
from litellm.constants import INTERNAL_CALL_ORIGIN_METADATA_KEY
|
||||
from litellm.litellm_core_utils.internal_call_metadata import (
|
||||
is_unbilled_non_inference_call,
|
||||
)
|
||||
|
||||
mock_job = MagicMock()
|
||||
mock_job.unified_object_id = "resp_test_billed"
|
||||
mock_job.created_by = "test-user"
|
||||
mock_job.id = "job-billed"
|
||||
mock_job.file_object = {"model": "gpt-5", "id": "resp_test_billed"}
|
||||
|
||||
mock_prisma_client.db.litellm_managedobjecttable.find_many = AsyncMock(
|
||||
return_value=[mock_job]
|
||||
)
|
||||
mock_prisma_client.db.litellm_managedobjecttable.update_many = AsyncMock(
|
||||
return_value=0
|
||||
)
|
||||
|
||||
mock_response = MagicMock()
|
||||
mock_response.status = "completed"
|
||||
|
||||
with patch("litellm.aget_responses", new_callable=AsyncMock) as mock_aget:
|
||||
mock_aget.return_value = mock_response
|
||||
await check_responses_cost_instance.check_responses_cost()
|
||||
|
||||
metadata = mock_aget.call_args[1]["litellm_metadata"]
|
||||
foreground_read = {"background": False}
|
||||
assert metadata[INTERNAL_CALL_ORIGIN_METADATA_KEY] == "background_response_cost_poll"
|
||||
assert is_unbilled_non_inference_call("aget_responses", metadata, foreground_read) is False
|
||||
assert is_unbilled_non_inference_call("aget_responses", None, foreground_read) is True
|
||||
|
|
|
|||
|
|
@ -201,6 +201,36 @@ def test_time_to_first_token_is_streaming_only():
|
|||
assert names == set(ALL_METRICS) - {TIME_TO_FIRST_TOKEN}
|
||||
|
||||
|
||||
def test_response_read_does_not_replay_the_generation_usage():
|
||||
"""A responses-management read returns the ORIGINAL generation's usage on the
|
||||
object it fetches. Recording it would add those tokens again on every poll, so
|
||||
the two usage-derived instruments are skipped while the duration ones, which
|
||||
describe the read itself, still fire."""
|
||||
metrics = _drive_success(InMemoryMetricReader(), call_type="aget_responses")
|
||||
|
||||
assert TOKEN_USAGE not in metrics
|
||||
assert TIME_PER_OUTPUT_TOKEN not in metrics
|
||||
assert OPERATION_DURATION in metrics
|
||||
assert RESPONSE_DURATION in metrics
|
||||
|
||||
|
||||
def test_background_response_read_still_records_usage():
|
||||
"""A background=true create returns no usage, so its completed read is the only
|
||||
place the generation's tokens are ever seen. Skipping it would lose them
|
||||
entirely rather than deduplicate them."""
|
||||
reader = InMemoryMetricReader()
|
||||
logger = _logger(reader, enable_metrics=True)
|
||||
kwargs, response_obj, start, end = _build_call(call_type="aget_responses")
|
||||
response_obj["background"] = True
|
||||
asyncio.run(logger.async_log_success_event(kwargs, response_obj, start, end))
|
||||
|
||||
metrics = _metrics_by_name(reader)
|
||||
by_type = {dp.attributes[TOKEN_TYPE]: dp for dp in metrics[TOKEN_USAGE]}
|
||||
assert by_type["input"].sum == PROMPT_TOKENS
|
||||
assert by_type["output"].sum == COMPLETION_TOKENS
|
||||
assert TIME_PER_OUTPUT_TOKEN in metrics
|
||||
|
||||
|
||||
def test_metrics_disabled_records_nothing():
|
||||
"""enable_metrics=False: the recorder is never built, so the injected reader
|
||||
sees no gen_ai.client.* series even though the success hook runs."""
|
||||
|
|
|
|||
|
|
@ -268,6 +268,27 @@ def test_vector_store_file_management_is_not_chat(call_type):
|
|||
assert resolve_operation(call_type).value == "litellm.vector_store_file_management"
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"call_type",
|
||||
[
|
||||
f"{prefix}{operation}"
|
||||
for operation in ("get_responses", "delete_responses", "cancel_responses", "list_input_items")
|
||||
for prefix in ("", "a")
|
||||
],
|
||||
)
|
||||
def test_responses_management_is_not_chat(call_type):
|
||||
"""Fetching, deleting or cancelling a stored response runs no inference, so it must not
|
||||
read as a chat completion: the retrieved object replays the original call's tokens and
|
||||
would inflate the chat series on every read. Regression test for LIT-5602."""
|
||||
assert resolve_operation(call_type) is GenAIOperation.LITELLM_RESPONSES_MANAGEMENT
|
||||
assert resolve_operation(call_type).value == "litellm.responses_management"
|
||||
|
||||
|
||||
def test_creating_a_response_is_still_chat():
|
||||
"""Guards the test above: ``/v1/responses`` itself is a chat completion."""
|
||||
assert resolve_operation("aresponses") is GenAIOperation.CHAT
|
||||
|
||||
|
||||
_NON_CHAT_ROUTES: Final = (
|
||||
("image_generation", GenAIOperation.GENERATE_CONTENT, GenAIOutputType.IMAGE),
|
||||
("speech", GenAIOperation.GENERATE_CONTENT, GenAIOutputType.SPEECH),
|
||||
|
|
|
|||
|
|
@ -6345,3 +6345,95 @@ class TestOpenTelemetryDatabaseSemconvAttributes(unittest.TestCase):
|
|||
span = self._service_span(ServiceTypes.DB, "get_data", None)
|
||||
self.assertEqual(span.attributes["db.system.name"], "postgresql")
|
||||
self.assertNotIn("server.address", span.attributes)
|
||||
|
||||
|
||||
class TestOpenTelemetryNonInferenceUsage(unittest.TestCase):
|
||||
"""Reading a stored response replays the usage of the call that created it, so emitting those
|
||||
token counts again on the read's span reports the same tokens a second time. Regression tests
|
||||
for LIT-5602, covering the legacy emitter that runs by default."""
|
||||
|
||||
USAGE = {"prompt_tokens": 4000, "completion_tokens": 2000, "total_tokens": 6000}
|
||||
TOKEN_KEYS = frozenset({"gen_ai.usage.input_tokens", "gen_ai.usage.output_tokens", "gen_ai.usage.total_tokens"})
|
||||
BACKGROUND_POLL = {"internal_call_origin": "background_response_cost_poll"}
|
||||
RESPONSE_OBJ = {"id": "resp_lit5602", "model": "gpt-4o", "usage": USAGE}
|
||||
BACKGROUND_RESPONSE_OBJ = {**RESPONSE_OBJ, "background": True}
|
||||
|
||||
def _kwargs(self, call_type, litellm_metadata=None):
|
||||
return {
|
||||
"model": "gpt-4o",
|
||||
"call_type": call_type,
|
||||
"optional_params": {},
|
||||
"litellm_params": {
|
||||
"custom_llm_provider": "openai",
|
||||
"litellm_metadata": litellm_metadata or {},
|
||||
},
|
||||
"standard_logging_object": {"id": "lit5602", "call_type": call_type, "metadata": {}},
|
||||
}
|
||||
|
||||
def _token_attributes_on_span(self, call_type, litellm_metadata=None, response_obj=None):
|
||||
otel = OpenTelemetry()
|
||||
mock_span = MagicMock()
|
||||
otel.set_attributes(
|
||||
span=mock_span,
|
||||
kwargs=self._kwargs(call_type, litellm_metadata),
|
||||
response_obj=response_obj or dict(self.RESPONSE_OBJ),
|
||||
)
|
||||
return {call[0][0] for call in mock_span.set_attribute.call_args_list if call[0][0] in self.TOKEN_KEYS}
|
||||
|
||||
def _token_histogram_calls(self, call_type, litellm_metadata=None, response_obj=None):
|
||||
otel = OpenTelemetry()
|
||||
otel._operation_duration_histogram = MagicMock()
|
||||
otel._token_usage_histogram = MagicMock()
|
||||
otel._cost_histogram = None
|
||||
now = datetime.now()
|
||||
otel._record_metrics(
|
||||
self._kwargs(call_type, litellm_metadata), response_obj or dict(self.RESPONSE_OBJ), now, now
|
||||
)
|
||||
return otel._token_usage_histogram.record.call_count
|
||||
|
||||
def _time_per_output_token_calls(self, call_type, litellm_metadata=None, response_obj=None):
|
||||
otel = OpenTelemetry()
|
||||
otel._time_per_output_token_histogram = MagicMock()
|
||||
now = datetime.now()
|
||||
otel._record_time_per_output_token_metric(
|
||||
self._kwargs(call_type, litellm_metadata), response_obj or dict(self.RESPONSE_OBJ), now, 1.0, {}
|
||||
)
|
||||
return otel._time_per_output_token_histogram.record.call_count
|
||||
|
||||
def test_inference_call_still_reports_its_tokens_on_the_span(self):
|
||||
self.assertEqual(self._token_attributes_on_span("acompletion"), set(self.TOKEN_KEYS))
|
||||
|
||||
def test_response_read_does_not_report_the_retrieved_tokens_on_the_span(self):
|
||||
self.assertEqual(self._token_attributes_on_span("aget_responses"), set())
|
||||
|
||||
def test_background_cost_poll_read_still_reports_its_tokens_on_the_span(self):
|
||||
self.assertEqual(self._token_attributes_on_span("aget_responses", self.BACKGROUND_POLL), set(self.TOKEN_KEYS))
|
||||
|
||||
def test_inference_call_still_records_the_token_usage_histogram(self):
|
||||
self.assertEqual(self._token_histogram_calls("acompletion"), 2)
|
||||
|
||||
def test_response_read_does_not_record_the_token_usage_histogram(self):
|
||||
self.assertEqual(self._token_histogram_calls("aget_responses"), 0)
|
||||
|
||||
def test_background_cost_poll_read_still_records_the_token_usage_histogram(self):
|
||||
self.assertEqual(self._token_histogram_calls("aget_responses", self.BACKGROUND_POLL), 2)
|
||||
|
||||
def test_background_response_read_still_reports_its_tokens_on_the_span(self):
|
||||
self.assertEqual(
|
||||
self._token_attributes_on_span("aget_responses", response_obj=self.BACKGROUND_RESPONSE_OBJ),
|
||||
set(self.TOKEN_KEYS),
|
||||
)
|
||||
|
||||
def test_background_response_read_still_records_the_token_usage_histogram(self):
|
||||
self.assertEqual(self._token_histogram_calls("aget_responses", response_obj=self.BACKGROUND_RESPONSE_OBJ), 2)
|
||||
|
||||
def test_inference_call_still_records_time_per_output_token(self):
|
||||
self.assertEqual(self._time_per_output_token_calls("acompletion"), 1)
|
||||
|
||||
def test_response_read_does_not_divide_its_latency_by_the_retrieved_token_count(self):
|
||||
self.assertEqual(self._time_per_output_token_calls("aget_responses"), 0)
|
||||
|
||||
def test_background_response_read_still_records_time_per_output_token(self):
|
||||
self.assertEqual(
|
||||
self._time_per_output_token_calls("aget_responses", response_obj=self.BACKGROUND_RESPONSE_OBJ), 1
|
||||
)
|
||||
|
|
|
|||
|
|
@ -5225,6 +5225,197 @@ async def test_restore_correlation_context_works_across_asyncio_task_boundary():
|
|||
session_id_var.set("")
|
||||
|
||||
|
||||
class TestNonInferenceCallTypesAreNotBilled:
|
||||
"""A retrieved response replays the usage of the call that created it, so pricing a read
|
||||
of it double bills the same tokens. Regression tests for LIT-5602."""
|
||||
|
||||
RETRIEVED_RESPONSE_USAGE = {"input_tokens": 4000, "output_tokens": 2000, "total_tokens": 6000}
|
||||
|
||||
BACKGROUND_POLL_METADATA = {"internal_call_origin": "background_response_cost_poll"}
|
||||
|
||||
def _logging_obj(self, call_type: str, litellm_metadata: dict | None = None):
|
||||
from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj
|
||||
|
||||
obj = LiteLLMLoggingObj(
|
||||
model="gpt-4o",
|
||||
messages=[],
|
||||
stream=False,
|
||||
call_type=call_type,
|
||||
start_time=time.time(),
|
||||
litellm_call_id=f"lit5602-{call_type}",
|
||||
function_id="fn-lit5602",
|
||||
)
|
||||
obj.update_environment_variables(
|
||||
model="gpt-4o",
|
||||
user="",
|
||||
optional_params={},
|
||||
litellm_params={
|
||||
"api_base": "",
|
||||
"custom_llm_provider": "openai",
|
||||
"litellm_metadata": litellm_metadata or {},
|
||||
},
|
||||
)
|
||||
return obj
|
||||
|
||||
def _retrieved_response(self, background: bool | None = None):
|
||||
from litellm.types.llms.openai import ResponsesAPIResponse
|
||||
|
||||
return ResponsesAPIResponse(
|
||||
id="resp_lit5602",
|
||||
created_at=1234567890,
|
||||
model="gpt-4o",
|
||||
output=[],
|
||||
usage=self.RETRIEVED_RESPONSE_USAGE,
|
||||
background=background,
|
||||
)
|
||||
|
||||
def test_creating_a_response_is_still_priced(self):
|
||||
"""Guards the tests below: the same response object must cost money on the create path."""
|
||||
cost = self._logging_obj("aresponses")._response_cost_calculator(result=self._retrieved_response())
|
||||
assert cost is not None and cost > 0
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"call_type",
|
||||
[
|
||||
"aget_responses",
|
||||
"adelete_responses",
|
||||
"acancel_responses",
|
||||
"alist_input_items",
|
||||
"avector_store_delete",
|
||||
"avector_store_file_content",
|
||||
"avector_store_file_delete",
|
||||
],
|
||||
)
|
||||
def test_read_and_management_calls_cost_nothing(self, call_type):
|
||||
cost = self._logging_obj(call_type)._response_cost_calculator(result=self._retrieved_response())
|
||||
assert cost == 0.0
|
||||
|
||||
def test_retrieved_usage_is_not_re_reported_in_standard_logging_payload(self):
|
||||
from litellm.litellm_core_utils.litellm_logging import (
|
||||
get_standard_logging_object_payload,
|
||||
)
|
||||
|
||||
from datetime import datetime
|
||||
|
||||
logging_obj = self._logging_obj("aget_responses")
|
||||
now = datetime.now()
|
||||
payload = get_standard_logging_object_payload(
|
||||
kwargs={
|
||||
"litellm_call_id": "lit5602-payload",
|
||||
"model": "gpt-4o",
|
||||
"call_type": "aget_responses",
|
||||
"litellm_params": {},
|
||||
},
|
||||
init_response_obj=self._retrieved_response(),
|
||||
start_time=now,
|
||||
end_time=now,
|
||||
logging_obj=logging_obj,
|
||||
status="success",
|
||||
)
|
||||
|
||||
assert payload is not None
|
||||
assert payload["prompt_tokens"] == 0
|
||||
assert payload["completion_tokens"] == 0
|
||||
assert payload["total_tokens"] == 0
|
||||
assert payload["response_cost"] == 0.0
|
||||
|
||||
def test_background_cost_poll_read_is_still_priced(self):
|
||||
"""A background create returns queued with no usage, so the poller's read carries the job's
|
||||
only billable usage. Zeroing it there means background jobs are never billed."""
|
||||
cost = self._logging_obj(
|
||||
"aget_responses", litellm_metadata=self.BACKGROUND_POLL_METADATA
|
||||
)._response_cost_calculator(result=self._retrieved_response())
|
||||
assert cost is not None and cost > 0
|
||||
|
||||
def test_background_cost_poll_reports_usage_in_standard_logging_payload(self):
|
||||
from datetime import datetime
|
||||
|
||||
from litellm.litellm_core_utils.litellm_logging import (
|
||||
get_standard_logging_object_payload,
|
||||
)
|
||||
|
||||
now = datetime.now()
|
||||
payload = get_standard_logging_object_payload(
|
||||
kwargs={
|
||||
"litellm_call_id": "lit5602-poll-payload",
|
||||
"model": "gpt-4o",
|
||||
"call_type": "aget_responses",
|
||||
"litellm_params": {"litellm_metadata": self.BACKGROUND_POLL_METADATA},
|
||||
},
|
||||
init_response_obj=self._retrieved_response(),
|
||||
start_time=now,
|
||||
end_time=now,
|
||||
logging_obj=self._logging_obj(
|
||||
"aget_responses", litellm_metadata=self.BACKGROUND_POLL_METADATA
|
||||
),
|
||||
status="success",
|
||||
)
|
||||
|
||||
assert payload is not None
|
||||
assert payload["total_tokens"] == 6000
|
||||
|
||||
def test_reading_a_background_response_is_still_priced(self):
|
||||
"""A background create answers queued with no usage at all, so whoever reads the finished
|
||||
job is the first and only caller to see its tokens. Zeroing that read bills the job nothing."""
|
||||
cost = self._logging_obj("aget_responses")._response_cost_calculator(
|
||||
result=self._retrieved_response(background=True)
|
||||
)
|
||||
assert cost is not None and cost > 0
|
||||
|
||||
def test_reading_a_background_response_reports_usage_in_standard_logging_payload(self):
|
||||
from datetime import datetime
|
||||
|
||||
from litellm.litellm_core_utils.litellm_logging import (
|
||||
get_standard_logging_object_payload,
|
||||
)
|
||||
|
||||
now = datetime.now()
|
||||
payload = get_standard_logging_object_payload(
|
||||
kwargs={
|
||||
"litellm_call_id": "lit5602-background-payload",
|
||||
"model": "gpt-4o",
|
||||
"call_type": "aget_responses",
|
||||
"litellm_params": {},
|
||||
},
|
||||
init_response_obj=self._retrieved_response(background=True),
|
||||
start_time=now,
|
||||
end_time=now,
|
||||
logging_obj=self._logging_obj("aget_responses"),
|
||||
status="success",
|
||||
)
|
||||
|
||||
assert payload is not None
|
||||
assert payload["total_tokens"] == 6000
|
||||
|
||||
def test_reading_a_foreground_response_is_still_free(self):
|
||||
"""Guards the test above against a blanket exemption: an explicit background=false read was
|
||||
already billed by its create and must stay at zero."""
|
||||
cost = self._logging_obj("aget_responses")._response_cost_calculator(
|
||||
result=self._retrieved_response(background=False)
|
||||
)
|
||||
assert cost == 0.0
|
||||
|
||||
def _read_call_messages(self):
|
||||
logging_obj, _ = litellm.utils.function_setup(
|
||||
original_function="aget_responses",
|
||||
rules_obj=litellm.utils.Rules(),
|
||||
start_time=time.time(),
|
||||
**{"litellm_call_id": "lit5602-setup", "response_id": "resp_lit5602"},
|
||||
)
|
||||
return logging_obj.model_call_details["messages"]
|
||||
|
||||
def test_read_calls_do_not_log_a_placeholder_chat_message(self):
|
||||
assert self._read_call_messages() == []
|
||||
|
||||
def test_read_call_messages_survive_a_logger_that_walks_them(self):
|
||||
"""Loggers reach into this value expecting a chat history and branch on it being a list.
|
||||
An empty list reads as no messages; a tuple matches no branch and crashes the success hook,
|
||||
and None is not iterable where other loggers walk it."""
|
||||
from litellm.integrations.lunary import parse_messages
|
||||
|
||||
assert parse_messages(self._read_call_messages()) == []
|
||||
|
||||
|
||||
def _build_success_payload(logging_obj, kwargs):
|
||||
import datetime
|
||||
|
||||
|
|
|
|||
|
|
@ -3274,6 +3274,82 @@ def test_user_traffic_carries_no_internal_call_origin():
|
|||
assert metadata["internal_call_origin"] is None
|
||||
|
||||
|
||||
def _spend_log_for_call_type(
|
||||
call_type: str, internal_call_origin: str | None = None, background: bool | None = None
|
||||
) -> dict:
|
||||
from litellm.types.llms.openai import ResponsesAPIResponse
|
||||
|
||||
return cast(
|
||||
dict,
|
||||
get_logging_payload(
|
||||
kwargs={
|
||||
"model": "gpt-4o",
|
||||
"call_type": call_type,
|
||||
"response_cost": 0.0,
|
||||
"litellm_params": {
|
||||
"metadata": {
|
||||
"user_api_key": "test-key",
|
||||
"internal_call_origin": internal_call_origin,
|
||||
}
|
||||
},
|
||||
},
|
||||
response_obj=ResponsesAPIResponse(
|
||||
id="resp_lit5602",
|
||||
created_at=1234567890,
|
||||
model="gpt-4o",
|
||||
output=[],
|
||||
usage={"input_tokens": 4000, "output_tokens": 2000, "total_tokens": 6000},
|
||||
background=background,
|
||||
),
|
||||
start_time=datetime.datetime.now(timezone.utc),
|
||||
end_time=datetime.datetime.now(timezone.utc),
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def test_spend_log_for_response_retrieval_does_not_replay_the_created_responses_tokens():
|
||||
"""A retrieved response carries the usage of the call that created it, so counting it again
|
||||
bills the same tokens twice. Regression test for LIT-5602."""
|
||||
payload = _spend_log_for_call_type("aget_responses")
|
||||
|
||||
assert payload["prompt_tokens"] == 0
|
||||
assert payload["completion_tokens"] == 0
|
||||
assert payload["total_tokens"] == 0
|
||||
assert payload["spend"] == 0.0
|
||||
|
||||
|
||||
def test_spend_log_for_background_response_cost_poll_counts_tokens():
|
||||
"""The poller's read is where a background job's usage first shows up, so dropping it there
|
||||
leaves the job unbilled forever."""
|
||||
payload = _spend_log_for_call_type("aget_responses", internal_call_origin="background_response_cost_poll")
|
||||
|
||||
assert payload["total_tokens"] == 6000
|
||||
|
||||
|
||||
def test_spend_log_for_background_response_retrieval_counts_tokens():
|
||||
"""A background create answers queued carrying no usage, so its retrieval is the first and only
|
||||
place the job's tokens are ever visible. Zeroing that read bills the whole job nothing on any
|
||||
proxy that is not running the enterprise cost poller."""
|
||||
payload = _spend_log_for_call_type("aget_responses", background=True)
|
||||
|
||||
assert payload["total_tokens"] == 6000
|
||||
|
||||
|
||||
def test_spend_log_for_foreground_response_retrieval_still_counts_nothing():
|
||||
"""Guards the test above against a blanket exemption: an explicit background=false read was
|
||||
already billed by its create and must stay at zero."""
|
||||
payload = _spend_log_for_call_type("aget_responses", background=False)
|
||||
|
||||
assert payload["total_tokens"] == 0
|
||||
|
||||
|
||||
def test_spend_log_for_response_creation_still_counts_tokens():
|
||||
"""Guards the test above: the same response object must still be counted on the create path."""
|
||||
payload = _spend_log_for_call_type("aresponses")
|
||||
|
||||
assert payload["total_tokens"] == 6000
|
||||
|
||||
|
||||
REDACTED_RESPONSE_PLACEHOLDER: Final = {"text": "redacted-by-litellm"}
|
||||
CONSTANT_ID_FROM_HASHED_PLACEHOLDER: Final = "00fcbef15a3b0097e14b0ca016ed30a0"
|
||||
|
||||
|
|
|
|||
|
|
@ -26,6 +26,7 @@ from litellm.proxy.common_request_processing import (
|
|||
_ClientDisconnectedBeforeFirstChunk,
|
||||
_extract_error_from_sse_chunk,
|
||||
_get_cost_breakdown_from_logging_obj,
|
||||
CostBreakdownHeaderValues,
|
||||
_has_attribute_error_in_chain,
|
||||
_is_azure_model_router_request,
|
||||
open_sse_before_first_byte,
|
||||
|
|
@ -5018,6 +5019,169 @@ class TestResponseCostHeaderForTypedDictResponses:
|
|||
assert fastapi_response.headers["x-litellm-response-cost"] == "0.00123"
|
||||
|
||||
|
||||
class TestCostHeadersForCallsPricedAtZero:
|
||||
"""
|
||||
Regression for LIT-5602. Pricing responses reads and vector-store management routes at
|
||||
zero dropped the entire x-litellm-response-cost family off those replies: the header
|
||||
build reads a falsy zero as "this response never recorded a cost" and filters it out,
|
||||
and a call that returns before pricing stores no cost breakdown for the component
|
||||
headers to read. A client parsing the cost off a read got a KeyError where it had
|
||||
previously been handed a number. Those calls now advertise the whole family at zero.
|
||||
"""
|
||||
|
||||
@staticmethod
|
||||
def _responses_read(*, background=False):
|
||||
from litellm.types.llms.openai import ResponsesAPIResponse
|
||||
|
||||
return ResponsesAPIResponse(
|
||||
id="resp_lit5602",
|
||||
created_at=0,
|
||||
model="gpt-4.1-mini",
|
||||
object="response",
|
||||
output=[],
|
||||
status="completed",
|
||||
background=background,
|
||||
usage={"input_tokens": 10, "output_tokens": 5, "total_tokens": 15},
|
||||
)
|
||||
|
||||
@staticmethod
|
||||
def _logging_obj(*, call_type, recovered_cost=0.0):
|
||||
logging_obj = MagicMock()
|
||||
logging_obj.litellm_call_id = "call-lit5602"
|
||||
logging_obj.call_type = call_type
|
||||
logging_obj.litellm_params = {}
|
||||
logging_obj.cost_breakdown = None
|
||||
logging_obj.model_call_details = {"response_cost": recovered_cost}
|
||||
logging_obj._response_cost_calculator = MagicMock(return_value=recovered_cost)
|
||||
logging_obj._enqueue_deferred_logging = None
|
||||
logging_obj._on_deferred_stream_complete = None
|
||||
return logging_obj
|
||||
|
||||
async def _drive(self, *, monkeypatch, response, logging_obj, route_type):
|
||||
import litellm.proxy.common_request_processing as crp
|
||||
from litellm.proxy._types import UserAPIKeyAuth as RealUserAPIKeyAuth
|
||||
|
||||
async def fake_route_request(**kwargs):
|
||||
async def _llm_call():
|
||||
return response
|
||||
|
||||
return _llm_call()
|
||||
|
||||
monkeypatch.setattr(crp, "route_request", fake_route_request)
|
||||
|
||||
async def fake_post_call_success_hook(data, user_api_key_dict, response):
|
||||
return response
|
||||
|
||||
proxy_logging_obj = MagicMock(spec=ProxyLogging)
|
||||
proxy_logging_obj.during_call_hook = AsyncMock(return_value=None)
|
||||
proxy_logging_obj.update_request_status = AsyncMock(return_value=None)
|
||||
proxy_logging_obj.post_call_response_headers_hook = AsyncMock(return_value={})
|
||||
proxy_logging_obj.post_call_success_hook = fake_post_call_success_hook
|
||||
|
||||
fastapi_response = Response()
|
||||
processing_obj = ProxyBaseLLMRequestProcessing(data={"litellm_logging_obj": logging_obj})
|
||||
|
||||
with patch.object(
|
||||
ProxyBaseLLMRequestProcessing, "_has_post_call_guardrails", return_value=False
|
||||
):
|
||||
await processing_obj.base_process_llm_request(
|
||||
request=MagicMock(spec=Request, headers={}),
|
||||
fastapi_response=fastapi_response,
|
||||
user_api_key_dict=RealUserAPIKeyAuth(api_key="sk-test"),
|
||||
route_type=route_type,
|
||||
proxy_logging_obj=proxy_logging_obj,
|
||||
general_settings={},
|
||||
proxy_config=MagicMock(spec=ProxyConfig),
|
||||
select_data_generator=None,
|
||||
llm_router=None,
|
||||
skip_pre_call_logic=True,
|
||||
)
|
||||
return fastapi_response
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_responses_read_emits_the_cost_header_family_at_zero(self, monkeypatch):
|
||||
fastapi_response = await self._drive(
|
||||
monkeypatch=monkeypatch,
|
||||
response=self._responses_read(),
|
||||
logging_obj=self._logging_obj(call_type="aget_responses"),
|
||||
route_type="aget_responses",
|
||||
)
|
||||
|
||||
assert fastapi_response.headers["x-litellm-response-cost"] == "0.0"
|
||||
for component in (
|
||||
"original",
|
||||
"discount-amount",
|
||||
"margin-amount",
|
||||
"margin-percent",
|
||||
"input",
|
||||
"output",
|
||||
"tool-usage",
|
||||
):
|
||||
assert fastapi_response.headers[f"x-litellm-response-cost-{component}"] == "0.0"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_reading_a_background_response_keeps_its_real_cost(self, monkeypatch):
|
||||
fastapi_response = await self._drive(
|
||||
monkeypatch=monkeypatch,
|
||||
response=self._responses_read(background=True),
|
||||
logging_obj=self._logging_obj(call_type="aget_responses", recovered_cost=0.00042),
|
||||
route_type="aget_responses",
|
||||
)
|
||||
|
||||
assert float(fastapi_response.headers["x-litellm-response-cost"]) == pytest.approx(0.00042)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_an_inference_call_without_a_recorded_cost_still_omits_the_header(self, monkeypatch):
|
||||
"""A chat completion has no zero-priced route, so a falsy cost there means the cost was
|
||||
never recorded and the header stays absent rather than advertising a made-up zero."""
|
||||
fastapi_response = await self._drive(
|
||||
monkeypatch=monkeypatch,
|
||||
response=SimpleNamespace(_hidden_params={}),
|
||||
logging_obj=self._logging_obj(call_type="acompletion"),
|
||||
route_type="acompletion",
|
||||
)
|
||||
|
||||
assert "x-litellm-response-cost" not in fastapi_response.headers
|
||||
|
||||
def test_cost_breakdown_reports_zero_components_for_a_call_priced_at_zero(self):
|
||||
breakdown = _get_cost_breakdown_from_logging_obj(
|
||||
litellm_logging_obj=self._logging_obj(call_type="aget_responses")
|
||||
)
|
||||
|
||||
assert breakdown.original_cost == 0.0
|
||||
assert breakdown.input_cost == 0.0
|
||||
assert breakdown.output_cost == 0.0
|
||||
assert breakdown.tool_usage_cost == 0.0
|
||||
|
||||
def test_cost_breakdown_stays_empty_for_an_inference_call(self):
|
||||
breakdown = _get_cost_breakdown_from_logging_obj(
|
||||
litellm_logging_obj=self._logging_obj(call_type="acompletion")
|
||||
)
|
||||
|
||||
assert breakdown == CostBreakdownHeaderValues()
|
||||
|
||||
def test_cost_breakdown_never_zeroes_the_split_under_a_real_total(self):
|
||||
"""Reading a background response prices normally, so a breakdown that has not landed by the
|
||||
time headers are built is reported as absent rather than as a zero split contradicting the
|
||||
real total alongside it."""
|
||||
breakdown = _get_cost_breakdown_from_logging_obj(
|
||||
litellm_logging_obj=self._logging_obj(call_type="aget_responses"),
|
||||
response_cost=1.96e-05,
|
||||
)
|
||||
|
||||
assert breakdown == CostBreakdownHeaderValues()
|
||||
|
||||
def test_cost_breakdown_reports_zero_components_under_a_zero_total(self):
|
||||
breakdown = _get_cost_breakdown_from_logging_obj(
|
||||
litellm_logging_obj=self._logging_obj(call_type="aget_responses"),
|
||||
response_cost=0.0,
|
||||
)
|
||||
|
||||
assert breakdown.original_cost == 0.0
|
||||
assert breakdown.input_cost == 0.0
|
||||
assert breakdown.output_cost == 0.0
|
||||
|
||||
|
||||
class TestPreCallWithFallbacksOnLocalRateLimit:
|
||||
|
||||
@pytest.mark.asyncio
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue