Commit graph

14 commits

Author SHA1 Message Date
ryan-crabbe-berri
2aac6e109d refactor(responses): cut the cost poller's docstrings back to the non-obvious why
The poller narrated its own straightforward behavior in seven multi-paragraph
docstrings, which the repo's comment policy rules out. Each is now the claim a
reader needs to avoid a wrong edit and nothing more. Also drops the deprecated
`Dict` and `Optional` aliases the file still used.

Claude-Session: https://claude.ai/code/session_01RHAjRxNhXTpKHeGMZ1nDKi
2026-09-15 00:37:54 +00:00
ryan-crabbe-berri
5aa1901249 test(responses): drive the poller tests through the injected router
The eight claim/release tests reached into `litellm.aget_responses` with
`patch`, which the test-quality gate flags and which couples them to an
import path rather than to the poller's own seam. Each job now carries a
LiteLLM-encoded provider id, so the read routes through the router the
fixture already injects and the assertions run against that mock.

Claude-Session: https://claude.ai/code/session_01RHAjRxNhXTpKHeGMZ1nDKi
2026-09-15 00:37:54 +00:00
ryan-crabbe-berri
492336a50b refactor(responses): give the background row store an injectable seam
The store lived inline in responses_api, so covering it meant patching five
proxy_server globals per test, which the test-quality gate counts as pinning
the test to the wiring rather than the behaviour. It is now a module-level
function that takes the managed-files hook as an argument, and the tests hand
it a fake directly.

The poller tests reach the same fetch through the router fixture that is
already injected, instead of patching litellm.aget_responses.

A response with no model_id now returns early rather than raising into the
caller's except block. Nothing is stored either way and the warning is
unchanged, so only the redundant second log line goes away.

Claude-Session: https://claude.ai/code/session_01RHAjRxNhXTpKHeGMZ1nDKi
2026-09-15 00:37:54 +00:00
ryan-crabbe-berri
49a8ce5e42 fix(responses): key a background response's managed row by the provider id
model_object_id is documented as "the id returned by the backend API
provider", and that is what batches and fine-tuning jobs store there. The
background responses create stored the advertised id in it instead, which
is encrypted with a fresh nonce on every call, so the row had no stable
handle on the generation it describes.

The cost poller now reads the provider id straight off the row. Rows
written before this still carry the advertised id there, and decrypting is
a no-op on an id that is already the provider's, so both shapes resolve
through the same call.

Claude-Session: https://claude.ai/code/session_01RHAjRxNhXTpKHeGMZ1nDKi
2026-09-15 00:37:54 +00:00
ryan-crabbe-berri
ffa52f2f6a refactor(responses): mark a billed row completed without copying the response onto it
The previous commit stored the finished ResponsesAPIResponse in file_object. That
duplicates content the provider still serves from its own copy, and the usage and
spend it was meant to preserve already land in LiteLLM_SpendLogs on every billed
call regardless of store_prompts_in_spend_logs, which gates only the messages and
response body columns.

The poller now writes status alone, as it did before. The write stays per job
rather than one bulk update so a single failure cannot strand the rest of the
cycle.

Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
2026-09-15 00:37:54 +00:00
ryan-crabbe-berri
0235dbd7f2 fix(responses): claim a background response before the read that bills it
Every pod and uvicorn worker schedules its own CheckResponsesCost against the
shared LiteLLM_ManagedObjectTable. The poller selected eligible rows, performed
the billed retrieval, and only then marked them completed in one bulk write, so
two pollers could select the same terminal response and both record a charge
before either completion update landed.

Each row is now claimed with a compare-and-swap on batch_processed before the
read, because the read is what prices the job: aget_responses stamped with the
poll origin writes the spend log itself, so there is no later point at which to
serialize. A row whose read raised, or whose provider status is still
non-terminal, releases its claim so a later cycle retries it rather than
retiring it unbilled. That is the failure #37050 fixed on the batch side.

A pod that dies between winning the claim and billing would otherwise strand the
row: it holds a claim nobody will release and its status never reaches terminal,
so every later cycle re-selects it and loses. The updated_at arm of the claim
takes such a row back after three poll cycles, and since updated_at is @updatedAt
a healthy in-flight claim written moments ago is never stolen.

The poller now also persists the finished response onto its managed row instead
of writing status alone, so the row carries the generation's usage rather than
the stale queued copy stored at create time.

Reuses the existing batch_processed column, so no migration. It already sits on
the shared table defaulted to false and was unused by response rows.

Claude-Session: https://claude.ai/code/session_01Hn5E8Jz1LjGLFyiYxBRcBW
2026-09-15 00:37:54 +00:00
jesus
bb81a9f9f1 fix(responses): bill a background response once via the cost poller, not on every read
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-15 00:37:54 +00:00
devin-ai-integration[bot]
4bf40c4e8d
fix(logging): stop billing and logging response reads as LLM calls (#36890)
* fix(logging): stop billing and logging response reads as LLM calls

Retrieving, deleting or cancelling a stored response, and vector store management calls, run through the same logging lifecycle as inference. A retrieved response replays the usage of the call that created it, so every read priced it again and wrote a second spend log row for the same tokens. Non-inference calls now cost 0, report no usage, log no placeholder chat message, and get a litellm.responses_management operation name instead of reading as chat.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep billing background response jobs after the poll

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(logging): use an empty list for read-call messages

A tuple matches no branch in the loggers that walk this value, so lunary's
parse_messages falls through to clean_message and raises AttributeError on the
success hook. An empty list reads as no messages everywhere: it satisfies the
isinstance(list) checks in newrelic, mlflow and datadog, iterates zero times in
traceloop and helicone, and is what StandardLoggingPayload.messages is typed to
hold. None would be type-legal too but is not iterable, so it trades one crash
for another in mlflow and traceloop.

* fix(otel): stop the legacy emitter reporting replayed tokens on response reads

The zeroing so far lands in the standard logging payload, which the legacy
OpenTelemetry emitter does not read for usage: it takes prompt, completion and
total tokens straight off the response object, so a retrieval span still carried
the token counts of the call that produced the response, and the token usage
histogram still recorded them. That emitter is the default, so the spend row said
zero while the trace said otherwise. The background cost poller keeps its counts,
the same exemption the pricing path already makes.

* fix(logging): keep billing a background response when its retrieval is read

A response created with background=true comes back queued and carries no usage, so
its create bills nothing. The retrieval that first sees the finished job is the only
place that job's tokens are ever visible, and pricing every read at zero therefore
loses the spend outright rather than deduplicating it. On a proxy without the
enterprise cost poller a background job ended up costing $0 end to end.

is_unbilled_non_inference_call now takes the response it is deciding about and treats
a background response the same way it already treats the poller's own read, which is
the same exemption seen from the other side. The legacy OpenTelemetry emitter's time
per output token metric picks up the read gate it was missing, so it stops dividing a
read's latency by the replayed completion token count.

* test(proxy): pass the read response to the non-inference predicate

The poller test called is_unbilled_non_inference_call with the pre-background signature, so it broke when the predicate gained the response it classifies. It now hands the predicate a foreground read, and asserts that the same read is free without the origin stamp, so the stamp is what the test proves.

* fix(otel): stop the v2 metrics recorder reporting replayed tokens on response reads

The v2 span builder sources usage from the standard logging payload, so the
earlier fix already zeroes it there. The metrics recorder reads response_obj
directly, so a responses-management read still recorded the original
generation's tokens into gen_ai.client.token.usage and divided generation time
by them for gen_ai.server.time_per_output_token.

The read still records operation and response duration, under the
litellm.responses_management operation, so it stays observable.

* fix(proxy): keep the response-cost headers on calls priced at zero

Pricing responses reads and vector-store management routes at zero dropped the whole
x-litellm-response-cost family off those replies. The header build reads a falsy zero as
a cost this response never recorded and filters it out, and a call that returns before
pricing stores no cost breakdown for the component headers to read, so a client parsing
the cost off a read got a KeyError where it had previously been handed a number.

Those calls now advertise the family at zero. Retrieving a background response, and the
cost poller's read of one, still report their real cost.

The params-taking form of the predicate moves from opentelemetry into
internal_call_metadata so the proxy header build and the OTEL recorders share one copy.

* fix(proxy): report a zero cost split only under a zero cost total

The component headers were filled from call-type membership alone, while the
total they sit beside keeps its real value when the read priced normally, so a
breakdown that had not landed by the time headers were built could advertise a
real total next to an all-zero split. The split is now reported as zero only
when the total agrees with it, and is otherwise left absent.

---------

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
2026-08-26 18:34:17 -07:00
mateo-berri
4f7d1fce3a fix(proxy): fall back to the SDK when a queued response's deployment is missing 2026-08-05 23:13:15 -07:00
Devin AI
4429742e83 fix(proxy): fetch background responses through the router in CheckResponsesCost
Closes #35131
2026-07-29 21:11:44 +00:00
Ishaan Jaffer
e8461b5b97
style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
Yuneng Jiang
48a68230c8
fix(test): update check_responses_cost tests for _expire_stale_rows
PR #25258 changed _cleanup_stale_managed_objects from update_many to
execute_raw via _expire_stale_rows, but the tests were not updated.
The tests now mock _expire_stale_rows on the instance and assert
update_many calls only for job completion, not stale cleanup.
2026-04-07 10:09:11 -07:00
Ishaan Jaff
1b96064600
fix(proxy): prevent OOM/Prisma connection loss from unbounded managed-object poll (#23472)
* fix(proxy): cap managed-object poll size + expire stale rows + kill-switch flag to prevent OOM/Prisma connection loss

* fix(constants): simplify PROXY_BATCH_POLLING_ENABLED readability

* docs+test: document new polling env vars, add pagination+stale-cleanup tests

* fix: exclude stale_expired from batch poll queries; fix update_many assertions in tests

* fix: scope stale cleanup to file_purpose, fix file_object mocks, add CheckBatchCost tests

* fix: avoid duplicate cost logging in fallback path; guard integer constants against zero/negative values

* fix: cache _has_batch_processed_column; guard cleanup from aborting poll; narrow fallback except

* fix: add complete/completed to primary query not_in; fix vacuous test assertion

- Primary find_many was missing "complete" and "completed" in its not_in
  filter, creating asymmetry with the fallback query. A job whose status
  was set to "complete" but whose batch_processed flag update failed would
  be silently re-fetched and re-processed every cycle, emitting duplicate
  cost logs.

- test_fallback_completion_update_omits_batch_processed patched
  _is_base64_encoded_unified_file_id to return None, causing an immediate
  continue — so update() was never called and the assertion looped over an
  empty list (vacuously true). Rewrote the test to mock the full
  completion pipeline, verify update() is called exactly once, and assert
  batch_processed is absent from the update data.

- Added symmetric test (primary path) proving batch_processed IS included
  when the column exists.

Made-with: Cursor
2026-03-13 11:01:40 -07:00
Sameer Kankute
7d0f41f437 Add cost tracking for responses api in background mode 2025-12-19 13:35:48 +05:30