* fix(langfuse): log token usage for /v1/responses calls
ResponseAPIUsage exposes input_tokens/output_tokens, so the Langfuse
generation logger read prompt_tokens/completion_tokens as 0 while cost
was still correct. Normalize the usage with
ResponseAPILoggingUtils._transform_response_api_usage_to_chat_usage
before extracting usage_details.
Resolves LIT-8238
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): adapt responses-api usage regression test to langfuse v4 spans
The v4 migration exports usage_details as an OTEL span attribute via an
in-memory exporter instead of a mocked generation() call. Assert on the
exported span's langfuse.observation.usage_details with the same token
counts.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(langfuse): log token usage for streamed /v1/responses calls
An assembled /v1/responses stream reaches the logger with its usage stored
as a plain dict of chat fields, which the attribute reads saw as zero. Route
both ResponseAPIUsage and that dict through one normalizer.
* fix(langfuse): keep generations whose usage dict fails validation; cover streaming in the e2e
A usage dict whose token counts are not integers falls back to the raw
object, so the generation is still logged with zero usage instead of being
dropped. The live e2e now runs streamed and non-streamed /v1/responses.
* ci: re-run the e2e gate
* test(langfuse): run the real prompt helper and label the responses usage e2e cases
Drop the _add_prompt_to_generation_params pass-through patch from the
Responses usage unit tests so the real helper runs with the empty metadata
they already pass. Declare a Subject on both e2e parametrizations
(observability, /v1/responses, openai, stream and nonstream).
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: yucheng-berri <293926326+yucheng-berri@users.noreply.github.com>