* fix(vertex_ai): translate /v1/responses batch rows through the Responses-to-Chat bridge
Vertex batch uploads treated every non-embeddings JSONL row as a chat
completions body, so a /v1/responses row lost its input and reached GCS
as a blank text part. Route detection now recognizes /v1/responses rows
and bridges them to chat through the same Responses-to-Chat bridge the
real-time path uses. That bridge call moves out of the Bedrock files
transformation into a shared helper both providers call, forwarding the
record's fields as sent, like real time, instead of validating them
against the SDK TypedDicts whose required keys clients omit.
* chore(batches): type the Vertex responses test helper and drop the quoted input cast
* fix(batches): translate developer messages to system on Vertex and Bedrock batch rows like real time
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* fix(logging): pass provider response headers to callbacks on every endpoint
Custom callbacks only received kwargs["response_headers"] for chat
completions. Responses, image generation and edit, speech, and
transcription calls either never recorded the provider's headers or
recorded them in one place and not the other.
Every handler now records the provider's httpx headers on the response's
hidden params as "headers" (raw) and "additional_headers" (processed,
with LiteLLM's own entries winning on a clash), and the logging object
derives model_call_details["response_headers"] from those hidden params
before cost calculation on the non-stream and both streaming success
paths, keeping a handler-set value authoritative. Binary speech responses
expose their hidden params to the standard logging payload, and the sync
OpenAI transcription request always fetches the raw response.
* test(images): point the legacy image and speech fakes at the raw response surface
Image generation now goes through the SDK's raw response so the provider headers can be read, and the speech binary response now carries hidden params. The unit fakes in the image generation, xinference, proxy provider, image edit, Vertex speech, and otel suites still pinned the old call surface and the old "no hidden params" assertion, so they read an uncalled mock or a fake response without headers.
* test(images): drop the rewritten mock comments and the generated edit PNGs
* test(images): move the llm-span test's image fake to the raw response surface
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
The new CircleCI tests pipeline (#42773) runs tests/unit under pytest-cov on
CPython 3.12.2, where coverage traces every line through sys.settrace. The two
tracemalloc peak comparisons in test_vertex_ai_files_streaming.py drive 8000-row
payloads through both pipelines and slow from ~10s to over 3 minutes under that
tracer, so both hit the 90s pytest-timeout on every run.
Mark them no_cover so pytest-cov pauses tracing for just these two. Their
assertions are unchanged and every other test in the file still reports coverage.
a seekable stream read from its current cursor shipped truncated base64
that azure rejected; read from offset 0 to match how httpx rewinds
multipart parts
- measure every file part and every embedded image field the request
carries instead of the caller's image arg, so a transform that drops,
filters, or adds parts and a JSON body that smuggles extra_body
references all bill what the provider actually receives
- read seekable streams from offset 0 and restore the cursor, matching
how httpx serializes the part
- measure input_image* base64 fields on azure image generation the
same as edits
- merge top-level litellm_params pricing kwargs over nested model_info
in _deployment_model_info so custom pricing works on non-router calls
- reject bool and non-positive width/height in the size fallback and
clamp negative reference_pixels to zero
- prices_tokens now checks rates for non-zero values; get_model_info
normalizes absent token rates to 0, which made is not None always
true and zero-billed per-pixel models that return usage
* fix(fireworks_ai): route firerouter short names and bill pass-through legs at the routed model's rates
fireworks_ai/firerouter and fireworks_ai/firerouter/<slug> resolve to
accounts/fireworks/routers/... instead of a models/ path, and the cost
calculator falls back to the routed model's own catalog entry before the
Fireworks size buckets so a Claude leg is no longer priced at $0
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(fireworks_ai): bill routed legs under the routed model's own provider
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(fireworks_ai): require the k suffix when parsing tiered input fields
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(params): carry stream_chunk_size through litellm_params instead of provider params
* test(integration): fence stream_chunk_size out of every provider request body
* test(bedrock): type parametrized stream chunk test params
* test(integration): drop the contracts manifest resurrected by the main merge
* test(bedrock): type the stream_chunk_size test helpers
* test(params): finish AGENTS.md typing pass on stream_chunk_size tests
* test(integration): drop the covers marker from the stream_chunk_size wire test
---------
Co-authored-by: shrey kharbanda <shreshth@berri.ai>
* fix(bedrock): send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse
Bedrock rejects outputConfig.textFormat on Opus 4.7 and 4.8 with
"output_config.format: Extra inputs are not permitted", and the AWS
model cards list structured outputs as not supported for both, so
their cost-map entries no longer claim supports_native_structured_output
and json_schema requests fall back to the json_tool_call tool.
Fixes#27846
* test(bedrock): assert Opus 4.7 and 4.8 inline the schema on Invoke, move the native case to Sonnet 4.6
---------
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
* chore(cost-map): remove models past their deprecation date
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(cost-calc): drop the empty parametrize left behind by the gemini web search removal
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(cost-map): drop merge base block left by conflict resolution
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(cost-calc): drop gemini image cost tests pinned on removed model
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(unit): make bedrock collector and secret scan timing tests deterministic
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(unit): count interpreter calls instead of wall clock in the secret scan scaling test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(unit): profile the secret scan with cProfile, restore the outer profiler and tighten the scaling bound
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: kerry <kerry@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): sign batch S3 requests with s3_access_key_id and s3_secret_access_key
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): keep S3 signer test additions scoped to new cases
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(bedrock): drop e2e suite changes from the S3 signing fix
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(bedrock): build S3 credentials directly from the s3_* pair so ambient AWS_* env never mixes in
Restores the split-identity e2e coverage
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>