When the pump finishes draining while the client is still connected,
billing is deferred to the proxy's post-response hook, which only fires
on a normally completed response. A client disconnect before the relay
consumed the queued tail tore the generator down past that hook, so the
request logged no spend at all. The relay teardown now dispatches the
stored deferred billing whenever it never reached the end-of-stream
sentinel.
Also drops the live pass_through_tests script: that CI job runs against
a fixed config with no Bedrock model or AWS credentials, so it could
only fail there. The scenario is covered by unit tests on the
relay/pump seam.
_image_sources had no test asserting what it extracts. The existing image tests
live on the Bedrock side and all use base64 without a media_type, which is the one
path the fix left unchanged, so both behaviors it does change went unverified: the
url shape reaching the guardrail at all, and base64 arriving as a data URI.
Against the pre-fix extractor the url case sees [] and the media_type case sees
['AAAA'] instead of ['data:image/png;base64,AAAA'].
The remaining three assert behavior the fix deliberately preserves -- bare base64
passed through, a file source yielding nothing, a malformed source dropped rather
than handed on for a consumer to choke on.
Each message carries a text block because a message with no text never reaches the
guardrail, which would make every source shape look equally dropped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- forward unrouted /gigachat/* requests with env credentials like other passthrough providers (the old fallback returned 400 on any request without a routed model, /gigachat/models included)
- fix basedpyright budget breaches across the gigachat provider, common_request_processing, and llm_passthrough_endpoints with real narrowing, no new suppressions
- add regression tests for the fallback target, auth header, and model-less endpoints
_input_item_provenance converted every input prefix, so an n-item request paid
for n+1 full conversions. It now converts each item once, glues consecutive
function_call items (plus their trailing-assistant context) into units so the
transform's tool_call merging is reproduced inside the unit conversion, and
verifies the unit concatenation against one full conversion, bailing to the
full-conversion fallback on any mismatch. Messages from multi-item units are
tainted, which keeps parallel tool calls patchable exactly like the old prefix
pass while unpredicted merges fall back safely.
A guardrail handing back a non-list structured_messages payload (the
HiddenLayer v2 evaluation dict) previously fell through the length-mismatch
fallback and 500ed converting the dict's keys as messages. The write-back is
now skipped for non-list payloads, restoring the previous no-write-back
behavior on the Responses surface.
Also refreshes the compresr texts-mirror docstring, which still claimed the
Responses translation cannot round-trip structured_messages.
Azure's chat completions validator rejects tool parameters carrying a
top-level anyOf/oneOf/allOf for every model family. AzureOpenAIConfig and
the o-series config now flatten them via the shared helper moved to
prompt_templates common_utils. Requests bridged to the Responses API for
gpt-5.4+ with reasoning active keep the union, which that surface accepts
- _write_back_structured_messages now patches only the rewritten rows back
into the original input items, so reasoning items (encrypted_content),
function_call ids, and web_search_call items survive a guardrail rewrite
verbatim; rewrites that cannot be row-mapped fall back to the previous
full conversion
- the responses bridge stream snapshot restores usage hidden in
_hidden_params when stream_options is unset, so converted fake streams
report real input_tokens instead of 0
The typed rewrite made the properties recursion call .get on every child,
so a malformed schema with a string or list property value raised
AttributeError where it previously passed through untouched.
A deployment carrying reasoning_effort in its litellm_params on the
/v1/messages passthrough mapped the effort to a legacy thinking block
whose budget_tokens was forwarded as is, so any request whose max_tokens
sat at or below that budget was rejected upstream with a 400. The mapped
budget now runs through the same cap the adaptive-to-legacy branch and
the chat path already use: it is clamped to max_tokens - 1, and dropped
with a warning when even the minimum budget cannot fit.
The cap helper becomes public since three call sites outside
AnthropicConfig use it.
GPT-5 and later accept a top-level anyOf natively and call tools better with it intact, so the flattening now runs only for the gpt-4, gpt-3.5, chatgpt-4o, o1, o3, and o4 families. Non-dict tool entries pass through untouched, a typeless root that carries properties counts as an object, and the bounded $ref walker is listed in the recursion detector allowlist.
Message-rewriting guardrails such as Headroom return their rewrite in
structured_messages and leave texts untouched. The responses guardrail
translation only mapped texts back, so compression never reached the
upstream request on /v1/responses while the retrieve tool still got
injected. Convert the returned messages back to Responses input (plus
instructions) the way the chat and Anthropic handlers already do, and
keep developer messages as input_text in the chat-to-responses bridge.
Resolves LIT-6494
The supported_endpoints passthrough had no way to keep ttl for an upstream
that honors it, so the deployment now opts in with
model_info.cache_control_ttl: true, injected into the config the same way
the providers.json constraint is for JSON providers
Rewrite the sanitizer without recursion (the code-quality gate rejects new
recursive functions) and only touch cache_control where the Messages API
defines it: the request, system blocks, tools, message content blocks, and
tool_result content. Application data such as tool_use.input and tool
input_schema is left untouched even when it contains a cache_control key
Reprice ten more retired xAI slugs (grok-3 and grok-3-mini families,
grok-4-1-fast) to the grok-4.3 rates they now bill at, with family-correct
deprecation dates. Restore cache_read_input_token_cost on the Bedrock Grok 4.6
entries so implicit cache hits bill at the cache-read rate while explicit
cachePoint stays unsupported. Drop the unsourced 1080p video rate and the
gemini/ live native-audio entry the Gemini API 404s on. Add Groq qwen3.8-27b
tool-use flags per Groq docs. Extend the xai and gemini tests to lock all of
this in