* fix(snowflake): normalize Cortex Claude request shapes
Co-authored-by: Kamron Javaherpour <kamron@kargo.com>
Co-authored-by: Oleksandr Kononov <oleks.konov@kargo.com>
* style(snowflake): format Cortex request transformations
* fix(snowflake): annotate Cortex wire payloads
* fix(snowflake): route Cortex content through the shared Anthropic converters
* fix(snowflake): surface Cortex prompt-cache usage and thinking blocks
Parse Cortex's Anthropic-dialect responses and SSE with Anthropic's own parser so cache_creation/cache_read counts, thinking blocks and signatures reach the caller. Restore thinking for every Claude model: Cortex documents extended thinking broadly and only adaptive thinking is 4.6-gated.
* fix(snowflake): echo signed thinking blocks on every assistant turn
The reference converter extends signed thinking blocks on each assistant turn, not just tool-call turns, so a replayed thinking-plus-text response keeps its signed block. Content-less thinking turns send no empty text block.
* fix(snowflake): preserve thinking list content
---------
Co-authored-by: Oleksandr Kononov <oleks.konov@kargo.com>
Keeps staging's response-id keying next to the pass-through failure-path helpers
and types the read-only raw_bytes parameters as Sequence[bytes] so the merged tree
stays inside the lint budgets
The block accepts output_cost_per_reasoning_token and cache_creation_input_token_cost. The generic
cost path and the DashScope calculator swap them in while a window is open, and unset keys keep the
standard rate. One shared TokenRates value replaces the DashScope-local copy, and
apply_off_peak_pricing takes and returns it.
A provider read timeout after the 200 was already committed on a streamed
/v1/messages request used to run the success logging path, so the failure
callbacks never fired and the failure metrics stayed flat. The pass-through
stream handler and the Bedrock relay iterator now dispatch the failure
handlers instead, with the usage and cost of the chunks already delivered
stashed on the logging object so the failure row still bills them.
Streaming /v1/messages against a model served through the chat-completions
bridge (every non-Anthropic provider other than OpenAI) minted its msg_ id
inside the stream wrapper, so the spend row landed under the provider's own
completion id and the caller could not find the call by the only id it saw.
The wrapper now mints the id once in its constructor and hands it to the
logging object, the same way the Responses-API bridge does.
The documented DISABLE_AIOHTTP_TRANSPORT env var already selects the httpx transport, so the extra module-global write was redundant. Types the monkeypatch fixture while here.
A streaming /v1/messages call against a non-Anthropic model is served an SSE
message_start frame carrying a msg_ id the adapter mints locally, since the
Responses API upstream only issues a resp_ id. That value never left the
adapter, so the spend row was keyed on the bridged response id and
GET /spend/logs?request_id=msg_... came back empty.
The adapter now hands the id it minted to the logging object, and the
/v1/messages logging path keys the row on it.
AmazonMoonshotConfig.transform_request called
_get_boto_credentials_from_optional_params purely for its side effect of
popping the aws_* keys off optional_params, then threw the result away. On
a box whose default AWS profile uses login_session without botocore[crt],
that call raises, so a bearer-token bedrock/invoke/moonshot.* deployment
still 500s with MissingDependencyException even after the rest of this
branch skips the chain.
It now filters the aws_* keys into a local dict the way the Qwen, OpenAI
and Claude 3 invoke transformations already do, so no credentials are
resolved and the caller's optional_params keeps the keys sign_request
reads afterwards.
Resolves two conflicts:
- tests/test_litellm/vector_stores/test_main.py: staging moved search() to a
RouterVectorStoreEmbeddingExecutor while this branch parametrized the same
test over query; keep both the executor assertions and the parametrize.
- tests/logging_callback_tests/test_bedrock_knowledgebase_hook.py: staging
carries a duplicate embedding_executor kwarg that makes the file a
SyntaxError; drop the trailing duplicate.