_get_or_start_block trusted the item_id -> block index map without checking
whether that block was still open, so a provider that reuses one item id for
a whole run and interleaves channels got a delta addressed to a stopped
block. Replaying reasoning, text, reasoning, text produced
content_block_delta index=0 after content_block_stop index=0, which is not a
valid Anthropic stream.
Treat the mapping as valid only while it points at the open block, and
rebind the item to a fresh block otherwise. Items registered through
response.output_item.added still reuse their block, since that block is the
open one while its deltas arrive.
Apodex Deep Research is what surfaced this: it labels every reasoning delta
of a run rs_<response_id> and every answer delta msg_<response_id>, so any
interleaving hits the stale mapping.
A Deep Research run streams two agents. The worker emits its chain of
thought on the `reasoning` channel and a draft answer on a channel-less
delta; the reporter emits its own reasoning plus the single `output_text`
delta that matches the final response.completed snapshot.
Only `output_text` was mapped, so 176 of 181 deltas in a sample run
surfaced as GenericEvent and the reasoning was effectively lost. Map the
`reasoning` channel to response.reasoning_summary_text.delta, which
LiteLLM already translates into an Anthropic thinking_delta on the
/v1/messages route Deep Research takes, and give it its own item id.
The channel-less deltas stay unclaimed on purpose: splicing the worker's
draft into the answer would corrupt the text. The remaining
response.swarm.* lifecycle events keep passing through, since
transform_streaming_response has no way to drop a chunk and run_finished
carries the final content.
A live stream against apodex-1-1-deep-research shows the event is real and
load-bearing: 181 of the 193 events are response.swarm.llm_delta and
response.output_text.delta never appears, so without the mapping the answer
text only arrives in the final response.completed snapshot.
Restores the transform with the provenance recorded in a docstring, and
covers it with the payload shape captured off the wire, including the
reasoning channel that carries 176 of those deltas and must not be mistaken
for the answer.
Cross-checked the provider against platform.apodex.ai/docs and a live
GET /v1/models call.
- apodex-1.1 and apodex-1.1-mini advertised 256K max output; /v1/models
reports 65536, and max_tokens is the legacy alias of max_output_tokens
- apodex-1-1-deep-discover is Responses-API-only; /v1/chat/completions
answers 400 unsupported_api for the Discover tiers
- core models do not support response_format, so state it explicitly
- transform_cancel_response_api_response carried Content-Encoding over to
a response whose body it had already replaced, so httpx tried to
decompress plain JSON on read. A non-JSON body (the gateway answers a
timed-out cancel with an HTML 504) also escaped as a pydantic
ValidationError instead of the provider error
- drop the undocumented response.swarm.llm_delta mapping
- only the Deep Research tiers default stream to true; the core models
follow OpenAI and default it to false
The JSON provider path applies one contract to a whole provider, which is wrong
for Apodex: its two model families take different parameters. Replaces the
providers.json entry with litellm/llms/apodex/, reverting the shared
openai_like machinery to its original state.
/v1/responses is now keyed off the model. Core models are a stateless subset,
so store is pinned false and previous_response_id / background are rejected
rather than passed upstream to fail with a 400. The deep research tiers keep
all three, so background survives a client disconnect.
/v1/messages resolves per model too. Apodex serves the protocol natively for
the core models only, so the deep research tiers get no native config and fall
back to translation instead of hitting a path that does not serve them.
Chat completions pin stream to false for both families, drop tool params on the
deep research tiers, and rename max_completion_tokens to max_tokens. The
responses config also stops inheriting OpenAI's OPENAI_API_KEY fallback, which
would otherwise forward an unrelated OpenAI key to Apodex.
Tests live under tests/test_litellm/llms/apodex/ and touch no existing test file.
Registers apodex via providers.json with /v1/chat/completions, /v1/responses
and native /v1/messages, plus price map entries for the two core models and
the six deep research tiers.
Apodex defaults `stream` to true on both /v1/chat/completions and /v1/responses,
so a non-streaming litellm call would get SSE back and fail to parse it. Adds a
`send_explicit_stream_false` special-handling flag that pins the field on the
wire, and rewrites the JSON provider param mapping to build its result instead
of mutating the caller's dict.
* fix(panw_prisma_airs): scan tool call args as plain text, not a tool_event
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(panw_prisma_airs): type the tool call argument extractor
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(panw_prisma_airs): cover tool call error fallback and dict masking paths
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): scan tool names with args and tolerate custom tool calls
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): scan tool call arguments that arrive already parsed
The tool call slice types arguments as a string, so a client posting parsed JSON
failed validation and the whole tool call, name included, read as unscannable and
was skipped without ever reaching AIRS. The OpenAI request path forwards
client-supplied tool_calls verbatim, so that shape is reachable.
Coerce non-string arguments instead of rejecting them, so the content is scanned.
* fix(panw_prisma_airs): route tool-block masked data by scan side, not by key name
Merging #37036 (already on staging) with this PR produces no conflict and a
silent bug. #37036 withholds prompt_masked_data on response-side tool blocks,
which was right while tool calls went out as a request-side tool_event: AIRS
reported the model's arguments under that key. This PR scans tool calls as
ordinary prompt/response text, so the side of the scan now decides which key
holds what. The model's arguments arrive under response_masked_data, already
covered by _CLIENT_HIDDEN_SCAN_FIELDS, and prompt_masked_data goes back to
being the caller's own input -- one of the audit fields LIT-5638 asks for.
Left as merged, a response-side tool block drops that field with nothing to
flag it.
- Tool-path block branch calls _build_error_detail without also_hide
- also_hide parameter removed; after this change it has no callers
- Regression test asserts both directions: model output withheld, caller
input preserved. It fails against the auto-merged combination.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(panw_prisma_airs): a wrong-typed tool name must not suppress the scan
_ToolCallFunctionSlice types name as str, and _get_tool_call_function turns any
ValidationError into (None, None), which _scan_tool_calls_for_guardrail reads as
an unscannable tool call and skips. So a client posting "name": 123 keeps its
arguments off the wire to AIRS entirely -- no error, no log, no block. The
OpenAI request path forwards client tool_calls verbatim, so this is reachable by
any caller holding a valid key.
_coerce_arguments already existed for exactly this failure mode on the sibling
field. Widening it to cover name closes the gap:
name='transfer_funds' AIRS called: 1x args scanned: True
name=123 (int) AIRS called: 0x args scanned: False <- before
name=123 (int) AIRS called: 1x args scanned: True <- after
Reported by Cursor Bugbot on fd9f6396e5.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This cell needs the websearch_interception callback and a declared search
backend, both listed in its own module docstring. The ephemeral e2e stack
ships neither, so the request falls through to the bedrock transformation
and takes the by-design 400 that tells you to enable interception.
The cell has never been green here: the error path merged about an hour and
a half before the cell did, and the last full suite to pass predates the
cell entirely. Skip it with the reason recorded so the run reports honestly
instead of carrying a permanent red, and unskip once the stack ships the
config the docstring already spells out.
Both providers reworded the error strings these two cells pinned, so the
suite went red without any behavior changing. Anthropic's auth error is now
"API key is invalid." rather than "invalid x-api-key", and OpenAI rejects an
empty upload with "This model does not support the format you provided.",
which names neither "file" nor "audio".
Assert the durable shape instead. The otel cell pins the machine-readable
authentication_error type plus a non-empty message, and the embedded JSON
still has to parse, which is what proves the attribute survived untruncated.
The transcription cell pins that the 400 relays the provider's own rejection
and is typed as a client input error, so a regression that swallows the
provider reason or returns a 500 still fails.
The google ai studio responses test still asserted tools == [], but the
transformation now pops empty tools and tool_choice before calling
completion, so assert the keys are absent.
test_openai_endpoints pinned claude-3-sonnet-20240229, which Bedrock has
retired; move it to us.anthropic.claude-sonnet-4-5-20250929-v1:0.
The AssemblyAI EU passthrough test depended on a credential that no longer
resolves in CI, and the US path plus the bad-key case already cover the
route; drop it rather than keep a permanently red case.
The playground, logs drawer and AI Hub modal moved off antd, so the specs
that reached for .ant-select, .ant-drawer-content, .ant-modal and
.ant-radio-button-wrapper no longer match anything and time out.
Address the same controls through their accessible role and name instead,
which holds across the component library swap and reads closer to what a
user does.
* fix(guardrails): return the full PANW AIRS scan response on blocked requests
The blocked-request error detail was assembled from a hardcoded allowlist, so audit fields like prompt_detection_details, prompt_masked_data, source, transaction_id and session_id never reached the client even though AIRS returned them.
Resolves LIT-5638
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(guardrails): drop redundant comment in AIRS error detail
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(panw_prisma_airs): withhold response_masked_data from the blocked-response error
The full AIRS passthrough also reached the response-side block path, where
response_masked_data carries the model's own generation. That branch is only
reached when mask_response_content is False, so the operator had explicitly
declined to deliver that text, and the error body handed it back anyway.
Withhold response_masked_data from the client-visible detail. prompt_masked_data
stays: it is the caller's own input and one of the fields the ticket asks for.
Every other AIRS field, including prompt_detection_details, source,
transaction_id and session_id, is unchanged.
* fix(panw_prisma_airs): withhold generated tool args from response-side blocks
_scan_tool_calls_for_guardrail calls AIRS with is_response=False because
tool_event is request-side in the AIRS schema, so AIRS returns the scanned
tool arguments under prompt_masked_data. When the tool calls being scanned
are the model's own output, that key holds generated content, and the
_CLIENT_HIDDEN_SCAN_FIELDS default (response_masked_data, empty on this
path) does not cover it. With the default mask_response_content=False the
block branch then shipped the model's masked tool arguments in the 400 --
the same content channel this PR closed for response_masked_data.
_build_error_detail takes an extra_hidden_fields argument so the withholding
stays in one place, and the tool-call block branch passes prompt_masked_data
when is_response is True. Request-side blocks are unchanged and still carry
prompt_masked_data, which is what LIT-5638 asks for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* style(panw_prisma_airs): apply ruff format
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Yucheng Zhu <yucheng@berri.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Probe the column before the scheduler registers CheckBatchCost, closing the window where a retrieve that decided the poller was inactive billed a batch the first poll cycle then billed again. Also drop narration docstrings and section banners from the new tests.
The retrieve path now defers a managed batch's accounting to CheckBatchCost, which
bills the key, team, and tags stored on the managed object row. The /v1/batches
create hook never persisted api_key or request_tags there (only the passthrough
creates did), so the poller attributed the cost to the user alone and the creating
key's spend stayed at zero.
* fix(ptu): stop a PTU deployment billing for grounded search
A PTU deployment is billed by the flat cost of its reserved capacity, so the
model write endpoints refuse a rate the caller supplies and zero the ones already
stored. search_context_cost_per_query escaped both: it holds its rates in a table
keyed by context size, and the guard only recognised a number as a price, so a
grounded request on a PTU deployment kept billing per search on top of the flat
cost.
A table now counts as a price when it holds a non-zero rate. It is zeroed in
place rather than emptied the way tiered_pricing is, because an absent table
means the provider's own default rate rather than free, so dropping it would
start a charge instead of stopping one. For the same reason an all-zero table is
not read as a price: it is how an operator expresses free.
* fix(ptu): zero the search rate on every PTU deployment
A deployment that never stored its own search table is the normal case, and an
absent table means the provider's default rate, so the zeroing has to be written
unconditionally the way the per-token zeros already are. Writing it only where a
table was already stored left the default path billing per grounded search, which
is the charge this set out to stop.
The predicate that reads a table is split out rather than recursing, since the
repo's recursion gate rejects an unignored recursive function and one level is
all a rate table needs.