* fix(caching): count tool_call cache_control marks in the injection census
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): remove cache census casts
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): only skip injection on message or content marks
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(caching): skip injection on messages whose tool calls carry marks
Reverts 9f08d8aef8. A default 5m mark injected on assistant text lands
before the client's 1h tool_use mark, which Anthropic rejects with a 400
because a 1h breakpoint must not follow a 5m one. Keeping the full census
in the skip check leaves the client's tool_call breakpoint as the only
one on that message.
* fix(caching): count every client tool_call cache mark in the breakpoint census
The census gated tool_call marks on type function and dict shape, so a client mark on a call without a type or with a string cache_control slipped past the count and injection overflowed the 4 breakpoint cap. Count any non-None tool_call mark, keep the server tool exclusion, and add integration cells for the capped surfaces, the yaml stand-down, Bedrock and Gemini, and router affinity
* test(integration): hold the upstream so the worker kill lands mid-burst
* refactor(caching): reuse the transform's server tool lookup in the breakpoint census
The census now calls the same helper the Anthropic transform uses to decide
whether a marked tool call becomes a server tool block, so the two cannot
drift apart. The owned-proxy burst test waits up to 90 seconds for the burst
to reach the wire before it kills a worker
* refactor(anthropic): move the server tool rebuild check under llms/anthropic
* test(integration): audit the tool call mark census across chat, messages, responses, and chaos
Adds the /audit cells for the breakpoint census on assistant tool_calls marks: the
Responses stream bridge, the OpenAI and Anthropic SDK clients, in-process Pydantic
messages, response cache twins, request-level points, a provider 401 on a capped request,
malformed provider_specific_fields and tool_call ids, null or empty points, a points
update mid-burst, a proxy restart mid-burst, and a worker SIGKILL that picks the worker
holding the burst's upstream connections
* test(integration): close the SDK clients the cache census cells open
---------
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>