* fix(guardrails): run end-of-stream post_call scan when the client disconnects mid-stream
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): close the guardrail stream chain in async_data_generator on client disconnect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): leave the raw upstream response to the shielded finalizer on client disconnect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): keep disconnect cleanup going when a streaming callback cleanup raises
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(proxy): assert the refund through a recorder instead of the mock
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): inspect tool calls released before a disconnect under incremental_diff and record a failed scan marker
The incremental_diff transform stream now scans tool calls it already released when the client disconnects, and a disconnect scan whose translation raises after the guardrail recorded success also records guardrail_failed_to_respond
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): pin that a text-only disconnect scan is not handed a tool_calls finish
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): cover disconnect scans on every streaming endpoint and client, plus outage, worker-kill and cache-hit cells
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): prove the cache-hit twin is served from the cache
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): type the disconnect-close streams so basedpyright stops reporting unknown arguments
* fix(guardrails): give the guardrail metadata cast a reason so the type discipline gate accepts it
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan released Messages and Responses tool calls on disconnect
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* style(guardrails): format the disconnect scan unit tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore(guardrails): drop mutable-ok markers that no longer suppress a rule
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan released Responses output after a finished item and end only in-flight Chat choices
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): type the disconnect scan test helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* refactor(guardrails): type request_data in the disconnect scan helpers
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): type request_data in the disconnect scan test doubles
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): pin that chat streams with no tool call in flight end as released
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): scan only released chunks on disconnect and skip it once a block owns the verdict
The disconnect scan now uses the chunks actually yielded to the client, copies them before scanning,
skips when a mid-stream block or HTTP error already settled the verdict, and the iterator wrapper
only closes hooks that are async generators so plain async iterator hooks keep working
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(guardrails): pin that a delivered guardrail error or final chunk settles the disconnect verdict
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(guardrails): close any hook iterator that exposes aclose when the stream ends early
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): accept a synchronous aclose on custom streaming hook iterators
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(proxy): swallow callback aclose errors at end of stream
A custom callback whose async_post_call_streaming_iterator_hook returns a
non-generator async iterator with a raising aclose() failed the finished
stream: content plus usage reached the client and then the stream surfaced
an error SSE with no [DONE], or aborted a post_call pipeline's buffering
loop into a 500 with an empty body. Wrap the aclose invocation in
_wrap_streaming_iterator_with_enrichment in try/except and log a warning
naming the callback and the cleanup error, matching close_guarded_stream
and _close_guarded_layers. Iteration-time hook exceptions still propagate.
* fix(proxy): log only the error type when a callback aclose raises
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>