Commit graph

4 commits

Author SHA1 Message Date
mateo-berri
c2f77fd358 refactor(a2a): resolve the relay's Entra hop bearer inside the a2a provider helper 2026-09-16 17:56:42 -07:00
mateo-berri
e243237a7c feat(a2a): reach Microsoft Foundry agents with Entra auth and versioned card discovery
Foundry serves its agent card only at agentCard/v1.0, accepts only an Entra ID
bearer, and defaults to a non-blocking send, so the A2A relay and the chat
completions route could not use it.

The relay gains an agent_card_path litellm_param plus agentCard/v1.0 as a third
discovery probe, mints a bearer from flat Entra fields on the agent
(tenant_id, client_id, client_secret, azure_ad_token, azure_username,
azure_password, azure_scope) for https://ai.azure.com/.default, and sends it on
the card fetch, message/send, message/stream, tasks/* and the chat bridge.

Chat completions look the registered agent up by its provider-stripped name so
its api_key and headers reach the request, tag every message with its kind, ask
for a blocking send, fall back to a blocking send when the registered card says
streaming: false, and fail the call on a JSON-RPC error inside a stream instead
of yielding an empty one. Entra fields stay out of the chat bridge's logged
parameters.

Resolves LIT-5122
2026-09-16 16:36:47 -07:00
Mateo Wang
4f7b20ec10
fix(guardrails): skip streaming guardrail rounds that re-scan cleared output (#39386)
* fix(guardrails): skip streaming guardrail rounds that re-scan cleared output

Streaming guardrails scanned the finished answer twice at end of stream
whenever the chunk count landed on a multiple of the sampling rate, ran
sampled rounds whose payload was identical to the previous one, and on
/v1/messages could scan an empty text before the first content chunk.
Every redundant round is a paid guardrail provider call.

Each endpoint handler now exposes a scan key describing what a round
would hand to apply_guardrail (the text so far, plus tool calls once the
stream has ended), and the unified streaming hook skips a sampled or
end-of-stream round whose key equals the last scanned one or carries
nothing to scan yet. Rounds that carry tool calls are never skipped.

* test(guardrails): expect one end-of-stream scan when the terminal chunk is sampled

Update sampled cadence expectations and use tuple-backed scan state

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yassin <yassin@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-02 18:25:28 -07:00
michelligabriele
cbe23d6184
fix(a2a): populate response usage in a2a chat transformation (#31980)
* fix(a2a): populate response usage in a2a chat transformation

* test(a2a): rename to avoid module basename collision with fastcrw's test_transformation.py
2026-07-03 09:28:36 +05:30