* test(mcp): exercise /mcp/proxy authorization against the real registry instead of patched manager methods
* fix(mcp): preserve proxy logging and authorization coverage
* test(mcp): respect the proxy FastAPI import boundary
Extract one immutable helper for reading the deployment's model_info off the
logging object, stop rebinding the model_info parameter inside ocr_cost, drop
the explanatory comment blocks, and move the OCR custom pricing regression
tests into tests/test_litellm/test_cost_calculator.py
The bridges derived prompt_cache_key as the first 64 chars of metadata.user_id.
Claude Code packs a JSON object into that field whose prefix is the per-install
device_id, so every session and subagent on one machine shared a single key,
and a plain end-user id pinned all of that user's conversations to one slot.
Parse the JSON and use session_id; send no key otherwise so the provider falls
back to its own prompt-prefix hashing. An explicit prompt_cache_key still wins.
Fixes#39145
* fix(otel v2): restore the Datadog auth span and the last-wins callback merge
Move @tracer.wrap() back onto user_api_key_auth so USE_DDTRACE=true emits the
auth span again, and let a failure entry's callback_vars take part in the
destination merge so the resolver picks the same account the runtime parser does
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(otel v2): drop docstrings from the two regression tests
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* ci: rerun proxy-infra after the flaky test_check_migration process-tree test
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
A translation that hands the hook no assembled response leaves the stream
as it is, and a rewrite the translation cannot rescan is released as the
original stream with a warning.
A Messages stream that ends with only tool_use blocks reaches the legacy
step without a texts key, while the non-streaming rescan of the same
response sends an empty list. Compare both as empty and keep the stream
when the hook left the tool calls alone.
Retrieval re-matches only the key, team, and model scopes, so a post_call
policy that reached a pending background response through a request-body
tag does not govern the completed response. Log that at submit, next to the
deferral, and cover the retrieval re-match with tag-scoped tests.
Register a HostedVLLMImageEditConfig so hosted_vllm/<model> deployments route POST /v1/images/edits to the vLLM-Omni OpenAI-compatible endpoint instead of failing with 'image edit is not supported for hosted_vllm' before any request is sent
Bedrock Guardrails resolved credentials and signed inline on the loop in
its three async paths, Titan embeddings signed each batch item inline in
the async loop, and Bedrock Mantle requests slipped past the off-loop
gate because BedrockMantleAuthMixin composes a BaseAWSLLM instead of
inheriting from it. Introduce the SignsRequestsWithAWS marker that both
BaseAWSLLM and the Mantle mixin carry so sign_request_off_loop_if_aws
covers Mantle, and move the guardrail and embedding signing into
asyncio.to_thread. Every new test fails at the previous tip.
* test(azure_sentinel): pin batch_size as a per-request bound under concurrent events
Adds a regression test to the mapped Azure Sentinel test file for the concurrency scenario from LIT-6920: 40 records logged concurrently at batch_size=5 while each ingestion request is still in flight. Asserts no request carries more than batch_size records, every record arrives exactly once in order, and the queue is empty afterwards. Runs for both the standard log queue and the audit log queue.
The test fails on the tree before #39880 (whole shared queue serialized per threshold send, then cleared) and passes on current staging. It is independent of the size-split coverage that #39880 added for LIT-5899.
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* test(azure_sentinel): gate the first send on events so later records provably arrive while it is in flight
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The per-chunk hook skipped pipeline-managed guardrails whenever the request
route was empty, while the gated stream could still fail to resolve a
translation and release the buffered stream ungoverned. The gate now needs a
translation resolved from the route, the iterator hook resolves it once and
hands it to the gated stream, and the ungoverned release branch is gone.
Turning the tier off, or switching to a classifier that cannot emit it, dropped
the flag and the pool but left plan_mode_min_tier naming a tier that is no longer
active. The backend rejects that on save, and the switch is disabled after a
classifier change, so the operator had no way to clear it.
Both paths now release the floor when it points at the cleared tier. An orphaned
keyword rule is left alone on purpose: getKeywordTierRulesError already names it
at the save gate, which is how a removed custom tier behaves.
A streaming step whose rewrite the executor threw away (a tool-call rewrite, a text rewrite
the translation cannot write back, or one the adapter refused) still marked its guardrail as
applied, so the header claimed an output the client never received. The step now returns
right after releasing the original chunks, which leaves the header as the merge base sent it