mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-05 02:41:56 +00:00
* fix(utils): resolve bedrock regional inference profiles to regional pricing in get_model_info (LIT-4056) (#32389) * fix(utils): resolve bedrock regional inference profiles to regional pricing in get_model_info (LIT-4056) * test(register_model): use a triple provider prefix as the unresolvable-key fixture get_model_info now resolves bedrock/bedrock/... like a routing prefix, so the double-prefix fixture stopped exercising the register_model fallback path. Lock the new double-prefix resolution in as a model-info regression test (cherry picked from commit734fd29e00) * fix(guardrails): walk Responses-API text taxonomy in shared content helpers (#32542) * fix(guardrails): walk Responses-API text taxonomy in shared content helpers Every guardrail sharing litellm/proxy/guardrails/_content_utils.py silently drops all text on the /v1/responses path. AIM turns it into a loud 422 ( {"error":"No messages in the request"}); every other guardrail (Lakera v2, Cato, Lasso, Repello, IBM, Azure Content Safety, enterprise secret detection) scans an empty payload and lets the request through unscanned. Three defects, all in _content_utils.py: 1. _iter_text_parts_in_content recognised only part.type == "text", but the Responses API uses input_text (request) and output_text (assistant). 2. _coerce_input_to_messages gated on "every item has a role key"; any Responses input list containing a function_call or function_call_output item failed the check and was wrapped as one opaque blob. 3. build_inspection_messages forwarded any role through, including a bare tool role missing tool_call_id, which validators like AIM's /fw/v1/analyze reject with a schema error. Fix walks the actual Responses item taxonomy (message, function_call, function_call_output, bare content parts and strings), recognises {text, input_text, output_text} everywhere, and coerces any role outside {system, user, assistant} to user in the outbound inspection payload. * style: ruff-format changed guardrail files * test(guardrails): cover function_call_output string form; drop em-dash in new docstring * fix(guardrails): map function_call_output straight to user role Avoids ever materialising a schema-invalid bare tool message. The downstream role-safety coercion in build_inspection_messages still guards genuinely caller-supplied non-standard roles (developer, function, custom values); add a regression test covering that path so the coercion has real coverage after this simplification. * test(guardrails): pin chat-completions tool-role coercion in build_inspection_messages * docs(test): soften AIM-specific claims in LIT-4294 test docstrings Ryan's review flagged that several test docstrings assert AIM's /fw/v1/analyze validates + rejects specific schema violations. That behavior is customer-reported in the LIT-4294 writeup, not directly verified by us. Rephrase to attribute the AIM 422 to the customer's writeup and describe the underlying constraint as the OpenAI chat schema; any downstream API that validates against that schema rejects the same shape. * refactor(guardrails): move unsupported-role coercion into AIM only The generic coercion in build_inspection_messages collapsed any role outside {system, user, assistant} to user for every caller of the helper. Combined with the pre-existing apply_redacted_messages_back write-back behavior in Lakera/AIM/Cato, that turned a loud OpenAI 400 on chat-completions tool-message masking into a silent semantic corruption of the outbound request (role tool with tool_call_id got rewritten to bare role user, dropping the assistant + tool_calls sibling). AIM specifically requires the coercion because its /fw/v1/analyze validates the payload against the OpenAI chat schema; other guardrails either do not validate roles or do their own reconstruction. Move the coercion to AimGuardrail._build_aim_inspection_messages so the shared helper keeps caller roles intact and no new cross-guardrail role corruption is introduced. The pre-existing apply_redacted_messages_back structural flatten remains as separate follow-up work. function_call_output items still synthesise role user in the shared helper because they have no natural role field, which is a different concern from coercing a caller-supplied role. * refactor(guardrails): preserve role fidelity in shared _content_utils Shared inspection helpers should extract text and preserve semantic role signals; role coercion for third-party schema safety stays inside the guardrail that needs it (AIM). Three shared-helper changes: - Bare content-part dicts (input_text/output_text) with an explicit role keep it; only role-less parts default to user. - Responses message items already had their role preserved; the behavior is now covered by an explicit test. - function_call_output items default to role tool (semantic equivalent of the chat-completions tool message shape) instead of role user, so Responses and chat completions produce symmetric inspection payloads. A caller-supplied role on the item is still preserved. AIM's schema-safe coercion in _build_aim_inspection_messages already handles the resulting role tool: it collapses to user before the POST to /fw/v1/analyze so AIM's OpenAI-schema validator does not reject the bare tool message (no tool_call_id can survive the flatten). Added a regression test in test_aim.py covering that path. (cherry picked from commite84a19acd5) * feat: add Meta Model API provider and muse-spark-1.1 (day-0) (#32701) (cherry picked from commitd82645d163) * fix(bedrock): keep mid-conversation system messages in place for Claude Invoke (#32578) Hoisting every role system entry into the top-level system field mutates the cache prefix whenever a client such as Claude Code appends a new mid-conversation system message, invalidating the prompt cache for the entire message history on Bedrock Invoke. Bedrock only rejects a system entry at messages.0, so hoist just the leading run and forward the rest in place (cherry picked from commitcc36d5469c) * feat(otel): emit the gen_ai.client.operation.exception event on failed LLM calls (#32655) * feat(otel): emit the gen_ai.client.operation.exception event on failed LLM calls The GenAI semantic conventions record failures of a GenAI client operation as a log-based event named gen_ai.client.operation.exception, carrying the exception.type / exception.message / exception.stacktrace trio at severity WARN and correlated to the failed span. OTel v2 never emitted it: a failed LLM call produced only the deprecated error.* span attributes, a generic exception span event without a stacktrace, and the stacktrace under the vendor key litellm.provider.error.stack_trace. Build the logs pipeline (LoggerProvider + console/OTLP log exporters mirroring the metrics plumbing) and record the event behind the enable_events flag, which until now was defined but consumed nowhere. An operator-configured LoggerProvider global is reused so the events ride their existing logs pipeline; an explicit NoOpLoggerProvider global is honored as an opt-out and builds no recorder at all. The existing span-side error surface (error.type, error.message, the exception span event, and the litellm.provider.error.* detail keys) is untouched for backwards compatibility. * fix(otel): always ride the semconv-required exception pair on the GenAI event Filtering the event attributes on truthiness conflated "absent" with "empty", so an empty exception.type or exception.message would have been dropped, leaving an event with neither semconv-required field. Build the attributes so the pair is unconditional and only the recommended stacktrace is omitted when the payload carries none. * docs(otel): document the events plumbing module in the package README * test(otel): cover the log exporter selection and logs endpoint normalization The new logs plumbing had no coverage for exporter-kind selection, the console fallback for an unrecognized kind, the /v1/logs signal-path rewriting that lets one OTEL_ENDPOINT serve every signal, or the simple-vs-batch processor split. (cherry picked from commit99b4c5ed3e) * fix(bedrock): gate in-place system role messages on model support for Claude Invoke (#32831) * fix(bedrock): gate in-place system role messages on model support for Claude Invoke * feat(bedrock): default unmapped Claude 4.8+ to in-place system role handling via fallback rule (cherry picked from commit5e23a5ab05) * fix(anthropic): translate adaptive thinking/effort to pre-4.6 model support (#32867) * fix(anthropic): translate adaptive thinking/effort to pre-4.6 model support AnthropicMessagesConfig now reshapes the 4.6+ adaptive-thinking interface (thinking:{type:adaptive} + output_config:{effort:...}) to whatever the routed model supports. Thinking-capable non-adaptive models (e.g. Haiku 4.5, Sonnet 4.5) get the effort translated to a legacy thinking budget_tokens. Models with no reasoning support have thinking/effort dropped under drop_params. And because adaptive thinking carries no budget while the legacy form must satisfy Anthropic's max_tokens > budget_tokens rule, the translated budget is capped below max_tokens, dropping thinking when max_tokens can't fit the minimum budget. 4.6+ models pass through untouched. This matters because clients like Claude Code speak native Anthropic /v1/messages and send the adaptive interface unconditionally, regardless of the routed model. The native passthrough previously only capability-gated the OpenAI-style reasoning_effort alias and forwarded native output_config/adaptive thinking raw, so a pre-4.6 model rejected it with "This model does not support the effort parameter" and the request failed. Claude Code already gets drop_params auto-set, so its requests now succeed. * test(anthropic): gate undersized-max_tokens thinking drop on drop_params; add edge tests Addresses review feedback on the max_tokens-too-small branch. Previously a thinking-capable model whose max_tokens could not fit the minimum thinking budget had thinking silently dropped regardless of drop_params, while a residual output_config field in the same call still raised when drop_params was off. Gate both consistently on drop_params: raise a clear error (naming max_tokens for the undersized case) when drop_params is off, drop otherwise. Claude Code gets drop_params auto-set, so it still succeeds. Adds tests for the undersized-max_tokens raise, the residual output_config raise, and the no-adaptive-interface passthrough on a non-adaptive model. * fix(anthropic): make adaptive-effort translation silent to avoid breaking provider strip contracts The previous raise-when-not-drop_params behavior broke existing bedrock and vertex messages tests: those providers already silently strip unsupported output_config for pre-4.6 models (issue #22797) with no drop_params required, and the shared parent transform raising pre-empted that. It also conflicted with the goal of keeping requests working rather than failing them. Make the reshape silent: translate effort to legacy thinking for thinking-capable models, drop thinking for non-reasoning models, and remove only the consumed effort key from output_config, leaving any residual (e.g. format) for provider subclasses (bedrock/vertex) to handle. No raise, no drop_params gating. This also resolves the review note about inconsistent drop_params handling by making every path uniform. Updates the tests to assert the silent behavior and residual output_config preservation. * fix(anthropic): handle output_config-capable but non-adaptive models (Opus 4.5) Greptile caught a real bug: the early-return guard treated supports_output_config as equivalent to supporting adaptive thinking. Claude Opus 4.5 advertises supports_output_config (it accepts output_config.effort) but is not adaptive, so it rejects thinking:{type:adaptive} with "adaptive thinking is not supported on this model". The guard early-returned for Opus 4.5 and forwarded the adaptive thinking block raw, reproducing the exact failure the fix is meant to prevent. thinking:{type:adaptive} and output_config.effort are independent capabilities. Only early-return for adaptive-thinking models. For a model that supports output_config.effort but is not adaptive, keep the native effort and drop only the unsupported adaptive thinking block. Verified live against Opus 4.5: the Claude Code payload now returns 200 instead of 400. Adds regression tests for Opus 4.5 with and without adaptive thinking. * fix(anthropic): translate adaptive thinking for effort-capable pre-4.6 models Claude Opus 4.5 advertises supports_output_config but not adaptive thinking, so the early-return guard forwarded thinking.type=adaptive raw and Anthropic rejected it. The guard now only skips true adaptive models; effort-only requests on effort-capable models still pass through untouched. The _map_reasoning_effort call is wrapped to surface unrecognized effort values as a clean 400, matching _translate_reasoning_effort_to_anthropic * fix(anthropic): fall back to legacy thinking when effort level unsupported Opus 4.5 accepts output_config.effort but only low/medium/high; Claude Code defaults to xhigh on newer models, so preserving that level raw gets rejected by Anthropic. Gate the native-effort passthrough on _validate_effort_for_model and fall through to the budget translation for unsupported levels * fix(anthropic): keep effort-only requests untouched for provider normalization The xhigh fall-through consumed effort-only requests on effort-capable models, breaking bedrock invoke's own normalization which clamps xhigh to the model's ceiling after the base transform runs (test_bedrock_messages_normalizes_output_config_effort_for_opus). Restrict the fall-through to requests that carry adaptive thinking; effort-only requests pass through so provider subclasses keep owning level clamping --------- Co-authored-by: Abhimanyu Kapur <38531241+akapur99@users.noreply.github.com> (cherry picked from commit3a62e5428f) * fix(bedrock): flag mapped Claude 4.8+ entries with supports_mid_conversation_system (#32882) Exact cost-map hits resolve before fallback-generalization rules, so the mapped Sonnet 5, Fable 5 and jp Opus 4.8 Bedrock entries bypassed the bedrock-anthropic-claude-mid-conversation-system rule and hoisted mid-conversation system messages, invalidating the prompt cache. (cherry picked from commitc15891fc98) * Merge pull request #32873 from BerriAI/litellm_fallback_rules_routing_split refactor(fallback-generalizations): split rules into routing and provider-neutral capability kinds (cherry picked from commit45d3644408) * Merge pull request #32874 from BerriAI/litellm_thread_provider_capability_probes fix(anthropic): thread real provider through capability probes instead of pinning anthropic (cherry picked from commitead7ad3804) * test: add /v1/messages to supported_endpoints schema enum (#32739) (cherry picked from commitbf02a4a47f) --------- Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: yucheng-berri <yucheng@berri.ai> Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: Yassin Kortam <yassin@berri.ai> Co-authored-by: Abhimanyu Kapur <38531241+akapur99@users.noreply.github.com> Co-authored-by: tin-berri <tin@berri.ai>
526 lines
19 KiB
Python
526 lines
19 KiB
Python
"""Tests for the shared guardrail content extraction helpers."""
|
|
|
|
from litellm.proxy.guardrails._content_utils import (
|
|
apply_redacted_messages_back,
|
|
build_inspection_messages,
|
|
has_non_string_content,
|
|
iter_message_text,
|
|
walk_user_text,
|
|
)
|
|
|
|
# ── iter_message_text ────────────────────────────────────────────────────────────
|
|
|
|
|
|
def test_iter_message_text_string_messages():
|
|
data = {
|
|
"messages": [
|
|
{"role": "user", "content": "hello"},
|
|
{"role": "assistant", "content": "hi"},
|
|
]
|
|
}
|
|
assert list(iter_message_text(data)) == ["hello", "hi"]
|
|
|
|
|
|
def test_iter_message_text_multimodal_list_content():
|
|
"""VERIA-11: list-format content must be inspected, not silently skipped."""
|
|
data = {
|
|
"messages": [
|
|
{
|
|
"role": "user",
|
|
"content": [
|
|
{"type": "text", "text": "AWS_KEY=AKIA..."},
|
|
{"type": "image_url", "image_url": {"url": "..."}},
|
|
{"type": "text", "text": "more secrets"},
|
|
],
|
|
}
|
|
]
|
|
}
|
|
assert list(iter_message_text(data)) == ["AWS_KEY=AKIA...", "more secrets"]
|
|
|
|
|
|
def test_iter_message_text_responses_api_string_input():
|
|
"""fniVO9-F: Responses-API ``input`` must be inspectable when ``messages`` absent."""
|
|
data = {"input": "tell me a secret"}
|
|
assert list(iter_message_text(data)) == ["tell me a secret"]
|
|
|
|
|
|
def test_iter_message_text_responses_api_list_input_messages():
|
|
data = {
|
|
"input": [
|
|
{"role": "user", "content": "first"},
|
|
{"role": "user", "content": "second"},
|
|
]
|
|
}
|
|
assert list(iter_message_text(data)) == ["first", "second"]
|
|
|
|
|
|
def test_iter_message_text_responses_api_list_input_content_parts():
|
|
data = {
|
|
"input": [
|
|
{"type": "text", "text": "alpha"},
|
|
{"type": "image_url", "image_url": {"url": "..."}},
|
|
{"type": "text", "text": "beta"},
|
|
]
|
|
}
|
|
assert list(iter_message_text(data)) == ["alpha", "beta"]
|
|
|
|
|
|
def test_iter_message_text_responses_api_list_input_mixed_dicts_and_strings():
|
|
"""Greptile P2: mixed-list ``input`` with content-part dicts AND bare
|
|
strings must yield every text fragment — read helpers used to truncate
|
|
the bare strings."""
|
|
data = {
|
|
"input": [
|
|
{"type": "text", "text": "from-dict"},
|
|
"from-bare-string",
|
|
{"type": "image_url", "image_url": {"url": "..."}},
|
|
"another-bare-string",
|
|
]
|
|
}
|
|
assert list(iter_message_text(data)) == [
|
|
"from-dict",
|
|
"from-bare-string",
|
|
"another-bare-string",
|
|
]
|
|
|
|
|
|
def test_iter_message_text_walks_messages_and_input_independently():
|
|
"""When both are present (rare), every fragment from either field is
|
|
inspected — a stricter guarantee than "first one wins"."""
|
|
data = {
|
|
"messages": [{"role": "user", "content": "msg-content"}],
|
|
"input": "input-content",
|
|
}
|
|
assert list(iter_message_text(data)) == ["msg-content", "input-content"]
|
|
|
|
|
|
def test_iter_message_text_empty_data():
|
|
assert list(iter_message_text({})) == []
|
|
assert list(iter_message_text({"messages": []})) == []
|
|
assert list(iter_message_text({"input": ""})) == []
|
|
|
|
|
|
def test_iter_message_text_responses_api_input_text_and_output_text_parts():
|
|
"""LIT-4294: Responses-API content parts use ``input_text`` (request) and
|
|
``output_text`` (assistant); reading only ``type == "text"`` skipped every
|
|
``/v1/responses`` body and every text guardrail was a no-op on that path."""
|
|
data = {
|
|
"input": [
|
|
{
|
|
"type": "message",
|
|
"role": "user",
|
|
"content": [{"type": "input_text", "text": "user text"}],
|
|
},
|
|
{
|
|
"type": "message",
|
|
"role": "assistant",
|
|
"content": [{"type": "output_text", "text": "assistant text"}],
|
|
},
|
|
]
|
|
}
|
|
assert list(iter_message_text(data)) == ["user text", "assistant text"]
|
|
|
|
|
|
def test_iter_message_text_responses_api_tool_call_taxonomy():
|
|
"""LIT-4294: a Responses ``input`` list freely mixes message items,
|
|
``function_call`` (no ``role``), and ``function_call_output`` items. The
|
|
old ``all(item has 'role')`` gate wrapped the whole list as one blob and
|
|
yielded nothing; every text fragment must be visited independently."""
|
|
data = {
|
|
"input": [
|
|
{
|
|
"type": "message",
|
|
"role": "user",
|
|
"content": [{"type": "input_text", "text": "hello"}],
|
|
},
|
|
{
|
|
"type": "function_call",
|
|
"call_id": "c1",
|
|
"name": "get_weather",
|
|
"arguments": "{}",
|
|
},
|
|
{
|
|
"type": "function_call_output",
|
|
"call_id": "c1",
|
|
"output": [{"type": "input_text", "text": "sunny"}],
|
|
},
|
|
]
|
|
}
|
|
assert list(iter_message_text(data)) == ["hello", "sunny"]
|
|
|
|
|
|
# ── walk_user_text ────────────────────────────────────────────────────────────
|
|
|
|
|
|
def test_walk_user_text_redacts_string_messages_in_place():
|
|
data = {
|
|
"messages": [
|
|
{"role": "user", "content": "leak: AKIAEXAMPLE"},
|
|
{"role": "assistant", "content": "ok"},
|
|
]
|
|
}
|
|
visited = walk_user_text(data, lambda s: s.replace("AKIAEXAMPLE", "[REDACTED]"))
|
|
assert visited == 2
|
|
assert data["messages"][0]["content"] == "leak: [REDACTED]"
|
|
assert data["messages"][1]["content"] == "ok"
|
|
|
|
|
|
def test_walk_user_text_redacts_multimodal_text_parts():
|
|
"""VERIA-11: list-content text parts must be mutable for in-place redaction."""
|
|
data = {
|
|
"messages": [
|
|
{
|
|
"role": "user",
|
|
"content": [
|
|
{"type": "text", "text": "AKIAEXAMPLE here"},
|
|
{"type": "image_url", "image_url": {"url": "..."}},
|
|
{"type": "text", "text": "no secret"},
|
|
],
|
|
}
|
|
]
|
|
}
|
|
visited = walk_user_text(data, lambda s: s.replace("AKIAEXAMPLE", "[REDACTED]"))
|
|
assert visited == 2
|
|
parts = data["messages"][0]["content"]
|
|
assert parts[0] == {"type": "text", "text": "[REDACTED] here"}
|
|
# Non-text part must be left untouched.
|
|
assert parts[1] == {"type": "image_url", "image_url": {"url": "..."}}
|
|
assert parts[2] == {"type": "text", "text": "no secret"}
|
|
|
|
|
|
def test_walk_user_text_redacts_responses_api_string_input():
|
|
data = {"input": "leak AKIAEXAMPLE"}
|
|
visited = walk_user_text(data, lambda s: s.replace("AKIAEXAMPLE", "[REDACTED]"))
|
|
assert visited == 1
|
|
assert data["input"] == "leak [REDACTED]"
|
|
|
|
|
|
def test_walk_user_text_redacts_responses_api_list_input():
|
|
data = {
|
|
"input": [
|
|
{"type": "text", "text": "AKIAEXAMPLE"},
|
|
{"type": "image_url", "image_url": {"url": "..."}},
|
|
]
|
|
}
|
|
visited = walk_user_text(data, lambda s: f"[redacted]{s}[/]")
|
|
assert visited == 1
|
|
assert data["input"][0] == {"type": "text", "text": "[redacted]AKIAEXAMPLE[/]"}
|
|
assert data["input"][1] == {"type": "image_url", "image_url": {"url": "..."}}
|
|
|
|
|
|
def test_walk_user_text_redacts_responses_input_text_and_output_text_parts():
|
|
"""LIT-4294: ``walk_user_text`` must recognise the Responses text-part
|
|
variants so masking guardrails (secret detection, PII) actually redact
|
|
``/v1/responses`` bodies instead of no-op'ing on them."""
|
|
data = {
|
|
"input": [
|
|
{
|
|
"type": "message",
|
|
"role": "user",
|
|
"content": [{"type": "input_text", "text": "AKIAEXAMPLE"}],
|
|
},
|
|
{
|
|
"type": "message",
|
|
"role": "assistant",
|
|
"content": [{"type": "output_text", "text": "AKIAEXAMPLE too"}],
|
|
},
|
|
]
|
|
}
|
|
visited = walk_user_text(data, lambda s: s.replace("AKIAEXAMPLE", "[REDACTED]"))
|
|
assert visited == 2
|
|
assert data["input"][0]["content"][0] == {
|
|
"type": "input_text",
|
|
"text": "[REDACTED]",
|
|
}
|
|
assert data["input"][1]["content"][0] == {
|
|
"type": "output_text",
|
|
"text": "[REDACTED] too",
|
|
}
|
|
|
|
|
|
def test_walk_user_text_redacts_function_call_output_text():
|
|
"""LIT-4294: tool-call round-trips carry secrets in
|
|
``function_call_output.output``; the redact walker must descend into it
|
|
while leaving ``function_call`` items (call_id, arguments) untouched."""
|
|
data = {
|
|
"input": [
|
|
{
|
|
"type": "message",
|
|
"role": "user",
|
|
"content": [{"type": "input_text", "text": "AKIAEXAMPLE user"}],
|
|
},
|
|
{
|
|
"type": "function_call",
|
|
"call_id": "c1",
|
|
"name": "get_weather",
|
|
"arguments": '{"AKIAEXAMPLE": 1}',
|
|
},
|
|
{
|
|
"type": "function_call_output",
|
|
"call_id": "c1",
|
|
"output": [{"type": "input_text", "text": "AKIAEXAMPLE tool"}],
|
|
},
|
|
]
|
|
}
|
|
visited = walk_user_text(data, lambda s: s.replace("AKIAEXAMPLE", "[REDACTED]"))
|
|
assert visited == 2
|
|
assert data["input"][0]["content"][0]["text"] == "[REDACTED] user"
|
|
assert data["input"][1] == {
|
|
"type": "function_call",
|
|
"call_id": "c1",
|
|
"name": "get_weather",
|
|
"arguments": '{"AKIAEXAMPLE": 1}',
|
|
}
|
|
assert data["input"][2]["output"][0]["text"] == "[REDACTED] tool"
|
|
|
|
|
|
def test_walk_user_text_redacts_function_call_output_string_output():
|
|
"""LIT-4294: ``function_call_output.output`` is also a plain string in
|
|
OpenAI's Responses spec; the redact walker must handle both forms."""
|
|
data = {
|
|
"input": [
|
|
{
|
|
"type": "function_call_output",
|
|
"call_id": "c1",
|
|
"output": "AKIAEXAMPLE tool",
|
|
},
|
|
]
|
|
}
|
|
visited = walk_user_text(data, lambda s: s.replace("AKIAEXAMPLE", "[REDACTED]"))
|
|
assert visited == 1
|
|
assert data["input"][0]["output"] == "[REDACTED] tool"
|
|
|
|
|
|
def test_walk_user_text_redacts_mixed_list_input():
|
|
"""Read and write helpers must agree on coverage — bare strings inside
|
|
a mixed ``input`` list are inspected by both."""
|
|
data = {
|
|
"input": [
|
|
{"type": "text", "text": "secret-one"},
|
|
"secret-two",
|
|
{"type": "image_url", "image_url": {"url": "..."}},
|
|
]
|
|
}
|
|
visited = walk_user_text(data, lambda s: f"<{s}>")
|
|
assert visited == 2
|
|
assert data["input"][0] == {"type": "text", "text": "<secret-one>"}
|
|
assert data["input"][1] == "<secret-two>"
|
|
assert data["input"][2] == {"type": "image_url", "image_url": {"url": "..."}}
|
|
|
|
|
|
# ── build_inspection_messages ─────────────────────────────────────────────────
|
|
|
|
|
|
def test_build_inspection_messages_chat_completion_passthrough():
|
|
data = {
|
|
"messages": [
|
|
{"role": "system", "content": "be helpful"},
|
|
{"role": "user", "content": "hi"},
|
|
]
|
|
}
|
|
assert build_inspection_messages(data) == [
|
|
{"role": "system", "content": "be helpful"},
|
|
{"role": "user", "content": "hi"},
|
|
]
|
|
|
|
|
|
def test_build_inspection_messages_joins_multimodal_text_parts():
|
|
data = {
|
|
"messages": [
|
|
{
|
|
"role": "user",
|
|
"content": [
|
|
{"type": "text", "text": "first part"},
|
|
{"type": "image_url", "image_url": {"url": "..."}},
|
|
{"type": "text", "text": "second part"},
|
|
],
|
|
}
|
|
]
|
|
}
|
|
assert build_inspection_messages(data) == [{"role": "user", "content": "first part\nsecond part"}]
|
|
|
|
|
|
def test_build_inspection_messages_lifts_responses_api_input():
|
|
"""fniVO9-F: ``input`` must be visible to hooks that POST messages to a remote API."""
|
|
data = {"input": "responses-api content"}
|
|
assert build_inspection_messages(data) == [{"role": "user", "content": "responses-api content"}]
|
|
|
|
|
|
def test_build_inspection_messages_drops_messages_with_no_text():
|
|
data = {
|
|
"messages": [
|
|
{"role": "user", "content": ""},
|
|
{
|
|
"role": "user",
|
|
"content": [{"type": "image_url", "image_url": {"url": "..."}}],
|
|
},
|
|
{"role": "user", "content": "kept"},
|
|
]
|
|
}
|
|
assert build_inspection_messages(data) == [{"role": "user", "content": "kept"}]
|
|
|
|
|
|
def test_build_inspection_messages_responses_api_tool_call_taxonomy():
|
|
"""LIT-4294: mixed Responses ``input`` (message + function_call +
|
|
function_call_output) must produce a non-empty inspection list. The
|
|
customer's writeup reproduced a 422 from AIM's ``/fw/v1/analyze``
|
|
(``No messages in the request``) when this synthesised list came back
|
|
empty; every other guardrail silently scanned nothing on the same
|
|
input."""
|
|
data = {
|
|
"input": [
|
|
{
|
|
"type": "message",
|
|
"role": "user",
|
|
"content": [{"type": "input_text", "text": "hello"}],
|
|
},
|
|
{
|
|
"type": "function_call",
|
|
"call_id": "c1",
|
|
"name": "get_weather",
|
|
"arguments": "{}",
|
|
},
|
|
{
|
|
"type": "function_call_output",
|
|
"call_id": "c1",
|
|
"output": [{"type": "input_text", "text": "sunny"}],
|
|
},
|
|
]
|
|
}
|
|
assert build_inspection_messages(data) == [
|
|
{"role": "user", "content": "hello"},
|
|
{"role": "tool", "content": "sunny"},
|
|
]
|
|
|
|
|
|
def test_build_inspection_messages_function_call_output_defaults_to_tool():
|
|
"""LIT-4294: a Responses ``function_call_output`` item is the semantic
|
|
equivalent of a chat-completions ``role: "tool"`` message, so the shared
|
|
helper synthesises ``role: "tool"`` when the item has no explicit role.
|
|
AIM's schema-safe coercion happens at the AIM call site, not here."""
|
|
data = {
|
|
"input": [
|
|
{
|
|
"type": "function_call_output",
|
|
"call_id": "c1",
|
|
"output": [{"type": "input_text", "text": "tool text"}],
|
|
},
|
|
]
|
|
}
|
|
assert build_inspection_messages(data) == [{"role": "tool", "content": "tool text"}]
|
|
|
|
|
|
def test_build_inspection_messages_function_call_output_preserves_explicit_role():
|
|
"""When ``function_call_output`` carries a caller-supplied ``role`` the
|
|
shared helper preserves it rather than synthesising ``tool``."""
|
|
data = {
|
|
"input": [
|
|
{
|
|
"type": "function_call_output",
|
|
"role": "assistant",
|
|
"call_id": "c1",
|
|
"output": [{"type": "input_text", "text": "tool text"}],
|
|
},
|
|
]
|
|
}
|
|
assert build_inspection_messages(data) == [{"role": "assistant", "content": "tool text"}]
|
|
|
|
|
|
def test_build_inspection_messages_bare_content_part_preserves_explicit_role():
|
|
"""A bare content-part dict with an explicit ``role`` keeps it. Only
|
|
absent roles get defaulted to ``user``."""
|
|
data = {
|
|
"input": [
|
|
{"type": "input_text", "text": "no role"},
|
|
{"type": "output_text", "role": "assistant", "text": "with role"},
|
|
]
|
|
}
|
|
assert build_inspection_messages(data) == [
|
|
{"role": "user", "content": "no role"},
|
|
{"role": "assistant", "content": "with role"},
|
|
]
|
|
|
|
|
|
def test_build_inspection_messages_message_item_preserves_role():
|
|
"""Responses message items carry a role explicitly; the shared helper
|
|
passes it through untouched."""
|
|
data = {
|
|
"input": [
|
|
{"type": "message", "role": "system", "content": [{"type": "input_text", "text": "sys"}]},
|
|
{"type": "message", "role": "assistant", "content": [{"type": "output_text", "text": "asst"}]},
|
|
]
|
|
}
|
|
assert build_inspection_messages(data) == [
|
|
{"role": "system", "content": "sys"},
|
|
{"role": "assistant", "content": "asst"},
|
|
]
|
|
|
|
|
|
def test_build_inspection_messages_empty_data():
|
|
assert build_inspection_messages({}) == []
|
|
assert build_inspection_messages({"messages": []}) == []
|
|
assert build_inspection_messages({"input": ""}) == []
|
|
|
|
|
|
# ── has_non_string_content ────────────────────────────────────────────────────
|
|
|
|
|
|
def test_has_non_string_content_string_messages():
|
|
data = {"messages": [{"role": "user", "content": "hello"}]}
|
|
assert has_non_string_content(data) is False
|
|
|
|
|
|
def test_has_non_string_content_multimodal_messages():
|
|
data = {"messages": [{"role": "user", "content": [{"type": "text", "text": "hi"}]}]}
|
|
assert has_non_string_content(data) is True
|
|
|
|
|
|
def test_has_non_string_content_responses_api_string_input():
|
|
assert has_non_string_content({"input": "plain string"}) is False
|
|
|
|
|
|
def test_has_non_string_content_responses_api_list_input():
|
|
assert has_non_string_content({"input": ["a", "b"]}) is True
|
|
|
|
|
|
def test_has_non_string_content_empty_data():
|
|
assert has_non_string_content({}) is False
|
|
assert has_non_string_content({"messages": []}) is False
|
|
assert has_non_string_content({"input": ""}) is False
|
|
|
|
|
|
# ── apply_redacted_messages_back ──────────────────────────────────────────────
|
|
|
|
|
|
def test_apply_redacted_messages_back_chat_completion():
|
|
data = {"messages": [{"role": "user", "content": "secret"}]}
|
|
apply_redacted_messages_back(data, [{"role": "user", "content": "[REDACTED]"}])
|
|
assert data["messages"] == [{"role": "user", "content": "[REDACTED]"}]
|
|
assert "input" not in data
|
|
|
|
|
|
def test_apply_redacted_messages_back_responses_api_string_input():
|
|
"""A Responses-API request reads ``data["input"]``; writing only to
|
|
``messages`` would let unredacted text reach the LLM."""
|
|
data = {"input": "secret payload"}
|
|
apply_redacted_messages_back(data, [{"role": "user", "content": "[REDACTED]"}])
|
|
assert data["input"] == "[REDACTED]"
|
|
|
|
|
|
def test_apply_redacted_messages_back_both_fields():
|
|
"""Defensive: when both fields are present, both are updated."""
|
|
data = {
|
|
"messages": [{"role": "user", "content": "old"}],
|
|
"input": "old",
|
|
}
|
|
apply_redacted_messages_back(data, [{"role": "user", "content": "[REDACTED]"}])
|
|
assert data["messages"] == [{"role": "user", "content": "[REDACTED]"}]
|
|
assert data["input"] == "[REDACTED]"
|
|
|
|
|
|
def test_apply_redacted_messages_back_skips_input_when_not_string():
|
|
"""List ``input`` (multimodal Responses-API) is left alone — the
|
|
multimodal-degrades-to-block guard runs upstream."""
|
|
data = {"input": [{"type": "text", "text": "leak"}]}
|
|
apply_redacted_messages_back(data, [{"role": "user", "content": "[REDACTED]"}])
|
|
assert data["input"] == [{"type": "text", "text": "leak"}]
|