litellm/tests/test_litellm/proxy/guardrails
devin-ai-integration[bot] 264b09ac8d
fix(responses): scan and mask top-level instructions with guardrails (#43629)
* fix(responses): scan and mask top-level instructions with guardrails

The Responses guardrail translation handler put a non-empty top-level instructions field into structured_messages as a system row but never into the flat texts list, so guardrails that scan texts skipped it, flat-text masking could not rewrite it, and PANW latest-only selection failed its alignment guard whenever instructions were present.

Seed texts with the instructions row, carry that offset into the flat-text write-back so a rewritten row lands on data["instructions"], and account for the leading row in the PANW Responses alignment.

Resolves LIT-8931

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): reject empty guardrail rewrites instead of forwarding raw input

An explicit texts=[] answer from a guardrail now fails the count check and
raises UnappliableRequestRewrite like any other misaligned rewrite; only a
missing texts key means no rewrite. Types the out-param as dict[str, object]
and adds integration coverage for instructions blocking, masking, empty
instructions, tool loops, latest-only and concurrent workers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the texts-replacing guardrail helper explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): honor skip_system_message_in_guardrail for instructions and system input items

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): cover skip_system_message_in_guardrail on the live proxy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): keep skipped rows through full-coverage rewrites and align latest-only with skip_system

Trust a guardrail's structured_messages_cover_full_request claim only when it
returns as many rows as the full normalized request, otherwise merge the scoped
rows back so skipped instructions and system items survive the write-back.
Make PANW's Responses reasoning alignment skip-aware so latest-only still picks
the latest user turn when system content is excluded from texts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): annotate new guardrail tests with return types

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(responses): treat an empty guardrail texts answer as no rewrite like chat completions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(responses): type the guardrail test doubles explicitly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: yucheng <yucheng@berri.ai>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 11:44:35 -07:00
..
guardrail_hooks fix(responses): scan and mask top-level instructions with guardrails (#43629) 2026-09-30 11:44:35 -07:00
test_auto_router_compression.py fix(router): resolve team-scoped auto-routers by their public name (#40432) 2026-09-09 15:53:05 -07:00
test_content_filter_path_traversal.py Litellm OSS Staging 010626 (#29422) 2026-06-01 21:42:51 -07:00
test_content_utils.py feat(guardrails): add inspect_embeddings toggle for AIM and Cato (#39918) 2026-09-05 17:15:46 -07:00
test_custom_code_security.py fix(guardrails): block private destinations in custom code http_request and bound guardrail execution time (#43280) 2026-09-26 12:57:41 -07:00
test_deferred_guardrail_logging.py fix(cost): resolve a missing 1h cache write rate after off-peak pricing 2026-09-14 21:46:33 -07:00
test_guardrail_coverage.py fix(enterprise): resolve openai_moderations model at call time and default to omni-moderation-latest 2026-09-18 23:06:47 +00:00
test_guardrail_endpoints.py fix(guardrails): block private destinations in custom code http_request and bound guardrail execution time (#43280) 2026-09-26 12:57:41 -07:00
test_guardrail_registry.py fix(guardrails): preserve Presidio output selection and restoration (#43401) 2026-09-28 17:34:38 -07:00
test_init_guardrails.py fix(guardrails): enable explicit PANW MCP output scanning (#43109) 2026-09-29 12:50:08 -07:00
test_llm_as_a_judge.py fix(guardrails): only honor the judge call-origin stamp on logging_only in llm_as_a_judge 2026-09-16 01:05:53 +00:00
test_mcp_jwt_signer.py feat(mcp): scan and pin upstream tool descriptions (#43283) 2026-09-28 18:38:49 -07:00
test_pillar_guardrails.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_prompt_security_guardrails.py fix(guardrails): stream Prompt Security post_call redactions in incremental_diff mode 2026-09-17 01:06:12 +00:00
test_qostodian_nexus_guardrail.py test: enforce F811 so a duplicate definition cannot silently replace the first 2026-08-21 12:06:19 -07:00
test_usage_endpoints.py Revert "refactor(guardrails): rename scoped-out evaluation status from not_run to skipped" 2026-09-14 23:46:09 +00:00
test_usage_tracking.py Revert "refactor(guardrails): rename scoped-out evaluation status from not_run to skipped" 2026-09-14 23:46:09 +00:00