mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-14 23:21:35 +00:00
Adds live coverage for three guardrails litellm implements itself, each driven against a real chat completion and asserting on the permitted/denied tool or the judge's verdict rather than a bare status code. tool_permission is registered allow-list style (one allow rule, default_action=deny, on_disallowed_action=block): an unlisted tool is rejected pre-call with the denied tool named, and the listed tool survives the guardrail and is called by the model. tool_policy is exercised through a key-scoped blocked-tool override set with POST /v1/tool/policy: the blocked tool is rejected pre-call while a sibling tool on the same key and the same guardrail still reaches the model, so the block is attributable to the policy rather than to the guardrail refusing all tool use. llm_as_a_judge scores the model's answer against a French-only criterion; an English answer is rejected with 422 and the failing verdict, and a French one comes back to the caller. The judge test does not stop at the response. That guardrail fails open on any internal error, and a fail-open returns an ordinary 200 that is identical to an approval: same status, same body, and the applied-guardrails header still names it. So both halves also read the guardrail's own run log at /guardrails/usage/logs, where an approval is recorded as `passed`, an intervention as `blocked`, and a fail-open as `flagged`. Without that leg the accept half would pass just as happily against a build where adjudication never ran. Also corrects the llm_as_a_judge registry row. It asked for a pre_call block, which the guardrail cannot do: it supports post_call only and the proxy rejects registering it at pre_call outright, so the row could never go green as written. Retargeted to post_call, which is what the guardrail actually enforces. The judge runs on openai/gpt-4.1 rather than the suite's usual gpt-5.5 because the guardrail hardcodes temperature=0 on its judge call, gpt-5.5 accepts only the default temperature, and the resulting error is swallowed into a fail-open, so gpt-5.5 can never adjudicate anything. Covers guardrail.tool_permission.pre_call.blocks, guardrail.tool_permission.pre_call.allows, guardrail.tool_policy.pre_call.blocks and guardrail.llm_as_a_judge.post_call.blocks. The guardrail params, the tool-policy override bodies and the chat helper live in a suite-local module so the shared harness is untouched. |
||
|---|---|---|
| .. | ||
| a2a | ||
| access_control | ||
| batches | ||
| claude_code | ||
| coverage_registry | ||
| guardrails | ||
| llm_translation | ||
| load | ||
| logging | ||
| management | ||
| mcp | ||
| other | ||
| quota_management | ||
| router | ||
| ui | ||
| CLAUDE.md | ||
| conftest.py | ||
| CONTRIBUTING.md | ||
| e2e_config.py | ||
| e2e_db.py | ||
| e2e_http.py | ||
| junit_properties.py | ||
| lifecycle.py | ||
| models.py | ||
| otel_client.py | ||
| proxy_client.py | ||
| pytest.ini | ||
| transport.py | ||