litellm/.github/e2e-stack
Yuneng Jiang 008c462a32
test(guardrails): cover the LiteLLM/Presidio integration contract
The presidio suites so far prove that masking happens. They do not prove that
the pieces we own still line up with what a real Presidio answers, and a mock
analyzer cannot show that: the mock resolves overlapping spans at analyze time,
returns recognition_metadata, and has no NER engine, so it agrees with our code
by construction.

Five cases against a live Presidio, each pinned to something LiteLLM does with
the response rather than to Presidio's own accuracy:

- the configured entity filter reaches the analyze request, so entity types a
  customer did not configure survive untouched
- every detection is replaced in place by a placeholder naming its own type,
  with the rest of the prompt byte for byte
- a PII entity configured BLOCK refuses the request, while clean text on the
  same guardrail is still served
- output_parse_pii numbers the placeholders from the analyze spans and restores
  the original values in the answer
- an analyzer the proxy cannot reach refuses the request rather than forwarding
  raw PII to the model

None of them assert a confidence score, so Presidio adding a recognizer or
changing its scoring cannot turn them red. The UI spec's two score assertions
went the same way: it now pins that a score is rendered per entity, not which
numbers the analyzer chose.

The analyzer and anonymizer endpoints are an internal deployment, so they are
handled like a provider API key. presidio_env is the only place that reads
them, both suites route their failure messages through its scrubber so a
response body that echoes an endpoint cannot publish it into a CI log, and both
presidio files stay off the GitHub Actions lane, which is what keeps the secret
out of GHA entirely.

Also reframes the two echo prompts as transcription rather than "repeat this
back". Asked to repeat placeholders, a model may refuse and explain itself
instead, which left the merged pre_call assertion measuring the model's mood.

Mutation-tested against the live stack: dropping the entities field, ignoring
the anonymized text, never raising on a blocked entity, skipping the unmask,
and swallowing the analyzer connection error each turn their own case red, 5 of
5 killed.
2026-09-07 17:38:56 -07:00
..
assert_tests_ran.py ci(e2e): run the access_control canary on harness changes and name failed tests 2026-09-05 18:46:51 -07:00
down.sh ci(e2e): run a PR's changed e2e tests three times behind a human-approved environment 2026-09-02 14:53:40 -07:00
secrets_to_env.py fix(e2e-changed): keep the gate off suites the stack cannot run 2026-09-05 21:03:50 -07:00
select_tests.py test(guardrails): cover the LiteLLM/Presidio integration contract 2026-09-07 17:38:56 -07:00
up.sh fix(e2e): wait for every gateway before using a new model and keep the network rerun 2026-09-05 16:10:40 -07:00