litellm/tests/test_litellm/llms/anthropic
yucheng-berri 587b8aca9b
feat(guardrails): add Compresr guardrail for query-aware context compression (#33295)
* feat(guardrails): add Compresr guardrail for query-aware context compression

Adds a first-class guardrail that compresses bulky message content (tool
outputs, RAG chunks, search results) through the Compresr API before the
request reaches the LLM, via the apply_guardrail / structured_messages hook
so it covers /chat/completions, /v1/messages, and /v1/responses (the latter
through the texts channel, mirrored only when the replacement is
unambiguous; anything ambiguous is left uncompressed).

Distinct from whole-conversation compressors:
- Query-aware: each message is compressed against the intent that produced
  it (a tool output against its originating tool call's name + arguments,
  resolved via tool_call_id; otherwise the last user message).
- Recoverable: each compressed message carries a hash marker and the request
  gains a compresr_retrieve tool, so the model can pull the original content
  back through the agentic loop when the compressed version is not enough.
  Originals are cached in-process, scoped to the caller's virtual-key hash
  plus the request's litellm_call_id, with a TTL and a per-call byte cap;
  recovery is skipped when no caller scope is available so one caller can
  never read another's originals. The store is per-process, so multi-worker
  deployments need sticky routing (or enable_retrieval=false).

Fail-closed by default (fail_open configurable), SSRF-validated api_base
(alternate IP-literal encodings included), cross-tenant-isolated recovery
store, and upstream errors redacted from client-facing responses. The
outbound client follows redirects and re-resolves DNS per request, so the
api_base host/IP checks are defense-in-depth, not a full SSRF guarantee;
this is documented as a known limitation. Requests where nothing was
actually compressed are returned untouched (same object identity) so
handlers skip the write-back. Auto-discovered via the guardrail_hooks
registry.

* fix(guardrails): cap Compresr recovery store total memory

The recovery store bounded bytes per call and entry count, but had no
aggregate cap: 256 tracked call ids at the 10 MiB per-call default could
retain ~2.5 GiB per worker. A flood of requests with distinct
x-litellm-call-id values and large compressible tool outputs could
exhaust a shared proxy worker.

Add a global byte budget (_MAX_TOTAL_STORE_BYTES, 256 MiB) across all
entries. A running total is maintained on every insert/eviction so the
cap is enforced without re-encoding the whole store on the request path;
oldest entries are evicted once the budget is exceeded, always keeping
the most-recent entry so recovery still works for the request populating
the store. +2 regression tests.

* fix(guardrails): gate and bound Compresr recovery loop

Two hardening fixes to the compresr_retrieve agentic loop:

1. Only run the loop when a retrieve call resolves to recovery state this
   guardrail actually created for the request. Previously the gate checked
   only that the caller-supplied tool list contained a compresr_retrieve
   function and that the model emitted a call, so a caller could define
   their own same-named tool and force an extra provider round-trip with
   nothing to recover. The plan now returns run_agentic_loop=False when no
   requested hash resolves.

2. Bound the follow-up against retrieval amplification: each distinct hash
   is expanded at most once (repeats get a short marker) and at most
   _MAX_RETRIEVALS_PER_LOOP calls are honored, so prompting the model to
   call compresr_retrieve many times with the same marker cannot balloon
   the follow-up. _retrieve_original now returns None on miss.

+3 regression tests; two existing security tests updated to assert the
stronger veto behavior (forged/cross-tenant hashes now stop the loop
entirely instead of returning a not-found follow-up).

* fix(guardrails): warn when Compresr recovery is skipped without auth scope

When enable_retrieval is on (the default) but the proxy has no per-key
auth, the request has no caller scope, so recovery is silently disabled:
content is compressed but the compresr_retrieve tool is never injected and
the originals are dropped, with no runtime indication. Emit a one-shot
call-time warning so operators can see recovery is being suppressed and
configure virtual-key auth. +1 regression test.

* style(guardrails): tighten Compresr guardrail comments

Condense the verbose multi-line inline comments and the api_base docstring
to concise form. No behavior change.

* fix(guardrails): keep injected tool on Responses API + bound recovery markers by byte cap

Two fixes for reviewer-flagged defects in the Compresr guardrail:

- Responses API: _merge_tools_after_guardrail iterated only over the
  request's original tools, dropping any tool a guardrail appended (the
  compresr_retrieve recovery tool) whenever the request already had tools.
  Keep the appended tools so recovery works on /v1/responses.

- Recovery markers: markers + originals were built for every compressed
  target before the per-call byte cap trimmed the store, so an evicted
  original left a marker the model could never retrieve. Attach recovery
  only while the store (existing entries under the same key + this call's
  originals) stays within the cap, so a shipped marker is always retrievable
  -- including on a later turn that reuses the store key.

Adds regression tests for both paths.

* refactor(guardrails): extract _existing_originals to keep apply_guardrail under the complexity gate

The byte-cap fix added a branch to apply_guardrail, tipping it past the
C901 complexity ceiling. Move the store lookup into a small helper; no
behavior change.

* fix(guardrails): harden Compresr SSRF blocklist, re-arm no-scope warning, tolerate odd tool shapes

* fix(guardrails): rerun input guardrails on Compresr retrieval follow-up

* chore: remove unrelated deepkeep files committed by mistake

---------

Co-authored-by: charafkamel <charafkamel@live.com>
2026-07-15 13:53:41 -07:00
..
batches test(batches): add 1:1 test file scaffold for batches component paths (#30529) 2026-06-29 09:22:58 +05:30
chat feat(guardrails): add Compresr guardrail for query-aware context compression (#33295) 2026-07-15 13:53:41 -07:00
experimental_pass_through fix(anthropic-adapter): drop empty content_block_delta events 2026-07-14 18:16:20 -07:00
files fix(batches): price anthropic passthrough message batches correctly in batch cost job (#32307) 2026-07-06 20:33:57 -07:00
messages fix(anthropic): require caller api_key and SSRF-validate api_base in advisor tool (#32093) 2026-07-04 12:06:09 -07:00
__init__.py test(batches): add 1:1 test file scaffold for batches component paths (#30529) 2026-06-29 09:22:58 +05:30
test_anthropic_common_utils.py fix(anthropic): thread real provider through capability probes instead of pinning anthropic 2026-07-11 12:15:59 -07:00
test_anthropic_count_tokens_transformation.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_anthropic_files_and_batches.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_anthropic_structured_output.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_azure_ai_cache_pricing.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_cost_calculation_dict_safety.py fix: completion_cost AttributeError on streaming Anthropic web_search responses (#26153) (#27346) 2026-06-10 21:20:11 -07:00
test_count_tokens_oauth.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_message_sanitization.py fix(anthropic, mcp): sanitize tool names to match Anthropic's [a-zA-Z0-9_-]{1,128} pattern (#26788) 2026-05-06 00:00:36 +00:00