* feat(guardrails/headroom): add CCR (compress-cache-retrieve) support via agentic loop
When Headroom's /v1/compress returns messages containing hash markers
(hash=[a-f0-9]{24}), inject a headroom_retrieve tool into the request.
When the LLM calls that tool, intercept via async_should_run_agentic_loop
and async_build_agentic_loop_plan, call GET /v1/retrieve/{hash} on the
Headroom sidecar, and replay the LLM with the original content as a tool
result -- all transparent to the caller.
* style: run ruff format on headroom guardrail and tests
* fix(guardrails/headroom): detect headroom_retrieve calls in both OpenAI and Anthropic response formats
* test(guardrails/headroom): add test for Anthropic content block format detection in CCR loop
* ci: trigger CI checks
* fix(guardrails/headroom): replace List/Dict with list/dict to fix UP006 ruff violations
* fix(guardrails/headroom): replace except Exception with except ValueError to fix BLE001
* fix(guardrails/headroom): add Responses API output format detection for CCR tool calls
* refactor(guardrails/headroom): extract format-specific helpers to fix C901 complexity
* fix(guardrails/headroom): scope CCR retrieval to hashes produced by current request
Previously any LLM-supplied hash in a headroom_retrieve tool call was
forwarded to the Headroom retrieve API, letting a crafted tool call
fetch arbitrary cached content. Validate the hash against the set
produced by compressing the current request's messages before calling
retrieve.
* fix(guardrails/headroom): track issued hashes server-side, fix Responses API replay shape
Hash validation now also checks an in-memory cache of hashes actually
returned by /v1/compress, not just whether the hash text appears
somewhere in the request's messages. The message-text check alone is
forgeable: an attacker can plant a hash-shaped string in their own
prompt and have it treated as valid.
Responses API follow-up now emits function_call/function_call_output
items keyed by call_id instead of chat-style assistant/tool messages,
since the Responses API does not accept the latter as input. Also
fixes call_id/id field priority when extracting tool calls from
Responses API output, since call_id (not id) is what must match
between the function_call and its output.
* fix(guardrails/headroom): drop redundant quoted type annotations
UP037 flags quotes on annotations that are already lazily evaluated
via `from __future__ import annotations`.
* test(guardrails/headroom): add missing pytest.mark.asyncio decorators
Functional under asyncio_mode=auto, but every other async test in the
file has the decorator for consistency.
* fix(guardrails/headroom): scope CCR hashes per call_id, fix Anthropic replay shape
Two real gaps found in review:
1. The instance-wide issued-hash cache combined with a message-text
check did not actually scope retrieval to the request that produced
the hash. A hash issued for request A stays in the shared cache
until TTL expiry, and the message-text check is satisfied by any
request whose own messages happen to echo that hash string. Request
B could plant A's hash in its own prompt and retrieve A's content.
Fixed by keying the issued-hash cache by litellm_call_id, matching
the pattern already used in compression_interception: a hash is
only honored when it was issued under the exact call_id resolving
for the current request.
2. The Anthropic Messages replay path fell through to the chat-style
assistant/tool-message builder, which Anthropic does not accept.
Anthropic requires the tool_use block echoed in an assistant message
paired with a tool_result block in a user message, keyed by
tool_use_id. Added a dedicated branch for this shape.
* docs: note proactive API-fragmentation helper convention
Add a bullet to the coding-conventions list: look for or add a shared
helper when logic branches on API surface (chat completions vs
Anthropic Messages vs Responses API), instead of duplicating
format-detection per module.
* fix(guardrails/headroom): fix Anthropic tool-shape detection, extract shared cross-API tool util
Live e2e testing against the real Anthropic API surfaced two bugs the
mocked unit tests couldn't catch because they used MagicMock responses
instead of realistic response shapes:
1. has_headroom_retrieve_tool only recognized OpenAI-shaped function
tools. By the time an Anthropic Messages response reaches the
agentic-loop gate, the tool this guardrail injected has already been
transformed into Anthropic's native shape (type: "custom", top-level
"name"), so the gate never fired for real Anthropic requests.
2. AnthropicMessagesResponse is a TypedDict, so real responses are
plain dicts at runtime, not objects with attribute access. The
extractors and format detectors used bare getattr(), which silently
returns nothing for dict responses instead of reading the actual
key.
Extracted the cross-API-surface tool-call extraction and tool-presence
check into litellm/litellm_core_utils/prompt_templates/factory.py
(get_tool_calls_from_response, has_tool_with_name) so this format
fragmentation is handled in one place instead of being duplicated
per-guardrail, and reused the existing repair-aware
parse_tool_call_arguments from common_utils instead of a naive
json.loads. headroom.py now delegates to these shared helpers.
Confirmed live against the real Anthropic API: the retrieve loop now
fires and successfully retrieves the correct hash's content through
the full compress -> tool-call -> retrieve -> replay round-trip.
* fix(guardrails/headroom): fix ruff-strict UP006/I001 budget violations
Use lowercase list/dict generics in the new factory.py tool-call
helpers instead of typing.List/Dict, drop the now-unused Tuple import
in headroom.py, and reorder the new factory import ahead of the
llms.custom_httpx import to satisfy import sorting.
* fix(guardrails/headroom): match Anthropic tools without a type field
Anthropic's documented client tool format is just name + input_schema;
type: "custom" is only one possible value, not a requirement. Match
any non-OpenAI-shaped tool on its top-level name instead of requiring
type == "custom".
|
||
|---|---|---|
| .. | ||
| fixtures | ||
| realtime | ||
| reasoning_effort_grid | ||
| test-skill | ||
| test_llm_response_utils | ||
| test_skills_data | ||
| base_audio_transcription_unit_tests.py | ||
| base_embedding_unit_tests.py | ||
| base_llm_unit_tests.py | ||
| base_rerank_unit_tests.py | ||
| conftest.py | ||
| dog.wav | ||
| duck.png | ||
| gettysburg.wav | ||
| guinea.png | ||
| log.xt | ||
| Readme.md | ||
| test_a2a.py | ||
| test_anthropic_completion.py | ||
| test_aws_base_llm.py | ||
| test_azure_agents.py | ||
| test_azure_ai.py | ||
| test_azure_o_series.py | ||
| test_azure_openai.py | ||
| test_bedrock_agentcore.py | ||
| test_bedrock_agents.py | ||
| test_bedrock_anthropic_regression.py | ||
| test_bedrock_common_utils.py | ||
| test_bedrock_completion.py | ||
| test_bedrock_dynamic_auth_params_unit_tests.py | ||
| test_bedrock_embedding.py | ||
| test_bedrock_embedding_pricing.py | ||
| test_bedrock_govcloud.py | ||
| test_bedrock_gpt_oss.py | ||
| test_bedrock_invoke_tests.py | ||
| test_bedrock_llama.py | ||
| test_bedrock_mantle.py | ||
| test_bedrock_moonshot.py | ||
| test_bedrock_nova_embedding.py | ||
| test_bedrock_nova_json.py | ||
| test_cloudflare.py | ||
| test_cohere.py | ||
| test_containers_api.py | ||
| test_convert_dict_to_image.py | ||
| test_crusoe.py | ||
| test_databricks.py | ||
| test_deepgram.py | ||
| test_deepseek_completion.py | ||
| test_elevenlabs.py | ||
| test_evals_api.py | ||
| test_fireworks_ai_translation.py | ||
| test_gemini.py | ||
| test_gemini_image_usage.py | ||
| test_gigachat.py | ||
| test_gpt4o_audio.py | ||
| test_groq.py | ||
| test_hosted_vllm_embedding_e2e.py | ||
| test_huggingface_chat_completion.py | ||
| test_hyperbolic.py | ||
| test_infinity.py | ||
| test_jina_ai.py | ||
| test_lambda_ai.py | ||
| test_langgraph.py | ||
| test_litellm_proxy_provider.py | ||
| test_minimax_tts.py | ||
| test_mistral_api.py | ||
| test_model_cost_map_resilience.py | ||
| test_morph.py | ||
| test_nvidia_nim.py | ||
| test_openai.py | ||
| test_openai_o1.py | ||
| test_openai_record_replay_proxy.py | ||
| test_openrouter.py | ||
| test_optional_params.py | ||
| test_perplexity_reasoning.py | ||
| test_prompt_caching.py | ||
| test_prompt_factory.py | ||
| test_replicate.py | ||
| test_rerank.py | ||
| test_router_llm_translation_tests.py | ||
| test_sambanova_chat_transformation.py | ||
| test_skills_api.py | ||
| test_skills_e2e.py | ||
| test_snowflake.py | ||
| test_text_completion.py | ||
| test_text_completion_unit_tests.py | ||
| test_together_ai.py | ||
| test_triton.py | ||
| test_unit_test_bedrock_invoke.py | ||
| test_v0.py | ||
| test_vcr_classification.py | ||
| test_vcr_conftest_common_banner.py | ||
| test_vcr_filters.py | ||
| test_vcr_redis_persister.py | ||
| test_voyage_ai.py | ||
| test_watsonx.py | ||
| test_xai.py | ||
Unit tests for individual LLM providers.
Name of the test file is the name of the LLM provider - e.g. test_openai.py is for OpenAI.
Redis-backed VCR cache
Every test in this directory is auto-decorated with @pytest.mark.vcr (via
conftest.py). The first time a test runs we hit the live provider and
record the HTTP exchange into Redis under
litellm:vcr:cassette:<test_id>. Every subsequent run within 24h replays
from Redis without touching the network. The 24h TTL means each new day's
first run records again, so upstream API drift surfaces within a day.
The persister, header scrubbing, and 2xx-only filtering are defined in
tests/_vcr_redis_persister.py. Files that already use respx (which
patches the same httpx transport vcrpy does) are excluded from the
auto-marker — see _RESPX_CONFLICTING_FILES in conftest.py.
The same VCR cache is used by other test directories that exercise live
provider APIs. The reusable conftest plumbing lives in
tests/_vcr_conftest_common.py and is wired into:
tests/llm_translation/tests/llm_responses_api_testing/tests/audio_tests/tests/batches_tests/tests/guardrails_tests/tests/image_gen_tests/tests/litellm_utils_tests/tests/local_testing/(coverslocal_testing_part1,local_testing_part2,litellm_router_testing,litellm_assistants_api_testing,langfuse_logging_unit_tests)tests/logging_callback_tests/tests/pass_through_unit_tests/tests/router_unit_tests/tests/unified_google_tests/
Test directories that run LiteLLM proxy in Docker (e.g. build_and_test,
proxy_logging_guardrails_model_info_tests, proxy_store_model_in_db_tests)
are intentionally not included: VCR.py patches the in-process httpx
transport, so it cannot intercept the LLM calls that originate inside the
Docker container.
Required environment
CASSETTE_REDIS_URL — separate Redis instance from the application
Redis (REDIS_URL/REDIS_HOST) so test cassettes are not flushed by
proxy tests. Provider credentials (ANTHROPIC_API_KEY, OPENAI_API_KEY,
AWS_*, etc.) are needed only on cache-miss (the daily re-record), not
on replay.
Flushing the cache
When you want the next run to re-record immediately instead of waiting for the 24h TTL:
make test-llm-translation-flush-vcr-cache
Disabling VCR
Skip the cache entirely (every call goes live, no recording):
LITELLM_VCR_DISABLE=1 uv run pytest tests/llm_translation/test_<file>.py