test_text_message_blocked_by_guardrail_no_ai_response classified the model's reply against a safe_markers keyword list to decide whether the guardrail had blocked the message. gpt-realtime words its refusal of the guardrail's "say exactly" voice prompt nondeterministically, so any new phrasing outside the list turned CI red on unrelated PRs; the list had already been extended in #28191, #28200 and #29477, and drifted again to "Sorry, I can't comply with that request" (11 of the 13 failed realtime_translation_testing runs since 2026-06-24, e.g. CircleCI job 2009316 on #32380). Record every frame the proxy sends to the backend through a RecordingBackendWebSocket wrapper and assert the invariant the product actually guarantees: the blocked phrase never reaches OpenAI, only the guardrail's own conversation.item.create and response.create are forwarded (the client's reflexive response.create is dropped), and the blocked phrase never appears in AI output. Replace the fixed 0.3s/3.0s sleeps with an event-driven wait for response.done; client frames are processed sequentially so no inter-message sleep is needed. Verified by mutation: disabling the response.create drop fails the response.create count assertion, and disabling the guardrail fails the guardrail_violation assertion. |
||
|---|---|---|
| .. | ||
| fixtures | ||
| realtime | ||
| reasoning_effort_grid | ||
| test-skill | ||
| test_llm_response_utils | ||
| test_skills_data | ||
| base_audio_transcription_unit_tests.py | ||
| base_embedding_unit_tests.py | ||
| base_llm_unit_tests.py | ||
| base_rerank_unit_tests.py | ||
| conftest.py | ||
| dog.wav | ||
| duck.png | ||
| gettysburg.wav | ||
| guinea.png | ||
| log.xt | ||
| Readme.md | ||
| test_a2a.py | ||
| test_anthropic_completion.py | ||
| test_aws_base_llm.py | ||
| test_azure_agents.py | ||
| test_azure_ai.py | ||
| test_azure_o_series.py | ||
| test_azure_openai.py | ||
| test_bedrock_agentcore.py | ||
| test_bedrock_agents.py | ||
| test_bedrock_anthropic_regression.py | ||
| test_bedrock_common_utils.py | ||
| test_bedrock_completion.py | ||
| test_bedrock_dynamic_auth_params_unit_tests.py | ||
| test_bedrock_embedding.py | ||
| test_bedrock_embedding_pricing.py | ||
| test_bedrock_govcloud.py | ||
| test_bedrock_gpt_oss.py | ||
| test_bedrock_invoke_tests.py | ||
| test_bedrock_llama.py | ||
| test_bedrock_mantle.py | ||
| test_bedrock_moonshot.py | ||
| test_bedrock_nova_embedding.py | ||
| test_bedrock_nova_json.py | ||
| test_cloudflare.py | ||
| test_cohere.py | ||
| test_containers_api.py | ||
| test_convert_dict_to_image.py | ||
| test_crusoe.py | ||
| test_databricks.py | ||
| test_deepgram.py | ||
| test_deepseek_completion.py | ||
| test_elevenlabs.py | ||
| test_evals_api.py | ||
| test_fireworks_ai_translation.py | ||
| test_gemini.py | ||
| test_gemini_image_usage.py | ||
| test_gigachat.py | ||
| test_gpt4o_audio.py | ||
| test_groq.py | ||
| test_hosted_vllm_embedding_e2e.py | ||
| test_huggingface_chat_completion.py | ||
| test_hyperbolic.py | ||
| test_infinity.py | ||
| test_jina_ai.py | ||
| test_lambda_ai.py | ||
| test_langgraph.py | ||
| test_litellm_proxy_provider.py | ||
| test_minimax_tts.py | ||
| test_mistral_api.py | ||
| test_model_cost_map_resilience.py | ||
| test_morph.py | ||
| test_nvidia_nim.py | ||
| test_openai.py | ||
| test_openai_o1.py | ||
| test_openai_record_replay_proxy.py | ||
| test_openrouter.py | ||
| test_optional_params.py | ||
| test_perplexity_reasoning.py | ||
| test_prompt_caching.py | ||
| test_prompt_factory.py | ||
| test_replicate.py | ||
| test_rerank.py | ||
| test_router_llm_translation_tests.py | ||
| test_sambanova_chat_transformation.py | ||
| test_skills_api.py | ||
| test_skills_e2e.py | ||
| test_snowflake.py | ||
| test_text_completion.py | ||
| test_text_completion_unit_tests.py | ||
| test_together_ai.py | ||
| test_triton.py | ||
| test_unit_test_bedrock_invoke.py | ||
| test_v0.py | ||
| test_vcr_classification.py | ||
| test_vcr_conftest_common_banner.py | ||
| test_vcr_filters.py | ||
| test_vcr_redis_persister.py | ||
| test_voyage_ai.py | ||
| test_watsonx.py | ||
| test_xai.py | ||
Unit tests for individual LLM providers.
Name of the test file is the name of the LLM provider - e.g. test_openai.py is for OpenAI.
Redis-backed VCR cache
Every test in this directory is auto-decorated with @pytest.mark.vcr (via
conftest.py). The first time a test runs we hit the live provider and
record the HTTP exchange into Redis under
litellm:vcr:cassette:<test_id>. Every subsequent run within 24h replays
from Redis without touching the network. The 24h TTL means each new day's
first run records again, so upstream API drift surfaces within a day.
The persister, header scrubbing, and 2xx-only filtering are defined in
tests/_vcr_redis_persister.py. Files that already use respx (which
patches the same httpx transport vcrpy does) are excluded from the
auto-marker — see _RESPX_CONFLICTING_FILES in conftest.py.
The same VCR cache is used by other test directories that exercise live
provider APIs. The reusable conftest plumbing lives in
tests/_vcr_conftest_common.py and is wired into:
tests/llm_translation/tests/llm_responses_api_testing/tests/audio_tests/tests/batches_tests/tests/guardrails_tests/tests/image_gen_tests/tests/litellm_utils_tests/tests/local_testing/(coverslocal_testing_part1,local_testing_part2,litellm_router_testing,litellm_assistants_api_testing,langfuse_logging_unit_tests)tests/logging_callback_tests/tests/pass_through_unit_tests/tests/router_unit_tests/tests/unified_google_tests/
Test directories that run LiteLLM proxy in Docker (e.g. build_and_test,
proxy_logging_guardrails_model_info_tests, proxy_store_model_in_db_tests)
are intentionally not included: VCR.py patches the in-process httpx
transport, so it cannot intercept the LLM calls that originate inside the
Docker container.
Required environment
CASSETTE_REDIS_URL — separate Redis instance from the application
Redis (REDIS_URL/REDIS_HOST) so test cassettes are not flushed by
proxy tests. Provider credentials (ANTHROPIC_API_KEY, OPENAI_API_KEY,
AWS_*, etc.) are needed only on cache-miss (the daily re-record), not
on replay.
Flushing the cache
When you want the next run to re-record immediately instead of waiting for the 24h TTL:
make test-llm-translation-flush-vcr-cache
Disabling VCR
Skip the cache entirely (every call goes live, no recording):
LITELLM_VCR_DISABLE=1 uv run pytest tests/llm_translation/test_<file>.py