* fix(streaming): word-sliced cache replay for stream=true cache hits * fix(streaming): align mypy and replay happy-path test with word-sliced cache replay * fix(streaming): short-circuit whitespace-only content in cache replay splitter * fix(streaming): emit tool_calls/function_call only on first replay slice * refactor(streaming): drop dead delattr guard in cache replay A non-None usage on the replay base object always lives in __pydantic_extra__ (it is attached via setattr earlier in the same function), so delattr can never raise here; the try/except AttributeError that silently swallowed a failure was dead defensive code that could only ever hide a real regression, so it is removed in both the async and sync generators. Also switches the new replay annotations from typing.List to the builtin list to satisfy the strict ruff UP006 gate and drops the unused PLR0915 noqa directives (the rule is not enabled in this repo's ruff config, so RUF100 flagged them). * fix(streaming): drop carried-over metadata from later cache replay slices The word-sliced cache replay deep-copies the full ModelResponseStream per slice, so reasoning_content, thinking_blocks, logprobs, enhancements, annotations and the rest of the per-message metadata rode on every slice, not just the first. Downstream handlers that accumulate streamed deltas would collect each one once per slice, e.g. duplicating a cached reasoning trace N times on a stream=true cache hit. Later slices are now rebuilt as a content-only delta with choice-level logprobs and enhancements stripped, so the whole metadata class stays on the first slice. Adds async (logprobs) and sync (reasoning_content/thinking_blocks/logprobs/ enhancements, plus annotations) regression tests --------- Co-authored-by: Mateo <277851410+mateo-berri@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| fixtures | ||
| realtime | ||
| reasoning_effort_grid | ||
| test-skill | ||
| test_llm_response_utils | ||
| test_skills_data | ||
| base_audio_transcription_unit_tests.py | ||
| base_embedding_unit_tests.py | ||
| base_llm_unit_tests.py | ||
| base_rerank_unit_tests.py | ||
| conftest.py | ||
| dog.wav | ||
| duck.png | ||
| gettysburg.wav | ||
| guinea.png | ||
| log.xt | ||
| Readme.md | ||
| test_a2a.py | ||
| test_anthropic_completion.py | ||
| test_aws_base_llm.py | ||
| test_azure_agents.py | ||
| test_azure_ai.py | ||
| test_azure_o_series.py | ||
| test_azure_openai.py | ||
| test_bedrock_agentcore.py | ||
| test_bedrock_agents.py | ||
| test_bedrock_anthropic_regression.py | ||
| test_bedrock_common_utils.py | ||
| test_bedrock_completion.py | ||
| test_bedrock_dynamic_auth_params_unit_tests.py | ||
| test_bedrock_embedding.py | ||
| test_bedrock_embedding_pricing.py | ||
| test_bedrock_govcloud.py | ||
| test_bedrock_gpt_oss.py | ||
| test_bedrock_invoke_tests.py | ||
| test_bedrock_llama.py | ||
| test_bedrock_mantle.py | ||
| test_bedrock_moonshot.py | ||
| test_bedrock_nova_embedding.py | ||
| test_bedrock_nova_json.py | ||
| test_cloudflare.py | ||
| test_cohere.py | ||
| test_containers_api.py | ||
| test_convert_dict_to_image.py | ||
| test_crusoe.py | ||
| test_databricks.py | ||
| test_deepgram.py | ||
| test_deepseek_completion.py | ||
| test_elevenlabs.py | ||
| test_evals_api.py | ||
| test_fireworks_ai_translation.py | ||
| test_gemini.py | ||
| test_gemini_image_usage.py | ||
| test_gigachat.py | ||
| test_gpt4o_audio.py | ||
| test_groq.py | ||
| test_hosted_vllm_embedding_e2e.py | ||
| test_huggingface_chat_completion.py | ||
| test_hyperbolic.py | ||
| test_infinity.py | ||
| test_jina_ai.py | ||
| test_lambda_ai.py | ||
| test_langgraph.py | ||
| test_litellm_proxy_provider.py | ||
| test_minimax_tts.py | ||
| test_mistral_api.py | ||
| test_model_cost_map_resilience.py | ||
| test_morph.py | ||
| test_nvidia_nim.py | ||
| test_openai.py | ||
| test_openai_o1.py | ||
| test_openai_record_replay_proxy.py | ||
| test_openrouter.py | ||
| test_optional_params.py | ||
| test_perplexity_reasoning.py | ||
| test_prompt_caching.py | ||
| test_prompt_factory.py | ||
| test_replicate.py | ||
| test_rerank.py | ||
| test_router_llm_translation_tests.py | ||
| test_sambanova_chat_transformation.py | ||
| test_skills_api.py | ||
| test_skills_e2e.py | ||
| test_snowflake.py | ||
| test_text_completion.py | ||
| test_text_completion_unit_tests.py | ||
| test_together_ai.py | ||
| test_triton.py | ||
| test_unit_test_bedrock_invoke.py | ||
| test_v0.py | ||
| test_vcr_classification.py | ||
| test_vcr_conftest_common_banner.py | ||
| test_vcr_filters.py | ||
| test_vcr_redis_persister.py | ||
| test_voyage_ai.py | ||
| test_watsonx.py | ||
| test_xai.py | ||
Unit tests for individual LLM providers.
Name of the test file is the name of the LLM provider - e.g. test_openai.py is for OpenAI.
Redis-backed VCR cache
Every test in this directory is auto-decorated with @pytest.mark.vcr (via
conftest.py). The first time a test runs we hit the live provider and
record the HTTP exchange into Redis under
litellm:vcr:cassette:<test_id>. Every subsequent run within 24h replays
from Redis without touching the network. The 24h TTL means each new day's
first run records again, so upstream API drift surfaces within a day.
The persister, header scrubbing, and 2xx-only filtering are defined in
tests/_vcr_redis_persister.py. Files that already use respx (which
patches the same httpx transport vcrpy does) are excluded from the
auto-marker — see _RESPX_CONFLICTING_FILES in conftest.py.
The same VCR cache is used by other test directories that exercise live
provider APIs. The reusable conftest plumbing lives in
tests/_vcr_conftest_common.py and is wired into:
tests/llm_translation/tests/llm_responses_api_testing/tests/audio_tests/tests/batches_tests/tests/guardrails_tests/tests/image_gen_tests/tests/litellm_utils_tests/tests/local_testing/(coverslocal_testing_part1,local_testing_part2,litellm_router_testing,litellm_assistants_api_testing,langfuse_logging_unit_tests)tests/logging_callback_tests/tests/pass_through_unit_tests/tests/router_unit_tests/tests/unified_google_tests/
Test directories that run LiteLLM proxy in Docker (e.g. build_and_test,
proxy_logging_guardrails_model_info_tests, proxy_store_model_in_db_tests)
are intentionally not included: VCR.py patches the in-process httpx
transport, so it cannot intercept the LLM calls that originate inside the
Docker container.
Required environment
CASSETTE_REDIS_URL — separate Redis instance from the application
Redis (REDIS_URL/REDIS_HOST) so test cassettes are not flushed by
proxy tests. Provider credentials (ANTHROPIC_API_KEY, OPENAI_API_KEY,
AWS_*, etc.) are needed only on cache-miss (the daily re-record), not
on replay.
Flushing the cache
When you want the next run to re-record immediately instead of waiting for the 24h TTL:
make test-llm-translation-flush-vcr-cache
Disabling VCR
Skip the cache entirely (every call goes live, no recording):
LITELLM_VCR_DISABLE=1 uv run pytest tests/llm_translation/test_<file>.py