mirror of
https://github.com/BerriAI/litellm.git
synced 2026-08-28 05:25:59 +00:00
* test(vcr): make Redis-backed cassettes replay deterministically across runs - Pin LITELLM_LOCAL_MODEL_COST_MAP=True in the shared VCR harness so the per-test importlib.reload(litellm) no longer fetches the model cost map from raw.githubusercontent.com. That live fetch was being recorded into cassettes; for tests that subsequently skip it was the only recorded episode, so the persister refused to save it (skipped tests don't persist) and the test re-recorded it live every run (MISS:NOT_PERSISTED). - Compare-time symmetric matcher tolerance for Google OAuth (ya29.*) tokens, observability/telemetry payloads, credential-exchange bodies, and volatile UUID/timestamp tokens, so existing cassettes select a recorded episode instead of growing past the 50-episode cap and re-recording live. - Don't record fire-and-forget telemetry (langfuse/arize/otel/...) into non-telemetry tests' cassettes. Several modules set litellm.success_callback at import time, so observability logging is globally enabled and an async flush from the background logging worker lands in an unrelated test's VCR window, saved as a spurious MISS:RECORDED (observed: a Langfuse batch from another completion landing on test_lowest_latency_routing_buffer). Such a request now passes through live (telemetry hosts aren't real-spend hosts); tests that actually assert on telemetry keep recording it. - Dedupe + cap the VCR diagnostic dump so the classification summary survives CircleCI's ~400KB step-output truncation. - Stabilize a non-deterministic rate-limit test body; mark AWS Secrets Manager lifecycle tests VCR-incompatible (uniquely-named secrets can't be replayed). - Mark test_router_text_completion_client VCR-incompatible: it fires 300 identical requests to verify async-client reuse, but vcrpy patches the HTTP transport so replay never exercises the real connection pool the test validates, and recording 300 near-identical episodes overflows the 50-episode cap (MISS:OVERFLOW every run). It hits a free mock endpoint. - Mark the Vertex AI MaaS Mistral OCR tests (vertex_ai/mistral-ocr-2505) VCR-incompatible: the MaaS model is not provisioned in the CI GCP project, so the live :rawPredict call fails and the test skips every run, leaving no cassette to record (MISS:NOT_PERSISTED every run). Sibling direct-Mistral and Azure OCR tests are unaffected and still replay from cache. * fix(tests/vcr): refresh cassette TTL on read so replayed cassettes don't expire The Redis VCR persister loaded cassettes with a plain GET, which does not touch the key's TTL. A cassette that is only ever replayed (HIT/NOOP, never re-recorded) therefore expired exactly 24h after its last *write*, no matter how often it was read. Whichever CI run happened to cross that boundary re-recorded the cassette live and surfaced a spurious VCR MISS on otherwise deterministic cassettes — the residual per-run flakiness floor (a different random subset of read-only cassettes expiring each run). Slide the expiry forward on every successful load (best-effort EXPIRE), so any cassette used at least once per TTL window stays alive indefinitely and the 2nd/3rd run of a day replays cleanly. * fix(tests/vcr): recover from spurious GET-None for existing cassette keys Under concurrent CI load, the persister's load GET was observed returning None for a cassette key that demonstrably existed on the (single, non- clustered) Redis master — an external monitor saw the key present with a healthy TTL at the same instant the in-process client read None. Because None is a valid GET result (not a RedisError), the retry-on-error client config never engaged, so the cassette re-recorded live (a phantom MISS:RECORDED); for flaky/networked tests the failed live call then triggered a pytest rerun, which is why a rotating subset of otherwise deterministic tests missed each run. On a None result, re-check EXISTS and re-read once. If the key really exists, use the recovered value and log [vcr-transient-miss-recovered] (also counted in cassette_cache_health). A genuinely absent key (a new cassette) still falls through to CassetteNotFoundError. * chore(tests/vcr): TEMP diagnostic for persistent-miss cassette load path Logs GET/EXISTS at load time for the three cassettes that re-record every run despite being present in Redis, to capture what the in-process client sees. To be reverted before merge. * chore(tests/vcr): write load diagnostic to Redis (truncation-proof) CI stdout truncates to the last ~400KB, dropping the early loaddbg lines for the alphabetically-first failing test. Push the load probe to a Redis list instead so it survives. To be reverted before merge. * fix(tests/vcr): don't drop stored telemetry episodes during cassette load Root cause of the residual per-run misses on present cassettes: vcrpy's Cassette._load() replays each *stored* interaction through Cassette.append(), which runs before_record_request on it — and a None return there silently drops that episode. The telemetry-leak suppressor (_should_drop_telemetry_record) returns None for telemetry requests, so when a non-telemetry-named test (or the alphabetically-first test in a worker, whose _current_test_nodeid is still empty) loaded a cassette containing a Langfuse ingestion episode, the episode was dropped on read — forcing an endless live re-record (a phantom MISS:RECORDED on a cassette that was demonstrably present in Redis). Verified by reproducing Cassette._load() against the real cassette: empty/non-telemetry nodeid -> 0 episodes survive; with the guard -> 1 survives. Fix: guard the suppressor with a thread-local set around Cassette._load (via a small idempotent monkeypatch), so the drop only ever stops *new* incidental telemetry from being recorded and never filters the existing cassette on read. Also drops the speculative GET-None recovery + its diagnostics from the previous commits: the load diagnostic showed GET returns the cassette bytes fine (get=1440B), so the persister never returned a spurious None — the loss happened later in vcrpy's append. The proven TTL-refresh-on-read fix is retained. * fix(tests/vcr): drop incidental telemetry export POSTs to stop rotating async-flush misses litellm's observability loggers flush on a background thread, so a Langfuse ingestion POST scheduled by one telemetry test can fire mid-way through a *later* telemetry-named test (after that test's own httpx mock has exited) and be recorded by VCR as a phantom episode — a non-deterministic MISS:RECORDED / PARTIAL that rotates onto a different telemetry test from run to run. Telemetry export POSTs are fire-and-forget; no test asserts on a *recorded* export response except the pass-through proxy test (which forwards a client POST to Langfuse ingestion and replays its 207). So _should_drop_telemetry_record now drops incidental export POSTs for every test except that one. Dropping returns None (live fire-and-forget, never stored), so it can only turn a phantom miss into a harmless live call, never the reverse; recorded read-back GETs that telemetry tests assert on are matched by method and left untouched. * fix(tests/vcr): restore assertion in test_banner_silent_when_vcr_disabled The assertion that the banner is suppressed when VCR is disabled was inadvertently moved into test_diagnostic_log_silent_when_no_dir when the diagnostic-log tests were added, leaving the disabled-VCR test verifying nothing. Co-authored-by: Yassin Kortam <yassin@berri.ai> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Yassin Kortam <yassin@berri.ai> |
||
|---|---|---|
| .. | ||
| .litellm_cache | ||
| auto_router | ||
| example_config_yaml | ||
| test_configs | ||
| test_model_response_typing | ||
| azure_fine_tune.jsonl | ||
| azure_speech.mp3 | ||
| batch_job_results_furniture.jsonl | ||
| cache_unit_tests.py | ||
| conftest.py | ||
| create_mock_standard_logging_payload.py | ||
| data_map.txt | ||
| eagle.wav | ||
| example.jsonl | ||
| gettysburg.wav | ||
| large_text.py | ||
| model_cost.json | ||
| openai_batch_completions.jsonl | ||
| openai_batch_completions_router.jsonl | ||
| speech_vertex.mp3 | ||
| stream_chunk_testdata.py | ||
| test_acompletion.py | ||
| test_acompletion_fallbacks.py | ||
| test_acooldowns_router.py | ||
| test_add_function_to_prompt.py | ||
| test_add_update_models.py | ||
| test_aim_guardrails.py | ||
| test_alangfuse.py | ||
| test_amazing_vertex_completion.py | ||
| test_anthropic_prompt_caching.py | ||
| test_arize_ai.py | ||
| test_arize_phoenix.py | ||
| test_assistants.py | ||
| test_async_fn.py | ||
| test_auth_utils.py | ||
| test_azure_anthropic_sync_post.py | ||
| test_azure_content_safety.py | ||
| test_azure_openai.py | ||
| test_azure_perf.py | ||
| test_basic_python_version.py | ||
| test_batch_completion_return_exceptions.py | ||
| test_batch_completions.py | ||
| test_blocked_user_list.py | ||
| test_braintrust.py | ||
| test_budget_manager.py | ||
| test_cache_preset_key.py | ||
| test_caching.py | ||
| test_caching_handler.py | ||
| test_caching_ssl.py | ||
| test_class.py | ||
| test_completion.py | ||
| test_completion_cost.py | ||
| test_completion_with_retries.py | ||
| test_config.py | ||
| test_cost_calc.py | ||
| test_custom_api_logger.py | ||
| test_custom_callback_input.py | ||
| test_custom_llm.py | ||
| test_custom_logger.py | ||
| test_disk_cache_unit_tests.py | ||
| test_docker_no_network_on_deploy.py | ||
| test_dual_cache.py | ||
| test_dynamic_rate_limit_handler.py | ||
| test_dynamodb_logs.py | ||
| test_embedding.py | ||
| test_exceptions.py | ||
| test_file_types.py | ||
| test_function_call_parsing.py | ||
| test_function_calling.py | ||
| test_function_setup.py | ||
| test_gcs_bucket.py | ||
| test_gcs_cache_unit_tests.py | ||
| test_gemini_reasoning_content.py | ||
| test_get_llm_provider.py | ||
| test_get_model_file.py | ||
| test_get_model_info.py | ||
| test_get_optional_params_embeddings.py | ||
| test_get_optional_params_functions_not_supported.py | ||
| test_google_ai_studio_gemini.py | ||
| test_guardrails_ai.py | ||
| test_helicone_integration.py | ||
| test_http_parsing_utils.py | ||
| test_img_resize.py | ||
| test_lakera_ai_prompt_injection.py | ||
| test_langchain_ChatLiteLLM.py | ||
| test_langsmith.py | ||
| test_least_busy_routing.py | ||
| test_litellm_max_budget.py | ||
| test_llm_guard.py | ||
| test_load_test_router_s3.py | ||
| test_loadtest_router.py | ||
| test_logfire.py | ||
| test_logging.py | ||
| test_longer_context_fallback.py | ||
| test_lowest_cost_routing.py | ||
| test_lowest_latency_routing.py | ||
| test_lunary.py | ||
| test_max_tpm_rpm_limiter.py | ||
| test_mem_leak.py | ||
| test_mem_usage.py | ||
| test_mock_request.py | ||
| test_model_alias_map.py | ||
| test_model_max_token_adjust.py | ||
| test_multiple_deployments.py | ||
| test_ollama.py | ||
| test_ollama_local.py | ||
| test_ollama_local_chat.py | ||
| test_openai_moderations_hook.py | ||
| test_opik.py | ||
| test_pass_through_endpoints.py | ||
| test_profiling_router.py | ||
| test_prometheus_service.py | ||
| test_prompt_caching.py | ||
| test_prompt_injection_detection.py | ||
| test_promptlayer_integration.py | ||
| test_provider_specific_config.py | ||
| test_pydantic.py | ||
| test_pydantic_namespaces.py | ||
| test_redis_batch_optimizations.py | ||
| test_register_model.py | ||
| test_responses_stream_cache_keys.py | ||
| test_router.py | ||
| test_router_auto_router.py | ||
| test_router_batch_completion.py | ||
| test_router_budget_limiter.py | ||
| test_router_caching.py | ||
| test_router_client_init.py | ||
| test_router_cooldown_handlers.py | ||
| test_router_custom_routing.py | ||
| test_router_debug_logs.py | ||
| test_router_fallback_handlers.py | ||
| test_router_fallbacks.py | ||
| test_router_get_deployments.py | ||
| test_router_max_parallel_requests.py | ||
| test_router_pattern_matching.py | ||
| test_router_retries.py | ||
| test_router_timeout.py | ||
| test_router_utils.py | ||
| test_router_with_fallbacks.py | ||
| test_rules.py | ||
| test_sagemaker.py | ||
| test_sagemaker_nova_integration.py | ||
| test_scheduler.py | ||
| test_secret_detect_hook.py | ||
| test_spend_calculate_endpoint.py | ||
| test_stream_chunk_builder.py | ||
| test_streaming.py | ||
| test_supabase_integration.py | ||
| test_team_config.py | ||
| test_text_completion.py | ||
| test_timeout.py | ||
| test_together_ai.py | ||
| test_tpm_rpm_routing_v2.py | ||
| test_traceloop.py | ||
| test_ui_sso_helper_utils.py | ||
| test_unit_test_caching.py | ||
| test_update_spend.py | ||
| test_validate_environment.py | ||
| test_wandb.py | ||
| user_cost.json | ||
| vertex_ai.jsonl | ||
| vertex_batch_completions.jsonl | ||
| vertex_key.json | ||
| whitelisted_bedrock_models.txt | ||