litellm/tests/code_coverage_tests
devin-ai-integration[bot] 68d9b8bbb8
fix(router): retry a /v1/messages stream the provider drops before the first content chunk (#44276)
* fix(router): retry a /v1/messages stream the provider drops before the first content chunk

A /v1/messages stream that the upstream closed before any content reached the
client answered an error event after a single attempt, so the router's
num_retries never applied to that drop. The pre-content failure is now retried
within the model group before the fallback chain runs, with the budget resolved
the way a failure raised before the stream opened resolves it: a retry policy
that names the error class, then the request's num_retries, then the
deployment's, then the router's. A drop after content reached the client keeps
surfacing the provider's error after one attempt.

Fixes #44238

* fix(router): hand a retry's non-retriable error to the fallback chain and type the retry helpers

A retry that failed before its stream opened with an error no retry covers raised straight to the
client, skipping a fallback the first attempt would have used. assert_never now comes from
typing_extensions so the router imports on Python 3.10, and the retry helpers read their kwargs
through typed narrowing instead of Mapping[str, Any]

* fix(router): cast the untyped router fallback defaults the stream retry gate reads

The retry gate passed the router's fallback attributes, declared without element types, to the
typed request override helper, which basedpyright counted as new unknown-argument errors

* fix(router): consult context_window_fallbacks when a retried /v1/messages stream overflows

A retry attempt raising ContextWindowExceededError reached the fallback chain inside its
mid-stream envelope, so only the regular fallbacks list matched. The fallback attempt now
unwraps it the way it unwraps a content policy error. The new router helpers are covered for
the router code coverage check with two direct-call tests and named covering tests

* fix(router): retry a 408 raised by a /v1/messages retry and honor deployment num_retries before the stream opens

* fix(router): attribute a retried /v1/messages stream to the deployment that served it and bound the retry-policy hold

* fix(router): retry /v1/messages error frames under their retry-policy class and keep the first drop's committed budget

An `event: error` frame that arrives before the first content delta now raises the exception class the pre-stream mapping gives an HTTP answer with the same status (429 RateLimitError, 500 and 529 InternalServerError, 503 ServiceUnavailableError, 504 Timeout), so a retry policy's per-class budget governs it the way it governs the error before the stream opened. The status the client sees is unchanged

A retry that lands on a sibling deployment keeps the budget the first drop committed to, read back from the request's attempted_retries and max_retries, instead of recomputing it from the new deployment's num_retries, matching the pre-stream retry loop

* refactor(anthropic): keep the error-frame exception mapping under llms and type the retry test helper

The status-to-exception mapping an `event: error` frame gets before the retry policy is consulted now lives next to the Anthropic error status map in llms/anthropic/common_utils.py, with its own unit test, and the two-deployment retry test helper takes explicit typed parameters instead of a bare dict and untyped kwargs

* refactor(anthropic): map an error frame's status with explicit returns on every path

* fix(router): map stream error frames through the pre-stream exception mapping

An overloaded `event: error` frame on a /v1/messages stream now raises the InternalServerError a 529 answer maps to, built by exception_type from the frame's own body, so one retry policy class governs the error before and after the first byte; a failed fallback after such a frame answers 500 like every other litellm path instead of the frame map's 503

A model_group_retry_policy that does not parse (a non-integer budget, an entry that is not a mapping) no longer fails every healthy stream of that group before its first attempt: the stream runs with no policy and the plain num_retries budget, with a warning naming the group

* fix(router): forward an error frame nothing can take over for as the provider sent it

A pre-content error frame whose class the retry policy grants no retry, with no fallback configured, raised an HTTP error only on the first attempt while the same frame after exhausted retries reached the client verbatim. Both now pass through as sent, the way the merge base forwarded every frame.

* test(integration): audit /v1/messages pre-content retry across routes and budgets

Adds the /audit cells for the pre-content stream retry: the native Anthropic route
(drops and error frames before content, HTTP rejections before the stream opens, SDK
sync and async, after-content and non-retriable controls, budget exhaustion, cache
twin, spend row and headers), the chat and responses bridges, the generic routes
(responses, chat, vllm pass-through, Gemini generateContent, fine-tuning jobs list),
owned two-worker proxies for router-level budgets, retry policies and fallbacks, and
two chaos cells (a worker killed mid burst, an outage on every first attempt). Shared
helpers for scripted Anthropic SSE upstreams and OpenAI-compatible wire replies live
in tests/integration/_support

---------

Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-10-05 22:53:39 +00:00
..
azure_client_usage_test.py fix - correctly re-use azure openai client 2025-03-18 09:51:28 -07:00
ban_constant_numbers.py Squashed commit of the following: (#9709) 2025-04-02 21:24:54 -07:00
ban_copy_deepcopy_kwargs.py Fix - using managed files w/ OTEL + UI - add model group alias on UI (#13171) 2025-07-31 21:22:04 -07:00
bedrock_pricing.py Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_ruff_dead_test_code 2026-08-24 09:46:56 -07:00
callback_manager_test.py (Refactor / QA) - Use LoggingCallbackManager to append callbacks and ensure no duplicate callbacks are added (#8112) 2025-01-30 19:35:50 -08:00
check_data_replace_usage.py Bug fix - String data: stripped from entire content in streamed Gemini responses (#9070) 2025-03-07 21:06:39 -08:00
check_e2e_no_raw_requests.py chore: consolidate CLAUDE.md into AGENTS.md 2026-09-19 02:30:35 +00:00
check_endpoint_coverage.py Revert "[Feature] Add /public/supported_endpoints endpoint" 2026-02-26 17:21:43 -08:00
check_fastuuid_usage.py code cov test script check_fastuuid_usage.py 2025-09-24 10:27:22 +09:00
check_get_model_cost_key_performance.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
check_guardrail_apply_decorator.py content filter test fix 2026-02-12 17:54:16 -08:00
check_licenses.py feat(proxy): embed enterprise LiteAdmin MCP in LiteLLM images (#44610) 2026-10-05 15:24:34 -07:00
check_migrations_no_data_rewrites.py fix(proxy-extras): build the SpendLogs indexes in the migration job instead of in migrations (#43948) 2026-10-01 14:12:22 -07:00
check_prisma_binary_cache.py fix(ci): bound setup steps so pytest always gets its full budget 2026-08-10 16:22:07 +00:00
check_provider_folders_documented.py feat(harness): add Harness.TOOL_LOOP, a minimal in-process tool-calling loop (#44391) 2026-10-03 17:33:01 +00:00
check_py310_typing_imports.py fix: keep litellm importable on Python 3.10 and guard 3.11-only typing imports in CI (#39448) 2026-09-02 18:27:19 -07:00
check_spanattributes_value_usage.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
check_unbounded_in_lists.py ci: fail on new unbounded SQL IN lists and add a Prisma chunking helper (#42629) 2026-09-26 13:40:44 -07:00
check_unsafe_enterprise_import.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
check_workflow_job_name_collisions.py fix(ci): keep an expression matrix directive out of the comparison 2026-09-06 02:59:16 -07:00
check_workflow_startup_safety.py fix(ci): report budgets the startup guard cannot resolve 2026-08-10 18:12:18 +00:00
code_qa_check_tests.py test: move tests/test_litellm root and small trees into tests/unit (#43186) 2026-09-25 11:30:43 -07:00
enforce_llms_folder_style.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
ensure_async_clients_test.py refactor(lens)!: rename internal engine code and API (#44034) 2026-10-01 11:02:14 -07:00
info_log_check.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
liccheck.ini feat(proxy): embed enterprise LiteAdmin MCP in LiteLLM images (#44610) 2026-10-05 15:24:34 -07:00
license_cache.json fix(deps): update python-multipart to >=0.0.20 in CI and test configs 2026-03-03 15:10:39 -03:00
litellm_logging_code_coverage.py docs(litellm_logging_code_coverage.py): fix check 2025-06-18 21:36:03 -07:00
log.txt Squashed commit of the following: (#9709) 2025-04-02 21:24:54 -07:00
memory_test.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
pass_through_code_coverage.py test: initial commit enforcing testing on all anthropic pass through … (#7794) 2025-01-15 22:02:35 -08:00
prevent_key_leaks_in_exceptions.py fix(main.py): fix key leak error when unknown provider given (#8556) 2025-02-15 14:02:55 -08:00
recursive_detector.py fix(guardrails): encrypt guardrail litellm_params secrets at rest (#43627) 2026-10-02 22:49:41 -07:00
router_code_coverage.py fix(router): retry a /v1/messages stream the provider drops before the first content chunk (#44276) 2026-10-05 22:53:39 +00:00
router_enforce_line_length.py Litellm router code coverage 3 (#6274) 2024-10-16 21:30:25 -07:00
test_aio_http_image_conversion.py fix img URL for tests 2025-11-22 09:41:15 -08:00
test_ban_set_verbose.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_chat_completion_imports.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_e2e_changed_gate.py ci: move Postgres, MCP and Redis suites to CircleCI integration (#44453) 2026-10-05 09:33:14 -07:00
test_e2e_idp_stack.py test(e2e): verify IdP readiness through real HTTP 2026-09-11 17:25:38 -07:00
test_e2e_junit_report.py test(e2e): typed per-test metadata for the e2e suite (#42044) 2026-09-30 21:03:21 -07:00
test_e2e_metadata.py test(e2e): typed per-test metadata for the e2e suite (#42044) 2026-09-30 21:03:21 -07:00
test_merge_smoke.py ci: add merge smoke checks workflow with loopback-only harness and 11 curated cases (#42709) 2026-09-23 11:01:08 -07:00
test_no_hardcoded_secrets.py test: run the 30 test files stranded in the second mirror (#37595) 2026-08-20 10:59:43 -07:00
test_provider_cache.py fix(e2e): own a shared fixture's deployment by the fixture's node, not the first test 2026-09-16 17:35:05 -07:00
test_provider_replay_harness.py test: relocate strict replay harness coverage 2026-09-14 17:20:58 -07:00
test_proxy_types_import.py style: run black formatter on files from main merge 2026-04-17 13:02:59 -07:00
test_router_strategy_async.py test: drop the cwd-relative sys.path.insert calls from the test suite (#37802) 2026-08-22 09:25:58 -07:00
test_workflow_job_name_collisions.py fix(ci): keep an expression matrix directive out of the comparison 2026-09-06 02:59:16 -07:00
unbounded_in_baseline.txt fix(proxy): share ownership permissions for spend logs and traces (#44239) 2026-10-03 01:38:54 +00:00
user_api_key_auth_code_coverage.py ci: move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests (#42903) 2026-09-24 22:59:11 +00:00