mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-11 03:38:38 +00:00
* test(integration): poll the partition lock witness off the event loop The witness poll ran a blocking psycopg connect and query on the test's event loop, so the partition DDL task only progressed during the 20ms sleeps and missed the 3s deadline on loaded runners. Polling through asyncio.to_thread lets the DDL reach the held lock concurrently, and the deadline is 10s since it only bounds how long the DDL takes to start waiting * test(e2e): retry the reasoning turn until the model emits a reasoning item Whether gpt-5.4-mini emits a reasoning item next to a forced function call is up to the model, and it skipped it on four scheduled runs while a same-SHA rerun passed. The first turn now retries up to three times before the unchanged reasoning-item assertion * test(e2e): retry a Bedrock Converse model once when it reports no cache tokens Opus on Bedrock Converse reported zero cache tokens on half of the scheduled runs while a same-SHA rerun passed. A model that comes back uncached is asked once more inside the 5 minute window, and the non-zero cache token assertion is unchanged * test(e2e): poll until a deleted stored response stops being retrievable Azure kept serving a just-deleted streamed response for a moment, so a single retrieve did not raise. Both delete tests now poll retrieve until it returns an error and still require a 4xx * test(e2e): give the Together prompt cache five primed attempts Together documents its prefix cache as best-effort, and three primed attempts came back with zero cached tokens once while a same-SHA rerun passed * test(e2e): allow 30s for the idle RSS reading at session start A loaded router replica took longer than 10s to answer the session-start memory read, which pytest reruns cannot recover. The RSS budget assertion is unchanged * test(e2e): scope the MCP submission check to its own card The register call timed out at 15s on a loaded stack and leaked a submitted server, so the retries matched two "3 passing, 1 failing" cards. Registration gets 60s and the check reads only the submitted server's card * test(router): wait for the primary success record before counting it The shadow fan-out test waited only for the two shadow success events and then asserted exactly one primary success, which the logging worker sometimes had not delivered yet. It needed a rerun on 5 of 59 main runs * test(proxy): give the spend-log cleanup run a 1s budget A 0.25s budget could expire before the first batch on a loaded xdist worker, leaving rows_deleted at 0. One second still stops the 50-batch, 5-second loop on the deadline, which is what the test proves * test(autorouter): release the slow token count at 2.05s instead of 2.3s The count only has to finish after a quarter of the 8s worker budget, and the planner gives it 3s, so releasing at 2.3s left 0.7s of slack that a loaded runner used up. 2.05s is still past the quarter mark with almost 1s of slack * test(xai): run the reasoning-effort tests on grok-4.6 with a prompt that needs reasoning grok-4.7 now returns reasoning tokens without any reasoning text, so reasoning_content was never set and both tests failed on every scheduled run. Called directly, grok-4.6 returned reasoning text 6 of 6 times on a step-by-step arithmetic prompt but only some of the time on a bare greeting * test(integration): give the two worker-kill chaos tests the owned-proxy time budget Each test spends about 50s on the burst and then up to graceful_stop_seconds() stopping its owned proxy, which overran the 90s default pytest timeout on loaded runners. They now use the same 2 * graceful_stop_seconds() + 120 budget as other owned-proxy tests * test(integration): compare regional image responses without their created timestamp Two identical generations that straddle a second boundary differ only in created, which failed the byte-for-byte comparison. Every other field and both spend rows are still compared * test(integration): count only the S3 logger's flush task The set of asyncio tasks created while the logger starts also picked up client close finalizers left by earlier tests, so the count ranged from 1 to 8. The test now counts the periodic_flush task it owns and cancels * test(integration): send guardrail timeout probes eight at a time All 44 probes went at once to a two-worker proxy, so on a loaded runner some guardrails used their whole 1s timeout before their request left the proxy, and the sink never saw them. A different set of providers failed on each run. Eight in flight still overlaps the waits * test(integration): count a killed Arize chaos worker as gone once it is a zombie psutil reports an unreaped zombie as still running, so the 10s death check failed whenever the supervisor was slow to reap the SIGKILLed worker, and the stop then overran the 90s default timeout. The test now accepts a zombie or a missing process and has the owned-proxy time budget * test(integration): send the Typesafe connection test its real model and endpoint The test passed the proxy alias as litellm_params.model, and /health/test_connection lays the request over the stored deployment, so the alias replaced the real typesafe/ model and the check failed with "LLM Provider NOT provided" on every run since it landed in #45481. It now sends the provider model, api_base and api_key directly * test(integration): keep DB-stored models out of the usage-routing Redis read test The owned proxy inherited store_model_in_db and the job's shared database, so a model another test left behind joined the router and added its cooldown key to the MGET the test compares exactly. Reproduced locally with one /model/new model present (3 failed), and green with model loading from the database turned off * test(vertex_ai): prove the batch upload streams by laziness instead of peak memory The two tracemalloc ratio tests flaked on unrelated PRs because peak memory on a shared xdist worker includes other threads' allocations and garbage from earlier tests. They are replaced by deterministic checks of the same property: the upload stream parses and maps a row only when it is pulled, so a body whose tail is not JSON yields its valid rows first, and a Path source reflects a row rewritten on disk after the upload started. Making the parse, the output, or the file read eager fails these tests * test(integration): keep the alias test_connection call as a known bug * test: type the counting mapper and the Bedrock rerun helper * test: annotate the new test locals as Final and build them as tuples * test(vertex_ai): prove pull-driven transforms with the garbage tail alone |
||
|---|---|---|
| .. | ||
| basic_messaging_non_streaming | ||
| basic_messaging_streaming | ||
| count_tokens | ||
| cron_vm | ||
| long_context_1m | ||
| passthrough | ||
| pdf_input | ||
| prompt_caching_1h | ||
| prompt_caching_5m | ||
| structured_outputs | ||
| thinking | ||
| thinking_with_tool_use | ||
| tool_search | ||
| tool_use | ||
| tool_use_streaming | ||
| vision | ||
| web_search | ||
| __init__.py | ||
| _basic_messaging.py | ||
| _compat_models.py | ||
| _env.py | ||
| _gpt_cells.py | ||
| _passthrough.py | ||
| cli_driver.py | ||
| conftest.py | ||
| http_probe.py | ||
| manifest.yaml | ||
| matrix_builder.py | ||
| pr_gate_version_resolver.py | ||
| rate_limiter.py | ||
| run_compat.sh | ||
| sample_compatibility-matrix.json | ||
| test_config.yaml | ||