mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-06 02:48:13 +00:00
* [Test] Add Azure async chat completion timeout test. WIP * Capture TTFT for /v1/messages streaming responses The pass-through streaming path for /v1/messages (Anthropic, Bedrock, Vertex AI, Azure AI, Minimax) logged completion_start_time only after the entire stream finished. async_success_handler then fell back to end_time, making TTFT equal to total duration or null in the UI and Prometheus. Record the timestamp of the first chunk in async_sse_wrapper and propagate it to model_call_details before the logging handler runs, so gen_ai.response.time_to_first_token reflects the real first-chunk latency. Fixes #25598 * [Refactor] Implement timeout resolution logic in completion function add fetch ``request_timeout`` from litellm_settings * remove stale test case * remove extra print statement * default request timeout value in constants to 600s to match timeout defaults handled in the proxy * fix request timeout if using default value from constants.py * update code structure, test cases * only override if the global timeout sets timeout to 6000s * update code structure, move hard coded values to const and make the reslve function readable by moving fallback logic to a seperate function * modify default timeout values, replacing hard coded ones with default values defined --------- Co-authored-by: harish876 <harishgokul01@gmail.com> Co-authored-by: Joaquin Hui Gomez <joaquinhuigomez@users.noreply.github.com>
46 lines
1.4 KiB
Python
46 lines
1.4 KiB
Python
"""
|
|
``_get_httpx_client`` + ``HTTPHandler.post`` (same pattern as Azure Anthropic sync path:
|
|
``_get_httpx_client(params={"timeout": ...})`` then ``post(..., timeout=...)``).
|
|
|
|
Uses https://httpbin.org/delay/10 with ``timeout=5`` — the handler must raise :class:`~litellm.exceptions.Timeout`
|
|
before the 10s delay completes. Skips if httpbin is unreachable.
|
|
|
|
Lives under ``local_testing`` (not ``make test-unit``).
|
|
"""
|
|
|
|
import json
|
|
import os
|
|
import sys
|
|
|
|
import httpx
|
|
import pytest
|
|
|
|
sys.path.insert(
|
|
0, os.path.abspath(os.path.join(os.path.dirname(__file__), "../.."))
|
|
)
|
|
|
|
from litellm.exceptions import Timeout as LitellmTimeout
|
|
from litellm.llms.custom_httpx.http_handler import _get_httpx_client
|
|
|
|
_HTTPBIN_DELAY_S = 10
|
|
_PER_REQUEST_TIMEOUT_S = 5.0
|
|
_CLIENT_DEFAULT_TIMEOUT_S = 60.0
|
|
|
|
|
|
def test_post_delay_exceeds_per_request_timeout_raises():
|
|
try:
|
|
httpx.get("https://httpbin.org/get", timeout=5.0)
|
|
except Exception as e:
|
|
pytest.skip(f"httpbin.org unreachable: {e}")
|
|
|
|
handler = _get_httpx_client(params={"timeout": _CLIENT_DEFAULT_TIMEOUT_S})
|
|
try:
|
|
with pytest.raises(LitellmTimeout):
|
|
handler.post(
|
|
f"https://httpbin.org/delay/{_HTTPBIN_DELAY_S}",
|
|
headers={"content-type": "application/json"},
|
|
data=json.dumps({"model": "claude", "messages": []}),
|
|
timeout=_PER_REQUEST_TIMEOUT_S,
|
|
)
|
|
finally:
|
|
handler.close()
|