litellm/tests/claude_code/basic_messaging_streaming/test_anthropic.py
Cursor Agent b24059a92f
claude_code compat: skip skipped reports; drop unreachable stream-events check
- conftest: pytest_runtest_makereport now early-returns on report.skipped
  so pytest.skip(...) inside a compat test body doesn't get recorded as
  a phantom 'fail' row via the not-failed/empty-collected branch.

- _basic_messaging: drop require_stream_events. The check (not outcome.events)
  cannot catch a buffering regression because cli_driver uses
  subprocess.run(capture_output=True), which only exposes the post-exit
  stdout blob — buffered-then-flushed and truly streamed responses are
  indistinguishable. The check was also unreachable as an independent
  failure path (empty events -> empty text -> the text check fires first).
  Update all five streaming callers and docstrings accordingly.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-17 06:51:02 +00:00

52 lines
2 KiB
Python

"""basic_messaging_streaming x Anthropic.
Drive the real `claude` CLI in headless `--output-format stream-json`
mode against a running LiteLLM proxy that routes to Anthropic, and
report the outcome via `compat_result`.
The CLI is run with `--print --output-format stream-json`, which streams
incremental events as the upstream produces tokens. The cell goes green
only when every Claude tier returns a non-empty reply.
Note: a true "did the proxy buffer the full response before flushing?"
check would require observing event arrival times on the wire, which
the `cli_driver` cannot do today — it consumes stdout via
`subprocess.run(capture_output=True)` after the process exits, so a
buffered-then-flushed response is indistinguishable from a truly
streamed one. That regression check belongs in a streaming-aware
driver; until then this cell verifies the same shape as the
non-streaming variant.
The (feature, provider) for this cell is inferred from the file path by
`tests/claude_code/conftest.py`:
tests/claude_code/basic_messaging_streaming/test_anthropic.py
^^^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^^
feature_id provider
The shared `run_basic_messaging_cell` helper fans the three Claude tiers
out in parallel inside this single test, with one
`compat_result.add(...)` entry per model so the matrix builder still
sees three rows for this (feature, provider).
"""
from __future__ import annotations
from tests.claude_code._basic_messaging import run_basic_messaging_cell
ANTHROPIC_MODELS = [
"claude-haiku-4-5",
"claude-sonnet-4-6",
"claude-opus-4-7",
]
def test_basic_messaging_streaming_anthropic(compat_result):
"""Drive the `claude` CLI against the LiteLLM proxy and assert a
non-empty streamed reply (one row per Claude tier).
"""
run_basic_messaging_cell(
compat_result=compat_result,
models=ANTHROPIC_MODELS,
prompt="Count from 1 to 5, one number per line.",
)