mirror of
https://github.com/BerriAI/litellm.git
synced 2026-09-26 01:12:21 +00:00
A test's JUnit report says whether it passed, never what it did or where a failing test died. This records that from the harness, so nothing about it is hand-written and it cannot drift from what the test actually ran
`@step("create team with a budget")` from the new tests/e2e/e2e_metadata.py goes on harness helpers, never on tests, and appends its label to the running test's step log in call order. The label is recorded before the wrapped call, so a helper that raises still leaves its own label last: a failing test's last step is where it died. Every public harness method that performs an action now carries one, 355 across the client modules, lifecycle, idp, the logging readers, migrations and the claude_code driver
Only the outermost step records, tracked per thread. Harness layers call each other (ResourceManager.key goes through ProxyClient.generate_key, a domain client wraps the shared ProxyClient), so every layer carries a label and the story still reads at the level the test called in at, one beat per action. A step above @contextmanager holds the guard through __enter__ and __exit__, so a context's cleanup never lands behind the step a test died on, and a bare generator function is refused at import because its body interleaves with its caller's. Consecutive duplicates collapse and the log caps at 50, so a poll loop is one beat rather than fifty. The wrapper is a frame, so the eight cleanup and retry warnings raised directly inside decorated helpers use stacklevel=2 + STEP_FRAMES to keep reporting at their caller
The log is emptied first thing in pytest_runtest_setup and attached from the existing pytest_runtest_makereport wrapper after setup and again after call, so a test that errors in a fixture keeps the steps recorded before the crash. Teardown does not attach: finalizer steps are cleanup. Each attach drops the item's earlier step entries, so the second attach and a --reruns 1 retry replace the story rather than doubling it
Steps ride out as repeated <property name="step"> entries behind the fixed package/covers/source prefix, which stays byte-identical. The project-releaser emitter already regroups them into the results JSON's steps array. test_junit_report.py runs real pytest with --junitxml against this conftest, in-process and under -n 2, and pins the passing, failing, setup-error, rerun and wide-scope-fixture cases on the parsed XML
158 lines
6.1 KiB
Python
158 lines
6.1 KiB
Python
"""Shared body for the `basic_messaging_*` × <provider> compat cells.
|
||
|
||
Every basic_messaging cell follows the same skeleton:
|
||
|
||
1. Read the proxy base URL + API key from env, fail-early if missing.
|
||
2. Fan the three Claude tiers out via `run_claude_models_parallel`.
|
||
3. Inspect each model's outcome and report one `compat_result` row per
|
||
model — `ClaudeCLIError`, non-zero exit, and empty assistant text
|
||
are all per-model fails; everything else is a per-model pass.
|
||
4. Surface a joined failure message via `pytest.fail(...)` so the
|
||
pytest run also goes red.
|
||
|
||
The streaming variant additionally passes `verify_streaming=True`,
|
||
which adds the `--include-partial-messages` CLI flag and asserts that
|
||
the proxy actually streamed the response (see the helper docstring for
|
||
the wire-level rationale).
|
||
|
||
The conftest infers `(feature_id, provider)` purely from the test file
|
||
path, so each per-provider file just declares its model list and calls
|
||
`run_basic_messaging_cell(...)`. This keeps all cell logic in one place
|
||
— a future tweak to the env-missing guard or the failure-loop shape
|
||
now propagates to every cell automatically.
|
||
|
||
The leading underscore in the filename is what keeps pytest from
|
||
collecting this module as a test file.
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
from typing import Any, Callable, Mapping, Sequence
|
||
|
||
import pytest
|
||
|
||
from e2e_metadata import step
|
||
|
||
from claude_code._env import require_proxy
|
||
from claude_code.cli_driver import (
|
||
ClaudeCLIError,
|
||
DriverResult,
|
||
failure_diagnostic,
|
||
run_claude_models_parallel,
|
||
)
|
||
|
||
|
||
ClaudeRunner = Callable[..., Mapping[str, DriverResult | ClaudeCLIError]]
|
||
|
||
# Floor on the number of `stream_event` records (with delta payloads)
|
||
# we expect to see when the proxy actually streams. With
|
||
# `--include-partial-messages`, the CLI emits one `stream_event` per
|
||
# raw upstream SSE event — a fully-streamed response produces many
|
||
# (`message_start`, multiple `content_block_delta`s, `content_block_stop`,
|
||
# `message_delta`, `message_stop`); a proxy that buffers the upstream
|
||
# and returns a single non-streaming chunk produces 0 or 1. Floor of 2
|
||
# is safely above the buffered case for any non-trivial reply, which
|
||
# is why the streaming cells use a "count from 1 to 5" style prompt.
|
||
MIN_STREAM_DELTA_EVENTS = 2
|
||
|
||
|
||
def _count_stream_event_deltas(events: Sequence[Mapping[str, Any]]) -> int:
|
||
"""Count `stream_event` records that carry an SSE event payload.
|
||
|
||
With `--include-partial-messages`, Claude Code wraps every upstream
|
||
SSE event in a `{"type": "stream_event", "event": {...}}` record.
|
||
A buffering proxy collapses the upstream stream into a single
|
||
non-streaming response, so these records vanish. Counting them
|
||
(rather than just `len(events)`) is the wire-level signal that
|
||
"did the proxy preserve streaming?" — independent of the `system`
|
||
/`assistant`/`result` boilerplate records the CLI always emits.
|
||
"""
|
||
count = 0
|
||
for event in events:
|
||
if event.get("type") != "stream_event":
|
||
continue
|
||
if isinstance(event.get("event"), Mapping):
|
||
count += 1
|
||
return count
|
||
|
||
|
||
@step("send a basic message via the claude CLI")
|
||
def run_basic_messaging_cell(
|
||
*,
|
||
compat_result,
|
||
models: Sequence[str],
|
||
prompt: str,
|
||
verify_streaming: bool = False,
|
||
env: Mapping[str, str] | None = None,
|
||
runner: ClaudeRunner = run_claude_models_parallel,
|
||
) -> None:
|
||
"""Run the shared `basic_messaging_*` × <provider> cell body.
|
||
|
||
When ``verify_streaming=True``, the cell additionally asserts that
|
||
the proxy streamed the response end-to-end. The check works by
|
||
passing ``--include-partial-messages`` to the `claude` CLI, which
|
||
causes it to emit one ``stream_event`` record per raw upstream SSE
|
||
event (``message_start``, ``content_block_delta``, ``message_stop``,
|
||
etc.). A proxy that buffers the upstream stream and returns a
|
||
single non-streaming response collapses those records to zero —
|
||
so a floor of ``MIN_STREAM_DELTA_EVENTS`` ``stream_event`` records
|
||
catches the buffering regression without needing a streaming-aware
|
||
driver.
|
||
|
||
This is the same shape of check as ``tool_use_streaming`` uses,
|
||
just keyed off the explicit partial-message flag so it works for
|
||
plain assistant replies (where the CLI would otherwise collapse a
|
||
streamed reply to a single ``assistant`` event in
|
||
``--print --output-format stream-json`` mode).
|
||
"""
|
||
base_url, api_key = require_proxy(compat_result, env=env)
|
||
|
||
extra_args: Sequence[str] = (
|
||
("--include-partial-messages",) if verify_streaming else ()
|
||
)
|
||
|
||
outcomes = runner(
|
||
models=models,
|
||
prompt=prompt,
|
||
base_url=base_url,
|
||
api_key=api_key,
|
||
extra_args=extra_args,
|
||
)
|
||
|
||
failures = []
|
||
for model in models:
|
||
outcome = outcomes[model]
|
||
if isinstance(outcome, ClaudeCLIError):
|
||
error = f"[{model}] {outcome}"
|
||
compat_result.add({"status": "fail", "error": error})
|
||
failures.append(error)
|
||
continue
|
||
|
||
if outcome.exit_code != 0:
|
||
error = f"[{model}] claude CLI failed: {failure_diagnostic(outcome)}"
|
||
compat_result.add({"status": "fail", "error": error})
|
||
failures.append(error)
|
||
continue
|
||
|
||
if not outcome.text.strip():
|
||
error = f"[{model}] claude returned empty assistant text"
|
||
compat_result.add({"status": "fail", "error": error})
|
||
failures.append(error)
|
||
continue
|
||
|
||
if verify_streaming:
|
||
stream_event_count = _count_stream_event_deltas(outcome.events)
|
||
if stream_event_count < MIN_STREAM_DELTA_EVENTS:
|
||
error = (
|
||
f"[{model}] only {stream_event_count} stream_event records "
|
||
f"observed (< {MIN_STREAM_DELTA_EVENTS}); proxy likely "
|
||
f"buffered the upstream response instead of streaming it"
|
||
)
|
||
compat_result.add({"status": "fail", "error": error})
|
||
failures.append(error)
|
||
continue
|
||
|
||
compat_result.add({"status": "pass"})
|
||
|
||
if failures:
|
||
pytest.fail("; ".join(failures), pytrace=False)
|