litellm/tests/e2e/coverage_registry
yucheng-berri 0223383d94
test(e2e): datadog log delivery for streamed routes, read back from the real datadog api (#33566)
* fix(e2e): make the datadog read-back find what DataDog actually indexes

Live verification of the merged #33604 against real DataDog (us5) exposed
three read-back defects that the local-sink tests could never see; all
three fixes are verified against the real API:

- Marker search: DataDog consumes the shipped JSON message into the
  event's attributes and leaves the indexed message EMPTY, so the
  full-text '"marker"' query matched nothing and every test failed with
  zero events. The query is now '*:*marker*', which scans all attributes
  (the marker sits in messages.content); verified to return exactly the
  event for the call.

- Rate limit: the Logs Search API budget is 2 requests per 10s org-wide
  (x-ratelimit-name logs_public_search_api). Polling at POLL_INTERVAL=5s
  sat exactly at the limit and the reader hard-failed on the first 429.
  Searches now pace at DD_SEARCH_INTERVAL (10s default) and a 429 backs
  off and retries up to 5 times; only non-429 failures stay hard fails.

- Envelope status: DataDog re-derives the indexed event status from the
  parsed payload's status attribute ('success') and normalizes it to its
  OK severity, so the assertion expects 'ok', not the shipped 'info'.

Live run: chat_completions and responses pass every assertion including
the exact response-cost cross-check; messages red-pins the LIT-4447
duplicate for real (one call -> two sync-sweep copies + one async batch
copy, same request id, confirmed in proxy debug logs). The duplicate is
race-dependent, so the pin flickers until #33589 lands.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(e2e): datadog log delivery for streamed chat, messages, and responses

Rewritten from the dd-sink version (original #33566) to judge delivery on
what real DataDog ingested, matching the merged #33604 conversion: the
dd_logs reader searches events back through the Logs Search API and the
assertions validate the indexed envelope (source:litellm tag, ok status)
and the StandardLoggingPayload fields under the event's attributes.

Each streamed test drives one STREAMED call per route, asserts the stream
actually streamed (event-stream content type, >0 chunks, no upstream error
event), then pins exactly one DataDog event whose payload records
stream=true, the aggregated token count, and a response_cost equal to the
/spend/logs row for the call - a stream's headers ship before its cost
exists, so the spend row is the cross-check anchor, and the spend row and
DataDog event must also agree on total_tokens.

Coverage registry: adds logging.datadog.stream.exports_metric exercised on
chat_completions, messages, and responses.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Update test_datadog_log_e2e.py

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 19:37:07 -07:00
..
__init__.py test(e2e): add coverage registry and collector (#32304) 2026-07-07 15:51:20 -04:00
collector.py ci: gate tests/e2e on zero basedpyright errors in pre-commit and lint CI 2026-07-11 10:25:22 -07:00
guardrail.yaml test(e2e): add coverage registry and collector (#32304) 2026-07-07 15:51:20 -04:00
llm_claude_code_compat.yaml fix(e2e/claude_code): unblock stage collection, align proxy env names, register compat models (#33433) 2026-07-16 11:05:31 -07:00
llm_conversational.yaml test(e2e): cover model-aware mid-conversation system handling on Bedrock Invoke /v1/messages 2026-07-11 17:08:16 -07:00
llm_nonconversational.yaml test(e2e): add coverage registry and collector (#32304) 2026-07-07 15:51:20 -04:00
logging.yaml test(e2e): datadog log delivery for streamed routes, read back from the real datadog api (#33566) 2026-07-16 19:37:07 -07:00
mcp.yaml test(e2e): add coverage registry and collector (#32304) 2026-07-07 15:51:20 -04:00
mgmt.yaml fix(ui): scope key models dropdown options to the key's team (#32382) 2026-07-07 18:54:19 -07:00
other.yaml test(e2e): add coverage registry and collector (#32304) 2026-07-07 15:51:20 -04:00
quota_management.yaml fix(e2e): name the route the spend_calculate registry cell actually exercises 2026-07-11 16:13:48 -07:00
README.md feat(e2e): emit structured E2E_RESULT lines for package status history (#33578) 2026-07-16 15:01:01 -07:00
registry.py ci: gate tests/e2e on zero basedpyright errors in pre-commit and lint CI 2026-07-11 10:25:22 -07:00
reliability.yaml fix(complexity_router): return empty dict from _classifier_call_metadata when metadata is absent (#33452) 2026-07-15 17:46:00 -07:00
schema.py fix(e2e/claude_code): unblock stage collection, align proxy env names, register compat models (#33433) 2026-07-16 11:05:31 -07:00
test_collector.py test: emit e2e coverage lines for loki (#32513) 2026-07-08 11:06:56 -07:00

e2e coverage registry

This directory is the denominator for e2e test coverage: the set of behaviors we want covered, one row per behavior, checked into the repo so coverage is a number we can track instead of a guess. It implements the plan in the "E2E Coverage Tracking" note; the naming grammar lives in tests/e2e/CLAUDE.md.

The model

A cell is one customer-noticeable behavior a single e2e test can assert pass/fail on, for example llm.chat_completions.bedrock_converse.tool_use.stream.works. Cells are grouped module > feature > test, with LLM cells split into Core LLMs and Non-Core LLMs for dashboarding. Each cell carries a tier (P0/P1/P2), a source, and a fail_before_fix flag.

The rows live in per-prefix YAML files (llm_*.yaml, mgmt.yaml, mcp.yaml, reliability.yaml, quota_management.yaml, logging.yaml, guardrail.yaml, other.yaml) and validate against the discriminated union in schema.py, so an LLM row cannot carry a guardrail field and vice versa. llm rows with subject_endpoint of chat_completions, messages, or responses roll up to Core LLMs; all other LLM endpoints roll up to Non-Core LLMs. LLM endpoint, route, and capability values are typed in schema.py, so new taxonomy values require an explicit schema change. logging and guardrail are two id-prefixes that roll up into the single Logging & Guardrails dashboard module.

A test declares what it covers with a marker:

@pytest.mark.covers("llm.chat_completions.openai.tool_use.stream.works")
def test_openai_streaming_tool_calls(self) -> None:
    ...

The number

collector.py diffs the registry against those markers and reports coverage per module. It is static: a collect-only pass reads the markers, so it runs no test and needs no live proxy. Whether a covered cell currently passes or fails is a separate, live concern.

cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector

Use --format loki after the e2e pytest run in the same Kubernetes job/pod to print structured stdout lines for Loki:

cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector --format loki --strict

This emits exactly one COVERAGE_TOTAL line and one COVERAGE_MODULE line per module in MODULE_ORDER, in that order. Loki uses log-safe module= labels from LOKI_MODULE_LABELS (core_llms, management_ui, etc.) so existing JSON and Prometheus consumers keep their human-readable module names unchanged.

Live pass/fail is separate: each finished pytest node prints an E2E_RESULT logfmt line (see tests/e2e/e2e_result_reporter.py and tests/e2e/grafana/status_history_panels.md). Coverage answers "is there a test for this cell?"; E2E_RESULT answers "did that run pass?"

The headline is overall coverage. The collector also lists markers that point at ids not in the registry, so a typo or an unenumerated behavior surfaces instead of being silently dropped.

Use strict mode in CI once existing draft markers are reconciled:

cd tests/e2e && PYTHONPATH=. python -m coverage_registry.collector --strict

Strict mode exits non-zero on @pytest.mark.covers(...) ids that are not checked into the registry. Add --fail-on-collection-errors when the job should also fail on pytest collection errors.

Status: this is a draft for review

The cells were enumerated from the codebase and the tiers are a first proposal. Known things to settle before treating the set as final:

  • tiers are proposed, not signed off; 125 P0 is a lot to prove fail-before-fix, so P0 may want tightening
  • a few cells need a support check or a prune (for example llm.embeddings.anthropic.* and reliability.perf.throughput.under_slo)
  • auth is covered in two places (other.auth.* and the mgmt authz assertions); the boundary needs a decision, and the auth cluster may deserve promotion to its own module
  • the P2 "niche" cells each stand in for a large tail of integrations/providers by design, so the denominator is deliberately P0-weighted rather than a full inventory