litellm/tests/claude_code/basic_messaging_streaming
mateo-berri 28cdbb4485 RALPH: compat matrix slice 5 - add remaining 5 v0 features (full 6x5 grid) (#26481, PRD #26476)
Slice 5 of the Claude Code Compatibility Matrix: extend the published
matrix from the 1x5 grid that landed in slice 2 to the full v0 6x5
grid described in the PRD's "Features in v0" section. After this
slice merges and the daily cron runs, the docs page reflects all six
v0 features against all five providers.

What landed:

- tests/claude_code/manifest.yaml
  Five new entries appended in PRD row order:
  basic_messaging_streaming, tool_use, prompt_caching_5m, vision,
  extended_thinking. The manifest is the row-order source of truth
  the matrix builder respects.

- tests/claude_code/<feature>/test_<provider>.py (25 new files)
  For each of the five new features, five per-provider test files
  modeled on slice 2's basic_messaging_non_streaming/. Each non-Azure
  file parametrizes over Haiku 4.5 / Sonnet 4.6 / Opus 4.7 and drives
  the real `claude` CLI through the driver with feature-specific
  options:

  * basic_messaging_streaming — count-1-to-5 prompt; asserts the
    stream-json wire actually emitted events plus a non-empty reply.
  * tool_use — `--allowed-tools Bash` plus an `echo pong` prompt;
    asserts a `tool_use` content block was emitted.
  * prompt_caching_5m — same baseline prompt as non-streaming, but
    asserts the upstream usage block reports
    cache_creation_input_tokens or cache_read_input_tokens > 0
    (Claude Code stamps cache_control on its system prompt by default,
    so a single live call surfaces it).
  * vision — decodes a checked-in 1x1 PNG (base64 const) into
    `tmp_path` and attaches it via `--image`; asserts a non-empty
    reply.
  * extended_thinking — sets `MAX_THINKING_TOKENS=4096`; asserts a
    `thinking` content block was emitted.

  All five Azure files report `not_applicable` with the standard
  reason: Azure OpenAI Service does not host Anthropic models.

- tests/claude_code/sample_compatibility-matrix.json
  Hand-authored 6x5 sample showing the realistic best-case outcome:
  4 pass + 1 not_applicable (Azure) per row.

- tests/claude_code/_builder_unit_tests/test_v0_layout.py
  New structural unit tests pinning the on-disk shape so future edits
  can't silently flip the matrix shape:
  * manifest lists all six v0 feature ids in PRD order
  * manifest lists all five v0 provider columns in PRD order
  * every (feature, provider) has a test file at the inferred path
  * every test file references all three required Claude tiers
  * every Azure test file is a `not_applicable` declaration

- tests/claude_code/_builder_unit_tests/test_matrix_builder.py
  Renamed the slice-2 1x5 golden test to
  test_build_matrix_6x5_grid_matches_published_sample and rebuilt
  its inputs to feed all six features. The golden file is now the
  6x5 sample.

Key decisions:

- Per-feature per-provider test bodies are deliberately duplicated
  (per the PRD: "Duplication across per-provider files is accepted").
  Each file is self-contained so a contributor touching one cell
  doesn't accidentally regress neighbors.
- Only Azure cells are marked `not_applicable` in this slice. Other
  combinations that turn out to genuinely not apply on the live cron
  run (e.g. a provider that doesn't support `thinking` for a tier)
  will be tightened to `not_applicable` reasons in a follow-up; for
  now they fail honestly, which the matrix renderer paints red.
- prompt_caching_5m's assertion (cache tokens > 0 in the usage block)
  exercises the path Claude Code customers care about: that the proxy
  preserves `cache_control` annotations end-to-end. It does not try
  to differentiate cache_creation vs cache_read across runs.
- The vision PNG fixture is generated at test time from a base64
  const rather than checked into git as a binary — keeps the diff
  text-only and avoids needing PIL or any image-generation library.

Tests: 31 -> 142 unit tests passing (no proxy / no `claude` CLI
required). Test counts:
  * 12 builder tests (was 11; +1 for 6x5 golden, the slice-2 1x5
    test was renamed in place)
  * 100 v0_layout structural tests (new)
  * 10 driver tests (unchanged)
  * 9 compat_result tests (unchanged)
  * 14 publisher unit tests (unchanged)
  * 8 PR-gate version-resolver tests (unchanged)
  * 6 CircleCI structural tests (unchanged)
The 90 per-cell tests under tests/claude_code/<feature>/ continue to
require a running proxy + `claude` CLI; they only run inside the
CircleCI PR gate or the daily-cron VM (both established in slices
3 and 4).

Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- The companion update to compatibility-matrix.json in the docs repo.
  After slice 4's daily-cron lands the App credentials, the cron run
  will replace the docs-side hand-authored JSON automatically; until
  then the slice-2 1x5 sample remains in the docs repo.

Notes for next iteration:
- The exact `claude` CLI flags for tool-allowlist (`--allowed-tools`),
  vision (`--image`), and extended thinking (`MAX_THINKING_TOKENS`)
  are best-guess from the current Claude Code surface; if the live
  PR-gate run reveals different flag names, tighten in place.
- Several non-Azure cells will likely need `not_applicable`
  declarations once the cron VM produces real outcomes (e.g.
  Bedrock Invoke + extended_thinking is uncertain). That refinement
  is an iteration-2 follow-up driven by data, not a blocker for this
  slice.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-25 05:24:22 +00:00
..
__init__.py RALPH: compat matrix slice 5 - add remaining 5 v0 features (full 6x5 grid) (#26481, PRD #26476) 2026-04-25 05:24:22 +00:00
test_anthropic.py RALPH: compat matrix slice 5 - add remaining 5 v0 features (full 6x5 grid) (#26481, PRD #26476) 2026-04-25 05:24:22 +00:00
test_azure.py RALPH: compat matrix slice 5 - add remaining 5 v0 features (full 6x5 grid) (#26481, PRD #26476) 2026-04-25 05:24:22 +00:00
test_bedrock_converse.py RALPH: compat matrix slice 5 - add remaining 5 v0 features (full 6x5 grid) (#26481, PRD #26476) 2026-04-25 05:24:22 +00:00
test_bedrock_invoke.py RALPH: compat matrix slice 5 - add remaining 5 v0 features (full 6x5 grid) (#26481, PRD #26476) 2026-04-25 05:24:22 +00:00
test_vertex_ai.py RALPH: compat matrix slice 5 - add remaining 5 v0 features (full 6x5 grid) (#26481, PRD #26476) 2026-04-25 05:24:22 +00:00