Address the Greptile concern that basic_messaging_streaming and
basic_messaging_non_streaming used the same implementation, so a proxy
that buffered the upstream stream would silently show green for the
streaming row.
The fix:
- _basic_messaging.run_basic_messaging_cell accepts verify_streaming=True,
which passes --include-partial-messages to the claude CLI. That flag
causes the CLI to emit one stream_event record per upstream SSE event
(message_start, content_block_delta, message_stop, ...). A buffering
proxy collapses the stream to a single non-streaming response, so
zero stream_event records are emitted.
- The cell rejects any model whose stream_event count is below
MIN_STREAM_DELTA_EVENTS (2) -- safely above the buffered case for any
non-trivial reply. Same all-must-pass shape as the existing
tool_use_streaming row.
- All five basic_messaging_streaming/test_*.py per-provider cells now
pass verify_streaming=True; the non-streaming variants are unchanged.
- New unit tests cover the helper, the partial-messages flag wiring,
the streamed/buffered branching, and the all-models-must-stream
contract.
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
- conftest: pytest_runtest_makereport now early-returns on report.skipped
so pytest.skip(...) inside a compat test body doesn't get recorded as
a phantom 'fail' row via the not-failed/empty-collected branch.
- _basic_messaging: drop require_stream_events. The check (not outcome.events)
cannot catch a buffering regression because cli_driver uses
subprocess.run(capture_output=True), which only exposes the post-exit
stdout blob — buffered-then-flushed and truly streamed responses are
indistinguishable. The check was also unreachable as an independent
failure path (empty events -> empty text -> the text check fires first).
Update all five streaming callers and docstrings accordingly.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
- run_claude_models_parallel: catch all exceptions in the per-model
worker and wrap unexpected ones into a ClaudeCLIError so the
documented 'errors as values' contract holds for OSError, ValueError,
etc., not just ClaudeCLIError. Without this, an unexpected raise in
any layer (rate limiter file I/O, infer_provider, etc.) abandons the
remaining models' results and crashes the calling test.
- test_run_claude_places_extra_args_before_prompt: drop the dead first
branch of the 'or' assertion — cmd[-3:] never matches that shape, so
the alternative was misleading dead code.
- basic_messaging_{non_streaming,streaming}/test_*.py: extract the
shared cell body into tests/claude_code/_basic_messaging.py.
Each per-provider file now declares its model list and calls
run_basic_messaging_cell(), eliminating ~700 lines of copy-paste
across 10 files. Updated _builder_unit_tests/test_v0_layout.py to
accept the helper-based pattern alongside direct run_claude() calls.
Co-authored-by: Yassin Kortam <yassin@berri.ai>
- Switch all per-cell tests from @pytest.mark.parametrize("model", ...)
(3 sequential invocations) to a single test that fans out to all 3
Claude tiers via run_claude_models_parallel. Per-cell wall time is now
bounded by the slowest model rather than the sum.
- Add 5 new v0 feature dirs (5 providers each, 25 new test files):
web_search, pdf_input, prompt_caching_1h,
tool_use_streaming, thinking_with_tool_use
Manifest expanded to match.
- Add cross-process token-bucket rate limiter (rate_limiter.py + tests)
so xdist workers stay under per-provider req/s limits during full-grid
runs. New env knobs: LITELLM_COMPAT_RATE_{ANTHROPIC,AZURE,VERTEX_AI,
BEDROCK_CONVERSE,BEDROCK_INVOKE}.
- conftest.py: write per-worker shards under <artifact>.shards/, merge
in the controller; preserve the "don't write empty artifact" guard so
unit-test runs don't clobber a real compat-results.json.
- Vertex test_config.yaml: route project/location through env so the
cron VM can target a different GCP project than the upstream default.
- Add run_compat.sh wrapper for binary-searching ideal req/s per
provider against compat-rate-limit-summary.json output.
Anthropic and Microsoft announced Claude Haiku 4.5, Sonnet 4.5/4.6, and
Opus 4.1/4.6/4.7 in Microsoft Foundry on 2025-11-18, so the matrix's
Azure column should exercise a real route through the LiteLLM proxy
rather than report not_applicable.
Foundry serves Claude on an Anthropic-shape /anthropic/v1/messages
endpoint (not the Azure OpenAI chat-completions route), and LiteLLM
already supports it via the azure_ai/claude-* provider prefix
(litellm/llms/azure_ai/anthropic/{handler,transformation,messages_transformation}.py).
- test_config.yaml: add 3 azure aliases pointing at azure_ai/claude-*
with AZURE_FOUNDRY_API_BASE / AZURE_FOUNDRY_API_KEY env
- 6x test_azure.py: replace not_applicable stubs with real run_claude
drivers, mirroring the existing test_vertex_ai.py shape exactly
- sample_compatibility-matrix.json: Azure cells flip to pass
- _builder_unit_tests: pin the new invariant (run_claude is used,
not_applicable is gone) and feed pass results across all 5 providers
in the 6x5 golden test
Slice 5 of the Claude Code Compatibility Matrix: extend the published
matrix from the 1x5 grid that landed in slice 2 to the full v0 6x5
grid described in the PRD's "Features in v0" section. After this
slice merges and the daily cron runs, the docs page reflects all six
v0 features against all five providers.
What landed:
- tests/claude_code/manifest.yaml
Five new entries appended in PRD row order:
basic_messaging_streaming, tool_use, prompt_caching_5m, vision,
extended_thinking. The manifest is the row-order source of truth
the matrix builder respects.
- tests/claude_code/<feature>/test_<provider>.py (25 new files)
For each of the five new features, five per-provider test files
modeled on slice 2's basic_messaging_non_streaming/. Each non-Azure
file parametrizes over Haiku 4.5 / Sonnet 4.6 / Opus 4.7 and drives
the real `claude` CLI through the driver with feature-specific
options:
* basic_messaging_streaming — count-1-to-5 prompt; asserts the
stream-json wire actually emitted events plus a non-empty reply.
* tool_use — `--allowed-tools Bash` plus an `echo pong` prompt;
asserts a `tool_use` content block was emitted.
* prompt_caching_5m — same baseline prompt as non-streaming, but
asserts the upstream usage block reports
cache_creation_input_tokens or cache_read_input_tokens > 0
(Claude Code stamps cache_control on its system prompt by default,
so a single live call surfaces it).
* vision — decodes a checked-in 1x1 PNG (base64 const) into
`tmp_path` and attaches it via `--image`; asserts a non-empty
reply.
* extended_thinking — sets `MAX_THINKING_TOKENS=4096`; asserts a
`thinking` content block was emitted.
All five Azure files report `not_applicable` with the standard
reason: Azure OpenAI Service does not host Anthropic models.
- tests/claude_code/sample_compatibility-matrix.json
Hand-authored 6x5 sample showing the realistic best-case outcome:
4 pass + 1 not_applicable (Azure) per row.
- tests/claude_code/_builder_unit_tests/test_v0_layout.py
New structural unit tests pinning the on-disk shape so future edits
can't silently flip the matrix shape:
* manifest lists all six v0 feature ids in PRD order
* manifest lists all five v0 provider columns in PRD order
* every (feature, provider) has a test file at the inferred path
* every test file references all three required Claude tiers
* every Azure test file is a `not_applicable` declaration
- tests/claude_code/_builder_unit_tests/test_matrix_builder.py
Renamed the slice-2 1x5 golden test to
test_build_matrix_6x5_grid_matches_published_sample and rebuilt
its inputs to feed all six features. The golden file is now the
6x5 sample.
Key decisions:
- Per-feature per-provider test bodies are deliberately duplicated
(per the PRD: "Duplication across per-provider files is accepted").
Each file is self-contained so a contributor touching one cell
doesn't accidentally regress neighbors.
- Only Azure cells are marked `not_applicable` in this slice. Other
combinations that turn out to genuinely not apply on the live cron
run (e.g. a provider that doesn't support `thinking` for a tier)
will be tightened to `not_applicable` reasons in a follow-up; for
now they fail honestly, which the matrix renderer paints red.
- prompt_caching_5m's assertion (cache tokens > 0 in the usage block)
exercises the path Claude Code customers care about: that the proxy
preserves `cache_control` annotations end-to-end. It does not try
to differentiate cache_creation vs cache_read across runs.
- The vision PNG fixture is generated at test time from a base64
const rather than checked into git as a binary — keeps the diff
text-only and avoids needing PIL or any image-generation library.
Tests: 31 -> 142 unit tests passing (no proxy / no `claude` CLI
required). Test counts:
* 12 builder tests (was 11; +1 for 6x5 golden, the slice-2 1x5
test was renamed in place)
* 100 v0_layout structural tests (new)
* 10 driver tests (unchanged)
* 9 compat_result tests (unchanged)
* 14 publisher unit tests (unchanged)
* 8 PR-gate version-resolver tests (unchanged)
* 6 CircleCI structural tests (unchanged)
The 90 per-cell tests under tests/claude_code/<feature>/ continue to
require a running proxy + `claude` CLI; they only run inside the
CircleCI PR gate or the daily-cron VM (both established in slices
3 and 4).
Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- The companion update to compatibility-matrix.json in the docs repo.
After slice 4's daily-cron lands the App credentials, the cron run
will replace the docs-side hand-authored JSON automatically; until
then the slice-2 1x5 sample remains in the docs repo.
Notes for next iteration:
- The exact `claude` CLI flags for tool-allowlist (`--allowed-tools`),
vision (`--image`), and extended thinking (`MAX_THINKING_TOKENS`)
are best-guess from the current Claude Code surface; if the live
PR-gate run reveals different flag names, tighten in place.
- Several non-Azure cells will likely need `not_applicable`
declarations once the cron VM produces real outcomes (e.g.
Bedrock Invoke + extended_thinking is uncertain). That refinement
is an iteration-2 follow-up driven by data, not a blocker for this
slice.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>