litellm/tests/claude_code/manifest.yaml
mateo-berri be65b4e23b feat(claude_code): rename thinking row + add 4 feature rows (15 total)
Matrix grows from 11 to 15 feature rows. All new tests collected + 180
unit tests still pass; smoke runs hit real LiteLLM bug surfaces on
bedrock_invoke, bedrock_converse, and vertex_ai (cells correctly red
in PR #142).

Rename
------
`extended_thinking` -> `thinking` (directory, manifest id+name, 5
test fn names, 5 docstrings, builder unit-test fixtures, sample JSON,
run_compat.sh). Existing test logic already covers both manual
(`thinking.type=enabled`, Haiku 4.5) and adaptive
(`thinking.type=adaptive`, Opus 4.7) shapes because Claude Code picks
the shape per model from `--effort max`; the name change just stops
the column from looking like a Claude 3.7 reference.

New rows
--------
- structured_outputs (5 files, CLI `--json-schema`). Claude Code
  synthesizes a single `StructuredOutput` tool from the schema and
  surfaces the tool_use input as `structured_output` on the trailing
  `result` event. Test ships its own `_validate_against_schema` so
  we don't take a jsonschema dep just for matrix surface.

- count_tokens (5 files, HTTP probe). POSTs the proxy's
  `/v1/messages/count_tokens` directly and asserts the response is
  `{input_tokens: positive int}`. No CLI hook exists for this
  endpoint; the test goes through the new http_probe helper instead.

- tool_search (5 files, HTTP probe). Sends
  `tools: [{type: tool_search_tool_regex_20251119, name:
  tool_search_tool_regex}]` and asserts the proxy doesn't 400. MCP
  fan-out via `--mcp-config` would also exercise the tool-search
  beta header path, but it's flaky w.r.t. Claude Code's internal
  tool-deferral threshold; the HTTP probe hits the actual bug surface
  (per-provider beta-header translation `advanced-tool-use-2025-11-20`
  vs `tool-search-tool-2025-10-19`).

- long_context_1m (5 files, CLI `--betas context-1m-2025-08-07
  --max-budget-usd 6`). A ~210k-token padded prompt over stdin
  exercises the 1M-context beta. Sonnet 4.6 + Opus 4.7 only --
  Haiku 4.5's window is 200k, so it's excluded from MODELS (not
  marked not_applicable) to keep the per-cell aggregator semantics
  intact. Prompt uses a document-style preamble + 8 cycling pangrams
  rather than repeating identical chunks; without that, Opus 4.7
  trips the safety filter mid-response with a Usage Policy refusal.
  `--max-budget-usd 6` is a runaway-loop guard, ~2x worst-case Opus
  per-cell spend.

New helper
----------
`tests/claude_code/http_probe.py`: shared `ProbeResult` dataclass
plus per-endpoint `probe_*` + `assert_*_shape` pairs for the
HTTP-probe rows. Uses httpx with `anthropic-version: 2023-06-01` and
a 30s timeout.
2026-05-16 20:37:01 +00:00

103 lines
4.5 KiB
YAML

# Claude Code Compatibility Matrix — feature manifest.
#
# Defines the row order of the matrix and maps each feature_id to its
# human-readable display name. Adding a new feature to the matrix is a
# three-step change:
# 1. Append an entry to `features:` below.
# 2. Create a directory `tests/claude_code/<feature_id>/`.
# 3. Add per-provider test files inside that directory.
#
# `feature_id` MUST match the directory name on disk; the test harness
# infers (feature, provider) for each test from its file path.
schema_version: "1"
# Provider column order in the rendered matrix.
providers:
- anthropic
- bedrock_invoke
- bedrock_converse
- vertex_ai
- azure
# Feature row order.
features:
- id: basic_messaging_non_streaming
name: Basic messaging (non-streaming)
- id: basic_messaging_streaming
name: Basic messaging (streaming)
- id: tool_use
name: Tool use
- id: prompt_caching_5m
name: Prompt caching (5m TTL)
- id: vision
name: Vision
- id: thinking
name: Thinking
# The single row covers both API shapes Anthropic exposes — manual
# `thinking: {type: "enabled", budget_tokens: N}` (Haiku 4.5) and
# `thinking: {type: "adaptive"}` (Opus 4.7); Sonnet 4.6 supports
# either and Claude Code picks per model. A break in either
# transformer surfaces as a red cell because all three tiers must
# pass for the cell to go green. The row was named
# `extended_thinking` historically; Anthropic's docs now reserve
# that name for the deprecated manual mode only, so the row was
# renamed to the feature-level "Thinking".
- id: tool_use_streaming
name: Tool use (streaming / fine-grained)
- id: thinking_with_tool_use
name: Extended thinking + tool use
- id: pdf_input
name: PDF document input
- id: prompt_caching_1h
name: Prompt caching (1h TTL)
- id: web_search
name: Web search (server tool)
- id: structured_outputs
name: Structured outputs
# Drives `claude --json-schema '<schema>'`. Implementation note:
# Claude Code translates `--json-schema` to a synthetic
# `StructuredOutput` tool whose `input_schema` is the user's
# schema, then surfaces the tool_use input as
# `structured_output: {...}` on the trailing `result` event.
# This row tests that proxy-side handling of that tool round-
# trips end-to-end. It does NOT test Anthropic's server-side
# `output_config.schema` parameter (a separate feature used
# internally by Claude Code for session-title generation) --
# `output_config` regressions surface in the HTTP-probe rows.
- id: count_tokens
name: count_tokens endpoint
# HTTP-probe row. Sends a direct POST to
# `{proxy}/v1/messages/count_tokens` for each Claude tier and
# asserts the response is shaped `{"input_tokens": <positive
# int>}`. The CLI uses this endpoint internally but never
# surfaces its result in stream-json, so the only way to test
# the proxy's handling of it is to hit it directly. LiteLLM has
# shipped fixes here (e.g. Claude Code release-notes 2.1.121
# "Vertex AI count_tokens returning 400 errors for proxy
# gateways"), which is exactly the regression class this row
# is meant to catch.
- id: tool_search
name: Tool search (MCP discovery)
# HTTP-probe row. Sends a request whose `tools` array includes
# a `tool_search_tool_regex_20251119` discovery tool and asserts
# the proxy + upstream accept it. This verifies LiteLLM's
# per-provider beta-header translation
# (`advanced-tool-use-2025-11-20` for Anthropic/Azure,
# `tool-search-tool-2025-10-19` for Vertex/Bedrock) is wired up.
# We deliberately don't try to trigger Claude Code's MCP-fan-out
# heuristic via `--mcp-config` -- that would couple the row to
# an internal behavior threshold that changes between Claude
# Code releases. The HTTP probe hits the bug surface LiteLLM
# has actually shipped fixes for (2.1.117, 2.1.72, 2.1.70 per
# the Claude Code release notes).
- id: long_context_1m
name: Long context (1M)
# Sends a ~210k-token padded prompt with the
# `context-1m-2025-08-07` beta header. Just-above the standard
# 200k context window so the request can only succeed when the
# beta header makes it all the way through the proxy to the
# upstream. Haiku 4.5 is intentionally omitted from this row's
# model list (its window is 200k); Sonnet 4.6 and Opus 4.7 are
# the only tiers exercised. Costs roughly $4/cell/run --
# tighten the prompt-token target if pricing changes meaningfully.