Commit graph

4 commits

Author SHA1 Message Date
mateo-berri
797a449daf RALPH: compat matrix slice 4 - daily cron VM publishes matrix to docs (#26480, PRD #26476)
Slice 4 of the Claude Code Compatibility Matrix: stand up the daily-cron
pipeline that publishes `compatibility-matrix.json` to the docs repo. After
this slice lands, the hand-authored matrix in the docs repo is replaced by
auto-generated output, and the docs page begins reflecting real test runs
against the latest stable LiteLLM release.

What landed:

- tests/claude_code/resolver.py
  Latest Stable LiteLLM Resolver. Calls the GitHub Releases API and
  returns the newest tag matching `v*-stable`. Sort is numeric on
  (major, minor, patch) so v1.10.0-stable correctly outranks
  v1.9.5-stable. Injectable `http_get` so tests run offline.

- tests/claude_code/publisher.py
  Daily-cron orchestrator. Resolves the latest stable tag, pulls
  `ghcr.io/berriai/litellm:<tag>`, starts it as the proxy, installs
  `@anthropic-ai/claude-code@latest`, runs `pytest tests/claude_code/`,
  invokes the Matrix JSON Builder, and direct-pushes
  `compatibility-matrix.json` to the docs repo's main branch using a
  GitHub App installation token (`DOCS_REPO_TOKEN`). Idempotent: a no-op
  if the JSON is byte-identical to what's already on main.

- tests/claude_code/_publisher_unit_tests/test_resolver.py
  test_publisher.py
  14 unit tests covering the small pure helpers — version sort,
  non-stable filtering, http-get injection, commit message determinism,
  Docker image-name builder, and the file allowlist that enforces the
  "only `compatibility-matrix.json` ever ships" guarantee. Per the PRD's
  "Testing Decisions" section, the publisher's full subprocess
  orchestration intentionally ships without a unit-test harness; the
  daily-cron failure surface is itself the test.

- .github/workflows/claude_code_compat_matrix.yml
  GitHub Actions workflow with three triggers (daily cron at 06:00 UTC,
  `release: published` filtered to `*-stable` tags, and
  `workflow_dispatch`). Mints a docs-repo installation token from a
  GitHub App scoped to `BerriAI/litellm-docs` only with `contents:
  write`, then runs the publisher.

- .gitignore
  Add `compatibility-matrix.json` (cron VM output).

Key decisions:

- "Isolated VM" is realized as a GitHub-hosted ubuntu-latest runner —
  every run gets a fresh ephemeral VM, and the always-latest Claude
  Code CLI is only ever installed inside that ephemeral environment,
  so a malicious or broken Claude Code release cannot affect the
  trusted PR-gate CI in CircleCI.
- File-level restriction on the GitHub App's broad `contents: write`
  scope is enforced by `select_files_to_commit` (script correctness),
  per the PRD's explicit acknowledgement that GitHub does not support
  file-path-scoped tokens.
- `release` runs are filtered to tags ending in `-stable` at the
  workflow level, so a `v1.84.0-rc1` release does not republish the
  matrix.
- Resolver and publisher live under `tests/claude_code/` alongside
  `matrix_builder.py` and `cli_driver.py` — production code that
  supports the test suite, kept colocated with it to match the slice
  1+2 layout.

Out of scope / blockers for next iteration:

- Provisioning the GitHub App itself (creating it under BerriAI's
  org, installing it on litellm-docs only, generating the private key
  and registering `COMPAT_MATRIX_APP_ID` / `COMPAT_MATRIX_APP_PRIVATE_KEY`
  as repo secrets) is an operator/infra step that cannot land via a
  code change in this repo.
- The first successful cron run is what removes the hand-authored
  `compatibility-matrix.json` from the docs repo and replaces it with
  generated output — that happens after this PR merges and the App is
  installed; not a code change here.

Tests: 34 -> 45 passing (added 7 resolver tests + 7 publisher helper
tests, all unit-only and offline). The 12 per-cell failures under
`tests/claude_code/basic_messaging_non_streaming/` remain by design —
they require a running proxy which the cron VM provides.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-06 23:27:14 +00:00
mateo-berri
e1fc6a8ffc RALPH: compat matrix slice 3 - wire PR gate in CircleCI (#26479, PRD #26476)
Slice 3 of the Claude Code Compatibility Matrix: wire the
`tests/claude_code/` suite into CircleCI as a pre-merge gate. A red
status on the new `claude_code_compat_pr_gate` job blocks merge into
the staging branch.

What landed:

- tests/claude_code/pr_gate_version_resolver.py
  The Claude Code PR-Gate Version Resolver described in the PRD's
  "Version resolvers" section. Queries the npm registry for
  `@anthropic-ai/claude-code` and returns the newest version whose
  publish timestamp is at least 3 days old. The 3-day window is a
  security review buffer: a malicious or broken Claude Code release
  has at least 72 hours to be detected before it can land in our PR
  gate. Importable function (with `metadata=` / `fetcher=` / `as_of=`
  injection seams for tests) and a `python -m ...` CLI for the CI step.

- tests/claude_code/test_config.yaml
  Proxy routing config that maps the per-cell aliases the tests use
  (`claude-haiku-4-5`, `claude-haiku-4-5-bedrock-invoke`, ...,
  `claude-opus-4-7-vertex`) to real upstream model ids on Anthropic /
  Bedrock (Invoke + Converse) / Vertex AI. Azure intentionally has no
  entries here because every Azure × claude-code cell is
  `not_applicable` (Azure OpenAI doesn't host Claude).

- .circleci/config.yml
  New `claude_code_compat_pr_gate` job. Pattern modeled on
  `proxy_e2e_anthropic_messages_tests` (load PR-built docker image,
  start postgres, mount config.yaml). New step in the middle:
  resolve the Claude Code version from the resolver, install Node 20
  via the machine image's preinstalled nvm, and `npm install -g
  @anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}` (pinned, never
  `latest`). Wired into `workflows.build_and_test` with a
  `requires: [build_docker_database_image]` gate and the same
  `*main_branches` filter the other proxy e2e job uses.

- tests/claude_code/_pr_gate_unit_tests/
  16 new unit tests:
  * 8 against the version resolver: boundary (>= 3d inclusive),
    empty / all-too-new metadata, semver-vs-publish-time tiebreak,
    custom min_age, fetcher injection, npm `time.created` /
    `time.modified` skipping.
  * 8 structural tests against `.circleci/config.yml`: job exists,
    is in the workflow, requires the docker image, invokes the
    resolver, install command is pinned (rejects unpinned `latest`),
    runs `tests/claude_code/`, mounts `test_config.yaml`, exports
    `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY`. Plus one
    regression test: the existing `proxy_e2e_anthropic_messages_tests`
    job is unchanged in shape (acceptance criterion).

Key decisions:

- "Newest version" in the resolver is by **publish time**, not by
  semver string ordering — if a patch lands on an older major after a
  newer release, the patched line is the eligible one. (Tested.)
- The resolver's CLI prints the announcement to stderr and the bare
  version to stdout, so the CI step can do
  `CLAUDE_CODE_VERSION=$(uv run python -m ...)` cleanly while still
  surfacing the selected version in the job log (acceptance criterion:
  "the selected Claude Code version is logged").
- The structural CircleCI tests live under `_pr_gate_unit_tests/` so
  the conftest path-inference hook skips them (the leading underscore
  is the existing convention from `_driver_unit_tests/` /
  `_builder_unit_tests/`); they don't pollute the matrix artifact.
- No `--no-verify` style supply-chain safety relaxation. Per the PRD,
  Claude Code's pinning is the 3-day publish-age window, not a fixed
  hash — by design, since the daily cron also pulls newer versions.

Tests: 47 -> 47 passing for the unit suite (16 new + 31 from slices
1 and 2). The end-to-end cells under `basic_messaging_non_streaming/`
require `LITELLM_PROXY_BASE_URL` / `LITELLM_PROXY_API_KEY` and a
running proxy + `claude` CLI; they only run inside the new CircleCI
job.

Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- No docs PR is needed for this slice — the gate produces a status
  check, not a published artifact. The compat matrix JSON the docs
  page consumes is published by the daily-cron job (a future slice),
  not by the PR gate.

Notes for next iteration:
- The daily cron / matrix publisher is the next slice. Several
  pieces this slice introduces (the `tests/claude_code/test_config.yaml`
  proxy config, the structure of the compat-results.json artifact)
  will be reused by it.
- The bedrock-converse / vertex_ai aliases in `test_config.yaml` use
  best-guess upstream model ids (`us.anthropic.claude-{tier}` and
  `vertex_ai/claude-{tier}`); the real ids may need to be tightened
  once the gate runs against live AWS / GCP credentials and we see
  what resolves.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-06 23:27:05 +00:00
mateo-berri
415cf3d6b1 RALPH: compat matrix slice 2 - add 4 provider columns for basic_messaging_non_streaming (#26478, PRD #26476)
Slice 2 of the Claude Code Compatibility Matrix: extend the tracer-bullet
cell from slice 1 across all four remaining provider columns for
basic_messaging_non_streaming. Proves the multi-provider, multi-model,
all-must-pass aggregation logic against a 1x5 grid that exercises every
status state.

What landed:

- tests/claude_code/basic_messaging_non_streaming/test_bedrock_invoke.py
- tests/claude_code/basic_messaging_non_streaming/test_bedrock_converse.py
- tests/claude_code/basic_messaging_non_streaming/test_vertex_ai.py
  Per-provider files modeled on test_anthropic.py: each parametrizes
  over Haiku 4.5 / Sonnet 4.6 / Opus 4.7 (the three Claude tiers
  required by the PRD), drives the real `claude` CLI through the
  driver, and reports pass/fail via `compat_result`. Per-cell error
  strings always include `[<model>]` so the docs tooltip can name the
  failing model when a cell goes red.

- tests/claude_code/basic_messaging_non_streaming/test_azure.py
  All three (Azure, Claude) cells report `not_applicable` with a
  reason: Azure OpenAI Service does not host Anthropic models. The
  test still parametrizes over the same three model ids so the test
  count per cell is uniform across columns, and a future "Azure adds
  Anthropic" announcement only requires flipping the body, not the
  parametrization.

- tests/claude_code/sample_compatibility-matrix.json
  Hand-authored 1x5 sample updated to reflect the slice 2 outcome:
  anthropic / bedrock_invoke / bedrock_converse / vertex_ai = pass,
  azure = not_applicable.

- tests/claude_code/_builder_unit_tests/test_matrix_builder.py
  Two new golden-file tests:
  1. 1x5 grid: feed the per-model results the four new test files
     produce on a real run; assert the builder output equals the
     hand-authored sample byte-for-byte.
  2. fail-with-model-named: feed pass/fail/pass for one cell and assert
     the cell aggregates to fail with the failing model id surfaced
     in the error string (acceptance criterion: "the error string
     identifies which model broke").

Key decisions:

- Duplication across the four per-provider files is accepted (per the
  PRD) rather than extracted into a helper. Each file is self-contained
  so a test author touching one provider doesn't accidentally regress
  the others.
- Per-provider model alias names: `claude-<tier>-<provider-suffix>`
  (e.g. `claude-haiku-4-5-bedrock-invoke`). These are the alias names
  the proxy operator wires up in the routing config; the test only
  knows the alias, the proxy knows the upstream model id and region.
- Azure is `not_applicable` rather than `not_tested` because the
  cell will never apply, not "we haven't gotten to it yet" - the two
  states are visually and semantically distinct in the rendered grid.
- Sample shows the realistic best-case outcome (4 pass + 1 NA). The
  React renderer's coverage of the `fail` and `not_tested` states is
  exercised by other cells in v1+, not the v0 sample.

Tests: 31 -> 34 passing (added 2 builder golden tests + 3 Azure
not_applicable parametrizations that pass without env vars).

Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- The companion update to compatibility-matrix.json in the docs repo.
  The hand-authored sample in this repo is the artifact the docs PR
  copies; opening that doc PR is the next step in this slice.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-06 23:27:05 +00:00
mateo-berri
6c573de426 RALPH: tracer-bullet for Claude Code compatibility matrix (#26477, PRD #26476)
Slice 1 of the Claude Code Compatibility Matrix: the thinnest end-to-end
path through every layer for a single (feature, provider) cell, so a
future docs page can render a real green cell sourced from a real test.

What landed in this repo:

- tests/claude_code/manifest.yaml — feature manifest with one entry
  (basic_messaging_non_streaming) plus the v0 provider column order.
- tests/claude_code/cli_driver.py — Claude Code CLI Driver. One entry
  point (run_claude); handles subprocess assembly, env overlay, stream-JSON
  parsing, and structured failure modes. `runner=` is a unit-test seam.
- tests/claude_code/conftest.py — `compat_result` fixture (tagged-union
  recorder) + pytest_runtest_makereport hook that infers (feature, provider)
  from the file path and writes a structured compat-results.json artifact.
- tests/claude_code/basic_messaging_non_streaming/test_anthropic.py — the
  one cell, parametrized over Haiku/Sonnet/Opus per the PRD's per-cell
  model rule.
- tests/claude_code/matrix_builder.py — pure-function builder from
  (manifest, results, run-metadata) to the v1 JSON schema. Aggregates per-
  model results into one cell (pass iff all pass). build_from_paths is the
  thin I/O wrapper for the publisher.
- tests/claude_code/sample_compatibility-matrix.json — hand-authored sample
  of the v1 JSON; copied to the docs repo by hand as part of this slice.
- Unit tests: 10 driver tests (mocked subprocess), 9 compat_result tests,
  10 matrix-builder golden-file tests. 29/29 pass.

Key decisions:

- (feature, provider) is inferred from file path, not declared in metadata —
  mirrors the PRD's "no drift" goal.
- Driver injects subprocess via a `runner` kwarg so unit tests don't need
  the real `claude` CLI; production callers leave it default.
- Builder is a pure function on Mappings/Sequences; load/write live in a
  thin `build_from_paths` wrapper. Golden-file tests pin the schema.
- `_driver_unit_tests/` and `_builder_unit_tests/` are prefixed with `_`
  so the conftest's path-inference hook skips them and they don't
  pollute the matrix artifact.
- `compat-results.json` added to .gitignore (CI-only output).

Out of scope per CLAUDE.md (docs live in BerriAI/litellm-docs):
- The MDX page `docs/tutorials/claude-code-compatibility` and the
  `<CompatibilityMatrix />` React component. The hand-authored
  compatibility-matrix.json (`sample_compatibility-matrix.json` in this
  repo) is the artifact those docs files will consume; opening that doc
  PR is the next step in this slice.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-06 23:27:05 +00:00