Commit graph

8 commits

Author SHA1 Message Date
mateo-berri
3e7ecdd152
fix(cron_vm): greptile — scrub provider secrets from pytest invocation env
Wrap the pytest call in run_daily.sh in `env -i` with the same minimal
allowlist the PR-gate already uses, mirroring the CircleCI step. The
systemd EnvironmentFile injects ANTHROPIC_API_KEY /
AWS_BEARER_TOKEN_BEDROCK / VERTEXAI_* / AZURE_FOUNDRY_* /
AGENT_SHIN_GITHUB_TOKEN / GITHUB_TOKEN into the script for the proxy to
consume; pytest inherits them by default but only needs the loopback
proxy URL/key. Scrubbing them closes two paths: PR-controlled test code
reading them out of os.environ, and model-directed Read tool calls
reaching /proc/<pytest-pid>/environ during PDF/vision cells. Add a pin
test mirroring the existing version-probe one.
2026-05-19 06:40:17 +00:00
Cursor Agent
5ad351bbd2
fix(cron_vm): veria — isolate $HOME and hide credential dotdirs from claude
The cron systemd unit's `ProtectHome=read-only` blocks writes to
/home/mateo but still allows reads. With `HOME=/home/mateo` forwarded
to the `claude` subprocess, a compromised @anthropic-ai/claude-code
release (running during the `claude --version` probe) — or a
model-directed `Read` tool call during a PDF cell (which passes
`--allowed-tools Read`) — could read host credential files like
~/.config/gh/hosts.yml (gh-host token), ~/.ssh/, or ~/.bash_history
and exfiltrate them.

Two complementary mitigations, addressing veria's exact recommendation:

1. Per-invocation isolated HOME for every `claude` subprocess:
   * cli_driver.py: drop HOME from _CLI_ENV_ALLOWLIST; create a
     fresh empty tmpdir under tempfile.gettempdir() (`PrivateTmp=true`
     keeps it on a service-private tmpfs) and pass it as HOME to
     each `claude` invocation. Cleaned up in a `finally` so
     timeouts and CLI-not-found don't leak tmpdirs.
   * run_daily.sh: the up-front `claude --version` probe also runs
     under $CLAUDE_PROBE_HOME (a per-run dir under ${WORKDIR}) so
     the probe can never reach the runtime user's real home; the
     existing `cleanup` trap removes ${WORKDIR}.
   * Closes the `os.path.expanduser('~/.config/gh/hosts.yml')`-style
     attack from a compromised CLI / model.

2. Filesystem-level hiding of credential dotdirs in the systemd unit:
   * Add `InaccessiblePaths=-/home/mateo/.config/gh -/home/mateo/.ssh
     -/home/mateo/.aws -/home/mateo/.docker -/home/mateo/.kube
     -/home/mateo/.gnupg`. The kernel hides these paths from every
     process in the unit's mount namespace, defeating the absolute-path
     attack (`Read('/home/mateo/.config/gh/...')`) that the per-
     invocation HOME override alone cannot block.
   * Drop `/home/mateo/.config/gh` from `ReadWritePaths=` (it's
     now hidden, and we pass GH_TOKEN inline to every `gh` call).
   * Pass GH_TOKEN inline to `gh repo clone` in run_daily.sh
     (was relying on host gh-cli config); the docs repo is public
     so this is a no-op functionally, but it lets us drop the
     ~/.config/gh dependency entirely.

Tests:
  * test_run_claude_uses_isolated_per_invocation_home: pin that the
    CLI subprocess never sees the parent's $HOME, and that the
    isolated HOME is a fresh tmpdir prefixed claude-cli-home-.
  * test_run_claude_isolated_home_is_distinct_per_invocation: pin that
    each call gets its own dir (no cross-call planting).
  * test_run_claude_isolated_home_cleaned_up_after_run / on_subprocess
    _failure: pin that the tmpdir is rm-rf'd on both the happy path
    and the timeout/CLI-error path.
  * test_version_probe_uses_isolated_home_not_runtime_user_home: pin
    that run_daily.sh's probe forwards $CLAUDE_PROBE_HOME, not
    ${HOME}, into its `env -i` block.
  * test_systemd_unit_credential_isolation.py (new): pin that
    InaccessiblePaths covers all credential dotdirs, that
    .config/gh is not under ReadWritePaths, and that ProtectHome
    stays at least read-only.

All 349 existing claude_code unit tests still pass.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-05-19 05:49:18 +00:00
mateo-berri
0cd506f24d
fix(cron_vm): veria — scrub provider secrets from claude --version probe env 2026-05-19 03:29:05 +00:00
Cursor Agent
c23b1bba8c
fix(cron): bugbot — paginate all release pages to pick highest-semver stable
The release pagination loop in run_daily.sh used to break the moment a
page contained any v*-stable tag. GitHub's /releases endpoint orders by
created_at, not semver, so a freshly-cut backport on an older series
(e.g. v1.80.1-stable published today) can appear on an earlier page than
a higher-versioned release (v1.83.0-stable published two weeks ago).
The early-break would silently pin the cron to the stale tag because
the higher-versioned release on a later page never made it into the
merged set the final sort_by consumed — and the cron would publish a
compatibility matrix against a stale LiteLLM version with no visible
signal that anything was wrong.

Keep the empty-page guard (so a quiet release feed still doesn't burn
through the full 5-page cap) but drop the broken early-break.

Tests live in tests/claude_code/_publisher_unit_tests/ (mirroring the
existing _driver_unit_tests / _builder_unit_tests / _pr_gate_unit_tests
naming convention already excluded from the PR-gate pytest run). They:
- Statically assert the buggy length>0 + break combo is not in the
  pagination loop body.
- Statically assert the empty-page guard is still in place.
- Drive the actual run_daily.sh resolution snippet with a fake curl
  whose page 1 contains a low-version backport stable and page 2
  contains the high-version stable, then assert that the high-version
  tag is the one resolved. This is the end-to-end regression test.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-05-17 21:39:36 +00:00
mateo-berri
bb60422e7d compat-matrix: replace publisher.py + resolver.py with a bash run_daily.sh
The previous iteration of this PR ported the populator to a Python
module (`publisher.py`) with 8 unit-tested pure helpers for branch
naming and PR-body rendering. After the docker code came out, the
GitHub App auth came out, and the worktree-vs-tempdir decision was
made, what was left was: 'git fetch + git checkout + uv sync + start
a subprocess + run pytest + clone docs + git commit + gh pr create'.
That's a bash script.

This commit replaces 800 lines of Python (publisher.py + resolver.py +
their unit tests) with a 297-line run_daily.sh and a 48-line
build_matrix.py whose only job is to be importable Python that can
call into the matrix_builder we already have. Net deletion: -822 lines.

Removed
-------

  * tests/claude_code/publisher.py — the full Python orchestrator.
    Every code path it had is now in run_daily.sh.
  * tests/claude_code/resolver.py — the GitHub Releases v*-stable
    resolver. Replaced by ~10 lines of jq inside run_daily.sh.
  * tests/claude_code/_publisher_unit_tests/ — both test files. The
    pure helpers they covered (commit message, file allowlist, branch
    name, PR title/body) only existed because publisher.py was Python.
    The bash equivalents are short heredoc strings.

Added
-----

  * tests/claude_code/cron_vm/run_daily.sh — the actual cron job,
    structured as numbered phases (resolve / worktree / proxy /
    pytest / build / publish) so journalctl output is readable.
  * tests/claude_code/cron_vm/build_matrix.py — a 48-line CLI that
    calls the existing matrix_builder.build_from_paths. Kept in
    Python because the builder itself is Python and well-tested.

Modified
--------

  * tests/claude_code/cron_vm/litellm-compat-matrix.service —
    ExecStart now invokes run_daily.sh instead of
    'python -m tests.claude_code.publisher'.
  * tests/claude_code/cron_vm/README.md — updated layout table,
    file roles, and operating commands to match.

Why this is the right shape
---------------------------

  * The failure mode at 06:00 UTC is 'read journalctl, see the literal
    failing command with its + prefix, copy-paste it into a shell to
    reproduce'. Bash makes that immediate; Python's subprocess.run
    output looks similar but the surrounding orchestration is harder
    to step through interactively.
  * Every operation the script does is already a shell command (git,
    uv, gh, jq, curl, pytest). The Python wrapper was translating
    between argv arrays and back.
  * The two pieces that genuinely benefit from being in a typed
    language are matrix_builder (already Python) and the resolver's
    semver sort (now done in jq, with the version_key tuple sort
    inline). 'Already Python' wins, 'tiny jq pipeline' wins.

What's preserved
----------------

  * Idempotency: same (litellm, claude, UTC date) -> same branch ->
    same PR. force-with-lease push, gh-pr-create no-op-on-exists.
  * Byte-identical-JSON early return (git diff --cached --quiet).
  * Per-feature status table in the PR body (jq pipeline mirroring
    the Python pr_body_for_matrix logic).
  * Persistent worktree approach so disk doesn't grow unboundedly.
  * Proxy bound to :4100 to avoid colliding with a developer's :4000.
  * SKIP_PUBLISH=1 and PYTEST_K=... operator escape hatches.
2026-05-06 23:27:14 +00:00
mateo-berri
449a44dba3 compat-matrix: run from a GCP VM via systemd; drop docker + GHA
The daily compat-matrix runs from the dedicated GCP VM
`litellm-compatibility-matrix-populator` rather than from a GitHub
Actions runner. The VM has no docker daemon, has `gh` already
authenticated against an account with `pull-requests: write` on
`BerriAI/litellm-docs`, and runs a long-lived litellm checkout we can
reuse across runs. That makes Docker, the GitHub App auth flow, and the
GHA workflow itself dead code.

Removed
-------

  * `.github/workflows/claude_code_compat_matrix.yml` — no longer
    triggers anything; the systemd timer in this PR owns the daily fire.
  * `docker_image_for_tag` + `DOCKER_IMAGE_BASE` constants and the two
    unit tests that covered them.
  * `_start_proxy(image, port)` / `_stop_proxy(container_id)` /
    `docker run` flow, replaced by direct `uv run litellm` subprocess
    management with a sigterm-the-process-group teardown.
  * `--skip-proxy` CLI flag (was only useful when the GHA workflow
    split docker-bringup from publish into separate jobs).
  * `docs_token` parameter and `DOCS_REPO_TOKEN` env var; `gh` on
    the VM is already authenticated, so we don't pass an explicit
    token through the publisher.

Added
-----

  * Persistent worktree flow in `publisher.py`. First run clones
    `BerriAI/litellm` into `~/litellm-cron-worktree/`; subsequent runs
    do `git fetch --tags && git checkout --force <stable-tag> &&
    uv sync --frozen`. Disk footprint is bounded because uv sync
    removes packages no longer pinned and `git clean -fdx -e .venv`
    wipes per-run cruft while keeping the venv around.
  * `tests/claude_code/cron_vm/` containing systemd units and a
    setup README:
    - `litellm-compat-matrix.service` (`Type=oneshot`, runs as the
      `mateo` user, sources `/etc/litellm-compat-matrix.env` for
      provider creds, hardened with `NoNewPrivileges` /
      `ProtectSystem=strict` / `PrivateTmp`);
    - `litellm-compat-matrix.timer` (`OnCalendar=*-*-* 06:00:00 UTC`,
      `Persistent=true` so a missed run fires when the VM is back up,
      `RandomizedDelaySec=10min`);
    - `.env.example` documenting the provider-credential surface;
    - `README.md` covering one-time install, daily operation,
      `journalctl` debugging, and the gotchas (proxy port `4100` to
      avoid colliding with a developer's `:4000`, `gh` token
      rotation, what to do after a Claude Code CLI upgrade).

Operator notes
--------------

  * The proxy now binds `:4100` by default so a developer SSH'd into
    the VM with their own `:4000` proxy isn't preempted by the cron.
  * The Claude Code CLI is exercised as-is from the system install;
    the populator does NOT `npm install` it. Operators upgrade the
    CLI by running `npm install -g @anthropic-ai/claude-code@latest`
    out of band, typically after watching a `--skip-publish` run to
    verify the matrix doesn't suddenly turn red.
  * 20 publisher unit tests pass (`pytest
    tests/claude_code/_publisher_unit_tests/`).
  * End-to-end validation on the VM happens after this PR lands as
    follow-up commits on the same branch — the systemd unit is
    `Type=oneshot` so a manual `systemctl start` reproduces the cron.
2026-05-06 23:27:14 +00:00
mateo-berri
0e65ce1495 compat-matrix: open a docs-repo PR instead of direct-pushing main
The daily Claude Code compatibility-matrix cron has been direct-pushing
`compatibility-matrix.json` to litellm-docs's main branch. Switch to
opening (or updating) a pull request so docs maintainers can review each
matrix update before it ships to readers.

Behavioural changes
-------------------

publisher.publish() now:
  * checks out a deterministic head branch
    (`compat-matrix/<litellm>-<claude>-<UTC-date>`) before staging the
    JSON, instead of committing on top of the docs branch directly;
  * `git push --force-with-lease` so a same-day rerun updates the
    existing branch (and therefore the existing PR), without
    overwriting any docs-maintainer fixup commit on the same branch;
  * shells out to `gh pr create` against `docs_repo` with a
    title/body that surfaces the resolved versions and a per-feature
    status summary, so reviewers can triage from the inbox;
  * treats 'a pull request for branch ... already exists' as success,
    so two cron runs on the same day produce one PR, not two.

Idempotency contract
--------------------

  * Same (litellm_version, claude_code_version, UTC date) -> same
    branch -> same PR. Verified by the new
    `test_pr_branch_name_is_deterministic_per_inputs` /
    `...changes_when_any_component_changes` tests.
  * Byte-identical JSON to the docs branch -> early-return before
    push, same as the previous direct-push path.
  * Empty version inputs are rejected up front so two distinct PRs
    can never silently collapse onto one branch.

Tests
-----

  * 8 new tests in `_publisher_unit_tests/test_publisher.py` cover
    `pr_branch_name`, `pr_title_for_matrix`, and `pr_body_for_matrix`
    (determinism, content, ordering, missing-provider rectangularity,
    empty-input rejection).
  * Existing 7 `commit_message_for_matrix` /
    `docker_image_for_tag` / `select_files_to_commit` tests are
    unchanged and still pass.

Workflow
--------

`.github/workflows/claude_code_compat_matrix.yml` updates only the
header doc comment to reflect that the GitHub App now needs
`pull-requests: write` in addition to `contents: write`. `gh` is
preinstalled on `ubuntu-latest` (also used by
`auto_update_price_and_context_window.yml`), so no install step is
needed.

Operator action required (one-time)
-----------------------------------

The compat-matrix GitHub App installation on `BerriAI/litellm-docs`
needs `pull-requests: write` added to its installation permissions
before the next cron run. Without it, the new `gh pr create` call
will fail with a 403; `compat-results.json` and
`compatibility-matrix.json` will still upload as workflow artifacts
for debugging.
2026-05-06 23:27:14 +00:00
mateo-berri
797a449daf RALPH: compat matrix slice 4 - daily cron VM publishes matrix to docs (#26480, PRD #26476)
Slice 4 of the Claude Code Compatibility Matrix: stand up the daily-cron
pipeline that publishes `compatibility-matrix.json` to the docs repo. After
this slice lands, the hand-authored matrix in the docs repo is replaced by
auto-generated output, and the docs page begins reflecting real test runs
against the latest stable LiteLLM release.

What landed:

- tests/claude_code/resolver.py
  Latest Stable LiteLLM Resolver. Calls the GitHub Releases API and
  returns the newest tag matching `v*-stable`. Sort is numeric on
  (major, minor, patch) so v1.10.0-stable correctly outranks
  v1.9.5-stable. Injectable `http_get` so tests run offline.

- tests/claude_code/publisher.py
  Daily-cron orchestrator. Resolves the latest stable tag, pulls
  `ghcr.io/berriai/litellm:<tag>`, starts it as the proxy, installs
  `@anthropic-ai/claude-code@latest`, runs `pytest tests/claude_code/`,
  invokes the Matrix JSON Builder, and direct-pushes
  `compatibility-matrix.json` to the docs repo's main branch using a
  GitHub App installation token (`DOCS_REPO_TOKEN`). Idempotent: a no-op
  if the JSON is byte-identical to what's already on main.

- tests/claude_code/_publisher_unit_tests/test_resolver.py
  test_publisher.py
  14 unit tests covering the small pure helpers — version sort,
  non-stable filtering, http-get injection, commit message determinism,
  Docker image-name builder, and the file allowlist that enforces the
  "only `compatibility-matrix.json` ever ships" guarantee. Per the PRD's
  "Testing Decisions" section, the publisher's full subprocess
  orchestration intentionally ships without a unit-test harness; the
  daily-cron failure surface is itself the test.

- .github/workflows/claude_code_compat_matrix.yml
  GitHub Actions workflow with three triggers (daily cron at 06:00 UTC,
  `release: published` filtered to `*-stable` tags, and
  `workflow_dispatch`). Mints a docs-repo installation token from a
  GitHub App scoped to `BerriAI/litellm-docs` only with `contents:
  write`, then runs the publisher.

- .gitignore
  Add `compatibility-matrix.json` (cron VM output).

Key decisions:

- "Isolated VM" is realized as a GitHub-hosted ubuntu-latest runner —
  every run gets a fresh ephemeral VM, and the always-latest Claude
  Code CLI is only ever installed inside that ephemeral environment,
  so a malicious or broken Claude Code release cannot affect the
  trusted PR-gate CI in CircleCI.
- File-level restriction on the GitHub App's broad `contents: write`
  scope is enforced by `select_files_to_commit` (script correctness),
  per the PRD's explicit acknowledgement that GitHub does not support
  file-path-scoped tokens.
- `release` runs are filtered to tags ending in `-stable` at the
  workflow level, so a `v1.84.0-rc1` release does not republish the
  matrix.
- Resolver and publisher live under `tests/claude_code/` alongside
  `matrix_builder.py` and `cli_driver.py` — production code that
  supports the test suite, kept colocated with it to match the slice
  1+2 layout.

Out of scope / blockers for next iteration:

- Provisioning the GitHub App itself (creating it under BerriAI's
  org, installing it on litellm-docs only, generating the private key
  and registering `COMPAT_MATRIX_APP_ID` / `COMPAT_MATRIX_APP_PRIVATE_KEY`
  as repo secrets) is an operator/infra step that cannot land via a
  code change in this repo.
- The first successful cron run is what removes the hand-authored
  `compatibility-matrix.json` from the docs repo and replaces it with
  generated output — that happens after this PR merges and the App is
  installed; not a code change here.

Tests: 34 -> 45 passing (added 7 resolver tests + 7 publisher helper
tests, all unit-only and offline). The 12 per-cell failures under
`tests/claude_code/basic_messaging_non_streaming/` remain by design —
they require a running proxy which the cron VM provides.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-06 23:27:14 +00:00