mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-09 03:18:44 +00:00
Matrix grows from 11 to 15 feature rows. All new tests collected + 180 unit tests still pass; smoke runs hit real LiteLLM bug surfaces on bedrock_invoke, bedrock_converse, and vertex_ai (cells correctly red in PR #142). Rename ------ `extended_thinking` -> `thinking` (directory, manifest id+name, 5 test fn names, 5 docstrings, builder unit-test fixtures, sample JSON, run_compat.sh). Existing test logic already covers both manual (`thinking.type=enabled`, Haiku 4.5) and adaptive (`thinking.type=adaptive`, Opus 4.7) shapes because Claude Code picks the shape per model from `--effort max`; the name change just stops the column from looking like a Claude 3.7 reference. New rows -------- - structured_outputs (5 files, CLI `--json-schema`). Claude Code synthesizes a single `StructuredOutput` tool from the schema and surfaces the tool_use input as `structured_output` on the trailing `result` event. Test ships its own `_validate_against_schema` so we don't take a jsonschema dep just for matrix surface. - count_tokens (5 files, HTTP probe). POSTs the proxy's `/v1/messages/count_tokens` directly and asserts the response is `{input_tokens: positive int}`. No CLI hook exists for this endpoint; the test goes through the new http_probe helper instead. - tool_search (5 files, HTTP probe). Sends `tools: [{type: tool_search_tool_regex_20251119, name: tool_search_tool_regex}]` and asserts the proxy doesn't 400. MCP fan-out via `--mcp-config` would also exercise the tool-search beta header path, but it's flaky w.r.t. Claude Code's internal tool-deferral threshold; the HTTP probe hits the actual bug surface (per-provider beta-header translation `advanced-tool-use-2025-11-20` vs `tool-search-tool-2025-10-19`). - long_context_1m (5 files, CLI `--betas context-1m-2025-08-07 --max-budget-usd 6`). A ~210k-token padded prompt over stdin exercises the 1M-context beta. Sonnet 4.6 + Opus 4.7 only -- Haiku 4.5's window is 200k, so it's excluded from MODELS (not marked not_applicable) to keep the per-cell aggregator semantics intact. Prompt uses a document-style preamble + 8 cycling pangrams rather than repeating identical chunks; without that, Opus 4.7 trips the safety filter mid-response with a Usage Policy refusal. `--max-budget-usd 6` is a runaway-loop guard, ~2x worst-case Opus per-cell spend. New helper ---------- `tests/claude_code/http_probe.py`: shared `ProbeResult` dataclass plus per-endpoint `probe_*` + `assert_*_shape` pairs for the HTTP-probe rows. Uses httpx with `anthropic-version: 2023-06-01` and a 30s timeout.
94 lines
3.2 KiB
Python
94 lines
3.2 KiB
Python
"""count_tokens x Bedrock (Invoke).
|
|
|
|
HTTP-probe row. Unlike the CLI-driven rows, this test never invokes
|
|
the `claude` CLI: it `POST`s directly to
|
|
`{proxy}/v1/messages/count_tokens` for each Claude tier and asserts
|
|
the response is shaped `{"input_tokens": <positive int>}`.
|
|
|
|
The (feature, provider) for this cell is inferred from the file path by
|
|
`tests/claude_code/conftest.py`:
|
|
|
|
tests/claude_code/count_tokens/test_bedrock_invoke.py
|
|
^^^^^^^^^^^^ ^^^^^^^^^^^^^^
|
|
feature_id provider
|
|
|
|
Why HTTP probe instead of CLI:
|
|
|
|
Claude Code calls `count_tokens` internally to compute budget /
|
|
context-window usage display, but the result is consumed by the CLI
|
|
in-process and never appears in stream-json events. There is no CLI
|
|
flag that emits the count to stdout in a way our existing
|
|
stream-json parser can pick up, so we can't test the endpoint round
|
|
trip through the CLI surface.
|
|
|
|
The proxy *is* expected to expose `/v1/messages/count_tokens` for
|
|
every Claude-style provider it routes to -- LiteLLM has historically
|
|
had provider-specific bugs in this endpoint (Vertex AI `count_tokens`
|
|
returned 400 to proxy gateways; see Claude Code release notes 2.1.121).
|
|
Treating it as a matrix row keeps regressions in the cron's daily
|
|
diff.
|
|
|
|
The cell goes red if *any* tier's probe fails the minimal shape
|
|
check; the matrix's per-cell aggregator handles that automatically.
|
|
Three tiers run sequentially because count_tokens is cheap (<100ms
|
|
per request typical) and the parallelization that matters for the
|
|
CLI rows isn't useful here.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
|
|
import pytest
|
|
|
|
from tests.claude_code.http_probe import (
|
|
assert_count_tokens_shape,
|
|
probe_count_tokens,
|
|
)
|
|
|
|
PROXY_BASE_URL_ENV = "LITELLM_PROXY_BASE_URL"
|
|
PROXY_API_KEY_ENV = "LITELLM_PROXY_API_KEY"
|
|
|
|
BEDROCK_INVOKE_MODELS = [
|
|
"claude-haiku-4-5-bedrock-invoke",
|
|
"claude-sonnet-4-6-bedrock-invoke",
|
|
"claude-opus-4-7-bedrock-invoke",
|
|
]
|
|
|
|
|
|
def test_count_tokens_bedrock_invoke(compat_result):
|
|
"""Probe `/v1/messages/count_tokens` for each Bedrock (Invoke) tier and
|
|
assert the response shape."""
|
|
base_url = os.environ.get(PROXY_BASE_URL_ENV)
|
|
api_key = os.environ.get(PROXY_API_KEY_ENV)
|
|
if not base_url or not api_key:
|
|
compat_result.set(
|
|
{
|
|
"status": "fail",
|
|
"error": (
|
|
f"missing required env: set {PROXY_BASE_URL_ENV} and "
|
|
f"{PROXY_API_KEY_ENV} to point at a running LiteLLM proxy"
|
|
),
|
|
}
|
|
)
|
|
pytest.fail(
|
|
f"{PROXY_BASE_URL_ENV} / {PROXY_API_KEY_ENV} not configured",
|
|
pytrace=False,
|
|
)
|
|
|
|
failures = []
|
|
for model in BEDROCK_INVOKE_MODELS:
|
|
result = probe_count_tokens(
|
|
base_url=base_url, api_key=api_key, model=model
|
|
)
|
|
shape_error = assert_count_tokens_shape(result)
|
|
if shape_error is not None:
|
|
error = f"[{model}] count_tokens probe failed: {shape_error}"
|
|
compat_result.add({"status": "fail", "error": error})
|
|
failures.append(error)
|
|
continue
|
|
|
|
compat_result.add({"status": "pass"})
|
|
|
|
if failures:
|
|
pytest.fail("; ".join(failures), pytrace=False)
|