Commit graph

39004 commits

Author SHA1 Message Date
Ishaan Jaff
b13deccdac
fix: ASCII-only AMI description (AWS rejects em-dash)
EC2 `ModifyImageAttribute` rejects "Character sets beyond ASCII are not
supported" when registering the AMI description. The whole AMI rolls back
on this error. Replace em-dash with hyphen.

LIT-2878
2026-05-06 15:36:29 -07:00
Ishaan Jaff
5847e7bf32
docs: bake real AMI id (ami-074a518157fe137b4) into example config
Built by `packer build` against the BYOC PoC account (us-west-2). Customers
running their own BYOC account should re-run the Packer build and replace
this value.

LIT-2878
2026-05-06 15:32:46 -07:00
Ishaan Jaff
26de269add
fix: rename env_creds in get_team_vm_config to satisfy mypy
`creds` was reassigned from `AwsCreds` to `Optional[AwsCreds]` in the
fallback branch — rename the second binding so mypy can narrow the type.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
b25dd8b716
fix: type-narrow session_id in _terminate_and_mark (mypy)
Defensive: if both `session_id` and `id` are missing, return early. The
str-cast keeps mypy happy when the row uses a UUID-typed column.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
2dc9b255f2
test: register pytest 'slow' marker for real-cloud agent_session tests
Avoids the PytestUnknownMarkWarning when collecting
`tests/test_litellm/proxy/agent_session_endpoints/vm_providers/test_ec2_provider_real.py`
without `-m slow`.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
88a6d0e242
fix: AMI build — disable apt cnf-update-db post-invoke hook + don't remap python3
The cnf-update-db post-invoke hook fails (exit 100) when we install a
non-default python3 alongside, because cnf-update-db imports apt_pkg which
is bound to /usr/bin/python3 -> python3.12. Disabling the hook makes apt
update + install idempotent for AMI builds.

Also stop remapping /usr/bin/python3 via update-alternatives — it breaks
Ubuntu's python-coupled apt tooling. Tools that need 3.13 invoke it
explicitly; the systemd unit already uses /usr/bin/python3.13.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
48c8ba1fe7
test: real-cloud EC2Provider tests — slow-marked, gated by BYOC env vars
Validations #3 (real boot), #11 (real BYOC fail-fast), and the real-cloud
piece of #8. Skipped by default; enable with `pytest -m slow` when the
`LITELLM_AGENT_AWS_*` + `LITELLM_TEST_*` env vars are set.

Each test wraps RunInstances in try/finally with TerminateInstances and
installs a 60-min process watchdog (per the AWS safety boundary) so a
hung test cannot leak an instance.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
0b3cb3040d
test: sweepers — bootstrap-timeout, heartbeat-loss, max-session-minutes (Validation #7, #9, #10)
Mocked Prisma client driving each sweeper through its happy path plus the
optimistic-lock branch (sweeper skips a session whose status changed
underneath it).

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
35f174b392
test: BYOC creds resolver — encrypt/decrypt round trip, cross-team isolation, env fallback
Covers validation #11 (no creds = fail-fast, no instance launched), #12
(two teams resolve to two distinct creds objects), and the env-var fallback
path used in local dev before the LiteLLM_AgentVMConfig table is populated.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
d519852165
test: EC2Provider unit tests — spot fallback, invalid-creds, creds-no-leak
Mocked tests using a fake boto3 client. Covers validations #5 (spot →
on-demand fallback), #8 (terminate idempotent on already-gone), #11
(invalid creds raise InvalidCredentialsError without launching anything),
and #13 (AWS keys never appear in log records, repr, or exception messages).

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
a149bd6bb7
test: NoopProvider lifecycle + idempotent terminate
LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
0b6be2ac1d
test: factory tests — provider swap is config-only, unknown values rejected (Validation #1, #6)
LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
10a5d25e95
test: scaffold test packages for agent_session_endpoints (LIT-2878) 2026-05-06 15:31:57 -07:00
Ishaan Jaff
ba2edee79d
docs: AMI build + boot-mode + cleanup playbook (infra/ami/README.md)
Covers: prerequisites (Packer + AWS profile), `packer init` + `packer
build` invocation, sharing the AMI cross-account via `ami_users`, the
`LITELLM_AGENT_MODE` boot-mode contract, the Epic C migration path, and
the leak-cleanup one-liner that targets `LitellmManagedBy=agent-vm-provider`.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
2535a437c6
feat: daemon stub — placeholder until Epic C (LIT-2879) ships the real daemon
Behaviour:
- reads runtime config from systemd's EnvironmentFile
- session mode: bootstrap → heartbeat every 30s, exit cleanly on HTTP 410
- warm mode: idle (real warm-pool hydrate lands in B2)
- redacts JWT in log output

Zero non-system deps (only `requests`, installed by the AMI builder).
Replaced wholesale by Epic C.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
03de8f56a6
feat: systemd unit for litellm-agent-runtime daemon
Reads runtime config from /etc/litellm-agent/runtime.env (written by
EC2 user-data, mode 600). Restart=on-failure with a 5s backoff so transient
network blips during bootstrap don't permanently kill the session.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
29667d22cd
feat: AMI runtime install script (node24, python3.13, git, gh, uv, bun)
Provisioner script driven by `litellm-agent-runtime.pkr.hcl`. Each tool is
pinned to a specific version and verified against a SHA-256 sidecar where
upstream provides one (uv) — see CLAUDE.md "CI Supply-Chain Safety".
The bun installer has no checksum sidecar, so we pin a version and pull the
artifact directly (not the install script).

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
0ea60840f6
feat: Packer config for litellm-agent-runtime AMI
Builds an Ubuntu 24.04 AMI with node 24, python 3.13, git, gh, uv, bun, and
the agent-runtime systemd unit autostarted on boot. The daemon honours
`LITELLM_AGENT_MODE` from EC2 user-data: `session` for cold-boot,
`warm` for warm-pool prewarming (B2).

Uses IMDSv2 only. Tags every resource Packer creates for easy cleanup.
`ami_users` lets us share the AMI cross-account without rebuilding.

Validation #2 (`packer build`) covers this file.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
82c302453d
docs: example config.yaml for agent_settings.vm_provider=ec2
Documents the agent_settings YAML shape consumed by `get_vm_provider`.
B0's AWS resource IDs are referenced via `default_ami_id: ami-CHANGEME`;
the user fills in the real AMI after running `packer build`.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
e06f7b2d56
feat: bootstrap_timeout / heartbeat / max-session sweepers for agent sessions
Three sweepers run on the same 30s tick:
- bootstrap_timeout — sessions stuck in `provisioning` past the timeout
- heartbeat_timeout — `ready` sessions whose daemon stopped checking in
- max_session_minutes — sessions older than the configured ceiling

Each sweeper:
- bounds its batch to 100 rows so a backlog doesn't stall the loop
- re-fetches the row (optimistic lock) before terminating so multiple
  proxy replicas don't double-terminate
- treats terminate failures as non-fatal (retry next tick)

Uses Prisma model methods (`find_many` / `find_unique` / `update`); no
raw SQL per project rules.

Validations covered: #7 (max_session_minutes), #9 (bootstrap_timeout), #10
(heartbeat_loss).

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
1bf4e6d385
feat: EC2Provider — boto3-backed VM provisioner with BYOC + spot fallback
One EC2 per session, launched in the team's AWS account using the team's
BYOC creds (decrypted at use, never logged). Spot first, on-demand fallback
when capacity unavailable. Per-session tags (litellm-session-id,
litellm-team-id, litellm-agent-id) for cleanup.

Safety:
- `set_stream_logger('botocore', WARNING)` so SigV4 payloads can't leak the
  access key into proxy logs (regression-tested in #13)
- creds enter via ProvisionContext, never leave this module
- `_safe_aws_error` formats ClientError without echoing the request payload
- `InvalidClientTokenId` / `SignatureDoesNotMatch` → fail-fast InvalidCredentialsError
- terminate is idempotent on already-gone instances

Validations covered: #5 (spot fallback), #8 (terminate idempotent), #11
(invalid creds fail-fast), #13 (creds never logged).

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
0203e72358
feat: BYOC AWS creds resolver — get_team_vm_config() + encrypt/decrypt helpers
Reads the team's BYOC AWS creds from `LiteLLM_AgentVMConfig` (decrypts
each field individually) and falls back to `LITELLM_AGENT_AWS_*` env vars
for local dev. Raises `InvalidCredentialsError` if neither path yields
creds (validation #11 fail-fast).

Falls back gracefully when the table doesn't exist yet (Epic G hasn't
shipped its migration), so this code can land before LIT-2891.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
6832c176f5
feat: get_vm_provider() factory keyed off agent_settings.vm_provider
Reads `agent_settings.vm_provider` and `agent_settings.<provider>` from
the loaded proxy config and builds the matching provider. Defaults to
`noop`. Unknown values raise `ValueError` listing the supported
providers so config typos surface fast.

Validation #1 (test_factory) covers this path. Validation #6 (provider
swap is config-only) is also exercised here.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
bdcda518a7
feat: NoopProvider for tests + agent_settings.vm_provider=noop config
In-memory `AgentVMProvider` used by the unit tests and as the default when
`agent_settings.vm_provider` is unset. The factory returns `NoopProvider`
when no AWS-backed provider is configured so the proxy boots cleanly without
AWS credentials.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
380cf11298
feat: AgentVMProvider abstraction + ProvisionContext / VMHandle / VMStatus
The pluggable VM-provider ABC for agent sessions. Per-session VMs are
provisioned via this interface; v1 implementation is EC2 (BYOC AWS).

Key types: `ProvisionContext` carries the team's AWS creds + EC2 overrides
through to the provider. `AwsCreds.__repr__` redacts secrets so we cannot
accidentally print them. `InvalidCredentialsError` (400) and
`ProvisionError` (500) are the user-facing error types.

LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
b3a6dd0f4c
feat: vm_providers package public re-exports (LIT-2878) 2026-05-06 15:31:57 -07:00
Ishaan Jaff
a6c10b75c4
feat: scaffold litellm/proxy/agent_session_endpoints package (LIT-2878) 2026-05-06 15:31:57 -07:00
Ishaan Jaff
2f6aac4d86
feat: migration for LiteLLM_AgentVMConfig BYOC AWS settings table
Picked up by `prisma migrate deploy` on proxy startup (the
`litellm-proxy-extras` package bundles its own migration directory).
LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
b65ebb586b
feat: add LiteLLM_AgentVMConfig table for BYOC AWS settings (proxy-extras schema)
Mirrors the schema in the published `litellm-proxy-extras` package so the
bundled migrations match what Prisma actually applies on proxy startup.
LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
f6ce16e281
feat: add LiteLLM_AgentVMConfig table for BYOC AWS settings (proxy schema)
Mirrors the root-level schema.prisma change so the bundled proxy schema
stays in sync. LIT-2878
2026-05-06 15:31:57 -07:00
Ishaan Jaff
530211c558
feat: add LiteLLM_AgentVMConfig table for BYOC AWS settings (root schema)
Per-team BYOC AWS config consumed by the agent-session EC2 VM provider.
`aws_creds_enc` is a JSON blob with each field individually encrypted via
`encrypt_value_helper` so a partial DB leak doesn't expose the secret.
Owned by Epic G's Settings UI (LIT-2891); consumed by Epic B's EC2 provider.

LIT-2878
2026-05-06 15:31:57 -07:00
ishaan-berri
bd1a05aed9
Fix MCP DB reload partial failures (#27314)
* Fix MCP database reload partial failures

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

* Avoid staged MCP registry exposure

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: oss-agent-shin <279349115+oss-agent-shin@users.noreply.github.com>
Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>
2026-05-06 15:18:18 -07:00
ishaan-berri
924c141843
Add new chat model metadata (#27313)
* add new model metadata

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

* address review feedback

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: oss-agent-shin <279349115+oss-agent-shin@users.noreply.github.com>
Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>
2026-05-06 15:15:21 -07:00
ishaan-berri
487479eff7
perf: cap Prometheus end-user metric cardinality with TTL + LRU eviction (#27272)
Co-authored-by: Yassin Kortam <yassinkortam@g.ucla.edu>
2026-05-06 13:35:13 -07:00
oss-agent-shin
c8e47dcb43
Fix early proxy request size enforcement (#27311)
* Add early proxy request size guard

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

* Address request size review feedback

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: oss-agent-shin <279349115+oss-agent-shin@users.noreply.github.com>
Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>
2026-05-06 12:29:11 -07:00
Dibyo Mukherjee
169c436684
Fix/member access group team (#27317)
* fix(auth): pass team_id in member-level model access check

_check_team_member_model_access calls _can_object_call_model without
team_id, so access groups defined via model_info.access_groups cannot
resolve for team-scoped DB models (their internal router name is
model_name_<team>_<uuid>, not the public name). The team-level check
already passes team_id; this mirrors that.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test(auth): add tests for member-level access group resolution with team_id

Eight tests covering _can_object_call_model and
_check_team_member_model_access with team-scoped DB models:

- access group resolves when team_id is passed
- access group fails without team_id (pre-fix behavior)
- literal model name still works with team_id (no regression)
- denied model still denied with team_id
- second model in group also reachable
- end-to-end member access via access group (mocked membership)
- end-to-end member denied for model not in allowed list
- no-override member inherits team-level check

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-06 12:05:22 -07:00
oss-agent-shin
d90cf56245
Fix SCIM user lookup filters (#27308)
* Fix SCIM Okta userName lookup

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

* fix scim user filter typing

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: oss-agent-shin <279349115+oss-agent-shin@users.noreply.github.com>
Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>
2026-05-06 11:58:47 -07:00
ishaan-berri
c92a08a307
Fix team member budget enforcement without user row (#27273)
* Fix team member budget enforcement without user row

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

* Clarify regenerated key budget repro

Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

---------

Co-authored-by: oss-agent-shin <279349115+oss-agent-shin@users.noreply.github.com>
Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>
2026-05-06 11:42:29 -07:00
Yassin Kortam
b1f577199a
fix(proxy): keep spend log cleanup running after batch failures and surface DB errors (#27303)
Co-authored-by: Yassin Kortam <yassinkortam@g.ucla.edu>
2026-05-06 18:39:15 +00:00
Mateo Wang
b83d11351f
proxy: hot-reload config YAML when --reload is set (#27274)
* proxy: hot-reload config YAML when --reload is set

Uvicorn's --reload only watches *.py by default, so editing the
--config YAML did not restart the proxy. _get_reload_options() now
extends reload_dirs/reload_includes with the config file's directory
and basename when --config is provided.

* proxy: qualify reload_includes with absolute config path

Address Greptile review on PR #27274. When the --config file lives
outside cwd, reload_includes previously stored only the basename, which
meant uvicorn/watchfiles would also reload on edits to any same-named
file inside cwd. Use the absolute config path as the include pattern in
that case so only the actual proxy config triggers a restart.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>

* fix(proxy): use basename for reload_includes config pattern

Uvicorn's resolve_reload_patterns() calls pathlib.Path.glob(), which
raises NotImplementedError on absolute patterns (uvicorn discussion
2156). Passing config_abs (an absolute path) when the config file lived
outside cwd crashed startup under --reload. The config_dir is already
added to reload_dirs, so using just the basename as the include pattern
is sufficient to match the specific config file.

* fix: make it reload app when yaml changes

* style: remove unneeded comments

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-05-06 16:06:58 +00:00
Yassin Kortam
bd1ea0252a
perf(proxy): run daily activity aggregation off the event loop (#27264)
Some checks are pending
Unit Tests: Proxy DB Operations / custom-logging (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / logging-misc (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / assert-shard-coverage (push) Waiting to run
Unit Tests: Proxy DB Operations / auth-checks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / budgets (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / db-and-spend (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / endpoints-and-responses (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / guardrails-hooks (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / jwt-and-keys (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / key-generation (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-runtime (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-server-core (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / schema-migration (push) Blocked by required conditions
Unit Tests: Proxy DB Operations / proxy-utils (push) Blocked by required conditions
Unit Tests: Security / security (push) Waiting to run
Unit Tests: Caching (Redis) / caching-redis (push) Waiting to run
Co-authored-by: Yassin Kortam <yassinkortam@g.ucla.edu>
2026-05-05 20:19:28 -07:00
ishaan-berri
c32ad90823
Fix Prometheus custom metadata label counts (#27268) (#27271)
* Fix Prometheus custom metadata label counts (#27268)

Co-authored-by: oss-agent-shin <279349115+oss-agent-shin@users.noreply.github.com>
Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>

* fix enterprise test: update positional label assertions to keyword args

prometheus_label_factory now calls .labels() with keyword arguments.
Update test_async_log_failure_event assertion to match.

---------

Co-authored-by: oss-agent-shin <ext-agent-shin@berri.ai>
Co-authored-by: oss-agent-shin <279349115+oss-agent-shin@users.noreply.github.com>
Co-authored-by: ishaan-berri <ishaan-berri@users.noreply.github.com>
2026-05-05 20:04:56 -07:00
ishaan-berri
e9fb29061a
Include model name + configured TPM/RPM in priority rate-limit 429 er… (#27216)
* Include model name + configured TPM/RPM in priority rate-limit 429 errors (#27215)

* Include model name + configured TPM/RPM in priority rate-limit 429 errors

The current 429 message ('Priority-based rate limit exceeded. Priority: prod,
Rate limit type: tokens, Remaining: -664145, Model saturation: 86.3%') doesn't
tell the operator which model was hit or what the configured limit is, so they
can't tell whether the priority allocation needs tuning or the model TPM is
just too small.

Add Model, Model TPM, and Model RPM to both the priority-based 429 and the
sibling Model-capacity 429 in dynamic_rate_limiter_v3._check_rate_limits.
Pure error-message change — no behavior or schema impact.

* test: assert priority 429 includes model name + configured TPM/RPM

Adds a regression test for the new fields in the priority-based 429 detail
('Model:', 'Model TPM:', 'Model RPM:'). Verified locally that the test
fails against the unpatched dynamic_rate_limiter_v3.py and passes after
the patch.

---------

Co-authored-by: shin-watcher <ext-agent-shin@berri.ai>

* Update litellm/proxy/hooks/dynamic_rate_limiter_v3.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Update litellm/proxy/hooks/dynamic_rate_limiter_v3.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: shin-watcher <ext-agent-shin@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-05-05 19:05:22 -07:00
Dennis Henry
73de892654
fix: replace user api key auth with authorization or cookie for mcp server creation (#27190)
* fix: replace user api key auth with authorization or cookie for mcp server creation

* updated tests
2026-05-05 18:36:22 -07:00
Michael-RZ-Berri
e75c7a312a
union x-litellm-tags with static team/key tags (#27247)
Co-authored-by: Michael Riad Zaky <michaelr@Mac.localdomain>
2026-05-05 17:46:42 -07:00
Mateo Wang
fdaa288607
ci(circleci): enable Rerun Failed Tests for all pytest jobs (#27155)
* ci(circleci): enable Rerun Failed Tests for all pytest suites

Migrated every pytest-based CircleCI job that uploads JUnit results to use
'circleci tests run' instead of invoking pytest directly. This is the
prerequisite for CircleCI's 'Rerun failed tests' feature to be available
on each job in the pipeline.

For each job:
- Glob test files via 'circleci tests glob' and pipe them into
  'circleci tests run --command="xargs ... pytest ..."' so the agent can
  feed the failed-test subset on rerun.
- Preserve all original pytest flags (parallelism, timeouts, retries,
  coverage, junit output paths).
- For jobs that previously lacked 'store_test_results' (proxy spend
  accuracy, proxy_build_from_pip, db_migration_disable_update_check),
  add the step so JUnit XML is uploaded and rerun is actually wired up.
- Replace the dynamic IGNORE_DIRS shell array in llm_translation_testing
  with a 'grep -v' filter on the glob output, matching the previous
  behavior of skipping tests/llm_translation/realtime.
- For 'build_and_test', glob 'tests/test_*.py' (top-level only) which
  matches the prior 'tests/*.py' shell glob; the long list of
  '--ignore=tests/<subdir>' flags was vestigial and is dropped.

Jobs already using 'circleci tests run' (local_testing_part1/2,
litellm_router_testing) are unchanged.

* fix(ci): convert classnames to file paths on rerun

CircleCI's Rerun Failed Tests sends each previously failed test as a
JUnit classname (e.g. 'tests.otel_tests.test_key_logging_callbacks'),
but pytest needs a file path. Without the awk preprocess step, rerun
runs fail with 'file or directory not found'.

Mirror the awk transform that local_testing_part1, local_testing_part2,
and litellm_router_testing already use, so rerun works in every job
that this PR migrated to 'circleci tests run'.

* ci: drop -x from OTEL pytest run so all failures are reported

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-05 17:27:09 -07:00
Sameer Kankute
fd7ff0f269
fix(hosted_vllm): normalize custom tools for chat completions (#25763)
* fix(hosted_vllm): normalize custom tools for chat completions

Convert custom tool definitions into OpenAI function tools before forwarding hosted_vllm chat requests to avoid provider-side validation failures. Add a regression test and include a local curl verification screenshot.

Made-with: Cursor

* Fix black issue

* Fix hosted vllm custom tool schema fallback

* fix black

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-05 17:27:02 -07:00
yuneng-jiang
9a338e1b6b
[Test] Tests: Stop parametrizing API keys into pytest test IDs (#27249)
Several tests parametrized over (model, api_key, ...) tuples or raw
token strings, causing pytest to embed those values in the test ID
and print them in CI logs. Refactored each affected test to keep the
same coverage without putting key material into parametrize.

- audio_tests/test_audio_speech.py: split env-var keys into separate
  azure/openai test functions sharing a helper; sync_mode parametrize
  preserved.
- audio_tests/test_whisper.py: split into openai_whisper /
  azure_whisper functions sharing a helper; response_format parametrize
  preserved.
- local_testing/test_embedding.py: single-case parametrize inlined.
- proxy_unit_tests/test_user_api_key_auth.py: 5 header parametrize
  cases split into 5 named tests sharing an _assert helper.
- proxy_unit_tests/test_proxy_utils.py: 4 api_key_value cases split
  into 4 named tests.
- test_litellm/proxy/auth/test_user_api_key_auth.py: 5 key-prefix
  cases (Bearer / Basic / lowercase bearer / raw / AWS SigV4) split
  into 5 named tests.

Verified: black clean; 14 refactored unit tests pass; pytest collects
audio/embedding tests with safe IDs (no key material in test IDs).
2026-05-05 17:21:18 -07:00
Sameer Kankute
e912e6d4ff
feat(audio_transcription): add NVIDIA Riva STT provider (#27185)
* feat(audio_transcription): add NVIDIA Riva STT provider

Adds nvidia_riva as a new audio transcription provider, supporting both
NVCF-hosted and self-hosted Riva ASR deployments via gRPC streaming.

- Auto-resamples input audio to 16 kHz mono LINEAR_PCM (soundfile + numpy,
  audioread fallback) so callers can send any common format.
- Maps OpenAI params: language (en -> en-US), response_format (text/json/
  verbose_json), timestamp_granularities=["word"] -> enable_word_time_offsets,
  word offsets converted ms -> s for verbose_json.
- Auth: NVCF when nvcf_function_id is set (SSL on by default), self-hosted
  otherwise (SSL off by default), with explicit use_ssl override.
- gRPC errors wrapped via NvidiaRivaException -> litellm exception classes.
- Optional deps gated behind [stt-nvidia-riva] extra (nvidia-riva-client,
  soundfile, audioread, numpy).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(nvidia_riva): address PR review feedback

- handler: forward call-level `timeout` to streaming_response_generator
  (kwarg-detected via inspect for older riva-client compat) so a stalled
  Riva server cannot block the caller indefinitely.
- audio_utils: spill bytes to a tempfile before audioread.audio_open;
  most audioread backends (FFmpeg, GStreamer) require a real filesystem
  path and previously raised TypeError on BytesIO, breaking the mp3/m4a
  fallback path.
- audio_utils: prefer soxr / scipy.signal.resample_poly for resampling
  (anti-aliased polyphase) when installed, falling back to linear only
  as a last resort. Avoids aliasing on 44.1/48 kHz -> 16 kHz downsamples.
- transformation: bare `es` now maps to es-ES (Castilian) instead of
  es-US, matching BCP-47 conventions.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: trigger CI re-run [stabilize loop 1/3]

* Update litellm/llms/nvidia_riva/audio_transcription/transformation.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* chore: trigger CI re-run [stabilize loop 1/3]

* fix code qa

* fix lint

* fix mypy

* fix mypy

* Fix NVIDIA Riva ASR service lookup

* Fix NVIDIA Riva transcription payload logging

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: oss-pr-review-agent-shin[bot] <281797381+oss-pr-review-agent-shin[bot]@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
2026-05-05 17:17:51 -07:00
Krrish Dholakia
454ce5073f
fix(anthropic, mcp): sanitize tool names to match Anthropic's [a-zA-Z0-9_-]{1,128} pattern (#26788)
* fix(anthropic, mcp): sanitize tool names to match Anthropic's `^[a-zA-Z0-9_-]{1,128}$`

Tool names with characters like `/` or `.` (commonly produced by the
OpenAPI -> MCP generator from `operationId`s such as
`actions/download-job-logs-for-workflow-run`) caused Anthropic to reject
requests with `tools.N.custom.name: String should match pattern
'^[a-zA-Z0-9_-]{1,128}$'`.

Two layers of fix:

1. Anthropic transformation: build a per-request forward map (original ->
   sanitized, disambiguated by suffix on collisions) and a reverse map
   (only for names actually rewritten). Forward map is applied to tool
   defs, `tool_choice`, and historical assistant tool_calls in messages.
   Reverse map is threaded through both the non-streaming and streaming
   response paths so callers continue to see their original tool names
   in `tool_use` blocks.

2. OpenAPI -> MCP generator: sanitize `operationId` (and the
   method+path fallback) at registration time so generated MCP tools are
   valid for any strict-name provider, not just Anthropic. The dashboard
   preview endpoint applies the same sanitization for parity.

Includes unit tests covering: collision disambiguation between
`foo_bar` and `foo/bar` in the same request, reverse-map only firing
for actually-rewritten names, message rewrite for historical tool_calls,
streaming chunk_parser reverse-mapping, and sanitization of OpenAPI
operationIds plus the preview endpoint output.

Made-with: Cursor

* fix(anthropic): build tool-name maps in transform_request, not optional_params

The previous patch stashed the per-request forward and reverse tool-name
maps under ``optional_params["_anthropic_tool_name_forward_map"]`` and
``optional_params["_anthropic_tool_name_map"]``. ``optional_params`` is
the dict that becomes the JSON body via ``data = {**optional_params}``,
so those internal keys leaked over the wire and Anthropic 400'd with:

  _anthropic_tool_name_forward_map: Extra inputs are not permitted

Worse, this meant *every* request whose tool list contained any name with
an invalid character (the exact case the patch was meant to fix) regressed
into a confusing meta-error pointing at LiteLLM's internal map instead of
the offending tool.

Fix: move all tool-name sanitization into ``transform_request``, which is
the single chokepoint already shared by ``AnthropicConfig``,
``AmazonAnthropicConfig`` (Bedrock invoke), ``VertexAIAnthropicConfig``,
and ``AzureAnthropicConfig`` (all call ``super().transform_request`` /
``AnthropicConfig.transform_request(self, ...)``). New static helper
``_sanitize_tool_names_in_request`` walks the already-Anthropic-shaped
``optional_params["tools"]`` (only ``type=="custom"`` entries -- hosted
tool names are reserved by Anthropic and must not be touched), builds
the per-request forward/reverse maps, and applies the forward map in
place to ``tools[*].name`` and ``tool_choice.name``. The reverse map is
stashed exclusively on ``litellm_params`` (which is never serialized to
a provider) under ``_anthropic_tool_name_map`` for the response paths
to consume.

Side effect of this restructure: ``map_openai_params`` is now a pure
OpenAI->Anthropic param translator with no side-channel state, which
matches its contract everywhere else in the codebase.

Tests: replaced the now-incorrect "stashes maps in optional_params"
tests with regressions that assert no underscore-prefixed keys appear
in either ``optional_params`` after ``map_openai_params`` or in the
final ``transform_request`` body. Added end-to-end coverage for:
sanitization in ``transform_request``, ``tool_choice`` rewriting,
historical ``tool_calls`` rewriting in messages, and hosted-tool
passthrough.

Made-with: Cursor

* fix(anthropic): always sanitize empty text content blocks

Anthropic 400s on `{"role": "user", "content": ""}` with:
  "messages: text content blocks must be non-empty"

LiteLLM already had `_sanitize_empty_text_content` to rewrite empty text
to a placeholder, but it was gated behind `litellm.modify_params=True`.
With that flag off (default), empty content from upstream agent
frameworks (e.g. pydantic-ai) flowed straight through and tripped the
Anthropic validator.

Fix:
- Always run `_sanitize_empty_text_content` at the top of
  `anthropic_messages_pt`, independent of `modify_params`. There is no
  way to "pass through" an empty text block, so this is non-optional.
  The richer tool-call sanitizations (Cases A/B/D, which actually
  mutate conversation structure) remain gated on `modify_params`.
- Extend `_sanitize_empty_text_content` to also handle list-of-blocks
  content (`[{"type": "text", "text": ""}]`), not just string content.

Adds 3 regression tests covering string content, list-of-blocks
content, and the no-op case (non-empty messages with modify_params off).

Made-with: Cursor

* fix(anthropic): drop dead tool-name forward-map params, fix mypy + caller-mutation

- remove unused `name_forward_map` param from `_map_tool_choice`,
  `_map_tool_helper`, `_map_tools` and the `_apply_anthropic_tool_name_forward`
  helper. Production sanitization runs in `_sanitize_tool_names_in_request`
  at `transform_request`; these params were never threaded through.
- handler.py: use `ANTHROPIC_TOOL_NAME_REVERSE_MAP_KEY` constant instead of
  the hardcoded `"_anthropic_tool_name_map"` string.
- fix mypy `"object" has no attribute "__iter__"` in
  `_rewrite_tool_names_in_messages` by guarding `tool_calls` with
  `isinstance(..., list)`.
- `_sanitize_tool_names_in_request`: build a new tools list with copy-on-
  change entries (and copy `tool_choice` on rewrite) so a caller reusing
  the same tool list/dicts across requests doesn't see its inputs
  permanently rewritten.
- doc-comment `_build_request_tool_name_maps` clarifying it operates on
  OpenAI-format tools (vs `_sanitize_tool_names_in_request` which runs
  on Anthropic-format tools post-`_map_tools`).
- tests: drop 3 tests pinning the now-removed param paths; add coverage
  for tool_calls + None function_call rewrite and caller-dict immutability.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(mcp): inherit stored credentials in test/tools/list for edit flow

When editing an existing MCP server, the Tool Configuration preview
calls POST /mcp-rest/test/tools/list with server_id but no credentials
(management API redacts them). The endpoint now calls
_inherit_credentials_from_existing_server() so stored bearer tokens
and OAuth2 M2M credentials are loaded from global_mcp_server_manager
automatically — tools load without re-entering credentials.

New servers (no server_id) and requests with explicit credentials are
unaffected (function is a no-op in both cases).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(mcp): show all tools in edit panel, not just allowed tools

Edit flow was passing externalTools (from GET /tools/list, filtered by
allowed_tools) to MCPToolConfiguration, disabling the internal hook.
Remove the external props so the internal hook fires via
POST /test/tools/list, which returns all tools unfiltered. Combined
with the credential inheritance fix, tools load automatically without
re-entering credentials and all tools are visible for re-configuration.

existingAllowedTools still pre-checks previously allowed tools.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Fix order-dependent collision in _build_anthropic_tool_name_maps

Use a two-pass approach: first pre-register all already-valid tool names
in the 'used' set, then sanitize/disambiguate names that need rewriting.
This ensures valid names always have priority regardless of input order,
preventing duplicate tool names on the wire when e.g. 'foo/bar' appears
before 'foo_bar' in the tool list.

Add regression test for the reversed ordering case.

* Fix OpenAPI tool name collision: disambiguate sanitized names with numeric suffixes

sanitize_openapi_tool_name replaces all invalid chars with '_', but when
two operationIds differ only by sanitized characters (e.g. 'foo/list' and
'foo.list' both become 'foo_list'), the second registration silently
overwrites the first in the tool registry.

Add collision disambiguation in register_tools_from_openapi that appends
_2, _3, ... suffixes when a sanitized name is already taken, mirroring
the existing logic in _build_anthropic_tool_name_maps.

* Fix preview endpoint missing collision disambiguation for tool names

Add used_names tracking and _2/_3 suffix disambiguation to
_preview_openapi_tools, matching the logic in register_tools_from_openapi.
Without this, two operationIds that sanitize to the same string (e.g.
'foo/list' and 'foo.list' both becoming 'foo_list') would show duplicate
names in the preview while registration would disambiguate them.

* Align preview HTTP method order with register_tools_from_openapi

The preview endpoint and register_tools_from_openapi both use
order-dependent collision disambiguation (_2, _3 suffixes). When the
iteration order differs, two operations on the same path with sanitized
names that collide get different suffixes in preview vs registration,
so the dashboard shows names that don't match what actually got
registered.

Also adds a regression test that fails on the swapped order.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>

* Skip duplicate originals in _build_anthropic_tool_name_maps

If the same invalid tool name appeared twice in original_names (e.g.
['foo/bar', 'foo/bar']), the second occurrence overwrote the forward
map entry with a freshly-suffixed name (foo_bar_2), leaving foo_bar
orphaned in 'used' with no reverse mapping. _sanitize_tool_names_in_request
then rewrote both tool entries to foo_bar_2, and Anthropic 400'd on
duplicate tool names.

Skip the rewrite if forward already has the original mapped.

Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
2026-05-06 00:00:36 +00:00