EC2 `ModifyImageAttribute` rejects "Character sets beyond ASCII are not
supported" when registering the AMI description. The whole AMI rolls back
on this error. Replace em-dash with hyphen.
LIT-2878
Built by `packer build` against the BYOC PoC account (us-west-2). Customers
running their own BYOC account should re-run the Packer build and replace
this value.
LIT-2878
`creds` was reassigned from `AwsCreds` to `Optional[AwsCreds]` in the
fallback branch — rename the second binding so mypy can narrow the type.
LIT-2878
Avoids the PytestUnknownMarkWarning when collecting
`tests/test_litellm/proxy/agent_session_endpoints/vm_providers/test_ec2_provider_real.py`
without `-m slow`.
LIT-2878
The cnf-update-db post-invoke hook fails (exit 100) when we install a
non-default python3 alongside, because cnf-update-db imports apt_pkg which
is bound to /usr/bin/python3 -> python3.12. Disabling the hook makes apt
update + install idempotent for AMI builds.
Also stop remapping /usr/bin/python3 via update-alternatives — it breaks
Ubuntu's python-coupled apt tooling. Tools that need 3.13 invoke it
explicitly; the systemd unit already uses /usr/bin/python3.13.
LIT-2878
Validations #3 (real boot), #11 (real BYOC fail-fast), and the real-cloud
piece of #8. Skipped by default; enable with `pytest -m slow` when the
`LITELLM_AGENT_AWS_*` + `LITELLM_TEST_*` env vars are set.
Each test wraps RunInstances in try/finally with TerminateInstances and
installs a 60-min process watchdog (per the AWS safety boundary) so a
hung test cannot leak an instance.
LIT-2878
Mocked Prisma client driving each sweeper through its happy path plus the
optimistic-lock branch (sweeper skips a session whose status changed
underneath it).
LIT-2878
Covers validation #11 (no creds = fail-fast, no instance launched), #12
(two teams resolve to two distinct creds objects), and the env-var fallback
path used in local dev before the LiteLLM_AgentVMConfig table is populated.
LIT-2878
Covers: prerequisites (Packer + AWS profile), `packer init` + `packer
build` invocation, sharing the AMI cross-account via `ami_users`, the
`LITELLM_AGENT_MODE` boot-mode contract, the Epic C migration path, and
the leak-cleanup one-liner that targets `LitellmManagedBy=agent-vm-provider`.
LIT-2878
Behaviour:
- reads runtime config from systemd's EnvironmentFile
- session mode: bootstrap → heartbeat every 30s, exit cleanly on HTTP 410
- warm mode: idle (real warm-pool hydrate lands in B2)
- redacts JWT in log output
Zero non-system deps (only `requests`, installed by the AMI builder).
Replaced wholesale by Epic C.
LIT-2878
Reads runtime config from /etc/litellm-agent/runtime.env (written by
EC2 user-data, mode 600). Restart=on-failure with a 5s backoff so transient
network blips during bootstrap don't permanently kill the session.
LIT-2878
Provisioner script driven by `litellm-agent-runtime.pkr.hcl`. Each tool is
pinned to a specific version and verified against a SHA-256 sidecar where
upstream provides one (uv) — see CLAUDE.md "CI Supply-Chain Safety".
The bun installer has no checksum sidecar, so we pin a version and pull the
artifact directly (not the install script).
LIT-2878
Builds an Ubuntu 24.04 AMI with node 24, python 3.13, git, gh, uv, bun, and
the agent-runtime systemd unit autostarted on boot. The daemon honours
`LITELLM_AGENT_MODE` from EC2 user-data: `session` for cold-boot,
`warm` for warm-pool prewarming (B2).
Uses IMDSv2 only. Tags every resource Packer creates for easy cleanup.
`ami_users` lets us share the AMI cross-account without rebuilding.
Validation #2 (`packer build`) covers this file.
LIT-2878
Documents the agent_settings YAML shape consumed by `get_vm_provider`.
B0's AWS resource IDs are referenced via `default_ami_id: ami-CHANGEME`;
the user fills in the real AMI after running `packer build`.
LIT-2878
Three sweepers run on the same 30s tick:
- bootstrap_timeout — sessions stuck in `provisioning` past the timeout
- heartbeat_timeout — `ready` sessions whose daemon stopped checking in
- max_session_minutes — sessions older than the configured ceiling
Each sweeper:
- bounds its batch to 100 rows so a backlog doesn't stall the loop
- re-fetches the row (optimistic lock) before terminating so multiple
proxy replicas don't double-terminate
- treats terminate failures as non-fatal (retry next tick)
Uses Prisma model methods (`find_many` / `find_unique` / `update`); no
raw SQL per project rules.
Validations covered: #7 (max_session_minutes), #9 (bootstrap_timeout), #10
(heartbeat_loss).
LIT-2878
One EC2 per session, launched in the team's AWS account using the team's
BYOC creds (decrypted at use, never logged). Spot first, on-demand fallback
when capacity unavailable. Per-session tags (litellm-session-id,
litellm-team-id, litellm-agent-id) for cleanup.
Safety:
- `set_stream_logger('botocore', WARNING)` so SigV4 payloads can't leak the
access key into proxy logs (regression-tested in #13)
- creds enter via ProvisionContext, never leave this module
- `_safe_aws_error` formats ClientError without echoing the request payload
- `InvalidClientTokenId` / `SignatureDoesNotMatch` → fail-fast InvalidCredentialsError
- terminate is idempotent on already-gone instances
Validations covered: #5 (spot fallback), #8 (terminate idempotent), #11
(invalid creds fail-fast), #13 (creds never logged).
LIT-2878
Reads the team's BYOC AWS creds from `LiteLLM_AgentVMConfig` (decrypts
each field individually) and falls back to `LITELLM_AGENT_AWS_*` env vars
for local dev. Raises `InvalidCredentialsError` if neither path yields
creds (validation #11 fail-fast).
Falls back gracefully when the table doesn't exist yet (Epic G hasn't
shipped its migration), so this code can land before LIT-2891.
LIT-2878
Reads `agent_settings.vm_provider` and `agent_settings.<provider>` from
the loaded proxy config and builds the matching provider. Defaults to
`noop`. Unknown values raise `ValueError` listing the supported
providers so config typos surface fast.
Validation #1 (test_factory) covers this path. Validation #6 (provider
swap is config-only) is also exercised here.
LIT-2878
In-memory `AgentVMProvider` used by the unit tests and as the default when
`agent_settings.vm_provider` is unset. The factory returns `NoopProvider`
when no AWS-backed provider is configured so the proxy boots cleanly without
AWS credentials.
LIT-2878
The pluggable VM-provider ABC for agent sessions. Per-session VMs are
provisioned via this interface; v1 implementation is EC2 (BYOC AWS).
Key types: `ProvisionContext` carries the team's AWS creds + EC2 overrides
through to the provider. `AwsCreds.__repr__` redacts secrets so we cannot
accidentally print them. `InvalidCredentialsError` (400) and
`ProvisionError` (500) are the user-facing error types.
LIT-2878
Mirrors the schema in the published `litellm-proxy-extras` package so the
bundled migrations match what Prisma actually applies on proxy startup.
LIT-2878
Per-team BYOC AWS config consumed by the agent-session EC2 VM provider.
`aws_creds_enc` is a JSON blob with each field individually encrypted via
`encrypt_value_helper` so a partial DB leak doesn't expose the secret.
Owned by Epic G's Settings UI (LIT-2891); consumed by Epic B's EC2 provider.
LIT-2878
* fix(auth): pass team_id in member-level model access check
_check_team_member_model_access calls _can_object_call_model without
team_id, so access groups defined via model_info.access_groups cannot
resolve for team-scoped DB models (their internal router name is
model_name_<team>_<uuid>, not the public name). The team-level check
already passes team_id; this mirrors that.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* test(auth): add tests for member-level access group resolution with team_id
Eight tests covering _can_object_call_model and
_check_team_member_model_access with team-scoped DB models:
- access group resolves when team_id is passed
- access group fails without team_id (pre-fix behavior)
- literal model name still works with team_id (no regression)
- denied model still denied with team_id
- second model in group also reachable
- end-to-end member access via access group (mocked membership)
- end-to-end member denied for model not in allowed list
- no-override member inherits team-level check
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* proxy: hot-reload config YAML when --reload is set
Uvicorn's --reload only watches *.py by default, so editing the
--config YAML did not restart the proxy. _get_reload_options() now
extends reload_dirs/reload_includes with the config file's directory
and basename when --config is provided.
* proxy: qualify reload_includes with absolute config path
Address Greptile review on PR #27274. When the --config file lives
outside cwd, reload_includes previously stored only the basename, which
meant uvicorn/watchfiles would also reload on edits to any same-named
file inside cwd. Use the absolute config path as the include pattern in
that case so only the actual proxy config triggers a restart.
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* fix(proxy): use basename for reload_includes config pattern
Uvicorn's resolve_reload_patterns() calls pathlib.Path.glob(), which
raises NotImplementedError on absolute patterns (uvicorn discussion
2156). Passing config_abs (an absolute path) when the config file lived
outside cwd crashed startup under --reload. The config_dir is already
added to reload_dirs, so using just the basename as the include pattern
is sufficient to match the specific config file.
* fix: make it reload app when yaml changes
* style: remove unneeded comments
---------
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Mateo Wang <mateo-berri@users.noreply.github.com>
* Include model name + configured TPM/RPM in priority rate-limit 429 errors (#27215)
* Include model name + configured TPM/RPM in priority rate-limit 429 errors
The current 429 message ('Priority-based rate limit exceeded. Priority: prod,
Rate limit type: tokens, Remaining: -664145, Model saturation: 86.3%') doesn't
tell the operator which model was hit or what the configured limit is, so they
can't tell whether the priority allocation needs tuning or the model TPM is
just too small.
Add Model, Model TPM, and Model RPM to both the priority-based 429 and the
sibling Model-capacity 429 in dynamic_rate_limiter_v3._check_rate_limits.
Pure error-message change — no behavior or schema impact.
* test: assert priority 429 includes model name + configured TPM/RPM
Adds a regression test for the new fields in the priority-based 429 detail
('Model:', 'Model TPM:', 'Model RPM:'). Verified locally that the test
fails against the unpatched dynamic_rate_limiter_v3.py and passes after
the patch.
---------
Co-authored-by: shin-watcher <ext-agent-shin@berri.ai>
* Update litellm/proxy/hooks/dynamic_rate_limiter_v3.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Update litellm/proxy/hooks/dynamic_rate_limiter_v3.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: shin-watcher <ext-agent-shin@berri.ai>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* ci(circleci): enable Rerun Failed Tests for all pytest suites
Migrated every pytest-based CircleCI job that uploads JUnit results to use
'circleci tests run' instead of invoking pytest directly. This is the
prerequisite for CircleCI's 'Rerun failed tests' feature to be available
on each job in the pipeline.
For each job:
- Glob test files via 'circleci tests glob' and pipe them into
'circleci tests run --command="xargs ... pytest ..."' so the agent can
feed the failed-test subset on rerun.
- Preserve all original pytest flags (parallelism, timeouts, retries,
coverage, junit output paths).
- For jobs that previously lacked 'store_test_results' (proxy spend
accuracy, proxy_build_from_pip, db_migration_disable_update_check),
add the step so JUnit XML is uploaded and rerun is actually wired up.
- Replace the dynamic IGNORE_DIRS shell array in llm_translation_testing
with a 'grep -v' filter on the glob output, matching the previous
behavior of skipping tests/llm_translation/realtime.
- For 'build_and_test', glob 'tests/test_*.py' (top-level only) which
matches the prior 'tests/*.py' shell glob; the long list of
'--ignore=tests/<subdir>' flags was vestigial and is dropped.
Jobs already using 'circleci tests run' (local_testing_part1/2,
litellm_router_testing) are unchanged.
* fix(ci): convert classnames to file paths on rerun
CircleCI's Rerun Failed Tests sends each previously failed test as a
JUnit classname (e.g. 'tests.otel_tests.test_key_logging_callbacks'),
but pytest needs a file path. Without the awk preprocess step, rerun
runs fail with 'file or directory not found'.
Mirror the awk transform that local_testing_part1, local_testing_part2,
and litellm_router_testing already use, so rerun works in every job
that this PR migrated to 'circleci tests run'.
* ci: drop -x from OTEL pytest run so all failures are reported
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>