Behaviour:
- reads runtime config from systemd's EnvironmentFile
- session mode: bootstrap → heartbeat every 30s, exit cleanly on HTTP 410
- warm mode: idle (real warm-pool hydrate lands in B2)
- redacts JWT in log output
Zero non-system deps (only `requests`, installed by the AMI builder).
Replaced wholesale by Epic C.
LIT-2878
Reads runtime config from /etc/litellm-agent/runtime.env (written by
EC2 user-data, mode 600). Restart=on-failure with a 5s backoff so transient
network blips during bootstrap don't permanently kill the session.
LIT-2878
Provisioner script driven by `litellm-agent-runtime.pkr.hcl`. Each tool is
pinned to a specific version and verified against a SHA-256 sidecar where
upstream provides one (uv) — see CLAUDE.md "CI Supply-Chain Safety".
The bun installer has no checksum sidecar, so we pin a version and pull the
artifact directly (not the install script).
LIT-2878
Builds an Ubuntu 24.04 AMI with node 24, python 3.13, git, gh, uv, bun, and
the agent-runtime systemd unit autostarted on boot. The daemon honours
`LITELLM_AGENT_MODE` from EC2 user-data: `session` for cold-boot,
`warm` for warm-pool prewarming (B2).
Uses IMDSv2 only. Tags every resource Packer creates for easy cleanup.
`ami_users` lets us share the AMI cross-account without rebuilding.
Validation #2 (`packer build`) covers this file.
LIT-2878
Documents the agent_settings YAML shape consumed by `get_vm_provider`.
B0's AWS resource IDs are referenced via `default_ami_id: ami-CHANGEME`;
the user fills in the real AMI after running `packer build`.
LIT-2878
Three sweepers run on the same 30s tick:
- bootstrap_timeout — sessions stuck in `provisioning` past the timeout
- heartbeat_timeout — `ready` sessions whose daemon stopped checking in
- max_session_minutes — sessions older than the configured ceiling
Each sweeper:
- bounds its batch to 100 rows so a backlog doesn't stall the loop
- re-fetches the row (optimistic lock) before terminating so multiple
proxy replicas don't double-terminate
- treats terminate failures as non-fatal (retry next tick)
Uses Prisma model methods (`find_many` / `find_unique` / `update`); no
raw SQL per project rules.
Validations covered: #7 (max_session_minutes), #9 (bootstrap_timeout), #10
(heartbeat_loss).
LIT-2878
One EC2 per session, launched in the team's AWS account using the team's
BYOC creds (decrypted at use, never logged). Spot first, on-demand fallback
when capacity unavailable. Per-session tags (litellm-session-id,
litellm-team-id, litellm-agent-id) for cleanup.
Safety:
- `set_stream_logger('botocore', WARNING)` so SigV4 payloads can't leak the
access key into proxy logs (regression-tested in #13)
- creds enter via ProvisionContext, never leave this module
- `_safe_aws_error` formats ClientError without echoing the request payload
- `InvalidClientTokenId` / `SignatureDoesNotMatch` → fail-fast InvalidCredentialsError
- terminate is idempotent on already-gone instances
Validations covered: #5 (spot fallback), #8 (terminate idempotent), #11
(invalid creds fail-fast), #13 (creds never logged).
LIT-2878
Reads the team's BYOC AWS creds from `LiteLLM_AgentVMConfig` (decrypts
each field individually) and falls back to `LITELLM_AGENT_AWS_*` env vars
for local dev. Raises `InvalidCredentialsError` if neither path yields
creds (validation #11 fail-fast).
Falls back gracefully when the table doesn't exist yet (Epic G hasn't
shipped its migration), so this code can land before LIT-2891.
LIT-2878
Reads `agent_settings.vm_provider` and `agent_settings.<provider>` from
the loaded proxy config and builds the matching provider. Defaults to
`noop`. Unknown values raise `ValueError` listing the supported
providers so config typos surface fast.
Validation #1 (test_factory) covers this path. Validation #6 (provider
swap is config-only) is also exercised here.
LIT-2878
In-memory `AgentVMProvider` used by the unit tests and as the default when
`agent_settings.vm_provider` is unset. The factory returns `NoopProvider`
when no AWS-backed provider is configured so the proxy boots cleanly without
AWS credentials.
LIT-2878
The pluggable VM-provider ABC for agent sessions. Per-session VMs are
provisioned via this interface; v1 implementation is EC2 (BYOC AWS).
Key types: `ProvisionContext` carries the team's AWS creds + EC2 overrides
through to the provider. `AwsCreds.__repr__` redacts secrets so we cannot
accidentally print them. `InvalidCredentialsError` (400) and
`ProvisionError` (500) are the user-facing error types.
LIT-2878
Mirrors the schema in the published `litellm-proxy-extras` package so the
bundled migrations match what Prisma actually applies on proxy startup.
LIT-2878
Per-team BYOC AWS config consumed by the agent-session EC2 VM provider.
`aws_creds_enc` is a JSON blob with each field individually encrypted via
`encrypt_value_helper` so a partial DB leak doesn't expose the secret.
Owned by Epic G's Settings UI (LIT-2891); consumed by Epic B's EC2 provider.
LIT-2878
LIT-2890 / B2's long-poll heartbeat calls find_worker_by_jwt on every
request, which filters by worker_jwt_hash. Without an index that's a
full table scan per heartbeat once the worker pool grows past a handful
of rows. Adding the index now so it lands with the table create rather
than as a follow-up migration during B2 ramp-up.
Covers all four resolution paths and the security-critical default:
* LITELLM_CLOUD_AGENT_PROXY_BASE_URL wins regardless of forwarded
headers (and trailing slashes get stripped).
* X-Forwarded-Host is IGNORED unless LITELLM_TRUST_PROXY_HEADERS=1.
A forged header pointing at attacker.example must not show up.
* Trust-flag opt-in honors X-Forwarded-Host + X-Forwarded-Proto,
but x-forwarded-proto alone (no host) doesn't poison the response.
* Last-resort fallback to localhost:4000 when no Host header.
Resolution order for the URL embedded in the install one-liner is now:
1. LITELLM_CLOUD_AGENT_PROXY_BASE_URL env var (operator-configured,
fully trusted) — recommended for production.
2. X-Forwarded-Host / X-Forwarded-Proto, ONLY when the operator
opts in via LITELLM_TRUST_PROXY_HEADERS=1.
3. The request's direct Host header — safe by default because it
reflects the actual TCP destination, not an attacker-supplied hop.
Previously any authenticated caller could forge X-Forwarded-Host to
embed an attacker-controlled URL in the install command. If a second
operator ran that command, the worker would send its raw pair token to
the attacker's host, who could then call POST /v2/agent-workers/register
and gain a long-lived worker JWT.
Also adds structured logging on /register failures (invalid / replayed
/ expired tokens) so operators running the proxy behind a WAF / fail2ban
can detect abuse at the network layer (the proxy itself doesn't ship a
built-in per-IP limiter).
Locks the production-safe default for LITELLM_CLOUD_AGENT_MOCK_AWS so a
future revert to the unsafe "1" default trips CI. Also covers the
strict-string parsing — typos like "true" / "yes" must NOT silently
flip on the mock.
Previously the test-connection endpoint defaulted to mock-on, so a fresh
production proxy would silently return a synthetic success for any
non-empty AWS access key. Operators saving incorrect credentials would
only discover the failure later when VMs failed to launch.
Default is now "0" — operators must set LITELLM_CLOUD_AGENT_MOCK_AWS=1
explicitly to opt into the mock path during local development.
Also adds an inline comment on _build_update_payload's `is not None`
guard so a future contributor doesn't silently drop `False` / `0`
updates by switching to truthy comparison.
Adds three regression cases that fail if a future change drops the
shlex.quote() pass on proxy_url, raw_token, or install_script_url. The
existing simple-input cases still pass unchanged because shlex.quote
returns alnum/colon/slash/dot strings verbatim.
shlex.quote() the install_script_url, proxy_url, and raw_token before
interpolating into the curl-pipe-sh one-liner. Without quoting, a
proxy_url containing spaces or shell metacharacters (e.g. via a misconfig
or the X-Forwarded-Host issue Greptile also flagged) could produce a
malformed or exploitable command on the worker box.
Greptile P2: a misbehaving or adversarial server returning
'Retry-After: 9999999' could stall the SDK indefinitely. Cap the
honored delay at MAX_RETRY_AFTER_MS (60s).
Greptile P1: the reconnects counter accumulated across the entire
stream lifetime — for a long-running stream with several transient
drops over hours, the budget would be exhausted even though every
individual reconnect succeeded. Now the counter tracks *consecutive*
failures: once a connection delivers at least one new event, the
counter resets to zero on the next drop, so only a sustained outage
trips sse_reconnect_exhausted.
Greptile P1: wait() polled indefinitely with no escape hatch — a stuck
or partitioned server would hang the caller forever. Now accepts
{ signal, timeoutMs, pollMs } and throws LiteLLMAgentError with codes
wait_aborted / wait_timeout. Forwards the signal to the underlying
requestJson call and to the inter-poll sleep.
Greptile P2: runFromInfo was exported but had no internal callers
(agent.ts and session.ts both build Run directly). Removed to shrink
the surface and avoid leaking resolveClient as a construction detail.
Address Greptile P1/P2 feedback:
- Quote {CALLBACK_URL} in user-data heredoc and SSM curl command so URLs
with shell-special characters (& or ?) do not break the curl call.
- Replace while/else fallthrough with explicit ssm_online flag and early
return when the SSM agent never comes online, so we exit cleanly
instead of crashing inside send_command with TargetNotConnected.
@ant-design/icons isn't a direct dependency of the dashboard. Use a
text glyph (▸/▾) instead — keeps the toggle visible without adding
a runtime import.
Optional-chaining dayjs(...).fromNow?.() was a TS error because the
relativeTime plugin wasn't loaded. Use relativeOrAbsolute() from the
shared helper instead.
Centralizes dayjs.extend(relativeTime) so fromNow() is typed and
loaded across the agents components. relativeOrAbsolute() falls back
to '—' for null/invalid timestamps so callers don't have to
re-implement the guard.