mirror of
https://github.com/BerriAI/litellm.git
synced 2026-10-06 02:48:13 +00:00
* feat(proxy): offload spend tracking to a pod-local spend worker sidecar py-spy on the gateway showed the post-response _PROXY_track_cost_callback, spend-log and DBSpendUpdateWriter work running on the inference workers' event loop, so a DB or Redis stall backed up the request path. When LITELLM_SPEND_WORKER_ENABLED=true, _ProxyDBLogger serializes one compact typed SpendEvent per success and hands it to a SpendEventProducer that ships it over a unix socket (default) or loopback-only TCP to a sidecar started as `python -m gateway.spend_worker`. The sidecar runs the unchanged _ProxyDBLogger pipeline against the pod's PgBouncer (pooled_database_url). When the sidecar is unreachable, the buffer is full, or the gateway shuts down with events still queued or in flight, the producer applies LITELLM_SPEND_WORKER_ON_UNAVAILABLE (fallback in-process, or drop). The sidecar half-closes producers on SIGTERM and drains, the producer treats EOF as unavailable, and the gateway flushes buffered spend counters on shutdown. The sidecar honors LITELLM_LOG so its writes are visible in its own process log. Helm: both charts gain an opt-in spend-worker sidecar container sharing an emptyDir socket dir, and the componentized chart's HPA uses a ContainerResource CPU metric scoped to the gateway container so sidecar CPU does not drive inference scaling. * feat(terraform): opt-in spend-worker sidecar for the AWS and GCP gateway stacks Adds spend_worker_* inputs to both modules. On ECS Fargate the sidecar is a second, non-essential container in the gateway task; on Cloud Run it is a second container in the gateway service. Both listen on loopback TCP, share the gateway's DB/Redis/secret env, and set LITELLM_JOB_ROLE=spend_worker. Disabled by default. Plan-only tests cover both, and the terraform CI workflow now runs the gcp module too Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(proxy): retrieve a completed batch in the in-process spend path test The base now defers cost tracking for batches that are still in flight, so an in_progress batch never reaches update_database Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(proxy): rename the spend worker sidecar to collector Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): run the collector from the installed litellm package and finish in-flight fallbacks on shutdown The sidecar command becomes python -m litellm.proxy.collector so the classic image, whose runtime stage copies only the installed package, can run it. The module now assembles DATABASE_URL and the pod-local pgbouncer URL itself, replacing gateway/collector.py The componentized collector sidecar inherits gateway.volumeMounts so custom CA mounts reach it. SpendEventProducer shields an in-progress fallback from the writer task cancellation so close() no longer loses an event already handed to the in-process pipeline Helpers used across modules (address_argument, should_store_prompts_and_responses_in_spend_logs, flush_spend_counters_on_shutdown) become public so the change adds no reportPrivateUsage errors Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * ci(terraform): drop the gcp job duplicated by the aws/gcp matrix Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(collector): keep metrics env off the classic sidecar and reject shared loopback ports The classic chart no longer hands PROMETHEUS_METRICS_PORT and the billing metrics env to the collector container, and gives it the same /.npm scratch mount as the proxy on a read-only root. AWS and GCP now refuse a plan where the spend collector and the metrics sidecar bind the same loopback port. A regression test drives a sidecar crash mid-stream on asyncio and uvloop and checks no event is billed by both the sidecar and the in-process fallback; the producer docstring spells out why a failed drain() cannot double count Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * style(proxy): format pooled_database_url after the pgbouncer rebase Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): keep the cache-hit preset key and survive dead producers on collector drain Cache hits updated the logging object after the early return, so the offloaded spend event carried preset_cache_key=None and the collector re-hashed reconstructed kwargs. Also guard write_eof() against producer transports uvloop already closed so one dead connection cannot abort the drain Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(terraform): keep the gcp collector port off the metrics sidecar health port Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): collector connects to Postgres directly under IAM or Entra token auth Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(proxy): mark the collector's DATABASE_URL as pooled when it uses the pod's pgbouncer Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
272 lines
8.7 KiB
YAML
272 lines
8.7 KiB
YAML
suite: test collector sidecar
|
|
templates:
|
|
- deployment.yaml
|
|
- hpa.yaml
|
|
- configmap-litellm.yaml
|
|
tests:
|
|
- it: should run the proxy alone with no collector env by default
|
|
template: deployment.yaml
|
|
asserts:
|
|
- lengthEqual:
|
|
path: spec.template.spec.containers
|
|
count: 1
|
|
- notContains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_ENABLED
|
|
value: "true"
|
|
- notContains:
|
|
path: spec.template.spec.volumes
|
|
content:
|
|
name: collector-socket
|
|
any: true
|
|
|
|
- it: should add the sidecar on the same image and point both containers at the unix socket
|
|
template: deployment.yaml
|
|
set:
|
|
image.tag: test
|
|
db.connectionPool.enabled: true
|
|
collector.enabled: true
|
|
collector.resources:
|
|
requests:
|
|
cpu: 500m
|
|
memory: 1Gi
|
|
limits:
|
|
cpu: "1"
|
|
memory: 2Gi
|
|
asserts:
|
|
- lengthEqual:
|
|
path: spec.template.spec.containers
|
|
count: 2
|
|
- equal:
|
|
path: spec.template.spec.containers[1].name
|
|
value: litellm-collector
|
|
- equal:
|
|
path: spec.template.spec.containers[1].image
|
|
value: ghcr.io/berriai/litellm:test
|
|
- equal:
|
|
path: spec.template.spec.containers[1].command
|
|
value: [python, -m, litellm.proxy.collector]
|
|
- equal:
|
|
path: spec.template.spec.containers[1].resources.requests.cpu
|
|
value: 500m
|
|
- equal:
|
|
path: spec.template.spec.containers[1].resources.limits.memory
|
|
value: 2Gi
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_ENABLED
|
|
value: "true"
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_ADDRESS
|
|
value: unix:///var/run/litellm/collector.sock
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_BUFFER_SIZE
|
|
value: "1000"
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_ON_UNAVAILABLE
|
|
value: fallback
|
|
- notContains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_JOB_ROLE
|
|
value: collector
|
|
- contains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: LITELLM_JOB_ROLE
|
|
value: collector
|
|
- contains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_ADDRESS
|
|
value: unix:///var/run/litellm/collector.sock
|
|
- contains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: CONFIG_FILE_PATH
|
|
value: /etc/litellm/config.yaml
|
|
- contains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: DATABASE_HOST
|
|
value: RELEASE-NAME-postgresql
|
|
- contains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: DATABASE_PASSWORD
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: RELEASE-NAME-litellm-dbcredentials
|
|
key: password
|
|
- contains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: LITELLM_PGBOUNCER_ENABLED
|
|
value: "true"
|
|
- contains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: LITELLM_PGBOUNCER_MAX_DB_CONNECTIONS
|
|
value: "20"
|
|
- contains:
|
|
path: spec.template.spec.containers[0].volumeMounts
|
|
content:
|
|
name: collector-socket
|
|
mountPath: /var/run/litellm
|
|
- contains:
|
|
path: spec.template.spec.containers[1].volumeMounts
|
|
content:
|
|
name: collector-socket
|
|
mountPath: /var/run/litellm
|
|
- contains:
|
|
path: spec.template.spec.containers[1].volumeMounts
|
|
content:
|
|
name: litellm-config
|
|
mountPath: /etc/litellm/config.yaml
|
|
subPath: config.yaml
|
|
- contains:
|
|
path: spec.template.spec.volumes
|
|
content:
|
|
name: collector-socket
|
|
emptyDir:
|
|
sizeLimit: 1Mi
|
|
|
|
- it: should skip the socket volume and pass the policy through on tcp transport
|
|
template: deployment.yaml
|
|
set:
|
|
collector.enabled: true
|
|
collector.address: tcp://127.0.0.1:4100
|
|
collector.onUnavailable: drop
|
|
collector.bufferSize: 50
|
|
envVars:
|
|
CONFIG_FILE_PATH: /custom/config.yaml
|
|
asserts:
|
|
- lengthEqual:
|
|
path: spec.template.spec.containers
|
|
count: 2
|
|
- notContains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: CONFIG_FILE_PATH
|
|
value: /etc/litellm/config.yaml
|
|
- contains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: CONFIG_FILE_PATH
|
|
value: /custom/config.yaml
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_ADDRESS
|
|
value: tcp://127.0.0.1:4100
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_ON_UNAVAILABLE
|
|
value: drop
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_COLLECTOR_BUFFER_SIZE
|
|
value: "50"
|
|
- notContains:
|
|
path: spec.template.spec.volumes
|
|
content:
|
|
name: collector-socket
|
|
any: true
|
|
|
|
- it: should keep metrics and billing env on the proxy container only
|
|
template: deployment.yaml
|
|
set:
|
|
collector.enabled: true
|
|
metricsServer.enabled: true
|
|
metricsServer.port: 9090
|
|
billingMetrics.enabled: true
|
|
billingMetrics.endpoint: https://metering.example.com
|
|
billingMetrics.secretName: billing-mtls
|
|
asserts:
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: PROMETHEUS_METRICS_PORT
|
|
value: "9090"
|
|
- contains:
|
|
path: spec.template.spec.containers[0].env
|
|
content:
|
|
name: LITELLM_BILLING_METRICS_ENDPOINT
|
|
value: https://metering.example.com
|
|
- notContains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: PROMETHEUS_METRICS_PORT
|
|
any: true
|
|
- notContains:
|
|
path: spec.template.spec.containers[1].env
|
|
content:
|
|
name: LITELLM_BILLING_METRICS_ENDPOINT
|
|
any: true
|
|
- notContains:
|
|
path: spec.template.spec.containers[1].volumeMounts
|
|
content:
|
|
name: billing-metrics-mtls
|
|
any: true
|
|
|
|
- it: should give the sidecar the same scratch mounts as the proxy on a read-only root
|
|
template: deployment.yaml
|
|
set:
|
|
collector.enabled: true
|
|
securityContext.readOnlyRootFilesystem: true
|
|
asserts:
|
|
- contains:
|
|
path: spec.template.spec.containers[1].volumeMounts
|
|
content:
|
|
name: npm
|
|
mountPath: /.npm
|
|
- contains:
|
|
path: spec.template.spec.containers[1].volumeMounts
|
|
content:
|
|
name: cache
|
|
mountPath: /.cache
|
|
- contains:
|
|
path: spec.template.spec.containers[1].volumeMounts
|
|
content:
|
|
name: tmp
|
|
mountPath: /tmp
|
|
|
|
- it: should keep the pod-wide cpu metric unless asked to scale on the proxy container
|
|
template: hpa.yaml
|
|
set:
|
|
autoscaling.enabled: true
|
|
collector.enabled: true
|
|
asserts:
|
|
- equal: { path: "spec.metrics[0].type", value: Resource }
|
|
- equal: { path: "spec.metrics[0].resource.name", value: cpu }
|
|
|
|
- it: should scale on the proxy container's cpu only when opted in
|
|
template: hpa.yaml
|
|
set:
|
|
autoscaling.enabled: true
|
|
collector.enabled: true
|
|
collector.scaleOnProxyContainerCpu: true
|
|
asserts:
|
|
- equal: { path: "spec.metrics[0].type", value: ContainerResource }
|
|
- equal: { path: "spec.metrics[0].containerResource.name", value: cpu }
|
|
- equal: { path: "spec.metrics[0].containerResource.container", value: litellm }
|
|
- equal: { path: "spec.metrics[0].containerResource.target.averageUtilization", value: 60 }
|
|
- isNull: { path: "spec.metrics[0].resource" }
|
|
|
|
- it: should not switch to the container metric while the sidecar is off
|
|
template: hpa.yaml
|
|
set:
|
|
autoscaling.enabled: true
|
|
collector.scaleOnProxyContainerCpu: true
|
|
asserts:
|
|
- equal: { path: "spec.metrics[0].type", value: Resource }
|