` development images must match both the gateway commit and release identity. Build from source for the worker host's native architecture
+
+After upgrading the gateway, update the worker image and redeploy it while keeping its proxy URL and token. Existing containers do not update automatically. If an investigation reports a worker compatibility error, update the image before retrying
+
+For deployments managed with Compose, download `compose.yaml` and provide `LITELLM_URL`, `LENS_WORKER_TOKEN`, and an explicit `LENS_WORKER_IMAGE` in a private environment file:
```bash
docker compose --env-file /path/to/lens.env -f compose.yaml up -d
```
-Developers can build locally with `LENS_WORKER_IMAGE=litellm-lens-worker:local docker compose -f deploy/lens/compose.yaml -f deploy/lens/compose.build.yaml up -d --build`
+To work on Lens itself, `make lens-dev` runs the proxy, a worker from source and the hot-reload dashboard together; set `LENS_DEV_PROXY_PORT` / `LENS_DEV_UI_PORT` to move them off 4000/3000. For a local container build, set `LENS_WORKER_IMAGE=litellm-lens-worker:local` and `LITELLM_RELEASE_TAG` to the gateway's release tag, then use `docker compose -f deploy/lens/compose.yaml -f deploy/lens/compose.build.yaml up -d --build`
The generated command gives the worker 1 GiB of temporary memory-backed storage, shared across parallel reviews. Change `size=1g` in the Docker command or set `LENS_WORKER_TMP_SIZE` with Compose to fit your server and workload. A storage failure marks the scan as failed, cleans up temporary traces, and leaves the worker available for other scans; it does not silently truncate the review. Existing workers must be recreated with the new image and mount options
@@ -92,6 +172,38 @@ curl "$LITELLM_URL/lens/$LENS_ID/runs/$BATCH_ID" -H "Authorization: Bearer $LITE
Creation queues the first batch. Posting to `/lens/{id}/runs` queues another, or returns the existing active batch. The run response contains its ID under `jobs[0].id`. Poll the batch URL for status, findings and assessments. List responses omit large result payloads; request a batch to retrieve them. Supply an optional complete `settings` object on the runs POST for a one-off override; the saved lens stays unchanged. Selection accepts `team_id`, exact `filters`, and opaque `execution_ids` returned by `/lens/preview/sample`. Preview accepts `offset` and `as_of` to keep the time window fixed while paging. Feedback uses `PATCH /lens/{id}/findings/{finding_id}` with `status` and `reason`
+## Local development
+
+`make lens-dev ARGS=--seed` starts the full dev stack. The live dashboard is at `http://localhost:3000/ui/lens/`, with login at `http://localhost:3000/ui/login/`. Next.js forwards API requests to the proxy on port 4000, so login and navigation stay in the live UI and edits hot-reload
+
+The default is Next.js dev with no production build (`LENS_DEV_BUILD_UI=0`). Set `LENS_DEV_BUILD_UI=1` when you also want a fresh static dashboard at `http://localhost:4000/ui/`. Build output goes to `.lens-dev/logs/ui-build.log`; a failed build stops startup. Both modes keep the live dashboard on port 3000. Startup checks the live login route before seeding and fails with the UI log path if Next.js exits. `LENS_DEV_STARTUP_TIMEOUT_SECONDS` controls startup readiness retries (default 300; `LENS_DEV_READINESS_REQUEST_TIMEOUT_SECONDS` caps each HTTP probe, default 5)
+
+For local fixture data, run `make lens-dev ARGS=--seed`. Use `make lens-dev ARGS="--seed large"` for 2,000 fixture copies, over one million spans and linked request logs. To seed a running stack without restarting it, use `make lens-dev ARGS="--seed-only --seed large --copies 100"`. The default profile replays one copy of every checked-in capture through authenticated `/v1/traces`, including failures, retries, streaming and multiple agent frameworks. Large seeds use the same parser and compressed ClickHouse writer in batches of four copies, and write matching request logs to PostgreSQL. The first and last batches verify linked spend totals through the proxy
+
+Seeds append fresh IDs on every invocation and spread copies over recent timestamps. Restarts without `SEED` do not add data. Lens excludes activity received in the last two minutes, so wait two minutes after seeding before checking investigation previews. `LENS_DEV_SEED_COPIES` overrides total copies, and `LENS_DEV_SEED_BATCH_COPIES` overrides copies per bulk insert (default 4, about 2,000 spans). Start with four or fewer on a constrained machine. Larger batches still respect the existing ClickHouse insert size limit; each capture is decoded separately within the OTLP safety budget. Large seeds test data volume and pagination, rather than concurrent ingestion throughput or review accuracy. They can use substantial disk space; adjust `--copies` for your machine. Seeding expects the generated local tracing configuration. The old `run_tracing_proxy_local.sh --seed` command forwards to Lens dev, using its ports and saved master key
+
+Local ingestion limits are explicit and configurable. Set OTLP and ClickHouse variables before starting the proxy and seeder so both processes use the same settings. Invalid, zero and negative values fail instead of silently falling back. Changing these limits does not require rebuilding Rust
+
+| Environment variable | Default | Controls |
+| --- | --- | --- |
+| `LENS_DEV_SEED_COPIES` | 1 default, 2000 large | Total fixture copies |
+| `LENS_DEV_SEED_BATCH_COPIES` | 4 | Copies per bulk insert |
+| `LENS_DEV_SEED_TIMEOUT_SECONDS` | 120 | Seeder HTTP timeout |
+| `OTLP_MAX_BODY_BYTES` | 16777216 | HTTP body and decompressed payload bytes |
+| `OTLP_MAX_CONCURRENT_INGESTS` | 2 | Concurrent proxy ingestion requests |
+| `OTLP_MAX_ATTRIBUTE_VALUE_BYTES` | 65536 | Stored attribute/content bytes |
+| `OTLP_MAX_DECODE_DEPTH` | 32 | Nested decode depth |
+| `OTLP_MAX_DECODE_NODES` | 65536 | JSON values or protobuf fields per export |
+| `OTLP_MAX_SPANS` | 4096 | Spans per export |
+| `OTLP_MAX_ATTRIBUTES` | 256 | Attributes per resource, scope, span, event or link |
+| `OTLP_MAX_EVENTS` | 256 | Events per span |
+| `OTLP_MAX_LINKS` | 256 | Links per span |
+| `OTLP_MAX_DECODED_SPAN_BYTES` | 16777216 | Decoded span allocation budget |
+| `CLICKHOUSE_TRACE_MAX_INSERT_BYTES` | 67108864 | Encoded trace or spend insert bytes |
+| `CLICKHOUSE_INSERT_TIMEOUT_SECONDS` | 30 | ClickHouse insert HTTP timeout |
+
+The wire parsers also enforce their library recursion limits (128 levels for JSON, 100 for protobuf). Raising the configured depth does not remove those parser limits. Bulk seeding parses each capture separately, keeping the per-export limits distinct from the bulk insert limit. Use smaller batches if an insert exceeds its byte budget. For example, `LENS_DEV_SEED_COPIES=100 LENS_DEV_SEED_BATCH_COPIES=2 make lens-dev ARGS="--seed large"`
+
## Quality evaluation
Run the checked-in cases against a configured real model. Expected labels are used only for scoring, never passed to the model. Dev and held-out cases include missing outcomes, failed tools, recovery, handoffs, unsupported claims, repeated work, long evidence and prompt injection. The background option adds clean arithmetic traces to test rare-issue discovery at scale; those repeated synthetic cases do not establish accuracy on every production workload
@@ -115,3 +227,19 @@ The Lens API now uses `/lens` instead of `/engine`, list responses use `lenses`,
Stop workers and let active scans finish before upgrading. Deploy proxy instances together: older proxies cannot use the renamed database tables. The schema migration renames the three Lens tables and the run-history identifier column in place, preserving saved investigations, findings, history, worker credentials, and billing assignments. Existing migration files retain their original names and checksums
Upgrades using `--use_prisma_db_push` stop before schema changes if any legacy Lens table exists, preventing Prisma from dropping saved data. Apply `litellm-proxy-extras/litellm_proxy_extras/migrations/20261001100000_rename_lens/migration.sql` to the configured database schema before retrying. Deployments already using migration history can instead start without `--use_prisma_db_push` to apply the shipped migration normally. Fresh databases and databases already using the renamed tables can continue using database push
+
+
+## Release compatibility
+
+Gateway and worker builds carry the same `LITELLM_RELEASE_TAG`. A worker announces its release and protocol before claiming an investigation. A mismatch returns HTTP 409 with the required image, leaving queued investigations untouched. During a rolling upgrade, workers wait for a gateway from their release
+
+The dashboard reads its image from the running gateway. `LENS_WORKER_IMAGE` overrides the registry/image for private deployments. Set an explicit `LENS_WORKER_IMAGE` for worker-only Compose. Verify that the image exists and matches the gateway before deploying it
+
+For source development, use `make lens-dev`, which gives the proxy and source worker the same commit identity. For custom containers, build both from the same checkout with `--build-arg LITELLM_RELEASE_TAG=sha-$(git rev-parse HEAD)` and set the proxy's `LENS_WORKER_IMAGE` to the worker image you built. An unlabelled custom build refuses worker setup and claims instead of guessing from the Python package version. Normal package-index installations use their installed release version
+
+The hourly development pipeline pins all component images to the same selected commit and publishes its chart only after every build and worker smoke test succeeds. The public commit-tagged worker workflow publishes to `ghcr.io/berriai/litellm-lens-worker-dev` on Lens-related changes, so an arbitrary `main` commit may require building your own pair; do not substitute the newest available worker
+
+
+## Worker dependencies
+
+The worker uses the same digest-pinned Wolfi base and Python version as the component images. Python dependencies and their hashes are locked in `deploy/lens/requirements.lock`. To update them, edit `deploy/lens/requirements.in`, then run `uv pip compile --universal --python-version 3.13 --generate-hashes --no-emit-index-url deploy/lens/requirements.in -o deploy/lens/requirements.lock`. The image installs only the locked wheels with hash verification. CI builds and scans both native architectures
diff --git a/deploy/lens/compose.build.yaml b/deploy/lens/compose.build.yaml
index e4237d8de23..52d59a84a79 100644
--- a/deploy/lens/compose.build.yaml
+++ b/deploy/lens/compose.build.yaml
@@ -3,4 +3,6 @@ services:
build:
context: ../..
dockerfile: deploy/lens/Dockerfile
+ args:
+ LITELLM_RELEASE_TAG: ${LITELLM_RELEASE_TAG:?Set the release tag used by the gateway}
image: litellm-lens-worker:local
diff --git a/deploy/lens/compose.yaml b/deploy/lens/compose.yaml
index d41cb8eb203..aa915fef663 100644
--- a/deploy/lens/compose.yaml
+++ b/deploy/lens/compose.yaml
@@ -1,6 +1,6 @@
services:
lens-worker:
- image: ${LENS_WORKER_IMAGE:-ghcr.io/berriai/litellm-lens-worker@sha256:a8e8731d954916594eea462969946b9292fb771681ff515a9fd296b53f856c77}
+ image: ${LENS_WORKER_IMAGE:-${LITELLM_VERSION:+ghcr.io/berriai/litellm-lens-worker:v}${LITELLM_VERSION:-}}
environment:
LITELLM_URL: ${LITELLM_URL:?Set the URL reachable from this container}
LENS_WORKER_TOKEN: ${LENS_WORKER_TOKEN:?Create a worker credential in the Lens UI}
diff --git a/deploy/lens/config.yaml b/deploy/lens/config.yaml
new file mode 100644
index 00000000000..cb12a2b0919
--- /dev/null
+++ b/deploy/lens/config.yaml
@@ -0,0 +1,7 @@
+general_settings:
+ master_key: os.environ/LITELLM_MASTER_KEY
+ tracing:
+ store:
+ type: clickhouse
+ url: os.environ/CLICKHOUSE_URL
+ retention_days: 14
diff --git a/deploy/lens/requirements.in b/deploy/lens/requirements.in
new file mode 100644
index 00000000000..3122d7bd6f2
--- /dev/null
+++ b/deploy/lens/requirements.in
@@ -0,0 +1,2 @@
+httpx==0.28.1
+pydantic==2.13.4
diff --git a/deploy/lens/requirements.lock b/deploy/lens/requirements.lock
new file mode 100644
index 00000000000..a895b6d645e
--- /dev/null
+++ b/deploy/lens/requirements.lock
@@ -0,0 +1,172 @@
+# This file was autogenerated by uv via the following command:
+# uv pip compile --universal --python-version 3.13 --generate-hashes --no-emit-index-url deploy/lens/requirements.in -o deploy/lens/requirements.lock
+annotated-types==0.8.0 \
+ --hash=sha256:13b2beaad985e05e2d6407ee4c4f35590b11f8d693a258a561055cac8f64cab7 \
+ --hash=sha256:f072f4d804ea359e4eaf198b1af7a8b0943881a87f31bb764f8bf219bb9419e0
+ # via pydantic
+anyio==4.15.1 \
+ --hash=sha256:6152fdbbf9a77fdec97731721bebf7c4c44f7c29b424b0065826173efc7ed101 \
+ --hash=sha256:9f28306018cbd6d329e64a36d58256edff76dd996fe423bc957326e578b82a94
+ # via httpx
+certifi==2026.7.22 \
+ --hash=sha256:62f22742b58a1a33014a2b6b706588a8d7e2a88ae7bd1a6ebe8c992928483775 \
+ --hash=sha256:741e2c3b351ddf169a738da9f2c048608ff7f2c5cc02f1ebc6b118bb090d5d55
+ # via
+ # httpcore
+ # httpx
+h11==0.16.0 \
+ --hash=sha256:4e35b956cf45792e4caa5885e69fba00bdbc6ffafbfa020300e549b208ee5ff1 \
+ --hash=sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86
+ # via httpcore
+httpcore==1.0.9 \
+ --hash=sha256:2d400746a40668fc9dec9810239072b40b4484b640a8c38fd654a024c7a1bf55 \
+ --hash=sha256:6e34463af53fd2ab5d807f399a9b45ea31c3dfa2276f15a2c3f00afff6e176e8
+ # via httpx
+httpx==0.28.1 \
+ --hash=sha256:75e98c5f16b0f35b567856f597f06ff2270a374470a5c2392242528e3e3e42fc \
+ --hash=sha256:d909fcccc110f8c7faf814ca82a9a4d816bc5a6dbfea25d6591d6985b8ba59ad
+ # via -r deploy/lens/requirements.in
+idna==3.20 \
+ --hash=sha256:a7db850025b95ded1eae8a46181a1a6c56c92c96f0e2b005d9ff8dc0210cab44 \
+ --hash=sha256:ab7ae7122974553370f0bdb919e1a960b2cd1bc1ef0276416d896db81c14582c
+ # via
+ # anyio
+ # httpx
+pydantic==2.13.4 \
+ --hash=sha256:45a282cde31d808236fd7ea9d919b128653c8b38b393d1c4ab335c62924d9aba \
+ --hash=sha256:c40756b57adaa8b1efeeced5c196f3f3b7c435f90e84ea7f443901bec8099ef6
+ # via -r deploy/lens/requirements.in
+pydantic-core==2.46.4 \
+ --hash=sha256:00c603d540afdd6b80eb39f078f33ebd46211f02f33e34a32d9f053bba711de0 \
+ --hash=sha256:0186750b482eefa11d7f435892b09c5c606193ef3375bcf94aa00ae6bfb66262 \
+ --hash=sha256:041bde0a48fd37cf71cab1c9d56d3e8625a3793fef1f7dd232b3ff37e978ecda \
+ --hash=sha256:0c563b08bca408dc7f65f700633d8442fffb2421fc47b8101377e9fd65051ff0 \
+ --hash=sha256:0cbe8b01f948de4286c74cdd6c667aceb38f5c1e26f0693b3983d9d74887c65e \
+ --hash=sha256:0ce40cd7b21210e99342afafbd4d0f76d784eb5b1d60f3bdc566be4983c6c73b \
+ --hash=sha256:0e96592440881c74a213e5ad528e2b24d3d4f940de2766bed9010ab1d9e51594 \
+ --hash=sha256:10e17cbb10a330363733efc4d7c4d0dd827ac0909b8f6a6542298fed1ea62f29 \
+ --hash=sha256:133878133d271ade3d41d1bfb2a45ec38dbdbda40bc065921c6b04e4630127e2 \
+ --hash=sha256:14d4edf427bdcf950a8a02d7cb44a08614388dd6e1bdcbf4f67504fa7887da9c \
+ --hash=sha256:14f4c5d6db102bd796a627bbb3a17b4cf4574b9ae861d8b7c9a9661c6dd3362d \
+ --hash=sha256:17299feefe090f2caa5b8e37222bb5f663e4935a8bfa6931d4102e5df1a9f398 \
+ --hash=sha256:184c081504d17f1c1066e430e117142b2c77d9448a97f7b65c6ac9fd9aee238d \
+ --hash=sha256:18e5ceec2ab67e6d5f1a9085e5a24c9c4e2ac4545730bfe668680bca05e555f3 \
+ --hash=sha256:19e51f073cd3df251856a8a4189fbdf1de4012c3ebacfb1884f94f1eb406079f \
+ --hash=sha256:1a7dd0b3ee80d90150e3495a3a13ac34dbcbfd4f012996a6a1d8900e91b5c0fb \
+ --hash=sha256:1d8ba486450b14f3b1d63bc521d410ec7565e52f887b9fb671791886436a42f7 \
+ --hash=sha256:2108ba5c1c1eca18030634489dc544844144ee36357f2f9f780b93e7ddbb44b5 \
+ --hash=sha256:228ee9bae8bef5b1e97ec58302f80357c37199e0d0a99174e138d28e6957b9d9 \
+ --hash=sha256:23ace664830ee0bfe014a0c7bc248b1f7f25ed7ad103852c317624a1083af462 \
+ --hash=sha256:2412e734dcb48da14d4e4006b82b46b74f2518b8a26ee7e58c6844a6cd6d03c4 \
+ --hash=sha256:29c61fc04a3d840155ff08e475a04809278972fe6aef51e2720554e96367e34b \
+ --hash=sha256:2f84c03c8607173d16b5a854ec68a2f9079ae03237a54fb506d13af47e1d018d \
+ --hash=sha256:3009f12e4e90b7f88b4f9adb1b0c4a3d58fe7820f3238c190047209d148026df \
+ --hash=sha256:3245406455a5d98187ec35530fd772b1d799b26667980872c8d4614991e2c4a2 \
+ --hash=sha256:3447661d99f75a3683a4cf5c87da72f2161964611864dbbeac7fbb118bb4bfc0 \
+ --hash=sha256:372429a130e469c9cd698925ce5fc50940b7a1336b0d82038e63d5bbc4edc519 \
+ --hash=sha256:395aebd9183f9d112f569aeb5b2214d1a10a33bec8456447f7fbdfa51d38d4cd \
+ --hash=sha256:3a233125ac121aa3ffba9a2b59edfc4a985a76092dc8279586ab4b71390875e7 \
+ --hash=sha256:3be77f45df024d789a672ae34f8b06fb346c4f9f46ea714956660ea4862e89ac \
+ --hash=sha256:3bf92c5d0e00fefaab325a4d27828fe6b6e2a21848686b5b60d2d9eeb09d76c6 \
+ --hash=sha256:3ecbc122d18468d06ca279dc26a8c2e2d5acb10943bb35e36ae92096dc3b5565 \
+ --hash=sha256:3fb702cd90b0446a3a1c5e470bfa0dd23c0233b676a9099ddcc964fa6ca13898 \
+ --hash=sha256:428e04521a40150c85216fc8b85e8d39fece235a9cf5e383761238c7fa9b96fb \
+ --hash=sha256:432c179df7874eeb73307aad2df0755e1ae0efa61ff0ea89b93e194411ae3928 \
+ --hash=sha256:4a05d69cba51d852c5c3e92758653245a50c0b646ced0cf05bd793ed592839d6 \
+ --hash=sha256:4c63ebc82684aa89d9a3bcbd13d515b3be44250dc68dd3bd81526c1cb31286c3 \
+ --hash=sha256:4fc73cb559bdb54b1134a706a2802a4cddd27a0633f5abb7e53056268751ac6a \
+ --hash=sha256:4fcbe087dbc2068af7eda3aa87634eba216dbda64d1ae73c8684b621d33f6596 \
+ --hash=sha256:56cb4851bcaf3d117eddcef4fe66afd750a50274b0da8e22be256d10e5611987 \
+ --hash=sha256:5855698a4856556d86e8e6cd8434bc3ac0314ee8e12089ae0e143f64c6256e4e \
+ --hash=sha256:5a4330cdbc57162e4b3aa303f588ba752257694c9c9be3e7ebb11b4aca659b5d \
+ --hash=sha256:5b712b53160b79a5850310b912a5ef8e57e56947c8ad690c227f5c9d7e561712 \
+ --hash=sha256:5d5902252db0d3cedf8d4a1bc68f70eeb430f7e4c7104c8c476753519b423008 \
+ --hash=sha256:617d7e2ca7dcb8c5cf6bcb8c59b8832c94b36196bbf1cbd1bfb56ed341905edd \
+ --hash=sha256:62f875393d7f270851f20523dd2e29f082bcc82292d66db2b64ea71f64b6e1c1 \
+ --hash=sha256:633147d34cf4550417f12e2b1a0383973bdf5cdfde212cb09e9a581cf10820be \
+ --hash=sha256:66ce7632c22d837c95301830e111ad0128a32b8207533b60896a96c4915192ea \
+ --hash=sha256:6b3ace8194b0e5204818c92802dcdca7fc6d88aabbb799d7c795540d9cd6d292 \
+ --hash=sha256:6f2eeda33a839975441c86a4119e1383c50b47faf0cbb5176985565c6bb02c33 \
+ --hash=sha256:7027560ee92211647d0d34e3f7cd6f50da56399d26a9c8ad0da286d3869a53f3 \
+ --hash=sha256:7283d57845ecf5a163403eb0702dfc220cc4fbdd18919cb5ccea4f95ee1cdab4 \
+ --hash=sha256:7a5f930472650a82629163023e630d160863fce524c616f4e5186e5de9d9a49b \
+ --hash=sha256:7bfb192b3f4b9e8a89b6277b6ce787564f62cfd272055f6e685726b111dc7826 \
+ --hash=sha256:811ff8e9c313ab425368bcbb36e5c4ebd7108c2bbf4e4089cfbb0b01eff63fac \
+ --hash=sha256:8233f2947cf85404441fd7e0085f53b10c93e0ee78611099b5c7237e36aacbf7 \
+ --hash=sha256:82cf5301172168103724d49a1444d3378cb20cdee30b116a1bd6031236298a5d \
+ --hash=sha256:8358a950c8909158e3df31538a7e4edc2d7265a7c54b47f0864d9e5bae9dcebf \
+ --hash=sha256:85bb3611ff1802f3ee7fdd7dbff26b56f343fb432d57a4728fdd49b6ef35e2f4 \
+ --hash=sha256:86e1a4418c6cd97d60c95c71164158eaf7324fae7b0923264016baa993eba6fc \
+ --hash=sha256:8b9bab013d1c7a79d3501ff86d0bc9c31bf587db4551677b96bec07df78c6b15 \
+ --hash=sha256:8c5dac79fa1614d1e06ca695109c6105923bd9c7d1d6c918d4e637b7e6b32fd3 \
+ --hash=sha256:8d0820e8192167f80d88d64038e609c31452eeca865b4e1d9950a27a4609b00b \
+ --hash=sha256:8daafc69c93ee8a0204506a3b6b30f586ef54028f52aeeeb5c4cfc5184fd5914 \
+ --hash=sha256:9037063db01f09b09e237c282b6792bd4da634b5402c4e7f0c61effed7701a04 \
+ --hash=sha256:905a0ed8ea6f2d61c1738835f99b699348d7857379083e5fc497fa0c967a407c \
+ --hash=sha256:90884113d8b48f760e9587002789ddd741e76ab9f89518cd1e43b1f1a52ec44b \
+ --hash=sha256:91a06d2e259ecfbd8c901d70c3c507900458498142b3026a296b7de4d1322cc9 \
+ --hash=sha256:926c9541b14b12b1681dca8a0b75feb510b06c6341b70a8e500c2fdcff837cce \
+ --hash=sha256:9401557acd873c3a7f3eb9383edef8ac4968f9510e340f4808d427e75667e7b4 \
+ --hash=sha256:9551187363ffc0de2a00b2e47c25aeaeb1020b69b668762966df15fc5659dd5a \
+ --hash=sha256:962ccbab7b642487b1d8b7df90ef677e03134cf1fd8880bf698649b22a69371f \
+ --hash=sha256:97e7cf2be5c77b7d1a9713a05605d49460d02c6078d38d8bef3cbe323c548424 \
+ --hash=sha256:9aa768456404a8bf48a4406685ac2bec8e72b62c69313734fa3b73cf33b3a894 \
+ --hash=sha256:9bc519fbf2b7578398853d815009ae5e4d4603d12f4e3f91da8c06852d3da3e9 \
+ --hash=sha256:9d56801be94b86a9da183e5f3766e6310752b99ff647e38b09a9500d88e46e76 \
+ --hash=sha256:9f444c499b3eefd3a92e348059471ea0c3a6e303d9c1cec09fa748fd9f895201 \
+ --hash=sha256:9fa8ae11da9e2b3126c6426f147e0fba88d96d65921799bb30c6abd1cb2c97fb \
+ --hash=sha256:a0f62d0a58f4e7da165457e995725421e0064f2255d8eccebc49f41bbc23b109 \
+ --hash=sha256:a396dcc17e5a0b164dbe026896245a4fa9ff402edca1dff0be3d53a517f74de4 \
+ --hash=sha256:aaa2a54443eff1950ba5ddc6b6ccda0d9c84a364276a62f969bdf2a390650848 \
+ --hash=sha256:ad785e92e6dc634c21555edc8bd6b64957ab844541bcb96a1366c202951ae526 \
+ --hash=sha256:af8244b2bef6aaad6d92cda81372de7f8c8d36c9f0c3ea36e827c60e7d9467a0 \
+ --hash=sha256:b078afbc25f3a1436c7a1d2cd3e322497ee99615ba97c563566fdf46aff1ee01 \
+ --hash=sha256:b2f69dec1725e79a012d920df1707de5caf7ed5e08f3be4435e25803efc47458 \
+ --hash=sha256:b8458003118a712e66286df6a707db01c52c0f52f7db8e4a38f0da1d3b94fc4e \
+ --hash=sha256:bb63e0198ca18aad131c089b9204c23079c3afa95487e561f4c522d519e55aba \
+ --hash=sha256:bfec22eab3c8cc2ceec0248aec886624116dc079afa027ecc8ad4a7e62010f8a \
+ --hash=sha256:c1747f85cee84c26985853c6f3d9bd3e75da5212912443fa111c113b9c246f39 \
+ --hash=sha256:c1b3f518abeca3aa13c712fd202306e145abf59a18b094a6bafb2d2bbf59192c \
+ --hash=sha256:c50f2528cf200c5eed56faf3f4e22fcd5f38c157a8b78576e6ba3168ec35f000 \
+ --hash=sha256:c68fcd102d71ea85c5b2dfac3f4f8476eff42a9e078fd5faefff6d145063536b \
+ --hash=sha256:c7a7bd4e39e8e4c12c39cd480356842b6a8a06e41b23a55a5e3e191718838ddf \
+ --hash=sha256:c94f0688e7b8d0a67abf40e57a7eaaecd17cc9586706a31b76c031f63df052b4 \
+ --hash=sha256:cbaf13819775b7f769bf4a1f066cb6df7a28d4480081a589828ef190226881cd \
+ --hash=sha256:cd2213145bcc2ba85884d0ac63d222fece9209678f77b9b4d76f054c561adb28 \
+ --hash=sha256:ce5c1d2a8b27468f433ca974829c44060b8097eedc39933e3c206a90ee49c4a9 \
+ --hash=sha256:d396ec2b979760aaf3218e76c24e65bd0aca24983298653b3a9d7a45f9e47b30 \
+ --hash=sha256:d51026d73fcfd93610abc7b27789c26b313920fcfb20e27462d74a7f8b06e983 \
+ --hash=sha256:d80ee3d731373b24cebbc10d689ca4ee1875caf0d5703a245db18efd4dd37fc1 \
+ --hash=sha256:d995260fdf4e1db774581b4900e0f832abe3c7c84996726bbc161b19c8f29e76 \
+ --hash=sha256:da4b951fe36dc7c3a1ccb4e3cd1747c3542b8c9ceede8fc86cae054e764485f5 \
+ --hash=sha256:daa27d92c36f24388fe3ad306b174781c747627f134452e4f128ea00ce1fe8c4 \
+ --hash=sha256:db06ffe51636ffe9ca531fe9023dd64bdd794be8754cb5df57c5498ae5b518a7 \
+ --hash=sha256:e0d65b8c354be7fb5f720c3caa8bc940bc2d20ce749c8e06135f07f8ed95dd7c \
+ --hash=sha256:e68b7a074f65a2fd746c52a7ce6142ab7006074ac269ace0c25cd8ba171f8066 \
+ --hash=sha256:e739fee756ba1010f8bcccb534252e85a35fe45ae92c295a06059ce58b74ccd3 \
+ --hash=sha256:e846ae7835bf0703ae43f534ab79a867146dadd59dc9ca5c8b53d5c8f7c9ef02 \
+ --hash=sha256:e9c26f834c65f5752f3f06cb08cb86a913ceb7274d0db6e267808a708b46bc89 \
+ --hash=sha256:ea793e075b70290d89d8142074262885d3f7da19634845135751bd6344f73b50 \
+ --hash=sha256:f027324c56cd5406ca49c124b0db10e56c69064fec039acc571c29020cc87c76 \
+ --hash=sha256:f13a646d65d09fbf1bc6b3a9635d30095c8e7e5cc419ff35ecc563c5fd04cd49 \
+ --hash=sha256:f47286a97f0bc9b8859519809077b91b2cefe4ae47fcbf5e466a009c1c5d742b \
+ --hash=sha256:f747929cf940cddb5b3668a390056ddd5ba2e5010615ea2dcf4f9c4f3ab8791d \
+ --hash=sha256:f99626688942fb746e545232e7726926f3be91b5975f8b55327665fafda991c7 \
+ --hash=sha256:f9fa868638bf362d3d138ea55829cefb3d5f4b0d7f142234382a15e2485dbec4 \
+ --hash=sha256:fbdb89b3e1c94a30cc5edfce477c6e6a5dc4d8f84665b455c27582f211a1c72c \
+ --hash=sha256:fc010ab034c8c7452522748bf937df58020d256ccae0874463d1f4d01758af8e \
+ --hash=sha256:fc3e9034a63de20e15e8ade85358bc6efc614008cab72898b4b4952bea0509ff \
+ --hash=sha256:fd8b3d9fd264be37976686c7f65cd52a83f5e84f4bfd2adf9c1d469676bbb6ae
+ # via pydantic
+typing-extensions==4.16.0 \
+ --hash=sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8 \
+ --hash=sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5
+ # via
+ # anyio
+ # pydantic
+ # pydantic-core
+ # typing-inspection
+typing-inspection==0.4.4 \
+ --hash=sha256:547274fa6b0a561ccf549cc9524b999a578e737d015d8709d021f9d0d13bea47 \
+ --hash=sha256:65b8397ba37ccbce054456aaccddfc91e6e3083c92824df348d96ca832f3f147
+ # via pydantic
diff --git a/deploy/lens/stack.yaml b/deploy/lens/stack.yaml
new file mode 100644
index 00000000000..ab559e27b19
--- /dev/null
+++ b/deploy/lens/stack.yaml
@@ -0,0 +1,91 @@
+name: litellm-lens
+
+services:
+ litellm:
+ image: ghcr.io/berriai/litellm:${LITELLM_VERSION:?Set LITELLM_VERSION to a published release, without the v prefix}
+ entrypoint:
+ - python3
+ - -c
+ - |
+ import os, sys
+ from urllib.parse import quote
+ postgres_password = quote(os.environ["POSTGRES_PASSWORD"], safe="")
+ clickhouse_password = quote(os.environ["CLICKHOUSE_PASSWORD"], safe="")
+ os.environ["DATABASE_URL"] = f"postgresql://litellm:{postgres_password}@db:5432/litellm"
+ os.environ["CLICKHOUSE_URL"] = f"http://default:{clickhouse_password}@clickhouse:8123"
+ os.execv("docker/prod_entrypoint.sh", ["docker/prod_entrypoint.sh", *sys.argv[1:]])
+ command: ["--config", "/app/lens-config.yaml", "--port", "4000"]
+ environment:
+ LITELLM_MASTER_KEY: ${LITELLM_MASTER_KEY:?Set a strong master key}
+ LITELLM_SALT_KEY: ${LITELLM_SALT_KEY:?Set a permanent encryption key and keep it across upgrades}
+ POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?Set a permanent database password}
+ STORE_MODEL_IN_DB: "True"
+ CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:?Set a permanent ClickHouse password}
+ LENS_WORKER_IMAGE: ghcr.io/berriai/litellm-lens-worker:v${LITELLM_VERSION}
+ volumes:
+ - ./config.yaml:/app/lens-config.yaml:ro
+ ports:
+ - "127.0.0.1:${LITELLM_PORT:-4000}:4000"
+ networks: [proxy, storage]
+ depends_on:
+ db:
+ condition: service_healthy
+ clickhouse:
+ condition: service_healthy
+ restart: unless-stopped
+
+ lens-worker:
+ profiles: [lens]
+ image: ghcr.io/berriai/litellm-lens-worker:v${LITELLM_VERSION}
+ environment:
+ LITELLM_URL: http://litellm:4000
+ LENS_WORKER_TOKEN: ${LENS_WORKER_TOKEN:-}
+ depends_on: [litellm]
+ networks: [proxy]
+ restart: unless-stopped
+ read_only: true
+ tmpfs:
+ - /tmp:rw,noexec,nosuid,size=${LENS_WORKER_TMP_SIZE:-1g}
+ cap_drop: [ALL]
+ security_opt: [no-new-privileges:true]
+
+ db:
+ image: postgres:16
+ environment:
+ POSTGRES_DB: litellm
+ POSTGRES_USER: litellm
+ POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
+ networks: [storage]
+ volumes:
+ - postgres_data:/var/lib/postgresql/data
+ healthcheck:
+ test: ["CMD-SHELL", "pg_isready -U litellm -d litellm"]
+ interval: 5s
+ timeout: 5s
+ retries: 20
+ restart: unless-stopped
+
+ clickhouse:
+ image: clickhouse/clickhouse-server:26.9.6.6
+ environment:
+ CLICKHOUSE_USER: default
+ CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD}
+ CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: "1"
+ volumes:
+ - clickhouse_data:/var/lib/clickhouse
+ healthcheck:
+ test: ["CMD", "clickhouse-client", "--user", "default", "--password", "${CLICKHOUSE_PASSWORD}", "--query", "SELECT 1"]
+ interval: 5s
+ timeout: 5s
+ retries: 20
+ restart: unless-stopped
+ networks: [storage]
+
+networks:
+ proxy:
+ storage:
+ internal: true
+
+volumes:
+ postgres_data:
+ clickhouse_data:
diff --git a/docker-compose.hardened.yml b/docker-compose.hardened.yml
index 31d0c2e9ef2..84a23faa054 100644
--- a/docker-compose.hardened.yml
+++ b/docker-compose.hardened.yml
@@ -6,8 +6,6 @@ services:
context: .
dockerfile: docker/Dockerfile.non_root
target: runtime
- args:
- PROXY_EXTRAS_SOURCE: "local"
depends_on:
- squid
user: "101:101"
diff --git a/docker-compose.liteadmin.yml b/docker-compose.liteadmin.yml
new file mode 100644
index 00000000000..a66846429de
--- /dev/null
+++ b/docker-compose.liteadmin.yml
@@ -0,0 +1,44 @@
+services:
+ litellm:
+ image: ${LITELLM_IMAGE:?Set the native-enabled gateway image}
+ environment:
+ LITELLM_ADMIN_AGENT_URL: http://liteadmin:10000
+ ADMIN_AGENT_SERVICE_TOKEN: ${ADMIN_AGENT_SERVICE_TOKEN:?Set a shared worker token}
+ PROXY_BASE_URL: ${LITELLM_PUBLIC_URL:?Set the existing HTTPS gateway URL}
+
+ liteadmin:
+ image: ${LITELLM_IMAGE:?Set the same native-enabled image used by the gateway}
+ command: ["--admin-agent"]
+ restart: unless-stopped
+ init: true
+ read_only: true
+ cap_drop: [ALL]
+ security_opt: [no-new-privileges:true]
+ stop_grace_period: 75s
+ environment:
+ CONNECTION_AUTH_MODE: native
+ LITELLM_BASE_URL: ${LITELLM_PUBLIC_URL:?Set the existing HTTPS gateway URL}
+ LITELLM_MODEL: ${LITELLM_ADMIN_MODEL:?Set a gateway model with tool support}
+ SLACK_BOT_TOKEN: ${SLACK_BOT_TOKEN:?Install the Slack app}
+ SLACK_APP_TOKEN: ${SLACK_APP_TOKEN:?Enable Socket Mode}
+ SLACK_WORKSPACE_ID: ${SLACK_WORKSPACE_ID:?Set the Slack workspace ID}
+ ADMIN_AGENT_SERVICE_TOKEN: ${ADMIN_AGENT_SERVICE_TOKEN:?Set a shared worker token}
+ CREDENTIAL_ENCRYPTION_KEY: ${CREDENTIAL_ENCRYPTION_KEY:?Set a persistent Fernet key}
+ STATE_DB: /var/data/events.sqlite3
+ ADMIN_READ_ONLY: ${ADMIN_READ_ONLY:-false}
+ OPENAI_AGENTS_DISABLE_TRACING: "1"
+ volumes:
+ - liteadmin_state:/var/data
+ tmpfs:
+ - /tmp:rw,noexec,nosuid,size=64m
+ healthcheck:
+ test: ["CMD", "/opt/liteadmin/bin/python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:10000/readyz', timeout=3)"]
+ interval: 30s
+ timeout: 5s
+ start_period: 30s
+ depends_on:
+ litellm:
+ condition: service_healthy
+
+volumes:
+ liteadmin_state:
diff --git a/docker/Dockerfile.database b/docker/Dockerfile.database
index 61b6faae691..3309fdd5341 100644
--- a/docker/Dockerfile.database
+++ b/docker/Dockerfile.database
@@ -113,6 +113,8 @@ RUN sed -i 's/\r$//' docker/entrypoint.sh && chmod +x docker/entrypoint.sh && \
sed -i 's/\r$//' docker/prod_entrypoint.sh && chmod +x docker/prod_entrypoint.sh
FROM $LITELLM_RUNTIME_IMAGE AS runtime
+ARG LITELLM_RELEASE_TAG=""
+ENV LITELLM_RELEASE_TAG=${LITELLM_RELEASE_TAG}
USER root
diff --git a/docker/Dockerfile.non_root b/docker/Dockerfile.non_root
index eca12855afa..bafd1af46d1 100644
--- a/docker/Dockerfile.non_root
+++ b/docker/Dockerfile.non_root
@@ -3,7 +3,6 @@
# Base images
ARG LITELLM_BUILD_IMAGE=cgr.dev/chainguard/wolfi-base@sha256:1d95114038f76513a9ace6fca107d5582b08c65981f81f61cb56bf7fd2ef216d
ARG LITELLM_RUNTIME_IMAGE=cgr.dev/chainguard/wolfi-base@sha256:1d95114038f76513a9ace6fca107d5582b08c65981f81f61cb56bf7fd2ef216d
-ARG PROXY_EXTRAS_SOURCE=published
ARG UV_IMAGE=ghcr.io/astral-sh/uv:0.11.7@sha256:240fb85ab0f263ef12f492d8476aa3a2e4e1e333f7d67fbdd923d00a506a516a
# Pinned by digest like the other base images; bump explicitly on Node upgrades.
ARG UI_BUILD_IMAGE=node:24.19-alpine3.24@sha256:d32cdf619f63fe0471182d08996dd516c6275bb5fd31ae06e55a570bd9e1ad43
@@ -44,7 +43,6 @@ COPY ui/litellm-dashboard/ ./
RUN npm run build
FROM $LITELLM_BUILD_IMAGE AS builder
-ARG PROXY_EXTRAS_SOURCE
WORKDIR /app
USER root
@@ -107,26 +105,14 @@ RUN mkdir -p /var/lib/litellm/ui /var/lib/litellm/assets && \
touch /var/lib/litellm/ui/.litellm_ui_ready
RUN --mount=type=cache,target=/app/.cache/uv,id=litellm-uv-cache \
- if [ "$PROXY_EXTRAS_SOURCE" = "published" ]; then \
- uv sync --frozen --no-default-groups --no-editable \
- --extra proxy \
- --extra proxy-runtime \
- --extra extra_proxy \
- --extra semantic-router \
- --extra saml \
- --extra bedrock-realtime \
- --python python3.13 \
- --no-sources-package litellm-proxy-extras; \
- else \
- uv sync --frozen --no-default-groups --no-editable \
- --extra proxy \
- --extra proxy-runtime \
- --extra extra_proxy \
- --extra semantic-router \
- --extra saml \
- --extra bedrock-realtime \
- --python python3.13; \
- fi
+ uv sync --frozen --no-default-groups --no-editable \
+ --extra proxy \
+ --extra proxy-runtime \
+ --extra extra_proxy \
+ --extra semantic-router \
+ --extra saml \
+ --extra bedrock-realtime \
+ --python python3.13
RUN HOME=/opt/prisma XDG_CACHE_HOME=/opt/prisma/.cache PRISMA_BINARY_CACHE_DIR=/opt/prisma/binaries \
npm_config_cache=/root/.npm \
@@ -136,7 +122,8 @@ RUN sed -i 's/\r$//' docker/entrypoint.sh && chmod +x docker/entrypoint.sh && \
sed -i 's/\r$//' docker/prod_entrypoint.sh && chmod +x docker/prod_entrypoint.sh
FROM $LITELLM_RUNTIME_IMAGE AS runtime
-ARG PROXY_EXTRAS_SOURCE
+ARG LITELLM_RELEASE_TAG=""
+ENV LITELLM_RELEASE_TAG=${LITELLM_RELEASE_TAG}
WORKDIR /app
USER root
diff --git a/docker/docker-compose.quickstart.yml b/docker/docker-compose.quickstart.yml
index 11631603a72..a1d47e323ff 100644
--- a/docker/docker-compose.quickstart.yml
+++ b/docker/docker-compose.quickstart.yml
@@ -13,11 +13,13 @@ services:
litellm:
image: docker.litellm.ai/berriai/litellm:main-stable
ports:
- - "4000:4000"
+ # LITELLM_BIND is empty by default, so this stays "4000:4000". The quickstart
+ # script sets it to "127.0.0.1:" so new installs listen on this machine only.
+ - "${LITELLM_BIND:-}${LITELLM_PORT:-4000}:4000"
environment:
LITELLM_MASTER_KEY: ${LITELLM_MASTER_KEY:?set it in .env - see the header of this file}
LITELLM_SALT_KEY: ${LITELLM_SALT_KEY:?set it in .env - see the header of this file}
- DATABASE_URL: postgresql://litellm:litellm@db:5432/litellm
+ DATABASE_URL: postgresql://litellm:${POSTGRES_PASSWORD:-litellm}@db:5432/litellm
STORE_MODEL_IN_DB: "True"
depends_on:
db:
@@ -27,7 +29,7 @@ services:
image: postgres:16
environment:
POSTGRES_USER: litellm
- POSTGRES_PASSWORD: litellm
+ POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:-litellm}
POSTGRES_DB: litellm
healthcheck:
test: ["CMD-SHELL", "pg_isready -U litellm"]
diff --git a/docker/docker-compose.tracing.yml b/docker/docker-compose.tracing.yml
index b39fc8f4561..c8d90fbc0ae 100644
--- a/docker/docker-compose.tracing.yml
+++ b/docker/docker-compose.tracing.yml
@@ -5,16 +5,19 @@ services:
build:
context: ..
target: runtime
+ args:
+ LITELLM_RELEASE_TAG: ${LITELLM_RELEASE_TAG:-}
command: ["--config", "/app/tracing-config.yaml", "--port", "4000"]
environment:
- LITELLM_MASTER_KEY: local-tracing-master-key
+ LITELLM_MASTER_KEY: sk-1234
+ LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY: "true"
LITELLM_SALT_KEY: sk-local-tracing-salt-key
DATABASE_URL: postgresql://litellm:litellm@db:5432/litellm
STORE_MODEL_IN_DB: "True"
CLICKHOUSE_URL: http://default:local-tracing@clickhouse:8123
- CLICKHOUSE_READER_URL: http://default:local-tracing@clickhouse:8123
CLICKHOUSE_DATABASE: litellm
OPENAI_API_KEY: ${OPENAI_API_KEY:-}
+ LENS_WORKER_IMAGE: ${LENS_WORKER_IMAGE:-}
volumes:
- ./tracing-config.yaml:/app/tracing-config.yaml:ro
ports:
diff --git a/docker/prod_entrypoint.sh b/docker/prod_entrypoint.sh
index 630eb6b065b..4386be65a32 100644
--- a/docker/prod_entrypoint.sh
+++ b/docker/prod_entrypoint.sh
@@ -1,5 +1,11 @@
#!/bin/sh
+if [ "$1" = "--admin-agent" ]; then
+ shift
+ export CONNECTION_AUTH_MODE=native
+ exec /opt/liteadmin/bin/litellm-admin-agent --web "$@"
+fi
+
case "$USE_DDTRACE" in
[Tt][Rr][Uu][Ee])
export DD_TRACE_OPENAI_ENABLED="False"
diff --git a/docker/tracing-config.yaml b/docker/tracing-config.yaml
index 03637cfa9fb..d8e3759641f 100644
--- a/docker/tracing-config.yaml
+++ b/docker/tracing-config.yaml
@@ -7,4 +7,7 @@ model_list:
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
tracing:
- store: clickhouse
+ store:
+ type: clickhouse
+ url: os.environ/CLICKHOUSE_URL
+ retention_days: 14
diff --git a/enterprise/enterprise_hooks/blocked_user_list.py b/enterprise/enterprise_hooks/blocked_user_list.py
index a032ea7662d..dfaf91ea081 100644
--- a/enterprise/enterprise_hooks/blocked_user_list.py
+++ b/enterprise/enterprise_hooks/blocked_user_list.py
@@ -7,15 +7,19 @@
## This accepts a list of user id's for whom calls will be rejected
-from typing import Optional, Literal
-import litellm
-from litellm.proxy.utils import PrismaClient
-from litellm.caching.caching import DualCache
-from litellm.proxy._types import UserAPIKeyAuth, LiteLLM_EndUserTable
-from litellm.integrations.custom_logger import CustomLogger
-from litellm._logging import verbose_proxy_logger
+from typing import Literal, Optional
+
from fastapi import HTTPException
+import litellm
+from litellm._internal_context import with_service_target
+from litellm._logging import verbose_proxy_logger
+from litellm.caching.caching import DualCache
+from litellm.integrations.custom_logger import CustomLogger
+from litellm.proxy._types import LiteLLM_EndUserTable, UserAPIKeyAuth
+from litellm.proxy.common_utils.user_api_key_cache import AUTH_OBJECTS_TARGET
+from litellm.proxy.utils import PrismaClient
+
class _ENTERPRISE_BlockedUserList(CustomLogger):
enforces_request_content: bool = True
@@ -54,6 +58,7 @@ class _ENTERPRISE_BlockedUserList(CustomLogger):
if litellm.set_verbose is True:
print(print_statement) # noqa
+ @with_service_target(AUTH_OBJECTS_TARGET)
async def async_pre_call_hook(
self,
user_api_key_dict: UserAPIKeyAuth,
diff --git a/enterprise/litellm_enterprise/enterprise_callbacks/send_emails/base_email.py b/enterprise/litellm_enterprise/enterprise_callbacks/send_emails/base_email.py
index 6e33d9f1bf3..a29f0a1b43a 100644
--- a/enterprise/litellm_enterprise/enterprise_callbacks/send_emails/base_email.py
+++ b/enterprise/litellm_enterprise/enterprise_callbacks/send_emails/base_email.py
@@ -6,7 +6,7 @@ Base class for sending emails to user after creating keys or invite links
import html
import json
import os
-from typing import List, Literal, Optional
+from typing import Final, List, Literal, Optional
from litellm_enterprise.types.enterprise_callbacks.send_emails import (
EmailEvent,
@@ -15,6 +15,7 @@ from litellm_enterprise.types.enterprise_callbacks.send_emails import (
SendKeyRotatedEmailEvent,
)
+from litellm._internal_context import with_service_target
from litellm._logging import verbose_proxy_logger
from litellm.caching.caching import DualCache
from litellm.constants import (
@@ -48,6 +49,8 @@ from litellm.proxy._types import (
from litellm.secret_managers.main import get_secret_bool
from litellm.types.integrations.slack_alerting import LITELLM_LOGO_URL
+_BUDGET_ALERT_CLAIMS_TARGET: Final = "budget_alert_claims"
+
def _max_budget_alert_id(user_info: CallInfo) -> str:
if user_info.event_group == Litellm_EntityType.TEAM_MEMBER:
@@ -437,6 +440,7 @@ class BaseEmailLogger(CustomLogger):
html_body=email_html_content,
)
+ @with_service_target(_BUDGET_ALERT_CLAIMS_TARGET)
async def budget_alerts(
self,
type: Literal[
@@ -606,6 +610,7 @@ class BaseEmailLogger(CustomLogger):
await self._release_budget_alert_claim(_cache, _cache_key)
return
+ @with_service_target(_BUDGET_ALERT_CLAIMS_TARGET)
async def _handle_multi_threshold_max_budget_alert(
self,
user_info: CallInfo,
@@ -691,6 +696,7 @@ class BaseEmailLogger(CustomLogger):
)
await self._release_budget_alert_claim(_cache, _cache_key)
+ @with_service_target(_BUDGET_ALERT_CLAIMS_TARGET)
async def _release_budget_alert_claim(self, cache: DualCache, cache_key: str) -> None:
try:
await cache.async_delete_cache(key=cache_key)
diff --git a/enterprise/litellm_enterprise/enterprise_callbacks/send_emails/endpoints.py b/enterprise/litellm_enterprise/enterprise_callbacks/send_emails/endpoints.py
index 1ab173a915a..cf22488edcb 100644
--- a/enterprise/litellm_enterprise/enterprise_callbacks/send_emails/endpoints.py
+++ b/enterprise/litellm_enterprise/enterprise_callbacks/send_emails/endpoints.py
@@ -17,6 +17,7 @@ from litellm_enterprise.types.enterprise_callbacks.send_emails import (
from litellm._logging import verbose_proxy_logger
from litellm.proxy._types import UserAPIKeyAuth
from litellm.proxy.auth.user_api_key_auth import user_api_key_auth
+from litellm.proxy.db.db_span import db_span
router = APIRouter()
@@ -94,16 +95,17 @@ async def _save_email_settings(prisma_client, settings: Dict[str, bool]):
json_settings = json.dumps(general_settings, default=str)
# Save updated general settings
- await prisma_client.db.litellm_config.upsert(
- where={"param_name": "general_settings"},
- data={
- "create": {
- "param_name": "general_settings",
- "param_value": json_settings,
+ async with db_span("save_email_settings", "LiteLLM_Config"):
+ await prisma_client.db.litellm_config.upsert(
+ where={"param_name": "general_settings"},
+ data={
+ "create": {
+ "param_name": "general_settings",
+ "param_value": json_settings,
+ },
+ "update": {"param_value": json_settings},
},
- "update": {"param_value": json_settings},
- },
- )
+ )
except Exception as e:
raise HTTPException(
status_code=500,
diff --git a/enterprise/litellm_enterprise/proxy/enterprise_routes.py b/enterprise/litellm_enterprise/proxy/enterprise_routes.py
index ec37c049809..a76b8d01f0e 100644
--- a/enterprise/litellm_enterprise/proxy/enterprise_routes.py
+++ b/enterprise/litellm_enterprise/proxy/enterprise_routes.py
@@ -6,6 +6,7 @@ from litellm_enterprise.enterprise_callbacks.send_emails.endpoints import (
from . import ui_crud_endpoints # side-effect: registers extra UI settings
from .audit_logging_endpoints import router as audit_logging_router
+from .liteadmin import router as liteadmin_router
from .management_endpoints import management_endpoints_router
from .utils import _should_block_robots
@@ -14,6 +15,7 @@ __all__ = ["router", "ui_crud_endpoints"]
router = APIRouter()
router.include_router(email_events_router)
router.include_router(audit_logging_router)
+router.include_router(liteadmin_router)
router.include_router(management_endpoints_router)
diff --git a/enterprise/litellm_enterprise/proxy/hooks/managed_files.py b/enterprise/litellm_enterprise/proxy/hooks/managed_files.py
index 21bf7abdc2e..e0a94612646 100644
--- a/enterprise/litellm_enterprise/proxy/hooks/managed_files.py
+++ b/enterprise/litellm_enterprise/proxy/hooks/managed_files.py
@@ -3,7 +3,7 @@
import base64
import json
-from collections.abc import Mapping, Sequence
+from collections.abc import Iterator, Mapping, Sequence
from types import MappingProxyType
from typing import (
TYPE_CHECKING,
@@ -26,6 +26,7 @@ from pydantic import ValidationError
import litellm
from litellm import Router, verbose_logger
+from litellm._internal_context import with_service_target
from litellm._uuid import uuid
from litellm.caching.caching import DualCache
from litellm.constants import MAX_FILE_LIST_LIMIT
@@ -144,6 +145,7 @@ def _parse_managed_file_object(raw_file_object: object, unified_file_id: str) ->
class _ManagedFileRow(Protocol):
unified_file_id: str
file_object: OpenAIFileObject
+ flat_model_file_ids: Sequence[str]
storage_backend: Optional[str]
storage_url: Optional[str]
created_by: Optional[str]
@@ -201,6 +203,16 @@ def _managed_file_table(prisma_client: PrismaClient) -> _ManagedFileTableActions
return prisma_client.db.litellm_managedfiletable
+def _iter_provider_file_id_pairs(
+ rows: Sequence[_ManagedFileRow],
+ requested_provider_file_ids: frozenset[str],
+) -> Iterator[tuple[str, str]]:
+ for row in rows:
+ for provider_file_id in row.flat_model_file_ids:
+ if provider_file_id in requested_provider_file_ids:
+ yield provider_file_id, row.unified_file_id
+
+
def _managed_object_table(prisma_client: PrismaClient) -> _ManagedObjectTableActions:
return prisma_client.db.litellm_managedobjecttable
@@ -218,6 +230,9 @@ def _storage_metadata_of(file_object: OpenAIFileObject | None) -> Mapping[str, s
)
+_MANAGED_FILES_TARGET: Final = "managed_files"
+
+
class _PROXY_LiteLLMManagedFiles(CustomLogger, BaseFileEndpoints):
# Class variables or attributes
def __init__(self, internal_usage_cache: InternalUsageCache, prisma_client: PrismaClient):
@@ -231,6 +246,7 @@ class _PROXY_LiteLLMManagedFiles(CustomLogger, BaseFileEndpoints):
return PrometheusLogger.get_instance()
+ @with_service_target(_MANAGED_FILES_TARGET)
async def store_unified_file_id(
self,
file_id: str,
@@ -314,6 +330,7 @@ class _PROXY_LiteLLMManagedFiles(CustomLogger, BaseFileEndpoints):
verbose_logger.warning(f"could not resolve org for managed object attribution: {e}")
return None
+ @with_service_target(_MANAGED_FILES_TARGET)
async def store_unified_object_id(
self,
unified_object_id: str,
@@ -401,6 +418,7 @@ class _PROXY_LiteLLMManagedFiles(CustomLogger, BaseFileEndpoints):
},
)
+ @with_service_target(_MANAGED_FILES_TARGET)
async def get_unified_file_id(
self, file_id: str, litellm_parent_otel_span: Optional[Span] = None
) -> Optional[LiteLLM_ManagedFileTable]:
@@ -423,6 +441,7 @@ class _PROXY_LiteLLMManagedFiles(CustomLogger, BaseFileEndpoints):
return LiteLLM_ManagedFileTable.model_validate(db_object.model_dump())
return None
+ @with_service_target(_MANAGED_FILES_TARGET)
async def delete_unified_file_id(
self, file_id: str, litellm_parent_otel_span: Optional[Span] = None
) -> OpenAIFileObject:
@@ -710,6 +729,39 @@ class _PROXY_LiteLLMManagedFiles(CustomLogger, BaseFileEndpoints):
return None
return batch_obj
+ async def get_unified_file_ids_for_provider_file_ids(
+ self,
+ provider_file_ids: Sequence[str],
+ user_api_key_dict: UserAPIKeyAuth,
+ ) -> Mapping[str, str]:
+ if not provider_file_ids:
+ return MappingProxyType({})
+
+ unique_provider_file_ids: Final = tuple(dict.fromkeys(provider_file_ids))
+ owner_filter: Final = build_owner_filter(user_api_key_dict)
+ if owner_filter is None:
+ return MappingProxyType({})
+
+ provider_file_ids_list: Final = [ # mutable-ok: Prisma hasSome requires a list
+ provider_file_id for provider_file_id in unique_provider_file_ids
+ ]
+ rows: Final = await _managed_file_table(self.prisma_client).find_many(
+ where={ # mutable-ok: Prisma requires a plain dictionary for where
+ **owner_filter,
+ "flat_model_file_ids": { # mutable-ok: Prisma requires a plain filter dictionary
+ "hasSome": provider_file_ids_list,
+ },
+ }
+ )
+ return MappingProxyType(
+ dict(
+ _iter_provider_file_id_pairs(
+ rows,
+ frozenset(unique_provider_file_ids),
+ )
+ )
+ )
+
async def get_user_created_file_ids(
self, user_api_key_dict: UserAPIKeyAuth, model_object_ids: List[str]
) -> List[OpenAIFileObject]:
diff --git a/enterprise/litellm_enterprise/proxy/liteadmin.py b/enterprise/litellm_enterprise/proxy/liteadmin.py
new file mode 100644
index 00000000000..6a9110f1460
--- /dev/null
+++ b/enterprise/litellm_enterprise/proxy/liteadmin.py
@@ -0,0 +1,283 @@
+from __future__ import annotations
+
+import hashlib
+import hmac
+import html
+import os
+import re
+import secrets
+from collections.abc import Awaitable, Callable
+from dataclasses import dataclass
+from datetime import datetime, timedelta, timezone
+from typing import Annotated, Final
+from urllib.parse import urlencode, urlsplit
+
+import httpx
+from fastapi import APIRouter, Depends, HTTPException, Request
+from fastapi.responses import HTMLResponse, RedirectResponse, Response
+from pydantic import BaseModel, ConfigDict, Field, SecretStr, TypeAdapter, ValidationError
+
+from litellm.llms.custom_httpx.http_handler import get_async_httpx_client
+from litellm.proxy._experimental.mcp_server.oauth_utils import get_request_base_url
+from litellm.proxy._types import LiteLLM_UserTable, LitellmUserRoles, UserAPIKeyAuth
+from litellm.types.proxy.auth.auth_checks import UserNotFoundError
+
+router: Final = APIRouter()
+_PREFIX: Final = "/liteadmin/slack/connect/"
+_COOKIE: Final = "__Host-litellm-slack-connect-"
+_HEADERS: Final = {
+ "Cache-Control": "no-store",
+ "Referrer-Policy": "same-origin",
+ "X-Frame-Options": "DENY",
+ "X-Content-Type-Options": "nosniff",
+ "Content-Security-Policy": "default-src 'none'; style-src 'unsafe-inline'; form-action 'self'; frame-ancestors 'none'; base-uri 'none'",
+}
+
+
+class LinkDetails(BaseModel):
+ model_config = ConfigDict(frozen=True, strict=True, extra="forbid")
+ workspace_id: str = Field(min_length=1, max_length=64)
+ slack_user_id: str = Field(min_length=1, max_length=64)
+ email: str = Field(min_length=1, max_length=320)
+
+
+class AdminSession(BaseModel):
+ model_config = ConfigDict(frozen=True)
+ user_id: str
+ credential: SecretStr
+ expires_at: float
+
+
+@dataclass(frozen=True, slots=True)
+class NativeAdminContext:
+ worker_url: str
+ service_token: SecretStr
+ client: httpx.AsyncClient
+ session_user: Callable[[Request], Awaitable[str | None]]
+ load_user: Callable[[str], Awaitable[LiteLLM_UserTable | None]]
+ mint_session: Callable[[LiteLLM_UserTable], AdminSession]
+
+ async def worker_request(self, token: str, session: AdminSession | None = None) -> httpx.Response:
+ if re.fullmatch(r"[A-Za-z0-9_-]{43}", token) is None:
+ raise HTTPException(410, "Connection link expired. Send connect in Slack for a new link")
+ try:
+ response: Final = await self.client.request(
+ "GET" if session is None else "POST",
+ f"{self.worker_url}/internal/liteadmin/links/{token}",
+ headers={"X-LiteLLM-Admin-Agent-Token": self.service_token.get_secret_value()},
+ json=None
+ if session is None
+ else {
+ "user_id": session.user_id,
+ "credential": session.credential.get_secret_value(),
+ "expires_at": session.expires_at,
+ },
+ timeout=15,
+ follow_redirects=False,
+ )
+ except httpx.HTTPError:
+ raise HTTPException(503, "LiteAdmin is temporarily unavailable") from None
+ if response.status_code == 410:
+ raise HTTPException(410, "Connection link expired. Send connect in Slack for a new link")
+ if response.status_code == 403:
+ raise HTTPException(403, "Connect your own active LiteLLM proxy-admin account with the same email as Slack")
+ if response.status_code != 200:
+ raise HTTPException(503, "LiteAdmin could not verify this connection")
+ return response
+
+ async def details(self, token: str) -> LinkDetails:
+ response: Final = await self.worker_request(token)
+ try:
+ return LinkDetails.model_validate_json(response.content)
+ except ValidationError:
+ raise HTTPException(503, "LiteAdmin could not verify this connection") from None
+
+ async def admin(self, user_id: str, details: LinkDetails) -> LiteLLM_UserTable:
+ user: Final = await self.load_user(user_id)
+ if (
+ user is None
+ or user.user_role != LitellmUserRoles.PROXY_ADMIN.value
+ or not user.user_email
+ or user.user_email.strip().casefold() != details.email.strip().casefold()
+ ):
+ raise HTTPException(403, "Connect your own active LiteLLM proxy-admin account with the same email as Slack")
+ return user
+
+
+def _page(title: str, body: str) -> HTMLResponse:
+ return HTMLResponse(
+ f''
+ f'{html.escape(title)}'
+ ""
+ f"{html.escape(title)}
{body}",
+ headers=_HEADERS,
+ )
+
+
+def _cookie_name(token: str) -> str:
+ return _COOKIE + hashlib.sha256(token.encode()).hexdigest()[:16]
+
+
+async def _session_user(request: Request) -> str | None:
+ from litellm.proxy._experimental.mcp_server.byok_oauth_endpoints import (
+ get_authenticated_browser_user_id,
+ )
+
+ return await get_authenticated_browser_user_id(request)
+
+
+async def _load_user(user_id: str) -> LiteLLM_UserTable | None:
+ from litellm.proxy.auth.auth_checks import get_user_object
+ from litellm.proxy.proxy_server import prisma_client, user_api_key_cache
+
+ if prisma_client is None:
+ raise HTTPException(503, "LiteAdmin requires a database")
+ try:
+ return await get_user_object(
+ user_id=user_id,
+ prisma_client=prisma_client,
+ user_api_key_cache=user_api_key_cache,
+ user_id_upsert=False,
+ check_db_only=True,
+ )
+ except UserNotFoundError:
+ return None
+ except Exception:
+ raise HTTPException(503, "LiteAdmin could not verify your current permissions") from None
+
+
+def mint_admin_session(user: LiteLLM_UserTable) -> AdminSession:
+ from litellm.proxy.auth.auth_checks import LITELLM_SESSION_TOKEN_PREFIX
+ from litellm.proxy.common_utils.encrypt_decrypt_utils import encrypt_bearer_token
+
+ expires: Final = datetime.now(timezone.utc) + timedelta(hours=24)
+ auth: Final = UserAPIKeyAuth(
+ token="liteadmin-" + secrets.token_urlsafe(24),
+ key_name="LiteAdmin Slack",
+ key_alias="LiteAdmin Slack",
+ user_id=user.user_id,
+ user_role=LitellmUserRoles.PROXY_ADMIN,
+ models=TypeAdapter(list[str]).validate_python(user.model_dump().get("models", [])),
+ expires=expires,
+ is_session_token=True,
+ )
+ return AdminSession(
+ user_id=user.user_id,
+ credential=SecretStr(
+ encrypt_bearer_token(auth.model_dump_json(exclude_none=True), LITELLM_SESSION_TOKEN_PREFIX)
+ ),
+ expires_at=expires.timestamp(),
+ )
+
+
+def validate_native_configuration(
+ worker_url: str, service_token: str, enterprise: bool, database_available: bool
+) -> None:
+ if not worker_url:
+ raise HTTPException(404, "LiteAdmin Slack is not enabled")
+ if not enterprise:
+ raise HTTPException(403, "LiteAdmin Slack requires LiteLLM Enterprise")
+ if not database_available:
+ raise HTTPException(503, "LiteAdmin requires a database")
+ try:
+ parsed: Final = urlsplit(worker_url)
+ port: Final = parsed.port
+ except ValueError:
+ raise HTTPException(503, "LiteAdmin worker configuration is invalid") from None
+ if (
+ parsed.scheme not in {"http", "https"}
+ or not parsed.hostname
+ or port == 0
+ or parsed.username
+ or parsed.password
+ or parsed.path
+ or parsed.query
+ or parsed.fragment
+ or len(service_token) < 32
+ or any(character.isspace() for character in service_token)
+ ):
+ raise HTTPException(503, "LiteAdmin worker configuration is invalid")
+
+
+async def native_admin_context() -> NativeAdminContext:
+ from litellm.proxy.proxy_server import premium_user, prisma_client
+
+ worker_url: Final = os.getenv("LITELLM_ADMIN_AGENT_URL", "").rstrip("/")
+ service_token: Final = os.getenv("ADMIN_AGENT_SERVICE_TOKEN", "")
+ validate_native_configuration(worker_url, service_token, premium_user is True, prisma_client is not None)
+ client: Final = get_async_httpx_client(
+ llm_provider="liteadmin_native", params={"timeout": 15.0, "follow_redirects": False}
+ ).client
+ return NativeAdminContext(
+ worker_url, SecretStr(service_token), client, _session_user, _load_user, mint_admin_session
+ )
+
+
+@router.get(_PREFIX + "{token}", include_in_schema=False, response_class=HTMLResponse)
+async def connect_page(
+ request: Request,
+ token: str,
+ context: Annotated[NativeAdminContext, Depends(native_admin_context)],
+) -> Response:
+ details: Final = await context.details(token)
+ base_url: Final = get_request_base_url(request)
+ parsed_base: Final = urlsplit(base_url)
+ if parsed_base.scheme != "https":
+ raise HTTPException(400, "LiteAdmin account connections require HTTPS")
+ user_id: Final = await context.session_user(request)
+ if user_id is None:
+ return RedirectResponse(
+ base_url + "/sso/key/generate?" + urlencode({"return_to": parsed_base.path + _PREFIX + token}),
+ status_code=303,
+ headers=_HEADERS,
+ )
+ await context.admin(user_id, details)
+ csrf: Final = secrets.token_urlsafe(32)
+ page: Final = _page(
+ "Connect LiteAdmin to Slack",
+ f"Connect {html.escape(details.email)} to LiteAdmin in your Slack workspace?
"
+ "Model requests and administrative actions will use your own LiteLLM account and current permissions
"
+ f''
+ "This connection lasts 24 hours. Send disconnect in Slack to remove the saved session
",
+ )
+ page.set_cookie(_cookie_name(token), csrf, max_age=600, secure=True, httponly=True, samesite="strict", path="/")
+ return page
+
+
+@router.post(_PREFIX + "{token}", include_in_schema=False, response_class=HTMLResponse)
+async def connect_account(
+ request: Request,
+ token: str,
+ context: Annotated[NativeAdminContext, Depends(native_admin_context)],
+) -> Response:
+ base_url: Final = get_request_base_url(request)
+ parsed_base: Final = urlsplit(base_url)
+ origin: Final = f"{parsed_base.scheme}://{parsed_base.netloc}"
+ if parsed_base.scheme != "https" or request.headers.get("Origin") != origin:
+ raise HTTPException(403, "Reopen your private Slack connection link")
+ if request.headers.get("Content-Type", "").split(";", 1)[0] != "application/x-www-form-urlencoded":
+ raise HTTPException(400, "Expected a connection form")
+ form: Final = await request.form(max_fields=1, max_files=0, max_part_size=1024)
+ supplied: Final = form.get("csrf")
+ expected: Final = request.cookies.get(_cookie_name(token), "")
+ if (
+ not isinstance(supplied, str)
+ or len(expected) != 43
+ or len(supplied) != 43
+ or not hmac.compare_digest(supplied.encode(), expected.encode())
+ ):
+ raise HTTPException(403, "Reopen your private Slack connection link")
+ user_id: Final = await context.session_user(request)
+ if user_id is None:
+ raise HTTPException(401, "Your login expired. Reopen your private Slack connection link")
+ details: Final = await context.details(token)
+ user: Final = await context.admin(user_id, details)
+ await context.worker_request(token, context.mint_session(user))
+ page: Final = _page(
+ "Account connected", "Return to Slack and ask LiteAdmin to list your teams or check a budget
"
+ )
+ page.delete_cookie(_cookie_name(token), path="/", secure=True, httponly=True, samesite="strict")
+ return page
diff --git a/enterprise/pyproject.toml b/enterprise/pyproject.toml
index 74cedb9d84d..43aa5a1f728 100644
--- a/enterprise/pyproject.toml
+++ b/enterprise/pyproject.toml
@@ -1,6 +1,6 @@
[project]
name = "litellm-enterprise"
-version = "0.1.72"
+version = "0.1.73"
description = "Package for LiteLLM Enterprise features"
readme = "README.md"
requires-python = ">=3.9"
@@ -26,7 +26,7 @@ required-version = ">=0.10.9"
module-root = ""
[tool.commitizen]
-version = "0.1.72"
+version = "0.1.73"
version_files = [
"pyproject.toml:^version",
"../pyproject.toml:litellm-enterprise==",
diff --git a/gateway/main.py b/gateway/main.py
index 61b885b27e4..fb4ae830808 100644
--- a/gateway/main.py
+++ b/gateway/main.py
@@ -9,9 +9,13 @@ Run with:
uvicorn gateway.main:app --host 0.0.0.0 --port 4000
"""
+from collections.abc import AsyncGenerator, Mapping
from contextlib import asynccontextmanager
+from typing import Final
-from fastapi.routing import Mount
+from starlette.applications import Starlette
+from starlette.routing import Mount
+from starlette.types import Lifespan
# Assemble DATABASE_URL (+ DATABASE_URL_READ_REPLICA) from the discrete
# DATABASE_* env vars before proxy_server imports spin up Prisma. Handles
@@ -54,14 +58,16 @@ def _is_gateway_route(route) -> bool:
# register routes. A module-load filter would miss routes added during
# startup; running inside the lifespan, after the inner __aenter__, catches
# them while still completing before uvicorn opens the listener.
-_proxy_lifespan = app.router.lifespan_context
+_proxy_lifespan: Final = app.router.lifespan_context
@asynccontextmanager
-async def _gateway_lifespan(app_):
- async with _proxy_lifespan(app_):
+async def _gateway_lifespan(
+ app_: Starlette, lifespan: Lifespan[Starlette] = _proxy_lifespan
+) -> AsyncGenerator[Mapping[str, object], None]:
+ async with lifespan(app_) as state:
app_.router.routes = [r for r in app_.router.routes if _is_gateway_route(r)]
- yield
+ yield state if state is not None else {}
app.router.lifespan_context = _gateway_lifespan
diff --git a/gateway/routes/allowlist.py b/gateway/routes/allowlist.py
index 6e91f5486d0..fc11c059c85 100644
--- a/gateway/routes/allowlist.py
+++ b/gateway/routes/allowlist.py
@@ -1,7 +1,7 @@
"""Path allowlist for the gateway component.
The gateway exposes the LLM data-plane surface: chat/completions, embeddings,
-audio, batches, files, fine-tuning, rerank, ocr, rag, video, search, image,
+audio, batches, files, fine-tuning, rerank, decisions, ocr, rag, video, search, image,
responses, vector stores, passthrough providers, realtime websockets, MCP
tool-call endpoints, and operational endpoints (/health, /metrics, and the
/debug/memory/summary read of the serving worker's RSS).
@@ -60,6 +60,8 @@ GATEWAY_PATH_PREFIXES: tuple[str, ...] = (
"/v1/rerank",
"/v2/rerank",
"/rerank",
+ "/v1/decisions",
+ "/decisions",
"/v1/ocr",
"/ocr",
"/v1/rag/",
diff --git a/helm/litellm-helm/templates/deployment.yaml b/helm/litellm-helm/templates/deployment.yaml
index cf7b3f8a38d..299d41e2019 100644
--- a/helm/litellm-helm/templates/deployment.yaml
+++ b/helm/litellm-helm/templates/deployment.yaml
@@ -57,6 +57,19 @@ spec:
imagePullPolicy: {{ .Values.image.pullPolicy }}
env:
{{- include "litellm.proxyEnv" . | nindent 12 }}
+ {{- if .Values.liteadmin.enabled }}
+ - name: LITELLM_ADMIN_AGENT_URL
+ value: {{ printf "http://%s-liteadmin:10000" (include "litellm.fullname" . | trunc 53 | trimSuffix "-") | quote }}
+ - name: ADMIN_AGENT_SERVICE_TOKEN
+ valueFrom:
+ secretKeyRef:
+ name: {{ required "liteadmin.existingSecret is required" .Values.liteadmin.existingSecret }}
+ key: ADMIN_AGENT_SERVICE_TOKEN
+ {{- if not (hasKey (default dict .Values.envVars) "PROXY_BASE_URL") }}
+ - name: PROXY_BASE_URL
+ value: {{ required "liteadmin.gatewayUrl is required" .Values.liteadmin.gatewayUrl | quote }}
+ {{- end }}
+ {{- end }}
{{- include "litellm.proxyMetricsEnv" . | nindent 12 }}
{{- if .Values.collector.enabled }}
{{- include "litellm.collectorEnv" . | nindent 12 }}
diff --git a/helm/litellm-helm/templates/liteadmin.yaml b/helm/litellm-helm/templates/liteadmin.yaml
new file mode 100644
index 00000000000..711edaf6913
--- /dev/null
+++ b/helm/litellm-helm/templates/liteadmin.yaml
@@ -0,0 +1,112 @@
+{{- if .Values.liteadmin.enabled }}
+{{- $name := printf "%s-liteadmin" (include "litellm.fullname" . | trunc 53 | trimSuffix "-") }}
+{{- $secret := required "liteadmin.existingSecret is required" .Values.liteadmin.existingSecret }}
+apiVersion: apps/v1
+kind: Deployment
+metadata:
+ name: {{ $name }}
+spec:
+ replicas: 1
+ strategy:
+ type: Recreate
+ selector:
+ matchLabels:
+ app.kubernetes.io/name: {{ $name }}
+ app.kubernetes.io/instance: {{ .Release.Name }}
+ template:
+ metadata:
+ labels:
+ app.kubernetes.io/name: {{ $name }}
+ app.kubernetes.io/instance: {{ .Release.Name }}
+ spec:
+ automountServiceAccountToken: false
+ terminationGracePeriodSeconds: 75
+ {{- with .Values.imagePullSecrets }}
+ imagePullSecrets:
+ {{- toYaml . | nindent 8 }}
+ {{- end }}
+ securityContext:
+ runAsUser: 10001
+ runAsGroup: 10001
+ fsGroup: 10001
+ runAsNonRoot: true
+ containers:
+ - name: liteadmin
+ image: "{{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}"
+ imagePullPolicy: {{ .Values.image.pullPolicy }}
+ args: ["--admin-agent"]
+ securityContext:
+ allowPrivilegeEscalation: false
+ readOnlyRootFilesystem: true
+ capabilities:
+ drop: [ALL]
+ envFrom:
+ - secretRef:
+ name: {{ $secret }}
+ env:
+ - name: CONNECTION_AUTH_MODE
+ value: native
+ - name: LITELLM_BASE_URL
+ value: {{ required "liteadmin.gatewayUrl is required" .Values.liteadmin.gatewayUrl | quote }}
+ - name: LITELLM_MODEL
+ value: {{ required "liteadmin.model is required" .Values.liteadmin.model | quote }}
+ - name: STATE_DB
+ value: /var/data/events.sqlite3
+ - name: ADMIN_READ_ONLY
+ value: {{ .Values.liteadmin.readOnly | quote }}
+ - name: OPENAI_AGENTS_DISABLE_TRACING
+ value: "1"
+ ports:
+ - name: health
+ containerPort: 10000
+ readinessProbe:
+ httpGet:
+ path: /readyz
+ port: health
+ periodSeconds: 15
+ livenessProbe:
+ httpGet:
+ path: /healthz
+ port: health
+ periodSeconds: 30
+ resources:
+ {{- toYaml .Values.liteadmin.resources | nindent 12 }}
+ volumeMounts:
+ - name: state
+ mountPath: /var/data
+ - name: tmp
+ mountPath: /tmp
+ volumes:
+ - name: state
+ persistentVolumeClaim:
+ claimName: {{ $name }}
+ - name: tmp
+ emptyDir:
+ sizeLimit: 64Mi
+---
+apiVersion: v1
+kind: Service
+metadata:
+ name: {{ $name }}
+spec:
+ type: ClusterIP
+ selector:
+ app.kubernetes.io/name: {{ $name }}
+ app.kubernetes.io/instance: {{ .Release.Name }}
+ ports:
+ - port: 10000
+ targetPort: health
+---
+apiVersion: v1
+kind: PersistentVolumeClaim
+metadata:
+ name: {{ $name }}
+spec:
+ accessModes: [ReadWriteOnce]
+ {{- with .Values.liteadmin.storageClassName }}
+ storageClassName: {{ . | quote }}
+ {{- end }}
+ resources:
+ requests:
+ storage: {{ .Values.liteadmin.storageSize }}
+{{- end }}
diff --git a/helm/litellm-helm/tests/migrations-job_tests.yaml b/helm/litellm-helm/tests/migrations-job_tests.yaml
index 1fe545636d4..dd4276ac60f 100644
--- a/helm/litellm-helm/tests/migrations-job_tests.yaml
+++ b/helm/litellm-helm/tests/migrations-job_tests.yaml
@@ -112,6 +112,24 @@ tests:
name: CUSTOM_VAR
value: "custom_value"
+ - it: should override a user-supplied DISABLE_SCHEMA_UPDATE so the Job always migrates
+ template: migrations-job.yaml
+ set:
+ envVars:
+ DISABLE_SCHEMA_UPDATE: "true"
+ migrationJob:
+ enabled: true
+ asserts:
+ # The Job is what owns the schema, so it renders its own
+ # DISABLE_SCHEMA_UPDATE=false after envVars and extraEnvVars. Kubernetes
+ # takes the last value for a duplicated name, so the user's "true" cannot
+ # leave the schema unmigrated. Skipping migrations is migrationJob.enabled.
+ - equal:
+ path: spec.template.spec.containers[0].env[-1]
+ value:
+ name: DISABLE_SCHEMA_UPDATE
+ value: "false"
+
- it: should not include DATABASE_URL when deployStandalone is false
template: migrations-job.yaml
set:
diff --git a/helm/litellm-helm/values.yaml b/helm/litellm-helm/values.yaml
index fcee331a5aa..83dbb3c5aa0 100644
--- a/helm/litellm-helm/values.yaml
+++ b/helm/litellm-helm/values.yaml
@@ -3,6 +3,20 @@
# Declare variables to be passed into your templates.
replicaCount: 1
+liteadmin:
+ enabled: false
+ existingSecret: ""
+ gatewayUrl: ""
+ model: ""
+ readOnly: false
+ storageSize: 1Gi
+ storageClassName: ""
+ resources:
+ requests:
+ cpu: 100m
+ memory: 256Mi
+ limits:
+ memory: 1Gi
# numWorkers: 2
image:
@@ -545,7 +559,6 @@ redis:
# Prisma migration job settings
migrationJob:
enabled: true # Enable or disable the schema migration Job
- retries: 3 # Number of retries for the Job in case of failure
backoffLimit: 4 # Backoff limit for Job restarts
# Wall-clock budget for the whole Job, shared across every `backoffLimit`
# retry rather than granted per attempt. Without it a migration that blocks
@@ -554,7 +567,6 @@ migrationJob:
# stop reconciling the whole chart until someone deletes the Job by hand.
# Set to null to opt out and restore the unbounded behaviour.
activeDeadlineSeconds: 1800
- disableSchemaUpdate: false # Skip schema migrations for specific environments. When True, the job will exit with code 0.
# Optional service account for the migration job.
# Only used when migrationJob.hooks.helm.enabled=true and serviceAccount.create=true.
# In that case, pre-install/pre-upgrade hooks run before normal resources, so this defaults to "default".
diff --git a/helm/litellm/templates/_helpers.tpl b/helm/litellm/templates/_helpers.tpl
index 20fd1a722dc..eb7433c279a 100644
--- a/helm/litellm/templates/_helpers.tpl
+++ b/helm/litellm/templates/_helpers.tpl
@@ -471,6 +471,24 @@ Directory of the collector's unix socket, shared by the gateway and
collector containers through an emptyDir. Empty when the sidecar is off
or gateway.collector.address is a tcp://127.0.0.1: address.
*/}}
+{{- define "litellm.lensWorker.image" -}}
+{{- if .Values.lensWorker.image.digest -}}
+{{- if not (regexMatch "^sha256:[0-9a-f]{64}$" .Values.lensWorker.image.digest) -}}
+{{- fail "lensWorker.image.digest must be sha256 followed by 64 lowercase hex characters" -}}
+{{- end -}}
+{{- printf "%s@%s" .Values.lensWorker.image.repository .Values.lensWorker.image.digest -}}
+{{- else -}}
+{{- $backendTag := .Values.backend.image.tag | default .Chart.AppVersion -}}
+{{- $releaseTag := ternary (printf "v%s" $backendTag) $backendTag (regexMatch "^[0-9]" $backendTag) -}}
+{{- $tag := .Values.lensWorker.image.tag | default $releaseTag -}}
+{{- $repository := .Values.lensWorker.image.repository -}}
+{{- if and (hasPrefix "sha-" $tag) (eq $repository "ghcr.io/berriai/litellm-lens-worker") -}}
+{{- $repository = "ghcr.io/berriai/litellm-lens-worker-dev" -}}
+{{- end -}}
+{{- printf "%s:%s" $repository $tag -}}
+{{- end -}}
+{{- end -}}
+
{{- define "litellm.gateway.collectorSocketDir" -}}
{{- if and .Values.gateway.collector.enabled (hasPrefix "unix://" .Values.gateway.collector.address) -}}
{{- dir (trimPrefix "unix://" .Values.gateway.collector.address) -}}
diff --git a/helm/litellm/templates/backend/deployment.yaml b/helm/litellm/templates/backend/deployment.yaml
index 3eb64e5528c..5d3be1439bd 100644
--- a/helm/litellm/templates/backend/deployment.yaml
+++ b/helm/litellm/templates/backend/deployment.yaml
@@ -57,6 +57,8 @@ spec:
containerPort: 4001
protocol: TCP
env:
+ - name: LENS_WORKER_IMAGE
+ value: {{ include "litellm.lensWorker.image" . | quote }}
{{- include "litellm.serverEnv" (dict "root" $ "component" .Values.backend) | nindent 12 }}
{{- if .Values.gateway.config.create }}
- name: CONFIG_FILE_PATH
diff --git a/helm/litellm/templates/lens/deployment.yaml b/helm/litellm/templates/lens/deployment.yaml
new file mode 100644
index 00000000000..787581b9ad1
--- /dev/null
+++ b/helm/litellm/templates/lens/deployment.yaml
@@ -0,0 +1,72 @@
+{{- if .Values.lensWorker.enabled }}
+apiVersion: apps/v1
+kind: Deployment
+metadata:
+ name: {{ include "litellm.fullname" . }}-lens-worker
+ labels:
+ {{- include "litellm.commonLabels" . | nindent 4 }}
+ app.kubernetes.io/component: lens-worker
+spec:
+ replicas: {{ .Values.lensWorker.replicaCount }}
+ selector:
+ matchLabels:
+ app.kubernetes.io/instance: {{ .Release.Name }}
+ app.kubernetes.io/component: lens-worker
+ template:
+ metadata:
+ labels:
+ {{- include "litellm.commonLabels" . | nindent 8 }}
+ app.kubernetes.io/component: lens-worker
+ spec:
+ automountServiceAccountToken: false
+ {{- with .Values.imagePullSecrets }}
+ imagePullSecrets:
+ {{- toYaml . | nindent 8 }}
+ {{- end }}
+ securityContext:
+ runAsNonRoot: true
+ runAsUser: 65532
+ runAsGroup: 65532
+ fsGroup: 65532
+ seccompProfile:
+ type: RuntimeDefault
+ containers:
+ - name: lens-worker
+ image: {{ include "litellm.lensWorker.image" . | quote }}
+ imagePullPolicy: {{ .Values.lensWorker.image.pullPolicy }}
+ securityContext:
+ allowPrivilegeEscalation: false
+ readOnlyRootFilesystem: true
+ capabilities:
+ drop: [ALL]
+ env:
+ - name: LITELLM_URL
+ value: {{ .Values.lensWorker.url | default (printf "http://%s:%v" (include "litellm.backend.fullname" .) .Values.backend.service.port) | quote }}
+ - name: LENS_WORKER_TOKEN
+ valueFrom:
+ secretKeyRef:
+ name: {{ required "lensWorker.tokenSecret.name must reference a Lens worker token" .Values.lensWorker.tokenSecret.name | quote }}
+ key: {{ .Values.lensWorker.tokenSecret.key | quote }}
+ resources:
+ {{- toYaml .Values.lensWorker.resources | nindent 12 }}
+ volumeMounts:
+ - name: tmp
+ mountPath: /tmp
+ volumes:
+ - name: tmp
+ emptyDir:
+ medium: Memory
+ sizeLimit: {{ .Values.lensWorker.tmpSizeLimit }}
+ {{- with .Values.lensWorker.nodeSelector }}
+ nodeSelector:
+ {{- toYaml . | nindent 8 }}
+ {{- end }}
+ {{- with .Values.lensWorker.tolerations }}
+ tolerations:
+ {{- toYaml . | nindent 8 }}
+ {{- end }}
+ {{- with .Values.lensWorker.affinity }}
+ affinity:
+ {{- toYaml . | nindent 8 }}
+ {{- end }}
+{{- end }}
diff --git a/helm/litellm/tests/lens_worker_tests.yaml b/helm/litellm/tests/lens_worker_tests.yaml
new file mode 100644
index 00000000000..9a83a4af6a4
--- /dev/null
+++ b/helm/litellm/tests/lens_worker_tests.yaml
@@ -0,0 +1,174 @@
+suite: Lens worker release and credentials
+templates:
+ - lens/deployment.yaml
+ - backend/deployment.yaml
+ - gateway/configmap.yaml
+values:
+ - ./values/required.yaml
+tests:
+ - it: installs the development package for a source commit
+ template: lens/deployment.yaml
+ set:
+ backend.image.tag: sha-0123456789abcdef
+ lensWorker.enabled: true
+ lensWorker.tokenSecret.name: lens-credential
+ asserts:
+ - equal:
+ path: spec.template.spec.containers[0].image
+ value: ghcr.io/berriai/litellm-lens-worker-dev:sha-0123456789abcdef
+ - it: advertises the development package for standalone source workers
+ template: backend/deployment.yaml
+ set:
+ backend.image.tag: sha-0123456789abcdef
+ asserts:
+ - contains:
+ path: spec.template.spec.containers[0].env
+ content:
+ name: LENS_WORKER_IMAGE
+ value: ghcr.io/berriai/litellm-lens-worker-dev:sha-0123456789abcdef
+ - it: preserves an explicit private source image repository
+ template: lens/deployment.yaml
+ set:
+ backend.image.tag: sha-0123456789abcdef
+ lensWorker.enabled: true
+ lensWorker.tokenSecret.name: lens-credential
+ lensWorker.image.repository: registry.example/lens-worker
+ asserts:
+ - equal:
+ path: spec.template.spec.containers[0].image
+ value: registry.example/lens-worker:sha-0123456789abcdef
+ - it: pins the worker to its approved digest even when its tag changes
+ template: lens/deployment.yaml
+ set:
+ lensWorker.enabled: true
+ lensWorker.tokenSecret.name: lens-credential
+ lensWorker.image.tag: replaced-release
+ lensWorker.image.digest: sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
+ asserts:
+ - equal:
+ path: spec.template.spec.containers[0].image
+ value: ghcr.io/berriai/litellm-lens-worker@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
+ - it: advertises the approved digest to standalone installers
+ template: backend/deployment.yaml
+ set:
+ lensWorker.image.tag: replaced-release
+ lensWorker.image.digest: sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
+ asserts:
+ - contains:
+ path: spec.template.spec.containers[0].env
+ content:
+ name: LENS_WORKER_IMAGE
+ value: ghcr.io/berriai/litellm-lens-worker@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
+ - it: refuses a malformed digest instead of falling back to the tag
+ template: backend/deployment.yaml
+ set:
+ lensWorker.image.digest: sha256:invalid
+ asserts:
+ - failedTemplate:
+ errorMessage: lensWorker.image.digest must be sha256 followed by 64 lowercase hex characters
+ - it: keeps the worker opt in
+ template: lens/deployment.yaml
+ asserts:
+ - hasDocuments:
+ count: 0
+ - it: requires a limited worker credential when enabled
+ template: lens/deployment.yaml
+ set:
+ lensWorker.enabled: true
+ asserts:
+ - failedTemplate:
+ errorMessage: lensWorker.tokenSecret.name must reference a Lens worker token
+ - it: uses the chart release and a secret without granting Kubernetes access
+ template: lens/deployment.yaml
+ chart:
+ appVersion: v1.2.3
+ set:
+ lensWorker.enabled: true
+ lensWorker.tokenSecret.name: lens-credential
+ asserts:
+ - equal:
+ path: spec.template.spec.containers[0].image
+ value: ghcr.io/berriai/litellm-lens-worker:v1.2.3
+ - equal:
+ path: spec.template.spec.containers[0].env[1].valueFrom.secretKeyRef
+ value:
+ name: lens-credential
+ key: token
+ - equal:
+ path: spec.template.spec.automountServiceAccountToken
+ value: false
+ - equal:
+ path: spec.template.spec.containers[0].securityContext.readOnlyRootFilesystem
+ value: true
+ - equal:
+ path: spec.template.spec.volumes[0].emptyDir
+ value:
+ medium: Memory
+ sizeLimit: 1Gi
+ - it: advertises the same private dev image to standalone installers
+ template: backend/deployment.yaml
+ set:
+ lensWorker.image.repository: registry.example/lens-worker
+ lensWorker.image.tag: branch-main-1234567
+ asserts:
+ - contains:
+ path: spec.template.spec.containers[0].env
+ content:
+ name: LENS_WORKER_IMAGE
+ value: registry.example/lens-worker:branch-main-1234567
+ - it: supports an external gateway and a registry override
+ template: lens/deployment.yaml
+ set:
+ lensWorker.enabled: true
+ lensWorker.tokenSecret.name: lens-credential
+ lensWorker.url: https://gateway.example/proxy
+ lensWorker.image.repository: registry.example/lens-worker
+ lensWorker.image.tag: branch-main-1234567
+ asserts:
+ - equal:
+ path: spec.template.spec.containers[0].image
+ value: registry.example/lens-worker:branch-main-1234567
+ - equal:
+ path: spec.template.spec.containers[0].env[0].value
+ value: https://gateway.example/proxy
+ - it: prefixes a numeric chart release with v
+ template: lens/deployment.yaml
+ chart:
+ appVersion: 1.2.3-rc.4
+ set:
+ lensWorker.enabled: true
+ lensWorker.tokenSecret.name: lens-credential
+ asserts:
+ - equal:
+ path: spec.template.spec.containers[0].image
+ value: ghcr.io/berriai/litellm-lens-worker:v1.2.3-rc.4
+ - it: follows a backend image override when no worker tag is set
+ template: lens/deployment.yaml
+ set:
+ backend.image.tag: branch-main-1234567
+ lensWorker.enabled: true
+ lensWorker.tokenSecret.name: lens-credential
+ asserts:
+ - equal:
+ path: spec.template.spec.containers[0].image
+ value: ghcr.io/berriai/litellm-lens-worker:branch-main-1234567
+ - it: recommends the overridden backend release for standalone installers
+ template: backend/deployment.yaml
+ set:
+ backend.image.tag: v1.2.3-dev.4
+ asserts:
+ - contains:
+ path: spec.template.spec.containers[0].env
+ content:
+ name: LENS_WORKER_IMAGE
+ value: ghcr.io/berriai/litellm-lens-worker:v1.2.3-dev.4
+ - it: normalizes a numeric backend tag to the published worker tag
+ template: lens/deployment.yaml
+ set:
+ backend.image.tag: 1.2.3-dev.4
+ lensWorker.enabled: true
+ lensWorker.tokenSecret.name: lens-credential
+ asserts:
+ - equal:
+ path: spec.template.spec.containers[0].image
+ value: ghcr.io/berriai/litellm-lens-worker:v1.2.3-dev.4
diff --git a/helm/litellm/values.yaml b/helm/litellm/values.yaml
index 2c0c7151a32..cf3334f6156 100644
--- a/helm/litellm/values.yaml
+++ b/helm/litellm/values.yaml
@@ -629,3 +629,26 @@ ui:
affinity: {}
# Same shape as gateway.topologySpreadConstraints.
topologySpreadConstraints: []
+
+lensWorker:
+ enabled: false
+ replicaCount: 1
+ image:
+ repository: ghcr.io/berriai/litellm-lens-worker
+ tag: ""
+ digest: ""
+ pullPolicy: IfNotPresent
+ tokenSecret:
+ name: ""
+ key: token
+ url: ""
+ tmpSizeLimit: 1Gi
+ resources:
+ requests:
+ cpu: 100m
+ memory: 256Mi
+ limits:
+ memory: 2Gi
+ nodeSelector: {}
+ tolerations: []
+ affinity: {}
diff --git a/litellm-proxy-extras/litellm_proxy_extras/migration_lock.py b/litellm-proxy-extras/litellm_proxy_extras/migration_lock.py
index e4ccbe585a9..bea5e36fd18 100644
--- a/litellm-proxy-extras/litellm_proxy_extras/migration_lock.py
+++ b/litellm-proxy-extras/litellm_proxy_extras/migration_lock.py
@@ -87,3 +87,21 @@ def migration_lock(database_url: str) -> Generator[MigrationCoordinator, None, N
f"Timed out waiting for another v2 migration resolver after {wait_seconds}s. "
f"Check the running migration or increase {MIGRATION_LOCK_TIMEOUT_ENV_VAR}."
)
+
+
+@contextmanager
+def held_migration_lock(connection: "psycopg.Connection[tuple[object, ...]]") -> Generator[bool, None, None]:
+ """A session-level, non-blocking hold of the migration coordinator lock on an autocommit
+ connection, for DDL that cannot run inside a transaction (`CREATE INDEX CONCURRENTLY`).
+ Yields whether the lock was acquired; a v2 resolver or another migration job's index build
+ holding it yields False. Released on exit."""
+ from psycopg.rows import class_row
+
+ with connection.cursor(row_factory=class_row(_LockResult)) as cursor:
+ row: Final = cursor.execute("SELECT pg_try_advisory_lock(%s) AS acquired", (MIGRATION_LOCK_KEY,)).fetchone()
+ acquired: Final = row is not None and row.acquired
+ try:
+ yield acquired
+ finally:
+ if acquired:
+ connection.execute("SELECT pg_advisory_unlock(%s)", (MIGRATION_LOCK_KEY,))
diff --git a/litellm-proxy-extras/litellm_proxy_extras/migration_recovery.py b/litellm-proxy-extras/litellm_proxy_extras/migration_recovery.py
index 9202317c776..5a55b35b255 100644
--- a/litellm-proxy-extras/litellm_proxy_extras/migration_recovery.py
+++ b/litellm-proxy-extras/litellm_proxy_extras/migration_recovery.py
@@ -1,4 +1,5 @@
import hashlib
+import re
import subprocess
from collections.abc import Mapping
from dataclasses import dataclass
@@ -156,3 +157,48 @@ def baseline_current_schema(
"review any feature-specific backfill requirements.",
len(migrations),
)
+
+
+_LINE_COMMENT_RE: Final = re.compile(r"--[^\n]*")
+_BLOCK_COMMENT_RE: Final = re.compile(r"/\*.*?\*/", re.DOTALL)
+_NO_OP_STATEMENT_RE: Final = re.compile(r"^\s*SELECT\s+1\s*$", re.IGNORECASE)
+
+
+def is_inert_migration(script: str) -> bool:
+ """Whether a migration file changes nothing: only comments and `SELECT 1`, so
+ applying it can neither repeat nor skip a database change."""
+ stripped: Final = _LINE_COMMENT_RE.sub("", _BLOCK_COMMENT_RE.sub("", script))
+ return all(not part.strip() or _NO_OP_STATEMENT_RE.match(part) for part in stripped.split(";"))
+
+
+def roll_back_failed_inert_migration(coordinator: MigrationCoordinator, schema: str, migration: Path) -> bool:
+ """Roll back the failed ledger row of a migration whose file in this build is inert,
+ so `migrate deploy` applies the inert file on its next pass. The row records an
+ earlier build's attempt at SQL this build no longer ships (an index now built by the
+ migration job), so no database change can be repeated or skipped by replaying
+ the empty file. The caller commits this checkpoint before the next Prisma command.
+ """
+ from psycopg import sql
+
+ if not is_inert_migration(migration.read_text(encoding="utf-8")):
+ return False
+ coordinator.acquire_prisma_lock()
+ records: Final = _migration_records(coordinator.connection, schema, migration)
+ unfinished: Final = tuple(record for record in records if not record.finished)
+ if len(unfinished) != 1:
+ return False
+ result: Final = coordinator.connection.execute(
+ sql.SQL(
+ "UPDATE {} SET rolled_back_at = current_timestamp "
+ "WHERE id = %s AND finished_at IS NULL AND rolled_back_at IS NULL"
+ ).format(sql.Identifier(schema, "_prisma_migrations")),
+ (unfinished[0].id,),
+ )
+ if result.rowcount != 1:
+ raise RuntimeError("Could not roll back the failed inert migration history row; rerun the database setup.")
+ logger.info(
+ "Rolled back the failed history row of %s: this build ships it as an inert migration, "
+ "its index is built by the migration job",
+ migration.parent.name,
+ )
+ return True
diff --git a/litellm-proxy-extras/litellm_proxy_extras/migrations/20260823000000_add_spend_logs_api_key_starttime_index/migration.sql b/litellm-proxy-extras/litellm_proxy_extras/migrations/20260823000000_add_spend_logs_api_key_starttime_index/migration.sql
index 9a061aaed43..a2bec81ca00 100644
--- a/litellm-proxy-extras/litellm_proxy_extras/migrations/20260823000000_add_spend_logs_api_key_starttime_index/migration.sql
+++ b/litellm-proxy-extras/litellm_proxy_extras/migrations/20260823000000_add_spend_logs_api_key_starttime_index/migration.sql
@@ -1,2 +1,6 @@
--- CreateIndex
-CREATE INDEX IF NOT EXISTS "LiteLLM_SpendLogs_api_key_startTime_idx" ON "LiteLLM_SpendLogs"("api_key", "startTime");
+-- The (api_key, startTime) index on LiteLLM_SpendLogs is built after migrate deploy,
+-- through litellm_proxy_extras/request_log_indexes.py: concurrently on a plain table and
+-- per partition on a partitioned one. The migration job builds it; a serving proxy that
+-- ran the migrations itself builds it in the background once it serves. A migration
+-- cannot do either without blocking spend-log writes or failing on a partitioned table.
+SELECT 1;
diff --git a/litellm-proxy-extras/litellm_proxy_extras/migrations/20260831120001_spend_logs_litellm_call_id_index/migration.sql b/litellm-proxy-extras/litellm_proxy_extras/migrations/20260831120001_spend_logs_litellm_call_id_index/migration.sql
index 62ad5c42ba7..7eba7fc9b97 100644
--- a/litellm-proxy-extras/litellm_proxy_extras/migrations/20260831120001_spend_logs_litellm_call_id_index/migration.sql
+++ b/litellm-proxy-extras/litellm_proxy_extras/migrations/20260831120001_spend_logs_litellm_call_id_index/migration.sql
@@ -1,12 +1,6 @@
--- CreateIndex (CONCURRENTLY)
---
--- Disclaimer:
--- - CREATE INDEX CONCURRENTLY cannot run inside a transaction. This migration must stay a
--- single statement so Prisma Migrate on PostgreSQL can apply it outside a transaction.
--- - Builds are slower and use more I/O than a blocking CREATE INDEX; if the build is
--- interrupted, Postgres may leave an INVALID index that must be dropped and recreated.
--- - Do not edit this file after it has been applied to any database: Prisma checksums
--- migrations; add a new migration instead.
--- - Requires PostgreSQL that supports CONCURRENTLY with IF NOT EXISTS (use a new migration
--- without IF NOT EXISTS if you must support older versions).
-CREATE INDEX CONCURRENTLY IF NOT EXISTS "LiteLLM_SpendLogs_litellm_call_id_idx" ON "LiteLLM_SpendLogs"("litellm_call_id");
+-- The litellm_call_id index on LiteLLM_SpendLogs is built after migrate deploy, through
+-- litellm_proxy_extras/request_log_indexes.py: concurrently on a plain table and per
+-- partition on a partitioned one. The migration job builds it; a serving proxy that ran
+-- the migrations itself builds it in the background once it serves. Postgres refuses
+-- CREATE INDEX CONCURRENTLY on a partitioned parent, so this migration no longer runs it.
+SELECT 1;
diff --git a/litellm-proxy-extras/litellm_proxy_extras/migrations/20260915000000_add_background_interaction_settlement/migration.sql b/litellm-proxy-extras/litellm_proxy_extras/migrations/20260915000000_add_background_interaction_settlement/migration.sql
new file mode 100644
index 00000000000..94d5e98f2a7
--- /dev/null
+++ b/litellm-proxy-extras/litellm_proxy_extras/migrations/20260915000000_add_background_interaction_settlement/migration.sql
@@ -0,0 +1,16 @@
+-- CreateTable
+CREATE TABLE IF NOT EXISTS "LiteLLM_BackgroundInteractionSettlement" (
+ "interaction_id" TEXT NOT NULL,
+ "custom_llm_provider" TEXT NOT NULL,
+ "create_context" JSONB NOT NULL,
+ "created_at" TIMESTAMP(3) NOT NULL DEFAULT CURRENT_TIMESTAMP,
+ "claimed_at" TIMESTAMP(3),
+ "claimed_by" TEXT,
+ "settled_at" TIMESTAMP(3),
+ "outcome" TEXT,
+
+ CONSTRAINT "LiteLLM_BackgroundInteractionSettlement_pkey" PRIMARY KEY ("interaction_id")
+);
+
+-- CreateIndex
+CREATE INDEX IF NOT EXISTS "idx_background_interaction_settlement_claimed_at" ON "LiteLLM_BackgroundInteractionSettlement"("claimed_at");
diff --git a/litellm-proxy-extras/litellm_proxy_extras/migrations/20261001200000_add_autorouter_daily_spend/migration.sql b/litellm-proxy-extras/litellm_proxy_extras/migrations/20261001200000_add_autorouter_daily_spend/migration.sql
new file mode 100644
index 00000000000..ce166b4df45
--- /dev/null
+++ b/litellm-proxy-extras/litellm_proxy_extras/migrations/20261001200000_add_autorouter_daily_spend/migration.sql
@@ -0,0 +1,17 @@
+CREATE TABLE IF NOT EXISTS "LiteLLM_AutoRouterDailySpend" (
+ "date" TEXT NOT NULL,
+ "api_key" TEXT NOT NULL,
+ "user_id" TEXT NOT NULL,
+ "router_name" TEXT NOT NULL,
+ "router_type" TEXT NOT NULL,
+ "turns" INTEGER NOT NULL DEFAULT 0,
+ "spend" DOUBLE PRECISION NOT NULL DEFAULT 0,
+ "saved_spend" DOUBLE PRECISION NOT NULL DEFAULT 0,
+ "savings_estimated_turns" INTEGER NOT NULL DEFAULT 0,
+ "savings_estimated_actual_spend" DOUBLE PRECISION NOT NULL DEFAULT 0,
+ "savings_estimated_saved_spend" DOUBLE PRECISION NOT NULL DEFAULT 0,
+ "classifier_cost" DOUBLE PRECISION NOT NULL DEFAULT 0,
+ "classifier_cost_recorded_turns" INTEGER NOT NULL DEFAULT 0,
+
+ CONSTRAINT "LiteLLM_AutoRouterDailySpend_pkey" PRIMARY KEY ("date", "api_key", "user_id", "router_name", "router_type")
+);
diff --git a/litellm-proxy-extras/litellm_proxy_extras/migrations/20261002220000_lens_worker_scope_index/migration.sql b/litellm-proxy-extras/litellm_proxy_extras/migrations/20261002220000_lens_worker_scope_index/migration.sql
new file mode 100644
index 00000000000..124e5713994
--- /dev/null
+++ b/litellm-proxy-extras/litellm_proxy_extras/migrations/20261002220000_lens_worker_scope_index/migration.sql
@@ -0,0 +1,3 @@
+CREATE INDEX IF NOT EXISTS "LiteLLM_LensWorker_active_scope_idx"
+ON "LiteLLM_LensWorker" USING GIN ((data->'scope') jsonb_path_ops)
+WHERE data @> '{"revoked": false}'::jsonb;
diff --git a/litellm-proxy-extras/litellm_proxy_extras/migrations/20261003000000_add_managed_file_flat_ids_gin_index/migration.sql b/litellm-proxy-extras/litellm_proxy_extras/migrations/20261003000000_add_managed_file_flat_ids_gin_index/migration.sql
new file mode 100644
index 00000000000..b222cc57dab
--- /dev/null
+++ b/litellm-proxy-extras/litellm_proxy_extras/migrations/20261003000000_add_managed_file_flat_ids_gin_index/migration.sql
@@ -0,0 +1,12 @@
+-- CreateIndex (CONCURRENTLY)
+--
+-- Disclaimer:
+-- - CREATE INDEX CONCURRENTLY cannot run inside a transaction. This migration must stay a
+-- single statement so Prisma Migrate on PostgreSQL can apply it outside a transaction.
+-- - Builds are slower and use more I/O than a blocking CREATE INDEX; if the build is
+-- interrupted, Postgres may leave an INVALID index that must be dropped and recreated.
+-- - Do not edit this file after it has been applied to any database: Prisma checksums
+-- migrations; add a new migration instead.
+-- - Requires PostgreSQL that supports CONCURRENTLY with IF NOT EXISTS (use a new migration
+-- without IF NOT EXISTS if you must support older versions).
+CREATE INDEX CONCURRENTLY IF NOT EXISTS "LiteLLM_ManagedFileTable_flat_model_file_ids_idx" ON "LiteLLM_ManagedFileTable" USING GIN ("flat_model_file_ids");
diff --git a/litellm-proxy-extras/litellm_proxy_extras/request_log_indexes.py b/litellm-proxy-extras/litellm_proxy_extras/request_log_indexes.py
new file mode 100644
index 00000000000..6c31e8364a9
--- /dev/null
+++ b/litellm-proxy-extras/litellm_proxy_extras/request_log_indexes.py
@@ -0,0 +1,463 @@
+"""The request-log indexes built after `prisma migrate deploy` instead of by a migration:
+by the migration job, or by a serving proxy that ran the migrations itself (in the
+background, once it serves).
+
+A migration cannot build them: a plain `CREATE INDEX` blocks spend-log inserts for the
+whole build, and `CREATE INDEX CONCURRENTLY` is refused on a partitioned parent
+(db_scripts/partition_spend_logs.sql). `REQUEST_LOG_INDEXES` is the one list to extend;
+names match what Prisma derives from the `@@index` declarations in schema.prisma, so an
+index a database already has is recognized and never rebuilt.
+"""
+
+import hashlib
+import random
+import re
+import time
+from collections.abc import Callable
+from dataclasses import dataclass
+from typing import TYPE_CHECKING, Final
+
+from litellm_proxy_extras._logging import logger
+from litellm_proxy_extras.migration_lock import held_migration_lock
+
+if TYPE_CHECKING:
+ import psycopg
+ from psycopg import sql
+
+
+@dataclass(frozen=True, slots=True)
+class RequestLogIndex:
+ """One index the migration job owns: the table, the exact Prisma index name and the
+ column list as it would be written after `ON `."""
+
+ table: str
+ name: str
+ definition: str
+
+ @property
+ def columns(self) -> tuple[str, ...]:
+ return tuple(re.findall(r'"([^"]+)"', self.definition))
+
+ def partition_index_name(self, partition: str) -> str:
+ """The child index name for one partition, built the way Postgres names the
+ children of a partitioned index, and kept within the 63 byte identifier limit."""
+ name: Final = f"{partition}_{self.name.removeprefix(f'{self.table}_')}"
+ if len(name.encode()) <= _IDENTIFIER_MAX_BYTES:
+ return name
+ digest: Final = hashlib.sha256(name.encode()).hexdigest()[:_DIGEST_LENGTH]
+ budget: Final = _IDENTIFIER_MAX_BYTES - _DIGEST_LENGTH - 1
+ kept: Final = next(name[:length] for length in range(len(name), 0, -1) if len(name[:length].encode()) <= budget)
+ return f"{kept}_{digest}"
+
+
+REQUEST_LOG_INDEXES: Final = (
+ RequestLogIndex("LiteLLM_SpendLogs", "LiteLLM_SpendLogs_api_key_startTime_idx", '("api_key", "startTime")'),
+ RequestLogIndex("LiteLLM_SpendLogs", "LiteLLM_SpendLogs_litellm_call_id_idx", '("litellm_call_id")'),
+)
+
+_IDENTIFIER_MAX_BYTES: Final = 63
+_DDL_LOCK_TIMEOUT: Final = "200ms"
+_DDL_LOCK_ATTEMPTS: Final = 10
+_DDL_RETRY_BASE_SECONDS: Final = 0.25
+_DDL_RETRY_MAX_SECONDS: Final = 8.0
+_LOCK_HANDOVER_SECONDS: Final = 2.0
+_DIGEST_LENGTH: Final = 8
+_CREATE_INDEX_STATEMENT: Final = re.compile(
+ r'^\s*CREATE\s+(?:UNIQUE\s+)?INDEX\s+(?:CONCURRENTLY\s+)?(?:IF\s+NOT\s+EXISTS\s+)?"(?P[^"]+)"\s+ON\b',
+ re.IGNORECASE,
+)
+_TABLE_KIND_SQL: Final = "SELECT c.relkind = 'p' AS partitioned FROM pg_class c WHERE c.oid = to_regclass(%s)"
+_CHILDREN_WITHOUT_THE_INDEX_SQL: Final = (
+ "SELECT child.relname AS name, n.nspname AS schema, child.relkind = 'p' AS partitioned "
+ "FROM pg_inherits i JOIN pg_class child ON child.oid = i.inhrelid "
+ "JOIN pg_namespace n ON n.oid = child.relnamespace "
+ "WHERE i.inhparent = to_regclass(%s) AND NOT EXISTS ("
+ "SELECT 1 FROM pg_inherits attached JOIN pg_index x ON x.indexrelid = attached.inhrelid "
+ "WHERE attached.inhparent = to_regclass(%s) AND x.indrelid = child.oid) "
+ "ORDER BY child.relname"
+)
+_EQUIVALENT_INDEXES_SQL: Final = (
+ "SELECT i.relname AS name, x.indisvalid AS valid "
+ "FROM pg_index x JOIN pg_class i ON i.oid = x.indexrelid JOIN pg_am am ON am.oid = i.relam "
+ "WHERE x.indrelid = to_regclass(%s) AND i.relname <> %s AND am.amname = 'btree' AND NOT x.indisunique "
+ "AND x.indexprs IS NULL AND x.indpred IS NULL AND x.indnkeyatts = x.indnatts "
+ "AND NOT EXISTS (SELECT 1 FROM unnest(x.indoption::int2[]) o WHERE o <> 0) "
+ "AND NOT EXISTS (SELECT 1 FROM unnest(x.indclass::oid[]) c JOIN pg_opclass oc ON oc.oid = c WHERE NOT oc.opcdefault) "
+ "AND NOT EXISTS (SELECT 1 FROM unnest(x.indcollation::oid[]) WITH ORDINALITY c(coll, ord) "
+ "JOIN unnest(x.indkey::int2[]) WITH ORDINALITY k(attnum, ord) ON k.ord = c.ord "
+ "JOIN pg_attribute a ON a.attrelid = x.indrelid AND a.attnum = k.attnum "
+ "WHERE c.coll <> 0 AND c.coll <> a.attcollation) "
+ "AND (SELECT array_agg(a.attname::text ORDER BY k.ord) FROM unnest(x.indkey::int2[]) WITH ORDINALITY k(attnum, ord) "
+ "JOIN pg_attribute a ON a.attrelid = x.indrelid AND a.attnum = k.attnum) = %s::text[] "
+ "AND NOT EXISTS (SELECT 1 FROM pg_inherits WHERE inhrelid = x.indexrelid) "
+ "ORDER BY x.indisvalid DESC, i.relname"
+)
+_INDEX_STATE_SQL: Final = (
+ 'SELECT x.indisvalid AS valid, t.relname AS "table" '
+ "FROM pg_index x JOIN pg_class t ON t.oid = x.indrelid WHERE x.indexrelid = to_regclass(%s)"
+)
+
+
+@dataclass(frozen=True, slots=True)
+class _Relation:
+ name: str
+ schema: str
+ partitioned: bool
+
+
+@dataclass(frozen=True, slots=True)
+class _IndexState:
+ valid: bool
+ table: str
+
+
+@dataclass(frozen=True, slots=True)
+class _EquivalentIndex:
+ name: str
+ valid: bool
+
+
+@dataclass(frozen=True, slots=True)
+class _TableKind:
+ partitioned: bool
+
+
+def filter_request_log_index_diff(diff_sql: str, indexes: tuple[RequestLogIndex, ...] = REQUEST_LOG_INDEXES) -> str:
+ """The `prisma migrate diff` script without the statements that create a migration-job-owned
+ index, which the schema declares and the migrations deliberately do not build."""
+ names: Final = frozenset(index.name for index in indexes)
+ statements: Final = diff_sql.split(";")
+ kept: Final = tuple(statement for statement in statements if not _creates_one_of(statement, names))
+ return ";".join(kept) if any(part.strip() for part in kept) else ""
+
+
+def _creates_one_of(statement: str, names: frozenset[str]) -> bool:
+ match: Final = _CREATE_INDEX_STATEMENT.match(_without_comments(statement))
+ return match is not None and match["index"] in names
+
+
+def _without_comments(statement: str) -> str:
+ return "\n".join(line for line in statement.splitlines() if not line.lstrip().startswith("--"))
+
+
+def _connect(database_url: str) -> "psycopg.Connection[tuple[object, ...]]":
+ import psycopg
+
+ return psycopg.connect(database_url, connect_timeout=10, autocommit=True)
+
+
+def ensure_request_log_indexes(
+ database_url: str,
+ schema: str,
+ indexes: tuple[RequestLogIndex, ...] = REQUEST_LOG_INDEXES,
+ connect: "Callable[[str], psycopg.Connection[tuple[object, ...]]]" = _connect,
+) -> bool:
+ """Build every listed index that is missing or invalid. Each build step runs under
+ the migration coordinator lock, held per statement so a resolver booting on another
+ replica gets in between partitions rather than waiting for the whole table. Any
+ failure is logged and left for the next index build; the result says whether
+ every index ended up valid. Never raises."""
+ import psycopg
+
+ try:
+ with connect(database_url) as connection:
+ connection.execute("SET statement_timeout = 0")
+ results: Final = tuple(_ensure_index(connection, schema, index) for index in indexes)
+ except psycopg.Error as exc:
+ logger.warning("Could not build the request-log indexes, leaving them for the next index build: %s", exc)
+ return False
+ if not all(results):
+ logger.warning("Some request-log indexes are not in place yet, leaving them for the next index build")
+ return False
+ logger.info("Request-log indexes are all in place")
+ return True
+
+
+def _under_migration_lock(connection: "psycopg.Connection[tuple[object, ...]]", step: Callable[[], bool]) -> bool:
+ with held_migration_lock(connection) as held:
+ if not held:
+ logger.info(
+ "Another process holds the migration lock, leaving the request-log indexes to the next index build"
+ )
+ return False
+ return step()
+
+
+def _with_bounded_lock(
+ connection: "psycopg.Connection[tuple[object, ...]]", step: Callable[[], bool], what: str
+) -> bool:
+ """Run `step` under the migration lock with a short lock_timeout, so a DDL statement that has to wait for open
+ transactions holds new writes back for at most that long; retry with capped exponential backoff, holding the
+ migration lock per attempt only and releasing it while sleeping. False when another process holds the migration
+ lock or every attempt timed out."""
+ import psycopg
+ from psycopg import sql
+
+ for attempt in range(_DDL_LOCK_ATTEMPTS):
+ if attempt:
+ time.sleep(min(_DDL_RETRY_MAX_SECONDS, _DDL_RETRY_BASE_SECONDS * 2.0**attempt) * random.uniform(0.5, 1.0))
+ connection.execute(sql.SQL("SET lock_timeout = {}").format(sql.Literal(_DDL_LOCK_TIMEOUT)))
+ try:
+ return _under_migration_lock(connection, step)
+ except psycopg.errors.LockNotAvailable:
+ logger.info("Waiting for open transactions before %s", what)
+ finally:
+ connection.execute("SET lock_timeout = 0")
+ logger.warning(
+ "Could not get the lock for %s without holding writes back, leaving it for the next index build", what
+ )
+ return False
+
+
+def _ensure_index(connection: "psycopg.Connection[tuple[object, ...]]", schema: str, index: RequestLogIndex) -> bool:
+ from psycopg.rows import class_row
+
+ with connection.cursor(row_factory=class_row(_TableKind)) as cursor:
+ table: Final = cursor.execute(_TABLE_KIND_SQL, (_regclass_name(connection, schema, index.table),)).fetchone()
+ if table is None:
+ logger.info("Table %s does not exist yet, skipping index %s", index.table, index.name)
+ return True
+ if table.partitioned:
+ return build_index_on_partitioned_table(connection, schema, index)
+ return _build_leaf_index(connection, schema, index.table, index.name, index)
+
+
+def _regclass_name(connection: "psycopg.Connection[tuple[object, ...]]", schema: str, name: str) -> str:
+ from psycopg import sql
+
+ return sql.Identifier(schema, name).as_string(connection)
+
+
+def _create_index_statement(
+ connection: "psycopg.Connection[tuple[object, ...]]", prefix: "sql.Composed", definition: str
+) -> bytes:
+ return (prefix.as_string(connection) + definition).encode()
+
+
+def _index_state(connection: "psycopg.Connection[tuple[object, ...]]", schema: str, index: str) -> "_IndexState | None":
+ from psycopg.rows import class_row
+
+ with connection.cursor(row_factory=class_row(_IndexState)) as cursor:
+ return cursor.execute(_INDEX_STATE_SQL, (_regclass_name(connection, schema, index),)).fetchone()
+
+
+def _equivalent_indexes(
+ connection: "psycopg.Connection[tuple[object, ...]]",
+ schema: str,
+ table: str,
+ name: str,
+ index: RequestLogIndex,
+) -> tuple[_EquivalentIndex, ...]:
+ """The indexes on `table` other than `name` with the same definition: default btree
+ over the same columns in the same order, no expression, predicate, DESC or custom
+ opclass or collation, and not attached under a partitioned index. Valid ones first."""
+ from psycopg.rows import class_row
+
+ with connection.cursor(row_factory=class_row(_EquivalentIndex)) as cursor:
+ return tuple(
+ cursor.execute(
+ _EQUIVALENT_INDEXES_SQL, (_regclass_name(connection, schema, table), name, list(index.columns))
+ ).fetchall()
+ )
+
+
+def _adopt_equivalent_index(
+ connection: "psycopg.Connection[tuple[object, ...]]",
+ schema: str,
+ table: str,
+ name: str,
+ index: RequestLogIndex,
+) -> bool:
+ """Rename a valid index of the same definition under another name (an operator's
+ hand-built copy, say) to the name this code expects, instead of building a second
+ one. RENAME on an index is a catalog change that lets writes through."""
+ from psycopg import sql
+
+ equivalent: Final = next(
+ (found for found in _equivalent_indexes(connection, schema, table, name, index) if found.valid), None
+ )
+ if equivalent is None:
+ return False
+ logger.info(
+ "Renaming the equivalent index %s on %s to %s instead of building a second one", equivalent.name, table, name
+ )
+ connection.execute(
+ sql.SQL("ALTER INDEX {} RENAME TO {}").format(sql.Identifier(schema, equivalent.name), sql.Identifier(name))
+ )
+ return True
+
+
+def _report_second_copies(
+ connection: "psycopg.Connection[tuple[object, ...]]",
+ schema: str,
+ table: str,
+ name: str,
+ index: RequestLogIndex,
+ concurrently: bool,
+) -> None:
+ """Log every other index of the same definition with the statement that removes it.
+ Dropping is the operator's call: a second copy costs writes and disk, never results."""
+ from psycopg import sql
+
+ drop: Final = "DROP INDEX CONCURRENTLY" if concurrently else "DROP INDEX"
+ for copy in _equivalent_indexes(connection, schema, table, name, index):
+ logger.warning(
+ "Index %s on %s is a second copy of %s and only costs writes and disk; remove it with: %s %s",
+ copy.name,
+ table,
+ name,
+ drop,
+ sql.Identifier(schema, copy.name).as_string(connection),
+ )
+
+
+def _children_without_the_index(
+ connection: "psycopg.Connection[tuple[object, ...]]", schema: str, table: str, index: str
+) -> tuple[_Relation, ...]:
+ from psycopg.rows import class_row
+
+ with connection.cursor(row_factory=class_row(_Relation)) as cursor:
+ return tuple(
+ cursor.execute(
+ _CHILDREN_WITHOUT_THE_INDEX_SQL,
+ (_regclass_name(connection, schema, table), _regclass_name(connection, schema, index)),
+ ).fetchall()
+ )
+
+
+def _build_leaf_index(
+ connection: "psycopg.Connection[tuple[object, ...]]",
+ schema: str,
+ table: str,
+ name: str,
+ index: RequestLogIndex,
+) -> bool:
+ """Build one plain table's or partition's index with CONCURRENTLY so writes keep
+ flowing. The catalog is read under the migration lock, so a replica that saw an
+ invalid index before the lock finds the valid one another replica just built and
+ leaves it. An invalid index left by an interrupted build is dropped and rebuilt; a
+ valid index of the same definition under another name is renamed rather than
+ duplicated; an index of that name on another table is a collision this code will
+ not touch."""
+ from psycopg import sql
+
+ def build() -> bool:
+ existing: Final = _index_state(connection, schema, name)
+ if existing is not None and existing.table != table:
+ logger.warning(
+ "Index %s already exists on %s rather than %s, leaving it alone", name, existing.table, table
+ )
+ return False
+ if existing is not None and existing.valid:
+ return True
+ if existing is not None:
+ logger.info("Dropping the invalid index %s left by an interrupted build on %s", name, table)
+ connection.execute(sql.SQL("DROP INDEX CONCURRENTLY {}").format(sql.Identifier(schema, name)))
+ elif _adopt_equivalent_index(connection, schema, table, name, index):
+ return True
+ logger.info("Building index %s on %s concurrently", name, table)
+ prefix: Final = sql.SQL("CREATE INDEX CONCURRENTLY IF NOT EXISTS {} ON {} ").format(
+ sql.Identifier(name), sql.Identifier(schema, table)
+ )
+ connection.execute(_create_index_statement(connection, prefix, index.definition))
+ built: Final = _index_state(connection, schema, name)
+ return built is not None and built.valid
+
+ current: Final = _index_state(connection, schema, name)
+ if current is None or not current.valid or current.table != table:
+ if not _under_migration_lock(connection, build):
+ return False
+ time.sleep(_LOCK_HANDOVER_SECONDS)
+ _report_second_copies(connection, schema, table, name, index, concurrently=True)
+ return True
+
+
+def build_index_on_partitioned_table(
+ connection: "psycopg.Connection[tuple[object, ...]]",
+ schema: str,
+ index: RequestLogIndex,
+ table: "str | None" = None,
+ name: "str | None" = None,
+) -> bool:
+ """Build the index the way Postgres allows on a partitioned parent: a metadata-only
+ parent index ON ONLY the parent, one CONCURRENTLY build per partition, and ATTACH
+ PARTITION for each child. Partitions that are themselves partitioned get the same
+ treatment one level down. Every step checks the catalog before acting, so an
+ interrupted run resumes where it stopped and a second run finds nothing to do; a
+ parent or child index of the same definition under another name is renamed and
+ used rather than duplicated. The connection must be in autocommit mode. True when
+ the parent index ends up valid."""
+
+ parent_table: Final = index.table if table is None else table
+ parent_index: Final = index.name if name is None else name
+ existing: Final = _index_state(connection, schema, parent_index)
+ if existing is not None and existing.table != parent_table:
+ logger.warning(
+ "Index %s already exists on %s rather than %s, leaving it alone", parent_index, existing.table, parent_table
+ )
+ return False
+ if existing is None and not _with_bounded_lock(
+ connection,
+ lambda: (
+ _adopt_equivalent_index(connection, schema, parent_table, parent_index, index)
+ or _create_parent_index(connection, schema, parent_index, parent_table, index)
+ ),
+ f"creating the parent index {parent_index}",
+ ):
+ return False
+ children: Final = _children_without_the_index(connection, schema, parent_table, parent_index)
+ if not all(_attach_child_index(connection, schema, parent_index, child, index) for child in children):
+ return False
+ final: Final = _index_state(connection, schema, parent_index)
+ if final is None or not final.valid:
+ return False
+ _report_second_copies(connection, schema, parent_table, parent_index, index, concurrently=False)
+ return True
+
+
+def _create_parent_index(
+ connection: "psycopg.Connection[tuple[object, ...]]",
+ schema: str,
+ name: str,
+ table: str,
+ index: RequestLogIndex,
+) -> bool:
+ """Create the metadata-only parent index. The caller bounds Postgres's SHARE lock wait on the parent."""
+ from psycopg import sql
+
+ prefix: Final = sql.SQL("CREATE INDEX IF NOT EXISTS {} ON ONLY {} ").format(
+ sql.Identifier(name), sql.Identifier(schema, table)
+ )
+ statement: Final = _create_index_statement(connection, prefix, index.definition)
+ connection.execute(statement)
+ return True
+
+
+def _attach_child_index(
+ connection: "psycopg.Connection[tuple[object, ...]]",
+ schema: str,
+ parent_index: str,
+ child: _Relation,
+ index: RequestLogIndex,
+) -> bool:
+ from psycopg import sql
+
+ child_index: Final = index.partition_index_name(child.name)
+ built: Final = (
+ build_index_on_partitioned_table(connection, child.schema, index, child.name, child_index)
+ if child.partitioned
+ else _build_leaf_index(connection, child.schema, child.name, child_index, index)
+ )
+ if not built:
+ return False
+
+ def attach() -> bool:
+ connection.execute(
+ sql.SQL("ALTER INDEX {} ATTACH PARTITION {}").format(
+ sql.Identifier(schema, parent_index), sql.Identifier(child.schema, child_index)
+ )
+ )
+ logger.info("Attached index %s on partition %s to %s", child_index, child.name, parent_index)
+ return True
+
+ return _with_bounded_lock(connection, attach, f"attaching {child_index}")
diff --git a/litellm-proxy-extras/litellm_proxy_extras/schema.prisma b/litellm-proxy-extras/litellm_proxy_extras/schema.prisma
index 6f285e9dc39..cf76b764350 100644
--- a/litellm-proxy-extras/litellm_proxy_extras/schema.prisma
+++ b/litellm-proxy-extras/litellm_proxy_extras/schema.prisma
@@ -1144,6 +1144,7 @@ model LiteLLM_ManagedFileTable {
updated_by String?
@@index([unified_file_id])
+ @@index([flat_model_file_ids], type: Gin)
@@index([team_id, created_at(sort: Desc)])
}
@@ -1744,6 +1745,27 @@ model LiteLLM_AutoRouterUserSession {
@@index([user_id, last_turn_at], map: "idx_autorouter_user_session_user_last_turn")
}
+// Auto-routed requests per UTC request day and router: the selected-day money behind the
+// auto-router usage view. Written in the same statement as the session rollup, so a day row
+// and its session row never disagree; corrected in the same transaction as late baselines.
+model LiteLLM_AutoRouterDailySpend {
+ date String
+ api_key String
+ user_id String
+ router_name String
+ router_type String
+ turns Int @default(0)
+ spend Float @default(0)
+ saved_spend Float @default(0)
+ savings_estimated_turns Int @default(0)
+ savings_estimated_actual_spend Float @default(0)
+ savings_estimated_saved_spend Float @default(0)
+ classifier_cost Float @default(0)
+ classifier_cost_recorded_turns Int @default(0)
+
+ @@id([date, api_key, user_id, router_name, router_type])
+}
+
// Shadow eval: evaluation of an auto-router against one or more keys' live traffic, in
// either direction. forward duplicates the requests the keys did not route through the
// router through it, answering whether they should adopt it; reverse duplicates the
@@ -1895,6 +1917,22 @@ model LiteLLM_WorkflowMessage {
@@index([run_id])
}
+// Pending billing settlements for background interactions, keyed by the
+// interaction id so any replica can settle one that another replica created.
+// `claimed_at` is the exactly-once gate: the first conditional update wins.
+model LiteLLM_BackgroundInteractionSettlement {
+ interaction_id String @id
+ custom_llm_provider String
+ create_context Json
+ created_at DateTime @default(now())
+ claimed_at DateTime?
+ claimed_by String?
+ settled_at DateTime?
+ outcome String?
+
+ @@index([claimed_at], map: "idx_background_interaction_settlement_claimed_at")
+}
+
model LiteLLM_Lens {
id String @id
version Int @default(0)
diff --git a/litellm-proxy-extras/litellm_proxy_extras/utils.py b/litellm-proxy-extras/litellm_proxy_extras/utils.py
index 2f74df63367..1fd292b8137 100644
--- a/litellm-proxy-extras/litellm_proxy_extras/utils.py
+++ b/litellm-proxy-extras/litellm_proxy_extras/utils.py
@@ -1,3 +1,4 @@
+import functools
import glob
import os
import random
@@ -5,14 +6,17 @@ import re
import shutil
import subprocess
import tempfile
+import threading
import time
from collections.abc import Callable
from dataclasses import dataclass, replace
from pathlib import Path
-from typing import TYPE_CHECKING, Final, Optional
+from typing import TYPE_CHECKING, Final, Optional, Union
+from urllib.parse import unquote, urlsplit
from litellm_proxy_extras import prisma_toolchain
from litellm_proxy_extras._logging import logger
+from litellm_proxy_extras.migration_lock import held_migration_lock
from litellm_proxy_extras.prisma_toolchain import (
PRISMA_COMMAND_TIMEOUT_ENV_VAR,
PRISMA_MIGRATE_DEPLOY_TIMEOUT_ENV_VAR,
@@ -24,6 +28,7 @@ from litellm_proxy_extras.replica_identity import (
REPLICA_IDENTITY_FULL_ENV_VAR,
apply_replica_identity_full,
)
+from litellm_proxy_extras.request_log_indexes import ensure_request_log_indexes, filter_request_log_index_diff
if TYPE_CHECKING:
import psycopg
@@ -75,6 +80,23 @@ class _InvalidIndex:
table_size: str
MAX_MIGRATE_DEPLOY_ATTEMPTS = 4
+LIBPQ_URL_PARAMS: Final = frozenset(
+ {
+ "sslmode",
+ "sslcert",
+ "sslkey",
+ "sslrootcert",
+ "sslpassword",
+ "application_name",
+ "connect_timeout",
+ "client_encoding",
+ "options",
+ "service",
+ "gssencmode",
+ "krbsrvname",
+ "target_session_attrs",
+ }
+)
@dataclass(frozen=True)
@@ -182,6 +204,66 @@ def _max_migration_timestamp(names) -> int:
return max(_migration_timestamp(n) for n in names)
+_REDACTED: Final = "REDACTED"
+_PASSWORD_QUERY_KEYS: Final = frozenset(("password", "sslpassword"))
+
+
+@functools.cache
+def _secret_shape_redactor() -> Callable[[str], str]:
+ try:
+ from litellm._logging import redact_secrets
+ except ImportError:
+ return lambda text: text
+ return redact_secrets
+
+
+def _url_passwords(url: str) -> frozenset[str]:
+ try:
+ parts: Final = urlsplit(url)
+ except ValueError:
+ return frozenset()
+ query_pairs: Final = tuple(pair.partition("=") for pair in parts.query.split("&"))
+ raw_query_passwords: Final = tuple(
+ value for key, separator, value in query_pairs if separator and key.lower() in _PASSWORD_QUERY_KEYS
+ )
+ raw_passwords: Final = ((parts.password,) if parts.password else ()) + raw_query_passwords
+ return frozenset(password for password in raw_passwords + tuple(map(unquote, raw_passwords)) if password)
+
+
+def _configured_database_passwords() -> frozenset[str]:
+ database_url: Final = os.getenv("DATABASE_URL")
+ direct_url: Final = os.getenv("DIRECT_URL")
+ database_passwords: Final = _url_passwords(database_url) if database_url else frozenset()
+ direct_passwords: Final = _url_passwords(direct_url) if direct_url else frozenset()
+ return database_passwords | direct_passwords
+
+
+def _redact_credentials(text: str) -> str:
+ """Mask configured database passwords before passing the text to LiteLLM redaction."""
+ passwords: Final = sorted(_configured_database_passwords(), key=len, reverse=True)
+ alternation: Final = "|".join(re.escape(password) for password in passwords)
+ password_pattern: Final = (
+ re.compile(rf"(?P:|password=)(?:{alternation})(?=@|&|$|[\s'\"\]),])", re.IGNORECASE)
+ if passwords
+ else None
+ )
+ result: Final = password_pattern.sub(rf"\g{_REDACTED}", text) if password_pattern is not None else text
+ return _secret_shape_redactor()(result)
+
+
+def _redacted_command(command: object) -> Union[str, tuple[str, ...], list[str]]:
+ if isinstance(command, tuple):
+ return tuple(_redact_credentials(str(argument)) for argument in command)
+ if isinstance(command, list):
+ return [_redact_credentials(str(argument)) for argument in command]
+ return _redact_credentials(str(command))
+
+
+def _redact_command_error(error: subprocess.CalledProcessError) -> str:
+ redacted_command: Final = _redacted_command(error.cmd)
+ return str(subprocess.CalledProcessError(error.returncode, redacted_command))
+
+
def _get_prisma_command() -> str:
"""Get the Prisma command to use, bypassing Python wrapper in offline mode."""
if str_to_bool(os.getenv("PRISMA_OFFLINE_MODE")):
@@ -295,7 +377,8 @@ class ProxyExtrasDBManager:
return False
except subprocess.CalledProcessError as e:
logger.warning(
- f"Error creating baseline migration: {e}, {e.stderr}, {e.stdout}"
+ f"Error creating baseline migration: {_redact_command_error(e)}, "
+ f"{_redact_credentials(str(e.stderr))}, {_redact_credentials(str(e.stdout))}"
)
raise e
@@ -333,9 +416,8 @@ class ProxyExtrasDBManager:
pass
@staticmethod
- def _failed_migration_logs(migration_name: str) -> Optional[str]:
- """Return failed migration logs, or None if the ledger is unavailable."""
- database_url = os.getenv("DATABASE_URL")
+ def _read_migration_ledger(query: str, params: tuple[str, ...]) -> "tuple[object, ...] | None":
+ database_url: Final = os.getenv("DATABASE_URL")
if not database_url:
return None
@@ -344,28 +426,37 @@ class ProxyExtrasDBManager:
except ImportError:
return None
- cleaned_url = ProxyExtrasDBManager._strip_prisma_query_params(database_url)
- ledger_table = psycopg.sql.SQL("{}.{}").format(
- psycopg.sql.Identifier(
- ProxyExtrasDBManager._prisma_schema_param(database_url) or "public"
- ),
+ cleaned_url: Final = ProxyExtrasDBManager._strip_prisma_query_params(database_url)
+ ledger_table: Final = psycopg.sql.SQL("{}.{}").format(
+ psycopg.sql.Identifier(ProxyExtrasDBManager._prisma_schema_param(database_url) or "public"),
psycopg.sql.Identifier("_prisma_migrations"),
)
try:
- with psycopg.connect(
- cleaned_url, connect_timeout=10, autocommit=True
- ) as conn:
- row = conn.execute(
- psycopg.sql.SQL(
- "SELECT logs FROM {} "
- "WHERE migration_name = %s AND finished_at IS NULL "
- "AND rolled_back_at IS NULL"
- ).format(ledger_table),
- (migration_name,),
- ).fetchone()
+ with psycopg.connect(cleaned_url, connect_timeout=10, autocommit=True) as conn:
+ row: Final = conn.execute(psycopg.sql.SQL(query).format(ledger_table), params).fetchone()
except (psycopg.OperationalError, psycopg.DatabaseError):
return None
- return (row[0] or "") if row else ""
+ return tuple(row) if row is not None else ()
+
+ @staticmethod
+ def _failed_migration_logs(migration_name: str, started_at: str) -> Optional[str]:
+ row: Final = ProxyExtrasDBManager._read_migration_ledger(
+ "SELECT logs FROM {} WHERE migration_name = %s AND started_at = %s::timestamptz "
+ "AND finished_at IS NULL AND rolled_back_at IS NULL",
+ (migration_name, started_at),
+ )
+ if row is None:
+ return None
+ return row[0] if row and isinstance(row[0], str) else ""
+
+ @staticmethod
+ def _failed_migration_recovered(migration_name: str, started_at: str) -> bool:
+ row: Final = ProxyExtrasDBManager._read_migration_ledger(
+ "SELECT 1 FROM {} WHERE migration_name = %s AND started_at = %s::timestamptz "
+ "AND (finished_at IS NOT NULL OR rolled_back_at IS NOT NULL)",
+ (migration_name, started_at),
+ )
+ return bool(row)
@staticmethod
def _resolve_specific_migration(migration_name: str):
@@ -433,6 +524,21 @@ class ProxyExtrasDBManager:
return True
return False
+ @staticmethod
+ def _filter_migration_job_owned_drift(diff_sql: str, partitioned: bool | None = None) -> str:
+ """The drift script without the indexes the migration job builds (the schema
+ declares them, the migrations deliberately do not) and, when LiteLLM_SpendLogs
+ is partitioned, without its primary-key rewrite and partitioning artifacts."""
+ without_indexes: Final = filter_request_log_index_diff(diff_sql)
+ is_partitioned: Final = ProxyExtrasDBManager.spend_logs_is_partitioned() if partitioned is None else partitioned
+ if not is_partitioned:
+ return without_indexes
+ logger.info(
+ "LiteLLM_SpendLogs is partitioned; removed its primary-key "
+ "rewrite and partitioning artifacts from the drift script"
+ )
+ return filter_partitioned_spend_logs_diff(without_indexes)
+
@staticmethod
def _resolve_all_migrations(
migrations_dir: str, schema_path: str, mark_all_applied: bool = True
@@ -513,21 +619,14 @@ class ProxyExtrasDBManager:
return
logger.info(f"Migration diff created at {diff_sql_path}")
- if ProxyExtrasDBManager.spend_logs_is_partitioned():
- filtered_sql = filter_partitioned_spend_logs_diff(
- diff_sql_path.read_text()
- )
- diff_sql_path.write_text(filtered_sql)
- logger.info(
- "LiteLLM_SpendLogs is partitioned; removed its primary-key "
- "rewrite and partitioning artifacts from the drift script"
- )
- if not filtered_sql.strip():
- logger.info("Drift script is empty after filtering; nothing to apply")
- if not mark_all_applied:
- return
- ProxyExtrasDBManager._mark_migrations_applied(migrations_dir)
+ filtered_sql: Final = ProxyExtrasDBManager._filter_migration_job_owned_drift(diff_sql_path.read_text())
+ diff_sql_path.write_text(filtered_sql)
+ if not filtered_sql.strip():
+ logger.info("Drift script is empty after filtering; nothing to apply")
+ if not mark_all_applied:
return
+ ProxyExtrasDBManager._mark_migrations_applied(migrations_dir)
+ return
# 2. Run prisma db execute to apply the migration
applied_ok = False
@@ -678,30 +777,43 @@ class ProxyExtrasDBManager:
@staticmethod
def _strip_prisma_query_params(url: str) -> str:
- """Remove Prisma-specific query params (connection_limit, pool_timeout,
- schema, etc.) from DATABASE_URL so psycopg can parse it."""
+ """Rewrite a Prisma-dialect URL for libpq: drop the Prisma-only params
+ (connection_limit, pool_timeout, schema, pgbouncer, sslaccept, ...) and
+ translate Prisma's TLS params back, since libpq reads ``sslcert`` as a
+ client certificate where Prisma reads it as the CA."""
from urllib.parse import parse_qsl, quote, urlencode, urlparse, urlunparse
- parsed = urlparse(url)
+ parsed: Final = urlparse(url)
if not parsed.query:
return url
- libpq_params = {
- "sslmode",
- "sslcert",
- "sslkey",
- "sslrootcert",
- "sslpassword",
- "application_name",
- "connect_timeout",
- "client_encoding",
- "options",
- "service",
- "gssencmode",
- "krbsrvname",
- "target_session_attrs",
- }
- kept = [(k, v) for k, v in parse_qsl(parsed.query) if k in libpq_params]
- return urlunparse(parsed._replace(query=urlencode(kept, quote_via=quote)))
+ pairs: Final = tuple(parse_qsl(parsed.query))
+ kept: Final = tuple((k, v) for k, v in pairs if k in LIBPQ_URL_PARAMS)
+ sslaccept: Final = next((v for k, v in pairs if k == "sslaccept"), None)
+ libpq_pairs: Final = ProxyExtrasDBManager._libpq_tls_params(kept, sslaccept)
+ return urlunparse(parsed._replace(query=urlencode(libpq_pairs, quote_via=quote)))
+
+ @staticmethod
+ def _libpq_tls_params(
+ pairs: "tuple[tuple[str, str], ...]", sslaccept: "str | None"
+ ) -> "tuple[tuple[str, str], ...]":
+ """Undo ``translate_libpq_ssl_params``. Prisma's ``sslcert`` is the CA and
+ ``sslaccept=strict`` checks chain and hostname, which libpq only does in
+ ``sslmode=verify-full``, so strict becomes ``sslrootcert`` plus
+ ``verify-full`` whatever ``sslmode`` said (``disable`` stays off). Prisma
+ defaults an absent ``sslaccept`` to ``accept_invalid_certs`` and anything
+ else to strict. Without strict it checks nothing, so the CA is dropped and
+ ``sslmode`` is kept as is: libpq only verifies when a root cert is present.
+ A URL that also carries ``sslkey`` is libpq's own client-certificate form
+ and is kept."""
+ keys: Final = frozenset(k for k, _ in pairs)
+ if "sslcert" not in keys or "sslkey" in keys:
+ return pairs
+ sslmode: Final = next((v for k, v in pairs if k == "sslmode"), None)
+ rest: Final = tuple((k, v) for k, v in pairs if k not in ("sslcert", "sslmode"))
+ if sslaccept in (None, "accept_invalid_certs") or sslmode == "disable":
+ return rest if sslmode is None else rest + (("sslmode", sslmode),)
+ root_cert: Final = tuple(("sslrootcert", v) for k, v in pairs if k == "sslcert" and "sslrootcert" not in keys)
+ return rest + root_cert + (("sslmode", "verify-full"),)
@staticmethod
def _warn_if_db_ahead_of_head(migrations_dir: str) -> None:
@@ -800,7 +912,7 @@ class ProxyExtrasDBManager:
conn.execute(statement)
except psycopg.Error as e:
logger.warning(
- "Could not repair invalid index %s.%s, will retry on the next startup. "
+ "Could not repair invalid index %s.%s, will retry on the next database setup run. "
"If this keeps happening, run `%s` by hand as the index owner. Error: %s",
index.schema,
index.name,
@@ -811,16 +923,21 @@ class ProxyExtrasDBManager:
logger.info("%s invalid index %s.%s", action, index.schema, index.name)
@staticmethod
- def repair_invalid_indexes(lock_timeout: str = "30s") -> bool:
+ def repair_invalid_indexes(
+ lock_timeout: str = "30s",
+ repair: "Callable[[psycopg.Connection[tuple[str, str, str]], _InvalidIndex], None] | None" = None,
+ ) -> bool:
"""Rebuild LiteLLM indexes an interrupted CREATE INDEX CONCURRENTLY left
INVALID (a migration deadlock between replicas is the usual cause; the
retried migration skips them because of IF NOT EXISTS). Never raises:
returns True when no invalid index remains, False when the repair was
- skipped or failed and will be retried on the next startup. Looks in the
+ skipped or failed and will be retried on the next database setup run. Looks in the
schema DATABASE_URL names, the only URL Prisma migrates through, but
connects over DIRECT_URL when set: the session settings, the advisory
lock and REINDEX CONCURRENTLY all need one server session, which a
- transaction pooler does not give."""
+ transaction pooler does not give. Each rebuild holds the migration
+ coordinator lock on its own, like the migration job's index build, so a resolver
+ booting on another replica waits for one index at most."""
prisma_url: Final = os.getenv("DATABASE_URL")
if not prisma_url:
return False
@@ -856,20 +973,53 @@ class ProxyExtrasDBManager:
if lock_row is None or not lock_row[0]:
logger.info("Another replica is already rebuilding the invalid indexes, skipping")
return False
- for index in ProxyExtrasDBManager._invalid_litellm_indexes(conn, schema):
- ProxyExtrasDBManager._repair_index(conn, index)
+ repair_one: Final = repair or ProxyExtrasDBManager._repair_index
+ repaired: Final = all(
+ ProxyExtrasDBManager._repair_under_migration_lock(conn, schema, index, repair_one)
+ for index in found
+ )
+ if not repaired:
+ return False
remaining: Final = ProxyExtrasDBManager._invalid_litellm_indexes(conn, schema)
except psycopg.Error as e:
- logger.warning("Could not check for invalid indexes, will retry on the next startup. Error: %s", e)
+ logger.warning(
+ "Could not check for invalid indexes, will retry on the next database setup run. Error: %s", e
+ )
return False
return not remaining
+ @staticmethod
+ def _repair_under_migration_lock(
+ conn: "psycopg.Connection[tuple[str, str, str]]",
+ schema: str,
+ index: _InvalidIndex,
+ repair: "Callable[[psycopg.Connection[tuple[str, str, str]], _InvalidIndex], None]",
+ ) -> bool:
+ """Rebuild one index under the migration coordinator lock, skipping it when a
+ migration job finished or dropped it in the meantime. False when another process
+ holds the lock, so the check waits for the next database setup run."""
+ with held_migration_lock(conn) as held:
+ if not held:
+ logger.info(
+ "Another process is building indexes under the migration lock, leaving the "
+ "invalid index check to the next database setup run"
+ )
+ return False
+ still_invalid: Final = ProxyExtrasDBManager._invalid_litellm_indexes(conn, schema)
+ if any(found.schema == index.schema and found.name == index.name for found in still_invalid):
+ repair(conn, index)
+ return True
+
@staticmethod
def _setup_database_v2(use_migrate: bool) -> bool:
if not use_migrate:
return ProxyExtrasDBManager._run_database_v2(False)
from litellm_proxy_extras.migration_lock import migration_environment, migration_lock
- from litellm_proxy_extras.migration_recovery import baseline_current_schema, recover_completed_migration
+ from litellm_proxy_extras.migration_recovery import (
+ baseline_current_schema,
+ recover_completed_migration,
+ roll_back_failed_inert_migration,
+ )
database_url: Final = os.environ.get("DATABASE_URL")
if not database_url:
@@ -884,7 +1034,9 @@ class ProxyExtrasDBManager:
if not migration.is_file():
return False
with migration_lock(lock_url) as coordinator:
- return recover_completed_migration(coordinator, schema, migration)
+ return recover_completed_migration(coordinator, schema, migration) or roll_back_failed_inert_migration(
+ coordinator, schema, migration
+ )
def baseline_existing(migrations_dir: str) -> None:
with migration_lock(lock_url) as coordinator:
@@ -1021,6 +1173,11 @@ class ProxyExtrasDBManager:
return match.group(1) if match else None
return None
+ @staticmethod
+ def _v2_failed_migration_started_at(stderr: str, migration_name: str) -> "str | None":
+ match: Final = re.search(rf"`{re.escape(migration_name)}` migration started at ([^\r\n]+?) failed", stderr)
+ return match.group(1) if match else None
+
@staticmethod
def _v2_roll_back_migration_best_effort(migration_name: str) -> None:
from litellm_proxy_extras.migration_lock import migration_environment
@@ -1049,8 +1206,11 @@ class ProxyExtrasDBManager:
if "P3009" in stderr:
migration_name = ProxyExtrasDBManager._v2_failed_migration_name(stderr)
- if migration_name:
- ledger_logs = ProxyExtrasDBManager._failed_migration_logs(migration_name)
+ started_at: Final = (
+ ProxyExtrasDBManager._v2_failed_migration_started_at(stderr, migration_name) if migration_name else None
+ )
+ if migration_name and started_at:
+ ledger_logs: Final = ProxyExtrasDBManager._failed_migration_logs(migration_name, started_at)
if ledger_logs and _MIGRATION_DEADLOCK_MARKER in ledger_logs:
logger.info(
"Migration %s failed in a concurrent migrate deploy "
@@ -1059,6 +1219,14 @@ class ProxyExtrasDBManager:
)
ProxyExtrasDBManager._v2_roll_back_migration_best_effort(migration_name)
return budget.spend()
+ if ProxyExtrasDBManager._failed_migration_recovered(migration_name, started_at):
+ logger.info(
+ "Migration %s started at %s was already rolled back or completed by a concurrent "
+ "migrate deploy, retrying",
+ migration_name,
+ started_at,
+ )
+ return budget.spend()
raise RuntimeError(
"Migration completion could not be verified. LiteLLM startup has stopped.\n\n"
f"Prisma migration history (migration name and start time):\n{stderr}\n\n"
@@ -1177,13 +1345,16 @@ class ProxyExtrasDBManager:
)
@staticmethod
- def setup_database(
- use_migrate: bool = False, use_v2_resolver: bool = False
- ) -> bool:
+ def setup_database(use_migrate: bool = False, use_v2_resolver: bool = False) -> bool:
"""
Set up the database using either prisma migrate or prisma db push
Uses migrations from litellm-proxy-extras package
+ The request-log indexes in `REQUEST_LOG_INDEXES` are not built here: the
+ migration job builds them through `run_migration_job`, and a serving proxy that
+ ran the migrations itself starts them through `start_request_log_index_build`
+ once it is ready to serve.
+
Args:
use_migrate: Whether to use prisma migrate instead of db push
use_v2_resolver: Opt into the v2 migration resolver (safer during
@@ -1200,10 +1371,48 @@ class ProxyExtrasDBManager:
migrated = ProxyExtrasDBManager._run_migrations(
use_migrate=use_migrate, use_v2_resolver=use_v2_resolver
)
- if migrated:
- ProxyExtrasDBManager.repair_invalid_indexes()
- ProxyExtrasDBManager.apply_replica_identity_full_if_requested()
- return migrated
+ if not migrated:
+ return False
+ ProxyExtrasDBManager.repair_invalid_indexes()
+ ProxyExtrasDBManager.apply_replica_identity_full_if_requested()
+ return True
+
+ @staticmethod
+ def build_request_log_indexes(build: Callable[[str, str], bool] = ensure_request_log_indexes) -> bool:
+ """Build the indexes in `REQUEST_LOG_INDEXES` on the writer, in the schema the
+ migrations target. Idempotent and never raises; False when an index is still
+ missing or invalid, so the migration job reports it and gets rerun instead of
+ leaving the table unindexed until the next deploy."""
+ database_url: Final = os.environ.get("DATABASE_URL")
+ if not database_url:
+ return True
+ direct_url: Final = ProxyExtrasDBManager._strip_prisma_query_params(
+ os.environ.get("DIRECT_URL") or database_url
+ )
+ schema: Final = ProxyExtrasDBManager._prisma_schema_param(database_url) or "public"
+ return build(direct_url, schema)
+
+ @staticmethod
+ def run_migration_job(
+ use_migrate: bool = False,
+ use_v2_resolver: bool = False,
+ setup: Callable[[bool, bool], bool] = setup_database,
+ build: Callable[[], bool] = build_request_log_indexes,
+ ) -> bool:
+ """The migration job's whole run: `setup_database`, then the request-log indexes,
+ built synchronously so the job exits only once they are in place. False when the
+ migrations failed or an index could not be built, so the Job is rerun."""
+ return setup(use_migrate, use_v2_resolver) and build()
+
+ @staticmethod
+ def start_request_log_index_build(build: Callable[[], bool] = build_request_log_indexes) -> threading.Thread:
+ """A serving proxy that ran the migrations itself (schema updates not disabled)
+ builds the request-log indexes on a daemon thread, so a long build never delays
+ readiness. A build that could not finish is logged and picked up by the next boot
+ or the migration job."""
+ thread: Final = threading.Thread(target=build, name="litellm-request-log-indexes", daemon=True)
+ thread.start()
+ return thread
@staticmethod
def _run_migrations(use_migrate: bool, use_v2_resolver: bool) -> bool:
@@ -1247,15 +1456,16 @@ class ProxyExtrasDBManager:
logger.info("✅ Post-migration sanity check completed")
return True
except subprocess.CalledProcessError as e:
- logger.info(f"prisma db error: {e.stderr}, e: {e.stdout}")
- if "P3009" in e.stderr:
+ stderr: Final = str(e.stderr or "")
+ logger.info(f"prisma db error: {stderr}, e: {e.stdout}")
+ if "P3009" in stderr:
# Extract the failed migration name from the error message
migration_match = re.search(
- r"`(\d+_.*)` migration", e.stderr
+ r"`(\d+_.*)` migration", stderr
)
if migration_match:
failed_migration = migration_match.group(1)
- if ProxyExtrasDBManager._is_idempotent_error(e.stderr):
+ if ProxyExtrasDBManager._is_idempotent_error(stderr):
logger.info(
f"Migration {failed_migration} failed due to idempotent error (e.g., column already exists), resolving as applied"
)
@@ -1311,8 +1521,8 @@ class ProxyExtrasDBManager:
f"✅ Migration {failed_migration} marked as rolled back... retrying"
)
elif (
- "P3005" in e.stderr
- and "database schema is not empty" in e.stderr
+ "P3005" in stderr
+ and "database schema is not empty" in stderr
):
logger.info(
"Database schema is not empty, creating baseline migration. In read-only file system, please set an environment variable `LITELLM_MIGRATION_DIR` to a writable directory to enable migrations. Learn more - https://docs.litellm.ai/docs/proxy/prod#read-only-file-system"
@@ -1326,13 +1536,13 @@ class ProxyExtrasDBManager:
)
logger.info("✅ All migrations resolved.")
return True
- elif "P3018" in e.stderr:
+ elif "P3018" in stderr:
# Check if this is a permission error or idempotent error
- if ProxyExtrasDBManager._is_permission_error(e.stderr):
+ if ProxyExtrasDBManager._is_permission_error(stderr):
# Permission errors should NOT be marked as applied
# Extract migration name for logging
migration_match = re.search(
- r"Migration name: (\d+_.*)", e.stderr
+ r"Migration name: (\d+_.*)", stderr
)
migration_name = (
migration_match.group(1)
@@ -1342,7 +1552,7 @@ class ProxyExtrasDBManager:
logger.error(
f"❌ Migration {migration_name} failed due to insufficient permissions. "
- f"Please check database user privileges. Error: {e.stderr}"
+ f"Please check database user privileges. Error: {stderr}"
)
# Mark as rolled back and exit with error
@@ -1365,7 +1575,7 @@ class ProxyExtrasDBManager:
f"was NOT applied. Please grant necessary database permissions and retry."
) from e
- elif ProxyExtrasDBManager._is_idempotent_error(e.stderr):
+ elif ProxyExtrasDBManager._is_idempotent_error(stderr):
# Idempotent errors mean the migration has effectively been applied
logger.info(
"Migration failed due to idempotent error (e.g., column already exists), "
@@ -1373,7 +1583,7 @@ class ProxyExtrasDBManager:
)
# Extract the migration name from the error message
migration_match = re.search(
- r"Migration name: (\d+_.*)", e.stderr
+ r"Migration name: (\d+_.*)", stderr
)
if migration_match:
migration_name = migration_match.group(1)
@@ -1422,9 +1632,14 @@ class ProxyExtrasDBManager:
logger.warning(
f"P3018 error encountered but could not classify "
f"as permission or idempotent error. "
- f"Error: {e.stderr}"
+ f"Error: {stderr}"
)
raise
+ else:
+ logger.error(
+ "prisma migrate deploy failed with an error the resolver does not handle: "
+ f"{_redact_credentials(stderr)}"
+ )
else:
if ProxyExtrasDBManager.spend_logs_is_partitioned():
raise RuntimeError(PARTITIONED_SPEND_LOGS_PUSH_ERROR)
@@ -1439,7 +1654,7 @@ class ProxyExtrasDBManager:
)
return True
except subprocess.TimeoutExpired:
- logger.warning(
+ logger.error(
"Attempt %s timed out. Raise %s if this database needs longer to apply its schema.",
attempt + 1,
PRISMA_MIGRATE_DEPLOY_TIMEOUT_ENV_VAR if use_migrate else PRISMA_COMMAND_TIMEOUT_ENV_VAR,
@@ -1452,7 +1667,12 @@ class ProxyExtrasDBManager:
if attempts_left > 0
else ""
)
- logger.info(f"The process failed to execute. Details: {e}.{retry_msg}")
+ stderr_detail: Final = (
+ f" stderr: {_redact_credentials(str(e.stderr))}" if e.stderr else ""
+ )
+ logger.error(
+ f"The process failed to execute. Details: {_redact_command_error(e)}.{stderr_detail}{retry_msg}"
+ )
time.sleep(random.randrange(5, 15))
finally:
os.chdir(original_dir)
diff --git a/litellm-proxy-extras/pyproject.toml b/litellm-proxy-extras/pyproject.toml
index 2e2f3f2ce5a..79549a88cd9 100644
--- a/litellm-proxy-extras/pyproject.toml
+++ b/litellm-proxy-extras/pyproject.toml
@@ -1,6 +1,6 @@
[project]
name = "litellm-proxy-extras"
-version = "0.4.103"
+version = "0.4.105"
description = "Additional files for the LiteLLM Proxy. Reduces the size of the main litellm package."
readme = "README.md"
requires-python = ">=3.9"
@@ -30,7 +30,7 @@ required-version = ">=0.10.9"
module-root = ""
[tool.commitizen]
-version = "0.4.103"
+version = "0.4.105"
version_files = [
"pyproject.toml:^version",
"../pyproject.toml:litellm-proxy-extras==",
diff --git a/litellm-rust/Cargo.lock b/litellm-rust/Cargo.lock
index 0ad05d99e76..348beeb3813 100644
--- a/litellm-rust/Cargo.lock
+++ b/litellm-rust/Cargo.lock
@@ -97,6 +97,53 @@ version = "1.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "03918c3dbd7701a85c6b9887732e2921175f26c350b4563841d0958c21d57e6d"
+[[package]]
+name = "askama"
+version = "0.16.1"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "6024d73179f43f15ccd2b881bfea6fee7f3a46ec53f33b52210dea749ebebaa4"
+dependencies = [
+ "askama_macros",
+ "itoa",
+ "percent-encoding",
+ "serde",
+ "serde_json",
+]
+
+[[package]]
+name = "askama_derive"
+version = "0.16.1"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "071ee5ebf2138e3ad180e0aacf6940c2cab5e6d8333741d9925c7bee2b153f39"
+dependencies = [
+ "askama_parser",
+ "memchr",
+ "proc-macro2",
+ "quote",
+ "rustc-hash",
+ "syn 3.0.6",
+]
+
+[[package]]
+name = "askama_macros"
+version = "0.16.1"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "643e1c7cbb6aec1d920332fe51a7c0d8219e273dcb8602db03f5263e4d16487b"
+dependencies = [
+ "askama_derive",
+]
+
+[[package]]
+name = "askama_parser"
+version = "0.16.1"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "2c5ae75772275d268b03ab8bdccdd12117b6169ee23256942b34e46c9f476583"
+dependencies = [
+ "rustc-hash",
+ "unicode-ident",
+ "winnow 1.0.4",
+]
+
[[package]]
name = "asn1-rs"
version = "0.7.2"
@@ -4038,6 +4085,26 @@ dependencies = [
"strum",
]
+[[package]]
+name = "litellm-migrate"
+version = "0.1.0"
+dependencies = [
+ "litellm-migrate-macros",
+ "rstest",
+]
+
+[[package]]
+name = "litellm-migrate-macros"
+version = "0.1.0"
+dependencies = [
+ "proc-macro2",
+ "quote",
+ "rstest",
+ "syn 2.0.119",
+ "tempfile",
+ "thiserror 2.0.19",
+]
+
[[package]]
name = "litellm-model-catalog"
version = "0.1.0"
@@ -4086,9 +4153,12 @@ dependencies = [
"litellm-secrets",
"litellm-secrets-aws",
"litellm-secrets-types",
+ "litellm-storage-clickhouse",
"litellm-token-counter",
"litellm-traces",
+ "litellm-traces-clickhouse",
"litellm-tracing",
+ "prost",
"pyo3",
"pyo3-async-runtimes",
"qdrant-client",
@@ -4288,6 +4358,21 @@ dependencies = [
"veil",
]
+[[package]]
+name = "litellm-storage-clickhouse"
+version = "0.1.0"
+dependencies = [
+ "flate2",
+ "litellm-http",
+ "rstest",
+ "serde",
+ "serde_json",
+ "thiserror 2.0.19",
+ "tokio",
+ "url",
+ "wiremock",
+]
+
[[package]]
name = "litellm-testkit"
version = "0.1.0"
@@ -4368,20 +4453,67 @@ dependencies = [
name = "litellm-traces"
version = "0.1.0"
dependencies = [
+ "askama",
"base64 0.22.1",
- "flate2",
- "litellm-http",
+ "criterion",
+ "indexmap 2.14.0",
+ "litellm-llms-types",
+ "macro_rules_attribute",
"opentelemetry-proto",
"prost",
"rstest",
+ "schemars 1.2.2",
+ "serde",
+ "serde_json",
+ "strum",
+ "thiserror 2.0.19",
+ "time",
+]
+
+[[package]]
+name = "litellm-traces-cache"
+version = "0.1.0"
+dependencies = [
+ "litellm-traces",
+ "moka",
+ "rstest",
+ "serde_json",
+ "sha2 0.10.9",
+ "thiserror 2.0.19",
+ "tokio",
+]
+
+[[package]]
+name = "litellm-traces-clickhouse"
+version = "0.1.0"
+dependencies = [
+ "askama",
+ "base64 0.22.1",
+ "flate2",
+ "futures-util",
+ "hmac 0.12.1",
+ "itertools 0.14.0",
+ "jsonschema",
+ "litellm-http",
+ "litellm-migrate",
+ "litellm-storage-clickhouse",
+ "litellm-traces",
+ "litellm-traces-cache",
+ "macro_rules_attribute",
+ "moka",
+ "rstest",
+ "schemars 1.2.2",
"serde",
"serde_json",
"sha2 0.10.9",
+ "strum",
"testcontainers-modules",
"thiserror 2.0.19",
"time",
"tokio",
+ "tracing",
"url",
+ "wiremock",
]
[[package]]
@@ -4804,6 +4936,7 @@ dependencies = [
"js-sys",
"pin-project-lite",
"thiserror 2.0.19",
+ "tracing",
]
[[package]]
@@ -4818,6 +4951,8 @@ dependencies = [
"opentelemetry_sdk 0.33.0",
"prost",
"serde",
+ "tonic",
+ "tonic-prost",
]
[[package]]
diff --git a/litellm-rust/Cargo.toml b/litellm-rust/Cargo.toml
index 257a47268e4..b0766f11e87 100644
--- a/litellm-rust/Cargo.toml
+++ b/litellm-rust/Cargo.toml
@@ -13,6 +13,11 @@ litellm-config = { path = "crates/config" }
litellm-router = { path = "crates/router" }
litellm-tracing = { path = "crates/tracing" }
litellm-traces = { path = "crates/traces" }
+litellm-traces-cache = { path = "crates/traces-cache" }
+litellm-traces-clickhouse = { path = "crates/traces-clickhouse" }
+litellm-storage-clickhouse = { path = "crates/storage-clickhouse" }
+litellm-migrate = { path = "crates/migrate" }
+litellm-migrate-macros = { path = "crates/migrate-macros" }
litellm-core = { path = "crates/core" }
litellm-gateway-mcp = { path = "crates/gateway-mcp" }
litellm-gateway = { path = "crates/gateway" }
@@ -62,6 +67,7 @@ litellm-token-counter-tiktoken = { path = "crates/token-counter-tiktoken" }
litellm-host-python = { path = "crates/host-python" }
litellm-python-compat = { path = "crates/python-compat" }
+askama = { version = "0.16.1", default-features = false, features = ["derive", "std"] }
tracing = "0.1"
axum = { version = "0.8.9", default-features = false, features = ["http1", "tokio", "multipart"] }
axum-login = "0.18.0"
@@ -81,6 +87,7 @@ reqwest = { version = "0.12", default-features = false, features = ["json", "mul
qdrant-client = { version = "1.19.0", default-features = false }
uuid = { version = "1", features = ["v4"] }
rstest = "0.26.1"
+wiremock = "0.6.5"
rstest_reuse = "0.7.0"
rustls = { version = "0.23", default-features = false, features = ["ring", "std", "tls12"] }
rustify = "=0.7.0"
@@ -91,7 +98,10 @@ serde = { version = "1.0", features = ["derive"] }
serde_json = { version = "1.0", features = ["float_roundtrip"] }
serde_with = { version = "=3.16.1", default-features = false, features = ["std", "macros"] }
sha2 = "0.10"
+syn = { version = "2", default-features = false }
sqlx = { version = "0.9.0", default-features = false, features = ["json", "macros", "postgres", "runtime-tokio", "chrono", "tls-rustls-ring-native-roots"] }
+proc-macro2 = "1"
+quote = "1"
subtle = "2"
thiserror = "2.0"
tokenizers = { version = "0.23.1", default-features = false, features = ["onig"] }
@@ -115,6 +125,8 @@ time = { version = "0.3.53", features = ["parsing"] }
criterion = "0.8.2"
fancy-regex = "0.19.2"
veil = "0.3.0"
+prost = "0.14.4"
+opentelemetry-proto = "0.33"
[profile.release]
opt-level = 3
diff --git a/litellm-rust/crates/cache-azure-blob/Cargo.toml b/litellm-rust/crates/cache-azure-blob/Cargo.toml
index 5bdfa16ef53..c28cb90d84a 100644
--- a/litellm-rust/crates/cache-azure-blob/Cargo.toml
+++ b/litellm-rust/crates/cache-azure-blob/Cargo.toml
@@ -26,4 +26,4 @@ litellm-cache-testing.workspace = true
rstest.workspace = true
serde_json.workspace = true
tokio = { workspace = true, features = ["macros", "rt-multi-thread"] }
-wiremock = "0.6.5"
+wiremock.workspace = true
diff --git a/litellm-rust/crates/cache-gcs/Cargo.toml b/litellm-rust/crates/cache-gcs/Cargo.toml
index 1a06683e615..91630879cbe 100644
--- a/litellm-rust/crates/cache-gcs/Cargo.toml
+++ b/litellm-rust/crates/cache-gcs/Cargo.toml
@@ -21,4 +21,4 @@ litellm-cache-testing.workspace = true
rstest.workspace = true
serde_json.workspace = true
tokio.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
diff --git a/litellm-rust/crates/cache-response/Cargo.toml b/litellm-rust/crates/cache-response/Cargo.toml
index 1379573e505..869c40a12ab 100644
--- a/litellm-rust/crates/cache-response/Cargo.toml
+++ b/litellm-rust/crates/cache-response/Cargo.toml
@@ -21,4 +21,4 @@ redis = "1.7.0"
redis-test = "1.0.4"
rstest.workspace = true
tokio.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
diff --git a/litellm-rust/crates/cache-s3/Cargo.toml b/litellm-rust/crates/cache-s3/Cargo.toml
index 680f2da8215..eb3a2fff1ac 100644
--- a/litellm-rust/crates/cache-s3/Cargo.toml
+++ b/litellm-rust/crates/cache-s3/Cargo.toml
@@ -23,6 +23,6 @@ tokio.workspace = true
litellm-http = { workspace = true, features = ["test-support"] }
litellm-cache-testing.workspace = true
rstest.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
serde_json.workspace = true
tokio = { workspace = true, features = ["macros", "rt-multi-thread"] }
diff --git a/litellm-rust/crates/config/src/lib.rs b/litellm-rust/crates/config/src/lib.rs
index e7010941c3d..fed60ab1a4f 100644
--- a/litellm-rust/crates/config/src/lib.rs
+++ b/litellm-rust/crates/config/src/lib.rs
@@ -12,7 +12,10 @@ use serde::Deserialize;
pub use error::Error;
pub use mcp::{McpAuth, McpServer, McpTransport};
pub use model::{LiteLlmParams, Model};
-pub use settings::{GeneralSettings, LiteLlmSettings, RouterSettings};
+pub use settings::{
+ ClickHouseStoreSettings, GeneralSettings, LiteLlmSettings, RouterSettings, TracingSettings,
+ TracingStoreSettings,
+};
pub use value::{AdditionalFields, Flag, NumberOrString, Object, OneOrMany, Value};
#[derive(Clone, Default, Deserialize)]
diff --git a/litellm-rust/crates/config/src/settings.rs b/litellm-rust/crates/config/src/settings.rs
index b6b97475eda..b1ead35e55b 100644
--- a/litellm-rust/crates/config/src/settings.rs
+++ b/litellm-rust/crates/config/src/settings.rs
@@ -5,6 +5,47 @@ use serde::Deserialize;
use crate::{AdditionalFields, Flag, NumberOrString, Object, OneOrMany, Value};
+#[derive(Clone, Debug, Deserialize)]
+#[serde(rename_all = "lowercase")]
+pub enum TracingStoreKind {
+ Clickhouse,
+}
+
+#[derive(Clone, Deserialize)]
+#[serde(deny_unknown_fields)]
+pub struct ClickHouseStoreSettings {
+ #[serde(rename = "type")]
+ pub kind: TracingStoreKind,
+ pub url: Option,
+ pub database: Option,
+ pub retention_days: Option,
+}
+
+impl fmt::Debug for ClickHouseStoreSettings {
+ fn fmt(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result {
+ formatter
+ .debug_struct("ClickHouseStoreSettings")
+ .field("kind", &self.kind)
+ .field("database", &self.database)
+ .field("retention_days", &self.retention_days)
+ .finish()
+ }
+}
+
+#[derive(Clone, Debug, Deserialize)]
+#[serde(untagged)]
+pub enum TracingStoreSettings {
+ ClickHouse(ClickHouseStoreSettings),
+}
+
+#[derive(Clone, Default, Debug, Deserialize)]
+#[serde(default)]
+pub struct TracingSettings {
+ pub store: Option,
+ #[serde(flatten)]
+ pub additional_fields: AdditionalFields,
+}
+
#[derive(Clone, Deserialize)]
#[serde(default)]
pub struct GeneralSettings {
@@ -14,6 +55,7 @@ pub struct GeneralSettings {
pub admission_queue_timeout_seconds: f64,
pub master_key: Option,
pub database_url: Option,
+ pub tracing: Option,
pub database_connection_pool_limit: Option,
pub database_connection_timeout: Option,
pub database_connect_timeout: Option,
@@ -50,6 +92,7 @@ impl Default for GeneralSettings {
admission_queue_timeout_seconds: 1.0,
master_key: None,
database_url: None,
+ tracing: None,
database_connection_pool_limit: Some(10),
database_connection_timeout: Some(60.0),
database_connect_timeout: None,
@@ -97,6 +140,7 @@ impl fmt::Debug for GeneralSettings {
)
.field("master_key", &self.master_key)
.field("database_url", &self.database_url)
+ .field("tracing", &self.tracing)
.field("store_model_in_db", &self.store_model_in_db)
.field("additional_fields", &self.additional_fields.keys())
.finish_non_exhaustive()
diff --git a/litellm-rust/crates/config/tests/config.rs b/litellm-rust/crates/config/tests/config.rs
index 447aa9e1d2a..ab9f3403a01 100644
--- a/litellm-rust/crates/config/tests/config.rs
+++ b/litellm-rust/crates/config/tests/config.rs
@@ -1,4 +1,4 @@
-use litellm_config::{Config, Error, Flag, NumberOrString};
+use litellm_config::{Config, Error, Flag, NumberOrString, TracingStoreSettings};
use rstest::{fixture, rstest};
use tempfile::TempDir;
@@ -113,6 +113,54 @@ fn missing_general_settings_has_no_master_key() {
assert!(config.general_settings.master_key.is_none());
}
+#[test]
+fn tracing_settings_are_typed_and_redact_the_url() {
+ let config = Config::from_yaml(
+ "general_settings:\n tracing:\n store:\n type: clickhouse\n url: https://writer:password@example.com\n database: analytics\n retention_days: 7\n",
+ )
+ .unwrap();
+ let tracing = config.general_settings.tracing.as_ref().unwrap();
+ let Some(TracingStoreSettings::ClickHouse(store)) = tracing.store.as_ref() else {
+ panic!("expected ClickHouse tracing store")
+ };
+ assert_eq!(
+ store.url.as_ref().unwrap().expose(),
+ "https://writer:password@example.com"
+ );
+ assert_eq!(store.database.as_deref(), Some("analytics"));
+ assert_eq!(store.retention_days, Some(NumberOrString::Number(7.0)));
+ assert!(!format!("{config:?}").contains("password"));
+}
+
+#[test]
+fn tracing_settings_accept_environment_references() {
+ let config = Config::from_yaml(
+ "general_settings:\n tracing:\n store:\n type: clickhouse\n url: os.environ/CLICKHOUSE_URL\n retention_days: os.environ/RETENTION_DAYS\n",
+ )
+ .unwrap();
+ let Some(TracingStoreSettings::ClickHouse(store)) =
+ config.general_settings.tracing.unwrap().store
+ else {
+ panic!("expected ClickHouse tracing store")
+ };
+ assert_eq!(
+ store.retention_days,
+ Some(NumberOrString::String(
+ "os.environ/RETENTION_DAYS".to_owned()
+ ))
+ );
+}
+
+#[test]
+fn tracing_settings_reject_string_store() {
+ assert!(Config::from_yaml("general_settings:\n tracing:\n store: clickhouse\n").is_err());
+}
+
+#[test]
+fn tracing_settings_reject_removed_reader_configuration() {
+ assert!(Config::from_yaml("general_settings:\n tracing:\n store:\n type: clickhouse\n reader_url: http://localhost:8123\n").is_err());
+}
+
#[rstest]
fn empty_config_matches_python_defaults() {
let config = Config::from_yaml("{}").unwrap();
diff --git a/litellm-rust/crates/core/Cargo.toml b/litellm-rust/crates/core/Cargo.toml
index 8410aff1d6a..85362fd90d2 100644
--- a/litellm-rust/crates/core/Cargo.toml
+++ b/litellm-rust/crates/core/Cargo.toml
@@ -47,4 +47,4 @@ litellm-host-native.workspace = true
litellm-llms = { workspace = true, features = ["test-support"] }
rstest.workspace = true
rstest_reuse.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
diff --git a/litellm-rust/crates/gateway-inference/Cargo.toml b/litellm-rust/crates/gateway-inference/Cargo.toml
index c854f0ea1ad..4544152d059 100644
--- a/litellm-rust/crates/gateway-inference/Cargo.toml
+++ b/litellm-rust/crates/gateway-inference/Cargo.toml
@@ -30,4 +30,4 @@ futures-util.workspace = true
tokio = { workspace = true, features = ["io-util"] }
rstest.workspace = true
tower = { version = "0.5.3", features = ["util"] }
-wiremock = "0.6.5"
+wiremock.workspace = true
diff --git a/litellm-rust/crates/host-python/src/conversion_cache.rs b/litellm-rust/crates/host-python/src/conversion_cache.rs
new file mode 100644
index 00000000000..c78ed42bcee
--- /dev/null
+++ b/litellm-rust/crates/host-python/src/conversion_cache.rs
@@ -0,0 +1,57 @@
+use std::collections::{HashMap, hash_map::Entry};
+
+use pyo3::prelude::*;
+
+pub struct ToPythonCache<'a, 'py, T> {
+ entries: HashMap)>,
+}
+
+impl Default for ToPythonCache<'_, '_, T> {
+ fn default() -> Self {
+ Self {
+ entries: HashMap::new(),
+ }
+ }
+}
+
+impl<'a, 'py, T> ToPythonCache<'a, 'py, T> {
+ pub fn get_or_try_insert_with(
+ &mut self,
+ value: &'a T,
+ convert: impl FnOnce(&'a T) -> PyResult>,
+ ) -> PyResult<&Bound<'py, PyAny>> {
+ let identity = std::ptr::from_ref(value) as usize;
+ let entry = match self.entries.entry(identity) {
+ Entry::Occupied(entry) => entry.into_mut(),
+ Entry::Vacant(entry) => entry.insert((value, convert(value)?)),
+ };
+ Ok(&entry.1)
+ }
+}
+
+pub struct FromPythonCache<'py, T> {
+ entries: HashMap, T)>,
+}
+
+impl Default for FromPythonCache<'_, T> {
+ fn default() -> Self {
+ Self {
+ entries: HashMap::new(),
+ }
+ }
+}
+
+impl<'py, T> FromPythonCache<'py, T> {
+ pub fn get_or_try_insert_with(
+ &mut self,
+ value: &Bound<'py, PyAny>,
+ convert: impl FnOnce(&Bound<'py, PyAny>) -> PyResult,
+ ) -> PyResult<&T> {
+ let identity = value.as_ptr() as usize;
+ let entry = match self.entries.entry(identity) {
+ Entry::Occupied(entry) => entry.into_mut(),
+ Entry::Vacant(entry) => entry.insert((value.clone(), convert(value)?)),
+ };
+ Ok(&entry.1)
+ }
+}
diff --git a/litellm-rust/crates/host-python/src/lib.rs b/litellm-rust/crates/host-python/src/lib.rs
index 4de404e3624..00543f64085 100644
--- a/litellm-rust/crates/host-python/src/lib.rs
+++ b/litellm-rust/crates/host-python/src/lib.rs
@@ -5,6 +5,7 @@
mod argument;
mod binding;
+mod conversion_cache;
mod driver;
mod error;
mod file_reader;
@@ -20,6 +21,7 @@ mod services;
pub use argument::lookup;
pub use binding::PythonBinding;
+pub use conversion_cache::{FromPythonCache, ToPythonCache};
pub use driver::{CallOptions, run_call};
pub use error::{InvokeError, missing_state};
pub use file_reader::{FileContent, PythonFileReader, py_bytes};
diff --git a/litellm-rust/crates/host-python/tests/conversion_cache.rs b/litellm-rust/crates/host-python/tests/conversion_cache.rs
new file mode 100644
index 00000000000..70ad838e001
--- /dev/null
+++ b/litellm-rust/crates/host-python/tests/conversion_cache.rs
@@ -0,0 +1,121 @@
+use std::{cell::Cell, rc::Rc};
+
+use litellm_host_python::{FromPythonCache, Pythonized, ToPythonCache};
+use pyo3::{exceptions::PyValueError, prelude::*, types::PyDict};
+use rstest::{fixture, rstest};
+
+#[fixture]
+fn python() {
+ Python::initialize();
+}
+
+#[rstest]
+fn rust_identity_reuses_python_objects_without_merging_equal_values(#[from(python)] _python: ()) {
+ Python::attach(|py| {
+ let original = Rc::new(vec![1, 2]);
+ let cloned = original.clone();
+ let equal = Rc::new(vec![1, 2]);
+ let mut cache = ToPythonCache::default();
+ let first = cache
+ .get_or_try_insert_with(original.as_ref(), |value| {
+ Pythonized(value).into_pyobject(py)
+ })
+ .unwrap()
+ .clone();
+ let second = cache
+ .get_or_try_insert_with(cloned.as_ref(), |_| panic!("must reuse conversion"))
+ .unwrap()
+ .clone();
+ let third = cache
+ .get_or_try_insert_with(equal.as_ref(), |value| Pythonized(value).into_pyobject(py))
+ .unwrap();
+ assert!(first.is(&second));
+ assert!(!first.is(third));
+ assert!(first.eq(third).unwrap());
+ });
+}
+
+#[rstest]
+fn python_identity_reuses_rust_values_without_merging_equal_objects(#[from(python)] _python: ()) {
+ Python::attach(|py| {
+ let original = PyDict::new(py);
+ original.set_item("value", 1).unwrap();
+ let equal = original.copy().unwrap();
+ let calls = Cell::new(0);
+ let mut cache = FromPythonCache::default();
+ let convert = |value: &Bound<'_, PyAny>| {
+ calls.set(calls.get() + 1);
+ value.get_item("value")?.extract::().map(Rc::new)
+ };
+ let first = cache
+ .get_or_try_insert_with(original.as_any(), convert)
+ .unwrap()
+ .clone();
+ let second = cache
+ .get_or_try_insert_with(original.as_any(), convert)
+ .unwrap()
+ .clone();
+ let third = cache
+ .get_or_try_insert_with(equal.as_any(), convert)
+ .unwrap();
+ assert!(Rc::ptr_eq(&first, &second));
+ assert!(!Rc::ptr_eq(&first, third));
+ assert_eq!(&first, third);
+ assert_eq!(calls.get(), 2);
+ });
+}
+
+#[rstest]
+fn python_sources_stay_alive_until_the_cache_is_dropped(#[from(python)] _python: ()) {
+ Python::attach(|py| {
+ let value = py
+ .eval(pyo3::ffi::c_str!("type('Tracked', (), {})()"), None, None)
+ .unwrap();
+ let weak = py
+ .import("weakref")
+ .unwrap()
+ .call_method1("ref", (&value,))
+ .unwrap();
+ let mut cache = FromPythonCache::default();
+ cache.get_or_try_insert_with(&value, |_| Ok(42)).unwrap();
+ drop(value);
+ assert!(!weak.call0().unwrap().is_none());
+ drop(cache);
+ assert!(weak.call0().unwrap().is_none());
+ });
+}
+
+#[rstest]
+#[case::to_python(true)]
+#[case::from_python(false)]
+fn failed_conversions_preserve_exceptions_and_can_be_retried(
+ #[from(python)] _python: (),
+ #[case] to_python: bool,
+) {
+ Python::attach(|py| {
+ let failure = PyValueError::new_err("conversion failed");
+ if to_python {
+ let source = vec![1, 2];
+ let mut cache = ToPythonCache::default();
+ let error = cache
+ .get_or_try_insert_with(&source, |_| Err(failure.clone_ref(py)))
+ .unwrap_err();
+ assert!(error.value(py).is(failure.value(py)));
+ let result = cache
+ .get_or_try_insert_with(&source, |value| Pythonized(value).into_pyobject(py))
+ .unwrap();
+ assert_eq!(result.extract::>().unwrap(), source);
+ } else {
+ let source = PyDict::new(py).into_any();
+ let mut cache = FromPythonCache::default();
+ let error = cache
+ .get_or_try_insert_with(&source, |_| Err(failure.clone_ref(py)))
+ .unwrap_err();
+ assert!(error.value(py).is(failure.value(py)));
+ assert_eq!(
+ *cache.get_or_try_insert_with(&source, |_| Ok(42)).unwrap(),
+ 42
+ );
+ }
+ });
+}
diff --git a/litellm-rust/crates/migrate-macros/Cargo.toml b/litellm-rust/crates/migrate-macros/Cargo.toml
new file mode 100644
index 00000000000..5cd68415ca2
--- /dev/null
+++ b/litellm-rust/crates/migrate-macros/Cargo.toml
@@ -0,0 +1,19 @@
+[package]
+name = "litellm-migrate-macros"
+version = "0.1.0"
+edition.workspace = true
+license.workspace = true
+repository.workspace = true
+
+[lib]
+proc-macro = true
+
+[dependencies]
+proc-macro2.workspace = true
+quote.workspace = true
+syn = { workspace = true, features = ["parsing", "printing", "proc-macro"] }
+thiserror.workspace = true
+
+[dev-dependencies]
+rstest.workspace = true
+tempfile.workspace = true
diff --git a/litellm-rust/crates/migrate-macros/src/error.rs b/litellm-rust/crates/migrate-macros/src/error.rs
new file mode 100644
index 00000000000..9833009517b
--- /dev/null
+++ b/litellm-rust/crates/migrate-macros/src/error.rs
@@ -0,0 +1,21 @@
+use std::io;
+
+#[derive(Debug, thiserror::Error)]
+pub enum Error {
+ #[error("could not read migrations directory `{path}`")]
+ ReadDirectory {
+ path: String,
+ #[source]
+ source: io::Error,
+ },
+ #[error(
+ "migration name `{name}` must be `_.sql` with a `[a-z0-9_]` description"
+ )]
+ InvalidName { name: String },
+ #[error("migration version `{version}` is declared more than once")]
+ DuplicateVersion { version: u64 },
+ #[error("migrations directory `{path}` contains no migrations")]
+ Empty { path: String },
+ #[error("migration path `{path}` is not valid UTF-8")]
+ NonUtf8Path { path: String },
+}
diff --git a/litellm-rust/crates/migrate-macros/src/lib.rs b/litellm-rust/crates/migrate-macros/src/lib.rs
new file mode 100644
index 00000000000..501f59e6fc2
--- /dev/null
+++ b/litellm-rust/crates/migrate-macros/src/lib.rs
@@ -0,0 +1,199 @@
+mod error;
+
+use std::path::{Path, PathBuf};
+
+use error::Error;
+use proc_macro::TokenStream;
+use quote::quote;
+use syn::LitStr;
+
+struct Entry {
+ version: u64,
+ description: String,
+ path: PathBuf,
+}
+
+fn resolve(dir: &Path) -> Result, Error> {
+ let mut entries = Vec::new();
+ let files = std::fs::read_dir(dir).map_err(|source| Error::ReadDirectory {
+ path: dir.display().to_string(),
+ source,
+ })?;
+ for file in files {
+ let file = file.map_err(|source| Error::ReadDirectory {
+ path: dir.display().to_string(),
+ source,
+ })?;
+ let path = file.path();
+ let name = path
+ .file_name()
+ .and_then(|name| name.to_str())
+ .ok_or_else(|| Error::NonUtf8Path {
+ path: path.display().to_string(),
+ })?
+ .to_owned();
+ let invalid = || Error::InvalidName { name: name.clone() };
+ let stem = name
+ .strip_suffix(".sql")
+ .filter(|_| file.file_type().is_ok_and(|kind| kind.is_file()))
+ .and_then(|stem| stem.split_once('_'))
+ .filter(|(version, description)| {
+ !version.is_empty()
+ && version.bytes().all(|b| b.is_ascii_digit())
+ && !description.is_empty()
+ && description
+ .bytes()
+ .all(|b| b.is_ascii_lowercase() || b.is_ascii_digit() || b == b'_')
+ })
+ .ok_or_else(invalid)?;
+ let version = stem.0.parse::().map_err(|_| invalid())?;
+ entries.push(Entry {
+ version,
+ description: stem.1.to_owned(),
+ path,
+ });
+ }
+ if entries.is_empty() {
+ return Err(Error::Empty {
+ path: dir.display().to_string(),
+ });
+ }
+ entries.sort_by_key(|entry| entry.version);
+ for pair in entries.windows(2) {
+ if pair[0].version == pair[1].version {
+ return Err(Error::DuplicateVersion {
+ version: pair[0].version,
+ });
+ }
+ }
+ Ok(entries)
+}
+
+fn resolve_input(lit: &LitStr) -> Result, Error> {
+ let root = std::env::var("CARGO_MANIFEST_DIR")
+ .map(PathBuf::from)
+ .unwrap_or_default();
+ let dir = root.join(lit.value());
+ let dir = dir.canonicalize().map_err(|source| Error::ReadDirectory {
+ path: dir.display().to_string(),
+ source,
+ })?;
+ if dir.to_str().is_none() {
+ return Err(Error::NonUtf8Path {
+ path: dir.display().to_string(),
+ });
+ }
+ resolve(&dir)
+}
+
+#[proc_macro]
+pub fn migrate(input: TokenStream) -> TokenStream {
+ let lit = syn::parse_macro_input!(input as LitStr);
+ match resolve_input(&lit) {
+ Ok(entries) => {
+ let migrations = entries.iter().map(|entry| {
+ let version = entry.version;
+ let description = &entry.description;
+ let path = entry
+ .path
+ .to_str()
+ .expect("canonical migration path is UTF-8");
+ quote! {
+ ::litellm_migrate::Migration {
+ version: #version,
+ description: #description,
+ sql: ::core::include_str!(#path),
+ }
+ }
+ });
+ quote! { &[#(#migrations),*] }.into()
+ }
+ Err(err) => syn::Error::new(lit.span(), err).to_compile_error().into(),
+ }
+}
+
+#[cfg(test)]
+mod tests {
+ use std::fs;
+
+ use rstest::rstest;
+ use tempfile::TempDir;
+
+ use super::{Error, resolve};
+
+ fn migrations_dir(files: &[&str]) -> TempDir {
+ let dir = TempDir::new().expect("tempdir");
+ for file in files {
+ fs::write(dir.path().join(file), "SELECT 1").expect("write fixture");
+ }
+ dir
+ }
+
+ #[rstest]
+ fn orders_versions_numerically() {
+ let dir = migrations_dir(&["10_tenth.sql", "2_second.sql", "1_first.sql"]);
+ let entries = resolve(dir.path()).expect("resolves");
+ let versions: Vec = entries.iter().map(|entry| entry.version).collect();
+ let descriptions: Vec<&str> = entries
+ .iter()
+ .map(|entry| entry.description.as_str())
+ .collect();
+ assert_eq!(versions, [1, 2, 10]);
+ assert_eq!(descriptions, ["first", "second", "tenth"]);
+ }
+
+ #[rstest]
+ #[case::dash_in_version(&["0001-dash.sql"])]
+ #[case::not_sql(&["notes.txt"])]
+ #[case::empty_description(&["0001_.sql"])]
+ #[case::non_digit_version(&["x_name.sql"])]
+ #[case::uppercase_description(&["0001_Upper.sql"])]
+ #[case::no_underscore(&["0001.sql"])]
+ #[case::plus_sign_version(&["+10_add.sql"])]
+ fn rejects_invalid_names(#[case] files: &[&str]) {
+ let dir = migrations_dir(files);
+ assert!(matches!(
+ resolve(dir.path()),
+ Err(Error::InvalidName { .. })
+ ));
+ }
+
+ #[rstest]
+ fn rejects_subdirectories() {
+ let dir = migrations_dir(&["0001_a.sql"]);
+ fs::create_dir(dir.path().join("0002_b.sql")).expect("subdir");
+ assert!(matches!(
+ resolve(dir.path()),
+ Err(Error::InvalidName { .. })
+ ));
+ }
+
+ #[cfg(unix)]
+ #[rstest]
+ fn rejects_symlinks() {
+ let dir = migrations_dir(&["0001_a.sql"]);
+ let target = TempDir::new().expect("tempdir");
+ let target_file = target.path().join("real.sql");
+ fs::write(&target_file, "SELECT 2").expect("write fixture");
+ std::os::unix::fs::symlink(&target_file, dir.path().join("0002_b.sql")).expect("symlink");
+ assert!(matches!(
+ resolve(dir.path()),
+ Err(Error::InvalidName { .. })
+ ));
+ }
+
+ #[rstest]
+ fn rejects_duplicate_versions() {
+ let dir = migrations_dir(&["0001_a.sql", "1_b.sql"]);
+ assert!(matches!(
+ resolve(dir.path()),
+ Err(Error::DuplicateVersion { version: 1 })
+ ));
+ }
+
+ #[rstest]
+ fn rejects_empty_directory() {
+ let dir = migrations_dir(&[]);
+ assert!(matches!(resolve(dir.path()), Err(Error::Empty { .. })));
+ }
+}
diff --git a/litellm-rust/crates/migrate/Cargo.toml b/litellm-rust/crates/migrate/Cargo.toml
new file mode 100644
index 00000000000..bb1ecaa3128
--- /dev/null
+++ b/litellm-rust/crates/migrate/Cargo.toml
@@ -0,0 +1,12 @@
+[package]
+name = "litellm-migrate"
+version = "0.1.0"
+edition.workspace = true
+license.workspace = true
+repository.workspace = true
+
+[dependencies]
+litellm-migrate-macros.workspace = true
+
+[dev-dependencies]
+rstest.workspace = true
diff --git a/litellm-rust/crates/migrate/README.md b/litellm-rust/crates/migrate/README.md
new file mode 100644
index 00000000000..4817029451c
--- /dev/null
+++ b/litellm-rust/crates/migrate/README.md
@@ -0,0 +1,5 @@
+# Migrations
+
+`litellm-migrate` exports the `Migration` struct and the `migrate!` macro that embeds a directory of `_.sql` files at compile time, sorted by numeric version
+
+The crate does not apply or track migrations; callers decide how and when the embedded SQL runs
diff --git a/litellm-rust/crates/migrate/src/lib.rs b/litellm-rust/crates/migrate/src/lib.rs
new file mode 100644
index 00000000000..f4e065e1b53
--- /dev/null
+++ b/litellm-rust/crates/migrate/src/lib.rs
@@ -0,0 +1,8 @@
+pub use litellm_migrate_macros::migrate;
+
+#[derive(Debug, Clone, Copy, PartialEq, Eq)]
+pub struct Migration {
+ pub version: u64,
+ pub description: &'static str,
+ pub sql: &'static str,
+}
diff --git a/litellm-rust/crates/migrate/tests/fixtures/migrations/10_tenth.sql b/litellm-rust/crates/migrate/tests/fixtures/migrations/10_tenth.sql
new file mode 100644
index 00000000000..31807719e9c
--- /dev/null
+++ b/litellm-rust/crates/migrate/tests/fixtures/migrations/10_tenth.sql
@@ -0,0 +1 @@
+SELECT 10;
diff --git a/litellm-rust/crates/migrate/tests/fixtures/migrations/1_first.sql b/litellm-rust/crates/migrate/tests/fixtures/migrations/1_first.sql
new file mode 100644
index 00000000000..e0ac49d1ecf
--- /dev/null
+++ b/litellm-rust/crates/migrate/tests/fixtures/migrations/1_first.sql
@@ -0,0 +1 @@
+SELECT 1;
diff --git a/litellm-rust/crates/migrate/tests/fixtures/migrations/2_second.sql b/litellm-rust/crates/migrate/tests/fixtures/migrations/2_second.sql
new file mode 100644
index 00000000000..e7f8100648d
--- /dev/null
+++ b/litellm-rust/crates/migrate/tests/fixtures/migrations/2_second.sql
@@ -0,0 +1 @@
+SELECT 2;
diff --git a/litellm-rust/crates/migrate/tests/migrate.rs b/litellm-rust/crates/migrate/tests/migrate.rs
new file mode 100644
index 00000000000..61c80351cf4
--- /dev/null
+++ b/litellm-rust/crates/migrate/tests/migrate.rs
@@ -0,0 +1,21 @@
+use litellm_migrate::Migration;
+use rstest::rstest;
+
+const MIGRATIONS: &[Migration] = litellm_migrate::migrate!("tests/fixtures/migrations");
+
+#[rstest]
+#[case::first(0, 1, "first", include_str!("fixtures/migrations/1_first.sql"))]
+#[case::second(1, 2, "second", include_str!("fixtures/migrations/2_second.sql"))]
+#[case::tenth(2, 10, "tenth", include_str!("fixtures/migrations/10_tenth.sql"))]
+fn embeds_every_file_sorted_by_numeric_version(
+ #[case] index: usize,
+ #[case] version: u64,
+ #[case] description: &str,
+ #[case] sql: &str,
+) {
+ assert_eq!(MIGRATIONS.len(), 3);
+ let migration = &MIGRATIONS[index];
+ assert_eq!(migration.version, version);
+ assert_eq!(migration.description, description);
+ assert_eq!(migration.sql, sql);
+}
diff --git a/litellm-rust/crates/model-catalog/src/model_info.rs b/litellm-rust/crates/model-catalog/src/model_info.rs
index 380f6713d7a..96dec84de84 100644
--- a/litellm-rust/crates/model-catalog/src/model_info.rs
+++ b/litellm-rust/crates/model-catalog/src/model_info.rs
@@ -467,6 +467,10 @@ pub struct ModelInfo {
#[serde(skip_serializing_if = "Option::is_none")]
pub supports_audio_output: Option,
#[serde(skip_serializing_if = "Option::is_none")]
+ pub supports_bedrock_runtime_chat_completions_response_format: Option,
+ #[serde(skip_serializing_if = "Option::is_none")]
+ pub supports_bedrock_runtime_chat_completions_tools_with_reasoning: Option,
+ #[serde(skip_serializing_if = "Option::is_none")]
pub supports_computer_use: Option,
#[serde(skip_serializing_if = "Option::is_none")]
pub supports_embedding_image_input: Option,
diff --git a/litellm-rust/crates/python-bridge/Cargo.toml b/litellm-rust/crates/python-bridge/Cargo.toml
index 99c95632bb3..c3a86009111 100644
--- a/litellm-rust/crates/python-bridge/Cargo.toml
+++ b/litellm-rust/crates/python-bridge/Cargo.toml
@@ -22,6 +22,8 @@ tiktoken = ["litellm-token-counter/tiktoken"]
fancy-regex.workspace = true
litellm-tracing.workspace = true
litellm-traces.workspace = true
+litellm-traces-clickhouse.workspace = true
+litellm-storage-clickhouse.workspace = true
litellm-host.workspace = true
bytes.workspace = true
futures-util.workspace = true
@@ -51,6 +53,7 @@ litellm-llms-types.workspace = true
litellm-host-python.workspace = true
litellm-token-counter = { path = "../token-counter", default-features = false }
pyo3.workspace = true
+prost.workspace = true
pyo3-async-runtimes.workspace = true
reqwest.workspace = true
redis = { version = "1.7.0", features = ["tls-rustls"] }
@@ -72,7 +75,7 @@ futures-util.workspace = true
rstest.workspace = true
sha2.workspace = true
tokio-tungstenite.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
aws-sdk-secretsmanager = "1.117.0"
[[bench]]
diff --git a/litellm-rust/crates/python-bridge/src/lib.rs b/litellm-rust/crates/python-bridge/src/lib.rs
index 0d4df996552..6659be5160e 100644
--- a/litellm-rust/crates/python-bridge/src/lib.rs
+++ b/litellm-rust/crates/python-bridge/src/lib.rs
@@ -44,7 +44,9 @@ mod _native {
#[pymodule_export]
use crate::routes::token_counter::TokenCounter;
#[pymodule_export]
- use crate::routes::traces::{NativeTraceStorage, trace_decode_otlp};
+ use crate::routes::traces::{
+ NativeTraceConfig, NativeTraceStorage, trace_encode_error, trace_span_rows,
+ };
#[cfg(feature = "huggingface")]
#[pymodule_export]
use crate::tokenizer::HuggingFaceEncoding;
@@ -109,8 +111,10 @@ mod tests {
"aresponses",
"ResponsesWebSocketConnection",
"NativeDiagnosticProcessor",
+ "NativeTraceConfig",
"NativeTraceStorage",
- "trace_decode_otlp",
+ "trace_encode_error",
+ "trace_span_rows",
"TokenCounter",
"Tokenizer",
"gil_stats",
diff --git a/litellm-rust/crates/python-bridge/src/routes/traces.rs b/litellm-rust/crates/python-bridge/src/routes/traces.rs
index 2e7a6b178a8..dd5c6b9860f 100644
--- a/litellm-rust/crates/python-bridge/src/routes/traces.rs
+++ b/litellm-rust/crates/python-bridge/src/routes/traces.rs
@@ -1,71 +1,145 @@
use std::collections::BTreeMap;
use litellm_http::ClientVariant;
-use litellm_traces::{Connection, Error, InsertTable, Parameter, ReadQuery};
+use litellm_traces::{QueryScope, ReadQuery, Tenant, query::named::ReadAccessParams};
+use litellm_traces_clickhouse::{Config, Error, InsertTable, Parameter, QueryReaders};
+use prost::Message;
use pyo3::{
exceptions::{PyOverflowError, PyRuntimeError, PyValueError},
prelude::*,
+ types::PyBytes,
};
+#[derive(Message)]
+struct OtlpErrorStatus {
+ #[prost(int32, tag = "1")]
+ code: i32,
+ #[prost(string, tag = "2")]
+ message: String,
+}
+
+#[pyfunction]
+pub fn trace_encode_error<'py>(py: Python<'py>, message: &str) -> Bound<'py, PyBytes> {
+ let status = OtlpErrorStatus {
+ code: 0,
+ message: message.to_owned(),
+ };
+ PyBytes::new(py, &status.encode_to_vec())
+}
+
fn map_error(error: Error) -> PyErr {
+ map_error_ref(&error)
+}
+
+fn map_error_ref(error: &Error) -> PyErr {
+ use litellm_storage_clickhouse::Error as StorageError;
+
match error {
+ Error::Decode(litellm_traces::Error::TooLarge)
+ | Error::InsertTooLarge
+ | Error::ReadTooLarge => PyOverflowError::new_err(error.to_string()),
Error::InvalidRow
+ | Error::InvalidLimit(_)
| Error::InvalidTable
+ | Error::InvalidCursor(_)
+ | Error::AmbiguousTrace
+ | Error::TraceChanged
+ | Error::Decode(_)
| Error::InvalidSchema
- | Error::EmptySql
- | Error::InvalidQuery => PyValueError::new_err(error.to_string()),
- Error::InsertTooLarge => PyOverflowError::new_err(error.to_string()),
- Error::InvalidUrl
- | Error::QueryFailed(_)
- | Error::InsertFailed(_)
+ | Error::InvalidQuery
+ | Error::InvalidParameters
+ | Error::InvalidScope => PyValueError::new_err(error.to_string()),
+ Error::Task
| Error::SchemaFailed(_)
- | Error::ResponseTooLarge
- | Error::InvalidResponse
- | Error::Transport => PyRuntimeError::new_err(error.to_string()),
+ | Error::SchemaTransport
+ | Error::MissingSecret
+ | Error::Busy
+ | Error::ProvisionFailed(_)
+ | Error::ProvisionTransport
+ | Error::InvalidResponse => PyRuntimeError::new_err(error.to_string()),
+ Error::Cached(source) => map_error_ref(source),
+ Error::Storage(source) => match source {
+ StorageError::InvalidRow
+ | StorageError::InvalidLimit(_)
+ | StorageError::InvalidTable
+ | StorageError::InvalidSchema
+ | StorageError::EmptySql
+ | StorageError::InvalidParameters
+ | StorageError::InvalidQuery => PyValueError::new_err(error.to_string()),
+ StorageError::InsertTooLarge => PyOverflowError::new_err(error.to_string()),
+ StorageError::InvalidUrl
+ | StorageError::QueryFailed(_)
+ | StorageError::InsertFailed(_)
+ | StorageError::SchemaFailed(_)
+ | StorageError::ResponseTooLarge
+ | StorageError::InvalidResponse
+ | StorageError::Transport => PyRuntimeError::new_err(error.to_string()),
+ },
+ }
+}
+
+fn map_sql_error(error: Error) -> PyErr {
+ match error {
+ Error::Storage(litellm_storage_clickhouse::Error::QueryFailed(400 | 404)) => {
+ PyValueError::new_err(error.to_string())
+ }
+ error => map_error(error),
+ }
+}
+
+#[pyclass(frozen)]
+pub struct NativeTraceConfig {
+ inner: Config,
+}
+
+#[pymethods]
+impl NativeTraceConfig {
+ #[new]
+ fn new(
+ database: String,
+ url: &str,
+ retention_days: u32,
+ max_attribute_value_bytes: usize,
+ ) -> PyResult {
+ Ok(Self {
+ inner: Config::new(database, url, retention_days, max_attribute_value_bytes)
+ .map_err(map_error)?,
+ })
}
}
#[pyclass]
pub struct NativeTraceStorage {
- database: String,
- writer: Connection,
- reader: Option,
+ config: Config,
+ query_readers: QueryReaders,
}
#[pymethods]
impl NativeTraceStorage {
#[new]
- #[pyo3(signature = (database, url, reader_url = None))]
- fn new(database: String, url: &str, reader_url: Option<&str>) -> PyResult {
- litellm_traces::schema_statements(&database, 1, 1).map_err(map_error)?;
+ fn new(config: PyRef<'_, NativeTraceConfig>) -> PyResult {
Ok(Self {
- writer: Connection::writer(url).map_err(map_error)?,
- reader: reader_url
- .map(|value| Connection::reader(value, &database))
- .transpose()
- .map_err(map_error)?,
- database,
+ query_readers: QueryReaders::new(
+ config.inner.storage().writer().clone(),
+ config.inner.storage().database().to_owned(),
+ ),
+ config: config.inner.clone(),
})
}
- fn ensure_schema<'py>(
- &self,
- py: Python<'py>,
- trace_retention_days: u32,
- spend_log_retention_days: u32,
- ) -> PyResult> {
+ fn ensure_schema<'py>(&self, py: Python<'py>) -> PyResult> {
let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
- let connection = self.writer.clone();
- let database = self.database.clone();
+ let connection = self.config.storage().writer().clone();
+ let database = self.config.storage().database().to_owned();
+ let retention_days = self.config.retention_days();
crate::execution::run_async(
py,
async move {
- litellm_traces::ensure_schema(
+ litellm_traces_clickhouse::ensure_schema(
&client,
&connection,
&database,
- trace_retention_days,
- spend_log_retention_days,
+ retention_days,
)
.await
},
@@ -83,37 +157,226 @@ impl NativeTraceStorage {
) -> PyResult> {
let table = InsertTable::parse(table).map_err(map_error)?;
let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
- let connection = self.writer.clone();
- let database = self.database.clone();
+ let connection = self.config.storage().writer().clone();
+ let database = self.config.storage().database().to_owned();
crate::execution::run_async(
py,
async move {
- litellm_traces::insert_rows(&client, &connection, &database, table, rows).await
+ litellm_traces_clickhouse::insert_rows(&client, &connection, &database, table, rows)
+ .await
},
map_error,
)
}
- fn lens_query<'py>(
+ fn ingest<'py>(
&self,
py: Python<'py>,
- name: &str,
- #[pyo3(from_py_with = litellm_host_python::from_py_argument)] parameters: BTreeMap<
- String,
- Parameter,
- >,
+ payload: &[u8],
+ content_type: Option,
+ #[pyo3(from_py_with = litellm_host_python::from_py_argument)] tenant: Tenant,
) -> PyResult> {
- let query = litellm_traces::LensQuery::parse(name).map_err(map_error)?;
- let connection = self.reader.clone().ok_or_else(|| {
- PyRuntimeError::new_err("Trace reads require a separate ClickHouse reader URL")
- })?;
+ let payload = payload.to_vec();
+ let max_value_bytes = self.config.max_attribute_value_bytes();
+ let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
+ let connection = self.config.storage().writer().clone();
+ let database = self.config.storage().database().to_owned();
+ crate::execution::run_async(
+ py,
+ async move {
+ let rows = tokio::task::spawn_blocking(move || {
+ litellm_traces::decode_otlp(&payload, content_type.as_deref()).map(|spans| {
+ litellm_traces_clickhouse::span_rows(spans, &tenant, max_value_bytes)
+ })
+ })
+ .await
+ .map_err(|_| Error::Task)??;
+ let count = rows.len();
+ litellm_traces_clickhouse::insert_shared_rows(
+ &client,
+ &connection,
+ &database,
+ InsertTable::OtelTraces,
+ rows,
+ )
+ .await?;
+ Ok(count)
+ },
+ map_error,
+ )
+ }
+
+ #[pyo3(signature = (scope, start_ms, end_ms, cursor, limit))]
+ fn list_traces<'py>(
+ &self,
+ py: Python<'py>,
+ #[pyo3(from_py_with = litellm_host_python::from_py_argument)] scope: ReadAccessParams,
+ start_ms: i64,
+ end_ms: i64,
+ cursor: Option,
+ limit: u32,
+ ) -> PyResult> {
+ let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
+ let connection = self.config.storage().reader().clone();
+ crate::execution::run_async(
+ py,
+ async move {
+ litellm_traces_clickhouse::list_traces(
+ &client,
+ &connection,
+ &scope,
+ start_ms,
+ end_ms,
+ cursor.as_deref(),
+ limit,
+ )
+ .await
+ },
+ map_error,
+ )
+ }
+
+ #[pyo3(signature = (trace_id, scope, trace_ref, cursor=None, page_size=None))]
+ fn get_trace<'py>(
+ &self,
+ py: Python<'py>,
+ trace_id: String,
+ #[pyo3(from_py_with = litellm_host_python::from_py_argument)] scope: ReadAccessParams,
+ trace_ref: String,
+ cursor: Option,
+ page_size: Option,
+ ) -> PyResult> {
+ let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
+ let connection = self.config.storage().reader().clone();
+ crate::execution::run_async(
+ py,
+ async move {
+ if let Some(page_size) = page_size {
+ litellm_traces_clickhouse::get_trace_page(
+ &client,
+ &connection,
+ &scope,
+ &trace_id,
+ &trace_ref,
+ cursor.as_deref(),
+ page_size,
+ )
+ .await
+ } else if cursor.is_some() {
+ Err(Error::InvalidParameters)
+ } else {
+ litellm_traces_clickhouse::get_trace(
+ &client,
+ &connection,
+ &scope,
+ &trace_id,
+ &trace_ref,
+ )
+ .await
+ }
+ },
+ map_error,
+ )
+ }
+
+ fn get_span<'py>(
+ &self,
+ py: Python<'py>,
+ trace_id: String,
+ span_id: String,
+ #[pyo3(from_py_with = litellm_host_python::from_py_argument)] scope: ReadAccessParams,
+ trace_ref: String,
+ ) -> PyResult> {
+ let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
+ let connection = self.config.storage().reader().clone();
+ crate::execution::run_async(
+ py,
+ async move {
+ litellm_traces_clickhouse::get_span(
+ &client,
+ &connection,
+ &scope,
+ &trace_id,
+ &span_id,
+ &trace_ref,
+ )
+ .await
+ },
+ map_error,
+ )
+ }
+
+ #[pyo3(signature = (trace_id, span_id, scope, trace_ref, cursor))]
+ fn get_span_error<'py>(
+ &self,
+ py: Python<'py>,
+ trace_id: String,
+ span_id: String,
+ #[pyo3(from_py_with = litellm_host_python::from_py_argument)] scope: ReadAccessParams,
+ trace_ref: String,
+ cursor: Option,
+ ) -> PyResult> {
+ let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
+ let connection = self.config.storage().reader().clone();
+ crate::execution::run_async(
+ py,
+ async move {
+ litellm_traces_clickhouse::get_span_error(
+ &client,
+ &connection,
+ &scope,
+ &trace_id,
+ &span_id,
+ &trace_ref,
+ cursor.as_deref(),
+ )
+ .await
+ },
+ map_error,
+ )
+ }
+
+ fn query_sql<'py>(
+ &self,
+ py: Python<'py>,
+ sql: String,
+ #[pyo3(from_py_with = litellm_host_python::from_py_argument)] scope: QueryScope,
+ secret: String,
+ ) -> PyResult> {
+ if sql.trim().is_empty() {
+ return Err(map_error(
+ litellm_storage_clickhouse::Error::EmptySql.into(),
+ ));
+ }
+ let readers = self.query_readers.clone();
let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
crate::execution::run_async(
py,
async move {
- litellm_traces::execute_read(&client, &connection, query.sql(), ¶meters).await
+ let _permit = readers.acquire()?;
+ let connection = readers.connection(&client, &scope, &secret).await?;
+ litellm_traces_clickhouse::query_sql(&client, &connection, &sql).await
},
- map_error,
+ map_sql_error,
+ )
+ }
+
+ fn query_help<'py>(
+ &self,
+ py: Python<'py>,
+ #[pyo3(from_py_with = litellm_host_python::from_py_argument)] scope: QueryScope,
+ secret: String,
+ ) -> PyResult> {
+ let readers = self.query_readers.clone();
+ let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
+ crate::execution::run_async(
+ py,
+ async move {
+ let _permit = readers.acquire()?;
+ let connection = readers.connection(&client, &scope, &secret).await?;
+ litellm_traces_clickhouse::query_help(&client, &connection).await
+ },
+ map_sql_error,
)
}
@@ -126,41 +389,126 @@ impl NativeTraceStorage {
Parameter,
>,
) -> PyResult> {
- let query = ReadQuery::parse(query).map_err(map_error)?;
- let connection = self.reader.clone().ok_or_else(|| {
- PyRuntimeError::new_err("Trace reads require a separate ClickHouse reader URL")
- })?;
+ let query =
+ ReadQuery::parse(query).map_err(|error| PyValueError::new_err(error.to_string()))?;
+ let connection = self.config.storage().reader().clone();
let client = crate::http::host_client(py, ClientVariant::NoRedirect)?;
crate::execution::run_async(
py,
async move {
- litellm_traces::execute_named_read(&client, &connection, query, ¶meters).await
+ litellm_traces_clickhouse::execute_named_read(
+ &client,
+ &connection,
+ query,
+ ¶meters,
+ )
+ .await
},
map_error,
)
}
}
+/// The `otel_traces` rows an export would be stored as, without writing them.
#[pyfunction]
-pub fn trace_decode_otlp<'py>(
+pub fn trace_span_rows<'py>(
py: Python<'py>,
body: &[u8],
content_type: Option<&str>,
- content_encoding: Option<&str>,
- max_decompressed_bytes: usize,
+ #[pyo3(from_py_with = litellm_host_python::from_py_argument)] tenant: Tenant,
+ max_attribute_value_bytes: usize,
) -> PyResult> {
- let spans = py
+ let rows = py
.detach(|| {
- litellm_traces::decode_otlp(
- body,
- content_type,
- content_encoding,
- max_decompressed_bytes,
- )
+ litellm_traces::decode_otlp(body, content_type).map(|spans| {
+ litellm_traces_clickhouse::span_rows(spans, &tenant, max_attribute_value_bytes)
+ })
})
- .map_err(|error| match error {
- litellm_traces::DecodeError::TooLarge => PyOverflowError::new_err(error.to_string()),
- _ => PyValueError::new_err(error.to_string()),
- })?;
- litellm_host_python::Pythonized(spans).into_pyobject(py)
+ .map_err(|error| map_error(error.into()))?;
+ litellm_host_python::Pythonized(rows).into_pyobject(py)
+}
+
+#[cfg(test)]
+mod tests {
+ use super::*;
+ use rstest::rstest;
+
+ #[rstest]
+ #[case::row(Error::InvalidRow, "ValueError")]
+ #[case::insert_limit(Error::InvalidLimit("CLICKHOUSE_TRACE_MAX_INSERT_BYTES"), "ValueError")]
+ #[case::insert_timeout(
+ Error::Storage(litellm_storage_clickhouse::Error::InvalidLimit(
+ "CLICKHOUSE_INSERT_TIMEOUT_SECONDS"
+ )),
+ "ValueError"
+ )]
+ #[case::insert_budget(Error::InsertTooLarge, "OverflowError")]
+ #[case::scope(Error::InvalidScope, "ValueError")]
+ #[case::schema(Error::SchemaFailed(503), "RuntimeError")]
+ #[case::reader(Error::MissingSecret, "RuntimeError")]
+ #[case::storage(
+ Error::Storage(litellm_storage_clickhouse::Error::InvalidUrl),
+ "RuntimeError"
+ )]
+ #[case::cached_scope(Error::Cached(std::sync::Arc::new(Error::InvalidScope)), "ValueError")]
+ fn trace_failures_preserve_public_exception_types(
+ #[case] error: Error,
+ #[case] exception_name: &str,
+ ) {
+ Python::initialize();
+ Python::attach(|py| {
+ let message = error.to_string();
+ let exception = map_error(error);
+ assert_eq!(exception.get_type(py).name().unwrap(), exception_name);
+ assert_eq!(
+ exception.value(py).str().unwrap().to_str().unwrap(),
+ message
+ );
+ });
+ }
+
+ #[rstest]
+ #[case::invalid_sql(400, "ValueError")]
+ #[case::missing_table(404, "ValueError")]
+ #[case::unavailable(503, "RuntimeError")]
+ fn wrapped_query_status_preserves_public_exception_type(
+ #[case] status: u16,
+ #[case] exception_name: &str,
+ ) {
+ Python::initialize();
+ Python::attach(|py| {
+ let error = Error::Storage(litellm_storage_clickhouse::Error::QueryFailed(status));
+ let message = error.to_string();
+ let exception = map_sql_error(error);
+ assert_eq!(exception.get_type(py).name().unwrap(), exception_name);
+ assert_eq!(
+ exception.value(py).str().unwrap().to_str().unwrap(),
+ message
+ );
+ });
+ }
+
+ #[rstest]
+ #[case::decode_budget(Error::Decode(litellm_traces::Error::TooLarge), "OverflowError")]
+ #[case::invalid_export(Error::Decode(litellm_traces::Error::InvalidPayload), "ValueError")]
+ #[case::invalid_decode_limit(
+ Error::Decode(litellm_traces::Error::InvalidLimit("OTLP_MAX_SPANS")),
+ "ValueError"
+ )]
+ #[case::cursor(Error::InvalidCursor("trace"), "ValueError")]
+ #[case::ambiguous(Error::AmbiguousTrace, "ValueError")]
+ #[case::changed_snapshot(Error::TraceChanged, "ValueError")]
+ #[case::read_budget(Error::ReadTooLarge, "OverflowError")]
+ fn trace_read_and_ingest_failures_preserve_public_exception_types(
+ #[case] error: Error,
+ #[case] exception_name: &str,
+ ) {
+ Python::initialize();
+ Python::attach(|py| {
+ assert_eq!(
+ map_error(error).get_type(py).name().unwrap(),
+ exception_name
+ );
+ });
+ }
}
diff --git a/litellm-rust/crates/secrets-aws/Cargo.toml b/litellm-rust/crates/secrets-aws/Cargo.toml
index e7a394bd247..5d3bd413484 100644
--- a/litellm-rust/crates/secrets-aws/Cargo.toml
+++ b/litellm-rust/crates/secrets-aws/Cargo.toml
@@ -21,5 +21,5 @@ aws-credential-types = "1.3.0"
base64.workspace = true
rstest.workspace = true
tokio.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
tempfile = "3"
diff --git a/litellm-rust/crates/secrets-azure/Cargo.toml b/litellm-rust/crates/secrets-azure/Cargo.toml
index efdf681e2bc..7ec03fb98da 100644
--- a/litellm-rust/crates/secrets-azure/Cargo.toml
+++ b/litellm-rust/crates/secrets-azure/Cargo.toml
@@ -20,7 +20,7 @@ percent-encoding = "2.3"
[dev-dependencies]
litellm-http = { workspace = true, features = ["test-support"] }
-wiremock = "0.6.5"
+wiremock.workspace = true
rstest.workspace = true
serde_json.workspace = true
sha2.workspace = true
diff --git a/litellm-rust/crates/secrets-cyberark/Cargo.toml b/litellm-rust/crates/secrets-cyberark/Cargo.toml
index 0a91c61ade9..f630d5857d8 100644
--- a/litellm-rust/crates/secrets-cyberark/Cargo.toml
+++ b/litellm-rust/crates/secrets-cyberark/Cargo.toml
@@ -25,6 +25,6 @@ rcgen = "0.14.10"
rstest.workspace = true
tempfile = "3.27.0"
tokio.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
serde.workspace = true
serde_json.workspace = true
diff --git a/litellm-rust/crates/secrets-google/Cargo.toml b/litellm-rust/crates/secrets-google/Cargo.toml
index 208b5ddd03f..3ce14fe7a12 100644
--- a/litellm-rust/crates/secrets-google/Cargo.toml
+++ b/litellm-rust/crates/secrets-google/Cargo.toml
@@ -28,4 +28,4 @@ reqwest.workspace = true
litellm-http = { workspace = true, features = ["test-support"] }
google-cloud-auth.workspace = true
rstest.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
diff --git a/litellm-rust/crates/secrets-hashicorp/Cargo.toml b/litellm-rust/crates/secrets-hashicorp/Cargo.toml
index c049ba127e5..7dd3d3c674f 100644
--- a/litellm-rust/crates/secrets-hashicorp/Cargo.toml
+++ b/litellm-rust/crates/secrets-hashicorp/Cargo.toml
@@ -21,4 +21,4 @@ veil.workspace = true
rstest.workspace = true
tempfile = "3"
tokio.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
diff --git a/litellm-rust/crates/secrets/Cargo.toml b/litellm-rust/crates/secrets/Cargo.toml
index f855a8a64a6..3655ce8bbc2 100644
--- a/litellm-rust/crates/secrets/Cargo.toml
+++ b/litellm-rust/crates/secrets/Cargo.toml
@@ -36,7 +36,7 @@ tokio = { workspace = true, features = ["fs"] }
[dev-dependencies]
litellm-http = { workspace = true, features = ["test-support"] }
rstest.workspace = true
-wiremock = "0.6.5"
+wiremock.workspace = true
tempfile = "3"
aws-sdk-kms = "1.120.0"
google-cloud-kms-v1 = "1.14.0"
diff --git a/litellm-rust/crates/storage-clickhouse/AGENTS.md b/litellm-rust/crates/storage-clickhouse/AGENTS.md
new file mode 100644
index 00000000000..959ffdffb88
--- /dev/null
+++ b/litellm-rust/crates/storage-clickhouse/AGENTS.md
@@ -0,0 +1,5 @@
+# ClickHouse storage
+
+`litellm-storage-clickhouse` exports `Storage`, a writer and bounded reader derived from one ClickHouse URL and database. It also exports bounded HTTP read and insert execution
+
+The crate has no trace tables, OTLP types, or named trace queries. `litellm-traces-clickhouse` supplies those rules and uses this storage for both trace rows and spend rows
diff --git a/litellm-rust/crates/storage-clickhouse/Cargo.toml b/litellm-rust/crates/storage-clickhouse/Cargo.toml
new file mode 100644
index 00000000000..acce941f2c9
--- /dev/null
+++ b/litellm-rust/crates/storage-clickhouse/Cargo.toml
@@ -0,0 +1,21 @@
+[package]
+name = "litellm-storage-clickhouse"
+version = "0.1.0"
+description = "Shared ClickHouse connection and HTTP storage for LiteLLM features"
+edition.workspace = true
+license.workspace = true
+repository.workspace = true
+
+[dependencies]
+flate2.workspace = true
+litellm-http.workspace = true
+serde.workspace = true
+serde_json.workspace = true
+thiserror.workspace = true
+url.workspace = true
+
+[dev-dependencies]
+litellm-http = { workspace = true, features = ["test-support"] }
+rstest.workspace = true
+tokio.workspace = true
+wiremock.workspace = true
diff --git a/litellm-rust/crates/storage-clickhouse/src/error.rs b/litellm-rust/crates/storage-clickhouse/src/error.rs
new file mode 100644
index 00000000000..acec9b91675
--- /dev/null
+++ b/litellm-rust/crates/storage-clickhouse/src/error.rs
@@ -0,0 +1,33 @@
+#[derive(Debug, thiserror::Error)]
+pub enum Error {
+ #[error("invalid ClickHouse insert row")]
+ InvalidRow,
+ #[error("{0} must be a positive integer")]
+ InvalidLimit(&'static str),
+ #[error("invalid ClickHouse insert table")]
+ InvalidTable,
+ #[error("invalid ClickHouse HTTP URL")]
+ InvalidUrl,
+ #[error("database must be a nonempty SQL identifier and retention must be positive")]
+ InvalidSchema,
+ #[error("SQL query must not be empty")]
+ EmptySql,
+ #[error("invalid ClickHouse query parameters")]
+ InvalidParameters,
+ #[error("unknown ClickHouse read query")]
+ InvalidQuery,
+ #[error("ClickHouse query failed with HTTP status {0}")]
+ QueryFailed(u16),
+ #[error("ClickHouse insert failed with HTTP status {0}")]
+ InsertFailed(u16),
+ #[error("ClickHouse insert exceeds the encoded size limit")]
+ InsertTooLarge,
+ #[error("ClickHouse schema setup failed with HTTP status {0}")]
+ SchemaFailed(u16),
+ #[error("ClickHouse query exceeded the response size limit")]
+ ResponseTooLarge,
+ #[error("ClickHouse returned an invalid or failed JSON query response")]
+ InvalidResponse,
+ #[error("ClickHouse query transport failed")]
+ Transport,
+}
diff --git a/litellm-rust/crates/storage-clickhouse/src/insert.rs b/litellm-rust/crates/storage-clickhouse/src/insert.rs
new file mode 100644
index 00000000000..1cdad1897c1
--- /dev/null
+++ b/litellm-rust/crates/storage-clickhouse/src/insert.rs
@@ -0,0 +1,99 @@
+use std::{io::Write, time::Duration};
+
+use flate2::{Compression, write::GzEncoder};
+use litellm_http::Client;
+
+use crate::{Connection, Error, valid_identifier};
+
+fn insert_timeout() -> Result {
+ let name = "CLICKHOUSE_INSERT_TIMEOUT_SECONDS";
+ match std::env::var(name) {
+ Ok(value) => value
+ .parse::()
+ .ok()
+ .filter(|value| *value > 0)
+ .map(Duration::from_secs)
+ .ok_or(Error::InvalidLimit(name)),
+ Err(std::env::VarError::NotPresent) => Ok(Duration::from_secs(30)),
+ Err(_) => Err(Error::InvalidLimit(name)),
+ }
+}
+
+pub async fn insert_encoded_rows(
+ client: &Client,
+ connection: &Connection,
+ database: &str,
+ table: &str,
+ token: &str,
+ encoded: &str,
+) -> Result<(), Error> {
+ if !valid_identifier(database) {
+ return Err(Error::InvalidSchema);
+ }
+ if !valid_identifier(table) {
+ return Err(Error::InvalidTable);
+ }
+ let mut encoder = GzEncoder::new(Vec::new(), Compression::default());
+ encoder
+ .write_all(encoded.as_bytes())
+ .map_err(|_| Error::InvalidRow)?;
+ let body = encoder.finish().map_err(|_| Error::InvalidRow)?;
+ insert_compressed_rows(client, connection, database, table, token, body).await
+}
+
+pub async fn insert_compressed_rows(
+ client: &Client,
+ connection: &Connection,
+ database: &str,
+ table: &str,
+ token: &str,
+ body: Vec,
+) -> Result<(), Error> {
+ if !valid_identifier(database) {
+ return Err(Error::InvalidSchema);
+ }
+ if !valid_identifier(table) {
+ return Err(Error::InvalidTable);
+ }
+ let mut url = connection.url().clone();
+ let existing_pairs: Vec<(String, String)> = url
+ .query_pairs()
+ .filter(|(key, _)| {
+ !matches!(
+ key.as_ref(),
+ "query"
+ | "async_insert"
+ | "async_insert_deduplicate"
+ | "wait_for_async_insert"
+ | "input_format_skip_unknown_fields"
+ | "date_time_input_format"
+ )
+ })
+ .map(|(key, value)| (key.into_owned(), value.into_owned()))
+ .collect();
+ url.query_pairs_mut()
+ .clear()
+ .extend_pairs(existing_pairs)
+ .append_pair(
+ "query",
+ &format!("INSERT INTO `{database}`.{} FORMAT JSONEachRow", table),
+ )
+ .append_pair("insert_deduplication_token", token)
+ .append_pair("async_insert", "1")
+ .append_pair("async_insert_deduplicate", "1")
+ .append_pair("wait_for_async_insert", "1")
+ .append_pair("input_format_skip_unknown_fields", "0")
+ .append_pair("date_time_input_format", "best_effort");
+ let response = client
+ .post(url)
+ .timeout(insert_timeout()?)
+ .header("Content-Encoding", "gzip")
+ .body(body)
+ .send()
+ .await
+ .map_err(|_| Error::Transport)?;
+ if !response.status().is_success() {
+ return Err(Error::InsertFailed(response.status().as_u16()));
+ }
+ Ok(())
+}
diff --git a/litellm-rust/crates/storage-clickhouse/src/lib.rs b/litellm-rust/crates/storage-clickhouse/src/lib.rs
new file mode 100644
index 00000000000..7ab2aa9bc0a
--- /dev/null
+++ b/litellm-rust/crates/storage-clickhouse/src/lib.rs
@@ -0,0 +1,125 @@
+mod error;
+mod insert;
+mod read;
+
+pub use error::Error;
+pub use insert::{insert_compressed_rows, insert_encoded_rows};
+pub use read::{Parameter, Query, READ_LIMITS, ReadLimits, execute_read, fetch, fetch_json};
+use url::Url;
+
+#[derive(Clone)]
+pub struct Connection {
+ url: Url,
+}
+
+impl Connection {
+ pub fn parse(value: &str) -> Result {
+ let url = Url::parse(value).map_err(|_| Error::InvalidUrl)?;
+ if !matches!(url.scheme(), "http" | "https") || url.host().is_none() {
+ return Err(Error::InvalidUrl);
+ }
+ Ok(Self { url })
+ }
+
+ pub fn configured(
+ url: &str,
+ database: &str,
+ user: &str,
+ password: &str,
+ ) -> Result {
+ let mut connection = Self::parse(url)?;
+ connection
+ .url
+ .set_username(user)
+ .map_err(|_| Error::InvalidUrl)?;
+ connection
+ .url
+ .set_password(Some(password))
+ .map_err(|_| Error::InvalidUrl)?;
+ let pairs: Vec<_> = connection
+ .url
+ .query_pairs()
+ .filter(|(key, _)| !matches!(key.as_ref(), "database" | "user" | "password"))
+ .map(|(key, value)| (key.into_owned(), value.into_owned()))
+ .collect();
+ connection
+ .url
+ .query_pairs_mut()
+ .clear()
+ .extend_pairs(pairs)
+ .append_pair("database", database);
+ Ok(connection)
+ }
+
+ pub fn writer(url: &str) -> Result {
+ let mut connection = Self::parse(url)?;
+ let pairs: Vec<_> = connection
+ .url
+ .query_pairs()
+ .filter(|(key, _)| !matches!(key.as_ref(), "database" | "readonly" | "query"))
+ .map(|(key, value)| (key.into_owned(), value.into_owned()))
+ .collect();
+ connection.url.query_pairs_mut().clear().extend_pairs(pairs);
+ Ok(connection)
+ }
+
+ pub fn reader(url: &str, database: &str) -> Result {
+ let mut connection = Self::parse(url)?;
+ let pairs: Vec<_> = connection
+ .url
+ .query_pairs()
+ .filter(|(key, _)| key != "database")
+ .map(|(key, value)| (key.into_owned(), value.into_owned()))
+ .collect();
+ connection
+ .url
+ .query_pairs_mut()
+ .clear()
+ .extend_pairs(pairs)
+ .append_pair("database", database);
+ Ok(connection)
+ }
+
+ pub fn url(&self) -> &Url {
+ &self.url
+ }
+}
+
+#[derive(Clone)]
+pub struct Storage {
+ database: String,
+ writer: Connection,
+ reader: Connection,
+}
+
+impl Storage {
+ pub fn new(database: String, url: &str) -> Result {
+ if !valid_identifier(&database) {
+ return Err(Error::InvalidSchema);
+ }
+ Ok(Self {
+ writer: Connection::writer(url)?,
+ reader: Connection::reader(url, &database)?,
+ database,
+ })
+ }
+
+ pub fn database(&self) -> &str {
+ &self.database
+ }
+
+ pub fn writer(&self) -> &Connection {
+ &self.writer
+ }
+
+ pub fn reader(&self) -> &Connection {
+ &self.reader
+ }
+}
+
+pub(crate) fn valid_identifier(value: &str) -> bool {
+ !value.is_empty()
+ && value
+ .bytes()
+ .all(|c| c.is_ascii_alphanumeric() || c == b'_')
+}
diff --git a/litellm-rust/crates/traces/src/sql.rs b/litellm-rust/crates/storage-clickhouse/src/read.rs
similarity index 60%
rename from litellm-rust/crates/traces/src/sql.rs
rename to litellm-rust/crates/storage-clickhouse/src/read.rs
index 8346e06cb71..c4bfdef393a 100644
--- a/litellm-rust/crates/traces/src/sql.rs
+++ b/litellm-rust/crates/storage-clickhouse/src/read.rs
@@ -1,46 +1,30 @@
use std::{collections::BTreeMap, time::Duration};
-use serde::Deserialize;
-
use litellm_http::Client;
+use serde::{Deserialize, Serialize, de::DeserializeOwned};
use crate::{Connection, Error};
-const MAX_RESPONSE_BYTES: usize = 4 * 1024 * 1024;
-
-pub enum ReadQuery {
- ListTraces,
- TraceSpans,
- SpanDetail,
- SpendByResponseIds,
+#[derive(Clone, Copy, Debug, Eq, PartialEq)]
+pub struct ReadLimits {
+ pub result_rows: u64,
+ pub response_bytes: usize,
+ pub execution_seconds: u64,
}
-impl ReadQuery {
- pub fn parse(value: &str) -> Result {
- match value {
- "list_traces" => Ok(Self::ListTraces),
- "trace_spans" => Ok(Self::TraceSpans),
- "span_detail" => Ok(Self::SpanDetail),
- "spend_by_response_ids" => Ok(Self::SpendByResponseIds),
- _ => Err(Error::InvalidQuery),
- }
- }
+pub const READ_LIMITS: ReadLimits = ReadLimits {
+ result_rows: 1000,
+ response_bytes: 4 * 1024 * 1024,
+ execution_seconds: 10,
+};
- fn sql(&self) -> &'static str {
- match self {
- Self::ListTraces => include_str!("../query/list_traces.sql"),
- Self::TraceSpans => include_str!("../query/trace_spans.sql"),
- Self::SpanDetail => include_str!("../query/span_detail.sql"),
- Self::SpendByResponseIds => include_str!("../query/spend_by_response_ids.sql"),
- }
- }
-}
-
-#[derive(Debug, Deserialize)]
+#[derive(Debug, Deserialize, Serialize)]
#[serde(untagged)]
pub enum Parameter {
Text(String),
Integer(i64),
+ Unsigned(u64),
+ Float(f64),
Strings(Vec),
}
@@ -49,6 +33,8 @@ impl Parameter {
match self {
Self::Text(value) => escaped(value),
Self::Integer(value) => value.to_string(),
+ Self::Unsigned(value) => value.to_string(),
+ Self::Float(value) => value.to_string(),
Self::Strings(values) => format!(
"[{}]",
values
@@ -103,9 +89,12 @@ pub async fn execute_read(
.clear()
.extend_pairs(existing_pairs)
.append_pair("readonly", "1")
- .append_pair("max_result_rows", "1000")
+ .append_pair("max_result_rows", &READ_LIMITS.result_rows.to_string())
.append_pair("result_overflow_mode", "throw")
- .append_pair("max_execution_time", "10")
+ .append_pair(
+ "max_execution_time",
+ &READ_LIMITS.execution_seconds.to_string(),
+ )
.append_pair("wait_end_of_query", "1")
.append_pair("default_format", "JSON");
@@ -121,12 +110,19 @@ pub async fn execute_read(
.body(sql.to_owned());
let mut response = request.send().await.map_err(|_| Error::Transport)?;
if !response.status().is_success() {
+ if response
+ .headers()
+ .get("x-clickhouse-exception-code")
+ .is_some_and(|code| code == "396")
+ {
+ return Err(Error::ResponseTooLarge);
+ }
return Err(Error::QueryFailed(response.status().as_u16()));
}
let mut body = Vec::new();
while let Some(chunk) = response.chunk().await.map_err(|_| Error::Transport)? {
- if body.len() + chunk.len() > MAX_RESPONSE_BYTES {
+ if body.len() + chunk.len() > READ_LIMITS.response_bytes {
return Err(Error::ResponseTooLarge);
}
body.extend_from_slice(&chunk);
@@ -141,36 +137,44 @@ pub async fn execute_read(
String::from_utf8(body).map_err(|_| Error::InvalidResponse)
}
-#[derive(Clone, Copy)]
-pub enum LensQuery {
- Sample,
- Content,
- Evidence,
+pub trait Query {
+ type Params: Serialize;
+ type Row: DeserializeOwned;
+
+ const SQL: &'static str;
}
-impl LensQuery {
- pub fn parse(name: &str) -> Result {
- match name {
- "sample" => Ok(Self::Sample),
- "content" => Ok(Self::Content),
- "evidence" => Ok(Self::Evidence),
- _ => Err(Error::InvalidQuery),
- }
- }
- pub fn sql(self) -> &'static str {
- match self {
- Self::Sample => include_str!("../query/lens_sample.sql"),
- Self::Content => include_str!("../query/lens_content.sql"),
- Self::Evidence => include_str!("../query/lens_evidence.sql"),
- }
- }
+#[derive(Deserialize)]
+struct Rows {
+ data: Vec,
}
-pub async fn execute_named_read(
+fn parameters(params: &T) -> Result, Error> {
+ let value = serde_json::to_value(params).map_err(|_| Error::InvalidParameters)?;
+ serde_json::from_value(value).map_err(|_| Error::InvalidParameters)
+}
+
+pub async fn fetch(
client: &Client,
connection: &Connection,
- query: ReadQuery,
- parameters: &BTreeMap,
-) -> Result {
- execute_read(client, connection, query.sql(), parameters).await
+ params: &Q::Params,
+) -> Result, Error> {
+ let body = execute_read(client, connection, Q::SQL, ¶meters(params)?).await?;
+ decode_rows::(&body)
+}
+
+pub async fn fetch_json(
+ client: &Client,
+ connection: &Connection,
+ params: &Q::Params,
+) -> Result {
+ let body = execute_read(client, connection, Q::SQL, ¶meters(params)?).await?;
+ decode_rows::(&body)?;
+ Ok(body)
+}
+
+fn decode_rows(body: &str) -> Result, Error> {
+ serde_json::from_str::>(body)
+ .map(|rows| rows.data)
+ .map_err(|_| Error::InvalidResponse)
}
diff --git a/litellm-rust/crates/storage-clickhouse/tests/connection.rs b/litellm-rust/crates/storage-clickhouse/tests/connection.rs
new file mode 100644
index 00000000000..e371718259a
--- /dev/null
+++ b/litellm-rust/crates/storage-clickhouse/tests/connection.rs
@@ -0,0 +1,39 @@
+use litellm_storage_clickhouse::{Connection, Storage};
+use rstest::rstest;
+
+#[rstest]
+#[case::http("http://localhost:8123", true)]
+#[case::https("https://localhost:8443", true)]
+#[case::tcp("tcp://localhost:9000", false)]
+#[case::missing_host("http://", false)]
+fn accepts_only_clickhouse_http_urls(#[case] value: &str, #[case] expected: bool) {
+ assert_eq!(Connection::parse(value).is_ok(), expected);
+}
+
+#[test]
+fn storage_uses_one_url_for_writes_and_bounded_reads() {
+ let storage =
+ Storage::new("litellm".to_owned(), "http://localhost:8123").expect("valid ClickHouse URLs");
+
+ assert_eq!(storage.database(), "litellm");
+ assert_eq!(storage.writer().url().host_str(), Some("localhost"));
+ assert_eq!(storage.writer().url().port(), Some(8123));
+ assert_eq!(storage.reader().url().port(), Some(8123));
+ assert_eq!(
+ storage
+ .reader()
+ .url()
+ .query_pairs()
+ .find(|(key, _)| key == "database")
+ .unwrap()
+ .1,
+ "litellm"
+ );
+}
+
+#[rstest]
+#[case::empty("")]
+#[case::injection("db; DROP DATABASE default")]
+fn storage_rejects_invalid_database(#[case] database: &str) {
+ assert!(Storage::new(database.to_owned(), "http://localhost:8123").is_err());
+}
diff --git a/litellm-rust/crates/storage-clickhouse/tests/transport.rs b/litellm-rust/crates/storage-clickhouse/tests/transport.rs
new file mode 100644
index 00000000000..f70fe52e1de
--- /dev/null
+++ b/litellm-rust/crates/storage-clickhouse/tests/transport.rs
@@ -0,0 +1,193 @@
+use std::collections::BTreeMap;
+
+use litellm_http::Client;
+use litellm_storage_clickhouse::{Connection, Error, Query, execute_read, insert_encoded_rows};
+use rstest::rstest;
+
+#[rstest]
+#[case::invalid_database("db; DROP DATABASE default", "spend_logs", true)]
+#[case::invalid_table("litellm", "spend_logs; DROP TABLE otel_traces", false)]
+#[tokio::test]
+async fn insert_rejects_invalid_identifiers(
+ #[case] database: &str,
+ #[case] table: &str,
+ #[case] invalid_database: bool,
+) {
+ let client = Client::no_redirect_for_test();
+ let connection = Connection::writer("http://localhost:8123").expect("valid URL");
+ let result = insert_encoded_rows(&client, &connection, database, table, "token", "{}").await;
+
+ assert!(matches!(&result, Err(Error::InvalidSchema)) == invalid_database);
+ assert!(matches!(&result, Err(Error::InvalidTable)) == !invalid_database);
+}
+
+#[rstest]
+#[tokio::test]
+async fn read_rejects_empty_sql() {
+ let client = Client::no_redirect_for_test();
+ let connection = Connection::reader("http://localhost:8123", "litellm").expect("valid URL");
+
+ assert!(matches!(
+ execute_read(&client, &connection, " ", &BTreeMap::new()).await,
+ Err(Error::EmptySql)
+ ));
+}
+
+#[derive(serde::Serialize)]
+struct QueryParams {
+ signed: i64,
+ unsigned: u64,
+ float: f64,
+ text: String,
+ strings: Vec,
+}
+
+#[derive(Debug, serde::Deserialize, PartialEq)]
+struct QueryRow {
+ answer: String,
+}
+
+struct TypedQuery;
+
+impl litellm_storage_clickhouse::Query for TypedQuery {
+ type Params = QueryParams;
+ type Row = QueryRow;
+ const SQL: &'static str = "SELECT typed_parameters";
+}
+
+#[rstest]
+#[case::valid(
+ r#"{"meta":[],"data":[{"answer":"ok"}],"rows":1,"statistics":{"elapsed":0.1}}"#,
+ true
+)]
+#[case::wrong_type(r#"{"data":[{"answer":1}]}"#, false)]
+#[case::missing_column(r#"{"data":[{}]}"#, false)]
+#[case::exception(r#"{"data":[],"exception":"failed"}"#, false)]
+#[tokio::test]
+async fn typed_fetch_encodes_parameters_and_validates_rows(
+ #[case] body: &str,
+ #[case] valid: bool,
+) {
+ use litellm_storage_clickhouse::{fetch, fetch_json};
+ use wiremock::{
+ Mock, MockServer, ResponseTemplate,
+ matchers::{body_string, query_param},
+ };
+
+ let server = MockServer::start().await;
+ Mock::given(body_string(TypedQuery::SQL))
+ .and(query_param("param_signed", i64::MIN.to_string()))
+ .and(query_param("param_unsigned", u64::MAX.to_string()))
+ .and(query_param("param_float", "12.5"))
+ .and(query_param("param_text", "line\\nbreak"))
+ .and(query_param("param_strings", "['a\\'b','雪']"))
+ .and(query_param("readonly", "1"))
+ .and(query_param("max_result_rows", "1000"))
+ .respond_with(ResponseTemplate::new(200).set_body_string(body))
+ .expect(2)
+ .mount(&server)
+ .await;
+ let client = Client::no_redirect_for_test();
+ let connection = Connection::parse(&server.uri()).unwrap();
+ let params = QueryParams {
+ signed: i64::MIN,
+ unsigned: u64::MAX,
+ float: 12.5,
+ text: "line\nbreak".into(),
+ strings: vec!["a'b".into(), "雪".into()],
+ };
+ let rows = fetch::(&client, &connection, ¶ms).await;
+ let envelope = fetch_json::(&client, &connection, ¶ms).await;
+ if valid {
+ assert_eq!(
+ rows.unwrap(),
+ vec![QueryRow {
+ answer: "ok".into()
+ }]
+ );
+ assert_eq!(envelope.unwrap(), body);
+ } else {
+ assert!(matches!(rows, Err(Error::InvalidResponse)));
+ assert!(matches!(envelope, Err(Error::InvalidResponse)));
+ }
+}
+
+#[rstest]
+#[case::result_limit("396", true)]
+#[case::memory_limit("241", false)]
+#[case::timeout("159", false)]
+#[case::unknown("", false)]
+#[tokio::test]
+async fn server_result_limits_allow_smaller_pages_without_retrying_other_failures(
+ #[case] code: &str,
+ #[case] result_limit: bool,
+) {
+ use wiremock::{Mock, MockServer, ResponseTemplate, matchers::method};
+ let server = MockServer::start().await;
+ Mock::given(method("POST"))
+ .respond_with(ResponseTemplate::new(500).insert_header("X-ClickHouse-Exception-Code", code))
+ .expect(1)
+ .mount(&server)
+ .await;
+ let connection = Connection::parse(&server.uri()).unwrap();
+ let error = execute_read(
+ &Client::no_redirect_for_test(),
+ &connection,
+ "SELECT 1",
+ &BTreeMap::new(),
+ )
+ .await
+ .unwrap_err();
+ if result_limit {
+ assert!(matches!(error, Error::ResponseTooLarge));
+ } else {
+ assert!(matches!(error, Error::QueryFailed(500)));
+ }
+}
+
+#[test]
+fn insert_timeout_environment_controls_transport() {
+ for value in ["1", "3", "0", "invalid"] {
+ let result = std::process::Command::new(std::env::current_exe().unwrap())
+ .args(["--exact", "insert_timeout_environment_child"])
+ .env("LITELLM_TEST_INSERT_TIMEOUT", value)
+ .env("CLICKHOUSE_INSERT_TIMEOUT_SECONDS", value)
+ .output()
+ .unwrap();
+ assert!(
+ result.status.success(),
+ "{}",
+ String::from_utf8_lossy(&result.stdout)
+ );
+ }
+}
+
+#[tokio::test]
+async fn insert_timeout_environment_child() {
+ use wiremock::{Mock, MockServer, ResponseTemplate, matchers::method};
+ let Ok(value) = std::env::var("LITELLM_TEST_INSERT_TIMEOUT") else {
+ return;
+ };
+ let server = MockServer::start().await;
+ Mock::given(method("POST"))
+ .respond_with(ResponseTemplate::new(200).set_delay(std::time::Duration::from_millis(1500)))
+ .mount(&server)
+ .await;
+ let result = insert_encoded_rows(
+ &Client::no_redirect_for_test(),
+ &Connection::parse(&server.uri()).unwrap(),
+ "traces",
+ "otel_traces",
+ "token",
+ "{}",
+ )
+ .await;
+ match value.as_str() {
+ "1" => assert!(matches!(result, Err(Error::Transport))),
+ "3" => assert!(result.is_ok()),
+ _ => assert!(matches!(
+ result,
+ Err(Error::InvalidLimit("CLICKHOUSE_INSERT_TIMEOUT_SECONDS"))
+ )),
+ }
+}
diff --git a/litellm-rust/crates/traces-cache/AGENTS.md b/litellm-rust/crates/traces-cache/AGENTS.md
new file mode 100644
index 00000000000..bfaea42d901
--- /dev/null
+++ b/litellm-rust/crates/traces-cache/AGENTS.md
@@ -0,0 +1,5 @@
+Own resolved-trace snapshot storage, cache identity, weighting, and expiry
+Depend on trace domain types, never storage, HTTP, or Python
+Preserve the full source and authorization scope in every cache key
+Keep snapshots immutable and expose borrowed data
+Keep cursor formats and database reads in their existing owners
diff --git a/litellm-rust/crates/traces-cache/Cargo.toml b/litellm-rust/crates/traces-cache/Cargo.toml
new file mode 100644
index 00000000000..3e58ea4c965
--- /dev/null
+++ b/litellm-rust/crates/traces-cache/Cargo.toml
@@ -0,0 +1,17 @@
+[package]
+name = "litellm-traces-cache"
+version = "0.1.0"
+edition.workspace = true
+license.workspace = true
+repository.workspace = true
+
+[dependencies]
+litellm-traces.workspace = true
+moka.workspace = true
+serde_json.workspace = true
+sha2.workspace = true
+thiserror.workspace = true
+
+[dev-dependencies]
+rstest.workspace = true
+tokio.workspace = true
diff --git a/litellm-rust/crates/traces-cache/src/error.rs b/litellm-rust/crates/traces-cache/src/error.rs
new file mode 100644
index 00000000000..ac5f19369fe
--- /dev/null
+++ b/litellm-rust/crates/traces-cache/src/error.rs
@@ -0,0 +1,7 @@
+#[derive(Debug, thiserror::Error)]
+pub enum Error {
+ #[error("trace snapshot serialization failed")]
+ Serialization(#[from] serde_json::Error),
+ #[error("trace snapshot exceeds the size limit")]
+ ReadTooLarge,
+}
diff --git a/litellm-rust/crates/traces-cache/src/lib.rs b/litellm-rust/crates/traces-cache/src/lib.rs
new file mode 100644
index 00000000000..7dfd3bf32a1
--- /dev/null
+++ b/litellm-rust/crates/traces-cache/src/lib.rs
@@ -0,0 +1,166 @@
+use std::{sync::Arc, time::Duration};
+
+use litellm_traces::{Trace, query::named::ReadAccessParams};
+use moka::future::Cache;
+use sha2::{Digest, Sha256};
+
+mod error;
+
+pub use error::Error;
+
+#[derive(Clone, Eq, Hash, PartialEq)]
+pub struct SnapshotKey(String);
+
+impl SnapshotKey {
+ pub fn new(
+ source: &str,
+ access: &ReadAccessParams,
+ trace_id: &str,
+ trace_ref: &str,
+ snapshot_ms: u64,
+ ) -> Result {
+ let encoded = serde_json::to_vec(&(source, access, trace_id, trace_ref, snapshot_ms))?;
+ Ok(Self(format!("{:x}", Sha256::digest(encoded))))
+ }
+}
+
+pub struct Snapshot {
+ trace: Trace,
+ version: String,
+ weight: u32,
+}
+
+impl Snapshot {
+ pub fn trace(&self) -> &Trace {
+ &self.trace
+ }
+
+ pub fn version(&self) -> &str {
+ &self.version
+ }
+}
+
+pub struct SnapshotCache {
+ entries: Cache>,
+ max_graph_bytes: usize,
+}
+
+impl SnapshotCache {
+ pub fn new(max_graph_bytes: usize, ttl: Duration) -> Self {
+ Self {
+ entries: Cache::builder()
+ .max_capacity((max_graph_bytes as u64).saturating_mul(2))
+ .weigher(|_: &SnapshotKey, snapshot: &Arc| snapshot.weight)
+ .time_to_live(ttl)
+ .build(),
+ max_graph_bytes,
+ }
+ }
+
+ pub async fn get(&self, key: &SnapshotKey) -> Option> {
+ self.entries.get(key).await
+ }
+
+ pub async fn insert(&self, key: SnapshotKey, trace: Trace) -> Result, Error> {
+ let encoded = serde_json::to_vec(&trace)?;
+ if encoded.len() > self.max_graph_bytes {
+ return Err(Error::ReadTooLarge);
+ }
+
+ let span_ids: Vec<&str> = trace
+ .spans
+ .iter()
+ .map(|span| span.span_id.as_str())
+ .collect();
+
+ let version = format!("{:x}", Sha256::digest(serde_json::to_vec(&span_ids)?));
+
+ let snapshot = Arc::new(Snapshot {
+ trace,
+ version,
+ weight: u32::try_from(encoded.len().saturating_mul(2)).unwrap_or(u32::MAX),
+ });
+
+ self.entries.insert(key, Arc::clone(&snapshot)).await;
+ Ok(snapshot)
+ }
+}
+
+#[cfg(test)]
+mod tests {
+ use litellm_traces::{
+ SpanStatus,
+ query::named::{SpendByResponseIdsRow, TraceSpansRow},
+ resolve_trace,
+ };
+
+ use super::*;
+
+ fn trace(span_id: &str) -> Trace {
+ let rows = [TraceSpansRow {
+ trace_id: String::new(),
+ span_id: span_id.into(),
+ parent_span_id: String::new(),
+ name: "run".into(),
+ kind: litellm_traces::ObservationType::Agent,
+ wrapper_candidate: false,
+ agent: "agent".into(),
+ framework: String::new(),
+ status: SpanStatus::Ok,
+ status_message: String::new(),
+ error_truncated: false,
+ start_ns: 1_790_742_989_000_000_000,
+ duration_ns: 10_000_000,
+ service: "agent-demo".into(),
+ input_preview: format!("input of {span_id}"),
+ model: String::new(),
+ input_tokens: 0,
+ output_tokens: 0,
+ litellm_request_id: String::new(),
+ call_keys: Vec::new(),
+ call_evidence: None,
+ tool_call_id: String::new(),
+ team_id: String::new(),
+ api_key_hash: String::new(),
+ user_id: String::new(),
+ }];
+ resolve_trace("trace", "ref", &rows, &[] as &[SpendByResponseIdsRow])
+ .expect("fixture should resolve")
+ }
+
+ fn key(suffix: &str) -> SnapshotKey {
+ SnapshotKey::new(
+ "source",
+ &ReadAccessParams {
+ all_teams: false,
+ user_id: String::new(),
+ team_ids: vec!["team".into()],
+ },
+ suffix,
+ "ref",
+ 100,
+ )
+ .unwrap()
+ }
+
+ #[tokio::test]
+ async fn weighted_capacity_bounds_retained_snapshots() {
+ let limit = ["first", "second", "third"]
+ .iter()
+ .map(|span_id| serde_json::to_vec(&trace(span_id)).unwrap().len())
+ .max()
+ .unwrap();
+ let cache = SnapshotCache::new(limit, Duration::from_secs(120));
+
+ for (key, span_id) in [
+ (key("a"), "first"),
+ (key("b"), "second"),
+ (key("c"), "third"),
+ ] {
+ cache.insert(key, trace(span_id)).await.unwrap();
+ }
+
+ cache.entries.run_pending_tasks().await;
+ assert!(cache.entries.weighted_size() <= (limit as u64) * 2);
+ }
+}
diff --git a/litellm-rust/crates/traces-cache/tests/snapshots.rs b/litellm-rust/crates/traces-cache/tests/snapshots.rs
new file mode 100644
index 00000000000..de199ae0301
--- /dev/null
+++ b/litellm-rust/crates/traces-cache/tests/snapshots.rs
@@ -0,0 +1,191 @@
+use std::time::Duration;
+
+use litellm_traces::{
+ SpanStatus, Trace,
+ query::named::{ReadAccessParams, SpendByResponseIdsRow, TraceSpansRow},
+ resolve_trace,
+};
+use litellm_traces_cache::{Error, SnapshotCache, SnapshotKey};
+use rstest::{fixture, rstest};
+
+const T0: i64 = 1_790_742_989_000_000_000;
+const MS: i64 = 1_000_000;
+const TTL: Duration = Duration::from_secs(120);
+
+fn row(span_id: &str, parent: &str, name: &str, kind: &str, agent: &str) -> TraceSpansRow {
+ TraceSpansRow {
+ trace_id: String::new(),
+ span_id: span_id.into(),
+ parent_span_id: parent.into(),
+ name: name.into(),
+ kind: kind.parse().unwrap(),
+ wrapper_candidate: false,
+ agent: agent.into(),
+ framework: String::new(),
+ status: SpanStatus::Ok,
+ status_message: String::new(),
+ error_truncated: false,
+ start_ns: T0,
+ duration_ns: 10 * MS as u64,
+ service: "agent-demo".into(),
+ input_preview: format!("input of {name}"),
+ model: String::new(),
+ input_tokens: 0,
+ output_tokens: 0,
+ litellm_request_id: String::new(),
+ call_keys: Vec::new(),
+ call_evidence: None,
+ tool_call_id: String::new(),
+ team_id: String::new(),
+ api_key_hash: String::new(),
+ user_id: String::new(),
+ }
+}
+
+fn access() -> ReadAccessParams {
+ ReadAccessParams {
+ all_teams: false,
+ user_id: String::new(),
+ team_ids: vec!["team".into()],
+ }
+}
+
+fn key(
+ source: &str,
+ access: &ReadAccessParams,
+ trace_id: &str,
+ trace_ref: &str,
+ ms: u64,
+) -> SnapshotKey {
+ SnapshotKey::new(source, access, trace_id, trace_ref, ms).unwrap()
+}
+
+#[fixture]
+fn trace() -> Trace {
+ resolve_trace(
+ "trace",
+ "ref",
+ &[row("root", "", "run", "agent", "agent")],
+ &[] as &[SpendByResponseIdsRow],
+ )
+ .expect("fixture should resolve")
+}
+
+#[rstest]
+#[case::different_team(false, "", "other-team")]
+#[case::different_user(false, "other-user", "team")]
+#[case::different_scope(true, "", "team")]
+#[tokio::test]
+async fn cached_trace_is_isolated_by_access_scope(
+ trace: Trace,
+ #[case] all_teams: bool,
+ #[case] user_id: &str,
+ #[case] team_id: &str,
+) {
+ let cache = SnapshotCache::new(1024 * 1024, TTL);
+ let stored = key("source", &access(), "trace", "ref", 100);
+
+ cache.insert(stored.clone(), trace.clone()).await.unwrap();
+
+ let other_access = ReadAccessParams {
+ all_teams,
+ user_id: user_id.into(),
+ team_ids: vec![team_id.into()],
+ };
+ let other = key("source", &other_access, "trace", "ref", 100);
+
+ assert!(cache.get(&other).await.is_none());
+ let cached = cache.get(&stored).await.unwrap();
+ assert_eq!(cached.trace(), &trace);
+}
+
+#[rstest]
+#[case::different_source("other-source", "trace", "ref", 100)]
+#[case::different_trace_id("source", "other-trace", "ref", 100)]
+#[case::different_trace_ref("source", "trace", "other-ref", 100)]
+#[case::different_snapshot_ms("source", "trace", "ref", 200)]
+#[tokio::test]
+async fn cached_trace_is_isolated_by_key_fields(
+ trace: Trace,
+ #[case] source: &str,
+ #[case] trace_id: &str,
+ #[case] trace_ref: &str,
+ #[case] snapshot_ms: u64,
+) {
+ let cache = SnapshotCache::new(1024 * 1024, TTL);
+ let stored = key("source", &access(), "trace", "ref", 100);
+
+ cache.insert(stored.clone(), trace.clone()).await.unwrap();
+
+ let other = key(source, &access(), trace_id, trace_ref, snapshot_ms);
+ assert!(cache.get(&other).await.is_none());
+ assert!(cache.get(&stored).await.is_some());
+}
+
+#[rstest]
+#[tokio::test]
+async fn snapshot_at_the_size_limit_is_accepted(trace: Trace) {
+ let size = serde_json::to_vec(&trace).unwrap().len();
+ let cache = SnapshotCache::new(size, TTL);
+ let stored = key("source", &access(), "trace", "ref", 100);
+
+ cache.insert(stored.clone(), trace).await.unwrap();
+ assert!(cache.get(&stored).await.is_some());
+}
+
+#[rstest]
+#[tokio::test]
+async fn snapshot_one_byte_over_the_size_limit_is_rejected(trace: Trace) {
+ let size = serde_json::to_vec(&trace).unwrap().len();
+ let cache = SnapshotCache::new(size - 1, TTL);
+ let stored = key("source", &access(), "trace", "ref", 100);
+
+ assert!(matches!(
+ cache.insert(stored.clone(), trace).await,
+ Err(Error::ReadTooLarge)
+ ));
+ assert!(cache.get(&stored).await.is_none());
+}
+
+#[rstest]
+#[case::same_ids(&["root", "child"], &["root", "child"], true)]
+#[case::different_ids(&["root", "child"], &["root", "other"], false)]
+#[tokio::test]
+async fn snapshot_version_tracks_the_ordered_span_ids(
+ #[case] first_ids: &[&str],
+ #[case] second_ids: &[&str],
+ #[case] equal: bool,
+) {
+ let build = |ids: &[&str]| -> Trace {
+ let rows: Vec = ids
+ .iter()
+ .map(|span_id| row(span_id, "", "run", "agent", "agent"))
+ .collect();
+ resolve_trace("trace", "ref", &rows, &[] as &[SpendByResponseIdsRow])
+ .expect("fixture should resolve")
+ };
+ let cache = SnapshotCache::new(1024 * 1024, TTL);
+
+ let first = cache
+ .insert(key("source", &access(), "a", "ref", 100), build(first_ids))
+ .await
+ .unwrap();
+ let second = cache
+ .insert(key("source", &access(), "b", "ref", 100), build(second_ids))
+ .await
+ .unwrap();
+
+ assert_eq!(first.version() == second.version(), equal);
+}
+
+#[rstest]
+#[tokio::test]
+async fn snapshots_expire_after_the_ttl(trace: Trace) {
+ let cache = SnapshotCache::new(1024 * 1024, Duration::from_millis(50));
+ let stored = key("source", &access(), "trace", "ref", 100);
+
+ cache.insert(stored.clone(), trace).await.unwrap();
+ tokio::time::sleep(Duration::from_millis(200)).await;
+
+ assert!(cache.get(&stored).await.is_none());
+}
diff --git a/litellm-rust/crates/traces-clickhouse/AGENTS.md b/litellm-rust/crates/traces-clickhouse/AGENTS.md
new file mode 100644
index 00000000000..ae0c1c8eeb9
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/AGENTS.md
@@ -0,0 +1,7 @@
+- Own trace schema, row encoding, SQL query adapters and reader provisioning; consume domain types from `litellm-traces`
+- Keep generic ClickHouse connections and HTTP execution in `litellm-storage-clickhouse`; keep PyO3 conversion in `python-bridge`
+- Keep schema definitions only in `migrations/NNNN_description.sql`, embedded by `litellm_migrate::migrate!`
+- Require typed query parameters and SELECT-only readers with server-side limits and tenant isolation
+- Bound insert time and encoded bytes; preserve shared values and explicit retry deduplication
+- Test storage behavior through the public API against ClickHouse
+- Expose one top-level `Error` enum in `src/error.rs`; own trace failures and wrap storage errors with `#[from]` or `#[source]`
diff --git a/litellm-rust/crates/traces-clickhouse/Cargo.toml b/litellm-rust/crates/traces-clickhouse/Cargo.toml
new file mode 100644
index 00000000000..6b95eb149d3
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/Cargo.toml
@@ -0,0 +1,46 @@
+[package]
+name = "litellm-traces-clickhouse"
+version = "0.1.0"
+edition.workspace = true
+license.workspace = true
+repository.workspace = true
+
+[features]
+schema = ["dep:schemars", "litellm-traces/schema"]
+
+[dependencies]
+macro_rules_attribute.workspace = true
+schemars = { workspace = true, optional = true }
+askama.workspace = true
+base64.workspace = true
+flate2.workspace = true
+futures-util.workspace = true
+hmac = "0.12.1"
+itertools = "0.14.0"
+litellm-http.workspace = true
+litellm-migrate.workspace = true
+litellm-storage-clickhouse.workspace = true
+litellm-traces.workspace = true
+litellm-traces-cache.workspace = true
+moka.workspace = true
+serde.workspace = true
+serde_json.workspace = true
+sha2.workspace = true
+strum.workspace = true
+thiserror.workspace = true
+time = { workspace = true, features = ["formatting"] }
+tokio.workspace = true
+tracing.workspace = true
+url.workspace = true
+
+[dev-dependencies]
+jsonschema = { version = "0.55.1", default-features = false }
+litellm-http = { workspace = true, features = ["test-support"] }
+rstest.workspace = true
+testcontainers-modules = { version = "0.15.0", features = ["clickhouse"] }
+wiremock.workspace = true
+
+[[bin]]
+name = "export-traces-clickhouse-schema"
+path = "src/bin/export_schema.rs"
+required-features = ["schema"]
diff --git a/litellm-rust/crates/traces-clickhouse/build.rs b/litellm-rust/crates/traces-clickhouse/build.rs
new file mode 100644
index 00000000000..3a8149ef075
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/build.rs
@@ -0,0 +1,3 @@
+fn main() {
+ println!("cargo:rerun-if-changed=migrations");
+}
diff --git a/litellm-rust/crates/traces/migrations/0001_otel_traces.sql b/litellm-rust/crates/traces-clickhouse/migrations/0001_otel_traces.sql
similarity index 94%
rename from litellm-rust/crates/traces/migrations/0001_otel_traces.sql
rename to litellm-rust/crates/traces-clickhouse/migrations/0001_otel_traces.sql
index d8e0184b5a3..fb5eaa367d7 100644
--- a/litellm-rust/crates/traces/migrations/0001_otel_traces.sql
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0001_otel_traces.sql
@@ -38,10 +38,11 @@ CREATE TABLE IF NOT EXISTS {database}.otel_traces
Input String CODEC(ZSTD(3)),
Output String CODEC(ZSTD(3)),
InputPreview String DEFAULT substring(Input, 1, 240),
+ EngineReceivedMs UInt64 DEFAULT 0,
INDEX idx_trace_id TraceId TYPE bloom_filter(0.001) GRANULARITY 1,
INDEX idx_req_id LiteLLMRequestId TYPE bloom_filter(0.01) GRANULARITY 1
)
ENGINE = MergeTree
PARTITION BY toDate(Timestamp)
ORDER BY (TeamId, ServiceName, toDateTime(Timestamp), TraceId)
-SETTINGS ttl_only_drop_parts = 1, non_replicated_deduplication_window = 1000
+SETTINGS ttl_only_drop_parts = 1, materialize_ttl_recalculate_only = 1, non_replicated_deduplication_window = 1000
diff --git a/litellm-rust/crates/traces/migrations/0005_otel_traces_ttl.sql b/litellm-rust/crates/traces-clickhouse/migrations/0002_otel_traces_ttl.sql
similarity index 60%
rename from litellm-rust/crates/traces/migrations/0005_otel_traces_ttl.sql
rename to litellm-rust/crates/traces-clickhouse/migrations/0002_otel_traces_ttl.sql
index 4ac597b8902..7402634b7e1 100644
--- a/litellm-rust/crates/traces/migrations/0005_otel_traces_ttl.sql
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0002_otel_traces_ttl.sql
@@ -1 +1 @@
-ALTER TABLE {database}.otel_traces MODIFY TTL toDateTime(Timestamp) + INTERVAL {trace_retention_days} DAY
+ALTER TABLE {database}.otel_traces MODIFY TTL toDateTime(Timestamp) + INTERVAL {retention_days} DAY
diff --git a/litellm-rust/crates/traces/migrations/0002_agent_traces.sql b/litellm-rust/crates/traces-clickhouse/migrations/0003_agent_traces.sql
similarity index 93%
rename from litellm-rust/crates/traces/migrations/0002_agent_traces.sql
rename to litellm-rust/crates/traces-clickhouse/migrations/0003_agent_traces.sql
index 0c3547872bb..821cc2f3723 100644
--- a/litellm-rust/crates/traces/migrations/0002_agent_traces.sql
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0003_agent_traces.sql
@@ -22,4 +22,4 @@ CREATE TABLE IF NOT EXISTS {database}.agent_traces_by_key
)
ENGINE = AggregatingMergeTree
ORDER BY (TeamId, ApiKeyHash, TraceId)
-SETTINGS non_replicated_deduplication_window = 1000
+SETTINGS materialize_ttl_recalculate_only = 1, non_replicated_deduplication_window = 1000
diff --git a/litellm-rust/crates/traces/migrations/0006_agent_traces_ttl.sql b/litellm-rust/crates/traces-clickhouse/migrations/0004_agent_traces_ttl.sql
similarity index 57%
rename from litellm-rust/crates/traces/migrations/0006_agent_traces_ttl.sql
rename to litellm-rust/crates/traces-clickhouse/migrations/0004_agent_traces_ttl.sql
index 8681f0622a4..70147f95d0e 100644
--- a/litellm-rust/crates/traces/migrations/0006_agent_traces_ttl.sql
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0004_agent_traces_ttl.sql
@@ -1 +1 @@
-ALTER TABLE {database}.agent_traces_by_key MODIFY TTL toDateTime(StartTs) + INTERVAL {trace_retention_days} DAY
+ALTER TABLE {database}.agent_traces_by_key MODIFY TTL toDateTime(StartTs) + INTERVAL {retention_days} DAY
diff --git a/litellm-rust/crates/traces/migrations/0003_agent_traces_mv.sql b/litellm-rust/crates/traces-clickhouse/migrations/0005_agent_traces_mv.sql
similarity index 100%
rename from litellm-rust/crates/traces/migrations/0003_agent_traces_mv.sql
rename to litellm-rust/crates/traces-clickhouse/migrations/0005_agent_traces_mv.sql
diff --git a/litellm-rust/crates/traces/migrations/0004_spend_logs.sql b/litellm-rust/crates/traces-clickhouse/migrations/0006_spend_logs.sql
similarity index 94%
rename from litellm-rust/crates/traces/migrations/0004_spend_logs.sql
rename to litellm-rust/crates/traces-clickhouse/migrations/0006_spend_logs.sql
index a14930f438f..44f7959b2bf 100644
--- a/litellm-rust/crates/traces/migrations/0004_spend_logs.sql
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0006_spend_logs.sql
@@ -34,9 +34,11 @@ CREATE TABLE IF NOT EXISTS {database}.spend_logs
metadata String CODEC(ZSTD(3)),
messages String CODEC(ZSTD(3)),
response String CODEC(ZSTD(3)),
+ EngineReceivedMs UInt64 DEFAULT 0,
INDEX idx_response_id response_id TYPE bloom_filter(0.001) GRANULARITY 1,
INDEX idx_trace_id trace_id TYPE bloom_filter(0.001) GRANULARITY 1
)
ENGINE = ReplacingMergeTree(end_time)
PARTITION BY toYYYYMM(start_time)
ORDER BY (team_id, start_time, request_id)
+SETTINGS materialize_ttl_recalculate_only = 1
diff --git a/litellm-rust/crates/traces/migrations/0007_spend_logs_ttl.sql b/litellm-rust/crates/traces-clickhouse/migrations/0007_spend_logs_ttl.sql
similarity index 58%
rename from litellm-rust/crates/traces/migrations/0007_spend_logs_ttl.sql
rename to litellm-rust/crates/traces-clickhouse/migrations/0007_spend_logs_ttl.sql
index 131573927ac..d9b1a2403b4 100644
--- a/litellm-rust/crates/traces/migrations/0007_spend_logs_ttl.sql
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0007_spend_logs_ttl.sql
@@ -1 +1 @@
-ALTER TABLE {database}.spend_logs MODIFY TTL toDateTime(start_time) + INTERVAL {spend_log_retention_days} DAY
+ALTER TABLE {database}.spend_logs MODIFY TTL toDateTime(start_time) + INTERVAL {retention_days} DAY
diff --git a/litellm-rust/crates/traces-clickhouse/migrations/0008_trace_user.sql b/litellm-rust/crates/traces-clickhouse/migrations/0008_trace_user.sql
new file mode 100644
index 00000000000..845c93aea21
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0008_trace_user.sql
@@ -0,0 +1,2 @@
+ALTER TABLE {database}.otel_traces
+ ADD COLUMN IF NOT EXISTS UserId String DEFAULT ''
diff --git a/litellm-rust/crates/traces-clickhouse/migrations/0009_trace_rollup_ownership.sql b/litellm-rust/crates/traces-clickhouse/migrations/0009_trace_rollup_ownership.sql
new file mode 100644
index 00000000000..fd696349cf5
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0009_trace_rollup_ownership.sql
@@ -0,0 +1,3 @@
+ALTER TABLE {database}.agent_traces_by_key
+ ADD COLUMN IF NOT EXISTS UserIds SimpleAggregateFunction(groupUniqArrayArray, Array(String)) DEFAULT [],
+ ADD COLUMN IF NOT EXISTS IdentifiedLlmCount SimpleAggregateFunction(sum, UInt64) DEFAULT 0
diff --git a/litellm-rust/crates/traces-clickhouse/migrations/0010_trace_cost_completeness.sql b/litellm-rust/crates/traces-clickhouse/migrations/0010_trace_cost_completeness.sql
new file mode 100644
index 00000000000..af87a6bf40b
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0010_trace_cost_completeness.sql
@@ -0,0 +1,22 @@
+ALTER TABLE {database}.agent_traces_by_key_mv MODIFY QUERY
+SELECT
+ TeamId, ApiKeyHash, TraceId, groupUniqArray(UserId) AS UserIds,
+ min(Timestamp) AS StartTs,
+ max(Timestamp + toIntervalNanosecond(Duration)) AS EndTs,
+ any(ServiceName) AS ServiceName,
+ anyLastIf(toNullable(SpanName), ParentSpanId = '') AS RootName,
+ anyLastIf(toNullable(InputPreview), ParentSpanId = '') AS RootInput,
+ anyLastIf(toNullable(StatusCode), ParentSpanId = '') AS RootStatus,
+ count() AS SpanCount,
+ countIf(ObservationType = 'agent') AS AgentCount,
+ countIf(ObservationType = 'llm') AS LlmCount,
+ countIf(ObservationType = 'llm' AND LiteLLMRequestId != '') AS IdentifiedLlmCount,
+ countIf(ObservationType = 'tool') AS ToolCount,
+ countIf(StatusCode = 'STATUS_CODE_ERROR') AS ErrorCount,
+ sum(InputTokens) AS InputTokens,
+ sum(OutputTokens) AS OutputTokens,
+ groupUniqArrayIf(toString(Model), Model != '') AS Models,
+ groupUniqArrayIf(SpanName, ObservationType = 'agent') AS AgentNames,
+ groupArrayIf(LiteLLMRequestId, ObservationType = 'llm' OR LiteLLMRequestId != '') AS RequestIds
+FROM {database}.otel_traces
+GROUP BY TeamId, ApiKeyHash, TraceId
diff --git a/litellm-rust/crates/traces-clickhouse/migrations/0011_otel_traces_framework.sql b/litellm-rust/crates/traces-clickhouse/migrations/0011_otel_traces_framework.sql
new file mode 100644
index 00000000000..1d6c2c83769
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0011_otel_traces_framework.sql
@@ -0,0 +1 @@
+ALTER TABLE {database}.otel_traces ADD COLUMN IF NOT EXISTS Framework LowCardinality(String) AFTER AgentName
diff --git a/litellm-rust/crates/traces-clickhouse/migrations/0012_otel_traces_call_evidence.sql b/litellm-rust/crates/traces-clickhouse/migrations/0012_otel_traces_call_evidence.sql
new file mode 100644
index 00000000000..538567d1cdf
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0012_otel_traces_call_evidence.sql
@@ -0,0 +1,5 @@
+ALTER TABLE {database}.otel_traces
+ ADD COLUMN IF NOT EXISTS WrapperCandidate Bool DEFAULT false AFTER ObservationType,
+ ADD COLUMN IF NOT EXISTS CallKeys Array(String) DEFAULT [] AFTER LiteLLMRequestId,
+ ADD COLUMN IF NOT EXISTS CallEvidence LowCardinality(String) DEFAULT '' AFTER CallKeys,
+ ADD COLUMN IF NOT EXISTS ToolCallId String DEFAULT '' AFTER Output
diff --git a/litellm-rust/crates/traces-clickhouse/migrations/0013_otel_traces_agent_metadata.sql b/litellm-rust/crates/traces-clickhouse/migrations/0013_otel_traces_agent_metadata.sql
new file mode 100644
index 00000000000..fc2e4f790df
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0013_otel_traces_agent_metadata.sql
@@ -0,0 +1,2 @@
+ALTER TABLE {database}.otel_traces
+ ADD COLUMN IF NOT EXISTS AgentMetadata String DEFAULT '{}' CODEC(ZSTD(3))
diff --git a/litellm-rust/crates/traces-clickhouse/migrations/0014_spend_unknown_cost.sql b/litellm-rust/crates/traces-clickhouse/migrations/0014_spend_unknown_cost.sql
new file mode 100644
index 00000000000..6b8ac8e414b
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0014_spend_unknown_cost.sql
@@ -0,0 +1 @@
+ALTER TABLE {database}.spend_logs MODIFY COLUMN spend Nullable(Float64) DEFAULT NULL
diff --git a/litellm-rust/crates/traces-clickhouse/migrations/0015_spend_gateway_call_id.sql b/litellm-rust/crates/traces-clickhouse/migrations/0015_spend_gateway_call_id.sql
new file mode 100644
index 00000000000..2febb9e8f24
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/migrations/0015_spend_gateway_call_id.sql
@@ -0,0 +1,6 @@
+ALTER TABLE {database}.spend_logs
+ ADD COLUMN IF NOT EXISTS litellm_call_id String DEFAULT '' AFTER response_id,
+ ADD INDEX IF NOT EXISTS idx_litellm_call_id litellm_call_id
+ TYPE bloom_filter(0.001) GRANULARITY 1,
+ ADD INDEX IF NOT EXISTS idx_request_id request_id
+ TYPE bloom_filter(0.001) GRANULARITY 1
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/correlated_calls.sql b/litellm-rust/crates/traces-clickhouse/query/help/correlated_calls.sql
new file mode 100644
index 00000000000..c4a6932c39b
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/correlated_calls.sql
@@ -0,0 +1,14 @@
+SELECT t.TraceId, t.SpanId, s.request_id, s.spend, s.metadata
+FROM otel_traces AS t
+INNER JOIN (
+ SELECT *
+ FROM spend_logs FINAL
+ WHERE start_time >= now() - INTERVAL 1 DAY
+) AS s
+ ON t.LiteLLMRequestId = s.response_id
+ AND t.TeamId = s.team_id
+ AND ((t.UserId != '' AND t.UserId = s.user)
+ OR (t.ApiKeyHash != '' AND t.ApiKeyHash = s.api_key))
+WHERE t.Timestamp >= now() - INTERVAL 1 DAY
+ AND t.LiteLLMRequestId != ''
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/custom_metadata.sql b/litellm-rust/crates/traces-clickhouse/query/help/custom_metadata.sql
new file mode 100644
index 00000000000..6b6ff531349
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/custom_metadata.sql
@@ -0,0 +1,8 @@
+SELECT
+ request_id, response_id, model, spend, JSONExtractString(metadata, 'project') AS project
+FROM spend_logs FINAL
+WHERE start_time >= now() - INTERVAL 1 DAY
+ AND JSONHas(metadata, 'project')
+ AND JSONExtractString(metadata, 'project') = 'example'
+ORDER BY start_time DESC
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/discover_keys.sql b/litellm-rust/crates/traces-clickhouse/query/help/discover_keys.sql
new file mode 100644
index 00000000000..4e09539adb5
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/discover_keys.sql
@@ -0,0 +1,6 @@
+SELECT
+ DISTINCT arrayJoin(JSONExtractKeys(metadata)) AS key
+FROM spend_logs FINAL
+WHERE start_time >= now() - INTERVAL 30 DAY
+ORDER BY key
+LIMIT 200
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/failed_spans.sql b/litellm-rust/crates/traces-clickhouse/query/help/failed_spans.sql
new file mode 100644
index 00000000000..b0f3cc413c1
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/failed_spans.sql
@@ -0,0 +1,7 @@
+SELECT TeamId AS team, ApiKeyHash AS api_key, TraceId AS trace_id,
+ SpanId AS span_id, StatusMessage AS message
+FROM otel_traces
+WHERE Timestamp >= now() - INTERVAL 1 DAY
+ AND StatusCode = 'STATUS_CODE_ERROR'
+ORDER BY Timestamp DESC, team, api_key, trace_id, span_id
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/metadata_filter.sql b/litellm-rust/crates/traces-clickhouse/query/help/metadata_filter.sql
new file mode 100644
index 00000000000..d4106586a6a
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/metadata_filter.sql
@@ -0,0 +1,8 @@
+SELECT team_id AS team, api_key, request_id, spend,
+ JSONExtractString(metadata, 'labels', 'priority') AS priority
+FROM spend_logs FINAL
+WHERE start_time >= now() - INTERVAL 1 DAY
+ AND JSONHas(metadata, 'labels', 'priority')
+ AND JSONExtractString(metadata, 'labels', 'priority') = 'high'
+ORDER BY team, api_key, request_id
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/model_spend.sql b/litellm-rust/crates/traces-clickhouse/query/help/model_spend.sql
new file mode 100644
index 00000000000..e8911480b22
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/model_spend.sql
@@ -0,0 +1,17 @@
+SELECT
+ team_id, model, requests, unknown_cost_requests,
+ if(unknown_cost_requests = 0, recorded_spend, NULL) AS spend,
+ input_tokens, output_tokens
+FROM (
+ SELECT
+ team_id, model, count() AS requests,
+ countIf(isNull(spend) OR NOT isFinite(spend)) AS unknown_cost_requests,
+ sum(spend) AS recorded_spend,
+ sum(prompt_tokens) AS input_tokens,
+ sum(completion_tokens) AS output_tokens
+ FROM spend_logs FINAL
+ WHERE start_time >= now() - INTERVAL 1 DAY
+ GROUP BY team_id, model
+)
+ORDER BY team_id, model
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/nested_metadata.sql b/litellm-rust/crates/traces-clickhouse/query/help/nested_metadata.sql
new file mode 100644
index 00000000000..cccec2177a0
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/nested_metadata.sql
@@ -0,0 +1,8 @@
+SELECT
+ request_id,
+ JSONType(metadata, 'labels', 'priority') AS type,
+ JSONExtractRaw(metadata, 'labels', 'priority') AS value
+FROM spend_logs FINAL
+WHERE start_time >= now() - INTERVAL 1 DAY
+ AND JSONHas(metadata, 'labels', 'priority')
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/recent_spans.sql b/litellm-rust/crates/traces-clickhouse/query/help/recent_spans.sql
new file mode 100644
index 00000000000..c1e88560110
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/recent_spans.sql
@@ -0,0 +1,8 @@
+SELECT
+ TraceId, SpanId, Model, InputTokens, OutputTokens,
+ Duration / 1000000 AS duration_ms
+FROM otel_traces
+WHERE Timestamp >= now() - INTERVAL 1 DAY
+ AND ObservationType = 'llm'
+ORDER BY Timestamp DESC
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/recent_spend.sql b/litellm-rust/crates/traces-clickhouse/query/help/recent_spend.sql
new file mode 100644
index 00000000000..809b1bd44f0
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/recent_spend.sql
@@ -0,0 +1,7 @@
+SELECT
+ request_id, response_id, trace_id, span_id, model, spend,
+ prompt_tokens, completion_tokens, status
+FROM spend_logs FINAL
+WHERE start_time >= now() - INTERVAL 1 DAY
+ORDER BY start_time DESC, request_id
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/trace_spend.sql b/litellm-rust/crates/traces-clickhouse/query/help/trace_spend.sql
new file mode 100644
index 00000000000..3f0ec16186d
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/trace_spend.sql
@@ -0,0 +1,10 @@
+SELECT
+ team_id, api_key, trace_id, count() AS requests,
+ countIf(isNull(spend) OR NOT isFinite(spend)) AS unknown_cost_requests,
+ if(unknown_cost_requests = 0, sum(spend), NULL) AS recorded_spend
+FROM spend_logs FINAL
+WHERE start_time >= now() - INTERVAL 1 DAY
+ AND trace_id != ''
+GROUP BY team_id, api_key, trace_id
+ORDER BY team_id, api_key, trace_id
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/trace_summary.sql b/litellm-rust/crates/traces-clickhouse/query/help/trace_summary.sql
new file mode 100644
index 00000000000..7a5dcaf10ff
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/trace_summary.sql
@@ -0,0 +1,12 @@
+SELECT TeamId AS team, ApiKeyHash AS api_key, TraceId AS trace_id,
+ ifNull(any(RootName), '') AS name,
+ toUInt32(sum(SpanCount)) AS spans,
+ toUInt32(sum(LlmCount)) AS llm_calls,
+ toUInt32(sum(ErrorCount)) AS errors,
+ toUInt32(sum(InputTokens)) AS input_tokens,
+ toUInt32(sum(OutputTokens)) AS output_tokens
+FROM agent_traces_by_key
+GROUP BY TeamId, ApiKeyHash, TraceId
+HAVING min(StartTs) >= now() - INTERVAL 1 DAY
+ORDER BY team, api_key, trace_id
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/help/unmatched_spans.sql b/litellm-rust/crates/traces-clickhouse/query/help/unmatched_spans.sql
new file mode 100644
index 00000000000..d5ad0fbb87e
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/help/unmatched_spans.sql
@@ -0,0 +1,18 @@
+SELECT
+ t.TraceId, t.SpanId, t.Model, t.LiteLLMRequestId,
+ t.InputTokens, t.OutputTokens
+FROM otel_traces AS t
+LEFT ANTI JOIN (
+ SELECT *
+ FROM spend_logs FINAL
+ WHERE start_time >= now() - INTERVAL 1 DAY
+) AS s
+ ON t.TeamId = s.team_id
+ AND ((t.UserId != '' AND t.UserId = s.user)
+ OR (t.ApiKeyHash != '' AND t.ApiKeyHash = s.api_key))
+ AND t.LiteLLMRequestId != ''
+ AND (t.LiteLLMRequestId = s.response_id OR t.LiteLLMRequestId = s.request_id)
+WHERE t.Timestamp >= now() - INTERVAL 1 DAY
+ AND t.ObservationType = 'llm'
+ORDER BY t.Timestamp DESC, t.SpanId
+LIMIT 100
diff --git a/litellm-rust/crates/traces-clickhouse/query/lens_agents.sql b/litellm-rust/crates/traces-clickhouse/query/lens_agents.sql
new file mode 100644
index 00000000000..fbdd578f8e7
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/lens_agents.sql
@@ -0,0 +1,6 @@
+SELECT DISTINCT AgentName AS agent_name
+FROM otel_traces
+WHERE AgentName != ''
+ AND ({all_teams:UInt8}=1 OR TeamId={team:String})
+ AND ({key_hash:String}='' OR ApiKeyHash={key_hash:String})
+ORDER BY agent_name
diff --git a/litellm-rust/crates/traces-clickhouse/query/lens_availability.sql b/litellm-rust/crates/traces-clickhouse/query/lens_availability.sql
new file mode 100644
index 00000000000..8d350dd1779
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/lens_availability.sql
@@ -0,0 +1,8 @@
+SELECT
+ EXISTS(SELECT 1 FROM otel_traces
+ WHERE ({all_teams:UInt8}=1 OR TeamId={team:String})
+ AND ({key_hash:String}='' OR ApiKeyHash={key_hash:String})) AS traces,
+ EXISTS(SELECT 1 FROM spend_logs
+ WHERE ({all_teams:UInt8}=1 OR team_id={team:String})
+ AND ({key_hash:String}='' OR api_key={key_hash:String})
+ AND NOT JSONExtractBool(metadata,'litellm_lens_internal')) AS requests
diff --git a/litellm-rust/crates/traces/query/lens_content.sql b/litellm-rust/crates/traces-clickhouse/query/lens_content.sql
similarity index 100%
rename from litellm-rust/crates/traces/query/lens_content.sql
rename to litellm-rust/crates/traces-clickhouse/query/lens_content.sql
diff --git a/litellm-rust/crates/traces/query/lens_evidence.sql b/litellm-rust/crates/traces-clickhouse/query/lens_evidence.sql
similarity index 100%
rename from litellm-rust/crates/traces/query/lens_evidence.sql
rename to litellm-rust/crates/traces-clickhouse/query/lens_evidence.sql
diff --git a/litellm-rust/crates/traces/query/lens_sample.sql b/litellm-rust/crates/traces-clickhouse/query/lens_sample.sql
similarity index 97%
rename from litellm-rust/crates/traces/query/lens_sample.sql
rename to litellm-rust/crates/traces-clickhouse/query/lens_sample.sql
index 1fc9c964a6f..92086c33c13 100644
--- a/litellm-rust/crates/traces/query/lens_sample.sql
+++ b/litellm-rust/crates/traces-clickhouse/query/lens_sample.sql
@@ -28,6 +28,7 @@ SELECT *, selection_key FROM (
GROUP BY TeamId,ApiKeyHash,TraceId
HAVING max(EngineReceivedMs) < {end:UInt64}
AND max(toUnixTimestamp64Milli(Timestamp)+toInt64(intDiv(Duration,1000000))) < {end:UInt64}
+ AND ({agent_name:String}='' OR countIf(AgentName={agent_name:String}) > 0)
AND countIf(arrayAll((k,v) -> ResourceAttributes[k]=v OR SpanAttributes[k]=v,
{filter_keys:Array(String)},{filter_values:Array(String)})
AND ({service:String}='' OR ServiceName={service:String})) > 0
@@ -48,6 +49,7 @@ SELECT *, selection_key FROM (
OR JSONExtractString(metadata,'requester_metadata',k)=v OR (k='tag' AND has(request_tags,v)),
{filter_keys:Array(String)},{filter_values:Array(String)})
AND ({service:String}='' OR model_group={service:String})
+ AND {agent_name:String}=''
AND NOT JSONExtractBool(metadata,'litellm_lens_internal')
AND ({source:String}!='both' OR (team_id,api_key,response_id) NOT IN (
SELECT TeamId,ApiKeyHash,LiteLLMRequestId FROM otel_traces
diff --git a/litellm-rust/crates/traces-clickhouse/query/list_traces.sql b/litellm-rust/crates/traces-clickhouse/query/list_traces.sql
new file mode 100644
index 00000000000..c52adf7ef49
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/list_traces.sql
@@ -0,0 +1,48 @@
+WITH page AS (
+SELECT TraceId AS trace_id,
+ hex(SHA256(concat(TeamId, char(0), ApiKeyHash, char(0), TraceId))) AS trace_ref,
+ if(length(groupUniqArrayArray(UserIds)) = 1, arrayElement(groupUniqArrayArray(UserIds), 1), '') AS user_id, TeamId AS team_id, ApiKeyHash AS api_key_hash,
+ ifNull(any(RootName), '') AS name, any(ServiceName) AS service,
+ ifNull(any(RootInput), '') AS input_preview, ifNull(any(RootStatus), '') AS status,
+ toUnixTimestamp64Milli(min(StartTs)) AS start_ms,
+ min(StartTs) AS trace_start, max(EndTs) AS trace_end,
+ dateDiff('millisecond', min(StartTs), max(EndTs)) AS duration_ms,
+ sum(SpanCount) AS span_count,
+ sum(AgentCount) AS agent_invocations,
+ sum(LlmCount) AS llm_calls, sum(ToolCount) AS tool_calls,
+ sum(InputTokens) AS input_tokens, sum(OutputTokens) AS output_tokens,
+ groupUniqArrayArray(Models) AS models, sum(ErrorCount) AS error_count,
+ arrayDistinct(if(sum(IdentifiedLlmCount) != sum(LlmCount),
+ arrayConcat(groupArrayArray(RequestIds), ['']),
+ groupArrayArray(RequestIds))) AS request_ids
+FROM agent_traces_by_key
+WHERE ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND UserIds = [{user_id:String}])
+ OR has({team_ids:Array(String)}, TeamId))
+GROUP BY TeamId, ApiKeyHash, TraceId
+HAVING min(StartTs) >= fromUnixTimestamp64Milli({start_ms:Int64})
+ AND min(StartTs) < fromUnixTimestamp64Milli({end_ms:Int64})
+ AND ({cursor_ms:Int64} = 0 OR (toUnixTimestamp64Milli(min(StartTs)), trace_ref)
+ < ({cursor_ms:Int64}, {cursor_trace_id:String}))
+ORDER BY start_ms DESC, trace_ref DESC
+LIMIT {limit:UInt32}
+)
+SELECT page.* EXCEPT (trace_start, trace_end),
+ identities.agent_names AS agent_names, identities.agent_count AS agent_count,
+ identities.frameworks AS frameworks
+FROM page
+LEFT JOIN (
+ SELECT TeamId, ApiKeyHash, TraceId,
+ arraySort(groupUniqArrayIf(AgentName, AgentName != '')) AS agent_names,
+ arraySort(groupUniqArrayIf(toString(Framework), Framework != '')) AS frameworks,
+ uniqExactIf(if(AgentName = '', SpanName, AgentName), ObservationType = 'agent') AS agent_count
+ FROM otel_traces
+ WHERE Timestamp >= (SELECT min(trace_start) FROM page)
+ AND Timestamp <= (SELECT max(trace_end) FROM page)
+ AND TraceId IN (SELECT trace_id FROM page)
+ AND (TeamId, ApiKeyHash, TraceId) IN (SELECT team_id, api_key_hash, trace_id FROM page)
+ GROUP BY TeamId, ApiKeyHash, TraceId
+) AS identities
+ON page.team_id = identities.TeamId AND page.api_key_hash = identities.ApiKeyHash
+ AND page.trace_id = identities.TraceId
+ORDER BY page.start_ms DESC, page.trace_ref DESC
diff --git a/litellm-rust/crates/traces-clickhouse/query/span_detail.sql b/litellm-rust/crates/traces-clickhouse/query/span_detail.sql
new file mode 100644
index 00000000000..7db742ea3ee
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/span_detail.sql
@@ -0,0 +1,24 @@
+SELECT o.SpanId AS span_id, o.Input AS input,
+ if(o.Output = '' AND o.ObservationType = 'agent', answer.output, o.Output) AS output,
+ o.SpanAttributes AS attributes
+FROM otel_traces AS o
+LEFT JOIN (
+ SELECT TeamId, ApiKeyHash, ParentSpanId AS parent_span_id, argMax(Output, Timestamp) AS output
+ FROM otel_traces
+ WHERE TraceId = {trace_id:String} AND ParentSpanId = {span_id:String}
+ AND ObservationType = 'llm' AND Output != ''
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND UserId = {user_id:String})
+ OR has({team_ids:Array(String)}, TeamId))
+ AND ({trace_ref:String} = '' OR
+ hex(SHA256(concat(TeamId, char(0), ApiKeyHash, char(0), TraceId))) = {trace_ref:String})
+ GROUP BY TeamId, ApiKeyHash, ParentSpanId
+) AS answer ON answer.parent_span_id = o.SpanId
+ AND answer.TeamId = o.TeamId AND answer.ApiKeyHash = o.ApiKeyHash
+WHERE o.TraceId = {trace_id:String} AND o.SpanId = {span_id:String}
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND o.UserId = {user_id:String})
+ OR has({team_ids:Array(String)}, o.TeamId))
+ AND ({trace_ref:String} = '' OR
+ hex(SHA256(concat(o.TeamId, char(0), o.ApiKeyHash, char(0), o.TraceId))) = {trace_ref:String})
+LIMIT 1
diff --git a/litellm-rust/crates/traces-clickhouse/query/span_error.sql b/litellm-rust/crates/traces-clickhouse/query/span_error.sql
new file mode 100644
index 00000000000..e1226c4d23c
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/span_error.sql
@@ -0,0 +1,14 @@
+SELECT SpanId AS span_id,
+ substringUTF8(StatusMessage, {error_offset:UInt64} + 1, 16384) AS message,
+ lengthUTF8(StatusMessage) AS total_chars,
+ hex(SHA256(StatusMessage)) AS version
+FROM otel_traces
+WHERE TraceId = {trace_id:String} AND SpanId = {span_id:String}
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND UserId = {user_id:String})
+ OR has({team_ids:Array(String)}, TeamId))
+ AND ({trace_ref:String} = '' OR
+ hex(SHA256(concat(TeamId, char(0), ApiKeyHash, char(0), TraceId))) = {trace_ref:String})
+ AND ({error_version:String} = '' OR hex(SHA256(StatusMessage)) = {error_version:String})
+ORDER BY Timestamp, EngineReceivedMs, StatusMessage
+LIMIT 1
diff --git a/litellm-rust/crates/traces-clickhouse/query/spend_batch.sql b/litellm-rust/crates/traces-clickhouse/query/spend_batch.sql
new file mode 100644
index 00000000000..3918286a61f
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/spend_batch.sql
@@ -0,0 +1,28 @@
+SELECT * FROM (
+SELECT request_id, litellm_call_id, response_id, upstream_response_id, trace_id, span_id, team_id, api_key, user, spend,
+ toUnixTimestamp64Milli(start_time) AS start_ms
+FROM (
+ SELECT *,
+ -- A chat request served through the Responses API returns the upstream `resp_` id to the
+ -- client but logs LiteLLM's managed `resp_` id, which embeds it.
+ if(startsWith(response_id, 'resp_'),
+ extract(tryBase64Decode(substring(response_id, 6)), 'response_id:([^;]+)'),
+ '') AS upstream_response_id
+ FROM spend_logs FINAL
+ WHERE start_time >= fromUnixTimestamp64Milli({start_ms:Int64})
+ AND start_time < fromUnixTimestamp64Milli({end_ms:Int64})
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND user = {user_id:String})
+ OR has({team_ids:Array(String)}, team_id))
+)
+WHERE response_id IN {response_ids:Array(String)}
+ OR upstream_response_id IN {response_ids:Array(String)}
+ OR litellm_call_id IN {request_ids:Array(String)}
+ OR (litellm_call_id = '' AND request_id IN {request_ids:Array(String)})
+ OR (trace_id != '' AND trace_id IN {trace_ids:Array(String)})
+ORDER BY start_time DESC
+)
+WHERE {has_cursor:UInt8} = 0
+ OR (team_id, start_ms, request_id) > ({after_team:String}, {after_ms:Int64}, {after_id:String})
+ORDER BY team_id, start_ms, request_id
+LIMIT {page_size:UInt32}
diff --git a/litellm-rust/crates/traces-clickhouse/query/spend_by_response_ids.sql b/litellm-rust/crates/traces-clickhouse/query/spend_by_response_ids.sql
new file mode 100644
index 00000000000..da64dafbc39
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/spend_by_response_ids.sql
@@ -0,0 +1,22 @@
+SELECT request_id, litellm_call_id, response_id, upstream_response_id, trace_id, span_id, team_id, api_key, user, spend,
+ toUnixTimestamp64Milli(start_time) AS start_ms
+FROM (
+ SELECT *,
+ -- A chat request served through the Responses API returns the upstream `resp_` id to the
+ -- client but logs LiteLLM's managed `resp_` id, which embeds it.
+ if(startsWith(response_id, 'resp_'),
+ extract(tryBase64Decode(substring(response_id, 6)), 'response_id:([^;]+)'),
+ '') AS upstream_response_id
+ FROM spend_logs FINAL
+ WHERE start_time >= fromUnixTimestamp64Milli({start_ms:Int64})
+ AND start_time < fromUnixTimestamp64Milli({end_ms:Int64})
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND user = {user_id:String})
+ OR has({team_ids:Array(String)}, team_id))
+)
+WHERE response_id IN {response_ids:Array(String)}
+ OR upstream_response_id IN {response_ids:Array(String)}
+ OR litellm_call_id IN {request_ids:Array(String)}
+ OR (litellm_call_id = '' AND request_id IN {request_ids:Array(String)})
+ OR (trace_id != '' AND trace_id IN {trace_ids:Array(String)})
+ORDER BY start_time DESC
diff --git a/litellm-rust/crates/traces-clickhouse/query/trace_identity.sql b/litellm-rust/crates/traces-clickhouse/query/trace_identity.sql
new file mode 100644
index 00000000000..e3881b150b7
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/trace_identity.sql
@@ -0,0 +1,8 @@
+SELECT hex(SHA256(concat(TeamId, char(0), ApiKeyHash, char(0), TraceId))) AS trace_ref
+FROM otel_traces
+WHERE TraceId = {trace_id:String}
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND UserId = {user_id:String})
+ OR has({team_ids:Array(String)}, TeamId))
+GROUP BY TeamId, ApiKeyHash, TraceId
+LIMIT 2
diff --git a/litellm-rust/crates/traces-clickhouse/query/trace_list_span_batch.sql b/litellm-rust/crates/traces-clickhouse/query/trace_list_span_batch.sql
new file mode 100644
index 00000000000..679edfdec2e
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/trace_list_span_batch.sql
@@ -0,0 +1,31 @@
+SELECT * FROM (
+SELECT o.TraceId AS trace_id, o.SpanId AS span_id, o.ParentSpanId AS parent_span_id, o.SpanName AS name,
+ o.ObservationType AS type, toUInt8(o.WrapperCandidate) AS wrapper_candidate, o.AgentName AS agent,
+ o.Framework AS framework, o.StatusCode AS status,
+ substringUTF8(o.StatusMessage, 1, 128) AS status_message,
+ lengthUTF8(o.StatusMessage) > 128 AS error_truncated,
+ toUnixTimestamp64Nano(o.Timestamp) AS start_ns, o.Duration AS duration_ns,
+ o.ServiceName AS service, o.InputPreview AS input_preview, o.Model AS model,
+ o.InputTokens AS input_tokens, o.OutputTokens AS output_tokens,
+ o.LiteLLMRequestId AS litellm_request_id,
+ o.CallKeys AS call_keys, o.CallEvidence AS call_evidence,
+ -- Rows written before ToolCallId keep the call id only in their attributes.
+ if(o.ToolCallId != '' OR o.ObservationType != 'tool', o.ToolCallId,
+ coalesce(nullIf(o.SpanAttributes['gen_ai.tool.call.id'], ''), nullIf(o.SpanAttributes['tool.id'], ''), ''))
+ AS tool_call_id,
+ o.UserId AS user_id, o.TeamId AS team_id, o.ApiKeyHash AS api_key_hash
+FROM otel_traces AS o
+WHERE o.Timestamp >= fromUnixTimestamp64Milli({start_ms:Int64})
+ AND o.Timestamp < fromUnixTimestamp64Milli({end_ms:Int64})
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND o.UserId = {user_id:String})
+ OR has({team_ids:Array(String)}, o.TeamId))
+ AND hex(SHA256(concat(o.TeamId, char(0), o.ApiKeyHash, char(0), o.TraceId))) IN {trace_refs:Array(String)}
+ AND o.EngineReceivedMs <= {snapshot_ms:UInt64}
+ORDER BY o.Timestamp, o.EngineReceivedMs, o.StatusMessage
+LIMIT 1 BY o.TeamId, o.ApiKeyHash, o.TraceId, o.SpanId
+
+)
+WHERE (team_id, api_key_hash, trace_id, span_id) > ({after_team:String}, {after_key:String}, {after_trace:String}, {after_span:String})
+ORDER BY team_id, api_key_hash, trace_id, span_id
+LIMIT {page_size:UInt32}
diff --git a/litellm-rust/crates/traces-clickhouse/query/trace_page_spans.sql b/litellm-rust/crates/traces-clickhouse/query/trace_page_spans.sql
new file mode 100644
index 00000000000..b81bba61e7c
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/trace_page_spans.sql
@@ -0,0 +1,24 @@
+SELECT o.TraceId AS trace_id, o.SpanId AS span_id, o.ParentSpanId AS parent_span_id, o.SpanName AS name,
+ o.ObservationType AS type, toUInt8(o.WrapperCandidate) AS wrapper_candidate, o.AgentName AS agent,
+ o.Framework AS framework, o.StatusCode AS status,
+ substringUTF8(o.StatusMessage, 1, 128) AS status_message,
+ lengthUTF8(o.StatusMessage) > 128 AS error_truncated,
+ toUnixTimestamp64Nano(o.Timestamp) AS start_ns, o.Duration AS duration_ns,
+ o.ServiceName AS service, o.InputPreview AS input_preview, o.Model AS model,
+ o.InputTokens AS input_tokens, o.OutputTokens AS output_tokens,
+ o.LiteLLMRequestId AS litellm_request_id,
+ o.CallKeys AS call_keys, o.CallEvidence AS call_evidence,
+ -- Rows written before ToolCallId keep the call id only in their attributes.
+ if(o.ToolCallId != '' OR o.ObservationType != 'tool', o.ToolCallId,
+ coalesce(nullIf(o.SpanAttributes['gen_ai.tool.call.id'], ''), nullIf(o.SpanAttributes['tool.id'], ''), ''))
+ AS tool_call_id,
+ o.UserId AS user_id, o.TeamId AS team_id, o.ApiKeyHash AS api_key_hash
+FROM otel_traces AS o
+WHERE o.Timestamp >= fromUnixTimestamp64Milli({start_ms:Int64})
+ AND o.Timestamp < fromUnixTimestamp64Milli({end_ms:Int64})
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND o.UserId = {user_id:String})
+ OR has({team_ids:Array(String)}, o.TeamId))
+ AND hex(SHA256(concat(o.TeamId, char(0), o.ApiKeyHash, char(0), o.TraceId))) IN {trace_refs:Array(String)}
+ORDER BY o.Timestamp, o.EngineReceivedMs, o.StatusMessage
+LIMIT 1 BY o.TeamId, o.ApiKeyHash, o.TraceId, o.SpanId
diff --git a/litellm-rust/crates/traces-clickhouse/query/trace_span_batch.sql b/litellm-rust/crates/traces-clickhouse/query/trace_span_batch.sql
new file mode 100644
index 00000000000..967ef2fcf1f
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/trace_span_batch.sql
@@ -0,0 +1,30 @@
+SELECT * FROM (
+SELECT o.TraceId AS trace_id, o.SpanId AS span_id, o.ParentSpanId AS parent_span_id, o.SpanName AS name,
+ o.ObservationType AS type, toUInt8(o.WrapperCandidate) AS wrapper_candidate, o.AgentName AS agent,
+ o.Framework AS framework, o.StatusCode AS status,
+ substringUTF8(o.StatusMessage, 1, 128) AS status_message,
+ lengthUTF8(o.StatusMessage) > 128 AS error_truncated,
+ toUnixTimestamp64Nano(o.Timestamp) AS start_ns, o.Duration AS duration_ns,
+ o.ServiceName AS service, o.InputPreview AS input_preview, o.Model AS model,
+ o.InputTokens AS input_tokens, o.OutputTokens AS output_tokens,
+ o.LiteLLMRequestId AS litellm_request_id,
+ o.CallKeys AS call_keys, o.CallEvidence AS call_evidence,
+ -- Rows written before ToolCallId keep the call id only in their attributes.
+ if(o.ToolCallId != '' OR o.ObservationType != 'tool', o.ToolCallId,
+ coalesce(nullIf(o.SpanAttributes['gen_ai.tool.call.id'], ''), nullIf(o.SpanAttributes['tool.id'], ''), ''))
+ AS tool_call_id,
+ o.UserId AS user_id, o.TeamId AS team_id, o.ApiKeyHash AS api_key_hash
+FROM otel_traces AS o
+WHERE o.TraceId = {trace_id:String}
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND o.UserId = {user_id:String})
+ OR has({team_ids:Array(String)}, o.TeamId))
+ AND ({trace_ref:String} = '' OR
+ hex(SHA256(concat(o.TeamId, char(0), o.ApiKeyHash, char(0), o.TraceId))) = {trace_ref:String})
+ AND o.EngineReceivedMs <= {snapshot_ms:UInt64}
+ORDER BY o.Timestamp, o.EngineReceivedMs, o.StatusMessage
+LIMIT 1 BY o.SpanId
+)
+WHERE span_id > {after_span_id:String}
+ORDER BY span_id
+LIMIT {page_size:UInt32}
diff --git a/litellm-rust/crates/traces-clickhouse/query/trace_spans.sql b/litellm-rust/crates/traces-clickhouse/query/trace_spans.sql
new file mode 100644
index 00000000000..2e0ac4f6dfb
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/query/trace_spans.sql
@@ -0,0 +1,24 @@
+SELECT o.TraceId AS trace_id, o.SpanId AS span_id, o.ParentSpanId AS parent_span_id, o.SpanName AS name,
+ o.ObservationType AS type, toUInt8(o.WrapperCandidate) AS wrapper_candidate, o.AgentName AS agent,
+ o.Framework AS framework, o.StatusCode AS status,
+ substringUTF8(o.StatusMessage, 1, 128) AS status_message,
+ lengthUTF8(o.StatusMessage) > 128 AS error_truncated,
+ toUnixTimestamp64Nano(o.Timestamp) AS start_ns, o.Duration AS duration_ns,
+ o.ServiceName AS service, o.InputPreview AS input_preview, o.Model AS model,
+ o.InputTokens AS input_tokens, o.OutputTokens AS output_tokens,
+ o.LiteLLMRequestId AS litellm_request_id,
+ o.CallKeys AS call_keys, o.CallEvidence AS call_evidence,
+ -- Rows written before ToolCallId keep the call id only in their attributes.
+ if(o.ToolCallId != '' OR o.ObservationType != 'tool', o.ToolCallId,
+ coalesce(nullIf(o.SpanAttributes['gen_ai.tool.call.id'], ''), nullIf(o.SpanAttributes['tool.id'], ''), ''))
+ AS tool_call_id,
+ o.UserId AS user_id, o.TeamId AS team_id, o.ApiKeyHash AS api_key_hash
+FROM otel_traces AS o
+WHERE o.TraceId = {trace_id:String}
+ AND ({all_teams:UInt8} = 1
+ OR ({user_id:String} != '' AND o.UserId = {user_id:String})
+ OR has({team_ids:Array(String)}, o.TeamId))
+ AND ({trace_ref:String} = '' OR
+ hex(SHA256(concat(o.TeamId, char(0), o.ApiKeyHash, char(0), o.TraceId))) = {trace_ref:String})
+ORDER BY o.Timestamp, o.EngineReceivedMs, o.StatusMessage
+LIMIT 1 BY o.SpanId
diff --git a/litellm-rust/crates/traces-clickhouse/src/bin/export_schema.rs b/litellm-rust/crates/traces-clickhouse/src/bin/export_schema.rs
new file mode 100644
index 00000000000..23910f73f72
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/bin/export_schema.rs
@@ -0,0 +1,6 @@
+fn main() {
+ println!(
+ "{}",
+ serde_json::to_string_pretty(&litellm_traces_clickhouse::wire_schema::schemas()).unwrap()
+ );
+}
diff --git a/litellm-rust/crates/traces-clickhouse/src/config.rs b/litellm-rust/crates/traces-clickhouse/src/config.rs
new file mode 100644
index 00000000000..dccd1c368f6
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/config.rs
@@ -0,0 +1,38 @@
+use crate::Error;
+use litellm_storage_clickhouse::Storage;
+
+#[derive(Clone)]
+pub struct Config {
+ storage: Storage,
+ retention_days: u32,
+ max_attribute_value_bytes: usize,
+}
+
+impl Config {
+ pub fn new(
+ database: String,
+ url: &str,
+ retention_days: u32,
+ max_attribute_value_bytes: usize,
+ ) -> Result {
+ super::schema_statements(&database, retention_days)?;
+ Ok(Self {
+ storage: Storage::new(database, url)?,
+ retention_days,
+ max_attribute_value_bytes,
+ })
+ }
+
+ pub fn storage(&self) -> &Storage {
+ &self.storage
+ }
+
+ pub fn retention_days(&self) -> u32 {
+ self.retention_days
+ }
+
+ /// Stored span attribute and payload values longer than this are truncated with a marker.
+ pub fn max_attribute_value_bytes(&self) -> usize {
+ self.max_attribute_value_bytes
+ }
+}
diff --git a/litellm-rust/crates/traces-clickhouse/src/error.rs b/litellm-rust/crates/traces-clickhouse/src/error.rs
new file mode 100644
index 00000000000..78df9a2121c
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/error.rs
@@ -0,0 +1,60 @@
+#[derive(Debug, thiserror::Error)]
+pub enum Error {
+ #[error("invalid ClickHouse insert row")]
+ InvalidRow,
+ #[error("{0} must be a positive integer")]
+ InvalidLimit(&'static str),
+ #[error("invalid ClickHouse insert table")]
+ InvalidTable,
+ #[error("database must be a nonempty SQL identifier and retention must be positive")]
+ InvalidSchema,
+ #[error("unknown ClickHouse read query")]
+ InvalidQuery,
+ #[error("invalid ClickHouse query parameters")]
+ InvalidParameters,
+ #[error("ClickHouse returned an invalid or failed JSON query response")]
+ InvalidResponse,
+ #[error("ClickHouse insert exceeds the encoded size limit")]
+ InsertTooLarge,
+ #[error("Trace exceeds the interactive read budget; use a filtered trace query")]
+ ReadTooLarge,
+ #[error("ClickHouse schema setup failed with HTTP status {0}")]
+ SchemaFailed(u16),
+ #[error("ClickHouse schema setup transport failed")]
+ SchemaTransport,
+ #[error("trace SQL queries require a configured proxy master key")]
+ MissingSecret,
+ #[error("invalid trace query scope")]
+ InvalidScope,
+ #[error("trace SQL query concurrency limit exceeded")]
+ Busy,
+ #[error(
+ "ClickHouse reader provisioning failed with HTTP status {0}; the configured connection must be allowed to manage users, row policies, and SELECT grants on the trace tables"
+ )]
+ ProvisionFailed(u16),
+ #[error("ClickHouse reader provisioning transport failed")]
+ ProvisionTransport,
+ #[error("Invalid {0} cursor")]
+ InvalidCursor(&'static str),
+ #[error("Multiple traces have this ID; provide trace_ref")]
+ AmbiguousTrace,
+ #[error("Trace changed while paging; refresh the trace to continue")]
+ TraceChanged,
+ #[error(transparent)]
+ Decode(#[from] litellm_traces::Error),
+ #[error("trace ingestion task failed")]
+ Task,
+ #[error(transparent)]
+ Storage(#[from] litellm_storage_clickhouse::Error),
+ #[error(transparent)]
+ Cached(#[from] std::sync::Arc),
+}
+
+impl From for Error {
+ fn from(error: litellm_traces_cache::Error) -> Self {
+ match error {
+ litellm_traces_cache::Error::Serialization(_) => Self::InvalidResponse,
+ litellm_traces_cache::Error::ReadTooLarge => Self::ReadTooLarge,
+ }
+ }
+}
diff --git a/litellm-rust/crates/traces-clickhouse/src/insert.rs b/litellm-rust/crates/traces-clickhouse/src/insert.rs
new file mode 100644
index 00000000000..9e6452844f4
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/insert.rs
@@ -0,0 +1,322 @@
+use std::{
+ borrow::Cow,
+ collections::BTreeMap,
+ io::{BufWriter, Write},
+};
+
+use serde::{Serialize, Serializer, ser::SerializeMap};
+
+use flate2::{Compression, write::GzEncoder};
+use litellm_http::Client;
+use serde_json::Value;
+use sha2::{Digest, Sha256};
+use time::{OffsetDateTime, format_description::well_known::Rfc3339};
+
+use super::{Connection, Error};
+use litellm_traces::Shared;
+
+fn max_insert_bytes() -> Result {
+ let name = "CLICKHOUSE_TRACE_MAX_INSERT_BYTES";
+ match std::env::var(name) {
+ Ok(value) => value
+ .parse::()
+ .ok()
+ .filter(|value| *value > 0)
+ .ok_or(Error::InvalidLimit(name)),
+ Err(std::env::VarError::NotPresent) => Ok(64 * 1024 * 1024),
+ Err(_) => Err(Error::InvalidLimit(name)),
+ }
+}
+
+pub type InsertRow = BTreeMap>;
+
+pub enum InsertTable {
+ OtelTraces,
+ SpendLogs,
+}
+
+impl InsertTable {
+ pub fn parse(value: &str) -> Result {
+ match value {
+ "otel_traces" => Ok(Self::OtelTraces),
+ "spend_logs" => Ok(Self::SpendLogs),
+ _ => Err(Error::InvalidTable),
+ }
+ }
+
+ fn name(&self) -> &'static str {
+ match self {
+ Self::OtelTraces => "otel_traces",
+ Self::SpendLogs => "spend_logs",
+ }
+ }
+}
+
+pub async fn insert_rows(
+ client: &Client,
+ connection: &Connection,
+ database: &str,
+ table: InsertTable,
+ rows: Vec>,
+) -> Result<(), Error> {
+ insert_shared_rows(client, connection, database, table, shared_rows(rows)).await
+}
+
+pub async fn insert_shared_rows(
+ client: &Client,
+ connection: &Connection,
+ database: &str,
+ table: InsertTable,
+ rows: Vec,
+) -> Result<(), Error> {
+ if rows.is_empty() {
+ return Ok(());
+ }
+ let received_ms = (OffsetDateTime::now_utc().unix_timestamp_nanos() / 1_000_000) as u64;
+ let (token, body) = prepare_insert(&rows, received_ms, max_insert_bytes()?)?;
+ litellm_storage_clickhouse::insert_compressed_rows(
+ client,
+ connection,
+ database,
+ table.name(),
+ &token,
+ body,
+ )
+ .await
+ .map_err(Error::from)
+}
+
+fn shared_rows(rows: Vec>) -> Vec {
+ rows.into_iter()
+ .map(|row| {
+ row.into_iter()
+ .map(|(key, value)| (key, Shared::new(value)))
+ .collect()
+ })
+ .collect()
+}
+
+pub fn encode_rows(rows: Vec>) -> Result {
+ let body = write_rows(&shared_rows(rows), None, Vec::new(), usize::MAX)?;
+ String::from_utf8(body).map_err(|_| Error::InvalidRow)
+}
+
+fn prepare_insert(
+ rows: &[InsertRow],
+ received_ms: u64,
+ limit: usize,
+) -> Result<(String, Vec), Error> {
+ let hash = write_rows(rows, None, HashWriter(Sha256::new()), limit)?;
+ let token = format!("{:x}", hash.0.finalize());
+ let encoder = write_rows(
+ rows,
+ Some(received_ms),
+ BufWriter::new(GzEncoder::new(Vec::new(), Compression::default())),
+ limit,
+ )?;
+ let body = encoder
+ .into_inner()
+ .map_err(|_| Error::InvalidRow)?
+ .finish()
+ .map_err(|_| Error::InvalidRow)?;
+ Ok((token, body))
+}
+
+struct HashWriter(Sha256);
+
+impl Write for HashWriter {
+ fn write(&mut self, bytes: &[u8]) -> std::io::Result {
+ self.0.update(bytes);
+ Ok(bytes.len())
+ }
+
+ fn flush(&mut self) -> std::io::Result<()> {
+ Ok(())
+ }
+}
+
+struct LimitedWriter {
+ inner: W,
+ remaining: usize,
+ exceeded: bool,
+}
+
+impl Write for LimitedWriter {
+ fn write(&mut self, bytes: &[u8]) -> std::io::Result {
+ if bytes.len() > self.remaining {
+ self.exceeded = true;
+ return Err(std::io::Error::other(Error::InsertTooLarge));
+ }
+ let written = self.inner.write(bytes)?;
+ self.remaining -= written;
+ Ok(written)
+ }
+
+ fn flush(&mut self) -> std::io::Result<()> {
+ self.inner.flush()
+ }
+}
+
+fn write_rows(
+ rows: &[InsertRow],
+ received_ms: Option,
+ writer: W,
+ limit: usize,
+) -> Result {
+ let mut writer = LimitedWriter {
+ inner: writer,
+ remaining: limit,
+ exceeded: false,
+ };
+ for (index, row) in rows.iter().enumerate() {
+ let result = (|| {
+ if index != 0 {
+ writer.write_all(b"\n").map_err(serde_json::Error::io)?;
+ }
+ serde_json::to_writer(&mut writer, &EncodedRow { row, received_ms })
+ })();
+ if result.is_err() {
+ return Err(if writer.exceeded {
+ Error::InsertTooLarge
+ } else {
+ Error::InvalidRow
+ });
+ }
+ }
+ Ok(writer.inner)
+}
+
+struct EncodedRow<'a> {
+ row: &'a InsertRow,
+ received_ms: Option,
+}
+
+impl Serialize for EncodedRow<'_> {
+ fn serialize(&self, serializer: S) -> Result {
+ let mut map = serializer.serialize_map(None)?;
+ let mut received_ms = self.received_ms;
+ for (name, value) in self.row {
+ if name.as_str() >= "EngineReceivedMs"
+ && let Some(timestamp) = received_ms.take()
+ {
+ map.serialize_entry("EngineReceivedMs", ×tamp)?;
+ }
+ if name == "EngineReceivedMs" && self.received_ms.is_some() {
+ continue;
+ }
+ let value = insert_value(name, value).map_err(serde::ser::Error::custom)?;
+ map.serialize_entry(name, &value)?;
+ }
+ if let Some(timestamp) = received_ms {
+ map.serialize_entry("EngineReceivedMs", ×tamp)?;
+ }
+ map.end()
+ }
+}
+
+fn insert_value<'a>(name: &str, value: &'a Value) -> Result, Error> {
+ let multiplier = match name {
+ "Timestamp" => 1,
+ "start_time" | "end_time" | "completion_start_time" => 1_000_000,
+ _ => return Ok(Cow::Borrowed(value)),
+ };
+ if name == "completion_start_time" && value.is_null() {
+ return Ok(Cow::Borrowed(value));
+ }
+ let timestamp = value.as_i64().ok_or(Error::InvalidRow)?;
+ let datetime = OffsetDateTime::from_unix_timestamp_nanos(i128::from(timestamp) * multiplier)
+ .map_err(|_| Error::InvalidRow)?;
+ datetime
+ .format(&Rfc3339)
+ .map(|value| Cow::Owned(Value::String(value)))
+ .map_err(|_| Error::InvalidRow)
+}
+
+#[cfg(test)]
+mod tests {
+ use std::collections::BTreeMap;
+
+ use rstest::rstest;
+ use serde_json::json;
+
+ use super::Error;
+ use super::{shared_rows, write_rows};
+
+ #[rstest]
+ fn encoded_limit_counts_utf8_bytes_across_rows() {
+ let rows = shared_rows(vec![
+ BTreeMap::from([("Input".to_owned(), json!("雪"))]),
+ BTreeMap::from([("Input".to_owned(), json!("雪"))]),
+ ]);
+ let encoded = write_rows(&rows, None, Vec::new(), usize::MAX).expect("valid rows");
+
+ assert!(write_rows(&rows, None, Vec::new(), encoded.len()).is_ok());
+ assert!(matches!(
+ write_rows(&rows, None, Vec::new(), encoded.len() - 1),
+ Err(Error::InsertTooLarge)
+ ));
+ }
+
+ #[rstest]
+ #[case::absent(None)]
+ #[case::submitted(Some(123))]
+ fn streamed_insert_preserves_token_and_stamps_receive_time(#[case] submitted: Option) {
+ use flate2::read::GzDecoder;
+ use sha2::{Digest, Sha256};
+ use std::io::Read;
+ let mut row = BTreeMap::from([
+ ("ApiKeyHash".into(), json!("key")),
+ ("ResourceAttributes".into(), json!({"message": "雪\n\""})),
+ ("Timestamp".into(), json!(1_234_567_890)),
+ ]);
+ if let Some(value) = submitted {
+ row.insert("EngineReceivedMs".into(), json!(value));
+ }
+ let legacy = match submitted {
+ Some(_) => {
+ "{\"ApiKeyHash\":\"key\",\"EngineReceivedMs\":123,\"ResourceAttributes\":{\"message\":\"雪\\n\\\"\"},\"Timestamp\":\"1970-01-01T00:00:01.23456789Z\"}"
+ }
+ None => {
+ "{\"ApiKeyHash\":\"key\",\"ResourceAttributes\":{\"message\":\"雪\\n\\\"\"},\"Timestamp\":\"1970-01-01T00:00:01.23456789Z\"}"
+ }
+ };
+ let rows = shared_rows(vec![row.clone(), row]);
+ let (token, body) = super::prepare_insert(&rows, 456, 4096).unwrap();
+ assert_eq!(
+ token,
+ format!("{:x}", Sha256::digest(format!("{legacy}\n{legacy}")))
+ );
+ let mut decoded = String::new();
+ GzDecoder::new(body.as_slice())
+ .read_to_string(&mut decoded)
+ .unwrap();
+ let expected = json!({
+ "ApiKeyHash": "key", "EngineReceivedMs": 456,
+ "ResourceAttributes": {"message": "雪\n\""},
+ "Timestamp": "1970-01-01T00:00:01.23456789Z",
+ });
+ assert_eq!(
+ decoded
+ .lines()
+ .map(|line| serde_json::from_str::(line).unwrap())
+ .collect::>(),
+ vec![expected.clone(), expected]
+ );
+ assert_eq!(
+ rows[0]
+ .get("EngineReceivedMs")
+ .map(|value| value.as_u64().unwrap()),
+ submitted
+ );
+ }
+
+ #[rstest]
+ fn stamped_insert_enforces_the_encoded_limit() {
+ let rows = shared_rows(vec![BTreeMap::new()]);
+ assert!(super::prepare_insert(&rows, 1, 22).is_ok());
+ assert!(matches!(
+ super::prepare_insert(&rows, 1, 21),
+ Err(Error::InsertTooLarge)
+ ));
+ }
+}
diff --git a/litellm-rust/crates/traces-clickhouse/src/lib.rs b/litellm-rust/crates/traces-clickhouse/src/lib.rs
new file mode 100644
index 00000000000..d83708d27f1
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/lib.rs
@@ -0,0 +1,40 @@
+macro_rules_attribute::attribute_alias! {
+ #[apply(wire_type)] =
+ #[derive(serde::Serialize, serde::Deserialize)]
+ #[cfg_attr(feature = "schema", derive(schemars::JsonSchema))];
+ #[apply(response_type)] =
+ #[derive(serde::Serialize)]
+ #[cfg_attr(feature = "schema", derive(schemars::JsonSchema))];
+ #[apply(request_type)] =
+ #[derive(serde::Deserialize)]
+ #[cfg_attr(feature = "schema", derive(schemars::JsonSchema))];
+}
+
+mod config;
+mod error;
+mod insert;
+pub mod query;
+mod query_access;
+mod reads;
+mod schema;
+mod span_batches;
+mod span_row;
+mod sql;
+mod table;
+#[cfg(feature = "schema")]
+pub mod wire_schema;
+
+pub use config::Config;
+pub use error::Error;
+pub use insert::{InsertRow, InsertTable, encode_rows, insert_rows, insert_shared_rows};
+pub use litellm_storage_clickhouse::{Connection, Parameter};
+pub use litellm_traces::{QueryScope, ReadQuery};
+pub use query::{QueryHelp, execute_read, query_help, query_sql};
+pub use query_access::QueryReaders;
+pub use reads::{get_span, get_span_error, get_trace, get_trace_page, list_traces};
+pub use schema::{
+ NORMALIZED_FIELD_DEFINITIONS, NormalizedFieldDefinition, ensure_schema, schema_statements,
+};
+pub use span_row::span_rows;
+pub use sql::execute_named_read;
+pub use table::TraceTable;
diff --git a/litellm-rust/crates/traces-clickhouse/src/query.rs b/litellm-rust/crates/traces-clickhouse/src/query.rs
new file mode 100644
index 00000000000..d8f7041386c
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/query.rs
@@ -0,0 +1,597 @@
+use std::collections::{BTreeMap, BTreeSet};
+
+use crate::TraceTable;
+use futures_util::{
+ StreamExt,
+ stream::{self, TryStreamExt},
+};
+use litellm_http::Client;
+use litellm_traces::query::guide::{Example, QueryGuide, Section};
+use serde::{Deserialize, Serialize, Serializer};
+use serde_json::Value;
+use strum::IntoEnumIterator;
+
+use super::{
+ Connection, Error, NORMALIZED_FIELD_DEFINITIONS, NormalizedFieldDefinition, Parameter,
+ query_access::READER_LIMITS,
+};
+
+mod guide;
+pub mod lens;
+pub mod named;
+mod number;
+
+const SAMPLE_ROWS: usize = 200;
+const MAX_FIELDS: usize = 200;
+const MAX_DEPTH: usize = 16;
+const METADATA_SQL: &str = "SELECT metadata FROM spend_logs FINAL \
+ WHERE start_time >= now() - INTERVAL 7 DAY AND length(metadata) <= 8192 \
+ LIMIT 201";
+const METADATA_SCOPE: &str = "Up to 200 unordered rows from the last 7 days, excluding metadata larger than 8192 bytes; up to 200 paths and 16 levels. Missing paths may exist outside this sample. Array indexes are 1-based and describe sampled positions, not a fixed schema";
+const ATTRIBUTE_SCOPE: &str = "Distinct keys from up to 200 unordered spans in the last 7 days; up to 200 keys per map. Missing keys may exist outside this sample";
+
+#[derive(Deserialize)]
+struct Rows {
+ data: Vec,
+}
+
+#[derive(Deserialize)]
+struct MetadataRow {
+ metadata: String,
+}
+
+#[macro_rules_attribute::apply(request_type)]
+struct AttributeRow {
+ key: String,
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[derive(Clone, Debug, Eq, Ord, PartialEq, PartialOrd)]
+#[serde(untagged)]
+enum PathPart {
+ Key(String),
+ Index(usize),
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[derive(Clone, Copy, Debug, Eq, Ord, PartialEq, PartialOrd, strum::Display)]
+#[serde(rename_all = "lowercase")]
+#[strum(serialize_all = "lowercase")]
+#[cfg_attr(feature = "schema", schemars(rename = "MetadataValueType"))]
+enum JsonKind {
+ Array,
+ Boolean,
+ Integer,
+ Null,
+ Number,
+ Object,
+ String,
+}
+
+impl JsonKind {
+ fn of(value: &Value) -> Self {
+ match value {
+ Value::Null => Self::Null,
+ Value::Bool(_) => Self::Boolean,
+ Value::Number(number) if number.is_i64() || number.is_u64() => Self::Integer,
+ Value::Number(_) => Self::Number,
+ Value::String(_) => Self::String,
+ Value::Array(_) => Self::Array,
+ Value::Object(_) => Self::Object,
+ }
+ }
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[derive(Clone, Copy, Debug, strum::Display)]
+enum MapValueType {
+ String,
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryMetadataField"))]
+struct MetadataField {
+ path: Vec,
+ types: BTreeSet,
+ expression: String,
+}
+
+#[macro_rules_attribute::apply(wire_type)]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryColumn"))]
+struct ColumnSchema {
+ name: String,
+ #[serde(rename = "type")]
+ kind: String,
+ #[serde(flatten)]
+ details: BTreeMap,
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryTable"))]
+struct TableSchema {
+ name: TraceTable,
+ columns: Vec,
+}
+
+trait Unobserved {
+ fn unobserved() -> Self;
+}
+
+enum Discovery {
+ Observed(T),
+ Unavailable(String),
+}
+
+#[cfg(feature = "schema")]
+impl schemars::JsonSchema for Discovery {
+ fn schema_name() -> std::borrow::Cow<'static, str> {
+ format!("Discovery{}", T::schema_name()).into()
+ }
+
+ fn json_schema(generator: &mut schemars::SchemaGenerator) -> schemars::Schema {
+ let mut schema = T::json_schema(generator);
+ schema
+ .as_object_mut()
+ .unwrap()
+ .get_mut("properties")
+ .unwrap()
+ .as_object_mut()
+ .unwrap()
+ .insert(
+ "error".into(),
+ serde_json::json!({"type": ["string", "null"], "default": null}),
+ );
+ schema
+ }
+}
+
+#[cfg(feature = "schema")]
+pub(crate) fn help_schema() -> schemars::Schema {
+ schemars::generate::SchemaSettings::draft2020_12()
+ .for_serialize()
+ .with_transform(litellm_traces::schema::integer_bounds)
+ .into_generator()
+ .into_root_schema_for::()
+}
+
+impl Serialize for Discovery {
+ fn serialize(&self, serializer: S) -> Result {
+ #[derive(Serialize)]
+ struct Unavailable<'a, T> {
+ #[serde(flatten)]
+ sample: T,
+ error: &'a str,
+ }
+ match self {
+ Self::Observed(sample) => sample.serialize(serializer),
+ Self::Unavailable(error) => Unavailable {
+ sample: T::unobserved(),
+ error,
+ }
+ .serialize(serializer),
+ }
+ }
+}
+
+#[macro_rules_attribute::apply(response_type)]
+struct MetadataSample {
+ fields: Vec,
+ sampled_rows: usize,
+ invalid_json_rows: usize,
+ truncated: bool,
+}
+
+impl Unobserved for MetadataSample {
+ fn unobserved() -> Self {
+ Self {
+ fields: Vec::new(),
+ sampled_rows: 0,
+ invalid_json_rows: 0,
+ truncated: true,
+ }
+ }
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryMetadata"))]
+struct MetadataCatalog {
+ table: TraceTable,
+ column: &'static str,
+ #[serde(flatten)]
+ discovery: Discovery,
+ sample_sql: &'static str,
+ scope: &'static str,
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryAttributeField"))]
+struct AttributeField {
+ key: String,
+ #[serde(rename = "type")]
+ kind: MapValueType,
+ expression: String,
+}
+
+#[macro_rules_attribute::apply(response_type)]
+struct AttributeSample {
+ fields: Vec,
+ truncated: bool,
+}
+
+impl Unobserved for AttributeSample {
+ fn unobserved() -> Self {
+ Self {
+ fields: Vec::new(),
+ truncated: true,
+ }
+ }
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryAttributes"))]
+struct AttributeCatalog {
+ table: TraceTable,
+ column: &'static str,
+ #[serde(flatten)]
+ discovery: Discovery,
+ discovery_sql: String,
+ scope: &'static str,
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryNormalizedField"))]
+struct NormalizedField {
+ table: TraceTable,
+ name: &'static str,
+ column: &'static str,
+ #[serde(rename = "type")]
+ kind: &'static str,
+ meaning: &'static str,
+}
+
+impl From<&NormalizedFieldDefinition> for NormalizedField {
+ fn from(field: &NormalizedFieldDefinition) -> Self {
+ Self {
+ table: TraceTable::OtelTraces,
+ name: field.name,
+ column: field.clickhouse_column,
+ kind: field.clickhouse_type,
+ meaning: field.meaning,
+ }
+ }
+}
+
+#[macro_rules_attribute::apply(response_type)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryRelationship"))]
+struct Relationship {
+ left: &'static str,
+ right: &'static str,
+ additional_predicates: &'static str,
+ meaning: &'static str,
+}
+
+const RELATIONSHIPS: [Relationship; 1] = [Relationship {
+ left: "otel_traces.LiteLLMRequestId",
+ right: "spend_logs.response_id",
+ additional_predicates: "otel_traces.TeamId = spend_logs.team_id AND ((otel_traces.UserId != '' AND otel_traces.UserId = spend_logs.user) OR (otel_traces.ApiKeyHash != '' AND otel_traces.ApiKeyHash = spend_logs.api_key))",
+ meaning: "LiteLLMRequestId contains the first normalized request or provider response ID. This relationship matches response IDs only; CallKeys retains all typed identifiers. Cached requests can share response_id; joins may return multiple spend rows",
+}];
+
+#[macro_rules_attribute::apply(response_type)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+#[cfg_attr(feature = "schema", schemars(rename = "TraceQueryHelp"))]
+pub struct QueryHelp {
+ dialect: &'static str,
+ access: &'static str,
+ response: &'static str,
+ tables: Vec,
+ normalized_fields: Vec,
+ metadata: MetadataCatalog,
+ attributes: Vec,
+ relationships: &'static [Relationship],
+ #[cfg_attr(feature = "schema", schemars(with = "Vec"))]
+ examples: [Example; 12],
+ #[cfg_attr(feature = "schema", schemars(with = "Vec"))]
+ gotchas: [String; 13],
+ guide: String,
+}
+
+pub async fn execute_read(
+ client: &Client,
+ connection: &Connection,
+ sql: &str,
+ parameters: &BTreeMap,
+) -> Result {
+ litellm_storage_clickhouse::execute_read(client, connection, sql, parameters)
+ .await
+ .map_err(Error::from)
+}
+
+pub async fn query_sql(
+ client: &Client,
+ connection: &Connection,
+ sql: &str,
+) -> Result {
+ execute_read(client, connection, sql, &BTreeMap::new()).await
+}
+
+async fn rows(
+ client: &Client,
+ connection: &Connection,
+ sql: &str,
+) -> Result, Error> {
+ let body = query_sql(client, connection, sql).await?;
+ serde_json::from_str::>(&body)
+ .map(|result| result.data)
+ .map_err(|_| Error::InvalidResponse)
+}
+
+fn literal(value: &str) -> String {
+ format!("'{}'", value.replace('\\', "\\\\").replace('\'', "\\'"))
+}
+
+fn metadata_expression(path: &[PathPart]) -> String {
+ let arguments = path
+ .iter()
+ .map(|part| match part {
+ PathPart::Key(key) => literal(key),
+ PathPart::Index(index) => index.to_string(),
+ })
+ .collect::>()
+ .join(", ");
+ format!("JSONExtractRaw(metadata, {arguments})")
+}
+
+fn discover(
+ value: &Value,
+ path: Vec,
+ fields: &mut BTreeMap, BTreeSet>,
+) -> bool {
+ if path.len() > MAX_DEPTH || (fields.len() >= MAX_FIELDS && !fields.contains_key(&path)) {
+ return true;
+ }
+ if !path.is_empty() {
+ fields
+ .entry(path.clone())
+ .or_default()
+ .insert(JsonKind::of(value));
+ }
+ match value {
+ Value::Object(object) => object.iter().fold(false, |limited, (key, value)| {
+ let child = path
+ .iter()
+ .cloned()
+ .chain([PathPart::Key(key.clone())])
+ .collect();
+ discover(value, child, fields) | limited
+ }),
+ Value::Array(array) => array
+ .iter()
+ .enumerate()
+ .fold(false, |limited, (index, value)| {
+ let child = path
+ .iter()
+ .cloned()
+ .chain([PathPart::Index(index + 1)])
+ .collect();
+ discover(value, child, fields) | limited
+ }),
+ _ => false,
+ }
+}
+
+fn metadata_sample(sample: &[MetadataRow]) -> MetadataSample {
+ let (fields, limited, invalid_rows) = sample.iter().take(SAMPLE_ROWS).fold(
+ (BTreeMap::new(), sample.len() > SAMPLE_ROWS, 0),
+ |(fields, limited, invalid_rows), row| match serde_json::from_str::(&row.metadata) {
+ Ok(value) => {
+ let mut fields = fields;
+ let limited = limited | discover(&value, Vec::new(), &mut fields);
+ (fields, limited, invalid_rows)
+ }
+ Err(_) => (fields, limited, invalid_rows + 1),
+ },
+ );
+ let fields: Vec<_> = fields
+ .into_iter()
+ .map(|(path, types)| MetadataField {
+ expression: metadata_expression(&path),
+ path,
+ types,
+ })
+ .collect();
+ MetadataSample {
+ fields,
+ sampled_rows: sample.len().min(SAMPLE_ROWS),
+ invalid_json_rows: invalid_rows,
+ truncated: limited,
+ }
+}
+
+pub async fn query_help(client: &Client, connection: &Connection) -> Result {
+ let tables = stream::iter(TraceTable::iter())
+ .then(|table| async move {
+ Ok::<_, Error>(TableSchema {
+ name: table,
+ columns: rows::(
+ client,
+ connection,
+ &format!("DESCRIBE TABLE {table}"),
+ )
+ .await?,
+ })
+ })
+ .try_collect::>()
+ .await?;
+ let metadata = MetadataCatalog {
+ table: TraceTable::SpendLogs,
+ column: "metadata",
+ discovery: match rows::(client, connection, METADATA_SQL).await {
+ Ok(sample) => Discovery::Observed(metadata_sample(&sample)),
+ Err(error) => Discovery::Unavailable(error.to_string()),
+ },
+ sample_sql: METADATA_SQL,
+ scope: METADATA_SCOPE,
+ };
+ let attributes = stream::iter(["SpanAttributes", "ResourceAttributes"])
+ .then(|column| async move {
+ let sql = format!(
+ "SELECT DISTINCT arrayJoin(mapKeys({column})) AS key FROM \
+ (SELECT {column} FROM otel_traces WHERE Timestamp >= now() - INTERVAL 7 DAY \
+ LIMIT 200) ORDER BY key LIMIT 201"
+ );
+ let discovery = match rows::(client, connection, &sql).await {
+ Ok(keys) => Discovery::Observed(AttributeSample {
+ truncated: keys.len() > MAX_FIELDS,
+ fields: keys
+ .into_iter()
+ .take(MAX_FIELDS)
+ .map(|row| AttributeField {
+ expression: format!("{column}[{}]", literal(&row.key)),
+ key: row.key,
+ kind: MapValueType::String,
+ })
+ .collect(),
+ }),
+ Err(error) => Discovery::Unavailable(error.to_string()),
+ };
+ AttributeCatalog {
+ table: TraceTable::OtelTraces,
+ column,
+ discovery,
+ discovery_sql: sql,
+ scope: ATTRIBUTE_SCOPE,
+ }
+ })
+ .collect::>()
+ .await;
+ let guide = guide::QueryGuide {
+ tables: &tables,
+ normalized_fields: &NORMALIZED_FIELD_DEFINITIONS,
+ metadata: &metadata,
+ attributes: &attributes,
+ limits: &READER_LIMITS,
+ };
+ let bodies = guide.sections()?;
+ let sections = [
+ "Live ClickHouse schema",
+ "Normalized span fields",
+ "Observed LLM call metadata",
+ "Observed span and resource attributes",
+ ]
+ .into_iter()
+ .zip(&bodies)
+ .map(|(title, body)| Section { title, body })
+ .collect::>();
+ let examples = guide.examples()?;
+ let gotchas = guide.gotchas()?;
+ let rendered = QueryGuide {
+ sections: §ions,
+ examples: &examples,
+ gotchas: &gotchas,
+ }
+ .render()
+ .map_err(|_| Error::InvalidResponse)?;
+ Ok(QueryHelp {
+ dialect: "ClickHouse SQL",
+ access: "Request-log visibility enforced by ClickHouse row policies; proxy admins see all rows, users see their own rows and permitted teams",
+ response: "ClickHouse JSON envelope: meta, data, rows, statistics; 64-bit integers may be strings",
+ examples,
+ gotchas,
+ guide: rendered,
+ normalized_fields: NORMALIZED_FIELD_DEFINITIONS
+ .iter()
+ .map(NormalizedField::from)
+ .collect(),
+ relationships: &RELATIONSHIPS,
+ tables,
+ metadata,
+ attributes,
+ })
+}
+
+#[cfg(test)]
+mod tests {
+ use super::*;
+ use rstest::rstest;
+ use serde_json::json;
+
+ #[cfg(feature = "schema")]
+ #[rstest]
+ #[case::observed(false)]
+ #[case::unavailable(true)]
+ fn discovery_serialization_matches_its_schema(#[case] unavailable: bool) {
+ let discovery = if unavailable {
+ Discovery::Unavailable("discovery failed".into())
+ } else {
+ Discovery::Observed(MetadataSample::unobserved())
+ };
+ let catalog = MetadataCatalog {
+ table: TraceTable::SpendLogs,
+ column: "metadata",
+ discovery,
+ sample_sql: METADATA_SQL,
+ scope: METADATA_SCOPE,
+ };
+ let schema = schemars::generate::SchemaSettings::draft2020_12()
+ .for_serialize()
+ .into_generator()
+ .into_root_schema_for::();
+ let serialized = serde_json::to_value(&catalog).unwrap();
+ assert!(jsonschema::is_valid(schema.as_value(), &serialized));
+ assert_eq!(serialized.get("error").is_some(), unavailable);
+ assert!(serialized["fields"].is_array());
+ }
+
+ #[rstest]
+ fn metadata_discovery_preserves_mixed_types_and_reports_invalid_rows() {
+ let sample = [
+ MetadataRow {
+ metadata: r#"{"x": 1}"#.into(),
+ },
+ MetadataRow {
+ metadata: r#"{"x": "one"}"#.into(),
+ },
+ MetadataRow {
+ metadata: "invalid".into(),
+ },
+ ];
+ let catalog = json!(metadata_sample(&sample));
+ assert_eq!(
+ catalog["fields"],
+ json!([{
+ "path": ["x"], "types": ["integer", "string"], "expression": "JSONExtractRaw(metadata, 'x')"
+ }])
+ );
+ assert_eq!(catalog["invalid_json_rows"], 1);
+ assert_eq!(catalog["sampled_rows"], sample.len());
+ }
+
+ #[rstest]
+ #[case::rows(SAMPLE_ROWS + 1, 1)]
+ #[case::paths(1, MAX_FIELDS + 1)]
+ fn metadata_discovery_reports_truncation(#[case] row_count: usize, #[case] field_count: usize) {
+ let metadata: BTreeMap<_, _> = (0..field_count)
+ .map(|index| (format!("field{index}"), index))
+ .collect();
+ let sample: Vec<_> = (0..row_count)
+ .map(|_| MetadataRow {
+ metadata: json!(metadata).to_string(),
+ })
+ .collect();
+ let catalog = json!(metadata_sample(&sample));
+ assert_eq!(catalog["truncated"], true);
+ assert_eq!(catalog["sampled_rows"], row_count.min(SAMPLE_ROWS));
+ assert_eq!(
+ catalog["fields"].as_array().unwrap().len(),
+ field_count.min(MAX_FIELDS)
+ );
+ }
+}
diff --git a/litellm-rust/crates/traces-clickhouse/src/query/guide.rs b/litellm-rust/crates/traces-clickhouse/src/query/guide.rs
new file mode 100644
index 00000000000..6755558a974
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/query/guide.rs
@@ -0,0 +1,143 @@
+use askama::Template;
+use litellm_traces::query::guide::Example;
+
+use super::{AttributeCatalog, Discovery, MetadataCatalog, TableSchema};
+use crate::{Error, NormalizedFieldDefinition, query_access::ReaderLimits};
+
+#[derive(Template)]
+#[template(path = "query_help.jinja", escape = "none", blocks = [
+ "live_schema",
+ "normalized_fields",
+ "metadata",
+ "attributes",
+ "recent_spans_name",
+ "recent_spans_sql",
+ "custom_metadata_name",
+ "custom_metadata_sql",
+ "nested_metadata_name",
+ "nested_metadata_sql",
+ "correlated_calls_name",
+ "correlated_calls_sql",
+ "discover_keys_name",
+ "discover_keys_sql",
+ "recent_spend_name",
+ "recent_spend_sql",
+ "model_spend_name",
+ "model_spend_sql",
+ "trace_spend_name",
+ "trace_spend_sql",
+ "unmatched_spans_name",
+ "unmatched_spans_sql",
+ "trace_summary_name",
+ "trace_summary_sql",
+ "failed_spans_name",
+ "failed_spans_sql",
+ "metadata_filter_name",
+ "metadata_filter_sql",
+ "missing_spend",
+ "partial_spend",
+ "time_window",
+ "reader_limits",
+ "reader_profile",
+ "output_format",
+ "json_values",
+ "map_values",
+ "literal_keys",
+ "time_units",
+ "spend_totals",
+ "trace_rollups",
+ "sampling",
+])]
+pub(super) struct QueryGuide<'a> {
+ pub tables: &'a [TableSchema],
+ pub normalized_fields: &'a [NormalizedFieldDefinition],
+ pub metadata: &'a MetadataCatalog,
+ pub attributes: &'a [AttributeCatalog],
+ pub limits: &'a ReaderLimits,
+}
+
+impl QueryGuide<'_> {
+ pub fn sections(&self) -> Result<[String; 4], Error> {
+ Ok([
+ render(&self.as_live_schema())?,
+ render(&self.as_normalized_fields())?,
+ render(&self.as_metadata())?,
+ render(&self.as_attributes())?,
+ ])
+ }
+
+ pub fn examples(&self) -> Result<[Example; 12], Error> {
+ Ok([
+ Example {
+ name: render(&self.as_recent_spans_name())?,
+ sql: render(&self.as_recent_spans_sql())?,
+ },
+ Example {
+ name: render(&self.as_custom_metadata_name())?,
+ sql: render(&self.as_custom_metadata_sql())?,
+ },
+ Example {
+ name: render(&self.as_nested_metadata_name())?,
+ sql: render(&self.as_nested_metadata_sql())?,
+ },
+ Example {
+ name: render(&self.as_correlated_calls_name())?,
+ sql: render(&self.as_correlated_calls_sql())?,
+ },
+ Example {
+ name: render(&self.as_discover_keys_name())?,
+ sql: render(&self.as_discover_keys_sql())?,
+ },
+ Example {
+ name: render(&self.as_recent_spend_name())?,
+ sql: render(&self.as_recent_spend_sql())?,
+ },
+ Example {
+ name: render(&self.as_model_spend_name())?,
+ sql: render(&self.as_model_spend_sql())?,
+ },
+ Example {
+ name: render(&self.as_trace_spend_name())?,
+ sql: render(&self.as_trace_spend_sql())?,
+ },
+ Example {
+ name: render(&self.as_unmatched_spans_name())?,
+ sql: render(&self.as_unmatched_spans_sql())?,
+ },
+ Example {
+ name: render(&self.as_trace_summary_name())?,
+ sql: render(&self.as_trace_summary_sql())?,
+ },
+ Example {
+ name: render(&self.as_failed_spans_name())?,
+ sql: render(&self.as_failed_spans_sql())?,
+ },
+ Example {
+ name: render(&self.as_metadata_filter_name())?,
+ sql: render(&self.as_metadata_filter_sql())?,
+ },
+ ])
+ }
+
+ pub fn gotchas(&self) -> Result<[String; 13], Error> {
+ Ok([
+ render(&self.as_time_window())?,
+ render(&self.as_reader_limits())?,
+ render(&self.as_reader_profile())?,
+ render(&self.as_output_format())?,
+ render(&self.as_json_values())?,
+ render(&self.as_map_values())?,
+ render(&self.as_literal_keys())?,
+ render(&self.as_time_units())?,
+ render(&self.as_spend_totals())?,
+ render(&self.as_missing_spend())?,
+ render(&self.as_partial_spend())?,
+ render(&self.as_trace_rollups())?,
+ render(&self.as_sampling())?,
+ ])
+ }
+}
+
+pub(super) fn render(template: &impl Template) -> Result {
+ template.render().map_err(|_| Error::InvalidResponse)
+}
diff --git a/litellm-rust/crates/traces-clickhouse/src/query/lens.rs b/litellm-rust/crates/traces-clickhouse/src/query/lens.rs
new file mode 100644
index 00000000000..ff30f127000
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/query/lens.rs
@@ -0,0 +1,268 @@
+use litellm_storage_clickhouse::Query;
+
+pub const LENS_QUERIES: [litellm_traces::ReadQuery; 5] = [
+ litellm_traces::ReadQuery::Availability,
+ litellm_traces::ReadQuery::Agents,
+ litellm_traces::ReadQuery::Sample,
+ litellm_traces::ReadQuery::Content,
+ litellm_traces::ReadQuery::Evidence,
+];
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[serde(rename_all = "lowercase")]
+pub enum ExecutionSource {
+ Traces,
+ Requests,
+ Both,
+}
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[serde(rename_all = "lowercase")]
+pub enum ContentSource {
+ Traces,
+ Requests,
+}
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[cfg_attr(feature = "schema", schemars(deny_unknown_fields))]
+pub struct LensAccessParams {
+ #[serde(
+ deserialize_with = "super::number::boolean",
+ serialize_with = "litellm_traces::wire::serialize_flag"
+ )]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "litellm_traces::schema::flag")
+ )]
+ pub all_teams: bool,
+ pub team: String,
+ pub key_hash: String,
+}
+
+pub struct LensAvailability;
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[serde(deny_unknown_fields)]
+pub struct LensAvailabilityParams {
+ #[serde(flatten)]
+ pub access: LensAccessParams,
+}
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[cfg_attr(feature = "schema", schemars(rename = "ActivityAvailability"))]
+pub struct LensAvailabilityRow {
+ #[serde(default, deserialize_with = "super::number::flag")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "crate::wire_schema::boolean_flag")
+ )]
+ pub traces: u8,
+ #[serde(default, deserialize_with = "super::number::flag")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "crate::wire_schema::boolean_flag")
+ )]
+ pub requests: u8,
+}
+
+impl Query for LensAvailability {
+ type Params = LensAvailabilityParams;
+ type Row = LensAvailabilityRow;
+
+ const SQL: &'static str = include_str!("../../query/lens_availability.sql");
+}
+
+pub struct LensAgents;
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[serde(deny_unknown_fields)]
+pub struct LensAgentsParams {
+ #[serde(flatten)]
+ pub access: LensAccessParams,
+}
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[cfg_attr(feature = "schema", schemars(rename = "AgentRow"))]
+pub struct LensAgentsRow {
+ pub agent_name: String,
+}
+
+impl Query for LensAgents {
+ type Params = LensAgentsParams;
+ type Row = LensAgentsRow;
+
+ const SQL: &'static str = include_str!("../../query/lens_agents.sql");
+}
+
+pub struct LensSample;
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[serde(deny_unknown_fields)]
+pub struct LensSampleParams {
+ #[serde(flatten)]
+ pub access: LensAccessParams,
+ pub source: ExecutionSource,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub start: u64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub end: u64,
+ pub agent_name: String,
+ pub service: String,
+ pub filter_keys: Vec,
+ pub filter_values: Vec,
+ pub selected_team: String,
+ pub execution_ids: Vec,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub sample_cap: u64,
+ #[serde(deserialize_with = "super::number::percent")]
+ #[cfg_attr(feature = "schema", schemars(range(min = 0, max = 100)))]
+ pub sample_percent: f64,
+ #[serde(deserialize_with = "super::number::flag")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "litellm_traces::schema::flag")
+ )]
+ pub preview: u8,
+ pub after: String,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub limit: u32,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub offset: u64,
+}
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[cfg_attr(feature = "schema", schemars(rename = "ExecutionRow"))]
+pub struct LensSampleRow {
+ pub source: ContentSource,
+ pub trace_id: String,
+ pub team_id: String,
+ #[serde(default)]
+ pub trace_ref: String,
+ pub name: String,
+ pub start_time: String,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "crate::wire_schema::u64_number")
+ )]
+ pub span_count: u64,
+ #[serde(deserialize_with = "super::number::flag")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "crate::wire_schema::flag_number")
+ )]
+ pub root_seen: u8,
+ #[serde(default)]
+ pub service: String,
+ #[serde(default)]
+ pub attributes: Vec<(String, String)>,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "crate::wire_schema::u64_number")
+ )]
+ pub eligible: u64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ #[cfg_attr(feature = "schema", schemars(skip))]
+ pub position: u64,
+ #[serde(default, deserialize_with = "super::number::deserialize")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "crate::wire_schema::selected")
+ )]
+ pub selected: f64,
+ #[serde(default)]
+ pub selection_key: String,
+}
+
+impl Query for LensSample {
+ type Params = LensSampleParams;
+ type Row = LensSampleRow;
+
+ const SQL: &'static str = include_str!("../../query/lens_sample.sql");
+}
+
+pub struct LensContent;
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[serde(deny_unknown_fields)]
+pub struct LensContentParams {
+ #[serde(flatten)]
+ pub access: LensAccessParams,
+ pub source: ContentSource,
+ pub id: String,
+ pub record_team: String,
+ pub trace_ref: String,
+ pub cursor: String,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub offset: u32,
+}
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[cfg_attr(feature = "schema", schemars(rename = "PartRow"))]
+pub struct LensContentRow {
+ pub span_id: String,
+ pub parent_span_id: String,
+ pub name: String,
+ pub kind: String,
+ pub content: String,
+ #[serde(deserialize_with = "super::number::flag")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "crate::wire_schema::flag_number")
+ )]
+ pub truncated: u8,
+}
+
+impl Query for LensContent {
+ type Params = LensContentParams;
+ type Row = LensContentRow;
+
+ const SQL: &'static str = include_str!("../../query/lens_content.sql");
+}
+
+pub struct LensEvidence;
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[serde(deny_unknown_fields)]
+pub struct LensEvidenceParams {
+ #[serde(flatten)]
+ pub access: LensAccessParams,
+ pub source: ContentSource,
+ pub id: String,
+ pub record_team: String,
+ pub trace_ref: String,
+ pub span: String,
+ pub quote: String,
+}
+
+#[macro_rules_attribute::apply(wire_type)]
+#[derive(Debug)]
+#[cfg_attr(feature = "schema", schemars(rename = "CountRow"))]
+pub struct LensEvidenceRow {
+ #[serde(deserialize_with = "super::number::deserialize")]
+ #[cfg_attr(
+ feature = "schema",
+ schemars(schema_with = "crate::wire_schema::u64_number")
+ )]
+ pub count: u64,
+}
+
+impl Query for LensEvidence {
+ type Params = LensEvidenceParams;
+ type Row = LensEvidenceRow;
+
+ const SQL: &'static str = include_str!("../../query/lens_evidence.sql");
+}
diff --git a/litellm-rust/crates/traces-clickhouse/src/query/named.rs b/litellm-rust/crates/traces-clickhouse/src/query/named.rs
new file mode 100644
index 00000000000..beb4b42f76d
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/query/named.rs
@@ -0,0 +1,409 @@
+use litellm_storage_clickhouse::Query;
+use litellm_traces::query::named as contracts;
+use serde::{Deserialize, Serialize};
+
+pub use contracts::ReadAccessParams;
+
+#[derive(Deserialize, Serialize)]
+#[serde(remote = "contracts::ListTracesParams")]
+struct ListTracesParamsEncoding {
+ #[serde(flatten)]
+ pub access: contracts::ReadAccessParams,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub start_ms: i64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub end_ms: i64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub cursor_ms: i64,
+ pub cursor_trace_id: String,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub limit: u32,
+}
+
+#[derive(Debug, Deserialize, Serialize)]
+pub struct ListTracesParams(
+ #[serde(with = "ListTracesParamsEncoding")] pub contracts::ListTracesParams,
+);
+
+impl From for ListTracesParams {
+ fn from(value: contracts::ListTracesParams) -> Self {
+ Self(value)
+ }
+}
+
+#[derive(Deserialize, Serialize)]
+#[serde(remote = "contracts::ListTracesRow")]
+struct ListTracesRowEncoding {
+ pub trace_id: String,
+ pub trace_ref: String,
+ pub team_id: String,
+ pub api_key_hash: String,
+ pub user_id: String,
+ pub name: String,
+ pub service: String,
+ pub input_preview: String,
+ #[serde(serialize_with = "litellm_traces::wire::serialize_status")]
+ pub status: litellm_traces::SpanStatus,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub start_ms: i64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub duration_ms: i64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub span_count: u64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub agent_count: u64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub agent_invocations: u64,
+ #[serde(default)]
+ pub agent_names: Vec,
+ #[serde(default)]
+ pub frameworks: Vec,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub llm_calls: u64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub tool_calls: u64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub input_tokens: u64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub output_tokens: u64,
+ pub models: Vec,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub error_count: u64,
+ pub request_ids: Vec,
+}
+
+#[derive(Debug, Deserialize, Serialize)]
+pub struct ListTracesRow(#[serde(with = "ListTracesRowEncoding")] pub contracts::ListTracesRow);
+
+pub use contracts::TraceSpansParams;
+
+#[derive(Deserialize, Serialize)]
+#[serde(remote = "contracts::TraceSpansRow")]
+struct TraceSpansRowEncoding {
+ #[serde(default)]
+ pub trace_id: String,
+ pub span_id: String,
+ pub parent_span_id: String,
+ pub name: String,
+ #[serde(rename = "type")]
+ pub kind: litellm_traces::ObservationType,
+ #[serde(
+ default,
+ deserialize_with = "super::number::boolean",
+ serialize_with = "litellm_traces::wire::serialize_flag"
+ )]
+ pub wrapper_candidate: bool,
+ pub agent: String,
+ #[serde(default)]
+ pub framework: String,
+ #[serde(serialize_with = "litellm_traces::wire::serialize_status")]
+ pub status: litellm_traces::SpanStatus,
+ pub status_message: String,
+ #[serde(
+ deserialize_with = "super::number::boolean",
+ serialize_with = "litellm_traces::wire::serialize_flag"
+ )]
+ pub error_truncated: bool,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub start_ns: i64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub duration_ns: u64,
+ pub service: String,
+ pub input_preview: String,
+ pub model: String,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub input_tokens: u32,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub output_tokens: u32,
+ pub litellm_request_id: String,
+ #[serde(default)]
+ pub call_keys: Vec,
+ #[serde(
+ default,
+ deserialize_with = "litellm_traces::wire::evidence",
+ serialize_with = "litellm_traces::wire::serialize_evidence"
+ )]
+ pub call_evidence: Option,
+ #[serde(default)]
+ pub tool_call_id: String,
+ pub team_id: String,
+ pub api_key_hash: String,
+ pub user_id: String,
+}
+
+#[derive(Debug, Deserialize, Serialize)]
+pub struct TraceSpansRow(#[serde(with = "TraceSpansRowEncoding")] pub contracts::TraceSpansRow);
+
+pub use contracts::SpanDetailParams;
+
+pub use contracts::SpanDetailRow;
+
+#[derive(Deserialize, Serialize)]
+#[serde(remote = "contracts::SpanErrorParams")]
+struct SpanErrorParamsEncoding {
+ #[serde(flatten)]
+ pub access: contracts::ReadAccessParams,
+ pub trace_id: String,
+ pub trace_ref: String,
+ pub span_id: String,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub error_offset: u64,
+ pub error_version: String,
+}
+
+#[derive(Debug, Deserialize, Serialize)]
+pub struct SpanErrorParams(
+ #[serde(with = "SpanErrorParamsEncoding")] pub contracts::SpanErrorParams,
+);
+
+impl From for SpanErrorParams {
+ fn from(value: contracts::SpanErrorParams) -> Self {
+ Self(value)
+ }
+}
+
+#[derive(Deserialize, Serialize)]
+#[serde(remote = "contracts::SpanErrorRow")]
+struct SpanErrorRowEncoding {
+ pub span_id: String,
+ pub message: String,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub total_chars: u64,
+ pub version: String,
+}
+
+#[derive(Debug, Deserialize, Serialize)]
+pub struct SpanErrorRow(#[serde(with = "SpanErrorRowEncoding")] pub contracts::SpanErrorRow);
+
+#[derive(Deserialize, Serialize)]
+#[serde(remote = "contracts::SpendByResponseIdsParams")]
+struct SpendByResponseIdsParamsEncoding {
+ #[serde(flatten)]
+ pub access: contracts::ReadAccessParams,
+ pub response_ids: Vec,
+ pub request_ids: Vec,
+ pub trace_ids: Vec,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub start_ms: i64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub end_ms: i64,
+}
+
+#[derive(Debug, Deserialize, Serialize)]
+pub struct SpendByResponseIdsParams(
+ #[serde(with = "SpendByResponseIdsParamsEncoding")] pub contracts::SpendByResponseIdsParams,
+);
+
+impl From for SpendByResponseIdsParams {
+ fn from(value: contracts::SpendByResponseIdsParams) -> Self {
+ Self(value)
+ }
+}
+
+#[derive(Deserialize, Serialize)]
+#[serde(remote = "contracts::SpendByResponseIdsRow")]
+struct SpendByResponseIdsRowEncoding {
+ pub request_id: String,
+ pub litellm_call_id: String,
+ pub response_id: String,
+ pub upstream_response_id: String,
+ pub trace_id: String,
+ pub span_id: String,
+ pub team_id: String,
+ pub api_key: String,
+ pub user: String,
+ #[serde(deserialize_with = "super::number::optional_finite")]
+ pub spend: Option,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub start_ms: i64,
+}
+
+#[derive(Debug, Deserialize, Serialize)]
+pub struct SpendByResponseIdsRow(
+ #[serde(with = "SpendByResponseIdsRowEncoding")] pub contracts::SpendByResponseIdsRow,
+);
+
+pub struct ListTraces;
+
+impl Query for ListTraces {
+ type Params = ListTracesParams;
+ type Row = ListTracesRow;
+
+ const SQL: &'static str = include_str!("../../query/list_traces.sql");
+}
+
+#[derive(Deserialize, Serialize)]
+#[serde(remote = "contracts::TracePageSpansParams")]
+struct TracePageSpansParamsEncoding {
+ #[serde(flatten)]
+ pub access: contracts::ReadAccessParams,
+ pub trace_refs: Vec,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub start_ms: i64,
+ #[serde(deserialize_with = "super::number::deserialize")]
+ pub end_ms: i64,
+}
+
+#[derive(Debug, Deserialize, Serialize)]
+pub struct TracePageSpansParams(
+ #[serde(with = "TracePageSpansParamsEncoding")] pub contracts::TracePageSpansParams,
+);
+
+impl From for TracePageSpansParams {
+ fn from(value: contracts::TracePageSpansParams) -> Self {
+ Self(value)
+ }
+}
+
+pub struct TracePageSpans;
+
+impl Query for TracePageSpans {
+ type Params = TracePageSpansParams;
+ type Row = TraceSpansRow;
+
+ const SQL: &'static str = include_str!("../../query/trace_page_spans.sql");
+}
+
+pub struct TraceSpans;
+
+impl Query for TraceSpans {
+ type Params = TraceSpansParams;
+ type Row = TraceSpansRow;
+
+ const SQL: &'static str = include_str!("../../query/trace_spans.sql");
+}
+
+pub struct SpanDetail;
+
+impl Query for SpanDetail {
+ type Params = SpanDetailParams;
+ type Row = SpanDetailRow;
+
+ const SQL: &'static str = include_str!("../../query/span_detail.sql");
+}
+
+pub struct SpanError;
+
+impl Query for SpanError {
+ type Params = SpanErrorParams;
+ type Row = SpanErrorRow;
+
+ const SQL: &'static str = include_str!("../../query/span_error.sql");
+}
+
+pub struct SpendByResponseIds;
+
+impl Query for SpendByResponseIds {
+ type Params = SpendByResponseIdsParams;
+ type Row = SpendByResponseIdsRow;
+
+ const SQL: &'static str = include_str!("../../query/spend_by_response_ids.sql");
+}
+
+pub use contracts::{TraceIdentityParams, TraceIdentityRow};
+
+pub struct TraceIdentity;
+
+impl Query for TraceIdentity {
+ type Params = TraceIdentityParams;
+ type Row = TraceIdentityRow;
+ const SQL: &'static str = include_str!("../../query/trace_identity.sql");
+}
+
+#[cfg(test)]
+mod tests {
+ use super::*;
+ use rstest::rstest;
+ use serde_json::{Value, json};
+
+ fn round_trip(wire: Value, quoted: bool) {
+ let encoded = Value::Object(
+ wire.as_object()
+ .unwrap()
+ .iter()
+ .map(|(name, value)| {
+ let encoded = if quoted && value.is_number() && name != "all_teams" {
+ json!(value.to_string())
+ } else {
+ value.clone()
+ };
+ (name.clone(), encoded)
+ })
+ .collect(),
+ );
+ let decoded: T = serde_json::from_value(encoded).unwrap();
+ assert_eq!(serde_json::to_value(decoded).unwrap(), wire);
+ }
+
+ #[rstest]
+ #[case::unquoted(false)]
+ #[case::quoted(true)]
+ fn rows_decode_into_neutral_contracts(#[case] quoted: bool) {
+ round_trip::(
+ json!({"trace_id": "trace", "trace_ref": "ref", "team_id": "team", "api_key_hash": "key", "user_id": "user", "name": "agent", "service": "service", "input_preview": "input", "status": "STATUS_CODE_OK", "start_ms": -1, "duration_ms": 20, "span_count": u64::MAX, "agent_count": 1, "agent_invocations": 2, "agent_names": ["agent"], "frameworks": ["claude-agent-sdk"], "llm_calls": 3, "tool_calls": 4, "input_tokens": 5, "output_tokens": 6, "models": ["model"], "error_count": 0, "request_ids": ["request"]}),
+ quoted,
+ );
+ round_trip::(
+ json!({"trace_id": "trace", "span_id": "span", "parent_span_id": "parent", "name": "agent", "type": "agent", "wrapper_candidate": 1, "agent": "agent", "framework": "claude-agent-sdk", "status": "STATUS_CODE_ERROR", "status_message": "error", "error_truncated": 1, "start_ns": -1, "duration_ns": u64::MAX, "service": "service", "input_preview": "input", "model": "model", "input_tokens": u32::MAX, "output_tokens": 6, "litellm_request_id": "request", "call_keys": ["provider_response:request"], "call_evidence": "complete", "tool_call_id": "call", "team_id": "team", "api_key_hash": "key", "user_id": "user"}),
+ quoted,
+ );
+ round_trip::(
+ json!({"span_id": "span", "input": "input", "output": "output", "attributes": {"count": "42"}}),
+ quoted,
+ );
+ round_trip::(
+ json!({"span_id": "span", "message": "error", "total_chars": u64::MAX, "version": "version"}),
+ quoted,
+ );
+ round_trip::(
+ json!({"request_id": "request", "litellm_call_id": "gateway", "response_id": "response", "upstream_response_id": "upstream", "trace_id": "trace", "span_id": "span", "team_id": "team", "api_key": "key", "user": "user", "spend": 0.125, "start_ms": -1}),
+ quoted,
+ );
+ }
+
+ #[rstest]
+ #[case::unquoted(false)]
+ #[case::quoted(true)]
+ fn parameters_preserve_flattened_multi_team_access(#[case] quoted: bool) {
+ round_trip::(
+ json!({"all_teams": 0, "user_id": "user", "team_ids": ["team-a", "team-b"], "start_ms": -1, "end_ms": 10, "cursor_ms": 0, "cursor_trace_id": "", "limit": u32::MAX}),
+ quoted,
+ );
+ round_trip::(
+ json!({"all_teams": 0, "user_id": "", "team_ids": [], "trace_id": "trace", "trace_ref": "ref", "span_id": "span", "error_offset": u64::MAX, "error_version": "version"}),
+ quoted,
+ );
+ round_trip::(
+ json!({"all_teams": 0, "user_id": "user", "team_ids": ["team-a", "team-b"], "response_ids": ["response"], "request_ids": ["request"], "trace_ids": ["trace"], "start_ms": -1, "end_ms": 10}),
+ quoted,
+ );
+ }
+ #[rstest]
+ #[case::unknown(json!(null), None)]
+ #[case::free(json!(0), Some(0.0))]
+ #[case::paid(json!("0.125"), Some(0.125))]
+ fn spend_rows_preserve_unknown_and_known_cost(
+ #[case] cost: serde_json::Value,
+ #[case] expected: Option,
+ ) {
+ let row: SpendByResponseIdsRow = serde_json::from_value(json!({
+ "request_id": "request", "litellm_call_id": "gateway", "response_id": "response", "upstream_response_id": "",
+ "trace_id": "trace", "span_id": "span", "team_id": "team", "api_key": "key",
+ "user": "user", "spend": cost, "start_ms": 0
+ }))
+ .unwrap();
+ assert_eq!(row.0.spend, expected);
+ }
+ #[rstest]
+ #[case::nan(json!("NaN"))]
+ #[case::infinity(json!("1e999"))]
+ #[case::boolean(json!(true))]
+ fn spend_rows_reject_invalid_cost(#[case] cost: serde_json::Value) {
+ let row = serde_json::from_value::(json!({
+ "request_id": "request", "litellm_call_id": "gateway", "response_id": "response", "upstream_response_id": "",
+ "trace_id": "trace", "span_id": "span", "team_id": "team", "api_key": "key",
+ "user": "user", "spend": cost, "start_ms": 0
+ }));
+ assert!(row.is_err());
+ }
+}
diff --git a/litellm-rust/crates/traces-clickhouse/src/query/number.rs b/litellm-rust/crates/traces-clickhouse/src/query/number.rs
new file mode 100644
index 00000000000..9283903fee1
--- /dev/null
+++ b/litellm-rust/crates/traces-clickhouse/src/query/number.rs
@@ -0,0 +1,128 @@
+use serde::{Deserialize, Deserializer, de::DeserializeOwned};
+
+pub(super) fn deserialize<'de, D, T>(deserializer: D) -> Result
+where
+ D: Deserializer<'de>,
+ T: DeserializeOwned,
+{
+ #[derive(Deserialize)]
+ #[serde(untagged)]
+ enum Number {
+ Quoted(String),
+ Unquoted(serde_json::Number),
+ }
+ match Number::deserialize(deserializer)? {
+ Number::Quoted(value) => serde_json::from_str(&value),
+ Number::Unquoted(value) => serde_json::from_value(serde_json::Value::Number(value)),
+ }
+ .map_err(serde::de::Error::custom)
+}
+
+pub(super) fn optional_finite<'de, D: Deserializer<'de>>(
+ deserializer: D,
+) -> Result