From 49bb01125b0d3012baafad5c4cd2862c0b95dec1 Mon Sep 17 00:00:00 2001 From: Michael Chen Date: Wed, 20 May 2026 23:38:51 +0800 Subject: [PATCH 1/4] cookbook: add LiteLLM Proxy + SpendGuard runtime budget guardrails MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds a cookbook notebook (cookbook/logging_observability/ LiteLLM_Proxy_SpendGuard.ipynb) demonstrating how to wire SpendGuard's budget-gating CustomLogger into the LiteLLM proxy. What SpendGuard adds on top of LiteLLM's built-in spend tracking: - Pre-call reserve: contract-DSL evaluator decides ALLOW / DENY / DEGRADE / REQUIRE_APPROVAL BEFORE the upstream provider is hit. Budget exhausted → HTTP 403; provider invoice clock never starts. - Post-call commit with reconciliation: end-of-stream reconciler reads response.usage.completion_tokens and commits the real cost. - Signed append-only audit chain: every decision lands as a CloudEvent in canonical_events with the LiteLLM-specific 12-field enrichment (litellm_call_id, model, team_id, pricing_version, ...). - Fail-closed default; SPENDGUARD_LITELLM_FAIL_OPEN=1 dev override. Pure CustomLogger callback — no LiteLLM source changes. Companion to the existing Langfuse/Arize/Lunary notebooks. SpendGuard sits in the same observability slot but as a fail-closed budget control gate, not just a tracer. PyPI: https://pypi.org/project/spendguard-sdk/0.3.0/ Repo: https://github.com/m24927605/agentic-spendguard Signed-off-by: Michael Chen --- .../LiteLLM_Proxy_SpendGuard.ipynb | 234 ++++++++++++++++++ 1 file changed, 234 insertions(+) create mode 100644 cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb diff --git a/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb b/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb new file mode 100644 index 00000000000..ff2cfb320a9 --- /dev/null +++ b/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb @@ -0,0 +1,234 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Budget Guardrails — LiteLLM Proxy + SpendGuard\n", + "\n", + "This notebook demonstrates how to use the LiteLLM Proxy with [SpendGuard](https://github.com/m24927605/agentic-spendguard) to gate every `/v1/chat/completions` call against a budget BEFORE the upstream provider is hit.\n", + "\n", + "**What SpendGuard adds on top of LiteLLM's own spend tracking:**\n", + "- **Pre-call reserve** — SpendGuard's contract-DSL evaluator decides ALLOW / DENY / DEGRADE / REQUIRE_APPROVAL before LiteLLM forwards to the provider. Budget exhausted? HTTP 403, provider invoice clock never starts.\n", + "- **Post-call commit with reconciliation** — End-of-stream reconciler reads `response.usage.completion_tokens` and commits the real cost (not the worst-case estimator).\n", + "- **Signed append-only audit chain** — Every decision lands as a CloudEvent in `canonical_events`, with the 12-field LiteLLM-specific enrichment (`litellm_call_id`, `model`, `team_id`, `pricing_version`, etc) for forensics.\n", + "- **Fail-closed default** — Sidecar unreachable → HTTP 503; the request is denied, not forwarded.\n", + "\n", + "SpendGuard wires in as a standard LiteLLM `CustomLogger` callback. No LiteLLM source change required." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## 1. Setup LiteLLM Proxy + SpendGuard\n", + "\n", + "### 1.1 Install the SpendGuard SDK\n", + "\n", + "```bash\n", + "pip install 'spendguard-sdk[litellm]==0.3.0'\n", + "```\n", + "\n", + "The `litellm` extra pulls `litellm[proxy]` transitively so the proxy CLI works out of the box.\n", + "\n", + "### 1.2 Run the SpendGuard sidecar\n", + "\n", + "The SpendGuard sidecar is a Rust binary that runs alongside your LiteLLM proxy and exposes a Unix domain socket for the callback to talk to. The fastest path is the runnable example:\n", + "\n", + "```bash\n", + "git clone https://github.com/m24927605/agentic-spendguard\n", + "cd agentic-spendguard/examples/litellm-proxy-composite\n", + "docker compose up --build\n", + "```\n", + "\n", + "This brings up postgres + sidecar + your LiteLLM proxy in 3 containers." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### 1.3 Write your operator callback module\n", + "\n", + "Save as `spendguard_callback.py` next to your `proxy_config.yaml`:\n", + "\n", + "```python\n", + "import os\n", + "from spendguard._proto.spendguard.common.v1 import common_pb2\n", + "from spendguard.integrations.litellm import (\n", + " BudgetBinding, ResolverContext, _LoopBoundCallback,\n", + ")\n", + "\n", + "_UNIT = common_pb2.UnitRef(unit_id=os.environ['SPENDGUARD_UNIT_ID'],\n", + " token_kind='output_token', model_family='gpt-4')\n", + "_PRICING = common_pb2.PricingFreeze(\n", + " pricing_version=os.environ['SPENDGUARD_PRICING_VERSION'],\n", + " price_snapshot_hash=bytes.fromhex(os.environ['SPENDGUARD_PRICE_SNAPSHOT_HASH_HEX']),\n", + " fx_rate_version=os.environ['SPENDGUARD_FX_RATE_VERSION'],\n", + " unit_conversion_version=os.environ['SPENDGUARD_UNIT_CONVERSION_VERSION'],\n", + ")\n", + "_BINDING = BudgetBinding(\n", + " budget_id=os.environ['SPENDGUARD_BUDGET_ID'],\n", + " window_instance_id=os.environ['SPENDGUARD_WINDOW_INSTANCE_ID'],\n", + " unit=_UNIT, pricing=_PRICING,\n", + ")\n", + "\n", + "def _resolve(ctx: ResolverContext) -> BudgetBinding | None:\n", + " # In production: dispatch by ctx.user_api_key_dict.team_id\n", + " return _BINDING\n", + "\n", + "def _estimate(ctx):\n", + " return [common_pb2.BudgetClaim(\n", + " budget_id=_BINDING.budget_id, unit=_UNIT, amount_atomic='50',\n", + " direction=common_pb2.BudgetClaim.DEBIT,\n", + " window_instance_id=_BINDING.window_instance_id,\n", + " )]\n", + "\n", + "def _reconcile(ctx, response):\n", + " tokens = int(getattr(getattr(response, 'usage', None), 'completion_tokens', 0) or 0)\n", + " return [common_pb2.BudgetClaim(\n", + " budget_id=_BINDING.budget_id, unit=_UNIT, amount_atomic=str(max(tokens, 1)),\n", + " direction=common_pb2.BudgetClaim.DEBIT,\n", + " window_instance_id=_BINDING.window_instance_id,\n", + " )]\n", + "\n", + "handler_instance = _LoopBoundCallback(\n", + " socket_path=os.environ['SPENDGUARD_SIDECAR_UDS'],\n", + " tenant_id=os.environ['SPENDGUARD_TENANT_ID'],\n", + " budget_resolver=_resolve,\n", + " claim_estimator=_estimate,\n", + " claim_reconciler=_reconcile,\n", + ")\n", + "```" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### 1.4 Update proxy_config.yaml\n", + "\n", + "```yaml\n", + "model_list:\n", + " - model_name: gpt-4o-mini\n", + " litellm_params:\n", + " model: openai/gpt-4o-mini\n", + " api_key: os.environ/OPENAI_API_KEY\n", + "\n", + "litellm_settings:\n", + " callbacks: [\"spendguard_callback.handler_instance\"]\n", + " drop_params: true\n", + "\n", + "general_settings:\n", + " master_key: os.environ/LITELLM_MASTER_KEY\n", + "```\n", + "\n", + "Set `PYTHONPATH` to include the directory of `spendguard_callback.py` before launching the proxy." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### 1.5 Launch the proxy\n", + "\n", + "```bash\n", + "export SPENDGUARD_SIDECAR_UDS=/var/run/spendguard/adapter.sock\n", + "export SPENDGUARD_TENANT_ID=...\n", + "export SPENDGUARD_BUDGET_ID=...\n", + "# ... (full env-var list at github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/PROXY_RECIPE.md)\n", + "\n", + "python -m litellm.proxy.proxy_cli --config proxy_config.yaml --port 4000\n", + "```" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## 2. Make LLM Requests to the Gated Proxy\n", + "\n", + "Standard OpenAI client — no SpendGuard SDK in your app code. The gating is invisible until the budget hits a limit." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import openai\n", + "\n", + "client = openai.OpenAI(\n", + " base_url=\"http://localhost:4000\",\n", + " api_key=\"sk-demo-key\", # LITELLM_MASTER_KEY\n", + ")\n", + "\n", + "# ALLOW path — SpendGuard reserves + commits normally.\n", + "response = client.chat.completions.create(\n", + " model=\"gpt-4o-mini\",\n", + " messages=[{\"role\": \"user\", \"content\": \"hello\"}],\n", + ")\n", + "print(response.choices[0].message.content)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## 3. Verify the Audit Chain\n", + "\n", + "Every call writes a signed CloudEvent to `canonical_events`. Cross-table join with LiteLLM's own `LiteLLM_SpendLogs` (if you have a SpendLogs database configured):\n", + "\n", + "```sql\n", + "SELECT\n", + " ce.event_type,\n", + " ce.tenant_id,\n", + " cost_advisor_safe_decode_payload(ce.payload_json)->'spendguard'->>'litellm_call_id'\n", + " AS litellm_call_id,\n", + " cost_advisor_safe_decode_payload(ce.payload_json)->'spendguard'->>'model' AS model,\n", + " cost_advisor_safe_decode_payload(ce.payload_json)->'spendguard'->>'team_id' AS team_id,\n", + " cost_advisor_safe_decode_payload(ce.payload_json)->>'final_decision' AS decision\n", + "FROM canonical_events ce\n", + "WHERE ce.event_type IN ('spendguard.audit.decision', 'spendguard.audit.outcome')\n", + " AND ce.event_time > NOW() - INTERVAL '1 hour'\n", + "ORDER BY ce.event_time DESC;\n", + "```\n", + "\n", + "The `spendguard` JSONB sub-object on each row carries the 12-field LiteLLM-specific enrichment per [DESIGN §8.2a](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/DESIGN.md)." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## 4. Going Further\n", + "\n", + "- **Multi-tenant via virtual keys:** the operator template above is single-team. To dispatch per LiteLLM team, inspect `ctx.user_api_key_dict.team_id` in `_resolve` and look up the binding in your control plane. Full template: [PROXY_RECIPE.md §2](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/PROXY_RECIPE.md).\n", + "- **Streaming:** SpendGuard's `_async_log_success_streaming` reconciles end-of-stream `usage.completion_tokens` automatically. No app code change.\n", + "- **Direct (non-proxy) async callers:** `from spendguard.integrations.litellm import SpendGuardDirectAcompletion` wraps `litellm.acompletion()`. Sync `litellm.completion()` is not supported; route via the SpendGuard egress proxy for that.\n", + "- **Fail-open dev override:** `SPENDGUARD_LITELLM_FAIL_OPEN=1` lets calls through when the sidecar is unreachable (development only).\n", + "\n", + "Runnable end-to-end demo:\n", + "```bash\n", + "git clone https://github.com/m24927605/agentic-spendguard\n", + "cd agentic-spendguard\n", + "make demo-up DEMO_MODE=litellm_real # 4-step ALLOW + DENY + STREAM + MULTI-TEAM\n", + "make demo-up DEMO_MODE=litellm_deny # 3 fail-closed sub-steps\n", + "```" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python" + } + }, + "nbformat": 4, + "nbformat_minor": 2 +} From aa4ae5ff409cc1ec1f1fe01441ca21ced59cecbf Mon Sep 17 00:00:00 2001 From: Michael Chen Date: Thu, 21 May 2026 00:27:11 +0800 Subject: [PATCH 2/4] =?UTF-8?q?cookbook:=20address=20Greptile=20feedback?= =?UTF-8?q?=20=E2=80=94=20document=20amount=5Fatomic=20unit=20semantics?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds inline comment explaining that `amount_atomic='50'` is a worst-case pre-call cost in the BudgetBinding's unit (with `unit=output_token`, this means "reserve 50 output tokens upfront"). Production code should derive this from a pricing table; the reservation is reconciled to real `usage.completion_tokens` in `_reconcile` so any overshoot is refunded. Resolves Greptile review comment on the original PR. Signed-off-by: Michael Chen --- .../LiteLLM_Proxy_SpendGuard.ipynb | 27 ++++++++++++------- 1 file changed, 17 insertions(+), 10 deletions(-) diff --git a/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb b/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb index ff2cfb320a9..eb06f0e4861 100644 --- a/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb +++ b/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb @@ -4,15 +4,15 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "## Budget Guardrails — LiteLLM Proxy + SpendGuard\n", + "## Budget Guardrails \u2014 LiteLLM Proxy + SpendGuard\n", "\n", "This notebook demonstrates how to use the LiteLLM Proxy with [SpendGuard](https://github.com/m24927605/agentic-spendguard) to gate every `/v1/chat/completions` call against a budget BEFORE the upstream provider is hit.\n", "\n", "**What SpendGuard adds on top of LiteLLM's own spend tracking:**\n", - "- **Pre-call reserve** — SpendGuard's contract-DSL evaluator decides ALLOW / DENY / DEGRADE / REQUIRE_APPROVAL before LiteLLM forwards to the provider. Budget exhausted? HTTP 403, provider invoice clock never starts.\n", - "- **Post-call commit with reconciliation** — End-of-stream reconciler reads `response.usage.completion_tokens` and commits the real cost (not the worst-case estimator).\n", - "- **Signed append-only audit chain** — Every decision lands as a CloudEvent in `canonical_events`, with the 12-field LiteLLM-specific enrichment (`litellm_call_id`, `model`, `team_id`, `pricing_version`, etc) for forensics.\n", - "- **Fail-closed default** — Sidecar unreachable → HTTP 503; the request is denied, not forwarded.\n", + "- **Pre-call reserve** \u2014 SpendGuard's contract-DSL evaluator decides ALLOW / DENY / DEGRADE / REQUIRE_APPROVAL before LiteLLM forwards to the provider. Budget exhausted? HTTP 403, provider invoice clock never starts.\n", + "- **Post-call commit with reconciliation** \u2014 End-of-stream reconciler reads `response.usage.completion_tokens` and commits the real cost (not the worst-case estimator).\n", + "- **Signed append-only audit chain** \u2014 Every decision lands as a CloudEvent in `canonical_events`, with the 12-field LiteLLM-specific enrichment (`litellm_call_id`, `model`, `team_id`, `pricing_version`, etc) for forensics.\n", + "- **Fail-closed default** \u2014 Sidecar unreachable \u2192 HTTP 503; the request is denied, not forwarded.\n", "\n", "SpendGuard wires in as a standard LiteLLM `CustomLogger` callback. No LiteLLM source change required." ] @@ -78,6 +78,13 @@ " return _BINDING\n", "\n", "def _estimate(ctx):\n", + " # `amount_atomic` is the worst-case pre-call cost expressed in\n", + " # the BudgetBinding's unit. With `unit=output_token` above,\n", + " # '50' means \"reserve 50 output tokens upfront\". Production code\n", + " # should derive this from your pricing table \u2014 e.g. estimated\n", + " # max response tokens for the model. The reservation is\n", + " # reconciled to the real `usage.completion_tokens` in\n", + " # `_reconcile` below, and any overshoot is refunded.\n", " return [common_pb2.BudgetClaim(\n", " budget_id=_BINDING.budget_id, unit=_UNIT, amount_atomic='50',\n", " direction=common_pb2.BudgetClaim.DEBIT,\n", @@ -148,7 +155,7 @@ "source": [ "## 2. Make LLM Requests to the Gated Proxy\n", "\n", - "Standard OpenAI client — no SpendGuard SDK in your app code. The gating is invisible until the budget hits a limit." + "Standard OpenAI client \u2014 no SpendGuard SDK in your app code. The gating is invisible until the budget hits a limit." ] }, { @@ -164,7 +171,7 @@ " api_key=\"sk-demo-key\", # LITELLM_MASTER_KEY\n", ")\n", "\n", - "# ALLOW path — SpendGuard reserves + commits normally.\n", + "# ALLOW path \u2014 SpendGuard reserves + commits normally.\n", "response = client.chat.completions.create(\n", " model=\"gpt-4o-mini\",\n", " messages=[{\"role\": \"user\", \"content\": \"hello\"}],\n", @@ -195,7 +202,7 @@ "ORDER BY ce.event_time DESC;\n", "```\n", "\n", - "The `spendguard` JSONB sub-object on each row carries the 12-field LiteLLM-specific enrichment per [DESIGN §8.2a](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/DESIGN.md)." + "The `spendguard` JSONB sub-object on each row carries the 12-field LiteLLM-specific enrichment per [DESIGN \u00a78.2a](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/DESIGN.md)." ] }, { @@ -204,7 +211,7 @@ "source": [ "## 4. Going Further\n", "\n", - "- **Multi-tenant via virtual keys:** the operator template above is single-team. To dispatch per LiteLLM team, inspect `ctx.user_api_key_dict.team_id` in `_resolve` and look up the binding in your control plane. Full template: [PROXY_RECIPE.md §2](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/PROXY_RECIPE.md).\n", + "- **Multi-tenant via virtual keys:** the operator template above is single-team. To dispatch per LiteLLM team, inspect `ctx.user_api_key_dict.team_id` in `_resolve` and look up the binding in your control plane. Full template: [PROXY_RECIPE.md \u00a72](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/PROXY_RECIPE.md).\n", "- **Streaming:** SpendGuard's `_async_log_success_streaming` reconciles end-of-stream `usage.completion_tokens` automatically. No app code change.\n", "- **Direct (non-proxy) async callers:** `from spendguard.integrations.litellm import SpendGuardDirectAcompletion` wraps `litellm.acompletion()`. Sync `litellm.completion()` is not supported; route via the SpendGuard egress proxy for that.\n", "- **Fail-open dev override:** `SPENDGUARD_LITELLM_FAIL_OPEN=1` lets calls through when the sidecar is unreachable (development only).\n", @@ -231,4 +238,4 @@ }, "nbformat": 4, "nbformat_minor": 2 -} +} \ No newline at end of file From 0e175746f2c60fde3bb036e91a23d52a356eadd4 Mon Sep 17 00:00:00 2001 From: Michael Chen Date: Thu, 21 May 2026 00:56:46 +0800 Subject: [PATCH 3/4] chore: retrigger CI under litellm_internal_staging base PR was originally opened against `main` which trips the guard-main-branch workflow (external fork PRs must target a different branch). Retargeted via `gh pr edit --base litellm_internal_staging` per the workflow's instruction; this empty commit forces CI to re-evaluate against the new base. The "Verify PR source branch" gate only fires for PRs to `main`, so it shouldn't run on this PR anymore. Signed-off-by: Michael Chen From a002242fd6018b88cc0d986ae88cc7ede09ece54 Mon Sep 17 00:00:00 2001 From: Michael Date: Sat, 23 May 2026 03:19:35 +0800 Subject: [PATCH 4/4] =?UTF-8?q?cookbook:=20address=20Greptile=20P1=20?= =?UTF-8?q?=E2=80=94=20third-party=20integration=20banner?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Greptile flagged the cookbook as self-promotional since SpendGuard is authored by the PR contributor. This adds an explicit banner at the top of the notebook making the third-party / non-BerriAI-maintained status unambiguous to readers, matching the disclosure pattern of the other vendor-authored cookbooks in this directory (Arize / Langfuse / Lunary). Co-Authored-By: Claude Opus 4.7 (1M context) --- cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb | 2 ++ 1 file changed, 2 insertions(+) diff --git a/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb b/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb index eb06f0e4861..39770acd6ae 100644 --- a/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb +++ b/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb @@ -4,6 +4,8 @@ "cell_type": "markdown", "metadata": {}, "source": [ + "> **Third-party integration.** Not maintained by BerriAI. SpendGuard is an open-source project ([github.com/m24927605/agentic-spendguard](https://github.com/m24927605/agentic-spendguard), Apache 2.0) authored by the contributor of this cookbook. Like the Arize, Langfuse, and Lunary cookbooks in this same directory, this notebook is provided by the integration vendor.\n", + "\n", "## Budget Guardrails \u2014 LiteLLM Proxy + SpendGuard\n", "\n", "This notebook demonstrates how to use the LiteLLM Proxy with [SpendGuard](https://github.com/m24927605/agentic-spendguard) to gate every `/v1/chat/completions` call against a budget BEFORE the upstream provider is hit.\n",