From aa4ae5ff409cc1ec1f1fe01441ca21ced59cecbf Mon Sep 17 00:00:00 2001 From: Michael Chen Date: Thu, 21 May 2026 00:27:11 +0800 Subject: [PATCH] =?UTF-8?q?cookbook:=20address=20Greptile=20feedback=20?= =?UTF-8?q?=E2=80=94=20document=20amount=5Fatomic=20unit=20semantics?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds inline comment explaining that `amount_atomic='50'` is a worst-case pre-call cost in the BudgetBinding's unit (with `unit=output_token`, this means "reserve 50 output tokens upfront"). Production code should derive this from a pricing table; the reservation is reconciled to real `usage.completion_tokens` in `_reconcile` so any overshoot is refunded. Resolves Greptile review comment on the original PR. Signed-off-by: Michael Chen --- .../LiteLLM_Proxy_SpendGuard.ipynb | 27 ++++++++++++------- 1 file changed, 17 insertions(+), 10 deletions(-) diff --git a/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb b/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb index ff2cfb320a9..eb06f0e4861 100644 --- a/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb +++ b/cookbook/logging_observability/LiteLLM_Proxy_SpendGuard.ipynb @@ -4,15 +4,15 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "## Budget Guardrails — LiteLLM Proxy + SpendGuard\n", + "## Budget Guardrails \u2014 LiteLLM Proxy + SpendGuard\n", "\n", "This notebook demonstrates how to use the LiteLLM Proxy with [SpendGuard](https://github.com/m24927605/agentic-spendguard) to gate every `/v1/chat/completions` call against a budget BEFORE the upstream provider is hit.\n", "\n", "**What SpendGuard adds on top of LiteLLM's own spend tracking:**\n", - "- **Pre-call reserve** — SpendGuard's contract-DSL evaluator decides ALLOW / DENY / DEGRADE / REQUIRE_APPROVAL before LiteLLM forwards to the provider. Budget exhausted? HTTP 403, provider invoice clock never starts.\n", - "- **Post-call commit with reconciliation** — End-of-stream reconciler reads `response.usage.completion_tokens` and commits the real cost (not the worst-case estimator).\n", - "- **Signed append-only audit chain** — Every decision lands as a CloudEvent in `canonical_events`, with the 12-field LiteLLM-specific enrichment (`litellm_call_id`, `model`, `team_id`, `pricing_version`, etc) for forensics.\n", - "- **Fail-closed default** — Sidecar unreachable → HTTP 503; the request is denied, not forwarded.\n", + "- **Pre-call reserve** \u2014 SpendGuard's contract-DSL evaluator decides ALLOW / DENY / DEGRADE / REQUIRE_APPROVAL before LiteLLM forwards to the provider. Budget exhausted? HTTP 403, provider invoice clock never starts.\n", + "- **Post-call commit with reconciliation** \u2014 End-of-stream reconciler reads `response.usage.completion_tokens` and commits the real cost (not the worst-case estimator).\n", + "- **Signed append-only audit chain** \u2014 Every decision lands as a CloudEvent in `canonical_events`, with the 12-field LiteLLM-specific enrichment (`litellm_call_id`, `model`, `team_id`, `pricing_version`, etc) for forensics.\n", + "- **Fail-closed default** \u2014 Sidecar unreachable \u2192 HTTP 503; the request is denied, not forwarded.\n", "\n", "SpendGuard wires in as a standard LiteLLM `CustomLogger` callback. No LiteLLM source change required." ] @@ -78,6 +78,13 @@ " return _BINDING\n", "\n", "def _estimate(ctx):\n", + " # `amount_atomic` is the worst-case pre-call cost expressed in\n", + " # the BudgetBinding's unit. With `unit=output_token` above,\n", + " # '50' means \"reserve 50 output tokens upfront\". Production code\n", + " # should derive this from your pricing table \u2014 e.g. estimated\n", + " # max response tokens for the model. The reservation is\n", + " # reconciled to the real `usage.completion_tokens` in\n", + " # `_reconcile` below, and any overshoot is refunded.\n", " return [common_pb2.BudgetClaim(\n", " budget_id=_BINDING.budget_id, unit=_UNIT, amount_atomic='50',\n", " direction=common_pb2.BudgetClaim.DEBIT,\n", @@ -148,7 +155,7 @@ "source": [ "## 2. Make LLM Requests to the Gated Proxy\n", "\n", - "Standard OpenAI client — no SpendGuard SDK in your app code. The gating is invisible until the budget hits a limit." + "Standard OpenAI client \u2014 no SpendGuard SDK in your app code. The gating is invisible until the budget hits a limit." ] }, { @@ -164,7 +171,7 @@ " api_key=\"sk-demo-key\", # LITELLM_MASTER_KEY\n", ")\n", "\n", - "# ALLOW path — SpendGuard reserves + commits normally.\n", + "# ALLOW path \u2014 SpendGuard reserves + commits normally.\n", "response = client.chat.completions.create(\n", " model=\"gpt-4o-mini\",\n", " messages=[{\"role\": \"user\", \"content\": \"hello\"}],\n", @@ -195,7 +202,7 @@ "ORDER BY ce.event_time DESC;\n", "```\n", "\n", - "The `spendguard` JSONB sub-object on each row carries the 12-field LiteLLM-specific enrichment per [DESIGN §8.2a](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/DESIGN.md)." + "The `spendguard` JSONB sub-object on each row carries the 12-field LiteLLM-specific enrichment per [DESIGN \u00a78.2a](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/DESIGN.md)." ] }, { @@ -204,7 +211,7 @@ "source": [ "## 4. Going Further\n", "\n", - "- **Multi-tenant via virtual keys:** the operator template above is single-team. To dispatch per LiteLLM team, inspect `ctx.user_api_key_dict.team_id` in `_resolve` and look up the binding in your control plane. Full template: [PROXY_RECIPE.md §2](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/PROXY_RECIPE.md).\n", + "- **Multi-tenant via virtual keys:** the operator template above is single-team. To dispatch per LiteLLM team, inspect `ctx.user_api_key_dict.team_id` in `_resolve` and look up the binding in your control plane. Full template: [PROXY_RECIPE.md \u00a72](https://github.com/m24927605/agentic-spendguard/blob/main/docs/specs/litellm-integration/PROXY_RECIPE.md).\n", "- **Streaming:** SpendGuard's `_async_log_success_streaming` reconciles end-of-stream `usage.completion_tokens` automatically. No app code change.\n", "- **Direct (non-proxy) async callers:** `from spendguard.integrations.litellm import SpendGuardDirectAcompletion` wraps `litellm.acompletion()`. Sync `litellm.completion()` is not supported; route via the SpendGuard egress proxy for that.\n", "- **Fail-open dev override:** `SPENDGUARD_LITELLM_FAIL_OPEN=1` lets calls through when the sidecar is unreachable (development only).\n", @@ -231,4 +238,4 @@ }, "nbformat": 4, "nbformat_minor": 2 -} +} \ No newline at end of file