From d62ab8f185ad44b5cc7ee4f31364dfed77c87bd1 Mon Sep 17 00:00:00 2001 From: Fabro Date: Thu, 21 May 2026 19:10:15 -0400 Subject: [PATCH] =?UTF-8?q?checkpoint=20=E2=9A=92=EF=B8=8F=20Generated=20w?= =?UTF-8?q?ith=20[Fabro](https://fabro.sh)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- run.json | 247 ++++++++++++++-- stages/005-implement@1/diff.patch | 279 ++++++++++++++++++ stages/005-implement@1/status.json | 6 + stages/006-simplify_opus@1/prompt.md | 156 ++++++++++ stages/006-simplify_opus@1/provider_used.json | 5 + stages/006-simplify_opus@1/response.md | 30 ++ 6 files changed, 703 insertions(+), 20 deletions(-) create mode 100644 stages/005-implement@1/diff.patch create mode 100644 stages/005-implement@1/status.json create mode 100644 stages/006-simplify_opus@1/prompt.md create mode 100644 stages/006-simplify_opus@1/provider_used.json create mode 100644 stages/006-simplify_opus@1/response.md diff --git a/run.json b/run.json index a2e62c376..de4f16877 100644 --- a/run.json +++ b/run.json @@ -516,7 +516,7 @@ "kind": "running" }, "status_updated_at": "2026-05-21T22:36:05.837824Z", - "last_event_at": "2026-05-21T22:59:35.494606Z", + "last_event_at": "2026-05-21T23:10:14.247857Z", "pending_control": null, "checkpoints": [ { @@ -764,9 +764,9 @@ } }, { - "seq": 0, + "seq": 252, "checkpoint": { - "timestamp": "2026-05-21T22:59:35.632255Z", + "timestamp": "2026-05-21T22:59:40.211057Z", "current_node": "implement", "completed_nodes": [ "start", @@ -777,31 +777,163 @@ ], "node_retries": {}, "context_values": { - "graph.rankdir": "LR", - "internal.node_visit_count": 1, - "thread.toolchain.current_node": "preflight_compile", - "thread.preflight_lint.current_node": "implement", - "outcome": "succeeded", - "last_response": "All changes implemented as specified in the plan:\n\n**1. `lib/crates/fabro-model/src/billing.rs`** ✅\n- Added `BilledTokenCounts::token_counts(&self) -> TokenCounts` (drops `total_tokens`/`total_usd_m", - "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126", + "internal.retry_count.preflight_lint": 0, "internal.thread_id": "preflight_lint", + "last_response": "All changes implemented as specified in the plan:\n\n**1. `lib/crates/fabro-model/src/billing.rs`** ✅\n- Added `BilledTokenCounts::token_counts(&self) -> TokenCounts` (drops `total_tokens`/`total_usd_m", + "outcome": "succeeded", + "current_node": "implement", + "graph.rankdir": "LR", + "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126", + "internal.node_visit_count": 1, + "internal.retry_count.implement": 0, + "thread.preflight_lint.current_node": "implement", + "response.implement": "All changes implemented as specified in the plan:\n\n**1. `lib/crates/fabro-model/src/billing.rs`** ✅\n- Added `BilledTokenCounts::token_counts(&self) -> TokenCounts` (drops `total_tokens`/`total_usd_micros`).\n- Added `Catalog::price_tokens(&self, model, tokens) -> Option` that mirrors the cost lines of `billed_model_usage_from_llm`.\n\n**2. `lib/crates/fabro-workflow/src/billing_rollup.rs`** ✅\n- Changed `billing_rollup_from_projection` signature to take `catalog: Option<&Catalog>`.\n- Added private helper `stage_usage_with_cost` that clones `stage.usage` and prices it when `total_usd_micros.is_none()` and both `catalog` and `stage.model` are available.\n- Uses `priced` in place of `&stage.usage` for the `is_zero` check, `row.billing.add_counts`, `totals.add_counts`, and `model_entry.billing.add_counts`.\n- Updated existing tests to pass `None`; added a new `rollup_prices_in_flight_stage_usage_using_catalog` test (in-flight stage with no completion, non-zero usage, builtin model → `Some(..)` cost on stage row, totals, and `by_model`).\n\n**3. Call sites** ✅\n- `fabro-server/src/server/handler/billing.rs:82` — binds `let catalog = state.catalog();` and passes `Some(&catalog)`.\n- `fabro-server/src/server.rs` (2 aggregate-billing sites) — pass `None`.\n- `fabro-workflow/src/pipeline/finalize.rs` (4 sites) — pass `None`.\n\n**Verification** ✅\n- `cargo nextest run -p fabro-workflow billing_rollup` — 4 tests pass (new + 3 updated).\n- `cargo nextest run -p fabro-server billing` — 7 tests pass.\n- `cargo nextest run -p fabro-model billing` — 23 tests pass.\n- `cargo check --workspace --all-targets` — clean.\n- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` — clean.\n- `cargo +nightly-2026-04-14 fmt --check --all` — clean.\n- Broader test sweep on `fabro-workflow` + `fabro-server`: 1674 passing; the 2 unrelated `*graph*svg*` failures are pre-existing graphviz subprocess env issues (confirmed by stashing my changes and reproducing).", "graph.goal": "# Plan: Compute LLM cost on-read for in-flight stages\n\n## Context\n\nOn the run billing page (`/runs/{id}/billing`), an active stage shows token\nusage but no dollar cost — cost renders as `—` until the stage completes.\n\nRoot cause: while a stage runs, `AgentMessage` events carry usage built by\n`billed_token_counts_from_llm` (`fabro-workflow/src/outcome.rs:43`), which\nhard-codes `total_usd_micros: None`. Dollar cost is only computed by\n`billed_model_usage_from_llm` (`outcome.rs:14`) — which needs the pricing\n`Catalog` — and that runs only on `StageCompleted`/`PromptCompleted`/`StageFailed`.\nSo an in-flight stage's `StageProjection.usage.total_usd_micros` stays `None`.\n\nFix: price stages whose cost is `None` when the billing rollup is built for a\nread request, using the model + token counts already in the projection. The\nwire contract is unchanged (`total_usd_micros` is already nullable everywhere)\nand the frontend already renders whatever value comes back — no UI change.\n\n## Decisions\n\n- **Price any stage with `total_usd_micros == None`**, not just in-flight ones.\n Completed stages with unpriceable providers (`BillingPolicy::None`) return\n `None` again — harmless; no need to thread `StageState`.\n- **No \"estimated\" label.** Cost-so-far is exact for tokens consumed so far,\n matching the already-unlabeled live token counts and ticking runtime.\n- **Aggregate billing stays finalized-only.** The `BillingAccumulator` call\n sites pass `None` so a run's running estimate is never folded into org-wide\n totals (avoids double-count when the run later finalizes).\n- Per-stage rows, `totals`, and `by_model` are all priced from the same source\n so the billing page stays internally consistent.\n\n## Changes\n\n### 1. `lib/crates/fabro-model/src/billing.rs`\n\n- Add `BilledTokenCounts::token_counts(&self) -> TokenCounts` — drops\n `total_tokens`/`total_usd_micros`, keeps the five disjoint buckets.\n- Add `Catalog::price_tokens(&self, model: &ModelRef, tokens: &TokenCounts) -> Option`\n next to `pricing_for`/`billing_facts_for`. Body mirrors the cost lines of\n `billed_model_usage_from_llm`: build `ModelBillingFacts` via\n `billing_facts_for`, assemble `ModelBillingInput { ModelUsage { model, tokens }, facts }`,\n then `pricing_for(model).and_then(|p| p.bill(&input)).map(|a| a.0)`. Returns\n `None` when the provider has no billing policy.\n\n### 2. `lib/crates/fabro-workflow/src/billing_rollup.rs`\n\n- Change signature to\n `billing_rollup_from_projection(projection: &RunProjection, catalog: Option<&Catalog>)`.\n- Add a module-private helper `stage_usage_with_cost(catalog, stage) -> BilledTokenCounts`:\n clone `stage.usage`; if `total_usd_micros.is_none()` and both `catalog` and\n `stage.model` are present, set it via `catalog.price_tokens(model, &usage.token_counts())`.\n- In the loop, compute `priced` once per stage and use it in place of\n `&stage.usage` for the `is_zero` check, `row.billing.add_counts`,\n `totals.add_counts`, and `model_entry.billing.add_counts`.\n- Update the existing tests to pass `None`; add one new test: an in-flight\n stage (no `completion`, non-zero `usage` with `total_usd_micros: None`, a\n builtin `model`) yields `Some(..)` cost on the stage row and in `totals` when\n called with `Some(Catalog::builtin())`.\n\n### 3. Call sites of `billing_rollup_from_projection`\n\n- `lib/crates/fabro-server/src/server/handler/billing.rs:82` — bind\n `let catalog = state.catalog();` (returns `Arc`) and pass\n `Some(&catalog)`.\n- `lib/crates/fabro-server/src/server.rs` (2 aggregate-billing sites) — pass `None`.\n- `lib/crates/fabro-workflow/src/pipeline/finalize.rs` (4 sites) — pass `None`\n (stages already priced at completion; pricing would be a no-op anyway).\n\nNo changes to `fabro-api.yaml`, the generated clients, or `apps/fabro-web`.\n\n## Out of scope / known limitation\n\nIn-flight **prompt** stages have no `model` until `PromptCompleted` (only\n`AgentMessage` sets `stage.model` mid-run), so they still show `—` while\nrunning. Acceptable: prompt stages are a single short LLM call. The bug report\nconcerns agent stages, where `model` is available.\n\n## Verification\n\n- `cargo nextest run -p fabro-workflow billing_rollup` — new + updated unit tests pass.\n- `cargo nextest run -p fabro-server billing` — handler conformance still passes.\n- `cargo build --workspace` and `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings`.\n- Manual: `fabro server start` + `cd apps/fabro-web && bun run dev`, start a\n workflow with an agent stage, open `/runs/{id}/billing` mid-run — the active\n stage row and totals show a non-`—` dollar amount that grows with tokens.\n", "internal.work_dir": "/home/daytona/workspace/fabro", "thread.start.current_node": "toolchain", - "internal.retry_count.toolchain": 0, - "internal.retry_count.implement": 0, "failure_class": "", + "internal.retry_count.toolchain": 0, + "internal.fidelity": "compact", + "internal.run_id": "01KS6AW28FZVV4M2EHBJA7JMNP", + "failure_signature": "", + "thread.toolchain.current_node": "preflight_compile", + "graph.model_stylesheet": "\n * { model: claude-opus-4-7; }\n ", "last_stage": "implement", + "thread.preflight_compile.current_node": "preflight_lint", + "internal.retry_count.preflight_compile": 0, + "internal.retry_count.start": 0 + }, + "node_outcomes": { + "preflight_lint": { + "status": "succeeded", + "context_updates": { + "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126" + }, + "notes": "Script completed: cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", + "usage": null + }, + "start": { + "status": "succeeded", + "usage": null + }, + "toolchain": { + "status": "succeeded", + "context_updates": { + "command.output": "blob://sha256/fc14b2ba2d770e5cd3169df7a29525c962adfc4cfa3097b9098c63ebd61a748c" + }, + "notes": "Script completed: command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1", + "usage": null + }, + "implement": { + "status": "succeeded", + "context_updates": { + "response.implement": "All changes implemented as specified in the plan:\n\n**1. `lib/crates/fabro-model/src/billing.rs`** ✅\n- Added `BilledTokenCounts::token_counts(&self) -> TokenCounts` (drops `total_tokens`/`total_usd_micros`).\n- Added `Catalog::price_tokens(&self, model, tokens) -> Option` that mirrors the cost lines of `billed_model_usage_from_llm`.\n\n**2. `lib/crates/fabro-workflow/src/billing_rollup.rs`** ✅\n- Changed `billing_rollup_from_projection` signature to take `catalog: Option<&Catalog>`.\n- Added private helper `stage_usage_with_cost` that clones `stage.usage` and prices it when `total_usd_micros.is_none()` and both `catalog` and `stage.model` are available.\n- Uses `priced` in place of `&stage.usage` for the `is_zero` check, `row.billing.add_counts`, `totals.add_counts`, and `model_entry.billing.add_counts`.\n- Updated existing tests to pass `None`; added a new `rollup_prices_in_flight_stage_usage_using_catalog` test (in-flight stage with no completion, non-zero usage, builtin model → `Some(..)` cost on stage row, totals, and `by_model`).\n\n**3. Call sites** ✅\n- `fabro-server/src/server/handler/billing.rs:82` — binds `let catalog = state.catalog();` and passes `Some(&catalog)`.\n- `fabro-server/src/server.rs` (2 aggregate-billing sites) — pass `None`.\n- `fabro-workflow/src/pipeline/finalize.rs` (4 sites) — pass `None`.\n\n**Verification** ✅\n- `cargo nextest run -p fabro-workflow billing_rollup` — 4 tests pass (new + 3 updated).\n- `cargo nextest run -p fabro-server billing` — 7 tests pass.\n- `cargo nextest run -p fabro-model billing` — 23 tests pass.\n- `cargo check --workspace --all-targets` — clean.\n- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` — clean.\n- `cargo +nightly-2026-04-14 fmt --check --all` — clean.\n- Broader test sweep on `fabro-workflow` + `fabro-server`: 1674 passing; the 2 unrelated `*graph*svg*` failures are pre-existing graphviz subprocess env issues (confirmed by stashing my changes and reproducing).", + "last_response": "All changes implemented as specified in the plan:\n\n**1. `lib/crates/fabro-model/src/billing.rs`** ✅\n- Added `BilledTokenCounts::token_counts(&self) -> TokenCounts` (drops `total_tokens`/`total_usd_m", + "last_stage": "implement" + }, + "notes": "Stage completed: implement", + "usage": { + "input": { + "usage": { + "model": { + "provider": "anthropic", + "model_id": "claude-opus-4-7" + }, + "tokens": { + "input_tokens": 92896, + "output_tokens": 16017, + "reasoning_tokens": 0, + "cache_read_tokens": 4690781, + "cache_write_tokens": 101072 + } + }, + "facts": { + "algorithm": "anthropic", + "cache_write_5m_tokens": 101072, + "cache_write_1h_tokens": 0 + } + }, + "total_usd_micros": 3841995 + }, + "files_touched": [ + "/home/daytona/workspace/fabro/lib/crates/fabro-model/src/billing.rs", + "/home/daytona/workspace/fabro/lib/crates/fabro-server/src/server.rs", + "/home/daytona/workspace/fabro/lib/crates/fabro-server/src/server/handler/billing.rs", + "/home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/billing_rollup.rs", + "/home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/pipeline/finalize.rs" + ] + }, + "preflight_compile": { + "status": "succeeded", + "context_updates": { + "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126" + }, + "notes": "Script completed: cargo check -q --workspace 2>&1", + "usage": null + } + }, + "next_node_id": "simplify_opus", + "git_commit_sha": "01ee56ee8d0ff489ecf62f786fc099bde268fd57", + "node_visits": { + "preflight_lint": 1, + "preflight_compile": 1, + "start": 1, + "implement": 1, + "toolchain": 1 + } + }, + "diff": { + "patch": "diff --git a/lib/crates/fabro-model/src/billing.rs b/lib/crates/fabro-model/src/billing.rs\nindex 29cad2335..f06cad06a 100644\n--- a/lib/crates/fabro-model/src/billing.rs\n+++ b/lib/crates/fabro-model/src/billing.rs\n@@ -358,6 +358,19 @@ impl BilledTokenCounts {\n }\n }\n \n+ /// Returns the five disjoint per-call token buckets, dropping the derived\n+ /// `total_tokens` sum and the optional `total_usd_micros` cost.\n+ #[must_use]\n+ pub fn token_counts(&self) -> TokenCounts {\n+ TokenCounts {\n+ input_tokens: self.input_tokens,\n+ output_tokens: self.output_tokens,\n+ reasoning_tokens: self.reasoning_tokens,\n+ cache_read_tokens: self.cache_read_tokens,\n+ cache_write_tokens: self.cache_write_tokens,\n+ }\n+ }\n+\n pub fn add_counts(&mut self, source: &Self) {\n self.input_tokens += source.input_tokens;\n self.output_tokens += source.output_tokens;\n@@ -435,6 +448,27 @@ impl Catalog {\n self.provider(&model_ref.provider)\n .and_then(|provider| ModelBillingFacts::for_policy(provider.billing_policy, tokens))\n }\n+\n+ /// Price a partial token sample for `model` using catalog pricing.\n+ ///\n+ /// Returns `None` when the provider has no billing policy, the model is\n+ /// unknown, or the pricing algorithm cannot produce a result for the given\n+ /// tokens. Used by read-side rollups so in-flight stages can show an\n+ /// exact cost for the tokens consumed so far.\n+ #[must_use]\n+ pub fn price_tokens(&self, model: &ModelRef, tokens: &TokenCounts) -> Option {\n+ let facts = self.billing_facts_for(model, tokens)?;\n+ let input = ModelBillingInput {\n+ usage: ModelUsage {\n+ model: model.clone(),\n+ tokens: tokens.clone(),\n+ },\n+ facts,\n+ };\n+ self.pricing_for(model)\n+ .and_then(|pricing| pricing.bill(&input))\n+ .map(|amount| amount.0)\n+ }\n }\n \n fn costs_for_speed(\ndiff --git a/lib/crates/fabro-server/src/server.rs b/lib/crates/fabro-server/src/server.rs\nindex b5258fa65..731c3f926 100644\n--- a/lib/crates/fabro-server/src/server.rs\n+++ b/lib/crates/fabro-server/src/server.rs\n@@ -3238,7 +3238,7 @@ async fn execute_run_in_process(state: Arc, run_id: RunId) {\n .expect(\"aggregate_billing lock poisoned\");\n accumulate_billing_rollup(\n &mut agg,\n- &fabro_workflow::billing_rollup_from_projection(projection),\n+ &fabro_workflow::billing_rollup_from_projection(projection, None),\n );\n }\n }\n@@ -3523,7 +3523,7 @@ async fn execute_run_subprocess(state: Arc, run_id: RunId) {\n .expect(\"aggregate_billing lock poisoned\");\n accumulate_billing_rollup(\n &mut agg,\n- &fabro_workflow::billing_rollup_from_projection(&final_state),\n+ &fabro_workflow::billing_rollup_from_projection(&final_state, None),\n );\n }\n \ndiff --git a/lib/crates/fabro-server/src/server/handler/billing.rs b/lib/crates/fabro-server/src/server/handler/billing.rs\nindex 6eee9bbfb..cbaf7cc3e 100644\n--- a/lib/crates/fabro-server/src/server/handler/billing.rs\n+++ b/lib/crates/fabro-server/src/server/handler/billing.rs\n@@ -79,7 +79,8 @@ async fn get_run_billing(\n };\n let projection = cached.projection;\n \n- let rollup = fabro_workflow::billing_rollup_from_projection(&projection);\n+ let catalog = state.catalog();\n+ let rollup = fabro_workflow::billing_rollup_from_projection(&projection, Some(&catalog));\n let by_model = rollup\n .by_model\n .iter()\ndiff --git a/lib/crates/fabro-workflow/src/billing_rollup.rs b/lib/crates/fabro-workflow/src/billing_rollup.rs\nindex a91b06061..89791caa7 100644\n--- a/lib/crates/fabro-workflow/src/billing_rollup.rs\n+++ b/lib/crates/fabro-workflow/src/billing_rollup.rs\n@@ -1,6 +1,7 @@\n use std::collections::HashMap;\n \n-use fabro_types::{BilledTokenCounts, ModelRef, RunProjection};\n+use fabro_model::Catalog;\n+use fabro_types::{BilledTokenCounts, ModelRef, RunProjection, StageProjection};\n \n #[derive(Debug, Clone, PartialEq)]\n pub struct ProjectionBillingStage {\n@@ -34,7 +35,10 @@ impl ProjectionBillingRollup {\n }\n \n #[must_use]\n-pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionBillingRollup {\n+pub fn billing_rollup_from_projection(\n+ projection: &RunProjection,\n+ catalog: Option<&Catalog>,\n+) -> ProjectionBillingRollup {\n let mut stage_indices = HashMap::::new();\n let mut stages = Vec::::new();\n let mut by_model = HashMap::::new();\n@@ -46,7 +50,8 @@ pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionB\n if is_boundary_stage(projection, stage_id.node_id()) {\n continue;\n }\n- if stage.completion.is_none() && stage.duration_ms.is_none() && stage.usage.is_zero() {\n+ let priced = stage_usage_with_cost(catalog, stage);\n+ if stage.completion.is_none() && stage.duration_ms.is_none() && priced.is_zero() {\n continue;\n }\n \n@@ -68,10 +73,10 @@ pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionB\n runtime_ms = runtime_ms.saturating_add(duration_ms);\n }\n \n- if !stage.usage.is_zero() {\n+ if !priced.is_zero() {\n billed_visit_count += 1;\n- row.billing.add_counts(&stage.usage);\n- totals.add_counts(&stage.usage);\n+ row.billing.add_counts(&priced);\n+ totals.add_counts(&priced);\n \n if let Some(model) = &stage.model {\n row.model = Some(model.clone());\n@@ -84,7 +89,7 @@ pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionB\n billing: BilledTokenCounts::default(),\n });\n model_entry.stages += 1;\n- model_entry.billing.add_counts(&stage.usage);\n+ model_entry.billing.add_counts(&priced);\n }\n }\n }\n@@ -113,6 +118,16 @@ pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionB\n }\n }\n \n+fn stage_usage_with_cost(catalog: Option<&Catalog>, stage: &StageProjection) -> BilledTokenCounts {\n+ let mut usage = stage.usage.clone();\n+ if usage.total_usd_micros.is_none() {\n+ if let (Some(catalog), Some(model)) = (catalog, stage.model.as_ref()) {\n+ usage.total_usd_micros = catalog.price_tokens(model, &usage.token_counts());\n+ }\n+ }\n+ usage\n+}\n+\n fn is_boundary_stage(projection: &RunProjection, node_id: &str) -> bool {\n projection\n .spec()\n@@ -126,6 +141,7 @@ fn is_boundary_stage(projection: &RunProjection, node_id: &str) -> bool {\n mod tests {\n use std::collections::HashMap;\n \n+ use fabro_model::{Catalog, ModelRef, ProviderId};\n use fabro_types::{\n AttrValue, BilledModelUsage, BilledTokenCounts, Graph, Node, RunProjection, RunSpec,\n StageCompletion, StageOutcome, WorkflowSettings, first_event_seq, fixtures,\n@@ -190,7 +206,7 @@ mod tests {\n timestamp: chrono::Utc::now(),\n });\n \n- let rollup = billing_rollup_from_projection(&projection);\n+ let rollup = billing_rollup_from_projection(&projection, None);\n \n assert_eq!(rollup.stages.len(), 1);\n assert_eq!(rollup.stages[0].node_id, \"verify\");\n@@ -233,7 +249,7 @@ mod tests {\n timestamp: chrono::Utc::now(),\n });\n \n- let rollup = billing_rollup_from_projection(&projection);\n+ let rollup = billing_rollup_from_projection(&projection, None);\n \n assert_eq!(rollup.stages.len(), 1);\n assert_eq!(rollup.stages[0].node_id, \"build\");\n@@ -266,12 +282,48 @@ mod tests {\n timestamp: chrono::Utc::now(),\n });\n \n- let rollup = billing_rollup_from_projection(&projection);\n+ let rollup = billing_rollup_from_projection(&projection, None);\n \n assert_eq!(rollup.stages.len(), 0);\n assert_eq!(rollup.runtime_ms, 0);\n }\n \n+ #[test]\n+ fn rollup_prices_in_flight_stage_usage_using_catalog() {\n+ let mut projection = test_projection();\n+ let model = ModelRef {\n+ provider: ProviderId::openai(),\n+ model_id: \"gpt-5.4\".to_string(),\n+ speed: None,\n+ };\n+ let stage = projection.stage_entry(\"agent\", 1, first_event_seq(1));\n+ stage.started_at = Some(chrono::Utc::now());\n+ stage.usage = BilledTokenCounts {\n+ input_tokens: 500_000,\n+ output_tokens: 125_000,\n+ total_tokens: 625_000,\n+ ..BilledTokenCounts::default()\n+ };\n+ stage.model = Some(model.clone());\n+\n+ let priced = billing_rollup_from_projection(&projection, Some(Catalog::builtin()));\n+ let unpriced = billing_rollup_from_projection(&projection, None);\n+\n+ assert_eq!(priced.stages.len(), 1);\n+ assert_eq!(priced.stages[0].node_id, \"agent\");\n+ let stage_cost = priced.stages[0].billing.total_usd_micros;\n+ assert!(\n+ stage_cost.is_some_and(|cost| cost > 0),\n+ \"expected priced stage cost, got {stage_cost:?}\"\n+ );\n+ assert_eq!(priced.totals.total_usd_micros, stage_cost);\n+ assert_eq!(priced.by_model.len(), 1);\n+ assert_eq!(priced.by_model[0].billing.total_usd_micros, stage_cost);\n+ assert_eq!(unpriced.stages.len(), 1);\n+ assert_eq!(unpriced.stages[0].billing.total_usd_micros, None);\n+ assert_eq!(unpriced.totals.total_usd_micros, None);\n+ }\n+\n fn run_spec_with_boundary_nodes() -> RunSpec {\n let mut graph = Graph::new(\"test\");\n graph.nodes.insert(\"start\".to_string(), {\ndiff --git a/lib/crates/fabro-workflow/src/pipeline/finalize.rs b/lib/crates/fabro-workflow/src/pipeline/finalize.rs\nindex 5345bfb4a..a7f4619db 100644\n--- a/lib/crates/fabro-workflow/src/pipeline/finalize.rs\n+++ b/lib/crates/fabro-workflow/src/pipeline/finalize.rs\n@@ -82,7 +82,7 @@ pub(crate) async fn build_conclusion_from_store(\n .unwrap_or_default();\n let projection_billing = projection\n .as_ref()\n- .map(billing_rollup_from_projection)\n+ .map(|projection| billing_rollup_from_projection(projection, None))\n .unwrap_or_default();\n let checkpoint = projection\n .as_ref()\n@@ -438,7 +438,7 @@ async fn compute_final_patch(\n }\n \n pub(crate) fn billing_from_projection(projection: &RunProjection) -> Option {\n- billing_rollup_from_projection(projection).billing_if_present()\n+ billing_rollup_from_projection(projection, None).billing_if_present()\n }\n \n pub(crate) fn build_terminal_event(\n@@ -554,7 +554,7 @@ pub async fn finalize(executed: Executed, options: &FinalizeOptions) -> Result TokenCounts` — drops\n `total_tokens`/`total_usd_micros`, keeps the five disjoint buckets.\n- Add `Catalog::price_tokens(&self, model: &ModelRef, tokens: &TokenCounts) -> Option`\n next to `pricing_for`/`billing_facts_for`. Body mirrors the cost lines of\n `billed_model_usage_from_llm`: build `ModelBillingFacts` via\n `billing_facts_for`, assemble `ModelBillingInput { ModelUsage { model, tokens }, facts }`,\n then `pricing_for(model).and_then(|p| p.bill(&input)).map(|a| a.0)`. Returns\n `None` when the provider has no billing policy.\n\n### 2. `lib/crates/fabro-workflow/src/billing_rollup.rs`\n\n- Change signature to\n `billing_rollup_from_projection(projection: &RunProjection, catalog: Option<&Catalog>)`.\n- Add a module-private helper `stage_usage_with_cost(catalog, stage) -> BilledTokenCounts`:\n clone `stage.usage`; if `total_usd_micros.is_none()` and both `catalog` and\n `stage.model` are present, set it via `catalog.price_tokens(model, &usage.token_counts())`.\n- In the loop, compute `priced` once per stage and use it in place of\n `&stage.usage` for the `is_zero` check, `row.billing.add_counts`,\n `totals.add_counts`, and `model_entry.billing.add_counts`.\n- Update the existing tests to pass `None`; add one new test: an in-flight\n stage (no `completion`, non-zero `usage` with `total_usd_micros: None`, a\n builtin `model`) yields `Some(..)` cost on the stage row and in `totals` when\n called with `Some(Catalog::builtin())`.\n\n### 3. Call sites of `billing_rollup_from_projection`\n\n- `lib/crates/fabro-server/src/server/handler/billing.rs:82` — bind\n `let catalog = state.catalog();` (returns `Arc`) and pass\n `Some(&catalog)`.\n- `lib/crates/fabro-server/src/server.rs` (2 aggregate-billing sites) — pass `None`.\n- `lib/crates/fabro-workflow/src/pipeline/finalize.rs` (4 sites) — pass `None`\n (stages already priced at completion; pricing would be a no-op anyway).\n\nNo changes to `fabro-api.yaml`, the generated clients, or `apps/fabro-web`.\n\n## Out of scope / known limitation\n\nIn-flight **prompt** stages have no `model` until `PromptCompleted` (only\n`AgentMessage` sets `stage.model` mid-run), so they still show `—` while\nrunning. Acceptable: prompt stages are a single short LLM call. The bug report\nconcerns agent stages, where `model` is available.\n\n## Verification\n\n- `cargo nextest run -p fabro-workflow billing_rollup` — new + updated unit tests pass.\n- `cargo nextest run -p fabro-server billing` — handler conformance still passes.\n- `cargo build --workspace` and `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings`.\n- Manual: `fabro server start` + `cd apps/fabro-web && bun run dev`, start a\n workflow with an agent stage, open `/runs/{id}/billing` mid-run — the active\n stage row and totals show a non-`—` dollar amount that grows with tokens.\n", + "thread.start.current_node": "toolchain", + "thread.implement.current_node": "simplify_opus", + "last_stage": "simplify_opus", "internal.retry_count.start": 0, "failure_signature": "", + "graph.model_stylesheet": "\n * { model: claude-opus-4-7; }\n ", + "current_node": "simplify_opus", + "internal.retry_count.preflight_lint": 0, + "thread.preflight_lint.current_node": "implement", + "internal.fidelity": "compact", + "last_response": "All clean. Here's the summary:\n\n## Summary\n\nThree review agents (reuse, quality, efficiency) ran in parallel against the diff. Aggregated findings:\n\n### Fixed\n\n1. **Renamed `BilledTokenCounts::token_c", + "thread.preflight_compile.current_node": "preflight_lint", + "response.simplify_opus": "All clean. Here's the summary:\n\n## Summary\n\nThree review agents (reuse, quality, efficiency) ran in parallel against the diff. Aggregated findings:\n\n### Fixed\n\n1. **Renamed `BilledTokenCounts::token_counts()` → `to_token_counts()`** (`fabro-model/src/billing.rs`)\n - Quality finding: original name read like a field accessor and hid the fact that the conversion is lossy (drops `total_tokens` and `total_usd_micros`). The `to_` prefix is the established Rust convention for lossy conversions.\n\n2. **Moved pricing logic into `BilledTokenCounts::ensure_priced(catalog, model)`** (`fabro-model/src/billing.rs` + `fabro-workflow/src/billing_rollup.rs`)\n - Quality finding: the previous `stage_usage_with_cost` helper cloned `BilledTokenCounts` across a crate boundary and directly poked the private `total_usd_micros` field, leaking knowledge of how to price a `BilledTokenCounts` out of `fabro-model`. The new method encapsulates that logic next to `add_billed_usage` / `replace_with_billed_usage`, and the rollup-side helper is now a 3-line wrapper.\n\n### Considered and skipped\n\n- **Reuse #1 (refactor `billed_model_usage_from_llm` to use `price_tokens`)** — the \"duplicated\" pricing tail is 3 lines and the surrounding error-vs-option shapes don't compose cleanly. Refactoring adds more indirection than it removes.\n- **`Option<&Catalog>` parameter sprawl** — 6 of 7 call sites pass `None`, but splitting into two functions doubles the public surface for marginal benefit; the single signature is fine.\n- **`pricing_for` memoization across stages** — real but minor (a few `String` clones per stage on a polled endpoint). Worth doing if the path ever shows up in profiles; not justified at the cost of HashMap setup for this fix.\n- **`Cow::Borrowed` for the usage clone** — `BilledTokenCounts` is ~64 bytes of POD; no heap allocation.\n- **Bare `i64` return from `price_tokens`** — matches the established `total_usd_micros: Option` convention used throughout the codebase.\n\n### Verified\n\n- `cargo nextest run -p fabro-workflow billing_rollup` — 4/4 pass (including the new in-flight pricing test).\n- `cargo nextest run -p fabro-model billed_token_counts` — 5/5 pass.\n- `cargo nextest run -p fabro-server billing` — 7/7 pass.\n- `cargo build --workspace` — clean.\n- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` — clean.\n- `cargo +nightly-2026-04-14 fmt --check --all` — clean.", + "outcome": "succeeded", + "command.output": "blob://sha256/12ae32cb1ec02d01eda3581b127c1fee3b0dc53572ed6baf239721a03d82e126", + "internal.thread_id": "implement", + "internal.retry_count.toolchain": 0, + "internal.work_dir": "/home/daytona/workspace/fabro", + "internal.retry_count.simplify_opus": 0, + "failure_class": "", "response.implement": "All changes implemented as specified in the plan:\n\n**1. `lib/crates/fabro-model/src/billing.rs`** ✅\n- Added `BilledTokenCounts::token_counts(&self) -> TokenCounts` (drops `total_tokens`/`total_usd_micros`).\n- Added `Catalog::price_tokens(&self, model, tokens) -> Option` that mirrors the cost lines of `billed_model_usage_from_llm`.\n\n**2. `lib/crates/fabro-workflow/src/billing_rollup.rs`** ✅\n- Changed `billing_rollup_from_projection` signature to take `catalog: Option<&Catalog>`.\n- Added private helper `stage_usage_with_cost` that clones `stage.usage` and prices it when `total_usd_micros.is_none()` and both `catalog` and `stage.model` are available.\n- Uses `priced` in place of `&stage.usage` for the `is_zero` check, `row.billing.add_counts`, `totals.add_counts`, and `model_entry.billing.add_counts`.\n- Updated existing tests to pass `None`; added a new `rollup_prices_in_flight_stage_usage_using_catalog` test (in-flight stage with no completion, non-zero usage, builtin model → `Some(..)` cost on stage row, totals, and `by_model`).\n\n**3. Call sites** ✅\n- `fabro-server/src/server/handler/billing.rs:82` — binds `let catalog = state.catalog();` and passes `Some(&catalog)`.\n- `fabro-server/src/server.rs` (2 aggregate-billing sites) — pass `None`.\n- `fabro-workflow/src/pipeline/finalize.rs` (4 sites) — pass `None`.\n\n**Verification** ✅\n- `cargo nextest run -p fabro-workflow billing_rollup` — 4 tests pass (new + 3 updated).\n- `cargo nextest run -p fabro-server billing` — 7 tests pass.\n- `cargo nextest run -p fabro-model billing` — 23 tests pass.\n- `cargo check --workspace --all-targets` — clean.\n- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` — clean.\n- `cargo +nightly-2026-04-14 fmt --check --all` — clean.\n- Broader test sweep on `fabro-workflow` + `fabro-server`: 1674 passing; the 2 unrelated `*graph*svg*` failures are pre-existing graphviz subprocess env issues (confirmed by stashing my changes and reproducing).", "internal.run_id": "01KS6AW28FZVV4M2EHBJA7JMNP", - "graph.model_stylesheet": "\n * { model: claude-opus-4-7; }\n ", - "current_node": "implement", - "internal.retry_count.preflight_lint": 0, - "internal.fidelity": "compact", "internal.retry_count.preflight_compile": 0, - "thread.preflight_compile.current_node": "preflight_lint" + "internal.retry_count.implement": 0 }, "node_outcomes": { "toolchain": { @@ -870,11 +1002,48 @@ }, "notes": "Script completed: cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1", "usage": null + }, + "simplify_opus": { + "status": "succeeded", + "context_updates": { + "response.simplify_opus": "All clean. Here's the summary:\n\n## Summary\n\nThree review agents (reuse, quality, efficiency) ran in parallel against the diff. Aggregated findings:\n\n### Fixed\n\n1. **Renamed `BilledTokenCounts::token_counts()` → `to_token_counts()`** (`fabro-model/src/billing.rs`)\n - Quality finding: original name read like a field accessor and hid the fact that the conversion is lossy (drops `total_tokens` and `total_usd_micros`). The `to_` prefix is the established Rust convention for lossy conversions.\n\n2. **Moved pricing logic into `BilledTokenCounts::ensure_priced(catalog, model)`** (`fabro-model/src/billing.rs` + `fabro-workflow/src/billing_rollup.rs`)\n - Quality finding: the previous `stage_usage_with_cost` helper cloned `BilledTokenCounts` across a crate boundary and directly poked the private `total_usd_micros` field, leaking knowledge of how to price a `BilledTokenCounts` out of `fabro-model`. The new method encapsulates that logic next to `add_billed_usage` / `replace_with_billed_usage`, and the rollup-side helper is now a 3-line wrapper.\n\n### Considered and skipped\n\n- **Reuse #1 (refactor `billed_model_usage_from_llm` to use `price_tokens`)** — the \"duplicated\" pricing tail is 3 lines and the surrounding error-vs-option shapes don't compose cleanly. Refactoring adds more indirection than it removes.\n- **`Option<&Catalog>` parameter sprawl** — 6 of 7 call sites pass `None`, but splitting into two functions doubles the public surface for marginal benefit; the single signature is fine.\n- **`pricing_for` memoization across stages** — real but minor (a few `String` clones per stage on a polled endpoint). Worth doing if the path ever shows up in profiles; not justified at the cost of HashMap setup for this fix.\n- **`Cow::Borrowed` for the usage clone** — `BilledTokenCounts` is ~64 bytes of POD; no heap allocation.\n- **Bare `i64` return from `price_tokens`** — matches the established `total_usd_micros: Option` convention used throughout the codebase.\n\n### Verified\n\n- `cargo nextest run -p fabro-workflow billing_rollup` — 4/4 pass (including the new in-flight pricing test).\n- `cargo nextest run -p fabro-model billed_token_counts` — 5/5 pass.\n- `cargo nextest run -p fabro-server billing` — 7/7 pass.\n- `cargo build --workspace` — clean.\n- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` — clean.\n- `cargo +nightly-2026-04-14 fmt --check --all` — clean.", + "last_stage": "simplify_opus", + "last_response": "All clean. Here's the summary:\n\n## Summary\n\nThree review agents (reuse, quality, efficiency) ran in parallel against the diff. Aggregated findings:\n\n### Fixed\n\n1. **Renamed `BilledTokenCounts::token_c" + }, + "notes": "Stage completed: simplify_opus", + "usage": { + "input": { + "usage": { + "model": { + "provider": "anthropic", + "model_id": "claude-opus-4-7" + }, + "tokens": { + "input_tokens": 55570, + "output_tokens": 17369, + "reasoning_tokens": 0, + "cache_read_tokens": 1126033, + "cache_write_tokens": 66184 + } + }, + "facts": { + "algorithm": "anthropic", + "cache_write_5m_tokens": 66184, + "cache_write_1h_tokens": 0 + } + }, + "total_usd_micros": 1688741 + }, + "files_touched": [ + "/home/daytona/workspace/fabro/lib/crates/fabro-model/src/billing.rs", + "/home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/billing_rollup.rs" + ] } }, - "next_node_id": "simplify_opus", + "next_node_id": "simplify_gpt", "node_visits": { "implement": 1, + "simplify_opus": 1, "preflight_compile": 1, "preflight_lint": 1, "start": 1, @@ -909,7 +1078,12 @@ "first_event_seq": 50, "prompt": null, "response": null, - "completion": null, + "completion": { + "outcome": "succeeded", + "notes": "Stage completed: implement", + "failure_reason": null, + "timestamp": "2026-05-21T22:59:35.631657Z" + }, "provider_used": { "mode": "agent", "provider": "anthropic", @@ -922,6 +1096,7 @@ "output": null, "started_at": "2026-05-21T22:46:11.047830Z", "handler": "agent", + "duration_ms": 804582, "usage": { "input_tokens": 92896, "output_tokens": 16017, @@ -935,7 +1110,7 @@ "provider": "anthropic", "model_id": "claude-opus-4-7" }, - "state": "running" + "state": "succeeded" }, "preflight_compile@1": { "first_event_seq": 30, @@ -1094,6 +1269,38 @@ "cache_write_tokens": 0 }, "state": "succeeded" + }, + "simplify_opus@1": { + "first_event_seq": 255, + "prompt": null, + "response": null, + "completion": null, + "provider_used": { + "mode": "agent", + "provider": "anthropic", + "model": "claude-opus-4-7" + }, + "diff": null, + "script_invocation": null, + "script_timing": null, + "parallel_results": null, + "output": null, + "started_at": "2026-05-21T22:59:40.212631Z", + "handler": "agent", + "usage": { + "input_tokens": 55570, + "output_tokens": 17369, + "total_tokens": 1265156, + "reasoning_tokens": 0, + "cache_read_tokens": 1126033, + "cache_write_tokens": 66184, + "total_usd_micros": 1688741 + }, + "model": { + "provider": "anthropic", + "model_id": "claude-opus-4-7" + }, + "state": "running" } } } \ No newline at end of file diff --git a/stages/005-implement@1/diff.patch b/stages/005-implement@1/diff.patch new file mode 100644 index 000000000..3148e8922 --- /dev/null +++ b/stages/005-implement@1/diff.patch @@ -0,0 +1,279 @@ +diff --git a/lib/crates/fabro-model/src/billing.rs b/lib/crates/fabro-model/src/billing.rs +index 29cad2335..f06cad06a 100644 +--- a/lib/crates/fabro-model/src/billing.rs ++++ b/lib/crates/fabro-model/src/billing.rs +@@ -358,6 +358,19 @@ impl BilledTokenCounts { + } + } + ++ /// Returns the five disjoint per-call token buckets, dropping the derived ++ /// `total_tokens` sum and the optional `total_usd_micros` cost. ++ #[must_use] ++ pub fn token_counts(&self) -> TokenCounts { ++ TokenCounts { ++ input_tokens: self.input_tokens, ++ output_tokens: self.output_tokens, ++ reasoning_tokens: self.reasoning_tokens, ++ cache_read_tokens: self.cache_read_tokens, ++ cache_write_tokens: self.cache_write_tokens, ++ } ++ } ++ + pub fn add_counts(&mut self, source: &Self) { + self.input_tokens += source.input_tokens; + self.output_tokens += source.output_tokens; +@@ -435,6 +448,27 @@ impl Catalog { + self.provider(&model_ref.provider) + .and_then(|provider| ModelBillingFacts::for_policy(provider.billing_policy, tokens)) + } ++ ++ /// Price a partial token sample for `model` using catalog pricing. ++ /// ++ /// Returns `None` when the provider has no billing policy, the model is ++ /// unknown, or the pricing algorithm cannot produce a result for the given ++ /// tokens. Used by read-side rollups so in-flight stages can show an ++ /// exact cost for the tokens consumed so far. ++ #[must_use] ++ pub fn price_tokens(&self, model: &ModelRef, tokens: &TokenCounts) -> Option { ++ let facts = self.billing_facts_for(model, tokens)?; ++ let input = ModelBillingInput { ++ usage: ModelUsage { ++ model: model.clone(), ++ tokens: tokens.clone(), ++ }, ++ facts, ++ }; ++ self.pricing_for(model) ++ .and_then(|pricing| pricing.bill(&input)) ++ .map(|amount| amount.0) ++ } + } + + fn costs_for_speed( +diff --git a/lib/crates/fabro-server/src/server.rs b/lib/crates/fabro-server/src/server.rs +index b5258fa65..731c3f926 100644 +--- a/lib/crates/fabro-server/src/server.rs ++++ b/lib/crates/fabro-server/src/server.rs +@@ -3238,7 +3238,7 @@ async fn execute_run_in_process(state: Arc, run_id: RunId) { + .expect("aggregate_billing lock poisoned"); + accumulate_billing_rollup( + &mut agg, +- &fabro_workflow::billing_rollup_from_projection(projection), ++ &fabro_workflow::billing_rollup_from_projection(projection, None), + ); + } + } +@@ -3523,7 +3523,7 @@ async fn execute_run_subprocess(state: Arc, run_id: RunId) { + .expect("aggregate_billing lock poisoned"); + accumulate_billing_rollup( + &mut agg, +- &fabro_workflow::billing_rollup_from_projection(&final_state), ++ &fabro_workflow::billing_rollup_from_projection(&final_state, None), + ); + } + +diff --git a/lib/crates/fabro-server/src/server/handler/billing.rs b/lib/crates/fabro-server/src/server/handler/billing.rs +index 6eee9bbfb..cbaf7cc3e 100644 +--- a/lib/crates/fabro-server/src/server/handler/billing.rs ++++ b/lib/crates/fabro-server/src/server/handler/billing.rs +@@ -79,7 +79,8 @@ async fn get_run_billing( + }; + let projection = cached.projection; + +- let rollup = fabro_workflow::billing_rollup_from_projection(&projection); ++ let catalog = state.catalog(); ++ let rollup = fabro_workflow::billing_rollup_from_projection(&projection, Some(&catalog)); + let by_model = rollup + .by_model + .iter() +diff --git a/lib/crates/fabro-workflow/src/billing_rollup.rs b/lib/crates/fabro-workflow/src/billing_rollup.rs +index a91b06061..89791caa7 100644 +--- a/lib/crates/fabro-workflow/src/billing_rollup.rs ++++ b/lib/crates/fabro-workflow/src/billing_rollup.rs +@@ -1,6 +1,7 @@ + use std::collections::HashMap; + +-use fabro_types::{BilledTokenCounts, ModelRef, RunProjection}; ++use fabro_model::Catalog; ++use fabro_types::{BilledTokenCounts, ModelRef, RunProjection, StageProjection}; + + #[derive(Debug, Clone, PartialEq)] + pub struct ProjectionBillingStage { +@@ -34,7 +35,10 @@ impl ProjectionBillingRollup { + } + + #[must_use] +-pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionBillingRollup { ++pub fn billing_rollup_from_projection( ++ projection: &RunProjection, ++ catalog: Option<&Catalog>, ++) -> ProjectionBillingRollup { + let mut stage_indices = HashMap::::new(); + let mut stages = Vec::::new(); + let mut by_model = HashMap::::new(); +@@ -46,7 +50,8 @@ pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionB + if is_boundary_stage(projection, stage_id.node_id()) { + continue; + } +- if stage.completion.is_none() && stage.duration_ms.is_none() && stage.usage.is_zero() { ++ let priced = stage_usage_with_cost(catalog, stage); ++ if stage.completion.is_none() && stage.duration_ms.is_none() && priced.is_zero() { + continue; + } + +@@ -68,10 +73,10 @@ pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionB + runtime_ms = runtime_ms.saturating_add(duration_ms); + } + +- if !stage.usage.is_zero() { ++ if !priced.is_zero() { + billed_visit_count += 1; +- row.billing.add_counts(&stage.usage); +- totals.add_counts(&stage.usage); ++ row.billing.add_counts(&priced); ++ totals.add_counts(&priced); + + if let Some(model) = &stage.model { + row.model = Some(model.clone()); +@@ -84,7 +89,7 @@ pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionB + billing: BilledTokenCounts::default(), + }); + model_entry.stages += 1; +- model_entry.billing.add_counts(&stage.usage); ++ model_entry.billing.add_counts(&priced); + } + } + } +@@ -113,6 +118,16 @@ pub fn billing_rollup_from_projection(projection: &RunProjection) -> ProjectionB + } + } + ++fn stage_usage_with_cost(catalog: Option<&Catalog>, stage: &StageProjection) -> BilledTokenCounts { ++ let mut usage = stage.usage.clone(); ++ if usage.total_usd_micros.is_none() { ++ if let (Some(catalog), Some(model)) = (catalog, stage.model.as_ref()) { ++ usage.total_usd_micros = catalog.price_tokens(model, &usage.token_counts()); ++ } ++ } ++ usage ++} ++ + fn is_boundary_stage(projection: &RunProjection, node_id: &str) -> bool { + projection + .spec() +@@ -126,6 +141,7 @@ fn is_boundary_stage(projection: &RunProjection, node_id: &str) -> bool { + mod tests { + use std::collections::HashMap; + ++ use fabro_model::{Catalog, ModelRef, ProviderId}; + use fabro_types::{ + AttrValue, BilledModelUsage, BilledTokenCounts, Graph, Node, RunProjection, RunSpec, + StageCompletion, StageOutcome, WorkflowSettings, first_event_seq, fixtures, +@@ -190,7 +206,7 @@ mod tests { + timestamp: chrono::Utc::now(), + }); + +- let rollup = billing_rollup_from_projection(&projection); ++ let rollup = billing_rollup_from_projection(&projection, None); + + assert_eq!(rollup.stages.len(), 1); + assert_eq!(rollup.stages[0].node_id, "verify"); +@@ -233,7 +249,7 @@ mod tests { + timestamp: chrono::Utc::now(), + }); + +- let rollup = billing_rollup_from_projection(&projection); ++ let rollup = billing_rollup_from_projection(&projection, None); + + assert_eq!(rollup.stages.len(), 1); + assert_eq!(rollup.stages[0].node_id, "build"); +@@ -266,12 +282,48 @@ mod tests { + timestamp: chrono::Utc::now(), + }); + +- let rollup = billing_rollup_from_projection(&projection); ++ let rollup = billing_rollup_from_projection(&projection, None); + + assert_eq!(rollup.stages.len(), 0); + assert_eq!(rollup.runtime_ms, 0); + } + ++ #[test] ++ fn rollup_prices_in_flight_stage_usage_using_catalog() { ++ let mut projection = test_projection(); ++ let model = ModelRef { ++ provider: ProviderId::openai(), ++ model_id: "gpt-5.4".to_string(), ++ speed: None, ++ }; ++ let stage = projection.stage_entry("agent", 1, first_event_seq(1)); ++ stage.started_at = Some(chrono::Utc::now()); ++ stage.usage = BilledTokenCounts { ++ input_tokens: 500_000, ++ output_tokens: 125_000, ++ total_tokens: 625_000, ++ ..BilledTokenCounts::default() ++ }; ++ stage.model = Some(model.clone()); ++ ++ let priced = billing_rollup_from_projection(&projection, Some(Catalog::builtin())); ++ let unpriced = billing_rollup_from_projection(&projection, None); ++ ++ assert_eq!(priced.stages.len(), 1); ++ assert_eq!(priced.stages[0].node_id, "agent"); ++ let stage_cost = priced.stages[0].billing.total_usd_micros; ++ assert!( ++ stage_cost.is_some_and(|cost| cost > 0), ++ "expected priced stage cost, got {stage_cost:?}" ++ ); ++ assert_eq!(priced.totals.total_usd_micros, stage_cost); ++ assert_eq!(priced.by_model.len(), 1); ++ assert_eq!(priced.by_model[0].billing.total_usd_micros, stage_cost); ++ assert_eq!(unpriced.stages.len(), 1); ++ assert_eq!(unpriced.stages[0].billing.total_usd_micros, None); ++ assert_eq!(unpriced.totals.total_usd_micros, None); ++ } ++ + fn run_spec_with_boundary_nodes() -> RunSpec { + let mut graph = Graph::new("test"); + graph.nodes.insert("start".to_string(), { +diff --git a/lib/crates/fabro-workflow/src/pipeline/finalize.rs b/lib/crates/fabro-workflow/src/pipeline/finalize.rs +index 5345bfb4a..a7f4619db 100644 +--- a/lib/crates/fabro-workflow/src/pipeline/finalize.rs ++++ b/lib/crates/fabro-workflow/src/pipeline/finalize.rs +@@ -82,7 +82,7 @@ pub(crate) async fn build_conclusion_from_store( + .unwrap_or_default(); + let projection_billing = projection + .as_ref() +- .map(billing_rollup_from_projection) ++ .map(|projection| billing_rollup_from_projection(projection, None)) + .unwrap_or_default(); + let checkpoint = projection + .as_ref() +@@ -438,7 +438,7 @@ async fn compute_final_patch( + } + + pub(crate) fn billing_from_projection(projection: &RunProjection) -> Option { +- billing_rollup_from_projection(projection).billing_if_present() ++ billing_rollup_from_projection(projection, None).billing_if_present() + } + + pub(crate) fn build_terminal_event( +@@ -554,7 +554,7 @@ pub async fn finalize(executed: Executed, options: &FinalizeOptions) -> Result TokenCounts` — drops + `total_tokens`/`total_usd_micros`, keeps the five disjoint buckets. +- Add `Catalog::price_tokens(&self, model: &ModelRef, tokens: &TokenCounts) -> Option` + next to `pricing_for`/`billing_facts_for`. Body mirrors the cost lines of + `billed_model_usage_from_llm`: build `ModelBillingFacts` via + `billing_facts_for`, assemble `ModelBillingInput { ModelUsage { model, tokens }, facts }`, + then `pricing_for(model).and_then(|p| p.bill(&input)).map(|a| a.0)`. Returns + `None` when the provider has no billing policy. + +### 2. `lib/crates/fabro-workflow/src/billing_rollup.rs` + +- Change signature to + `billing_rollup_from_projection(projection: &RunProjection, catalog: Option<&Catalog>)`. +- Add a module-private helper `stage_usage_with_cost(catalog, stage) -> BilledTokenCounts`: + clone `stage.usage`; if `total_usd_micros.is_none()` and both `catalog` and + `stage.model` are present, set it via `catalog.price_tokens(model, &usage.token_counts())`. +- In the loop, compute `priced` once per stage and use it in place of + `&stage.usage` for the `is_zero` check, `row.billing.add_counts`, + `totals.add_counts`, and `model_entry.billing.add_counts`. +- Update the existing tests to pass `None`; add one new test: an in-flight + stage (no `completion`, non-zero `usage` with `total_usd_micros: None`, a + builtin `model`) yields `Some(..)` cost on the stage row and in `totals` when + called with `Some(Catalog::builtin())`. + +### 3. Call sites of `billing_rollup_from_projection` + +- `lib/crates/fabro-server/src/server/handler/billing.rs:82` — bind + `let catalog = state.catalog();` (returns `Arc`) and pass + `Some(&catalog)`. +- `lib/crates/fabro-server/src/server.rs` (2 aggregate-billing sites) — pass `None`. +- `lib/crates/fabro-workflow/src/pipeline/finalize.rs` (4 sites) — pass `None` + (stages already priced at completion; pricing would be a no-op anyway). + +No changes to `fabro-api.yaml`, the generated clients, or `apps/fabro-web`. + +## Out of scope / known limitation + +In-flight **prompt** stages have no `model` until `PromptCompleted` (only +`AgentMessage` sets `stage.model` mid-run), so they still show `—` while +running. Acceptable: prompt stages are a single short LLM call. The bug report +concerns agent stages, where `model` is available. + +## Verification + +- `cargo nextest run -p fabro-workflow billing_rollup` — new + updated unit tests pass. +- `cargo nextest run -p fabro-server billing` — handler conformance still passes. +- `cargo build --workspace` and `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings`. +- Manual: `fabro server start` + `cd apps/fabro-web && bun run dev`, start a + workflow with an agent stage, open `/runs/{id}/billing` mid-run — the active + stage row and totals show a non-`—` dollar amount that grows with tokens. + + +## Completed stages +- **toolchain**: succeeded + - Script: `command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1` + - Output: + ``` + cargo 1.95.0 (f2d3ce0bd 2026-03-21) + ``` +- **preflight_compile**: succeeded + - Script: `cargo check -q --workspace 2>&1` + - Output: (empty) +- **preflight_lint**: succeeded + - Script: `cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1` + - Output: (empty) +- **implement**: succeeded + - Model: claude-opus-4-7, 92.9k tokens in / 16.0k out + - Files: /home/daytona/workspace/fabro/lib/crates/fabro-model/src/billing.rs, /home/daytona/workspace/fabro/lib/crates/fabro-server/src/server.rs, /home/daytona/workspace/fabro/lib/crates/fabro-server/src/server/handler/billing.rs, /home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/billing_rollup.rs, /home/daytona/workspace/fabro/lib/crates/fabro-workflow/src/pipeline/finalize.rs + + +# Simplify: Code Review and Cleanup + +Review changes vs. origin for reuse, quality, and efficiency. Fix any issues found. + +## Phase 1: Identify Changes + +Run git diff (or git diff HEAD if there are staged changes) to see what changed. If there are no git changes, review the most recently modified files that the user mentioned or that you edited earlier in this conversation. + +## Phase 2: Launch Three Review Agents in Parallel + +Use the Agent tool to launch all three agents concurrently in a single message. Pass each agent the full diff so it has the complete context. + +### Agent 1: Code Reuse Review + +For each change: + +1. Search for existing utilities and helpers that could replace newly written code. Use Grep to find similar patterns elsewhere in the codebase — common locations are utility directories, shared modules, and files adjacent to the changed ones. +2. Flag any new function that duplicates existing functionality. Suggest the existing function to use instead. +3. Flag any inline logic that could use an existing utility — hand-rolled string manipulation, manual path handling, custom environment checks, ad-hoc type guards, and similar patterns are common candidates. + +Note: This is a greenfield app, so focus on maximizing simplicity and don't worry about changing things to achieve it. + +### Agent 2: Code Quality Review + +Review the same changes for hacky patterns: + +1. Redundant state: state that duplicates existing state, cached values that could be derived, observers/effects that could be direct calls +2. Parameter sprawl: adding new parameters to a function instead of generalizing or restructuring existing ones +3. Copy-paste with slight variation: near-duplicate code blocks that should be unified with a shared abstraction +4. Leaky abstractions: exposing internal details that should be encapsulated, or breaking existing abstraction boundaries +5. Stringly-typed code: using raw strings where constants, enums (string unions), or branded types already exist in the codebase + +Note: This is a greenfield app, so be aggressive in optimizing quality. + +### Agent 3: Efficiency Review + +Review the same changes for efficiency: + +1. Unnecessary work: redundant computations, repeated file reads, duplicate network/API calls, N+1 patterns +2. Missed concurrency: independent operations run sequentially when they could run in parallel +3. Hot-path bloat: new blocking work added to startup or per-request/per-render hot paths +4. Unnecessary existence checks: pre-checking file/resource existence before operating (TOCTOU anti-pattern) — operate directly and handle the error +5. Memory: unbounded data structures, missing cleanup, event listener leaks +6. Overly broad operations: reading entire files when only a portion is needed, loading all items when filtering for one + +## Phase 3: Fix Issues + +Wait for all three agents to complete. Aggregate their findings and fix each issue directly. If a finding is a false positive or not worth addressing, note it and move on — do not argue with the finding, just skip it. + +When done, briefly summarize what was fixed (or confirm the code was already clean). \ No newline at end of file diff --git a/stages/006-simplify_opus@1/provider_used.json b/stages/006-simplify_opus@1/provider_used.json new file mode 100644 index 000000000..672cc50e5 --- /dev/null +++ b/stages/006-simplify_opus@1/provider_used.json @@ -0,0 +1,5 @@ +{ + "mode": "agent", + "provider": "anthropic", + "model": "claude-opus-4-7" +} \ No newline at end of file diff --git a/stages/006-simplify_opus@1/response.md b/stages/006-simplify_opus@1/response.md new file mode 100644 index 000000000..7136a7bad --- /dev/null +++ b/stages/006-simplify_opus@1/response.md @@ -0,0 +1,30 @@ +All clean. Here's the summary: + +## Summary + +Three review agents (reuse, quality, efficiency) ran in parallel against the diff. Aggregated findings: + +### Fixed + +1. **Renamed `BilledTokenCounts::token_counts()` → `to_token_counts()`** (`fabro-model/src/billing.rs`) + - Quality finding: original name read like a field accessor and hid the fact that the conversion is lossy (drops `total_tokens` and `total_usd_micros`). The `to_` prefix is the established Rust convention for lossy conversions. + +2. **Moved pricing logic into `BilledTokenCounts::ensure_priced(catalog, model)`** (`fabro-model/src/billing.rs` + `fabro-workflow/src/billing_rollup.rs`) + - Quality finding: the previous `stage_usage_with_cost` helper cloned `BilledTokenCounts` across a crate boundary and directly poked the private `total_usd_micros` field, leaking knowledge of how to price a `BilledTokenCounts` out of `fabro-model`. The new method encapsulates that logic next to `add_billed_usage` / `replace_with_billed_usage`, and the rollup-side helper is now a 3-line wrapper. + +### Considered and skipped + +- **Reuse #1 (refactor `billed_model_usage_from_llm` to use `price_tokens`)** — the "duplicated" pricing tail is 3 lines and the surrounding error-vs-option shapes don't compose cleanly. Refactoring adds more indirection than it removes. +- **`Option<&Catalog>` parameter sprawl** — 6 of 7 call sites pass `None`, but splitting into two functions doubles the public surface for marginal benefit; the single signature is fine. +- **`pricing_for` memoization across stages** — real but minor (a few `String` clones per stage on a polled endpoint). Worth doing if the path ever shows up in profiles; not justified at the cost of HashMap setup for this fix. +- **`Cow::Borrowed` for the usage clone** — `BilledTokenCounts` is ~64 bytes of POD; no heap allocation. +- **Bare `i64` return from `price_tokens`** — matches the established `total_usd_micros: Option` convention used throughout the codebase. + +### Verified + +- `cargo nextest run -p fabro-workflow billing_rollup` — 4/4 pass (including the new in-flight pricing test). +- `cargo nextest run -p fabro-model billed_token_counts` — 5/5 pass. +- `cargo nextest run -p fabro-server billing` — 7/7 pass. +- `cargo build --workspace` — clean. +- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` — clean. +- `cargo +nightly-2026-04-14 fmt --check --all` — clean. \ No newline at end of file