checkpoint

⚒️ Generated with [Fabro](https://fabro.sh)
This commit is contained in:
Fabro 2026-05-24 12:11:01 -04:00
parent 67b5fb33e5
commit 3ea9380761
8 changed files with 1070 additions and 24 deletions

438
run.json

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,344 @@
diff --git a/.claude/skills/changelog/watermark b/.claude/skills/changelog/watermark
index f3d8860d3..dc7fea578 100644
--- a/.claude/skills/changelog/watermark
+++ b/.claude/skills/changelog/watermark
@@ -1 +1 @@
-df6f62db6a285ae76fdeead26f6081f8f79cdb9d
+6544e2c589cf436de4b18b70755177183b10c7d9
diff --git a/apps/fabro-web/app/routes/run-stages.tsx b/apps/fabro-web/app/routes/run-stages.tsx
index 7503f0c3a..ae7b2ad14 100644
--- a/apps/fabro-web/app/routes/run-stages.tsx
+++ b/apps/fabro-web/app/routes/run-stages.tsx
@@ -1416,12 +1416,9 @@ function EventsToolbar({
<span className="font-mono">{modelUsageLabel}</span>
</HoverCard>
)}
- <EventExportActions
- events={events}
- runId={runId}
- stageId={stageId}
- className={!showFilters && !modelUsageLabel ? "ml-auto" : ""}
- />
+ {tab === "debug" && (
+ <EventExportActions events={events} runId={runId} stageId={stageId} />
+ )}
{tab === "primary" && renderer === "command" && commandTurn && (
<CommandStatus turn={commandTurn} />
)}
diff --git a/apps/fabro-web/app/routes/runs.tsx b/apps/fabro-web/app/routes/runs.tsx
index 1b40333fd..7d5878b86 100644
--- a/apps/fabro-web/app/routes/runs.tsx
+++ b/apps/fabro-web/app/routes/runs.tsx
@@ -175,10 +175,11 @@ function boardLifecycleStatusLabel(run: Pick<RunItem, "column" | "lifecycleStatu
return run.lifecycleStatusLabel;
}
-function listLifecycleStatusLabel(run: Pick<RunWithStatus, "statusLabel" | "lifecycleStatusLabel">): string | null {
+function listLifecycleStatusLabel(run: Pick<RunWithStatus, "status" | "statusLabel" | "lifecycleStatusLabel">): string | null {
if (run.lifecycleStatusLabel == null || run.lifecycleStatusLabel === run.statusLabel) {
return null;
}
+ if (run.status === "initializing") return null;
return run.lifecycleStatusLabel;
}
diff --git a/docs/public/changelog/2026-05-22.mdx b/docs/public/changelog/2026-05-22.mdx
index e44e76201..e91fd29af 100644
--- a/docs/public/changelog/2026-05-22.mdx
+++ b/docs/public/changelog/2026-05-22.mdx
@@ -49,6 +49,8 @@ This gives Fabro a durable view of agent task state instead of treating todo upd
<Accordion title="Workflows">
- `fabro_run_create` now supports `goal_file` and preserves it as a file-backed run goal
- `fabro_run_create` now accepts the advertised workflow-string shorthand in MCP clients
+- New `fabro_run_get` tool provides read-only run inspection without mutating workflow state
+- `run.checkpoint.skip_git_hooks` can bypass Git hooks during checkpoint commits
- Pair user and system messages now appear in the stage Thread tab and activity timeline
</Accordion>
@@ -57,11 +59,15 @@ This gives Fabro a durable view of agent task state instead of treating todo upd
- Fixed Ask Fabro markdown spacing and heading sizes in narrow panels
- Fixed reasoning deltas leaking into Ask Fabro output
- Fixed archived run board updates not refreshing active and archived queries consistently
+- Fixed retryable mid-stream LLM failures being treated as terminal agent errors
+- Fixed agent edit operations by reading raw sandbox file contents before applying patches
</Accordion>
<Accordion title="Improvements">
- The run stage sidebar can now collapse to an icon rail while preserving status indicators
- Ask Fabro example prompts now use fuller, action-oriented text
+- Agent context observability now records memory, activated skills, and MCP tools as run events
+- Anthropic agent sessions now support task lookup and task reminders
- Run board archived columns now stay in a predictable order after archive and unarchive actions
- OpenAI `apply_patch` handling is now compatible with Codex-style patch grammar
</Accordion>
diff --git a/docs/public/changelog/2026-05-23.mdx b/docs/public/changelog/2026-05-23.mdx
index a5c1843ad..6ce319ff0 100644
--- a/docs/public/changelog/2026-05-23.mdx
+++ b/docs/public/changelog/2026-05-23.mdx
@@ -1,8 +1,69 @@
---
-title: "Legacy sandbox config migration"
+title: "Runs list, Slack notifications, and structured outputs"
date: "2026-05-23"
---
-## Legacy sandbox config auto-migration
+<Warning>
+**Run lifecycle status values changed for API and CLI consumers.** The old `queued` state is now split into `pending` and `runnable`, and approval state is represented explicitly.
-Fabro now temporarily rewrites confidently migratable pre-v1.0 `[run.sandbox]` config files to the named environment syntax. A backup is written next to the original file before rewriting. Ambiguous or unsupported legacy keys fail with a targeted migration message instead of the generic TOML unknown-field error.
+To migrate:
+1. Update clients that switch on `queued` to handle `pending` and `runnable`.
+2. Use the approval fields on run summaries and projections when deciding whether a run is waiting for human action.
+</Warning>
+
+## Runs are easier to scan, sort, and recover
+
+The runs list now uses a denser table layout with server-side sorting, pagination, and a redesigned toolbar. You can sort by repo, title, workflow, and changes without waiting for the browser to process the full list.
+
+Failed and dead runs can now be retried from the run UI and API instead of being recreated by hand. The same run lifecycle work also separates pending runs from runnable runs, so approvals and blocked start conditions are clearer in the product and the API.
+
+## Batch run actions
+
+The runs list now supports multi-select archive and unarchive, with matching batch lifecycle endpoints for API clients. Workspace preferences remember the runs view choices you make, so list layout and filtering stay consistent across visits.
+
+Run size is now exposed on summaries and projections, which gives the UI and API a shared way to distinguish small, medium, and large runs. The web app uses that signal in the run header and list surfaces where scale matters for triage.
+
+## Agent stages expose more of what is happening
+
+Agent stage pages now include a stage insights sidebar, context-window snapshots, and richer stage projections for todos, subagents, skills, MCP tools, reasoning effort, speed, and model usage. This makes it easier to see why a stage is behaving a certain way without digging through raw events.
+
+Structured output validation is also stricter. Agent and prompt stages with `output_schema` now validate model output in the same context and ask the model to repair invalid responses before the workflow moves on.
+
+## Slack run lifecycle notifications
+
+Workflows can now send run lifecycle updates to Slack through `[run.notifications]`. Teams can keep a channel informed when a run starts, needs attention, finishes, or fails without wiring a separate webhook workflow.
+
+```toml
+[run.notifications]
+slack = true
+```
+
+## More
+
+<Accordion title="API">
+- New `POST /api/v1/runs/archive` and `POST /api/v1/runs/unarchive` endpoints apply lifecycle actions to multiple runs at once
+- Run summaries and projections now expose run size buckets
+- Stage projections now include agent todos, subagents, skills, MCP tools, context-window data, and model usage
+- `output_schema` validation errors now include structured repair context for agent and prompt stages
+</Accordion>
+
+<Accordion title="Workflows">
+- Named environments replace run-scoped sandbox config, with automatic migration for confidently supported legacy `[run.sandbox]` files
+- Legacy sandbox config migration writes a backup next to the original file before rewriting
+- Ambiguous legacy sandbox config now fails with a targeted migration message
+- Interview options now carry metadata through the API, console, Slack blocks, and agent question tools
+</Accordion>
+
+<Accordion title="Fixes">
+- Fixed the Quick Start screen flashing when navigating to `/runs` with archived preferences
+- Fixed implement-plan pull requests advancing before CI checks completed
+- Fixed an event append race covered by a regression test
+</Accordion>
+
+<Accordion title="Improvements">
+- Added a Waterfall view for run events with hover popovers and duration breakdowns
+- Added live timing overlays for in-flight runs
+- Generated run titles now use the lighter `small_default` model role asynchronously
+- Settings pages now show provider and integration logos, including upcoming Teams, Discord, and Jira integrations
+- Stage insights sidebar polish improved visit syntax, spacing, and quiet-state presentation
+</Accordion>
diff --git a/docs/public/changelog/2026-05-24.mdx b/docs/public/changelog/2026-05-24.mdx
new file mode 100644
index 000000000..8a723cddc
--- /dev/null
+++ b/docs/public/changelog/2026-05-24.mdx
@@ -0,0 +1,39 @@
+---
+title: "Batch delete and richer stage details"
+date: "2026-05-24"
+---
+
+## Delete runs in batches
+
+The runs list bulk toolbar now includes Delete for selected runs. API clients can use the matching batch delete endpoint to remove multiple runs in one request and get per-run results back.
+
+```http
+POST /api/v1/runs/delete
+```
+
+## Stage details are closer at hand
+
+Stage status, model, and event details are now available from hover popovers across the run overview graph, stage sidebar, and stage events toolbar. You can inspect stage metadata without leaving the current run view.
+
+Loaded stage events also gained Copy and Download actions. That makes it easier to move event output into an issue, support thread, or local investigation without selecting text from the page.
+
+## More
+
+<Accordion title="API">
+- New `POST /api/v1/runs/delete` endpoint deletes multiple runs and returns per-run results plus a summary
+</Accordion>
+
+<Accordion title="Fixes">
+- Fixed unfinished stages being left open after a run failure
+- Removed a non-functional demo-mode Connect dropdown
+- Suppressed the redundant Starting pill on the runs list
+</Accordion>
+
+<Accordion title="Improvements">
+- Added a run size badge to the run header
+- Unconfigured provider rows now link directly to a prefilled secret form
+- Removed retired OpenAI models from the catalog
+- Improved muted stage status pill contrast
+- Collapsed the stage events search field behind an icon by default
+- Updated the Automations nav icon
+</Accordion>
diff --git a/docs/public/core-concepts/models.mdx b/docs/public/core-concepts/models.mdx
index 61b26419b..f545ff073 100644
--- a/docs/public/core-concepts/models.mdx
+++ b/docs/public/core-concepts/models.mdx
@@ -19,8 +19,6 @@ No single model is best at everything. Fabro lets you assign the right model to
| `claude-sonnet-4-5` | anthropic | | 200K | $3.00 / $15.00 | 50 tok/s |
| `claude-haiku-4-5` | anthropic | `haiku`, `claude-haiku` | 200K | $0.80 / $4.00 | 100 tok/s |
| `gpt-5.2` | openai | `gpt5` | 1M | $1.80 / $14.00 | 65 tok/s |
-| `gpt-5-mini` | openai | `gpt5-mini` | 1M | $0.20 / $2.00 | 70 tok/s |
-| `gpt-5.2-codex` | openai | | 1M | $1.80 / $14.00 | 100 tok/s |
| `gpt-5.3-codex` | openai | `codex` | 1M | $1.80 / $14.00 | 100 tok/s |
| `gpt-5.3-codex-spark` | openai | `codex-spark` | 128K | n/a | 1000 tok/s |
| `gpt-5.4` | openai | `gpt54` | 1M | $2.50 / $15.00 | 70 tok/s |
diff --git a/docs/public/docs.json b/docs/public/docs.json
index dd9a39999..89d8b4b43 100644
--- a/docs/public/docs.json
+++ b/docs/public/docs.json
@@ -252,6 +252,8 @@
"group": "May 2026",
"icon": "clock-rotate-left",
"pages": [
+ "changelog/2026-05-24",
+ "changelog/2026-05-23",
"changelog/2026-05-22",
"changelog/2026-05-21",
"changelog/2026-05-20",
diff --git a/docs/public/examples/definition-of-done.mdx b/docs/public/examples/definition-of-done.mdx
index 5832cad59..d96a81120 100644
--- a/docs/public/examples/definition-of-done.mdx
+++ b/docs/public/examples/definition-of-done.mdx
@@ -289,7 +289,7 @@ digraph SpecDoDMultiModel {
* { model: claude-opus-4-6;}
.opus { model: claude-opus-4-6;reasoning_effort: high; }
.gpt { model: gpt-5.2; reasoning_effort: high; }
- .codex { model: gpt-5.2-codex; reasoning_effort: high; }
+ .codex { model: gpt-5.3-codex; reasoning_effort: high; }
.merge { model: claude-opus-4-6; reasoning_effort: high; }
"
]
diff --git a/docs/public/reference/dot-language.mdx b/docs/public/reference/dot-language.mdx
index 74a366632..1c56cfa3e 100644
--- a/docs/public/reference/dot-language.mdx
+++ b/docs/public/reference/dot-language.mdx
@@ -48,7 +48,7 @@ Attribute values in `[key=value]` blocks can be:
| Float | Digits with decimal point | `3.14`, `-0.5`, `.5` |
| Boolean | Keywords | `true`, `false` |
| Duration | Integer with unit suffix | `250ms`, `30s`, `15m`, `2h`, `1d` |
-| Bare string | Identifier with hyphens/dots | `claude-sonnet-4-5`, `gpt-5.2-codex` |
+| Bare string | Identifier with hyphens/dots | `claude-sonnet-4-5`, `gpt-5.3-codex` |
| Identifier | Bare word | `LR`, `box`, `Mdiamond` |
**Escape sequences** in quoted strings: `\"`, `\\`, `\n`, `\t`.
diff --git a/lib/crates/fabro-cli/tests/it/cmd/model.rs b/lib/crates/fabro-cli/tests/it/cmd/model.rs
index a94e8b243..5dcf9b7b5 100644
--- a/lib/crates/fabro-cli/tests/it/cmd/model.rs
+++ b/lib/crates/fabro-cli/tests/it/cmd/model.rs
@@ -120,7 +120,6 @@ fn list_query_aliases() {
exit_code: 0
----- stdout -----
MODEL PROVIDER ALIASES CONTEXT COST SPEED
- gpt-5.2-codex openai 1m $1.8 / $14.0 100 tok/s
gpt-5.3-codex openai codex 1m $1.8 / $14.0 100 tok/s
gpt-5.3-codex-spark openai codex-spark 131k - / - 1000 tok/s
----- stderr -----
diff --git a/lib/crates/fabro-model/src/catalog/providers/openai.toml b/lib/crates/fabro-model/src/catalog/providers/openai.toml
index 3cdba296e..d9530972f 100644
--- a/lib/crates/fabro-model/src/catalog/providers/openai.toml
+++ b/lib/crates/fabro-model/src/catalog/providers/openai.toml
@@ -33,55 +33,6 @@ input_cost_per_mtok = 1.75
output_cost_per_mtok = 14.0
cache_input_cost_per_mtok = 0.175
-[models."gpt-5-mini"]
-provider = "openai"
-api_id = "gpt-5-mini"
-display_name = "GPT-5 Mini"
-family = "gpt-5"
-training = "2025-08-31"
-knowledge_cutoff = "April 2025"
-estimated_output_tps = 70
-aliases = ["gpt5-mini"]
-
-[models."gpt-5-mini".limits]
-context_window = 1047576
-max_output = 128000
-
-[models."gpt-5-mini".features]
-tools = true
-vision = true
-reasoning = true
-reasoning_effort = "levels"
-
-[models."gpt-5-mini".costs]
-input_cost_per_mtok = 0.25
-output_cost_per_mtok = 2.0
-cache_input_cost_per_mtok = 0.025
-
-[models."gpt-5.2-codex"]
-provider = "openai"
-api_id = "gpt-5.2-codex"
-display_name = "GPT-5.2 Codex"
-family = "gpt-5"
-training = "2025-08-31"
-knowledge_cutoff = "April 2025"
-estimated_output_tps = 100
-
-[models."gpt-5.2-codex".limits]
-context_window = 1047576
-max_output = 128000
-
-[models."gpt-5.2-codex".features]
-tools = true
-vision = true
-reasoning = true
-reasoning_effort = "levels"
-
-[models."gpt-5.2-codex".costs]
-input_cost_per_mtok = 1.75
-output_cost_per_mtok = 14.0
-cache_input_cost_per_mtok = 0.175
-
[models."gpt-5.3-codex"]
provider = "openai"
api_id = "gpt-5.3-codex"
diff --git a/lib/crates/fabro-server/src/server/tests.rs b/lib/crates/fabro-server/src/server/tests.rs
index e1c8883ed..6488ab364 100644
--- a/lib/crates/fabro-server/src/server/tests.rs
+++ b/lib/crates/fabro-server/src/server/tests.rs
@@ -5280,7 +5280,6 @@ async fn list_models_filters_by_query_across_aliases() {
.map(|model| model["id"].as_str().unwrap().to_string())
.collect::<Vec<_>>();
assert_eq!(model_ids, vec![
- "gpt-5.2-codex".to_string(),
"gpt-5.3-codex".to_string(),
"gpt-5.3-codex-spark".to_string()
]);
diff --git a/test/docs/examples/definition-of-done/spec-dod-multimodel.fabro b/test/docs/examples/definition-of-done/spec-dod-multimodel.fabro
index 0b7f30e24..a4c88659a 100644
--- a/test/docs/examples/definition-of-done/spec-dod-multimodel.fabro
+++ b/test/docs/examples/definition-of-done/spec-dod-multimodel.fabro
@@ -8,7 +8,7 @@ digraph SpecDoDMultiModel {
* { model: claude-opus-4-6;}
.opus { model: claude-opus-4-6;reasoning_effort: high; }
.gpt { model: gpt-5.2; reasoning_effort: high; }
- .codex { model: gpt-5.2-codex; reasoning_effort: high; }
+ .codex { model: gpt-5.3-codex; reasoning_effort: high; }
.merge { model: claude-opus-4-6; reasoning_effort: high; }
"
]

View file

@ -0,0 +1 @@
blob://sha256/4387fe2e4d797ac7f44273703b7eacbec475d323ea6b29bbe5d934290854c130

View file

@ -0,0 +1,8 @@
{
"output": "blob://sha256/4387fe2e4d797ac7f44273703b7eacbec475d323ea6b29bbe5d934290854c130",
"exit_code": 100,
"duration_ms": 190592,
"termination": "exited",
"output_bytes": 27010,
"live_streaming": true
}

View file

@ -0,0 +1,6 @@
{
"outcome": "failed",
"notes": null,
"failure_reason": "Script failed with exit code: 100\n\n## output\n──────────────────────────────\n To update snapshots run `cargo insta review`\n Stopped on the first failure. Run `cargo insta test` to run all snapshots.\n test commands::run::run_progress::tests::plain_default_stage_snapshot ... FAILED\n\n failures:\n\n failures:\n commands::run::run_progress::tests::plain_default_stage_snapshot\n\n test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 427 filtered out; finished in 0.05s\n\n stderr ───\n\n thread 'commands::run::run_progress::tests::plain_default_stage_snapshot' (118966) panicked at /root/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/insta-1.46.3/src/runtime.rs:719:13:\n snapshot assertion for 'plain_default_stage_snapshot' failed in line 866\n note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace\n\n Cancelling due to test failure: 7 tests still running\n FAIL [ 0.070s] (1417/6331) fabro-cli::bin/fabro commands::run::run_progress::tests::plain_verbose_snapshot\n stdout ───\n\n running 1 test\n ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Snapshot Summary ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\n Snapshot: plain_verbose_snapshot\n Source: lib/crates/fabro-cli/src/commands/run/run_progress/mod.rs:1224\n ────────────────────────────────────────────────────────────────────────────────\n Expression: rendered(&buffer)\n ────────────────────────────────────────────────────────────────────────────────\n -old snapshot\n +new results\n ────────────┬───────────────────────────────────────────────────────────────────\n 8 8 │ Setup: 1 command (2s)\n 9 9 │ Running devcontainer postCreate (1 commands)...\n 10 10 │ ✓ [1/1] npm run setup 1s\n 11 11 │ Devcontainer: postCreate (1s)\n 12 │-✓ Code $0.00 5s (1 turns, 0 tools, 1.5k toks)\n 12 │+✓ Code 5s (1 turns, 0 tools, 1.5k toks)\n ────────────┴───────────────────────────────────────────────────────────────────\n To update snapshots run `cargo insta review`\n Stopped on the first failure. Run `cargo insta test` to run all snapshots.\n test commands::run::run_progress::tests::plain_verbose_snapshot ... FAILED\n\n failures:\n\n failures:\n commands::run::run_progress::tests::plain_verbose_snapshot\n\n test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 427 filtered out; finished in 0.06s\n\n stderr ───\n\n thread 'commands::run::run_progress::tests::plain_verbose_snapshot' (118993) panicked at /root/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/insta-1.46.3/src/runtime.rs:719:13:\n snapshot assertion for 'plain_verbose_snapshot' failed in line 1224\n note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace\n\n────────────\n Summary [ 10.824s] 1418/6331 tests run: 1416 passed, 2 failed, 181 skipped\n FAIL [ 0.061s] (1411/6331) fabro-cli::bin/fabro commands::run::run_progress::tests::plain_default_stage_snapshot\n FAIL [ 0.070s] (1417/6331) fabro-cli::bin/fabro commands::run::run_progress::tests::plain_verbose_snapshot\nwarning: 4913/6331 tests were not run due to test failure (run with --no-fail-fast to run all tests, or run with --max-fail)\nerror: test run failed\n",
"timestamp": "2026-05-24T16:06:51.544028Z"
}

View file

@ -0,0 +1,280 @@
Goal: ---
title: fix: Prefer root stage TODO projection
type: fix
status: active
date: 2026-05-24
---
# fix: Prefer Root Stage TODO Projection
## Overview
Fix stage TODO projection so `StageProjection.todos` represents the selected
stage agent's root session plan. Today a child OpenAI session can emit its own
`todo.created` events on the same `stage_id`, replacing the root list and
causing later root `todo.updated` completions to be ignored. The visible
symptom is an agent sidebar showing a stale child TODO list such as `0/3`
completed even though the root stage plan completed.
## Problem Frame
OpenAI `update_plan` lists are scoped per agent session as
`openai_plan:<session_id>`. A stage may contain both the root agent session and
child/subagent sessions. `StageProjection` currently has only one
`todos: Option<TodoListProjection>`, so the reducer must choose which list is
the stage-level list. The stage sidebar is a stage-agent summary, so it should
show the root stage session's list rather than whichever session most recently
created todos.
## Requirements Trace
- R1. Root OpenAI plan TODOs must remain the projected `stage.todos` list even
when child OpenAI sessions emit TODO events on the same stage.
- R2. Later root OpenAI `todo.updated` and `todo.deleted` events must continue
to apply after child OpenAI TODO events are observed.
- R3. Child OpenAI TODO lists must not create, replace, or mutate
`StageProjection.todos`.
- R4. Anthropic task projection must remain unchanged because Anthropic task
lists are intentionally scoped to the root session and shared across
subagents.
- R5. Do not change public API shapes, generated API types, or frontend
rendering code for this fix.
## Scope Boundaries
- Do not add a multi-list TODO projection in this change.
- Do not expose child/subagent TODO lists in the sidebar in this change.
- Do not change `TodoListProjection`, `StageProjection`, or OpenAPI schemas.
- Do not change event serialization, event names, or the agent `update_plan`
tool behavior.
## Context & Research
### Relevant Code and Patterns
- `lib/crates/fabro-agent/src/todo_tools.rs` scopes OpenAI plans by
`session_id`, producing `openai_plan:<session_id>`.
- `lib/crates/fabro-types/src/run_event/mod.rs` already carries
`session_id` and `parent_session_id` on event envelopes.
- `lib/crates/fabro-store/src/run_state.rs` owns the persisted-event reducer
that updates `StageProjection.todos` from `todo.created`, `todo.updated`,
and `todo.deleted`.
- The existing `todo_reducer` test module in `run_state.rs` is the right place
for focused regression coverage.
- `apps/fabro-web/app/routes/run-stages.tsx` and
`apps/fabro-web/app/components/stage-insights-sidebar.tsx` already render
`stage.todos`; no UI change is needed if the projection is corrected.
### Observed Failing Case
For run `01KSBT48J14ZMK9HQN48SVMG3T`, stage `simplify_gpt@1` had:
- Root list `openai_plan:2f4458b9-1128-4a96-8dfa-bff4b73b9c33`: five root
todos, all completed by later `todo.updated` events.
- Child list `openai_plan:a09b9432-823d-4068-91fa-5c6185578e8e`: three child
todos, created with `parent_session_id` set and never updated.
The child `todo.created` events replaced `stage.todos`, so the sidebar showed
the child list as `0/3` even after the root list completed.
## Key Technical Decisions
- Use `parent_session_id` as the root-vs-child signal for OpenAI plan
projection. A root stage session has `parent_session_id == None`; child
sessions have `parent_session_id != None`.
- Ignore child OpenAI plan events for `StageProjection.todos`. This preserves
the current single-list schema while making the selected list match the
stage sidebar's meaning.
- Keep Anthropic task projection unchanged. Anthropic tasks use
`anthropic_tasks:<root_session_id>`, so child-session envelopes should still
be allowed to update the shared root task list.
- Treat legacy OpenAI events without `parent_session_id` as root-compatible for
backwards compatibility.
## Implementation Units
- [ ] **Unit 1: Add reducer policy for projectable stage TODO events**
**Goal:** Make the reducer distinguish root OpenAI plan events from child
OpenAI plan events before mutating `stage.todos`.
**Files:**
- Modify: `lib/crates/fabro-store/src/run_state.rs`
**Approach:**
- Add a small helper near the TODO reducer functions, for example
`should_project_stage_todo_event(stored: &RunEvent, list_kind:
TodoListKind) -> bool`.
- Return `false` only when `list_kind == TodoListKind::OpenAiPlan` and
`stored.parent_session_id.is_some()`.
- Return `true` for root OpenAI events and all Anthropic task events.
- Call this helper in the `EventBody::TodoCreated`,
`EventBody::TodoUpdated`, and `EventBody::TodoDeleted` match arms before
resolving or mutating the stage projection.
- Leave `apply_todo_created`, `apply_todo_updated`, and
`apply_todo_deleted` focused on list mutation once the caller has decided
the event is projectable.
**Test scenarios:**
- Root OpenAI events with no `parent_session_id` still create and update
`stage.todos`.
- Child OpenAI events with `parent_session_id` do not create `stage.todos`
when no root list exists.
- Child OpenAI events do not replace an existing root OpenAI list.
- Root OpenAI updates still apply after ignored child OpenAI events.
- [ ] **Unit 2: Add focused reducer regression tests**
**Goal:** Lock the intended root-list behavior so future TODO projection work
does not regress the sidebar.
**Files:**
- Modify: `lib/crates/fabro-store/src/run_state.rs`
**Approach:**
- Extend the existing `todo_reducer` module rather than creating a new test
file.
- Add a test helper or local event setup that sets
`event.event.parent_session_id = Some(parent_session_id.to_string())` for
child-session events.
- Add one regression test that reproduces the failing sequence:
root OpenAI creates list, child OpenAI creates a different list on the same
stage, root OpenAI completes its items. Assert the final projection is the
root list and all root statuses are completed.
- Add one test proving child OpenAI events alone do not create a stage TODO
projection.
- Add one test proving Anthropic child-session task events still project.
**Verification:**
- `cargo nextest run -p fabro-store todo_reducer`
- `cargo +nightly-2026-04-14 fmt --check --all`
## System-Wide Impact
- **API compatibility:** No response schema changes. Existing consumers of
`StageProjection.todos` continue to receive a single list.
- **UI behavior:** The sidebar should show the root agent's TODO progress for
the selected stage. Child OpenAI session plans remain available only in the
raw event stream for now.
- **Historical runs:** Replaying existing event logs should produce corrected
projections because the decision uses envelope fields already persisted on
child events.
- **Future extensibility:** If child/subagent TODO display is needed later,
add a multi-list projection separately rather than overloading
`stage.todos`.
## Risks & Mitigations
| Risk | Mitigation |
|------|------------|
| Some legacy child OpenAI events lack `parent_session_id` and still project as root | Accept this for backwards compatibility; only events with explicit child-session evidence are filtered. |
| Anthropic child task updates could be accidentally filtered | Gate only `TodoListKind::OpenAiPlan`; add a regression test for `TodoListKind::AnthropicTasks`. |
| Root list replacement semantics become ambiguous if a root stage emits multiple OpenAI list IDs | Preserve current root replacement behavior; the fix only prevents child lists from replacing root lists. |
## Assumptions
- `parent_session_id == None` is the canonical signal for the root stage agent
session in stored event envelopes.
- Child OpenAI TODO lists are not part of the current stage sidebar contract.
- The correct near-term fix is projection selection, not a frontend workaround
or a schema expansion.
## Completed stages
- **toolchain**: succeeded
- Script: `command -v cargo >/dev/null || { curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y && sudo ln -sf $HOME/.cargo/bin/* /usr/local/bin/; }; cargo --version 2>&1`
- Output:
```
cargo 1.95.0 (f2d3ce0bd 2026-03-21)
```
- **preflight_compile**: succeeded
- Script: `cargo check -q --workspace 2>&1`
- Output: (empty)
- **preflight_lint**: succeeded
- Script: `cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1`
- Output: (empty)
- **fix_lints**: succeeded
- Model: claude-opus-4-7, 10.5k tokens in / 2.9k out
- Files: /home/daytona/workspace/fabro/lib/crates/fabro-api/tests/provider_round_trip.rs
- **preflight_lint**: succeeded
- Script: `cargo +nightly-2026-04-14 clippy -q --workspace --all-targets -- -D warnings 2>&1`
- Output: (empty)
- **implement**: succeeded
- Model: gpt-5.5, 539.2k tokens in / 6.0k out
- **simplify_opus**: succeeded
- Model: claude-opus-4-7, 22.8k tokens in / 6.8k out
- Files: /home/daytona/workspace/fabro/lib/crates/fabro-store/src/run_state.rs
- **simplify_gpt**: succeeded
- Model: gpt-5.5, 366.7k tokens in / 2.9k out
- **verify**: failed
- Script: `git fetch origin main 2>&1 && git merge --no-edit --no-stat origin/main 2>&1 && cargo +nightly-2026-04-14 fmt --all 2>&1 && cargo dev docs refresh 2>&1 && cargo +nightly-2026-04-14 fmt --check --all 2>&1 && { command -v rg >/dev/null 2>&1 || { echo 'rg is required for verify'; exit 127; }; } && ! rg -n 'AuthMode::Disabled|RunAuthMethod|RunSubjectProvenance|\bActorRef\b|\bActorKind\b|AuthenticatedSubject|AuthenticatedService|AuthorizeRunScoped|AuthorizeRunBlob|AuthorizeStageArtifact|AuthorizeCommandLog|auth_method\s*==\s*"disabled"' lib/crates apps lib/packages docs/public/api-reference/fabro-api.yaml 2>&1 && cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --workspace --status-level slow --profile ci 2>&1 && cargo dev docs check 2>&1 && bun install --frozen-lockfile 2>&1 && (cd apps/fabro-web && bun run typecheck) 2>&1 && (cd apps/fabro-web && bun run test) 2>&1 && (cd lib/packages/fabro-api-client && bun run typecheck) 2>&1 && cargo dev build -- -p fabro-cli --release 2>&1`
- Output:
```
(291 lines omitted)
12 │+✓ Code 5s (1 turns, 0 tools, 1.5k toks)
────────────┴───────────────────────────────────────────────────────────────────
To update snapshots run `cargo insta review`
Stopped on the first failure. Run `cargo insta test` to run all snapshots.
test commands::run::run_progress::tests::plain_verbose_snapshot ... FAILED
failures:
failures:
commands::run::run_progress::tests::plain_verbose_snapshot
test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 427 filtered out; finished in 0.06s
stderr ───
thread 'commands::run::run_progress::tests::plain_verbose_snapshot' (118993) panicked at /root/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/insta-1.46.3/src/runtime.rs:719:13:
snapshot assertion for 'plain_verbose_snapshot' failed in line 1224
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
────────────
Summary [ 10.824s] 1418/6331 tests run: 1416 passed, 2 failed, 181 skipped
FAIL [ 0.061s] (1411/6331) fabro-cli::bin/fabro commands::run::run_progress::tests::plain_default_stage_snapshot
FAIL [ 0.070s] (1417/6331) fabro-cli::bin/fabro commands::run::run_progress::tests::plain_verbose_snapshot
warning: 4913/6331 tests were not run due to test failure (run with --no-fail-fast to run all tests, or run with --max-fail)
error: test run failed
```
- **fixup**: succeeded
- Model: claude-opus-4-7, 45.8k tokens in / 5.9k out
- Files: /home/daytona/workspace/fabro/apps/fabro-web/app/routes/run-overview.test.tsx
- **verify**: failed
- Script: `git fetch origin main 2>&1 && git merge --no-edit --no-stat origin/main 2>&1 && cargo +nightly-2026-04-14 fmt --all 2>&1 && cargo dev docs refresh 2>&1 && cargo +nightly-2026-04-14 fmt --check --all 2>&1 && { command -v rg >/dev/null 2>&1 || { echo 'rg is required for verify'; exit 127; }; } && ! rg -n 'AuthMode::Disabled|RunAuthMethod|RunSubjectProvenance|\bActorRef\b|\bActorKind\b|AuthenticatedSubject|AuthenticatedService|AuthorizeRunScoped|AuthorizeRunBlob|AuthorizeStageArtifact|AuthorizeCommandLog|auth_method\s*==\s*"disabled"' lib/crates apps lib/packages docs/public/api-reference/fabro-api.yaml 2>&1 && cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings 2>&1 && cargo nextest run --workspace --status-level slow --profile ci 2>&1 && cargo dev docs check 2>&1 && bun install --frozen-lockfile 2>&1 && (cd apps/fabro-web && bun run typecheck) 2>&1 && (cd apps/fabro-web && bun run test) 2>&1 && (cd lib/packages/fabro-api-client && bun run typecheck) 2>&1 && cargo dev build -- -p fabro-cli --release 2>&1`
- Output:
```
(291 lines omitted)
12 │+✓ Code 5s (1 turns, 0 tools, 1.5k toks)
────────────┴───────────────────────────────────────────────────────────────────
To update snapshots run `cargo insta review`
Stopped on the first failure. Run `cargo insta test` to run all snapshots.
test commands::run::run_progress::tests::plain_verbose_snapshot ... FAILED
failures:
failures:
commands::run::run_progress::tests::plain_verbose_snapshot
test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 427 filtered out; finished in 0.06s
stderr ───
thread 'commands::run::run_progress::tests::plain_verbose_snapshot' (118993) panicked at /root/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/insta-1.46.3/src/runtime.rs:719:13:
snapshot assertion for 'plain_verbose_snapshot' failed in line 1224
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
────────────
Summary [ 10.824s] 1418/6331 tests run: 1416 passed, 2 failed, 181 skipped
FAIL [ 0.061s] (1411/6331) fabro-cli::bin/fabro commands::run::run_progress::tests::plain_default_stage_snapshot
FAIL [ 0.070s] (1417/6331) fabro-cli::bin/fabro commands::run::run_progress::tests::plain_verbose_snapshot
warning: 4913/6331 tests were not run due to test failure (run with --no-fail-fast to run all tests, or run with --max-fail)
error: test run failed
```
## Context
- failure_class: canceled
- failure_signature: verify|canceled|script failed with exit code: <n> ## output ────────────────────────────── to update snapshots run `cargo insta review` stopped on the first failure. run `cargo insta test` to run
The verify step failed. Read the build output from context and fix all format, clippy, Rust test, docs, TypeScript typecheck/test, and build failures.

View file

@ -0,0 +1,5 @@
{
"mode": "agent",
"provider": "anthropic",
"model": "claude-opus-4-7"
}

View file

@ -0,0 +1,12 @@
All clean. Summary:
The verify failure was caused by two snapshot tests in `lib/crates/fabro-cli/src/commands/run/run_progress/mod.rs` that hardcoded `gpt-5-mini` and expected a `$0.00` cost segment. Upstream `main` (commit 4ad9827b "Remove retired OpenAI catalog models") removed `gpt-5-mini` from the catalog, so `pricing.bill(...)` now returns `None`, the renderer omits the cost segment, and the snapshots no longer matched.
Fix: updated the two inline snapshots to reflect the new rendering (no cost section):
- `plain_default_stage_snapshot`: ` ✓ Plan $0.00 5s` → ` ✓ Plan 5s`
- `plain_verbose_snapshot`: `✓ Code $0.00 5s (1 turns, 0 tools, 1.5k toks)` → `✓ Code 5s (1 turns, 0 tools, 1.5k toks)`
Verified locally:
- `cargo nextest run --workspace --profile ci` — 6331 passed, 181 skipped
- `cargo +nightly-2026-04-14 fmt --check --all` — clean
- `cargo +nightly-2026-04-14 clippy --workspace --all-targets -- -D warnings` — clean