fabro/docs/execution/checkpoints.mdx
Bryan Helmkamp 551d16b74c refactor: clean CREATE/START/RESUME separation
Resume now follows the same subprocess pattern as run: look up run
directory by ID prefix, validate checkpoint exists, clean stale
artifacts, reset status to Submitted, spawn _run_engine --resume, and
attach. This eliminates ~1600 lines of duplicated env/sandbox setup
from resume.rs.

Key changes:
- operations::start() and operations::resume() take run_dir instead
  of Persisted, loading state from disk internally
- run_engine() builds RunOptions from RunRecord on disk, so callers
  no longer extract record fields manually
- StartOptions flattened (no more nested InitOptions)
- FabroError::Precondition variant for start/resume guard checks
- _run_engine accepts --resume flag to dispatch to resume path
- operations::restore removed (no longer needed)
- Resume CLI stripped to just <RUN_ID> + --detach

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 13:47:07 -04:00

187 lines
8.4 KiB
Text

---
title: "Checkpoints"
description: "How Fabro uses Git to checkpoint and resume workflow runs"
---
Fabro checkpoints every workflow run using Git. After each node completes, Fabro commits the file changes and execution state so that interrupted runs can be resumed exactly where they left off. This happens automatically — no configuration required beyond running inside a Git repository.
## Two branches, two purposes
Each run creates two Git branches that work in tandem:
| Branch | Ref format | Contains |
|---|---|---|
| **Run branch** | `fabro/run/{run_id}` | File changes made by agents and commands — the actual work product |
| **Metadata branch** | `fabro/meta/{run_id}` | Checkpoint JSON, the workflow graph, run and start records, and offloaded artifacts |
The run branch is a regular Git branch that grows one commit per completed node. The metadata branch is an orphan branch (no shared history with your code) that stores structured data using Git's object database directly — no working tree needed.
### Run branch commits
After each node finishes, Fabro stages file changes and creates a commit on the run branch. Files matching `[checkpoint] exclude_globs` patterns (configured in [run.toml](/execution/run-configuration#checkpoint) or [server.toml](/administration/server-configuration#checkpoint-section)) are excluded from staging:
```
fabro(01JKXYZ...): plan (success)
Fabro-Run: 01JKXYZ...
Fabro-Completed: 2
Fabro-Checkpoint: a1b2c3d4...
```
The commit message follows a structured format:
| Part | Description |
|---|---|
| Subject line | `fabro({run_id}): {node_id} ({status})` |
| `Fabro-Run` trailer | The run ID |
| `Fabro-Completed` trailer | Number of completed nodes so far |
| `Fabro-Checkpoint` trailer | SHA of the corresponding commit on the metadata branch |
The `Fabro-Checkpoint` trailer links each run branch commit to its metadata branch commit, so you can navigate from file changes to the full execution state and back.
### Metadata branch
The metadata branch (`fabro/meta/{run_id}`) is an orphan branch that stores structured run data using Git's object storage directly (via `git2`). It is initialized at run start with:
- **`run.json`** — Run record: run ID, created_at, config, graph, workflow slug, working directory, host repo path, base branch, labels
- **`start.json`** — Start record: run ID, start time, run branch, base SHA
After each node, the metadata branch is updated with:
- **`checkpoint.json`** — Full execution state (see below)
- **`artifacts/*.json`** — Any offloaded artifact data (large context values over 100KB)
- **`nodes/{node_id}/`** — Per-node execution trace files (prompts, responses, status, diffs — files under 512KB from an allowlist)
## What's in a checkpoint
The `checkpoint.json` captures everything needed to resume a run:
| Field | Description |
|---|---|
| `timestamp` | When the checkpoint was created |
| `current_node` | The node that just completed |
| `next_node_id` | The next node the engine would execute |
| `completed_nodes` | Ordered list of all completed node IDs |
| `node_retries` | How many retry attempts each node has used |
| `node_outcomes` | Full outcome (status, context updates, usage) for each completed node |
| `context_values` | Snapshot of the entire [run context](/execution/context) |
| `git_commit_sha` | SHA of the run branch commit at this checkpoint |
| `loop_failure_signatures` | Failure signature counts for loop detection |
| `restart_failure_signatures` | Failure signature counts across loop-restart edges |
The checkpoint is also saved to `checkpoint.json` in the run directory for quick local access.
## Worktrees
Fabro uses Git worktrees to isolate workflow runs from your working directory. When a run starts in a clean Git repository:
1. Fabro records the current HEAD as the **base SHA**
2. Creates a new branch `fabro/run/{run_id}` at that SHA
3. Adds a worktree at `{run_dir}/worktree` on that branch
4. Changes into the worktree directory for the duration of the run
This means your original working directory stays untouched while the agent makes changes in the worktree. When the run completes, Fabro removes the worktree and restores your original directory.
<Note>
If the working directory has uncommitted changes, Fabro skips worktree setup and runs in place, logging a warning. Git checkpointing is disabled in this case.
</Note>
For Daytona sandboxes, the worktree is created inside the remote sandbox instead. The metadata branch is still written to the host repository so that runs can be resumed locally. Both the run branch and the metadata branch are pushed to origin after each checkpoint — the run branch is pushed from the sandbox, while the metadata branch is pushed from the host using a GitHub App installation token.
## Resuming a run
Resume an interrupted run from its checkpoint on disk:
```bash
fabro resume 01JKXYZ
```
Fabro looks up the run directory by ID prefix, loads `checkpoint.json` and `run.json` from the run directory, and spawns a new engine process to continue execution. No workflow file or override flags are needed — all configuration is read from the persisted run state.
<Accordion title="What happens during resume">
1. Fabro looks up the run directory by ID prefix
2. Validates that `checkpoint.json` exists and no engine process is already running
3. Cleans stale artifacts from the previous execution (conclusion, PID file, etc.)
4. Resets status to `Submitted` and spawns a new engine subprocess with `--resume`
5. The engine loads `run.json` and `checkpoint.json`, restores the full context, completed node list, retry counts, and failure signatures
6. If the checkpointed node used `full` fidelity, downgrades the first resumed node to `summary:high` (since the original conversation thread no longer exists in memory)
7. Continues execution from `next_node_id`
</Accordion>
## The checkpoint cycle
Here's the full sequence that runs after every node completes:
1. **Save checkpoint to disk** — Write `checkpoint.json` to the run directory
2. **Write metadata branch** — Serialize the checkpoint and any new artifacts to the metadata branch (shadow commit)
3. **Commit to run branch** — Stage all file changes, commit with structured trailers linking to the shadow commit SHA
4. **Update checkpoint** — Re-save `checkpoint.json` with the `git_commit_sha` field set
Steps 2-4 are best-effort — if any Git operation fails, the run continues and emits a `RunNotice` warning event. The disk checkpoint from step 1 is always available as a fallback.
## Inspecting run history
Because checkpoints are plain Git commits, you can inspect them with standard Git tools:
```bash
# View the commit log for a run
git log fabro/run/01JKXYZ... --oneline
# See what an agent changed at a specific node
git show fabro/run/01JKXYZ...
# Diff the full run against the starting point
git diff main..fabro/run/01JKXYZ...
# Read checkpoint data from the metadata branch
git show fabro/meta/01JKXYZ...:checkpoint.json | jq .current_node
```
## Rewinding to an earlier checkpoint
If a later stage goes off-track, you can rewind a run to an earlier checkpoint and resume from there instead of restarting the entire workflow:
```bash
# List the checkpoint timeline
fabro rewind <RUN_ID> --list
# Rewind to a specific checkpoint
fabro rewind <RUN_ID> plan@2
# Resume from the rewound point
fabro resume <RUN_ID>
```
See [`fabro rewind`](/reference/cli#fabro-rewind) for the full command reference.
## Forking a run
If you want to explore an alternate path from a checkpoint without losing the original run's history, use `fabro fork` instead of `fabro rewind`. Fork creates a new independent run branching from the target checkpoint — the original run stays intact.
```bash
# List checkpoints
fabro fork <RUN_ID> --list
# Fork from a specific checkpoint
fabro fork <RUN_ID> plan@2
# Resume the forked run
fabro resume <NEW_RUN_ID>
```
Use **rewind** when you want to redo a run from an earlier point (destructive — resets the original). Use **fork** when you want to try a different approach while keeping the original run as a reference.
See [`fabro fork`](/reference/cli#fabro-fork) for the full command reference.
## When checkpointing is active
Git checkpointing activates automatically when:
- The working directory is a clean Git repository (local and Docker sandboxes)
- The sandbox is Daytona (metadata branch on the host, commits inside the sandbox)
It is skipped when:
- The working directory has uncommitted changes
- The working directory is not a Git repository
- The run uses `--dry-run`