fabro/lib/components/fabro-workflow
Bryan Helmkamp 3e8b2ebfcc
Run the server's lifecycle over platform records instead of run_events
Step 4 of the legacy executor deletion, first commit of several: step 4
spans commits because the legacy event log and its consumers cannot go
in one compiling change. This commit moves every writer off `run_events`;
the reducer, `EventBody`, the Slate bridge and the API's event types
still exist for the readers the next commits port or delete.

Writers:
- The server records a run's lifecycle (submitted, runnable, starting,
  running, blocked, paused, control requests and effects, the terminal
  status), its title, parent link, archive state, notices and pull
  request state as platform records (`fabro_store::platform_records`),
  through the new `server::run_records` module. Every append wakes the
  projector and waits for its pass, so the read that follows a write
  holds the record.
- Pull request creation is recorded as `pull_request.requested`,
  `pull_request.created`, `pull_request.failed`, `pull_request.linked`
  and `pull_request.unlinked`; the projection folds them into the run's
  pull request and creation state.
- Answers to questions are recorded as `interview.answered` with the
  answering principal and the answer text; the interview adapter no
  longer posts legacy `interview.*` events (`QuestionSink` is now an
  optional observer).
- The worker (`fabro run __run-worker`) records its lifecycle, notices
  and pause state over `HttpPlatformRecords`; `HttpRunStore` for the
  legacy event log and the worker's `run_store` are gone.
- `persist_created_run` appends `run.created` and `run.submitted`.

Readers:
- A stream follower (`server::stream_follower`) follows each live run's
  stream (Petri events and platform records), folds lifecycle records
  into the in-memory run state, forwards items to the global attach
  broadcast, and syncs blocked and paused from the projection.
- Slack posts questions from the projection's pending interviews,
  finishes them on `interview.answered` or `question_expired`, and sends
  lifecycle notifications with `notification.sent` dedupe.
- `GET /runs/{id}/events` and the attach endpoints serve only the run
  stream; the per-event, per-stage and `POST /runs/{id}/events`
  endpoints and their tests are deleted.
- `Database::load_run_projection` reads the Petri projection only.

Deleted with the writers:
- The SQLite blob and run-history activation migrations and their
  legacy Slate imports (`legacy_blob_import`, `legacy_run_history_import`,
  the activation backup): a greenfield server has no Slate history to
  import, and the run-history verification refused to start a server
  whose runs have no legacy events.
- `fabro-workflow`'s `operations::archive` and `operations::run_store`.
- The server's legacy-event unit tests and the CLI's `HttpRunStore` tests.

The in-process answer transport is now set after the starting and
running records land, not gated on the live status still being
`Starting` (the records already moved it).

The manifest validation test for a `run.agent.mcps.<name>` catalog
reference now expects `unsupported.workflow_toml.run.agent.mcps.reference`:
Petri's Fabro frontend has no server catalog to resolve it against.

Legacy readers still fail their tests until the next commits: the
reducer and Slate tests in fabro-store, the fabro-workflow create tests
that read the run back through the legacy store, the CLI tests seeded
through `POST /runs/{id}/events`, the CLI's legacy attach and render
paths, the sessions API, the OpenAPI conformance test, and the web
fixtures.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 12:32:06 -04:00
..
src Run the server's lifecycle over platform records instead of run_events 2026-09-18 12:32:06 -04:00
Cargo.toml Delete fabro-validate and fabro-acp; validate on Petri's check 2026-09-18 11:32:04 -04:00
README.md fix(workflow): make publish failures terminal 2026-07-27 11:25:18 -04:00

fabro-workflow

A DOT-based pipeline runner for multi-stage AI workflows. Define workflows as Graphviz digraph files and execute them with pluggable handlers, conditional routing, human-in-the-loop gates, parallel branching, retry policies, and checkpoint-based recovery.

Key Concepts

  • Graph -- A directed graph parsed from DOT syntax containing nodes, edges, and attributes. The graph carries a goal describing the pipeline's purpose.
  • Node -- A workflow step. Graphviz shapes map to handler types (e.g., Mdiamond = start, Msquare = exit, box = agent, tab = prompt, diamond = conditional, hexagon = human gate, component = parallel).
  • Edge -- A connection between nodes with optional condition, label, weight, and fidelity attributes that control routing.
  • Handler -- An async trait implementation that executes a node and returns an Outcome. Built-in handlers include StartHandler, ExitHandler, AgentHandler, PromptHandler, ConditionalHandler, HumanHandler, ParallelHandler, FanInHandler, CommandHandler, and SubWorkflowHandler.
  • Outcome -- The result of executing a handler, carrying a StageOutcome (Success, Fail, PartialSuccess, Retry, Skipped), optional routing hints (preferred_label, suggested_next_ids), and context updates.
  • Context -- A thread-safe key-value store shared across pipeline stages, supporting snapshots and isolated cloning for parallel branches.
  • Interviewer -- A trait for human-in-the-loop interactions. Implementations include AutoApproveInterviewer, QueueInterviewer, CallbackInterviewer, ConsoleInterviewer, and RecordingInterviewer.
  • Checkpoint -- A serializable snapshot of execution state (completed nodes, context values) for crash recovery and resume.

Pipeline Definition

Pipelines are defined using Graphviz DOT syntax:

digraph MyPipeline {
    graph [goal="Implement and validate a feature"]
    rankdir=LR
    node [shape=box, timeout="900s"]

    start     [shape=Mdiamond, label="Start"]
    exit      [shape=Msquare, label="Exit"]
    plan      [label="Plan", prompt="Plan the implementation"]
    implement [label="Implement", prompt="Implement the plan"]
    validate  [label="Validate", prompt="Run tests"]
    gate      [shape=diamond, label="Tests passing?"]

    start -> plan -> implement -> validate -> gate
    gate -> exit      [label="Yes", condition="outcome=succeeded"]
    gate -> implement [label="No", condition="outcome!=succeeded"]
}

Usage

Parsing and Validating a Pipeline

use fabro_workflow::operations::{create, CreateOptions};

let dot_source = r#"digraph Simple {
    graph [goal="Run tests"]
    start [shape=Mdiamond]
    exit  [shape=Msquare]
    work  [shape=box, prompt="Run the test suite"]
    start -> work -> exit
}"#;

let validated = create(dot_source, CreateOptions::default())
    .expect("pipeline should parse");
validated.raise_on_errors().expect("pipeline should validate");
let (graph, _, _) = validated.into_parts();
assert_eq!(graph.name, "Simple");
assert_eq!(graph.goal(), "Run tests");

operations::create parses the DOT source, applies built-in transforms (variable expansion, stylesheet application, preamble injection), and returns diagnostics through Validated.

Running a Pipeline

use fabro_workflow::operations::start;
use fabro_workflow::pipeline;

// Use `operations::start(...)` for the full
// initialize -> execute -> conclude -> publish -> finalize flow.
// Use `pipeline::initialize(...)` + `pipeline::execute(...)` when you need partial lifecycle control.

Custom Handlers

Implement the Handler trait to add custom node behavior:

use arc_workflows::handler::Handler;
use arc_workflows::context::Context;
use arc_workflows::graph::{Graph, Node};
use arc_workflows::outcome::Outcome;
use arc_workflows::error::ArcError;
use async_trait::async_trait;
use std::path::Path;

struct MyHandler;

#[async_trait]
impl Handler for MyHandler {
    async fn execute(
        &self,
        node: &Node,
        context: &Context,
        graph: &Graph,
        run_dir: &Path,
    ) -> Result<Outcome, ArcError> {
        // Custom logic here
        Ok(Outcome::success())
    }
}

Model Stylesheets

CSS-like stylesheets control LLM model assignment with specificity-based cascading:

digraph Styled {
    graph [
        goal="Build feature",
        model_stylesheet="
            * { model: claude-sonnet-4-5;}
            .code { model: claude-opus-4-6; }
            #critical_review { model: gpt-5.2;}
        "
    ]
    // ...
}

Selectors by specificity: * (universal, 0) < shape (1) < .class (2) < #id (3). Explicit node attributes are never overridden.

Condition Expressions

Edge conditions use a simple expression syntax for routing:

outcome=succeeded
outcome!=failed
outcome=succeeded && context.tests_passed=true
my_flag

Clauses support =, !=, and bare key truthiness checks, joined with &&.

Human-in-the-Loop Gates

Nodes with shape=hexagon or type="human" pause execution for human input. Outgoing edge labels become selectable options, with accelerator key parsing for patterns like [A] Approve and F) Fix.

Parallel Execution

Nodes with shape=component fan out to branches concurrently. Branches receive isolated context forks, share the same sandbox checkout, and always finish before the workflow continues. Use max_parallel to limit concurrency; concurrent workspace writes are user-managed.

Checkpoints and Resume

The engine saves a checkpoint after each node. Resume from a checkpoint with engine.run_from_checkpoint(&graph, &config, &checkpoint).

Architecture

parser (DOT -> AST -> Graph)
  -> transform (variable expansion, stylesheet, preamble)
    -> validation (14 lint rules)
      -> engine (execution loop with retry, edge selection, goal gates)
        -> handler (pluggable node executors)
          -> interviewer (human-in-the-loop I/O)