Commit graph

5212 commits

Author SHA1 Message Date
Bryan Helmkamp
3e157d3356
Format the Petri test fixtures after the engine key removal
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:54:45 -04:00
Bryan Helmkamp
1f0dbd86ae
Delete fabro-hooks and the engine freeze check
`fabro-hooks` ran the legacy executor's hooks; Petri's Attractor steps
run Fabro's hooks now, so nothing in the workspace uses the crate. The
engine freeze (the CI workflow, the two scripts, and the AGENTS.md and
fabro-petri README sections) guarded the engine half of `fabro-workflow`,
which the previous commit deleted.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:47:56 -04:00
Bryan Helmkamp
d90a5d9cbb
Delete fabro-core and the engine half of fabro-workflow
Every run executes on Petri, so the in-process legacy executor goes:
`fabro-core` and, in `fabro-workflow`, the handlers, lifecycle, pipeline
execution, routing, retry, conditions, node handlers, steering, agent
memory, artifacts, checkpoints, command log, and the `start`, `resume`,
`retry`, `fork`, `rewind` and `timeline` operations. The two are deleted
together because the engine half of `fabro-workflow` was the only user of
`fabro-core` and `fabro-core` the only runtime of that half; neither
compiles without the other.

Kept in `fabro-workflow`, narrowed: the parse/transform/validate/persist
pipeline and `create`, `archive`, `validate` (workflow definitions still
come from DOT and settings); the run tools (`run_tools`, moved from
`handler/llm/fabro_tools.rs`) for Ask Fabro, `fabro exec` and Petri's
host tools; the pull request pipeline (`pull_request`, moved from
`pipeline/`, for the step 0 port); Run Files' diff helpers in
`sandbox_git`; `git_identity`, `usage_rollup`, `run_status`,
`run_materialization`, `web_search` and `workflow_bundle`.

Server: `RegistryFactoryOverride` becomes `execute_in_process`;
`RunAnswerTransport::InProcess` carries only the interviewer; the
interrupt endpoint answers 501 `interrupt_unsupported` and every pair
endpoint 501 `pair_unsupported` (status lists none); rewind, fork, retry
and timeline handlers and routes are removed; the command log is served
from the stage output blob; usage rollups accumulate from the settled
projection after an in-process run as after a worker exit.

Ported while here:

- `materialize_admitted_run` materializes the goal and drops a disabled
  pull request block, as the legacy materializer did.
- A run whose admitted graph has an agent or prompt node is refused at
  create when no LLM provider is ready (`fabro.model.no_ready_provider`);
  a workflow of commands and gates needs no model and is admitted.
- The projection's question type falls back on the options, as the
  interview adapter does, so a gate with edge-label options answers as
  multiple choice.

Tests: the server scenarios (lifecycle, run completion, SSE, helpers)
run in process on Petri and assert Petri's stage labels and stream
names; the reconcile tests assert Petri's relaunch semantics; legacy
unit tests of the deleted executor are removed; three server unit tests
the removal took with it are restored; the pair fixtures go with the
pair feature. Petri test fixtures no longer name `[workflow] engine`.

Still red after this commit, all legacy consumers the next steps
delete or port: fabro-store's Slate/reducer fixtures and fabro-types
legacy JSON tests (step 4); server unit tests over legacy run events
(retry endpoints, list_run_events, artifacts, per-event pause/unpause,
run history activation, legacy sandbox fixtures) (steps 3-4); CLI tests
that parse legacy event envelopes, the legacy `events`/`attach`/`diff`/
`dump`/`inspect` snapshots, `run rewind`/`run fork`, the ACP and
git-identity workflow tests, and the runner tests that drive the legacy
worker by hand (steps 3-4); the web app's Petri fixtures still carry
`engine` (regenerate with `FABRO_CAPTURE_PETRI_FIXTURES` in step 4).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:44:41 -04:00
Bryan Helmkamp
a36bea15d2
Remove the engine flag: every run is a Petri run
Delete `Engine`, `RunEngine`, `[workflow] engine`, `[server.execution]
engine`, `FABRO_SERVER_ENGINE` and `fabro server start --engine`. The run
spec records what Petri admitted as `admission: PetriAdmission`; the
create handler always admits through `Runtime::check`; `execute_run`
always launches the Petri worker (or executes in process under the test
override); the CLI runner takes only the Petri worker path, and its
legacy control arm, artifact uploader, signal pause handlers and
credential helpers go with it. The CLI's `attach` and `events` read the
run stream only.

Two gaps this surfaced are closed here: the check adapter binds the
server's run variables as Petri compile variables (`{{ vars.* }}` in a
prompt no longer fails admission), and deleting a run removes its Petri
records, lease, platform records, projection and stream.

Tests: the config engine tests are replaced (an engine key is unknown),
the API round-trip test covers `PetriAdmission`, the server and CLI
Petri scenarios drop their engine settings, and the API tests that read
legacy event names now read the run stream or the session events. The
remaining red tests are fixtures and scenarios of the legacy executor
and the legacy event store (`fabro-store` `slate` and `run_state`,
`fabro-types` legacy `run.created` JSON, the server's handler-registry
scenarios, the CLI dry-run snapshots), which the next steps of the F4.3
series delete or port.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:58:21 -04:00
Bryan Helmkamp
b5fdc015f0
Merge branch 'petri-integration-sandbox' into petri-integration
# Conflicts:
#	lib/apps/fabro-cli/tests/it/scenario/mod.rs
2026-09-18 09:11:19 -04:00
Bryan Helmkamp
d60744b4d6
Checkpoint and recover Petri workspaces inside Docker and Daytona sandboxes
A Petri run on Docker or Daytona keeps its workspace inside the scope's
sandbox. Fabro's hooks now take the environment Petri hands them at
`scope_acquired`, run `git` inside the scope through it (the path Petri's
own sandbox-placed hooks take), and commit each stage on the run branch
with the same message and trailers as the host path. The commit leaves
the sandbox as a Git bundle, created against the newest ancestor the
snapshot repository already holds, split into 8 MiB parts (the plugin
transport reads one file up to 16 MiB), read out through the
environment's file transfer, fetched into the bare snapshot repository
on the host and named there under the checkpoint's ref. The repository
holds every checkpoint whatever the provider, and the platform records
name the same commits. The host path is unchanged; both sites share one
runner and the same commands.

Recovery is split: `recovery::plan` decides, over the records and the
snapshot repository alone, what every live workspace must sit on and
reconciles a lost record; `recover` applies it to host workspaces on the
server, as before, and reports a sandbox workspace's target as deferred.
The worker's hooks read the same plan at the scope's first acquisition
after a resume and bring the sandbox workspace to it before any attempt
runs there: verified or reset in a retained sandbox that still holds the
commit, else restored from a bundle of the checkpoint written into the
sandbox. Petri replaces a lease's lost sandbox on Fabro's request
(`LostSandbox::Replace`), so a removed container comes back fresh and
restored.

The checkpoint records are written for every provider now. A Docker
variant of the in-process hooks test moves a 20 MiB file through the
split transfer; the same test runs on Daytona when live credentials are
present. Three CLI scenarios run on a Docker environment: every stage's
checkpoint published from the container, a retained container whose
workspace drifted reset on restart, and a removed container replaced and
restored from the snapshot.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:10:29 -04:00
Bryan Helmkamp
97f74c3ab0
Forward the Docker daemon selection and Daytona credentials to Petri workers
A Petri run's worker launches the sandbox-driver plugins itself, and the
Docker plugin forwards `DOCKER_HOST`, `DOCKER_TLS_VERIFY`,
`DOCKER_CERT_PATH`, `DOCKER_API_VERSION`, `DOCKER_CONFIG` and
`DOCKER_CONTEXT` from the process that launches it. They now cross the
worker's environment allowlist, so the worker's sandboxes go to the daemon
the server uses. The concern that kept them out, the legacy worker's own
Docker client picking up a daemon it was not meant to, is moot: the
legacy executor is being deleted. The same variables pass through the
test harness's isolation, so a developer's daemon selection reaches the
servers tests start.

Daytona's non-secret selectors, `DAYTONA_API_URL` and
`DAYTONA_ORGANIZATION_ID`, cross the allowlist too. The API key stays the
vault's: a Daytona run's worker command carries it the way the GitHub app
key travels, and the Daytona plugin reads it from the worker's process.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:10:29 -04:00
Bryan Helmkamp
2a000cc7d3
Move the Petri pins to 4d4bdd6
Petri's `scope_acquired` hook point and `SandboxOptions::lost_sandbox`,
which the sandbox checkpoints and their recovery build on.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:10:29 -04:00
Bryan Helmkamp
fc9b82aa7f
Note the worker's paused mirror and refused steers in VIEWS.md
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 07:57:19 -04:00
Bryan Helmkamp
4726eed214
Cover pause, steer and a paused resume on Petri through the real binary
Three scenarios over `petri.rs`'s harness: a pause between two command
stages holds the second until the unpause while the API says `paused`
with no pending control; a steer sent while the agent stage waits on a
tool reaches its session on the twin, which sees the text in its next
request, and the stream carries the `control.requested` record; a run
paused with its next stage held at admission, whose server and worker
then die, resumes paused, admits nothing until the unpause, and then
finishes. The harness helpers the sibling module needs are opened to it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 07:57:10 -04:00
Bryan Helmkamp
5093265efb
Wire pause, unpause and steer into the Petri worker
A Petri run answered only cancel and answers; pause, unpause and steer
were ignored with a warning. `fabro_petri::controls::RunControls` now
wraps Petri's `ControlService` per run: `engine::run` installs its pause
gate over the run's hooks, observes the run through it and wires it to
the coordinator, on a start and a resume alike, so a run paused when its
worker died resumes paused.

The worker's control channel takes a `WorkerControls` enum: the legacy
hub and pause flag, or the Petri run's controls. On Petri, `run.pause`
holds admission, `run.unpause` releases it once the record is durable,
and `run.steer` goes to the one live agent stage (Fabro's steer names no
stage); with none or several it is refused with a `run.notice` record.
The paused state is mirrored to Fabro's lifecycle as `run.paused` and
`run.unpaused` events, so the server's live status and the projection
follow Petri's own records.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 07:49:54 -04:00
Bryan Helmkamp
806ac90c98
Freeze the engine half of fabro-workflow to bug fixes
The integration plan's F4.1: once Fabro runs on Petri, the engine half of
fabro-workflow (handler/, lifecycle/, pipeline/execute, graph/routing,
node_handler, retry, condition, context, model_fallback) takes bug fixes
only, and new engine behaviour goes to Petri.

scripts/check-engine-freeze.sh holds the frozen path list, diffs the
branch against a base ref and exits 1 when any frozen file gained lines;
scripts/check-engine-freeze-test.sh proves that on a synthetic
repository. The Engine freeze workflow runs both on every pull request
that touches the crate's src, re-runs on label changes, and fails unless
the pull request carries the `bugfix` label. AGENTS.md and the
fabro-petri README name the freeze, the label and the script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 02:19:29 -04:00
Bryan Helmkamp
6e38ad27fb
Order position-keyed platform records by their firing, not their clock
A checkpoint record is stamped by the server and Petri's records by the
worker, so ordering the stream by recorded_at could put a stage's
checkpoint after the next stage's route. A platform record that carries
a Petri position now goes right after its firing's finish: before the
firing's first routing.resolved in the pass (the next firing's
visit.started hangs off that record), else after the firing's last
event, else, when the firing finished in an earlier pass, before the
first event of a later firing. The hook writes the record after the
driver appended the attempt's finish, but the driver's store writer
flushes on its own schedule, so the record can be committed before its
firing's step.finished; such a record is held back, with every platform
record after it, until the finish is in the stream, or the run finished.
The fold now remembers which firings finished. Unit tests cover the
placement, the earlier-pass case, an unpositioned record, the hold and
its release; the CLI read-back scenario is stable over eight runs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 02:15:53 -04:00
Bryan Helmkamp
5d93f20fb0
Merge branch 'petri-integration-tools' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/apps/fabro-cli/src/commands/run/petri_worker.rs
#	lib/apps/fabro-cli/tests/it/scenario/petri.rs
#	lib/apps/fabro-server/src/server/petri_runs.rs
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/src/lib.rs
2026-09-18 02:02:51 -04:00
Bryan Helmkamp
23c422f9a2
Cover the run tools inside a Petri run through the server and its worker
Three scenarios on the real binary: an agent creates a child run with
`fabro_run_create` from inside a Petri run and the child carries the
parent link; a `[[run.hooks]]` pre_tool_use hook blocks a run tool, the
model reads the reason, and Petri's record holds the report and the
denied call; a sub-agent calls an inherited run tool, recorded under the
parent stage naming the parent session.

The Petri scenario harness is shared: the server can start with extra
settings and vault entries, and the detached run takes extra arguments.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 01:52:46 -04:00
Bryan Helmkamp
49e70e481d
Register Fabro's run tools on a Petri run through the host tool capability
Plan item F3.4. `fabro_petri::host_tools` adapts Petri's `HostTools`
capability to `register_fabro_run_tools`: every native agent session of a
run gets the tools the legacy worker registers, bound to the worker's
client and the run id, so a child run a stage creates is parented to the
Petri run. The tools run under the run's tool hooks, are recorded under
the stage, and reach sub-agents through Pebble's inheritance.

`RuntimeSpec::run_tools` installs the capability; the worker sets it when
the run's settings enable `[run.agent] fabro_tools` and the worker token
carries `agent:run_tools`, the legacy worker's gate. The server's
in-process test path runs without them, like the legacy one.

The identity the tools need is the run id alone; no run tool records a
stage on an effect, so nothing derives Fabro's `node@visit` label. A
context for another run gets no tools.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 01:52:46 -04:00
Bryan Helmkamp
26429a7444
Merge branch 'petri-integration-api' into petri-integration
# Conflicts:
#	lib/apps/fabro-cli/tests/it/scenario/petri.rs
2026-09-18 01:39:53 -04:00
Bryan Helmkamp
bfc2abebab
Wait for the terminal lifecycle record before reading a stream's end
Fabro's terminal `run.lifecycle` record lands a moment after Petri's
`run.finished`: the worker exits, the server records the status, the
projector folds it. A CLI scenario that asserts on the end of the stream
now waits for that record instead of reading the stream as soon as the
runs row turns `succeeded`, which the projector writes from the engine's
finish alone.

The fabro-petri README names the projector's stream reader and commit
signal, the server's reconnect test with its fixture capture, and the
CLI scenarios that read a run back through the stream.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 01:25:40 -04:00
Bryan Helmkamp
1bfe62577d
Read a Petri run through the CLI from its stream
`run events` on a Petri run prints the run stream: raw, the envelope as
one JSON line per item; `--pretty`, Petri's events by `<subject>.<verb>`
with the stage's label (a visit's start and end with its elapsed time,
the route both ends of the edge, a fork's branches, a question with its
options and its answer, log lines, the agent's messages and tool calls,
the engine's finish) and the platform records by kind (the run's
creation, its lifecycle, a checkpoint's commit, a pull request, a
notice, who answered). `--follow` attaches from the last `stream_seq`
printed and reconnects from its cursor when the server ends the stream
before the run's terminal record.

`run attach` on a Petri run replays the stream through the progress
renderer (a new mapping from stream items onto the progress events the
renderer draws, sharing the coding-agent mapping with the legacy
envelope), follows it live from its cursor with the same reconnect, asks
a question the stream carries at the terminal, and exits with the status
the engine's finish or the terminal lifecycle record decides. `wait` and
`inspect` read the projection unchanged.

The CLI never names a Petri type: `PetriItem` reads the item as JSON
where Petri's contract keeps the event name, the subject and the parsed
progress payloads.

The CLI's Petri scenarios read the stream instead of the legacy events
(the lifecycle records, the question and who answered it, the expiry),
and three new ones cover a finished run through `events` (raw, tail,
and a `--pretty` snapshot), `attach`, `wait` and `inspect`; `attach`
answering a gate from the terminal; and `events --follow` to the run's
end.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 01:21:12 -04:00
Bryan Helmkamp
3941a24a28
Test the web app's Petri views over the captured scenario fixtures
The Pebble envelopes a step records are read as `CodingAgentEvent`s
(`{seq, stream_id, session_id, timestamp, event: {Variant}}`), so an
agent stage's chat shows the `UserInput` prompt and the assistant's
answer; a command step's exit status comes from its output; a user's
login names who answered a gate.

`lib/petri-stream.test.ts` checks the derivations over the hello, command,
parallel and gate fixtures: stream density, names, stage labels with the
fork's delegates skipped, the gate's question and answer with the
principal, the run phases from the lifecycle records, the notice between
the branches, the fork's branches from the projection, the edge a stage
took, and the envelopes. `run-events.test.tsx` checks the SWR keys a
stream item invalidates and the run filter on the coordinated stream.
`run-petri.render.test.tsx` renders the stage list, the chat, the
parallel children, the fan-in, the events list, the waterfall, the Q&A,
the decision and the platform records for each fixture.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 01:03:19 -04:00
Bryan Helmkamp
f22cf18c86
Render a Petri run's detail views from its projection and stream
The web app reads a Petri run through the run stream (`useRunStream`
pages `GET /runs/{id}/events` by `after`) and the projection, keyed on
`RunSpec.engine`, beside the legacy event path. `lib/petri-stream.ts`
holds the pure derivations `VIEWS.md` maps: the stage label of an item's
subject (lowering nodes skipped), the interview pairs from `parsed.question`
and the delivered `control.requested` with the `interview.answered`
principal, the run phases from the platform lifecycle records, the edge a
`route.applied` took, the fork's branches and the fan-in transcript from
the projection, the stage context from the final `step.finished`, the
Pebble envelopes of a step, and the debug rows the listings show.

The events route lists stream rows (named `<subject>.<verb>` or by the
platform kind, with the stage beside them and the raw item in the details
panel) and the waterfall takes its phases from the lifecycle records. The
stages route builds a Petri stage's turns from the projection's prompt and
response and the step's envelopes, its debug tab from the stage's items,
and hands the human, conditional, parallel and fan-in renderers the
derived data instead of events. The overview lists the platform records
(checkpoints with their commit, pull request, notices). The SSE
subscription invalidates SWR keys from a stream item's event name or
platform kind, ending on the terminal lifecycle record, and the cross-tab
dedupe keys a stream item by its run and `stream_seq`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:59:10 -04:00
Bryan Helmkamp
c9c4c27753
Merge branch 'petri-integration-hooks' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/apps/fabro-cli/src/commands/run/petri_worker.rs
#	lib/apps/fabro-cli/tests/it/scenario/petri.rs
#	lib/apps/fabro-server/src/server/petri_runs.rs
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/src/engine.rs
#	lib/components/fabro-petri/src/lib.rs
#	lib/components/fabro-store/src/platform_records.rs
2026-09-18 00:51:21 -04:00
Bryan Helmkamp
523f831a7a
Serve a Petri run's events as one stream with a stream_seq cursor
`GET /runs/{id}/events` and `GET /runs/{id}/attach` serve a Petri run's
public events and Fabro's platform records as one ordered stream in a
Fabro envelope (`RunStreamItem`: `run_id`, `stream_seq`, `kind`, `id`,
`recorded_at`, `item`), read from the projector's `petri_stream` table.
The cursor is `stream_seq` (`?after=`); the item's own identity (the
Petri `EventId` as `<log>/<seq>/<index>`, or the platform record's seq)
travels beside it for deduplication. A legacy run keeps its envelope on
the same endpoints; the OpenAPI response is the union of the two lists,
and the stream list reports Petri's `EVENT_CONTRACT_VERSION`.

The attached stream follows the projector's commit signal (a wake-up,
with a poll as the fallback) and ends after the platform record of the
run's terminal lifecycle transition, the analog of the legacy stream's
`run.completed`, or a bounded grace after the projection went terminal.

`RunSpec.engine` (`RunEngine`, `PetriAdmission`, `PetriGraphRef`) is
named in the spec and reuses the Rust types. `fabro-client` matches the
union and adds `list_run_stream`, `list_run_stream_page` and
`attach_run_stream`.

A server test attaches to a two-branch parallel run, disconnects once
both branches started, records a platform notice while both branch
scripts run, reconnects from the last `stream_seq`, and checks the
union is the whole stream: every item once, in order, no gap, no
duplicate, the notice between the branch events, and the same as the
paged listing. The Petri scenarios capture their settled projection and
stream as JSON fixtures for the web app under
`FABRO_CAPTURE_PETRI_FIXTURES`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:47:26 -04:00
Bryan Helmkamp
34139b1936
Prove the checkpoint hooks and the recovery protocol
In-process tests over the memory store: every finish is committed on the
run branch with its identity trailers and recorded with its commit, the
run-end hooks reach Petri's local service through Fabro's wrapper, a
stage that fails on its own terms is committed and its failure route runs
on the committed files, a failed checkpoint records `checkpoint_failed`
with no route taken and a restart reports the run failed, and a
`[[run.hooks]]` hook blocks an agent's tool call through the forwarded
service, with the model told why.

Real-binary scenarios crash the server and its worker with SIGKILL: after
a durable finish the stage's commit is not repeated and the interrupted
stage reruns on its snapshot; a crash held before the commit reruns the
stage once; a crash held after the commit but before its record
reconciles the record from the snapshot repository; a deleted workspace
is restored; a failure route sees the same committed files after a
crash; a failed checkpoint fails the run and a restart leaves it failed.

Recovery selects the executions `inspect_run` reports incomplete, and a
run whose coordinator log is still empty is left to the worker's resume.
The worker's platform record endpoints get an API test and the generated
TypeScript client.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:25:10 -04:00
Bryan Helmkamp
a6ac3f120e
Give a Petri question one identity across the adapter and the projection
The interview adapter derived its own question id from Petri's identity
and posted it on `interview.started`, while the projection over Petri's
records serves the pending question under Petri's `Question.id` with the
firing's stage label. The answer endpoint validates against the
projection, so an answer under the projection's id never reached the
adapter's wait.

The adapter now waits under Petri's id and labels the question's stage
through the projection's own rule: `stage_label`, `is_shown` and
`visit_of` move out of `start_visit` into shared functions, and the
adapter's observer derives each firing's `visit.started` through Petri's
`Projection`, as the projector does, so the label matches by
construction. The full Petri identity stays on `AskedQuestion`.

The legacy `interview.*` events are still posted, under Petri's id, for
the readers that follow the event stream rather than the projection: the
Slack service, `run attach`, the web app's Q&A renderer and the server's
answer claim. The store already derives the `interview.answered` platform
record from `interview.completed` for a Petri run, so who answered is
recorded under Petri's id with the answering principal.

The gate scenarios assert the new identity and encode the id as one path
segment, as the generated clients do. Projection tests cover an expired
question and an auto-approved answer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:19:30 -04:00
Bryan Helmkamp
01beea0a6c
Checkpoint a Petri run's stages and recover its workspaces on restart
Fabro's hooks on a Petri run wrap the hooks the runtime installed for
`[[run.hooks]]` and forward every point. In `prepare_result`, before the
finish is recorded, they commit the stage's files on the run branch of
its host workspace with Fabro's author identity and the run, execution,
firing and attempt as trailers, and publish the commit to a snapshot
repository beside the run's workspaces under a ref per checkpoint. A
stage that failed on its own terms is committed like a successful one; a
commit that fails is fatal: the outcome becomes a `checkpoint_failed`
failure, the run is cancelled through the coordinator handle, and the
transition refuses the firing's routes. In `transition` they write the
platform checkpoint record, keyed on the Petri position and the
checkpoint's operation identity, and a failed write is a recorded
problem.

On restart the server runs the recovery protocol before it relaunches a
worker: a run with a failed checkpoint is reported failed; otherwise
every live execution's last durable finish names the snapshot its
workspace is verified against, reset to, or restored from, with a lost
record reconciled from the snapshot repository, and a finish with no
snapshot fails the run rather than resume it on stale files.

The worker reaches the platform records over two new worker-scoped
endpoints; the server reaches the table directly. A test gate directory
lets the CLI scenarios hold a checkpoint at a named point.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:12:14 -04:00
Bryan Helmkamp
5dae891a98
Merge branch 'petri-integration-read' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/README.md
#	lib/components/fabro-petri/src/lib.rs
2026-09-17 23:48:36 -04:00
Bryan Helmkamp
ee4850689e
Name the per-run pass lock's type through an import
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:58:31 -04:00
Bryan Helmkamp
4a20acafc7
Run one projection pass at a time per run
The startup pass called the pass directly while a signalled pass could
run for the same run, so both read one committed stream sequence and
the second insert into the stream failed on its primary key, which
stopped the restarted server. Passes now take a per-run lock, and a run
whose startup pass fails is logged and left for its next signal instead
of stopping the server. A test races four passes, a signal and the
startup pass over one run and checks the stream stays contiguous.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:56:30 -04:00
Bryan Helmkamp
e0b546d465
Read the view tables where the run summary store keeps them
The projector takes two pools: the one Petri's records live in and the
one the view tables live in. In the server both are the one database;
a test fixture keeps the runs row, the platform records and the
projection tables in the run summary store's own pool, which the
projector was not reading, so a run projected in a test server folded
its Petri events before its run.created record. The startup run-history
verification checks only a Petri run's identity and legacy guard, since
its row is the projector's. An agent stage's response is the
response.<node> its outcome wrote into the run context, as the prompt
step writes it. The scenario tests assert each branch's own index.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:49:15 -04:00
Bryan Helmkamp
53b16b591d
Merge branch 'petri-integration-adapters' into petri-integration 2026-09-17 22:40:30 -04:00
Bryan Helmkamp
a316181a17
Give fabro-petri's workflow tests the same test timeout as the apps
The adapter tests run whole workflows on the host sandbox through the
plugin, and one calls the twin; under a full parallel run one of them was
killed at the default 3 s.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:35:20 -04:00
Bryan Helmkamp
7b466ad9f3
Drive the Petri projector from the server
The server holds one projector over its database and signals it after
each committed worker append, after each committed platform record
(through the run summary store's hook), at worker exit, and over every
Petri run at startup after the restart reconcile. A run executing in
the server process under the test override appends through the
projector's observing store, so it is signalled the same way. The
scenario tests read GET /runs/{id}/state after the view settles: the
hello prompt stage with its response, the command stage with its
output, and a two-branch parallel bundle whose branches are grouped
under the fork with the fork's results.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:32:23 -04:00
Bryan Helmkamp
e2bf05c0f0
Project a Petri run's records into Fabro's run view
The projection folds Petri's public events (replay_since over the run's
stored records) and Fabro's platform records into the RunProjection the
API serves, row by row as VIEWS.md maps them. The stage key is the
execution and firing; the StageId label is node@visit, made unique with
the execution when two child invocations would share one. A stage's
first_event_seq is the milliseconds from the run's creation to its
visit.started, so the view built live equals the view rebuilt from the
records whatever order two logs' records were committed in.

The projector is the view pass and its wake-up. Records first: an append
returns before any view work; a pass reads what is committed, folds the
items past the committed positions, and writes the projection document,
the ordered stream (one stream_seq per Petri event or platform record,
with the item's own identity beside it) and the narrowed runs row in one
later transaction. Signals coalesce per run, a lost signal costs only
latency, the startup pass folds every run the view trails, and a pass
that races a platform record leaves the view alone and runs again. A
torn tail holds the view where it stands and reports the run incomplete
with the replay's error; inspect_run decides completeness once the run
recorded its finish.

The tests build the view live for the hello bundle, a command workflow
and a two-branch parallel workflow and compare it with the rebuild; drop
every wake-up and catch up by a signal and by the startup pass; crash
between the record commit and the view transaction and apply only the
suffix; restart the projector over child executions; and hold at a torn
tail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:32:23 -04:00
Bryan Helmkamp
162791979c
Wake the projector after a Petri run's first event too
The run's creation commits its first event on the create path, not the
append path, so the platform record hook never fired for run.created.
The store now notifies after that commit as well, and a test proves a
Petri run's lifecycle events leave platform records beside them with one
wake-up per record while a legacy run leaves none.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:32:23 -04:00
Bryan Helmkamp
497281cdd7
Cover human gates in Petri runs through the questions API
The server scenario answers a gate in the in-process run through the
questions API and checks the branch it routed and the cleared pending
question. The CLI scenarios drive the real worker: a gate answered through
the API over the worker's control channel, two parallel gates each bound
to their own answer, and an unanswered gate that expires with its default
and records `interview.timeout`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:31:35 -04:00
Bryan Helmkamp
3aadc21739
Test the interview, secret, blob and home adapters through the engine
Integration tests in `fabro-petri` run workflows through `engine::run` on
the host sandbox: a gate answered under the posted question id, two
parallel gates each bound to their own answer, an expired question
completed as a timeout with the gate's default, an auto-approved run, a
cancelled run; a secret resolved from a vault into a command and masked in
every `petri_records` row; a command's large output round-tripped through
the `blobs` table under `blob://sha256/<hex>`; and the `hello` bundle on
the OpenAI twin with a model client over a vault that holds the key,
whose skills step searched the configured home.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:31:35 -04:00
Bryan Helmkamp
223e10ea20
Install Petri's interview, secret, blob and home adapters in a Fabro run
A Petri run in the worker, and in the server under its test override, now
gets Fabro's platform adapters instead of the standalone defaults:

- `fabro_petri::interview`: Petri's `Interviewer` over the questions API
  and the worker's control channel. A human gate's question is posted as
  the `interview.started` event a legacy stage emits, keyed by an id
  derived from Petri's identity (node, execution, firing, occurrence,
  ask), so the API, the web app and Slack list it; the answer posted to
  the questions endpoint reaches the control interviewer the adapter waits
  on and is mapped onto Petri's answer. An expiry the gate reports is
  completed as `interview.timeout`, a cancel as `interview.interrupted`,
  and an auto-approved run answers itself. The hook points the read side
  takes over are marked.
- `fabro_petri::secrets`: Petri's `SecretProvider` over the vault's token
  entries, so `{{ secrets.NAME }}` resolves at spawn and is masked in every
  record; a sensitive answer registers as a dynamic secret.
- `fabro_petri::blobs`: Petri's `OutputStore` over Fabro's `blobs` table,
  through the server's blob store or the worker's client.
- The Fabro home the server resolved travels to the worker as
  `--fabro-home`, so the skills step reads it whatever the worker's
  environment says.

`engine::RunRequest` takes the interviewer, its observers, the secret
provider and the blob table from the caller; `interviewer::Unattended` is
gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:31:35 -04:00
Bryan Helmkamp
abbc7ca11d
Add platform records and the Petri projection tables
A Petri run's own Fabro facts (its lifecycle before and after the
engine, a checkpoint commit, a pull request, a notification, a pairing)
are platform records in a table beside Petri's records, one typed enum
of kinds tagged on the wire, each keyed to a Petri stage where it
belongs to one and carrying the operation identity of the effect it
records. The run summary store derives the lifecycle kinds from the
legacy run events a Petri run still appends, in the event's
transaction, and calls a hook after the commit so the run's projector
can wake up.

Two more tables serve the projection that follows: the per-run
projection document with its committed positions, and the ordered
stream of everything the view consumed. The run summary store reads the
Petri projection back for the API and writes the narrowed runs row from
it without touching the legacy concurrency guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:46:21 -04:00
Bryan Helmkamp
a745d0b75d
Check a workflow version in memory with Runtime::check_source
`fabro_petri::check` materialized the version's bundle into a temporary
directory because `Runtime::check` read the workflow and its settings
files from disk. Petri now has `Runtime::check_source`, which takes the
workflow's repository-relative path, its text and a `FileSource`, so the
bundle goes into a `frontend::MapFiles` map instead: every file at its
bundle-relative path, `workflow.toml` beside the workflow, and
`.fabro/project.toml` at the root when the caller has one. Nothing is
written to disk, and the diagnostics name the bundle-relative paths
directly, with no root to strip.

The compile inputs are unchanged: the intent's inputs, the launch model
and provider, and `petri.repository` bound by Fabro itself (the launch's
path, or `null`). An entrypoint that is not one of the bundle's files is
now `CheckError::MissingEntrypoint`; `CheckError::Materialize` goes away.
`tempfile` becomes a dev-dependency, as only the tests use it.

New tests: a version whose `workflow.toml` names `engine = "petri"` is
admitted (the new Petri pin knows the key), an unknown `[workflow]` key
is refused with `unsupported.workflow_toml.key` and named in
`workflow.toml`, the project settings are read from the map, and a
missing entrypoint is an error.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:31:04 -04:00
Bryan Helmkamp
f4fe505bac
Move the Petri pin to 83345a8
Petri's `fabro-integration-p1` branch at 83345a8 adds
`Runtime::check_source`, the in-memory check entry point over a
`frontend::FileSource`, and makes `[workflow] engine` a known
`workflow.toml` key in its Fabro frontend, refusing any other unknown
`[workflow]` key with `unsupported.workflow_toml.key`.

Every `petri_*` entry moves from a0d2ceb to 83345a8. The lockfile changes
only the source line of the fifteen Petri packages; no other crate moves.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:31:04 -04:00
Bryan Helmkamp
16354186fa
Run a Petri run in the worker process over the HTTP store
When `fabro run __run-worker` finds its run's stored spec names Petri, the
new `petri_worker` module executes it through `fabro_petri::engine` over
`HttpRunStore`, leased for a launch id the worker mints and logs at start.
`--mode start` loads the admitted graphs through the client's blob read;
`--mode resume` continues the run from its records. The worker's existing
services carry over: the control channel's cancel and SIGTERM/SIGINT cancel
Petri's root invocation politely, a lost control channel cancels the run
and is reported once it settles, and pause, unpause and steer are received
and ignored with a warning until their adapters land. The model client
comes from the worker's catalog and vault snapshot for the providers whose
credentials resolve, and the lifecycle events (`run.starting`,
`run.running`, then `run.completed` or `run.failed`) go through the client
as the legacy worker's do.

Scenario tests against the real binary: a command-only Petri run executes
in the worker a foreground server launched, its records reach
`petri_records` over the HTTP store and its lease ends with the worker; and
a run whose server and worker are both killed mid-stage resumes in a new
worker after the server restarts, with one `run.completed`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:20:22 -04:00
Bryan Helmkamp
b5ebb472ae
Launch a worker for a Petri run and resume it after a server restart
`execute_run` no longer runs a Petri run in the server process by default:
it takes the subprocess path a legacy run takes, and `worker_exited` still
releases the worker's lease when the process ends. The in-process path
stays under the handler-registry test override, so the scenario tests need
no worker binary; it now honours the managed run's execution mode.

At startup, `reconcile_incomplete_runs_on_startup` hands a Petri run the
previous server left in flight (runnable, starting, running, blocked or
paused, with no cancel pending) back to a worker instead of failing it:
`PetriRuns::release_for_restart` ends the dead worker's lease from outside,
which fences it should it still be alive, the run is asked to start again
as a resume (`run.start_requested` with `resume`, then `run.runnable`, the
pair the API's resume appends), and the managed run is registered in
resume mode when Petri's store holds the run, else in start mode. Full
workspace recovery is the plan's F3.5 and is noted in the module docs.

Tests: the restart reconcile releases the lease, rewrites the history, and
launches the worker with `--mode resume`; a worker's HTTP store leases for
its launch id over the loopback server.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:20:22 -04:00
Bryan Helmkamp
1e2a205ec3
Carry Petri's sandbox plugin variables into workers and test servers
A Petri run executes in the worker process, which resolves the
sandbox-driver plugins itself. The `PETRI_SANDBOX_*` variables (plugin
paths, checksum overrides, dev mode, the Docker host address and the action
host image) now have `EnvVars` names, cross the worker's environment
allowlist with `PATH`, and pass through the test harness's isolation so a
developer's plugin override reaches the servers tests start and the workers
those servers launch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:20:22 -04:00
Bryan Helmkamp
a621fb72e1
Let the Petri engine start or resume a run over any store
`fabro_petri::engine` is now the one assembly the worker process and the
server share: `RunRequest` takes the run's store as `Arc<dyn RunStore>` and
an `Execution`, either `Start` with the admitted graphs or `Resume` from the
run's records through `host::resume_configured`, with the same interview
observer a start installs. A resume whose record has no root invocation is
refused with a named error instead of a panic in the host. The outcome is
mapped to a `Conclusion` (succeeded, or failed with Fabro's reason and a
message) so both callers record the same terminal event.

`admission::load_with` loads the admitted graphs through any blob read, so
a worker loads them through its client; `admission::load` over the server's
`BlobStore` delegates to it.

`HttpRunStore::for_worker` takes every lease for the worker's launch id,
whatever owner Petri minted for the run runtime, and the open logs the
owner. The module docs state the rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:20:22 -04:00
Bryan Helmkamp
832f39f704
Merge branch 'petri-integration-http' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/README.md
#	lib/components/fabro-petri/src/lib.rs
2026-09-17 20:37:03 -04:00
Bryan Helmkamp
b78c020298
Accept the CLI snapshots the engine flag and the listed diagnostics changed
The create error now names each validation diagnostic as `rule: message`
after "Validation failed", and `fabro server start --help` lists
`--engine`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:34:22 -04:00
Bryan Helmkamp
d2f0e70dca
Run a workflow through Petri in the server, behind the engine flag
When a run's engine is Petri, the create handler hands the bundle, inputs
and launch to Petri's check instead of the legacy compile, lint and model
pinning, refuses the run with the validation error the legacy validator
uses (Petri's codes as the rules, listed in the API detail), and records
the admission on the run spec. The Fabro graph the read side displays is
parsed without validation. The scheduler executes a Petri run in the server
process through fabro_petri::engine, appending only the run lifecycle
events the read side needs (run.starting, run.running, run.completed or
run.failed); no stage or agent event is projected yet.

Scenario tests run the hello bundle on the OpenAI twin under the version
flag and a command-only bundle under the server setting, check Petri's
record agrees, and cover the refusals for an unknown attribute, an
undeclared node and an unknown model.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:27:44 -04:00
Bryan Helmkamp
03309d4240
Let Petri compile and execute a Fabro run through fabro-petri
`check` materializes a workflow version's bundle into a temporary directory
(`Runtime::check` reads files from disk), lowers it with the run's inputs
and launch, and returns the admitted graphs or Petri's diagnostics in a
shape the server maps onto Fabro's. `admission` keeps the admitted graphs
in the blob store, named on the run spec and verified by digest on load.
`runtime` assembles the same Petri runtime at create and at execution: the
Fabro frontend with the server's settings layer, the Attractor step kinds,
the model client as the PebbleClient capability so admission pins every
model. `engine` runs the admitted graph in the server process over
SqliteRunStore under the Fabro run id, with the standalone defaults, an
interviewer that fails any question, and cancel on a token, and derives the
outcome from inspect_run over the run's record.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:27:44 -04:00
Bryan Helmkamp
6557a2f404
Check the HTTP run store against a loopback server
Petri's store conformance suite runs over `HttpRunStore` talking to an
axum listener on a loopback port. The suite opens runs under keys of
its own, while a key over the API is a Fabro run id the worker's token
names, so an adapter gives each suite key a fresh run with a token
minted for that run alone: the least a worker holds.

Three more tests cover what the suite cannot: the operator release
through the server's store turns the worker's handle stale; a
middleware swallows the reply of one committed append and the store's
resend leaves each record once; and two workers with owners of their
own never hold one run's lease at the same time.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:12:20 -04:00