Petri 639ce3e added `ControlService::interrupt_firing` and the `LiveTurns`
capability a host installs beside the pause hooks. Fabro now drives it:
`RunControls::interrupt(stage, text)` resolves its stage the way a steer
does (a label, a node name, or the run's one live agent stage) and stops
that stage's current model turn, keeping the session; the text, when
given, is the stage's next input. `engine::run` installs the live-turn set
as a runtime capability, so without it no interrupt could ever land.
The worker maps `run.interrupt` and `run.interrupt_then_steer`, both of
which now carry an optional `stage`, to that call. The control bus is
one-way, so a refusal is recorded the way a refused steer is: a
`run.notice` on the run's stream whose code says why (`no_live_turn` when
Petri refuses a stage with no turn in flight, `no_such_stage`,
`interrupt_refused`).
The server's `POST /runs/{id}/interrupt` and `POST /runs/{id}/steer` with
`interrupt=true` forward the control and answer 202, replacing the 501
`interrupt_unsupported` stub. The interrupt endpoint takes an optional
body (`stage`, `text`), refuses a finished run with 409
`run_not_interruptible`, and forwards an interrupt of a blocked run, since
an agent stage may be running a turn beside the question and the worker
judges each stage itself. `fabro events --pretty` prints the delivered
interrupt and the stage's `attractor.turn.interrupted` report.
Verified on the twin: `fabro steer --interrupt` during a long tool call
ends the turn, the text is the agent's next request, and the stream
carries the `$interrupt` record and the interrupted-turn report; an
interrupt of a gate stage is refused with `no_live_turn` and the gate's
question is untouched.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`SteerRunRequest` takes an optional `stage`: the label the projection
shows (`node@visit`, or `node/e<execution>@visit` when two executions
share one) or the node's name. The server passes it on the worker control
message; the worker's `RunControls` resolves a label to the live agent
firing and steers that firing, and a node name through Petri's own
live-stage index. Unnamed, the one-live-agent rule stays, and the refusal
now names the live stages by their labels. `fabro steer --stage` sets it.
A controls scenario runs two agent stages side by side, sees the unnamed
steer refused with both named, and steers each apart, one over the API
and one through the flag.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`RunSandboxInstance` carries `ready_duration_ms` from the root scope's
`scope.acquired` and `retained` from its `scope.released`, so the view
says how long the sandbox took and whether it still exists after the run.
The OpenAPI schema, the TypeScript client and the web sandbox tab's
overview show both; the host sandbox scenario asserts them on a real run.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri's retention decides whether a released workspace is kept or removed.
Fabro's lifecycle settings decide whether a sandbox keeps running after
the run and whether a delete may remove it; none asks for removal at the
run's end, and the sandbox tab, `fabro cp`, the run's delete and the
sandbox scenarios read the container after the run. So the mapping is
`Retention::Always` for every setting, named once as `engine::RETENTION`
with the reasoning, instead of a per-setting function that released a
finished sandbox.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
No store sits behind `[server.slatedb]` any more: the section leaves the
settings layer, the resolved server settings, the defaults, the API schema,
the TypeScript client, the install wizard and the docs, and `fabro install`
probes the bucket for the `artifacts/` prefix alone. A settings file that
still carries the section is rewritten once at startup by a temporary
migration that removes it with a backup beside the file.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The run and stage artifact listings, the download and the archive join the
artifacts the projection records, whose bytes are in the blob table, with
the ones uploaded to the artifact store, each once.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The commit that creates a workspace's run branch records `run.branch`
(the base commit, or the first checkpoint in a workspace with no history)
and `git.identity`. Every checkpoint record after the first carries the
stage's diff from its parent commit, with the patch as a text blob. After
the checkpoint record, the transition hook lists the stage's workspace
through the scope's environment, on the host and in a sandbox alike, and
collects every file under `[run.artifacts] include` into the blob table as
an `artifact.collected` record, skipping a file already collected under
the same path and digest. At the run's end the hooks diff the branch's
last checkpoint against its base in the snapshot repository and record
`run.diff`.
`engine::retention` maps the environment's lifecycle settings onto
Petri's workspace retention instead of always keeping every workspace:
`preserve`, `stop_on_terminal = false` and the local provider keep them,
anything else keeps a failed scope's only. The hooks docs no longer name a
redundant link target, so rustdoc passes with warnings denied.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri a5906f6 records where each scope's sandbox ran (`scope.acquired`,
`scope.failed`) and how its lease was released (`scope.released`). The
projection folds the root invocation's records into `Run.sandbox`:
`initializing` from `run.started`, `ready` with the `RunSandboxInstance`
(the provider, Petri's `host` as Fabro's `local`, the provider's id, the
image and snapshot, the working directory) from `scope.acquired`, `failed`
from `scope.failed`; the retention outcome is kept in the fold state, since
the view has no field for it. Ask Fabro reconnect and `sandbox cp`,
`preview` and `ssh` reach the run's sandbox again.
A local reconnect designates the recorded working directory again when the
host provider does not know the id: the provider mints a registry-only id
for a workspace path too long for a path-derived one, and that registry
belongs to the run's worker. The stream listing redacts its items the way
the attached stream does, so a client that pages after a stream sees the
same items.
Pins move to Petri a5906f6 (run format 6, engine log v11, event contract
4). The attach stream snapshot is re-recorded with the new record and a
filter for the host provider's minted ids; `sandbox cp` reads an upload
back through the run's workspace, which is no longer the target folder.
Server scenario tests prove the projected instance on the host and Docker
providers.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Step 4 of the legacy executor deletion, fourth commit: with no writer
and no reader left, the legacy event log goes.
- `fabro-types`: `run_event` (`EventBody`, `RunEvent` and every props
struct), `EventEnvelope` and the `RunEventDetail*` types are deleted.
What the projection and the API still use moves out of the event
vocabulary: `AgentEventProps`, `AgentSessionActivatedProps`,
`AgentToolsAvailableProps`, `StagePromptProps`, `SessionCapability`
and the coding event names to `agent_props`; `RunNoticeLevel` and
`RunNoticeCode` to `notice`; `InterviewOption` beside the question
types; `RunRunnableSource` beside the run status. `Checkpoint` is
what Fabro records for a Petri run: `timestamp`, `current_node`,
`git_commit_sha`; the conclusion's stage summaries derive from the
projection's stages instead of the checkpoint's node maps.
- `fabro-store`: the Slate bridge (`RunDatabase`, the Slate `Database`,
`keys`, `record`, `EventPayload`) and the reducer (`run_state`) are
deleted. `Database` is the blob table and the run summary store over
one pool; the blob store is SQLite only; the run summary store keeps
the `runs` row a projector writes and lists, and finds the pull
request creation candidates over `platform_records`; `build_summary`
and `projected_usage` live in `run_summary`. The SlateDB dependency
is gone. Test fixtures build the store from its two SQLite stores.
- `fabro-workflow`: the `event` module (the `Event` enum, its
conversion, sink, emitter, redaction, stored fields and names),
`runtime_store`, `StageScope` and the legacy seeding test helpers are
deleted; the tests that seeded legacy runs read platform records or
a projection instead.
- `fabro-sandbox` owns `GitRetryReason`.
- The server builds the store without an object store; the legacy
`POST /runs/{id}/events` tests go, an interrupt answers
`interrupt_unsupported` in the tests as it does in the handler, and
the tests that read a run back through the Slate handle read its
projection or its platform records. The projection folds a block
that lands while the run is paused as the pause's prior block, and a
pause or unpause clears the pending control it answers; a control
request's check-and-append holds a per-run lock so two concurrent
cancels record one request.
- The CLI's final output is the response of the last stage that
produced one; the workflow tests read completed nodes from the
succeeded stages.
- The spec's `RunCheckpoint` carries the three fields the type keeps.
Still failing until the next commits: the CLI tests that seed runs
through `POST /runs/{id}/events` or wait for legacy event names, and
the two Ask Fabro resume tests (the sandbox instance gap).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Step 4 of the legacy executor deletion, third commit: the legacy event
API and every reader of it go, so that the next commits can delete the
event log, its reducer and the types beneath them.
The API:
- `GET /runs/{id}/events` pages the run stream only
(`PaginatedRunStreamList` by `after`); the legacy `since_seq`,
`before_seq` and `order` cursors, the `oneOf` envelope, the legacy
`EventEnvelope`, `PaginatedEventList`, `RunEvent`, `EventSeq`,
`AppendEventResponse` and `RunEventDetailResponse` schemas,
`POST /runs/{id}/events`, `GET /runs/{id}/events/{seq}` and
`GET /runs/{id}/stages/{stageId}/events` are deleted. `GET
/runs/{id}/attach` and `GET /attach` frame `RunStreamItem`s only.
- The Rust and TypeScript clients regenerate; the removed models leave
the TypeScript package.
The readers:
- `fabro-client` drops the legacy run event listing, tail and attach
methods and `RunEventStream`; `list_run_stream_until` bounds a stream
read.
- `fabro-tool`'s `fabro_run_events` lists, searches and details the run
stream: `after` is the exclusive `stream_seq` cursor, `event_id` the
item's id, filters match the item's name and `recorded_at`.
- `fabro-dump` writes the stream to `events.jsonl`; `fabro dump` reads
it.
- The CLI's progress renderer keeps only what the run stream drives:
the legacy event conversion, the sandbox and setup displays and their
styles go. `fabro system events` prints stream items.
- The server's demo mode folds its agent fixture straight into the
session projection and answers the attach stub with a stream item;
the demo stage events endpoint is gone.
- The web app: every run is a Petri run. The legacy event hooks,
renderer props, stage popover summary, run phases derivation and
live-event payload handling are deleted or ported to `RunStreamItem`;
toasts and board refreshes read the stream's platform records.
- Tests: the legacy API round trips and pagination tests are deleted;
the CLI's MCP, attach and system event mocks serve stream pages; the
CLI test helpers read stream items.
Still failing until the later commits: the CLI tests seeded through
`POST /runs/{id}/events`, the server tests over the legacy store, and
the legacy type tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Step 4 of the legacy executor deletion, second commit. Ask Fabro's
sessions were the last writer of `run_events`: a session's creation, its
turns and their messages, tool calls and endings went into the run's
legacy event log, keyed by the run's sequence. They now have a log of
their own.
- `run_session_events` (migration `2026091802`): one row per session
event, numbered per session from 1, with the owning run, the turn, the
event name and its properties. `RunSessionEventStore` appends under the
write lock, lists a session from a sequence, names a session's owner
from its creation event, deletes a run's sessions with the run, and
publishes each committed event to its subscribers.
- `fabro_types::SessionEvent`: `seq`, `session_id`, `run_id`, `ts` and a
flattened body (`event` naming the kind, `properties` its fields), with
the same event names and property shapes the legacy events carried,
so the web app and the CLI read the same JSON. The property structs
move to `session_event`; `run_event::session` re-exports them under
their old names until the legacy event log goes.
- The API: `GET /sessions/{id}/events` pages `PaginatedSessionEventList`
by the session's own sequence, `GET /sessions/{id}/attach` replays and
streams `SessionEvent` frames (subscribed before the replay, so no
event falls between the two), the turn stream carries the same frames,
and an interrupt answers with the recorded event. The session
projection folds `SessionEvent`s; the legacy `find_session_owner` over
`run_events` is gone.
- The CLI's `run ask` and the web app's session stream read
`SessionEvent`; the web runtime no longer accepts the nested legacy
envelope shape.
The two session resume tests in the server keep failing for a reason
this commit does not touch: Ask Fabro reconnects to the run's sandbox
from the projection's sandbox instance, which the Petri projection does
not carry yet (`VIEWS.md`, the `scope.acquired` gap).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Step 4 of the legacy executor deletion, first commit of several: step 4
spans commits because the legacy event log and its consumers cannot go
in one compiling change. This commit moves every writer off `run_events`;
the reducer, `EventBody`, the Slate bridge and the API's event types
still exist for the readers the next commits port or delete.
Writers:
- The server records a run's lifecycle (submitted, runnable, starting,
running, blocked, paused, control requests and effects, the terminal
status), its title, parent link, archive state, notices and pull
request state as platform records (`fabro_store::platform_records`),
through the new `server::run_records` module. Every append wakes the
projector and waits for its pass, so the read that follows a write
holds the record.
- Pull request creation is recorded as `pull_request.requested`,
`pull_request.created`, `pull_request.failed`, `pull_request.linked`
and `pull_request.unlinked`; the projection folds them into the run's
pull request and creation state.
- Answers to questions are recorded as `interview.answered` with the
answering principal and the answer text; the interview adapter no
longer posts legacy `interview.*` events (`QuestionSink` is now an
optional observer).
- The worker (`fabro run __run-worker`) records its lifecycle, notices
and pause state over `HttpPlatformRecords`; `HttpRunStore` for the
legacy event log and the worker's `run_store` are gone.
- `persist_created_run` appends `run.created` and `run.submitted`.
Readers:
- A stream follower (`server::stream_follower`) follows each live run's
stream (Petri events and platform records), folds lifecycle records
into the in-memory run state, forwards items to the global attach
broadcast, and syncs blocked and paused from the projection.
- Slack posts questions from the projection's pending interviews,
finishes them on `interview.answered` or `question_expired`, and sends
lifecycle notifications with `notification.sent` dedupe.
- `GET /runs/{id}/events` and the attach endpoints serve only the run
stream; the per-event, per-stage and `POST /runs/{id}/events`
endpoints and their tests are deleted.
- `Database::load_run_projection` reads the Petri projection only.
Deleted with the writers:
- The SQLite blob and run-history activation migrations and their
legacy Slate imports (`legacy_blob_import`, `legacy_run_history_import`,
the activation backup): a greenfield server has no Slate history to
import, and the run-history verification refused to start a server
whose runs have no legacy events.
- `fabro-workflow`'s `operations::archive` and `operations::run_store`.
- The server's legacy-event unit tests and the CLI's `HttpRunStore` tests.
The in-process answer transport is now set after the starting and
running records land, not gated on the live status still being
`Starting` (the records already moved it).
The manifest validation test for a `run.agent.mcps.<name>` catalog
reference now expects `unsupported.workflow_toml.run.agent.mcps.reference`:
Petri's Fabro frontend has no server catalog to resolve it against.
Legacy readers still fail their tests until the next commits: the
reducer and Slate tests in fabro-store, the fabro-workflow create tests
that read the run back through the legacy store, the CLI tests seeded
through `POST /runs/{id}/events`, the CLI's legacy attach and render
paths, the sessions API, the OpenAPI conformance test, and the web
fixtures.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri judges a workflow at admission, so Fabro's lint rules go:
`fabro-validate` (its 35 rules and the `LintRule` trait) is deleted, and
with it `fabro-acp` (only a rule and two legacy executor tests used it),
the model-resolution transform, the legacy `create`, `compile_create_run`
and `materialize_create_run` stages, and `fabro-graphviz`'s `condition`
and `fidelity` modules. `Diagnostic`, `RelatedDiagnostic` and `Severity`
move to `fabro_types::diagnostic`, the one shape every diagnostic takes.
Validation is now the same question the create handler asks. A new
server module, `petri_check`, builds Petri's check request from a
workflow bundle and the run's settings (every workflow of the bundle at
its bundle-relative path, the inputs, the run variables, the launch),
runs the check, and maps the diagnostics; Fabro adds one rule of its
own, `fabro.model.no_ready_provider`, refusing a model node when no
provider is ready. Admission, the validate and preflight endpoints and
the offline `fabro validate` all go through it:
- `validate_prepared_manifest` runs Fabro's structural pass (parse and
transform, whose diagnostics stay) and then Petri's check, on the
blocking pool from the handlers;
- the offline `validate_manifest` checks with no model client and with
an unbound input as a warning (`CheckRequest.unbound_is_warning`), so
a workflow validates before its inputs exist; a collected workflow
before upload checks with unbound inputs as errors, as before;
- preflight resolves each LLM node's selector against the ready
providers and the catalog for its probe, as the deleted transform did,
and no longer probes a model Petri refused;
- the graph render endpoint needs only the structural pass;
- a run manifest now carries its `[run.goal] file`, which Petri reads
from the bundle as it does for a version.
The transforms keep the authored model selector (`sonnet` stays
`sonnet`): Petri pins the catalog model in the admitted graph, not in
the graph Fabro displays or in the settings snapshot. Tests assert that,
and the CLI's validate, preflight and graph snapshots carry Petri's
diagnostics (`attractor.no_start`, `attractor.undeclared_node`,
`attractor.bad_on_failure`, ...) in place of the lint rules' text.
Known gaps, Petri's side: a `workflow.toml` whose `[run.environment]`
names an environment the server catalog defines but the file does not
is refused (`unsupported.workflow_toml.run.environment`), as admission
already refused it; an unbound input inside an included template
partial is a render error rather than the unbound-input warning.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The legacy executor replayed a run from a checkpoint; Petri resumes a
run from its records instead, and the checkpoint timeline, rewind, fork
and retry were the operations that replay carried. The server dropped
their handlers with the executor; this removes the rest:
- the API spec's `/runs/{id}/retry`, `/rewind`, `/fork` and `/timeline`
paths with the `ForkRequest`, `ForkResponse`, `RewindRequest`,
`RewindResponse` and `TimelineEntryResponse` schemas, and the
generated TypeScript models;
- `fabro rewind` and `fabro fork` (with the checkpoint timeline printer
and the repo-origin check only they used), their reference pages and
the checkpoints guide's rewind and fork sections;
- `fabro-client`'s `rewind_run`, `fork_run` and `run_timeline`;
- the web app's Retry action on the run list and the run page.
`Run.retried_from` stays on the run type: a run that was retried before
the cutover would still name its source.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`fabro-hooks` ran the legacy executor's hooks; Petri's Attractor steps
run Fabro's hooks now, so nothing in the workspace uses the crate. The
engine freeze (the CI workflow, the two scripts, and the AGENTS.md and
fabro-petri README sections) guarded the engine half of `fabro-workflow`,
which the previous commit deleted.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every run executes on Petri, so the in-process legacy executor goes:
`fabro-core` and, in `fabro-workflow`, the handlers, lifecycle, pipeline
execution, routing, retry, conditions, node handlers, steering, agent
memory, artifacts, checkpoints, command log, and the `start`, `resume`,
`retry`, `fork`, `rewind` and `timeline` operations. The two are deleted
together because the engine half of `fabro-workflow` was the only user of
`fabro-core` and `fabro-core` the only runtime of that half; neither
compiles without the other.
Kept in `fabro-workflow`, narrowed: the parse/transform/validate/persist
pipeline and `create`, `archive`, `validate` (workflow definitions still
come from DOT and settings); the run tools (`run_tools`, moved from
`handler/llm/fabro_tools.rs`) for Ask Fabro, `fabro exec` and Petri's
host tools; the pull request pipeline (`pull_request`, moved from
`pipeline/`, for the step 0 port); Run Files' diff helpers in
`sandbox_git`; `git_identity`, `usage_rollup`, `run_status`,
`run_materialization`, `web_search` and `workflow_bundle`.
Server: `RegistryFactoryOverride` becomes `execute_in_process`;
`RunAnswerTransport::InProcess` carries only the interviewer; the
interrupt endpoint answers 501 `interrupt_unsupported` and every pair
endpoint 501 `pair_unsupported` (status lists none); rewind, fork, retry
and timeline handlers and routes are removed; the command log is served
from the stage output blob; usage rollups accumulate from the settled
projection after an in-process run as after a worker exit.
Ported while here:
- `materialize_admitted_run` materializes the goal and drops a disabled
pull request block, as the legacy materializer did.
- A run whose admitted graph has an agent or prompt node is refused at
create when no LLM provider is ready (`fabro.model.no_ready_provider`);
a workflow of commands and gates needs no model and is admitted.
- The projection's question type falls back on the options, as the
interview adapter does, so a gate with edge-label options answers as
multiple choice.
Tests: the server scenarios (lifecycle, run completion, SSE, helpers)
run in process on Petri and assert Petri's stage labels and stream
names; the reconcile tests assert Petri's relaunch semantics; legacy
unit tests of the deleted executor are removed; three server unit tests
the removal took with it are restored; the pair fixtures go with the
pair feature. Petri test fixtures no longer name `[workflow] engine`.
Still red after this commit, all legacy consumers the next steps
delete or port: fabro-store's Slate/reducer fixtures and fabro-types
legacy JSON tests (step 4); server unit tests over legacy run events
(retry endpoints, list_run_events, artifacts, per-event pause/unpause,
run history activation, legacy sandbox fixtures) (steps 3-4); CLI tests
that parse legacy event envelopes, the legacy `events`/`attach`/`diff`/
`dump`/`inspect` snapshots, `run rewind`/`run fork`, the ACP and
git-identity workflow tests, and the runner tests that drive the legacy
worker by hand (steps 3-4); the web app's Petri fixtures still carry
`engine` (regenerate with `FABRO_CAPTURE_PETRI_FIXTURES` in step 4).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Delete `Engine`, `RunEngine`, `[workflow] engine`, `[server.execution]
engine`, `FABRO_SERVER_ENGINE` and `fabro server start --engine`. The run
spec records what Petri admitted as `admission: PetriAdmission`; the
create handler always admits through `Runtime::check`; `execute_run`
always launches the Petri worker (or executes in process under the test
override); the CLI runner takes only the Petri worker path, and its
legacy control arm, artifact uploader, signal pause handlers and
credential helpers go with it. The CLI's `attach` and `events` read the
run stream only.
Two gaps this surfaced are closed here: the check adapter binds the
server's run variables as Petri compile variables (`{{ vars.* }}` in a
prompt no longer fails admission), and deleting a run removes its Petri
records, lease, platform records, projection and stream.
Tests: the config engine tests are replaced (an engine key is unknown),
the API round-trip test covers `PetriAdmission`, the server and CLI
Petri scenarios drop their engine settings, and the API tests that read
legacy event names now read the run stream or the session events. The
remaining red tests are fixtures and scenarios of the legacy executor
and the legacy event store (`fabro-store` `slate` and `run_state`,
`fabro-types` legacy `run.created` JSON, the server's handler-registry
scenarios, the CLI dry-run snapshots), which the next steps of the F4.3
series delete or port.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A Petri run's worker launches the sandbox-driver plugins itself, and the
Docker plugin forwards `DOCKER_HOST`, `DOCKER_TLS_VERIFY`,
`DOCKER_CERT_PATH`, `DOCKER_API_VERSION`, `DOCKER_CONFIG` and
`DOCKER_CONTEXT` from the process that launches it. They now cross the
worker's environment allowlist, so the worker's sandboxes go to the daemon
the server uses. The concern that kept them out, the legacy worker's own
Docker client picking up a daemon it was not meant to, is moot: the
legacy executor is being deleted. The same variables pass through the
test harness's isolation, so a developer's daemon selection reaches the
servers tests start.
Daytona's non-secret selectors, `DAYTONA_API_URL` and
`DAYTONA_ORGANIZATION_ID`, cross the allowlist too. The API key stays the
vault's: a Daytona run's worker command carries it the way the GitHub app
key travels, and the Daytona plugin reads it from the worker's process.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A Petri run answered only cancel and answers; pause, unpause and steer
were ignored with a warning. `fabro_petri::controls::RunControls` now
wraps Petri's `ControlService` per run: `engine::run` installs its pause
gate over the run's hooks, observes the run through it and wires it to
the coordinator, on a start and a resume alike, so a run paused when its
worker died resumes paused.
The worker's control channel takes a `WorkerControls` enum: the legacy
hub and pause flag, or the Petri run's controls. On Petri, `run.pause`
holds admission, `run.unpause` releases it once the record is durable,
and `run.steer` goes to the one live agent stage (Fabro's steer names no
stage); with none or several it is refused with a `run.notice` record.
The paused state is mirrored to Fabro's lifecycle as `run.paused` and
`run.unpaused` events, so the server's live status and the projection
follow Petri's own records.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Plan item F3.4. `fabro_petri::host_tools` adapts Petri's `HostTools`
capability to `register_fabro_run_tools`: every native agent session of a
run gets the tools the legacy worker registers, bound to the worker's
client and the run id, so a child run a stage creates is parented to the
Petri run. The tools run under the run's tool hooks, are recorded under
the stage, and reach sub-agents through Pebble's inheritance.
`RuntimeSpec::run_tools` installs the capability; the worker sets it when
the run's settings enable `[run.agent] fabro_tools` and the worker token
carries `agent:run_tools`, the legacy worker's gate. The server's
in-process test path runs without them, like the legacy one.
The identity the tools need is the run id alone; no run tool records a
stage on an effect, so nothing derives Fabro's `node@visit` label. A
context for another run gets no tools.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`GET /runs/{id}/events` and `GET /runs/{id}/attach` serve a Petri run's
public events and Fabro's platform records as one ordered stream in a
Fabro envelope (`RunStreamItem`: `run_id`, `stream_seq`, `kind`, `id`,
`recorded_at`, `item`), read from the projector's `petri_stream` table.
The cursor is `stream_seq` (`?after=`); the item's own identity (the
Petri `EventId` as `<log>/<seq>/<index>`, or the platform record's seq)
travels beside it for deduplication. A legacy run keeps its envelope on
the same endpoints; the OpenAPI response is the union of the two lists,
and the stream list reports Petri's `EVENT_CONTRACT_VERSION`.
The attached stream follows the projector's commit signal (a wake-up,
with a poll as the fallback) and ends after the platform record of the
run's terminal lifecycle transition, the analog of the legacy stream's
`run.completed`, or a bounded grace after the projection went terminal.
`RunSpec.engine` (`RunEngine`, `PetriAdmission`, `PetriGraphRef`) is
named in the spec and reuses the Rust types. `fabro-client` matches the
union and adds `list_run_stream`, `list_run_stream_page` and
`attach_run_stream`.
A server test attaches to a two-branch parallel run, disconnects once
both branches started, records a platform notice while both branch
scripts run, reconnects from the last `stream_seq`, and checks the
union is the whole stream: every item once, in order, no gap, no
duplicate, the notice between the branch events, and the same as the
paged listing. The Petri scenarios capture their settled projection and
stream as JSON fixtures for the web app under
`FABRO_CAPTURE_PETRI_FIXTURES`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
In-process tests over the memory store: every finish is committed on the
run branch with its identity trailers and recorded with its commit, the
run-end hooks reach Petri's local service through Fabro's wrapper, a
stage that fails on its own terms is committed and its failure route runs
on the committed files, a failed checkpoint records `checkpoint_failed`
with no route taken and a restart reports the run failed, and a
`[[run.hooks]]` hook blocks an agent's tool call through the forwarded
service, with the model told why.
Real-binary scenarios crash the server and its worker with SIGKILL: after
a durable finish the stage's commit is not repeated and the interrupted
stage reruns on its snapshot; a crash held before the commit reruns the
stage once; a crash held after the commit but before its record
reconciles the record from the snapshot repository; a deleted workspace
is restored; a failure route sees the same committed files after a
crash; a failed checkpoint fails the run and a restart leaves it failed.
Recovery selects the executions `inspect_run` reports incomplete, and a
run whose coordinator log is still empty is left to the worker's resume.
The worker's platform record endpoints get an API test and the generated
TypeScript client.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The interview adapter derived its own question id from Petri's identity
and posted it on `interview.started`, while the projection over Petri's
records serves the pending question under Petri's `Question.id` with the
firing's stage label. The answer endpoint validates against the
projection, so an answer under the projection's id never reached the
adapter's wait.
The adapter now waits under Petri's id and labels the question's stage
through the projection's own rule: `stage_label`, `is_shown` and
`visit_of` move out of `start_visit` into shared functions, and the
adapter's observer derives each firing's `visit.started` through Petri's
`Projection`, as the projector does, so the label matches by
construction. The full Petri identity stays on `AskedQuestion`.
The legacy `interview.*` events are still posted, under Petri's id, for
the readers that follow the event stream rather than the projection: the
Slack service, `run attach`, the web app's Q&A renderer and the server's
answer claim. The store already derives the `interview.answered` platform
record from `interview.completed` for a Petri run, so who answered is
recorded under Petri's id with the answering principal.
The gate scenarios assert the new identity and encode the id as one path
segment, as the generated clients do. Projection tests cover an expired
question and an auto-approved answer.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fabro's hooks on a Petri run wrap the hooks the runtime installed for
`[[run.hooks]]` and forward every point. In `prepare_result`, before the
finish is recorded, they commit the stage's files on the run branch of
its host workspace with Fabro's author identity and the run, execution,
firing and attempt as trailers, and publish the commit to a snapshot
repository beside the run's workspaces under a ref per checkpoint. A
stage that failed on its own terms is committed like a successful one; a
commit that fails is fatal: the outcome becomes a `checkpoint_failed`
failure, the run is cancelled through the coordinator handle, and the
transition refuses the firing's routes. In `transition` they write the
platform checkpoint record, keyed on the Petri position and the
checkpoint's operation identity, and a failed write is a recorded
problem.
On restart the server runs the recovery protocol before it relaunches a
worker: a run with a failed checkpoint is reported failed; otherwise
every live execution's last durable finish names the snapshot its
workspace is verified against, reset to, or restored from, with a lost
record reconciled from the snapshot repository, and a finish with no
snapshot fails the run rather than resume it on stale files.
The worker reaches the platform records over two new worker-scoped
endpoints; the server reaches the table directly. A test gate directory
lets the CLI scenarios hold a checkpoint at a named point.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The projector takes two pools: the one Petri's records live in and the
one the view tables live in. In the server both are the one database;
a test fixture keeps the runs row, the platform records and the
projection tables in the run summary store's own pool, which the
projector was not reading, so a run projected in a test server folded
its Petri events before its run.created record. The startup run-history
verification checks only a Petri run's identity and legacy guard, since
its row is the projector's. An agent stage's response is the
response.<node> its outcome wrote into the run context, as the prompt
step writes it. The scenario tests assert each branch's own index.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The server holds one projector over its database and signals it after
each committed worker append, after each committed platform record
(through the run summary store's hook), at worker exit, and over every
Petri run at startup after the restart reconcile. A run executing in
the server process under the test override appends through the
projector's observing store, so it is signalled the same way. The
scenario tests read GET /runs/{id}/state after the view settles: the
hello prompt stage with its response, the command stage with its
output, and a two-branch parallel bundle whose branches are grouped
under the fork with the fork's results.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The server scenario answers a gate in the in-process run through the
questions API and checks the branch it routed and the cleared pending
question. The CLI scenarios drive the real worker: a gate answered through
the API over the worker's control channel, two parallel gates each bound
to their own answer, and an unanswered gate that expires with its default
and records `interview.timeout`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A Petri run in the worker, and in the server under its test override, now
gets Fabro's platform adapters instead of the standalone defaults:
- `fabro_petri::interview`: Petri's `Interviewer` over the questions API
and the worker's control channel. A human gate's question is posted as
the `interview.started` event a legacy stage emits, keyed by an id
derived from Petri's identity (node, execution, firing, occurrence,
ask), so the API, the web app and Slack list it; the answer posted to
the questions endpoint reaches the control interviewer the adapter waits
on and is mapped onto Petri's answer. An expiry the gate reports is
completed as `interview.timeout`, a cancel as `interview.interrupted`,
and an auto-approved run answers itself. The hook points the read side
takes over are marked.
- `fabro_petri::secrets`: Petri's `SecretProvider` over the vault's token
entries, so `{{ secrets.NAME }}` resolves at spawn and is masked in every
record; a sensitive answer registers as a dynamic secret.
- `fabro_petri::blobs`: Petri's `OutputStore` over Fabro's `blobs` table,
through the server's blob store or the worker's client.
- The Fabro home the server resolved travels to the worker as
`--fabro-home`, so the skills step reads it whatever the worker's
environment says.
`engine::RunRequest` takes the interviewer, its observers, the secret
provider and the blob table from the caller; `interviewer::Unattended` is
gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`execute_run` no longer runs a Petri run in the server process by default:
it takes the subprocess path a legacy run takes, and `worker_exited` still
releases the worker's lease when the process ends. The in-process path
stays under the handler-registry test override, so the scenario tests need
no worker binary; it now honours the managed run's execution mode.
At startup, `reconcile_incomplete_runs_on_startup` hands a Petri run the
previous server left in flight (runnable, starting, running, blocked or
paused, with no cancel pending) back to a worker instead of failing it:
`PetriRuns::release_for_restart` ends the dead worker's lease from outside,
which fences it should it still be alive, the run is asked to start again
as a resume (`run.start_requested` with `resume`, then `run.runnable`, the
pair the API's resume appends), and the managed run is registered in
resume mode when Petri's store holds the run, else in start mode. Full
workspace recovery is the plan's F3.5 and is noted in the module docs.
Tests: the restart reconcile releases the lease, rewrites the history, and
launches the worker with `--mode resume`; a worker's HTTP store leases for
its launch id over the loopback server.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A Petri run executes in the worker process, which resolves the
sandbox-driver plugins itself. The `PETRI_SANDBOX_*` variables (plugin
paths, checksum overrides, dev mode, the Docker host address and the action
host image) now have `EnvVars` names, cross the worker's environment
allowlist with `PATH`, and pass through the test harness's isolation so a
developer's plugin override reaches the servers tests start and the workers
those servers launch.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
When a run's engine is Petri, the create handler hands the bundle, inputs
and launch to Petri's check instead of the legacy compile, lint and model
pinning, refuses the run with the validation error the legacy validator
uses (Petri's codes as the rules, listed in the API detail), and records
the admission on the run spec. The Fabro graph the read side displays is
parsed without validation. The scheduler executes a Petri run in the server
process through fabro_petri::engine, appending only the run lifecycle
events the read side needs (run.starting, run.running, run.completed or
run.failed); no stage or agent event is projected yet.
Scenario tests run the hello bundle on the OpenAI twin under the version
flag and a command-only bundle under the server setting, check Petri's
record agrees, and cover the refusals for an unknown attribute, an
undeclared node and an unknown model.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri's store conformance suite runs over `HttpRunStore` talking to an
axum listener on a loopback port. The suite opens runs under keys of
its own, while a key over the API is a Fabro run id the worker's token
names, so an adapter gives each suite key a fresh run with a token
minted for that run alone: the least a worker holds.
Three more tests cover what the suite cannot: the operator release
through the server's store turns the worker's handle stale; a
middleware swallows the reply of one committed append and the store's
resend leaves each record once; and two workers with owners of their
own never hold one run's lease at the same time.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The server answers the `/api/v1/runs/{id}/petri/*` endpoints from one
`SqliteRunStore` over its pool. `PetriRuns` in `AppState` keeps the
writer handle each worker opened, keyed by the run and the worker's
owner id, so the lease semantics stay the store's: the handle drops on
the worker's `release`, and every handle of a run drops when the server
observes the run's worker exit, in the subprocess wait path. Never by
timeout. A write from an owner with no held handle reopens only when
the lease row still names that owner, so a server restart or a lost
open reply recovers, and an owner the lease moved away from gets
`petri_stale_owner`.
Every endpoint is worker-scoped through the existing worker auth; a
new `RequireWorkerRunSegment` extractor covers the two-segment routes.
Store errors answer with a machine-readable code, the leased owner and
the conflict position under `meta`, and a backend failure's cause goes
to the server log rather than the worker.
A test drives a held worker through the scheduler, opens the run over
the API with its token, ends the worker, and sees the lease end.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The worker shape of the integration plan (F1.3) needs a run's worker to
reach the run's Petri records over the server's API. This adds the
contract: six worker-scoped endpoints under `/api/v1/runs/{id}/petri/`
(open, release, list and append records of one log, write and read a
blob), their request and response schemas, and the generated Rust and
TypeScript clients.
A store error needs more than a code: `petri_run_leased` names the
holding owner and `petri_record_conflict` names the refused position.
`ErrorResponseEntry` gains an optional `meta` object for such
code-specific members, `ApiError` can carry it, and the client's
`ApiFailure` parses it beside the code so a caller can act on it.
Records travel as `{seq, recorded_at, record}`, the store's own unit,
with `seq` and `recorded_at` as `uint64`. The log path segment is the
log id's text (`coordinator`, `resources`, `execution <n>`), which the
generated client percent-encodes. The blob write reuses
`WriteBlobResponse`, since Petri's digest is Fabro's blob hash.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A workflow version names its engine with `engine = "petri"` in the
`[workflow]` table of `workflow.toml`, and `[server.execution] engine`
(`FABRO_SERVER_ENGINE`, `--engine`) defaults it for every version that
names none. The choice, with what Petri admitted (the lowered root graph
and its children by blob and digest), is recorded on the run spec as
`RunEngine`, carried on `run.created`, and replayed into the projection.
A legacy run's spec omits the field, so existing specs decode unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Pebble main a39f43e26effdf99635eaf343f095c17157c9c93 (pebble #22) carries
an assistant turn's usage as Usage in the session record and moves the
record format to version 5. CodingRuntime::from_record refuses a record
in another format with UnsupportedRecord { version, supported } before
it reads the route. Fabro persists those records in SQLite for Ask Fabro
resume, and old runs get no migration, so a record written by an older
build is read back as stored and refused on the next turn.
Two tests pin that down. The store reads pebble's own version 4 fixture
back through get without a parse error and reports it unsupported. A
resumed Ask Fabro session whose stored record declares the previous
format fails its next turn with the agent_error code and the message
"session record format version 4 is not supported (this build requires
5)", runs no turn, and leaves the stored record in place.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Re-pin lithos-llm to 55add4596b861a0623d00c3a54aa5c147c8d504b and
pebble to c91810fe51aece80359b9cd8efea971af0c46925, where token usage
and cost travel together as Usage { tokens: TokenCounts, cost:
Option<Cost> }. Fabro now carries that one type everywhere it used to
carry BilledTokenCounts, BilledModelUsage, UsdMicros, or a token count
beside a cost_usd_micros.
fabro-types: billing.rs is usage.rs with ModelRef, ModelUsage { model,
usage }, sum_usage, and usage_is_empty; billing_rollup.rs is
usage_rollup.rs with ProjectionUsageStage, ProjectionUsageByModel,
ProjectionUsageRollup, and usage_rollup_from_projection. Every usage
field is named usage: StageProjection.usage and usage_by_model,
Outcome<Option<ModelUsage>>, stage.completed and stage.failed usage and
usage_by_model, prompt.completed usage, run.completed and run.failed
usage (total_usd_micros is gone), Conclusion.usage, StageSummary.usage,
Run.usage. RunSize buckets by Cost.
fabro-workflow: model_usage_from_llm prices tokens from the catalog with
a Catalog cost source, with_reported_cost keeps a provider cost, and the
pebble handler's stage_usage groups pebble's accounts by model and sums
rows with Usage::saturating_add, so a total has a cost only when every
priced part was priced. The store fold's live usage is the agent's
usage plus its descendants'.
API: the OpenAPI spec deletes BilledTokenCounts, BilledModelUsage,
CompletionUsage, CompletionCost, TokenUsage, and RunBillingSummary,
adds TokenCounts, Cost, Usage, and ModelUsage, and renames every
billing schema, property, tag, path, and operation to usage. fabro-api
reuses lithos-llm's and fabro-types' types through with_replacement,
with a round-trip test per replacement.
Old stored runs get no migration: their pebble events in the old shape
read back with zero usage, and their rebuilt projections lose agent
usage.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The demo agent stage's stored events now read as one pebble session: MCP
servers up and failed, skills, a subagent, a failover, a compaction, and a
written file, ending with ProcessingEnd. Demo mode serves the run state it
answered not_implemented to, with the agent stage carrying the coding
agent's fold of those events, so the stage sidebar renders them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>