- Keep plugin-era Daytona lease fingerprints: read only DAYTONA_API_URL and
DAYTONA_ORGANIZATION_ID (no URL alias, no placement target), and stop
forwarding DAYTONA_SERVER_URL and DAYTONA_TARGET to the worker.
- Take the Docker fingerprint and network from this process's DOCKER_HOST,
the endpoint the Docker client actually connects to; make the provider
configuration's fields private.
- Return an error instead of panicking when Petri supplies no Host registry.
- Run deletion reads the Daytona key only for a Daytona run, and a forced
or restarted delete goes on when the secret store fails, as it does for
every other prune failure.
- Stop putting DAYTONA_API_KEY in the worker's environment; the worker reads
it from the vault. Give the worker's Daytona client the shared HTTP client.
- Fork, rewind and retry no longer read the vault: a fork acquires no sandbox.
- Remove the dead worker plugin forwarding and document that runs execute
only on the built-in providers.
- Build every Petri runtime through providers::standard_runtime or
bare_runtime, with a Clippy lint against Runtime::standard/bare.
- Share the Docker require-or-skip policy in fabro-test, tighten the Host
scope assertion.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Load the Daytona key for fork and prune through one AppState method
instead of two copied vault reads, and pass the sandbox configuration
into runtime_spec rather than building it and overwriting it. The
worker reuses the CLI's process_env_var lookup.
Share one Docker availability check and the backend-requirement
variable through fabro-test, drop the built-in plugin path and pin
constants nothing reads any more, and let enabled_plugins() exclude the
bundled kinds itself. Refresh the comments and the spawn_env test that
still described built-in plugins.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Register lazy Host, Docker, and Daytona factories for Petri execution,
fork, and prune. Share server provider configuration, preserve lease
fingerprints, and source Daytona credentials from the vault.
Remove built-in plugin setup and skip gates; add a release-mode worker
and prune regression to catch the failure that blocked nightly builds.
Co-Authored-By: Codex <noreply@openai.com>
A worker-backed run's in-memory managed run settled only when the
worker process exited. GET /runs/{id} reads the stored summary, which
the projector ends at Petri's own `run.finished` record, a moment
before the worker stores Fabro's terminal lifecycle record and exits.
The delete precheck prefers the managed run, so a delete issued in that
window was refused with 409 "cannot remove active run". Against a real
worker the window hit about six times in ten.
The server sees both records before they are stored: `run.finished` on
the coordinator log through the worker's records endpoint, and the
terminal lifecycle record through the platform-records endpoint. The
managed run now settles at either, ahead of the store, so the view
never reports the run ended while the managed run still says running.
The stream follower no longer reopens a settled run with the records
that precede its terminal one, and the worker's exit keeps the settled
status: it reaps the process, records a missing terminal record as it
did, and takes the store's status only when the store ended the run
differently. The mapping from Petri's finish to the run's status is the
projection's own, shared with its fold.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The Docker recovery scenarios scraped the worker log for "sandbox
workspace brought to its durable snapshot"; since the host and sandbox
restores share one path, the line reads "workspace brought to its
durable snapshot" with the site as a field.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A live-server scenario runs a one-stage workflow on Docker whose command
writes a file into the workspace, opens an Ask Fabro session on the
finished run, and sends one turn. The session attaches to the container
Petri created, stopped at the run's end, starts it again, and its tool
reads the file inside it; the turn succeeds, the tool's output and the
model's reply carry the file's content, and the twin's follow-up request
shows the model read it from the tool. Ask Fabro's tool policy is
read-only, so the shell tool is hidden from the model and refused; the
turn reads the file with the `read_file` tool, scripted on the twin,
instead of a shell `cat`.
The scenario skips, and says why, without the Docker plugin or a daemon,
as the other Docker scenarios do.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The CLI artifact scenario seeded its run through the deleted upload
route; it now runs a real Petri workflow whose hooks collect the
artifacts, and the fabro artifact list and cp assertions read those.
A real command retry is not producible from a command node (a plain
failure or a timeout routes onward), so the retry dimension of the old
fixture goes; the stage, node, and retry filters, the tree copies, the
cross-stage ambiguity, and the filename collision stay covered. The
archive guard test drops its upload row (the blob write row covers an
octet-stream mutation). A dangling doc comment and two absolute paths
clippy flagged are fixed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The worker control bus was publish-only: the steer and interrupt
endpoints answered 202 once the control was forwarded, and a refusal
showed up only later as a `run.notice` on the run's stream.
A steer or an interrupt now carries a request id. The worker answers it
over the control stream it arrived on with `{request_id, outcome}`,
where the outcome is `delivered` (with the stage's label) or `refused`
(with the code and the reason). The server keeps the outstanding
requests in a registry and waits up to 5 s for the answer: the endpoint
answers 202 `{"outcome":"delivered","stage":…}`, 409 with the refusal's
code (`no_live_turn`, `no_such_stage`, `steer_refused`,
`interrupt_refused`) and message, or 202 `{"outcome":"pending"}` when
the worker gave no answer in time. The `run.notice` record on refusal
stays, under the same code, so a steer to a stage that is not running is
now `no_such_stage` there too. Pause and unpause are unchanged.
The in-process test path answers a steer or an interrupt from the run's
own controls at once. `FABRO_TEST_CONTROL_ACKS_MUTED=1` on the server
mutes the worker's answers, so a test can see the pending fallback.
`fabro steer` prints the worker's answer, and a refusal is its error.
The OpenAPI spec documents the 202 body and the 409 codes; the Rust and
TypeScript clients are regenerated.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Through a real server and its worker: a three-stage run forked at its
first stage continues with the other two on the first stage's restored
file and commits on the new run branch after the source's commits; a
retry of a run whose last stage failed transiently reruns that stage and
succeeds; a rewind archives and supersedes its source, and an archived
run is refused; the timeline lists every checkpoint record with its
commit; a fork at a checkpoint inside a parallel branch is refused and
creates nothing, while one after the join continues.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The worker's `run.notice` for an interrupt it could not deliver now reads
"Interrupt of stage `gate` refused: the stage has no model turn to
interrupt" (or "Interrupt refused: …" when the control named no stage),
so the web and the CLI can show which stage and why, with the reason as
Petri's `ControlError` spells it. The gate scenario asserts both messages.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri 639ce3e added `ControlService::interrupt_firing` and the `LiveTurns`
capability a host installs beside the pause hooks. Fabro now drives it:
`RunControls::interrupt(stage, text)` resolves its stage the way a steer
does (a label, a node name, or the run's one live agent stage) and stops
that stage's current model turn, keeping the session; the text, when
given, is the stage's next input. `engine::run` installs the live-turn set
as a runtime capability, so without it no interrupt could ever land.
The worker maps `run.interrupt` and `run.interrupt_then_steer`, both of
which now carry an optional `stage`, to that call. The control bus is
one-way, so a refusal is recorded the way a refused steer is: a
`run.notice` on the run's stream whose code says why (`no_live_turn` when
Petri refuses a stage with no turn in flight, `no_such_stage`,
`interrupt_refused`).
The server's `POST /runs/{id}/interrupt` and `POST /runs/{id}/steer` with
`interrupt=true` forward the control and answer 202, replacing the 501
`interrupt_unsupported` stub. The interrupt endpoint takes an optional
body (`stage`, `text`), refuses a finished run with 409
`run_not_interruptible`, and forwards an interrupt of a blocked run, since
an agent stage may be running a turn beside the question and the worker
judges each stage itself. `fabro events --pretty` prints the delivered
interrupt and the stage's `attractor.turn.interrupted` report.
Verified on the twin: `fabro steer --interrupt` during a long tool call
ends the turn, the text is the agent's next request, and the stream
carries the `$interrupt` record and the interrupted-turn report; an
interrupt of a gate stage is refused with `no_live_turn` and the gate's
question is untouched.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`run.branch` and `git.identity` are written by the checkpoint that creates
the run branch, before that firing's finish is appended, and carried no
position, so the stream ordered them by the millisecond clock: on either
side of the finish from one run to the next. Both now take that
checkpoint's stage position, and the existing ordering rule places them
after the firing's finish and before its routes, beside its checkpoint
record. The two CLI snapshots that had each recorded one of the two
orders now record the one order every run produces.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`SteerRunRequest` takes an optional `stage`: the label the projection
shows (`node@visit`, or `node/e<execution>@visit` when two executions
share one) or the node's name. The server passes it on the worker control
message; the worker's `RunControls` resolves a label to the live agent
firing and steers that firing, and a node name through Petri's own
live-stage index. Unnamed, the one-live-agent rule stays, and the refusal
now names the live stages by their labels. `fabro steer --stage` sets it.
A controls scenario runs two agent stages side by side, sees the unnamed
steer refused with both named, and steers each apart, one over the API
and one through the flag.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The worker's signal handlers pause and unpause the run through its
`RunControls`, the same path the server's pause and unpause take, so a
signal holds admission, records Petri's `run.paused` and the lifecycle
mirror, and releases it on the unpause. The legacy pause state the plan
named no longer exists in the tree; nothing was left to delete. A controls
scenario sends both signals to a real worker and reads the records back.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`run events --pretty` reads the flattened `git.identity` fields, and
shows a `run.diff` record as its summary and an `artifact.collected`
record as its path and size. The snapshot filters redact the base commit a
`Branch:` line names. The CLI snapshots now carry the `Base:` line, the
`run.branch` and `git.identity` stream items, the dry run's simulated
response and the two response files a dry run dumps.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
No store sits behind `[server.slatedb]` any more: the section leaves the
settings layer, the resolved server settings, the defaults, the API schema,
the TypeScript client, the install wizard and the docs, and `fabro install`
probes the bucket for the `artifacts/` prefix alone. A settings file that
still carries the section is rewritten once at startup by a temporary
migration that removes it with a backup beside the file.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Step 4 of the legacy executor deletion, third commit: the legacy event
API and every reader of it go, so that the next commits can delete the
event log, its reducer and the types beneath them.
The API:
- `GET /runs/{id}/events` pages the run stream only
(`PaginatedRunStreamList` by `after`); the legacy `since_seq`,
`before_seq` and `order` cursors, the `oneOf` envelope, the legacy
`EventEnvelope`, `PaginatedEventList`, `RunEvent`, `EventSeq`,
`AppendEventResponse` and `RunEventDetailResponse` schemas,
`POST /runs/{id}/events`, `GET /runs/{id}/events/{seq}` and
`GET /runs/{id}/stages/{stageId}/events` are deleted. `GET
/runs/{id}/attach` and `GET /attach` frame `RunStreamItem`s only.
- The Rust and TypeScript clients regenerate; the removed models leave
the TypeScript package.
The readers:
- `fabro-client` drops the legacy run event listing, tail and attach
methods and `RunEventStream`; `list_run_stream_until` bounds a stream
read.
- `fabro-tool`'s `fabro_run_events` lists, searches and details the run
stream: `after` is the exclusive `stream_seq` cursor, `event_id` the
item's id, filters match the item's name and `recorded_at`.
- `fabro-dump` writes the stream to `events.jsonl`; `fabro dump` reads
it.
- The CLI's progress renderer keeps only what the run stream drives:
the legacy event conversion, the sandbox and setup displays and their
styles go. `fabro system events` prints stream items.
- The server's demo mode folds its agent fixture straight into the
session projection and answers the attach stub with a stream item;
the demo stage events endpoint is gone.
- The web app: every run is a Petri run. The legacy event hooks,
renderer props, stage popover summary, run phases derivation and
live-event payload handling are deleted or ported to `RunStreamItem`;
toasts and board refreshes read the stream's platform records.
- Tests: the legacy API round trips and pagination tests are deleted;
the CLI's MCP, attach and system event mocks serve stream pages; the
CLI test helpers read stream items.
Still failing until the later commits: the CLI tests seeded through
`POST /runs/{id}/events`, the server tests over the legacy store, and
the legacy type tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Step 4 of the legacy executor deletion, first commit of several: step 4
spans commits because the legacy event log and its consumers cannot go
in one compiling change. This commit moves every writer off `run_events`;
the reducer, `EventBody`, the Slate bridge and the API's event types
still exist for the readers the next commits port or delete.
Writers:
- The server records a run's lifecycle (submitted, runnable, starting,
running, blocked, paused, control requests and effects, the terminal
status), its title, parent link, archive state, notices and pull
request state as platform records (`fabro_store::platform_records`),
through the new `server::run_records` module. Every append wakes the
projector and waits for its pass, so the read that follows a write
holds the record.
- Pull request creation is recorded as `pull_request.requested`,
`pull_request.created`, `pull_request.failed`, `pull_request.linked`
and `pull_request.unlinked`; the projection folds them into the run's
pull request and creation state.
- Answers to questions are recorded as `interview.answered` with the
answering principal and the answer text; the interview adapter no
longer posts legacy `interview.*` events (`QuestionSink` is now an
optional observer).
- The worker (`fabro run __run-worker`) records its lifecycle, notices
and pause state over `HttpPlatformRecords`; `HttpRunStore` for the
legacy event log and the worker's `run_store` are gone.
- `persist_created_run` appends `run.created` and `run.submitted`.
Readers:
- A stream follower (`server::stream_follower`) follows each live run's
stream (Petri events and platform records), folds lifecycle records
into the in-memory run state, forwards items to the global attach
broadcast, and syncs blocked and paused from the projection.
- Slack posts questions from the projection's pending interviews,
finishes them on `interview.answered` or `question_expired`, and sends
lifecycle notifications with `notification.sent` dedupe.
- `GET /runs/{id}/events` and the attach endpoints serve only the run
stream; the per-event, per-stage and `POST /runs/{id}/events`
endpoints and their tests are deleted.
- `Database::load_run_projection` reads the Petri projection only.
Deleted with the writers:
- The SQLite blob and run-history activation migrations and their
legacy Slate imports (`legacy_blob_import`, `legacy_run_history_import`,
the activation backup): a greenfield server has no Slate history to
import, and the run-history verification refused to start a server
whose runs have no legacy events.
- `fabro-workflow`'s `operations::archive` and `operations::run_store`.
- The server's legacy-event unit tests and the CLI's `HttpRunStore` tests.
The in-process answer transport is now set after the starting and
running records land, not gated on the live status still being
`Starting` (the records already moved it).
The manifest validation test for a `run.agent.mcps.<name>` catalog
reference now expects `unsupported.workflow_toml.run.agent.mcps.reference`:
Petri's Fabro frontend has no server catalog to resolve it against.
Legacy readers still fail their tests until the next commits: the
reducer and Slate tests in fabro-store, the fabro-workflow create tests
that read the run back through the legacy store, the CLI tests seeded
through `POST /runs/{id}/events`, the CLI's legacy attach and render
paths, the sessions API, the OpenAPI conformance test, and the web
fixtures.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every run executes on Petri, so the in-process legacy executor goes:
`fabro-core` and, in `fabro-workflow`, the handlers, lifecycle, pipeline
execution, routing, retry, conditions, node handlers, steering, agent
memory, artifacts, checkpoints, command log, and the `start`, `resume`,
`retry`, `fork`, `rewind` and `timeline` operations. The two are deleted
together because the engine half of `fabro-workflow` was the only user of
`fabro-core` and `fabro-core` the only runtime of that half; neither
compiles without the other.
Kept in `fabro-workflow`, narrowed: the parse/transform/validate/persist
pipeline and `create`, `archive`, `validate` (workflow definitions still
come from DOT and settings); the run tools (`run_tools`, moved from
`handler/llm/fabro_tools.rs`) for Ask Fabro, `fabro exec` and Petri's
host tools; the pull request pipeline (`pull_request`, moved from
`pipeline/`, for the step 0 port); Run Files' diff helpers in
`sandbox_git`; `git_identity`, `usage_rollup`, `run_status`,
`run_materialization`, `web_search` and `workflow_bundle`.
Server: `RegistryFactoryOverride` becomes `execute_in_process`;
`RunAnswerTransport::InProcess` carries only the interviewer; the
interrupt endpoint answers 501 `interrupt_unsupported` and every pair
endpoint 501 `pair_unsupported` (status lists none); rewind, fork, retry
and timeline handlers and routes are removed; the command log is served
from the stage output blob; usage rollups accumulate from the settled
projection after an in-process run as after a worker exit.
Ported while here:
- `materialize_admitted_run` materializes the goal and drops a disabled
pull request block, as the legacy materializer did.
- A run whose admitted graph has an agent or prompt node is refused at
create when no LLM provider is ready (`fabro.model.no_ready_provider`);
a workflow of commands and gates needs no model and is admitted.
- The projection's question type falls back on the options, as the
interview adapter does, so a gate with edge-label options answers as
multiple choice.
Tests: the server scenarios (lifecycle, run completion, SSE, helpers)
run in process on Petri and assert Petri's stage labels and stream
names; the reconcile tests assert Petri's relaunch semantics; legacy
unit tests of the deleted executor are removed; three server unit tests
the removal took with it are restored; the pair fixtures go with the
pair feature. Petri test fixtures no longer name `[workflow] engine`.
Still red after this commit, all legacy consumers the next steps
delete or port: fabro-store's Slate/reducer fixtures and fabro-types
legacy JSON tests (step 4); server unit tests over legacy run events
(retry endpoints, list_run_events, artifacts, per-event pause/unpause,
run history activation, legacy sandbox fixtures) (steps 3-4); CLI tests
that parse legacy event envelopes, the legacy `events`/`attach`/`diff`/
`dump`/`inspect` snapshots, `run rewind`/`run fork`, the ACP and
git-identity workflow tests, and the runner tests that drive the legacy
worker by hand (steps 3-4); the web app's Petri fixtures still carry
`engine` (regenerate with `FABRO_CAPTURE_PETRI_FIXTURES` in step 4).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Delete `Engine`, `RunEngine`, `[workflow] engine`, `[server.execution]
engine`, `FABRO_SERVER_ENGINE` and `fabro server start --engine`. The run
spec records what Petri admitted as `admission: PetriAdmission`; the
create handler always admits through `Runtime::check`; `execute_run`
always launches the Petri worker (or executes in process under the test
override); the CLI runner takes only the Petri worker path, and its
legacy control arm, artifact uploader, signal pause handlers and
credential helpers go with it. The CLI's `attach` and `events` read the
run stream only.
Two gaps this surfaced are closed here: the check adapter binds the
server's run variables as Petri compile variables (`{{ vars.* }}` in a
prompt no longer fails admission), and deleting a run removes its Petri
records, lease, platform records, projection and stream.
Tests: the config engine tests are replaced (an engine key is unknown),
the API round-trip test covers `PetriAdmission`, the server and CLI
Petri scenarios drop their engine settings, and the API tests that read
legacy event names now read the run stream or the session events. The
remaining red tests are fixtures and scenarios of the legacy executor
and the legacy event store (`fabro-store` `slate` and `run_state`,
`fabro-types` legacy `run.created` JSON, the server's handler-registry
scenarios, the CLI dry-run snapshots), which the next steps of the F4.3
series delete or port.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A Petri run on Docker or Daytona keeps its workspace inside the scope's
sandbox. Fabro's hooks now take the environment Petri hands them at
`scope_acquired`, run `git` inside the scope through it (the path Petri's
own sandbox-placed hooks take), and commit each stage on the run branch
with the same message and trailers as the host path. The commit leaves
the sandbox as a Git bundle, created against the newest ancestor the
snapshot repository already holds, split into 8 MiB parts (the plugin
transport reads one file up to 16 MiB), read out through the
environment's file transfer, fetched into the bare snapshot repository
on the host and named there under the checkpoint's ref. The repository
holds every checkpoint whatever the provider, and the platform records
name the same commits. The host path is unchanged; both sites share one
runner and the same commands.
Recovery is split: `recovery::plan` decides, over the records and the
snapshot repository alone, what every live workspace must sit on and
reconciles a lost record; `recover` applies it to host workspaces on the
server, as before, and reports a sandbox workspace's target as deferred.
The worker's hooks read the same plan at the scope's first acquisition
after a resume and bring the sandbox workspace to it before any attempt
runs there: verified or reset in a retained sandbox that still holds the
commit, else restored from a bundle of the checkpoint written into the
sandbox. Petri replaces a lease's lost sandbox on Fabro's request
(`LostSandbox::Replace`), so a removed container comes back fresh and
restored.
The checkpoint records are written for every provider now. A Docker
variant of the in-process hooks test moves a 20 MiB file through the
split transfer; the same test runs on Daytona when live credentials are
present. Three CLI scenarios run on a Docker environment: every stage's
checkpoint published from the container, a retained container whose
workspace drifted reset on restart, and a removed container replaced and
restored from the snapshot.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three scenarios over `petri.rs`'s harness: a pause between two command
stages holds the second until the unpause while the API says `paused`
with no pending control; a steer sent while the agent stage waits on a
tool reaches its session on the twin, which sees the text in its next
request, and the stream carries the `control.requested` record; a run
paused with its next stage held at admission, whose server and worker
then die, resumes paused, admits nothing until the unpause, and then
finishes. The harness helpers the sibling module needs are opened to it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three scenarios on the real binary: an agent creates a child run with
`fabro_run_create` from inside a Petri run and the child carries the
parent link; a `[[run.hooks]]` pre_tool_use hook blocks a run tool, the
model reads the reason, and Petri's record holds the report and the
denied call; a sub-agent calls an inherited run tool, recorded under the
parent stage naming the parent session.
The Petri scenario harness is shared: the server can start with extra
settings and vault entries, and the detached run takes extra arguments.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fabro's terminal `run.lifecycle` record lands a moment after Petri's
`run.finished`: the worker exits, the server records the status, the
projector folds it. A CLI scenario that asserts on the end of the stream
now waits for that record instead of reading the stream as soon as the
runs row turns `succeeded`, which the projector writes from the engine's
finish alone.
The fabro-petri README names the projector's stream reader and commit
signal, the server's reconnect test with its fixture capture, and the
CLI scenarios that read a run back through the stream.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`run events` on a Petri run prints the run stream: raw, the envelope as
one JSON line per item; `--pretty`, Petri's events by `<subject>.<verb>`
with the stage's label (a visit's start and end with its elapsed time,
the route both ends of the edge, a fork's branches, a question with its
options and its answer, log lines, the agent's messages and tool calls,
the engine's finish) and the platform records by kind (the run's
creation, its lifecycle, a checkpoint's commit, a pull request, a
notice, who answered). `--follow` attaches from the last `stream_seq`
printed and reconnects from its cursor when the server ends the stream
before the run's terminal record.
`run attach` on a Petri run replays the stream through the progress
renderer (a new mapping from stream items onto the progress events the
renderer draws, sharing the coding-agent mapping with the legacy
envelope), follows it live from its cursor with the same reconnect, asks
a question the stream carries at the terminal, and exits with the status
the engine's finish or the terminal lifecycle record decides. `wait` and
`inspect` read the projection unchanged.
The CLI never names a Petri type: `PetriItem` reads the item as JSON
where Petri's contract keeps the event name, the subject and the parsed
progress payloads.
The CLI's Petri scenarios read the stream instead of the legacy events
(the lifecycle records, the question and who answered it, the expiry),
and three new ones cover a finished run through `events` (raw, tail,
and a `--pretty` snapshot), `attach`, `wait` and `inspect`; `attach`
answering a gate from the terminal; and `events --follow` to the run's
end.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
In-process tests over the memory store: every finish is committed on the
run branch with its identity trailers and recorded with its commit, the
run-end hooks reach Petri's local service through Fabro's wrapper, a
stage that fails on its own terms is committed and its failure route runs
on the committed files, a failed checkpoint records `checkpoint_failed`
with no route taken and a restart reports the run failed, and a
`[[run.hooks]]` hook blocks an agent's tool call through the forwarded
service, with the model told why.
Real-binary scenarios crash the server and its worker with SIGKILL: after
a durable finish the stage's commit is not repeated and the interrupted
stage reruns on its snapshot; a crash held before the commit reruns the
stage once; a crash held after the commit but before its record
reconciles the record from the snapshot repository; a deleted workspace
is restored; a failure route sees the same committed files after a
crash; a failed checkpoint fails the run and a restart leaves it failed.
Recovery selects the executions `inspect_run` reports incomplete, and a
run whose coordinator log is still empty is left to the worker's resume.
The worker's platform record endpoints get an API test and the generated
TypeScript client.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The interview adapter derived its own question id from Petri's identity
and posted it on `interview.started`, while the projection over Petri's
records serves the pending question under Petri's `Question.id` with the
firing's stage label. The answer endpoint validates against the
projection, so an answer under the projection's id never reached the
adapter's wait.
The adapter now waits under Petri's id and labels the question's stage
through the projection's own rule: `stage_label`, `is_shown` and
`visit_of` move out of `start_visit` into shared functions, and the
adapter's observer derives each firing's `visit.started` through Petri's
`Projection`, as the projector does, so the label matches by
construction. The full Petri identity stays on `AskedQuestion`.
The legacy `interview.*` events are still posted, under Petri's id, for
the readers that follow the event stream rather than the projection: the
Slack service, `run attach`, the web app's Q&A renderer and the server's
answer claim. The store already derives the `interview.answered` platform
record from `interview.completed` for a Petri run, so who answered is
recorded under Petri's id with the answering principal.
The gate scenarios assert the new identity and encode the id as one path
segment, as the generated clients do. Projection tests cover an expired
question and an auto-approved answer.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The server scenario answers a gate in the in-process run through the
questions API and checks the branch it routed and the cleared pending
question. The CLI scenarios drive the real worker: a gate answered through
the API over the worker's control channel, two parallel gates each bound
to their own answer, and an unanswered gate that expires with its default
and records `interview.timeout`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
When `fabro run __run-worker` finds its run's stored spec names Petri, the
new `petri_worker` module executes it through `fabro_petri::engine` over
`HttpRunStore`, leased for a launch id the worker mints and logs at start.
`--mode start` loads the admitted graphs through the client's blob read;
`--mode resume` continues the run from its records. The worker's existing
services carry over: the control channel's cancel and SIGTERM/SIGINT cancel
Petri's root invocation politely, a lost control channel cancels the run
and is reported once it settles, and pause, unpause and steer are received
and ignored with a warning until their adapters land. The model client
comes from the worker's catalog and vault snapshot for the providers whose
credentials resolve, and the lifecycle events (`run.starting`,
`run.running`, then `run.completed` or `run.failed`) go through the client
as the legacy worker's do.
Scenario tests against the real binary: a command-only Petri run executes
in the worker a foreground server launched, its records reach
`petri_records` over the HTTP store and its lease ends with the worker; and
a run whose server and worker are both killed mid-stage resumes in a new
worker after the server restarts, with one `run.completed`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`ps --json` reports the digraph name as `workflow_graph_name` and reserves
`workflow_name` for an explicit `[workflow] name` (6a86ced77). That change
updated the ps tests but not this ignored e2e test, which still expected
the graph name under `workflow_name`. The test now asserts the contract
the ps tests assert: `workflow_name` is null for a bare graph file and
`workflow_graph_name` is the digraph name.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>