`GET /runs/{id}/events` and `GET /runs/{id}/attach` serve a Petri run's
public events and Fabro's platform records as one ordered stream in a
Fabro envelope (`RunStreamItem`: `run_id`, `stream_seq`, `kind`, `id`,
`recorded_at`, `item`), read from the projector's `petri_stream` table.
The cursor is `stream_seq` (`?after=`); the item's own identity (the
Petri `EventId` as `<log>/<seq>/<index>`, or the platform record's seq)
travels beside it for deduplication. A legacy run keeps its envelope on
the same endpoints; the OpenAPI response is the union of the two lists,
and the stream list reports Petri's `EVENT_CONTRACT_VERSION`.
The attached stream follows the projector's commit signal (a wake-up,
with a poll as the fallback) and ends after the platform record of the
run's terminal lifecycle transition, the analog of the legacy stream's
`run.completed`, or a bounded grace after the projection went terminal.
`RunSpec.engine` (`RunEngine`, `PetriAdmission`, `PetriGraphRef`) is
named in the spec and reuses the Rust types. `fabro-client` matches the
union and adds `list_run_stream`, `list_run_stream_page` and
`attach_run_stream`.
A server test attaches to a two-branch parallel run, disconnects once
both branches started, records a platform notice while both branch
scripts run, reconnects from the last `stream_seq`, and checks the
union is the whole stream: every item once, in order, no gap, no
duplicate, the notice between the branch events, and the same as the
paged listing. The Petri scenarios capture their settled projection and
stream as JSON fixtures for the web app under
`FABRO_CAPTURE_PETRI_FIXTURES`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|---|---|---|
| .. | ||
| src | ||
| tests | ||
| Cargo.toml | ||
| README.md | ||
| VIEWS.md | ||
fabro-petri
Fabro's adapters over Petri, the workflow engine Fabro runs its workflows on.
Layering rule
Only this crate imports Petri. The workspace Cargo.toml pins the Petri
packages by revision under petri_* keys, and fabro-petri is the only
member that lists them as dependencies. Every other Fabro crate reaches the
engine through what this crate exports. A Petri pin move is therefore a change
to this crate and the lockfile, nothing else.
What it holds
Every adapter the integration plan describes lands here.
SqliteRunStore: Petri'sRunStoreandRunLogsover Fabro's SQLite database, so a run's records live in Fabro's tables (petri_runsfor the run and its writer lease,petri_recordsfor every record of every log, and the sharedblobstable). The module docs state the lease and append rules.runtime: the Petri runtime Fabro assembles, the same way at create time and at execution: the Fabro frontend with the server's settings layer, the Attractor step kinds (real, or simulated for a dry run), the model client as thePebbleClientcapability, the Fabro home.check: Petri compiles at create time. The workflow version's bundle goes into an in-memory file map (frontend::MapFiles, laid out as the bundle:workflow.tomlbeside the workflow,.fabro/project.tomlat the root),Runtime::check_sourcelowers it with the run's inputs and launch, and the admitted graphs or Petri's diagnostics come back in a shape the server maps onto Fabro's. Nothing is written to disk.admission: the admitted graphs in Fabro's blob store, named on the run spec asRunEngine::Petri(PetriAdmission), verified by digest on load.engine: a run executed by Petri, started from its admitted graphs or resumed from its records, with the outcome read from the run's record throughinspect_runand mapped to the conclusion Fabro's read side records. The run's worker process runs it overHttpRunStore; the server runs it in its own process only under its test override, overSqliteRunStore. The caller supplies the interviewer, and the secret provider and blob table when it has them.interview: Petri'sInterviewerover Fabro's questions API and the worker's control channel. A question has one id in Fabro, Petri's own (gate#2): the projection lists it pending from thequestionrecord,GET /runs/{id}/questionsserves it, and the answer posted to/questions/{qid}/answeris validated against that pending record and reaches the worker's control interviewer over the control bus (or the in-process one directly) under the same id, mapped onto Petri's answer. The adapter still posts the legacyinterview.*events (through the worker's run event sink, or the run's database in the server process) with that id and the projection's stage label, for the readers that follow the event stream rather than the projection: Slack,run attachand the web app's Q&A renderer. The store derives theinterview.answeredplatform record, with the answering principal, frominterview.completed. An expired or cancelled question is completed asinterview.timeoutorinterview.interrupted; an auto-approved run answers itself.secrets: Petri'sSecretProviderover the vault's token entries, so a{{ secrets.NAME }}reference resolves at spawn into a command's environment and is masked in every record; a sensitive answer registers as a dynamic secret.blobs: Petri'sOutputStoreover Fabro'sblobstable, through the server'sBlobStoreor the worker's client, so a large stage value leaves the records for the table underblob://sha256/<hex>.HttpRunStore: the same store as a run's worker process reaches it, over the server's/api/v1/runs/{id}/petri/*endpoints with the worker's token. The server answers from itsSqliteRunStore, so the lease and the(log, seq)rule are the store's; this layer carries requests, resends a request whose reply was lost, maps the server's error codes back toStoreError, and, for a worker, takes every lease for the worker's launch id. The module docs state the rules.petri: the Petri store vocabulary re-exported for the server, which answers the worker endpoints from aSqliteRunStorewithout naming a Petri package in its own manifest.projection: the fold of a Petri run's public events (replay_sinceover its records) and Fabro's platform records (fabro-store'splatform_records) into theRunProjectionthe API serves, row by row asVIEWS.mdmaps them. The stage key is(execution, firing); theStageIdlabel isnode@visit, made unique with the execution when two child invocations would share one.projector: the view pass and its wake-up. Records first: Petri's append and a platform record's insert return before any view work; a pass reads what is committed, folds the items past the committed positions, and writes the projection document (petri_projection), the ordered stream (petri_stream, onestream_seqper Petri event or platform record) and the narrowedrunsrow in one later transaction. The server signals the projector after each committed worker append, after each committed platform record (the run summary store's hook), at worker exit and, over every Petri run, at startup. A run that executes in the server process goes throughProjector::observe_store, which signals after each append. A torn tail (a record Petri cannot read) holds the view where it stands and reports the run incomplete with the reason.- The platform adapters the plan adds after it: hooks and the run tools.
What the projection leaves default
VIEWS.md rows with no source yet, or whose source this crate does not read
yet, keep their default value in the projection: StageProjection.diff and
Conclusion.diff.patch (the checkpoint's patch_blob is not resolved),
Checkpoint's engine-derived maps (completed_nodes, node_retries,
context_values, node_outcomes, next_node_id), agent_tools,
permission_level, script_invocation and script_timing, a stage's
notes, StageCompletion details for a parsed.note, the sandbox instance
(the matrix's two gaps), Run.ask_fabro, an interview option's
description and preview, the pull request creation state, and the
run's notices, notifications and pairings (recorded, not shown).
A run goes to Petri when its workflow version's workflow.toml names
engine = "petri" in [workflow], or when the server's
[server.execution] engine (FABRO_SERVER_ENGINE, fabro server start --engine) says so for versions that name none. The server side of both
halves is fabro-server's server::petri_runs; the worker side is
fabro-cli's commands::run::petri_worker, which fabro run __run-worker
takes when the run's stored spec names Petri. After a server restart, a
Petri run left in flight goes back to a worker in --mode resume: the run
continues from its records, as Petri's own resume does, and full recovery
of the workspace to a durable snapshot is the plan's F3.5.
How it is tested
Integration tests live under tests/:
runs.rsruns thehellobundle in memory throughRuntime::standard()with the Fabro frontend and the model-free stub registry, then a command-only workflow on the host sandbox through the real step registry. Both skip, and say why, when thesandbox-driver-hostplugin executable is not onPATH(every run takes its scope's environment through it); the sandbox-plugins CI job requires them.check.rsadmits thehellobundle and round-trips its graph through the blob store, binds the launch, admits a version whoseworkflow.tomlnamesengine = "petri", reads the project settings from the map, and refuses an unknown attribute, an unknown[workflow]key (unsupported.workflow_toml.key, named inworkflow.toml) and, with a model client over the test catalog, an unknown model (attractor.model.unknown). No plugin is needed.sqlite_store.rsruns Petri's store conformance suite (petri_testkit::run_store::conformance) againstSqliteRunStore, plus the operator release, lease exclusivity, a crash between appends, and blob interoperation with Fabro'sBlobStore.interview.rsruns human gates through the engine assembly with the interview adapter over a control interviewer: a gate answered under the posted id, two parallel gates each bound to their own answer, an expiry with the gate's default, an auto-approved run, and a cancelled run.secrets.rsresolves a{{ secrets.NAME }}reference from a vault into a command's environment overSqliteRunStoreand checks the value is in nopetri_recordsrow while the masked output is.blobs.rsoffloads a command's large output to theblobstable and reads it back by theblob://sha256/<hex>reference a record carries.model.rsruns thehellobundle against the OpenAI twin with a model client over a vault that holds the key, and checks the skills step searched the configured Fabro home.
Those four need the host plugin like runs.rs does, and model.rs also
starts the twin.
projection.rsbuilds the view live (every append signals the projector) for thehellobundle on the stub registry, a command-only workflow and a two-branch parallel workflow, and checks it equals the view rebuilt from the records alone (projector::rebuild); catches a view up after every wake-up was dropped, by a signal and by the startup pass; recovers a crash between the record commit and the view transaction by applying only the missing suffix, with the positions andstream_seqcontinuing; runs two projectors over one store with child executions; and holds the view at a torn tail. All skip without the host plugin.
The conformance suite over HttpRunStore needs a server to talk to, so it
lives with the server's integration tests
(lib/apps/fabro-server/tests/it/api/petri_store.rs), which reach the suite
through this crate's test-support feature (fabro_petri::test_support).
Run them with:
ulimit -n 4096 && cargo nextest run -p fabro-petri
The server's end-to-end coverage is lib/apps/fabro-server/tests/it/scenario/petri.rs:
the hello bundle on the OpenAI twin, a command-only bundle and a
two-branch parallel bundle run to completion through the create handler and
the scheduler, in the server process under its test override, under the
version flag and under the server setting, with GET /runs/{id}/state
serving the projection over Petri's records; a human gate is answered
through the questions API; and Petri's diagnostics refuse a run at create.
The server's petri_runs unit tests cover the lease ending at worker exit
and the restart reconcile that relaunches a worker in resume mode.
The worker path is covered with the real binary in
lib/apps/fabro-cli/tests/it/scenario/petri.rs: a command-only Petri run
executes in the worker a foreground server launched, its records reach
petri_records over the HTTP store and its lease ends with the worker; and
a run whose server and worker are both killed mid-stage resumes in a new
worker after the server restarts, with one run.completed; a human gate in
the worker is answered through the questions API over the control channel;
two parallel gates each bind their own answer; and an unanswered gate
expires with its default.