Petri's main since #33 and #34: the launch goal, the simulated dry-run
provider, and lithoscomputer/petri#35, which keeps Fabro's dry runs on
the host workspace their checkpoints work. The pin moves to the merge
commit once #35 lands.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A worker-backed run's in-memory managed run settled only when the
worker process exited. GET /runs/{id} reads the stored summary, which
the projector ends at Petri's own `run.finished` record, a moment
before the worker stores Fabro's terminal lifecycle record and exits.
The delete precheck prefers the managed run, so a delete issued in that
window was refused with 409 "cannot remove active run". Against a real
worker the window hit about six times in ten.
The server sees both records before they are stored: `run.finished` on
the coordinator log through the worker's records endpoint, and the
terminal lifecycle record through the platform-records endpoint. The
managed run now settles at either, ahead of the store, so the view
never reports the run ended while the managed run still says running.
The stream follower no longer reopens a settled run with the records
that precede its terminal one, and the worker's exit keeps the settled
status: it reaps the process, records a missing terminal record as it
did, and takes the store's status only when the store ended the run
differently. The mapping from Petri's finish to the run's status is the
projection's own, shared with its fold.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
An intent's goal override, and a `[run.goal]` layer, reached the run's
display graph and its settings but not Petri's check, so the agent stages
executed with the workflow's own goal while the run showed the override.
The launch now carries the run's resolved goal (the settings' inline
`run.goal`, layered as the create path layers it) as Petri's
`petri.launch_goal` compile variable, which Petri binds over the bundle's
`[run] goal` and the graph's own `goal`, so admission's frozen plan carries
the goal the run shows.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri 911dbb5f adds the `petri.launch_goal` compile variable and lets a
settings goal override the graph's own. Re-pin to the merge commit once
lithoscomputer/petri#33 merges.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The stored view reports a Petri run ended as soon as its own run.finished
record is folded, which is before the server stores the terminal
lifecycle record and settles the managed run in its map. The delete
precheck reads that map, so a delete sent as soon as the API reports the
run ended can be refused as active instead of by the held lease. The
scenario now waits for the managed run to settle, through a test-support
accessor for its status, before it deletes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The Docker recovery scenarios scraped the worker log for "sandbox
workspace brought to its durable snapshot"; since the host and sandbox
restores share one path, the line reads "workspace brought to its
durable snapshot" with the site as a field.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Four doc comments the projection split cut in half, the scope ledger's
cache miss over an Option<Option>, the Pebble envelope passed by value,
and the Progress enum's large Pebble variant, now boxed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
RecoveryRequest::for_run and HooksSpec::for_run derived the author, the
identity source, the checkpoint settings and host_workspaces from the
run namespace with the same expressions. RunGitSettings, in checkpoint.rs
beside RunWorkspaces, is that derivation once; both specs carry it.
recovery::plan nested the "recorded, else found and reconciled" lookup
two loops deep and tracked a found flag over a tuple list. A Planner
holds the records, the workspace lookup and the recorded checkpoints,
and its methods read in order: targets, snapshot_of, reconcile_record,
newest. Candidate names the (key, sha) pair, and the Target struct that
duplicated its key's execution is gone; last_finish yields the
CheckpointKey itself, whose Display the failure reason now uses.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
projector.rs mixed the commit rule with things that are not the pass:
the ordering rule and its tests, the in-process signalling store, the
stream table's rows and reads, and six readers documented "for a test".
projector/mod.rs now holds the pass alone; order.rs, signalling.rs and
stream.rs hold the rest; the test readers (rebuild, stored_projection,
stored_stream, stored_platform_records, event_json) live in test_support
behind the test-support feature, as the test-support boundary rule asks,
and the two test callers reach them there. The fault injection and the
cache's test-only reader are gated the same way, so neither ships.
pass() reads as three steps: at_head, read_new_events (a NewEvents
struct in place of a tuple) and stream::stream_rows, which rebuild
shares instead of repeating the fold loop. PassReport::skipped and
PassReport::contended replace the zero-filled literals.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
stage_key formatted "<execution>:<firing>" and five places re-derived or
re-parsed that string: the engine, progress and platform folds, the
projector's ordering rule, and fork.rs's stage_labels, which split it
back apart. FiringKey is the fact itself, with of_event for the event
side and From<StagePosition> for the platform-record side. It still
serializes as "<execution>:<firing>", so the fold_json a stored view
holds keeps its shape and needs no migration; a unit test pins that.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fold_progress matched five payload kinds by string and walked each
payload's JSON by hand. Progress is now one internally tagged enum over
those kinds, with a struct per payload (PlannedRoute, OfferedTool,
ForkOccurrence) as the Attractor steps document them, and a test holds
each kind literal to the constant petri_attractor_steps exports, so a
rename there fails a test here instead of projecting nothing.
The fallback plan's original route carries its reasoning effort and
speed in the same lithos-llm types StageModelUsage holds, and the fold
now keeps them; it set both to None before. That is visible on the
stage's provider_used for an agent stage that ran under a fallback plan.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Four small duplications in the projection fold: a StageCompletion
literal built three times from an attempt's status, a StageModelUsage
literal built three times with no request controls, the pause and
unpause status arithmetic written four times across the coordinator and
lifecycle folds, and a bare Conclusion built beside the full one. Each is
now one function: completion() in the engine fold, StageModelUsage::new,
RunStatus::paused and RunStatus::unpaused beside blocked_reason (with
settle_control for the pending control they clear), and
Conclusion::outcome_only.
One case reads differently: a coordinator RunPaused that lands on a run
already paused behind a block now keeps that block for the unpause, as
the lifecycle Paused already did, instead of dropping it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
projection.rs held 1,834 lines: the fold state, the platform-record fold,
the lifecycle fold, the coordinator fold, the engine and view folds, the
progress and Pebble folds, the sandbox mapping, the tool mapping and the
model parsing, in one file. VIEWS.md is organised by source; the code now
is too. projection/mod.rs keeps RunView, FoldState and the helpers every
fold shares; platform.rs, coordinator.rs, engine.rs, progress.rs,
sandbox.rs and model.rs each hold one source's rows. A pure move: no
function body changed, only the visibility the cross-module calls need.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Eleven private hook methods returned Result<_, String> and rendered
their causes with the same collect_chain(...).join(": ") in fourteen
places. The error-handling strategy reserves String for rendered
projections. HookError keeps every failure with its source; it is
rendered once, by HookError::render, at the four Petri boundaries that
carry text: the adjusted outcome of a failed checkpoint, a transition's
problems, a scope acquisition's error, and the log. CheckpointKey gains
a Display so three messages stop spelling it out by hand. The rendered
text is the same as before, which the failed-checkpoint tests assert on.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
FabroHooks held twenty-one flat fields. Three of them were "a set read
from the store once, then kept current", in two shapes: a Mutex beside a
OnceCell<()> that had to be initialised first, and, for the restore plan,
a OnceCell<Mutex<_>> that could not be read uninitialised. The recorded
checkpoints and the collected artifacts now use the second shape too.
The checkpoint state (committed, recorded, last, branch), the artifact
state (globs, collected) and the scope state (environments, inherited
workspaces, per-workspace locks) are three structs with their own methods,
so each invariant lives in one type instead of in every caller.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
RunWorkspaces already ran every git command through a private Site
(a host path or a sandbox environment) but exposed each operation twice,
as commit/commit_in, matches/matches_in, has_commit/has_commit_in,
reset/reset_in, restore/restore_in and workspace_head/workspace_head_in.
The pairs propagated into every caller: recovery had bring_host_to and
bring_sandbox_to, the hooks had snapshot and snapshot_in_sandbox, and
restore_host and restore_sandbox. Site is now the public parameter, each
operation exists once, and the callers collapse to one function each.
The hooks resolve a scope's workspace and site in one place, site_of.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Seven crates' files each carried the same four-line lock function that
recovers a poisoned mutex. fabro_util::sync::lock is that function, once;
the copies in fabro-petri and fabro-server are gone. fabro-template's
helper panics on poison instead, a different policy, and is left as is.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The in-process Petri path persisted the run's terminal lifecycle
record, settled the projector, aggregated usage, and only then settled
the managed run. GET /runs/{id} reads the stored summary, so it reported
the run as ended while the delete precheck, which prefers the managed
run, still saw it running and refused the delete as active. The prune
scenario hit that window about once in thirty runs. Settle the managed
run right after the record is stored, before the view catches up.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two fabro-server scenarios run their container on the catalog image
ghcr.io/lithoscomputer/ubuntu-22.04:slim and wait about five seconds
for the run to finish; on a runner without the image the plugin's pull
takes longer than that. Pull it with the default runner image before
the suite, and give the Docker job the default runner image pull too,
since its fabro-petri Docker test runs that image.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The include error names the partial relative to the run's working
directory, so its `../` run is as long as that directory is deep: eight
on this machine's temp dir, three on the CI runner's. Collapse the run
to a token before the snapshot compares.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The detached cancel test waited for the run's status to read `failed`
and then asserted on the stored `run.lifecycle` record. The projection
concludes the run from Petri's `run.finished` coordinator record, and
the worker stores the platform's terminal lifecycle record a moment
later, so the read raced the write and the assertion failed about once
in thirty runs. Wait for the record itself, and print the stored events
and the run state when the assertion fails.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The host_plugin_ and docker_plugin_ variants ran each scenario under
Fabro's old plugin transport with the provider kinds `host` and
`docker-plugin`, which Petri's Fabro frontend rejects. Under Petri every
provider is already served by a sandbox-driver plugin, so those variants
test nothing distinct. A single docker_ variant replaces them: an
environment with provider `docker` on buildpack-deps:noble, created on
an isolated server, skipping without the sandbox-driver-docker
executable or a daemon with the image unless
FABRO_REQUIRE_SANDBOX_PLUGINS is set.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every Petri run takes its scope through a sandbox-driver plugin
executable that Petri finds on PATH, so the test jobs need
sandbox-driver-host and sandbox-driver-docker installed at the rev the
workspace pins. The three jobs share one from-source install through an
actions/cache entry keyed on the OS and the rev.
The Linux test job also pre-pulls Petri's default runner image, which
the suite's Docker scenarios leave to Petri: the plugin pulls it on
first use, but a 1 GiB pull inside a run's timeout is a flake.
The stdio plugin job was built for the deleted fabro-sandbox layer. It
becomes the Docker providers job: the `docker_` scenario variants and
the fabro-petri suite, with the fabro-sandbox and fabro-workflow steps
whose tests no longer exist removed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
CodeQL read the assertion messages as a log of the credentials'
Debug output. The test proves that output never holds the key, so the
messages added nothing.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri 12e8a17 merges origin/main into PR #30's branch and pins
sandbox-driver at 07600aa, the same revision this workspace moved to in
the last merge, so the lock links one copy of sandbox-driver again.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Clippy's absolute-paths lint on the rendered prune error, the sync
directory reads the prune test documents, and a redundant rustdoc link.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A live-server scenario runs a one-stage workflow on Docker whose command
writes a file into the workspace, opens an Ask Fabro session on the
finished run, and sends one turn. The session attaches to the container
Petri created, stopped at the run's end, starts it again, and its tool
reads the file inside it; the turn succeeds, the tool's output and the
model's reply carry the file's content, and the twin's follow-up request
shows the model read it from the tool. Ask Fabro's tool policy is
read-only, so the shell tool is hidden from the model and refused; the
turn reads the file with the `read_file` tool, scripted on the twin,
instead of a shell `cat`.
The scenario skips, and says why, without the Docker plugin or a daemon,
as the other Docker scenarios do.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Run deletion called the driver's `provider.delete(id)` under the run's
`petri.run` scope, a delete of Fabro's own over a sandbox whose lease
record Petri owns. It now goes the way `petri sandbox prune` goes:
`fabro_petri::prune` builds the run's Petri runtime over the server's
store (the run key, the run directory, the sandbox backend) and calls
Petri's prune, which opens the run for writing, checks each lease's
provider fingerprint, writes the delete intent and the tombstone beside
the run's other records, and lets each provider remove its managed
workspace, a host workspace included.
A run a live process holds answers 409 unless the delete is forced; a
lease Petri could not prune answers 409 with the problem text, or is
warned and skipped under force or a delete that already started. The
server drops the worker's handles before the prune, on the store
instance the prune opens, so the lease a stopped worker held is released
first. The projection reads only the coordinator and execution logs, so
the resource records change nothing it reports.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The worker's environment carried only the host, Docker and Daytona plugin
variables from the server's own environment, so a run on a third-party
provider kind never learned where its plugin was, although the server
kept `[server.sandbox.providers.<kind>]` plugin settings for its own
attach. The launch spec now derives `PETRI_SANDBOX_<KIND>_PLUGIN` and
`PETRI_SANDBOX_<KIND>_SHA256` from every enabled kind's plugin settings,
and `PETRI_SANDBOX_PLUGIN_DEV=1` when any of them sets `dev`, set after
the allowlist so the settings win over an ambient variable of the same
name and the allowlist stays the fallback.
The server's default plugin binary name is `sandbox-driver-<kind>`, the
name Petri looks up, since one executable serves both sides.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The grep test reads paths as the driver reports them for a resolved
absolute path; the local preflight check test gives the manifest a
source directory that exists; the Docker attach scenario accepts that
the in-process app has no daemon record for the run-tools client an Ask
Fabro turn builds after the sandbox attach, and asserts the turn got past
the sandbox. The inventory's lazy connection is boxed for clippy's
variant-size lint.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A server scenario runs a command workflow on Docker through Petri's
plugin (skipped without the plugin or a daemon), then reaches the
container without Petri: the sandbox tab describes it under its
`petri.run` label, Run Files writes, lists and reads a file in its
workspace after starting the stopped container, a preview URL opens to a
port in it, and an Ask Fabro turn runs against it through the OpenAI
twin. A unit test attaches through the ownership seam with a scripted
provider: the run's own container attaches, another run's and one that
carries only Fabro's retired `sh.fabro.*` labels are refused as not
owned.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Nothing imports it any more: the Pebble glue lives in
fabro-pebble-sandbox, the server reaches run sandboxes through
sandbox_access, and Petri creates every run sandbox. The crate, its
test-support, its integration tests and every dependency edge go with
it. The `[server.sandbox.providers.<kind>.plugin]` settings stay: the
server still launches a plugin executable through them to attach to a
sandbox of a non-bundled kind.
AGENTS.md names the new crate and the direct-access pattern in place of
`RunSandbox`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The sandbox handlers describe, list, download, upload, open a terminal
and build SSH and VNC access on the `Arc<dyn Sandbox>` the server attaches
to the run's record, with paths resolved against the recorded working
directory. Run Files holds the handle beside that directory and runs its
git through the driver's git facet and Fabro's exec policy;
fabro-workflow's sandbox git takes the same pair, and its
`GitCommandError` carries the driver's error. Run deletion deletes by id
through the provider scoped to the run's `petri.run` label, so a foreign
sandbox is refused and a designated host directory is left in place. Ask
Fabro wraps the attached, running handle. The legacy access shim is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`fabro-server/src/sandbox_access.rs` is the server's own path to a run's
sandbox: it connects the record's provider (the driver's Host, Docker and
Daytona providers in process, a plugin executable for any other kind),
keys ownership on the `petri.run` label Petri stamps on every sandbox it
creates, attaches by the recorded id, and for a host record designates
the recorded directory again when the id lives only in the worker's
registry. The Docker client resolves its endpoint from the same variables
Petri forwards to its plugin, so both meet on one daemon.
The doctor's Docker check and the Daytona credential probe move here with
`DaytonaCredentials`, and the `/sandboxes` inventory is rebuilt over the
driver's `list`, narrowed to sandboxes that carry Petri's run label.
Preflight asks the provider for its health instead of creating and
deleting a throwaway sandbox in Fabro's own shape, which no run uses; the
git retry policy behind the repository probe moves into run_manifest.
The callers still on fabro-sandbox's reconnect read their access through
a `legacy_provider_access` shim until they move.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Petri creates and owns every run sandbox through the sandbox driver, so
what Fabro still needs around a driver handle is the Pebble glue: the
Environment pebble's coding agent runs its tools through, the exec policy
(stop grace, working directory, StripAll, the termination mapping, the
redacted output tail), pebble's port routes over the driver's preview
URLs, the secret redactor, the path helpers, and a log rendering that
appends a failed command's redacted tail. This crate holds that glue,
moved from fabro-sandbox, over `Arc<dyn Sandbox>` plus a working
directory instead of `RunSandbox`, with a `MockSandbox` double behind
`test-support`.
`fabro exec` creates its host sandbox directly on the driver's Host
provider and activates it; Ask Fabro wraps the attached handle. Both
keep the provider alive beside the sandbox where the session's processes
are the provider's process groups.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fabro's own DOT parser was left with one job after create-time compile
moved to Petri: walking a workflow's file references for the bundler and
the workflow-version store, and reading a name, a goal and two counts.
Petri's frontend parses the same language, so the parser goes and a small
crate reads the graph through Petri's.
`fabro-dot` is that crate: `WorkflowGraph::parse` over
`petri_frontend_attractor::dot` and its semantic model (defaults applied,
subgraphs flattened, chains expanded), `references(position)` as the one
walker over the static-reference vocabulary (each reference with its node
and position, file references checked to be template-free), and
`normalize_for_graphviz`, the re-emit of Fabro DOT with dotted attribute
keys quoted, which the SVG render needs. It sits beside `fabro-petri`
rather than inside it because `fabro-petri` depends on `fabro-workflow`,
which depends on `fabro-workflow-version`: the version store cannot reach
`fabro-petri` without a cycle, and the bundler should not pull the engine
in to read a graph.
Deleted: `fabro-graphviz`'s lexer, grammar, AST, semantic pass and
`parse_ast` (1,829 lines, plus the `nom` dependency); the DOT model in
`fabro-types::graph` (`Graph`, `Node`, `Edge`, `AttrValue`,
`shape_to_handler_type`), with only `ReferenceKind` kept, moved to
`fabro_types::reference`; `fabro-template`'s `visit_graph_references` and
the `GraphReference`/`GraphPosition` types, with the template-syntax rule
(`validate_static_reference`) kept there; the pull-request body's DOT
fallback summary, which was unreachable because the DOT source only
travels with the run spec whose display graph the summary already reads.
`fabro-graphviz` is now the render alone, over `fabro-dot`.
Parity: the old and new walkers were run over every `.fabro` and `.dot`
file in the repository (118) before the deletion. Every reference set is
identical. Five files differ in what Petri reads more correctly: a
backslash before a newline inside a quoted string is a line continuation
(four files, inline prompt text only), and a node named only by an edge
counts as a node (`test/edge_only_node.fabro`, 3 nodes rather than 2, so
the `fabro validate` snapshot moves). The checked-in bundles' shapes and
references are pinned by a snapshot in `fabro-dot`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The rest of the change whose deletions the previous commit carries (its
`git add` stopped at an already-removed path): `fabro_types::RunGraph`
and the `fabro-petri` builder that reads it off the admitted graph, the
server's create, validate, preflight and render paths on Petri's check
alone, the consumers moved to the new shape, the OpenAPI `RunGraph`
schemas with their parity tests, the regenerated TS client, and the
docs naming Petri's diagnostic codes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every run is admitted by Petri, whose check lowers imports, file
references, templates and the model stylesheet, lints the workflow and
pins its models. Fabro then re-parsed the same workflow through its own
legacy pipeline (parse, transforms, structural validation) only to fill
`RunSpec.graph` for the read side. That second pass is gone: the run's
display graph is `fabro_types::RunGraph`, built once in `fabro-petri`
from the admitted graph's metadata (the workflow name and goal from the
graph params; each declared stage's label and handler kind; one edge per
routing arm as written, lowering artifacts left out), and stored on the
spec at create beside the DOT as `graph_source`.
Deleted: `fabro-workflow`'s `pipeline`, `transforms`, `file_resolver`,
`operations::{source, validate}`, `run_materialization` and the legacy
`compile_admitted_run`; the server's `compile_admitted`, the structural
manifest pass, `preflight_model` and the model probe `run_llm_check`
(Petri's admission raises `attractor.model.unknown`); `fabro-graphviz`'s
stylesheet parser; most of `fabro_types::graph` (the DOT model keeps
what the bundler, version registration and template walker read). The
DOT parser stays for the bundler and the SVG render.
`POST /validate`, `POST /preflight`, `POST /graph/render`, `fabro
validate` and `fabro preflight` run on Petri's check alone, so their
diagnostics carry Petri's codes (`attractor.unbound_input`,
`unsupported.template.unbound_input`) where Fabro's
`template_undefined_variable` and `goal_self_reference` were. A refused
workflow's summary still names the DOT as written. The OpenAPI `RunSpec`
schema gains `RunGraph`, `RunGraphNode` and `RunGraphEdge`, reused from
`fabro-types` with parity tests; the TS client is regenerated.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The CLI artifact scenario seeded its run through the deleted upload
route; it now runs a real Petri workflow whose hooks collect the
artifacts, and the fabro artifact list and cp assertions read those.
A real command retry is not producible from a command node (a plain
failure or a timeout routes onward), so the retry dimension of the old
fixture goes; the stage, node, and retry filters, the tree copies, the
cross-stage ambiguity, and the filename collision stay covered. The
archive guard test drops its upload row (the blob write row covers an
octet-stream mutation). A dangling doc comment and two absolute paths
clippy flagged are fixed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>