Commit graph

688 commits

Author SHA1 Message Date
Bryan Helmkamp
dc48197553
Add the artifact, run diff and branch workspace platform records
`artifact.collected {execution, firing, attempt, path, blob, bytes, digest}`
records one file a stage left in its workspace, with the bytes in the blob
table; `run.diff {base_sha, head_sha, diff_summary, patch_blob}` records
the run branch against its base; `run.branch` names the workspace the
branch was created in. `RunProjection.artifacts` lists the collected
files, and `parse_blob_ref_encoded` reads Petri's `#json` marker so a
reader knows whether a blob is text or a JSON value.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 16:00:11 -04:00
Bryan Helmkamp
2f4888c199
Describe the run stream where the docs described the legacy event log
`docs/internal/events-strategy.md` is now the run stream strategy: the
two logs (Petri's records and Fabro's platform records), the projector
that folds them and assigns `stream_seq`, how to record a fact Petri
cannot know, how to read the stream, and the Ask Fabro session log.
`docs/internal/events.md` (the 106-event catalog) and the event schema
v2 shape document described the deleted `EventBody` model and are
deleted; AGENTS.md routes to the strategy for platform records and
stream consumers. The testing strategy's `progress.jsonl` rules name
records and stream items instead, and the public API nav drops the
removed per-stage events endpoint.

The interview adapter's module docs and the fabro-petri README no
longer claim the adapter posts `interview.*` events: readers see a
question in Petri's own progress record, and the server records who
answered as the `interview.answered` platform record.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 14:45:13 -04:00
Bryan Helmkamp
4942788297
Drop the run_events table and narrow the runs row
Every run is a Petri run whose history is `petri_records` and
`platform_records`, so the legacy run event log has no reader left.
The new migration drops `run_events` (with its indexes), the two
one-time activation tables, and rebuilds `runs` without the columns
only that log wrote or read: `source_last_seq` and the six token and
file-count columns nothing read, as VIEWS.md records. Pre-cutover
development runs are discarded, as decided; the surviving columns of
existing rows are copied across.

fabro-db loses the three migration consts of the dropped schema and the
session-owner preflight that inspected `run_events`, and gains
`DROP_RUN_EVENTS_MIGRATION_SQL` so fixtures that install the runs
schema reach the production shape. The run summary upsert binds only
the surviving columns.

Tests: the `run_events` schema, query-plan and preflight tests are
deleted; `runs_schema_has_its_final_shape_without_the_legacy_event_log`
pins the final columns and indexes, and
`dropping_the_event_log_keeps_the_run_rows` migrates a database left by
an older binary and checks the run row survives.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 14:37:18 -04:00
Bryan Helmkamp
60b503322c
Delete the legacy run event log, its reducer and its types
Step 4 of the legacy executor deletion, fourth commit: with no writer
and no reader left, the legacy event log goes.

- `fabro-types`: `run_event` (`EventBody`, `RunEvent` and every props
  struct), `EventEnvelope` and the `RunEventDetail*` types are deleted.
  What the projection and the API still use moves out of the event
  vocabulary: `AgentEventProps`, `AgentSessionActivatedProps`,
  `AgentToolsAvailableProps`, `StagePromptProps`, `SessionCapability`
  and the coding event names to `agent_props`; `RunNoticeLevel` and
  `RunNoticeCode` to `notice`; `InterviewOption` beside the question
  types; `RunRunnableSource` beside the run status. `Checkpoint` is
  what Fabro records for a Petri run: `timestamp`, `current_node`,
  `git_commit_sha`; the conclusion's stage summaries derive from the
  projection's stages instead of the checkpoint's node maps.
- `fabro-store`: the Slate bridge (`RunDatabase`, the Slate `Database`,
  `keys`, `record`, `EventPayload`) and the reducer (`run_state`) are
  deleted. `Database` is the blob table and the run summary store over
  one pool; the blob store is SQLite only; the run summary store keeps
  the `runs` row a projector writes and lists, and finds the pull
  request creation candidates over `platform_records`; `build_summary`
  and `projected_usage` live in `run_summary`. The SlateDB dependency
  is gone. Test fixtures build the store from its two SQLite stores.
- `fabro-workflow`: the `event` module (the `Event` enum, its
  conversion, sink, emitter, redaction, stored fields and names),
  `runtime_store`, `StageScope` and the legacy seeding test helpers are
  deleted; the tests that seeded legacy runs read platform records or
  a projection instead.
- `fabro-sandbox` owns `GitRetryReason`.
- The server builds the store without an object store; the legacy
  `POST /runs/{id}/events` tests go, an interrupt answers
  `interrupt_unsupported` in the tests as it does in the handler, and
  the tests that read a run back through the Slate handle read its
  projection or its platform records. The projection folds a block
  that lands while the run is paused as the pause's prior block, and a
  pause or unpause clears the pending control it answers; a control
  request's check-and-append holds a per-run lock so two concurrent
  cancels record one request.
- The CLI's final output is the response of the last stage that
  produced one; the workflow tests read completed nodes from the
  succeeded stages.
- The spec's `RunCheckpoint` carries the three fields the type keeps.

Still failing until the next commits: the CLI tests that seed runs
through `POST /runs/{id}/events` or wait for legacy event names, and
the two Ask Fabro resume tests (the sandbox instance gap).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 14:09:17 -04:00
Bryan Helmkamp
bc31b8d23a
Serve the run stream as the only run event API
Step 4 of the legacy executor deletion, third commit: the legacy event
API and every reader of it go, so that the next commits can delete the
event log, its reducer and the types beneath them.

The API:
- `GET /runs/{id}/events` pages the run stream only
  (`PaginatedRunStreamList` by `after`); the legacy `since_seq`,
  `before_seq` and `order` cursors, the `oneOf` envelope, the legacy
  `EventEnvelope`, `PaginatedEventList`, `RunEvent`, `EventSeq`,
  `AppendEventResponse` and `RunEventDetailResponse` schemas,
  `POST /runs/{id}/events`, `GET /runs/{id}/events/{seq}` and
  `GET /runs/{id}/stages/{stageId}/events` are deleted. `GET
  /runs/{id}/attach` and `GET /attach` frame `RunStreamItem`s only.
- The Rust and TypeScript clients regenerate; the removed models leave
  the TypeScript package.

The readers:
- `fabro-client` drops the legacy run event listing, tail and attach
  methods and `RunEventStream`; `list_run_stream_until` bounds a stream
  read.
- `fabro-tool`'s `fabro_run_events` lists, searches and details the run
  stream: `after` is the exclusive `stream_seq` cursor, `event_id` the
  item's id, filters match the item's name and `recorded_at`.
- `fabro-dump` writes the stream to `events.jsonl`; `fabro dump` reads
  it.
- The CLI's progress renderer keeps only what the run stream drives:
  the legacy event conversion, the sandbox and setup displays and their
  styles go. `fabro system events` prints stream items.
- The server's demo mode folds its agent fixture straight into the
  session projection and answers the attach stub with a stream item;
  the demo stage events endpoint is gone.
- The web app: every run is a Petri run. The legacy event hooks,
  renderer props, stage popover summary, run phases derivation and
  live-event payload handling are deleted or ported to `RunStreamItem`;
  toasts and board refreshes read the stream's platform records.
- Tests: the legacy API round trips and pagination tests are deleted;
  the CLI's MCP, attach and system event mocks serve stream pages; the
  CLI test helpers read stream items.

Still failing until the later commits: the CLI tests seeded through
`POST /runs/{id}/events`, the server tests over the legacy store, and
the legacy type tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 13:36:00 -04:00
Bryan Helmkamp
34996d630f
Give Ask Fabro sessions their own event log
Step 4 of the legacy executor deletion, second commit. Ask Fabro's
sessions were the last writer of `run_events`: a session's creation, its
turns and their messages, tool calls and endings went into the run's
legacy event log, keyed by the run's sequence. They now have a log of
their own.

- `run_session_events` (migration `2026091802`): one row per session
  event, numbered per session from 1, with the owning run, the turn, the
  event name and its properties. `RunSessionEventStore` appends under the
  write lock, lists a session from a sequence, names a session's owner
  from its creation event, deletes a run's sessions with the run, and
  publishes each committed event to its subscribers.
- `fabro_types::SessionEvent`: `seq`, `session_id`, `run_id`, `ts` and a
  flattened body (`event` naming the kind, `properties` its fields), with
  the same event names and property shapes the legacy events carried,
  so the web app and the CLI read the same JSON. The property structs
  move to `session_event`; `run_event::session` re-exports them under
  their old names until the legacy event log goes.
- The API: `GET /sessions/{id}/events` pages `PaginatedSessionEventList`
  by the session's own sequence, `GET /sessions/{id}/attach` replays and
  streams `SessionEvent` frames (subscribed before the replay, so no
  event falls between the two), the turn stream carries the same frames,
  and an interrupt answers with the recorded event. The session
  projection folds `SessionEvent`s; the legacy `find_session_owner` over
  `run_events` is gone.
- The CLI's `run ask` and the web app's session stream read
  `SessionEvent`; the web runtime no longer accepts the nested legacy
  envelope shape.

The two session resume tests in the server keep failing for a reason
this commit does not touch: Ask Fabro reconnects to the run's sandbox
from the projection's sandbox instance, which the Petri projection does
not carry yet (`VIEWS.md`, the `scope.acquired` gap).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 12:47:43 -04:00
Bryan Helmkamp
3e8b2ebfcc
Run the server's lifecycle over platform records instead of run_events
Step 4 of the legacy executor deletion, first commit of several: step 4
spans commits because the legacy event log and its consumers cannot go
in one compiling change. This commit moves every writer off `run_events`;
the reducer, `EventBody`, the Slate bridge and the API's event types
still exist for the readers the next commits port or delete.

Writers:
- The server records a run's lifecycle (submitted, runnable, starting,
  running, blocked, paused, control requests and effects, the terminal
  status), its title, parent link, archive state, notices and pull
  request state as platform records (`fabro_store::platform_records`),
  through the new `server::run_records` module. Every append wakes the
  projector and waits for its pass, so the read that follows a write
  holds the record.
- Pull request creation is recorded as `pull_request.requested`,
  `pull_request.created`, `pull_request.failed`, `pull_request.linked`
  and `pull_request.unlinked`; the projection folds them into the run's
  pull request and creation state.
- Answers to questions are recorded as `interview.answered` with the
  answering principal and the answer text; the interview adapter no
  longer posts legacy `interview.*` events (`QuestionSink` is now an
  optional observer).
- The worker (`fabro run __run-worker`) records its lifecycle, notices
  and pause state over `HttpPlatformRecords`; `HttpRunStore` for the
  legacy event log and the worker's `run_store` are gone.
- `persist_created_run` appends `run.created` and `run.submitted`.

Readers:
- A stream follower (`server::stream_follower`) follows each live run's
  stream (Petri events and platform records), folds lifecycle records
  into the in-memory run state, forwards items to the global attach
  broadcast, and syncs blocked and paused from the projection.
- Slack posts questions from the projection's pending interviews,
  finishes them on `interview.answered` or `question_expired`, and sends
  lifecycle notifications with `notification.sent` dedupe.
- `GET /runs/{id}/events` and the attach endpoints serve only the run
  stream; the per-event, per-stage and `POST /runs/{id}/events`
  endpoints and their tests are deleted.
- `Database::load_run_projection` reads the Petri projection only.

Deleted with the writers:
- The SQLite blob and run-history activation migrations and their
  legacy Slate imports (`legacy_blob_import`, `legacy_run_history_import`,
  the activation backup): a greenfield server has no Slate history to
  import, and the run-history verification refused to start a server
  whose runs have no legacy events.
- `fabro-workflow`'s `operations::archive` and `operations::run_store`.
- The server's legacy-event unit tests and the CLI's `HttpRunStore` tests.

The in-process answer transport is now set after the starting and
running records land, not gated on the live status still being
`Starting` (the records already moved it).

The manifest validation test for a `run.agent.mcps.<name>` catalog
reference now expects `unsupported.workflow_toml.run.agent.mcps.reference`:
Petri's Fabro frontend has no server catalog to resolve it against.

Legacy readers still fail their tests until the next commits: the
reducer and Slate tests in fabro-store, the fabro-workflow create tests
that read the run back through the legacy store, the CLI tests seeded
through `POST /runs/{id}/events`, the CLI's legacy attach and render
paths, the sessions API, the OpenAPI conformance test, and the web
fixtures.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 12:32:06 -04:00
Bryan Helmkamp
af38d68942
Delete fabro-validate and fabro-acp; validate on Petri's check
Petri judges a workflow at admission, so Fabro's lint rules go:
`fabro-validate` (its 35 rules and the `LintRule` trait) is deleted, and
with it `fabro-acp` (only a rule and two legacy executor tests used it),
the model-resolution transform, the legacy `create`, `compile_create_run`
and `materialize_create_run` stages, and `fabro-graphviz`'s `condition`
and `fidelity` modules. `Diagnostic`, `RelatedDiagnostic` and `Severity`
move to `fabro_types::diagnostic`, the one shape every diagnostic takes.

Validation is now the same question the create handler asks. A new
server module, `petri_check`, builds Petri's check request from a
workflow bundle and the run's settings (every workflow of the bundle at
its bundle-relative path, the inputs, the run variables, the launch),
runs the check, and maps the diagnostics; Fabro adds one rule of its
own, `fabro.model.no_ready_provider`, refusing a model node when no
provider is ready. Admission, the validate and preflight endpoints and
the offline `fabro validate` all go through it:

- `validate_prepared_manifest` runs Fabro's structural pass (parse and
  transform, whose diagnostics stay) and then Petri's check, on the
  blocking pool from the handlers;
- the offline `validate_manifest` checks with no model client and with
  an unbound input as a warning (`CheckRequest.unbound_is_warning`), so
  a workflow validates before its inputs exist; a collected workflow
  before upload checks with unbound inputs as errors, as before;
- preflight resolves each LLM node's selector against the ready
  providers and the catalog for its probe, as the deleted transform did,
  and no longer probes a model Petri refused;
- the graph render endpoint needs only the structural pass;
- a run manifest now carries its `[run.goal] file`, which Petri reads
  from the bundle as it does for a version.

The transforms keep the authored model selector (`sonnet` stays
`sonnet`): Petri pins the catalog model in the admitted graph, not in
the graph Fabro displays or in the settings snapshot. Tests assert that,
and the CLI's validate, preflight and graph snapshots carry Petri's
diagnostics (`attractor.no_start`, `attractor.undeclared_node`,
`attractor.bad_on_failure`, ...) in place of the lint rules' text.

Known gaps, Petri's side: a `workflow.toml` whose `[run.environment]`
names an environment the server catalog defines but the file does not
is refused (`unsupported.workflow_toml.run.environment`), as admission
already refused it; an unbound input inside an included template
partial is a render error rather than the unbound-input warning.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 11:32:04 -04:00
Bryan Helmkamp
5df733a22b
Move fabro-mcp's pebble mapping and test client into fabro-cli
`fabro-mcp` held two things after the legacy executor went: the mapping
from Fabro's MCP server settings to the servers pebble starts, which
only `fabro exec` still uses, and a stdio MCP client the tests of
Fabro's own MCP server speak through. The mapping is now
`fabro-cli`'s `mcp_servers` module and the client its test support's
`McpStdioTestClient`; the crate is deleted. Its `config` module was a
re-export of `fabro_types::settings::run`, which callers import directly.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:57:07 -04:00
Bryan Helmkamp
1f0dbd86ae
Delete fabro-hooks and the engine freeze check
`fabro-hooks` ran the legacy executor's hooks; Petri's Attractor steps
run Fabro's hooks now, so nothing in the workspace uses the crate. The
engine freeze (the CI workflow, the two scripts, and the AGENTS.md and
fabro-petri README sections) guarded the engine half of `fabro-workflow`,
which the previous commit deleted.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:47:56 -04:00
Bryan Helmkamp
d90a5d9cbb
Delete fabro-core and the engine half of fabro-workflow
Every run executes on Petri, so the in-process legacy executor goes:
`fabro-core` and, in `fabro-workflow`, the handlers, lifecycle, pipeline
execution, routing, retry, conditions, node handlers, steering, agent
memory, artifacts, checkpoints, command log, and the `start`, `resume`,
`retry`, `fork`, `rewind` and `timeline` operations. The two are deleted
together because the engine half of `fabro-workflow` was the only user of
`fabro-core` and `fabro-core` the only runtime of that half; neither
compiles without the other.

Kept in `fabro-workflow`, narrowed: the parse/transform/validate/persist
pipeline and `create`, `archive`, `validate` (workflow definitions still
come from DOT and settings); the run tools (`run_tools`, moved from
`handler/llm/fabro_tools.rs`) for Ask Fabro, `fabro exec` and Petri's
host tools; the pull request pipeline (`pull_request`, moved from
`pipeline/`, for the step 0 port); Run Files' diff helpers in
`sandbox_git`; `git_identity`, `usage_rollup`, `run_status`,
`run_materialization`, `web_search` and `workflow_bundle`.

Server: `RegistryFactoryOverride` becomes `execute_in_process`;
`RunAnswerTransport::InProcess` carries only the interviewer; the
interrupt endpoint answers 501 `interrupt_unsupported` and every pair
endpoint 501 `pair_unsupported` (status lists none); rewind, fork, retry
and timeline handlers and routes are removed; the command log is served
from the stage output blob; usage rollups accumulate from the settled
projection after an in-process run as after a worker exit.

Ported while here:

- `materialize_admitted_run` materializes the goal and drops a disabled
  pull request block, as the legacy materializer did.
- A run whose admitted graph has an agent or prompt node is refused at
  create when no LLM provider is ready (`fabro.model.no_ready_provider`);
  a workflow of commands and gates needs no model and is admitted.
- The projection's question type falls back on the options, as the
  interview adapter does, so a gate with edge-label options answers as
  multiple choice.

Tests: the server scenarios (lifecycle, run completion, SSE, helpers)
run in process on Petri and assert Petri's stage labels and stream
names; the reconcile tests assert Petri's relaunch semantics; legacy
unit tests of the deleted executor are removed; three server unit tests
the removal took with it are restored; the pair fixtures go with the
pair feature. Petri test fixtures no longer name `[workflow] engine`.

Still red after this commit, all legacy consumers the next steps
delete or port: fabro-store's Slate/reducer fixtures and fabro-types
legacy JSON tests (step 4); server unit tests over legacy run events
(retry endpoints, list_run_events, artifacts, per-event pause/unpause,
run history activation, legacy sandbox fixtures) (steps 3-4); CLI tests
that parse legacy event envelopes, the legacy `events`/`attach`/`diff`/
`dump`/`inspect` snapshots, `run rewind`/`run fork`, the ACP and
git-identity workflow tests, and the runner tests that drive the legacy
worker by hand (steps 3-4); the web app's Petri fixtures still carry
`engine` (regenerate with `FABRO_CAPTURE_PETRI_FIXTURES` in step 4).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 10:44:41 -04:00
Bryan Helmkamp
a36bea15d2
Remove the engine flag: every run is a Petri run
Delete `Engine`, `RunEngine`, `[workflow] engine`, `[server.execution]
engine`, `FABRO_SERVER_ENGINE` and `fabro server start --engine`. The run
spec records what Petri admitted as `admission: PetriAdmission`; the
create handler always admits through `Runtime::check`; `execute_run`
always launches the Petri worker (or executes in process under the test
override); the CLI runner takes only the Petri worker path, and its
legacy control arm, artifact uploader, signal pause handlers and
credential helpers go with it. The CLI's `attach` and `events` read the
run stream only.

Two gaps this surfaced are closed here: the check adapter binds the
server's run variables as Petri compile variables (`{{ vars.* }}` in a
prompt no longer fails admission), and deleting a run removes its Petri
records, lease, platform records, projection and stream.

Tests: the config engine tests are replaced (an engine key is unknown),
the API round-trip test covers `PetriAdmission`, the server and CLI
Petri scenarios drop their engine settings, and the API tests that read
legacy event names now read the run stream or the session events. The
remaining red tests are fixtures and scenarios of the legacy executor
and the legacy event store (`fabro-store` `slate` and `run_state`,
`fabro-types` legacy `run.created` JSON, the server's handler-registry
scenarios, the CLI dry-run snapshots), which the next steps of the F4.3
series delete or port.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:58:21 -04:00
Bryan Helmkamp
b5fdc015f0
Merge branch 'petri-integration-sandbox' into petri-integration
# Conflicts:
#	lib/apps/fabro-cli/tests/it/scenario/mod.rs
2026-09-18 09:11:19 -04:00
Bryan Helmkamp
d60744b4d6
Checkpoint and recover Petri workspaces inside Docker and Daytona sandboxes
A Petri run on Docker or Daytona keeps its workspace inside the scope's
sandbox. Fabro's hooks now take the environment Petri hands them at
`scope_acquired`, run `git` inside the scope through it (the path Petri's
own sandbox-placed hooks take), and commit each stage on the run branch
with the same message and trailers as the host path. The commit leaves
the sandbox as a Git bundle, created against the newest ancestor the
snapshot repository already holds, split into 8 MiB parts (the plugin
transport reads one file up to 16 MiB), read out through the
environment's file transfer, fetched into the bare snapshot repository
on the host and named there under the checkpoint's ref. The repository
holds every checkpoint whatever the provider, and the platform records
name the same commits. The host path is unchanged; both sites share one
runner and the same commands.

Recovery is split: `recovery::plan` decides, over the records and the
snapshot repository alone, what every live workspace must sit on and
reconciles a lost record; `recover` applies it to host workspaces on the
server, as before, and reports a sandbox workspace's target as deferred.
The worker's hooks read the same plan at the scope's first acquisition
after a resume and bring the sandbox workspace to it before any attempt
runs there: verified or reset in a retained sandbox that still holds the
commit, else restored from a bundle of the checkpoint written into the
sandbox. Petri replaces a lease's lost sandbox on Fabro's request
(`LostSandbox::Replace`), so a removed container comes back fresh and
restored.

The checkpoint records are written for every provider now. A Docker
variant of the in-process hooks test moves a 20 MiB file through the
split transfer; the same test runs on Daytona when live credentials are
present. Three CLI scenarios run on a Docker environment: every stage's
checkpoint published from the container, a retained container whose
workspace drifted reset on restart, and a removed container replaced and
restored from the snapshot.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 09:10:29 -04:00
Bryan Helmkamp
fc9b82aa7f
Note the worker's paused mirror and refused steers in VIEWS.md
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 07:57:19 -04:00
Bryan Helmkamp
5093265efb
Wire pause, unpause and steer into the Petri worker
A Petri run answered only cancel and answers; pause, unpause and steer
were ignored with a warning. `fabro_petri::controls::RunControls` now
wraps Petri's `ControlService` per run: `engine::run` installs its pause
gate over the run's hooks, observes the run through it and wires it to
the coordinator, on a start and a resume alike, so a run paused when its
worker died resumes paused.

The worker's control channel takes a `WorkerControls` enum: the legacy
hub and pause flag, or the Petri run's controls. On Petri, `run.pause`
holds admission, `run.unpause` releases it once the record is durable,
and `run.steer` goes to the one live agent stage (Fabro's steer names no
stage); with none or several it is refused with a `run.notice` record.
The paused state is mirrored to Fabro's lifecycle as `run.paused` and
`run.unpaused` events, so the server's live status and the projection
follow Petri's own records.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 07:49:54 -04:00
Bryan Helmkamp
806ac90c98
Freeze the engine half of fabro-workflow to bug fixes
The integration plan's F4.1: once Fabro runs on Petri, the engine half of
fabro-workflow (handler/, lifecycle/, pipeline/execute, graph/routing,
node_handler, retry, condition, context, model_fallback) takes bug fixes
only, and new engine behaviour goes to Petri.

scripts/check-engine-freeze.sh holds the frozen path list, diffs the
branch against a base ref and exits 1 when any frozen file gained lines;
scripts/check-engine-freeze-test.sh proves that on a synthetic
repository. The Engine freeze workflow runs both on every pull request
that touches the crate's src, re-runs on label changes, and fails unless
the pull request carries the `bugfix` label. AGENTS.md and the
fabro-petri README name the freeze, the label and the script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 02:19:29 -04:00
Bryan Helmkamp
6e38ad27fb
Order position-keyed platform records by their firing, not their clock
A checkpoint record is stamped by the server and Petri's records by the
worker, so ordering the stream by recorded_at could put a stage's
checkpoint after the next stage's route. A platform record that carries
a Petri position now goes right after its firing's finish: before the
firing's first routing.resolved in the pass (the next firing's
visit.started hangs off that record), else after the firing's last
event, else, when the firing finished in an earlier pass, before the
first event of a later firing. The hook writes the record after the
driver appended the attempt's finish, but the driver's store writer
flushes on its own schedule, so the record can be committed before its
firing's step.finished; such a record is held back, with every platform
record after it, until the finish is in the stream, or the run finished.
The fold now remembers which firings finished. Unit tests cover the
placement, the earlier-pass case, an unpositioned record, the hold and
its release; the CLI read-back scenario is stable over eight runs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 02:15:53 -04:00
Bryan Helmkamp
5d93f20fb0
Merge branch 'petri-integration-tools' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/apps/fabro-cli/src/commands/run/petri_worker.rs
#	lib/apps/fabro-cli/tests/it/scenario/petri.rs
#	lib/apps/fabro-server/src/server/petri_runs.rs
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/src/lib.rs
2026-09-18 02:02:51 -04:00
Bryan Helmkamp
49e70e481d
Register Fabro's run tools on a Petri run through the host tool capability
Plan item F3.4. `fabro_petri::host_tools` adapts Petri's `HostTools`
capability to `register_fabro_run_tools`: every native agent session of a
run gets the tools the legacy worker registers, bound to the worker's
client and the run id, so a child run a stage creates is parented to the
Petri run. The tools run under the run's tool hooks, are recorded under
the stage, and reach sub-agents through Pebble's inheritance.

`RuntimeSpec::run_tools` installs the capability; the worker sets it when
the run's settings enable `[run.agent] fabro_tools` and the worker token
carries `agent:run_tools`, the legacy worker's gate. The server's
in-process test path runs without them, like the legacy one.

The identity the tools need is the run id alone; no run tool records a
stage on an effect, so nothing derives Fabro's `node@visit` label. A
context for another run gets no tools.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 01:52:46 -04:00
Bryan Helmkamp
26429a7444
Merge branch 'petri-integration-api' into petri-integration
# Conflicts:
#	lib/apps/fabro-cli/tests/it/scenario/petri.rs
2026-09-18 01:39:53 -04:00
Bryan Helmkamp
bfc2abebab
Wait for the terminal lifecycle record before reading a stream's end
Fabro's terminal `run.lifecycle` record lands a moment after Petri's
`run.finished`: the worker exits, the server records the status, the
projector folds it. A CLI scenario that asserts on the end of the stream
now waits for that record instead of reading the stream as soon as the
runs row turns `succeeded`, which the projector writes from the engine's
finish alone.

The fabro-petri README names the projector's stream reader and commit
signal, the server's reconnect test with its fixture capture, and the
CLI scenarios that read a run back through the stream.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 01:25:40 -04:00
Bryan Helmkamp
c9c4c27753
Merge branch 'petri-integration-hooks' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/apps/fabro-cli/src/commands/run/petri_worker.rs
#	lib/apps/fabro-cli/tests/it/scenario/petri.rs
#	lib/apps/fabro-server/src/server/petri_runs.rs
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/src/engine.rs
#	lib/components/fabro-petri/src/lib.rs
#	lib/components/fabro-store/src/platform_records.rs
2026-09-18 00:51:21 -04:00
Bryan Helmkamp
523f831a7a
Serve a Petri run's events as one stream with a stream_seq cursor
`GET /runs/{id}/events` and `GET /runs/{id}/attach` serve a Petri run's
public events and Fabro's platform records as one ordered stream in a
Fabro envelope (`RunStreamItem`: `run_id`, `stream_seq`, `kind`, `id`,
`recorded_at`, `item`), read from the projector's `petri_stream` table.
The cursor is `stream_seq` (`?after=`); the item's own identity (the
Petri `EventId` as `<log>/<seq>/<index>`, or the platform record's seq)
travels beside it for deduplication. A legacy run keeps its envelope on
the same endpoints; the OpenAPI response is the union of the two lists,
and the stream list reports Petri's `EVENT_CONTRACT_VERSION`.

The attached stream follows the projector's commit signal (a wake-up,
with a poll as the fallback) and ends after the platform record of the
run's terminal lifecycle transition, the analog of the legacy stream's
`run.completed`, or a bounded grace after the projection went terminal.

`RunSpec.engine` (`RunEngine`, `PetriAdmission`, `PetriGraphRef`) is
named in the spec and reuses the Rust types. `fabro-client` matches the
union and adds `list_run_stream`, `list_run_stream_page` and
`attach_run_stream`.

A server test attaches to a two-branch parallel run, disconnects once
both branches started, records a platform notice while both branch
scripts run, reconnects from the last `stream_seq`, and checks the
union is the whole stream: every item once, in order, no gap, no
duplicate, the notice between the branch events, and the same as the
paged listing. The Petri scenarios capture their settled projection and
stream as JSON fixtures for the web app under
`FABRO_CAPTURE_PETRI_FIXTURES`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:47:26 -04:00
Bryan Helmkamp
34139b1936
Prove the checkpoint hooks and the recovery protocol
In-process tests over the memory store: every finish is committed on the
run branch with its identity trailers and recorded with its commit, the
run-end hooks reach Petri's local service through Fabro's wrapper, a
stage that fails on its own terms is committed and its failure route runs
on the committed files, a failed checkpoint records `checkpoint_failed`
with no route taken and a restart reports the run failed, and a
`[[run.hooks]]` hook blocks an agent's tool call through the forwarded
service, with the model told why.

Real-binary scenarios crash the server and its worker with SIGKILL: after
a durable finish the stage's commit is not repeated and the interrupted
stage reruns on its snapshot; a crash held before the commit reruns the
stage once; a crash held after the commit but before its record
reconciles the record from the snapshot repository; a deleted workspace
is restored; a failure route sees the same committed files after a
crash; a failed checkpoint fails the run and a restart leaves it failed.

Recovery selects the executions `inspect_run` reports incomplete, and a
run whose coordinator log is still empty is left to the worker's resume.
The worker's platform record endpoints get an API test and the generated
TypeScript client.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:25:10 -04:00
Bryan Helmkamp
a6ac3f120e
Give a Petri question one identity across the adapter and the projection
The interview adapter derived its own question id from Petri's identity
and posted it on `interview.started`, while the projection over Petri's
records serves the pending question under Petri's `Question.id` with the
firing's stage label. The answer endpoint validates against the
projection, so an answer under the projection's id never reached the
adapter's wait.

The adapter now waits under Petri's id and labels the question's stage
through the projection's own rule: `stage_label`, `is_shown` and
`visit_of` move out of `start_visit` into shared functions, and the
adapter's observer derives each firing's `visit.started` through Petri's
`Projection`, as the projector does, so the label matches by
construction. The full Petri identity stays on `AskedQuestion`.

The legacy `interview.*` events are still posted, under Petri's id, for
the readers that follow the event stream rather than the projection: the
Slack service, `run attach`, the web app's Q&A renderer and the server's
answer claim. The store already derives the `interview.answered` platform
record from `interview.completed` for a Petri run, so who answered is
recorded under Petri's id with the answering principal.

The gate scenarios assert the new identity and encode the id as one path
segment, as the generated clients do. Projection tests cover an expired
question and an auto-approved answer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:19:30 -04:00
Bryan Helmkamp
01beea0a6c
Checkpoint a Petri run's stages and recover its workspaces on restart
Fabro's hooks on a Petri run wrap the hooks the runtime installed for
`[[run.hooks]]` and forward every point. In `prepare_result`, before the
finish is recorded, they commit the stage's files on the run branch of
its host workspace with Fabro's author identity and the run, execution,
firing and attempt as trailers, and publish the commit to a snapshot
repository beside the run's workspaces under a ref per checkpoint. A
stage that failed on its own terms is committed like a successful one; a
commit that fails is fatal: the outcome becomes a `checkpoint_failed`
failure, the run is cancelled through the coordinator handle, and the
transition refuses the firing's routes. In `transition` they write the
platform checkpoint record, keyed on the Petri position and the
checkpoint's operation identity, and a failed write is a recorded
problem.

On restart the server runs the recovery protocol before it relaunches a
worker: a run with a failed checkpoint is reported failed; otherwise
every live execution's last durable finish names the snapshot its
workspace is verified against, reset to, or restored from, with a lost
record reconciled from the snapshot repository, and a finish with no
snapshot fails the run rather than resume it on stale files.

The worker reaches the platform records over two new worker-scoped
endpoints; the server reaches the table directly. A test gate directory
lets the CLI scenarios hold a checkpoint at a named point.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:12:14 -04:00
Bryan Helmkamp
5dae891a98
Merge branch 'petri-integration-read' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/README.md
#	lib/components/fabro-petri/src/lib.rs
2026-09-17 23:48:36 -04:00
Bryan Helmkamp
ee4850689e
Name the per-run pass lock's type through an import
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:58:31 -04:00
Bryan Helmkamp
4a20acafc7
Run one projection pass at a time per run
The startup pass called the pass directly while a signalled pass could
run for the same run, so both read one committed stream sequence and
the second insert into the stream failed on its primary key, which
stopped the restarted server. Passes now take a per-run lock, and a run
whose startup pass fails is logged and left for its next signal instead
of stopping the server. A test races four passes, a signal and the
startup pass over one run and checks the stream stays contiguous.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:56:30 -04:00
Bryan Helmkamp
e0b546d465
Read the view tables where the run summary store keeps them
The projector takes two pools: the one Petri's records live in and the
one the view tables live in. In the server both are the one database;
a test fixture keeps the runs row, the platform records and the
projection tables in the run summary store's own pool, which the
projector was not reading, so a run projected in a test server folded
its Petri events before its run.created record. The startup run-history
verification checks only a Petri run's identity and legacy guard, since
its row is the projector's. An agent stage's response is the
response.<node> its outcome wrote into the run context, as the prompt
step writes it. The scenario tests assert each branch's own index.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:49:15 -04:00
Bryan Helmkamp
53b16b591d
Merge branch 'petri-integration-adapters' into petri-integration 2026-09-17 22:40:30 -04:00
Bryan Helmkamp
e2bf05c0f0
Project a Petri run's records into Fabro's run view
The projection folds Petri's public events (replay_since over the run's
stored records) and Fabro's platform records into the RunProjection the
API serves, row by row as VIEWS.md maps them. The stage key is the
execution and firing; the StageId label is node@visit, made unique with
the execution when two child invocations would share one. A stage's
first_event_seq is the milliseconds from the run's creation to its
visit.started, so the view built live equals the view rebuilt from the
records whatever order two logs' records were committed in.

The projector is the view pass and its wake-up. Records first: an append
returns before any view work; a pass reads what is committed, folds the
items past the committed positions, and writes the projection document,
the ordered stream (one stream_seq per Petri event or platform record,
with the item's own identity beside it) and the narrowed runs row in one
later transaction. Signals coalesce per run, a lost signal costs only
latency, the startup pass folds every run the view trails, and a pass
that races a platform record leaves the view alone and runs again. A
torn tail holds the view where it stands and reports the run incomplete
with the replay's error; inspect_run decides completeness once the run
recorded its finish.

The tests build the view live for the hello bundle, a command workflow
and a two-branch parallel workflow and compare it with the rebuild; drop
every wake-up and catch up by a signal and by the startup pass; crash
between the record commit and the view transaction and apply only the
suffix; restart the projector over child executions; and hold at a torn
tail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:32:23 -04:00
Bryan Helmkamp
162791979c
Wake the projector after a Petri run's first event too
The run's creation commits its first event on the create path, not the
append path, so the platform record hook never fired for run.created.
The store now notifies after that commit as well, and a test proves a
Petri run's lifecycle events leave platform records beside them with one
wake-up per record while a legacy run leaves none.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:32:23 -04:00
Bryan Helmkamp
3aadc21739
Test the interview, secret, blob and home adapters through the engine
Integration tests in `fabro-petri` run workflows through `engine::run` on
the host sandbox: a gate answered under the posted question id, two
parallel gates each bound to their own answer, an expired question
completed as a timeout with the gate's default, an auto-approved run, a
cancelled run; a secret resolved from a vault into a command and masked in
every `petri_records` row; a command's large output round-tripped through
the `blobs` table under `blob://sha256/<hex>`; and the `hello` bundle on
the OpenAI twin with a model client over a vault that holds the key,
whose skills step searched the configured home.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:31:35 -04:00
Bryan Helmkamp
223e10ea20
Install Petri's interview, secret, blob and home adapters in a Fabro run
A Petri run in the worker, and in the server under its test override, now
gets Fabro's platform adapters instead of the standalone defaults:

- `fabro_petri::interview`: Petri's `Interviewer` over the questions API
  and the worker's control channel. A human gate's question is posted as
  the `interview.started` event a legacy stage emits, keyed by an id
  derived from Petri's identity (node, execution, firing, occurrence,
  ask), so the API, the web app and Slack list it; the answer posted to
  the questions endpoint reaches the control interviewer the adapter waits
  on and is mapped onto Petri's answer. An expiry the gate reports is
  completed as `interview.timeout`, a cancel as `interview.interrupted`,
  and an auto-approved run answers itself. The hook points the read side
  takes over are marked.
- `fabro_petri::secrets`: Petri's `SecretProvider` over the vault's token
  entries, so `{{ secrets.NAME }}` resolves at spawn and is masked in every
  record; a sensitive answer registers as a dynamic secret.
- `fabro_petri::blobs`: Petri's `OutputStore` over Fabro's `blobs` table,
  through the server's blob store or the worker's client.
- The Fabro home the server resolved travels to the worker as
  `--fabro-home`, so the skills step reads it whatever the worker's
  environment says.

`engine::RunRequest` takes the interviewer, its observers, the secret
provider and the blob table from the caller; `interviewer::Unattended` is
gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 22:31:35 -04:00
Bryan Helmkamp
abbc7ca11d
Add platform records and the Petri projection tables
A Petri run's own Fabro facts (its lifecycle before and after the
engine, a checkpoint commit, a pull request, a notification, a pairing)
are platform records in a table beside Petri's records, one typed enum
of kinds tagged on the wire, each keyed to a Petri stage where it
belongs to one and carrying the operation identity of the effect it
records. The run summary store derives the lifecycle kinds from the
legacy run events a Petri run still appends, in the event's
transaction, and calls a hook after the commit so the run's projector
can wake up.

Two more tables serve the projection that follows: the per-run
projection document with its committed positions, and the ordered
stream of everything the view consumed. The run summary store reads the
Petri projection back for the API and writes the narrowed runs row from
it without touching the legacy concurrency guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:46:21 -04:00
Bryan Helmkamp
a745d0b75d
Check a workflow version in memory with Runtime::check_source
`fabro_petri::check` materialized the version's bundle into a temporary
directory because `Runtime::check` read the workflow and its settings
files from disk. Petri now has `Runtime::check_source`, which takes the
workflow's repository-relative path, its text and a `FileSource`, so the
bundle goes into a `frontend::MapFiles` map instead: every file at its
bundle-relative path, `workflow.toml` beside the workflow, and
`.fabro/project.toml` at the root when the caller has one. Nothing is
written to disk, and the diagnostics name the bundle-relative paths
directly, with no root to strip.

The compile inputs are unchanged: the intent's inputs, the launch model
and provider, and `petri.repository` bound by Fabro itself (the launch's
path, or `null`). An entrypoint that is not one of the bundle's files is
now `CheckError::MissingEntrypoint`; `CheckError::Materialize` goes away.
`tempfile` becomes a dev-dependency, as only the tests use it.

New tests: a version whose `workflow.toml` names `engine = "petri"` is
admitted (the new Petri pin knows the key), an unknown `[workflow]` key
is refused with `unsupported.workflow_toml.key` and named in
`workflow.toml`, the project settings are read from the map, and a
missing entrypoint is an error.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:31:04 -04:00
Bryan Helmkamp
a621fb72e1
Let the Petri engine start or resume a run over any store
`fabro_petri::engine` is now the one assembly the worker process and the
server share: `RunRequest` takes the run's store as `Arc<dyn RunStore>` and
an `Execution`, either `Start` with the admitted graphs or `Resume` from the
run's records through `host::resume_configured`, with the same interview
observer a start installs. A resume whose record has no root invocation is
refused with a named error instead of a panic in the host. The outcome is
mapped to a `Conclusion` (succeeded, or failed with Fabro's reason and a
message) so both callers record the same terminal event.

`admission::load_with` loads the admitted graphs through any blob read, so
a worker loads them through its client; `admission::load` over the server's
`BlobStore` delegates to it.

`HttpRunStore::for_worker` takes every lease for the worker's launch id,
whatever owner Petri minted for the run runtime, and the open logs the
owner. The module docs state the rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 21:20:22 -04:00
Bryan Helmkamp
832f39f704
Merge branch 'petri-integration-http' into petri-integration
# Conflicts:
#	Cargo.lock
#	lib/components/fabro-petri/Cargo.toml
#	lib/components/fabro-petri/README.md
#	lib/components/fabro-petri/src/lib.rs
2026-09-17 20:37:03 -04:00
Bryan Helmkamp
d2f0e70dca
Run a workflow through Petri in the server, behind the engine flag
When a run's engine is Petri, the create handler hands the bundle, inputs
and launch to Petri's check instead of the legacy compile, lint and model
pinning, refuses the run with the validation error the legacy validator
uses (Petri's codes as the rules, listed in the API detail), and records
the admission on the run spec. The Fabro graph the read side displays is
parsed without validation. The scheduler executes a Petri run in the server
process through fabro_petri::engine, appending only the run lifecycle
events the read side needs (run.starting, run.running, run.completed or
run.failed); no stage or agent event is projected yet.

Scenario tests run the hello bundle on the OpenAI twin under the version
flag and a command-only bundle under the server setting, check Petri's
record agrees, and cover the refusals for an unknown attribute, an
undeclared node and an unknown model.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:27:44 -04:00
Bryan Helmkamp
03309d4240
Let Petri compile and execute a Fabro run through fabro-petri
`check` materializes a workflow version's bundle into a temporary directory
(`Runtime::check` reads files from disk), lowers it with the run's inputs
and launch, and returns the admitted graphs or Petri's diagnostics in a
shape the server maps onto Fabro's. `admission` keeps the admitted graphs
in the blob store, named on the run spec and verified by digest on load.
`runtime` assembles the same Petri runtime at create and at execution: the
Fabro frontend with the server's settings layer, the Attractor step kinds,
the model client as the PebbleClient capability so admission pins every
model. `engine` runs the admitted graph in the server process over
SqliteRunStore under the Fabro run id, with the standalone defaults, an
interviewer that fails any question, and cancel on a token, and derives the
outcome from inspect_run over the run's record.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:27:44 -04:00
Bryan Helmkamp
25d47ebcd3
Implement Petri's run store over the server's API
`HttpRunStore` is Petri's `RunStore` and `RunLogs` as a worker process
reaches them: over `fabro_client::Client` with the worker's token,
against the server's SQLite store. A run key is a Fabro run id, the
`{id}` of every request, which is the plan's rule that Petri's run key
is Fabro's run id.

The lease rules are the store's. A same-owner reopen shares the live
handle in the process, and the server makes a same-owner reopen after
a lost reply the same lease. Dropping the last handle of an owner sends
`release` on the current Tokio runtime, and the store awaits every such
release before its next open, so a drop followed by an open observes
it. The server's worker-exit release is the backstop.

A reply that never arrives, a transport error or the client's request
timeout, is retried by resending the same request up to three times.
Every request is idempotent on the server, so that is safe; a reply
that did arrive is never retried. Each server error code maps back to
its `StoreError` variant, with the leased owner and the conflict
position read from `meta`.

`fabro_petri::petri` re-exports the store vocabulary for the server,
and the `test-support` feature re-exports Petri's test kit so the
server's tests can run the conformance suite over the wire. Both keep
this crate the one place that names a Petri package.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:12:20 -04:00
Bryan Helmkamp
c383b6a70b
Add the engine flag and record the engine on the run spec
A workflow version names its engine with `engine = "petri"` in the
`[workflow]` table of `workflow.toml`, and `[server.execution] engine`
(`FABRO_SERVER_ENGINE`, `--engine`) defaults it for every version that
names none. The choice, with what Petri admitted (the lowered root graph
and its children by blob and digest), is recorded on the run spec as
`RunEngine`, carried on `run.created`, and replayed into the projection.
A legacy run's spec omits the field, so existing specs decode unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 20:02:47 -04:00
Bryan Helmkamp
55a145f5a4
Add the Fabro-on-Petri view coverage matrix
Plan item F2.1: every Fabro view of a run, the Petri event or platform
record that supplies each fact, and the identity it is keyed on. Ends
with the two completeness checks (every EVENTS.md family, every Fabro
view) and the gaps table.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 19:38:25 -04:00
Bryan Helmkamp
3ffe7e00ce
Implement Petri's run store over Fabro's SQLite database
`SqliteRunStore` implements Petri's `RunStore` and `RunLogs` on the pool
Fabro's other stores share. A run's existence and writer lease live in
`petri_runs`; every record of every log lives in `petri_records`, keyed by
(run, log, seq) with the record stored as JSON and read back unchanged;
blobs share the `blobs` table with `BlobStore`. The lease is taken
idempotently per owner, ends when the last handle drops or when an
operator releases it, and never by timeout; every write checks it inside
its own transaction. An append is one `BEGIN IMMEDIATE` transaction per
batch: a repeated record is accepted, a different record at a taken seq or
a seq past the head is a conflict that stores nothing.

Petri's store conformance suite passes against it, with the operator
release, lease exclusivity, a crash between appends, and blob
interoperation checked beside it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 19:30:30 -04:00
Bryan Helmkamp
7eb5ca502c
Add the fabro-petri crate and pin the Petri packages
Fabro runs its workflows on Petri. The six Petri packages and the testkit
are pinned by revision in the workspace manifest under `petri_*` keys, and
`fabro-petri` is the one crate that depends on them. The crate's tests run
the `hello` bundle in memory on the stub registry and a command-only
workflow on the host sandbox; both skip without the sandbox-driver host
plugin, and the sandbox-plugins CI job requires it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 19:23:30 -04:00
Bryan Helmkamp
27f16f89c4
Delete fabro's own pricing: lithos-llm prices every response once
lithos-llm attaches a cost to every response at the client: the codec
keeps a provider-reported cost when the provider supplies one, and the
resolver fills the catalog's price for the route when it does not.
Pebble records that priced usage on every assistant turn and sums it,
so each AssistantMessage on the stream, and the store fold's live stage
usage, already carries the cost. Fabro's catalog re-pricing of the same
tokens was redundant, and is gone.

model_usage_from_llm, with_reported_cost, and every estimate_cost call
in fabro are deleted. The pebble handler's stage_usage groups pebble's
accounts by route and sums them with Usage::saturating_add, keeping the
cost and source pebble carried, so the terminal stage.completed usage is
the live fold's sum; it no longer fails when the catalog does not know a
provider. A one-shot prompt stage records the response's own usage and
cost as lithos-llm returned it. The per-model price cards in fabro-llm's
API module stay.

Tests: the pebble handler sums Catalog and Provider costs per row and
leaves a row's and the total's cost unknown once an answer was unpriced;
the store fold shows the same tokens and cost live and at completion,
and None at both for an unpriced answer; the agent integration test
compares the whole completed Usage with the live fold, cost included;
a one-shot prompt stage on a mocked OpenAI-compatible provider reports a
Catalog cost that lithos-llm's resolver attached, with no fabro pricing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:46:44 -06:00
Bryan Helmkamp
7d5e696c18
Merge pull request #873 from fabro-sh/sandbox-driver-64c14b8
Pin sandbox-driver main 64c14b8 and surface dropped Daytona output
2026-09-14 16:31:07 -04:00
Bryan Helmkamp
c98785d84a
Pin pebble a39f43e and refuse an older stored session record clearly
Pebble main a39f43e26effdf99635eaf343f095c17157c9c93 (pebble #22) carries
an assistant turn's usage as Usage in the session record and moves the
record format to version 5. CodingRuntime::from_record refuses a record
in another format with UnsupportedRecord { version, supported } before
it reads the route. Fabro persists those records in SQLite for Ask Fabro
resume, and old runs get no migration, so a record written by an older
build is read back as stored and refused on the next turn.

Two tests pin that down. The store reads pebble's own version 4 fixture
back through get without a parse error and reports it unsupported. A
resumed Ask Fabro session whose stored record declares the previous
format fails its next turn with the agent_error code and the message
"session record format version 4 is not supported (this build requires
5)", runs no turn, and leaves the stored record in place.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 12:58:31 -06:00