fabro/lib
Bryan Helmkamp 34139b1936
Prove the checkpoint hooks and the recovery protocol
In-process tests over the memory store: every finish is committed on the
run branch with its identity trailers and recorded with its commit, the
run-end hooks reach Petri's local service through Fabro's wrapper, a
stage that fails on its own terms is committed and its failure route runs
on the committed files, a failed checkpoint records `checkpoint_failed`
with no route taken and a restart reports the run failed, and a
`[[run.hooks]]` hook blocks an agent's tool call through the forwarded
service, with the model told why.

Real-binary scenarios crash the server and its worker with SIGKILL: after
a durable finish the stage's commit is not repeated and the interrupted
stage reruns on its snapshot; a crash held before the commit reruns the
stage once; a crash held after the commit but before its record
reconciles the record from the snapshot repository; a deleted workspace
is restored; a failure route sees the same committed files after a
crash; a failed checkpoint fails the run and a restart leaves it failed.

Recovery selects the executions `inspect_run` reports incomplete, and a
run whose coordinator log is still empty is left to the worker's resume.
The worker's platform record endpoints get an API test and the generated
TypeScript client.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 00:25:10 -04:00
..
apps Prove the checkpoint hooks and the recovery protocol 2026-09-18 00:25:10 -04:00
components Prove the checkpoint hooks and the recovery protocol 2026-09-18 00:25:10 -04:00
foundation Checkpoint a Petri run's stages and recover its workspaces on restart 2026-09-18 00:12:14 -04:00
packages/fabro-api-client Prove the checkpoint hooks and the recovery protocol 2026-09-18 00:25:10 -04:00