Petri now ends a run's lifetime at its first failed store write: it
records nothing after it, fails no firing for it, and returns
CoordinatorError::StoreFailed. The run is not over; the next lifetime
resumes it from what the store holds. Fabro read that error as an
unfinished run and failed it.
- engine: RunError::StoreFailed, returned without reading the record
back, and Conclusion::Interrupted for it.
- worker: an interrupted run gets no terminal lifecycle record; the
worker exits with EX_TEMPFAIL (75, the new ExitClass::Interrupted).
- server: WorkerExit carries the exit code. An interrupted worker's run
goes back to the scheduler in resume mode through the relaunch a
restart takes (lease release, recovery, start_requested + runnable),
now shared with reconcile_on_startup. The in-process path does the
same. A run is resumed at most MAX_STORE_INTERRUPTIONS (3) times per
server; the next interruption fails it. A pending cancel, a run that
ended or was deleted, and a shutdown also end it as before.
Tests: an engine run over a store whose first lease write fails is
interrupted with no finish, and a resume finishes it; the server
relaunches an interrupted worker in resume mode, fails the run after the
bound, and fails a worker that exits 1 as before; exit code 75.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Petri now refuses to resume a run whose creation a crash cut short (the
key is stored, the root invocation is not) with HostError::NotStarted,
and starts it again when the host runs it under the same key. The engine
used its own guard, check_resumable, which failed the run with
NothingToResume instead.
Execution::Resume now carries the admitted graphs, and a resume Petri
answers with NotStarted starts the run from them. The worker loads the
graphs in resume mode too, as does the server's in-process path. The
guard and RunError::NothingToResume are gone. New test:
a_resume_of_a_run_that_never_started_starts_it_again.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clears RUSTSEC-2026-0185 (fixed in 0.11.15). quinn is only in the lock
through reqwest's optional http3 feature, which Fabro doesn't enable, so
this only stops lockfile scanners from flagging it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rewrite the sandbox credential store only when the token source mints a
new generation, and skip the refresh loop for static tokens. Build the
run's read-only token source once in the worker and share it between the
workspace fetch and stage Git access. Drop unreachable branches in
for_run, reuse the shared contents-permission check and constants, and
fold the duplicated token resolution and test setup into helpers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pin Petri at e46845b, the merge of its executor layer change. The stage
credential layer is now a Petri `SpawnEnv` applied with
`EnvHandle::with_spawn_env`, so Petri forwards every other environment
method and applies the layer to one-shot containers as well as processes.
A container gets the managed GITHUB_TOKEN, but no credential store refresh
or Git helper configuration: the store lives in the scope's sandbox, which
the container does not share.
The managed GITHUB_TOKEN now replaces one set in the workflow environment,
an ACP agent's environment or the sandbox's own, as Fabro's stage
environment did before Petri. The token carries exactly the access the run
declares; a stage that needs other access changes its declaration.
The new pin also keeps a timeout as the failure a partial success came
from: `PartialSuccess.underlying` is now an `UnderlyingFailure`, so the
projection reports "the step timed out" for a partial success converted
from a timeout. The run format moves from 7 to 8, which the attach JSON
snapshot records; runs stored before this pin are refused, as with earlier
format changes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Creating a pull request for a finished run read the projection's final
patch field directly. Petri stores that patch in the blob table and the
field holds a blob reference, so the description model was handed the
reference string instead of the diff, and the empty-diff check could
never fire.
Move the reference resolution out of Files Changed into a shared
final_patch::load helper and use it for both readers. Pull request input
extraction now reads the real patch, judges emptiness by its text, and
reports a missing or unreadable blob instead of describing a placeholder.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The create, validate and preflight paths each wrapped their launch with
the run's --model and --provider flags before building the check
request, so a new caller could build a check without them. The check
request now takes the flags as a required argument and binds them onto
the launch itself, and unit tests cover the flags, a provider-only flag,
and no flags.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Petri now treats `petri.launch_model` and `petri.launch_provider` as the
model a run's flags ask for, above the file layers and the graph's
defaults. A host's last-resort default moved to `petri.default_model` and
`petri.default_provider`. Pin Petri at the merge of that change and bind
to it: the explicit `--model`/`--provider` flags go to the launch
variables, and the model the settings resolved (or the catalog default)
goes to the default variables.
`Launch` now names the two pairs `model`/`provider` and
`default_model`/`default_provider`, matching Petri, in place of the
separate override fields.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The host provider manages no networking and refuses any policy but its
default, so sending the resolved AllowAll to local runs failed every
non-dry-run local execution. Apply the run's policy only on container
backends; dry runs, which always use the host backend, are covered by the
same check.
Also fold the Docker environment test helper into one that takes a typed
network mode, share the probe setup between the live Docker and Daytona
network tests, count canary hits per mode, and bind the run environment
once in the worker.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Retry now starts the workflow over, and a local-folder run commits its
checkpoints, so the retry scenario checks that every stage commits again
under the retry's run id. The scenario for retrying without Git
checkpoints goes: local runs have them now, and the retry scenario
covers starting over.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Checkpoints were only enabled for runs with a GitHub source, so local
folder runs, empty Local runs and dry runs stopped committing. `fabro
diff` then failed for them, and their checkpoint, run branch and diff
records disappeared from the event stream.
A run whose workspace is on the host now commits checkpoints there
again, without pushing, as on main. Docker and Daytona runs with no
GitHub source still record execution checkpoints without Git commits,
so a sandbox image without `git` cannot fail the run.
The scenario tests for crash recovery go back to asserting commits. A
workspace deleted while the run is down now fails the resumed run,
since the server keeps no copy to restore it from.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The worker minted one read token at launch and a fresh push token for
every checkpoint. The read token expired an hour into a run, so a
workspace acquired later fetched with a dead credential. Minting per
push put every push in GitHub's token-replication window, where a token
minted moments earlier is rejected with 404 "Repository not found".
The worker now keeps two InstallationTokenSource caches for the run, a
read-only one for fetches and a contents: write one for pushes. Each
fetch and push resolves through its source, which reuses one token until
it nears expiry and then mints the next. Petri's RunSource asks a
SourceCredentials provider on every fetch instead of holding a fixed
credential.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Share one credential header helper between the in-sandbox fetch and
the run branch push
- Keep only the target branch, goal, and model on the publisher instead
of a full run spec copy
- Load the worker's LLM catalog once, and mint the read token only when
the run checks something out
- Pass the source explicitly to checkpoint fetch helpers, dropping
unreachable branches, and reuse has_object in has_commit
- Move the run patch into the publication instead of cloning it, and
build it only when a publisher exists
- Add test fixture helpers for file sources and recording publishers
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pushing the run branch and opening the pull request now happen in the
run's worker, in Fabro's run_finished hook, after the last stage and
before the run's terminal record, as the legacy publish step did. A
failed push or pull request fails the run with publish_failed instead of
leaving a warning on a run that already succeeded.
fabro-petri gains a RunPublisher the hooks call for a successful run with
its run branch, final commit, snapshot repository and patch; the worker's
GitHub publisher pushes from the snapshot repository with a push token it
mints at that moment, opens the pull request its settings ask for, and
records it. The worker resolves the server's GitHub credentials itself
for both the read-only checkout token and the push token, so the server
no longer hands it a clone credential or publishes after the run.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Since the Petri cutover, a run with a GitHub target started in an empty
workspace, nothing pushed its run branch to GitHub, and nothing asked for
the automatic pull request when it succeeded.
Fabro's hooks now check a fresh run's GitHub target out inside the
sandbox when Petri hands them the scope, before the first stage: the
workspace fetches the selected commit, tag or branch at the run's clone
depth, with a read-only token the server resolves at each worker launch.
The worker scrubs the token from its environment at startup and presents
it only to the fetch, so it never lands in the repository or its remote.
The files belong to the sandbox user, so git accepts them.
The same checkout seeds the workspace's snapshot repository with the
starting commit. Checkpoint bundles from a shallow clone then import, a
stage's own commits never make a bundle carry the source's history, a
restore into a fresh sandbox fetches the base again and applies the run's
commits, and a fork carries the base with its checkpoints.
When a successful GitHub-target run ends, the server pushes its final
commit from the snapshot repository to fabro/run/<id> with its own write
credentials, then, when the run changed files and asks for one, records
the pull request request for the existing creation supervisor. A failed
push or request is a warning notice on the run. The manual pull request
endpoint shares the request step.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Retry forked the source run at its last checkpoint and reran the failed
stage. It now creates a new run from the source's saved spec and starts
the workflow from the beginning in a fresh workspace, as retry did
before the Petri cutover. The new run records `retried_from` and no
`fork_source_ref`.
Retry no longer needs a checkpoint, a retained workspace, or a published
run branch, so it works for any terminal run that is not archived. To
continue from where a run stopped, fork it at a checkpoint.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>