The server is generally available, so drop the private early-access
warning from the deployment, self-host, Railway, and server operations
pages.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- stop_lock_holder returns the stopped pid as Option<u32> instead of a
LockHolder enum that only wrapped it
- stop_previous_worker uses with_context, and one run_scratch helper
replaces three spellings of the run's scratch path
- relaunch builds its runnable record through run_records::runnable,
moved out of the lifecycle handler so both callers share it
- WorkerExit derives success from its exit code instead of storing both
- the scenario tests share one worker-pid lookup and wait loop
- small readability fixes in the engine's resume arm and the resume test
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A worker leads a process group of its own, so it outlives a server
crash. The restarted server released the run's lease from outside and
launched a resume while that worker could still be running: a worker
whose lease is released keeps acting until its next write, beside its
successor. A delete likewise dropped the lease without knowing the
worker was gone.
Now each worker holds a lock on `worker.lock` in its run's scratch
directory for its whole life, taken before anything else. It is a POSIX
record lock: the kernel frees it only when the worker exits, and names
the process that holds it. Before the server ends a lease from outside
(the relaunch after a restart or a store interruption, and a delete), it
kills whatever process still holds the lock, with its process group, and
waits until the lock is free. A worker that finds the lock held does not
start.
- fabro-proc: ProcessLock::try_hold and stop_lock_holder, tested with
this test binary as the holding process.
- fabro-config: RunScratch::worker_lock_path.
- New scenario: a worker that outlives the server is gone before the
resume's worker launches, and the run succeeds once. It fails without
the server-side stop.
A host stage process runs in a process group of its own and still
outlives its killed worker, as it did at a worker crash.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Petri now ends a run's lifetime at its first failed store write: it
records nothing after it, fails no firing for it, and returns
CoordinatorError::StoreFailed. The run is not over; the next lifetime
resumes it from what the store holds. Fabro read that error as an
unfinished run and failed it.
- engine: RunError::StoreFailed, returned without reading the record
back, and Conclusion::Interrupted for it.
- worker: an interrupted run gets no terminal lifecycle record; the
worker exits with EX_TEMPFAIL (75, the new ExitClass::Interrupted).
- server: WorkerExit carries the exit code. An interrupted worker's run
goes back to the scheduler in resume mode through the relaunch a
restart takes (lease release, recovery, start_requested + runnable),
now shared with reconcile_on_startup. The in-process path does the
same. A run is resumed at most MAX_STORE_INTERRUPTIONS (3) times per
server; the next interruption fails it. A pending cancel, a run that
ended or was deleted, and a shutdown also end it as before.
Tests: an engine run over a store whose first lease write fails is
interrupted with no finish, and a resume finishes it; the server
relaunches an interrupted worker in resume mode, fails the run after the
bound, and fails a worker that exits 1 as before; exit code 75.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Petri now refuses to resume a run whose creation a crash cut short (the
key is stored, the root invocation is not) with HostError::NotStarted,
and starts it again when the host runs it under the same key. The engine
used its own guard, check_resumable, which failed the run with
NothingToResume instead.
Execution::Resume now carries the admitted graphs, and a resume Petri
answers with NotStarted starts the run from them. The worker loads the
graphs in resume mode too, as does the server's in-process path. The
guard and RunError::NothingToResume are gone. New test:
a_resume_of_a_run_that_never_started_starts_it_again.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clears RUSTSEC-2026-0185 (fixed in 0.11.15). quinn is only in the lock
through reqwest's optional http3 feature, which Fabro doesn't enable, so
this only stops lockfile scanners from flagging it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rewrite the sandbox credential store only when the token source mints a
new generation, and skip the refresh loop for static tokens. Build the
run's read-only token source once in the worker and share it between the
workspace fetch and stage Git access. Drop unreachable branches in
for_run, reuse the shared contents-permission check and constants, and
fold the duplicated token resolution and test setup into helpers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pin Petri at e46845b, the merge of its executor layer change. The stage
credential layer is now a Petri `SpawnEnv` applied with
`EnvHandle::with_spawn_env`, so Petri forwards every other environment
method and applies the layer to one-shot containers as well as processes.
A container gets the managed GITHUB_TOKEN, but no credential store refresh
or Git helper configuration: the store lives in the scope's sandbox, which
the container does not share.
The managed GITHUB_TOKEN now replaces one set in the workflow environment,
an ACP agent's environment or the sandbox's own, as Fabro's stage
environment did before Petri. The token carries exactly the access the run
declares; a stage that needs other access changes its declaration.
The new pin also keeps a timeout as the failure a partial success came
from: `PartialSuccess.underlying` is now an `UnderlyingFailure`, so the
projection reports "the step timed out" for a partial success converted
from a timeout. The run format moves from 7 to 8, which the attach JSON
snapshot records; runs stored before this pin are refused, as with earlier
format changes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Creating a pull request for a finished run read the projection's final
patch field directly. Petri stores that patch in the blob table and the
field holds a blob reference, so the description model was handed the
reference string instead of the diff, and the empty-diff check could
never fire.
Move the reference resolution out of Files Changed into a shared
final_patch::load helper and use it for both readers. Pull request input
extraction now reads the real patch, judges emptiness by its text, and
reports a missing or unreadable blob instead of describing a placeholder.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The create, validate and preflight paths each wrapped their launch with
the run's --model and --provider flags before building the check
request, so a new caller could build a check without them. The check
request now takes the flags as a required argument and binds them onto
the launch itself, and unit tests cover the flags, a provider-only flag,
and no flags.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Petri now treats `petri.launch_model` and `petri.launch_provider` as the
model a run's flags ask for, above the file layers and the graph's
defaults. A host's last-resort default moved to `petri.default_model` and
`petri.default_provider`. Pin Petri at the merge of that change and bind
to it: the explicit `--model`/`--provider` flags go to the launch
variables, and the model the settings resolved (or the catalog default)
goes to the default variables.
`Launch` now names the two pairs `model`/`provider` and
`default_model`/`default_provider`, matching Petri, in place of the
separate override fields.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The host provider manages no networking and refuses any policy but its
default, so sending the resolved AllowAll to local runs failed every
non-dry-run local execution. Apply the run's policy only on container
backends; dry runs, which always use the host backend, are covered by the
same check.
Also fold the Docker environment test helper into one that takes a typed
network mode, share the probe setup between the live Docker and Daytona
network tests, count canary hits per mode, and bind the run environment
once in the worker.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>