Commit graph

5444 commits

Author SHA1 Message Date
Scott Werner
c5f0296e1a
Merge pull request #905 from fabro-sh/petri-no-jump
Drop the jump route kind from the Petri stream reader
2026-10-07 10:50:46 -04:00
fabro-releases[bot]
f7ca525b7d Bump version to 0.379.0-nightly.0 2026-10-07 09:40:07 +00:00
Scott Werner
eac82ee71c
Merge pull request #907 from fabro-sh/petri-store-faults
Some checks are pending
Rust / Test (macOS) (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
Handle Petri store faults: start again, resume after a failed write, stop surviving workers
2026-10-06 16:43:06 -04:00
Scott Werner
74cbdc48df
Merge pull request #931 from fabro-sh/docs/remove-early-access-banner
Remove early-access banner from server docs
2026-10-06 16:32:07 -04:00
Scott Werner
bc6bfeae67 Remove early-access banner from server docs
The server is generally available, so drop the private early-access
warning from the deployment, self-host, Railway, and server operations
pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 16:16:58 -04:00
Scott Werner
e889f9978d Simplify the store-fault handling
- stop_lock_holder returns the stopped pid as Option<u32> instead of a
  LockHolder enum that only wrapped it
- stop_previous_worker uses with_context, and one run_scratch helper
  replaces three spellings of the run's scratch path
- relaunch builds its runnable record through run_records::runnable,
  moved out of the lifecycle handler so both callers share it
- WorkerExit derives success from its exit code instead of storing both
- the scenario tests share one worker-pid lookup and wait loop
- small readability fixes in the engine's resume arm and the resume test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 16:07:36 -04:00
Scott Werner
f12065e6ae Adapt recovery to current resume callers and API errors 2026-10-06 12:29:34 -04:00
Bryan Helmkamp
eb39e15781 Describe store interruptions and the worker lock in the Petri README
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Bryan Helmkamp
cb54a0df90 Stop a surviving worker before ending its lease from outside
A worker leads a process group of its own, so it outlives a server
crash. The restarted server released the run's lease from outside and
launched a resume while that worker could still be running: a worker
whose lease is released keeps acting until its next write, beside its
successor. A delete likewise dropped the lease without knowing the
worker was gone.

Now each worker holds a lock on `worker.lock` in its run's scratch
directory for its whole life, taken before anything else. It is a POSIX
record lock: the kernel frees it only when the worker exits, and names
the process that holds it. Before the server ends a lease from outside
(the relaunch after a restart or a store interruption, and a delete), it
kills whatever process still holds the lock, with its process group, and
waits until the lock is free. A worker that finds the lock held does not
start.

- fabro-proc: ProcessLock::try_hold and stop_lock_holder, tested with
  this test binary as the holding process.
- fabro-config: RunScratch::worker_lock_path.
- New scenario: a worker that outlives the server is gone before the
  resume's worker launches, and the run succeeds once. It fails without
  the server-side stop.

A host stage process runs in a process group of its own and still
outlives its killed worker, as it did at a worker crash.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Bryan Helmkamp
37116758fa Resume a Petri run whose store failed, as after a crash
Petri now ends a run's lifetime at its first failed store write: it
records nothing after it, fails no firing for it, and returns
CoordinatorError::StoreFailed. The run is not over; the next lifetime
resumes it from what the store holds. Fabro read that error as an
unfinished run and failed it.

- engine: RunError::StoreFailed, returned without reading the record
  back, and Conclusion::Interrupted for it.
- worker: an interrupted run gets no terminal lifecycle record; the
  worker exits with EX_TEMPFAIL (75, the new ExitClass::Interrupted).
- server: WorkerExit carries the exit code. An interrupted worker's run
  goes back to the scheduler in resume mode through the relaunch a
  restart takes (lease release, recovery, start_requested + runnable),
  now shared with reconcile_on_startup. The in-process path does the
  same. A run is resumed at most MAX_STORE_INTERRUPTIONS (3) times per
  server; the next interruption fails it. A pending cancel, a run that
  ended or was deleted, and a shutdown also end it as before.

Tests: an engine run over a store whose first lease write fails is
interrupted with no finish, and a resume finishes it; the server
relaunches an interrupted worker in resume mode, fails the run after the
bound, and fails a worker that exits 1 as before; exit code 75.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Bryan Helmkamp
3c395f9e6e Start a Petri run again when its creation was cut short
Petri now refuses to resume a run whose creation a crash cut short (the
key is stored, the root invocation is not) with HostError::NotStarted,
and starts it again when the host runs it under the same key. The engine
used its own guard, check_resumable, which failed the run with
NothingToResume instead.

Execution::Resume now carries the admitted graphs, and a resume Petri
answers with NotStarted starts the run from them. The worker loads the
graphs in resume mode too, as does the server's in-process path. The
guard and RunError::NothingToResume are gone. New test:
a_resume_of_a_run_that_never_started_starts_it_again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Scott Werner
d89d0c577b
Merge pull request #930 from fabro-sh/codex/partial-success-regression
Restore regression coverage for partial-success failure projection
2026-10-06 12:23:03 -04:00
Scott Werner
e04ba37e25 Test the cause retained by partial-success projection 2026-10-06 12:07:28 -04:00
Scott Werner
789ca8427b
Merge pull request #926 from fabro-sh/bump-lithos-llm-fd42e6b
Bump lithos-llm and Pebble to current main
2026-10-06 10:28:28 -04:00
Scott Werner
8be46335f3
Merge pull request #929 from fabro-sh/deps/quinn-proto-0.11.18
deps: bump quinn-proto to 0.11.18
2026-10-06 10:26:15 -04:00
Scott Werner
6d3859f4ec Merge remote-tracking branch 'origin/main' into bump-lithos-llm-fd42e6b 2026-10-06 09:13:53 -04:00
Scott Werner
ddf4b4db8d deps: bump quinn-proto to 0.11.18
Clears RUSTSEC-2026-0185 (fixed in 0.11.15). quinn is only in the lock
through reqwest's optional http3 feature, which Fabro doesn't enable, so
this only stops lockfile scanners from flagging it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 09:13:32 -04:00
Scott Werner
ee4169ff11
Merge pull request #899 from fabro-sh/ci/internal-deps-nightly
ci: add a nightly internal dependency update
2026-10-06 09:02:47 -04:00
fabro-releases[bot]
64b9d88159 Bump version to 0.378.0-nightly.0 2026-10-06 09:36:59 +00:00
Scott Werner
2e4055c624
Merge pull request #924 from fabro-sh/codex/restore-stage-git-auth
Some checks are pending
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
Restore renewable GitHub credentials to workflow stages
2026-10-05 22:25:09 -04:00
Scott Werner
1f4712b845 Simplify stage credential renewal and share the read token source
Rewrite the sandbox credential store only when the token source mints a
new generation, and skip the refresh loop for static tokens. Build the
run's read-only token source once in the worker and share it between the
workspace fetch and stage Git access. Drop unreachable branches in
for_run, reuse the shared contents-permission check and constants, and
fold the duplicated token resolution and test setup into helpers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 21:58:19 -04:00
Scott Werner
9e4af50863 Move stage credentials to Petri's spawn env and let the managed token win
Pin Petri at e46845b, the merge of its executor layer change. The stage
credential layer is now a Petri `SpawnEnv` applied with
`EnvHandle::with_spawn_env`, so Petri forwards every other environment
method and applies the layer to one-shot containers as well as processes.

A container gets the managed GITHUB_TOKEN, but no credential store refresh
or Git helper configuration: the store lives in the scope's sandbox, which
the container does not share.

The managed GITHUB_TOKEN now replaces one set in the workflow environment,
an ACP agent's environment or the sandbox's own, as Fabro's stage
environment did before Petri. The token carries exactly the access the run
declares; a stage that needs other access changes its declaration.

The new pin also keeps a timeout as the failure a partial success came
from: `PartialSuccess.underlying` is now an `UnderlyingFailure`, so the
projection reports "the step timed out" for a partial success converted
from a timeout. The run format moves from 7 to 8, which the attach JSON
snapshot records; runs stored before this pin are refused, as with earlier
format changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 19:26:08 -04:00
Scott Werner
8db70e072f Merge origin/main into codex/restore-stage-git-auth 2026-10-05 19:25:39 -04:00
Bryan Helmkamp
241f3eac39
Bump Pebble to current main 2026-10-05 09:18:43 -04:00
Bryan Helmkamp
85be2b192b
Bump lithos-llm and adapt retry and attachment APIs 2026-10-05 09:12:28 -04:00
fabro-releases[bot]
7fc0edbf81 Bump version to 0.375.0-nightly.0
Some checks failed
Rust / Format (push) Has been cancelled
Rust / Clippy (push) Has been cancelled
Rust / Rustdoc (push) Has been cancelled
Rust / Generated Docs (push) Has been cancelled
Rust / Test (Linux) (push) Has been cancelled
Rust / Sandbox providers (Docker) (push) Has been cancelled
Rust / Test (macOS) (push) Has been cancelled
Rust / Process titles (musl) (push) Has been cancelled
2026-10-03 09:33:23 +00:00
Scott Werner
063ee15ca8
Merge pull request #918 from fabro-sh/codex/durable-final-patch
Some checks failed
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
TypeScript / Typecheck (push) Has been cancelled
TypeScript / Test (push) Has been cancelled
TypeScript / Build (push) Has been cancelled
Restore Files Changed from saved final patch blobs
2026-10-02 16:31:32 -04:00
Scott Werner
3dd0c20aa4 Restore renewable GitHub credentials to workflow stages 2026-10-02 15:36:27 -04:00
Scott Werner
b6de8bf8d3 Resolve the saved final patch when creating a pull request afterwards
Creating a pull request for a finished run read the projection's final
patch field directly. Petri stores that patch in the blob table and the
field holds a blob reference, so the description model was handed the
reference string instead of the diff, and the empty-diff check could
never fire.

Move the reference resolution out of Files Changed into a shared
final_patch::load helper and use it for both readers. Pull request input
extraction now reads the real patch, judges emptiness by its text, and
reports a missing or unreadable blob instead of describing a placeholder.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 14:49:40 -04:00
Scott Werner
824b5a2238
Merge pull request #919 from fabro-sh/codex/cli-model-overrides
Apply CLI model and provider overrides to agent execution
2026-10-02 13:43:14 -04:00
Scott Werner
fae391c4dc Bind CLI model flags in one place when building Petri's check
The create, validate and preflight paths each wrapped their launch with
the run's --model and --provider flags before building the check
request, so a new caller could build a check without them. The check
request now takes the flags as a required argument and binds them onto
the launch itself, and unit tests cover the flags, a provider-only flag,
and no flags.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 13:10:52 -04:00
Scott Werner
b401ac0a10 Bind CLI model flags as Petri's launch model and settings as its default
Petri now treats `petri.launch_model` and `petri.launch_provider` as the
model a run's flags ask for, above the file layers and the graph's
defaults. A host's last-resort default moved to `petri.default_model` and
`petri.default_provider`. Pin Petri at the merge of that change and bind
to it: the explicit `--model`/`--provider` flags go to the launch
variables, and the model the settings resolved (or the catalog default)
goes to the default variables.

`Launch` now names the two pairs `model`/`provider` and
`default_model`/`default_provider`, matching Petri, in place of the
separate override fields.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 11:28:27 -04:00
Scott Werner
4e71554773 Apply CLI model and provider overrides to agent execution 2026-10-02 11:28:16 -04:00
Scott Werner
4c4bf3df1f
Merge pull request #922 from fabro-sh/codex/enforce-network-block
Enforce blocked networking for Docker and Daytona runs
2026-10-02 11:21:59 -04:00
Scott Werner
500cd25814 Keep network policy off host runs and tidy the network tests
The host provider manages no networking and refuses any policy but its
default, so sending the resolved AllowAll to local runs failed every
non-dry-run local execution. Apply the run's policy only on container
backends; dry runs, which always use the host backend, are covered by the
same check.

Also fold the Docker environment test helper into one that takes a typed
network mode, share the probe setup between the live Docker and Daytona
network tests, count canary hits per mode, and bind the run environment
once in the worker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 10:17:17 -04:00
fabro-releases[bot]
bfe5b33b6a Bump version to 0.374.0-nightly.0 2026-10-02 09:36:50 +00:00
Scott Werner
cd5f56a4ed Enforce resolved network policy for Docker and Daytona runs 2026-10-01 17:12:03 -04:00
Scott Werner
a1ee6e5f47
Merge pull request #921 from fabro-sh/codex/fix-attach-test-deadlines
Some checks are pending
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
TypeScript / Typecheck (push) Waiting to run
TypeScript / Test (push) Waiting to run
TypeScript / Build (push) Waiting to run
Fix attach test deadlines around workflow completion
2026-10-01 17:02:22 -04:00
Scott Werner
3c942d8a79 Fix attach test deadlines around workflow completion 2026-10-01 16:33:44 -04:00
Scott Werner
82156442a8 Resolve durable final patch blobs for Files Changed 2026-10-01 14:42:33 -04:00
Scott Werner
e7b4859038
Merge pull request #913 from fabro-sh/restore-github-run-publication
Keep GitHub checkout, checkpoints, and pushes in the run sandbox
2026-10-01 13:52:49 -04:00
Scott Werner
5f22936437 Expect checkpoint commits from a retried local run
Retry now starts the workflow over, and a local-folder run commits its
checkpoints, so the retry scenario checks that every stage commits again
under the retry's run id. The scenario for retrying without Git
checkpoints goes: local runs have them now, and the retry scenario
covers starting over.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:10:13 -04:00
Scott Werner
6e0ef2c402 Keep Git checkpoints for runs on host workspaces
Checkpoints were only enabled for runs with a GitHub source, so local
folder runs, empty Local runs and dry runs stopped committing. `fabro
diff` then failed for them, and their checkpoint, run branch and diff
records disappeared from the event stream.

A run whose workspace is on the host now commits checkpoints there
again, without pushing, as on main. Docker and Daytona runs with no
GitHub source still record execution checkpoints without Git commits,
so a sandbox image without `git` cannot fail the run.

The scenario tests for crash recovery go back to asserting commits. A
workspace deleted while the run is down now fails the resumed run,
since the server keeps no copy to restore it from.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:01:53 -04:00
Scott Werner
23591d550f Reuse cached GitHub installation tokens for run fetches and pushes
The worker minted one read token at launch and a fresh push token for
every checkpoint. The read token expired an hour into a run, so a
workspace acquired later fetched with a dead credential. Minting per
push put every push in GitHub's token-replication window, where a token
minted moments earlier is rejected with 404 "Repository not found".

The worker now keeps two InstallationTokenSource caches for the run, a
read-only one for fetches and a contents: write one for pushes. Each
fetch and push resolves through its source, which reuses one token until
it nears expiry and then mints the next. Petri's RunSource asks a
SourceCredentials provider on every fetch instead of holding a fixed
credential.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:01:13 -04:00
Scott Werner
aad954693e Keep checkpoints and run branch pushes in the sandbox 2026-10-01 13:01:13 -04:00
Scott Werner
886065957c Simplify GitHub checkout and run publication
- Share one credential header helper between the in-sandbox fetch and
  the run branch push
- Keep only the target branch, goal, and model on the publisher instead
  of a full run spec copy
- Load the worker's LLM catalog once, and mint the read token only when
  the run checks something out
- Pass the source explicitly to checkpoint fetch helpers, dropping
  unreachable branches, and reuse has_object in has_commit
- Move the run patch into the publication instead of cloning it, and
  build it only when a publisher exists
- Add test fixture helpers for file sources and recording publishers

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:00:26 -04:00
Scott Werner
0250b40586 Publish a successful run from its run_finished hook
Pushing the run branch and opening the pull request now happen in the
run's worker, in Fabro's run_finished hook, after the last stage and
before the run's terminal record, as the legacy publish step did. A
failed push or pull request fails the run with publish_failed instead of
leaving a warning on a run that already succeeded.

fabro-petri gains a RunPublisher the hooks call for a successful run with
its run branch, final commit, snapshot repository and patch; the worker's
GitHub publisher pushes from the snapshot repository with a push token it
mints at that moment, opens the pull request its settings ask for, and
records it. The worker resolves the server's GitHub credentials itself
for both the read-only checkout token and the push token, so the server
no longer hands it a clone credential or publishes after the run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:00:26 -04:00
Scott Werner
5c6195a352 Check GitHub targets out in the sandbox and publish their run branch
Since the Petri cutover, a run with a GitHub target started in an empty
workspace, nothing pushed its run branch to GitHub, and nothing asked for
the automatic pull request when it succeeded.

Fabro's hooks now check a fresh run's GitHub target out inside the
sandbox when Petri hands them the scope, before the first stage: the
workspace fetches the selected commit, tag or branch at the run's clone
depth, with a read-only token the server resolves at each worker launch.
The worker scrubs the token from its environment at startup and presents
it only to the fetch, so it never lands in the repository or its remote.
The files belong to the sandbox user, so git accepts them.

The same checkout seeds the workspace's snapshot repository with the
starting commit. Checkpoint bundles from a shallow clone then import, a
stage's own commits never make a bundle carry the source's history, a
restore into a fresh sandbox fetches the base again and applies the run's
commits, and a fork carries the base with its checkpoints.

When a successful GitHub-target run ends, the server pushes its final
commit from the snapshot repository to fabro/run/<id> with its own write
credentials, then, when the run changed files and asks for one, records
the pull request request for the existing creation supervisor. A failed
push or request is a warning notice on the run. The manual pull request
endpoint shares the request step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:00:26 -04:00
Scott Werner
865e95383b
Merge pull request #917 from fabro-sh/retry-from-start
Retry a finished run from the start
2026-10-01 12:59:35 -04:00
Scott Werner
3c62d851fe Retry a finished run from the start
Retry forked the source run at its last checkpoint and reran the failed
stage. It now creates a new run from the source's saved spec and starts
the workflow from the beginning in a fresh workspace, as retry did
before the Petri cutover. The new run records `retried_from` and no
`fork_source_ref`.

Retry no longer needs a checkpoint, a retained workspace, or a published
run branch, so it works for any terminal run that is not archived. To
continue from where a run stopped, fork it at a checkpoint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 11:14:45 -04:00