Commit graph

5473 commits

Author SHA1 Message Date
Scott Werner
0c279ea284
Merge pull request #944 from fabro-sh/bump-internal-deps-daytona-acp
Some checks are pending
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
Run ACP agents on Daytona
2026-10-09 13:47:39 -04:00
Scott Werner
ae5e825a44 Track the latest Petri and graphviz-sys
Move Petri to current main, which picks up streamed stdin for Daytona
from sandbox-driver, reports how an ACP agent exited, extracts
checkouts as the sandbox user, and bumps lithos-llm and Pebble. Move
graphviz-sys to its latest commit (CI and README only). The Windows
crates return to main's newer versions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 13:31:56 -04:00
Scott Werner
782190d177 Run ACP agents on Daytona
Move sandbox-driver to the commit that streams stdin into Daytona
commands, and the Daytona SDK to the fix for stdout/stderr markers that
leaked into the session log stream. Together they let an ACP agent start
on Daytona and keep its JSON-RPC output intact. The agents docs no longer
say ACP fails on Daytona.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 12:39:10 -04:00
Scott Werner
5b3689c132
Merge pull request #939 from fabro-sh/required-run-finalization
Commit required publication before reporting run completion
2026-10-09 12:38:30 -04:00
Scott Werner
e9c35911bb Require finalization for every Fabro run
FabroHooks always declares required run finalization instead of only
when the run publishes or checkpoints. Every run takes one path, the
best-effort diff branch in run_finished goes away, and a fork always
declares what its worker's hooks will declare on resume, so the fork no
longer carries its source's flag across.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 12:08:28 -04:00
Scott Werner
b8281d6c1f Settle managed runs through one rule
ManagedRun::settle owns the rule that a run's first terminal status
sticks, and the lifecycle fold, the finish settle, fail_managed_run and
the in-process finish all go through it. One release_managed_run
replaces the two helpers that released a run's live state and also
frees its scheduler slot. The engine reads the finish through the
projection's finished_status, so a cancelled finish carrying a
checkpoint failure, and a failed publication, are mapped in one place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 11:59:19 -04:00
Scott Werner
04309977e5 Commit checkpoint failures through required finalization
A failed checkpoint cancels the run, so Petri recorded it as cancelled
and the worker overrode its own outcome in memory. Runs that checkpoint
now declare required finalization, and finalize_run rejects with
checkpoint_failed before publishing. The projection and the engine
outcome report a cancelled finish carrying that failure as a workflow
failure with the checkpoint's message, and the in-memory override is
gone. A run whose checkpoint failed is never published.

Retry a host failure that storage rejects with a short backoff, so a
brief storage fault does not leave the run active and holding its
scheduler slot. A failure that never commits still leaves the run's
status alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 11:00:01 -04:00
fabro-releases[bot]
fbed0c7ae3 Bump version to 0.381.0-nightly.0 2026-10-09 09:40:02 +00:00
Scott Werner
b4aea6128d Tidy host failure handling and restore main's lockfile versions
Share the managed-run cleanup between persist_run_failure and
fail_managed_run, return the committed projection from
commit_host_failure, and keep the run's status when the failure cannot
be stored. Restore the cancelled-launch message and keep the spawn
error's cause in the launch failure.

Rebuild Cargo.lock from main with only the Petri bump so the Windows
crates stay on their newer versions. Import the projection module
rather than the function, wrap long tracing calls, and document the
finalization test fixture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 15:38:10 -04:00
Scott Werner
af41c7c66c Simplify required run finalization
Name the publish failure once, settle host failures through one flat
helper, share the append-then-finish path for in-process runs, reuse
the source state already read when checking a fork, and drop the
unused finished-run status shim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 12:45:25 -04:00
Scott Werner
599b674ec1 Build required publication on Petri main's finalization
Track Petri main now that required run finalization has merged there,
read workflow execution from the engine state, resume the failed-publication
regression with its admitted workflow, and keep the worker lifecycle helper
usable for rejected appends.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 12:45:25 -04:00
Scott Werner
3443a923c0 Carry required publication into fork declarations
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 12:45:25 -04:00
Scott Werner
59f9088b04 Preserve committed outcomes during host failure cleanup
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 12:45:24 -04:00
Scott Werner
21d4db6766 Commit publication failures through required run finalization 2026-10-08 12:45:24 -04:00
Scott Werner
36f8b61b60
Merge pull request #941 from fabro-sh/codex/petri-test-server-port
Some checks are pending
Rust / Test (macOS) (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
Bind Petri scenario servers to OS-assigned ports
2026-10-08 10:50:56 -04:00
Scott Werner
d3d8893bf9 Bind Petri scenario servers to OS-assigned ports 2026-10-08 10:33:38 -04:00
fabro-releases[bot]
b609f6e015 Bump version to 0.380.0-nightly.0 2026-10-08 09:43:20 +00:00
Scott Werner
fa688cad50
Merge pull request #934 from fabro-sh/codex/macos-dry-run-timeout
Some checks failed
Rust / Clippy (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Rust / Format (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
TypeScript / Typecheck (push) Has been cancelled
TypeScript / Test (push) Has been cancelled
TypeScript / Build (push) Has been cancelled
Skip question polling for auto-approved CLI attach
2026-10-07 17:29:00 -04:00
Scott Werner
8b773cdd00
Merge pull request #788 from chunga-ict/fix/schema-mismatch-error-message
Explain schema mismatches as a version skew
2026-10-07 16:33:03 -04:00
Scott Werner
ee86a497d5 Update runner test for schema mismatch diagnostics 2026-10-07 16:21:51 -04:00
Scott Werner
e1932f2c42 Preserve the cause of schema mismatch errors 2026-10-07 16:07:59 -04:00
Scott Werner
a21d3b47c4
Merge pull request #935 from fabro-sh/codex/bump-internal-deps-20261007
Bump internal Git dependencies to latest revisions
2026-10-07 13:17:06 -04:00
Scott Werner
637fab367c Keep dependency bump within nightly updater scope 2026-10-07 11:56:13 -04:00
Scott Werner
08ffe2cc19
Merge pull request #923 from seniorquico/docs/local-run-checkout
Fix docs related to Local runs
2026-10-07 11:27:06 -04:00
Scott Werner
fb79d0bc74 Clarify source and execution directories for Local runs 2026-10-07 11:12:00 -04:00
Scott Werner
c5f0296e1a
Merge pull request #905 from fabro-sh/petri-no-jump
Drop the jump route kind from the Petri stream reader
2026-10-07 10:50:46 -04:00
Scott Werner
170a1212b7 Bump internal Git dependencies to latest revisions 2026-10-07 10:44:37 -04:00
Scott Werner
c1ce4428c2 Skip question polling for auto-approved CLI attach 2026-10-07 08:37:12 -04:00
fabro-releases[bot]
f7ca525b7d Bump version to 0.379.0-nightly.0 2026-10-07 09:40:07 +00:00
Scott Werner
eac82ee71c
Merge pull request #907 from fabro-sh/petri-store-faults
Some checks are pending
Rust / Test (macOS) (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
Handle Petri store faults: start again, resume after a failed write, stop surviving workers
2026-10-06 16:43:06 -04:00
Scott Werner
74cbdc48df
Merge pull request #931 from fabro-sh/docs/remove-early-access-banner
Remove early-access banner from server docs
2026-10-06 16:32:07 -04:00
Scott Werner
bc6bfeae67 Remove early-access banner from server docs
The server is generally available, so drop the private early-access
warning from the deployment, self-host, Railway, and server operations
pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 16:16:58 -04:00
Scott Werner
e889f9978d Simplify the store-fault handling
- stop_lock_holder returns the stopped pid as Option<u32> instead of a
  LockHolder enum that only wrapped it
- stop_previous_worker uses with_context, and one run_scratch helper
  replaces three spellings of the run's scratch path
- relaunch builds its runnable record through run_records::runnable,
  moved out of the lifecycle handler so both callers share it
- WorkerExit derives success from its exit code instead of storing both
- the scenario tests share one worker-pid lookup and wait loop
- small readability fixes in the engine's resume arm and the resume test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 16:07:36 -04:00
Scott Werner
f12065e6ae Adapt recovery to current resume callers and API errors 2026-10-06 12:29:34 -04:00
Bryan Helmkamp
eb39e15781 Describe store interruptions and the worker lock in the Petri README
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Bryan Helmkamp
cb54a0df90 Stop a surviving worker before ending its lease from outside
A worker leads a process group of its own, so it outlives a server
crash. The restarted server released the run's lease from outside and
launched a resume while that worker could still be running: a worker
whose lease is released keeps acting until its next write, beside its
successor. A delete likewise dropped the lease without knowing the
worker was gone.

Now each worker holds a lock on `worker.lock` in its run's scratch
directory for its whole life, taken before anything else. It is a POSIX
record lock: the kernel frees it only when the worker exits, and names
the process that holds it. Before the server ends a lease from outside
(the relaunch after a restart or a store interruption, and a delete), it
kills whatever process still holds the lock, with its process group, and
waits until the lock is free. A worker that finds the lock held does not
start.

- fabro-proc: ProcessLock::try_hold and stop_lock_holder, tested with
  this test binary as the holding process.
- fabro-config: RunScratch::worker_lock_path.
- New scenario: a worker that outlives the server is gone before the
  resume's worker launches, and the run succeeds once. It fails without
  the server-side stop.

A host stage process runs in a process group of its own and still
outlives its killed worker, as it did at a worker crash.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Bryan Helmkamp
37116758fa Resume a Petri run whose store failed, as after a crash
Petri now ends a run's lifetime at its first failed store write: it
records nothing after it, fails no firing for it, and returns
CoordinatorError::StoreFailed. The run is not over; the next lifetime
resumes it from what the store holds. Fabro read that error as an
unfinished run and failed it.

- engine: RunError::StoreFailed, returned without reading the record
  back, and Conclusion::Interrupted for it.
- worker: an interrupted run gets no terminal lifecycle record; the
  worker exits with EX_TEMPFAIL (75, the new ExitClass::Interrupted).
- server: WorkerExit carries the exit code. An interrupted worker's run
  goes back to the scheduler in resume mode through the relaunch a
  restart takes (lease release, recovery, start_requested + runnable),
  now shared with reconcile_on_startup. The in-process path does the
  same. A run is resumed at most MAX_STORE_INTERRUPTIONS (3) times per
  server; the next interruption fails it. A pending cancel, a run that
  ended or was deleted, and a shutdown also end it as before.

Tests: an engine run over a store whose first lease write fails is
interrupted with no finish, and a resume finishes it; the server
relaunches an interrupted worker in resume mode, fails the run after the
bound, and fails a worker that exits 1 as before; exit code 75.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Bryan Helmkamp
3c395f9e6e Start a Petri run again when its creation was cut short
Petri now refuses to resume a run whose creation a crash cut short (the
key is stored, the root invocation is not) with HostError::NotStarted,
and starts it again when the host runs it under the same key. The engine
used its own guard, check_resumable, which failed the run with
NothingToResume instead.

Execution::Resume now carries the admitted graphs, and a resume Petri
answers with NotStarted starts the run from them. The worker loads the
graphs in resume mode too, as does the server's in-process path. The
guard and RunError::NothingToResume are gone. New test:
a_resume_of_a_run_that_never_started_starts_it_again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:29:34 -04:00
Scott Werner
d89d0c577b
Merge pull request #930 from fabro-sh/codex/partial-success-regression
Restore regression coverage for partial-success failure projection
2026-10-06 12:23:03 -04:00
Scott Werner
e04ba37e25 Test the cause retained by partial-success projection 2026-10-06 12:07:28 -04:00
Scott Werner
789ca8427b
Merge pull request #926 from fabro-sh/bump-lithos-llm-fd42e6b
Bump lithos-llm and Pebble to current main
2026-10-06 10:28:28 -04:00
Scott Werner
8be46335f3
Merge pull request #929 from fabro-sh/deps/quinn-proto-0.11.18
deps: bump quinn-proto to 0.11.18
2026-10-06 10:26:15 -04:00
Scott Werner
6d3859f4ec Merge remote-tracking branch 'origin/main' into bump-lithos-llm-fd42e6b 2026-10-06 09:13:53 -04:00
Scott Werner
ddf4b4db8d deps: bump quinn-proto to 0.11.18
Clears RUSTSEC-2026-0185 (fixed in 0.11.15). quinn is only in the lock
through reqwest's optional http3 feature, which Fabro doesn't enable, so
this only stops lockfile scanners from flagging it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 09:13:32 -04:00
Scott Werner
ee4169ff11
Merge pull request #899 from fabro-sh/ci/internal-deps-nightly
ci: add a nightly internal dependency update
2026-10-06 09:02:47 -04:00
fabro-releases[bot]
64b9d88159 Bump version to 0.378.0-nightly.0 2026-10-06 09:36:59 +00:00
Scott Werner
2e4055c624
Merge pull request #924 from fabro-sh/codex/restore-stage-git-auth
Some checks are pending
Rust / Format (push) Waiting to run
Rust / Clippy (push) Waiting to run
Rust / Rustdoc (push) Waiting to run
Rust / Generated Docs (push) Waiting to run
Rust / Test (Linux) (push) Waiting to run
Rust / Sandbox providers (Docker) (push) Waiting to run
Rust / Test (macOS) (push) Waiting to run
Rust / Process titles (musl) (push) Waiting to run
Restore renewable GitHub credentials to workflow stages
2026-10-05 22:25:09 -04:00
Scott Werner
1f4712b845 Simplify stage credential renewal and share the read token source
Rewrite the sandbox credential store only when the token source mints a
new generation, and skip the refresh loop for static tokens. Build the
run's read-only token source once in the worker and share it between the
workspace fetch and stage Git access. Drop unreachable branches in
for_run, reuse the shared contents-permission check and constants, and
fold the duplicated token resolution and test setup into helpers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 21:58:19 -04:00
Scott Werner
9e4af50863 Move stage credentials to Petri's spawn env and let the managed token win
Pin Petri at e46845b, the merge of its executor layer change. The stage
credential layer is now a Petri `SpawnEnv` applied with
`EnvHandle::with_spawn_env`, so Petri forwards every other environment
method and applies the layer to one-shot containers as well as processes.

A container gets the managed GITHUB_TOKEN, but no credential store refresh
or Git helper configuration: the store lives in the scope's sandbox, which
the container does not share.

The managed GITHUB_TOKEN now replaces one set in the workflow environment,
an ACP agent's environment or the sandbox's own, as Fabro's stage
environment did before Petri. The token carries exactly the access the run
declares; a stage that needs other access changes its declaration.

The new pin also keeps a timeout as the failure a partial success came
from: `PartialSuccess.underlying` is now an `UnderlyingFailure`, so the
projection reports "the step timed out" for a partial success converted
from a timeout. The run format moves from 7 to 8, which the attach JSON
snapshot records; runs stored before this pin are refused, as with earlier
format changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 19:26:08 -04:00
Scott Werner
8db70e072f Merge origin/main into codex/restore-stage-git-auth 2026-10-05 19:25:39 -04:00