mirror of
https://github.com/fabro-sh/fabro.git
synced 2026-10-07 03:00:29 +00:00
Keep checkpoints and run branch pushes in the sandbox
This commit is contained in:
parent
886065957c
commit
aad954693e
19 changed files with 1115 additions and 1914 deletions
32
AGENTS.md
32
AGENTS.md
|
|
@ -34,13 +34,15 @@ macOS note: if `cargo nextest run` fails with `Too many open files (os error 24)
|
|||
sandbox driver or Petri's `start` checkout: when a fresh run's scope is
|
||||
acquired, `fabro-petri`'s `RunWorkspaces::check_out_source` fetches the
|
||||
target's revision inside the sandbox with a read-only token the worker
|
||||
mints from the server's GitHub credentials, and seeds the workspace's
|
||||
snapshot repository with that commit so checkpoint bundles from a shallow
|
||||
clone import. In the `run_finished` hook, before the terminal record, a
|
||||
successful run's worker pushes the final checkpoint to `fabro/run/<id>`
|
||||
from the snapshot repository and opens the pull request its settings ask
|
||||
for (`fabro-cli/src/commands/run/publish.rs`); a failure fails the run
|
||||
with `publish_failed`. `CloneRequest` still
|
||||
mints from the server's GitHub credentials. Checkpoints and pushes execute
|
||||
in that same workspace, pushing `fabro/run/<id>` after each checkpoint;
|
||||
no server snapshot repository or Git bundle is maintained. Final publication
|
||||
retries the push and opens the configured pull request; a failure fails the
|
||||
run with `publish_failed` (`fabro-cli/src/commands/run/publish.rs`). Ordinary
|
||||
resume requires the retained workspace. A GitHub-backed fork fetches the
|
||||
source run branch inside its new sandbox. Empty and local-folder targets
|
||||
keep execution metadata without automatic Git checkpoints; retry starts a
|
||||
fresh execution from the saved spec. `CloneRequest` still
|
||||
travels beside the sandbox spec so the run record names the origin and
|
||||
branch; the sandbox layer refuses a request that asks it to clone.
|
||||
Preflight and `fabro exec` initialize sandboxes with `CloneRequest::none()`,
|
||||
|
|
@ -154,13 +156,15 @@ Fabro is an AI-powered workflow orchestration platform. Workflows are defined as
|
|||
sandbox driver or Petri's `start` checkout: when a fresh run's scope is
|
||||
acquired, `fabro-petri`'s `RunWorkspaces::check_out_source` fetches the
|
||||
target's revision inside the sandbox with a read-only token the worker
|
||||
mints from the server's GitHub credentials, and seeds the workspace's
|
||||
snapshot repository with that commit so checkpoint bundles from a shallow
|
||||
clone import. In the `run_finished` hook, before the terminal record, a
|
||||
successful run's worker pushes the final checkpoint to `fabro/run/<id>`
|
||||
from the snapshot repository and opens the pull request its settings ask
|
||||
for (`fabro-cli/src/commands/run/publish.rs`); a failure fails the run
|
||||
with `publish_failed`. `CloneRequest` still
|
||||
mints from the server's GitHub credentials. Checkpoints and pushes execute
|
||||
in that same workspace, pushing `fabro/run/<id>` after each checkpoint;
|
||||
no server snapshot repository or Git bundle is maintained. Final publication
|
||||
retries the push and opens the configured pull request; a failure fails the
|
||||
run with `publish_failed` (`fabro-cli/src/commands/run/publish.rs`). Ordinary
|
||||
resume requires the retained workspace. A GitHub-backed fork fetches the
|
||||
source run branch inside its new sandbox. Empty and local-folder targets
|
||||
keep execution metadata without automatic Git checkpoints; retry starts a
|
||||
fresh execution from the saved spec. `CloneRequest` still
|
||||
travels beside the sandbox spec so the run record names the origin and
|
||||
branch; the sandbox layer refuses a request that asks it to clone.
|
||||
Preflight and `fabro exec` initialize sandboxes with `CloneRequest::none()`,
|
||||
|
|
|
|||
|
|
@ -2437,7 +2437,10 @@ paths:
|
|||
Creates a new run from a checkpoint of a terminal source run and
|
||||
starts it, then archives the source run and records
|
||||
`run.superseded_by` on it. Returns 207 when the new run was created
|
||||
but the source archive step failed.
|
||||
but the source archive step failed. Requires a GitHub target with
|
||||
run-branch creation and pushes enabled. The checkpoint must leave
|
||||
work to execute: the final terminal checkpoint is refused; select
|
||||
an earlier checkpoint or retry the workflow from the beginning.
|
||||
parameters:
|
||||
- $ref: "#/components/parameters/RunId"
|
||||
requestBody:
|
||||
|
|
@ -2498,7 +2501,12 @@ paths:
|
|||
position and continues from there in a fresh workspace restored to
|
||||
the checkpoint's commit. The source run is left untouched. A
|
||||
checkpoint inside a parallel branch cannot be forked at; fork at the
|
||||
parallel stage instead.
|
||||
parallel stage instead. Requires a GitHub target with run-branch
|
||||
creation and pushes enabled; the checkpoint commit must be available
|
||||
on the source run's published branch. The final terminal checkpoint
|
||||
is refused because it leaves no work to acquire a workspace. For a
|
||||
completed run, select an earlier checkpoint explicitly or retry the
|
||||
workflow from the beginning.
|
||||
parameters:
|
||||
- $ref: "#/components/parameters/RunId"
|
||||
requestBody:
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@ title: "Checkpoints"
|
|||
description: "How Fabro uses Git to checkpoint and resume workflow runs"
|
||||
---
|
||||
|
||||
Fabro checkpoints every workflow run using Git plus the durable run store. After each node completes, Fabro commits file changes to the run branch and records execution state in the durable event stream so that interrupted runs can be resumed exactly where they left off. This happens automatically — no configuration required beyond running inside a Git repository.
|
||||
Fabro records every workflow run in the durable run store. For GitHub targets with run branches enabled, it also commits file changes inside the run workspace after each node completes. Empty and local-folder targets keep execution records without automatic Git checkpoint commits, even when the supplied folder contains a Git repository.
|
||||
|
||||
## Code and execution history
|
||||
|
||||
|
|
@ -71,22 +71,15 @@ The checkpoint projection captures the execution state needed to resume a run:
|
|||
|
||||
The durable run store also keeps the current checkpoint so `resume`, `inspect`, and API reads do not need to rely on scratch files.
|
||||
|
||||
## Worktrees
|
||||
## Run workspaces
|
||||
|
||||
Fabro uses Git worktrees to isolate workflow runs from your working directory. When a local run starts in a Git repository:
|
||||
A GitHub target is checked out inside the run workspace. For Docker and Daytona, checkout, checkpoint commits, diffs, and pushes execute inside the sandbox. The host provider performs those operations in its local run workspace. Fabro keeps no second checkout, bare snapshot repository, or Git checkpoint bundle on the server for a remote sandbox.
|
||||
|
||||
1. Fabro records the current HEAD as the **base SHA**
|
||||
2. Creates a new branch `fabro/run/{run_id}` at that SHA
|
||||
3. Adds a worktree at `{run_dir}/worktree` on that branch
|
||||
4. Changes into the worktree directory for the duration of the run
|
||||
When pushing is configured, Fabro pushes `fabro/run/{run_id}` after each checkpoint and again before successful completion. Execution records and artifact payloads remain in the run store and CAS.
|
||||
|
||||
This means your original working directory stays untouched while the agent makes changes in the worktree. When the run completes, Fabro removes the worktree and restores your original directory.
|
||||
Ordinary resume requires the original workspace to survive. Git-backed workspaces reset to their recorded checkpoint; non-Git workspaces continue with their surviving files. Losing the workspace does not trigger automatic reconstruction or sandbox replacement.
|
||||
|
||||
<Note>
|
||||
If the working directory has uncommitted changes, the worktree starts from committed `HEAD` and those uncommitted changes are not included. Fabro logs a warning so you can commit, stash, or run explicitly in place when that is what you want.
|
||||
</Note>
|
||||
|
||||
For Docker and Daytona sandboxes, a GitHub target is checked out inside the sandbox and checkpoint Git operations run there; each checkpoint commit reaches the server as a Git bundle. Before a successful run finishes, Fabro pushes the run branch to origin when pushing is configured.
|
||||
A GitHub-backed fork or rewind fetches the source run's published branch inside a fresh workspace and checks out the selected checkpoint. This requires run-branch pushes to be enabled and the commit to be available on origin. Select a checkpoint with remaining work: the final terminal checkpoint is refused because no stage remains to acquire a workspace. A retry starts the workflow from the beginning using its saved specification, without requiring the previous workspace or any Git checkpoint.
|
||||
|
||||
## Resuming a run
|
||||
|
||||
|
|
@ -123,7 +116,7 @@ After a node completes, Fabro:
|
|||
3. Collects the code diff.
|
||||
4. Emits a checkpoint event with execution state and the code commit SHA. The run store persists this event and updates the projection.
|
||||
|
||||
A checkpoint commit failure stops execution. A diff failure emits a warning notice. A failed final push or pull request marks a successful run as failed with `publish_failed`.
|
||||
A checkpoint commit failure stops execution. A stage diff failure emits a warning notice. An intermediate push failure warns and is retried at later checkpoints and final publication. Failure to prepare required final publication, push the final commit, or open the pull request marks an otherwise successful run as failed with `publish_failed`.
|
||||
|
||||
## Inspecting run history
|
||||
|
||||
|
|
@ -151,12 +144,12 @@ fabro dump 01JKXYZ --output ./run-dump
|
|||
|
||||
Git checkpointing activates automatically when:
|
||||
|
||||
- The run uses a Git repository and checkpointing has not been explicitly disabled
|
||||
- Local runs can create a Git worktree under the run scratch directory
|
||||
- Docker or Daytona can clone the configured GitHub origin into the sandbox
|
||||
- The run targets a GitHub repository with cloning enabled
|
||||
- Run branches are enabled
|
||||
- The run is not a dry run
|
||||
|
||||
It is skipped when:
|
||||
|
||||
- The working directory is not a Git repository
|
||||
- The target is an empty workspace or a local folder
|
||||
- The run uses `--dry-run`
|
||||
- The run is explicitly started in place with checkpointing disabled
|
||||
- Cloning or run branches are disabled
|
||||
|
|
|
|||
|
|
@ -29,7 +29,7 @@ The rest of this page describes the `app` strategy, which is required for browse
|
|||
|---|---|
|
||||
| **OAuth login** | Users sign in to the web UI with their GitHub account |
|
||||
| **Private repo cloning** | Daytona and Docker sandboxes fetch private repositories using short-lived, read-only Installation Access Tokens |
|
||||
| **Run branch pushing** | Before a successful run finishes, Fabro pushes the run branch to origin |
|
||||
| **Run branch pushing** | Fabro pushes from the sandbox after each checkpoint and again before a successful run finishes |
|
||||
| **Auto-PR** | When `[run.pull_request] enabled = true` in the [run config](/execution/run-configuration#runpull_request), Fabro opens a PR from the agent's working branch after a successful run |
|
||||
| **Auto-merge** | When `[run.pull_request] auto_merge = true`, Fabro enables GitHub's auto-merge on created PRs so they merge automatically once required checks pass |
|
||||
| **Sandbox GITHUB_TOKEN** | When `[run.integrations.github.permissions]` are declared at any layer (workflow, project, or user settings), Fabro mints a scoped Installation Access Token and injects it as `GITHUB_TOKEN` in the sandbox |
|
||||
|
|
@ -318,9 +318,11 @@ Every commit a run creates is authored and committed by the run's GitHub credent
|
|||
|
||||
### Checkpoint pushing
|
||||
|
||||
After each workflow stage, Fabro [checkpoints](/execution/checkpoints) the workspace on the run branch, `fabro/run/<run-id>`, and moves the commit to the server. After a successful run's last stage and before the run finishes, Fabro pushes the final commit to that branch on origin with an Installation Access Token with `contents: write`, minted for the push, so a long run never pushes with an expired token. The sandbox never holds a credential that can push. `[run.run_branch] push = false` keeps the branch on the server.
|
||||
After each workflow stage, Fabro [checkpoints](/execution/checkpoints) the workspace on the run branch, `fabro/run/<run-id>`, and pushes it directly from that workspace. Fabro keeps no second Git repository or checkpoint bundles on the server. In App mode, the worker mints a fresh Installation Access Token with `contents: write` for each checkpoint push and passes it only to the sandbox Git command. The token is not saved in the repository's remote URL or configuration. `[run.run_branch] push = false` keeps the branch only in the run workspace.
|
||||
|
||||
When pull request creation is enabled and the run changed files, Fabro then checks that GitHub reports the run branch at the exact final commit and opens the pull request. A failed push, branch check, or PR creation marks the run as failed with `publish_failed`; the terminal run event is emitted only after this step finishes.
|
||||
An intermediate push failure logs a warning; later checkpoints and final publication retry the push. After a successful run's last stage and before the run finishes, Fabro awaits one final push.
|
||||
|
||||
When pull request creation is enabled and the run changed files, Fabro then checks that GitHub reports the run branch at the exact final commit and opens the pull request. A failed final push, publication preparation, branch check, or PR creation marks the run as failed with `publish_failed`; the terminal run event is emitted only after this step finishes.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
|
|
|
|||
|
|
@ -6,17 +6,16 @@
|
|||
//! server named (`FABRO_CONFIG`), the App key the server hands the worker,
|
||||
//! or `GITHUB_TOKEN` from the worker's vault snapshot. From them it mints a
|
||||
//! read-only token for the checkout inside the sandbox when the run starts,
|
||||
//! and a push token when the run ends, so a long run never pushes with an
|
||||
//! expired token.
|
||||
//! and a fresh push token for each checkpoint, so a long run never pushes with
|
||||
//! an expired token.
|
||||
//!
|
||||
//! Publication runs in Fabro's `run_finished` hook, after the last stage and
|
||||
//! before the run's terminal record, as the legacy publish step did: the
|
||||
//! final checkpoint is pushed from the run's snapshot repository to the run
|
||||
//! final checkpoint is pushed from inside the sandbox to the run
|
||||
//! branch on GitHub, and, when the run changed files and its settings ask
|
||||
//! for one, a pull request is opened and recorded. A failure fails the run
|
||||
//! with `publish_failed`.
|
||||
|
||||
use std::path::Path;
|
||||
use std::sync::Arc;
|
||||
use std::time::Duration;
|
||||
|
||||
|
|
@ -27,6 +26,7 @@ use fabro_config::ServerSettingsBuilder;
|
|||
use fabro_github::{GitCloneCredentials, GitHubContext, GitHubCredentials};
|
||||
use fabro_llm::credentials::{CredentialProvider, readiness};
|
||||
use fabro_llm::lithos_catalog::Catalog;
|
||||
use fabro_petri::checkpoint::Site;
|
||||
use fabro_petri::hooks::{Publication, RunPublisher};
|
||||
use fabro_petri::platform_records::PlatformRecords;
|
||||
use fabro_petri::source::SourceCredential;
|
||||
|
|
@ -37,7 +37,6 @@ use fabro_types::settings::server::GithubIntegrationStrategy;
|
|||
use fabro_types::{GitHubRepositorySlug, RunId, RunSpec, RunTarget};
|
||||
use fabro_vault::Vault;
|
||||
use fabro_workflow::pull_request::{self, AutoMergeOptions, OpenPullRequestRequest};
|
||||
use tokio::process::Command;
|
||||
use tokio::time;
|
||||
use tracing::warn;
|
||||
|
||||
|
|
@ -146,6 +145,7 @@ impl GitHubPublisher {
|
|||
) -> Option<Self> {
|
||||
let settings = &spec.settings.run;
|
||||
if settings.execution.mode == RunMode::DryRun
|
||||
|| !settings.clone.enabled
|
||||
|| !settings.run_branch.enabled
|
||||
|| !settings.run_branch.push
|
||||
{
|
||||
|
|
@ -238,7 +238,7 @@ impl GitHubPublisher {
|
|||
|
||||
#[async_trait::async_trait]
|
||||
impl RunPublisher for GitHubPublisher {
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String> {
|
||||
async fn push(&self, site: &Site, branch: &str, sha: &str) -> Result<(), String> {
|
||||
let credentials = self.credentials.as_ref().ok_or_else(|| {
|
||||
"pushing the run branch requires the server's GitHub credentials".to_string()
|
||||
})?;
|
||||
|
|
@ -251,47 +251,63 @@ impl RunPublisher for GitHubPublisher {
|
|||
)
|
||||
.await
|
||||
.map_err(|err| format!("no push credential for {}: {err:#}", self.repository))?;
|
||||
push(&self.repository, &push_credentials, publication).await?;
|
||||
push(&self.repository, &push_credentials, site, branch, sha).await
|
||||
}
|
||||
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String> {
|
||||
self.push(
|
||||
&publication.site,
|
||||
&publication.run_branch,
|
||||
&publication.head_sha,
|
||||
)
|
||||
.await?;
|
||||
let Some(settings) = &self.pull_request else {
|
||||
return Ok(());
|
||||
};
|
||||
if publication.patch.trim().is_empty() {
|
||||
return Ok(());
|
||||
}
|
||||
let credentials = self
|
||||
.credentials
|
||||
.as_ref()
|
||||
.ok_or_else(|| "GitHub credentials are unavailable".to_string())?;
|
||||
let base_url = fabro_github::github_api_base_url();
|
||||
let context = GitHubContext::new(credentials, &base_url);
|
||||
Box::pin(self.open_pull_request(context, settings, publication)).await
|
||||
}
|
||||
}
|
||||
|
||||
/// Push the run's final commit from its snapshot repository to the run
|
||||
/// Push a checkpoint from inside its workspace to the run
|
||||
/// branch on GitHub, retrying a failure that may be a token still
|
||||
/// replicating. The credential reaches `git` as an HTTP header for the
|
||||
/// repository alone and never appears in the error.
|
||||
async fn push(
|
||||
repository: &GitHubRepositorySlug,
|
||||
credentials: &GitCloneCredentials,
|
||||
publication: &Publication,
|
||||
site: &Site,
|
||||
branch: &str,
|
||||
sha: &str,
|
||||
) -> Result<(), String> {
|
||||
let mut url = repository.https_url();
|
||||
url.push_str(".git");
|
||||
let env = encode(credentials)
|
||||
.map(|credential| credential.header_env(&url))
|
||||
.unwrap_or_default();
|
||||
let refspec = format!(
|
||||
"{}:refs/heads/{}",
|
||||
publication.head_sha, publication.run_branch
|
||||
);
|
||||
let refspec = format!("{sha}:refs/heads/{branch}");
|
||||
let mut last = String::new();
|
||||
for attempt in 1..=PUSH_ATTEMPTS {
|
||||
let output = git_push(&publication.snapshot_repository, &url, &refspec, &env).await?;
|
||||
if output.status.success() {
|
||||
return Ok(());
|
||||
match site.push(&url, &refspec, &env, PUSH_TIMEOUT).await {
|
||||
Ok(()) => return Ok(()),
|
||||
Err(error) => {
|
||||
last = error.to_string().replace(credentials.password(), "***");
|
||||
if let Some(encoded) = encode(credentials) {
|
||||
last = last.replace(encoded.encoded(), "***");
|
||||
}
|
||||
}
|
||||
}
|
||||
last = String::from_utf8_lossy(&output.stderr)
|
||||
.trim()
|
||||
.replace(credentials.password(), "***");
|
||||
warn!(
|
||||
attempt,
|
||||
branch = publication.run_branch,
|
||||
branch,
|
||||
error = last,
|
||||
"pushing the run branch failed"
|
||||
);
|
||||
|
|
@ -300,34 +316,10 @@ async fn push(
|
|||
}
|
||||
}
|
||||
Err(format!(
|
||||
"the run branch {} could not be pushed to {repository}: {last}",
|
||||
publication.run_branch
|
||||
"the run branch {branch} could not be pushed to {repository}: {last}"
|
||||
))
|
||||
}
|
||||
|
||||
/// `git push` of `refspec` to `url` from `directory` with `env` added,
|
||||
/// non-interactive, bounded.
|
||||
async fn git_push(
|
||||
directory: &Path,
|
||||
url: &str,
|
||||
refspec: &str,
|
||||
env: &[(String, String)],
|
||||
) -> Result<std::process::Output, String> {
|
||||
let mut command = Command::new("git");
|
||||
command
|
||||
.args(["push", "--quiet", url, refspec])
|
||||
.current_dir(directory)
|
||||
.envs(env.iter().map(|(key, value)| (key, value)))
|
||||
.env("GIT_TERMINAL_PROMPT", "0")
|
||||
.stdin(std::process::Stdio::null())
|
||||
.kill_on_drop(true);
|
||||
match time::timeout(PUSH_TIMEOUT, command.output()).await {
|
||||
Ok(Ok(output)) => Ok(output),
|
||||
Ok(Err(err)) => Err(format!("git could not run: {err}")),
|
||||
Err(_) => Err(format!("git push timed out after {PUSH_TIMEOUT:?}")),
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use fabro_llm::credentials::NoCredentials;
|
||||
|
|
|
|||
|
|
@ -286,38 +286,15 @@ impl RunningServer {
|
|||
}
|
||||
}
|
||||
|
||||
/// The one host workspace of the run, and the commits on its run
|
||||
/// branch, oldest first, as `(subject, key)`.
|
||||
fn workspace_commits(&self, run_id: &str) -> (PathBuf, Vec<(String, Option<CheckpointKey>)>) {
|
||||
/// The one host workspace, without assuming the run uses Git.
|
||||
fn workspace_path(&self, run_id: &str) -> PathBuf {
|
||||
let scopes = self.petri_run_dir(run_id).join("scopes");
|
||||
let mut workspaces: Vec<PathBuf> = std::fs::read_dir(&scopes)
|
||||
.expect("the scopes directory lists")
|
||||
.map(|entry| entry.expect("an entry reads").path().join("work"))
|
||||
let workspaces: Vec<_> = std::fs::read_dir(scopes)
|
||||
.expect("the scopes directory exists")
|
||||
.map(|entry| entry.expect("the scope entry reads").path().join("work"))
|
||||
.collect();
|
||||
assert_eq!(workspaces.len(), 1, "one workspace: {workspaces:?}");
|
||||
let workspace = workspaces.remove(0);
|
||||
let output = Command::new("git")
|
||||
.args(["log", "--reverse", "--format=%s%x00%B%x1e"])
|
||||
.current_dir(&workspace)
|
||||
.output()
|
||||
.expect("git runs");
|
||||
assert!(
|
||||
output.status.success(),
|
||||
"git log failed: {}",
|
||||
String::from_utf8_lossy(&output.stderr)
|
||||
);
|
||||
let log = String::from_utf8_lossy(&output.stdout).into_owned();
|
||||
let commits = log
|
||||
.split('\u{1e}')
|
||||
.filter(|entry| !entry.trim().is_empty())
|
||||
.map(|entry| {
|
||||
let mut parts = entry.trim_start().splitn(2, '\0');
|
||||
let subject = parts.next().unwrap_or_default().to_string();
|
||||
let body = parts.next().unwrap_or_default();
|
||||
(subject, CheckpointKey::from_message(body))
|
||||
})
|
||||
.collect();
|
||||
(workspace, commits)
|
||||
assert_eq!(workspaces.len(), 1);
|
||||
workspaces[0].clone()
|
||||
}
|
||||
|
||||
/// The run's checkpoint records, in seq order, as `(node position, sha)`.
|
||||
|
|
@ -1161,19 +1138,16 @@ async fn a_finished_petri_run_reads_back_through_the_cli() {
|
|||
[CLOCK] ▶ start
|
||||
[CLOCK] │ checkout: [TEMP_DIR]/petri-workspace is not a Git repository; the workspace starts empty
|
||||
[CLOCK] ✓ start [DURATION]
|
||||
[CLOCK] Branch: fabro/run/[ULID] from [SHA]
|
||||
[CLOCK] Git identity: Fabro <noreply@fabro.sh> default
|
||||
[CLOCK] ⎘ Checkpoint [SHA]
|
||||
[CLOCK] ⎘ Checkpoint (no commit)
|
||||
[CLOCK] ▶ say
|
||||
[CLOCK] start → say continue
|
||||
[CLOCK] │ hello from petri
|
||||
[CLOCK] ✓ say [DURATION]
|
||||
[CLOCK] ⎘ Checkpoint [SHA]
|
||||
[CLOCK] ⎘ Checkpoint (no commit)
|
||||
[CLOCK] ▶ exit
|
||||
[CLOCK] say → exit continue
|
||||
[CLOCK] ✓ exit [DURATION]
|
||||
[CLOCK] ⎘ Checkpoint [SHA]
|
||||
[CLOCK] Diff: +0 -0 in 0 file(s)
|
||||
[CLOCK] ⎘ Checkpoint (no commit)
|
||||
[CLOCK] ✓ SUCCEEDED [DURATION]
|
||||
[CLOCK] · succeeded
|
||||
----- stderr -----
|
||||
|
|
@ -1334,7 +1308,7 @@ fn three_stage_bundle(context: &fabro_test::TestContext, gate: &Path) -> PathBuf
|
|||
"digraph Stages {{\n graph [goal=\"Three stages\", default_max_retries=0]\n start \
|
||||
[shape=Mdiamond]\n exit [shape=Msquare]\n one [shape=parallelogram, script=\"echo \
|
||||
one > one.txt\"]\n two [shape=parallelogram, script=\"echo run >> two.log; {}; \
|
||||
echo two > two.txt\"]\n three [shape=parallelogram, script=\"test \\\"$(cat \
|
||||
echo two > two.txt\"]\n three [shape=parallelogram, goal_gate=true, script=\"test \\\"$(cat \
|
||||
one.txt)\\\" = one && test \\\"$(cat two.txt)\\\" = two && cp two.log \
|
||||
three.log\"]\n start -> one -> two -> three -> exit\n}}\n",
|
||||
wait_for(gate)
|
||||
|
|
@ -1389,27 +1363,10 @@ pub(super) async fn wait_for_success(server: &RunningServer, run_id: &str) {
|
|||
);
|
||||
}
|
||||
|
||||
/// The subjects of the commits on the run branch.
|
||||
fn subjects(commits: &[(String, Option<CheckpointKey>)]) -> Vec<&str> {
|
||||
commits
|
||||
.iter()
|
||||
.map(|(subject, _)| subject.as_str())
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// The commit subjects one run of the three-stage bundle produces.
|
||||
fn three_stage_subjects(run_id: &str) -> Vec<String> {
|
||||
["start", "one", "two", "three", "exit"]
|
||||
.iter()
|
||||
.map(|node| format!("fabro({run_id}): {node} (success)"))
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// A worker killed after a stage's finish is durable: on the restart the
|
||||
/// stage's commit is not repeated, the stage in flight reruns on the
|
||||
/// snapshot (its partial output gone), and the next stage sees both.
|
||||
/// Durable execution finishes survive restart. The in-flight stage reruns
|
||||
/// in its surviving workspace, including its partial non-Git output.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_crash_after_a_durable_finish_keeps_its_one_commit() {
|
||||
async fn a_non_git_crash_resumes_execution_in_the_surviving_workspace() {
|
||||
let context = test_context!();
|
||||
let mut server = RunningServer::start().await;
|
||||
let gate = context.temp_dir.join("two.gate");
|
||||
|
|
@ -1428,49 +1385,45 @@ async fn a_crash_after_a_durable_finish_keeps_its_one_commit() {
|
|||
std::fs::write(&gate, "go").expect("the gate opens");
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let (path, commits) = server.workspace_commits(&run_id);
|
||||
assert_eq!(subjects(&commits), three_stage_subjects(&run_id));
|
||||
let path = server.workspace_path(&run_id);
|
||||
assert!(!path.join(".git").exists());
|
||||
assert_eq!(
|
||||
std::fs::read_to_string(path.join("three.log")).expect("three copied the log"),
|
||||
"run\n",
|
||||
"the crashed attempt's partial output was reset before the rerun; server log:\n{}",
|
||||
"run\nrun\n",
|
||||
"non-Git recovery preserves the interrupted attempt's files; server log:\n{}",
|
||||
server.stderr_text()
|
||||
);
|
||||
let checkpoints = server.checkpoints(&run_id).await;
|
||||
assert_eq!(checkpoints.len(), 5, "{checkpoints:?}");
|
||||
let keys: Vec<Option<CheckpointKey>> = checkpoints.iter().map(|(key, _)| Some(*key)).collect();
|
||||
let committed: Vec<Option<CheckpointKey>> = commits.iter().map(|(_, key)| *key).collect();
|
||||
assert_eq!(keys, committed);
|
||||
assert!(server.checkpoints(&run_id).await.is_empty());
|
||||
assert!(!server.petri_run_dir(&run_id).join("snapshots").exists());
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A worker killed in `prepare_result` before the commit lands: the finish
|
||||
/// is not durable, the stage reruns once, and one commit exists for it.
|
||||
/// A worker killed before its durable finish reruns that stage once.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_crash_before_the_commit_lands_reruns_the_stage_once() {
|
||||
async fn a_crash_before_a_durable_finish_reruns_the_stage_once() {
|
||||
let context = test_context!();
|
||||
let mut server = RunningServer::start().await;
|
||||
let gate = context.temp_dir.join("two.gate");
|
||||
std::fs::write(&gate, "open").expect("the script gate is open from the start");
|
||||
let workspace = three_stage_bundle(&context, &gate);
|
||||
server.hold("commit", "two");
|
||||
server.hold("prepare", "two");
|
||||
let run_id = run_detached(&context, &server, &workspace);
|
||||
|
||||
wait_for_status(&server, &run_id, &["running"]).await;
|
||||
let worker = wait_for_worker(&run_id);
|
||||
server.wait_until_held(&run_id, "commit", "two");
|
||||
server.wait_until_held(&run_id, "prepare", "two");
|
||||
crash(&mut server, worker, None);
|
||||
|
||||
server.release("commit", "two");
|
||||
server.release("prepare", "two");
|
||||
server.launch().await;
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let (path, commits) = server.workspace_commits(&run_id);
|
||||
assert_eq!(subjects(&commits), three_stage_subjects(&run_id));
|
||||
let path = server.workspace_path(&run_id);
|
||||
assert!(!path.join(".git").exists());
|
||||
assert_eq!(
|
||||
std::fs::read_to_string(path.join("three.log")).expect("three copied the log"),
|
||||
"run\n",
|
||||
"the stage reran once, on the snapshot before it"
|
||||
"run\nrun\n",
|
||||
"the stage reran once in the existing workspace"
|
||||
);
|
||||
let store = server.petri_store().await;
|
||||
let outcome = engine::outcome_of(&store, &run_id)
|
||||
|
|
@ -1481,11 +1434,10 @@ async fn a_crash_before_the_commit_lands_reruns_the_stage_once() {
|
|||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A worker killed after the commit and its durable finish but before the
|
||||
/// platform record: the restart reconciles the record from the snapshot
|
||||
/// repository, the stage does not rerun, and one commit exists for it.
|
||||
/// A durable finish before its platform record does not rerun the stage;
|
||||
/// the resumed transition writes the missing metadata.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_crash_before_the_record_reconciles_it_from_the_run_branch() {
|
||||
async fn a_crash_before_the_platform_record_replays_the_record() {
|
||||
let context = test_context!();
|
||||
let mut server = RunningServer::start().await;
|
||||
let gate = context.temp_dir.join("two.gate");
|
||||
|
|
@ -1498,35 +1450,30 @@ async fn a_crash_before_the_record_reconciles_it_from_the_run_branch() {
|
|||
let worker = wait_for_worker(&run_id);
|
||||
server.wait_until_held(&run_id, "record", "two");
|
||||
let before = server.checkpoints(&run_id).await;
|
||||
assert_eq!(before.len(), 2, "start and one are recorded: {before:?}");
|
||||
assert!(
|
||||
before.is_empty(),
|
||||
"non-Git checkpoints do not carry commit SHAs"
|
||||
);
|
||||
crash(&mut server, worker, None);
|
||||
|
||||
server.release("record", "two");
|
||||
server.launch().await;
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let (path, commits) = server.workspace_commits(&run_id);
|
||||
assert_eq!(subjects(&commits), three_stage_subjects(&run_id));
|
||||
let path = server.workspace_path(&run_id);
|
||||
assert!(!path.join(".git").exists());
|
||||
assert_eq!(
|
||||
std::fs::read_to_string(path.join("three.log")).expect("three copied the log"),
|
||||
"run\n",
|
||||
"the stage with a durable finish did not rerun"
|
||||
);
|
||||
let checkpoints = server.checkpoints(&run_id).await;
|
||||
assert_eq!(checkpoints.len(), 5, "{checkpoints:?}");
|
||||
let keys: Vec<Option<CheckpointKey>> = checkpoints.iter().map(|(key, _)| Some(*key)).collect();
|
||||
let committed: Vec<Option<CheckpointKey>> = commits.iter().map(|(_, key)| *key).collect();
|
||||
assert_eq!(
|
||||
keys, committed,
|
||||
"the reconciled record names the one commit"
|
||||
);
|
||||
assert!(server.checkpoints(&run_id).await.is_empty());
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A workspace deleted while the run is down is restored from the snapshot
|
||||
/// repository, and the next stage sees the checkpoint's files.
|
||||
/// Execution records cannot reconstruct a deleted non-Git workspace.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_deleted_workspace_is_restored_from_its_snapshot() {
|
||||
async fn a_non_git_run_cannot_recover_deleted_files() {
|
||||
let context = test_context!();
|
||||
let mut server = RunningServer::start().await;
|
||||
let gate = context.temp_dir.join("two.gate");
|
||||
|
|
@ -1537,30 +1484,27 @@ async fn a_deleted_workspace_is_restored_from_its_snapshot() {
|
|||
let worker = wait_for_worker(&run_id);
|
||||
wait_until_gate_is_polled(&gate);
|
||||
crash(&mut server, worker, Some(&gate));
|
||||
let (path, commits) = server.workspace_commits(&run_id);
|
||||
assert_eq!(subjects(&commits), three_stage_subjects(&run_id)[..2]);
|
||||
let path = server.workspace_path(&run_id);
|
||||
assert!(path.join("one.txt").exists());
|
||||
std::fs::remove_dir_all(&path).expect("the workspace is deleted");
|
||||
|
||||
server.launch().await;
|
||||
wait_until_gate_is_polled(&gate);
|
||||
std::fs::write(&gate, "go").expect("the gate opens");
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let (restored, commits) = server.workspace_commits(&run_id);
|
||||
assert_eq!(restored, path);
|
||||
assert_eq!(subjects(&commits), three_stage_subjects(&run_id));
|
||||
assert_eq!(
|
||||
std::fs::read_to_string(restored.join("one.txt")).expect("one.txt was restored"),
|
||||
"one\n"
|
||||
wait_for_status(&server, &run_id, &["failed", "succeeded"]).await,
|
||||
"failed"
|
||||
);
|
||||
assert!(
|
||||
!path.join("one.txt").exists(),
|
||||
"deleted files cannot be reconstructed"
|
||||
);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A stage that fails on its own terms routes to its failure edge on the
|
||||
/// committed files, and after a crash once the failure is durable the
|
||||
/// route reruns on the same files.
|
||||
/// A failure route resumes with the files left in the existing workspace.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_failure_route_sees_the_same_committed_files_after_a_crash() {
|
||||
async fn a_failure_route_uses_the_surviving_files_after_a_crash() {
|
||||
let context = test_context!();
|
||||
let mut server = RunningServer::start().await;
|
||||
let gate = context.temp_dir.join("fix.gate");
|
||||
|
|
@ -1589,94 +1533,52 @@ async fn a_failure_route_sees_the_same_committed_files_after_a_crash() {
|
|||
std::fs::write(&gate, "go").expect("the gate opens");
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let (path, commits) = server.workspace_commits(&run_id);
|
||||
assert_eq!(subjects(&commits), vec![
|
||||
format!("fabro({run_id}): start (success)"),
|
||||
format!("fabro({run_id}): work (failure)"),
|
||||
format!("fabro({run_id}): fix (success)"),
|
||||
format!("fabro({run_id}): exit (success)"),
|
||||
]);
|
||||
let path = server.workspace_path(&run_id);
|
||||
assert_eq!(
|
||||
std::fs::read_to_string(path.join("out.txt")).expect("out.txt"),
|
||||
"partial\nfixed\n"
|
||||
);
|
||||
assert_eq!(
|
||||
std::fs::read_to_string(path.join("fix.log")).expect("fix.log"),
|
||||
"run\n",
|
||||
"the route saw the failed stage's files, not its own interrupted attempt's"
|
||||
"run\nrun\n",
|
||||
"the route continued with the existing files, including partial output"
|
||||
);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A checkpoint commit that fails ends the run: `checkpoint_failed` is
|
||||
/// recorded, no route runs, the run is reported failed, and a restart
|
||||
/// leaves it failed without launching a worker.
|
||||
/// A failed non-Git execution stays terminal across a server restart.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_failed_checkpoint_fails_the_run_and_a_restart_leaves_it_failed() {
|
||||
async fn a_failed_non_git_run_stays_failed_after_a_restart() {
|
||||
let context = test_context!();
|
||||
let mut server = RunningServer::start().await;
|
||||
let workspace = write_petri_workflow(
|
||||
&context,
|
||||
"digraph Wreck {\n graph [goal=\"Wreck the repository\", default_max_retries=0]\n start \
|
||||
[shape=Mdiamond]\n exit [shape=Msquare]\n wreck [shape=parallelogram, script=\"rm -rf \
|
||||
.git && echo garbage > .git\"]\n next [shape=parallelogram, script=\"echo next > \
|
||||
next.txt\"]\n fix [shape=parallelogram, script=\"echo fix > fix.txt\"]\n start -> wreck \
|
||||
-> next -> exit\n wreck -> fix [condition=\"outcome=failed\"]\n fix -> exit\n}\n",
|
||||
r#"digraph Failure {
|
||||
graph [goal="Fail", default_max_retries=0]
|
||||
start [shape=Mdiamond]
|
||||
fail [shape=parallelogram, script="exit 1", goal_gate=true]
|
||||
exit [shape=Msquare]
|
||||
start -> fail -> exit
|
||||
}"#,
|
||||
);
|
||||
let run_id = run_detached(&context, &server, &workspace);
|
||||
let status = wait_for_status(&server, &run_id, &["succeeded", "failed"]).await;
|
||||
let run = run_json(&server, &format!("runs/{run_id}")).await;
|
||||
assert_eq!(status, "failed", "run: {run}");
|
||||
// The run's failure travels on the stream as the platform record of
|
||||
// its terminal lifecycle transition, with the failure's message as the
|
||||
// reason.
|
||||
let failures: Vec<String> = settled_stream(&server, &run_id)
|
||||
.await
|
||||
.iter()
|
||||
.filter_map(|line| {
|
||||
let record = &line["item"]["record"];
|
||||
(line["kind"] == "platform"
|
||||
&& record["kind"] == "run.lifecycle"
|
||||
&& record["transition"] == "failed")
|
||||
.then(|| record["reason"].as_str().unwrap_or_default().to_string())
|
||||
})
|
||||
.collect();
|
||||
assert_eq!(failures.len(), 1, "{failures:?}");
|
||||
assert!(
|
||||
failures[0].contains("checkpoint commit of `wreck` failed"),
|
||||
"{failures:?}"
|
||||
);
|
||||
let scopes = server.petri_run_dir(&run_id).join("scopes");
|
||||
let work = std::fs::read_dir(&scopes)
|
||||
.expect("the scopes directory lists")
|
||||
.map(|entry| entry.expect("an entry reads").path().join("work"))
|
||||
.next()
|
||||
.expect("one workspace");
|
||||
assert!(!work.join("next.txt").exists(), "no route ran");
|
||||
assert!(!work.join("fix.txt").exists(), "no route ran");
|
||||
|
||||
let store = server.petri_store().await;
|
||||
let outcome = engine::outcome_of(&store, &run_id)
|
||||
.await
|
||||
.expect("the run's Petri record inspects");
|
||||
assert_ne!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
let checkpoints = server.checkpoints(&run_id).await;
|
||||
assert_eq!(
|
||||
checkpoints.len(),
|
||||
1,
|
||||
"only start was recorded: {checkpoints:?}"
|
||||
wait_for_status(&server, &run_id, &["succeeded", "failed"]).await,
|
||||
"failed"
|
||||
);
|
||||
|
||||
// The restart finds the run terminal and launches nothing for it.
|
||||
assert!(server.checkpoints(&run_id).await.is_empty());
|
||||
let deadline = Instant::now() + Duration::from_secs(10);
|
||||
while worker_pid(&run_id).is_some() {
|
||||
assert!(
|
||||
Instant::now() < deadline,
|
||||
"the terminal worker did not exit"
|
||||
);
|
||||
tokio::time::sleep(POLL).await;
|
||||
}
|
||||
server.kill();
|
||||
server.launch().await;
|
||||
assert_eq!(run_status(&server, &run_id).await, "failed");
|
||||
std::thread::sleep(Duration::from_secs(1));
|
||||
assert_eq!(
|
||||
worker_pid(&run_id),
|
||||
None,
|
||||
"no worker was launched for the failed run"
|
||||
);
|
||||
assert_eq!(worker_pid(&run_id), None);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
|
|
@ -1797,3 +1699,43 @@ async fn built_in_host_runs_and_prunes_without_plugins() {
|
|||
assert!(!scope.exists(), "prune removed the managed Host workspace");
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// Retry works without a Git checkpoint and starts the workflow again.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_retry_without_git_checkpoints_starts_the_workflow_again() {
|
||||
let context = test_context!();
|
||||
let server = RunningServer::start().await;
|
||||
let counter = context.temp_dir.join("retry-runs.txt");
|
||||
let workspace = write_petri_workflow(
|
||||
&context,
|
||||
&format!(
|
||||
r#"digraph Retry {{
|
||||
graph [goal="Retry from the start", default_max_retries=0]
|
||||
start [shape=Mdiamond]
|
||||
write [shape=parallelogram, script="echo run >> '{}'", goal_gate=true]
|
||||
exit [shape=Msquare]
|
||||
start -> write -> exit
|
||||
}}"#,
|
||||
counter.display()
|
||||
),
|
||||
);
|
||||
let original = run_detached(&context, &server, &workspace);
|
||||
wait_for_success(&server, &original).await;
|
||||
assert!(server.checkpoints(&original).await.is_empty());
|
||||
let response = fabro_test::test_http_client()
|
||||
.post(format!(
|
||||
"{}/api/v1/runs/{original}/retry",
|
||||
server.api_base_url
|
||||
))
|
||||
.bearer_auth(TEST_DEV_TOKEN)
|
||||
.send()
|
||||
.await
|
||||
.unwrap();
|
||||
let retried = expect_reqwest_json(response, fabro_http::StatusCode::CREATED, "retry").await;
|
||||
let run_id = retried["id"].as_str().expect("the retry run id");
|
||||
assert_ne!(run_id, original);
|
||||
wait_for_success(&server, run_id).await;
|
||||
assert_eq!(std::fs::read_to_string(counter).unwrap(), "run\nrun\n");
|
||||
assert!(server.checkpoints(run_id).await.is_empty());
|
||||
server.shutdown();
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1,9 +1,6 @@
|
|||
//! Petri runs on a Docker environment through a real server and its
|
||||
//! worker: the workspace lives inside the run's container, every stage's
|
||||
//! checkpoint is committed there and published to the snapshot repository
|
||||
//! on the host, and a restart brings the container's workspace back to the
|
||||
//! snapshot its durable state names, in the retained container or in a
|
||||
//! fresh one when the old one is gone.
|
||||
//! worker: a non-Git workspace survives a worker restart in its retained
|
||||
//! container. A deleted container fails resume; Fabro does not reconstruct it.
|
||||
//!
|
||||
//! An Ask Fabro session on a finished Docker run attaches to the container
|
||||
//! Petri created and reads a file the workflow wrote there, its model the
|
||||
|
|
@ -19,10 +16,9 @@
|
|||
)]
|
||||
|
||||
use std::env;
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::path::PathBuf;
|
||||
use std::process::{Command, Stdio};
|
||||
|
||||
use fabro_petri::checkpoint::CheckpointKey;
|
||||
use fabro_static::EnvVars;
|
||||
use fabro_test::{
|
||||
TwinScenario, TwinScenarios, TwinToolCall, expect_reqwest_json, test_context, twin_openai,
|
||||
|
|
@ -129,62 +125,6 @@ fn cleanup(run_id: &str) {
|
|||
}
|
||||
}
|
||||
|
||||
/// The snapshot repository of the run's one workspace, on the host.
|
||||
fn snapshot_repository(server: &RunningServer, run_id: &str) -> PathBuf {
|
||||
let snapshots = server.petri_run_dir(run_id).join("snapshots");
|
||||
let mut repositories: Vec<PathBuf> = std::fs::read_dir(&snapshots)
|
||||
.expect("the snapshots directory lists")
|
||||
.map(|entry| entry.expect("an entry reads").path())
|
||||
.filter(|path| path.extension().is_some_and(|extension| extension == "git"))
|
||||
.collect();
|
||||
assert_eq!(repositories.len(), 1, "one workspace: {repositories:?}");
|
||||
repositories.remove(0)
|
||||
}
|
||||
|
||||
/// `git` in a repository on the host, its stdout.
|
||||
fn git(repository: &Path, args: &[&str]) -> String {
|
||||
let output = Command::new("git")
|
||||
.args(args)
|
||||
.current_dir(repository)
|
||||
.output()
|
||||
.expect("git runs");
|
||||
assert!(
|
||||
output.status.success(),
|
||||
"git {args:?} failed: {}",
|
||||
String::from_utf8_lossy(&output.stderr)
|
||||
);
|
||||
String::from_utf8_lossy(&output.stdout).trim().to_owned()
|
||||
}
|
||||
|
||||
/// The commits the snapshot repository holds, oldest first, as
|
||||
/// `(sha, subject, key)`.
|
||||
fn snapshot_commits(repository: &Path) -> Vec<(String, String, Option<CheckpointKey>)> {
|
||||
let log = git(repository, &[
|
||||
"log",
|
||||
"--topo-order",
|
||||
"--reverse",
|
||||
"--all",
|
||||
"--format=%H%x00%s%x00%B%x1e",
|
||||
]);
|
||||
log.split('\u{1e}')
|
||||
.filter(|entry| !entry.trim().is_empty())
|
||||
.map(|entry| {
|
||||
let mut parts = entry.trim_start().splitn(3, '\0');
|
||||
let sha = parts.next().unwrap_or_default().to_string();
|
||||
let subject = parts.next().unwrap_or_default().to_string();
|
||||
let body = parts.next().unwrap_or_default();
|
||||
(sha, subject, CheckpointKey::from_message(body))
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn subjects(commits: &[(String, String, Option<CheckpointKey>)]) -> Vec<&str> {
|
||||
commits
|
||||
.iter()
|
||||
.map(|(_, subject, _)| subject.as_str())
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Two command stages: `one` writes a file; `two` checks it is the one
|
||||
/// `one` wrote, that nothing else is in the workspace beside the
|
||||
/// repository's own files, and writes another.
|
||||
|
|
@ -198,62 +138,9 @@ fn two_stage_bundle(context: &fabro_test::TestContext) -> PathBuf {
|
|||
)
|
||||
}
|
||||
|
||||
/// The commit subjects one run of the two-stage bundle produces.
|
||||
fn two_stage_subjects(run_id: &str) -> Vec<String> {
|
||||
["start", "one", "two", "exit"]
|
||||
.iter()
|
||||
.map(|node| format!("fabro({run_id}): {node} (success)"))
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// The checkpoint records and the published snapshots name the same
|
||||
/// commits, and the two stages' trees hold their files.
|
||||
fn assert_snapshots_complete(server: &RunningServer, run_id: &str, repository: &Path) {
|
||||
let commits = snapshot_commits(repository);
|
||||
assert_eq!(subjects(&commits), two_stage_subjects(run_id));
|
||||
let (one, _, _) = &commits[1];
|
||||
let (two, _, _) = &commits[2];
|
||||
assert_eq!(git(repository, &["show", &format!("{one}:one.txt")]), "one");
|
||||
assert_eq!(git(repository, &["show", &format!("{two}:two.txt")]), "two");
|
||||
let refs = git(repository, &[
|
||||
"for-each-ref",
|
||||
"--format=%(objectname)",
|
||||
"refs/checkpoints/",
|
||||
]);
|
||||
let mut published: Vec<&str> = refs.lines().collect();
|
||||
published.sort_unstable();
|
||||
let checkpoints = futures_lite_block_on(server.checkpoints(run_id));
|
||||
let mut recorded: Vec<&str> = checkpoints.iter().map(|(_, sha)| sha.as_str()).collect();
|
||||
recorded.sort_unstable();
|
||||
assert_eq!(
|
||||
published, recorded,
|
||||
"every record names a published snapshot"
|
||||
);
|
||||
assert_eq!(checkpoints.len(), 4, "{checkpoints:?}");
|
||||
}
|
||||
|
||||
/// Wait on a future from a synchronous helper inside a multi-thread test.
|
||||
fn futures_lite_block_on<T>(future: impl std::future::Future<Output = T>) -> T {
|
||||
tokio::task::block_in_place(|| tokio::runtime::Handle::current().block_on(future))
|
||||
}
|
||||
|
||||
/// What the worker's log says it did to the sandbox workspace at resume.
|
||||
fn restore_actions(server: &RunningServer, run_id: &str) -> Vec<String> {
|
||||
let log = std::fs::read_to_string(server.worker_log(run_id)).unwrap_or_default();
|
||||
log.lines()
|
||||
.filter(|line| line.contains("workspace brought to its durable snapshot"))
|
||||
.filter_map(|line| {
|
||||
line.split_whitespace()
|
||||
.find_map(|word| word.strip_prefix("action=").map(str::to_owned))
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// A run on Docker: every stage is committed inside the container, each
|
||||
/// checkpoint is published to the snapshot repository on the host, and
|
||||
/// nothing of the workspace is on the host.
|
||||
/// A non-Git Docker run keeps files only inside its container.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_docker_run_publishes_every_stages_checkpoint_from_the_container() {
|
||||
async fn a_non_git_docker_run_keeps_its_workspace_in_the_container() {
|
||||
if !fabro_test::docker_available() {
|
||||
return;
|
||||
}
|
||||
|
|
@ -265,8 +152,8 @@ async fn a_docker_run_publishes_every_stages_checkpoint_from_the_container() {
|
|||
]);
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let repository = snapshot_repository(&server, &run_id);
|
||||
assert_snapshots_complete(&server, &run_id, &repository);
|
||||
assert!(server.checkpoints(&run_id).await.is_empty());
|
||||
assert!(!server.petri_run_dir(&run_id).join("snapshots").exists());
|
||||
assert!(
|
||||
!server.petri_run_dir(&run_id).join("scopes").exists(),
|
||||
"no workspace is on the host"
|
||||
|
|
@ -279,12 +166,9 @@ async fn a_docker_run_publishes_every_stages_checkpoint_from_the_container() {
|
|||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A worker killed after the first stage's durable finish, with the
|
||||
/// container's workspace changed behind Petri's back: the restart resumes
|
||||
/// on the retained container, the workspace is reset to the snapshot, and
|
||||
/// the second stage sees the first stage's files and nothing else.
|
||||
/// A non-Git run resumes in the retained container with its existing files.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_retained_container_whose_workspace_drifted_is_reset_on_restart() {
|
||||
async fn a_non_git_run_resumes_in_its_retained_container() {
|
||||
if !fabro_test::docker_available() {
|
||||
return;
|
||||
}
|
||||
|
|
@ -300,10 +184,7 @@ async fn a_retained_container_whose_workspace_drifted_is_reset_on_restart() {
|
|||
let worker = wait_for_worker(&run_id);
|
||||
server.wait_until_held(&run_id, "record", "one");
|
||||
let container = container_of(&run_id).expect("the run's container exists");
|
||||
docker_exec(
|
||||
&container,
|
||||
"echo junk > one.txt && echo stray > stray.txt && git status --porcelain",
|
||||
);
|
||||
docker_exec(&container, "echo marker > marker.txt && test ! -d .git");
|
||||
crash(&mut server, worker, None);
|
||||
|
||||
server.release("record", "one");
|
||||
|
|
@ -317,19 +198,17 @@ async fn a_retained_container_whose_workspace_drifted_is_reset_on_restart() {
|
|||
Some(container.as_str()),
|
||||
"the run continued in its retained container"
|
||||
);
|
||||
assert_eq!(restore_actions(&server, &run_id), vec!["Reset".to_string()]);
|
||||
let repository = snapshot_repository(&server, &run_id);
|
||||
assert_snapshots_complete(&server, &run_id, &repository);
|
||||
|
||||
assert!(server.checkpoints(&run_id).await.is_empty());
|
||||
assert!(!server.petri_run_dir(&run_id).join("snapshots").exists());
|
||||
cleanup(&run_id);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A worker killed after the first stage's durable finish, with the
|
||||
/// container removed while the run is down: the restart gets a fresh
|
||||
/// container, the workspace is restored into it from the snapshot
|
||||
/// repository, and the second stage sees the first stage's files.
|
||||
/// A deleted container fails resume instead of silently creating an empty
|
||||
/// replacement workspace.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_lost_container_is_replaced_and_its_workspace_restored_from_the_snapshot() {
|
||||
async fn a_lost_container_is_not_replaced_on_resume() {
|
||||
if !fabro_test::docker_available() {
|
||||
return;
|
||||
}
|
||||
|
|
@ -351,15 +230,16 @@ async fn a_lost_container_is_replaced_and_its_workspace_restored_from_the_snapsh
|
|||
|
||||
server.release("record", "one");
|
||||
server.launch().await;
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let fresh = container_of(&run_id).expect("a fresh container was created");
|
||||
assert_ne!(fresh, container);
|
||||
assert_eq!(restore_actions(&server, &run_id), vec![
|
||||
"Restored".to_string()
|
||||
]);
|
||||
let repository = snapshot_repository(&server, &run_id);
|
||||
assert_snapshots_complete(&server, &run_id, &repository);
|
||||
assert_eq!(
|
||||
wait_for_status(&server, &run_id, &["failed", "succeeded"]).await,
|
||||
"failed"
|
||||
);
|
||||
assert_eq!(
|
||||
container_of(&run_id),
|
||||
None,
|
||||
"resume must not replace the lost container"
|
||||
);
|
||||
assert!(!server.petri_run_dir(&run_id).join("snapshots").exists());
|
||||
cleanup(&run_id);
|
||||
server.shutdown();
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1,13 +1,10 @@
|
|||
//! Fork, rewind, retry and the timeline over Petri runs, through a real
|
||||
//! server and its worker subprocess.
|
||||
//!
|
||||
//! The harness is `petri.rs`'s: a foreground server on disk storage, a run
|
||||
//! created and started with `fabro run --detach`, executed by the worker
|
||||
//! the server launches over the HTTP run store. Every checkpoint of a run
|
||||
//! is a commit on its run branch and a `checkpoint` platform record at its
|
||||
//! Petri position; the timeline lists them, and a fork seeds a new run from
|
||||
//! the source's records up to one of them, whose worker restores the
|
||||
//! checkpoint's files into a fresh workspace and continues from there.
|
||||
//! Non-Git runs have execution checkpoints, can retry from the start, and
|
||||
//! refuse fork/rewind because no published checkpoint can seed a workspace.
|
||||
//! Git-backed forks are covered by fabro-petri's sandbox/remote integration
|
||||
//! test.
|
||||
|
||||
#![expect(
|
||||
clippy::disallowed_methods,
|
||||
|
|
@ -15,7 +12,7 @@
|
|||
)]
|
||||
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::process::{Command, Output};
|
||||
use std::process::Output;
|
||||
|
||||
use fabro_test::test_context;
|
||||
|
||||
|
|
@ -100,7 +97,7 @@ async fn timeline(server: &RunningServer, run_id: &str) -> serde_json::Value {
|
|||
}
|
||||
|
||||
/// The timeline's entries as `(node, execution, firing, attempt, sha)`.
|
||||
fn entries(timeline: &serde_json::Value) -> Vec<(String, u64, u64, u64, String)> {
|
||||
fn entries(timeline: &serde_json::Value) -> Vec<(String, u64, u64, u64, Option<String>)> {
|
||||
timeline["entries"]
|
||||
.as_array()
|
||||
.expect("the timeline has entries")
|
||||
|
|
@ -111,10 +108,7 @@ fn entries(timeline: &serde_json::Value) -> Vec<(String, u64, u64, u64, String)>
|
|||
entry["execution"].as_u64().expect("an execution"),
|
||||
entry["firing"].as_u64().expect("a firing"),
|
||||
entry["attempt"].as_u64().expect("an attempt"),
|
||||
entry["run_commit_sha"]
|
||||
.as_str()
|
||||
.expect("a checkpoint commit")
|
||||
.to_string(),
|
||||
entry["run_commit_sha"].as_str().map(str::to_owned),
|
||||
)
|
||||
})
|
||||
.collect()
|
||||
|
|
@ -138,121 +132,30 @@ fn workspace(server: &RunningServer, run_id: &str) -> PathBuf {
|
|||
workspaces.remove(0)
|
||||
}
|
||||
|
||||
/// The subjects of the commits on the workspace's branch, oldest first.
|
||||
fn commit_subjects(workspace: &Path) -> Vec<String> {
|
||||
let output = Command::new("git")
|
||||
.args(["log", "--reverse", "--format=%s"])
|
||||
.current_dir(workspace)
|
||||
.output()
|
||||
.expect("git runs");
|
||||
assert!(
|
||||
output.status.success(),
|
||||
"git log failed: {}",
|
||||
String::from_utf8_lossy(&output.stderr)
|
||||
);
|
||||
String::from_utf8_lossy(&output.stdout)
|
||||
.lines()
|
||||
.map(ToOwned::to_owned)
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn current_branch(workspace: &Path) -> String {
|
||||
let output = Command::new("git")
|
||||
.args(["rev-parse", "--abbrev-ref", "HEAD"])
|
||||
.current_dir(workspace)
|
||||
.output()
|
||||
.expect("git runs");
|
||||
String::from_utf8_lossy(&output.stdout).trim().to_string()
|
||||
}
|
||||
|
||||
fn read(workspace: &Path, name: &str) -> String {
|
||||
std::fs::read_to_string(workspace.join(name))
|
||||
.unwrap_or_else(|err| panic!("{name} in {}: {err}", workspace.display()))
|
||||
}
|
||||
|
||||
/// A three-stage run forked at its first stage's checkpoint continues with
|
||||
/// the other two in a new run: the new run keeps the source's records up to
|
||||
/// `one`, its workspace holds `one`'s file restored from the checkpoint,
|
||||
/// and `two` and `three` run on it and commit on the new run branch. The
|
||||
/// new run says where it came from.
|
||||
/// A local-folder run cannot fork without a published checkpoint.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_fork_at_the_first_stage_continues_with_the_rest_on_its_files() {
|
||||
async fn a_non_git_fork_is_refused_without_creating_a_run() {
|
||||
let context = test_context!();
|
||||
let server = RunningServer::start().await;
|
||||
let bundle = write_petri_workflow(&context, &three_stage_dot());
|
||||
let source = run_detached(&context, &server, &bundle);
|
||||
wait_for_success(&server, &source).await;
|
||||
let source_timeline = timeline(&server, &source).await;
|
||||
assert_eq!(nodes(&source_timeline), [
|
||||
"start", "one", "two", "three", "exit"
|
||||
]);
|
||||
let source_entries = entries(&source_timeline);
|
||||
let (_, one_execution, one_firing, _, one_sha) = source_entries[1].clone();
|
||||
|
||||
let forked = cli_json(&context, &server, &["fork", &source, "one", "--json"]);
|
||||
assert_eq!(forked["source_run_id"], source);
|
||||
assert_eq!(forked["target"], "@2");
|
||||
assert_eq!(forked["execution"].as_u64(), Some(one_execution));
|
||||
assert_eq!(forked["firing"].as_u64(), Some(one_firing));
|
||||
assert_eq!(forked["checkpoint_sha"], one_sha);
|
||||
assert_eq!(forked["rerun_last"], false);
|
||||
let fork = forked["new_run_id"]
|
||||
.as_str()
|
||||
.expect("the new run id")
|
||||
.to_string();
|
||||
assert_ne!(fork, source);
|
||||
wait_for_success(&server, &fork).await;
|
||||
|
||||
// The fork's timeline: the source's checkpoints up to `one`, then its own.
|
||||
let fork_timeline = timeline(&server, &fork).await;
|
||||
assert_eq!(nodes(&fork_timeline), [
|
||||
"start", "one", "two", "three", "exit"
|
||||
]);
|
||||
let fork_entries = entries(&fork_timeline);
|
||||
assert_eq!(fork_entries[..2], source_entries[..2]);
|
||||
assert_ne!(fork_entries[2].4, source_entries[2].4);
|
||||
assert_eq!(fork_timeline["forked_from"]["source_run_id"], source);
|
||||
assert_eq!(
|
||||
fork_timeline["forked_from"]["execution"].as_u64(),
|
||||
Some(one_execution)
|
||||
);
|
||||
assert_eq!(
|
||||
fork_timeline["forked_from"]["firing"].as_u64(),
|
||||
Some(one_firing)
|
||||
);
|
||||
assert_eq!(fork_timeline["forked_from"]["rerun_last"], false);
|
||||
assert!(source_timeline["forked_from"].is_null());
|
||||
|
||||
// The fork's workspace: `one.txt` restored from the checkpoint, the
|
||||
// rest made by the fork's own stages, on the fork's run branch after the
|
||||
// source's commits.
|
||||
let fork_workspace = workspace(&server, &fork);
|
||||
assert_eq!(read(&fork_workspace, "one.txt"), "one\n");
|
||||
assert_eq!(read(&fork_workspace, "two.txt"), "two\n");
|
||||
assert_eq!(read(&fork_workspace, "three.txt"), "three\n");
|
||||
assert_eq!(current_branch(&fork_workspace), format!("fabro/run/{fork}"));
|
||||
assert_eq!(commit_subjects(&fork_workspace), [
|
||||
format!("fabro({source}): start (success)"),
|
||||
format!("fabro({source}): one (success)"),
|
||||
format!("fabro({fork}): two (success)"),
|
||||
format!("fabro({fork}): three (success)"),
|
||||
format!("fabro({fork}): exit (success)"),
|
||||
]);
|
||||
|
||||
// The projection names the origin on both sides.
|
||||
let state = run_json(&server, &format!("runs/{fork}/state")).await;
|
||||
assert_eq!(state["forked_from"]["source_run_id"], source);
|
||||
assert_eq!(state["spec"]["fork_source_ref"]["source_run_id"], source);
|
||||
assert_eq!(state["spec"]["fork_source_ref"]["checkpoint_sha"], one_sha);
|
||||
assert!(state["retried_from"].is_null());
|
||||
let refused = cli(&context, &server, &["fork", &source, "one", "--json"]);
|
||||
assert!(!refused.status.success());
|
||||
assert!(String::from_utf8_lossy(&refused.stderr).contains("GitHub-backed run"));
|
||||
let summary = run_json(&server, &format!("runs/{source}")).await;
|
||||
assert!(summary["superseded_by"].is_null());
|
||||
assert_eq!(summary["lifecycle"]["archived"], false);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A retry starts the workflow over in a new run: every stage runs again
|
||||
/// in a fresh workspace, so a transient failure passes the second time.
|
||||
/// A retry starts the workflow again in a new workspace. A transient
|
||||
/// failure can succeed on this second execution.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_retry_starts_over_and_succeeds_when_the_failure_was_transient() {
|
||||
let context = test_context!();
|
||||
|
|
@ -284,13 +187,7 @@ async fn a_retry_starts_over_and_succeeds_when_the_failure_was_transient() {
|
|||
assert_eq!(read(&retry_workspace, "one.txt"), "one\n");
|
||||
assert_eq!(read(&retry_workspace, "flaky.txt"), "flaky\n");
|
||||
assert_eq!(read(&retry_workspace, "three.txt"), "three\n");
|
||||
assert_eq!(commit_subjects(&retry_workspace), [
|
||||
format!("fabro({retry}): start (success)"),
|
||||
format!("fabro({retry}): one (success)"),
|
||||
format!("fabro({retry}): flaky (success)"),
|
||||
format!("fabro({retry}): three (success)"),
|
||||
format!("fabro({retry}): exit (success)"),
|
||||
]);
|
||||
assert!(!retry_workspace.join(".git").exists());
|
||||
let state = run_json(&server, &format!("runs/{retry}/state")).await;
|
||||
assert_eq!(state["retried_from"], source);
|
||||
assert!(state["spec"]["fork_source_ref"].is_null());
|
||||
|
|
@ -303,106 +200,40 @@ async fn a_retry_starts_over_and_succeeds_when_the_failure_was_transient() {
|
|||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A rewind forks the run and supersedes it: the source is archived and
|
||||
/// names the run that replaced it.
|
||||
/// Refusing rewind must leave the original run available and unarchived.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_rewind_supersedes_its_source() {
|
||||
async fn a_non_git_rewind_is_refused_without_archiving_its_source() {
|
||||
let context = test_context!();
|
||||
let server = RunningServer::start().await;
|
||||
let bundle = write_petri_workflow(&context, &three_stage_dot());
|
||||
let source = run_detached(&context, &server, &bundle);
|
||||
wait_for_success(&server, &source).await;
|
||||
|
||||
// Without a target the command lists the timeline instead.
|
||||
let listed = cli_json(&context, &server, &["rewind", &source, "--json"]);
|
||||
assert_eq!(
|
||||
listed["entries"]
|
||||
.as_array()
|
||||
.map(Vec::len)
|
||||
.expect("a timeline"),
|
||||
5
|
||||
);
|
||||
|
||||
let rewound = cli_json(&context, &server, &["rewind", &source, "@3", "--json"]);
|
||||
assert_eq!(rewound["source_run_id"], source);
|
||||
assert_eq!(rewound["target"], "@3");
|
||||
assert_eq!(rewound["archived"], true);
|
||||
assert!(rewound["archive_error"].is_null());
|
||||
assert_eq!(rewound["status"], 200);
|
||||
let replacement = rewound["new_run_id"]
|
||||
.as_str()
|
||||
.expect("the new run id")
|
||||
.to_string();
|
||||
|
||||
let summary = run_json(&server, &format!("runs/{source}")).await;
|
||||
assert_eq!(summary["superseded_by"], replacement);
|
||||
assert_eq!(summary["lifecycle"]["archived"], true);
|
||||
let state = run_json(&server, &format!("runs/{source}/state")).await;
|
||||
assert_eq!(state["superseded_by"], replacement);
|
||||
|
||||
wait_for_success(&server, &replacement).await;
|
||||
assert_eq!(nodes(&timeline(&server, &replacement).await), [
|
||||
"start", "one", "two", "three", "exit"
|
||||
]);
|
||||
let replacement_workspace = workspace(&server, &replacement);
|
||||
assert_eq!(read(&replacement_workspace, "two.txt"), "two\n");
|
||||
assert_eq!(commit_subjects(&replacement_workspace), [
|
||||
format!("fabro({source}): start (success)"),
|
||||
format!("fabro({source}): one (success)"),
|
||||
format!("fabro({source}): two (success)"),
|
||||
format!("fabro({replacement}): three (success)"),
|
||||
format!("fabro({replacement}): exit (success)"),
|
||||
]);
|
||||
|
||||
// An archived run is neither forked nor rewound again.
|
||||
let refused = cli(&context, &server, &["fork", &source, "@2"]);
|
||||
assert_eq!(listed["entries"].as_array().unwrap().len(), 5);
|
||||
let refused = cli(&context, &server, &["rewind", &source, "@3", "--json"]);
|
||||
assert!(!refused.status.success());
|
||||
assert!(
|
||||
String::from_utf8_lossy(&refused.stderr).contains("archived"),
|
||||
"stderr:\n{}",
|
||||
String::from_utf8_lossy(&refused.stderr)
|
||||
);
|
||||
assert!(String::from_utf8_lossy(&refused.stderr).contains("GitHub-backed run"));
|
||||
let summary = run_json(&server, &format!("runs/{source}")).await;
|
||||
assert!(summary["superseded_by"].is_null());
|
||||
assert_eq!(summary["lifecycle"]["archived"], false);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// The timeline lists every checkpoint the run recorded, in order, with
|
||||
/// its position and commit, as the CLI prints it and as the API serves it.
|
||||
/// its position and optional commit, in both the CLI and API.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn the_timeline_lists_every_checkpoint_with_its_commit() {
|
||||
async fn the_timeline_lists_non_git_checkpoints_without_commit_shas() {
|
||||
let context = test_context!();
|
||||
let server = RunningServer::start().await;
|
||||
let bundle = write_petri_workflow(&context, &three_stage_dot());
|
||||
let run_id = run_detached(&context, &server, &bundle);
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let recorded = server.checkpoints(&run_id).await;
|
||||
assert!(server.checkpoints(&run_id).await.is_empty());
|
||||
let listed = cli_json(&context, &server, &["timeline", &run_id, "--json"]);
|
||||
let listed_entries: Vec<(u64, u64, u64, String)> = listed["entries"]
|
||||
.as_array()
|
||||
.expect("entries")
|
||||
.iter()
|
||||
.map(|entry| {
|
||||
(
|
||||
entry["execution"].as_u64().expect("execution"),
|
||||
entry["firing"].as_u64().expect("firing"),
|
||||
entry["attempt"].as_u64().expect("attempt"),
|
||||
entry["run_commit_sha"].as_str().expect("sha").to_string(),
|
||||
)
|
||||
})
|
||||
.collect();
|
||||
let recorded_entries: Vec<(u64, u64, u64, String)> = recorded
|
||||
.iter()
|
||||
.map(|(key, sha)| {
|
||||
(
|
||||
key.execution,
|
||||
key.firing,
|
||||
u64::from(key.attempt),
|
||||
sha.clone(),
|
||||
)
|
||||
})
|
||||
.collect();
|
||||
assert_eq!(listed_entries, recorded_entries);
|
||||
let listed_entries = entries(&listed);
|
||||
assert_eq!(listed_entries.len(), 5);
|
||||
assert!(listed_entries.iter().all(|entry| entry.4.is_none()));
|
||||
let ordinals: Vec<u64> = listed["entries"]
|
||||
.as_array()
|
||||
.expect("entries")
|
||||
|
|
@ -437,7 +268,7 @@ async fn the_timeline_lists_every_checkpoint_with_its_commit() {
|
|||
/// A checkpoint inside a parallel branch is not a fork position: Petri
|
||||
/// refuses it, and the refusal says why. The join, in the root, is.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_fork_inside_a_parallel_branch_is_refused() {
|
||||
async fn a_non_git_parallel_run_cannot_be_forked() {
|
||||
let context = test_context!();
|
||||
let server = RunningServer::start().await;
|
||||
let bundle = write_petri_workflow(&context, ¶llel_dot());
|
||||
|
|
@ -459,7 +290,7 @@ async fn a_fork_inside_a_parallel_branch_is_refused() {
|
|||
);
|
||||
let stderr = String::from_utf8_lossy(&refused.stderr);
|
||||
assert!(
|
||||
stderr.contains("inside a child invocation cannot be forked"),
|
||||
stderr.contains("GitHub-backed run"),
|
||||
"the refusal names the branch:\n{stderr}"
|
||||
);
|
||||
// Nothing was started for it.
|
||||
|
|
@ -473,25 +304,7 @@ async fn a_fork_inside_a_parallel_branch_is_refused() {
|
|||
.collect();
|
||||
assert!(ids.iter().all(|id| *id == source), "runs: {ids:?}");
|
||||
|
||||
// A fork at the stage after the join continues from the root.
|
||||
let done = all
|
||||
.iter()
|
||||
.zip(1_u64..)
|
||||
.find(|((node, ..), _)| node == "done")
|
||||
.map(|(_, ordinal)| ordinal)
|
||||
.expect("`done` has a checkpoint");
|
||||
let forked = cli_json(&context, &server, &[
|
||||
"fork",
|
||||
&source,
|
||||
&format!("@{done}"),
|
||||
"--json",
|
||||
]);
|
||||
let fork = forked["new_run_id"]
|
||||
.as_str()
|
||||
.expect("the new run id")
|
||||
.to_string();
|
||||
wait_for_success(&server, &fork).await;
|
||||
let fork_workspace = workspace(&server, &fork);
|
||||
assert_eq!(read(&fork_workspace, "done.txt"), "done\n");
|
||||
let refused = cli(&context, &server, &["fork", &source, "done"]);
|
||||
assert!(!refused.status.success());
|
||||
server.shutdown();
|
||||
}
|
||||
|
|
|
|||
|
|
@ -5,8 +5,8 @@
|
|||
//! the stages the projector folded them onto. A fork resolves a target on
|
||||
//! that timeline (`@ordinal`, a node, or `node@visit`; the latest checkpoint
|
||||
//! by default), creates the new run's row (`fabro_workflow::operations`),
|
||||
//! seeds its records, checkpoints, snapshots and run branch from the source
|
||||
//! (`fabro_petri::fork`), and queues it in resume mode, so its worker
|
||||
//! seeds its execution records, checkpoint metadata and run branch from the
|
||||
//! source (`fabro_petri::fork`), and queues it in resume mode, so its worker
|
||||
//! acquires a fresh workspace, restores the checkpoint's commit into it and
|
||||
//! continues from the position. A rewind is a fork of a terminal run that
|
||||
//! archives the source and records `run.superseded` on it. A retry is not a
|
||||
|
|
@ -27,7 +27,7 @@ use fabro_petri::fork::{self as petri_fork, ForkError, ForkRequest};
|
|||
use fabro_petri::petri::RunStore;
|
||||
use fabro_petri::platform_records::SqlitePlatformRecords;
|
||||
use fabro_store::{PlatformRecordKind, RunProjection};
|
||||
use fabro_types::{FailureReason, Principal, RunId};
|
||||
use fabro_types::{FailureReason, Principal, RunId, RunTarget};
|
||||
use fabro_util::error as error_util;
|
||||
use fabro_workflow::Error as WorkflowError;
|
||||
use fabro_workflow::operations::{self, ForkTarget, ResolvedForkTarget, RunTimeline};
|
||||
|
|
@ -270,6 +270,14 @@ async fn fork_at(
|
|||
ForkKind::Rewind => operations::ensure_rewindable(&source, &id),
|
||||
}
|
||||
.map_err(workflow_operation_error)?;
|
||||
if !matches!(source.spec.target, Some(RunTarget::Git(_)))
|
||||
|| !source.spec.settings.run.run_branch.enabled
|
||||
|| !source.spec.settings.run.run_branch.push
|
||||
{
|
||||
return Err(ApiError::bad_request(
|
||||
"forking requires a GitHub-backed run with checkpoint pushes enabled",
|
||||
));
|
||||
}
|
||||
let timeline = timeline(state, id).await?;
|
||||
let entry = timeline
|
||||
.resolve_or_latest(target.as_ref())
|
||||
|
|
@ -287,7 +295,6 @@ async fn fork_at(
|
|||
|
||||
let new_run_id = RunId::new();
|
||||
let storage = Storage::new(state.server_storage_dir());
|
||||
let source_run_dir = storage.run_scratch(&id).root().to_path_buf();
|
||||
let run_dir = storage.run_scratch(&new_run_id).root().to_path_buf();
|
||||
operations::persist_forked_run(state.store_ref().as_ref(), &operations::ForkedRunInput {
|
||||
source: &source,
|
||||
|
|
@ -304,7 +311,6 @@ async fn fork_at(
|
|||
let seeded = petri_fork::fork(ForkRequest {
|
||||
source: id,
|
||||
fork: new_run_id,
|
||||
source_run_dir: source_run_dir.join("petri"),
|
||||
fork_run_dir: run_dir.join("petri"),
|
||||
store,
|
||||
records: Arc::new(SqlitePlatformRecords::new(Arc::clone(
|
||||
|
|
|
|||
|
|
@ -585,12 +585,10 @@ pub(crate) async fn reconcile_on_startup(
|
|||
let mode = if held {
|
||||
let request = RecoveryRequest::for_run(
|
||||
run_id,
|
||||
run_dir.join("petri"),
|
||||
Arc::new(SqliteRunStore::new(state.db_pool.clone())),
|
||||
Arc::new(SqlitePlatformRecords::new(Arc::clone(
|
||||
&state.stores.run_summaries,
|
||||
))),
|
||||
&run_state.spec.settings.run,
|
||||
);
|
||||
match recovery::recover(request)
|
||||
.await
|
||||
|
|
@ -601,7 +599,7 @@ pub(crate) async fn reconcile_on_startup(
|
|||
info!(
|
||||
run_id = %run_id,
|
||||
workspaces = workspaces.len(),
|
||||
"Petri run's workspaces match its durable state"
|
||||
"Petri recovery plan ready; the worker will verify retained workspaces"
|
||||
);
|
||||
RunExecutionMode::Resume
|
||||
}
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -195,12 +195,8 @@ pub async fn run(request: RunRequest) -> Result<RunOutcome, RunError> {
|
|||
if backend == SandboxBackend::Daytona {
|
||||
options.sandbox.daytona_resources = daytona_resources(&request.resources)?;
|
||||
}
|
||||
// Fabro's hooks restore a sandbox workspace from its snapshots at the
|
||||
// scope's acquisition, so a lease whose sandbox is gone gets a fresh
|
||||
// one instead of failing the run.
|
||||
if request.hooks.is_some() && backend != SandboxBackend::Host {
|
||||
options.sandbox.lost_sandbox = LostSandbox::Replace;
|
||||
}
|
||||
// A normal resume requires its original sandbox to survive.
|
||||
options.sandbox.lost_sandbox = LostSandbox::Refuse;
|
||||
let resumed = matches!(request.execution, Execution::Resume);
|
||||
let mut runtime = request
|
||||
.runtime
|
||||
|
|
|
|||
|
|
@ -1,5 +1,5 @@
|
|||
//! Forking a Fabro run at a checkpoint: the seam over Petri's
|
||||
//! `host::fork_from` that rewind, fork and retry are built on.
|
||||
//! `host::fork_from` that rewind and fork are built on.
|
||||
//!
|
||||
//! Fabro's checkpoint record ties a Petri position `(execution, firing)` to
|
||||
//! a Git commit. A fork seeds a new run from the source's records up to such
|
||||
|
|
@ -12,14 +12,10 @@
|
|||
//! `run.started` whose `forked_from` names the source and the position. No
|
||||
//! sandbox lease is carried over, so the resume acquires the position
|
||||
//! execution's scopes fresh.
|
||||
//! 2. The source's checkpoint records for every attempt the fork kept are
|
||||
//! written again under the new run, at their positions, and the snapshot
|
||||
//! repository of every workspace they name is seeded with those checkpoints'
|
||||
//! refs alone, fetched from the source's repository under the source's run
|
||||
//! scratch. That is what the resume's recovery plan
|
||||
//! ([`crate::recovery::plan`]) reads: the last durable finish of the
|
||||
//! position execution names the snapshot the fresh workspace is restored to
|
||||
//! at `scope_acquired`, on the host and in a sandbox alike.
|
||||
//! 2. The source's checkpoint records for retained attempts are copied under
|
||||
//! the new run. At first scope acquisition the worker fetches the source
|
||||
//! run's published branch from GitHub and checks out the selected commit. No
|
||||
//! repository is copied or stored on the server.
|
||||
//! 3. The fork's `run.branch` record names the run branch the restore creates
|
||||
//! (`fabro/run/<new id>`) and the commit it starts from, with the
|
||||
//! `git.identity` beside it, both at the checkpoint's position, so the hooks
|
||||
|
|
@ -31,7 +27,6 @@
|
|||
|
||||
use std::collections::{BTreeMap, BTreeSet};
|
||||
use std::path::PathBuf;
|
||||
use std::process::Stdio;
|
||||
use std::sync::Arc;
|
||||
|
||||
use fabro_checkpoint::author::GitAuthor;
|
||||
|
|
@ -50,11 +45,9 @@ use petri_execution::{
|
|||
use petri_runtime::RunOptions;
|
||||
use petri_runtime::ir::FiringId;
|
||||
use petri_store::StoreError;
|
||||
use tokio::fs;
|
||||
use tokio::process::Command;
|
||||
use tracing::{debug, info};
|
||||
|
||||
use crate::checkpoint::{CheckpointKey, RunWorkspaces, SOURCE_REF};
|
||||
use crate::checkpoint::CheckpointKey;
|
||||
use crate::platform_records::{PlatformRecordError, PlatformRecords};
|
||||
use crate::projection::FoldState;
|
||||
use crate::projector::ProjectError;
|
||||
|
|
@ -63,26 +56,23 @@ use crate::providers::{self, SandboxProviderConfig};
|
|||
/// One fork to seed.
|
||||
pub struct ForkRequest {
|
||||
/// The run whose records are copied.
|
||||
pub source: RunId,
|
||||
pub source: RunId,
|
||||
/// The new run's id: its Petri run key and its own run scratch.
|
||||
pub fork: RunId,
|
||||
/// The source's Petri run directory (its scratch's `petri`), where its
|
||||
/// snapshot repositories are.
|
||||
pub source_run_dir: PathBuf,
|
||||
/// The fork's Petri run directory, where its snapshot repositories go.
|
||||
pub fork_run_dir: PathBuf,
|
||||
pub fork: RunId,
|
||||
/// The fork's Petri run directory, used by Petri.
|
||||
pub fork_run_dir: PathBuf,
|
||||
/// The server's run store: the source is read from it, the fork is
|
||||
/// written into it.
|
||||
pub store: Arc<dyn RunStore>,
|
||||
pub store: Arc<dyn RunStore>,
|
||||
/// The platform records of both runs.
|
||||
pub records: Arc<dyn PlatformRecords>,
|
||||
pub records: Arc<dyn PlatformRecords>,
|
||||
/// The position the source's records are kept up to.
|
||||
pub position: ForkPosition,
|
||||
pub position: ForkPosition,
|
||||
/// Whether the position's firing runs again (a retry of a failed
|
||||
/// stage) instead of keeping its finish.
|
||||
pub rerun_last: bool,
|
||||
pub rerun_last: bool,
|
||||
/// The run's settings, for its Git author and checkpoint settings.
|
||||
pub settings: RunNamespace,
|
||||
pub settings: RunNamespace,
|
||||
}
|
||||
|
||||
/// A seeded fork, not yet resumed.
|
||||
|
|
@ -125,17 +115,13 @@ pub enum ForkError {
|
|||
Log(#[source] CoordinatorStoreError),
|
||||
#[error("the checkpoint records could not be read or written")]
|
||||
Records(#[source] PlatformRecordError),
|
||||
#[error("the snapshot repository for `{workspace}` could not be seeded: {detail}")]
|
||||
Snapshots {
|
||||
workspace: String,
|
||||
detail: String,
|
||||
},
|
||||
}
|
||||
|
||||
/// Refuse a position Petri would refuse, before anything is written for
|
||||
/// the fork: an execution the source does not have, or one inside a child
|
||||
/// invocation (a branch of a parallel node), whose caller's firing is live
|
||||
/// at every position inside it. The messages are Petri's own.
|
||||
/// at every position inside it. Also refuse the final terminal checkpoint:
|
||||
/// it leaves no work that would acquire a workspace for publication.
|
||||
pub async fn check(
|
||||
store: &dyn RunStore,
|
||||
source: RunId,
|
||||
|
|
@ -160,13 +146,33 @@ pub async fn check(
|
|||
position.execution
|
||||
)));
|
||||
}
|
||||
let inspection = inspect::inspect_run(&*logs)
|
||||
.await
|
||||
.map_err(ForkError::Inspect)?;
|
||||
if inspection
|
||||
.executions
|
||||
.iter()
|
||||
.find(|execution| execution.execution == position.execution)
|
||||
.and_then(|execution| execution.engine.as_ref())
|
||||
.is_some_and(|engine| {
|
||||
matches!(engine.exit, Some(inspect::ExitInspection::Terminal { .. }))
|
||||
&& engine
|
||||
.attempts
|
||||
.last()
|
||||
.is_some_and(|attempt| attempt.firing == position.firing.raw())
|
||||
})
|
||||
{
|
||||
return Err(ForkError::Refused(
|
||||
"the terminal checkpoint has no remaining work to acquire a sandbox; select an earlier checkpoint or retry the workflow from the start".to_string(),
|
||||
));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Seed the fork: Petri's records, then the kept checkpoints, their
|
||||
/// snapshots and the run branch. The new run must not exist in the store
|
||||
/// yet.
|
||||
/// Seed the fork: Petri's records, the kept checkpoints and the run branch. The
|
||||
/// new run must not exist in the store yet.
|
||||
pub async fn fork(request: ForkRequest) -> Result<Forked, ForkError> {
|
||||
check(request.store.as_ref(), request.source, request.position).await?;
|
||||
let source_key = RunKey::new(request.source.to_string());
|
||||
let fork_key = RunKey::new(request.fork.to_string());
|
||||
let source_logs = request
|
||||
|
|
@ -272,8 +278,6 @@ pub async fn fork(request: ForkRequest) -> Result<Forked, ForkError> {
|
|||
});
|
||||
}
|
||||
|
||||
// The snapshot repositories: one per workspace the kept checkpoints
|
||||
// name, holding those checkpoints' refs alone.
|
||||
let author = request
|
||||
.settings
|
||||
.git
|
||||
|
|
@ -281,30 +285,6 @@ pub async fn fork(request: ForkRequest) -> Result<Forked, ForkError> {
|
|||
.as_ref()
|
||||
.map(GitAuthor::from)
|
||||
.unwrap_or_default();
|
||||
let source_workspaces = RunWorkspaces::new(
|
||||
request.source_run_dir.clone(),
|
||||
request.source.to_string(),
|
||||
author.clone(),
|
||||
&request.settings.checkpoint,
|
||||
);
|
||||
let fork_workspaces = RunWorkspaces::new(
|
||||
request.fork_run_dir.clone(),
|
||||
request.fork.to_string(),
|
||||
author.clone(),
|
||||
&request.settings.checkpoint,
|
||||
);
|
||||
let mut by_workspace: BTreeMap<String, Vec<CheckpointKey>> = BTreeMap::new();
|
||||
for kept in &checkpoints {
|
||||
if let Some(workspace) = &kept.workspace {
|
||||
by_workspace
|
||||
.entry(workspace.clone())
|
||||
.or_default()
|
||||
.push(kept.key);
|
||||
}
|
||||
}
|
||||
for (workspace, keys) in &by_workspace {
|
||||
seed_snapshots(&source_workspaces, &fork_workspaces, workspace, keys).await?;
|
||||
}
|
||||
|
||||
// The run branch the restore creates, from the checkpoint the fork
|
||||
// starts on, and the identity that authors the fork's commits.
|
||||
|
|
@ -315,7 +295,7 @@ pub async fn fork(request: ForkRequest) -> Result<Forked, ForkError> {
|
|||
firing: start.key.firing,
|
||||
};
|
||||
let branch = PlatformRecord::RunBranch(RunBranchRecord {
|
||||
run_branch: Some(fork_workspaces.run_branch()),
|
||||
run_branch: Some(format!("fabro/run/{}", request.fork)),
|
||||
base_sha: Some(start.sha.clone()),
|
||||
workspace: start.workspace.clone(),
|
||||
});
|
||||
|
|
@ -377,88 +357,6 @@ fn checkpoint_key(record: &CheckpointRecord) -> Option<CheckpointKey> {
|
|||
})
|
||||
}
|
||||
|
||||
/// Create the fork's bare snapshot repository for `workspace` and fetch the
|
||||
/// kept checkpoints' refs into it from the source's.
|
||||
async fn seed_snapshots(
|
||||
source: &RunWorkspaces,
|
||||
fork: &RunWorkspaces,
|
||||
workspace: &str,
|
||||
keys: &[CheckpointKey],
|
||||
) -> Result<(), ForkError> {
|
||||
let failed = |detail: String| ForkError::Snapshots {
|
||||
workspace: workspace.to_string(),
|
||||
detail,
|
||||
};
|
||||
let source_repository = source.snapshot_repository(workspace);
|
||||
if !fs::try_exists(&source_repository).await.unwrap_or(false) {
|
||||
return Err(failed(format!(
|
||||
"the source run has no snapshot repository at {}",
|
||||
source_repository.display()
|
||||
)));
|
||||
}
|
||||
let repository = fork.snapshot_repository(workspace);
|
||||
fs::create_dir_all(&repository).await.map_err(|error| {
|
||||
failed(format!(
|
||||
"{} could not be created: {error}",
|
||||
repository.display()
|
||||
))
|
||||
})?;
|
||||
git(&repository, &["init", "-q", "--bare"])
|
||||
.await
|
||||
.map_err(failed)?;
|
||||
let mut args = vec![
|
||||
"fetch".to_string(),
|
||||
"-q".to_string(),
|
||||
source_repository.to_string_lossy().into_owned(),
|
||||
];
|
||||
for key in keys {
|
||||
let name = key.snapshot_ref();
|
||||
args.push(format!("+{name}:{name}"));
|
||||
}
|
||||
// The commit a checked-out workspace started from goes with its
|
||||
// checkpoints: the fork's bundles and restores are cut against it.
|
||||
if git(&source_repository, &[
|
||||
"rev-parse",
|
||||
"-q",
|
||||
"--verify",
|
||||
SOURCE_REF,
|
||||
])
|
||||
.await
|
||||
.is_ok()
|
||||
{
|
||||
args.push(format!("+{SOURCE_REF}:{SOURCE_REF}"));
|
||||
}
|
||||
git(&repository, &args).await.map_err(failed)?;
|
||||
debug!(
|
||||
workspace,
|
||||
refs = keys.len(),
|
||||
repository = %repository.display(),
|
||||
"the fork's snapshot repository is seeded"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Run `git` in `repository`; a non-zero exit is the error's detail.
|
||||
async fn git<S: AsRef<str>>(repository: &std::path::Path, args: &[S]) -> Result<(), String> {
|
||||
let output = Command::new("git")
|
||||
.args(args.iter().map(AsRef::as_ref))
|
||||
.current_dir(repository)
|
||||
.stdin(Stdio::null())
|
||||
.output()
|
||||
.await
|
||||
.map_err(|error| format!("git could not run: {error}"))?;
|
||||
if output.status.success() {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(format!(
|
||||
"git {} failed ({}): {}",
|
||||
args.first().map_or("", AsRef::as_ref),
|
||||
output.status,
|
||||
String::from_utf8_lossy(&output.stderr).trim()
|
||||
))
|
||||
}
|
||||
}
|
||||
|
||||
/// The position a checkpoint's execution and firing name, in Petri's ids.
|
||||
#[must_use]
|
||||
pub fn position(execution: u64, firing: u64) -> ForkPosition {
|
||||
|
|
|
|||
|
|
@ -11,16 +11,18 @@
|
|||
//! embedding host does. Its own work, at each point:
|
||||
//!
|
||||
//! - `prepare_result`: the checkpoint commit, before the `StepFinished` record
|
||||
//! is appended, so a durable finish implies a durable snapshot. A stage that
|
||||
//! failed on its own terms is committed like a successful one; only a
|
||||
//! cancelled attempt is not. A failed commit is fatal to the run: the outcome
|
||||
//! becomes a failure of class `checkpoint_failed`, the run is cancelled
|
||||
//! through the coordinator handle, and `transition` refuses the firing's
|
||||
//! routes, so no route is taken. The commit that creates the run branch also
|
||||
//! records where it started: the `run.branch` platform record (the branch
|
||||
//! name and the base commit) and the `git.identity` record (who authors the
|
||||
//! commits, and where that identity came from), both at that checkpoint's
|
||||
//! stage position, so the stream orders them with the firing's finish.
|
||||
//! is appended. Git-backed runs push that commit from the sandbox; a failed
|
||||
//! push warns and is retried at the next checkpoint and at final publication.
|
||||
//! A stage that failed on its own terms is committed like a successful one;
|
||||
//! only a cancelled attempt is not. A failed commit is fatal to the run: the
|
||||
//! outcome becomes a failure of class `checkpoint_failed`, the run is
|
||||
//! cancelled through the coordinator handle, and `transition` refuses the
|
||||
//! firing's routes, so no route is taken. The commit that creates the run
|
||||
//! branch also records where it started: the `run.branch` platform record
|
||||
//! (the branch name and the base commit) and the `git.identity` record (who
|
||||
//! authors the commits, and where that identity came from), both at that
|
||||
//! checkpoint's stage position, so the stream orders them with the firing's
|
||||
//! finish.
|
||||
//! - `transition`: the platform checkpoint record, keyed on the Petri position
|
||||
//! and the checkpoint's operation identity, with the stage's diff from its
|
||||
//! parent commit (`diff_summary`, and the patch as a blob); then the stage's
|
||||
|
|
@ -36,10 +38,10 @@
|
|||
//! record; then the forwarded point, so the local service runs `run_complete`
|
||||
//! and `run_failed` with the sandbox in place.
|
||||
//! - `scope_acquired`: a fresh run's Git target checked out into the workspace
|
||||
//! from inside the scope, with its snapshot repository seeded with the
|
||||
//! starting commit ([`crate::source`]); a resumed run's workspace brought to
|
||||
//! its snapshot instead (see below). A checkout that fails fails the scope's
|
||||
//! firings with the reason.
|
||||
//! from inside the scope ([`crate::source`]); a resumed run uses its
|
||||
//! surviving workspace, while an explicit fork fetches the source run's
|
||||
//! branch from GitHub. A checkout that fails fails the scope's firings with
|
||||
//! the reason.
|
||||
//! - `scope_released`: forwarded, so the local service runs `sandbox_cleanup`
|
||||
//! with the sandbox in place. Fabro's own end-of-run work (the terminal
|
||||
//! lifecycle event, notifications on it) is the run lifecycle path's, on the
|
||||
|
|
@ -66,17 +68,12 @@
|
|||
//! Petri's host backend keeps under the run directory (`crate::checkpoint`).
|
||||
//! On Docker or Daytona the workspace lives inside the scope's sandbox: the
|
||||
//! hooks keep the environment Petri hands them at `scope_acquired`, run
|
||||
//! `git` inside the scope through it, and move the commit out as a bundle
|
||||
//! into the same snapshot repository the host path pushes to. Artifacts are
|
||||
//! read out through the same environment on every provider. The same
|
||||
//! point is where a resumed run brings a workspace to the snapshot its
|
||||
//! durable state names, before the first attempt runs in it: verified,
|
||||
//! reset, or, in a fresh sandbox (Petri replaces a lost one on Fabro's
|
||||
//! request), restored from a bundle of the checkpoint. The plan is
|
||||
//! [`recovery::plan`], the one the server applied to host workspaces before
|
||||
//! it relaunched the worker; a host workspace is verified here, unless the
|
||||
//! run is a fork whose fresh workspace nothing restored yet
|
||||
//! ([`crate::fork`]), which is restored from the seeded snapshot repository.
|
||||
//! `git` inside the scope through it. Checkpoints, diffs, fetches and pushes
|
||||
//! all use this environment. Only execution metadata, patches and selected
|
||||
//! artifacts are persisted on the server. Normal resume requires the original
|
||||
//! workspace; a fork acquires a new sandbox and fetches its checkpoint from
|
||||
//! GitHub. Empty and local-folder targets record execution checkpoints without
|
||||
//! Git commits.
|
||||
|
||||
use std::collections::{BTreeMap, HashMap, HashSet};
|
||||
use std::path::{Path, PathBuf};
|
||||
|
|
@ -113,6 +110,7 @@ use crate::checkpoint::{
|
|||
CHECKPOINT_FAILED_CLASS, CheckpointError, CheckpointKey, EXCLUDE_DIRS, RunGitSettings,
|
||||
RunWorkspaces, Site, Snapshot, WorkspaceDiff,
|
||||
};
|
||||
use crate::fork::{self, ForkError};
|
||||
use crate::platform_records::{PlatformRecordError, PlatformRecords};
|
||||
use crate::recovery::{self, Plan, RecoveryError, RestoreTarget};
|
||||
use crate::source::RunSource;
|
||||
|
|
@ -203,6 +201,8 @@ pub enum HookError {
|
|||
#[source]
|
||||
source: EnvError,
|
||||
},
|
||||
#[error("the fork origin could not be read")]
|
||||
Fork(#[source] ForkError),
|
||||
#[error("the run's restore plan could not be read")]
|
||||
Plan(#[source] RecoveryError),
|
||||
#[error("{reason}")]
|
||||
|
|
@ -242,14 +242,14 @@ pub struct HooksSpec {
|
|||
}
|
||||
|
||||
/// What a successful run hands its publisher when it ends: the run branch,
|
||||
/// the commit it ends on, the snapshot repository that holds it, and the
|
||||
/// the commit it ends on, the workspace that holds it, and the
|
||||
/// run's patch against the commit the branch started from.
|
||||
#[derive(Clone, Debug)]
|
||||
pub struct Publication {
|
||||
pub run_branch: String,
|
||||
pub head_sha: String,
|
||||
pub snapshot_repository: PathBuf,
|
||||
pub patch: String,
|
||||
pub run_branch: String,
|
||||
pub head_sha: String,
|
||||
pub site: Site,
|
||||
pub patch: String,
|
||||
}
|
||||
|
||||
/// The platform's end-of-run publication: what a successful run's work does
|
||||
|
|
@ -258,6 +258,10 @@ pub struct Publication {
|
|||
/// message, as a failed publish did on the legacy executor.
|
||||
#[async_trait::async_trait]
|
||||
pub trait RunPublisher: Send + Sync {
|
||||
/// Push one checkpoint. A failure is retried by later checkpoints and
|
||||
/// final publication, which must succeed for the run to succeed.
|
||||
async fn push(&self, site: &Site, branch: &str, sha: &str) -> Result<(), String>;
|
||||
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String>;
|
||||
}
|
||||
|
||||
|
|
@ -427,34 +431,36 @@ impl ScopeEnvs {
|
|||
|
||||
/// Fabro's `ExecutionHooks`, around the hooks the runtime installed.
|
||||
pub struct FabroHooks {
|
||||
inner: Arc<dyn ExecutionHooks>,
|
||||
run_id: RunId,
|
||||
records: Arc<dyn PlatformRecords>,
|
||||
inner: Arc<dyn ExecutionHooks>,
|
||||
run_id: RunId,
|
||||
records: Arc<dyn PlatformRecords>,
|
||||
/// Where diff patches go; `None` records summaries alone.
|
||||
blobs: Option<Arc<dyn Blobs>>,
|
||||
artifact_writer: Arc<dyn ArtifactWriter>,
|
||||
workspaces: RunWorkspaces,
|
||||
lookup: WorkspaceLookup,
|
||||
identity: GitIdentity,
|
||||
host_workspaces: bool,
|
||||
test_gates: Option<PathBuf>,
|
||||
handle: OnceLock<CoordinatorHandle>,
|
||||
checkpoints: CheckpointLedger,
|
||||
artifacts: ArtifactLedger,
|
||||
scopes: ScopeEnvs,
|
||||
blobs: Option<Arc<dyn Blobs>>,
|
||||
artifact_writer: Arc<dyn ArtifactWriter>,
|
||||
workspaces: RunWorkspaces,
|
||||
checkpoint_enabled: bool,
|
||||
sites: Mutex<HashMap<String, Site>>,
|
||||
lookup: WorkspaceLookup,
|
||||
identity: GitIdentity,
|
||||
host_workspaces: bool,
|
||||
test_gates: Option<PathBuf>,
|
||||
handle: OnceLock<CoordinatorHandle>,
|
||||
checkpoints: CheckpointLedger,
|
||||
artifacts: ArtifactLedger,
|
||||
scopes: ScopeEnvs,
|
||||
/// The checkpoint failure that ended the run, when one did.
|
||||
failure: Mutex<Option<String>>,
|
||||
publisher: Option<Arc<dyn RunPublisher>>,
|
||||
failure: Mutex<Option<String>>,
|
||||
publisher: Option<Arc<dyn RunPublisher>>,
|
||||
/// Why the run's publication failed, when it did.
|
||||
publish_failure: Mutex<Option<String>>,
|
||||
publish_failure: Mutex<Option<String>>,
|
||||
/// Whether the run continues from its records: a sandbox workspace is
|
||||
/// then brought to its snapshot when its scope is first acquired.
|
||||
resumed: bool,
|
||||
resumed: bool,
|
||||
/// The snapshot every live sandbox workspace must sit on before work
|
||||
/// resumes in it, read once from the records; an entry leaves when it
|
||||
/// is applied.
|
||||
restore: OnceCell<Mutex<BTreeMap<String, RestoreTarget>>>,
|
||||
store: Arc<dyn RunStore>,
|
||||
restore: OnceCell<Mutex<BTreeMap<String, Vec<RestoreTarget>>>>,
|
||||
store: Arc<dyn RunStore>,
|
||||
}
|
||||
|
||||
impl FabroHooks {
|
||||
|
|
@ -479,6 +485,7 @@ impl FabroHooks {
|
|||
email: spec.git.author.email.clone(),
|
||||
source: spec.git.identity_source,
|
||||
};
|
||||
let checkpoint_enabled = spec.source.is_some() && spec.git.enabled;
|
||||
let workspaces = RunWorkspaces::new(
|
||||
run_dir,
|
||||
run_id.to_string(),
|
||||
|
|
@ -493,6 +500,8 @@ impl FabroHooks {
|
|||
blobs,
|
||||
artifact_writer: spec.artifact_writer,
|
||||
workspaces,
|
||||
checkpoint_enabled,
|
||||
sites: Mutex::default(),
|
||||
lookup: WorkspaceLookup::new(Arc::clone(&store), run_key),
|
||||
identity,
|
||||
host_workspaces: spec.git.host_workspaces,
|
||||
|
|
@ -630,6 +639,10 @@ impl FabroHooks {
|
|||
status: &Status,
|
||||
origin: ResultOrigin,
|
||||
) -> Result<Option<Note>, HookError> {
|
||||
self.gate("prepare", node).await;
|
||||
if !self.checkpoint_enabled {
|
||||
return Ok(None);
|
||||
}
|
||||
let Some((workspace, site)) = self.site_of(context, scope).await? else {
|
||||
// A skipped node or a driver-made outcome may precede the scope's
|
||||
// environment; nothing of the stage's exists to snapshot.
|
||||
|
|
@ -669,6 +682,22 @@ impl FabroHooks {
|
|||
"checkpoint committed"
|
||||
);
|
||||
self.committed(key, &workspace, &snapshot).await;
|
||||
// Only the workspace that owns the run branch publishes it.
|
||||
// Isolated child workspaces must not race to replace its head.
|
||||
if let Some(publisher) = &self.publisher {
|
||||
if self
|
||||
.stored_branch()
|
||||
.await?
|
||||
.is_some_and(|branch| branch.workspace.as_deref() == Some(&workspace))
|
||||
{
|
||||
if let Err(message) = publisher
|
||||
.push(&site, &self.workspaces.run_branch(), &snapshot.sha)
|
||||
.await
|
||||
{
|
||||
warn!(run_id = %self.run_id, error = %message, "checkpoint push failed; final publication will retry");
|
||||
}
|
||||
}
|
||||
}
|
||||
Ok(Some(Note::new(
|
||||
CHECKPOINT_NOTE,
|
||||
json!({
|
||||
|
|
@ -789,17 +818,15 @@ impl FabroHooks {
|
|||
|
||||
/// The restore plan of a resumed run, read once: what every live
|
||||
/// sandbox workspace must be brought to at its first acquisition.
|
||||
async fn restore_targets(&self) -> Result<&Mutex<BTreeMap<String, RestoreTarget>>, HookError> {
|
||||
async fn restore_targets(
|
||||
&self,
|
||||
) -> Result<&Mutex<BTreeMap<String, Vec<RestoreTarget>>>, HookError> {
|
||||
self.restore
|
||||
.get_or_try_init(|| async {
|
||||
let plan = recovery::plan(
|
||||
Arc::clone(&self.store),
|
||||
self.records.as_ref(),
|
||||
&self.run_id,
|
||||
&self.workspaces,
|
||||
)
|
||||
.await
|
||||
.map_err(HookError::Plan)?;
|
||||
let plan =
|
||||
recovery::plan(Arc::clone(&self.store), self.records.as_ref(), &self.run_id)
|
||||
.await
|
||||
.map_err(HookError::Plan)?;
|
||||
match plan {
|
||||
Plan::Resume { targets } => Ok(Mutex::new(targets)),
|
||||
Plan::Start => Ok(Mutex::default()),
|
||||
|
|
@ -845,19 +872,57 @@ impl FabroHooks {
|
|||
}
|
||||
}
|
||||
|
||||
/// Bring a workspace to the snapshot the resumed run's durable state
|
||||
/// names, once, at its first acquisition. After a restart the server
|
||||
/// already brought a host workspace there, so this verifies; a fork's
|
||||
/// fresh workspace is restored here from the snapshot repository the
|
||||
/// fork seeded; a sandbox workspace is only reachable here.
|
||||
/// Reset a surviving workspace to its recorded checkpoint. An explicit
|
||||
/// fork may fetch the source run's branch into a new workspace.
|
||||
async fn restore(&self, workspace: &str, site: &Site) -> Result<(), HookError> {
|
||||
let targets = self.restore_targets().await?;
|
||||
let target = sync::lock(targets).remove(workspace);
|
||||
let Some(target) = target else {
|
||||
let Some(targets) = target else {
|
||||
return Ok(());
|
||||
};
|
||||
let Some(mut target) = targets.first().cloned() else {
|
||||
return Ok(());
|
||||
};
|
||||
let serialized = self.scopes.lock_for(workspace);
|
||||
let _held = serialized.lock().await;
|
||||
if !self
|
||||
.workspaces
|
||||
.has_commit(site, &target.sha)
|
||||
.await
|
||||
.map_err(HookError::Find)?
|
||||
{
|
||||
let origin = fork::origin_of(self.store.as_ref(), self.run_id)
|
||||
.await
|
||||
.map_err(HookError::Fork)?;
|
||||
if let Some(origin) = origin {
|
||||
// An explicit fork may acquire a fresh workspace. Normal
|
||||
// resumes never reconstruct a lost repository.
|
||||
if self
|
||||
.workspaces
|
||||
.head(site)
|
||||
.await
|
||||
.map_err(HookError::Find)?
|
||||
.is_none()
|
||||
{
|
||||
self.workspaces
|
||||
.restore_fork(site, &origin.source_run_id.to_string(), &target.sha)
|
||||
.await
|
||||
.map_err(HookError::Find)?;
|
||||
}
|
||||
}
|
||||
}
|
||||
// Platform records can arrive out of commit order from parallel stages.
|
||||
// Choose by Git ancestry here, where the actual repository is available.
|
||||
for candidate in targets.iter().skip(1) {
|
||||
if self
|
||||
.workspaces
|
||||
.is_ancestor(site, &target.sha, &candidate.sha)
|
||||
.await
|
||||
.map_err(HookError::Find)?
|
||||
{
|
||||
target = candidate.clone();
|
||||
}
|
||||
}
|
||||
let action = recovery::bring_to(&self.workspaces, site, workspace, &target)
|
||||
.await
|
||||
.map_err(HookError::Restore)?;
|
||||
|
|
@ -884,6 +949,41 @@ impl FabroHooks {
|
|||
if sync::lock(recorded).contains(&key) {
|
||||
return Ok(());
|
||||
}
|
||||
if !self.checkpoint_enabled {
|
||||
self.records
|
||||
.append(
|
||||
&self.run_id,
|
||||
&PlatformRecord::Checkpoint(CheckpointRecord {
|
||||
execution: key.execution,
|
||||
firing: key.firing,
|
||||
attempt: Some(key.attempt),
|
||||
workspace: None,
|
||||
git_commit_sha: None,
|
||||
diff_summary: None,
|
||||
patch_blob: None,
|
||||
operation: Some(key.operation()),
|
||||
}),
|
||||
Some(StagePosition {
|
||||
execution: key.execution,
|
||||
firing: key.firing,
|
||||
}),
|
||||
)
|
||||
.await
|
||||
.map_err(|source| HookError::Write {
|
||||
kind: "checkpoint",
|
||||
source,
|
||||
})?;
|
||||
sync::lock(recorded).insert(key);
|
||||
return Ok(());
|
||||
}
|
||||
let site = self
|
||||
.site_of(context, scope)
|
||||
.await?
|
||||
.ok_or(HookError::NoWorkspace {
|
||||
scope,
|
||||
execution: context.execution,
|
||||
})?
|
||||
.1;
|
||||
let (workspace, sha) = if let Some(committed) = self.checkpoints.commit_of(key) {
|
||||
committed
|
||||
} else {
|
||||
|
|
@ -894,14 +994,14 @@ impl FabroHooks {
|
|||
};
|
||||
let serialized = self.scopes.lock_for(&workspace);
|
||||
let held = serialized.lock().await;
|
||||
let found = self.workspaces.find(&workspace, key).await;
|
||||
let found = self.workspaces.find(&site, key).await;
|
||||
drop(held);
|
||||
let sha = found
|
||||
.map_err(HookError::Find)?
|
||||
.ok_or(HookError::NoCommit { key })?;
|
||||
(workspace, sha)
|
||||
};
|
||||
let (diff_summary, patch_blob) = match self.stage_diff(&workspace, &sha).await {
|
||||
let (diff_summary, patch_blob) = match self.stage_diff(&site, &sha).await {
|
||||
Ok(diff) => diff,
|
||||
Err(error) => {
|
||||
// The record still names the commit; the diff is a view.
|
||||
|
|
@ -950,7 +1050,7 @@ impl FabroHooks {
|
|||
/// the diff is not empty.
|
||||
async fn stage_diff(
|
||||
&self,
|
||||
workspace: &str,
|
||||
site: &Site,
|
||||
sha: &str,
|
||||
) -> Result<(Option<DiffSummary>, Option<BlobHash>), HookError> {
|
||||
let failed = |source| HookError::Diff {
|
||||
|
|
@ -959,7 +1059,7 @@ impl FabroHooks {
|
|||
};
|
||||
let parent = self
|
||||
.workspaces
|
||||
.commit_parent(workspace, sha)
|
||||
.commit_parent(site, sha)
|
||||
.await
|
||||
.map_err(failed)?;
|
||||
let Some(parent) = parent else {
|
||||
|
|
@ -967,7 +1067,7 @@ impl FabroHooks {
|
|||
};
|
||||
let diff = self
|
||||
.workspaces
|
||||
.diff(workspace, Some(&parent), sha)
|
||||
.diff(site, Some(&parent), sha)
|
||||
.await
|
||||
.map_err(failed)?;
|
||||
let patch_blob = self.patch_blob(&diff).await?;
|
||||
|
|
@ -1201,7 +1301,7 @@ impl FabroHooks {
|
|||
}
|
||||
|
||||
/// The run's diff: the run branch's last checkpoint against the base
|
||||
/// the branch started from, in the snapshot repository on this host.
|
||||
/// the branch started from, computed inside the workspace.
|
||||
/// Nothing is recorded for a run that never created its branch or
|
||||
/// never checkpointed.
|
||||
async fn record_run_diff(&self) -> Result<Option<Publication>, HookError> {
|
||||
|
|
@ -1217,7 +1317,7 @@ impl FabroHooks {
|
|||
let Some(base_sha) = branch.base_sha.clone() else {
|
||||
return Ok(None);
|
||||
};
|
||||
let Some((workspace, head_sha)) = self.checkpoints.last() else {
|
||||
let Some((workspace, _)) = self.checkpoints.last() else {
|
||||
debug!(run_id = %self.run_id, "no checkpoint is recorded; no run diff");
|
||||
return Ok(None);
|
||||
};
|
||||
|
|
@ -1225,9 +1325,39 @@ impl FabroHooks {
|
|||
// in; a last checkpoint elsewhere (a nested invocation's workspace)
|
||||
// is not this branch's head.
|
||||
let workspace = branch.workspace.clone().unwrap_or(workspace);
|
||||
let site = sync::lock(&self.sites)
|
||||
.get(&workspace)
|
||||
.cloned()
|
||||
.ok_or_else(|| HookError::Unresumable {
|
||||
reason: format!(
|
||||
"the run branch workspace {workspace} is unavailable for publication"
|
||||
),
|
||||
})?;
|
||||
let stored = self
|
||||
.records
|
||||
.read_kind(&self.run_id, PlatformRecordKind::Checkpoint)
|
||||
.await
|
||||
.map_err(|source| HookError::Read {
|
||||
kind: "checkpoint",
|
||||
source,
|
||||
})?;
|
||||
let head_sha = stored
|
||||
.into_iter()
|
||||
.rev()
|
||||
.find_map(|stored| match stored.record {
|
||||
PlatformRecord::Checkpoint(record)
|
||||
if record.workspace.as_deref() == Some(&workspace) =>
|
||||
{
|
||||
record.git_commit_sha
|
||||
}
|
||||
_ => None,
|
||||
})
|
||||
.ok_or_else(|| HookError::Unresumable {
|
||||
reason: "the run branch has no checkpoint".to_string(),
|
||||
})?;
|
||||
let diff = self
|
||||
.workspaces
|
||||
.diff(&workspace, Some(&base_sha), &head_sha)
|
||||
.diff(&site, Some(&base_sha), &head_sha)
|
||||
.await
|
||||
.map_err(|source| HookError::Diff {
|
||||
what: "run's diff",
|
||||
|
|
@ -1241,7 +1371,7 @@ impl FabroHooks {
|
|||
.map(|run_branch| Publication {
|
||||
run_branch,
|
||||
head_sha: head_sha.clone(),
|
||||
snapshot_repository: self.workspaces.snapshot_repository(&workspace),
|
||||
site,
|
||||
patch: diff.patch,
|
||||
});
|
||||
let record = PlatformRecord::RunDiff(RunDiffRecord {
|
||||
|
|
@ -1500,13 +1630,23 @@ impl ExecutionHooks for FabroHooks {
|
|||
let publication = match self.record_run_diff().await {
|
||||
Ok(publication) => publication,
|
||||
Err(error) => {
|
||||
warn!(run_id = %self.run_id, error = %error.render(), "the run's diff was not recorded");
|
||||
let message = error.render();
|
||||
warn!(run_id = %self.run_id, error = %message, "the run's diff was not recorded");
|
||||
if self.publisher.is_some() && finished.status == RunStatus::Success {
|
||||
*sync::lock(&self.publish_failure) = Some(message);
|
||||
}
|
||||
None
|
||||
}
|
||||
};
|
||||
if let (Some(publisher), Some(publication)) = (&self.publisher, &publication) {
|
||||
if let Some(publisher) = &self.publisher {
|
||||
if finished.status == RunStatus::Success && self.checkpoint_failure().is_none() {
|
||||
self.publish(publisher.as_ref(), publication).await;
|
||||
if let Some(publication) = &publication {
|
||||
self.publish(publisher.as_ref(), publication).await;
|
||||
} else if self.publish_failure().is_none() {
|
||||
*sync::lock(&self.publish_failure) = Some(
|
||||
"the run has no recorded branch and checkpoint to publish".to_string(),
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
self.inner.run_finished(context, finished).await
|
||||
|
|
@ -1542,12 +1682,13 @@ impl ExecutionHooks for FabroHooks {
|
|||
} else {
|
||||
Site::Sandbox(Arc::clone(&acquired.env))
|
||||
};
|
||||
if !self.resumed {
|
||||
return self.check_out_source(&workspace, &site).await;
|
||||
sync::lock(&self.sites).insert(workspace.clone(), site.clone());
|
||||
if self.resumed {
|
||||
self.restore(&workspace, &site)
|
||||
.await
|
||||
.map_err(|error| ScopeAcquiredError::new(error.render()))?;
|
||||
}
|
||||
self.restore(&workspace, &site)
|
||||
.await
|
||||
.map_err(|error| ScopeAcquiredError::new(error.render()))
|
||||
self.check_out_source(&workspace, &site).await
|
||||
}
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -41,14 +41,11 @@
|
|||
//! `prepare_result` and its platform record in `transition`, around Petri's
|
||||
//! own hook service for `[[run.hooks]]`;
|
||||
//! - [`checkpoint`]: the Git snapshots of a run's workspaces, on the host or
|
||||
//! inside a Docker or Daytona sandbox, and the snapshot repository they are
|
||||
//! published to;
|
||||
//! inside a Docker or Daytona sandbox, with pushes from that same workspace;
|
||||
//! - [`source`]: where a GitHub target's workspace is checked out from, the
|
||||
//! revision, depth and read credential its in-sandbox fetch uses;
|
||||
//! - [`recovery`]: the resume-on-restart protocol, which brings every live
|
||||
//! workspace to the snapshot its durable state names: a host workspace before
|
||||
//! the run goes back to a worker, a sandbox workspace in the worker when its
|
||||
//! scope is acquired;
|
||||
//! - [`recovery`]: recovery planning from execution records; the worker resets
|
||||
//! surviving Git workspaces when their scopes are acquired;
|
||||
//! - [`platform_records`]: Fabro's platform records as the adapters reach them,
|
||||
//! in the server's database or over its API from a worker;
|
||||
//! - [`host_tools`]: Fabro's run tools on every native agent session of a run,
|
||||
|
|
@ -56,9 +53,8 @@
|
|||
//! - [`controls`]: the controls Fabro drives on a live run (pause, unpause,
|
||||
//! steer, cancel), over Petri's control service;
|
||||
//! - [`fork`]: a run seeded from another's records up to a checkpoint's
|
||||
//! position, over Petri's `host::fork_from`, with the kept checkpoints, their
|
||||
//! snapshots and the run branch carried over: what rewind, fork and retry are
|
||||
//! built on;
|
||||
//! position, over Petri's `host::fork_from`, with checkpoint metadata and the
|
||||
//! run branch carried over; the new sandbox fetches the published code;
|
||||
//! - [`prune`]: a run's sandboxes deleted through Petri's lease ledger, as
|
||||
//! `petri sandbox prune` deletes them, when Fabro deletes the run.
|
||||
//!
|
||||
|
|
|
|||
|
|
@ -1,54 +1,21 @@
|
|||
//! Resume on restart: the recovery protocol whose rule is that the
|
||||
//! workspace a resumed stage sees matches Petri's durable execution state
|
||||
//! (the integration plan's F3.5).
|
||||
//!
|
||||
//! For a Petri run the server finds in flight at startup, once the previous
|
||||
//! worker's lease is released, [`recover`] reads the durable execution
|
||||
//! state through `inspect_run` and decides:
|
||||
//!
|
||||
//! - a run with a `checkpoint_failed` finish anywhere is reported failed and
|
||||
//! not resumed: a failed checkpoint cancelled it, and nothing of it is
|
||||
//! reconciled;
|
||||
//! - otherwise, for every live execution, the last durable finish names the
|
||||
//! snapshot its workspace must sit on: the checkpoint record's commit, or,
|
||||
//! when the record was lost to the crash, the commit found by its key in the
|
||||
//! workspace's snapshot repository or history, which is then recorded again;
|
||||
//! - a workspace that survives is verified to sit on that commit, unchanged, or
|
||||
//! reset to it; a workspace that is gone, or a fresh one with no history (a
|
||||
//! fork's first acquisition), is restored from the run's snapshot repository;
|
||||
//! - a durable finish with no snapshot fails the run with a named error rather
|
||||
//! than resume it on stale files.
|
||||
//!
|
||||
//! Every child invocation's scope has its own snapshots, keyed by
|
||||
//! execution; a nested invocation that inherits its caller's sandbox shares
|
||||
//! the caller's workspace, and the workspace is brought to the newest of
|
||||
//! the live executions' snapshots on it.
|
||||
//!
|
||||
//! The decision is [`plan`], over the records and the snapshot repository
|
||||
//! alone, both on this host whatever the provider. Applying it differs: a
|
||||
//! host workspace is brought to its snapshot here, before the worker is
|
||||
//! relaunched; a Docker or Daytona workspace lives inside a sandbox only
|
||||
//! the worker's run reaches, so its target is deferred, and the worker's
|
||||
//! hooks read the same plan and apply it through the scope's environment
|
||||
//! at `scope_acquired`, before the first attempt runs there
|
||||
//! ([`bring_to`]).
|
||||
//! Resume execution from durable records using the surviving workspace.
|
||||
//! Git checkpoints are reset in place when present. Missing workspaces are
|
||||
//! never reconstructed here; explicit forks fetch their source run on GitHub
|
||||
//! when the worker acquires the new sandbox. Runs without Git checkpoints
|
||||
//! retain execution metadata without gaining workspace backup guarantees.
|
||||
|
||||
use std::collections::BTreeMap;
|
||||
use std::path::PathBuf;
|
||||
use std::sync::Arc;
|
||||
|
||||
use fabro_store::platform_records::CheckpointRecord;
|
||||
use fabro_store::{PlatformRecord, PlatformRecordKind, StagePosition};
|
||||
use fabro_store::{PlatformRecord, PlatformRecordKind};
|
||||
use fabro_types::RunId;
|
||||
use fabro_types::settings::run::RunNamespace;
|
||||
use petri_execution::host::{self, HostError};
|
||||
use petri_execution::inspect::{self, ExecutionInspection, InspectError};
|
||||
use petri_execution::{Access, InvocationId, RunKey, RunStore};
|
||||
use petri_store::StoreError;
|
||||
use tracing::info;
|
||||
|
||||
use crate::checkpoint::{
|
||||
CHECKPOINT_FAILED_CLASS, CheckpointError, CheckpointKey, RunGitSettings, RunWorkspaces, Site,
|
||||
CHECKPOINT_FAILED_CLASS, CheckpointError, CheckpointKey, RunWorkspaces, Site,
|
||||
};
|
||||
use crate::platform_records::{PlatformRecordError, PlatformRecords};
|
||||
use crate::workspace::{WorkspaceLookup, WorkspaceLookupError};
|
||||
|
|
@ -57,11 +24,8 @@ use crate::workspace::{WorkspaceLookup, WorkspaceLookupError};
|
|||
/// and its Git settings.
|
||||
pub struct RecoveryRequest {
|
||||
pub run_id: RunId,
|
||||
/// The run directory Petri ran under (the run's `petri` scratch).
|
||||
pub run_dir: PathBuf,
|
||||
pub store: Arc<dyn RunStore>,
|
||||
pub records: Arc<dyn PlatformRecords>,
|
||||
pub git: RunGitSettings,
|
||||
}
|
||||
|
||||
impl RecoveryRequest {
|
||||
|
|
@ -69,17 +33,13 @@ impl RecoveryRequest {
|
|||
#[must_use]
|
||||
pub fn for_run(
|
||||
run_id: RunId,
|
||||
run_dir: PathBuf,
|
||||
store: Arc<dyn RunStore>,
|
||||
records: Arc<dyn PlatformRecords>,
|
||||
settings: &RunNamespace,
|
||||
) -> Self {
|
||||
Self {
|
||||
run_id,
|
||||
run_dir,
|
||||
store,
|
||||
records,
|
||||
git: RunGitSettings::from(settings),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -91,8 +51,6 @@ pub enum WorkspaceAction {
|
|||
Verified,
|
||||
/// It was brought back to the snapshot.
|
||||
Reset,
|
||||
/// It was gone and was recreated from the snapshot repository.
|
||||
Restored,
|
||||
/// It lives in a sandbox this process does not reach: the worker's
|
||||
/// hooks bring it to the snapshot when its scope is acquired.
|
||||
Deferred,
|
||||
|
|
@ -112,7 +70,7 @@ pub enum Plan {
|
|||
Start,
|
||||
/// The run continues; each live workspace, by id, and its snapshot.
|
||||
Resume {
|
||||
targets: BTreeMap<String, RestoreTarget>,
|
||||
targets: BTreeMap<String, Vec<RestoreTarget>>,
|
||||
},
|
||||
/// The run cannot continue and is reported failed.
|
||||
Failed { reason: String },
|
||||
|
|
@ -162,13 +120,11 @@ pub enum RecoveryError {
|
|||
}
|
||||
|
||||
/// Decide how the run continues: the snapshot every live workspace must sit
|
||||
/// on, from the records and the snapshot repository, with a lost record
|
||||
/// reconciled from the repository. Nothing is touched.
|
||||
/// on, from execution and checkpoint records. Nothing is touched.
|
||||
pub async fn plan(
|
||||
store: Arc<dyn RunStore>,
|
||||
records: &dyn PlatformRecords,
|
||||
run_id: &RunId,
|
||||
workspaces: &RunWorkspaces,
|
||||
) -> Result<Plan, RecoveryError> {
|
||||
let key = RunKey::new(run_id.to_string());
|
||||
let logs = match store.open(&key, Access::Read).await {
|
||||
|
|
@ -204,29 +160,19 @@ pub async fn plan(
|
|||
return Ok(Plan::Failed { reason: failed });
|
||||
}
|
||||
let planner = Planner {
|
||||
records,
|
||||
run_id,
|
||||
workspaces,
|
||||
lookup: WorkspaceLookup::new(store, key),
|
||||
lookup: WorkspaceLookup::new(store, key),
|
||||
recorded: recorded_checkpoints(records, run_id).await?,
|
||||
};
|
||||
planner.targets(&inspection.executions).await
|
||||
}
|
||||
|
||||
/// Decide how the run continues, and bring its host workspaces to their
|
||||
/// snapshots; a sandbox workspace's target is deferred to the worker.
|
||||
/// Decide how the run continues. All Git work is deferred to the worker
|
||||
/// using the acquired workspace, including for the local provider.
|
||||
pub async fn recover(request: RecoveryRequest) -> Result<Recovery, RecoveryError> {
|
||||
let workspaces = RunWorkspaces::new(
|
||||
request.run_dir.clone(),
|
||||
request.run_id.to_string(),
|
||||
request.git.author.clone(),
|
||||
&request.git.checkpoint,
|
||||
);
|
||||
let targets = match plan(
|
||||
Arc::clone(&request.store),
|
||||
&*request.records,
|
||||
&request.run_id,
|
||||
&workspaces,
|
||||
)
|
||||
.await?
|
||||
{
|
||||
|
|
@ -235,62 +181,34 @@ pub async fn recover(request: RecoveryRequest) -> Result<Recovery, RecoveryError
|
|||
Plan::Resume { targets } => targets,
|
||||
};
|
||||
|
||||
let mut recovered = Vec::new();
|
||||
for (workspace, target) in targets {
|
||||
let action = if request.git.host_workspaces {
|
||||
bring_to(
|
||||
&workspaces,
|
||||
&workspaces.host(&workspace),
|
||||
&workspace,
|
||||
&target,
|
||||
)
|
||||
.await?
|
||||
} else {
|
||||
WorkspaceAction::Deferred
|
||||
};
|
||||
info!(
|
||||
run_id = %request.run_id,
|
||||
workspace,
|
||||
sha = target.sha,
|
||||
action = ?action,
|
||||
"workspace's durable snapshot decided"
|
||||
);
|
||||
recovered.push(RecoveredWorkspace {
|
||||
workspace,
|
||||
sha: target.sha,
|
||||
action,
|
||||
});
|
||||
}
|
||||
let recovered = targets
|
||||
.into_iter()
|
||||
.flat_map(|(workspace, targets)| {
|
||||
targets.into_iter().map(move |target| RecoveredWorkspace {
|
||||
workspace: workspace.clone(),
|
||||
sha: target.sha,
|
||||
action: WorkspaceAction::Deferred,
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
Ok(Recovery::Resume {
|
||||
workspaces: recovered,
|
||||
})
|
||||
}
|
||||
|
||||
/// The snapshot one live execution's last durable finish names on a
|
||||
/// workspace.
|
||||
struct Candidate {
|
||||
key: CheckpointKey,
|
||||
sha: String,
|
||||
/// The decision over one run's durable records.
|
||||
struct Planner {
|
||||
lookup: WorkspaceLookup,
|
||||
recorded: BTreeMap<CheckpointKey, (Option<String>, String)>,
|
||||
}
|
||||
|
||||
/// The decision over one run's records and snapshot repository.
|
||||
struct Planner<'a> {
|
||||
records: &'a dyn PlatformRecords,
|
||||
run_id: &'a RunId,
|
||||
workspaces: &'a RunWorkspaces,
|
||||
lookup: WorkspaceLookup,
|
||||
/// The run's checkpoint records by key: the workspace they name and
|
||||
/// the commit.
|
||||
recorded: BTreeMap<CheckpointKey, (Option<String>, String)>,
|
||||
}
|
||||
|
||||
impl Planner<'_> {
|
||||
impl Planner {
|
||||
/// The snapshot each live execution's workspace must sit on. A live
|
||||
/// execution is one whose log records no exit: `inspect_run` reports
|
||||
/// it as incomplete. A workspace several live executions share is
|
||||
/// brought to the newest of their snapshots.
|
||||
async fn targets(&self, executions: &[ExecutionInspection]) -> Result<Plan, RecoveryError> {
|
||||
let mut candidates: BTreeMap<String, Vec<Candidate>> = BTreeMap::new();
|
||||
let mut candidates: BTreeMap<String, Vec<RestoreTarget>> = BTreeMap::new();
|
||||
for execution in executions
|
||||
.iter()
|
||||
.filter(|execution| execution.status == "incomplete")
|
||||
|
|
@ -306,134 +224,26 @@ impl Planner<'_> {
|
|||
if owned.is_empty() {
|
||||
continue;
|
||||
}
|
||||
let mut found = false;
|
||||
for workspace in owned {
|
||||
if let Some(sha) = self.snapshot_of(&workspace, key).await? {
|
||||
found = true;
|
||||
candidates
|
||||
.entry(workspace)
|
||||
.or_default()
|
||||
.push(Candidate { key, sha });
|
||||
if let Some((recorded_workspace, sha)) = self.recorded.get(&key) {
|
||||
if recorded_workspace
|
||||
.as_deref()
|
||||
.is_none_or(|id| id == workspace)
|
||||
{
|
||||
candidates
|
||||
.entry(workspace)
|
||||
.or_default()
|
||||
.push(RestoreTarget {
|
||||
key,
|
||||
sha: sha.clone(),
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
return Ok(Plan::Failed {
|
||||
reason: format!(
|
||||
"no checkpoint snapshot exists for the last durable finish of {key}; the \
|
||||
run cannot resume on stale files"
|
||||
),
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
let mut targets = BTreeMap::new();
|
||||
for (workspace, candidates) in candidates {
|
||||
let target = self.newest(&workspace, &candidates).await?;
|
||||
targets.insert(workspace, target);
|
||||
}
|
||||
Ok(Plan::Resume { targets })
|
||||
}
|
||||
|
||||
/// The snapshot of `key` in `workspace`: the commit its record names,
|
||||
/// when the record names this workspace or none; else the commit found
|
||||
/// by its key in the workspace's snapshot repository or history, which
|
||||
/// is then recorded again for the record the crash lost. `None` when no
|
||||
/// snapshot exists.
|
||||
async fn snapshot_of(
|
||||
&self,
|
||||
workspace: &str,
|
||||
key: CheckpointKey,
|
||||
) -> Result<Option<String>, RecoveryError> {
|
||||
if let Some((recorded_workspace, sha)) = self.recorded.get(&key) {
|
||||
if recorded_workspace
|
||||
.as_deref()
|
||||
.is_none_or(|recorded| recorded == workspace)
|
||||
{
|
||||
return Ok(Some(sha.clone()));
|
||||
}
|
||||
}
|
||||
let found = self
|
||||
.workspaces
|
||||
.find(workspace, key)
|
||||
.await
|
||||
.map_err(|source| RecoveryError::Workspace {
|
||||
workspace: workspace.to_string(),
|
||||
source,
|
||||
})?;
|
||||
if let Some(sha) = &found {
|
||||
self.reconcile_record(key, workspace, sha).await?;
|
||||
}
|
||||
Ok(found)
|
||||
}
|
||||
|
||||
/// Write the record a crash lost, from the commit found by its key.
|
||||
async fn reconcile_record(
|
||||
&self,
|
||||
key: CheckpointKey,
|
||||
workspace: &str,
|
||||
sha: &str,
|
||||
) -> Result<(), RecoveryError> {
|
||||
info!(
|
||||
run_id = %self.run_id,
|
||||
execution = key.execution,
|
||||
firing = key.firing,
|
||||
attempt = key.attempt,
|
||||
sha,
|
||||
"checkpoint record reconciled from the run branch"
|
||||
);
|
||||
let record = PlatformRecord::Checkpoint(CheckpointRecord {
|
||||
execution: key.execution,
|
||||
firing: key.firing,
|
||||
attempt: Some(key.attempt),
|
||||
workspace: Some(workspace.to_string()),
|
||||
git_commit_sha: Some(sha.to_string()),
|
||||
diff_summary: None,
|
||||
patch_blob: None,
|
||||
operation: Some(key.operation()),
|
||||
});
|
||||
self.records
|
||||
.append(
|
||||
self.run_id,
|
||||
&record,
|
||||
Some(StagePosition {
|
||||
execution: key.execution,
|
||||
firing: key.firing,
|
||||
}),
|
||||
)
|
||||
.await
|
||||
.map_err(RecoveryError::Records)?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Of the snapshots live executions name on one workspace, the one
|
||||
/// every other descends from, else the last named. Two executions that
|
||||
/// name the same commit share it under the first one's key.
|
||||
async fn newest(
|
||||
&self,
|
||||
workspace: &str,
|
||||
candidates: &[Candidate],
|
||||
) -> Result<RestoreTarget, RecoveryError> {
|
||||
let mut chosen = &candidates[0];
|
||||
for candidate in &candidates[1..] {
|
||||
if self
|
||||
.workspaces
|
||||
.is_ancestor(workspace, &chosen.sha, &candidate.sha)
|
||||
.await
|
||||
.map_err(|source| RecoveryError::Workspace {
|
||||
workspace: workspace.to_string(),
|
||||
source,
|
||||
})?
|
||||
{
|
||||
chosen = candidate;
|
||||
}
|
||||
}
|
||||
let chosen = candidates
|
||||
.iter()
|
||||
.find(|candidate| candidate.sha == chosen.sha)
|
||||
.unwrap_or(chosen);
|
||||
Ok(RestoreTarget {
|
||||
key: chosen.key,
|
||||
sha: chosen.sha.clone(),
|
||||
Ok(Plan::Resume {
|
||||
targets: candidates,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
|
@ -500,11 +310,8 @@ async fn recorded_checkpoints(
|
|||
Ok(recorded)
|
||||
}
|
||||
|
||||
/// Verify, reset or restore the workspace at `site` onto its target: a
|
||||
/// workspace that still holds the commit is verified or reset in place; a
|
||||
/// gone one, a fresh directory with no history (a fork's first
|
||||
/// acquisition), or a sandbox whose repository lost the commit, is restored
|
||||
/// from the snapshot repository (a bundle of the snapshot, into a sandbox).
|
||||
/// Verify or reset an existing workspace. A missing commit is an error;
|
||||
/// this operation never reconstructs a lost workspace.
|
||||
pub async fn bring_to(
|
||||
workspaces: &RunWorkspaces,
|
||||
site: &Site,
|
||||
|
|
@ -530,9 +337,7 @@ pub async fn bring_to(
|
|||
workspaces.reset(site, &target.sha).await.map_err(failed)?;
|
||||
return Ok(WorkspaceAction::Reset);
|
||||
}
|
||||
workspaces
|
||||
.restore(site, workspace, target.key, &target.sha)
|
||||
.await
|
||||
.map_err(failed)?;
|
||||
Ok(WorkspaceAction::Restored)
|
||||
Err(failed(CheckpointError::MissingCommit {
|
||||
sha: target.sha.clone(),
|
||||
}))
|
||||
}
|
||||
|
|
|
|||
|
|
@ -6,10 +6,7 @@
|
|||
//! scope's environment is acquired for a fresh run, the hooks fetch the
|
||||
//! revision into the workspace from inside the scope, so the files belong to
|
||||
//! the user every later command runs as and nothing is copied in from this
|
||||
//! host (see [`crate::hooks`]). The same fetch seeds the workspace's snapshot
|
||||
//! repository with the starting commit, so every checkpoint leaves the
|
||||
//! sandbox as a bundle of the run's own commits alone, and a shallow clone's
|
||||
//! missing history is never needed. Petri's own `start` checkout is not used
|
||||
//! host (see [`crate::hooks`]). Petri's own `start` checkout is not used
|
||||
//! for a Git target: the run binds no repository for it.
|
||||
//!
|
||||
//! The credential reaches one `git` command at a time through its
|
||||
|
|
|
|||
|
|
@ -23,19 +23,23 @@ use fabro_petri::artifacts::StoreArtifactWriter;
|
|||
use fabro_petri::blobs::Blobs;
|
||||
use fabro_petri::check::{self, Bundle, CheckRequest, Launch};
|
||||
use fabro_petri::checkpoint::{
|
||||
CHECKPOINT_FAILED_CLASS, CheckpointKey, RunGitSettings, RunWorkspaces, SOURCE_REF, Site,
|
||||
CHECKPOINT_FAILED_CLASS, CheckpointKey, RunGitSettings, RunWorkspaces, Site,
|
||||
};
|
||||
use fabro_petri::controls::RunControls;
|
||||
use fabro_petri::engine::{self, Execution, RunRequest, RunStatus};
|
||||
use fabro_petri::fork;
|
||||
use fabro_petri::hooks::{HooksSpec, Publication, RunPublisher};
|
||||
use fabro_petri::platform_records::PlatformRecords;
|
||||
use fabro_petri::platform_records::{PlatformRecordError, PlatformRecords};
|
||||
use fabro_petri::providers::{DaytonaCredentials, SandboxProviderConfig};
|
||||
use fabro_petri::prune::{self, PruneRequest};
|
||||
use fabro_petri::recovery::{self, Recovery, RecoveryRequest};
|
||||
use fabro_petri::runtime::RuntimeSpec;
|
||||
use fabro_petri::source::{RunSource, SourceRevision};
|
||||
use fabro_petri::test_support::{MemoryBlobs, MemoryPlatformRecords};
|
||||
use fabro_store::{ArtifactStore, PlatformRecord, PlatformRecordKind};
|
||||
use fabro_types::settings::run::{EnvironmentResourcesSettings, RunCheckpointSettings};
|
||||
use fabro_types::settings::run::{
|
||||
EnvironmentResourcesSettings, RunCheckpointSettings, RunNamespace,
|
||||
};
|
||||
use fabro_types::{GitIdentitySource, RunId, SandboxProviderKind};
|
||||
use object_store::local::LocalFileSystem;
|
||||
use petri_execution::inspect::{self, RunInspection};
|
||||
|
|
@ -98,11 +102,12 @@ struct Harness {
|
|||
sandboxed: bool,
|
||||
/// What a successful run's work does when it ends.
|
||||
publisher: Option<Arc<dyn RunPublisher>>,
|
||||
fail_run_diff: bool,
|
||||
_root: tempfile::TempDir,
|
||||
}
|
||||
|
||||
impl Harness {
|
||||
fn new() -> Self {
|
||||
async fn new() -> Self {
|
||||
let root = tempfile::tempdir().expect("a temp dir");
|
||||
let artifact_root = root.path().join("artifacts");
|
||||
std::fs::create_dir(&artifact_root).expect("the isolated artifact directory creates");
|
||||
|
|
@ -113,6 +118,7 @@ impl Harness {
|
|||
),
|
||||
"captures-test",
|
||||
);
|
||||
let (origin, _) = upstream(&root.path().join("fixture"), 1).await;
|
||||
Self {
|
||||
artifact_store,
|
||||
run_id: RunId::new(),
|
||||
|
|
@ -121,16 +127,21 @@ impl Harness {
|
|||
records: Arc::new(MemoryPlatformRecords::new()),
|
||||
blobs: Arc::new(MemoryBlobs::new()),
|
||||
artifacts: Vec::new(),
|
||||
source: None,
|
||||
source: Some(file_source(&origin, "main", None)),
|
||||
sandboxed: false,
|
||||
publisher: None,
|
||||
fail_run_diff: false,
|
||||
_root: root,
|
||||
}
|
||||
}
|
||||
|
||||
fn hooks(&self, provider: &SandboxProviderKind) -> HooksSpec {
|
||||
HooksSpec {
|
||||
records: Arc::clone(&self.records) as Arc<dyn PlatformRecords>,
|
||||
records: if self.fail_run_diff {
|
||||
Arc::new(RejectRunDiff(self.records.clone()))
|
||||
} else {
|
||||
self.records.clone()
|
||||
},
|
||||
git: RunGitSettings {
|
||||
host_workspaces: *provider == SandboxProviderKind::LOCAL && !self.sandboxed,
|
||||
..RunGitSettings::default()
|
||||
|
|
@ -156,6 +167,16 @@ impl Harness {
|
|||
provider: SandboxProviderKind,
|
||||
workflow: &str,
|
||||
settings: &str,
|
||||
) -> engine::RunOutcome {
|
||||
self.execute_on(provider, workflow, settings, false).await
|
||||
}
|
||||
|
||||
async fn execute_on(
|
||||
&self,
|
||||
provider: SandboxProviderKind,
|
||||
workflow: &str,
|
||||
settings: &str,
|
||||
resumed: bool,
|
||||
) -> engine::RunOutcome {
|
||||
let (interviewer, observers) = no_questions();
|
||||
let hooks = self.hooks(&provider);
|
||||
|
|
@ -169,7 +190,11 @@ impl Harness {
|
|||
let request = RunRequest {
|
||||
run_id: self.run_id.to_string(),
|
||||
run_dir: self.run_dir.clone(),
|
||||
execution: Execution::Start(admit(workflow, settings)),
|
||||
execution: if resumed {
|
||||
Execution::Resume
|
||||
} else {
|
||||
Execution::Start(admit(workflow, settings))
|
||||
},
|
||||
store: Arc::clone(&self.store) as Arc<dyn petri_store::RunStore>,
|
||||
runtime: RuntimeSpec {
|
||||
sandbox,
|
||||
|
|
@ -188,27 +213,6 @@ impl Harness {
|
|||
engine::run(request).await.expect("the run executes")
|
||||
}
|
||||
|
||||
/// The commits the snapshot repository of `workspace` holds, oldest
|
||||
/// first, as `(sha, subject, key)`: every checkpoint's history, whatever
|
||||
/// site committed it.
|
||||
async fn snapshot_commits(
|
||||
&self,
|
||||
workspace: &str,
|
||||
) -> Vec<(String, String, Option<CheckpointKey>)> {
|
||||
let repository = self.workspaces().snapshot_repository(workspace);
|
||||
// Topological, so the linear run history reads parents first even
|
||||
// when commits share a timestamp.
|
||||
let log = git(&repository, &[
|
||||
"log",
|
||||
"--topo-order",
|
||||
"--reverse",
|
||||
"--all",
|
||||
"--format=%H%x00%s%x00%B%x1e",
|
||||
])
|
||||
.await;
|
||||
parse_log(&log)
|
||||
}
|
||||
|
||||
async fn inspection(&self) -> RunInspection {
|
||||
let logs = self
|
||||
.store
|
||||
|
|
@ -264,10 +268,8 @@ impl Harness {
|
|||
async fn recover(&self) -> Recovery {
|
||||
recovery::recover(RecoveryRequest {
|
||||
run_id: self.run_id,
|
||||
run_dir: self.run_dir.clone(),
|
||||
store: Arc::clone(&self.store) as Arc<dyn petri_store::RunStore>,
|
||||
records: Arc::clone(&self.records) as Arc<dyn PlatformRecords>,
|
||||
git: RunGitSettings::default(),
|
||||
})
|
||||
.await
|
||||
.expect("recovery decides")
|
||||
|
|
@ -316,6 +318,7 @@ fn parse_log(log: &str) -> Vec<(String, String, Option<CheckpointKey>)> {
|
|||
let body = parts.next().unwrap_or_default();
|
||||
(sha, subject, CheckpointKey::from_message(body))
|
||||
})
|
||||
.filter(|(_, _, key)| key.is_some())
|
||||
.collect()
|
||||
}
|
||||
|
||||
|
|
@ -333,9 +336,10 @@ fn stages(inspection: &RunInspection) -> Vec<(String, String)> {
|
|||
/// A dry run keeps checkpointable local workspaces even when the selected
|
||||
/// environment would require Docker or Daytona. Its command is simulated.
|
||||
#[tokio::test]
|
||||
async fn dry_runs_use_local_workspaces_and_checkpoint_without_sandbox_credentials() {
|
||||
async fn dry_runs_use_local_workspaces_without_git_or_sandbox_credentials() {
|
||||
for provider in [SandboxProviderKind::DOCKER, SandboxProviderKind::DAYTONA] {
|
||||
let harness = Harness::new();
|
||||
let mut harness = Harness::new().await;
|
||||
harness.source = None;
|
||||
let workflow = workflow(
|
||||
r#" write [shape=parallelogram, script="touch should-not-exist; exit 1"]"#,
|
||||
" start -> write -> exit",
|
||||
|
|
@ -379,10 +383,10 @@ async fn dry_runs_use_local_workspaces_and_checkpoint_without_sandbox_credential
|
|||
);
|
||||
assert_eq!(
|
||||
harness.checkpoints().len(),
|
||||
3,
|
||||
"{provider}: every stage checkpoints"
|
||||
0,
|
||||
"{provider}: dry runs do not create Git checkpoints"
|
||||
);
|
||||
assert_eq!(harness.snapshot_commits(&workspace).await.len(), 3);
|
||||
assert!(!harness.workspace_path(&workspace).join(".git").exists());
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -391,7 +395,7 @@ async fn dry_runs_use_local_workspaces_and_checkpoint_without_sandbox_credential
|
|||
/// reached Petri's local service through Fabro's wrapper.
|
||||
#[tokio::test]
|
||||
async fn every_finish_is_committed_and_recorded() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" write [shape=parallelogram, script=\"echo one > out.txt\"]\n check \
|
||||
[shape=parallelogram, script=\"test \\\"$(cat out.txt)\\\" = one\"]",
|
||||
|
|
@ -442,13 +446,7 @@ async fn every_finish_is_committed_and_recorded() {
|
|||
"{checkpoints:?}"
|
||||
);
|
||||
|
||||
// The snapshot repository holds every checkpoint.
|
||||
let published = harness
|
||||
.workspaces()
|
||||
.published(&workspace)
|
||||
.await
|
||||
.expect("the snapshots list");
|
||||
assert_eq!(published.len(), 4);
|
||||
assert!(!harness.run_dir.join("snapshots").exists());
|
||||
|
||||
// `run_complete` and `sandbox_cleanup` ran through the forwarded
|
||||
// service, with the sandbox in place.
|
||||
|
|
@ -465,7 +463,7 @@ async fn every_finish_is_committed_and_recorded() {
|
|||
/// the end.
|
||||
#[tokio::test]
|
||||
async fn artifacts_the_branch_and_the_diffs_are_recorded() {
|
||||
let mut harness = Harness::new();
|
||||
let mut harness = Harness::new().await;
|
||||
harness.artifacts = vec!["assets/**".to_string()];
|
||||
let workflow = workflow(
|
||||
" write [shape=parallelogram, script=\"mkdir -p assets && printf one > \
|
||||
|
|
@ -544,8 +542,15 @@ async fn artifacts_the_branch_and_the_diffs_are_recorded() {
|
|||
let checkpoints = harness.checkpoints();
|
||||
assert_eq!(
|
||||
branches[0].base_sha.as_deref(),
|
||||
Some(checkpoints[0].1.as_str()),
|
||||
"a branch in a fresh repository starts from its first checkpoint"
|
||||
Some(
|
||||
git(&harness.workspace_path(&workspace), &[
|
||||
"rev-parse",
|
||||
&format!("{}^", checkpoints[0].1)
|
||||
])
|
||||
.await
|
||||
.as_str()
|
||||
),
|
||||
"the branch starts at the upstream commit"
|
||||
);
|
||||
let identities: Vec<_> = records
|
||||
.iter()
|
||||
|
|
@ -571,7 +576,8 @@ async fn artifacts_the_branch_and_the_diffs_are_recorded() {
|
|||
})
|
||||
.collect();
|
||||
assert_eq!(diffs.len(), 5, "{diffs:?}");
|
||||
assert_eq!(diffs[0], (None, false));
|
||||
assert_eq!(diffs[0].0.unwrap().files_changed, 0);
|
||||
assert!(!diffs[0].1);
|
||||
let write = diffs[1].0.expect("the write diff");
|
||||
assert_eq!(
|
||||
(write.files_changed, write.additions, write.deletions),
|
||||
|
|
@ -645,7 +651,7 @@ async fn checkpoint_nodes(harness: &Harness) -> Vec<(String, u64)> {
|
|||
/// and its failure route runs on the committed files.
|
||||
#[tokio::test]
|
||||
async fn a_failed_stage_is_committed_and_its_route_sees_the_files() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" work [shape=parallelogram, script=\"echo partial > out.txt; exit 1\"]\n fix \
|
||||
[shape=parallelogram, script=\"test \\\"$(cat out.txt)\\\" = partial && echo fixed >> \
|
||||
|
|
@ -685,7 +691,7 @@ async fn a_failed_stage_is_committed_and_its_route_sees_the_files() {
|
|||
/// checkpoint's error, and a restart reports it failed without resuming.
|
||||
#[tokio::test]
|
||||
async fn a_failed_checkpoint_ends_the_run_with_no_route() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" wreck [shape=parallelogram, script=\"rm -rf .git && echo garbage > .git && echo wrecked \
|
||||
> out.txt\"]\n next [shape=parallelogram, script=\"echo next > next.txt\"]\n fix \
|
||||
|
|
@ -748,7 +754,7 @@ async fn a_failed_checkpoint_ends_the_run_with_no_route() {
|
|||
/// starts over, and a run that finished has nothing to bring back.
|
||||
#[tokio::test]
|
||||
async fn recovery_starts_an_unknown_run_and_resumes_a_finished_one() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
assert_eq!(harness.recover().await, Recovery::Start);
|
||||
|
||||
let workflow = workflow(
|
||||
|
|
@ -815,7 +821,7 @@ async fn a_run_hook_blocks_a_tool_effect_through_the_forwarded_service() {
|
|||
.expect("the model client builds")
|
||||
.expect("openai is eligible");
|
||||
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let (interviewer, observers) = no_questions();
|
||||
let workflow = format!(
|
||||
"digraph Hooks {{\n graph [backend=\"api\", goal=\"Check the tool hooks\", \
|
||||
|
|
@ -892,7 +898,7 @@ async fn the_records_name_the_root_invocations_workspace() {
|
|||
use fabro_petri::workspace::WorkspaceLookup;
|
||||
use petri_execution::InvocationId;
|
||||
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" write [shape=parallelogram, script=\"echo one > out.txt\"]",
|
||||
" start -> write -> exit",
|
||||
|
|
@ -923,7 +929,7 @@ async fn the_records_name_the_root_invocations_workspace() {
|
|||
/// problem.
|
||||
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
|
||||
async fn parallel_branches_checkpoint_the_shared_workspace_in_turn() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" fork [shape=component]\n a [shape=parallelogram, script=\"echo a > a.txt\"]\n b \
|
||||
[shape=parallelogram, script=\"echo b > b.txt\"]\n merge [shape=tripleoctagon]\n check \
|
||||
|
|
@ -968,6 +974,19 @@ async fn parallel_branches_checkpoint_the_shared_workspace_in_turn() {
|
|||
);
|
||||
assert_eq!(inspection.executions.len(), 3, "the root and two branches");
|
||||
assert_eq!(harness.workspace().await, "invocation-0-scope-0");
|
||||
let child = recorded.iter().find(|key| key.execution != 0).unwrap();
|
||||
let refused = fork::check(
|
||||
harness.store.as_ref(),
|
||||
harness.run_id,
|
||||
fork::position(child.execution, child.firing),
|
||||
)
|
||||
.await
|
||||
.unwrap_err();
|
||||
assert!(
|
||||
refused
|
||||
.to_string()
|
||||
.contains("inside a child invocation cannot be forked")
|
||||
);
|
||||
}
|
||||
|
||||
/// On Docker the workspace lives inside the container: every finished
|
||||
|
|
@ -996,78 +1015,44 @@ async fn a_daytona_run_commits_inside_the_sandbox_and_publishes_every_checkpoint
|
|||
assert_sandbox_run_publishes_every_checkpoint(SandboxProviderKind::DAYTONA).await;
|
||||
}
|
||||
|
||||
/// Bytes of incompressible data the first stage writes: past the plugin
|
||||
/// transport's 16 MiB cap on one file read, so its bundle leaves the
|
||||
/// sandbox in more than one part.
|
||||
const LARGE_FILE_BYTES: usize = 20 * 1024 * 1024;
|
||||
|
||||
/// A two-stage run on `provider`, whose workspace lives inside a sandbox:
|
||||
/// nothing of it is on the host, every checkpoint is published, and the
|
||||
/// bundles carried the stages' files, a large one in parts.
|
||||
/// Git commits and pushes use the acquired sandbox, without a host repository.
|
||||
async fn assert_sandbox_run_publishes_every_checkpoint(provider: SandboxProviderKind) {
|
||||
let harness = Harness::new();
|
||||
let mut harness = Harness::new().await;
|
||||
harness.source = Some(RunSource {
|
||||
origin: "https://github.com/octocat/Hello-World.git".to_string(),
|
||||
revision: SourceRevision::Branch("master".to_string()),
|
||||
branch: "master".to_string(),
|
||||
depth: Some(1),
|
||||
credential: None,
|
||||
});
|
||||
let publisher = RecordingPublisher::new(None);
|
||||
harness.publisher = Some(publisher.clone());
|
||||
let workflow = workflow(
|
||||
&format!(
|
||||
" write [shape=parallelogram, script=\"echo one > out.txt && head -c \
|
||||
{LARGE_FILE_BYTES} /dev/urandom > large.bin\"]\n check [shape=parallelogram, \
|
||||
script=\"test \\\"$(cat out.txt)\\\" = one && git log --format=%s | head -1 | grep -q \
|
||||
write\"]"
|
||||
),
|
||||
r#" write [shape=parallelogram, script="echo one > out.txt"]
|
||||
check [shape=parallelogram, script="test -f out.txt && git log --format=%s | head -1 | grep -q write"]"#,
|
||||
" start -> write -> check -> exit",
|
||||
);
|
||||
let outcome = harness.run_on(provider, &workflow, SETTINGS).await;
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
assert!(outcome.complete, "{:?}", outcome.incomplete);
|
||||
|
||||
let checkpoints = harness.checkpoints();
|
||||
assert_eq!(checkpoints.len(), 4, "{checkpoints:?}");
|
||||
let workspace = "invocation-0-scope-0";
|
||||
assert!(
|
||||
!harness.workspaces().workspace_exists(workspace).await,
|
||||
"nothing of the workspace is on the host"
|
||||
);
|
||||
let published = harness
|
||||
.workspaces()
|
||||
.published(workspace)
|
||||
let outcome = harness.run_on(provider.clone(), &workflow, SETTINGS).await;
|
||||
if provider == SandboxProviderKind::DOCKER {
|
||||
prune::prune(PruneRequest {
|
||||
sandbox: SandboxProviderConfig::from_lookup(None, |name| env::var(name).ok()),
|
||||
run_id: harness.run_id.to_string(),
|
||||
run_dir: harness.run_dir.clone(),
|
||||
store: harness.store.clone(),
|
||||
provider,
|
||||
})
|
||||
.await
|
||||
.expect("the snapshot repository lists");
|
||||
let mut by_key: Vec<(CheckpointKey, String)> = published
|
||||
.iter()
|
||||
.map(|snapshot| (snapshot.key, snapshot.sha.clone()))
|
||||
.collect();
|
||||
let mut recorded = checkpoints.clone();
|
||||
recorded.sort();
|
||||
by_key.sort();
|
||||
assert_eq!(by_key, recorded, "every record names a published snapshot");
|
||||
|
||||
let commits = harness.snapshot_commits(workspace).await;
|
||||
let subjects: Vec<&str> = commits
|
||||
.iter()
|
||||
.map(|(_, subject, _)| subject.as_str())
|
||||
.collect();
|
||||
let run_id = harness.run_id.to_string();
|
||||
assert_eq!(subjects, vec![
|
||||
format!("fabro({run_id}): start (success)"),
|
||||
format!("fabro({run_id}): write (success)"),
|
||||
format!("fabro({run_id}): check (success)"),
|
||||
format!("fabro({run_id}): exit (success)"),
|
||||
]);
|
||||
let (write_sha, _, _) = &commits[1];
|
||||
let repository = harness.workspaces().snapshot_repository(workspace);
|
||||
assert_eq!(
|
||||
git(&repository, &["show", &format!("{write_sha}:out.txt")]).await,
|
||||
"one",
|
||||
"the bundle carried the stage's files"
|
||||
);
|
||||
assert_eq!(
|
||||
git(&repository, &[
|
||||
"cat-file",
|
||||
"-s",
|
||||
&format!("{write_sha}:large.bin")
|
||||
])
|
||||
.await,
|
||||
LARGE_FILE_BYTES.to_string(),
|
||||
"the large file came through the split transfer whole"
|
||||
.expect("the test container is pruned");
|
||||
}
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
assert_eq!(harness.checkpoints().len(), 4);
|
||||
assert_eq!(publisher.pushed.lock().unwrap().len(), 4);
|
||||
assert!(!harness.run_dir.join("snapshots").exists());
|
||||
assert!(
|
||||
!harness
|
||||
.workspaces()
|
||||
.workspace_exists("invocation-0-scope-0")
|
||||
.await
|
||||
);
|
||||
}
|
||||
|
||||
|
|
@ -1118,7 +1103,7 @@ async fn upstream(root: &Path, commits: usize) -> (PathBuf, String) {
|
|||
#[tokio::test]
|
||||
async fn a_git_source_is_checked_out_shallow_and_checkpoints_build_on_its_commit() {
|
||||
for sandboxed in [false, true] {
|
||||
let mut harness = Harness::new();
|
||||
let mut harness = Harness::new().await;
|
||||
let (origin, head) = upstream(&harness.run_dir.with_file_name("upstream"), 3).await;
|
||||
harness.sandboxed = sandboxed;
|
||||
harness.source = Some(file_source(&origin, "main", Some(1)));
|
||||
|
|
@ -1139,16 +1124,8 @@ async fn a_git_source_is_checked_out_shallow_and_checkpoints_build_on_its_commit
|
|||
|
||||
let workspace = harness.workspace().await;
|
||||
let workspaces = harness.workspaces();
|
||||
assert_eq!(
|
||||
workspaces
|
||||
.source_base(&workspace)
|
||||
.await
|
||||
.expect("the base reads"),
|
||||
Some(head.clone()),
|
||||
"sandboxed={sandboxed}: the snapshot repository is seeded with the starting commit"
|
||||
);
|
||||
let repository = workspaces.snapshot_repository(&workspace);
|
||||
assert_eq!(git(&repository, &["rev-parse", SOURCE_REF]).await, head);
|
||||
assert!(!harness.run_dir.join("snapshots").exists());
|
||||
let repository = workspaces.workspace_path(&workspace);
|
||||
let checkpoints = harness.checkpoints();
|
||||
let (_, last) = checkpoints.last().expect("a checkpoint was recorded");
|
||||
assert_eq!(
|
||||
|
|
@ -1228,6 +1205,7 @@ async fn an_unavailable_revision_fails_the_checkout() {
|
|||
|
||||
/// A publisher that records what it was handed and answers as told.
|
||||
struct RecordingPublisher {
|
||||
pushed: std::sync::Mutex<Vec<(Site, String, String)>>,
|
||||
published: std::sync::Mutex<Vec<Publication>>,
|
||||
fail: Option<String>,
|
||||
}
|
||||
|
|
@ -1235,6 +1213,7 @@ struct RecordingPublisher {
|
|||
impl RecordingPublisher {
|
||||
fn new(fail: Option<&str>) -> Arc<Self> {
|
||||
Arc::new(Self {
|
||||
pushed: std::sync::Mutex::default(),
|
||||
published: std::sync::Mutex::default(),
|
||||
fail: fail.map(str::to_string),
|
||||
})
|
||||
|
|
@ -1254,6 +1233,13 @@ fn file_source(origin: &Path, branch: &str, depth: Option<u32>) -> RunSource {
|
|||
|
||||
#[async_trait::async_trait]
|
||||
impl RunPublisher for RecordingPublisher {
|
||||
async fn push(&self, site: &Site, branch: &str, sha: &str) -> Result<(), String> {
|
||||
self.pushed
|
||||
.lock()
|
||||
.unwrap()
|
||||
.push((site.clone(), branch.to_string(), sha.to_string()));
|
||||
Ok(())
|
||||
}
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String> {
|
||||
self.published
|
||||
.lock()
|
||||
|
|
@ -1269,7 +1255,7 @@ async fn published_run(
|
|||
attributes: &str,
|
||||
publisher: &Arc<RecordingPublisher>,
|
||||
) -> (Harness, engine::RunOutcome) {
|
||||
let mut harness = Harness::new();
|
||||
let mut harness = Harness::new().await;
|
||||
let (origin, _) = upstream(&harness.run_dir.with_file_name("upstream"), 2).await;
|
||||
harness.source = Some(file_source(&origin, "main", Some(1)));
|
||||
harness.publisher = Some(Arc::clone(publisher) as Arc<dyn RunPublisher>);
|
||||
|
|
@ -1302,7 +1288,10 @@ async fn a_successful_run_is_published_with_its_branch_head_and_patch() {
|
|||
let (_, last) = harness.checkpoints().last().cloned().expect("a checkpoint");
|
||||
assert_eq!(publication.head_sha, last);
|
||||
assert_eq!(
|
||||
git(&publication.snapshot_repository, &["cat-file", "-t", &last]).await,
|
||||
git(&harness.workspace_path(&harness.workspace().await), &[
|
||||
"cat-file", "-t", &last
|
||||
])
|
||||
.await,
|
||||
"commit"
|
||||
);
|
||||
assert!(
|
||||
|
|
@ -1331,3 +1320,275 @@ async fn a_failed_run_is_not_published() {
|
|||
assert!(!outcome.publish_failed);
|
||||
assert!(publisher.published.lock().unwrap().is_empty());
|
||||
}
|
||||
|
||||
struct OriginPublisher {
|
||||
origin: String,
|
||||
pushed: std::sync::Mutex<Vec<String>>,
|
||||
}
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl RunPublisher for OriginPublisher {
|
||||
async fn push(&self, site: &Site, branch: &str, sha: &str) -> Result<(), String> {
|
||||
assert!(
|
||||
matches!(site, Site::Sandbox(_)),
|
||||
"push runs through the sandbox environment"
|
||||
);
|
||||
site.push(
|
||||
&self.origin,
|
||||
&format!("{sha}:refs/heads/{branch}"),
|
||||
&[],
|
||||
std::time::Duration::from_secs(30),
|
||||
)
|
||||
.await
|
||||
.map_err(|error| error.to_string())?;
|
||||
self.pushed.lock().unwrap().push(sha.to_owned());
|
||||
Ok(())
|
||||
}
|
||||
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String> {
|
||||
self.push(
|
||||
&publication.site,
|
||||
&publication.run_branch,
|
||||
&publication.head_sha,
|
||||
)
|
||||
.await
|
||||
}
|
||||
}
|
||||
|
||||
/// The shallow checkout regression: a fork fetches a published checkpoint
|
||||
/// directly from the origin inside its new sandbox, after the old workspace
|
||||
/// has gone. No server Git refs or bundles participate.
|
||||
#[tokio::test]
|
||||
async fn a_shallow_run_pushes_each_checkpoint_and_its_fork_fetches_from_the_origin() {
|
||||
assert_shallow_fork(1).await;
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn a_terminal_checkpoint_is_refused_before_creating_a_fork() {
|
||||
assert_shallow_fork(3).await;
|
||||
}
|
||||
|
||||
async fn assert_shallow_fork(checkpoint_index: usize) {
|
||||
use fabro_petri::fork::{self, ForkRequest};
|
||||
let mut original = Harness::new().await;
|
||||
let (origin, _) = upstream(&original.run_dir.with_file_name("remote"), 5).await;
|
||||
original.source = Some(file_source(&origin, "main", Some(1)));
|
||||
original.sandboxed = true;
|
||||
let publisher = Arc::new(OriginPublisher {
|
||||
origin: format!("file://{}", origin.display()),
|
||||
pushed: std::sync::Mutex::default(),
|
||||
});
|
||||
original.publisher = Some(publisher.clone());
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="printf 'durable bytes\nsecond line\n' > result.txt"]
|
||||
verify [shape=parallelogram, script="test -f result.txt && test -f README.md && cat result.txt"]"#,
|
||||
" start -> write -> verify -> exit",
|
||||
);
|
||||
let outcome = original.run(&graph, SETTINGS).await;
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
let checkpoints = original.checkpoints();
|
||||
let pushed = publisher.pushed.lock().unwrap().clone();
|
||||
assert_eq!(
|
||||
&pushed[..checkpoints.len()],
|
||||
checkpoints
|
||||
.iter()
|
||||
.map(|(_, sha)| sha.clone())
|
||||
.collect::<Vec<_>>()
|
||||
);
|
||||
assert!(!original.run_dir.join("snapshots").exists());
|
||||
|
||||
let mut forked = Harness::new().await;
|
||||
forked.store = original.store.clone();
|
||||
forked.records = original.records.clone();
|
||||
forked.source = original.source.clone();
|
||||
forked.sandboxed = true;
|
||||
forked.publisher = Some(publisher.clone());
|
||||
let (key, checkpoint_sha) = &checkpoints[checkpoint_index];
|
||||
if checkpoint_index == checkpoints.len() - 1 {
|
||||
let refused = fork::check(
|
||||
original.store.as_ref(),
|
||||
original.run_id,
|
||||
fork::position(key.execution, key.firing),
|
||||
)
|
||||
.await
|
||||
.expect_err("terminal checkpoint is refused");
|
||||
assert!(
|
||||
refused
|
||||
.to_string()
|
||||
.contains("terminal checkpoint has no remaining work")
|
||||
);
|
||||
assert!(
|
||||
!forked.run_dir.exists(),
|
||||
"refuse before creating a new run or sandbox"
|
||||
);
|
||||
return;
|
||||
}
|
||||
let seeded = fork::fork(ForkRequest {
|
||||
source: original.run_id,
|
||||
fork: forked.run_id,
|
||||
fork_run_dir: forked.run_dir.clone(),
|
||||
store: forked.store.clone(),
|
||||
records: forked.records.clone(),
|
||||
position: fork::position(key.execution, key.firing),
|
||||
rerun_last: false,
|
||||
settings: RunNamespace::default(),
|
||||
})
|
||||
.await
|
||||
.expect("the fork is seeded");
|
||||
assert_eq!(
|
||||
seeded.start.expect("the selected checkpoint").sha,
|
||||
*checkpoint_sha
|
||||
);
|
||||
fs::remove_dir_all(&original.run_dir)
|
||||
.await
|
||||
.expect("the original workspace is deleted");
|
||||
let outcome = forked
|
||||
.execute_on(SandboxProviderKind::LOCAL, &graph, SETTINGS, true)
|
||||
.await;
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
let path = forked.workspace_path(&forked.workspace().await);
|
||||
if let Some((_, first_new)) = forked.checkpoints().get(checkpoint_index + 1) {
|
||||
assert_eq!(
|
||||
git(&path, &["rev-parse", &format!("{first_new}^")]).await,
|
||||
*checkpoint_sha,
|
||||
"the fork continues from the selected checkpoint, not the source run's final head"
|
||||
);
|
||||
} else {
|
||||
assert_eq!(git(&path, &["rev-parse", "HEAD"]).await, *checkpoint_sha);
|
||||
}
|
||||
assert_eq!(
|
||||
fs::read(path.join("result.txt"))
|
||||
.await
|
||||
.expect("the recovered file reads"),
|
||||
b"durable bytes\nsecond line\n"
|
||||
);
|
||||
assert_eq!(
|
||||
git(&path, &["branch", "--show-current"]).await,
|
||||
format!("fabro/run/{}", forked.run_id)
|
||||
);
|
||||
assert_eq!(
|
||||
git(&origin, &[
|
||||
"rev-parse",
|
||||
&format!("refs/heads/fabro/run/{}", forked.run_id)
|
||||
])
|
||||
.await,
|
||||
git(&path, &["rev-parse", "HEAD"]).await
|
||||
);
|
||||
assert!(!forked.run_dir.join("snapshots").exists());
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn an_empty_workspace_keeps_execution_records_without_git_checkpoints() {
|
||||
let mut harness = Harness::new().await;
|
||||
harness.source = None;
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="echo data > result.txt"]"#,
|
||||
" start -> write -> exit",
|
||||
);
|
||||
let outcome = harness.run(&graph, SETTINGS).await;
|
||||
assert_eq!(outcome.status, RunStatus::Success);
|
||||
assert!(harness.checkpoints().is_empty());
|
||||
let path = harness.workspace_path(&harness.workspace().await);
|
||||
assert!(!path.join(".git").exists());
|
||||
let records = harness
|
||||
.records
|
||||
.read_kind(&harness.run_id, PlatformRecordKind::Checkpoint)
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(records.len(), 3);
|
||||
assert!(!harness.run_dir.join("snapshots").exists());
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn ordinary_recovery_refuses_a_missing_workspace() {
|
||||
use fabro_petri::recovery::RestoreTarget;
|
||||
let harness = Harness::new().await;
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="echo data > result.txt"]"#,
|
||||
" start -> write -> exit",
|
||||
);
|
||||
assert_eq!(
|
||||
harness.run(&graph, SETTINGS).await.status,
|
||||
RunStatus::Success
|
||||
);
|
||||
let workspace = harness.workspace().await;
|
||||
let path = harness.workspace_path(&workspace);
|
||||
let (key, sha) = harness.checkpoints().last().unwrap().clone();
|
||||
fs::remove_dir_all(&path).await.unwrap();
|
||||
let result = recovery::bring_to(
|
||||
&harness.workspaces(),
|
||||
&Site::Host(path.clone()),
|
||||
&workspace,
|
||||
&RestoreTarget { key, sha },
|
||||
)
|
||||
.await;
|
||||
assert!(result.is_err());
|
||||
assert!(
|
||||
!path.exists(),
|
||||
"normal resume does not recreate the checkout"
|
||||
);
|
||||
}
|
||||
|
||||
struct RejectRunDiff(Arc<MemoryPlatformRecords>);
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl PlatformRecords for RejectRunDiff {
|
||||
async fn append(
|
||||
&self,
|
||||
run_id: &RunId,
|
||||
record: &PlatformRecord,
|
||||
position: Option<fabro_store::StagePosition>,
|
||||
) -> Result<fabro_store::StoredPlatformRecord, PlatformRecordError> {
|
||||
if matches!(record, PlatformRecord::RunDiff(_)) {
|
||||
return Err(PlatformRecordError::Store(fabro_store::Error::Io(
|
||||
std::io::Error::other("run diff unavailable"),
|
||||
)));
|
||||
}
|
||||
self.0.append(run_id, record, position).await
|
||||
}
|
||||
async fn read_kind(
|
||||
&self,
|
||||
run_id: &RunId,
|
||||
kind: PlatformRecordKind,
|
||||
) -> Result<Vec<fabro_store::StoredPlatformRecord>, PlatformRecordError> {
|
||||
self.0.read_kind(run_id, kind).await
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn a_run_diff_failure_cannot_silently_skip_publication() {
|
||||
let mut harness = Harness::new().await;
|
||||
harness.fail_run_diff = true;
|
||||
let publisher = RecordingPublisher::new(None);
|
||||
harness.publisher = Some(publisher.clone());
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="echo data > result.txt"]"#,
|
||||
" start -> write -> exit",
|
||||
);
|
||||
let outcome = harness.run(&graph, SETTINGS).await;
|
||||
assert_eq!(outcome.status, RunStatus::Failed, "{outcome:?}");
|
||||
assert!(outcome.publish_failed);
|
||||
assert!(outcome.failure.unwrap().contains("run diff unavailable"));
|
||||
assert!(publisher.published.lock().unwrap().is_empty());
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn a_local_repository_without_a_git_target_gets_no_automatic_commits() {
|
||||
let mut harness = Harness::new().await;
|
||||
harness.source = None;
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="git init -q && echo data > result.txt && git add result.txt && git -c user.name=User -c user.email=user@example.com commit -qm user-commit"]"#,
|
||||
" start -> write -> exit",
|
||||
);
|
||||
assert_eq!(
|
||||
harness.run(&graph, SETTINGS).await.status,
|
||||
RunStatus::Success
|
||||
);
|
||||
let path = harness.workspace_path(&harness.workspace().await);
|
||||
assert_eq!(git(&path, &["rev-list", "--count", "HEAD"]).await, "1");
|
||||
assert_eq!(
|
||||
git(&path, &["log", "--format=%s", "-1"]).await,
|
||||
"user-commit"
|
||||
);
|
||||
assert!(harness.checkpoints().is_empty());
|
||||
}
|
||||
|
|
|
|||
16
lib/packages/fabro-api-client/src/api/runs-api.ts
generated
16
lib/packages/fabro-api-client/src/api/runs-api.ts
generated
|
|
@ -546,7 +546,7 @@ export const RunsApiAxiosParamCreator = function (configuration?: Configuration)
|
|||
};
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead.
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead. Requires a GitHub target with run-branch creation and pushes enabled; the checkpoint commit must be available on the source run\'s published branch. The final terminal checkpoint is refused because it leaves no work to acquire a workspace. For a completed run, select an earlier checkpoint explicitly or retry the workflow from the beginning.
|
||||
* @summary Fork Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {ForkRequest} [forkRequest]
|
||||
|
|
@ -1204,7 +1204,7 @@ export const RunsApiAxiosParamCreator = function (configuration?: Configuration)
|
|||
};
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed.
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed. Requires a GitHub target with run-branch creation and pushes enabled. The checkpoint must leave work to execute: the final terminal checkpoint is refused; select an earlier checkpoint or retry the workflow from the beginning.
|
||||
* @summary Rewind Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {RewindRequest} [rewindRequest]
|
||||
|
|
@ -1732,7 +1732,7 @@ export const RunsApiFp = function(configuration?: Configuration) {
|
|||
return (axios, basePath) => createRequestFunction(localVarAxiosArgs, globalAxios, BASE_PATH, configuration)(axios, localVarOperationServerBasePath || basePath);
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead.
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead. Requires a GitHub target with run-branch creation and pushes enabled; the checkpoint commit must be available on the source run\'s published branch. The final terminal checkpoint is refused because it leaves no work to acquire a workspace. For a completed run, select an earlier checkpoint explicitly or retry the workflow from the beginning.
|
||||
* @summary Fork Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {ForkRequest} [forkRequest]
|
||||
|
|
@ -1938,7 +1938,7 @@ export const RunsApiFp = function(configuration?: Configuration) {
|
|||
return (axios, basePath) => createRequestFunction(localVarAxiosArgs, globalAxios, BASE_PATH, configuration)(axios, localVarOperationServerBasePath || basePath);
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed.
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed. Requires a GitHub target with run-branch creation and pushes enabled. The checkpoint must leave work to execute: the final terminal checkpoint is refused; select an earlier checkpoint or retry the workflow from the beginning.
|
||||
* @summary Rewind Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {RewindRequest} [rewindRequest]
|
||||
|
|
@ -2180,7 +2180,7 @@ export const RunsApiFactory = function (configuration?: Configuration, basePath?
|
|||
return localVarFp.denyRun(id, denyRunRequest, options).then((request) => request(axios, basePath));
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead.
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead. Requires a GitHub target with run-branch creation and pushes enabled; the checkpoint commit must be available on the source run\'s published branch. The final terminal checkpoint is refused because it leaves no work to acquire a workspace. For a completed run, select an earlier checkpoint explicitly or retry the workflow from the beginning.
|
||||
* @summary Fork Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {ForkRequest} [forkRequest]
|
||||
|
|
@ -2341,7 +2341,7 @@ export const RunsApiFactory = function (configuration?: Configuration, basePath?
|
|||
return localVarFp.retryRun(id, options).then((request) => request(axios, basePath));
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed.
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed. Requires a GitHub target with run-branch creation and pushes enabled. The checkpoint must leave work to execute: the final terminal checkpoint is refused; select an earlier checkpoint or retry the workflow from the beginning.
|
||||
* @summary Rewind Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {RewindRequest} [rewindRequest]
|
||||
|
|
@ -2565,7 +2565,7 @@ export class RunsApi extends BaseAPI {
|
|||
}
|
||||
|
||||
/**
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead.
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead. Requires a GitHub target with run-branch creation and pushes enabled; the checkpoint commit must be available on the source run\'s published branch. The final terminal checkpoint is refused because it leaves no work to acquire a workspace. For a completed run, select an earlier checkpoint explicitly or retry the workflow from the beginning.
|
||||
* @summary Fork Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {ForkRequest} [forkRequest]
|
||||
|
|
@ -2741,7 +2741,7 @@ export class RunsApi extends BaseAPI {
|
|||
}
|
||||
|
||||
/**
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed.
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed. Requires a GitHub target with run-branch creation and pushes enabled. The checkpoint must leave work to execute: the final terminal checkpoint is refused; select an earlier checkpoint or retry the workflow from the beginning.
|
||||
* @summary Rewind Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {RewindRequest} [rewindRequest]
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue