mirror of
https://github.com/fabro-sh/fabro.git
synced 2026-10-08 03:10:26 +00:00
Merge pull request #913 from fabro-sh/restore-github-run-publication
Keep GitHub checkout, checkpoints, and pushes in the run sandbox
This commit is contained in:
commit
e7b4859038
22 changed files with 2334 additions and 1570 deletions
50
AGENTS.md
50
AGENTS.md
|
|
@ -30,12 +30,25 @@ macOS note: if `cargo nextest run` fails with `Too many open files (os error 24)
|
|||
### Docker sandbox provider
|
||||
- Docker is the default runtime sandbox provider from `defaults.toml`. The Fabro process must have a working Docker client environment (`DOCKER_HOST`, socket access, Docker Desktop behavior, TLS settings, groups/permissions, and any remote daemon policy are operator responsibilities).
|
||||
- The packaged compose service mounts `/var/run/docker.sock` so the server can create sibling run containers on the host daemon. This is host-root-equivalent under Docker's security model; only use it in the trusted, single-tenant deployment model described by the sandbox code/docs.
|
||||
- Fabro no longer clones a repository into a sandbox: the engine prepares
|
||||
every run's checkout. `CloneRequest` still travels beside the sandbox spec
|
||||
so the run record names the origin and branch; fabro validates it (a pin
|
||||
needs a branch, a non-GitHub origin needs `skip_clone`) and refuses a
|
||||
request that asks for a clone. Preflight and `fabro exec` initialize
|
||||
sandboxes with `CloneRequest::none()`, which creates an empty workspace.
|
||||
- A GitHub target's workspace is checked out by Fabro's hooks, not by the
|
||||
sandbox driver or Petri's `start` checkout: when a fresh run's scope is
|
||||
acquired, `fabro-petri`'s `RunWorkspaces::check_out_source` fetches the
|
||||
target's revision inside the sandbox with a read-only token the worker
|
||||
mints from the server's GitHub credentials. Checkpoints and pushes execute
|
||||
in that same workspace, pushing `fabro/run/<id>` after each checkpoint;
|
||||
no server snapshot repository or Git bundle is maintained. Final publication
|
||||
retries the push and opens the configured pull request; a failure fails the
|
||||
run with `publish_failed` (`fabro-cli/src/commands/run/publish.rs`). Ordinary
|
||||
resume requires the retained workspace. A GitHub-backed fork fetches the
|
||||
source run branch inside its new sandbox. Runs on host workspaces (local
|
||||
folders, empty Local targets, dry runs) still commit checkpoints in that
|
||||
workspace, without pushing; Docker and Daytona runs with no GitHub target
|
||||
keep execution metadata without Git commits. Retry starts a fresh execution
|
||||
from the saved spec. `CloneRequest` still
|
||||
travels beside the sandbox spec so the run record names the origin and
|
||||
branch; the sandbox layer refuses a request that asks it to clone.
|
||||
Preflight and `fabro exec` initialize sandboxes with `CloneRequest::none()`,
|
||||
which creates an empty workspace.
|
||||
|
||||
### Release automation
|
||||
- `cargo dev release` — creates the next stable release tag. Use `cargo dev release --nightly` for a nightly prerelease. Use `--dry-run` to print planned commands without mutating git or running Cargo, `--skip-tests` only after running the release-mode smoke yourself, and `--release-date YYYY-MM-DD` or `FABRO_RELEASE_DATE` for deterministic version computation.
|
||||
|
|
@ -141,12 +154,25 @@ Fabro is an AI-powered workflow orchestration platform. Workflows are defined as
|
|||
### Docker sandbox provider
|
||||
- Docker is the default runtime sandbox provider from `defaults.toml`. The Fabro process must have a working Docker client environment (`DOCKER_HOST`, socket access, Docker Desktop behavior, TLS settings, groups/permissions, and any remote daemon policy are operator responsibilities).
|
||||
- The packaged compose service mounts `/var/run/docker.sock` so the server can create sibling run containers on the host daemon. This is host-root-equivalent under Docker's security model; only use it in the trusted, single-tenant deployment model described by the sandbox code/docs.
|
||||
- Fabro no longer clones a repository into a sandbox: the engine prepares
|
||||
every run's checkout. `CloneRequest` still travels beside the sandbox spec
|
||||
so the run record names the origin and branch; fabro validates it (a pin
|
||||
needs a branch, a non-GitHub origin needs `skip_clone`) and refuses a
|
||||
request that asks for a clone. Preflight and `fabro exec` initialize
|
||||
sandboxes with `CloneRequest::none()`, which creates an empty workspace.
|
||||
- A GitHub target's workspace is checked out by Fabro's hooks, not by the
|
||||
sandbox driver or Petri's `start` checkout: when a fresh run's scope is
|
||||
acquired, `fabro-petri`'s `RunWorkspaces::check_out_source` fetches the
|
||||
target's revision inside the sandbox with a read-only token the worker
|
||||
mints from the server's GitHub credentials. Checkpoints and pushes execute
|
||||
in that same workspace, pushing `fabro/run/<id>` after each checkpoint;
|
||||
no server snapshot repository or Git bundle is maintained. Final publication
|
||||
retries the push and opens the configured pull request; a failure fails the
|
||||
run with `publish_failed` (`fabro-cli/src/commands/run/publish.rs`). Ordinary
|
||||
resume requires the retained workspace. A GitHub-backed fork fetches the
|
||||
source run branch inside its new sandbox. Runs on host workspaces (local
|
||||
folders, empty Local targets, dry runs) still commit checkpoints in that
|
||||
workspace, without pushing; Docker and Daytona runs with no GitHub target
|
||||
keep execution metadata without Git commits. Retry starts a fresh execution
|
||||
from the saved spec. `CloneRequest` still
|
||||
travels beside the sandbox spec so the run record names the origin and
|
||||
branch; the sandbox layer refuses a request that asks it to clone.
|
||||
Preflight and `fabro exec` initialize sandboxes with `CloneRequest::none()`,
|
||||
which creates an empty workspace.
|
||||
|
||||
### Release automation
|
||||
- `cargo dev release` — creates the next stable release tag. Use `cargo dev release --nightly` for a nightly prerelease. Use `--dry-run` to print planned commands without mutating git or running Cargo, `--skip-tests` only after running the release-mode smoke yourself, and `--release-date YYYY-MM-DD` or `FABRO_RELEASE_DATE` for deterministic version computation.
|
||||
|
|
|
|||
|
|
@ -2437,7 +2437,10 @@ paths:
|
|||
Creates a new run from a checkpoint of a terminal source run and
|
||||
starts it, then archives the source run and records
|
||||
`run.superseded_by` on it. Returns 207 when the new run was created
|
||||
but the source archive step failed.
|
||||
but the source archive step failed. Requires a GitHub target with
|
||||
run-branch creation and pushes enabled. The checkpoint must leave
|
||||
work to execute: the final terminal checkpoint is refused; select
|
||||
an earlier checkpoint or retry the workflow from the beginning.
|
||||
parameters:
|
||||
- $ref: "#/components/parameters/RunId"
|
||||
requestBody:
|
||||
|
|
@ -2498,7 +2501,12 @@ paths:
|
|||
position and continues from there in a fresh workspace restored to
|
||||
the checkpoint's commit. The source run is left untouched. A
|
||||
checkpoint inside a parallel branch cannot be forked at; fork at the
|
||||
parallel stage instead.
|
||||
parallel stage instead. Requires a GitHub target with run-branch
|
||||
creation and pushes enabled; the checkpoint commit must be available
|
||||
on the source run's published branch. The final terminal checkpoint
|
||||
is refused because it leaves no work to acquire a workspace. For a
|
||||
completed run, select an earlier checkpoint explicitly or retry the
|
||||
workflow from the beginning.
|
||||
parameters:
|
||||
- $ref: "#/components/parameters/RunId"
|
||||
requestBody:
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@ title: "Checkpoints"
|
|||
description: "How Fabro uses Git to checkpoint and resume workflow runs"
|
||||
---
|
||||
|
||||
Fabro checkpoints every workflow run using Git plus the durable run store. After each node completes, Fabro commits file changes to the run branch and records execution state in the durable event stream so that interrupted runs can be resumed exactly where they left off. This happens automatically — no configuration required beyond running inside a Git repository.
|
||||
Fabro records every workflow run in the durable run store. With run branches enabled, it also commits file changes inside the run workspace after each node completes. GitHub targets push those commits to the run branch; local runs keep them in their workspace. A local-folder run works in a clone of the folder's committed `HEAD`, so uncommitted changes are not included, and a folder that is not a Git repository starts as an empty workspace.
|
||||
|
||||
## Code and execution history
|
||||
|
||||
|
|
@ -71,22 +71,15 @@ The checkpoint projection captures the execution state needed to resume a run:
|
|||
|
||||
The durable run store also keeps the current checkpoint so `resume`, `inspect`, and API reads do not need to rely on scratch files.
|
||||
|
||||
## Worktrees
|
||||
## Run workspaces
|
||||
|
||||
Fabro uses Git worktrees to isolate workflow runs from your working directory. When a local run starts in a Git repository:
|
||||
A GitHub target is checked out inside the run workspace. For Docker and Daytona, checkout, checkpoint commits, diffs, and pushes execute inside the sandbox. The host provider performs those operations in its local run workspace. Fabro keeps no second checkout, bare snapshot repository, or Git checkpoint bundle on the server for a remote sandbox.
|
||||
|
||||
1. Fabro records the current HEAD as the **base SHA**
|
||||
2. Creates a new branch `fabro/run/{run_id}` at that SHA
|
||||
3. Adds a worktree at `{run_dir}/worktree` on that branch
|
||||
4. Changes into the worktree directory for the duration of the run
|
||||
When pushing is configured, Fabro pushes `fabro/run/{run_id}` after each checkpoint and again before successful completion. Execution records and artifact payloads remain in the run store and CAS.
|
||||
|
||||
This means your original working directory stays untouched while the agent makes changes in the worktree. When the run completes, Fabro removes the worktree and restores your original directory.
|
||||
Ordinary resume requires the original workspace to survive. Git-backed workspaces reset to their recorded checkpoint; non-Git workspaces continue with their surviving files. Losing the workspace does not trigger automatic reconstruction or sandbox replacement.
|
||||
|
||||
<Note>
|
||||
If the working directory has uncommitted changes, the worktree starts from committed `HEAD` and those uncommitted changes are not included. Fabro logs a warning so you can commit, stash, or run explicitly in place when that is what you want.
|
||||
</Note>
|
||||
|
||||
For Docker and Daytona sandboxes, the repository is cloned into the sandbox and checkpoint Git operations run there. The run branch is pushed to origin from the sandbox after each checkpoint when pushing is configured.
|
||||
A GitHub-backed fork or rewind fetches the source run's published branch inside a fresh workspace and checks out the selected checkpoint. This requires run-branch pushes to be enabled and the commit to be available on origin. Select a checkpoint with remaining work: the final terminal checkpoint is refused because no stage remains to acquire a workspace. A retry starts the workflow from the beginning using its saved specification, without requiring the previous workspace or any Git checkpoint.
|
||||
|
||||
## Resuming a run
|
||||
|
||||
|
|
@ -120,10 +113,10 @@ After a node completes, Fabro:
|
|||
|
||||
1. Stores offloaded context payloads in CAS.
|
||||
2. Creates a code checkpoint commit when Git checkpointing is enabled.
|
||||
3. Pushes the run branch when configured and collects the code diff.
|
||||
3. Collects the code diff.
|
||||
4. Emits a checkpoint event with execution state and the code commit SHA. The run store persists this event and updates the projection.
|
||||
|
||||
A checkpoint commit failure stops execution. Intermediate push and diff failures emit warning notices. A required final publish failure marks the run as failed.
|
||||
A checkpoint commit failure stops execution. A stage diff failure emits a warning notice. An intermediate push failure warns and is retried at later checkpoints and final publication. Failure to prepare required final publication, push the final commit, or open the pull request marks an otherwise successful run as failed with `publish_failed`.
|
||||
|
||||
## Inspecting run history
|
||||
|
||||
|
|
@ -149,14 +142,12 @@ fabro dump 01JKXYZ --output ./run-dump
|
|||
|
||||
## When checkpointing is active
|
||||
|
||||
Git checkpointing activates automatically when:
|
||||
Git checkpointing activates automatically when run branches are enabled and either:
|
||||
|
||||
- The run uses a Git repository and checkpointing has not been explicitly disabled
|
||||
- Local runs can create a Git worktree under the run scratch directory
|
||||
- Docker or Daytona can clone the configured GitHub origin into the sandbox
|
||||
- The run targets a GitHub repository with cloning enabled, or
|
||||
- The run's workspace is on the host: a Local environment, or any `--dry-run`
|
||||
|
||||
It is skipped when:
|
||||
|
||||
- The working directory is not a Git repository
|
||||
- The run uses `--dry-run`
|
||||
- The run is explicitly started in place with checkpointing disabled
|
||||
- Run branches are disabled
|
||||
- A Docker or Daytona run has no GitHub target, or its cloning is disabled
|
||||
|
|
|
|||
|
|
@ -28,8 +28,8 @@ The rest of this page describes the `app` strategy, which is required for browse
|
|||
| Feature | How it's used |
|
||||
|---|---|
|
||||
| **OAuth login** | Users sign in to the web UI with their GitHub account |
|
||||
| **Private repo cloning** | Daytona and Docker sandboxes clone private repositories using short-lived Installation Access Tokens |
|
||||
| **Checkpoint pushing** | After each workflow stage, Fabro pushes the run branch back to origin from inside the sandbox |
|
||||
| **Private repo cloning** | Daytona and Docker sandboxes fetch private repositories using short-lived, read-only Installation Access Tokens |
|
||||
| **Run branch pushing** | Fabro pushes from the sandbox after each checkpoint and again before a successful run finishes |
|
||||
| **Auto-PR** | When `[run.pull_request] enabled = true` in the [run config](/execution/run-configuration#runpull_request), Fabro opens a PR from the agent's working branch after a successful run |
|
||||
| **Auto-merge** | When `[run.pull_request] auto_merge = true`, Fabro enables GitHub's auto-merge on created PRs so they merge automatically once required checks pass |
|
||||
| **Sandbox GITHUB_TOKEN** | When `[run.integrations.github.permissions]` are declared at any layer (workflow, project, or user settings), Fabro mints a scoped Installation Access Token and injects it as `GITHUB_TOKEN` in the sandbox |
|
||||
|
|
@ -216,16 +216,15 @@ An empty `allowed_usernames` list rejects all users.
|
|||
|
||||
### Repository cloning in sandboxes
|
||||
|
||||
When a workflow runs in a remote sandbox (Daytona or Docker), Fabro clones the current repository into the sandbox using the GitHub App:
|
||||
When a run targets a GitHub repository, its workspace is checked out inside the sandbox (Docker or Daytona) before the first stage runs:
|
||||
|
||||
1. Fabro detects the local repository's `origin` remote URL and current branch
|
||||
2. SSH URLs (e.g. `git@github.com:owner/repo.git`) are converted to HTTPS
|
||||
3. Fabro signs a short-lived JWT using the App ID and private key (RS256, 10-minute validity)
|
||||
4. Using the JWT, Fabro looks up the GitHub App installation for the repository (`GET /repos/\{owner\}/\{repo\}/installation`)
|
||||
5. Fabro requests a scoped Installation Access Token with `contents: write` permission on the specific repository
|
||||
6. The sandbox clones via HTTPS using `x-access-token` as the username and the token as the password
|
||||
1. When the run's worker starts, it signs a short-lived JWT using the App ID and private key (RS256, 10-minute validity)
|
||||
2. Using the JWT, Fabro looks up the GitHub App installation for the repository (`GET /repos/\{owner\}/\{repo\}/installation`)
|
||||
3. Fabro requests a scoped Installation Access Token with `contents: read` permission on the specific repository
|
||||
4. Inside the sandbox, Fabro fetches the selected revision from `https://github.com/<owner>/<repo>` at the run's `[run.clone] depth`, presenting the token as an HTTP header on that one command, and checks out the working branch
|
||||
5. The workspace's `origin` is the plain HTTPS URL: the token is never written into the repository, its configuration, or its remote
|
||||
|
||||
For public repositories, the clone works without credentials. The token is still generated because it's needed for pushing checkpoints.
|
||||
The files belong to the user the sandbox runs commands as. For public repositories the fetch works without credentials when none are configured.
|
||||
|
||||
#### Git targets for run intents
|
||||
|
||||
|
|
@ -319,11 +318,11 @@ Every commit a run creates is authored and committed by the run's GitHub credent
|
|||
|
||||
### Checkpoint pushing
|
||||
|
||||
After each workflow stage, Fabro [checkpoints](/execution/checkpoints) by pushing the run branch to origin. Before a successful run becomes terminal, the publish stage pushes the final commit again and treats failure as a run failure. Inside remote sandboxes, the git remote URL is configured with the Installation Access Token for authenticated pushing.
|
||||
After each workflow stage, Fabro [checkpoints](/execution/checkpoints) the workspace on the run branch, `fabro/run/<run-id>`, and pushes it directly from that workspace. Fabro keeps no second Git repository or checkpoint bundles on the server. In App mode, the worker pushes with an Installation Access Token with `contents: write`, reusing one token until it nears expiry and then minting the next, and passes it only to the sandbox Git command. The token is not saved in the repository's remote URL or configuration. `[run.run_branch] push = false` keeps the branch only in the run workspace.
|
||||
|
||||
When pull request creation is enabled, Fabro then checks that GitHub reports the run branch at the exact final commit before opening the PR. A failed final push, branch check, or PR creation marks the run as failed with `publish_failed`; the terminal run event is emitted only after this step finishes.
|
||||
An intermediate push failure logs a warning; later checkpoints and final publication retry the push. After a successful run's last stage and before the run finishes, Fabro awaits one final push.
|
||||
|
||||
For long-running workflows, Fabro refreshes the token before each push since Installation Access Tokens are short-lived (typically 1 hour).
|
||||
When pull request creation is enabled and the run changed files, Fabro then checks that GitHub reports the run branch at the exact final commit and opens the pull request. A failed final push, publication preparation, branch check, or PR creation marks the run as failed with `publish_failed`; the terminal run event is emitted only after this step finishes.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
|
|
|
|||
|
|
@ -113,6 +113,7 @@ walkdir.workspace = true
|
|||
rmcp = { workspace = true, features = ["client", "transport-child-process"] }
|
||||
fabro-build-support = { path = "../../foundation/build-support" }
|
||||
fabro-server = { path = "../fabro-server", features = ["test-support"] }
|
||||
fabro-github = { path = "../../components/fabro-github", features = ["test-support"] }
|
||||
fabro-petri = { path = "../../components/fabro-petri", features = ["test-support"] }
|
||||
fabro-workflow = { path = "../../components/fabro-workflow", features = ["test-support"] }
|
||||
fabro-types = { path = "../../foundation/fabro-types", features = ["clap", "test-support"] }
|
||||
|
|
|
|||
|
|
@ -23,6 +23,7 @@ pub(crate) mod overrides;
|
|||
pub(crate) mod petri_stream;
|
||||
mod petri_worker;
|
||||
pub(crate) mod preview;
|
||||
mod publish;
|
||||
mod remote_workflow;
|
||||
mod resolution;
|
||||
pub(crate) mod resume;
|
||||
|
|
|
|||
|
|
@ -71,17 +71,19 @@ use fabro_auth::VaultCredentialSource;
|
|||
use fabro_client::{Client, ServerTarget};
|
||||
use fabro_interview::{ControlInterviewer, WorkerControlMessage, WorkerControlOutcome};
|
||||
use fabro_llm::credentials::{CredentialProvider, readiness};
|
||||
use fabro_llm::lithos_catalog::Catalog;
|
||||
use fabro_petri::artifacts::ClientArtifactWriter;
|
||||
use fabro_petri::blobs::ClientBlobs;
|
||||
use fabro_petri::controls::{RunControls, SteerError};
|
||||
use fabro_petri::engine::{self, Conclusion, Execution, RunRequest};
|
||||
use fabro_petri::hooks::HooksSpec;
|
||||
use fabro_petri::hooks::{HooksSpec, RunPublisher};
|
||||
use fabro_petri::interview::{Approval, FabroInterviewer};
|
||||
use fabro_petri::petri::OwnerId;
|
||||
use fabro_petri::platform_records::{HttpPlatformRecords, PlatformRecords};
|
||||
use fabro_petri::providers::{DaytonaCredentials, SandboxProviderConfig};
|
||||
use fabro_petri::runtime::{self, RuntimeSpec};
|
||||
use fabro_petri::secrets::VaultSecrets;
|
||||
use fabro_petri::source::RunSource;
|
||||
use fabro_petri::{HttpRunStore, admission};
|
||||
use fabro_static::EnvVars;
|
||||
use fabro_store::RunProjection;
|
||||
|
|
@ -98,6 +100,7 @@ use tokio::task::JoinHandle;
|
|||
use tokio_util::sync::CancellationToken;
|
||||
use tracing::{info, warn};
|
||||
|
||||
use super::publish;
|
||||
use super::runner::{self, WorkerTitlePhase};
|
||||
use crate::args::RunWorkerMode;
|
||||
use crate::command_context;
|
||||
|
|
@ -166,7 +169,10 @@ pub(super) async fn execute(worker: PetriWorker<'_>) -> Result<()> {
|
|||
let vault = runner::load_worker_vault(worker.storage_dir).await?;
|
||||
let secrets = VaultSecrets::from_vault(&*vault.read().await);
|
||||
let run_tools = run_tool_services(&worker);
|
||||
let catalog =
|
||||
command_context::load_cli_catalog().context("failed to build worker LLM catalog")?;
|
||||
let runtime = runtime_spec(
|
||||
catalog.clone(),
|
||||
&vault,
|
||||
&worker.run_state,
|
||||
worker.fabro_home.clone(),
|
||||
|
|
@ -199,11 +205,41 @@ pub(super) async fn execute(worker: PetriWorker<'_>) -> Result<()> {
|
|||
}
|
||||
runner::set_worker_title(&run_id, WorkerTitlePhase::Running);
|
||||
|
||||
// A GitHub target is fetched into its sandbox with a read-only token and
|
||||
// published with a push token, each from a token source over the
|
||||
// server's credentials that the run keeps for its whole life.
|
||||
let github = match publish::github_credentials(&*vault.read().await) {
|
||||
Ok(credentials) => credentials,
|
||||
Err(err) => {
|
||||
warn!(run_id = %run_id, error = %err, "GitHub credentials are unavailable to the worker");
|
||||
None
|
||||
}
|
||||
};
|
||||
let mut source = RunSource::for_run(
|
||||
worker.run_state.spec.target.as_ref(),
|
||||
&worker.run_state.spec.settings.run,
|
||||
None,
|
||||
);
|
||||
if let Some(source) = &mut source {
|
||||
source.credentials = publish::source_credentials(&worker.run_state.spec, github.as_ref());
|
||||
}
|
||||
let publisher = publish::GitHubPublisher::for_run(
|
||||
run_id,
|
||||
&worker.run_state.spec,
|
||||
github,
|
||||
Arc::new(VaultCredentialSource::new(Arc::clone(&vault))),
|
||||
Arc::new(catalog),
|
||||
Arc::clone(&records),
|
||||
worker.client.clone_for_reuse(),
|
||||
)
|
||||
.map(|publisher| Arc::new(publisher) as Arc<dyn RunPublisher>);
|
||||
let hooks = HooksSpec::for_run(
|
||||
Arc::clone(&records),
|
||||
&worker.run_state.spec.settings.run,
|
||||
Arc::new(ClientArtifactWriter::new(worker.client.clone_for_reuse())),
|
||||
)
|
||||
.with_source(source)
|
||||
.with_publisher(publisher)
|
||||
.with_test_gates(test_checkpoint_gates());
|
||||
let request = RunRequest {
|
||||
run_id: run_id.to_string(),
|
||||
|
|
@ -607,13 +643,12 @@ fn run_tool_services(worker: &PetriWorker<'_>) -> Option<FabroRunToolServices> {
|
|||
/// the providers whose credentials resolve, the run's mode, the Fabro
|
||||
/// home the server named, and the run tools when the run has them.
|
||||
async fn runtime_spec(
|
||||
catalog: Catalog,
|
||||
vault: &Arc<AsyncRwLock<Vault>>,
|
||||
run_state: &RunProjection,
|
||||
fabro_home: Option<PathBuf>,
|
||||
run_tools: Option<FabroRunToolServices>,
|
||||
) -> Result<RuntimeSpec> {
|
||||
let catalog =
|
||||
command_context::load_cli_catalog().context("failed to build worker LLM catalog")?;
|
||||
let credentials: Arc<dyn CredentialProvider> =
|
||||
Arc::new(VaultCredentialSource::new(Arc::clone(vault)));
|
||||
let ready = readiness(catalog.enabled_providers(), credentials.as_ref()).await;
|
||||
|
|
|
|||
447
lib/apps/fabro-cli/src/commands/run/publish.rs
Normal file
447
lib/apps/fabro-cli/src/commands/run/publish.rs
Normal file
|
|
@ -0,0 +1,447 @@
|
|||
//! A GitHub-target run's repository work in its worker: the read credential
|
||||
//! its workspaces are fetched with, and its publication when it succeeds.
|
||||
//!
|
||||
//! The worker resolves the server's GitHub credentials itself, as the
|
||||
//! legacy worker did: the strategy and App id from the server settings the
|
||||
//! server named (`FABRO_CONFIG`), the App key the server hands the worker,
|
||||
//! or `GITHUB_TOKEN` from the worker's vault snapshot. From them it keeps two
|
||||
//! cached token sources for the run: a read-only one for fetches inside the
|
||||
//! sandbox and a `contents: write` one for pushes. Each fetch and push asks
|
||||
//! its source, which reuses one installation token until it nears expiry and
|
||||
//! then mints the next. Reuse matters: GitHub can reject a token minted
|
||||
//! moments earlier, before it has replicated, so minting per push turns a
|
||||
//! long run's pushes into repeated failures. Re-minting near expiry keeps a
|
||||
//! run of any length on a live token.
|
||||
//!
|
||||
//! Publication runs in Fabro's `run_finished` hook, after the last stage and
|
||||
//! before the run's terminal record, as the legacy publish step did: the
|
||||
//! final checkpoint is pushed from inside the sandbox to the run
|
||||
//! branch on GitHub, and, when the run changed files and its settings ask
|
||||
//! for one, a pull request is opened and recorded. A failure fails the run
|
||||
//! with `publish_failed`.
|
||||
|
||||
use std::sync::Arc;
|
||||
use std::time::Duration;
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use base64::Engine as _;
|
||||
use base64::engine::general_purpose::STANDARD as BASE64_STANDARD;
|
||||
use fabro_config::ServerSettingsBuilder;
|
||||
use fabro_github::token_source::{InstallationTokenSource, SecretString};
|
||||
use fabro_github::{GitHubContext, GitHubCredentials};
|
||||
use fabro_llm::credentials::{self, CredentialProvider};
|
||||
use fabro_llm::lithos_catalog::Catalog;
|
||||
use fabro_petri::checkpoint::Site;
|
||||
use fabro_petri::hooks::{Publication, RunPublisher};
|
||||
use fabro_petri::platform_records::PlatformRecords;
|
||||
use fabro_petri::source::{SourceCredential, SourceCredentials};
|
||||
use fabro_static::EnvVars;
|
||||
use fabro_store::platform_records::{PlatformRecord, PullRequestCreatedRecord};
|
||||
use fabro_types::settings::run::{PullRequestSettings, RunMode};
|
||||
use fabro_types::settings::server::GithubIntegrationStrategy;
|
||||
use fabro_types::{GitHubRepositorySlug, RunId, RunSpec, RunTarget};
|
||||
use fabro_vault::Vault;
|
||||
use fabro_workflow::pull_request::{self, AutoMergeOptions, OpenPullRequestRequest};
|
||||
use tokio::time;
|
||||
use tracing::warn;
|
||||
|
||||
/// How long one push to the repository may take.
|
||||
const PUSH_TIMEOUT: Duration = Duration::from_mins(5);
|
||||
/// Attempts at the push, all with the one token resolved for it: a freshly
|
||||
/// minted token can take a moment to reach every GitHub replica.
|
||||
const PUSH_ATTEMPTS: u32 = 3;
|
||||
const PUSH_RETRY_DELAY: Duration = Duration::from_secs(2);
|
||||
|
||||
/// The server's GitHub credentials as the worker reaches them, or `None`
|
||||
/// when none are configured.
|
||||
pub(super) fn github_credentials(vault: &Vault) -> Result<Option<GitHubCredentials>> {
|
||||
let settings = ServerSettingsBuilder::load_default().context("loading the server settings")?;
|
||||
let github = &settings.server.integrations.github;
|
||||
match github.strategy {
|
||||
GithubIntegrationStrategy::App => {
|
||||
GitHubCredentials::from_env_with_slug(github.app_id.as_deref(), github.slug.as_deref())
|
||||
.map_err(anyhow::Error::msg)
|
||||
}
|
||||
GithubIntegrationStrategy::Token => {
|
||||
let Some(token) = vault
|
||||
.get(EnvVars::GITHUB_TOKEN)
|
||||
.map(str::trim)
|
||||
.filter(|token| !token.is_empty())
|
||||
else {
|
||||
return Ok(None);
|
||||
};
|
||||
fabro_github::validate_static_github_token(token)?;
|
||||
Ok(Some(GitHubCredentials::Pat(token.to_string())))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The run's GitHub repository, when its target names one.
|
||||
fn repository(spec: &RunSpec) -> Option<GitHubRepositorySlug> {
|
||||
let Some(RunTarget::Git(target)) = spec.target.as_ref() else {
|
||||
return None;
|
||||
};
|
||||
Some(target.clone().validate().ok()?.repository().clone())
|
||||
}
|
||||
|
||||
/// A cached token source for the run's repository with `permissions`, or
|
||||
/// `None` when the run has no GitHub target or no credentials resolve.
|
||||
fn token_source(
|
||||
spec: &RunSpec,
|
||||
credentials: Option<&GitHubCredentials>,
|
||||
permissions: serde_json::Value,
|
||||
) -> Option<Arc<InstallationTokenSource>> {
|
||||
let repository = repository(spec)?;
|
||||
let credentials = credentials?;
|
||||
match InstallationTokenSource::for_repository(
|
||||
credentials,
|
||||
repository.owner().to_string(),
|
||||
repository.repo().to_string(),
|
||||
permissions,
|
||||
) {
|
||||
Ok(source) => Some(source),
|
||||
Err(err) => {
|
||||
warn!(repository = %repository, error = %format!("{err:#}"), "no GitHub token source for the run's repository");
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The read-only credentials the run's workspaces are fetched with. `None`
|
||||
/// when the run has no GitHub target or no credentials resolve; a public
|
||||
/// repository is then fetched anonymously.
|
||||
pub(super) fn source_credentials(
|
||||
spec: &RunSpec,
|
||||
credentials: Option<&GitHubCredentials>,
|
||||
) -> Option<Arc<dyn SourceCredentials>> {
|
||||
let tokens = token_source(spec, credentials, serde_json::json!({ "contents": "read" }))?;
|
||||
Some(Arc::new(ReadCredentials(tokens)))
|
||||
}
|
||||
|
||||
/// Fetch credentials resolved from the run's read-only token source.
|
||||
struct ReadCredentials(Arc<InstallationTokenSource>);
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl SourceCredentials for ReadCredentials {
|
||||
async fn credential(&self) -> Option<SourceCredential> {
|
||||
match self.0.resolve().await {
|
||||
Ok(resolved) => basic(&resolved.token),
|
||||
Err(err) => {
|
||||
warn!(error = %format!("{err:#}"), "no read credential for the run's repository; it is fetched anonymously");
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The HTTP basic credential `git` presents for an installation token or
|
||||
/// personal access token.
|
||||
fn basic(token: &SecretString) -> Option<SourceCredential> {
|
||||
SourceCredential::from_encoded(
|
||||
BASE64_STANDARD.encode(format!("x-access-token:{}", token.expose())),
|
||||
)
|
||||
}
|
||||
|
||||
/// A successful GitHub-target run's publication: the run branch pushed, and
|
||||
/// the pull request its settings ask for opened.
|
||||
pub(super) struct GitHubPublisher {
|
||||
run_id: RunId,
|
||||
repository: GitHubRepositorySlug,
|
||||
/// The branch the target names: the pull request's base.
|
||||
base_branch: String,
|
||||
goal: String,
|
||||
/// The run's model, when its settings name one.
|
||||
model: Option<String>,
|
||||
credentials: Option<GitHubCredentials>,
|
||||
/// The run's `contents: write` token source; `None` without credentials.
|
||||
push_tokens: Option<Arc<InstallationTokenSource>>,
|
||||
pull_request: Option<PullRequestSettings>,
|
||||
llm_source: Arc<dyn CredentialProvider>,
|
||||
catalog: Arc<Catalog>,
|
||||
records: Arc<dyn PlatformRecords>,
|
||||
client: fabro_client::Client,
|
||||
}
|
||||
|
||||
impl GitHubPublisher {
|
||||
/// The publisher of a run whose target is a GitHub repository and whose
|
||||
/// run branch is pushed; `None` for a dry run or any other run.
|
||||
pub(super) fn for_run(
|
||||
run_id: RunId,
|
||||
spec: &RunSpec,
|
||||
credentials: Option<GitHubCredentials>,
|
||||
llm_source: Arc<dyn CredentialProvider>,
|
||||
catalog: Arc<Catalog>,
|
||||
records: Arc<dyn PlatformRecords>,
|
||||
client: fabro_client::Client,
|
||||
) -> Option<Self> {
|
||||
let settings = &spec.settings.run;
|
||||
if settings.execution.mode == RunMode::DryRun
|
||||
|| !settings.clone.enabled
|
||||
|| !settings.run_branch.enabled
|
||||
|| !settings.run_branch.push
|
||||
{
|
||||
return None;
|
||||
}
|
||||
let Some(RunTarget::Git(target)) = spec.target.as_ref() else {
|
||||
return None;
|
||||
};
|
||||
let repository = repository(spec)?;
|
||||
let push_tokens = token_source(
|
||||
spec,
|
||||
credentials.as_ref(),
|
||||
serde_json::json!({ "contents": "write" }),
|
||||
);
|
||||
Some(Self {
|
||||
run_id,
|
||||
repository,
|
||||
base_branch: target.branch.clone(),
|
||||
goal: spec.graph.goal.clone(),
|
||||
model: settings.model.name.clone(),
|
||||
credentials,
|
||||
push_tokens,
|
||||
pull_request: settings
|
||||
.pull_request
|
||||
.clone()
|
||||
.filter(|pull_request| pull_request.enabled),
|
||||
llm_source,
|
||||
catalog,
|
||||
records,
|
||||
client,
|
||||
})
|
||||
}
|
||||
|
||||
/// The model that writes the pull request: the run's, or the catalog's
|
||||
/// default among the providers whose credentials resolve.
|
||||
async fn model(&self) -> Option<String> {
|
||||
if let Some(model) = &self.model {
|
||||
return Some(model.clone());
|
||||
}
|
||||
let ready =
|
||||
credentials::readiness(self.catalog.enabled_providers(), self.llm_source.as_ref())
|
||||
.await;
|
||||
self.catalog
|
||||
.default_offering_for(&ready.ready)
|
||||
.map(|entry| entry.model.id().to_string())
|
||||
}
|
||||
|
||||
async fn open_pull_request(
|
||||
&self,
|
||||
context: GitHubContext<'_>,
|
||||
settings: &PullRequestSettings,
|
||||
publication: &Publication,
|
||||
) -> Result<(), String> {
|
||||
let model = self
|
||||
.model()
|
||||
.await
|
||||
.ok_or_else(|| "no LLM model is available to write the pull request".to_string())?;
|
||||
// The run so far, for the pull request's details: best effort.
|
||||
let run_state = self.client.get_run_state(&self.run_id).await.ok();
|
||||
let origin_url = self.repository.https_url();
|
||||
let created = pull_request::open_pull_request(OpenPullRequestRequest {
|
||||
github: context,
|
||||
origin_url: &origin_url,
|
||||
base_branch: &self.base_branch,
|
||||
head_branch: &publication.run_branch,
|
||||
expected_head_sha: &publication.head_sha,
|
||||
goal: &self.goal,
|
||||
diff: &publication.patch,
|
||||
model: &model,
|
||||
draft: settings.draft,
|
||||
auto_merge: settings.auto_merge.then_some(AutoMergeOptions {
|
||||
merge_strategy: settings.merge_strategy,
|
||||
}),
|
||||
llm_source: Arc::clone(&self.llm_source),
|
||||
catalog: Arc::clone(&self.catalog),
|
||||
conclusion: None,
|
||||
run_state: run_state.as_ref(),
|
||||
})
|
||||
.await
|
||||
.map_err(|err| format!("failed to create pull request: {err}"))?;
|
||||
let link = &created.link;
|
||||
let record = PlatformRecord::PullRequestCreated(PullRequestCreatedRecord {
|
||||
number: link.number,
|
||||
owner: link.owner.clone(),
|
||||
repo: link.repo.clone(),
|
||||
html_url: link.html_url(),
|
||||
head_sha: Some(publication.head_sha.clone()),
|
||||
draft: settings.draft,
|
||||
operation: None,
|
||||
});
|
||||
self.records
|
||||
.append(&self.run_id, &record, None)
|
||||
.await
|
||||
.map_err(|err| format!("the pull request was opened but not recorded: {err}"))?;
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl RunPublisher for GitHubPublisher {
|
||||
async fn push(&self, site: &Site, branch: &str, sha: &str) -> Result<(), String> {
|
||||
let tokens = self.push_tokens.as_ref().ok_or_else(|| {
|
||||
"pushing the run branch requires the server's GitHub credentials".to_string()
|
||||
})?;
|
||||
let resolved = tokens
|
||||
.resolve()
|
||||
.await
|
||||
.map_err(|err| format!("no push credential for {}: {err:#}", self.repository))?;
|
||||
push(&self.repository, &resolved.token, site, branch, sha).await
|
||||
}
|
||||
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String> {
|
||||
self.push(
|
||||
&publication.site,
|
||||
&publication.run_branch,
|
||||
&publication.head_sha,
|
||||
)
|
||||
.await?;
|
||||
let Some(settings) = &self.pull_request else {
|
||||
return Ok(());
|
||||
};
|
||||
if publication.patch.trim().is_empty() {
|
||||
return Ok(());
|
||||
}
|
||||
let credentials = self
|
||||
.credentials
|
||||
.as_ref()
|
||||
.ok_or_else(|| "GitHub credentials are unavailable".to_string())?;
|
||||
let base_url = fabro_github::github_api_base_url();
|
||||
let context = GitHubContext::new(credentials, &base_url);
|
||||
Box::pin(self.open_pull_request(context, settings, publication)).await
|
||||
}
|
||||
}
|
||||
|
||||
/// Push a checkpoint from inside its workspace to the run
|
||||
/// branch on GitHub, retrying a failure that may be a token still
|
||||
/// replicating. The token reaches `git` as an HTTP header for the
|
||||
/// repository alone and never appears in the error.
|
||||
async fn push(
|
||||
repository: &GitHubRepositorySlug,
|
||||
token: &SecretString,
|
||||
site: &Site,
|
||||
branch: &str,
|
||||
sha: &str,
|
||||
) -> Result<(), String> {
|
||||
let mut url = repository.https_url();
|
||||
url.push_str(".git");
|
||||
let credential = basic(token);
|
||||
let env = credential
|
||||
.as_ref()
|
||||
.map(|credential| credential.header_env(&url))
|
||||
.unwrap_or_default();
|
||||
let refspec = format!("{sha}:refs/heads/{branch}");
|
||||
let mut last = String::new();
|
||||
for attempt in 1..=PUSH_ATTEMPTS {
|
||||
match site.push(&url, &refspec, &env, PUSH_TIMEOUT).await {
|
||||
Ok(()) => return Ok(()),
|
||||
Err(error) => {
|
||||
last = error.to_string().replace(token.expose(), "***");
|
||||
if let Some(credential) = &credential {
|
||||
last = last.replace(credential.encoded(), "***");
|
||||
}
|
||||
}
|
||||
}
|
||||
warn!(
|
||||
attempt,
|
||||
branch,
|
||||
error = last,
|
||||
"pushing the run branch failed"
|
||||
);
|
||||
if attempt < PUSH_ATTEMPTS {
|
||||
time::sleep(PUSH_RETRY_DELAY).await;
|
||||
}
|
||||
}
|
||||
Err(format!(
|
||||
"the run branch {branch} could not be pushed to {repository}: {last}"
|
||||
))
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use std::sync::atomic::{AtomicUsize, Ordering};
|
||||
|
||||
use chrono::Utc;
|
||||
use fabro_github::InstallationToken;
|
||||
use fabro_github::test_support::{InstallationTokenMinter, installation_token_source};
|
||||
use fabro_llm::credentials::NoCredentials;
|
||||
use fabro_llm::test_support;
|
||||
use fabro_petri::test_support::MemoryPlatformRecords;
|
||||
use fabro_types::GitRunTarget;
|
||||
use fabro_types::test_support::test_run_spec;
|
||||
|
||||
use super::*;
|
||||
|
||||
fn spec() -> RunSpec {
|
||||
let mut spec = test_run_spec();
|
||||
spec.target = Some(RunTarget::Git(GitRunTarget {
|
||||
repo: "acme/widgets".to_string(),
|
||||
branch: "main".to_string(),
|
||||
tag: None,
|
||||
sha: None,
|
||||
}));
|
||||
spec.settings.run.run_branch.enabled = true;
|
||||
spec.settings.run.run_branch.push = true;
|
||||
spec
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn only_a_pushed_github_target_run_is_published() {
|
||||
let publishes = |spec: &RunSpec| {
|
||||
GitHubPublisher::for_run(
|
||||
RunId::new(),
|
||||
spec,
|
||||
None,
|
||||
Arc::new(NoCredentials),
|
||||
Arc::new(test_support::test_catalog()),
|
||||
Arc::new(MemoryPlatformRecords::new()),
|
||||
fabro_client::Client::new_no_proxy("http://127.0.0.1:9").unwrap(),
|
||||
)
|
||||
.is_some()
|
||||
};
|
||||
assert!(publishes(&spec()));
|
||||
|
||||
let mut dry = spec();
|
||||
dry.settings.run.execution.mode = RunMode::DryRun;
|
||||
assert!(!publishes(&dry));
|
||||
|
||||
let mut unpushed = spec();
|
||||
unpushed.settings.run.run_branch.push = false;
|
||||
assert!(!publishes(&unpushed));
|
||||
|
||||
let mut empty = spec();
|
||||
empty.target = Some(RunTarget::None {});
|
||||
assert!(!publishes(&empty));
|
||||
}
|
||||
|
||||
/// Mints `token-<n>` for its n-th mint, valid for an hour.
|
||||
#[derive(Default)]
|
||||
struct CountingMinter(AtomicUsize);
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl InstallationTokenMinter for CountingMinter {
|
||||
async fn mint(&self) -> anyhow::Result<InstallationToken> {
|
||||
let n = self.0.fetch_add(1, Ordering::SeqCst) + 1;
|
||||
Ok(InstallationToken {
|
||||
token: format!("token-{n}"),
|
||||
expires_at: Utc::now() + chrono::Duration::hours(1),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn fetches_reuse_one_cached_read_token() {
|
||||
let minter = Arc::new(CountingMinter::default());
|
||||
let credentials = ReadCredentials(installation_token_source(
|
||||
"acme/widgets",
|
||||
Arc::clone(&minter) as Arc<dyn InstallationTokenMinter>,
|
||||
));
|
||||
let first = credentials.credential().await.unwrap();
|
||||
let second = credentials.credential().await.unwrap();
|
||||
assert_eq!(first, second);
|
||||
assert_eq!(minter.0.load(Ordering::SeqCst), 1);
|
||||
assert_eq!(
|
||||
first.encoded(),
|
||||
BASE64_STANDARD.encode("x-access-token:token-1")
|
||||
);
|
||||
}
|
||||
}
|
||||
|
|
@ -1334,7 +1334,7 @@ fn three_stage_bundle(context: &fabro_test::TestContext, gate: &Path) -> PathBuf
|
|||
"digraph Stages {{\n graph [goal=\"Three stages\", default_max_retries=0]\n start \
|
||||
[shape=Mdiamond]\n exit [shape=Msquare]\n one [shape=parallelogram, script=\"echo \
|
||||
one > one.txt\"]\n two [shape=parallelogram, script=\"echo run >> two.log; {}; \
|
||||
echo two > two.txt\"]\n three [shape=parallelogram, script=\"test \\\"$(cat \
|
||||
echo two > two.txt\"]\n three [shape=parallelogram, goal_gate=true, script=\"test \\\"$(cat \
|
||||
one.txt)\\\" = one && test \\\"$(cat two.txt)\\\" = two && cp two.log \
|
||||
three.log\"]\n start -> one -> two -> three -> exit\n}}\n",
|
||||
wait_for(gate)
|
||||
|
|
@ -1397,9 +1397,10 @@ fn subjects(commits: &[(String, Option<CheckpointKey>)]) -> Vec<&str> {
|
|||
.collect()
|
||||
}
|
||||
|
||||
/// The commit subjects one run of the three-stage bundle produces.
|
||||
/// The commit subjects one run of the three-stage bundle produces; its goal
|
||||
/// gate on `three` adds the `goal_check` stage.
|
||||
fn three_stage_subjects(run_id: &str) -> Vec<String> {
|
||||
["start", "one", "two", "three", "exit"]
|
||||
["start", "one", "two", "three", "goal_check", "exit"]
|
||||
.iter()
|
||||
.map(|node| format!("fabro({run_id}): {node} (success)"))
|
||||
.collect()
|
||||
|
|
@ -1437,7 +1438,7 @@ async fn a_crash_after_a_durable_finish_keeps_its_one_commit() {
|
|||
server.stderr_text()
|
||||
);
|
||||
let checkpoints = server.checkpoints(&run_id).await;
|
||||
assert_eq!(checkpoints.len(), 5, "{checkpoints:?}");
|
||||
assert_eq!(checkpoints.len(), 6, "{checkpoints:?}");
|
||||
let keys: Vec<Option<CheckpointKey>> = checkpoints.iter().map(|(key, _)| Some(*key)).collect();
|
||||
let committed: Vec<Option<CheckpointKey>> = commits.iter().map(|(_, key)| *key).collect();
|
||||
assert_eq!(keys, committed);
|
||||
|
|
@ -1513,7 +1514,7 @@ async fn a_crash_before_the_record_reconciles_it_from_the_run_branch() {
|
|||
"the stage with a durable finish did not rerun"
|
||||
);
|
||||
let checkpoints = server.checkpoints(&run_id).await;
|
||||
assert_eq!(checkpoints.len(), 5, "{checkpoints:?}");
|
||||
assert_eq!(checkpoints.len(), 6, "{checkpoints:?}");
|
||||
let keys: Vec<Option<CheckpointKey>> = checkpoints.iter().map(|(key, _)| Some(*key)).collect();
|
||||
let committed: Vec<Option<CheckpointKey>> = commits.iter().map(|(_, key)| *key).collect();
|
||||
assert_eq!(
|
||||
|
|
@ -1523,10 +1524,10 @@ async fn a_crash_before_the_record_reconciles_it_from_the_run_branch() {
|
|||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A workspace deleted while the run is down is restored from the snapshot
|
||||
/// repository, and the next stage sees the checkpoint's files.
|
||||
/// A workspace deleted while the run is down is not reconstructed: the
|
||||
/// server keeps no copy of the repository, so the resumed run fails.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_deleted_workspace_is_restored_from_its_snapshot() {
|
||||
async fn a_deleted_workspace_fails_the_resumed_run() {
|
||||
let context = test_context!();
|
||||
let mut server = RunningServer::start().await;
|
||||
let gate = context.temp_dir.join("two.gate");
|
||||
|
|
@ -1542,17 +1543,11 @@ async fn a_deleted_workspace_is_restored_from_its_snapshot() {
|
|||
std::fs::remove_dir_all(&path).expect("the workspace is deleted");
|
||||
|
||||
server.launch().await;
|
||||
wait_until_gate_is_polled(&gate);
|
||||
std::fs::write(&gate, "go").expect("the gate opens");
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let (restored, commits) = server.workspace_commits(&run_id);
|
||||
assert_eq!(restored, path);
|
||||
assert_eq!(subjects(&commits), three_stage_subjects(&run_id));
|
||||
assert_eq!(
|
||||
std::fs::read_to_string(restored.join("one.txt")).expect("one.txt was restored"),
|
||||
"one\n"
|
||||
wait_for_status(&server, &run_id, &["failed", "succeeded"]).await,
|
||||
"failed"
|
||||
);
|
||||
assert!(!path.exists(), "the deleted workspace is not recreated");
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -1,9 +1,6 @@
|
|||
//! Petri runs on a Docker environment through a real server and its
|
||||
//! worker: the workspace lives inside the run's container, every stage's
|
||||
//! checkpoint is committed there and published to the snapshot repository
|
||||
//! on the host, and a restart brings the container's workspace back to the
|
||||
//! snapshot its durable state names, in the retained container or in a
|
||||
//! fresh one when the old one is gone.
|
||||
//! worker: a non-Git workspace survives a worker restart in its retained
|
||||
//! container. A deleted container fails resume; Fabro does not reconstruct it.
|
||||
//!
|
||||
//! An Ask Fabro session on a finished Docker run attaches to the container
|
||||
//! Petri created and reads a file the workflow wrote there, its model the
|
||||
|
|
@ -19,10 +16,9 @@
|
|||
)]
|
||||
|
||||
use std::env;
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::path::PathBuf;
|
||||
use std::process::{Command, Stdio};
|
||||
|
||||
use fabro_petri::checkpoint::CheckpointKey;
|
||||
use fabro_static::EnvVars;
|
||||
use fabro_test::{
|
||||
TwinScenario, TwinScenarios, TwinToolCall, expect_reqwest_json, test_context, twin_openai,
|
||||
|
|
@ -129,62 +125,6 @@ fn cleanup(run_id: &str) {
|
|||
}
|
||||
}
|
||||
|
||||
/// The snapshot repository of the run's one workspace, on the host.
|
||||
fn snapshot_repository(server: &RunningServer, run_id: &str) -> PathBuf {
|
||||
let snapshots = server.petri_run_dir(run_id).join("snapshots");
|
||||
let mut repositories: Vec<PathBuf> = std::fs::read_dir(&snapshots)
|
||||
.expect("the snapshots directory lists")
|
||||
.map(|entry| entry.expect("an entry reads").path())
|
||||
.filter(|path| path.extension().is_some_and(|extension| extension == "git"))
|
||||
.collect();
|
||||
assert_eq!(repositories.len(), 1, "one workspace: {repositories:?}");
|
||||
repositories.remove(0)
|
||||
}
|
||||
|
||||
/// `git` in a repository on the host, its stdout.
|
||||
fn git(repository: &Path, args: &[&str]) -> String {
|
||||
let output = Command::new("git")
|
||||
.args(args)
|
||||
.current_dir(repository)
|
||||
.output()
|
||||
.expect("git runs");
|
||||
assert!(
|
||||
output.status.success(),
|
||||
"git {args:?} failed: {}",
|
||||
String::from_utf8_lossy(&output.stderr)
|
||||
);
|
||||
String::from_utf8_lossy(&output.stdout).trim().to_owned()
|
||||
}
|
||||
|
||||
/// The commits the snapshot repository holds, oldest first, as
|
||||
/// `(sha, subject, key)`.
|
||||
fn snapshot_commits(repository: &Path) -> Vec<(String, String, Option<CheckpointKey>)> {
|
||||
let log = git(repository, &[
|
||||
"log",
|
||||
"--topo-order",
|
||||
"--reverse",
|
||||
"--all",
|
||||
"--format=%H%x00%s%x00%B%x1e",
|
||||
]);
|
||||
log.split('\u{1e}')
|
||||
.filter(|entry| !entry.trim().is_empty())
|
||||
.map(|entry| {
|
||||
let mut parts = entry.trim_start().splitn(3, '\0');
|
||||
let sha = parts.next().unwrap_or_default().to_string();
|
||||
let subject = parts.next().unwrap_or_default().to_string();
|
||||
let body = parts.next().unwrap_or_default();
|
||||
(sha, subject, CheckpointKey::from_message(body))
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
fn subjects(commits: &[(String, String, Option<CheckpointKey>)]) -> Vec<&str> {
|
||||
commits
|
||||
.iter()
|
||||
.map(|(_, subject, _)| subject.as_str())
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Two command stages: `one` writes a file; `two` checks it is the one
|
||||
/// `one` wrote, that nothing else is in the workspace beside the
|
||||
/// repository's own files, and writes another.
|
||||
|
|
@ -198,62 +138,9 @@ fn two_stage_bundle(context: &fabro_test::TestContext) -> PathBuf {
|
|||
)
|
||||
}
|
||||
|
||||
/// The commit subjects one run of the two-stage bundle produces.
|
||||
fn two_stage_subjects(run_id: &str) -> Vec<String> {
|
||||
["start", "one", "two", "exit"]
|
||||
.iter()
|
||||
.map(|node| format!("fabro({run_id}): {node} (success)"))
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// The checkpoint records and the published snapshots name the same
|
||||
/// commits, and the two stages' trees hold their files.
|
||||
fn assert_snapshots_complete(server: &RunningServer, run_id: &str, repository: &Path) {
|
||||
let commits = snapshot_commits(repository);
|
||||
assert_eq!(subjects(&commits), two_stage_subjects(run_id));
|
||||
let (one, _, _) = &commits[1];
|
||||
let (two, _, _) = &commits[2];
|
||||
assert_eq!(git(repository, &["show", &format!("{one}:one.txt")]), "one");
|
||||
assert_eq!(git(repository, &["show", &format!("{two}:two.txt")]), "two");
|
||||
let refs = git(repository, &[
|
||||
"for-each-ref",
|
||||
"--format=%(objectname)",
|
||||
"refs/checkpoints/",
|
||||
]);
|
||||
let mut published: Vec<&str> = refs.lines().collect();
|
||||
published.sort_unstable();
|
||||
let checkpoints = futures_lite_block_on(server.checkpoints(run_id));
|
||||
let mut recorded: Vec<&str> = checkpoints.iter().map(|(_, sha)| sha.as_str()).collect();
|
||||
recorded.sort_unstable();
|
||||
assert_eq!(
|
||||
published, recorded,
|
||||
"every record names a published snapshot"
|
||||
);
|
||||
assert_eq!(checkpoints.len(), 4, "{checkpoints:?}");
|
||||
}
|
||||
|
||||
/// Wait on a future from a synchronous helper inside a multi-thread test.
|
||||
fn futures_lite_block_on<T>(future: impl std::future::Future<Output = T>) -> T {
|
||||
tokio::task::block_in_place(|| tokio::runtime::Handle::current().block_on(future))
|
||||
}
|
||||
|
||||
/// What the worker's log says it did to the sandbox workspace at resume.
|
||||
fn restore_actions(server: &RunningServer, run_id: &str) -> Vec<String> {
|
||||
let log = std::fs::read_to_string(server.worker_log(run_id)).unwrap_or_default();
|
||||
log.lines()
|
||||
.filter(|line| line.contains("workspace brought to its durable snapshot"))
|
||||
.filter_map(|line| {
|
||||
line.split_whitespace()
|
||||
.find_map(|word| word.strip_prefix("action=").map(str::to_owned))
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// A run on Docker: every stage is committed inside the container, each
|
||||
/// checkpoint is published to the snapshot repository on the host, and
|
||||
/// nothing of the workspace is on the host.
|
||||
/// A non-Git Docker run keeps files only inside its container.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_docker_run_publishes_every_stages_checkpoint_from_the_container() {
|
||||
async fn a_non_git_docker_run_keeps_its_workspace_in_the_container() {
|
||||
if !fabro_test::docker_available() {
|
||||
return;
|
||||
}
|
||||
|
|
@ -265,8 +152,8 @@ async fn a_docker_run_publishes_every_stages_checkpoint_from_the_container() {
|
|||
]);
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let repository = snapshot_repository(&server, &run_id);
|
||||
assert_snapshots_complete(&server, &run_id, &repository);
|
||||
assert!(server.checkpoints(&run_id).await.is_empty());
|
||||
assert!(!server.petri_run_dir(&run_id).join("snapshots").exists());
|
||||
assert!(
|
||||
!server.petri_run_dir(&run_id).join("scopes").exists(),
|
||||
"no workspace is on the host"
|
||||
|
|
@ -279,12 +166,9 @@ async fn a_docker_run_publishes_every_stages_checkpoint_from_the_container() {
|
|||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A worker killed after the first stage's durable finish, with the
|
||||
/// container's workspace changed behind Petri's back: the restart resumes
|
||||
/// on the retained container, the workspace is reset to the snapshot, and
|
||||
/// the second stage sees the first stage's files and nothing else.
|
||||
/// A non-Git run resumes in the retained container with its existing files.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_retained_container_whose_workspace_drifted_is_reset_on_restart() {
|
||||
async fn a_non_git_run_resumes_in_its_retained_container() {
|
||||
if !fabro_test::docker_available() {
|
||||
return;
|
||||
}
|
||||
|
|
@ -300,10 +184,7 @@ async fn a_retained_container_whose_workspace_drifted_is_reset_on_restart() {
|
|||
let worker = wait_for_worker(&run_id);
|
||||
server.wait_until_held(&run_id, "record", "one");
|
||||
let container = container_of(&run_id).expect("the run's container exists");
|
||||
docker_exec(
|
||||
&container,
|
||||
"echo junk > one.txt && echo stray > stray.txt && git status --porcelain",
|
||||
);
|
||||
docker_exec(&container, "echo marker > marker.txt && test ! -d .git");
|
||||
crash(&mut server, worker, None);
|
||||
|
||||
server.release("record", "one");
|
||||
|
|
@ -317,19 +198,17 @@ async fn a_retained_container_whose_workspace_drifted_is_reset_on_restart() {
|
|||
Some(container.as_str()),
|
||||
"the run continued in its retained container"
|
||||
);
|
||||
assert_eq!(restore_actions(&server, &run_id), vec!["Reset".to_string()]);
|
||||
let repository = snapshot_repository(&server, &run_id);
|
||||
assert_snapshots_complete(&server, &run_id, &repository);
|
||||
|
||||
assert!(server.checkpoints(&run_id).await.is_empty());
|
||||
assert!(!server.petri_run_dir(&run_id).join("snapshots").exists());
|
||||
cleanup(&run_id);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A worker killed after the first stage's durable finish, with the
|
||||
/// container removed while the run is down: the restart gets a fresh
|
||||
/// container, the workspace is restored into it from the snapshot
|
||||
/// repository, and the second stage sees the first stage's files.
|
||||
/// A deleted container fails resume instead of silently creating an empty
|
||||
/// replacement workspace.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_lost_container_is_replaced_and_its_workspace_restored_from_the_snapshot() {
|
||||
async fn a_lost_container_is_not_replaced_on_resume() {
|
||||
if !fabro_test::docker_available() {
|
||||
return;
|
||||
}
|
||||
|
|
@ -351,15 +230,16 @@ async fn a_lost_container_is_replaced_and_its_workspace_restored_from_the_snapsh
|
|||
|
||||
server.release("record", "one");
|
||||
server.launch().await;
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let fresh = container_of(&run_id).expect("a fresh container was created");
|
||||
assert_ne!(fresh, container);
|
||||
assert_eq!(restore_actions(&server, &run_id), vec![
|
||||
"Restored".to_string()
|
||||
]);
|
||||
let repository = snapshot_repository(&server, &run_id);
|
||||
assert_snapshots_complete(&server, &run_id, &repository);
|
||||
assert_eq!(
|
||||
wait_for_status(&server, &run_id, &["failed", "succeeded"]).await,
|
||||
"failed"
|
||||
);
|
||||
assert_eq!(
|
||||
container_of(&run_id),
|
||||
None,
|
||||
"resume must not replace the lost container"
|
||||
);
|
||||
assert!(!server.petri_run_dir(&run_id).join("snapshots").exists());
|
||||
cleanup(&run_id);
|
||||
server.shutdown();
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1,13 +1,11 @@
|
|||
//! Fork, rewind, retry and the timeline over Petri runs, through a real
|
||||
//! server and its worker subprocess.
|
||||
//!
|
||||
//! The harness is `petri.rs`'s: a foreground server on disk storage, a run
|
||||
//! created and started with `fabro run --detach`, executed by the worker
|
||||
//! the server launches over the HTTP run store. Every checkpoint of a run
|
||||
//! is a commit on its run branch and a `checkpoint` platform record at its
|
||||
//! Petri position; the timeline lists them, and a fork seeds a new run from
|
||||
//! the source's records up to one of them, whose worker restores the
|
||||
//! checkpoint's files into a fresh workspace and continues from there.
|
||||
//! Local-folder runs are checkpointed in their workspace and can retry from
|
||||
//! the start, but refuse fork and rewind: only a published run branch can
|
||||
//! seed a new workspace.
|
||||
//! Git-backed forks are covered by fabro-petri's sandbox/remote integration
|
||||
//! test.
|
||||
|
||||
#![expect(
|
||||
clippy::disallowed_methods,
|
||||
|
|
@ -100,7 +98,7 @@ async fn timeline(server: &RunningServer, run_id: &str) -> serde_json::Value {
|
|||
}
|
||||
|
||||
/// The timeline's entries as `(node, execution, firing, attempt, sha)`.
|
||||
fn entries(timeline: &serde_json::Value) -> Vec<(String, u64, u64, u64, String)> {
|
||||
fn entries(timeline: &serde_json::Value) -> Vec<(String, u64, u64, u64, Option<String>)> {
|
||||
timeline["entries"]
|
||||
.as_array()
|
||||
.expect("the timeline has entries")
|
||||
|
|
@ -111,10 +109,7 @@ fn entries(timeline: &serde_json::Value) -> Vec<(String, u64, u64, u64, String)>
|
|||
entry["execution"].as_u64().expect("an execution"),
|
||||
entry["firing"].as_u64().expect("a firing"),
|
||||
entry["attempt"].as_u64().expect("an attempt"),
|
||||
entry["run_commit_sha"]
|
||||
.as_str()
|
||||
.expect("a checkpoint commit")
|
||||
.to_string(),
|
||||
entry["run_commit_sha"].as_str().map(str::to_owned),
|
||||
)
|
||||
})
|
||||
.collect()
|
||||
|
|
@ -156,103 +151,30 @@ fn commit_subjects(workspace: &Path) -> Vec<String> {
|
|||
.collect()
|
||||
}
|
||||
|
||||
fn current_branch(workspace: &Path) -> String {
|
||||
let output = Command::new("git")
|
||||
.args(["rev-parse", "--abbrev-ref", "HEAD"])
|
||||
.current_dir(workspace)
|
||||
.output()
|
||||
.expect("git runs");
|
||||
String::from_utf8_lossy(&output.stdout).trim().to_string()
|
||||
}
|
||||
|
||||
fn read(workspace: &Path, name: &str) -> String {
|
||||
std::fs::read_to_string(workspace.join(name))
|
||||
.unwrap_or_else(|err| panic!("{name} in {}: {err}", workspace.display()))
|
||||
}
|
||||
|
||||
/// A three-stage run forked at its first stage's checkpoint continues with
|
||||
/// the other two in a new run: the new run keeps the source's records up to
|
||||
/// `one`, its workspace holds `one`'s file restored from the checkpoint,
|
||||
/// and `two` and `three` run on it and commit on the new run branch. The
|
||||
/// new run says where it came from.
|
||||
/// A local-folder run cannot fork without a published checkpoint.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_fork_at_the_first_stage_continues_with_the_rest_on_its_files() {
|
||||
async fn a_local_folder_fork_is_refused_without_creating_a_run() {
|
||||
let context = test_context!();
|
||||
let server = RunningServer::start().await;
|
||||
let bundle = write_petri_workflow(&context, &three_stage_dot());
|
||||
let source = run_detached(&context, &server, &bundle);
|
||||
wait_for_success(&server, &source).await;
|
||||
let source_timeline = timeline(&server, &source).await;
|
||||
assert_eq!(nodes(&source_timeline), [
|
||||
"start", "one", "two", "three", "exit"
|
||||
]);
|
||||
let source_entries = entries(&source_timeline);
|
||||
let (_, one_execution, one_firing, _, one_sha) = source_entries[1].clone();
|
||||
|
||||
let forked = cli_json(&context, &server, &["fork", &source, "one", "--json"]);
|
||||
assert_eq!(forked["source_run_id"], source);
|
||||
assert_eq!(forked["target"], "@2");
|
||||
assert_eq!(forked["execution"].as_u64(), Some(one_execution));
|
||||
assert_eq!(forked["firing"].as_u64(), Some(one_firing));
|
||||
assert_eq!(forked["checkpoint_sha"], one_sha);
|
||||
assert_eq!(forked["rerun_last"], false);
|
||||
let fork = forked["new_run_id"]
|
||||
.as_str()
|
||||
.expect("the new run id")
|
||||
.to_string();
|
||||
assert_ne!(fork, source);
|
||||
wait_for_success(&server, &fork).await;
|
||||
|
||||
// The fork's timeline: the source's checkpoints up to `one`, then its own.
|
||||
let fork_timeline = timeline(&server, &fork).await;
|
||||
assert_eq!(nodes(&fork_timeline), [
|
||||
"start", "one", "two", "three", "exit"
|
||||
]);
|
||||
let fork_entries = entries(&fork_timeline);
|
||||
assert_eq!(fork_entries[..2], source_entries[..2]);
|
||||
assert_ne!(fork_entries[2].4, source_entries[2].4);
|
||||
assert_eq!(fork_timeline["forked_from"]["source_run_id"], source);
|
||||
assert_eq!(
|
||||
fork_timeline["forked_from"]["execution"].as_u64(),
|
||||
Some(one_execution)
|
||||
);
|
||||
assert_eq!(
|
||||
fork_timeline["forked_from"]["firing"].as_u64(),
|
||||
Some(one_firing)
|
||||
);
|
||||
assert_eq!(fork_timeline["forked_from"]["rerun_last"], false);
|
||||
assert!(source_timeline["forked_from"].is_null());
|
||||
|
||||
// The fork's workspace: `one.txt` restored from the checkpoint, the
|
||||
// rest made by the fork's own stages, on the fork's run branch after the
|
||||
// source's commits.
|
||||
let fork_workspace = workspace(&server, &fork);
|
||||
assert_eq!(read(&fork_workspace, "one.txt"), "one\n");
|
||||
assert_eq!(read(&fork_workspace, "two.txt"), "two\n");
|
||||
assert_eq!(read(&fork_workspace, "three.txt"), "three\n");
|
||||
assert_eq!(current_branch(&fork_workspace), format!("fabro/run/{fork}"));
|
||||
assert_eq!(commit_subjects(&fork_workspace), [
|
||||
format!("fabro({source}): start (success)"),
|
||||
format!("fabro({source}): one (success)"),
|
||||
format!("fabro({fork}): two (success)"),
|
||||
format!("fabro({fork}): three (success)"),
|
||||
format!("fabro({fork}): exit (success)"),
|
||||
]);
|
||||
|
||||
// The projection names the origin on both sides.
|
||||
let state = run_json(&server, &format!("runs/{fork}/state")).await;
|
||||
assert_eq!(state["forked_from"]["source_run_id"], source);
|
||||
assert_eq!(state["spec"]["fork_source_ref"]["source_run_id"], source);
|
||||
assert_eq!(state["spec"]["fork_source_ref"]["checkpoint_sha"], one_sha);
|
||||
assert!(state["retried_from"].is_null());
|
||||
let refused = cli(&context, &server, &["fork", &source, "one", "--json"]);
|
||||
assert!(!refused.status.success());
|
||||
assert!(String::from_utf8_lossy(&refused.stderr).contains("GitHub-backed run"));
|
||||
let summary = run_json(&server, &format!("runs/{source}")).await;
|
||||
assert!(summary["superseded_by"].is_null());
|
||||
assert_eq!(summary["lifecycle"]["archived"], false);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A retry starts the workflow over in a new run: every stage runs again
|
||||
/// in a fresh workspace, so a transient failure passes the second time.
|
||||
/// A retry starts the workflow again in a new workspace. A transient
|
||||
/// failure can succeed on this second execution.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_retry_starts_over_and_succeeds_when_the_failure_was_transient() {
|
||||
let context = test_context!();
|
||||
|
|
@ -303,65 +225,22 @@ async fn a_retry_starts_over_and_succeeds_when_the_failure_was_transient() {
|
|||
server.shutdown();
|
||||
}
|
||||
|
||||
/// A rewind forks the run and supersedes it: the source is archived and
|
||||
/// names the run that replaced it.
|
||||
/// Refusing rewind must leave the original run available and unarchived.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_rewind_supersedes_its_source() {
|
||||
async fn a_local_folder_rewind_is_refused_without_archiving_its_source() {
|
||||
let context = test_context!();
|
||||
let server = RunningServer::start().await;
|
||||
let bundle = write_petri_workflow(&context, &three_stage_dot());
|
||||
let source = run_detached(&context, &server, &bundle);
|
||||
wait_for_success(&server, &source).await;
|
||||
|
||||
// Without a target the command lists the timeline instead.
|
||||
let listed = cli_json(&context, &server, &["rewind", &source, "--json"]);
|
||||
assert_eq!(
|
||||
listed["entries"]
|
||||
.as_array()
|
||||
.map(Vec::len)
|
||||
.expect("a timeline"),
|
||||
5
|
||||
);
|
||||
|
||||
let rewound = cli_json(&context, &server, &["rewind", &source, "@3", "--json"]);
|
||||
assert_eq!(rewound["source_run_id"], source);
|
||||
assert_eq!(rewound["target"], "@3");
|
||||
assert_eq!(rewound["archived"], true);
|
||||
assert!(rewound["archive_error"].is_null());
|
||||
assert_eq!(rewound["status"], 200);
|
||||
let replacement = rewound["new_run_id"]
|
||||
.as_str()
|
||||
.expect("the new run id")
|
||||
.to_string();
|
||||
|
||||
let summary = run_json(&server, &format!("runs/{source}")).await;
|
||||
assert_eq!(summary["superseded_by"], replacement);
|
||||
assert_eq!(summary["lifecycle"]["archived"], true);
|
||||
let state = run_json(&server, &format!("runs/{source}/state")).await;
|
||||
assert_eq!(state["superseded_by"], replacement);
|
||||
|
||||
wait_for_success(&server, &replacement).await;
|
||||
assert_eq!(nodes(&timeline(&server, &replacement).await), [
|
||||
"start", "one", "two", "three", "exit"
|
||||
]);
|
||||
let replacement_workspace = workspace(&server, &replacement);
|
||||
assert_eq!(read(&replacement_workspace, "two.txt"), "two\n");
|
||||
assert_eq!(commit_subjects(&replacement_workspace), [
|
||||
format!("fabro({source}): start (success)"),
|
||||
format!("fabro({source}): one (success)"),
|
||||
format!("fabro({source}): two (success)"),
|
||||
format!("fabro({replacement}): three (success)"),
|
||||
format!("fabro({replacement}): exit (success)"),
|
||||
]);
|
||||
|
||||
// An archived run is neither forked nor rewound again.
|
||||
let refused = cli(&context, &server, &["fork", &source, "@2"]);
|
||||
assert_eq!(listed["entries"].as_array().unwrap().len(), 5);
|
||||
let refused = cli(&context, &server, &["rewind", &source, "@3", "--json"]);
|
||||
assert!(!refused.status.success());
|
||||
assert!(
|
||||
String::from_utf8_lossy(&refused.stderr).contains("archived"),
|
||||
"stderr:\n{}",
|
||||
String::from_utf8_lossy(&refused.stderr)
|
||||
);
|
||||
assert!(String::from_utf8_lossy(&refused.stderr).contains("GitHub-backed run"));
|
||||
let summary = run_json(&server, &format!("runs/{source}")).await;
|
||||
assert!(summary["superseded_by"].is_null());
|
||||
assert_eq!(summary["lifecycle"]["archived"], false);
|
||||
server.shutdown();
|
||||
}
|
||||
|
||||
|
|
@ -375,33 +254,18 @@ async fn the_timeline_lists_every_checkpoint_with_its_commit() {
|
|||
let run_id = run_detached(&context, &server, &bundle);
|
||||
wait_for_success(&server, &run_id).await;
|
||||
|
||||
let recorded = server.checkpoints(&run_id).await;
|
||||
let recorded: Vec<(u64, u64, u64, Option<String>)> = server
|
||||
.checkpoints(&run_id)
|
||||
.await
|
||||
.into_iter()
|
||||
.map(|(key, sha)| (key.execution, key.firing, u64::from(key.attempt), Some(sha)))
|
||||
.collect();
|
||||
let listed = cli_json(&context, &server, &["timeline", &run_id, "--json"]);
|
||||
let listed_entries: Vec<(u64, u64, u64, String)> = listed["entries"]
|
||||
.as_array()
|
||||
.expect("entries")
|
||||
.iter()
|
||||
.map(|entry| {
|
||||
(
|
||||
entry["execution"].as_u64().expect("execution"),
|
||||
entry["firing"].as_u64().expect("firing"),
|
||||
entry["attempt"].as_u64().expect("attempt"),
|
||||
entry["run_commit_sha"].as_str().expect("sha").to_string(),
|
||||
)
|
||||
})
|
||||
let listed_entries: Vec<(u64, u64, u64, Option<String>)> = entries(&listed)
|
||||
.into_iter()
|
||||
.map(|(_, execution, firing, attempt, sha)| (execution, firing, attempt, sha))
|
||||
.collect();
|
||||
let recorded_entries: Vec<(u64, u64, u64, String)> = recorded
|
||||
.iter()
|
||||
.map(|(key, sha)| {
|
||||
(
|
||||
key.execution,
|
||||
key.firing,
|
||||
u64::from(key.attempt),
|
||||
sha.clone(),
|
||||
)
|
||||
})
|
||||
.collect();
|
||||
assert_eq!(listed_entries, recorded_entries);
|
||||
assert_eq!(listed_entries, recorded);
|
||||
assert_eq!(listed_entries.len(), 5);
|
||||
let ordinals: Vec<u64> = listed["entries"]
|
||||
.as_array()
|
||||
|
|
@ -437,7 +301,7 @@ async fn the_timeline_lists_every_checkpoint_with_its_commit() {
|
|||
/// A checkpoint inside a parallel branch is not a fork position: Petri
|
||||
/// refuses it, and the refusal says why. The join, in the root, is.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn a_fork_inside_a_parallel_branch_is_refused() {
|
||||
async fn a_local_folder_parallel_run_cannot_be_forked() {
|
||||
let context = test_context!();
|
||||
let server = RunningServer::start().await;
|
||||
let bundle = write_petri_workflow(&context, ¶llel_dot());
|
||||
|
|
@ -459,7 +323,7 @@ async fn a_fork_inside_a_parallel_branch_is_refused() {
|
|||
);
|
||||
let stderr = String::from_utf8_lossy(&refused.stderr);
|
||||
assert!(
|
||||
stderr.contains("inside a child invocation cannot be forked"),
|
||||
stderr.contains("GitHub-backed run"),
|
||||
"the refusal names the branch:\n{stderr}"
|
||||
);
|
||||
// Nothing was started for it.
|
||||
|
|
@ -473,25 +337,7 @@ async fn a_fork_inside_a_parallel_branch_is_refused() {
|
|||
.collect();
|
||||
assert!(ids.iter().all(|id| *id == source), "runs: {ids:?}");
|
||||
|
||||
// A fork at the stage after the join continues from the root.
|
||||
let done = all
|
||||
.iter()
|
||||
.zip(1_u64..)
|
||||
.find(|((node, ..), _)| node == "done")
|
||||
.map(|(_, ordinal)| ordinal)
|
||||
.expect("`done` has a checkpoint");
|
||||
let forked = cli_json(&context, &server, &[
|
||||
"fork",
|
||||
&source,
|
||||
&format!("@{done}"),
|
||||
"--json",
|
||||
]);
|
||||
let fork = forked["new_run_id"]
|
||||
.as_str()
|
||||
.expect("the new run id")
|
||||
.to_string();
|
||||
wait_for_success(&server, &fork).await;
|
||||
let fork_workspace = workspace(&server, &fork);
|
||||
assert_eq!(read(&fork_workspace, "done.txt"), "done\n");
|
||||
let refused = cli(&context, &server, &["fork", &source, "done"]);
|
||||
assert!(!refused.status.success());
|
||||
server.shutdown();
|
||||
}
|
||||
|
|
|
|||
|
|
@ -5,8 +5,8 @@
|
|||
//! the stages the projector folded them onto. A fork resolves a target on
|
||||
//! that timeline (`@ordinal`, a node, or `node@visit`; the latest checkpoint
|
||||
//! by default), creates the new run's row (`fabro_workflow::operations`),
|
||||
//! seeds its records, checkpoints, snapshots and run branch from the source
|
||||
//! (`fabro_petri::fork`), and queues it in resume mode, so its worker
|
||||
//! seeds its execution records, checkpoint metadata and run branch from the
|
||||
//! source (`fabro_petri::fork`), and queues it in resume mode, so its worker
|
||||
//! acquires a fresh workspace, restores the checkpoint's commit into it and
|
||||
//! continues from the position. A rewind is a fork of a terminal run that
|
||||
//! archives the source and records `run.superseded` on it. A retry is not a
|
||||
|
|
@ -27,7 +27,7 @@ use fabro_petri::fork::{self as petri_fork, ForkError, ForkRequest};
|
|||
use fabro_petri::petri::RunStore;
|
||||
use fabro_petri::platform_records::SqlitePlatformRecords;
|
||||
use fabro_store::{PlatformRecordKind, RunProjection};
|
||||
use fabro_types::{FailureReason, Principal, RunId};
|
||||
use fabro_types::{FailureReason, Principal, RunId, RunTarget};
|
||||
use fabro_util::error as error_util;
|
||||
use fabro_workflow::Error as WorkflowError;
|
||||
use fabro_workflow::operations::{self, ForkTarget, ResolvedForkTarget, RunTimeline};
|
||||
|
|
@ -270,6 +270,14 @@ async fn fork_at(
|
|||
ForkKind::Rewind => operations::ensure_rewindable(&source, &id),
|
||||
}
|
||||
.map_err(workflow_operation_error)?;
|
||||
if !matches!(source.spec.target, Some(RunTarget::Git(_)))
|
||||
|| !source.spec.settings.run.run_branch.enabled
|
||||
|| !source.spec.settings.run.run_branch.push
|
||||
{
|
||||
return Err(ApiError::bad_request(
|
||||
"forking requires a GitHub-backed run with checkpoint pushes enabled",
|
||||
));
|
||||
}
|
||||
let timeline = timeline(state, id).await?;
|
||||
let entry = timeline
|
||||
.resolve_or_latest(target.as_ref())
|
||||
|
|
@ -287,7 +295,6 @@ async fn fork_at(
|
|||
|
||||
let new_run_id = RunId::new();
|
||||
let storage = Storage::new(state.server_storage_dir());
|
||||
let source_run_dir = storage.run_scratch(&id).root().to_path_buf();
|
||||
let run_dir = storage.run_scratch(&new_run_id).root().to_path_buf();
|
||||
operations::persist_forked_run(state.store_ref().as_ref(), &operations::ForkedRunInput {
|
||||
source: &source,
|
||||
|
|
@ -304,7 +311,6 @@ async fn fork_at(
|
|||
let seeded = petri_fork::fork(ForkRequest {
|
||||
source: id,
|
||||
fork: new_run_id,
|
||||
source_run_dir: source_run_dir.join("petri"),
|
||||
fork_run_dir: run_dir.join("petri"),
|
||||
store,
|
||||
records: Arc::new(SqlitePlatformRecords::new(Arc::clone(
|
||||
|
|
|
|||
|
|
@ -51,6 +51,7 @@ use fabro_petri::providers::SandboxProviderConfig;
|
|||
use fabro_petri::recovery::{self, Recovery, RecoveryRequest};
|
||||
use fabro_petri::runtime::{self, RuntimeSpec};
|
||||
use fabro_petri::secrets::VaultSecrets;
|
||||
use fabro_petri::source::RunSource;
|
||||
use fabro_petri::{SqliteRunStore, admission, projection, run_graph};
|
||||
use fabro_static::EnvVars;
|
||||
use fabro_store::platform_records::{RunLifecycleKind, RunLifecycleRecord};
|
||||
|
|
@ -460,13 +461,21 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
|
|||
let observers = vec![petri_interviewer.observer()];
|
||||
let (_, eligible) = state.resolve_llm_client_with_ready_ids().await;
|
||||
let dry_run = run_state.spec.settings.run.execution.mode == RunMode::DryRun;
|
||||
// The in-process path serves the server's tests: a Git target is
|
||||
// fetched without a credential, and nothing is published.
|
||||
let source = RunSource::for_run(
|
||||
run_state.spec.target.as_ref(),
|
||||
&run_state.spec.settings.run,
|
||||
None,
|
||||
);
|
||||
let hooks = HooksSpec::for_run(
|
||||
Arc::new(SqlitePlatformRecords::new(Arc::clone(
|
||||
&state.stores.run_summaries,
|
||||
))),
|
||||
&run_state.spec.settings.run,
|
||||
Arc::new(StoreArtifactWriter::new(state.artifact_store.clone())),
|
||||
);
|
||||
)
|
||||
.with_source(source);
|
||||
let runtime = runtime_spec(
|
||||
&state,
|
||||
&eligible,
|
||||
|
|
@ -576,12 +585,10 @@ pub(crate) async fn reconcile_on_startup(
|
|||
let mode = if held {
|
||||
let request = RecoveryRequest::for_run(
|
||||
run_id,
|
||||
run_dir.join("petri"),
|
||||
Arc::new(SqliteRunStore::new(state.db_pool.clone())),
|
||||
Arc::new(SqlitePlatformRecords::new(Arc::clone(
|
||||
&state.stores.run_summaries,
|
||||
))),
|
||||
&run_state.spec.settings.run,
|
||||
);
|
||||
match recovery::recover(request)
|
||||
.await
|
||||
|
|
@ -592,7 +599,7 @@ pub(crate) async fn reconcile_on_startup(
|
|||
info!(
|
||||
run_id = %run_id,
|
||||
workspaces = workspaces.len(),
|
||||
"Petri run's workspaces match its durable state"
|
||||
"Petri recovery plan ready; the worker will verify retained workspaces"
|
||||
);
|
||||
RunExecutionMode::Resume
|
||||
}
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -131,14 +131,16 @@ pub enum RunStatus {
|
|||
/// What the durable record says about the run once it ended.
|
||||
#[derive(Clone, Debug)]
|
||||
pub struct RunOutcome {
|
||||
pub status: RunStatus,
|
||||
pub status: RunStatus,
|
||||
/// The root invocation's failure message, when it failed.
|
||||
pub failure: Option<String>,
|
||||
pub failure: Option<String>,
|
||||
/// Whether the record is whole: the run recorded its finish and every
|
||||
/// log replays byte for byte.
|
||||
pub complete: bool,
|
||||
pub complete: bool,
|
||||
/// Every reason `complete` is false.
|
||||
pub incomplete: Vec<String>,
|
||||
pub incomplete: Vec<String>,
|
||||
/// Whether the run failed publishing its work after its last stage.
|
||||
pub publish_failed: bool,
|
||||
}
|
||||
|
||||
/// Why the run could not be executed or its outcome read.
|
||||
|
|
@ -193,12 +195,8 @@ pub async fn run(request: RunRequest) -> Result<RunOutcome, RunError> {
|
|||
if backend == SandboxBackend::Daytona {
|
||||
options.sandbox.daytona_resources = daytona_resources(&request.resources)?;
|
||||
}
|
||||
// Fabro's hooks restore a sandbox workspace from its snapshots at the
|
||||
// scope's acquisition, so a lease whose sandbox is gone gets a fresh
|
||||
// one instead of failing the run.
|
||||
if request.hooks.is_some() && backend != SandboxBackend::Host {
|
||||
options.sandbox.lost_sandbox = LostSandbox::Replace;
|
||||
}
|
||||
// A normal resume requires its original sandbox to survive.
|
||||
options.sandbox.lost_sandbox = LostSandbox::Refuse;
|
||||
let resumed = matches!(request.execution, Execution::Resume);
|
||||
let mut runtime = request
|
||||
.runtime
|
||||
|
|
@ -301,6 +299,16 @@ pub async fn run(request: RunRequest) -> Result<RunOutcome, RunError> {
|
|||
outcome.status = RunStatus::Failed;
|
||||
outcome.failure = Some(failure);
|
||||
}
|
||||
// A successful run whose publication failed is a failed run: its work
|
||||
// did not reach where the settings sent it.
|
||||
if let Some(failure) = fabro_hooks
|
||||
.as_ref()
|
||||
.and_then(|hooks| hooks.publish_failure())
|
||||
{
|
||||
outcome.status = RunStatus::Failed;
|
||||
outcome.failure = Some(failure);
|
||||
outcome.publish_failed = true;
|
||||
}
|
||||
Ok(outcome)
|
||||
}
|
||||
|
||||
|
|
@ -360,6 +368,7 @@ pub fn conclusion(result: &Result<RunOutcome, RunError>) -> Conclusion {
|
|||
Ok(outcome) => {
|
||||
let reason = match outcome.status {
|
||||
RunStatus::Cancelled => FailureReason::Cancelled,
|
||||
RunStatus::Failed if outcome.publish_failed => FailureReason::PublishFailed,
|
||||
RunStatus::Success | RunStatus::Failed => FailureReason::WorkflowError,
|
||||
};
|
||||
Conclusion::Failed {
|
||||
|
|
@ -493,6 +502,7 @@ fn outcome(
|
|||
failure,
|
||||
complete: inspection.complete,
|
||||
incomplete: inspection.incomplete,
|
||||
publish_failed: false,
|
||||
})
|
||||
}
|
||||
|
||||
|
|
@ -531,9 +541,21 @@ mod tests {
|
|||
} else {
|
||||
vec!["execution 0 did not finish".to_string()]
|
||||
},
|
||||
publish_failed: false,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_failed_publication_concludes_publish_failed() {
|
||||
let mut outcome = outcome_with(RunStatus::Failed, Some("the push was rejected"), true);
|
||||
outcome.publish_failed = true;
|
||||
let Conclusion::Failed { reason, message } = conclusion(&Ok(outcome)) else {
|
||||
panic!("the run failed");
|
||||
};
|
||||
assert_eq!(reason, FailureReason::PublishFailed);
|
||||
assert!(message.contains("the push was rejected"), "{message}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_whole_successful_record_concludes_succeeded() {
|
||||
assert_eq!(
|
||||
|
|
|
|||
|
|
@ -1,5 +1,5 @@
|
|||
//! Forking a Fabro run at a checkpoint: the seam over Petri's
|
||||
//! `host::fork_from` that rewind, fork and retry are built on.
|
||||
//! `host::fork_from` that rewind and fork are built on.
|
||||
//!
|
||||
//! Fabro's checkpoint record ties a Petri position `(execution, firing)` to
|
||||
//! a Git commit. A fork seeds a new run from the source's records up to such
|
||||
|
|
@ -12,14 +12,10 @@
|
|||
//! `run.started` whose `forked_from` names the source and the position. No
|
||||
//! sandbox lease is carried over, so the resume acquires the position
|
||||
//! execution's scopes fresh.
|
||||
//! 2. The source's checkpoint records for every attempt the fork kept are
|
||||
//! written again under the new run, at their positions, and the snapshot
|
||||
//! repository of every workspace they name is seeded with those checkpoints'
|
||||
//! refs alone, fetched from the source's repository under the source's run
|
||||
//! scratch. That is what the resume's recovery plan
|
||||
//! ([`crate::recovery::plan`]) reads: the last durable finish of the
|
||||
//! position execution names the snapshot the fresh workspace is restored to
|
||||
//! at `scope_acquired`, on the host and in a sandbox alike.
|
||||
//! 2. The source's checkpoint records for retained attempts are copied under
|
||||
//! the new run. At first scope acquisition the worker fetches the source
|
||||
//! run's published branch from GitHub and checks out the selected commit. No
|
||||
//! repository is copied or stored on the server.
|
||||
//! 3. The fork's `run.branch` record names the run branch the restore creates
|
||||
//! (`fabro/run/<new id>`) and the commit it starts from, with the
|
||||
//! `git.identity` beside it, both at the checkpoint's position, so the hooks
|
||||
|
|
@ -31,7 +27,6 @@
|
|||
|
||||
use std::collections::{BTreeMap, BTreeSet};
|
||||
use std::path::PathBuf;
|
||||
use std::process::Stdio;
|
||||
use std::sync::Arc;
|
||||
|
||||
use fabro_checkpoint::author::GitAuthor;
|
||||
|
|
@ -50,11 +45,9 @@ use petri_execution::{
|
|||
use petri_runtime::RunOptions;
|
||||
use petri_runtime::ir::FiringId;
|
||||
use petri_store::StoreError;
|
||||
use tokio::fs;
|
||||
use tokio::process::Command;
|
||||
use tracing::{debug, info};
|
||||
|
||||
use crate::checkpoint::{CheckpointKey, RunWorkspaces};
|
||||
use crate::checkpoint::CheckpointKey;
|
||||
use crate::platform_records::{PlatformRecordError, PlatformRecords};
|
||||
use crate::projection::FoldState;
|
||||
use crate::projector::ProjectError;
|
||||
|
|
@ -63,26 +56,23 @@ use crate::providers::{self, SandboxProviderConfig};
|
|||
/// One fork to seed.
|
||||
pub struct ForkRequest {
|
||||
/// The run whose records are copied.
|
||||
pub source: RunId,
|
||||
pub source: RunId,
|
||||
/// The new run's id: its Petri run key and its own run scratch.
|
||||
pub fork: RunId,
|
||||
/// The source's Petri run directory (its scratch's `petri`), where its
|
||||
/// snapshot repositories are.
|
||||
pub source_run_dir: PathBuf,
|
||||
/// The fork's Petri run directory, where its snapshot repositories go.
|
||||
pub fork_run_dir: PathBuf,
|
||||
pub fork: RunId,
|
||||
/// The fork's Petri run directory, used by Petri.
|
||||
pub fork_run_dir: PathBuf,
|
||||
/// The server's run store: the source is read from it, the fork is
|
||||
/// written into it.
|
||||
pub store: Arc<dyn RunStore>,
|
||||
pub store: Arc<dyn RunStore>,
|
||||
/// The platform records of both runs.
|
||||
pub records: Arc<dyn PlatformRecords>,
|
||||
pub records: Arc<dyn PlatformRecords>,
|
||||
/// The position the source's records are kept up to.
|
||||
pub position: ForkPosition,
|
||||
pub position: ForkPosition,
|
||||
/// Whether the position's firing runs again (a retry of a failed
|
||||
/// stage) instead of keeping its finish.
|
||||
pub rerun_last: bool,
|
||||
pub rerun_last: bool,
|
||||
/// The run's settings, for its Git author and checkpoint settings.
|
||||
pub settings: RunNamespace,
|
||||
pub settings: RunNamespace,
|
||||
}
|
||||
|
||||
/// A seeded fork, not yet resumed.
|
||||
|
|
@ -125,17 +115,13 @@ pub enum ForkError {
|
|||
Log(#[source] CoordinatorStoreError),
|
||||
#[error("the checkpoint records could not be read or written")]
|
||||
Records(#[source] PlatformRecordError),
|
||||
#[error("the snapshot repository for `{workspace}` could not be seeded: {detail}")]
|
||||
Snapshots {
|
||||
workspace: String,
|
||||
detail: String,
|
||||
},
|
||||
}
|
||||
|
||||
/// Refuse a position Petri would refuse, before anything is written for
|
||||
/// the fork: an execution the source does not have, or one inside a child
|
||||
/// invocation (a branch of a parallel node), whose caller's firing is live
|
||||
/// at every position inside it. The messages are Petri's own.
|
||||
/// at every position inside it. Also refuse the final terminal checkpoint:
|
||||
/// it leaves no work that would acquire a workspace for publication.
|
||||
pub async fn check(
|
||||
store: &dyn RunStore,
|
||||
source: RunId,
|
||||
|
|
@ -160,13 +146,33 @@ pub async fn check(
|
|||
position.execution
|
||||
)));
|
||||
}
|
||||
let inspection = inspect::inspect_run(&*logs)
|
||||
.await
|
||||
.map_err(ForkError::Inspect)?;
|
||||
if inspection
|
||||
.executions
|
||||
.iter()
|
||||
.find(|execution| execution.execution == position.execution)
|
||||
.and_then(|execution| execution.engine.as_ref())
|
||||
.is_some_and(|engine| {
|
||||
matches!(engine.exit, Some(inspect::ExitInspection::Terminal { .. }))
|
||||
&& engine
|
||||
.attempts
|
||||
.last()
|
||||
.is_some_and(|attempt| attempt.firing == position.firing.raw())
|
||||
})
|
||||
{
|
||||
return Err(ForkError::Refused(
|
||||
"the terminal checkpoint has no remaining work to acquire a sandbox; select an earlier checkpoint or retry the workflow from the start".to_string(),
|
||||
));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Seed the fork: Petri's records, then the kept checkpoints, their
|
||||
/// snapshots and the run branch. The new run must not exist in the store
|
||||
/// yet.
|
||||
/// Seed the fork: Petri's records, the kept checkpoints and the run branch. The
|
||||
/// new run must not exist in the store yet.
|
||||
pub async fn fork(request: ForkRequest) -> Result<Forked, ForkError> {
|
||||
check(request.store.as_ref(), request.source, request.position).await?;
|
||||
let source_key = RunKey::new(request.source.to_string());
|
||||
let fork_key = RunKey::new(request.fork.to_string());
|
||||
let source_logs = request
|
||||
|
|
@ -272,8 +278,6 @@ pub async fn fork(request: ForkRequest) -> Result<Forked, ForkError> {
|
|||
});
|
||||
}
|
||||
|
||||
// The snapshot repositories: one per workspace the kept checkpoints
|
||||
// name, holding those checkpoints' refs alone.
|
||||
let author = request
|
||||
.settings
|
||||
.git
|
||||
|
|
@ -281,30 +285,6 @@ pub async fn fork(request: ForkRequest) -> Result<Forked, ForkError> {
|
|||
.as_ref()
|
||||
.map(GitAuthor::from)
|
||||
.unwrap_or_default();
|
||||
let source_workspaces = RunWorkspaces::new(
|
||||
request.source_run_dir.clone(),
|
||||
request.source.to_string(),
|
||||
author.clone(),
|
||||
&request.settings.checkpoint,
|
||||
);
|
||||
let fork_workspaces = RunWorkspaces::new(
|
||||
request.fork_run_dir.clone(),
|
||||
request.fork.to_string(),
|
||||
author.clone(),
|
||||
&request.settings.checkpoint,
|
||||
);
|
||||
let mut by_workspace: BTreeMap<String, Vec<CheckpointKey>> = BTreeMap::new();
|
||||
for kept in &checkpoints {
|
||||
if let Some(workspace) = &kept.workspace {
|
||||
by_workspace
|
||||
.entry(workspace.clone())
|
||||
.or_default()
|
||||
.push(kept.key);
|
||||
}
|
||||
}
|
||||
for (workspace, keys) in &by_workspace {
|
||||
seed_snapshots(&source_workspaces, &fork_workspaces, workspace, keys).await?;
|
||||
}
|
||||
|
||||
// The run branch the restore creates, from the checkpoint the fork
|
||||
// starts on, and the identity that authors the fork's commits.
|
||||
|
|
@ -315,7 +295,7 @@ pub async fn fork(request: ForkRequest) -> Result<Forked, ForkError> {
|
|||
firing: start.key.firing,
|
||||
};
|
||||
let branch = PlatformRecord::RunBranch(RunBranchRecord {
|
||||
run_branch: Some(fork_workspaces.run_branch()),
|
||||
run_branch: Some(format!("fabro/run/{}", request.fork)),
|
||||
base_sha: Some(start.sha.clone()),
|
||||
workspace: start.workspace.clone(),
|
||||
});
|
||||
|
|
@ -377,75 +357,6 @@ fn checkpoint_key(record: &CheckpointRecord) -> Option<CheckpointKey> {
|
|||
})
|
||||
}
|
||||
|
||||
/// Create the fork's bare snapshot repository for `workspace` and fetch the
|
||||
/// kept checkpoints' refs into it from the source's.
|
||||
async fn seed_snapshots(
|
||||
source: &RunWorkspaces,
|
||||
fork: &RunWorkspaces,
|
||||
workspace: &str,
|
||||
keys: &[CheckpointKey],
|
||||
) -> Result<(), ForkError> {
|
||||
let failed = |detail: String| ForkError::Snapshots {
|
||||
workspace: workspace.to_string(),
|
||||
detail,
|
||||
};
|
||||
let source_repository = source.snapshot_repository(workspace);
|
||||
if !fs::try_exists(&source_repository).await.unwrap_or(false) {
|
||||
return Err(failed(format!(
|
||||
"the source run has no snapshot repository at {}",
|
||||
source_repository.display()
|
||||
)));
|
||||
}
|
||||
let repository = fork.snapshot_repository(workspace);
|
||||
fs::create_dir_all(&repository).await.map_err(|error| {
|
||||
failed(format!(
|
||||
"{} could not be created: {error}",
|
||||
repository.display()
|
||||
))
|
||||
})?;
|
||||
git(&repository, &["init", "-q", "--bare"])
|
||||
.await
|
||||
.map_err(failed)?;
|
||||
let mut args = vec![
|
||||
"fetch".to_string(),
|
||||
"-q".to_string(),
|
||||
source_repository.to_string_lossy().into_owned(),
|
||||
];
|
||||
for key in keys {
|
||||
let name = key.snapshot_ref();
|
||||
args.push(format!("+{name}:{name}"));
|
||||
}
|
||||
git(&repository, &args).await.map_err(failed)?;
|
||||
debug!(
|
||||
workspace,
|
||||
refs = keys.len(),
|
||||
repository = %repository.display(),
|
||||
"the fork's snapshot repository is seeded"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Run `git` in `repository`; a non-zero exit is the error's detail.
|
||||
async fn git<S: AsRef<str>>(repository: &std::path::Path, args: &[S]) -> Result<(), String> {
|
||||
let output = Command::new("git")
|
||||
.args(args.iter().map(AsRef::as_ref))
|
||||
.current_dir(repository)
|
||||
.stdin(Stdio::null())
|
||||
.output()
|
||||
.await
|
||||
.map_err(|error| format!("git could not run: {error}"))?;
|
||||
if output.status.success() {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(format!(
|
||||
"git {} failed ({}): {}",
|
||||
args.first().map_or("", AsRef::as_ref),
|
||||
output.status,
|
||||
String::from_utf8_lossy(&output.stderr).trim()
|
||||
))
|
||||
}
|
||||
}
|
||||
|
||||
/// The position a checkpoint's execution and firing name, in Petri's ids.
|
||||
#[must_use]
|
||||
pub fn position(execution: u64, firing: u64) -> ForkPosition {
|
||||
|
|
|
|||
|
|
@ -11,16 +11,18 @@
|
|||
//! embedding host does. Its own work, at each point:
|
||||
//!
|
||||
//! - `prepare_result`: the checkpoint commit, before the `StepFinished` record
|
||||
//! is appended, so a durable finish implies a durable snapshot. A stage that
|
||||
//! failed on its own terms is committed like a successful one; only a
|
||||
//! cancelled attempt is not. A failed commit is fatal to the run: the outcome
|
||||
//! becomes a failure of class `checkpoint_failed`, the run is cancelled
|
||||
//! through the coordinator handle, and `transition` refuses the firing's
|
||||
//! routes, so no route is taken. The commit that creates the run branch also
|
||||
//! records where it started: the `run.branch` platform record (the branch
|
||||
//! name and the base commit) and the `git.identity` record (who authors the
|
||||
//! commits, and where that identity came from), both at that checkpoint's
|
||||
//! stage position, so the stream orders them with the firing's finish.
|
||||
//! is appended. Git-backed runs push that commit from the sandbox; a failed
|
||||
//! push warns and is retried at the next checkpoint and at final publication.
|
||||
//! A stage that failed on its own terms is committed like a successful one;
|
||||
//! only a cancelled attempt is not. A failed commit is fatal to the run: the
|
||||
//! outcome becomes a failure of class `checkpoint_failed`, the run is
|
||||
//! cancelled through the coordinator handle, and `transition` refuses the
|
||||
//! firing's routes, so no route is taken. The commit that creates the run
|
||||
//! branch also records where it started: the `run.branch` platform record
|
||||
//! (the branch name and the base commit) and the `git.identity` record (who
|
||||
//! authors the commits, and where that identity came from), both at that
|
||||
//! checkpoint's stage position, so the stream orders them with the firing's
|
||||
//! finish.
|
||||
//! - `transition`: the platform checkpoint record, keyed on the Petri position
|
||||
//! and the checkpoint's operation identity, with the stage's diff from its
|
||||
//! parent commit (`diff_summary`, and the patch as a blob); then the stage's
|
||||
|
|
@ -30,9 +32,16 @@
|
|||
//! was already collected earlier in the run. A failed write is a recorded
|
||||
//! problem on the transition, never a blocked route.
|
||||
//! - `run_finished`: the run's diff, its run branch against its base commit, as
|
||||
//! the `run.diff` platform record with the patch as a blob; then the
|
||||
//! forwarded point, so the local service runs `run_complete` and `run_failed`
|
||||
//! with the sandbox in place.
|
||||
//! the `run.diff` platform record with the patch as a blob; for a successful
|
||||
//! run, its publication ([`RunPublisher`]: the platform pushes the run branch
|
||||
//! and opens a pull request), whose failure fails the run before its terminal
|
||||
//! record; then the forwarded point, so the local service runs `run_complete`
|
||||
//! and `run_failed` with the sandbox in place.
|
||||
//! - `scope_acquired`: a fresh run's Git target checked out into the workspace
|
||||
//! from inside the scope ([`crate::source`]); a resumed run uses its
|
||||
//! surviving workspace, while an explicit fork fetches the source run's
|
||||
//! branch from GitHub. A checkout that fails fails the scope's firings with
|
||||
//! the reason.
|
||||
//! - `scope_released`: forwarded, so the local service runs `sandbox_cleanup`
|
||||
//! with the sandbox in place. Fabro's own end-of-run work (the terminal
|
||||
//! lifecycle event, notifications on it) is the run lifecycle path's, on the
|
||||
|
|
@ -59,17 +68,14 @@
|
|||
//! Petri's host backend keeps under the run directory (`crate::checkpoint`).
|
||||
//! On Docker or Daytona the workspace lives inside the scope's sandbox: the
|
||||
//! hooks keep the environment Petri hands them at `scope_acquired`, run
|
||||
//! `git` inside the scope through it, and move the commit out as a bundle
|
||||
//! into the same snapshot repository the host path pushes to. Artifacts are
|
||||
//! read out through the same environment on every provider. The same
|
||||
//! point is where a resumed run brings a workspace to the snapshot its
|
||||
//! durable state names, before the first attempt runs in it: verified,
|
||||
//! reset, or, in a fresh sandbox (Petri replaces a lost one on Fabro's
|
||||
//! request), restored from a bundle of the checkpoint. The plan is
|
||||
//! [`recovery::plan`], the one the server applied to host workspaces before
|
||||
//! it relaunched the worker; a host workspace is verified here, unless the
|
||||
//! run is a fork whose fresh workspace nothing restored yet
|
||||
//! ([`crate::fork`]), which is restored from the seeded snapshot repository.
|
||||
//! `git` inside the scope through it. Checkpoints, diffs, fetches and pushes
|
||||
//! all use this environment. Only execution metadata, patches and selected
|
||||
//! artifacts are persisted on the server. Normal resume requires the original
|
||||
//! workspace; a fork acquires a new sandbox and fetches its checkpoint from
|
||||
//! GitHub. A run with no GitHub source commits only when its workspace is on
|
||||
//! the host (a local-folder, empty or dry run); its checkpoints stay in that
|
||||
//! workspace. Docker and Daytona runs with no GitHub source record execution
|
||||
//! checkpoints without Git commits.
|
||||
|
||||
use std::collections::{BTreeMap, HashMap, HashSet};
|
||||
use std::path::{Path, PathBuf};
|
||||
|
|
@ -94,7 +100,7 @@ use petri_runtime::driver::lifecycle::{
|
|||
ScopeReleased, Transition, TransitionError, TransitionReport,
|
||||
};
|
||||
use petri_runtime::executor::{EnvError, ExecEnv};
|
||||
use petri_runtime::ir::{ExecutionId, FailureInfo, ScopeId, Status};
|
||||
use petri_runtime::ir::{ExecutionId, FailureInfo, RunStatus, ScopeId, Status};
|
||||
use serde_json::json;
|
||||
use tokio::sync::{Mutex as AsyncMutex, OnceCell};
|
||||
use tokio::{fs, time};
|
||||
|
|
@ -106,8 +112,10 @@ use crate::checkpoint::{
|
|||
CHECKPOINT_FAILED_CLASS, CheckpointError, CheckpointKey, EXCLUDE_DIRS, RunGitSettings,
|
||||
RunWorkspaces, Site, Snapshot, WorkspaceDiff,
|
||||
};
|
||||
use crate::fork::{self, ForkError};
|
||||
use crate::platform_records::{PlatformRecordError, PlatformRecords};
|
||||
use crate::recovery::{self, Plan, RecoveryError, RestoreTarget};
|
||||
use crate::source::RunSource;
|
||||
use crate::workspace::{self, WorkspaceLookup, WorkspaceLookupError};
|
||||
|
||||
/// The note kind the hooks record on a firing about its checkpoint.
|
||||
|
|
@ -195,6 +203,8 @@ pub enum HookError {
|
|||
#[source]
|
||||
source: EnvError,
|
||||
},
|
||||
#[error("the fork origin could not be read")]
|
||||
Fork(#[source] ForkError),
|
||||
#[error("the run's restore plan could not be read")]
|
||||
Plan(#[source] RecoveryError),
|
||||
#[error("{reason}")]
|
||||
|
|
@ -225,6 +235,36 @@ pub struct HooksSpec {
|
|||
pub test_gates: Option<PathBuf>,
|
||||
/// Where captured workspace files go.
|
||||
pub artifact_writer: Arc<dyn ArtifactWriter>,
|
||||
/// Where a Git target's workspaces are checked out from, when a fresh
|
||||
/// run first acquires them. `None` for a run with no remote repository.
|
||||
pub source: Option<RunSource>,
|
||||
/// What a successful run's work does when it ends; `None` publishes
|
||||
/// nothing.
|
||||
pub publisher: Option<Arc<dyn RunPublisher>>,
|
||||
}
|
||||
|
||||
/// What a successful run hands its publisher when it ends: the run branch,
|
||||
/// the commit it ends on, the workspace that holds it, and the
|
||||
/// run's patch against the commit the branch started from.
|
||||
#[derive(Clone, Debug)]
|
||||
pub struct Publication {
|
||||
pub run_branch: String,
|
||||
pub head_sha: String,
|
||||
pub site: Site,
|
||||
pub patch: String,
|
||||
}
|
||||
|
||||
/// The platform's end-of-run publication: what a successful run's work does
|
||||
/// after its last stage and before its terminal record, such as pushing the
|
||||
/// run branch and opening a pull request. An `Err` fails the run with the
|
||||
/// message, as a failed publish did on the legacy executor.
|
||||
#[async_trait::async_trait]
|
||||
pub trait RunPublisher: Send + Sync {
|
||||
/// Push one checkpoint. A failure is retried by later checkpoints and
|
||||
/// final publication, which must succeed for the run to succeed.
|
||||
async fn push(&self, site: &Site, branch: &str, sha: &str) -> Result<(), String>;
|
||||
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String>;
|
||||
}
|
||||
|
||||
impl HooksSpec {
|
||||
|
|
@ -242,9 +282,25 @@ impl HooksSpec {
|
|||
artifacts: settings.artifacts.include.clone(),
|
||||
test_gates: None,
|
||||
artifact_writer,
|
||||
source: None,
|
||||
publisher: None,
|
||||
}
|
||||
}
|
||||
|
||||
/// Publish a successful run's work through `publisher` when it ends.
|
||||
#[must_use]
|
||||
pub fn with_publisher(mut self, publisher: Option<Arc<dyn RunPublisher>>) -> Self {
|
||||
self.publisher = publisher;
|
||||
self
|
||||
}
|
||||
|
||||
/// Check a Git target's workspaces out from `source`.
|
||||
#[must_use]
|
||||
pub fn with_source(mut self, source: Option<RunSource>) -> Self {
|
||||
self.source = source;
|
||||
self
|
||||
}
|
||||
|
||||
#[must_use]
|
||||
pub fn with_test_gates(mut self, gates: Option<PathBuf>) -> Self {
|
||||
self.test_gates = gates;
|
||||
|
|
@ -377,31 +433,36 @@ impl ScopeEnvs {
|
|||
|
||||
/// Fabro's `ExecutionHooks`, around the hooks the runtime installed.
|
||||
pub struct FabroHooks {
|
||||
inner: Arc<dyn ExecutionHooks>,
|
||||
run_id: RunId,
|
||||
records: Arc<dyn PlatformRecords>,
|
||||
inner: Arc<dyn ExecutionHooks>,
|
||||
run_id: RunId,
|
||||
records: Arc<dyn PlatformRecords>,
|
||||
/// Where diff patches go; `None` records summaries alone.
|
||||
blobs: Option<Arc<dyn Blobs>>,
|
||||
artifact_writer: Arc<dyn ArtifactWriter>,
|
||||
workspaces: RunWorkspaces,
|
||||
lookup: WorkspaceLookup,
|
||||
identity: GitIdentity,
|
||||
host_workspaces: bool,
|
||||
test_gates: Option<PathBuf>,
|
||||
handle: OnceLock<CoordinatorHandle>,
|
||||
checkpoints: CheckpointLedger,
|
||||
artifacts: ArtifactLedger,
|
||||
scopes: ScopeEnvs,
|
||||
blobs: Option<Arc<dyn Blobs>>,
|
||||
artifact_writer: Arc<dyn ArtifactWriter>,
|
||||
workspaces: RunWorkspaces,
|
||||
checkpoint_enabled: bool,
|
||||
sites: Mutex<HashMap<String, Site>>,
|
||||
lookup: WorkspaceLookup,
|
||||
identity: GitIdentity,
|
||||
host_workspaces: bool,
|
||||
test_gates: Option<PathBuf>,
|
||||
handle: OnceLock<CoordinatorHandle>,
|
||||
checkpoints: CheckpointLedger,
|
||||
artifacts: ArtifactLedger,
|
||||
scopes: ScopeEnvs,
|
||||
/// The checkpoint failure that ended the run, when one did.
|
||||
failure: Mutex<Option<String>>,
|
||||
failure: Mutex<Option<String>>,
|
||||
publisher: Option<Arc<dyn RunPublisher>>,
|
||||
/// Why the run's publication failed, when it did.
|
||||
publish_failure: Mutex<Option<String>>,
|
||||
/// Whether the run continues from its records: a sandbox workspace is
|
||||
/// then brought to its snapshot when its scope is first acquired.
|
||||
resumed: bool,
|
||||
resumed: bool,
|
||||
/// The snapshot every live sandbox workspace must sit on before work
|
||||
/// resumes in it, read once from the records; an entry leaves when it
|
||||
/// is applied.
|
||||
restore: OnceCell<Mutex<BTreeMap<String, RestoreTarget>>>,
|
||||
store: Arc<dyn RunStore>,
|
||||
restore: OnceCell<Mutex<BTreeMap<String, Vec<RestoreTarget>>>>,
|
||||
store: Arc<dyn RunStore>,
|
||||
}
|
||||
|
||||
impl FabroHooks {
|
||||
|
|
@ -426,12 +487,18 @@ impl FabroHooks {
|
|||
email: spec.git.author.email.clone(),
|
||||
source: spec.git.identity_source,
|
||||
};
|
||||
// A host workspace commits checkpoints without a GitHub source (local
|
||||
// folders, empty Local targets, dry runs); a sandbox without one does
|
||||
// not, so an image without `git` cannot fail the run.
|
||||
let checkpoint_enabled =
|
||||
spec.git.enabled && (spec.source.is_some() || spec.git.host_workspaces);
|
||||
let workspaces = RunWorkspaces::new(
|
||||
run_dir,
|
||||
run_id.to_string(),
|
||||
spec.git.author,
|
||||
&spec.git.checkpoint,
|
||||
);
|
||||
)
|
||||
.with_source(spec.source);
|
||||
Self {
|
||||
inner,
|
||||
run_id,
|
||||
|
|
@ -439,6 +506,8 @@ impl FabroHooks {
|
|||
blobs,
|
||||
artifact_writer: spec.artifact_writer,
|
||||
workspaces,
|
||||
checkpoint_enabled,
|
||||
sites: Mutex::default(),
|
||||
lookup: WorkspaceLookup::new(Arc::clone(&store), run_key),
|
||||
identity,
|
||||
host_workspaces: spec.git.host_workspaces,
|
||||
|
|
@ -452,6 +521,8 @@ impl FabroHooks {
|
|||
},
|
||||
scopes: ScopeEnvs::default(),
|
||||
failure: Mutex::default(),
|
||||
publisher: spec.publisher,
|
||||
publish_failure: Mutex::default(),
|
||||
resumed,
|
||||
restore: OnceCell::new(),
|
||||
store,
|
||||
|
|
@ -473,6 +544,13 @@ impl FabroHooks {
|
|||
sync::lock(&self.failure).clone()
|
||||
}
|
||||
|
||||
/// Why the run's publication failed, when it did: the run then fails
|
||||
/// with this message.
|
||||
#[must_use]
|
||||
pub fn publish_failure(&self) -> Option<String> {
|
||||
sync::lock(&self.publish_failure).clone()
|
||||
}
|
||||
|
||||
/// The run's workspaces on this host, as the hooks reach them.
|
||||
#[must_use]
|
||||
pub fn workspaces(&self) -> &RunWorkspaces {
|
||||
|
|
@ -567,6 +645,10 @@ impl FabroHooks {
|
|||
status: &Status,
|
||||
origin: ResultOrigin,
|
||||
) -> Result<Option<Note>, HookError> {
|
||||
self.gate("prepare", node).await;
|
||||
if !self.checkpoint_enabled {
|
||||
return Ok(None);
|
||||
}
|
||||
let Some((workspace, site)) = self.site_of(context, scope).await? else {
|
||||
// A skipped node or a driver-made outcome may precede the scope's
|
||||
// environment; nothing of the stage's exists to snapshot.
|
||||
|
|
@ -606,6 +688,22 @@ impl FabroHooks {
|
|||
"checkpoint committed"
|
||||
);
|
||||
self.committed(key, &workspace, &snapshot).await;
|
||||
// Only the workspace that owns the run branch publishes it.
|
||||
// Isolated child workspaces must not race to replace its head.
|
||||
if let Some(publisher) = &self.publisher {
|
||||
if self
|
||||
.stored_branch()
|
||||
.await?
|
||||
.is_some_and(|branch| branch.workspace.as_deref() == Some(&workspace))
|
||||
{
|
||||
if let Err(message) = publisher
|
||||
.push(&site, &self.workspaces.run_branch(), &snapshot.sha)
|
||||
.await
|
||||
{
|
||||
warn!(run_id = %self.run_id, error = %message, "checkpoint push failed; final publication will retry");
|
||||
}
|
||||
}
|
||||
}
|
||||
Ok(Some(Note::new(
|
||||
CHECKPOINT_NOTE,
|
||||
json!({
|
||||
|
|
@ -726,17 +824,15 @@ impl FabroHooks {
|
|||
|
||||
/// The restore plan of a resumed run, read once: what every live
|
||||
/// sandbox workspace must be brought to at its first acquisition.
|
||||
async fn restore_targets(&self) -> Result<&Mutex<BTreeMap<String, RestoreTarget>>, HookError> {
|
||||
async fn restore_targets(
|
||||
&self,
|
||||
) -> Result<&Mutex<BTreeMap<String, Vec<RestoreTarget>>>, HookError> {
|
||||
self.restore
|
||||
.get_or_try_init(|| async {
|
||||
let plan = recovery::plan(
|
||||
Arc::clone(&self.store),
|
||||
self.records.as_ref(),
|
||||
&self.run_id,
|
||||
&self.workspaces,
|
||||
)
|
||||
.await
|
||||
.map_err(HookError::Plan)?;
|
||||
let plan =
|
||||
recovery::plan(Arc::clone(&self.store), self.records.as_ref(), &self.run_id)
|
||||
.await
|
||||
.map_err(HookError::Plan)?;
|
||||
match plan {
|
||||
Plan::Resume { targets } => Ok(Mutex::new(targets)),
|
||||
Plan::Start => Ok(Mutex::default()),
|
||||
|
|
@ -746,19 +842,93 @@ impl FabroHooks {
|
|||
.await
|
||||
}
|
||||
|
||||
/// Bring a workspace to the snapshot the resumed run's durable state
|
||||
/// names, once, at its first acquisition. After a restart the server
|
||||
/// already brought a host workspace there, so this verifies; a fork's
|
||||
/// fresh workspace is restored here from the snapshot repository the
|
||||
/// fork seeded; a sandbox workspace is only reachable here.
|
||||
/// A fresh run's workspace, checked out from the run's source before
|
||||
/// the first attempt runs in it. A failure fails the scope's firings
|
||||
/// with the reason.
|
||||
async fn check_out_source(
|
||||
&self,
|
||||
workspace: &str,
|
||||
site: &Site,
|
||||
) -> Result<(), ScopeAcquiredError> {
|
||||
let serialized = self.scopes.lock_for(workspace);
|
||||
let _held = serialized.lock().await;
|
||||
match self.workspaces.check_out_source(site, workspace).await {
|
||||
Ok(Some(sha)) => {
|
||||
info!(
|
||||
run_id = %self.run_id,
|
||||
workspace,
|
||||
sha,
|
||||
site = ?site,
|
||||
"workspace checked out from the run's repository"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
Ok(None) => Ok(()),
|
||||
Err(error) => {
|
||||
warn!(
|
||||
run_id = %self.run_id,
|
||||
workspace,
|
||||
error = %error,
|
||||
"the run's repository could not be checked out"
|
||||
);
|
||||
Err(ScopeAcquiredError::new(format!(
|
||||
"the run's repository could not be checked out: {error}"
|
||||
)))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Reset a surviving workspace to its recorded checkpoint. An explicit
|
||||
/// fork may fetch the source run's branch into a new workspace.
|
||||
async fn restore(&self, workspace: &str, site: &Site) -> Result<(), HookError> {
|
||||
let targets = self.restore_targets().await?;
|
||||
let target = sync::lock(targets).remove(workspace);
|
||||
let Some(target) = target else {
|
||||
let Some(targets) = target else {
|
||||
return Ok(());
|
||||
};
|
||||
let Some(mut target) = targets.first().cloned() else {
|
||||
return Ok(());
|
||||
};
|
||||
let serialized = self.scopes.lock_for(workspace);
|
||||
let _held = serialized.lock().await;
|
||||
if !self
|
||||
.workspaces
|
||||
.has_commit(site, &target.sha)
|
||||
.await
|
||||
.map_err(HookError::Find)?
|
||||
{
|
||||
let origin = fork::origin_of(self.store.as_ref(), self.run_id)
|
||||
.await
|
||||
.map_err(HookError::Fork)?;
|
||||
if let Some(origin) = origin {
|
||||
// An explicit fork may acquire a fresh workspace. Normal
|
||||
// resumes never reconstruct a lost repository.
|
||||
if self
|
||||
.workspaces
|
||||
.head(site)
|
||||
.await
|
||||
.map_err(HookError::Find)?
|
||||
.is_none()
|
||||
{
|
||||
self.workspaces
|
||||
.restore_fork(site, &origin.source_run_id.to_string(), &target.sha)
|
||||
.await
|
||||
.map_err(HookError::Find)?;
|
||||
}
|
||||
}
|
||||
}
|
||||
// Platform records can arrive out of commit order from parallel stages.
|
||||
// Choose by Git ancestry here, where the actual repository is available.
|
||||
for candidate in targets.iter().skip(1) {
|
||||
if self
|
||||
.workspaces
|
||||
.is_ancestor(site, &target.sha, &candidate.sha)
|
||||
.await
|
||||
.map_err(HookError::Find)?
|
||||
{
|
||||
target = candidate.clone();
|
||||
}
|
||||
}
|
||||
let action = recovery::bring_to(&self.workspaces, site, workspace, &target)
|
||||
.await
|
||||
.map_err(HookError::Restore)?;
|
||||
|
|
@ -785,6 +955,41 @@ impl FabroHooks {
|
|||
if sync::lock(recorded).contains(&key) {
|
||||
return Ok(());
|
||||
}
|
||||
if !self.checkpoint_enabled {
|
||||
self.records
|
||||
.append(
|
||||
&self.run_id,
|
||||
&PlatformRecord::Checkpoint(CheckpointRecord {
|
||||
execution: key.execution,
|
||||
firing: key.firing,
|
||||
attempt: Some(key.attempt),
|
||||
workspace: None,
|
||||
git_commit_sha: None,
|
||||
diff_summary: None,
|
||||
patch_blob: None,
|
||||
operation: Some(key.operation()),
|
||||
}),
|
||||
Some(StagePosition {
|
||||
execution: key.execution,
|
||||
firing: key.firing,
|
||||
}),
|
||||
)
|
||||
.await
|
||||
.map_err(|source| HookError::Write {
|
||||
kind: "checkpoint",
|
||||
source,
|
||||
})?;
|
||||
sync::lock(recorded).insert(key);
|
||||
return Ok(());
|
||||
}
|
||||
let site = self
|
||||
.site_of(context, scope)
|
||||
.await?
|
||||
.ok_or(HookError::NoWorkspace {
|
||||
scope,
|
||||
execution: context.execution,
|
||||
})?
|
||||
.1;
|
||||
let (workspace, sha) = if let Some(committed) = self.checkpoints.commit_of(key) {
|
||||
committed
|
||||
} else {
|
||||
|
|
@ -795,14 +1000,14 @@ impl FabroHooks {
|
|||
};
|
||||
let serialized = self.scopes.lock_for(&workspace);
|
||||
let held = serialized.lock().await;
|
||||
let found = self.workspaces.find(&workspace, key).await;
|
||||
let found = self.workspaces.find(&site, key).await;
|
||||
drop(held);
|
||||
let sha = found
|
||||
.map_err(HookError::Find)?
|
||||
.ok_or(HookError::NoCommit { key })?;
|
||||
(workspace, sha)
|
||||
};
|
||||
let (diff_summary, patch_blob) = match self.stage_diff(&workspace, &sha).await {
|
||||
let (diff_summary, patch_blob) = match self.stage_diff(&site, &sha).await {
|
||||
Ok(diff) => diff,
|
||||
Err(error) => {
|
||||
// The record still names the commit; the diff is a view.
|
||||
|
|
@ -851,7 +1056,7 @@ impl FabroHooks {
|
|||
/// the diff is not empty.
|
||||
async fn stage_diff(
|
||||
&self,
|
||||
workspace: &str,
|
||||
site: &Site,
|
||||
sha: &str,
|
||||
) -> Result<(Option<DiffSummary>, Option<BlobHash>), HookError> {
|
||||
let failed = |source| HookError::Diff {
|
||||
|
|
@ -860,7 +1065,7 @@ impl FabroHooks {
|
|||
};
|
||||
let parent = self
|
||||
.workspaces
|
||||
.commit_parent(workspace, sha)
|
||||
.commit_parent(site, sha)
|
||||
.await
|
||||
.map_err(failed)?;
|
||||
let Some(parent) = parent else {
|
||||
|
|
@ -868,7 +1073,7 @@ impl FabroHooks {
|
|||
};
|
||||
let diff = self
|
||||
.workspaces
|
||||
.diff(workspace, Some(&parent), sha)
|
||||
.diff(site, Some(&parent), sha)
|
||||
.await
|
||||
.map_err(failed)?;
|
||||
let patch_blob = self.patch_blob(&diff).await?;
|
||||
|
|
@ -1102,10 +1307,10 @@ impl FabroHooks {
|
|||
}
|
||||
|
||||
/// The run's diff: the run branch's last checkpoint against the base
|
||||
/// the branch started from, in the snapshot repository on this host.
|
||||
/// the branch started from, computed inside the workspace.
|
||||
/// Nothing is recorded for a run that never created its branch or
|
||||
/// never checkpointed.
|
||||
async fn record_run_diff(&self) -> Result<(), HookError> {
|
||||
async fn record_run_diff(&self) -> Result<Option<Publication>, HookError> {
|
||||
self.recorded_checkpoints().await?;
|
||||
let branch = match self.checkpoints.branch.get() {
|
||||
Some(branch) => Some(branch.clone()),
|
||||
|
|
@ -1113,28 +1318,68 @@ impl FabroHooks {
|
|||
};
|
||||
let Some(branch) = branch else {
|
||||
debug!(run_id = %self.run_id, "no run branch is recorded; no run diff");
|
||||
return Ok(());
|
||||
return Ok(None);
|
||||
};
|
||||
let Some(base_sha) = branch.base_sha.clone() else {
|
||||
return Ok(());
|
||||
return Ok(None);
|
||||
};
|
||||
let Some((workspace, head_sha)) = self.checkpoints.last() else {
|
||||
let Some((workspace, _)) = self.checkpoints.last() else {
|
||||
debug!(run_id = %self.run_id, "no checkpoint is recorded; no run diff");
|
||||
return Ok(());
|
||||
return Ok(None);
|
||||
};
|
||||
// The run's diff is measured in the workspace the branch started
|
||||
// in; a last checkpoint elsewhere (a nested invocation's workspace)
|
||||
// is not this branch's head.
|
||||
let workspace = branch.workspace.clone().unwrap_or(workspace);
|
||||
let site = sync::lock(&self.sites)
|
||||
.get(&workspace)
|
||||
.cloned()
|
||||
.ok_or_else(|| HookError::Unresumable {
|
||||
reason: format!(
|
||||
"the run branch workspace {workspace} is unavailable for publication"
|
||||
),
|
||||
})?;
|
||||
let stored = self
|
||||
.records
|
||||
.read_kind(&self.run_id, PlatformRecordKind::Checkpoint)
|
||||
.await
|
||||
.map_err(|source| HookError::Read {
|
||||
kind: "checkpoint",
|
||||
source,
|
||||
})?;
|
||||
let head_sha = stored
|
||||
.into_iter()
|
||||
.rev()
|
||||
.find_map(|stored| match stored.record {
|
||||
PlatformRecord::Checkpoint(record)
|
||||
if record.workspace.as_deref() == Some(&workspace) =>
|
||||
{
|
||||
record.git_commit_sha
|
||||
}
|
||||
_ => None,
|
||||
})
|
||||
.ok_or_else(|| HookError::Unresumable {
|
||||
reason: "the run branch has no checkpoint".to_string(),
|
||||
})?;
|
||||
let diff = self
|
||||
.workspaces
|
||||
.diff(&workspace, Some(&base_sha), &head_sha)
|
||||
.diff(&site, Some(&base_sha), &head_sha)
|
||||
.await
|
||||
.map_err(|source| HookError::Diff {
|
||||
what: "run's diff",
|
||||
source,
|
||||
})?;
|
||||
let patch_blob = self.patch_blob(&diff).await?;
|
||||
let publication = branch
|
||||
.run_branch
|
||||
.clone()
|
||||
.filter(|_| self.publisher.is_some())
|
||||
.map(|run_branch| Publication {
|
||||
run_branch,
|
||||
head_sha: head_sha.clone(),
|
||||
site,
|
||||
patch: diff.patch,
|
||||
});
|
||||
let record = PlatformRecord::RunDiff(RunDiffRecord {
|
||||
base_sha: Some(base_sha),
|
||||
head_sha: Some(head_sha),
|
||||
|
|
@ -1155,7 +1400,24 @@ impl FabroHooks {
|
|||
deletions = diff.summary.deletions,
|
||||
"run diff recorded"
|
||||
);
|
||||
Ok(())
|
||||
Ok(publication)
|
||||
}
|
||||
|
||||
/// Hand a successful run's work to the publisher, before the run's
|
||||
/// terminal record. A failure fails the run with its reason.
|
||||
async fn publish(&self, publisher: &dyn RunPublisher, publication: &Publication) {
|
||||
match publisher.publish(publication).await {
|
||||
Ok(()) => info!(
|
||||
run_id = %self.run_id,
|
||||
branch = publication.run_branch,
|
||||
sha = publication.head_sha,
|
||||
"run published"
|
||||
),
|
||||
Err(message) => {
|
||||
warn!(run_id = %self.run_id, error = %message, "the run's publication failed");
|
||||
*sync::lock(&self.publish_failure) = Some(message);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Hold at a test gate when one is set for this point and node.
|
||||
|
|
@ -1371,8 +1633,27 @@ impl ExecutionHooks for FabroHooks {
|
|||
failure = finished.failure.as_deref().unwrap_or(""),
|
||||
"Petri run finished; recording the run's diff and running the run-end hooks"
|
||||
);
|
||||
if let Err(error) = self.record_run_diff().await {
|
||||
warn!(run_id = %self.run_id, error = %error.render(), "the run's diff was not recorded");
|
||||
let publication = match self.record_run_diff().await {
|
||||
Ok(publication) => publication,
|
||||
Err(error) => {
|
||||
let message = error.render();
|
||||
warn!(run_id = %self.run_id, error = %message, "the run's diff was not recorded");
|
||||
if self.publisher.is_some() && finished.status == RunStatus::Success {
|
||||
*sync::lock(&self.publish_failure) = Some(message);
|
||||
}
|
||||
None
|
||||
}
|
||||
};
|
||||
if let Some(publisher) = &self.publisher {
|
||||
if finished.status == RunStatus::Success && self.checkpoint_failure().is_none() {
|
||||
if let Some(publication) = &publication {
|
||||
self.publish(publisher.as_ref(), publication).await;
|
||||
} else if self.publish_failure().is_none() {
|
||||
*sync::lock(&self.publish_failure) = Some(
|
||||
"the run has no recorded branch and checkpoint to publish".to_string(),
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
self.inner.run_finished(context, finished).await
|
||||
}
|
||||
|
|
@ -1402,17 +1683,18 @@ impl ExecutionHooks for FabroHooks {
|
|||
acquired.scope,
|
||||
(workspace.clone(), Arc::clone(&acquired.env)),
|
||||
);
|
||||
if !self.resumed {
|
||||
return Ok(());
|
||||
}
|
||||
let site = if self.host_workspaces {
|
||||
self.workspaces.host(&workspace)
|
||||
} else {
|
||||
Site::Sandbox(Arc::clone(&acquired.env))
|
||||
};
|
||||
self.restore(&workspace, &site)
|
||||
.await
|
||||
.map_err(|error| ScopeAcquiredError::new(error.render()))
|
||||
sync::lock(&self.sites).insert(workspace.clone(), site.clone());
|
||||
if self.resumed {
|
||||
self.restore(&workspace, &site)
|
||||
.await
|
||||
.map_err(|error| ScopeAcquiredError::new(error.render()))?;
|
||||
}
|
||||
self.check_out_source(&workspace, &site).await
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -1525,6 +1807,8 @@ mod tests {
|
|||
artifacts: vec!["assets/**".to_string()],
|
||||
test_gates: None,
|
||||
artifact_writer,
|
||||
source: None,
|
||||
publisher: None,
|
||||
},
|
||||
Arc::new(NoHooks),
|
||||
run_id,
|
||||
|
|
|
|||
|
|
@ -41,12 +41,11 @@
|
|||
//! `prepare_result` and its platform record in `transition`, around Petri's
|
||||
//! own hook service for `[[run.hooks]]`;
|
||||
//! - [`checkpoint`]: the Git snapshots of a run's workspaces, on the host or
|
||||
//! inside a Docker or Daytona sandbox, and the snapshot repository they are
|
||||
//! published to;
|
||||
//! - [`recovery`]: the resume-on-restart protocol, which brings every live
|
||||
//! workspace to the snapshot its durable state names: a host workspace before
|
||||
//! the run goes back to a worker, a sandbox workspace in the worker when its
|
||||
//! scope is acquired;
|
||||
//! inside a Docker or Daytona sandbox, with pushes from that same workspace;
|
||||
//! - [`source`]: where a GitHub target's workspace is checked out from, the
|
||||
//! revision, depth and read credential its in-sandbox fetch uses;
|
||||
//! - [`recovery`]: recovery planning from execution records; the worker resets
|
||||
//! surviving Git workspaces when their scopes are acquired;
|
||||
//! - [`platform_records`]: Fabro's platform records as the adapters reach them,
|
||||
//! in the server's database or over its API from a worker;
|
||||
//! - [`host_tools`]: Fabro's run tools on every native agent session of a run,
|
||||
|
|
@ -54,9 +53,8 @@
|
|||
//! - [`controls`]: the controls Fabro drives on a live run (pause, unpause,
|
||||
//! steer, cancel), over Petri's control service;
|
||||
//! - [`fork`]: a run seeded from another's records up to a checkpoint's
|
||||
//! position, over Petri's `host::fork_from`, with the kept checkpoints, their
|
||||
//! snapshots and the run branch carried over: what rewind, fork and retry are
|
||||
//! built on;
|
||||
//! position, over Petri's `host::fork_from`, with checkpoint metadata and the
|
||||
//! run branch carried over; the new sandbox fetches the published code;
|
||||
//! - [`prune`]: a run's sandboxes deleted through Petri's lease ledger, as
|
||||
//! `petri sandbox prune` deletes them, when Fabro deletes the run.
|
||||
//!
|
||||
|
|
@ -86,6 +84,7 @@ pub mod run_graph;
|
|||
pub mod run_store;
|
||||
pub mod runtime;
|
||||
pub mod secrets;
|
||||
pub mod source;
|
||||
#[cfg(feature = "test-support")]
|
||||
pub mod test_support;
|
||||
pub mod workspace;
|
||||
|
|
|
|||
|
|
@ -1,54 +1,21 @@
|
|||
//! Resume on restart: the recovery protocol whose rule is that the
|
||||
//! workspace a resumed stage sees matches Petri's durable execution state
|
||||
//! (the integration plan's F3.5).
|
||||
//!
|
||||
//! For a Petri run the server finds in flight at startup, once the previous
|
||||
//! worker's lease is released, [`recover`] reads the durable execution
|
||||
//! state through `inspect_run` and decides:
|
||||
//!
|
||||
//! - a run with a `checkpoint_failed` finish anywhere is reported failed and
|
||||
//! not resumed: a failed checkpoint cancelled it, and nothing of it is
|
||||
//! reconciled;
|
||||
//! - otherwise, for every live execution, the last durable finish names the
|
||||
//! snapshot its workspace must sit on: the checkpoint record's commit, or,
|
||||
//! when the record was lost to the crash, the commit found by its key in the
|
||||
//! workspace's snapshot repository or history, which is then recorded again;
|
||||
//! - a workspace that survives is verified to sit on that commit, unchanged, or
|
||||
//! reset to it; a workspace that is gone, or a fresh one with no history (a
|
||||
//! fork's first acquisition), is restored from the run's snapshot repository;
|
||||
//! - a durable finish with no snapshot fails the run with a named error rather
|
||||
//! than resume it on stale files.
|
||||
//!
|
||||
//! Every child invocation's scope has its own snapshots, keyed by
|
||||
//! execution; a nested invocation that inherits its caller's sandbox shares
|
||||
//! the caller's workspace, and the workspace is brought to the newest of
|
||||
//! the live executions' snapshots on it.
|
||||
//!
|
||||
//! The decision is [`plan`], over the records and the snapshot repository
|
||||
//! alone, both on this host whatever the provider. Applying it differs: a
|
||||
//! host workspace is brought to its snapshot here, before the worker is
|
||||
//! relaunched; a Docker or Daytona workspace lives inside a sandbox only
|
||||
//! the worker's run reaches, so its target is deferred, and the worker's
|
||||
//! hooks read the same plan and apply it through the scope's environment
|
||||
//! at `scope_acquired`, before the first attempt runs there
|
||||
//! ([`bring_to`]).
|
||||
//! Resume execution from durable records using the surviving workspace.
|
||||
//! Git checkpoints are reset in place when present. Missing workspaces are
|
||||
//! never reconstructed here; explicit forks fetch their source run on GitHub
|
||||
//! when the worker acquires the new sandbox. Runs without Git checkpoints
|
||||
//! retain execution metadata without gaining workspace backup guarantees.
|
||||
|
||||
use std::collections::BTreeMap;
|
||||
use std::path::PathBuf;
|
||||
use std::sync::Arc;
|
||||
|
||||
use fabro_store::platform_records::CheckpointRecord;
|
||||
use fabro_store::{PlatformRecord, PlatformRecordKind, StagePosition};
|
||||
use fabro_store::{PlatformRecord, PlatformRecordKind};
|
||||
use fabro_types::RunId;
|
||||
use fabro_types::settings::run::RunNamespace;
|
||||
use petri_execution::host::{self, HostError};
|
||||
use petri_execution::inspect::{self, ExecutionInspection, InspectError};
|
||||
use petri_execution::{Access, InvocationId, RunKey, RunStore};
|
||||
use petri_store::StoreError;
|
||||
use tracing::info;
|
||||
|
||||
use crate::checkpoint::{
|
||||
CHECKPOINT_FAILED_CLASS, CheckpointError, CheckpointKey, RunGitSettings, RunWorkspaces, Site,
|
||||
CHECKPOINT_FAILED_CLASS, CheckpointError, CheckpointKey, RunWorkspaces, Site,
|
||||
};
|
||||
use crate::platform_records::{PlatformRecordError, PlatformRecords};
|
||||
use crate::workspace::{WorkspaceLookup, WorkspaceLookupError};
|
||||
|
|
@ -57,11 +24,8 @@ use crate::workspace::{WorkspaceLookup, WorkspaceLookupError};
|
|||
/// and its Git settings.
|
||||
pub struct RecoveryRequest {
|
||||
pub run_id: RunId,
|
||||
/// The run directory Petri ran under (the run's `petri` scratch).
|
||||
pub run_dir: PathBuf,
|
||||
pub store: Arc<dyn RunStore>,
|
||||
pub records: Arc<dyn PlatformRecords>,
|
||||
pub git: RunGitSettings,
|
||||
}
|
||||
|
||||
impl RecoveryRequest {
|
||||
|
|
@ -69,17 +33,13 @@ impl RecoveryRequest {
|
|||
#[must_use]
|
||||
pub fn for_run(
|
||||
run_id: RunId,
|
||||
run_dir: PathBuf,
|
||||
store: Arc<dyn RunStore>,
|
||||
records: Arc<dyn PlatformRecords>,
|
||||
settings: &RunNamespace,
|
||||
) -> Self {
|
||||
Self {
|
||||
run_id,
|
||||
run_dir,
|
||||
store,
|
||||
records,
|
||||
git: RunGitSettings::from(settings),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -91,8 +51,6 @@ pub enum WorkspaceAction {
|
|||
Verified,
|
||||
/// It was brought back to the snapshot.
|
||||
Reset,
|
||||
/// It was gone and was recreated from the snapshot repository.
|
||||
Restored,
|
||||
/// It lives in a sandbox this process does not reach: the worker's
|
||||
/// hooks bring it to the snapshot when its scope is acquired.
|
||||
Deferred,
|
||||
|
|
@ -112,7 +70,7 @@ pub enum Plan {
|
|||
Start,
|
||||
/// The run continues; each live workspace, by id, and its snapshot.
|
||||
Resume {
|
||||
targets: BTreeMap<String, RestoreTarget>,
|
||||
targets: BTreeMap<String, Vec<RestoreTarget>>,
|
||||
},
|
||||
/// The run cannot continue and is reported failed.
|
||||
Failed { reason: String },
|
||||
|
|
@ -162,13 +120,11 @@ pub enum RecoveryError {
|
|||
}
|
||||
|
||||
/// Decide how the run continues: the snapshot every live workspace must sit
|
||||
/// on, from the records and the snapshot repository, with a lost record
|
||||
/// reconciled from the repository. Nothing is touched.
|
||||
/// on, from execution and checkpoint records. Nothing is touched.
|
||||
pub async fn plan(
|
||||
store: Arc<dyn RunStore>,
|
||||
records: &dyn PlatformRecords,
|
||||
run_id: &RunId,
|
||||
workspaces: &RunWorkspaces,
|
||||
) -> Result<Plan, RecoveryError> {
|
||||
let key = RunKey::new(run_id.to_string());
|
||||
let logs = match store.open(&key, Access::Read).await {
|
||||
|
|
@ -204,29 +160,19 @@ pub async fn plan(
|
|||
return Ok(Plan::Failed { reason: failed });
|
||||
}
|
||||
let planner = Planner {
|
||||
records,
|
||||
run_id,
|
||||
workspaces,
|
||||
lookup: WorkspaceLookup::new(store, key),
|
||||
lookup: WorkspaceLookup::new(store, key),
|
||||
recorded: recorded_checkpoints(records, run_id).await?,
|
||||
};
|
||||
planner.targets(&inspection.executions).await
|
||||
}
|
||||
|
||||
/// Decide how the run continues, and bring its host workspaces to their
|
||||
/// snapshots; a sandbox workspace's target is deferred to the worker.
|
||||
/// Decide how the run continues. All Git work is deferred to the worker
|
||||
/// using the acquired workspace, including for the local provider.
|
||||
pub async fn recover(request: RecoveryRequest) -> Result<Recovery, RecoveryError> {
|
||||
let workspaces = RunWorkspaces::new(
|
||||
request.run_dir.clone(),
|
||||
request.run_id.to_string(),
|
||||
request.git.author.clone(),
|
||||
&request.git.checkpoint,
|
||||
);
|
||||
let targets = match plan(
|
||||
Arc::clone(&request.store),
|
||||
&*request.records,
|
||||
&request.run_id,
|
||||
&workspaces,
|
||||
)
|
||||
.await?
|
||||
{
|
||||
|
|
@ -235,62 +181,34 @@ pub async fn recover(request: RecoveryRequest) -> Result<Recovery, RecoveryError
|
|||
Plan::Resume { targets } => targets,
|
||||
};
|
||||
|
||||
let mut recovered = Vec::new();
|
||||
for (workspace, target) in targets {
|
||||
let action = if request.git.host_workspaces {
|
||||
bring_to(
|
||||
&workspaces,
|
||||
&workspaces.host(&workspace),
|
||||
&workspace,
|
||||
&target,
|
||||
)
|
||||
.await?
|
||||
} else {
|
||||
WorkspaceAction::Deferred
|
||||
};
|
||||
info!(
|
||||
run_id = %request.run_id,
|
||||
workspace,
|
||||
sha = target.sha,
|
||||
action = ?action,
|
||||
"workspace's durable snapshot decided"
|
||||
);
|
||||
recovered.push(RecoveredWorkspace {
|
||||
workspace,
|
||||
sha: target.sha,
|
||||
action,
|
||||
});
|
||||
}
|
||||
let recovered = targets
|
||||
.into_iter()
|
||||
.flat_map(|(workspace, targets)| {
|
||||
targets.into_iter().map(move |target| RecoveredWorkspace {
|
||||
workspace: workspace.clone(),
|
||||
sha: target.sha,
|
||||
action: WorkspaceAction::Deferred,
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
Ok(Recovery::Resume {
|
||||
workspaces: recovered,
|
||||
})
|
||||
}
|
||||
|
||||
/// The snapshot one live execution's last durable finish names on a
|
||||
/// workspace.
|
||||
struct Candidate {
|
||||
key: CheckpointKey,
|
||||
sha: String,
|
||||
/// The decision over one run's durable records.
|
||||
struct Planner {
|
||||
lookup: WorkspaceLookup,
|
||||
recorded: BTreeMap<CheckpointKey, (Option<String>, String)>,
|
||||
}
|
||||
|
||||
/// The decision over one run's records and snapshot repository.
|
||||
struct Planner<'a> {
|
||||
records: &'a dyn PlatformRecords,
|
||||
run_id: &'a RunId,
|
||||
workspaces: &'a RunWorkspaces,
|
||||
lookup: WorkspaceLookup,
|
||||
/// The run's checkpoint records by key: the workspace they name and
|
||||
/// the commit.
|
||||
recorded: BTreeMap<CheckpointKey, (Option<String>, String)>,
|
||||
}
|
||||
|
||||
impl Planner<'_> {
|
||||
impl Planner {
|
||||
/// The snapshot each live execution's workspace must sit on. A live
|
||||
/// execution is one whose log records no exit: `inspect_run` reports
|
||||
/// it as incomplete. A workspace several live executions share is
|
||||
/// brought to the newest of their snapshots.
|
||||
async fn targets(&self, executions: &[ExecutionInspection]) -> Result<Plan, RecoveryError> {
|
||||
let mut candidates: BTreeMap<String, Vec<Candidate>> = BTreeMap::new();
|
||||
let mut candidates: BTreeMap<String, Vec<RestoreTarget>> = BTreeMap::new();
|
||||
for execution in executions
|
||||
.iter()
|
||||
.filter(|execution| execution.status == "incomplete")
|
||||
|
|
@ -306,134 +224,26 @@ impl Planner<'_> {
|
|||
if owned.is_empty() {
|
||||
continue;
|
||||
}
|
||||
let mut found = false;
|
||||
for workspace in owned {
|
||||
if let Some(sha) = self.snapshot_of(&workspace, key).await? {
|
||||
found = true;
|
||||
candidates
|
||||
.entry(workspace)
|
||||
.or_default()
|
||||
.push(Candidate { key, sha });
|
||||
if let Some((recorded_workspace, sha)) = self.recorded.get(&key) {
|
||||
if recorded_workspace
|
||||
.as_deref()
|
||||
.is_none_or(|id| id == workspace)
|
||||
{
|
||||
candidates
|
||||
.entry(workspace)
|
||||
.or_default()
|
||||
.push(RestoreTarget {
|
||||
key,
|
||||
sha: sha.clone(),
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
return Ok(Plan::Failed {
|
||||
reason: format!(
|
||||
"no checkpoint snapshot exists for the last durable finish of {key}; the \
|
||||
run cannot resume on stale files"
|
||||
),
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
let mut targets = BTreeMap::new();
|
||||
for (workspace, candidates) in candidates {
|
||||
let target = self.newest(&workspace, &candidates).await?;
|
||||
targets.insert(workspace, target);
|
||||
}
|
||||
Ok(Plan::Resume { targets })
|
||||
}
|
||||
|
||||
/// The snapshot of `key` in `workspace`: the commit its record names,
|
||||
/// when the record names this workspace or none; else the commit found
|
||||
/// by its key in the workspace's snapshot repository or history, which
|
||||
/// is then recorded again for the record the crash lost. `None` when no
|
||||
/// snapshot exists.
|
||||
async fn snapshot_of(
|
||||
&self,
|
||||
workspace: &str,
|
||||
key: CheckpointKey,
|
||||
) -> Result<Option<String>, RecoveryError> {
|
||||
if let Some((recorded_workspace, sha)) = self.recorded.get(&key) {
|
||||
if recorded_workspace
|
||||
.as_deref()
|
||||
.is_none_or(|recorded| recorded == workspace)
|
||||
{
|
||||
return Ok(Some(sha.clone()));
|
||||
}
|
||||
}
|
||||
let found = self
|
||||
.workspaces
|
||||
.find(workspace, key)
|
||||
.await
|
||||
.map_err(|source| RecoveryError::Workspace {
|
||||
workspace: workspace.to_string(),
|
||||
source,
|
||||
})?;
|
||||
if let Some(sha) = &found {
|
||||
self.reconcile_record(key, workspace, sha).await?;
|
||||
}
|
||||
Ok(found)
|
||||
}
|
||||
|
||||
/// Write the record a crash lost, from the commit found by its key.
|
||||
async fn reconcile_record(
|
||||
&self,
|
||||
key: CheckpointKey,
|
||||
workspace: &str,
|
||||
sha: &str,
|
||||
) -> Result<(), RecoveryError> {
|
||||
info!(
|
||||
run_id = %self.run_id,
|
||||
execution = key.execution,
|
||||
firing = key.firing,
|
||||
attempt = key.attempt,
|
||||
sha,
|
||||
"checkpoint record reconciled from the run branch"
|
||||
);
|
||||
let record = PlatformRecord::Checkpoint(CheckpointRecord {
|
||||
execution: key.execution,
|
||||
firing: key.firing,
|
||||
attempt: Some(key.attempt),
|
||||
workspace: Some(workspace.to_string()),
|
||||
git_commit_sha: Some(sha.to_string()),
|
||||
diff_summary: None,
|
||||
patch_blob: None,
|
||||
operation: Some(key.operation()),
|
||||
});
|
||||
self.records
|
||||
.append(
|
||||
self.run_id,
|
||||
&record,
|
||||
Some(StagePosition {
|
||||
execution: key.execution,
|
||||
firing: key.firing,
|
||||
}),
|
||||
)
|
||||
.await
|
||||
.map_err(RecoveryError::Records)?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Of the snapshots live executions name on one workspace, the one
|
||||
/// every other descends from, else the last named. Two executions that
|
||||
/// name the same commit share it under the first one's key.
|
||||
async fn newest(
|
||||
&self,
|
||||
workspace: &str,
|
||||
candidates: &[Candidate],
|
||||
) -> Result<RestoreTarget, RecoveryError> {
|
||||
let mut chosen = &candidates[0];
|
||||
for candidate in &candidates[1..] {
|
||||
if self
|
||||
.workspaces
|
||||
.is_ancestor(workspace, &chosen.sha, &candidate.sha)
|
||||
.await
|
||||
.map_err(|source| RecoveryError::Workspace {
|
||||
workspace: workspace.to_string(),
|
||||
source,
|
||||
})?
|
||||
{
|
||||
chosen = candidate;
|
||||
}
|
||||
}
|
||||
let chosen = candidates
|
||||
.iter()
|
||||
.find(|candidate| candidate.sha == chosen.sha)
|
||||
.unwrap_or(chosen);
|
||||
Ok(RestoreTarget {
|
||||
key: chosen.key,
|
||||
sha: chosen.sha.clone(),
|
||||
Ok(Plan::Resume {
|
||||
targets: candidates,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
|
@ -500,11 +310,8 @@ async fn recorded_checkpoints(
|
|||
Ok(recorded)
|
||||
}
|
||||
|
||||
/// Verify, reset or restore the workspace at `site` onto its target: a
|
||||
/// workspace that still holds the commit is verified or reset in place; a
|
||||
/// gone one, a fresh directory with no history (a fork's first
|
||||
/// acquisition), or a sandbox whose repository lost the commit, is restored
|
||||
/// from the snapshot repository (a bundle of the snapshot, into a sandbox).
|
||||
/// Verify or reset an existing workspace. A missing commit is an error;
|
||||
/// this operation never reconstructs a lost workspace.
|
||||
pub async fn bring_to(
|
||||
workspaces: &RunWorkspaces,
|
||||
site: &Site,
|
||||
|
|
@ -530,9 +337,7 @@ pub async fn bring_to(
|
|||
workspaces.reset(site, &target.sha).await.map_err(failed)?;
|
||||
return Ok(WorkspaceAction::Reset);
|
||||
}
|
||||
workspaces
|
||||
.restore(site, workspace, target.key, &target.sha)
|
||||
.await
|
||||
.map_err(failed)?;
|
||||
Ok(WorkspaceAction::Restored)
|
||||
Err(failed(CheckpointError::MissingCommit {
|
||||
sha: target.sha.clone(),
|
||||
}))
|
||||
}
|
||||
|
|
|
|||
282
lib/components/fabro-petri/src/source.rs
Normal file
282
lib/components/fabro-petri/src/source.rs
Normal file
|
|
@ -0,0 +1,282 @@
|
|||
//! Where a run's repository comes from when it lives on GitHub: the origin,
|
||||
//! the revision the run starts from, the working branch, the history depth,
|
||||
//! and the read credential the fetch presents.
|
||||
//!
|
||||
//! A Git target's workspace is checked out by Fabro, not by Petri: when a
|
||||
//! scope's environment is acquired for a fresh run, the hooks fetch the
|
||||
//! revision into the workspace from inside the scope, so the files belong to
|
||||
//! the user every later command runs as and nothing is copied in from this
|
||||
//! host (see [`crate::hooks`]). Petri's own `start` checkout is not used
|
||||
//! for a Git target: the run binds no repository for it.
|
||||
//!
|
||||
//! The credential reaches one `git` command at a time through its
|
||||
//! environment, as an HTTP header scoped to the origin. It is never written
|
||||
//! into the repository, its configuration, or its remote URL. Each fetch asks
|
||||
//! the run's [`SourceCredentials`] for it, so a workspace first acquired hours
|
||||
//! into a run still presents a live token.
|
||||
|
||||
use std::fmt;
|
||||
use std::sync::Arc;
|
||||
|
||||
use fabro_types::settings::run::{RunCloneSettings, RunMode, RunNamespace};
|
||||
use fabro_types::{GitRunTarget, RunTarget};
|
||||
|
||||
/// The revision a run starts from, in the order the target names it: an
|
||||
/// exact commit wins over a tag, and a tag over the branch head.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub enum SourceRevision {
|
||||
Commit(String),
|
||||
Tag(String),
|
||||
Branch(String),
|
||||
}
|
||||
|
||||
impl SourceRevision {
|
||||
/// What `git fetch` asks the origin for. A branch and a tag may share a
|
||||
/// name, so both are qualified.
|
||||
#[must_use]
|
||||
pub fn refspec(&self) -> String {
|
||||
match self {
|
||||
Self::Commit(sha) => sha.clone(),
|
||||
Self::Tag(tag) => format!("refs/tags/{tag}"),
|
||||
Self::Branch(branch) => format!("refs/heads/{branch}"),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The HTTP basic credential a fetch from the origin presents: the
|
||||
/// base64 of `username:password`, as the server resolved it for the run's
|
||||
/// repository.
|
||||
#[derive(Clone, PartialEq, Eq)]
|
||||
pub struct SourceCredential(String);
|
||||
|
||||
impl SourceCredential {
|
||||
/// A credential from its base64 `username:password` encoding; `None`
|
||||
/// for an empty one.
|
||||
#[must_use]
|
||||
pub fn from_encoded(encoded: impl Into<String>) -> Option<Self> {
|
||||
let encoded = encoded.into();
|
||||
let encoded = encoded.trim();
|
||||
(!encoded.is_empty()).then(|| Self(encoded.to_string()))
|
||||
}
|
||||
|
||||
/// The encoded value, for handing to a process that fetches.
|
||||
#[must_use]
|
||||
pub fn encoded(&self) -> &str {
|
||||
&self.0
|
||||
}
|
||||
|
||||
/// The environment that has `git` present this credential as an
|
||||
/// `Authorization` header to `url` alone.
|
||||
#[must_use]
|
||||
pub fn header_env(&self, url: &str) -> Vec<(String, String)> {
|
||||
vec![
|
||||
("GIT_CONFIG_COUNT".to_string(), "1".to_string()),
|
||||
(
|
||||
"GIT_CONFIG_KEY_0".to_string(),
|
||||
format!("http.{url}.extraheader"),
|
||||
),
|
||||
(
|
||||
"GIT_CONFIG_VALUE_0".to_string(),
|
||||
format!("AUTHORIZATION: basic {}", self.0),
|
||||
),
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
impl fmt::Debug for SourceCredential {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
f.write_str("SourceCredential(<redacted>)")
|
||||
}
|
||||
}
|
||||
|
||||
/// Where a run's fetches get their credential, asked once per fetch.
|
||||
#[async_trait::async_trait]
|
||||
pub trait SourceCredentials: Send + Sync {
|
||||
/// The credential the next fetch presents; `None` fetches anonymously.
|
||||
async fn credential(&self) -> Option<SourceCredential>;
|
||||
}
|
||||
|
||||
/// A fixed credential, presented by every fetch.
|
||||
#[async_trait::async_trait]
|
||||
impl SourceCredentials for SourceCredential {
|
||||
async fn credential(&self) -> Option<Self> {
|
||||
Some(self.clone())
|
||||
}
|
||||
}
|
||||
|
||||
/// A run's GitHub repository as its workspaces check it out.
|
||||
#[derive(Clone)]
|
||||
pub struct RunSource {
|
||||
/// The repository's HTTPS URL: the workspace's `origin`.
|
||||
pub origin: String,
|
||||
pub revision: SourceRevision,
|
||||
/// The branch the workspace stands on before the run branch is created
|
||||
/// from it.
|
||||
pub branch: String,
|
||||
/// Commits of history to fetch; `None` is the whole history.
|
||||
pub depth: Option<u32>,
|
||||
pub credentials: Option<Arc<dyn SourceCredentials>>,
|
||||
}
|
||||
|
||||
impl fmt::Debug for RunSource {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
f.debug_struct("RunSource")
|
||||
.field("origin", &self.origin)
|
||||
.field("revision", &self.revision)
|
||||
.field("branch", &self.branch)
|
||||
.field("depth", &self.depth)
|
||||
.field(
|
||||
"credentials",
|
||||
&self.credentials.as_ref().map(|_| "<redacted>"),
|
||||
)
|
||||
.finish()
|
||||
}
|
||||
}
|
||||
|
||||
impl RunSource {
|
||||
/// The source a Git target gives under the run's clone settings, or
|
||||
/// `None` when the run checks nothing out.
|
||||
#[must_use]
|
||||
pub fn for_target(
|
||||
target: &GitRunTarget,
|
||||
origin: String,
|
||||
clone: &RunCloneSettings,
|
||||
credentials: Option<Arc<dyn SourceCredentials>>,
|
||||
) -> Option<Self> {
|
||||
if !clone.enabled {
|
||||
return None;
|
||||
}
|
||||
let revision = if let Some(sha) = target.sha.clone().filter(|sha| !sha.is_empty()) {
|
||||
SourceRevision::Commit(sha)
|
||||
} else if let Some(tag) = target.tag.clone().filter(|tag| !tag.is_empty()) {
|
||||
SourceRevision::Tag(tag)
|
||||
} else {
|
||||
SourceRevision::Branch(target.branch.clone())
|
||||
};
|
||||
Some(Self {
|
||||
origin,
|
||||
revision,
|
||||
branch: target.branch.clone(),
|
||||
depth: clone
|
||||
.depth_limit()
|
||||
.and_then(|depth| u32::try_from(depth).ok()),
|
||||
credentials,
|
||||
})
|
||||
}
|
||||
|
||||
/// The source of a run whose target is a GitHub repository, under its
|
||||
/// settings: `None` for any other target, a dry run, or a run whose
|
||||
/// clone is disabled. A target that does not name a valid GitHub
|
||||
/// repository checks nothing out.
|
||||
#[must_use]
|
||||
pub fn for_run(
|
||||
target: Option<&RunTarget>,
|
||||
settings: &RunNamespace,
|
||||
credentials: Option<Arc<dyn SourceCredentials>>,
|
||||
) -> Option<Self> {
|
||||
let Some(RunTarget::Git(target)) = target else {
|
||||
return None;
|
||||
};
|
||||
if settings.execution.mode == RunMode::DryRun {
|
||||
return None;
|
||||
}
|
||||
let validated = target.clone().validate().ok()?;
|
||||
let origin = validated.repository().https_url();
|
||||
Self::for_target(validated.target(), origin, &settings.clone, credentials)
|
||||
}
|
||||
|
||||
/// The environment a `git` command that talks to the origin runs with:
|
||||
/// the credential as an `Authorization` header for the origin alone.
|
||||
pub async fn fetch_env(&self) -> Vec<(String, String)> {
|
||||
let Some(credentials) = &self.credentials else {
|
||||
return Vec::new();
|
||||
};
|
||||
credentials
|
||||
.credential()
|
||||
.await
|
||||
.map(|credential| credential.header_env(&self.origin))
|
||||
.unwrap_or_default()
|
||||
}
|
||||
|
||||
/// The `--depth` argument of a fetch, when the history is limited.
|
||||
#[must_use]
|
||||
pub fn depth_arg(&self) -> Option<String> {
|
||||
self.depth.map(|depth| format!("--depth={depth}"))
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn target(tag: Option<&str>, sha: Option<&str>) -> GitRunTarget {
|
||||
GitRunTarget {
|
||||
repo: "acme/widgets".to_string(),
|
||||
branch: "main".to_string(),
|
||||
tag: tag.map(str::to_string),
|
||||
sha: sha.map(str::to_string),
|
||||
}
|
||||
}
|
||||
|
||||
fn clone(enabled: bool, depth: i32) -> RunCloneSettings {
|
||||
RunCloneSettings { enabled, depth }
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_commit_wins_over_a_tag_and_a_tag_over_the_branch() {
|
||||
let origin = "https://github.com/acme/widgets".to_string();
|
||||
let settings = clone(true, 100);
|
||||
let pick = |tag, sha| {
|
||||
RunSource::for_target(&target(tag, sha), origin.clone(), &settings, None)
|
||||
.unwrap()
|
||||
.revision
|
||||
};
|
||||
assert_eq!(
|
||||
pick(Some("v1"), Some("abc")),
|
||||
SourceRevision::Commit("abc".to_string())
|
||||
);
|
||||
assert_eq!(
|
||||
pick(Some("v1"), None),
|
||||
SourceRevision::Tag("v1".to_string())
|
||||
);
|
||||
assert_eq!(pick(None, None), SourceRevision::Branch("main".to_string()));
|
||||
assert_eq!(pick(None, None).refspec(), "refs/heads/main");
|
||||
assert_eq!(pick(Some("v1"), None).refspec(), "refs/tags/v1");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_disabled_clone_checks_nothing_out_and_depth_zero_is_full_history() {
|
||||
let origin = "https://github.com/acme/widgets".to_string();
|
||||
assert!(
|
||||
RunSource::for_target(&target(None, None), origin.clone(), &clone(false, 1), None)
|
||||
.is_none()
|
||||
);
|
||||
let full =
|
||||
RunSource::for_target(&target(None, None), origin, &clone(true, 0), None).unwrap();
|
||||
assert_eq!(full.depth, None);
|
||||
assert_eq!(full.depth_arg(), None);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn the_credential_is_scoped_to_the_origin_and_never_printed() {
|
||||
let credential = SourceCredential::from_encoded("c2VjcmV0").unwrap();
|
||||
let source = RunSource::for_target(
|
||||
&target(None, None),
|
||||
"https://github.com/acme/widgets".to_string(),
|
||||
&clone(true, 1),
|
||||
Some(Arc::new(credential)),
|
||||
)
|
||||
.unwrap();
|
||||
let env = source.fetch_env().await;
|
||||
assert!(env.contains(&(
|
||||
"GIT_CONFIG_KEY_0".to_string(),
|
||||
"http.https://github.com/acme/widgets.extraheader".to_string()
|
||||
)));
|
||||
assert!(env.contains(&(
|
||||
"GIT_CONFIG_VALUE_0".to_string(),
|
||||
"AUTHORIZATION: basic c2VjcmV0".to_string()
|
||||
)));
|
||||
assert!(!format!("{source:?}").contains("c2VjcmV0"));
|
||||
assert!(SourceCredential::from_encoded(" ").is_none());
|
||||
}
|
||||
}
|
||||
|
|
@ -23,18 +23,23 @@ use fabro_petri::artifacts::StoreArtifactWriter;
|
|||
use fabro_petri::blobs::Blobs;
|
||||
use fabro_petri::check::{self, Bundle, CheckRequest, Launch};
|
||||
use fabro_petri::checkpoint::{
|
||||
CHECKPOINT_FAILED_CLASS, CheckpointKey, RunGitSettings, RunWorkspaces,
|
||||
CHECKPOINT_FAILED_CLASS, CheckpointKey, RunGitSettings, RunWorkspaces, Site,
|
||||
};
|
||||
use fabro_petri::controls::RunControls;
|
||||
use fabro_petri::engine::{self, Execution, RunRequest, RunStatus};
|
||||
use fabro_petri::hooks::HooksSpec;
|
||||
use fabro_petri::platform_records::PlatformRecords;
|
||||
use fabro_petri::fork;
|
||||
use fabro_petri::hooks::{HooksSpec, Publication, RunPublisher};
|
||||
use fabro_petri::platform_records::{PlatformRecordError, PlatformRecords};
|
||||
use fabro_petri::providers::{DaytonaCredentials, SandboxProviderConfig};
|
||||
use fabro_petri::prune::{self, PruneRequest};
|
||||
use fabro_petri::recovery::{self, Recovery, RecoveryRequest};
|
||||
use fabro_petri::runtime::RuntimeSpec;
|
||||
use fabro_petri::source::{RunSource, SourceRevision};
|
||||
use fabro_petri::test_support::{MemoryBlobs, MemoryPlatformRecords};
|
||||
use fabro_store::{ArtifactStore, PlatformRecord, PlatformRecordKind};
|
||||
use fabro_types::settings::run::{EnvironmentResourcesSettings, RunCheckpointSettings};
|
||||
use fabro_types::settings::run::{
|
||||
EnvironmentResourcesSettings, RunCheckpointSettings, RunNamespace,
|
||||
};
|
||||
use fabro_types::{GitIdentitySource, RunId, SandboxProviderKind};
|
||||
use object_store::local::LocalFileSystem;
|
||||
use petri_execution::inspect::{self, RunInspection};
|
||||
|
|
@ -90,11 +95,19 @@ struct Harness {
|
|||
/// The `[run.artifacts] include` patterns the hooks collect under.
|
||||
artifacts: Vec<String>,
|
||||
artifact_store: ArtifactStore,
|
||||
/// Where the run's workspaces are checked out from.
|
||||
source: Option<RunSource>,
|
||||
/// Treat the local provider's workspaces as a sandbox's: `git` runs
|
||||
/// through the scope's environment and checkpoints leave as bundles.
|
||||
sandboxed: bool,
|
||||
/// What a successful run's work does when it ends.
|
||||
publisher: Option<Arc<dyn RunPublisher>>,
|
||||
fail_run_diff: bool,
|
||||
_root: tempfile::TempDir,
|
||||
}
|
||||
|
||||
impl Harness {
|
||||
fn new() -> Self {
|
||||
async fn new() -> Self {
|
||||
let root = tempfile::tempdir().expect("a temp dir");
|
||||
let artifact_root = root.path().join("artifacts");
|
||||
std::fs::create_dir(&artifact_root).expect("the isolated artifact directory creates");
|
||||
|
|
@ -105,6 +118,7 @@ impl Harness {
|
|||
),
|
||||
"captures-test",
|
||||
);
|
||||
let (origin, _) = upstream(&root.path().join("fixture"), 1).await;
|
||||
Self {
|
||||
artifact_store,
|
||||
run_id: RunId::new(),
|
||||
|
|
@ -113,20 +127,30 @@ impl Harness {
|
|||
records: Arc::new(MemoryPlatformRecords::new()),
|
||||
blobs: Arc::new(MemoryBlobs::new()),
|
||||
artifacts: Vec::new(),
|
||||
source: Some(file_source(&origin, "main", None)),
|
||||
sandboxed: false,
|
||||
publisher: None,
|
||||
fail_run_diff: false,
|
||||
_root: root,
|
||||
}
|
||||
}
|
||||
|
||||
fn hooks(&self, provider: &SandboxProviderKind) -> HooksSpec {
|
||||
HooksSpec {
|
||||
records: Arc::clone(&self.records) as Arc<dyn PlatformRecords>,
|
||||
records: if self.fail_run_diff {
|
||||
Arc::new(RejectRunDiff(self.records.clone()))
|
||||
} else {
|
||||
self.records.clone()
|
||||
},
|
||||
git: RunGitSettings {
|
||||
host_workspaces: *provider == SandboxProviderKind::LOCAL,
|
||||
host_workspaces: *provider == SandboxProviderKind::LOCAL && !self.sandboxed,
|
||||
..RunGitSettings::default()
|
||||
},
|
||||
artifacts: self.artifacts.clone(),
|
||||
test_gates: None,
|
||||
artifact_writer: Arc::new(StoreArtifactWriter::new(self.artifact_store.clone())),
|
||||
source: self.source.clone(),
|
||||
publisher: self.publisher.clone(),
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -143,6 +167,16 @@ impl Harness {
|
|||
provider: SandboxProviderKind,
|
||||
workflow: &str,
|
||||
settings: &str,
|
||||
) -> engine::RunOutcome {
|
||||
self.execute_on(provider, workflow, settings, false).await
|
||||
}
|
||||
|
||||
async fn execute_on(
|
||||
&self,
|
||||
provider: SandboxProviderKind,
|
||||
workflow: &str,
|
||||
settings: &str,
|
||||
resumed: bool,
|
||||
) -> engine::RunOutcome {
|
||||
let (interviewer, observers) = no_questions();
|
||||
let hooks = self.hooks(&provider);
|
||||
|
|
@ -156,7 +190,11 @@ impl Harness {
|
|||
let request = RunRequest {
|
||||
run_id: self.run_id.to_string(),
|
||||
run_dir: self.run_dir.clone(),
|
||||
execution: Execution::Start(admit(workflow, settings)),
|
||||
execution: if resumed {
|
||||
Execution::Resume
|
||||
} else {
|
||||
Execution::Start(admit(workflow, settings))
|
||||
},
|
||||
store: Arc::clone(&self.store) as Arc<dyn petri_store::RunStore>,
|
||||
runtime: RuntimeSpec {
|
||||
sandbox,
|
||||
|
|
@ -175,27 +213,6 @@ impl Harness {
|
|||
engine::run(request).await.expect("the run executes")
|
||||
}
|
||||
|
||||
/// The commits the snapshot repository of `workspace` holds, oldest
|
||||
/// first, as `(sha, subject, key)`: every checkpoint's history, whatever
|
||||
/// site committed it.
|
||||
async fn snapshot_commits(
|
||||
&self,
|
||||
workspace: &str,
|
||||
) -> Vec<(String, String, Option<CheckpointKey>)> {
|
||||
let repository = self.workspaces().snapshot_repository(workspace);
|
||||
// Topological, so the linear run history reads parents first even
|
||||
// when commits share a timestamp.
|
||||
let log = git(&repository, &[
|
||||
"log",
|
||||
"--topo-order",
|
||||
"--reverse",
|
||||
"--all",
|
||||
"--format=%H%x00%s%x00%B%x1e",
|
||||
])
|
||||
.await;
|
||||
parse_log(&log)
|
||||
}
|
||||
|
||||
async fn inspection(&self) -> RunInspection {
|
||||
let logs = self
|
||||
.store
|
||||
|
|
@ -251,10 +268,8 @@ impl Harness {
|
|||
async fn recover(&self) -> Recovery {
|
||||
recovery::recover(RecoveryRequest {
|
||||
run_id: self.run_id,
|
||||
run_dir: self.run_dir.clone(),
|
||||
store: Arc::clone(&self.store) as Arc<dyn petri_store::RunStore>,
|
||||
records: Arc::clone(&self.records) as Arc<dyn PlatformRecords>,
|
||||
git: RunGitSettings::default(),
|
||||
})
|
||||
.await
|
||||
.expect("recovery decides")
|
||||
|
|
@ -303,6 +318,7 @@ fn parse_log(log: &str) -> Vec<(String, String, Option<CheckpointKey>)> {
|
|||
let body = parts.next().unwrap_or_default();
|
||||
(sha, subject, CheckpointKey::from_message(body))
|
||||
})
|
||||
.filter(|(_, _, key)| key.is_some())
|
||||
.collect()
|
||||
}
|
||||
|
||||
|
|
@ -322,7 +338,8 @@ fn stages(inspection: &RunInspection) -> Vec<(String, String)> {
|
|||
#[tokio::test]
|
||||
async fn dry_runs_use_local_workspaces_and_checkpoint_without_sandbox_credentials() {
|
||||
for provider in [SandboxProviderKind::DOCKER, SandboxProviderKind::DAYTONA] {
|
||||
let harness = Harness::new();
|
||||
let mut harness = Harness::new().await;
|
||||
harness.source = None;
|
||||
let workflow = workflow(
|
||||
r#" write [shape=parallelogram, script="touch should-not-exist; exit 1"]"#,
|
||||
" start -> write -> exit",
|
||||
|
|
@ -369,7 +386,7 @@ async fn dry_runs_use_local_workspaces_and_checkpoint_without_sandbox_credential
|
|||
3,
|
||||
"{provider}: every stage checkpoints"
|
||||
);
|
||||
assert_eq!(harness.snapshot_commits(&workspace).await.len(), 3);
|
||||
assert_eq!(commits(&harness.workspace_path(&workspace)).await.len(), 3);
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -378,7 +395,7 @@ async fn dry_runs_use_local_workspaces_and_checkpoint_without_sandbox_credential
|
|||
/// reached Petri's local service through Fabro's wrapper.
|
||||
#[tokio::test]
|
||||
async fn every_finish_is_committed_and_recorded() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" write [shape=parallelogram, script=\"echo one > out.txt\"]\n check \
|
||||
[shape=parallelogram, script=\"test \\\"$(cat out.txt)\\\" = one\"]",
|
||||
|
|
@ -429,13 +446,7 @@ async fn every_finish_is_committed_and_recorded() {
|
|||
"{checkpoints:?}"
|
||||
);
|
||||
|
||||
// The snapshot repository holds every checkpoint.
|
||||
let published = harness
|
||||
.workspaces()
|
||||
.published(&workspace)
|
||||
.await
|
||||
.expect("the snapshots list");
|
||||
assert_eq!(published.len(), 4);
|
||||
assert!(!harness.run_dir.join("snapshots").exists());
|
||||
|
||||
// `run_complete` and `sandbox_cleanup` ran through the forwarded
|
||||
// service, with the sandbox in place.
|
||||
|
|
@ -452,7 +463,7 @@ async fn every_finish_is_committed_and_recorded() {
|
|||
/// the end.
|
||||
#[tokio::test]
|
||||
async fn artifacts_the_branch_and_the_diffs_are_recorded() {
|
||||
let mut harness = Harness::new();
|
||||
let mut harness = Harness::new().await;
|
||||
harness.artifacts = vec!["assets/**".to_string()];
|
||||
let workflow = workflow(
|
||||
" write [shape=parallelogram, script=\"mkdir -p assets && printf one > \
|
||||
|
|
@ -531,8 +542,15 @@ async fn artifacts_the_branch_and_the_diffs_are_recorded() {
|
|||
let checkpoints = harness.checkpoints();
|
||||
assert_eq!(
|
||||
branches[0].base_sha.as_deref(),
|
||||
Some(checkpoints[0].1.as_str()),
|
||||
"a branch in a fresh repository starts from its first checkpoint"
|
||||
Some(
|
||||
git(&harness.workspace_path(&workspace), &[
|
||||
"rev-parse",
|
||||
&format!("{}^", checkpoints[0].1)
|
||||
])
|
||||
.await
|
||||
.as_str()
|
||||
),
|
||||
"the branch starts at the upstream commit"
|
||||
);
|
||||
let identities: Vec<_> = records
|
||||
.iter()
|
||||
|
|
@ -558,7 +576,8 @@ async fn artifacts_the_branch_and_the_diffs_are_recorded() {
|
|||
})
|
||||
.collect();
|
||||
assert_eq!(diffs.len(), 5, "{diffs:?}");
|
||||
assert_eq!(diffs[0], (None, false));
|
||||
assert_eq!(diffs[0].0.unwrap().files_changed, 0);
|
||||
assert!(!diffs[0].1);
|
||||
let write = diffs[1].0.expect("the write diff");
|
||||
assert_eq!(
|
||||
(write.files_changed, write.additions, write.deletions),
|
||||
|
|
@ -632,7 +651,7 @@ async fn checkpoint_nodes(harness: &Harness) -> Vec<(String, u64)> {
|
|||
/// and its failure route runs on the committed files.
|
||||
#[tokio::test]
|
||||
async fn a_failed_stage_is_committed_and_its_route_sees_the_files() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" work [shape=parallelogram, script=\"echo partial > out.txt; exit 1\"]\n fix \
|
||||
[shape=parallelogram, script=\"test \\\"$(cat out.txt)\\\" = partial && echo fixed >> \
|
||||
|
|
@ -672,7 +691,7 @@ async fn a_failed_stage_is_committed_and_its_route_sees_the_files() {
|
|||
/// checkpoint's error, and a restart reports it failed without resuming.
|
||||
#[tokio::test]
|
||||
async fn a_failed_checkpoint_ends_the_run_with_no_route() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" wreck [shape=parallelogram, script=\"rm -rf .git && echo garbage > .git && echo wrecked \
|
||||
> out.txt\"]\n next [shape=parallelogram, script=\"echo next > next.txt\"]\n fix \
|
||||
|
|
@ -735,7 +754,7 @@ async fn a_failed_checkpoint_ends_the_run_with_no_route() {
|
|||
/// starts over, and a run that finished has nothing to bring back.
|
||||
#[tokio::test]
|
||||
async fn recovery_starts_an_unknown_run_and_resumes_a_finished_one() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
assert_eq!(harness.recover().await, Recovery::Start);
|
||||
|
||||
let workflow = workflow(
|
||||
|
|
@ -802,7 +821,7 @@ async fn a_run_hook_blocks_a_tool_effect_through_the_forwarded_service() {
|
|||
.expect("the model client builds")
|
||||
.expect("openai is eligible");
|
||||
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let (interviewer, observers) = no_questions();
|
||||
let workflow = format!(
|
||||
"digraph Hooks {{\n graph [backend=\"api\", goal=\"Check the tool hooks\", \
|
||||
|
|
@ -879,7 +898,7 @@ async fn the_records_name_the_root_invocations_workspace() {
|
|||
use fabro_petri::workspace::WorkspaceLookup;
|
||||
use petri_execution::InvocationId;
|
||||
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" write [shape=parallelogram, script=\"echo one > out.txt\"]",
|
||||
" start -> write -> exit",
|
||||
|
|
@ -910,7 +929,7 @@ async fn the_records_name_the_root_invocations_workspace() {
|
|||
/// problem.
|
||||
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
|
||||
async fn parallel_branches_checkpoint_the_shared_workspace_in_turn() {
|
||||
let harness = Harness::new();
|
||||
let harness = Harness::new().await;
|
||||
let workflow = workflow(
|
||||
" fork [shape=component]\n a [shape=parallelogram, script=\"echo a > a.txt\"]\n b \
|
||||
[shape=parallelogram, script=\"echo b > b.txt\"]\n merge [shape=tripleoctagon]\n check \
|
||||
|
|
@ -955,6 +974,19 @@ async fn parallel_branches_checkpoint_the_shared_workspace_in_turn() {
|
|||
);
|
||||
assert_eq!(inspection.executions.len(), 3, "the root and two branches");
|
||||
assert_eq!(harness.workspace().await, "invocation-0-scope-0");
|
||||
let child = recorded.iter().find(|key| key.execution != 0).unwrap();
|
||||
let refused = fork::check(
|
||||
harness.store.as_ref(),
|
||||
harness.run_id,
|
||||
fork::position(child.execution, child.firing),
|
||||
)
|
||||
.await
|
||||
.unwrap_err();
|
||||
assert!(
|
||||
refused
|
||||
.to_string()
|
||||
.contains("inside a child invocation cannot be forked")
|
||||
);
|
||||
}
|
||||
|
||||
/// On Docker the workspace lives inside the container: every finished
|
||||
|
|
@ -983,77 +1015,559 @@ async fn a_daytona_run_commits_inside_the_sandbox_and_publishes_every_checkpoint
|
|||
assert_sandbox_run_publishes_every_checkpoint(SandboxProviderKind::DAYTONA).await;
|
||||
}
|
||||
|
||||
/// Bytes of incompressible data the first stage writes: past the plugin
|
||||
/// transport's 16 MiB cap on one file read, so its bundle leaves the
|
||||
/// sandbox in more than one part.
|
||||
const LARGE_FILE_BYTES: usize = 20 * 1024 * 1024;
|
||||
|
||||
/// A two-stage run on `provider`, whose workspace lives inside a sandbox:
|
||||
/// nothing of it is on the host, every checkpoint is published, and the
|
||||
/// bundles carried the stages' files, a large one in parts.
|
||||
/// Git commits and pushes use the acquired sandbox, without a host repository.
|
||||
async fn assert_sandbox_run_publishes_every_checkpoint(provider: SandboxProviderKind) {
|
||||
let harness = Harness::new();
|
||||
let mut harness = Harness::new().await;
|
||||
harness.source = Some(RunSource {
|
||||
origin: "https://github.com/octocat/Hello-World.git".to_string(),
|
||||
revision: SourceRevision::Branch("master".to_string()),
|
||||
branch: "master".to_string(),
|
||||
depth: Some(1),
|
||||
credentials: None,
|
||||
});
|
||||
let publisher = RecordingPublisher::new(None);
|
||||
harness.publisher = Some(publisher.clone());
|
||||
let workflow = workflow(
|
||||
&format!(
|
||||
" write [shape=parallelogram, script=\"echo one > out.txt && head -c \
|
||||
{LARGE_FILE_BYTES} /dev/urandom > large.bin\"]\n check [shape=parallelogram, \
|
||||
script=\"test \\\"$(cat out.txt)\\\" = one && git log --format=%s | head -1 | grep -q \
|
||||
write\"]"
|
||||
),
|
||||
r#" write [shape=parallelogram, script="echo one > out.txt"]
|
||||
check [shape=parallelogram, script="test -f out.txt && git log --format=%s | head -1 | grep -q write"]"#,
|
||||
" start -> write -> check -> exit",
|
||||
);
|
||||
let outcome = harness.run_on(provider, &workflow, SETTINGS).await;
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
assert!(outcome.complete, "{:?}", outcome.incomplete);
|
||||
|
||||
let checkpoints = harness.checkpoints();
|
||||
assert_eq!(checkpoints.len(), 4, "{checkpoints:?}");
|
||||
let workspace = "invocation-0-scope-0";
|
||||
assert!(
|
||||
!harness.workspaces().workspace_exists(workspace).await,
|
||||
"nothing of the workspace is on the host"
|
||||
);
|
||||
let published = harness
|
||||
.workspaces()
|
||||
.published(workspace)
|
||||
let outcome = harness.run_on(provider.clone(), &workflow, SETTINGS).await;
|
||||
if provider == SandboxProviderKind::DOCKER {
|
||||
prune::prune(PruneRequest {
|
||||
sandbox: SandboxProviderConfig::from_lookup(None, |name| env::var(name).ok()),
|
||||
run_id: harness.run_id.to_string(),
|
||||
run_dir: harness.run_dir.clone(),
|
||||
store: harness.store.clone(),
|
||||
provider,
|
||||
})
|
||||
.await
|
||||
.expect("the snapshot repository lists");
|
||||
let mut by_key: Vec<(CheckpointKey, String)> = published
|
||||
.iter()
|
||||
.map(|snapshot| (snapshot.key, snapshot.sha.clone()))
|
||||
.collect();
|
||||
let mut recorded = checkpoints.clone();
|
||||
recorded.sort();
|
||||
by_key.sort();
|
||||
assert_eq!(by_key, recorded, "every record names a published snapshot");
|
||||
|
||||
let commits = harness.snapshot_commits(workspace).await;
|
||||
let subjects: Vec<&str> = commits
|
||||
.iter()
|
||||
.map(|(_, subject, _)| subject.as_str())
|
||||
.collect();
|
||||
let run_id = harness.run_id.to_string();
|
||||
assert_eq!(subjects, vec![
|
||||
format!("fabro({run_id}): start (success)"),
|
||||
format!("fabro({run_id}): write (success)"),
|
||||
format!("fabro({run_id}): check (success)"),
|
||||
format!("fabro({run_id}): exit (success)"),
|
||||
]);
|
||||
let (write_sha, _, _) = &commits[1];
|
||||
let repository = harness.workspaces().snapshot_repository(workspace);
|
||||
assert_eq!(
|
||||
git(&repository, &["show", &format!("{write_sha}:out.txt")]).await,
|
||||
"one",
|
||||
"the bundle carried the stage's files"
|
||||
);
|
||||
assert_eq!(
|
||||
git(&repository, &[
|
||||
"cat-file",
|
||||
"-s",
|
||||
&format!("{write_sha}:large.bin")
|
||||
])
|
||||
.await,
|
||||
LARGE_FILE_BYTES.to_string(),
|
||||
"the large file came through the split transfer whole"
|
||||
.expect("the test container is pruned");
|
||||
}
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
assert_eq!(harness.checkpoints().len(), 4);
|
||||
assert_eq!(publisher.pushed.lock().unwrap().len(), 4);
|
||||
assert!(!harness.run_dir.join("snapshots").exists());
|
||||
assert!(
|
||||
!harness
|
||||
.workspaces()
|
||||
.workspace_exists("invocation-0-scope-0")
|
||||
.await
|
||||
);
|
||||
}
|
||||
|
||||
/// An upstream repository with `commits` commits on `main`, each changing
|
||||
/// `README.md`, and the commit `main` ends on.
|
||||
async fn upstream(root: &Path, commits: usize) -> (PathBuf, String) {
|
||||
let work = root.join("upstream-work");
|
||||
let bare = root.join("upstream.git");
|
||||
fs::create_dir_all(&work)
|
||||
.await
|
||||
.expect("the work tree creates");
|
||||
git(&work, &["init", "-q", "-b", "main"]).await;
|
||||
for index in 1..=commits {
|
||||
fs::write(work.join("README.md"), format!("revision {index}\n"))
|
||||
.await
|
||||
.expect("the file writes");
|
||||
git(&work, &["add", "README.md"]).await;
|
||||
git(&work, &[
|
||||
"-c",
|
||||
"user.name=Upstream",
|
||||
"-c",
|
||||
"user.email=upstream@example.com",
|
||||
"commit",
|
||||
"-q",
|
||||
"-m",
|
||||
&format!("revision {index}"),
|
||||
])
|
||||
.await;
|
||||
}
|
||||
git(root, &[
|
||||
"clone",
|
||||
"-q",
|
||||
"--bare",
|
||||
&work.to_string_lossy(),
|
||||
&bare.to_string_lossy(),
|
||||
])
|
||||
.await;
|
||||
let head = git(&work, &["rev-parse", "HEAD"]).await;
|
||||
(bare, head)
|
||||
}
|
||||
|
||||
/// A Git target's run starts from its repository at depth one: the stage
|
||||
/// sees the files and a shallow history, the snapshot repository is seeded
|
||||
/// with the starting commit, and every checkpoint builds on it (a commit the
|
||||
/// stage made itself included), whether the workspace is on the host or
|
||||
/// `git` runs through the scope's environment and checkpoints leave as
|
||||
/// bundles (which a shallow clone could not send whole).
|
||||
#[tokio::test]
|
||||
async fn a_git_source_is_checked_out_shallow_and_checkpoints_build_on_its_commit() {
|
||||
for sandboxed in [false, true] {
|
||||
let mut harness = Harness::new().await;
|
||||
let (origin, head) = upstream(&harness.run_dir.with_file_name("upstream"), 3).await;
|
||||
harness.sandboxed = sandboxed;
|
||||
harness.source = Some(file_source(&origin, "main", Some(1)));
|
||||
let workflow = workflow(
|
||||
" edit [shape=parallelogram, script=\"test \\\"$(cat README.md)\\\" = 'revision 3' && test \\\"$(git rev-parse --is-shallow-repository)\\\" = true && git rev-parse origin/main && echo edited >> README.md && git -c user.name=Agent -c \
|
||||
user.email=agent@example.com commit -q -am 'agent edit' && echo uncommitted > \
|
||||
notes.txt\"]",
|
||||
" start -> edit -> exit",
|
||||
);
|
||||
let outcome = harness
|
||||
.run_on(SandboxProviderKind::LOCAL, &workflow, SETTINGS)
|
||||
.await;
|
||||
assert_eq!(
|
||||
outcome.status,
|
||||
RunStatus::Success,
|
||||
"sandboxed={sandboxed}: {outcome:?}"
|
||||
);
|
||||
|
||||
let workspace = harness.workspace().await;
|
||||
let workspaces = harness.workspaces();
|
||||
assert!(!harness.run_dir.join("snapshots").exists());
|
||||
let repository = workspaces.workspace_path(&workspace);
|
||||
let checkpoints = harness.checkpoints();
|
||||
let (_, last) = checkpoints.last().expect("a checkpoint was recorded");
|
||||
assert_eq!(
|
||||
git(&repository, &["show", &format!("{last}:README.md")]).await,
|
||||
"revision 3\nedited",
|
||||
"sandboxed={sandboxed}: the last checkpoint carries the stage's edit"
|
||||
);
|
||||
let (_, first) = checkpoints.first().expect("a checkpoint was recorded");
|
||||
assert_eq!(
|
||||
git(&repository, &["rev-parse", &format!("{first}^")]).await,
|
||||
head,
|
||||
"sandboxed={sandboxed}: the run branch starts on the source's commit"
|
||||
);
|
||||
assert_eq!(
|
||||
git(&repository, &["show", &format!("{last}:notes.txt")]).await,
|
||||
"uncommitted",
|
||||
"sandboxed={sandboxed}: the checkpoint after the stage's own commit carries the rest"
|
||||
);
|
||||
assert_eq!(
|
||||
git(&repository, &["rev-list", "--count", last]).await,
|
||||
(checkpoints.len() + 2).to_string(),
|
||||
"sandboxed={sandboxed}: the snapshot holds the run's commits, the stage's own commit \
|
||||
among them, on the one starting commit"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/// A workspace that already holds a repository is not checked out again:
|
||||
/// the source is fetched once per fresh workspace.
|
||||
#[tokio::test]
|
||||
async fn a_prepared_workspace_is_left_as_it_is() {
|
||||
let root = tempfile::tempdir().expect("a temp dir");
|
||||
let (origin, _) = upstream(root.path(), 1).await;
|
||||
let workspaces = RunWorkspaces::new(
|
||||
root.path().join("run"),
|
||||
"run-1".to_string(),
|
||||
GitAuthor::default(),
|
||||
&RunCheckpointSettings::default(),
|
||||
)
|
||||
.with_source(Some(file_source(&origin, "main", None)));
|
||||
let path = root.path().join("prepared");
|
||||
fs::create_dir_all(&path)
|
||||
.await
|
||||
.expect("the workspace creates");
|
||||
git(&path, &["init", "-q"]).await;
|
||||
let site = Site::Host(path.clone());
|
||||
assert_eq!(
|
||||
workspaces
|
||||
.check_out_source(&site, "prepared")
|
||||
.await
|
||||
.expect("the check succeeds"),
|
||||
None
|
||||
);
|
||||
assert!(!path.join("README.md").exists());
|
||||
}
|
||||
|
||||
/// A revision the origin does not have fails the checkout with git's reason.
|
||||
#[tokio::test]
|
||||
async fn an_unavailable_revision_fails_the_checkout() {
|
||||
let root = tempfile::tempdir().expect("a temp dir");
|
||||
let (origin, _) = upstream(root.path(), 1).await;
|
||||
let workspaces = RunWorkspaces::new(
|
||||
root.path().join("run"),
|
||||
"run-1".to_string(),
|
||||
GitAuthor::default(),
|
||||
&RunCheckpointSettings::default(),
|
||||
)
|
||||
.with_source(Some(file_source(&origin, "missing", Some(1))));
|
||||
let path = root.path().join("fresh");
|
||||
let site = Site::Host(path);
|
||||
let error = workspaces
|
||||
.check_out_source(&site, "fresh")
|
||||
.await
|
||||
.expect_err("the branch does not exist");
|
||||
assert!(error.to_string().contains("git fetch failed"), "{error}");
|
||||
}
|
||||
|
||||
/// A publisher that records what it was handed and answers as told.
|
||||
struct RecordingPublisher {
|
||||
pushed: std::sync::Mutex<Vec<(Site, String, String)>>,
|
||||
published: std::sync::Mutex<Vec<Publication>>,
|
||||
fail: Option<String>,
|
||||
}
|
||||
|
||||
impl RecordingPublisher {
|
||||
fn new(fail: Option<&str>) -> Arc<Self> {
|
||||
Arc::new(Self {
|
||||
pushed: std::sync::Mutex::default(),
|
||||
published: std::sync::Mutex::default(),
|
||||
fail: fail.map(str::to_string),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// The source of a run checked out from the local repository `origin`.
|
||||
fn file_source(origin: &Path, branch: &str, depth: Option<u32>) -> RunSource {
|
||||
RunSource {
|
||||
origin: format!("file://{}", origin.display()),
|
||||
revision: SourceRevision::Branch(branch.to_string()),
|
||||
branch: branch.to_string(),
|
||||
depth,
|
||||
credentials: None,
|
||||
}
|
||||
}
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl RunPublisher for RecordingPublisher {
|
||||
async fn push(&self, site: &Site, branch: &str, sha: &str) -> Result<(), String> {
|
||||
self.pushed
|
||||
.lock()
|
||||
.unwrap()
|
||||
.push((site.clone(), branch.to_string(), sha.to_string()));
|
||||
Ok(())
|
||||
}
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String> {
|
||||
self.published
|
||||
.lock()
|
||||
.expect("the ledger locks")
|
||||
.push(publication.clone());
|
||||
self.fail.clone().map_or(Ok(()), Err)
|
||||
}
|
||||
}
|
||||
|
||||
/// A run checked out from `origin` on the local provider, whose one stage
|
||||
/// has the node attributes `attributes`, published through `publisher`.
|
||||
async fn published_run(
|
||||
attributes: &str,
|
||||
publisher: &Arc<RecordingPublisher>,
|
||||
) -> (Harness, engine::RunOutcome) {
|
||||
let mut harness = Harness::new().await;
|
||||
let (origin, _) = upstream(&harness.run_dir.with_file_name("upstream"), 2).await;
|
||||
harness.source = Some(file_source(&origin, "main", Some(1)));
|
||||
harness.publisher = Some(Arc::clone(publisher) as Arc<dyn RunPublisher>);
|
||||
let workflow = workflow(
|
||||
&format!(" edit [shape=parallelogram, {attributes}]"),
|
||||
" start -> edit -> exit",
|
||||
);
|
||||
let outcome = harness
|
||||
.run_on(SandboxProviderKind::LOCAL, &workflow, SETTINGS)
|
||||
.await;
|
||||
(harness, outcome)
|
||||
}
|
||||
|
||||
/// A successful run hands its publisher the run branch, the commit it ends
|
||||
/// on (held by the snapshot repository) and its patch, before the run ends.
|
||||
#[tokio::test]
|
||||
async fn a_successful_run_is_published_with_its_branch_head_and_patch() {
|
||||
let publisher = RecordingPublisher::new(None);
|
||||
let (harness, outcome) = published_run("script=\"echo edited >> README.md\"", &publisher).await;
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
assert!(!outcome.publish_failed);
|
||||
|
||||
let published = publisher.published.lock().unwrap().clone();
|
||||
assert_eq!(published.len(), 1, "published once");
|
||||
let publication = &published[0];
|
||||
assert_eq!(
|
||||
publication.run_branch,
|
||||
format!("fabro/run/{}", harness.run_id)
|
||||
);
|
||||
let (_, last) = harness.checkpoints().last().cloned().expect("a checkpoint");
|
||||
assert_eq!(publication.head_sha, last);
|
||||
assert_eq!(
|
||||
git(&harness.workspace_path(&harness.workspace().await), &[
|
||||
"cat-file", "-t", &last
|
||||
])
|
||||
.await,
|
||||
"commit"
|
||||
);
|
||||
assert!(
|
||||
publication.patch.contains("+edited"),
|
||||
"{}",
|
||||
publication.patch
|
||||
);
|
||||
}
|
||||
|
||||
/// A publication that fails fails the run, with the reason.
|
||||
#[tokio::test]
|
||||
async fn a_failed_publication_fails_the_run() {
|
||||
let publisher = RecordingPublisher::new(Some("the push was rejected"));
|
||||
let (_, outcome) = published_run("script=\"echo edited >> README.md\"", &publisher).await;
|
||||
assert_eq!(outcome.status, RunStatus::Failed, "{outcome:?}");
|
||||
assert!(outcome.publish_failed);
|
||||
assert_eq!(outcome.failure.as_deref(), Some("the push was rejected"));
|
||||
}
|
||||
|
||||
/// A run that fails (here, at a goal gate) is not published.
|
||||
#[tokio::test]
|
||||
async fn a_failed_run_is_not_published() {
|
||||
let publisher = RecordingPublisher::new(None);
|
||||
let (_, outcome) = published_run("script=\"exit 3\", goal_gate=true", &publisher).await;
|
||||
assert_eq!(outcome.status, RunStatus::Failed, "{outcome:?}");
|
||||
assert!(!outcome.publish_failed);
|
||||
assert!(publisher.published.lock().unwrap().is_empty());
|
||||
}
|
||||
|
||||
struct OriginPublisher {
|
||||
origin: String,
|
||||
pushed: std::sync::Mutex<Vec<String>>,
|
||||
}
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl RunPublisher for OriginPublisher {
|
||||
async fn push(&self, site: &Site, branch: &str, sha: &str) -> Result<(), String> {
|
||||
assert!(
|
||||
matches!(site, Site::Sandbox(_)),
|
||||
"push runs through the sandbox environment"
|
||||
);
|
||||
site.push(
|
||||
&self.origin,
|
||||
&format!("{sha}:refs/heads/{branch}"),
|
||||
&[],
|
||||
std::time::Duration::from_secs(30),
|
||||
)
|
||||
.await
|
||||
.map_err(|error| error.to_string())?;
|
||||
self.pushed.lock().unwrap().push(sha.to_owned());
|
||||
Ok(())
|
||||
}
|
||||
|
||||
async fn publish(&self, publication: &Publication) -> Result<(), String> {
|
||||
self.push(
|
||||
&publication.site,
|
||||
&publication.run_branch,
|
||||
&publication.head_sha,
|
||||
)
|
||||
.await
|
||||
}
|
||||
}
|
||||
|
||||
/// The shallow checkout regression: a fork fetches a published checkpoint
|
||||
/// directly from the origin inside its new sandbox, after the old workspace
|
||||
/// has gone. No server Git refs or bundles participate.
|
||||
#[tokio::test]
|
||||
async fn a_shallow_run_pushes_each_checkpoint_and_its_fork_fetches_from_the_origin() {
|
||||
assert_shallow_fork(1).await;
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn a_terminal_checkpoint_is_refused_before_creating_a_fork() {
|
||||
assert_shallow_fork(3).await;
|
||||
}
|
||||
|
||||
async fn assert_shallow_fork(checkpoint_index: usize) {
|
||||
use fabro_petri::fork::{self, ForkRequest};
|
||||
let mut original = Harness::new().await;
|
||||
let (origin, _) = upstream(&original.run_dir.with_file_name("remote"), 5).await;
|
||||
original.source = Some(file_source(&origin, "main", Some(1)));
|
||||
original.sandboxed = true;
|
||||
let publisher = Arc::new(OriginPublisher {
|
||||
origin: format!("file://{}", origin.display()),
|
||||
pushed: std::sync::Mutex::default(),
|
||||
});
|
||||
original.publisher = Some(publisher.clone());
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="printf 'durable bytes\nsecond line\n' > result.txt"]
|
||||
verify [shape=parallelogram, script="test -f result.txt && test -f README.md && cat result.txt"]"#,
|
||||
" start -> write -> verify -> exit",
|
||||
);
|
||||
let outcome = original.run(&graph, SETTINGS).await;
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
let checkpoints = original.checkpoints();
|
||||
let pushed = publisher.pushed.lock().unwrap().clone();
|
||||
assert_eq!(
|
||||
&pushed[..checkpoints.len()],
|
||||
checkpoints
|
||||
.iter()
|
||||
.map(|(_, sha)| sha.clone())
|
||||
.collect::<Vec<_>>()
|
||||
);
|
||||
assert!(!original.run_dir.join("snapshots").exists());
|
||||
|
||||
let mut forked = Harness::new().await;
|
||||
forked.store = original.store.clone();
|
||||
forked.records = original.records.clone();
|
||||
forked.source = original.source.clone();
|
||||
forked.sandboxed = true;
|
||||
forked.publisher = Some(publisher.clone());
|
||||
let (key, checkpoint_sha) = &checkpoints[checkpoint_index];
|
||||
if checkpoint_index == checkpoints.len() - 1 {
|
||||
let refused = fork::check(
|
||||
original.store.as_ref(),
|
||||
original.run_id,
|
||||
fork::position(key.execution, key.firing),
|
||||
)
|
||||
.await
|
||||
.expect_err("terminal checkpoint is refused");
|
||||
assert!(
|
||||
refused
|
||||
.to_string()
|
||||
.contains("terminal checkpoint has no remaining work")
|
||||
);
|
||||
assert!(
|
||||
!forked.run_dir.exists(),
|
||||
"refuse before creating a new run or sandbox"
|
||||
);
|
||||
return;
|
||||
}
|
||||
let seeded = fork::fork(ForkRequest {
|
||||
source: original.run_id,
|
||||
fork: forked.run_id,
|
||||
fork_run_dir: forked.run_dir.clone(),
|
||||
store: forked.store.clone(),
|
||||
records: forked.records.clone(),
|
||||
position: fork::position(key.execution, key.firing),
|
||||
rerun_last: false,
|
||||
settings: RunNamespace::default(),
|
||||
})
|
||||
.await
|
||||
.expect("the fork is seeded");
|
||||
assert_eq!(
|
||||
seeded.start.expect("the selected checkpoint").sha,
|
||||
*checkpoint_sha
|
||||
);
|
||||
fs::remove_dir_all(&original.run_dir)
|
||||
.await
|
||||
.expect("the original workspace is deleted");
|
||||
let outcome = forked
|
||||
.execute_on(SandboxProviderKind::LOCAL, &graph, SETTINGS, true)
|
||||
.await;
|
||||
assert_eq!(outcome.status, RunStatus::Success, "{outcome:?}");
|
||||
let path = forked.workspace_path(&forked.workspace().await);
|
||||
if let Some((_, first_new)) = forked.checkpoints().get(checkpoint_index + 1) {
|
||||
assert_eq!(
|
||||
git(&path, &["rev-parse", &format!("{first_new}^")]).await,
|
||||
*checkpoint_sha,
|
||||
"the fork continues from the selected checkpoint, not the source run's final head"
|
||||
);
|
||||
} else {
|
||||
assert_eq!(git(&path, &["rev-parse", "HEAD"]).await, *checkpoint_sha);
|
||||
}
|
||||
assert_eq!(
|
||||
fs::read(path.join("result.txt"))
|
||||
.await
|
||||
.expect("the recovered file reads"),
|
||||
b"durable bytes\nsecond line\n"
|
||||
);
|
||||
assert_eq!(
|
||||
git(&path, &["branch", "--show-current"]).await,
|
||||
format!("fabro/run/{}", forked.run_id)
|
||||
);
|
||||
assert_eq!(
|
||||
git(&origin, &[
|
||||
"rev-parse",
|
||||
&format!("refs/heads/fabro/run/{}", forked.run_id)
|
||||
])
|
||||
.await,
|
||||
git(&path, &["rev-parse", "HEAD"]).await
|
||||
);
|
||||
assert!(!forked.run_dir.join("snapshots").exists());
|
||||
}
|
||||
|
||||
/// A host workspace with no repository is initialized and checkpointed: every
|
||||
/// stage commits on the run branch, and nothing is kept on the server.
|
||||
#[tokio::test]
|
||||
async fn an_empty_host_workspace_is_initialized_and_checkpointed() {
|
||||
let mut harness = Harness::new().await;
|
||||
harness.source = None;
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="echo data > result.txt"]"#,
|
||||
" start -> write -> exit",
|
||||
);
|
||||
let outcome = harness.run(&graph, SETTINGS).await;
|
||||
assert_eq!(outcome.status, RunStatus::Success);
|
||||
assert_eq!(harness.checkpoints().len(), 3);
|
||||
let path = harness.workspace_path(&harness.workspace().await);
|
||||
assert_eq!(commits(&path).await.len(), 3);
|
||||
assert_eq!(
|
||||
git(&path, &["rev-parse", "--abbrev-ref", "HEAD"]).await,
|
||||
format!("fabro/run/{}", harness.run_id)
|
||||
);
|
||||
assert!(!harness.run_dir.join("snapshots").exists());
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn ordinary_recovery_refuses_a_missing_workspace() {
|
||||
use fabro_petri::recovery::RestoreTarget;
|
||||
let harness = Harness::new().await;
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="echo data > result.txt"]"#,
|
||||
" start -> write -> exit",
|
||||
);
|
||||
assert_eq!(
|
||||
harness.run(&graph, SETTINGS).await.status,
|
||||
RunStatus::Success
|
||||
);
|
||||
let workspace = harness.workspace().await;
|
||||
let path = harness.workspace_path(&workspace);
|
||||
let (key, sha) = harness.checkpoints().last().unwrap().clone();
|
||||
fs::remove_dir_all(&path).await.unwrap();
|
||||
let result = recovery::bring_to(
|
||||
&harness.workspaces(),
|
||||
&Site::Host(path.clone()),
|
||||
&workspace,
|
||||
&RestoreTarget { key, sha },
|
||||
)
|
||||
.await;
|
||||
assert!(result.is_err());
|
||||
assert!(
|
||||
!path.exists(),
|
||||
"normal resume does not recreate the checkout"
|
||||
);
|
||||
}
|
||||
|
||||
struct RejectRunDiff(Arc<MemoryPlatformRecords>);
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl PlatformRecords for RejectRunDiff {
|
||||
async fn append(
|
||||
&self,
|
||||
run_id: &RunId,
|
||||
record: &PlatformRecord,
|
||||
position: Option<fabro_store::StagePosition>,
|
||||
) -> Result<fabro_store::StoredPlatformRecord, PlatformRecordError> {
|
||||
if matches!(record, PlatformRecord::RunDiff(_)) {
|
||||
return Err(PlatformRecordError::Store(fabro_store::Error::Io(
|
||||
std::io::Error::other("run diff unavailable"),
|
||||
)));
|
||||
}
|
||||
self.0.append(run_id, record, position).await
|
||||
}
|
||||
async fn read_kind(
|
||||
&self,
|
||||
run_id: &RunId,
|
||||
kind: PlatformRecordKind,
|
||||
) -> Result<Vec<fabro_store::StoredPlatformRecord>, PlatformRecordError> {
|
||||
self.0.read_kind(run_id, kind).await
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn a_run_diff_failure_cannot_silently_skip_publication() {
|
||||
let mut harness = Harness::new().await;
|
||||
harness.fail_run_diff = true;
|
||||
let publisher = RecordingPublisher::new(None);
|
||||
harness.publisher = Some(publisher.clone());
|
||||
let graph = workflow(
|
||||
r#" write [shape=parallelogram, script="echo data > result.txt"]"#,
|
||||
" start -> write -> exit",
|
||||
);
|
||||
let outcome = harness.run(&graph, SETTINGS).await;
|
||||
assert_eq!(outcome.status, RunStatus::Failed, "{outcome:?}");
|
||||
assert!(outcome.publish_failed);
|
||||
assert!(outcome.failure.unwrap().contains("run diff unavailable"));
|
||||
assert!(publisher.published.lock().unwrap().is_empty());
|
||||
}
|
||||
|
|
|
|||
16
lib/packages/fabro-api-client/src/api/runs-api.ts
generated
16
lib/packages/fabro-api-client/src/api/runs-api.ts
generated
|
|
@ -546,7 +546,7 @@ export const RunsApiAxiosParamCreator = function (configuration?: Configuration)
|
|||
};
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead.
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead. Requires a GitHub target with run-branch creation and pushes enabled; the checkpoint commit must be available on the source run\'s published branch. The final terminal checkpoint is refused because it leaves no work to acquire a workspace. For a completed run, select an earlier checkpoint explicitly or retry the workflow from the beginning.
|
||||
* @summary Fork Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {ForkRequest} [forkRequest]
|
||||
|
|
@ -1204,7 +1204,7 @@ export const RunsApiAxiosParamCreator = function (configuration?: Configuration)
|
|||
};
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed.
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed. Requires a GitHub target with run-branch creation and pushes enabled. The checkpoint must leave work to execute: the final terminal checkpoint is refused; select an earlier checkpoint or retry the workflow from the beginning.
|
||||
* @summary Rewind Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {RewindRequest} [rewindRequest]
|
||||
|
|
@ -1732,7 +1732,7 @@ export const RunsApiFp = function(configuration?: Configuration) {
|
|||
return (axios, basePath) => createRequestFunction(localVarAxiosArgs, globalAxios, BASE_PATH, configuration)(axios, localVarOperationServerBasePath || basePath);
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead.
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead. Requires a GitHub target with run-branch creation and pushes enabled; the checkpoint commit must be available on the source run\'s published branch. The final terminal checkpoint is refused because it leaves no work to acquire a workspace. For a completed run, select an earlier checkpoint explicitly or retry the workflow from the beginning.
|
||||
* @summary Fork Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {ForkRequest} [forkRequest]
|
||||
|
|
@ -1938,7 +1938,7 @@ export const RunsApiFp = function(configuration?: Configuration) {
|
|||
return (axios, basePath) => createRequestFunction(localVarAxiosArgs, globalAxios, BASE_PATH, configuration)(axios, localVarOperationServerBasePath || basePath);
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed.
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed. Requires a GitHub target with run-branch creation and pushes enabled. The checkpoint must leave work to execute: the final terminal checkpoint is refused; select an earlier checkpoint or retry the workflow from the beginning.
|
||||
* @summary Rewind Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {RewindRequest} [rewindRequest]
|
||||
|
|
@ -2180,7 +2180,7 @@ export const RunsApiFactory = function (configuration?: Configuration, basePath?
|
|||
return localVarFp.denyRun(id, denyRunRequest, options).then((request) => request(axios, basePath));
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead.
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead. Requires a GitHub target with run-branch creation and pushes enabled; the checkpoint commit must be available on the source run\'s published branch. The final terminal checkpoint is refused because it leaves no work to acquire a workspace. For a completed run, select an earlier checkpoint explicitly or retry the workflow from the beginning.
|
||||
* @summary Fork Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {ForkRequest} [forkRequest]
|
||||
|
|
@ -2341,7 +2341,7 @@ export const RunsApiFactory = function (configuration?: Configuration, basePath?
|
|||
return localVarFp.retryRun(id, options).then((request) => request(axios, basePath));
|
||||
},
|
||||
/**
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed.
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed. Requires a GitHub target with run-branch creation and pushes enabled. The checkpoint must leave work to execute: the final terminal checkpoint is refused; select an earlier checkpoint or retry the workflow from the beginning.
|
||||
* @summary Rewind Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {RewindRequest} [rewindRequest]
|
||||
|
|
@ -2565,7 +2565,7 @@ export class RunsApi extends BaseAPI {
|
|||
}
|
||||
|
||||
/**
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead.
|
||||
* Creates a new run from a checkpoint of the source run and starts it. The new run holds the source\'s records up to the checkpoint\'s position and continues from there in a fresh workspace restored to the checkpoint\'s commit. The source run is left untouched. A checkpoint inside a parallel branch cannot be forked at; fork at the parallel stage instead. Requires a GitHub target with run-branch creation and pushes enabled; the checkpoint commit must be available on the source run\'s published branch. The final terminal checkpoint is refused because it leaves no work to acquire a workspace. For a completed run, select an earlier checkpoint explicitly or retry the workflow from the beginning.
|
||||
* @summary Fork Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {ForkRequest} [forkRequest]
|
||||
|
|
@ -2741,7 +2741,7 @@ export class RunsApi extends BaseAPI {
|
|||
}
|
||||
|
||||
/**
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed.
|
||||
* Creates a new run from a checkpoint of a terminal source run and starts it, then archives the source run and records `run.superseded_by` on it. Returns 207 when the new run was created but the source archive step failed. Requires a GitHub target with run-branch creation and pushes enabled. The checkpoint must leave work to execute: the final terminal checkpoint is refused; select an earlier checkpoint or retry the workflow from the beginning.
|
||||
* @summary Rewind Run
|
||||
* @param {string} id Unique run identifier (ULID).
|
||||
* @param {RewindRequest} [rewindRequest]
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue