Run the server's lifecycle over platform records instead of run_events

Step 4 of the legacy executor deletion, first commit of several: step 4
spans commits because the legacy event log and its consumers cannot go
in one compiling change. This commit moves every writer off `run_events`;
the reducer, `EventBody`, the Slate bridge and the API's event types
still exist for the readers the next commits port or delete.

Writers:
- The server records a run's lifecycle (submitted, runnable, starting,
  running, blocked, paused, control requests and effects, the terminal
  status), its title, parent link, archive state, notices and pull
  request state as platform records (`fabro_store::platform_records`),
  through the new `server::run_records` module. Every append wakes the
  projector and waits for its pass, so the read that follows a write
  holds the record.
- Pull request creation is recorded as `pull_request.requested`,
  `pull_request.created`, `pull_request.failed`, `pull_request.linked`
  and `pull_request.unlinked`; the projection folds them into the run's
  pull request and creation state.
- Answers to questions are recorded as `interview.answered` with the
  answering principal and the answer text; the interview adapter no
  longer posts legacy `interview.*` events (`QuestionSink` is now an
  optional observer).
- The worker (`fabro run __run-worker`) records its lifecycle, notices
  and pause state over `HttpPlatformRecords`; `HttpRunStore` for the
  legacy event log and the worker's `run_store` are gone.
- `persist_created_run` appends `run.created` and `run.submitted`.

Readers:
- A stream follower (`server::stream_follower`) follows each live run's
  stream (Petri events and platform records), folds lifecycle records
  into the in-memory run state, forwards items to the global attach
  broadcast, and syncs blocked and paused from the projection.
- Slack posts questions from the projection's pending interviews,
  finishes them on `interview.answered` or `question_expired`, and sends
  lifecycle notifications with `notification.sent` dedupe.
- `GET /runs/{id}/events` and the attach endpoints serve only the run
  stream; the per-event, per-stage and `POST /runs/{id}/events`
  endpoints and their tests are deleted.
- `Database::load_run_projection` reads the Petri projection only.

Deleted with the writers:
- The SQLite blob and run-history activation migrations and their
  legacy Slate imports (`legacy_blob_import`, `legacy_run_history_import`,
  the activation backup): a greenfield server has no Slate history to
  import, and the run-history verification refused to start a server
  whose runs have no legacy events.
- `fabro-workflow`'s `operations::archive` and `operations::run_store`.
- The server's legacy-event unit tests and the CLI's `HttpRunStore` tests.

The in-process answer transport is now set after the starting and
running records land, not gated on the live status still being
`Starting` (the records already moved it).

The manifest validation test for a `run.agent.mcps.<name>` catalog
reference now expects `unsupported.workflow_toml.run.agent.mcps.reference`:
Petri's Fabro frontend has no server catalog to resolve it against.

Legacy readers still fail their tests until the next commits: the
reducer and Slate tests in fabro-store, the fabro-workflow create tests
that read the run back through the legacy store, the CLI tests seeded
through `POST /runs/{id}/events`, the CLI's legacy attach and render
paths, the sessions API, the OpenAPI conformance test, and the web
fixtures.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
Bryan Helmkamp 2026-09-18 12:32:06 -04:00
parent 753c926072
commit 3e8b2ebfcc
No known key found for this signature in database
46 changed files with 1557 additions and 16600 deletions

View file

@ -20,7 +20,7 @@ The crate-local `src/migrations.rs` module is the registry. It imports numbered
Examples:
- `fabro-config` owns settings-file migrations.
- `fabro-server` owns server startup activation migrations for SQLite blob storage and run history.
- `fabro-db` owns the SQL schema migrations under `lib/foundation/fabro-db/migrations/`.
Keep migration APIs `pub(crate)` unless another crate genuinely orchestrates the migration.

View file

@ -12,9 +12,9 @@
//! every lease the run takes over the API names it. `--mode start` loads
//! the admitted graphs through the client's blob read and runs them;
//! `--mode resume` continues the run from its records. Either way the
//! worker appends the lifecycle events Fabro's read side needs
//! (`run.starting`, `run.running`, then `run.completed` or `run.failed`)
//! through the client, as the legacy worker does.
//! worker records the lifecycle transitions Fabro's read side needs
//! (`starting`, `running`, then `succeeded` or `failed`) as platform
//! records through the client.
//!
//! The server's controls arrive over the control channel and go to Petri
//! through [`PetriControls`]: cancel (and `SIGTERM`/`SIGINT`) fires one
@ -24,15 +24,15 @@
//! through the API continues; pause and unpause hold and release admission
//! through the run's [`RunControls`]; a steer goes to the run's one live
//! agent stage, or is refused with a `run.notice` record saying why. The
//! paused state is mirrored to Fabro's lifecycle as the legacy worker
//! reported it: a `run.paused` lifecycle event when admission is held and
//! `run.unpaused` when it is released, so the server's live status and the
//! projection agree with Petri's own `run.paused` and `run.unpaused`
//! records. A resumed run that was paused when its worker died comes back
//! paused, and the mirror reports that too. The interrupt and pair
//! controls have no Petri adapter yet and are ignored with a warning. A
//! control channel that is lost for good cancels the run the same way, and
//! the worker exits with that loss as its error once the run has settled.
//! paused state is mirrored to Fabro's lifecycle: a `paused` lifecycle
//! record when admission is held and `unpaused` when it is released, so
//! the server's live status and the projection agree with Petri's own
//! `run.paused` and `run.unpaused` records. A resumed run that was paused when
//! its worker died comes back paused, and the mirror reports that too. The
//! interrupt and pair controls have no Petri adapter yet and are ignored with a
//! warning. A control channel that is lost for good cancels the run the same
//! way, and the worker exits with that loss as its error once the run has
//! settled.
//!
//! Fabro's hooks ride the run with their platform records over the same
//! client: the checkpoint commit in the run's workspace, on the host or
@ -66,20 +66,21 @@ use fabro_petri::blobs::ClientBlobs;
use fabro_petri::controls::RunControls;
use fabro_petri::engine::{self, Conclusion, Execution, RunRequest};
use fabro_petri::hooks::HooksSpec;
use fabro_petri::interview::{Approval, EventSinkQuestions, FabroInterviewer};
use fabro_petri::interview::{Approval, FabroInterviewer};
use fabro_petri::petri::OwnerId;
use fabro_petri::platform_records::HttpPlatformRecords;
use fabro_petri::platform_records::{HttpPlatformRecords, PlatformRecords};
use fabro_petri::runtime::{self, RuntimeSpec};
use fabro_petri::secrets::VaultSecrets;
use fabro_petri::{HttpRunStore, admission};
use fabro_static::EnvVars;
use fabro_store::RunProjection;
use fabro_store::platform_records::{
PlatformRecord, RunLifecycleKind, RunLifecycleRecord, RunNoticeRecord,
};
use fabro_types::settings::run::{ApprovalMode, RunMode};
use fabro_types::{FailureReason, RunId, RunNoticeLevel, RunTiming, StageOutcome, SuccessReason};
use fabro_types::{FailureReason, RunId, RunNoticeLevel, RunStatus, SuccessReason};
use fabro_vault::Vault;
use fabro_workflow::Error as WorkflowError;
use fabro_workflow::event::{self as workflow_event, Event, RunEventSink};
use fabro_workflow::runtime_store::RunStoreHandle;
use fabro_workflow::services::FabroRunToolServices;
use tokio::sync::RwLock as AsyncRwLock;
use tokio::task::JoinHandle;
@ -95,9 +96,6 @@ pub(super) struct PetriWorker<'a> {
pub(super) run_id: RunId,
pub(super) target: ServerTarget,
pub(super) client: Client,
/// The legacy run store over the same client, which carries the
/// lifecycle events to the server with its retries.
pub(super) run_store: RunStoreHandle,
pub(super) run_state: RunProjection,
pub(super) storage_dir: &'a Path,
pub(super) run_dir: PathBuf,
@ -129,12 +127,15 @@ pub(super) async fn execute(worker: PetriWorker<'_>) -> Result<()> {
let cancel_token = CancellationToken::new();
runner::install_signal_handlers(cancel_token.clone())?;
let interviewer = Arc::new(ControlInterviewer::new());
let sink = RunEventSink::map(
runner::stamp_system_worker,
RunEventSink::backend(worker.run_store.clone()),
);
// Fabro's own records of the run, over the client.
let records: Arc<dyn PlatformRecords> =
Arc::new(HttpPlatformRecords::new(worker.client.clone_for_reuse()));
let controls = RunControls::new();
let petri_controls = Arc::new(PetriControls::new(run_id, controls.clone(), sink.clone()));
let petri_controls = Arc::new(PetriControls::new(
run_id,
controls.clone(),
Arc::clone(&records),
));
let mut control_manager = runner::spawn_worker_control_manager(
worker.target.clone(),
run_id,
@ -149,8 +150,7 @@ pub(super) async fn execute(worker: PetriWorker<'_>) -> Result<()> {
} else {
Approval::Prompt
};
let questions = Arc::new(EventSinkQuestions::new(sink.clone(), run_id));
let petri_interviewer = FabroInterviewer::new(interviewer, questions, approval);
let petri_interviewer = FabroInterviewer::new(interviewer, approval);
let observers = vec![petri_interviewer.observer()];
let vault = runner::load_worker_vault(worker.storage_dir).await?;
@ -181,16 +181,16 @@ pub(super) async fn execute(worker: PetriWorker<'_>) -> Result<()> {
};
let started = Instant::now();
for event in [Event::RunStarting, Event::RunRunning] {
workflow_event::append_event_to_sink(&sink, &run_id, &event).await?;
for transition in [
(RunLifecycleKind::Starting, RunStatus::Starting),
(RunLifecycleKind::Running, RunStatus::Running),
] {
lifecycle(&records, run_id, transition.0, transition.1, None).await?;
}
runner::set_worker_title(&run_id, WorkerTitlePhase::Running);
let hooks = HooksSpec::for_run(
Arc::new(HttpPlatformRecords::new(worker.client.clone_for_reuse())),
&worker.run_state.spec.settings.run,
)
.with_test_gates(test_checkpoint_gates());
let hooks = HooksSpec::for_run(Arc::clone(&records), &worker.run_state.spec.settings.run)
.with_test_gates(test_checkpoint_gates());
let request = RunRequest {
run_id: run_id.to_string(),
run_dir: worker.run_dir.join("petri"),
@ -216,7 +216,7 @@ pub(super) async fn execute(worker: PetriWorker<'_>) -> Result<()> {
))),
hooks: Some(hooks),
};
let paused_mirror = mirror_paused_state(run_id, &controls, sink.clone());
let paused_mirror = mirror_paused_state(run_id, &controls, Arc::clone(&records));
let run = Box::pin(engine::run(request));
tokio::pin!(run);
let mut control_lost = None;
@ -235,33 +235,31 @@ pub(super) async fn execute(worker: PetriWorker<'_>) -> Result<()> {
control_manager.finish();
paused_mirror.abort();
let timing = RunTiming {
wall_time_ms: u64::try_from(started.elapsed().as_millis()).unwrap_or(u64::MAX),
..RunTiming::default()
};
let (event, phase, failure) = match engine::conclusion(&result) {
info!(
run_id = %run_id,
elapsed_ms = u64::try_from(started.elapsed().as_millis()).unwrap_or(u64::MAX),
"Petri run ended"
);
let (record, phase, failure) = match engine::conclusion(&result) {
Conclusion::Succeeded => {
info!(run_id = %run_id, "Petri run completed");
(
Event::WorkflowRunCompleted {
timing,
artifact_count: 0,
status: StageOutcome::Succeeded.to_string(),
reason: SuccessReason::Completed,
final_git_commit_sha: None,
final_patch: None,
diff_summary: None,
usage: None,
},
(
RunLifecycleKind::Succeeded,
RunStatus::Succeeded {
reason: SuccessReason::Completed,
},
None,
),
WorkerTitlePhase::Succeeded,
None,
)
}
Conclusion::Failed { reason, message } => {
info!(run_id = %run_id, error = %message, "Petri run did not succeed");
let error = match reason {
FailureReason::Cancelled => WorkflowError::Cancelled,
_ => WorkflowError::engine(message.clone()),
let detail = match reason {
FailureReason::Cancelled => WorkflowError::Cancelled.to_string(),
_ => message.clone(),
};
let phase = if reason == FailureReason::Cancelled {
WorkerTitlePhase::Cancelled
@ -269,15 +267,17 @@ pub(super) async fn execute(worker: PetriWorker<'_>) -> Result<()> {
WorkerTitlePhase::Failed
};
(
Event::workflow_run_failed_from_error(
&error, timing, reason, None, None, None, None,
(
RunLifecycleKind::Failed,
RunStatus::Failed { reason },
Some(detail),
),
phase,
Some(message),
)
}
};
workflow_event::append_event_to_sink(&sink, &run_id, &event).await?;
lifecycle(&records, run_id, record.0, record.1, record.2).await?;
runner::set_worker_title(&run_id, phase);
if let Some(lost) = control_lost {
return Err(lost);
@ -295,15 +295,19 @@ pub(super) struct PetriControls {
run_id: RunId,
controls: RunControls,
/// Where a refused steer's notice goes.
sink: RunEventSink,
records: Arc<dyn PlatformRecords>,
}
impl PetriControls {
pub(super) fn new(run_id: RunId, controls: RunControls, sink: RunEventSink) -> Self {
pub(super) fn new(
run_id: RunId,
controls: RunControls,
records: Arc<dyn PlatformRecords>,
) -> Self {
Self {
run_id,
controls,
sink,
records,
}
}
@ -352,15 +356,12 @@ impl PetriControls {
/// A `run.notice` record on the run, so a refused control is visible in
/// the run's stream and not only in the worker's log.
async fn notice(&self, code: &str, message: String) {
let event = Event::RunNotice {
let record = PlatformRecord::RunNotice(RunNoticeRecord {
level: RunNoticeLevel::Warn,
code: code.to_string(),
message,
exec_output_tail: None,
};
if let Err(error) =
workflow_event::append_event_to_sink(&self.sink, &self.run_id, &event).await
{
});
if let Err(error) = self.records.append(&self.run_id, &record, None).await {
warn!(run_id = %self.run_id, error = %error, "the control notice was not recorded");
}
}
@ -382,14 +383,31 @@ fn control_name(message: &WorkerControlMessage) -> &'static str {
}
}
/// Mirror the run's paused state to Fabro's lifecycle: `run.paused` when
/// One lifecycle transition of the run, recorded through the client.
async fn lifecycle(
records: &Arc<dyn PlatformRecords>,
run_id: RunId,
transition: RunLifecycleKind,
status: RunStatus,
reason: Option<String>,
) -> Result<()> {
let mut record = RunLifecycleRecord::new(transition).with_status(status);
record.reason = reason;
records
.append(&run_id, &PlatformRecord::RunLifecycle(record), None)
.await
.with_context(|| format!("recording the run's {transition} transition"))?;
Ok(())
}
/// Mirror the run's paused state to Fabro's lifecycle: `paused` when
/// admission is held (a pause, or a resume that came back paused) and
/// `run.unpaused` when it is released, each once per change, with the
/// `unpaused` when it is released, each once per change, with the
/// worker's title alongside. Aborted with the run.
fn mirror_paused_state(
run_id: RunId,
controls: &RunControls,
sink: RunEventSink,
records: Arc<dyn PlatformRecords>,
) -> JoinHandle<()> {
let mut changes = controls.paused_changes();
tokio::spawn(async move {
@ -400,12 +418,13 @@ fn mirror_paused_state(
continue;
}
last = paused;
let (event, phase) = if paused {
(Event::RunPaused, WorkerTitlePhase::Paused)
let (transition, phase) = if paused {
(RunLifecycleKind::Paused, WorkerTitlePhase::Paused)
} else {
(Event::RunUnpaused, WorkerTitlePhase::Running)
(RunLifecycleKind::Unpaused, WorkerTitlePhase::Running)
};
if let Err(error) = workflow_event::append_event_to_sink(&sink, &run_id, &event).await {
let record = PlatformRecord::RunLifecycle(RunLifecycleRecord::new(transition));
if let Err(error) = records.append(&run_id, &record, None).await {
warn!(run_id = %run_id, error = %error, "the paused state was not reported");
}
runner::set_worker_title(&run_id, phase);

View file

@ -4,7 +4,6 @@ use std::sync::Arc;
use std::time::Duration;
use anyhow::{Context, Result, anyhow};
use async_trait::async_trait;
use fabro_client::ServerTarget;
use fabro_config::Storage;
use fabro_interview::{
@ -14,11 +13,9 @@ use fabro_interview::{
WorkerControlMessage,
};
use fabro_manifest::SuppliedWorkflowVersionPackager;
use fabro_store::{EventEnvelope, RunProjection, RunProjectionReducer};
use fabro_tool::fabro_client::ClientBackend;
use fabro_types::{BlobHash, Principal, RunEvent, RunId};
use fabro_types::RunId;
use fabro_vault::{SecretStore, Vault};
use fabro_workflow::runtime_store::{RunStoreBackend, RunStoreHandle};
use fabro_workflow::services::FabroRunToolServices;
use futures::{SinkExt, StreamExt};
use jsonwebtoken::dangerous::insecure_decode;
@ -29,7 +26,7 @@ use tokio::net::TcpStream;
use tokio::net::UnixStream;
#[cfg(unix)]
use tokio::signal::unix::{SignalKind, signal};
use tokio::sync::{Mutex, RwLock as AsyncRwLock, oneshot};
use tokio::sync::{RwLock as AsyncRwLock, oneshot};
use tokio::task::JoinHandle;
use tokio::time::{self, Instant, MissedTickBehavior};
use tokio_tungstenite::tungstenite::client::IntoClientRequest;
@ -43,12 +40,6 @@ use super::petri_worker::{self, PetriControls, PetriWorker};
use crate::args::RunWorkerMode;
use crate::server_client;
const RUN_STORE_RETRY_DELAYS: [Duration; 3] = [
Duration::from_millis(50),
Duration::from_millis(100),
Duration::from_millis(250),
];
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub(super) enum WorkerTitlePhase {
Start,
@ -74,16 +65,14 @@ pub(crate) async fn execute(
let target = server.parse::<ServerTarget>()?;
let client = server_client::connect_server_target_with_bearer(&target, worker_token).await?;
let run_store = HttpRunStore::connect(run_id, client.clone_for_reuse()).await?;
let run_state = run_store
.state()
let run_state = client
.get_run_state(&run_id)
.await
.with_context(|| format!("failed to load run state for {run_id}"))?;
Box::pin(petri_worker::execute(PetriWorker {
run_id,
target,
client,
run_store,
run_state,
storage_dir: &storage_dir,
run_dir,
@ -663,150 +652,6 @@ async fn apply_worker_control_message(
}
}
#[derive(Clone)]
struct HttpRunStore {
run_id: RunId,
client: server_client::Client,
state: Arc<Mutex<RunProjection>>,
events: Arc<Mutex<Option<Vec<EventEnvelope>>>>,
}
impl HttpRunStore {
async fn connect(run_id: RunId, client: server_client::Client) -> Result<RunStoreHandle> {
let state = client
.get_run_state(&run_id)
.await
.with_context(|| format!("failed to fetch run state for {run_id}"))?;
Ok(RunStoreHandle::new(Arc::new(Self {
run_id,
client,
state: Arc::new(Mutex::new(state)),
events: Arc::new(Mutex::new(None)),
})))
}
async fn with_retries<T, F, Fut>(&self, operation: &'static str, mut op: F) -> Result<T>
where
F: FnMut() -> Fut,
Fut: std::future::Future<Output = Result<T>>,
{
let mut last_error = None;
for attempt in 0..=RUN_STORE_RETRY_DELAYS.len() {
match op().await {
Ok(value) => return Ok(value),
Err(err) => last_error = Some(err),
}
if let Some(delay) = RUN_STORE_RETRY_DELAYS.get(attempt) {
time::sleep(*delay).await;
}
}
Err(last_error
.unwrap_or_else(|| anyhow!("run store operation failed"))
.context(format!(
"worker lost canonical run store during {operation}"
)))
}
async fn refresh_state_from_server(&self) -> Result<RunProjection> {
self.with_retries("refresh state", || {
let client = self.client.clone_for_reuse();
let run_id = self.run_id;
async move { client.get_run_state(&run_id).await }
})
.await
}
async fn apply_acknowledged_event(&self, seq: u32, event: &RunEvent) -> Result<()> {
let envelope = EventEnvelope {
seq,
event: event.clone(),
};
{
let mut state = self.state.lock().await;
if let Err(err) = state.apply_event(&envelope) {
tracing::warn!(run_id = %self.run_id, error = %err, "failed to apply acknowledged event to local run-state mirror; refreshing from server");
drop(state);
let refreshed = self.refresh_state_from_server().await?;
*self.state.lock().await = refreshed;
}
}
let mut events = self.events.lock().await;
if let Some(cached) = events.as_mut() {
cached.push(envelope);
}
Ok(())
}
}
#[async_trait]
impl RunStoreBackend for HttpRunStore {
async fn load_state(&self) -> Result<RunProjection> {
Ok(self.state.lock().await.clone())
}
async fn list_events(&self) -> Result<Vec<EventEnvelope>> {
let mut cached = self.events.lock().await;
if let Some(events) = cached.as_ref() {
return Ok(events.clone());
}
let events = self
.with_retries("list run events", || {
let client = self.client.clone_for_reuse();
let run_id = self.run_id;
async move { client.list_run_events(&run_id, None, None).await }
})
.await?;
*cached = Some(events.clone());
Ok(events)
}
async fn append_run_event(&self, event: &RunEvent) -> Result<()> {
let seq = Box::pin(self.with_retries("append run event", || {
let client = self.client.clone_for_reuse();
let run_id = self.run_id;
let event = event.clone();
async move { client.append_run_event(&run_id, &event).await }
}))
.await?;
// Both the sandbox lifecycle and the lithos event shapes grew this
// future past clippy's stack budget; box it once at the call.
Box::pin(self.apply_acknowledged_event(seq, event)).await
}
async fn write_blob(&self, data: &[u8]) -> Result<BlobHash> {
self.with_retries("write run blob", || {
let client = self.client.clone_for_reuse();
let run_id = self.run_id;
let data = data.to_vec();
async move { client.write_run_blob(&run_id, &data).await }
})
.await
}
async fn read_blob(&self, blob_hash: &BlobHash) -> Result<Option<bytes::Bytes>> {
self.with_retries("read run blob", || {
let client = self.client.clone_for_reuse();
let run_id = self.run_id;
let blob_hash = *blob_hash;
async move { client.read_run_blob(&run_id, &blob_hash).await }
})
.await
}
async fn read_run_log(&self) -> Result<Option<Vec<u8>>> {
self.with_retries("get run logs", || {
let client = self.client.clone_for_reuse();
let run_id = self.run_id;
async move { client.get_run_logs(&run_id).await }
})
.await
}
}
pub(super) fn set_worker_title(run_id: &RunId, phase: WorkerTitlePhase) {
fabro_proc::title_set(&worker_title(run_id, phase));
}
@ -832,15 +677,6 @@ fn worker_title(run_id: &RunId, phase: WorkerTitlePhase) -> String {
format!("fabro {short_id} {phase}")
}
pub(super) fn stamp_system_worker(mut event: RunEvent) -> RunEvent {
if event.actor.is_none() {
event.actor = Some(Principal::Worker {
run_id: event.run_id,
});
}
event
}
/// `SIGTERM` and `SIGINT` cancel the run, the way the server's cancel does.
pub(super) fn install_signal_handlers(cancel_token: CancellationToken) -> Result<()> {
#[cfg(unix)]
@ -873,16 +709,13 @@ mod tests {
use std::sync::Arc;
use std::time::Duration;
use chrono::Utc;
use fabro_client::ServerTarget;
use fabro_config::Storage;
use fabro_interview::{
AnswerValue, ControlInterviewer, Interviewer, Question, WorkerControlEnvelope,
};
use fabro_types::run_event::RunStatusTransitionProps;
use fabro_types::{AuthMethod, EventBody, IdpIdentity, Principal, QuestionType, fixtures};
use fabro_types::{QuestionType, fixtures};
use fabro_vault::{SecretType, Vault};
use fabro_workflow::event::RunEventSink;
use tokio::time;
use tokio_tungstenite::tungstenite::protocol::{Message as TestWebSocketMessage, Role};
use tokio_util::sync::CancellationToken;
@ -893,19 +726,17 @@ mod tests {
WorkerControls, WorkerTitlePhase, apply_worker_control_delivery_frame,
apply_worker_control_message, build_worker_control_stream_request,
connect_worker_control_stream, handle_worker_control_socket, initial_worker_title_phase,
load_worker_vault, next_worker_control_reconnect_backoff, stamp_system_worker,
worker_title,
load_worker_vault, next_worker_control_reconnect_backoff, worker_title,
};
use crate::args::RunWorkerMode;
/// A run's controls over a sink that keeps nothing: what the channel
/// A run's controls over records kept in memory: what the channel
/// tests drive.
fn test_controls() -> WorkerControls {
let sink = RunEventSink::callback(|_event| async move { Ok(()) });
Arc::new(PetriControls::new(
fixtures::RUN_1,
fabro_petri::controls::RunControls::new(),
sink,
Arc::new(fabro_petri::test_support::MemoryPlatformRecords::new()),
))
}
@ -952,32 +783,6 @@ mod tests {
.expect("test worker token should encode")
}
fn test_user_principal(login: &str) -> Principal {
Principal::user(
IdpIdentity::new("https://github.com", "12345").unwrap(),
login.to_string(),
AuthMethod::Github,
)
}
fn running_event(actor: Option<Principal>) -> fabro_types::RunEvent {
fabro_types::RunEvent {
id: "evt_1".to_string(),
ts: Utc::now(),
run_id: fixtures::RUN_1,
node_id: None,
node_label: None,
stage_id: None,
parallel_group_id: None,
parallel_branch_id: None,
session_id: None,
parent_session_id: None,
tool_call_id: None,
actor,
body: EventBody::RunRunning(RunStatusTransitionProps::default()),
}
}
#[test]
fn worker_title_uses_short_run_id_and_phase() {
let short_id: String = fixtures::RUN_1.to_string().chars().take(12).collect();
@ -1003,67 +808,6 @@ mod tests {
);
}
#[test]
fn stamp_system_worker_fills_missing_actor_only() {
let stamped = stamp_system_worker(running_event(None));
assert_eq!(
stamped.actor,
Some(Principal::Worker {
run_id: fixtures::RUN_1,
})
);
let existing_actor = test_user_principal("octocat");
let stamped = stamp_system_worker(running_event(Some(existing_actor.clone())));
assert_eq!(stamped.actor, Some(existing_actor));
}
#[tokio::test]
async fn worker_event_stamp_applies_to_all_fanout_sinks() {
let first = Arc::new(tokio::sync::Mutex::new(Vec::new()));
let second = Arc::new(tokio::sync::Mutex::new(Vec::new()));
let first_events = Arc::clone(&first);
let second_events = Arc::clone(&second);
let sink = RunEventSink::map(
stamp_system_worker,
RunEventSink::fanout(vec![
RunEventSink::callback(move |event| {
let first_events = Arc::clone(&first_events);
async move {
first_events.lock().await.push(event);
Ok(())
}
}),
RunEventSink::callback(move |event| {
let second_events = Arc::clone(&second_events);
async move {
second_events.lock().await.push(event);
Ok(())
}
}),
]),
);
let event = running_event(None);
sink.write_run_event(&event).await.unwrap();
let first = first.lock().await;
let second = second.lock().await;
assert_eq!(
first[0].actor,
Some(Principal::Worker {
run_id: fixtures::RUN_1,
})
);
assert_eq!(
second[0].actor,
Some(Principal::Worker {
run_id: fixtures::RUN_1,
})
);
}
#[tokio::test]
async fn worker_control_routes_answer_by_question_id() {
let interviewer = Arc::new(ControlInterviewer::new());

View file

@ -169,7 +169,8 @@ impl RunningServer {
.stdout(self.stderr_log())
.stderr(self.stderr_log());
let mut child = cmd.spawn().expect("the server spawns");
wait_for_http_ready(&self.api_base_url, &mut child).await;
let log_path = self.storage_dir.with_file_name("server.stderr.log");
wait_for_http_ready(&self.api_base_url, &mut child, &log_path).await;
self.child = Some(child);
}
@ -370,7 +371,7 @@ fn reserve_port() -> u16 {
.port()
}
async fn wait_for_http_ready(base_url: &str, child: &mut Child) {
async fn wait_for_http_ready(base_url: &str, child: &mut Child, log_path: &Path) {
let client = fabro_test::test_http_client();
let deadline = Instant::now() + Duration::from_secs(10);
loop {
@ -378,7 +379,13 @@ async fn wait_for_http_ready(base_url: &str, child: &mut Child) {
Ok(response) if response.status().is_success() => return,
Ok(_) | Err(_) if Instant::now() < deadline => {
if let Some(status) = child.try_wait().expect("the server polls") {
panic!("the server exited before it was ready with status {status}");
let log = std::fs::read_to_string(log_path).unwrap_or_default();
let tail = log.lines().rev().take(20).collect::<Vec<_>>();
panic!(
"the server exited before it was ready with status {status}; its log ends \
with:\n{}",
tail.into_iter().rev().collect::<Vec<_>>().join("\n")
);
}
tokio::time::sleep(Duration::from_millis(25)).await;
}

View file

@ -1,746 +0,0 @@
//! Fail-closed activation of SQLite blob storage.
//!
//! This compatibility bridge remains until at least 30 calendar days after
//! the first successful production activation, and until the cold-start,
//! warm-restart, production-observation, and backup-integrity evidence is
//! complete and Scott explicitly approves its removal. The date is an
//! eligibility floor, never an automatic deletion trigger.
use std::path::{Path, PathBuf};
use std::sync::Arc;
use std::time::Duration;
use object_store::ObjectStore;
use tokio::fs;
use tracing::{debug, info, warn};
use crate::migrations::sqlite_activation_backup::{self, BackupError};
use crate::migrations::sqlite_run_history_activation::BACKUP_SUFFIX as RUN_HISTORY_BACKUP_SUFFIX;
use crate::server::resource_sampler;
/// Earliest date this bridge becomes eligible for removal, assuming the first
/// production activation happens no earlier than this change ships. Removal
/// additionally requires the evidence and explicit approval described in the
/// module docs; the date alone never triggers deletion.
pub(crate) const REMOVAL_DEADLINE: &str = "2026-09-22";
const DISK_HEADROOM_BYTES: u64 = 64 * 1024 * 1024;
const BACKUP_SUFFIX: &str = ".pre-blob-activation.bak";
pub(crate) struct ActivatedBlobStorage {
pub(crate) store: Arc<fabro_store::Database>,
pub(crate) run_history_identity: fabro_store::LegacyRunHistorySourceIdentity,
}
impl std::fmt::Debug for ActivatedBlobStorage {
fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
formatter
.debug_struct("ActivatedBlobStorage")
.field("run_history_identity", &self.run_history_identity)
.finish_non_exhaustive()
}
}
#[derive(Debug, thiserror::Error)]
pub(crate) enum BlobActivationError {
#[error("canonicalizing the SQLite database path {path}")]
Canonicalize {
path: PathBuf,
#[source]
source: std::io::Error,
},
#[error("inventorying the legacy blob source")]
Inventory(#[source] fabro_store::LegacyBlobInventoryError),
#[error("identifying legacy run history before importing blobs")]
RunHistorySourceIdentity(#[source] fabro_store::LegacyRunHistorySourceIdentityError),
#[error(
"activation backup is missing at {path} while {existing_rows} of {legacy_rows} legacy blob rows are already present in SQLite"
)]
MissingBackupAfterImport {
path: PathBuf,
legacy_rows: u64,
existing_rows: u64,
},
#[error("reading SQLite file metadata at {path}")]
SqliteMetadata {
path: PathBuf,
#[source]
source: std::io::Error,
},
#[error("the blob activation disk requirement overflowed")]
DiskRequirementOverflow,
#[error(
"insufficient disk space for blob activation: {available_bytes} bytes available, {required_bytes} required"
)]
InsufficientDisk {
required_bytes: u64,
available_bytes: u64,
},
#[error(transparent)]
Backup(#[from] BackupError),
#[error("importing legacy blobs into SQLite")]
Import(#[source] Box<fabro_store::LegacyBlobImportError>),
#[error("verifying legacy and SQLite blobs")]
Verification(#[source] Box<fabro_store::LegacyBlobVerificationError>),
#[error("running the live SQLite integrity check")]
LiveIntegrity(#[source] sqlx::Error),
#[error("the live SQLite integrity check did not return exactly one ok result")]
LiveIntegrityFailed,
#[error("running the final SQLite WAL truncate checkpoint")]
FinalCheckpoint(#[source] sqlx::Error),
}
pub(crate) async fn activate_blob_storage(
database: &fabro_db::Database,
sqlite_path: &Path,
object_store: Arc<dyn ObjectStore>,
slatedb_prefix: String,
flush_interval: Duration,
cache_path: Option<PathBuf>,
) -> Result<ActivatedBlobStorage, BlobActivationError> {
let canonical_path = fs::canonicalize(sqlite_path).await.map_err(|source| {
BlobActivationError::Canonicalize {
path: sqlite_path.to_path_buf(),
source,
}
})?;
let backup_path = fabro_db::append_to_path(&canonical_path, BACKUP_SUFFIX);
info!(
database_path = %canonical_path.display(),
backup_path = %backup_path.display(),
"Starting SQLite blob storage activation"
);
let blob_store = Arc::new(fabro_store::BlobStore::new(database.clone_pool()));
let run_summary_store = Arc::new(fabro_store::RunSummaryStore::new(database.clone_pool()));
let store = Arc::new(fabro_store::Database::new(
object_store,
slatedb_prefix,
flush_interval,
cache_path,
Arc::clone(&blob_store),
run_summary_store,
));
let inventory = store
.legacy_blob_inventory(database.pool())
.await
.map_err(BlobActivationError::Inventory)?;
let run_history_identity = store
.legacy_run_history_source_identity()
.await
.map_err(BlobActivationError::RunHistorySourceIdentity)?;
let backup_exists = sqlite_activation_backup::backup_exists(&backup_path).await?;
if backup_exists {
sqlite_activation_backup::validate_backup(&backup_path).await?;
}
if !backup_exists && inventory.pending_rows < inventory.rows {
return Err(BlobActivationError::MissingBackupAfterImport {
path: backup_path,
legacy_rows: inventory.rows,
existing_rows: inventory.rows - inventory.pending_rows,
});
}
let backup_required = inventory.rows > 0 && !backup_exists;
let run_history_backup_path =
fabro_db::append_to_path(&canonical_path, RUN_HISTORY_BACKUP_SUFFIX);
let run_history_backup_exists =
sqlite_activation_backup::backup_exists(&run_history_backup_path).await?;
if run_history_backup_exists {
sqlite_activation_backup::validate_backup(&run_history_backup_path).await?;
}
let run_history_backup_required =
run_history_identity.events != 0 && !run_history_backup_exists;
// The resource sampler treats a path with no matching mount as an
// unsupported-but-benign condition (tmpfs or squashfs roots, network
// filesystems, an unreadable mount table), so the preflight does too:
// skipping the capacity check must not block a boot the import itself
// could complete.
if let Some(available_free_bytes) = resource_sampler::available_space_for_path(&canonical_path)
{
let sqlite_bytes = if backup_required || run_history_backup_required {
sqlite_file_set_bytes(&canonical_path).await?
} else {
0
};
let backup_reserve = if backup_required { sqlite_bytes } else { 0 };
let run_history_backup_reserve = if run_history_backup_required {
projected_sqlite_bytes(sqlite_bytes, inventory.pending_bytes)?
} else {
0
};
// Only the rows the import still has to copy need new space; rows
// already present in SQLite cost nothing on a warm restart. Reserve
// the projected post-import database size as well when run-history
// activation will immediately take its own full SQLite snapshot.
let required_free_bytes = compute_disk_preflight(
inventory.pending_bytes,
backup_reserve,
run_history_backup_reserve,
available_free_bytes,
)?;
debug!(
legacy_rows = inventory.rows,
legacy_bytes = inventory.bytes,
pending_rows = inventory.pending_rows,
pending_bytes = inventory.pending_bytes,
backup_required,
backup_reserve,
run_history_events = run_history_identity.events,
run_history_backup_required,
run_history_backup_reserve,
required_free_bytes,
available_free_bytes,
"Checked SQLite blob activation disk capacity"
);
} else {
warn!(
database_path = %canonical_path.display(),
"No filesystem mount matched the SQLite database path; skipping the blob activation disk preflight"
);
}
let retained_backup = if backup_exists {
Some(backup_path)
} else if backup_required {
sqlite_activation_backup::create_backup(database.pool(), &backup_path).await?;
Some(backup_path)
} else {
None
};
let import = store
.import_legacy_blobs_into(database.pool())
.await
.map_err(|source| BlobActivationError::Import(Box::new(source)))?;
// The import pass already validates every legacy digest and byte-compares
// every already-present row on each boot, so the independent verification
// sweep only needs to double-check boots that actually inserted rows.
let verification = if import.imported_rows > 0 {
Some(
store
.verify_legacy_blobs_in(database.pool())
.await
.map_err(|source| BlobActivationError::Verification(Box::new(source)))?,
)
} else {
None
};
validate_live_integrity(database.pool()).await?;
final_truncate_checkpoint(database.pool()).await?;
info!(
legacy_rows = inventory.rows,
legacy_bytes = inventory.bytes,
imported_rows = import.imported_rows,
existing_rows = import.existing_rows,
matched_rows = verification.as_ref().map(|report| report.matched_rows),
target_rows = verification.as_ref().map(|report| report.target_rows),
passive_checkpoints = import.passive_checkpoints,
backup_required,
backup_path = ?retained_backup,
run_history_backup_required,
removal_deadline = REMOVAL_DEADLINE,
"Activated SQLite blob storage"
);
Ok(ActivatedBlobStorage {
store,
run_history_identity,
})
}
/// Fail-closed disk capacity check; returns the required free bytes.
fn compute_disk_preflight(
pending_bytes: u64,
backup_reserve: u64,
run_history_backup_reserve: u64,
available_free_bytes: u64,
) -> Result<u64, BlobActivationError> {
let import_reserve = blob_import_reserve(pending_bytes)?;
let required_free_bytes = backup_reserve
.checked_add(import_reserve)
.and_then(|value| value.checked_add(run_history_backup_reserve))
.and_then(|value| value.checked_add(DISK_HEADROOM_BYTES))
.ok_or(BlobActivationError::DiskRequirementOverflow)?;
if available_free_bytes < required_free_bytes {
return Err(BlobActivationError::InsufficientDisk {
required_bytes: required_free_bytes,
available_bytes: available_free_bytes,
});
}
Ok(required_free_bytes)
}
fn projected_sqlite_bytes(
sqlite_bytes: u64,
pending_bytes: u64,
) -> Result<u64, BlobActivationError> {
sqlite_bytes
.checked_add(blob_import_reserve(pending_bytes)?)
.ok_or(BlobActivationError::DiskRequirementOverflow)
}
fn blob_import_reserve(pending_bytes: u64) -> Result<u64, BlobActivationError> {
let half = pending_bytes
.checked_add(1)
.ok_or(BlobActivationError::DiskRequirementOverflow)?
/ 2;
pending_bytes
.checked_add(half)
.ok_or(BlobActivationError::DiskRequirementOverflow)
}
async fn sqlite_file_set_bytes(path: &Path) -> Result<u64, BlobActivationError> {
let mut total = required_file_bytes(path).await?;
for suffix in ["-wal", "-shm"] {
let sibling = fabro_db::append_to_path(path, suffix);
let bytes = optional_file_bytes(&sibling).await?;
total = total
.checked_add(bytes)
.ok_or(BlobActivationError::DiskRequirementOverflow)?;
}
Ok(total)
}
async fn required_file_bytes(path: &Path) -> Result<u64, BlobActivationError> {
fs::metadata(path)
.await
.map(|metadata| metadata.len())
.map_err(|source| BlobActivationError::SqliteMetadata {
path: path.to_path_buf(),
source,
})
}
async fn optional_file_bytes(path: &Path) -> Result<u64, BlobActivationError> {
match fs::metadata(path).await {
Ok(metadata) => Ok(metadata.len()),
Err(source) if source.kind() == std::io::ErrorKind::NotFound => Ok(0),
Err(source) => Err(BlobActivationError::SqliteMetadata {
path: path.to_path_buf(),
source,
}),
}
}
async fn validate_live_integrity(pool: &sqlx::SqlitePool) -> Result<(), BlobActivationError> {
let ok = sqlite_activation_backup::integrity_check_is_ok(pool)
.await
.map_err(BlobActivationError::LiveIntegrity)?;
if !ok {
return Err(BlobActivationError::LiveIntegrityFailed);
}
Ok(())
}
async fn final_truncate_checkpoint(pool: &sqlx::SqlitePool) -> Result<(), BlobActivationError> {
let (busy, _, _): (i64, i64, i64) = sqlx::query_as("PRAGMA wal_checkpoint(TRUNCATE)")
.fetch_one(pool)
.await
.map_err(BlobActivationError::FinalCheckpoint)?;
if busy != 0 {
// A concurrent reader (a backup tool, a replication agent, an
// operator shell) can keep the WAL from truncating. An untruncated
// WAL threatens no data integrity, so it must not block startup; a
// later checkpoint truncates once the reader is gone.
warn!("The final SQLite WAL truncate checkpoint could not complete; continuing startup");
}
Ok(())
}
#[cfg(test)]
mod tests {
use std::sync::Arc;
use std::time::Duration;
use fabro_db::append_to_path;
use object_store::ObjectStore;
use object_store::memory::InMemory;
use tokio::fs;
use super::{
BACKUP_SUFFIX, BlobActivationError, DISK_HEADROOM_BYTES, activate_blob_storage,
compute_disk_preflight, final_truncate_checkpoint, projected_sqlite_bytes,
sqlite_file_set_bytes,
};
use crate::migrations::sqlite_activation_backup::{self, BackupError, create_backup};
type TestResult<T> = Result<T, Box<dyn std::error::Error>>;
#[test]
fn disk_preflight_passes_at_equality_and_fails_one_byte_below() {
let pending_bytes = 3;
let backup_reserve = 10;
let run_history_backup_reserve = 20;
let required =
backup_reserve + pending_bytes + 2 + run_history_backup_reserve + DISK_HEADROOM_BYTES;
let required_free_bytes = compute_disk_preflight(
pending_bytes,
backup_reserve,
run_history_backup_reserve,
required,
)
.expect("exact equality must pass");
assert_eq!(required_free_bytes, required);
let error = compute_disk_preflight(
pending_bytes,
backup_reserve,
run_history_backup_reserve,
required - 1,
)
.expect_err("one byte below must fail");
assert!(matches!(
error,
BlobActivationError::InsufficientDisk { .. }
));
}
#[test]
fn disk_preflight_requires_only_headroom_without_a_backup_reserve() {
let required_free_bytes =
compute_disk_preflight(2, 0, 0, u64::MAX).expect("available capacity should pass");
assert_eq!(required_free_bytes, 3 + DISK_HEADROOM_BYTES);
}
#[test]
fn disk_preflight_reserves_the_projected_post_import_database() {
let projected = projected_sqlite_bytes(10, 3).expect("the projection should fit");
assert_eq!(projected, 15);
let required = compute_disk_preflight(3, 0, projected, u64::MAX)
.expect("available capacity should pass");
assert_eq!(required, 20 + DISK_HEADROOM_BYTES);
}
#[test]
fn disk_preflight_fails_closed_on_overflow() {
let error = compute_disk_preflight(u64::MAX, 1, 0, u64::MAX)
.expect_err("overflow must fail closed");
assert!(matches!(
error,
BlobActivationError::DiskRequirementOverflow
));
}
#[tokio::test]
async fn disk_preflight_counts_the_sqlite_file_set_for_a_required_backup() -> TestResult<()> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
fs::write(&sqlite_path, [0_u8; 3]).await?;
fs::write(append_to_path(&sqlite_path, "-wal"), [0_u8; 5]).await?;
fs::write(append_to_path(&sqlite_path, "-shm"), [0_u8; 7]).await?;
assert_eq!(sqlite_file_set_bytes(&sqlite_path).await?, 15);
Ok(())
}
#[tokio::test]
async fn backup_is_private_integrity_clean_and_does_not_create_journal_siblings()
-> TestResult<()> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
let backup_path = append_to_path(&sqlite_path, BACKUP_SUFFIX);
sqlite_activation_backup::create_backup(database.pool(), &backup_path).await?;
sqlite_activation_backup::validate_backup(&backup_path).await?;
assert!(backup_path.is_file());
assert!(!append_to_path(&backup_path, "-wal").exists());
assert!(!append_to_path(&backup_path, "-shm").exists());
#[cfg(unix)]
{
use std::os::unix::fs::PermissionsExt as _;
assert_eq!(
std::fs::metadata(&backup_path)?.permissions().mode() & 0o077,
0
);
}
Ok(())
}
#[tokio::test]
async fn backup_publication_never_overwrites_an_existing_valid_backup() -> TestResult<()> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
let backup_path = append_to_path(&sqlite_path, BACKUP_SUFFIX);
sqlite_activation_backup::create_backup(database.pool(), &backup_path).await?;
let original = fs::read(&backup_path).await?;
sqlx::query("INSERT INTO blobs (hash, data) VALUES (?, ?)")
.bind(fabro_types::BlobHash::new(b"later").to_string())
.bind(b"later".as_slice())
.execute(database.pool())
.await?;
sqlite_activation_backup::create_backup(database.pool(), &backup_path).await?;
assert_eq!(fs::read(&backup_path).await?, original);
Ok(())
}
#[tokio::test]
async fn failed_backup_copy_never_publishes_a_destination() -> TestResult<()> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
let backup_path = append_to_path(&sqlite_path, BACKUP_SUFFIX);
database.pool().close().await;
let error = create_backup(database.pool(), &backup_path)
.await
.expect_err("a closed pool must fail backup creation");
assert!(matches!(
error,
BackupError::Stage(fabro_db::SnapshotStagingError::Write { .. })
));
assert!(!backup_path.exists());
Ok(())
}
#[tokio::test]
async fn cold_activation_and_warm_restart_share_verified_sqlite_blobs() -> TestResult<()> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
let object_store: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
let source = fabro_store::test_support::test_database(
Arc::clone(&object_store),
"activation-test",
Duration::from_millis(1),
None,
);
let legacy_bytes = b"legacy-blob";
let legacy_hash = fabro_store::test_support::put_legacy_blob(&source, legacy_bytes).await?;
drop(source);
let activation = activate_blob_storage(
&database,
&sqlite_path,
Arc::clone(&object_store),
"activation-test".to_string(),
Duration::from_millis(1),
None,
)
.await?;
let store = activation.store;
assert_eq!(
store.blobs().read(&legacy_hash).await?.as_deref(),
Some(legacy_bytes.as_slice())
);
let backup_path = append_to_path(&sqlite_path, BACKUP_SUFFIX);
let original_backup = fs::read(&backup_path).await?;
let run_id = fabro_types::RunId::new();
let writer = store.create_run(&run_id).await?;
let reader = store.open_run_reader(&run_id).await?;
let sqlite_only_bytes = b"written-after-activation";
let sqlite_only_hash = writer.write_blob(sqlite_only_bytes).await?;
assert_eq!(
reader.read_blob(&sqlite_only_hash).await?.as_deref(),
Some(sqlite_only_bytes.as_slice())
);
drop(reader);
drop(writer);
drop(store);
let warm = activate_blob_storage(
&database,
&sqlite_path,
object_store,
"activation-test".to_string(),
Duration::from_millis(1),
None,
)
.await?;
assert_eq!(fs::read(&backup_path).await?, original_backup);
assert_eq!(
warm.store.blobs().read(&sqlite_only_hash).await?.as_deref(),
Some(sqlite_only_bytes.as_slice())
);
Ok(())
}
#[tokio::test]
async fn missing_backup_after_prior_import_fails_closed() -> TestResult<()> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
let object_store: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
let source = fabro_store::test_support::test_database(
Arc::clone(&object_store),
"missing-backup-test",
Duration::from_millis(1),
None,
);
let bytes = b"already-imported";
let hash = fabro_store::test_support::put_legacy_blob(&source, bytes).await?;
sqlx::query("INSERT INTO blobs (hash, data) VALUES (?, ?)")
.bind(hash.to_string())
.bind(bytes.as_slice())
.execute(database.pool())
.await?;
drop(source);
let error = activate_blob_storage(
&database,
&sqlite_path,
object_store,
"missing-backup-test".to_string(),
Duration::from_millis(1),
None,
)
.await
.expect_err("startup must not move the pre-activation rollback boundary");
assert!(matches!(
error,
BlobActivationError::MissingBackupAfterImport {
legacy_rows: 1,
existing_rows: 1,
..
}
));
assert!(!append_to_path(&sqlite_path, BACKUP_SUFFIX).exists());
Ok(())
}
#[tokio::test]
async fn empty_inventory_skips_backup_and_serves_existing_sqlite_rows() -> TestResult<()> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
let bytes = b"sqlite-only";
let hash = fabro_types::BlobHash::new(bytes);
sqlx::query("INSERT INTO blobs (hash, data) VALUES (?, ?)")
.bind(hash.to_string())
.bind(bytes.as_slice())
.execute(database.pool())
.await?;
let activated = activate_blob_storage(
&database,
&sqlite_path,
Arc::new(InMemory::new()),
"empty-activation-test".to_string(),
Duration::from_millis(1),
None,
)
.await?;
assert!(!append_to_path(&sqlite_path, BACKUP_SUFFIX).exists());
assert_eq!(
activated.store.blobs().read(&hash).await?.as_deref(),
Some(bytes.as_slice())
);
Ok(())
}
#[tokio::test]
async fn busy_final_checkpoint_warns_and_does_not_fail_startup() -> TestResult<()> {
use sqlx::Connection as _;
use sqlx::sqlite::{
SqliteConnectOptions, SqliteConnection, SqliteJournalMode, SqlitePoolOptions,
};
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
sqlx::query("INSERT INTO blobs (hash, data) VALUES (?, ?)")
.bind(fabro_types::BlobHash::new(b"wal-content").to_string())
.bind(b"wal-content".as_slice())
.execute(database.pool())
.await?;
// A reader holding an open snapshot models a backup tool or operator
// shell that outlives the checkpoint's busy timeout.
let reader_options = SqliteConnectOptions::new()
.filename(&sqlite_path)
.read_only(true)
.create_if_missing(false);
let mut reader = SqliteConnection::connect_with(&reader_options).await?;
sqlx::query("BEGIN").execute(&mut reader).await?;
sqlx::query_scalar::<_, i64>("SELECT COUNT(*) FROM blobs")
.fetch_one(&mut reader)
.await?;
// A short busy timeout keeps the blocked truncate from stalling the
// test for the production pool's full five seconds.
let checkpoint_options = SqliteConnectOptions::new()
.filename(&sqlite_path)
.journal_mode(SqliteJournalMode::Wal)
.busy_timeout(Duration::from_millis(50))
.create_if_missing(false);
let checkpoint_pool = SqlitePoolOptions::new()
.max_connections(1)
.connect_with(checkpoint_options)
.await?;
final_truncate_checkpoint(&checkpoint_pool).await?;
// The reader really did block the truncate: the WAL was not reset.
let wal_bytes = fs::metadata(append_to_path(&sqlite_path, "-wal"))
.await?
.len();
assert!(wal_bytes > 0, "the WAL should remain untruncated");
drop(reader);
Ok(())
}
#[tokio::test]
async fn invalid_retained_backup_fails_before_importing() -> TestResult<()> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
let object_store: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
let source = fabro_store::test_support::test_database(
Arc::clone(&object_store),
"invalid-backup-test",
Duration::from_millis(1),
None,
);
fabro_store::test_support::put_legacy_blob(&source, b"must-not-import").await?;
drop(source);
let backup_path = append_to_path(&sqlite_path, BACKUP_SUFFIX);
fs::write(&backup_path, b"not a database").await?;
#[cfg(unix)]
{
use std::os::unix::fs::PermissionsExt as _;
fs::set_permissions(&backup_path, std::fs::Permissions::from_mode(0o600)).await?;
}
let error = activate_blob_storage(
&database,
&sqlite_path,
object_store,
"invalid-backup-test".to_string(),
Duration::from_millis(1),
None,
)
.await
.expect_err("an invalid retained backup must fail closed");
assert!(matches!(
error,
BlobActivationError::Backup(
BackupError::Integrity { .. } | BackupError::IntegrityFailed { .. }
)
));
let destination_rows: i64 = sqlx::query_scalar("SELECT COUNT(*) FROM blobs")
.fetch_one(database.pool())
.await?;
assert_eq!(destination_rows, 0);
Ok(())
}
}

View file

@ -1,603 +0,0 @@
//! Fail-closed activation of SQLite run history.
//!
//! This compatibility bridge remains for at least 30 days after the persisted
//! first-success timestamp, and until cold-start, warm-restart, production
//! observation, rollback-backup, deletion, and concurrent-reader evidence has
//! been accepted and Scott explicitly approves removal. The computed date is
//! an eligibility floor, never an automatic deletion trigger.
use std::path::{Path, PathBuf};
use chrono::{DateTime, Duration, Utc};
use tokio::fs;
use tracing::{info, warn};
use crate::migrations::sqlite_activation_backup::{self, BackupError};
pub(crate) const BACKUP_SUFFIX: &str = ".pre-run-history-activation.bak";
const REMOVAL_WINDOW: Duration = Duration::days(30);
#[derive(Clone, Debug, Eq, PartialEq)]
struct ActivationRecord {
source_fingerprint: Vec<u8>,
source_runs: u64,
source_events: u64,
activated_at_ms: i64,
}
#[derive(Debug, thiserror::Error)]
pub(crate) enum RunHistoryActivationError {
#[error("canonicalizing the SQLite database path {path}")]
Canonicalize {
path: PathBuf,
#[source]
source: std::io::Error,
},
#[error("reading the SQLite run-history activation state")]
ActivationState(#[source] sqlx::Error),
#[error("the persisted run-history activation marker does not match the legacy source")]
MarkerMismatch,
#[error("the persisted run-history activation marker contains an invalid count or timestamp")]
InvalidMarker,
#[error(
"SQLite contains {target_runs} run rows and {target_events} run events, but the legacy run-history source is empty and no activation marker exists"
)]
EmptySourceWithTarget {
target_runs: u64,
target_events: u64,
},
#[error(
"run-history activation backup is missing at {path} after SQLite import progress was recorded"
)]
MissingBackupAfterProgress { path: PathBuf },
#[error(transparent)]
Backup(#[from] BackupError),
#[error("importing legacy run history into SQLite")]
Import(#[source] Box<fabro_store::LegacyRunHistoryImportError>),
#[error("verifying legacy and SQLite run history")]
Verification(#[source] Box<fabro_store::LegacyRunHistoryVerificationError>),
#[error("running the live SQLite integrity check")]
LiveIntegrity(#[source] sqlx::Error),
#[error("the live SQLite integrity check failed")]
LiveIntegrityFailed,
#[error("persisting the SQLite run-history activation marker")]
PersistMarker(#[source] sqlx::Error),
#[error("the run-history activation timestamp is outside the supported range")]
InvalidActivationTimestamp,
#[error("running the final SQLite WAL truncate checkpoint")]
FinalCheckpoint(#[source] sqlx::Error),
#[error("a run-history activation count exceeds SQLite's integer range")]
CountOverflow,
}
pub(crate) async fn activate_run_history(
database: &fabro_db::Database,
sqlite_path: &Path,
store: &fabro_store::Database,
identity: &fabro_store::LegacyRunHistorySourceIdentity,
) -> Result<(), RunHistoryActivationError> {
let canonical_path = fs::canonicalize(sqlite_path).await.map_err(|source| {
RunHistoryActivationError::Canonicalize {
path: sqlite_path.to_path_buf(),
source,
}
})?;
let backup_path = fabro_db::append_to_path(&canonical_path, BACKUP_SUFFIX);
info!(
database_path = %canonical_path.display(),
backup_path = %backup_path.display(),
"Starting SQLite run-history activation"
);
let marker = read_activation_record(database.pool()).await?;
let (target_runs, target_events) = target_counts(database.pool()).await?;
if let Some(record) = &marker {
verify_marker(record, identity)?;
} else if identity.events == 0 && (target_runs != 0 || target_events != 0) {
return Err(RunHistoryActivationError::EmptySourceWithTarget {
target_runs,
target_events,
});
}
let backup_present = sqlite_activation_backup::backup_exists(&backup_path).await?;
if backup_present {
sqlite_activation_backup::validate_backup(&backup_path).await?;
}
let import_progress = target_events != 0 || marker.is_some();
if identity.events != 0 && import_progress && !backup_present {
return Err(RunHistoryActivationError::MissingBackupAfterProgress { path: backup_path });
}
let backup_required = identity.events != 0 && !backup_present;
if backup_required {
sqlite_activation_backup::create_backup(database.pool(), &backup_path).await?;
}
let import = store
.import_legacy_run_history_into(database.pool())
.await
.map_err(|source| RunHistoryActivationError::Import(Box::new(source)))?;
let verification = store
.verify_legacy_run_history_in(database.pool())
.await
.map_err(|source| RunHistoryActivationError::Verification(Box::new(source)))?;
validate_live_integrity(database.pool()).await?;
let activated_at_ms = marker.as_ref().map_or_else(
|| Utc::now().timestamp_millis(),
|record| record.activated_at_ms,
);
persist_activation_record(database.pool(), identity, activated_at_ms).await?;
final_truncate_checkpoint(database.pool()).await?;
let activated_at = DateTime::<Utc>::from_timestamp_millis(activated_at_ms)
.ok_or(RunHistoryActivationError::InvalidActivationTimestamp)?;
let removal_eligible_at = activated_at + REMOVAL_WINDOW;
info!(
source_runs = identity.runs,
source_events = identity.events,
imported_runs = import.imported_runs,
imported_events = import.imported_events,
existing_runs = import.verified_existing_runs,
existing_events = import.verified_existing_events,
tombstoned_source_runs = verification.tombstoned_source_runs,
tombstoned_source_events = verification.tombstoned_source_events,
target_runs = verification.target_runs,
target_events = verification.target_events,
sql_only_runs = verification.sql_only_runs,
sql_only_events = verification.sql_only_events,
backup_required,
backup_path = %backup_path.display(),
activated_at = %activated_at,
removal_eligible_at = %removal_eligible_at,
"Activated SQLite run history"
);
Ok(())
}
async fn read_activation_record(
pool: &sqlx::SqlitePool,
) -> Result<Option<ActivationRecord>, RunHistoryActivationError> {
let row = sqlx::query_as::<_, (Vec<u8>, i64, i64, i64)>(
r"
SELECT source_fingerprint, source_runs, source_events, activated_at_ms
FROM legacy_run_history_activation
WHERE singleton = 1
",
)
.fetch_optional(pool)
.await
.map_err(RunHistoryActivationError::ActivationState)?;
row.map(
|(source_fingerprint, source_runs, source_events, activated_at_ms)| {
Ok(ActivationRecord {
source_fingerprint,
source_runs: u64::try_from(source_runs)
.map_err(|_| RunHistoryActivationError::InvalidMarker)?,
source_events: u64::try_from(source_events)
.map_err(|_| RunHistoryActivationError::InvalidMarker)?,
activated_at_ms,
})
},
)
.transpose()
}
fn verify_marker(
marker: &ActivationRecord,
identity: &fabro_store::LegacyRunHistorySourceIdentity,
) -> Result<(), RunHistoryActivationError> {
if marker.activated_at_ms < 0 {
return Err(RunHistoryActivationError::InvalidMarker);
}
if marker.source_fingerprint.as_slice() != identity.fingerprint()
|| marker.source_runs != identity.runs
|| marker.source_events != identity.events
{
return Err(RunHistoryActivationError::MarkerMismatch);
}
Ok(())
}
async fn target_counts(pool: &sqlx::SqlitePool) -> Result<(u64, u64), RunHistoryActivationError> {
let (runs, events): (i64, i64) =
sqlx::query_as("SELECT (SELECT COUNT(*) FROM runs), (SELECT COUNT(*) FROM run_events)")
.fetch_one(pool)
.await
.map_err(RunHistoryActivationError::ActivationState)?;
Ok((
u64::try_from(runs).map_err(|_| RunHistoryActivationError::CountOverflow)?,
u64::try_from(events).map_err(|_| RunHistoryActivationError::CountOverflow)?,
))
}
async fn persist_activation_record(
pool: &sqlx::SqlitePool,
identity: &fabro_store::LegacyRunHistorySourceIdentity,
activated_at_ms: i64,
) -> Result<(), RunHistoryActivationError> {
let source_runs =
i64::try_from(identity.runs).map_err(|_| RunHistoryActivationError::CountOverflow)?;
let source_events =
i64::try_from(identity.events).map_err(|_| RunHistoryActivationError::CountOverflow)?;
// A pre-existing marker was already verified against `identity` above, so
// leaving it untouched on conflict keeps the original activation time.
sqlx::query(
r"
INSERT INTO legacy_run_history_activation (
singleton, source_fingerprint, source_runs, source_events, activated_at_ms
) VALUES (1, ?, ?, ?, ?)
ON CONFLICT(singleton) DO NOTHING
",
)
.bind(identity.fingerprint().as_slice())
.bind(source_runs)
.bind(source_events)
.bind(activated_at_ms)
.execute(pool)
.await
.map_err(RunHistoryActivationError::PersistMarker)?;
Ok(())
}
async fn validate_live_integrity(pool: &sqlx::SqlitePool) -> Result<(), RunHistoryActivationError> {
let ok = sqlite_activation_backup::integrity_check_is_ok(pool)
.await
.map_err(RunHistoryActivationError::LiveIntegrity)?;
if !ok {
return Err(RunHistoryActivationError::LiveIntegrityFailed);
}
Ok(())
}
async fn final_truncate_checkpoint(
pool: &sqlx::SqlitePool,
) -> Result<(), RunHistoryActivationError> {
let (busy, _, _): (i64, i64, i64) = sqlx::query_as("PRAGMA wal_checkpoint(TRUNCATE)")
.fetch_one(pool)
.await
.map_err(RunHistoryActivationError::FinalCheckpoint)?;
if busy != 0 {
// A concurrent reader can keep the WAL from truncating, but all
// activation data is already committed and remains durable in that
// WAL. A later checkpoint can truncate it after the reader exits.
warn!(
"The final SQLite run-history WAL truncate checkpoint could not complete; continuing startup"
);
}
Ok(())
}
#[cfg(test)]
mod tests {
use std::path::PathBuf;
use std::sync::Arc;
use std::time::Duration as StdDuration;
use chrono::{TimeZone as _, Utc};
use fabro_types::{Graph, RunId, WorkflowSettings, test_support};
use object_store::memory::InMemory;
use sqlx::Connection as _;
use tokio::fs;
use ulid::Ulid;
use super::{
BACKUP_SUFFIX, RunHistoryActivationError, activate_run_history, final_truncate_checkpoint,
read_activation_record,
};
type TestResult<T> = Result<T, Box<dyn std::error::Error>>;
struct TestContext {
_directory: tempfile::TempDir,
sqlite_path: PathBuf,
database: fabro_db::Database,
store: Arc<fabro_store::Database>,
}
impl TestContext {
async fn new(prefix: &str) -> TestResult<Self> {
let directory = tempfile::tempdir()?;
let sqlite_path = directory.path().join("fabro.sqlite3");
let database = fabro_db::Database::connect(&sqlite_path).await?;
database.migrate().await?;
let store = Arc::new(fabro_store::Database::new(
Arc::new(InMemory::new()),
prefix,
StdDuration::from_millis(1),
None,
Arc::new(fabro_store::BlobStore::new(database.clone_pool())),
Arc::new(fabro_store::RunSummaryStore::new(database.clone_pool())),
));
Ok(Self {
_directory: directory,
sqlite_path,
database,
store,
})
}
async fn put_event(
&self,
run_id: &RunId,
seq: u32,
event: &str,
properties: serde_json::Value,
) -> TestResult<()> {
let payload = serde_json::json!({
"id": format!("evt-{seq}-{event}"),
"ts": Utc
.timestamp_millis_opt(1_788_000_000_000 + i64::from(seq))
.single()
.unwrap()
.to_rfc3339(),
"run_id": run_id.to_string(),
"event": event,
"properties": properties,
});
fabro_store::test_support::put_legacy_run_event(&self.store, run_id, seq, &payload)
.await?;
Ok(())
}
async fn put_created(&self, run_id: &RunId) -> TestResult<()> {
self.put_event(
run_id,
1,
"run.created",
serde_json::json!({
"title": "Activation test",
"settings": WorkflowSettings::default(),
"graph": Graph::new("test"),
"workflow_slug": "test-workflow",
"labels": {},
"provenance": test_support::test_run_provenance(),
}),
)
.await
}
async fn source_identity(&self) -> TestResult<fabro_store::LegacyRunHistorySourceIdentity> {
Ok(self.store.legacy_run_history_source_identity().await?)
}
fn backup_path(&self) -> PathBuf {
fabro_db::append_to_path(&self.sqlite_path, BACKUP_SUFFIX)
}
}
fn run_id() -> RunId {
RunId::from(Ulid::from_parts(1_788_000_000_000, 1))
}
#[tokio::test]
async fn cold_activation_imports_and_warm_restart_preserves_marker() -> TestResult<()> {
let context = TestContext::new("cold-and-warm-run-activation").await?;
let run_id = run_id();
context.put_created(&run_id).await?;
context
.put_event(&run_id, 2, "run.submitted", serde_json::json!({}))
.await?;
let identity = context.source_identity().await?;
activate_run_history(
&context.database,
&context.sqlite_path,
&context.store,
&identity,
)
.await?;
let first_marker = read_activation_record(context.database.pool())
.await?
.unwrap();
assert_eq!(first_marker.source_runs, 1);
assert_eq!(first_marker.source_events, 2);
assert!(context.backup_path().is_file());
let backup_options = sqlx::sqlite::SqliteConnectOptions::new()
.filename(context.backup_path())
.read_only(true)
.create_if_missing(false);
let mut backup = sqlx::SqliteConnection::connect_with(&backup_options).await?;
assert_eq!(
sqlx::query_scalar::<_, i64>("SELECT COUNT(*) FROM run_events")
.fetch_one(&mut backup)
.await?,
0,
"the retained backup must capture the exact pre-import boundary"
);
assert_eq!(
sqlx::query_scalar::<_, i64>("SELECT COUNT(*) FROM run_events")
.fetch_one(context.database.pool())
.await?,
2
);
activate_run_history(
&context.database,
&context.sqlite_path,
&context.store,
&identity,
)
.await?;
assert_eq!(
read_activation_record(context.database.pool()).await?,
Some(first_marker)
);
Ok(())
}
#[tokio::test]
async fn busy_final_checkpoint_warns_and_does_not_fail_activation() -> TestResult<()> {
use sqlx::sqlite::{SqliteJournalMode, SqlitePoolOptions};
let context = TestContext::new("busy-final-run-checkpoint").await?;
sqlx::query("INSERT INTO blobs (hash, data) VALUES (?, ?)")
.bind(fabro_types::BlobHash::new(b"wal-content").to_string())
.bind(b"wal-content".as_slice())
.execute(context.database.pool())
.await?;
let reader_options = sqlx::sqlite::SqliteConnectOptions::new()
.filename(&context.sqlite_path)
.read_only(true)
.create_if_missing(false);
let mut reader = sqlx::SqliteConnection::connect_with(&reader_options).await?;
sqlx::query("BEGIN").execute(&mut reader).await?;
sqlx::query_scalar::<_, i64>("SELECT COUNT(*) FROM blobs")
.fetch_one(&mut reader)
.await?;
let checkpoint_options = sqlx::sqlite::SqliteConnectOptions::new()
.filename(&context.sqlite_path)
.journal_mode(SqliteJournalMode::Wal)
.busy_timeout(StdDuration::from_millis(50))
.create_if_missing(false);
let checkpoint_pool = SqlitePoolOptions::new()
.max_connections(1)
.connect_with(checkpoint_options)
.await?;
final_truncate_checkpoint(&checkpoint_pool).await?;
let wal_bytes = fs::metadata(fabro_db::append_to_path(&context.sqlite_path, "-wal"))
.await?
.len();
assert!(wal_bytes > 0, "the WAL should remain untruncated");
drop(reader);
Ok(())
}
#[tokio::test]
async fn changed_legacy_source_fails_before_mutating_sqlite() -> TestResult<()> {
let context = TestContext::new("changed-run-activation-source").await?;
let run_id = run_id();
context.put_created(&run_id).await?;
let original_identity = context.source_identity().await?;
activate_run_history(
&context.database,
&context.sqlite_path,
&context.store,
&original_identity,
)
.await?;
context
.put_event(&run_id, 2, "run.submitted", serde_json::json!({}))
.await?;
let changed_identity = context.source_identity().await?;
let error = activate_run_history(
&context.database,
&context.sqlite_path,
&context.store,
&changed_identity,
)
.await
.expect_err("the source identity must remain stable after activation");
assert!(matches!(error, RunHistoryActivationError::MarkerMismatch));
assert_eq!(
sqlx::query_scalar::<_, i64>("SELECT COUNT(*) FROM run_events")
.fetch_one(context.database.pool())
.await?,
1
);
Ok(())
}
#[tokio::test]
async fn empty_source_with_unmarked_target_fails_closed() -> TestResult<()> {
let context = TestContext::new("empty-source-with-target").await?;
let run_id = run_id();
context.put_created(&run_id).await?;
let identity = context.source_identity().await?;
activate_run_history(
&context.database,
&context.sqlite_path,
&context.store,
&identity,
)
.await?;
sqlx::query("DELETE FROM legacy_run_history_activation")
.execute(context.database.pool())
.await?;
let empty_store = Arc::new(fabro_store::Database::new(
Arc::new(InMemory::new()),
"empty-source",
StdDuration::from_millis(1),
None,
Arc::new(fabro_store::BlobStore::new(context.database.clone_pool())),
Arc::new(fabro_store::RunSummaryStore::new(
context.database.clone_pool(),
)),
));
let empty_identity = empty_store.legacy_run_history_source_identity().await?;
let error = activate_run_history(
&context.database,
&context.sqlite_path,
&empty_store,
&empty_identity,
)
.await
.expect_err("unmarked SQLite rows cannot be adopted from an empty source");
assert!(matches!(
error,
RunHistoryActivationError::EmptySourceWithTarget { .. }
));
Ok(())
}
#[tokio::test]
async fn missing_backup_after_import_progress_fails_closed() -> TestResult<()> {
let context = TestContext::new("missing-run-activation-backup").await?;
let run_id = run_id();
context.put_created(&run_id).await?;
context
.store
.import_legacy_run_history_into(context.database.pool())
.await?;
let identity = context.source_identity().await?;
let error = activate_run_history(
&context.database,
&context.sqlite_path,
&context.store,
&identity,
)
.await
.expect_err("partial import progress requires the retained backup");
assert!(matches!(
error,
RunHistoryActivationError::MissingBackupAfterProgress { .. }
));
Ok(())
}
#[tokio::test]
async fn empty_source_and_target_need_no_backup_on_cold_or_warm_start() -> TestResult<()> {
let context = TestContext::new("empty-run-activation").await?;
let identity = context.source_identity().await?;
activate_run_history(
&context.database,
&context.sqlite_path,
&context.store,
&identity,
)
.await?;
let marker = read_activation_record(context.database.pool())
.await?
.unwrap();
assert_eq!((marker.source_runs, marker.source_events), (0, 0));
assert!(!context.backup_path().exists());
activate_run_history(
&context.database,
&context.sqlite_path,
&context.store,
&identity,
)
.await?;
assert!(!context.backup_path().exists());
Ok(())
}
}

View file

@ -1,186 +0,0 @@
//! Pre-activation backup and integrity helpers shared by the SQLite
//! activation bridges.
//!
//! Every activation snapshots the live database to a private, integrity
//! checked backup file before importing legacy data, and re-validates any
//! backup it finds on a later start. This module owns that mechanism so the
//! blob and run-history bridges cannot drift apart.
use std::path::{Path, PathBuf};
use futures_util::TryStreamExt as _;
use sqlx::Connection as _;
use sqlx::sqlite::{SqliteConnectOptions, SqliteConnection};
use tokio::fs;
use tokio::task::{JoinError, spawn_blocking};
use tracing::debug;
const STAGING_SUFFIX: &str = ".tmp";
#[derive(Debug, thiserror::Error)]
pub(crate) enum BackupError {
#[error("reading activation backup metadata at {path}")]
Metadata {
path: PathBuf,
#[source]
source: std::io::Error,
},
#[error("activation backup is not a regular file at {path}")]
NotRegular { path: PathBuf },
#[error("activation backup permissions are not private at {path}")]
NotPrivate { path: PathBuf },
#[error("opening or checking activation backup integrity at {path}")]
Integrity {
path: PathBuf,
#[source]
source: sqlx::Error,
},
#[error("activation backup integrity check did not return exactly one ok result at {path}")]
IntegrityFailed { path: PathBuf },
#[error("staging the pre-activation SQLite backup")]
Stage(#[source] fabro_db::SnapshotStagingError),
#[error("joining the activation backup publication task")]
JoinPublication(#[source] JoinError),
#[error("publishing the activation backup at {path} without overwriting")]
Publish {
path: PathBuf,
#[source]
source: std::io::Error,
},
}
pub(crate) async fn backup_exists(path: &Path) -> Result<bool, BackupError> {
match fs::metadata(path).await {
Ok(_) => Ok(true),
Err(source) if source.kind() == std::io::ErrorKind::NotFound => Ok(false),
Err(source) => Err(BackupError::Metadata {
path: path.to_path_buf(),
source,
}),
}
}
/// Snapshots `pool` to `backup_path` without overwriting an existing file.
///
/// The snapshot is staged beside the target, validated, and then published
/// with an atomic no-clobber rename. A backup that another process published
/// concurrently is validated in place instead.
pub(crate) async fn create_backup(
pool: &sqlx::SqlitePool,
backup_path: &Path,
) -> Result<(), BackupError> {
let staging_path = fabro_db::append_to_path(backup_path, STAGING_SUFFIX);
fabro_db::write_snapshot_to_staging(pool, &staging_path)
.await
.map_err(BackupError::Stage)?;
validate_backup(&staging_path).await?;
let publish_staging = staging_path.clone();
let publish_backup = backup_path.to_path_buf();
let already_exists = spawn_blocking(move || {
let staging = tempfile::TempPath::try_from_path(publish_staging)?;
match staging.persist_noclobber(&publish_backup) {
Ok(()) => {
// Make the rename's directory entry durable: the retained
// backup is the documented rollback artifact, so it must not
// vanish in a crash after the import has already committed.
fabro_db::sync_parent_directory(&publish_backup)?;
Ok(false)
}
Err(error) if error.error.kind() == std::io::ErrorKind::AlreadyExists => Ok(true),
Err(error) => Err(error.error),
}
})
.await
.map_err(BackupError::JoinPublication)?
.map_err(|source| BackupError::Publish {
path: backup_path.to_path_buf(),
source,
})?;
// The staging copy was validated just before the atomic rename, so only a
// concurrently published file still needs its own validation.
if already_exists {
debug!(
backup_path = %backup_path.display(),
"Reusing concurrently published SQLite activation backup"
);
validate_backup(backup_path).await?;
}
Ok(())
}
/// Requires `path` to be a private regular file holding a SQLite database
/// whose `PRAGMA integrity_check` passes.
pub(crate) async fn validate_backup(path: &Path) -> Result<(), BackupError> {
let metadata = fs::symlink_metadata(path)
.await
.map_err(|source| BackupError::Metadata {
path: path.to_path_buf(),
source,
})?;
if !metadata.is_file() {
return Err(BackupError::NotRegular {
path: path.to_path_buf(),
});
}
validate_private_permissions(path, &metadata)?;
let options = SqliteConnectOptions::new()
.filename(path)
.read_only(true)
.immutable(true)
.create_if_missing(false);
let mut connection = SqliteConnection::connect_with(&options)
.await
.map_err(|source| BackupError::Integrity {
path: path.to_path_buf(),
source,
})?;
let ok = integrity_check_is_ok(&mut connection)
.await
.map_err(|source| BackupError::Integrity {
path: path.to_path_buf(),
source,
})?;
if !ok {
return Err(BackupError::IntegrityFailed {
path: path.to_path_buf(),
});
}
Ok(())
}
/// Returns whether `PRAGMA integrity_check` reports exactly one `ok` row.
pub(crate) async fn integrity_check_is_ok<'a, E>(executor: E) -> Result<bool, sqlx::Error>
where
E: sqlx::Executor<'a, Database = sqlx::Sqlite>,
{
let mut rows = sqlx::query_scalar::<_, String>("PRAGMA integrity_check").fetch(executor);
let first = rows.try_next().await?;
let second = rows.try_next().await?;
Ok(first.as_deref() == Some("ok") && second.is_none())
}
#[cfg(unix)]
fn validate_private_permissions(
path: &Path,
metadata: &std::fs::Metadata,
) -> Result<(), BackupError> {
use std::os::unix::fs::PermissionsExt as _;
if metadata.permissions().mode() & 0o077 != 0 {
return Err(BackupError::NotPrivate {
path: path.to_path_buf(),
});
}
Ok(())
}
#[cfg(not(unix))]
fn validate_private_permissions(
_path: &Path,
_metadata: &std::fs::Metadata,
) -> Result<(), BackupError> {
Ok(())
}

View file

@ -31,7 +31,6 @@ pub mod install;
mod interp;
pub mod jwt_auth;
pub mod manifest_validation;
mod migrations;
mod petri_check;
mod petri_runs;
mod principal_middleware;

View file

@ -1,9 +0,0 @@
#[path = "../migrations/sqlite_activation_backup.rs"]
mod sqlite_activation_backup;
#[path = "../migrations/2026082301_sqlite_blob_activation.rs"]
mod sqlite_blob_activation;
#[path = "../migrations/2026082801_sqlite_run_history_activation.rs"]
mod sqlite_run_history_activation;
pub(crate) use sqlite_blob_activation::activate_blob_storage;
pub(crate) use sqlite_run_history_activation::activate_run_history;

View file

@ -164,8 +164,8 @@ mod tests {
use fabro_config::daemon::ServerDaemon;
use fabro_petri::petri::RunStore as _;
use fabro_static::EnvVars;
use fabro_store::platform_records::{PlatformRecord, RunLifecycleKind, RunLifecycleRecord};
use fabro_types::{RunId, RunStatus, WorkflowPath, WorkflowVersion};
use fabro_workflow::event::{Event, append_event};
use serde_json::json;
use tokio::io::AsyncRead;
use tokio::sync::Notify;
@ -173,7 +173,9 @@ mod tests {
use tower::ServiceExt as _;
use super::*;
use crate::server::{AppState, reconcile_incomplete_runs_on_startup, spawn_scheduler};
use crate::server::{
AppState, reconcile_incomplete_runs_on_startup, run_records, spawn_scheduler,
};
use crate::test_support::{
TestAppStateBuilder, build_test_router, test_register_workflow_version,
test_secret_store_path, test_store_bundle,
@ -414,16 +416,17 @@ mod tests {
// The worker took the run as far as running and holds its lease;
// then the server died, so nothing released it.
let run_store = before
.stores
.runs
.open_run(&run_id)
for (transition, status) in [
(RunLifecycleKind::Starting, RunStatus::Starting),
(RunLifecycleKind::Running, RunStatus::Running),
] {
run_records::lifecycle(
&before,
run_id,
RunLifecycleRecord::new(transition).with_status(status),
)
.await
.expect("the run opens");
for event in [Event::RunStarting, Event::RunRunning] {
append_event(&run_store, &run_id, &event)
.await
.expect("the lifecycle event appends");
.expect("the lifecycle record appends");
}
let held = before
.petri_runs
@ -465,30 +468,33 @@ mod tests {
None,
"the previous worker's lease is released"
);
let reader = after
.stores
.runs
.open_run_reader(&run_id)
let run_state = run_records::projection(&after, run_id)
.await
.expect("the run opens for reading");
let run_state = reader.state().await.expect("the run state loads");
.expect("the run state loads")
.expect("the run projects");
assert_eq!(run_state.status, RunStatus::Runnable);
let names = reader
.list_events()
let transitions = after
.stores
.run_summaries
.platform_records()
.read(&run_id)
.await
.expect("the history lists")
.expect("the records list")
.into_iter()
.map(|envelope| envelope.event.event_name().to_string())
.filter_map(|stored| match stored.record {
PlatformRecord::RunLifecycle(record) => Some(record.transition),
_ => None,
})
.collect::<Vec<_>>();
assert_eq!(
&names[names.len() - 4..],
&transitions[transitions.len() - 4..],
[
"run.starting",
"run.running",
"run.start_requested",
"run.runnable"
RunLifecycleKind::Starting,
RunLifecycleKind::Running,
RunLifecycleKind::StartRequested,
RunLifecycleKind::Runnable
],
"{names:?}"
"{transitions:?}"
);
write_test_server_record(&after);

View file

@ -40,7 +40,7 @@ use crate::server::{
};
use crate::server_secrets::{ServerSecrets, process_env_snapshot};
use crate::startup::{resolve_startup, validate_startup_configuration};
use crate::{migrations, static_files};
use crate::static_files;
pub const DEFAULT_TCP_PORT: u16 = 32276;
type EnvLookup = Arc<dyn Fn(&str) -> Option<String> + Send + Sync>;
@ -748,25 +748,14 @@ where
} else {
None
};
let blob_activation = migrations::activate_blob_storage(
&database,
&sqlite_path,
let store = Arc::new(fabro_store::Database::new(
object_store,
slatedb_prefix,
flush_interval,
cache_path,
)
.await
.context("activating SQLite blob storage")?;
migrations::activate_run_history(
&database,
&sqlite_path,
&blob_activation.store,
&blob_activation.run_history_identity,
)
.await
.context("activating SQLite run history")?;
let store = blob_activation.store;
Arc::new(fabro_store::BlobStore::new(database.clone_pool())),
Arc::new(fabro_store::RunSummaryStore::new(database.clone_pool())),
));
// Refresh tokens now live in SQLite. Nothing reads the old records and no
// reaper collects them any more, so clear them out once rather than
// leaving them in the object store forever. Pending authorization codes

File diff suppressed because it is too large Load diff

View file

@ -664,8 +664,10 @@ mod tests {
assert_eq!(automation_ref.name.as_deref(), Some("Nightly"));
assert_eq!(automation_ref.trigger_id.as_deref(), Some("schedule"));
let run_id = runs[0].id;
let run_store = state.stores.runs.open_run_reader(&run_id).await.unwrap();
let projection = run_store.state().await.unwrap();
let projection = super::super::run_records::projection(&state, run_id)
.await
.unwrap()
.expect("the run projects");
assert!(projection.spec.workflow_version_id.is_some());
assert_eq!(
projection.spec.target,
@ -676,16 +678,24 @@ mod tests {
sha: Some("0123456789abcdef0123456789abcdef01234567".to_string()),
}))
);
assert_eq!(
run_store
.list_events()
.await
.unwrap()
.iter()
.filter(|event| event.event.event_name() == "run.start_requested")
.count(),
1
);
let start_requests = state
.stores
.run_summaries
.platform_records()
.read(&run_id)
.await
.unwrap()
.into_iter()
.filter(|stored| {
matches!(
&stored.record,
fabro_store::platform_records::PlatformRecord::RunLifecycle(record)
if record.transition
== fabro_store::platform_records::RunLifecycleKind::StartRequested
)
})
.count();
assert_eq!(start_requests, 1);
assert!(matches!(
state
.runs

View file

@ -103,14 +103,14 @@ async fn write_run_blob(
if let Some(response) = reject_if_archived(state.as_ref(), &id).await {
return response;
}
match state.stores.runs.open_run(&id).await {
Ok(run_store) => match run_store.write_blob(&body).await {
Ok(blob_hash) => Json(WriteBlobResponse { hash: blob_hash }).into_response(),
Err(err) => {
ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response()
}
},
Err(_) => ApiError::not_found("Run not found.").into_response(),
if let Err(err) = state.load_run_projection(&id).await {
return err.into_response();
}
match state.store_ref().blobs().write(&body).await {
Ok(blob_hash) => Json(WriteBlobResponse { hash: blob_hash }).into_response(),
Err(err) => {
ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response()
}
}
}
@ -118,15 +118,15 @@ async fn read_run_blob(
RequireRunBlob(id, blob_hash): RequireRunBlob,
State(state): State<Arc<AppState>>,
) -> Response {
match state.stores.runs.open_run_reader(&id).await {
Ok(run_store) => match run_store.read_blob(&blob_hash).await {
Ok(Some(bytes)) => octet_stream_response(bytes),
Ok(None) => ApiError::not_found("Blob not found.").into_response(),
Err(err) => {
ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response()
}
},
Err(_) => ApiError::not_found("Run not found.").into_response(),
if let Err(err) = state.load_run_projection(&id).await {
return err.into_response();
}
match state.store_ref().blobs().read(&blob_hash).await {
Ok(Some(bytes)) => octet_stream_response(bytes),
Ok(None) => ApiError::not_found("Blob not found.").into_response(),
Err(err) => {
ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response()
}
}
}

File diff suppressed because it is too large Load diff

View file

@ -2,6 +2,7 @@ use std::collections::HashSet;
use std::sync::Arc;
use chrono::Utc;
use fabro_store::platform_records::{PlatformRecord, RunLifecycleKind, RunLifecycleRecord};
use tokio::time::{Instant, sleep_until};
use super::super::{
@ -13,9 +14,9 @@ use super::super::{
RequireRunManagementTarget, RequiredUser, Response, Router, RunAnswerTransport,
RunControlAction, RunExecutionMode, RunId, RunRunnableSource, RunStatus, StartRunRequest,
State, StatusCode, Storage, WORKER_CANCEL_GRACE, WorkflowError, append_control_request,
clear_live_run_state, delete_run_internal, durable_run_status, load_pending_control,
managed_run, operations, parse_run_id_path, persist_cancelled_run_status, post,
reject_if_archived, update_live_run_from_event, workflow_event,
apply_lifecycle_to_managed_run, clear_live_run_state, delete_run_internal, durable_run_status,
load_pending_control, managed_run, parse_run_id_path, persist_cancelled_run_status, post,
reject_if_archived, run_records,
};
use crate::worker_runtime::WorkerRef;
@ -97,18 +98,7 @@ pub(in crate::server) async fn queue_run_start(
}
}
let Ok(run_store) = state.stores.runs.open_run(&id).await else {
return Err(ApiError::not_found("Run not found."));
};
let run_state = match run_store.state().await {
Ok(state) => state,
Err(err) => {
return Err(ApiError::new(
StatusCode::INTERNAL_SERVER_ERROR,
format!("Failed to load run state: {err}"),
));
}
};
let run_state = run_records::require_projection(state, id).await?;
if resume {
if run_state.current_checkpoint().is_none() {
@ -137,39 +127,27 @@ pub(in crate::server) async fn queue_run_start(
&actor,
Principal::Worker { run_id } if run_state.parent_id == Some(*run_id)
);
if let Err(err) =
workflow_event::append_event(&run_store, &id, &workflow_event::Event::RunStartRequested {
resume,
actor: Some(actor.clone()),
})
.await
{
return Err(ApiError::new(
StatusCode::INTERNAL_SERVER_ERROR,
err.to_string(),
));
}
let (next_status, next_event) = if approval_required {
(
RunStatus::Pending {
reason: PendingReason::ApprovalRequired,
},
workflow_event::Event::RunPending {
reason: PendingReason::ApprovalRequired,
actor: Some(actor),
},
)
let mut start_requested = RunLifecycleRecord::new(RunLifecycleKind::StartRequested);
start_requested.source = Some(if resume { "resume" } else { "start" }.to_string());
let next_status = if approval_required {
RunStatus::Pending {
reason: PendingReason::ApprovalRequired,
}
} else {
(RunStatus::Runnable, workflow_event::Event::RunRunnable {
source: RunRunnableSource::StartRequested,
actor: Some(actor),
})
RunStatus::Runnable
};
if let Err(err) = workflow_event::append_event(&run_store, &id, &next_event).await {
return Err(ApiError::new(
StatusCode::INTERNAL_SERVER_ERROR,
err.to_string(),
));
let next = if approval_required {
run_records::transition(RunLifecycleKind::Pending, next_status)
} else {
runnable(RunRunnableSource::StartRequested)
};
for record in [start_requested, next] {
if let Err(err) = run_records::lifecycle(state, id, record).await {
return Err(ApiError::new(
StatusCode::INTERNAL_SERVER_ERROR,
err.to_string(),
));
}
}
{
@ -208,36 +186,22 @@ async fn approve_run(
if let Some(response) = reject_if_archived(state.as_ref(), &id).await {
return response;
}
let Ok(run_store) = state.stores.runs.open_run(&id).await else {
return ApiError::not_found("Run not found.").into_response();
};
let run_state = match run_store.state().await {
Ok(state) => state,
Err(err) => {
return ApiError::new(
StatusCode::INTERNAL_SERVER_ERROR,
format!("Failed to load run state: {err}"),
)
.into_response();
}
let run_state = match run_records::require_projection(state.as_ref(), id).await {
Ok(run_state) => run_state,
Err(err) => return err.into_response(),
};
if !matches!(run_state.status, RunStatus::Pending {
reason: PendingReason::ApprovalRequired,
}) {
return ApiError::new(StatusCode::CONFLICT, "Run is not pending approval.").into_response();
}
let _ = user;
let actor = Some(Principal::User(user));
for event in [
workflow_event::Event::RunApproved {
actor: actor.clone(),
},
workflow_event::Event::RunRunnable {
source: RunRunnableSource::Approved,
actor,
},
for record in [
RunLifecycleRecord::new(RunLifecycleKind::Approved),
runnable(RunRunnableSource::Approved),
] {
if let Err(err) = workflow_event::append_event(&run_store, &id, &event).await {
if let Err(err) = run_records::lifecycle(state.as_ref(), id, record).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string())
.into_response();
}
@ -290,44 +254,27 @@ async fn deny_run(
let message = reason
.clone()
.unwrap_or_else(|| "Not approved for execution".to_string());
let Ok(run_store) = state.stores.runs.open_run(&id).await else {
return ApiError::not_found("Run not found.").into_response();
};
let run_state = match run_store.state().await {
Ok(state) => state,
Err(err) => {
return ApiError::new(
StatusCode::INTERNAL_SERVER_ERROR,
format!("Failed to load run state: {err}"),
)
.into_response();
}
let run_state = match run_records::require_projection(state.as_ref(), id).await {
Ok(run_state) => run_state,
Err(err) => return err.into_response(),
};
if !matches!(run_state.status, RunStatus::Pending {
reason: PendingReason::ApprovalRequired,
}) {
return ApiError::new(StatusCode::CONFLICT, "Run is not pending approval.").into_response();
}
let _ = user;
let actor = Some(Principal::User(user));
let denied_event = workflow_event::Event::RunDenied {
reason: reason.clone(),
actor,
};
if let Err(err) = workflow_event::append_event(&run_store, &id, &denied_event).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response();
}
let failure_event = workflow_event::Event::workflow_run_failed_from_error(
&WorkflowError::engine(message.clone()),
fabro_types::RunTiming::default(),
FailureReason::ApprovalDenied,
None,
None,
None,
None,
);
if let Err(err) = workflow_event::append_event(&run_store, &id, &failure_event).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response();
let mut denied = RunLifecycleRecord::new(RunLifecycleKind::Denied);
denied.reason.clone_from(&reason);
for record in [
denied,
run_records::failed(FailureReason::ApprovalDenied, message.clone()),
] {
if let Err(err) = run_records::lifecycle(state.as_ref(), id, record).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string())
.into_response();
}
}
{
@ -688,9 +635,11 @@ async fn pause_run(
}
}
PauseMode::AppendEvent => {
if let Some(response) = synchronous_transition(state.as_ref(), id, |events| {
events.push(workflow_event::Event::RunPaused);
})
if let Some(response) = synchronous_transition(
state.as_ref(),
id,
RunLifecycleRecord::new(RunLifecycleKind::Paused),
)
.await
{
return response;
@ -771,9 +720,11 @@ async fn unpause_run(
}
}
UnpauseMode::AppendEvent => {
if let Some(response) = synchronous_transition(state.as_ref(), id, |events| {
events.push(workflow_event::Event::RunUnpaused);
})
if let Some(response) = synchronous_transition(
state.as_ref(),
id,
RunLifecycleRecord::new(RunLifecycleKind::Unpaused),
)
.await
{
return response;
@ -1033,34 +984,54 @@ fn batch_result_failure(
}
}
/// Archive a terminal run, or unarchive one: idempotent either way, refused
/// with a precondition error when the run is not terminal.
async fn run_archive_operation(
state: &AppState,
id: &RunId,
actor: Option<Principal>,
action: ArchiveAction,
) -> Result<BatchRunLifecycleResultOutcome, WorkflowError> {
match action {
ArchiveAction::Archive => operations::archive(&state.stores.runs, id, actor)
.await
.map(|outcome| match outcome {
operations::ArchiveOutcome::Archived { .. } => {
BatchRunLifecycleResultOutcome::Archived
}
operations::ArchiveOutcome::AlreadyArchived => {
BatchRunLifecycleResultOutcome::AlreadyArchived
}
}),
ArchiveAction::Unarchive => operations::unarchive(&state.stores.runs, id, actor)
.await
.map(|outcome| match outcome {
operations::UnarchiveOutcome::Unarchived { .. } => {
BatchRunLifecycleResultOutcome::Unarchived
}
operations::UnarchiveOutcome::NotArchived { .. } => {
BatchRunLifecycleResultOutcome::NotArchived
}
}),
}
let _ = actor;
let projection = run_records::projection(state, *id)
.await
.map_err(|err| WorkflowError::engine(err.to_string()))?
.ok_or_else(|| WorkflowError::RunNotFound(id.to_string()))?;
let current = projection.status;
let terminal = matches!(
current,
RunStatus::Succeeded { .. } | RunStatus::Failed { .. } | RunStatus::Dead
);
let archived = projection.archived_at.is_some();
let record = match action {
ArchiveAction::Archive if archived => {
return Ok(BatchRunLifecycleResultOutcome::AlreadyArchived);
}
ArchiveAction::Archive if !terminal => {
return Err(WorkflowError::Precondition(format!(
"run {id} must be terminal (succeeded, failed, or dead) to archive; current \
status is {current}"
)));
}
ArchiveAction::Archive => PlatformRecord::RunArchived,
ArchiveAction::Unarchive if archived => PlatformRecord::RunUnarchived,
ArchiveAction::Unarchive if terminal => {
return Ok(BatchRunLifecycleResultOutcome::NotArchived);
}
ArchiveAction::Unarchive => {
return Err(WorkflowError::Precondition(format!(
"run {id} is not archived (status: {current}); nothing to unarchive"
)));
}
};
let outcome = match action {
ArchiveAction::Archive => BatchRunLifecycleResultOutcome::Archived,
ArchiveAction::Unarchive => BatchRunLifecycleResultOutcome::Unarchived,
};
run_records::append(state, *id, record)
.await
.map_err(|err| WorkflowError::engine(err.to_string()))?;
Ok(outcome)
}
fn archive_workflow_error_to_api_error(err: WorkflowError) -> ApiError {
@ -1087,32 +1058,26 @@ async fn archive_status_response(state: &AppState, id: RunId) -> Response {
run_response(state, id, StatusCode::OK).await
}
/// Persist a synchronous pause/unpause transition: append the caller-supplied
/// events to the run store and mirror the new status in the in-memory run map.
/// Returns `Some(Response)` on error, `None` on success.
/// The runnable transition, with what made the run runnable.
fn runnable(source: RunRunnableSource) -> RunLifecycleRecord {
let mut record = run_records::transition(RunLifecycleKind::Runnable, RunStatus::Runnable);
record.source = Some(<&'static str>::from(source).to_string());
record
}
/// Persist a synchronous pause/unpause transition: record it and mirror the
/// new status in the in-memory run map. Returns `Some(Response)` on error,
/// `None` on success.
async fn synchronous_transition(
state: &AppState,
id: RunId,
append_events: impl FnOnce(&mut Vec<workflow_event::Event>),
record: RunLifecycleRecord,
) -> Option<Response> {
let run_store = match state.stores.runs.open_run(&id).await {
Ok(run_store) => run_store,
Err(err) => {
return Some(
ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response(),
);
}
};
let mut events = Vec::new();
append_events(&mut events);
for event in events {
if let Err(err) = workflow_event::append_event(&run_store, &id, &event).await {
return Some(
ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response(),
);
}
let stored = workflow_event::to_run_event(&id, &event);
update_live_run_from_event(state, id, &stored);
if let Err(err) = run_records::lifecycle(state, id, record.clone()).await {
return Some(
ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response(),
);
}
apply_lifecycle_to_managed_run(state, id, &record);
None
}

View file

@ -2,12 +2,15 @@ use std::sync::Arc;
use std::time::Duration;
use axum::http::{HeaderValue, header};
use fabro_store::platform_records::{
PlatformRecord, PullRequestLinkedRecord, PullRequestRequestedRecord,
};
use super::super::{
ApiError, AppState, CloseRunPullRequestResponse, CreateRunPullRequestRequest, IntoResponse,
Json, LinkRunPullRequestRequest, MergeRunPullRequestRequest, MergeRunPullRequestResponse,
PullRequestLink, RequireRunScoped, Response, Router, RunId, State, StatusCode, get, post, warn,
workflow_event,
PullRequestLink, RequireRunScoped, Response, Router, RunId, State, StatusCode, get, post,
run_records, warn,
};
pub(super) fn routes() -> Router<Arc<AppState>> {
@ -316,9 +319,6 @@ async fn create_run_pull_request(
State(state): State<Arc<AppState>>,
Json(body): Json<CreateRunPullRequestRequest>,
) -> Response {
let Ok(run_store) = state.stores.runs.open_run(&id).await else {
return ApiError::not_found("Run not found.").into_response();
};
let run_state = match state.load_run_projection(&id).await {
Ok(run_state) => run_state,
Err(err) => return err.into_response(),
@ -354,26 +354,28 @@ async fn create_run_pull_request(
};
let _create_guard = state.pull_request_create_locks.lock(id).await;
let creation_id = fabro_types::PullRequestCreationId::new();
let event = workflow_event::Event::PullRequestCreationRequested {
creation_id,
model,
force: body.force,
// Under the create lock, the projection is the latest word on whether a
// pull request exists or a creation is already pending.
let run_state = match state.load_run_projection(&id).await {
Ok(run_state) => run_state,
Err(err) => return err.into_response(),
};
let appended = match workflow_event::append_event_if(&run_store, &id, &event, |projection| {
projection.pull_request.is_none()
&& !projection
.pull_request_creation
.as_ref()
.is_some_and(fabro_types::PullRequestCreation::is_pending)
})
.await
{
Ok(appended) => appended,
Err(err) => {
let appended = run_state.pull_request.is_none()
&& !run_state
.pull_request_creation
.as_ref()
.is_some_and(fabro_types::PullRequestCreation::is_pending);
if appended {
let record = PlatformRecord::PullRequestRequested(PullRequestRequestedRecord {
creation_id,
model,
force: body.force,
});
if let Err(err) = run_records::append(&state, id, record).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string())
.into_response();
}
};
}
let run_state = match state.load_run_projection(&id).await {
Ok(run_state) => run_state,
@ -455,13 +457,12 @@ async fn link_run_pull_request(
Ok(record) => record,
Err(err) => return err.into_response(),
};
let Ok(run_store) = state.stores.runs.open_run(&id).await else {
return ApiError::not_found("Run not found.").into_response();
};
let event = workflow_event::Event::PullRequestLinked {
pull_request: pull_request.clone(),
};
if let Err(err) = workflow_event::append_event(&run_store, &id, &event).await {
if let Err(err) = state.load_run_projection(&id).await {
return err.into_response();
}
let record =
PlatformRecord::PullRequestLinked(PullRequestLinkedRecord::from_link(&pull_request));
if let Err(err) = run_records::append(&state, id, record).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response();
}
@ -473,9 +474,6 @@ async fn unlink_run_pull_request(
State(state): State<Arc<AppState>>,
) -> Response {
let _create_guard = state.pull_request_create_locks.lock(id).await;
let Ok(run_store) = state.stores.runs.open_run(&id).await else {
return ApiError::not_found("Run not found.").into_response();
};
let run_state = match state.load_run_projection(&id).await {
Ok(run_state) => run_state,
Err(err) => return err.into_response(),
@ -488,10 +486,9 @@ async fn unlink_run_pull_request(
)
.into_response();
};
let event = workflow_event::Event::PullRequestUnlinked {
pull_request: pull_request.clone(),
};
if let Err(err) = workflow_event::append_event(&run_store, &id, &event).await {
let record =
PlatformRecord::PullRequestUnlinked(PullRequestLinkedRecord::from_link(&pull_request));
if let Err(err) = run_records::append(&state, id, record).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response();
}

View file

@ -23,6 +23,7 @@ use fabro_interview::AnswerSubmission;
use fabro_llm::Client as LlmClient;
use fabro_manifest::RunOverrideInput;
use fabro_static::EnvVars;
use fabro_store::platform_records::{PlatformRecord, RunParentRecord, RunTitleRecord};
use fabro_store::{
RunSummaryListQuery, RunSummarySort, RunSummarySortDirection, RunSummaryVisibility,
};
@ -32,8 +33,7 @@ use fabro_types::{
AutomationRef, ContextWindowStaleness, ManifestPath, Principal, Run, RunClientProvenance,
RunId, RunProvenance, RunServerProvenance, RunStatusKind, RunTarget, SandboxProviderKind,
StageContextWindow, StageContextWindowUnavailableReason, StageHandler, StageModelUsage,
StageProjection, SystemActorKind, ValidatedRunTarget, json_scalar_to_toml_value,
parse_blob_ref,
StageProjection, ValidatedRunTarget, json_scalar_to_toml_value, parse_blob_ref,
};
use fabro_util::error as error_util;
use fabro_util::version::FABRO_VERSION;
@ -50,8 +50,8 @@ use super::super::{
AppState, DeleteRunOutcome, ListResponse, RunExecutionMode, VariableError, answer_from_request,
api_question_from_pending_interview, clamp_page_limit, clamp_page_offset, default_page_limit,
delete_run_internal, load_pending_interview, managed_run, parse_run_id_path,
parse_stage_id_path, petri_runs, reject_if_archived, submit_pending_interview_answer,
workflow_event,
parse_stage_id_path, petri_runs, reject_if_archived, run_records,
submit_pending_interview_answer,
};
use crate::error::ApiError;
use crate::principal_middleware::{
@ -213,20 +213,12 @@ async fn link_run_parent(
.into_response();
}
let Ok(run_store) = state.stores.runs.open_run(&child_id).await else {
return ApiError::not_found("Run not found.").into_response();
};
if let Err(err) = workflow_event::append_event(
&run_store,
&child_id,
&workflow_event::Event::RunParentLinked {
previous_parent_id: child.parent_id,
parent_id,
actor: Some(actor),
},
)
.await
{
let _ = actor;
let record = PlatformRecord::RunParent(RunParentRecord {
parent_id: Some(parent_id),
previous_parent_id: child.parent_id,
});
if let Err(err) = run_records::append(&state, child_id, record).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response();
}
updated_run_response(&state, &child_id).await
@ -253,19 +245,12 @@ async fn unlink_run_parent(
.into_response();
};
let Ok(run_store) = state.stores.runs.open_run(&child_id).await else {
return ApiError::not_found("Run not found.").into_response();
};
if let Err(err) = workflow_event::append_event(
&run_store,
&child_id,
&workflow_event::Event::RunParentUnlinked {
previous_parent_id,
actor: Some(actor),
},
)
.await
{
let _ = actor;
let record = PlatformRecord::RunParent(RunParentRecord {
parent_id: None,
previous_parent_id: Some(previous_parent_id),
});
if let Err(err) = run_records::append(&state, child_id, record).await {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response();
}
updated_run_response(&state, &child_id).await
@ -491,19 +476,13 @@ async fn update_run(
.into_response();
}
let run_store = match state.stores.runs.open_run(&id).await {
Ok(run_store) => run_store,
Err(err) => {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string())
.into_response();
}
};
if let Err(err) =
workflow_event::append_event(&run_store, &id, &workflow_event::Event::RunTitleUpdated {
title,
actor: Some(Principal::User(subject.0)),
})
.await
let _ = subject;
if let Err(err) = run_records::append(
&state,
id,
PlatformRecord::RunTitle(RunTitleRecord { title }),
)
.await
{
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()).into_response();
}
@ -799,6 +778,9 @@ async fn finalize_created_run(
}
};
let created_at = created.run_id.created_at();
// The run's summary row is the projector's: wait for the pass that
// folds the run's first records before reading the run back.
state.petri_projector.settle(created.run_id).await;
let summary = match state
.stores
.run_summaries
@ -1138,28 +1120,29 @@ fn spawn_generated_title_task(task: GeneratedTitleTask) {
return;
}
let run_store = match task.state.stores.runs.open_run(&task.run_id).await {
Ok(store) => store,
// The generated title replaces the deterministic one only while the
// run still carries it: a title someone set meanwhile stays.
let current = match run_records::projection(&task.state, task.run_id).await {
Ok(Some(projection)) => projection.title().to_string(),
Ok(None) => return,
Err(err) => {
tracing::warn!(run_id = %task.run_id, error = %err, "Failed to open run store for title update");
tracing::warn!(run_id = %task.run_id, error = %err, "Failed to load the run for its title update");
return;
}
};
let expected_title = task.deterministic_title;
if let Err(err) = workflow_event::append_event_if(
&run_store,
&task.run_id,
&workflow_event::Event::RunTitleUpdated {
if current != task.deterministic_title {
return;
}
if let Err(err) = run_records::append(
&task.state,
task.run_id,
PlatformRecord::RunTitle(RunTitleRecord {
title: generated_title,
actor: Some(Principal::System {
system_kind: SystemActorKind::Engine,
}),
},
move |projection| projection.title().as_ref() == expected_title,
}),
)
.await
{
tracing::warn!(run_id = %task.run_id, error = %err, "Failed to append generated run title event");
tracing::warn!(run_id = %task.run_id, error = %err, "Failed to record the generated run title");
}
});
}
@ -1415,8 +1398,8 @@ async fn get_run_logs(
RequireRunScoped(id): RequireRunScoped,
State(state): State<Arc<AppState>>,
) -> Response {
if state.stores.runs.open_run_reader(&id).await.is_err() {
return ApiError::not_found("Run not found.").into_response();
if let Err(err) = state.load_run_projection(&id).await {
return err.into_response();
}
let path = Storage::new(state.server_storage_dir())

View file

@ -332,22 +332,12 @@ async fn get_system_repair_runs(
_auth: RequiredUser,
State(state): State<Arc<AppState>>,
) -> Response {
let issues = match state.stores.runs.list_unreadable_runs().await {
Ok(issues) => issues,
Err(err) => {
return ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string())
.into_response();
}
};
let total_count = to_i64(issues.len());
let runs = issues
.into_iter()
.map(|issue| SystemRepairRunIssue {
run_id: issue.run_id.to_string(),
created_at: issue.created_at,
error: issue.error,
})
.collect();
// A run's history is Petri's records and Fabro's platform records; a
// record the projector cannot read holds that run's view where it
// stands and is reported on the run, not repaired here.
let _ = state;
let runs: Vec<SystemRepairRunIssue> = Vec::new();
let total_count = to_i64(runs.len());
(
StatusCode::OK,

View file

@ -36,15 +36,16 @@ use fabro_interview::ControlInterviewer;
use fabro_petri::controls::RunControls;
use fabro_petri::engine::{self, Conclusion, Execution, RunRequest};
use fabro_petri::hooks::HooksSpec;
use fabro_petri::interview::{Approval, DatabaseQuestions, FabroInterviewer};
use fabro_petri::interview::{Approval, FabroInterviewer};
use fabro_petri::petri::StoreError;
use fabro_petri::platform_records::SqlitePlatformRecords;
use fabro_petri::recovery::{self, Recovery, RecoveryRequest};
use fabro_petri::runtime::{self, RuntimeSpec};
use fabro_petri::secrets::VaultSecrets;
use fabro_petri::{SqliteRunStore, admission};
use fabro_store::platform_records::{RunLifecycleKind, RunLifecycleRecord};
use fabro_types::settings::run::{ApprovalMode, RunMode};
use fabro_types::{PetriAdmission, RunId, RunRunnableSource, RunTarget, RunTiming, StageOutcome};
use fabro_types::{PetriAdmission, RunId, RunRunnableSource, RunTarget};
use fabro_util::error as error_util;
use fabro_workflow::Error as WorkflowError;
use fabro_workflow::run_status::{FailureReason, RunStatus, SuccessReason};
@ -53,7 +54,10 @@ use tokio::task;
use tokio_util::sync::CancellationToken;
use tracing::{error, info, warn};
use super::{AppState, RunAnswerTransport, RunExecutionMode, clear_live_run_state, workflow_event};
use super::{
AppState, RunAnswerTransport, RunExecutionMode, clear_live_run_state, run_records,
stream_follower,
};
use crate::petri_check;
use crate::petri_runs::PetriRuns;
use crate::run_compiler::{PreparedRun, RunCompilerError};
@ -189,28 +193,21 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
(run_dir, cancel, managed_run.execution_mode)
};
let run_store = match state.stores.runs.open_run(&run_id).await {
Ok(run_store) => run_store,
Err(err) => {
error!(run_id = %run_id, error = %err, "Failed to open run store");
stream_follower::follow_run(&state, run_id).await;
let run_state = match run_records::projection(&state, run_id).await {
Ok(Some(run_state)) => run_state,
Ok(None) => {
error!(run_id = %run_id, "Run not found at launch");
finish(
&state,
run_id,
RunStatus::Failed {
reason: FailureReason::WorkflowError,
},
Some(format!("Failed to open run store: {err}")),
Some("Run not found at launch".to_string()),
);
return;
}
};
tokio::spawn(super::forward_run_events_to_global(
Arc::clone(&state),
run_id,
run_store.subscribe(),
));
let run_state = match run_store.state().await {
Ok(run_state) => run_state,
Err(err) => {
error!(run_id = %run_id, error = %err, "Failed to load run state");
finish(
@ -242,7 +239,7 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
Ok(graphs) => Execution::Start(graphs),
Err(err) => {
let message = error_util::collect_chain(&err).join(": ");
fail_before_execution(&state, &run_store, run_id, &message).await;
fail_before_execution(&state, run_id, &message).await;
return;
}
}
@ -257,7 +254,6 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
let message = error_util::collect_chain(&err).join(": ");
fail_before_execution(
&state,
&run_store,
run_id,
&format!("the vault could not be read for the run: {message}"),
)
@ -266,35 +262,36 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
}
};
let started = Instant::now();
for event in [
workflow_event::Event::RunStarting,
workflow_event::Event::RunRunning,
for record in [
run_records::transition(RunLifecycleKind::Starting, RunStatus::Starting),
run_records::transition(RunLifecycleKind::Running, RunStatus::Running),
] {
if let Err(err) = workflow_event::append_event(&run_store, &run_id, &event).await {
error!(run_id = %run_id, error = %err, "Failed to persist run lifecycle event");
if let Err(err) = run_records::lifecycle(&state, run_id, record).await {
error!(run_id = %run_id, error = %err, "Failed to persist run lifecycle record");
finish(
&state,
run_id,
RunStatus::Failed {
reason: FailureReason::WorkflowError,
},
Some(format!("Failed to persist run lifecycle event: {err}")),
Some(format!("Failed to persist run lifecycle record: {err}")),
);
return;
}
}
// The answer endpoint reaches this interviewer directly, as it does
// for a legacy run in this process.
// The answer endpoint reaches this interviewer directly. The lifecycle
// records above already moved the live status to Running; a run that
// ended meanwhile (cancelled while starting) takes no transport.
let interviewer = Arc::new(ControlInterviewer::new());
{
let mut runs = state.runs.lock().expect("runs lock poisoned");
if let Some(managed_run) = runs.get_mut(&run_id) {
if managed_run.status == RunStatus::Starting {
managed_run.status = RunStatus::Running;
managed_run.answer_transport = Some(RunAnswerTransport::InProcess {
interviewer: Arc::clone(&interviewer),
});
}
if let Some(managed_run) = runs
.get_mut(&run_id)
.filter(|managed_run| !managed_run.status.is_terminal())
{
managed_run.answer_transport = Some(RunAnswerTransport::InProcess {
interviewer: Arc::clone(&interviewer),
});
}
}
let approval = if run_state.spec.settings.run.execution.approval == ApprovalMode::Auto {
@ -302,8 +299,7 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
} else {
Approval::Prompt
};
let questions = Arc::new(DatabaseQuestions::new(run_store.clone(), run_id));
let petri_interviewer = FabroInterviewer::new(interviewer, questions, approval);
let petri_interviewer = FabroInterviewer::new(interviewer, approval);
let observers = vec![petri_interviewer.observer()];
let (_, eligible) = state.resolve_llm_client_with_ready_ids().await;
let dry_run = run_state.spec.settings.run.execution.mode == RunMode::DryRun;
@ -333,11 +329,12 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
hooks: Some(hooks),
};
let result = Box::pin(engine::run(request)).await;
let timing = RunTiming {
wall_time_ms: u64::try_from(started.elapsed().as_millis()).unwrap_or(u64::MAX),
..RunTiming::default()
};
let (status, error, event) = match engine::conclusion(&result) {
info!(
run_id = %run_id,
elapsed_ms = u64::try_from(started.elapsed().as_millis()).unwrap_or(u64::MAX),
"Petri run ended"
);
let (status, error, record) = match engine::conclusion(&result) {
Conclusion::Succeeded => {
info!(run_id = %run_id, "Petri run completed");
(
@ -345,24 +342,15 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
reason: SuccessReason::Completed,
},
None,
workflow_event::Event::WorkflowRunCompleted {
timing,
artifact_count: 0,
status: StageOutcome::Succeeded.to_string(),
reason: SuccessReason::Completed,
final_git_commit_sha: None,
final_patch: None,
diff_summary: None,
usage: None,
},
run_records::succeeded(SuccessReason::Completed),
)
}
Conclusion::Failed { reason, message } => {
info!(run_id = %run_id, error = %message, "Petri run did not succeed");
failed(reason, message, timing)
failed(reason, message)
}
};
if let Err(err) = workflow_event::append_event(&run_store, &run_id, &event).await {
if let Err(err) = run_records::lifecycle(&state, run_id, record).await {
error!(run_id = %run_id, error = %err, "Failed to persist run outcome");
}
// The view trails the terminal record; the aggregate reads the settled
@ -397,7 +385,6 @@ pub(crate) async fn execute(state: Arc<AppState>, run_id: RunId) {
pub(crate) async fn reconcile_on_startup(
state: &Arc<AppState>,
run_id: RunId,
run_store: &fabro_store::RunDatabase,
run_state: &fabro_store::RunProjection,
) -> anyhow::Result<()> {
let key = PetriRuns::key(&run_id);
@ -442,9 +429,8 @@ pub(crate) async fn reconcile_on_startup(
error = %reason,
"Petri run left in flight by the previous server cannot resume; reporting it failed"
);
let (_, _, event) =
failed(FailureReason::WorkflowError, reason, RunTiming::default());
workflow_event::append_event(run_store, &run_id, &event).await?;
let (_, _, record) = failed(FailureReason::WorkflowError, reason);
run_records::lifecycle(state, run_id, record).await?;
return Ok(());
}
}
@ -457,17 +443,12 @@ pub(crate) async fn reconcile_on_startup(
mode = super::worker_mode_arg(mode),
"Petri run left in flight by the previous server; relaunching its worker"
);
for event in [
workflow_event::Event::RunStartRequested {
resume: true,
actor: None,
},
workflow_event::Event::RunRunnable {
source: RunRunnableSource::StartRequested,
actor: None,
},
] {
workflow_event::append_event(run_store, &run_id, &event).await?;
let mut start_requested = RunLifecycleRecord::new(RunLifecycleKind::StartRequested);
start_requested.source = Some("resume".to_string());
let mut runnable = run_records::transition(RunLifecycleKind::Runnable, RunStatus::Runnable);
runnable.source = Some(<&'static str>::from(RunRunnableSource::StartRequested).to_string());
for record in [start_requested, runnable] {
run_records::lifecycle(state, run_id, record).await?;
}
let mut runs = state.runs.lock().expect("runs lock poisoned");
runs.insert(
@ -483,39 +464,27 @@ pub(crate) async fn reconcile_on_startup(
Ok(())
}
/// The failed status, its message, and the `run.failed` event for it.
/// The failed status, its message, and the `failed` lifecycle record for it.
fn failed(
reason: FailureReason,
message: String,
timing: RunTiming,
) -> (RunStatus, Option<String>, workflow_event::Event) {
let error = match reason {
FailureReason::Cancelled => WorkflowError::Cancelled,
_ => WorkflowError::engine(message.clone()),
) -> (RunStatus, Option<String>, RunLifecycleRecord) {
let detail = match reason {
FailureReason::Cancelled => WorkflowError::Cancelled.to_string(),
_ => message.clone(),
};
(
RunStatus::Failed { reason },
Some(message),
workflow_event::Event::workflow_run_failed_from_error(
&error, timing, reason, None, None, None, None,
),
run_records::failed(reason, detail),
)
}
/// Record a failure that happened before Petri ran, then finish the run.
async fn fail_before_execution(
state: &Arc<AppState>,
run_store: &fabro_store::RunDatabase,
run_id: RunId,
message: &str,
) {
async fn fail_before_execution(state: &Arc<AppState>, run_id: RunId, message: &str) {
error!(run_id = %run_id, error = message, "Petri run cannot start");
let (status, error, event) = failed(
FailureReason::WorkflowError,
message.to_string(),
RunTiming::default(),
);
if let Err(err) = workflow_event::append_event(run_store, &run_id, &event).await {
let (status, error, record) = failed(FailureReason::WorkflowError, message.to_string());
if let Err(err) = run_records::lifecycle(state, run_id, record).await {
error!(run_id = %run_id, error = %err, "Failed to persist run failure status");
}
finish(state, run_id, status, error);

View file

@ -9,6 +9,9 @@ use std::collections::{BTreeSet, HashMap, HashSet};
use std::sync::Arc;
use std::time::Duration;
use fabro_store::platform_records::{
PlatformRecord, PullRequestCreatedRecord, PullRequestFailedRecord,
};
use fabro_types::{PullRequestCreation, PullRequestCreationId, RunId};
use tokio::task::{self, JoinHandle, JoinSet};
use tokio::time;
@ -17,7 +20,7 @@ use tracing::{Instrument as _, info_span, warn};
use super::handler::pull_requests::{
RunPrInputs, load_server_github_credentials, server_github_context,
};
use super::{AppState, pull_request, workflow_event};
use super::{AppState, pull_request, run_records};
const PULL_REQUEST_CREATION_TIMEOUT: Duration = Duration::from_mins(10);
const PULL_REQUEST_CREATION_SCAN_INTERVAL: Duration = Duration::from_secs(30);
@ -87,38 +90,28 @@ impl AppState {
.expect("pull request creation queue lock poisoned")
.pop()
}
#[cfg(test)]
pub(super) fn pull_request_creation_queue_len(&self) -> usize {
self.pull_request_creation_queue
.lock()
.expect("pull request creation queue lock poisoned")
.len()
}
#[cfg(test)]
pub(super) fn drain_pull_request_creation_queue(&self) -> Vec<RunId> {
let mut queue = self
.pull_request_creation_queue
.lock()
.expect("pull request creation queue lock poisoned");
std::iter::from_fn(|| queue.pop()).collect()
}
}
async fn append_pull_request_creation_failure(
run_store: &fabro_store::RunDatabase,
state: &AppState,
run_id: &RunId,
creation_id: PullRequestCreationId,
error: String,
) -> anyhow::Result<()> {
let event = workflow_event::Event::PullRequestFailed {
creation_id: Some(creation_id),
error,
};
workflow_event::append_event_if(run_store, run_id, &event, |projection| {
is_pending_creation(projection, creation_id)
})
let still_pending = run_records::projection(state, *run_id)
.await?
.is_some_and(|projection| is_pending_creation(&projection, creation_id));
if !still_pending {
return Ok(());
}
run_records::append(
state,
*run_id,
PlatformRecord::PullRequestFailed(PullRequestFailedRecord {
creation_id: Some(creation_id),
error,
}),
)
.await?;
Ok(())
}
@ -138,8 +131,7 @@ pub(in crate::server) async fn process_pull_request_creation(
run_id: RunId,
) -> anyhow::Result<()> {
let _create_guard = state.pull_request_create_locks.lock(run_id).await;
let run_store = state.stores.runs.open_run(&run_id).await?;
let Some(run_state) = state.stores.runs.load_run_projection(&run_id).await? else {
let Some(run_state) = run_records::projection(&state, run_id).await? else {
return Ok(());
};
let Some(creation) = run_state
@ -151,10 +143,10 @@ pub(in crate::server) async fn process_pull_request_creation(
return Ok(());
};
match attempt_pull_request_creation(&state, &run_store, &run_id, &run_state, &creation).await? {
match attempt_pull_request_creation(&state, &run_id, &run_state, &creation).await? {
Ok(()) => Ok(()),
Err(error) => {
append_pull_request_creation_failure(&run_store, &run_id, creation.id, error).await
append_pull_request_creation_failure(&state, &run_id, creation.id, error).await
}
}
}
@ -165,7 +157,6 @@ pub(in crate::server) async fn process_pull_request_creation(
/// infrastructure failure — nothing was recorded, so the supervisor may retry.
async fn attempt_pull_request_creation(
state: &AppState,
run_store: &fabro_store::RunDatabase,
run_id: &RunId,
run_state: &fabro_store::RunProjection,
creation: &PullRequestCreation,
@ -183,7 +174,6 @@ async fn attempt_pull_request_creation(
Err(err) => return Ok(Err(err.detail().to_string())),
};
let catalog = state.catalog();
let run_store_handle = run_store.clone().into();
let request = pull_request::OpenPullRequestRequest {
github,
origin_url: &inputs.normalized_origin,
@ -195,7 +185,6 @@ async fn attempt_pull_request_creation(
model: &creation.model,
draft: true,
auto_merge: None,
run_store: &run_store_handle,
llm_source: Arc::clone(&state.llm_source),
catalog,
conclusion: Some(inputs.conclusion),
@ -217,18 +206,28 @@ async fn attempt_pull_request_creation(
}
};
let event = workflow_event::Event::pull_request_created(
&created_pull_request.link,
&created_pull_request.base_branch,
&created_pull_request.head_branch,
inputs.final_git_sha,
&created_pull_request.title,
true,
);
workflow_event::append_event_if(run_store, run_id, &event, |projection| {
projection.pull_request.is_none() && is_pending_creation(projection, creation.id)
})
.await?;
let still_pending = run_records::projection(state, *run_id)
.await?
.is_some_and(|projection| {
projection.pull_request.is_none() && is_pending_creation(&projection, creation.id)
});
if still_pending {
let link = &created_pull_request.link;
run_records::append(
state,
*run_id,
PlatformRecord::PullRequestCreated(PullRequestCreatedRecord {
number: link.number,
owner: link.owner.clone(),
repo: link.repo.clone(),
html_url: link.html_url(),
head_sha: Some(inputs.final_git_sha.to_string()),
draft: true,
operation: None,
}),
)
.await?;
}
Ok(Ok(()))
}
@ -257,7 +256,7 @@ pub(super) async fn recover_pending_pull_request_creations(
if !can_dispatch(&run_id, active, failures) {
continue;
}
let projection = match state.stores.runs.load_run_projection(&run_id).await {
let projection = match run_records::projection(state, run_id).await {
Ok(Some(projection)) => projection,
Ok(None) => continue,
Err(error) => {

View file

@ -342,10 +342,6 @@ fn select_storage_disk<'a>(
.max_by_key(|disk| disk.mount_point.components().count())
}
pub(crate) fn available_space_for_path(storage_path: &Path) -> Option<u64> {
select_storage_disk(storage_path, &refreshed_disk_candidates()).map(|disk| disk.available_bytes)
}
fn percent(used: u64, total: u64) -> Option<f64> {
if total == 0 {
return None;

View file

@ -0,0 +1,96 @@
//! Fabro's own facts about a run, written and read by the server.
//!
//! A run's lifecycle before, beside and after the engine (its queue, an
//! approval, a control request, the terminal status Fabro reports), its
//! title and parent, its pull request and its notices are platform records
//! (`fabro_store::platform_records`), appended here and folded into the
//! run's projection by the Petri projector. Every append wakes the
//! projector; a caller that reads the run back right after waits for that
//! pass, so what it reads holds what it wrote.
use std::sync::Arc;
use anyhow::Context as _;
use axum::http::StatusCode;
use fabro_store::RunProjection;
use fabro_store::platform_records::{
PlatformRecord, RunLifecycleKind, RunLifecycleRecord, StoredPlatformRecord,
};
use fabro_types::{FailureReason, RunId, RunStatus, SuccessReason};
use super::AppState;
use crate::error::ApiError;
/// Append one record for the run, wake its projector and wait for the
/// pass that folds it: what the caller reads next holds the record.
pub(crate) async fn append(
state: &AppState,
run_id: RunId,
record: PlatformRecord,
) -> anyhow::Result<StoredPlatformRecord> {
let summaries = &state.stores.run_summaries;
let stored = summaries
.platform_records()
.append(&run_id, &record, None)
.await
.with_context(|| format!("appending a {} record for run {run_id}", record.kind()))?;
summaries.notify_platform_record(run_id);
state.petri_projector.settle(run_id).await;
Ok(stored)
}
/// Append one lifecycle transition for the run.
pub(crate) async fn lifecycle(
state: &AppState,
run_id: RunId,
record: RunLifecycleRecord,
) -> anyhow::Result<StoredPlatformRecord> {
append(state, run_id, PlatformRecord::RunLifecycle(record)).await
}
/// A transition that leads to `status`.
#[must_use]
pub(crate) fn transition(kind: RunLifecycleKind, status: RunStatus) -> RunLifecycleRecord {
RunLifecycleRecord::new(kind).with_status(status)
}
/// The run failed for `reason`, with `message` as the failure's detail.
#[must_use]
pub(crate) fn failed(reason: FailureReason, message: impl Into<String>) -> RunLifecycleRecord {
let mut record = transition(RunLifecycleKind::Failed, RunStatus::Failed { reason });
record.reason = Some(message.into());
record
}
/// The run succeeded for `reason`.
#[must_use]
pub(crate) fn succeeded(reason: SuccessReason) -> RunLifecycleRecord {
transition(RunLifecycleKind::Succeeded, RunStatus::Succeeded { reason })
}
/// The run's projection once every committed record is folded: the read
/// that follows a write.
pub(crate) async fn projection(
state: &AppState,
run_id: RunId,
) -> anyhow::Result<Option<Arc<RunProjection>>> {
state.petri_projector.settle(run_id).await;
state
.stores
.run_summaries
.load_petri_projection(&run_id)
.await
.with_context(|| format!("loading the projection of run {run_id}"))
}
/// [`projection`], as an API handler needs it: a missing run is the
/// canonical 404, a store failure a 500.
pub(crate) async fn require_projection(
state: &AppState,
run_id: RunId,
) -> Result<Arc<RunProjection>, ApiError> {
projection(state, run_id)
.await
.map_err(|err| ApiError::new(StatusCode::INTERNAL_SERVER_ERROR, err.to_string()))?
.ok_or_else(|| ApiError::not_found("Run not found."))
}

View file

@ -0,0 +1,164 @@
//! The server's own reader of every run's stream.
//!
//! The Petri projector commits a run's stream (Petri's events and Fabro's
//! platform records, one `stream_seq` each) and signals after each pass.
//! This follower reads what each pass committed and hands it to the
//! server's in-process consumers: the in-memory run map the scheduler and
//! the control handlers read, and the global broadcast the `/attach`
//! stream, the Slack service and any other subscriber take their items
//! from.
//!
//! A run is followed from the moment the server launches it
//! ([`StreamFollower::follow`]): the cursor starts at the stream's head
//! then, so nothing the run recorded before this process took charge of it
//! is replayed into the live state. A run the follower first sees by its
//! signal alone is followed from its head at that moment.
use std::collections::HashMap;
use std::sync::Arc;
use fabro_store::platform_records::PlatformRecord;
use fabro_types::{RunId, RunStatus, RunStreamItem};
use tokio::sync::Mutex;
use tokio::sync::broadcast::error::RecvError;
use tokio::task::JoinHandle;
use tracing::warn;
use super::{AppState, apply_lifecycle_to_managed_run, reconcile_live_interview_state};
/// How many stream items one read takes.
const BATCH_LIMIT: usize = 256;
/// The follower's cursors: the last `stream_seq` seen per run.
#[derive(Default)]
pub(crate) struct StreamFollower {
cursors: Mutex<HashMap<RunId, u64>>,
}
/// Follow `run_id` from the stream's current head.
pub(crate) async fn follow_run(state: &AppState, run_id: RunId) {
let head = match state.petri_projector.stream_head(run_id).await {
Ok(head) => head.unwrap_or(0),
Err(err) => {
warn!(run_id = %run_id, error = %err, "the run's stream head could not be read; following from its start");
0
}
};
state
.stream_follower
.cursors
.lock()
.await
.entry(run_id)
.or_insert(head);
}
/// Start the follower over the projector's signals.
pub(crate) fn spawn_stream_follower(state: Arc<AppState>) -> JoinHandle<()> {
let mut signals = state.petri_projector.subscribe();
tokio::spawn(async move {
loop {
let run_id = match signals.recv().await {
Ok(run_id) => run_id,
Err(RecvError::Lagged(_)) => continue,
Err(RecvError::Closed) => break,
};
if state.is_shutting_down() {
break;
}
catch_up(&state, run_id).await;
}
})
}
/// Read what the run's stream holds past the follower's cursor and hand it
/// to the live state and the global broadcast.
async fn catch_up(state: &Arc<AppState>, run_id: RunId) {
let mut cursor = {
let mut cursors = state.stream_follower.cursors.lock().await;
if let Some(cursor) = cursors.get(&run_id) {
*cursor
} else {
let head = match state.petri_projector.stream_head(run_id).await {
Ok(head) => head.unwrap_or(0),
Err(_) => return,
};
cursors.insert(run_id, head);
head
}
};
let mut saw_items = false;
loop {
let items = match state
.petri_projector
.stream_after(run_id, cursor, BATCH_LIMIT)
.await
{
Ok(items) => items,
Err(err) => {
warn!(run_id = %run_id, error = %err, "the run's stream could not be read");
return;
}
};
let drained = items.len() < BATCH_LIMIT;
for item in items {
cursor = item.stream_seq;
saw_items = true;
fold_into_live_state(state, run_id, &item);
let _ = state.global_event_tx.send(item);
}
if drained {
break;
}
}
state
.stream_follower
.cursors
.lock()
.await
.insert(run_id, cursor);
if saw_items {
sync_live_status_from_projection(state, run_id).await;
}
}
/// One item into the in-memory run: its lifecycle records fold into the
/// live status; a closed question releases its answer claim.
fn fold_into_live_state(state: &AppState, run_id: RunId, item: &RunStreamItem) {
if let Some(PlatformRecord::RunLifecycle(record)) = super::platform_record_of(item) {
apply_lifecycle_to_managed_run(state, run_id, &record);
}
let mut runs = state.runs.lock().expect("runs lock poisoned");
if let Some(managed_run) = runs.get_mut(&run_id) {
reconcile_live_interview_state(managed_run, item);
}
}
/// The blocked and paused substates of a live run come from Petri's own
/// records (a pending question, a held admission), which the projection
/// folds; the live status follows the projection there, and nowhere else.
async fn sync_live_status_from_projection(state: &AppState, run_id: RunId) {
let Ok(Some(projection)) = state
.stores
.run_summaries
.load_petri_projection(&run_id)
.await
else {
return;
};
let mut runs = state.runs.lock().expect("runs lock poisoned");
let Some(managed_run) = runs.get_mut(&run_id) else {
return;
};
let live = matches!(
managed_run.status,
RunStatus::Running | RunStatus::Blocked { .. } | RunStatus::Paused { .. }
);
let projected_live = matches!(
projection.status,
RunStatus::Running | RunStatus::Blocked { .. } | RunStatus::Paused { .. }
);
if live && projected_live && managed_run.status != projection.status {
managed_run.status = projection.status;
}
}

File diff suppressed because it is too large Load diff

View file

@ -1,187 +0,0 @@
use std::sync::Arc;
use std::time::Duration;
use axum::body::{Body, to_bytes};
use axum::http::{Request, StatusCode};
use chrono::{SecondsFormat, Utc};
use fabro_static::EnvVars;
use object_store::ObjectStore;
use object_store::memory::InMemory;
use tokio::sync::Barrier;
use tower::ServiceExt;
use crate::helpers::{MINIMAL_DOT, api, minimal_intent_json, response_json, test_settings};
fn app_with_store(
object_store: Arc<dyn ObjectStore>,
blobs: Arc<fabro_store::BlobStore>,
run_summaries: Arc<fabro_store::RunSummaryStore>,
) -> axum::Router {
let settings = test_settings();
let store = Arc::new(fabro_store::test_support::test_database_with_stores(
Arc::clone(&object_store),
"event-race",
Duration::from_millis(1),
None,
blobs,
run_summaries,
));
let artifact_store = fabro_store::ArtifactStore::new(object_store, "artifacts");
let state = fabro_server::test_support::TestAppStateBuilder::new()
.runtime_settings(settings.server_settings, settings.manifest_run_defaults)
.env_lookup(|_| None)
.vault_entries([(EnvVars::OPENAI_API_KEY, "test-openai-api-key")])
.store_bundle(store, artifact_store)
.build();
fabro_server::test_support::build_test_router(state)
}
async fn create_run(app: &axum::Router) -> (String, tempfile::TempDir) {
let workspace = tempfile::tempdir().expect("run target workspace should be created");
let request = Request::builder()
.method("POST")
.uri(api("/runs"))
.header("content-type", "application/json")
.body(Body::from(
serde_json::to_string(&minimal_intent_json(app, MINIMAL_DOT, workspace.path()).await)
.expect("intent should serialize"),
))
.expect("create-run request should build");
let body = response_json(
app.clone().oneshot(request).await.unwrap(),
StatusCode::CREATED,
"POST /api/v1/runs",
)
.await;
let run_id = body["id"]
.as_str()
.expect("create-run response should include an id")
.to_string();
(run_id, workspace)
}
fn append_stage_started_request(run_id: &str, index: usize) -> Request<Body> {
Request::builder()
.method("POST")
.uri(api(&format!("/runs/{run_id}/events")))
.header("content-type", "application/json")
.body(Body::from(
serde_json::to_string(&serde_json::json!({
"id": ulid::Ulid::new().to_string(),
"ts": Utc::now().to_rfc3339_opts(SecondsFormat::Millis, true),
"run_id": run_id,
"event": "stage.started",
"node_id": format!("race-{index}"),
"node_label": format!("Race {index}"),
"stage_id": format!("race-{index}@1"),
"actor": {
"kind": "worker",
"run_id": run_id,
},
"properties": {
"index": index,
"handler_type": "noop",
"attempt": 1,
"max_attempts": 1,
},
}))
.expect("event should serialize"),
))
.expect("append-event request should build")
}
async fn append_status_and_body(
app: axum::Router,
run_id: String,
index: usize,
barrier: Arc<Barrier>,
) -> (usize, StatusCode, String) {
barrier.wait().await;
let response = app
.oneshot(append_stage_started_request(&run_id, index))
.await
.expect("append-event response should execute");
let status = response.status();
let bytes = to_bytes(response.into_body(), usize::MAX)
.await
.expect("append-event body should buffer");
(index, status, String::from_utf8_lossy(&bytes).into_owned())
}
#[tokio::test(flavor = "multi_thread", worker_threads = 8)]
async fn concurrent_event_appends_after_restart_keep_projection_cache_contiguous() {
let object_store: Arc<dyn ObjectStore> = Arc::new(InMemory::new());
let blobs = fabro_store::test_support::test_blob_store();
let run_summaries = fabro_store::test_support::test_run_summary_store();
let first_app = app_with_store(
Arc::clone(&object_store),
Arc::clone(&blobs),
Arc::clone(&run_summaries),
);
let (run_id, _run_id_workspace) = create_run(&first_app).await;
tokio::time::sleep(Duration::from_millis(25)).await;
// Simulate a server restart: a fresh AppState opens the existing run with
// an empty active-run cache, so concurrent appends all race through the
// public event endpoint instead of sharing an already-open RunDatabase.
let restarted_app = app_with_store(object_store, blobs, run_summaries);
let appends = 64;
let barrier = Arc::new(Barrier::new(appends));
let mut tasks = Vec::with_capacity(appends);
for index in 0..appends {
tasks.push(tokio::spawn(append_status_and_body(
restarted_app.clone(),
run_id.clone(),
index,
Arc::clone(&barrier),
)));
}
let mut results = Vec::with_capacity(appends);
for task in tasks {
results.push(task.await.expect("append task should not panic"));
}
let failures = results
.iter()
.filter(|(_, status, _)| *status != StatusCode::OK)
.collect::<Vec<_>>();
assert!(
failures.is_empty(),
"all concurrent event appends should succeed, got failures: {failures:#?}"
);
let request = Request::builder()
.method("GET")
.uri(api(&format!("/runs/{run_id}/events")))
.body(Body::empty())
.expect("list-events request should build");
let body = response_json(
restarted_app.clone().oneshot(request).await.unwrap(),
StatusCode::OK,
format!("GET /api/v1/runs/{run_id}/events"),
)
.await;
let events = body["data"]
.as_array()
.expect("events response should include data");
let seqs = events
.iter()
.map(|event| {
event["seq"]
.as_u64()
.expect("event should include numeric seq")
})
.collect::<Vec<_>>();
assert_eq!(
events.len(),
appends + 2,
"every append should be durable; observed seqs: {seqs:?}"
);
assert_eq!(
seqs,
(1..=u64::try_from(appends + 2).unwrap()).collect::<Vec<_>>(),
"event seqs should be contiguous"
);
}

View file

@ -565,8 +565,13 @@ async fn invalid_mcp_server_id_is_bad_request() {
.await;
}
/// A `run.agent.mcps.<name>` entry that names a server catalog entry by
/// `id`: Petri's Fabro frontend reads `workflow.toml` itself and has no
/// server catalog to resolve the reference against, so the check refuses
/// it (`unsupported.workflow_toml.run.agent.mcps.reference`) until the
/// frontend takes the catalog. The run create path shares the gap.
#[tokio::test]
async fn created_mcp_server_can_be_referenced_by_manifest_validation() {
async fn manifest_validation_reports_a_catalog_mcp_reference_as_unsupported() {
let (app, _temp_dir, _mcp_dir) = mcp_server_app();
create_mcp_server(&app, "sentry", "Sentry").await;
@ -587,7 +592,18 @@ id = "sentry"
.expect("manifest validation should respond");
let body = response_json(response, StatusCode::OK, "POST /api/v1/validate").await;
assert_eq!(body["ok"], true);
assert_eq!(body["ok"], false, "{body}");
let rules: Vec<&str> = body["workflow"]["diagnostics"]
.as_array()
.expect("diagnostics")
.iter()
.filter_map(|diagnostic| diagnostic["rule"].as_str())
.collect();
assert_eq!(
rules,
vec!["unsupported.workflow_toml.run.agent.mcps.reference"],
"{body}"
);
}
#[tokio::test]

View file

@ -4,7 +4,6 @@ mod cli_auth_token;
mod compression;
mod docs;
mod environments;
mod events;
mod install;
mod install_openai_compatible;
mod mcp_servers;

View file

@ -358,20 +358,6 @@ async fn a_reconnecting_client_receives_every_stream_item_once_in_order() {
.count();
assert_eq!(finished, 1, "the stream ends with the run's finish");
// The legacy cursors are refused for a Petri run; the stream cursor is
// refused for nothing else.
let req = Request::builder()
.method("GET")
.uri(api(&format!("/runs/{run_id}/events?since_seq=1")))
.body(Body::empty())
.expect("events request should build");
let response = app
.clone()
.oneshot(req)
.await
.expect("events request routes");
assert_eq!(response.status(), StatusCode::BAD_REQUEST);
capture_fixture(&app, &run_id, "parallel", &projection).await;
}

View file

@ -98,12 +98,10 @@ use fabro_interview::{
Answer as LegacyAnswer, AnswerSubmission, AnswerValue, AutoApproveInterviewer,
ControlInterviewer, Interviewer as LegacyInterviewer, Question as LegacyQuestion,
};
use fabro_store::RunDatabase;
use fabro_types::{
InterviewOption, Principal, QuestionType, ReviewTarget, ReviewTargetKind, RunId, StageId,
InterviewOption, Principal, QuestionType, ReviewTarget, ReviewTargetKind, StageId,
SystemActorKind,
};
use fabro_workflow::event::{self as workflow_event, Event, RunEventSink};
use petri_execution::events::{Parsed, Projection, ViewEvent};
use petri_execution::{
CoordinatorRecord, ExecutionId, ExecutionObserver, InterviewError, InterviewReply,
@ -199,63 +197,6 @@ pub enum QuestionNotice {
}
impl QuestionNotice {
/// The run event the legacy `human` stage emits for the same fact.
#[must_use]
pub fn into_event(self) -> Event {
match self {
Self::Asked(asked) => Event::InterviewStarted {
question_id: asked.question_id,
question: asked.text,
stage: asked.stage,
question_type: asked.question_type.to_string(),
options: asked.options,
allow_freeform: asked.allow_freeform,
timeout_seconds: asked.timeout_seconds,
context_display: None,
review_target: asked.review_target,
},
Self::Answered {
question_id,
text,
answer,
actor,
duration_ms,
} => Event::InterviewCompleted {
actor: Some(actor),
question_id,
question: text,
answer,
duration_ms,
},
Self::Expired {
question_id,
text,
stage,
duration_ms,
} => Event::InterviewTimeout {
actor: None,
question_id,
question: text,
stage,
duration_ms,
},
Self::Interrupted {
question_id,
text,
stage,
reason,
duration_ms,
} => Event::InterviewInterrupted {
actor: None,
question_id,
question: text,
stage,
reason,
duration_ms,
},
}
}
fn question_id(&self) -> &str {
match self {
Self::Asked(asked) => &asked.question_id,
@ -266,52 +207,21 @@ impl QuestionNotice {
}
}
/// Where the adapter posts what happens to a question: the run's event
/// stream, whichever way the process reaches it.
/// Where the adapter reports what happens to a question, when something
/// observes it: a test's board. A run's own record of a question is
/// Petri's, and who answered it is the server's platform record.
#[async_trait::async_trait]
pub trait QuestionSink: Send + Sync {
async fn post(&self, notice: QuestionNotice) -> anyhow::Result<()>;
}
/// The worker's sink: the run event sink its lifecycle events go through.
pub struct EventSinkQuestions {
sink: RunEventSink,
run_id: RunId,
}
impl EventSinkQuestions {
#[must_use]
pub fn new(sink: RunEventSink, run_id: RunId) -> Self {
Self { sink, run_id }
}
}
/// A sink that drops every notice: the adapter with nothing observing it.
struct Unobserved;
#[async_trait::async_trait]
impl QuestionSink for EventSinkQuestions {
async fn post(&self, notice: QuestionNotice) -> anyhow::Result<()> {
workflow_event::append_event_to_sink(&self.sink, &self.run_id, &notice.into_event())
.await
.map_err(anyhow::Error::new)
}
}
/// The server's sink for a run in its own process: the run's database.
pub struct DatabaseQuestions {
store: RunDatabase,
run_id: RunId,
}
impl DatabaseQuestions {
#[must_use]
pub fn new(store: RunDatabase, run_id: RunId) -> Self {
Self { store, run_id }
}
}
#[async_trait::async_trait]
impl QuestionSink for DatabaseQuestions {
async fn post(&self, notice: QuestionNotice) -> anyhow::Result<()> {
workflow_event::append_event(&self.store, &self.run_id, &notice.into_event()).await
impl QuestionSink for Unobserved {
async fn post(&self, _notice: QuestionNotice) -> anyhow::Result<()> {
Ok(())
}
}
@ -415,22 +325,24 @@ pub struct FabroInterviewer {
}
impl FabroInterviewer {
/// Over the control interviewer the run's answers are delivered to,
/// and the sink its questions are posted through.
/// Over the control interviewer the run's answers are delivered to.
#[must_use]
pub fn new(
answers: Arc<ControlInterviewer>,
sink: Arc<dyn QuestionSink>,
approval: Approval,
) -> Self {
pub fn new(answers: Arc<ControlInterviewer>, approval: Approval) -> Self {
Self {
answers,
sink,
sink: Arc::new(Unobserved),
approval,
observed: Arc::new(Observed::default()),
}
}
/// Report what happens to each question to `sink` as well.
#[must_use]
pub fn with_sink(mut self, sink: Arc<dyn QuestionSink>) -> Self {
self.sink = sink;
self
}
/// The observer that labels each firing as the projection does and
/// sees a gate report a question's expiry. A run registers it ahead of
/// the interview dispatcher, so a question's stage is known when it is

View file

@ -36,11 +36,11 @@ use fabro_types::{
BlockedReason, CheckpointRecord as ViewCheckpoint, CodingAgentEvent, CodingEvent, Conclusion,
FailureCategory, FailureDetail, FailureReason, InterviewOption, InterviewQuestionRecord,
ModelRef, ModelUsage, ParallelBranchId, ParallelBranchResult, PendingInterviewRecord,
PullRequestLink, RunApproval, RunApprovalState, RunControlAction, RunDiff, RunFailure, RunId,
RunProjection, RunSandbox, RunSandboxPlan, RunStatus, RunTiming, SandboxProviderKind,
StageCompletion, StageHandler, StageId, StageInferenceProjection, StageModelUsage,
StageOutcome, StageProjection, StageState, StageTiming, StartRecord, SuccessReason,
first_event_seq, timing, usage_rollup,
PullRequestCreation, PullRequestCreationStatus, PullRequestLink, RunApproval, RunApprovalState,
RunControlAction, RunDiff, RunFailure, RunId, RunProjection, RunSandbox, RunSandboxPlan,
RunStatus, RunTiming, SandboxProviderKind, StageCompletion, StageHandler, StageId,
StageInferenceProjection, StageModelUsage, StageOutcome, StageProjection, StageState,
StageTiming, StartRecord, SuccessReason, first_event_seq, timing, usage_rollup,
};
use lithos_llm::catalog::{ModelId, ProviderId};
use lithos_llm::types::Usage;
@ -258,12 +258,62 @@ impl RunView {
},
});
}
PlatformRecord::PullRequestRequested(record) => {
projection.pull_request_creation = Some(PullRequestCreation {
id: record.creation_id,
status: PullRequestCreationStatus::Pending,
model: record.model.clone(),
force: record.force,
requested_at: at,
updated_at: at,
pull_request: None,
error: None,
});
}
PlatformRecord::PullRequestCreated(record) => {
projection.pull_request = Some(PullRequestLink {
let link = PullRequestLink {
owner: record.owner.clone(),
repo: record.repo.clone(),
number: record.number,
});
};
projection.pull_request = Some(link.clone());
if let Some(creation) = projection
.pull_request_creation
.as_mut()
.filter(|creation| creation.is_pending())
{
creation.succeed(link, at);
}
}
PlatformRecord::PullRequestFailed(record) => {
if let Some(creation) =
projection
.pull_request_creation
.as_mut()
.filter(|creation| {
creation.is_pending()
&& record
.creation_id
.is_none_or(|creation_id| creation_id == creation.id)
})
{
creation.fail(record.error.clone(), at);
}
}
PlatformRecord::PullRequestLinked(record) => {
let link = record.link();
projection.pull_request = Some(link.clone());
if let Some(creation) = projection
.pull_request_creation
.as_mut()
.filter(|creation| creation.is_pending())
{
creation.succeed(link, at);
}
}
PlatformRecord::PullRequestUnlinked(_) => {
projection.pull_request = None;
projection.pull_request_creation = None;
}
}
}
@ -1101,8 +1151,24 @@ fn fold_lifecycle(projection: &mut RunProjection, record: &RunLifecycleRecord, a
at,
);
}
Kind::Runnable
| Kind::Starting
Kind::Runnable => {
// A run left in flight by a restart goes back to the queue: the
// resume's `runnable` steps back from wherever the run stood.
let in_flight = matches!(
projection.status,
RunStatus::Starting
| RunStatus::Running
| RunStatus::Blocked { .. }
| RunStatus::Paused { .. }
);
if in_flight && record.status == Some(RunStatus::Runnable) {
projection.status = RunStatus::Runnable;
projection.status_updated_at = at;
} else if let Some(status) = record.status {
apply_status(projection, status, at);
}
}
Kind::Starting
| Kind::Running
| Kind::Blocked
| Kind::Unblocked

View file

@ -156,7 +156,7 @@ impl Gate {
}
fn interviewer(&self, approval: Approval) -> FabroInterviewer {
FabroInterviewer::new(Arc::clone(&self.control), self.board.clone(), approval)
FabroInterviewer::new(Arc::clone(&self.control), approval).with_sink(self.board.clone())
}
fn marker(&self, name: &str) -> bool {

View file

@ -960,11 +960,7 @@ impl GateRun {
Launch::default(),
&runtime,
);
let interviewer = FabroInterviewer::new(
Arc::new(ControlInterviewer::new()),
Arc::new(support::Silent),
approval,
);
let interviewer = FabroInterviewer::new(Arc::new(ControlInterviewer::new()), approval);
let store = self
.projector
.observe_store(Arc::new(SqliteRunStore::new(self.scenario.pool.clone())));

View file

@ -129,9 +129,9 @@ pub(crate) fn run_request(
pub(crate) fn no_questions(sink: Arc<dyn QuestionSink>) -> FabroInterviewer {
FabroInterviewer::new(
Arc::new(fabro_interview::ControlInterviewer::new()),
sink,
Approval::Prompt,
)
.with_sink(sink)
}
/// A sink that drops every notice.

View file

@ -40,6 +40,7 @@ impl SlateKey {
&self.0
}
#[cfg(test)]
pub(crate) fn segments(raw: &str) -> impl Iterator<Item = &str> {
raw.split(Self::SEP)
}
@ -53,38 +54,6 @@ impl AsRef<[u8]> for SlateKey {
// --- Construction ---
pub(crate) fn run_events_prefix(run_id: &RunId) -> SlateKey {
SlateKey::new("runs")
.with(run_id)
.with("events")
.into_prefix()
}
/// Prefix of the retired `runs/_index/by-start/<run_id>` catalog markers that
/// the legacy layout kept beside each run's events.
pub(crate) fn run_catalog_prefix() -> SlateKey {
run_catalog_root().into_prefix()
}
#[cfg(test)]
pub(crate) fn run_catalog_key(run_id: &RunId) -> SlateKey {
run_catalog_root().with(run_id)
}
/// Extracts the run id from a full catalog marker key, or `None` when the key
/// is not exactly `runs/_index/by-start/<run_id>`.
pub(crate) fn parse_run_catalog_key(raw: &str) -> Option<RunId> {
let segments = SlateKey::segments(raw).collect::<Vec<_>>();
let ["runs", "_index", "by-start", run_id] = segments.as_slice() else {
return None;
};
run_id.parse().ok()
}
fn run_catalog_root() -> SlateKey {
SlateKey::new("runs").with("_index").with("by-start")
}
// Sequence keys zero-pad `seq` to six digits so lexicographic key order
// matches numeric seq order through `MAX_EVENT_SEQ`. Seek-based event listing
// (`run_events_range`) depends on this invariant, so event allocation rejects
@ -116,10 +85,6 @@ pub(crate) fn run_events_range(run_id: &RunId, start_seq: u32) -> Range<SlateKey
run_event_seq_prefix(run_id, start_seq)..end
}
pub(crate) fn sessions_by_id_prefix() -> SlateKey {
SlateKey::new("sessions").with("by-id").into_prefix()
}
#[cfg(test)]
pub(crate) fn session_by_id_key(session_id: &fabro_types::SessionId) -> SlateKey {
SlateKey::new("sessions").with("by-id").with(session_id)

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -5,8 +5,6 @@ mod blob_store;
mod error;
mod keyed_mutex;
mod keys;
mod legacy_blob_import;
mod legacy_run_history_import;
pub mod platform_records;
#[cfg(test)]
mod record;
@ -35,15 +33,6 @@ pub use fabro_types::{
BlobHash, EventEnvelope, PendingInterviewRecord, Run, RunProjection, StageId, StageProjection,
};
pub use keyed_mutex::{KeyedMutex, KeyedMutexGuard};
pub use legacy_blob_import::{
LegacyBlobImportError, LegacyBlobImportReport, LegacyBlobInventory, LegacyBlobInventoryError,
LegacyBlobVerificationError, LegacyBlobVerificationReport,
};
pub use legacy_run_history_import::{
LegacyRunHistoryDiagnostics, LegacyRunHistoryImportError, LegacyRunHistoryImportReport,
LegacyRunHistorySourceIdentity, LegacyRunHistorySourceIdentityError,
LegacyRunHistoryVerificationError, LegacyRunHistoryVerificationReport,
};
pub use platform_records::{
PlatformRecord, PlatformRecordHook, PlatformRecordKind, PlatformRecordStore, StagePosition,
StoredPlatformRecord,

View file

@ -32,8 +32,8 @@ use fabro_types::run_event::{
RunNoticeLevel, RunPairStartedProps, RunRunnableSource, RunStartedProps, RunSupersededByProps,
};
use fabro_types::{
BlobHash, DiffSummary, EventBody, GitIdentity, PairId, PairTarget, Principal, RunControlAction,
RunEvent, RunId, RunSpec, RunStatus,
BlobHash, DiffSummary, EventBody, GitIdentity, PairId, PairTarget, Principal,
PullRequestCreationId, PullRequestLink, RunControlAction, RunEvent, RunId, RunSpec, RunStatus,
};
use serde::{Deserialize, Serialize};
use sqlx::sqlite::{SqliteConnection, SqliteRow};
@ -125,9 +125,21 @@ pub enum PlatformRecordKind {
#[serde(rename = "checkpoint")]
#[strum(serialize = "checkpoint")]
Checkpoint,
#[serde(rename = "pull_request.requested")]
#[strum(serialize = "pull_request.requested")]
PullRequestRequested,
#[serde(rename = "pull_request.created")]
#[strum(serialize = "pull_request.created")]
PullRequestCreated,
#[serde(rename = "pull_request.failed")]
#[strum(serialize = "pull_request.failed")]
PullRequestFailed,
#[serde(rename = "pull_request.linked")]
#[strum(serialize = "pull_request.linked")]
PullRequestLinked,
#[serde(rename = "pull_request.unlinked")]
#[strum(serialize = "pull_request.unlinked")]
PullRequestUnlinked,
#[serde(rename = "notification.sent")]
#[strum(serialize = "notification.sent")]
NotificationSent,
@ -176,8 +188,19 @@ pub enum PlatformRecord {
/// record, written after the commit succeeds.
#[serde(rename = "checkpoint")]
Checkpoint(CheckpointRecord),
/// A pull request was asked for: the supervisor creates it.
#[serde(rename = "pull_request.requested")]
PullRequestRequested(PullRequestRequestedRecord),
#[serde(rename = "pull_request.created")]
PullRequestCreated(PullRequestCreatedRecord),
/// The requested pull request could not be created.
#[serde(rename = "pull_request.failed")]
PullRequestFailed(PullRequestFailedRecord),
/// An existing pull request was linked to the run by hand.
#[serde(rename = "pull_request.linked")]
PullRequestLinked(PullRequestLinkedRecord),
#[serde(rename = "pull_request.unlinked")]
PullRequestUnlinked(PullRequestLinkedRecord),
#[serde(rename = "notification.sent")]
NotificationSent(NotificationSentRecord),
#[serde(rename = "run.paired")]
@ -200,7 +223,11 @@ impl PlatformRecord {
Self::RunBranch(_) => PlatformRecordKind::RunBranch,
Self::GitIdentity(_) => PlatformRecordKind::GitIdentity,
Self::Checkpoint(_) => PlatformRecordKind::Checkpoint,
Self::PullRequestRequested(_) => PlatformRecordKind::PullRequestRequested,
Self::PullRequestCreated(_) => PlatformRecordKind::PullRequestCreated,
Self::PullRequestFailed(_) => PlatformRecordKind::PullRequestFailed,
Self::PullRequestLinked(_) => PlatformRecordKind::PullRequestLinked,
Self::PullRequestUnlinked(_) => PlatformRecordKind::PullRequestUnlinked,
Self::NotificationSent(_) => PlatformRecordKind::NotificationSent,
Self::RunPaired(_) => PlatformRecordKind::RunPaired,
}
@ -225,6 +252,10 @@ impl PlatformRecord {
| Self::InterviewAnswered(_)
| Self::RunBranch(_)
| Self::GitIdentity(_)
| Self::PullRequestRequested(_)
| Self::PullRequestFailed(_)
| Self::PullRequestLinked(_)
| Self::PullRequestUnlinked(_)
| Self::RunPaired(_) => None,
}
}
@ -356,6 +387,12 @@ pub struct InterviewAnsweredRecord {
pub principal: Option<Principal>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub channel: Option<String>,
/// The question's text, for a reader that shows the answer beside it.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub text: Option<String>,
/// The answer as the person gave it, rendered as text.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub answer: Option<String>,
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
@ -407,6 +444,48 @@ pub struct PullRequestCreatedRecord {
pub operation: Option<OperationKey>,
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct PullRequestRequestedRecord {
pub creation_id: PullRequestCreationId,
pub model: String,
#[serde(default)]
pub force: bool,
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct PullRequestFailedRecord {
#[serde(default, skip_serializing_if = "Option::is_none")]
pub creation_id: Option<PullRequestCreationId>,
pub error: String,
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct PullRequestLinkedRecord {
pub owner: String,
pub repo: String,
pub number: u64,
}
impl PullRequestLinkedRecord {
#[must_use]
pub fn link(&self) -> PullRequestLink {
PullRequestLink {
owner: self.owner.clone(),
repo: self.repo.clone(),
number: self.number,
}
}
#[must_use]
pub fn from_link(link: &PullRequestLink) -> Self {
Self {
owner: link.owner.clone(),
repo: link.repo.clone(),
number: link.number,
}
}
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct NotificationSentRecord {
pub route: String,
@ -786,6 +865,8 @@ fn interview_answered_record(
question: props.question_id.clone(),
principal,
channel: None,
text: Some(props.question.clone()),
answer: Some(props.answer.clone()),
}
}
@ -848,6 +929,8 @@ mod tests {
question: "q-1".to_string(),
principal: None,
channel: Some("web".to_string()),
text: None,
answer: None,
})
}
PlatformRecordKind::RunBranch => PlatformRecord::RunBranch(RunBranchRecord {
@ -893,6 +976,33 @@ mod tests {
operation: None,
})
}
PlatformRecordKind::PullRequestRequested => {
PlatformRecord::PullRequestRequested(PullRequestRequestedRecord {
creation_id: PullRequestCreationId::new(),
model: "gpt-5.4".to_string(),
force: false,
})
}
PlatformRecordKind::PullRequestFailed => {
PlatformRecord::PullRequestFailed(PullRequestFailedRecord {
creation_id: None,
error: "no remote".to_string(),
})
}
PlatformRecordKind::PullRequestLinked => {
PlatformRecord::PullRequestLinked(PullRequestLinkedRecord {
owner: "fabro-sh".to_string(),
repo: "fabro".to_string(),
number: 7,
})
}
PlatformRecordKind::PullRequestUnlinked => {
PlatformRecord::PullRequestUnlinked(PullRequestLinkedRecord {
owner: "fabro-sh".to_string(),
repo: "fabro".to_string(),
number: 7,
})
}
PlatformRecordKind::NotificationSent => {
PlatformRecord::NotificationSent(NotificationSentRecord {
route: "slack".to_string(),

View file

@ -355,31 +355,6 @@ impl RunSummaryStore {
Ok(())
}
#[cfg(test)]
pub(crate) async fn test_mark_run_history_activated(&self) -> Result<()> {
sqlx::query(
r"
INSERT INTO legacy_run_history_activation (
singleton, source_fingerprint, source_runs, source_events, activated_at_ms
) VALUES (1, zeroblob(32), 0, 0, 1)
ON CONFLICT(singleton) DO NOTHING
",
)
.execute(&self.pool)
.await?;
Ok(())
}
#[cfg(test)]
pub(crate) async fn test_is_run_history_tombstoned(&self, run_id: &RunId) -> Result<bool> {
Ok(sqlx::query_scalar(
"SELECT EXISTS(SELECT 1 FROM legacy_run_history_deletions WHERE run_id = ?)",
)
.bind(run_id.to_string())
.fetch_one(&self.pool)
.await?)
}
pub(crate) async fn acquire(&self) -> Result<PoolConnection<Sqlite>> {
Ok(self.pool.acquire().await?)
}
@ -571,31 +546,11 @@ ON CONFLICT(singleton) DO NOTHING
Ok(Some(run_id))
}
pub(crate) async fn delete_canonical(&self, run_id: &RunId, deleted_at_ms: i64) -> Result<()> {
// This transaction reads the activation marker before it writes the
// tombstone and run deletion. A deferred SQLite transaction can fail
// immediately when that read transaction is upgraded while another
// writer is active, bypassing the configured busy timeout. Reserve
// the write lock up front so concurrent deletes wait normally.
pub(crate) async fn delete_canonical(&self, run_id: &RunId) -> Result<()> {
// Reserve the write lock up front: a deferred transaction upgraded
// while another writer is active can fail at once, bypassing the
// configured busy timeout, so concurrent deletes wait normally.
let mut transaction = self.pool.begin_with("BEGIN IMMEDIATE").await?;
let activated: bool = sqlx::query_scalar(
"SELECT EXISTS(SELECT 1 FROM legacy_run_history_activation WHERE singleton = 1)",
)
.fetch_one(&mut *transaction)
.await?;
if activated {
sqlx::query(
r"
INSERT INTO legacy_run_history_deletions (run_id, deleted_at_ms)
VALUES (?, ?)
ON CONFLICT(run_id) DO UPDATE SET deleted_at_ms = excluded.deleted_at_ms
",
)
.bind(run_id.to_string())
.bind(deleted_at_ms)
.execute(&mut *transaction)
.await?;
}
sqlx::query("DELETE FROM runs WHERE id = ?")
.bind(run_id.to_string())
.execute(&mut *transaction)
@ -851,71 +806,6 @@ impl RunSummaryStore {
decode_event_rows_with_json(&rows, run_id)
}
pub(crate) async fn insert_imported_run_on_connection(
connection: &mut SqliteConnection,
entry: &ProjectedRun,
) -> Result<()> {
let record = PreparedRunSummary::from_entry(entry);
ensure_entry_identity(entry, &record, entry.last_seq)?;
if !(1..=keys::MAX_EVENT_SEQ).contains(&entry.last_seq) {
return Err(Error::RunHeadMismatch {
run_id: entry.run_id.to_string(),
expected_last_seq: entry.last_seq,
actual_last_seq: None,
});
}
insert_run_on_connection(connection, &record).await
}
pub(crate) async fn insert_imported_event_on_connection(
connection: &mut SqliteConnection,
run_id: &RunId,
payload: &EventPayload,
envelope: &EventEnvelope,
event_json: &str,
) -> Result<()> {
payload.validate(run_id)?;
let decoded = RunEvent::try_from(payload)?;
if envelope.event != decoded || envelope.event.run_id != *run_id {
return Err(run_event_mismatch(run_id, envelope.seq, "event_json"));
}
if !(1..=keys::MAX_EVENT_SEQ).contains(&envelope.seq) {
return Err(run_event_mismatch(run_id, envelope.seq, "seq"));
}
insert_event_json_on_connection(connection, run_id, envelope, event_json).await
}
pub(crate) async fn verify_current_run_on_connection(
connection: &mut SqliteConnection,
entry: &ProjectedRun,
) -> Result<()> {
let record = PreparedRunSummary::from_entry(entry);
ensure_entry_identity(entry, &record, entry.last_seq)?;
let row = sqlx::query(
r"
SELECT id, source_last_seq, created_at_ms, started_at_ms, last_event_at_ms, completed_at_ms,
status, archived_at_ms, parent_id, title, workflow_slug, workflow_name,
repository_name, automation_id, diff_files_changed, diff_additions, diff_deletions,
input_tokens, output_tokens, reasoning_tokens, cache_read_tokens, cache_write_tokens,
total_usd_micros, summary_json
FROM runs
WHERE id = ?
",
)
.bind(entry.run_id.to_string())
.fetch_optional(connection)
.await?
.ok_or_else(|| Error::RunNotFound(entry.run_id.to_string()))?;
let run = &record.run;
verify_run_field(&row, run, "id", &run.id.to_string())?;
verify_run_field(&row, run, "source_last_seq", &i64::from(record.last_seq))?;
// The run's row is written by its projector from Petri's records
// and the platform records; the legacy fold knows the lifecycle
// alone, so only the identity and the legacy guard are checked here.
Ok(())
}
pub(crate) async fn list_events_from_with_limit_on_connection(
connection: &mut SqliteConnection,
run_id: &RunId,
@ -1228,20 +1118,6 @@ fn decode_event_rows_with_json(
.collect()
}
fn verify_run_field<T>(row: &SqliteRow, run: &Run, field: &'static str, expected: &T) -> Result<()>
where
T: for<'row> sqlx::Decode<'row, Sqlite> + sqlx::Type<Sqlite> + PartialEq,
{
let stored: T = row.try_get(field)?;
if &stored != expected {
return Err(Error::RunSummaryMismatch {
run_id: run.id.to_string(),
field,
});
}
Ok(())
}
fn decode_event_row(
row: &SqliteRow,
expected_run_id: &RunId,
@ -2184,11 +2060,9 @@ mod tests {
.await
.unwrap();
transaction.commit().await.unwrap();
store.test_mark_run_history_activated().await.unwrap();
let blocker = store.pool.begin_with("BEGIN IMMEDIATE").await.unwrap();
let contender = store.clone();
let delete = tokio::spawn(async move { contender.delete_canonical(&id, 2).await });
let delete = tokio::spawn(async move { contender.delete_canonical(&id).await });
time::sleep(Duration::from_millis(25)).await;
assert!(
!delete.is_finished(),
@ -2198,7 +2072,6 @@ mod tests {
blocker.commit().await.unwrap();
delete.await.unwrap().unwrap();
assert!(!store.contains(&id).await.unwrap());
assert!(store.test_is_run_history_tombstoned(&id).await.unwrap());
}
#[tokio::test]

View file

@ -238,24 +238,11 @@ impl Database {
Ok(())
}
/// The run's projection: for a legacy run the reducer's fold of its
/// events; for a Petri run the projection its projector last committed
/// over Petri's records and the platform records, falling back to the
/// legacy fold (the lifecycle alone) until the first view pass commits.
/// The run's projection: the one its projector last committed over
/// Petri's records and the platform records, or `None` before the
/// first view pass commits (or for no such run).
pub async fn load_run_projection(&self, run_id: &RunId) -> Result<Option<Arc<RunProjection>>> {
let legacy = if let Some(active) = self.get_active_run(run_id).await {
active.projection_snapshot().await?
} else {
match self.run_summary_store.load_projection(run_id).await {
Ok(projected) => projected.projection,
Err(Error::RunNotFound(_)) => return Ok(None),
Err(error) => return Err(error),
}
};
if let Some(petri) = self.run_summary_store.load_petri_projection(run_id).await? {
return Ok(Some(petri));
}
Ok(Some(legacy))
self.run_summary_store.load_petri_projection(run_id).await
}
/// Install the wake-up called after a platform record of a Petri run is
@ -277,9 +264,7 @@ impl Database {
Some(active) => Some(active.state_lock.lock().await),
None => None,
};
self.run_summary_store
.delete_canonical(run_id, Utc::now().timestamp_millis())
.await?;
self.run_summary_store.delete_canonical(run_id).await?;
active_runs.remove(run_id);
Ok(())
}
@ -865,67 +850,6 @@ mod tests {
assert_eq!(read.as_deref(), Some(shared_blob.as_slice()));
}
#[tokio::test]
async fn activated_delete_tombstone_prevents_legacy_history_resurrection() {
let (directory, summaries) = make_run_summary_store().await;
let (_object_store, store) = make_store_with_run_summaries(Arc::clone(&summaries));
let run_id = test_run_id("run-1");
let created = event_payload(
"run-1",
"2026-03-27T12:00:00Z",
"run.created",
&serde_json::json!({
"settings": sample_run_spec("run-1").settings,
"graph": sample_run_spec("run-1").graph,
"provenance": test_support::test_run_provenance(),
}),
);
store
.put_unvalidated_legacy_run_event(&run_id, 1, created.as_value())
.await
.unwrap();
let run = store.create_run(&run_id).await.unwrap();
run.append_event(&created).await.unwrap();
summaries.test_mark_run_history_activated().await.unwrap();
store.delete_run(&run_id).await.unwrap();
assert!(
summaries
.test_is_run_history_tombstoned(&run_id)
.await
.unwrap()
);
assert!(store.open_run(&run_id).await.is_err());
let database = fabro_db::Database::connect(directory.path().join("fabro.sqlite3"))
.await
.unwrap();
let report = store
.import_legacy_run_history_into(database.pool())
.await
.unwrap();
assert_eq!(report.tombstoned_source_runs, 1);
assert_eq!(report.tombstoned_source_events, 1);
assert!(store.open_run(&run_id).await.is_err());
let verification = store
.verify_legacy_run_history_in(database.pool())
.await
.unwrap();
assert_eq!(verification.tombstoned_source_runs, 1);
assert_eq!(verification.tombstoned_source_events, 1);
let recreated = store.create_run(&run_id).await.unwrap();
recreated.append_event(&created).await.unwrap();
assert!(
store
.verify_legacy_run_history_in(database.pool())
.await
.is_err(),
"a tombstone and live canonical data must fail verification"
);
}
#[tokio::test]
async fn missing_run_open_paths_return_run_not_found() {
let (_object_store, store) = make_store();

View file

@ -1,470 +0,0 @@
use fabro_store::Database;
use fabro_types::{Principal, RunId, RunStatus, TerminalStatus};
use super::run_store::map_open_run_error;
use crate::error::Error;
use crate::event::{self, Event};
/// The canonical "run is archived — mutation rejected" error message. Shared
/// by the operations layer, the CLI rewind precheck, and the server HTTP
/// guards so the user sees the same actionable guidance everywhere.
#[must_use]
pub fn archived_rejection_message(run_id: &RunId) -> String {
format!("run {run_id} is archived; run `fabro unarchive {run_id}` to restore it and try again")
}
/// Returns `Err(Error::Precondition)` when the given status represents an
/// archived run. Use this at any mutation entry point that would otherwise
/// transition or emit events against the run (rewind, resume, etc.).
pub fn ensure_not_archived(archived: bool, run_id: &RunId) -> Result<(), Error> {
if archived {
Err(Error::Precondition(archived_rejection_message(run_id)))
} else {
Ok(())
}
}
/// Outcome of an `archive` call.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum ArchiveOutcome {
/// Event was appended; projection marks the run archived.
Archived { prior_status: TerminalStatus },
/// Run was already archived; no event emitted.
AlreadyArchived,
}
/// Outcome of an `unarchive` call.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum UnarchiveOutcome {
/// Event was appended; projection clears archive metadata.
Unarchived { restored_status: TerminalStatus },
/// Run was terminal but not archived; no event emitted. Symmetric with
/// `ArchiveOutcome::AlreadyArchived`.
NotArchived { status: RunStatus },
}
/// Archive a terminal run. Idempotent if already archived.
pub async fn archive(
store: &Database,
run_id: &RunId,
actor: Option<Principal>,
) -> Result<ArchiveOutcome, Error> {
let run_store = store
.open_run(run_id)
.await
.map_err(|err| map_open_run_error(run_id, err))?;
let projection = run_store
.state()
.await
.map_err(|err| Error::engine(err.to_string()))?;
let current = projection.status;
if projection.archived_at.is_some() {
return Ok(ArchiveOutcome::AlreadyArchived);
}
if !matches!(
current,
RunStatus::Succeeded { .. } | RunStatus::Failed { .. } | RunStatus::Dead
) {
return Err(Error::Precondition(format!(
"run {run_id} must be terminal (succeeded, failed, or dead) to archive; \
current status is {current}"
)));
}
event::append_event(&run_store, run_id, &Event::RunArchived { actor })
.await
.map_err(|err| Error::engine(err.to_string()))?;
let prior_status = current.terminal_status().ok_or_else(|| {
Error::engine(format!(
"run {run_id} passed archive precondition but had non-terminal status {current}"
))
})?;
Ok(ArchiveOutcome::Archived { prior_status })
}
/// Unarchive a previously archived run, restoring its prior terminal status.
/// Idempotent on terminal-but-not-archived runs (returns `NotArchived` without
/// emitting an event).
pub async fn unarchive(
store: &Database,
run_id: &RunId,
actor: Option<Principal>,
) -> Result<UnarchiveOutcome, Error> {
let run_store = store
.open_run(run_id)
.await
.map_err(|err| map_open_run_error(run_id, err))?;
let projection = run_store
.state()
.await
.map_err(|err| Error::engine(err.to_string()))?;
let current = projection.status;
if projection.archived_at.is_some() {
event::append_event(&run_store, run_id, &Event::RunUnarchived { actor })
.await
.map_err(|err| Error::engine(err.to_string()))?;
let prior = current.terminal_status().ok_or_else(|| {
Error::engine(format!(
"run {run_id} is archived but has non-terminal status {current}"
))
})?;
return Ok(UnarchiveOutcome::Unarchived {
restored_status: prior,
});
}
if matches!(
current,
RunStatus::Succeeded { .. } | RunStatus::Failed { .. } | RunStatus::Dead
) {
return Ok(UnarchiveOutcome::NotArchived { status: current });
}
Err(Error::Precondition(format!(
"run {run_id} is not archived (status: {current}); nothing to unarchive"
)))
}
#[cfg(test)]
mod tests {
use std::sync::Arc;
use std::time::Duration;
use fabro_store::Database;
use fabro_types::{
FailureReason, PetriAdmission, RunId, SuccessReason, TerminalStatus, fixtures, test_support,
};
use object_store::memory::InMemory;
use super::*;
fn memory_store() -> Arc<Database> {
Arc::new(fabro_store::test_support::test_database(
Arc::new(InMemory::new()),
"",
Duration::from_millis(1),
None,
))
}
async fn seed_succeeded(store: &Database, run_id: &RunId) {
let run_store = store.create_run(run_id).await.unwrap();
seed_created(&run_store, run_id).await;
seed_runnable(&run_store, run_id).await;
event::append_event(&run_store, run_id, &Event::RunStarting)
.await
.unwrap();
event::append_event(&run_store, run_id, &Event::RunRunning)
.await
.unwrap();
event::append_event(&run_store, run_id, &Event::WorkflowRunCompleted {
timing: fabro_types::RunTiming::wall_only(10),
artifact_count: 0,
status: "succeeded".to_string(),
reason: SuccessReason::Completed,
final_git_commit_sha: None,
final_patch: None,
diff_summary: None,
usage: None,
})
.await
.unwrap();
}
async fn seed_failed(store: &Database, run_id: &RunId) {
let run_store = store.create_run(run_id).await.unwrap();
seed_created(&run_store, run_id).await;
seed_runnable(&run_store, run_id).await;
event::append_event(&run_store, run_id, &Event::RunStarting)
.await
.unwrap();
event::append_event(&run_store, run_id, &Event::RunRunning)
.await
.unwrap();
let failure_event = Event::workflow_run_failed_from_error(
&crate::error::Error::engine("boom"),
fabro_types::RunTiming::wall_only(10),
FailureReason::WorkflowError,
None,
None,
None,
None,
);
event::append_event(&run_store, run_id, &failure_event)
.await
.unwrap();
}
async fn seed_running(store: &Database, run_id: &RunId) {
let run_store = store.create_run(run_id).await.unwrap();
seed_created(&run_store, run_id).await;
seed_runnable(&run_store, run_id).await;
event::append_event(&run_store, run_id, &Event::RunStarting)
.await
.unwrap();
event::append_event(&run_store, run_id, &Event::RunRunning)
.await
.unwrap();
}
async fn seed_created(run_store: &fabro_store::RunDatabase, run_id: &RunId) {
event::append_event(run_store, run_id, &Event::RunCreated {
run_id: *run_id,
title: None,
settings: serde_json::to_value(fabro_types::WorkflowSettings::default())
.unwrap(),
graph: serde_json::to_value(fabro_types::Graph::new("test")).unwrap(),
workflow_source: None,
labels: std::collections::BTreeMap::default(),
source_directory: None,
workflow_slug: None,
workflow_version_id: None,
target: None,
automation: None,
provenance: test_support::test_run_provenance(),
spec_blob: None,
git: None,
fork_source_ref: None,
retried_from: None,
parent_id: None,
web_url: None,
admission: PetriAdmission::default(),
})
.await
.unwrap();
}
async fn seed_runnable(run_store: &fabro_store::RunDatabase, run_id: &RunId) {
event::append_event(run_store, run_id, &Event::RunRunnable {
source: fabro_types::RunRunnableSource::StartRequested,
actor: None,
})
.await
.unwrap();
}
async fn current_status(store: &Database, run_id: &RunId) -> RunStatus {
let run_store = store.open_run_reader(run_id).await.unwrap();
run_store.state().await.unwrap().status
}
async fn is_archived(store: &Database, run_id: &RunId) -> bool {
let run_store = store.open_run_reader(run_id).await.unwrap();
run_store.state().await.unwrap().archived_at.is_some()
}
async fn event_count(store: &Database, run_id: &RunId) -> usize {
let run_store = store.open_run_reader(run_id).await.unwrap();
run_store.list_events().await.unwrap().len()
}
#[tokio::test]
async fn archive_on_succeeded_emits_event_and_transitions_to_archived() {
let store = memory_store();
let run_id = fixtures::RUN_1;
seed_succeeded(&store, &run_id).await;
let outcome = archive(&store, &run_id, None).await.unwrap();
assert_eq!(outcome, ArchiveOutcome::Archived {
prior_status: TerminalStatus::Succeeded {
reason: SuccessReason::Completed,
},
});
assert_eq!(
current_status(&store, &run_id).await,
RunStatus::Succeeded {
reason: SuccessReason::Completed,
}
);
assert!(is_archived(&store, &run_id).await);
let projection = store
.open_run_reader(&run_id)
.await
.unwrap()
.state()
.await
.unwrap();
assert_eq!(projection.status, RunStatus::Succeeded {
reason: SuccessReason::Completed,
});
assert!(projection.archived_at.is_some());
}
#[tokio::test]
async fn archive_on_failed_captures_failed_as_prior_status() {
let store = memory_store();
let run_id = fixtures::RUN_2;
seed_failed(&store, &run_id).await;
let outcome = archive(&store, &run_id, None).await.unwrap();
assert_eq!(outcome, ArchiveOutcome::Archived {
prior_status: TerminalStatus::Failed {
reason: FailureReason::WorkflowError,
},
});
assert_eq!(current_status(&store, &run_id).await, RunStatus::Failed {
reason: FailureReason::WorkflowError,
});
assert!(is_archived(&store, &run_id).await);
}
#[tokio::test]
async fn archive_on_already_archived_is_idempotent_and_emits_no_event() {
let store = memory_store();
let run_id = fixtures::RUN_1;
seed_succeeded(&store, &run_id).await;
archive(&store, &run_id, None).await.unwrap();
let events_before = event_count(&store, &run_id).await;
let outcome = archive(&store, &run_id, None).await.unwrap();
let events_after = event_count(&store, &run_id).await;
assert_eq!(outcome, ArchiveOutcome::AlreadyArchived);
assert_eq!(events_before, events_after);
}
#[tokio::test]
async fn archive_on_running_rejects_with_precondition_error() {
let store = memory_store();
let run_id = fixtures::RUN_1;
seed_running(&store, &run_id).await;
let err = archive(&store, &run_id, None).await.unwrap_err();
let Error::Precondition(message) = err else {
panic!("expected Precondition, got {err:?}");
};
assert!(
message.contains("must be terminal"),
"message should explain terminal requirement, got: {message}"
);
}
#[tokio::test]
async fn unarchive_restores_succeeded_and_clears_prior_status() {
let store = memory_store();
let run_id = fixtures::RUN_1;
seed_succeeded(&store, &run_id).await;
archive(&store, &run_id, None).await.unwrap();
let outcome = unarchive(&store, &run_id, None).await.unwrap();
assert_eq!(outcome, UnarchiveOutcome::Unarchived {
restored_status: TerminalStatus::Succeeded {
reason: SuccessReason::Completed,
},
});
let projection = store
.open_run_reader(&run_id)
.await
.unwrap()
.state()
.await
.unwrap();
assert_eq!(projection.status, RunStatus::Succeeded {
reason: SuccessReason::Completed,
});
}
#[tokio::test]
async fn unarchive_restores_failed_when_prior_was_failed() {
let store = memory_store();
let run_id = fixtures::RUN_2;
seed_failed(&store, &run_id).await;
archive(&store, &run_id, None).await.unwrap();
let outcome = unarchive(&store, &run_id, None).await.unwrap();
assert_eq!(outcome, UnarchiveOutcome::Unarchived {
restored_status: TerminalStatus::Failed {
reason: FailureReason::WorkflowError,
},
});
assert_eq!(current_status(&store, &run_id).await, RunStatus::Failed {
reason: FailureReason::WorkflowError,
});
}
#[tokio::test]
async fn unarchive_on_terminal_non_archived_run_is_idempotent_no_op() {
let store = memory_store();
let run_id = fixtures::RUN_1;
seed_succeeded(&store, &run_id).await;
let events_before = event_count(&store, &run_id).await;
let outcome = unarchive(&store, &run_id, None).await.unwrap();
let events_after = event_count(&store, &run_id).await;
assert_eq!(outcome, UnarchiveOutcome::NotArchived {
status: RunStatus::Succeeded {
reason: SuccessReason::Completed,
},
});
assert_eq!(events_before, events_after);
}
#[tokio::test]
async fn archive_on_unknown_run_returns_run_not_found() {
let store = memory_store();
let run_id = fixtures::RUN_3;
let err = archive(&store, &run_id, None).await.unwrap_err();
assert!(
matches!(err, Error::RunNotFound(_)),
"expected RunNotFound, got {err:?}"
);
}
#[tokio::test]
async fn unarchive_on_unknown_run_returns_run_not_found() {
let store = memory_store();
let run_id = fixtures::RUN_3;
let err = unarchive(&store, &run_id, None).await.unwrap_err();
assert!(
matches!(err, Error::RunNotFound(_)),
"expected RunNotFound, got {err:?}"
);
}
#[tokio::test]
async fn unarchive_on_running_rejects_with_precondition_error() {
let store = memory_store();
let run_id = fixtures::RUN_1;
seed_running(&store, &run_id).await;
let err = unarchive(&store, &run_id, None).await.unwrap_err();
let Error::Precondition(message) = err else {
panic!("expected Precondition, got {err:?}");
};
assert!(
message.contains("not archived"),
"message should explain run is not archived, got: {message}"
);
}
#[tokio::test]
async fn archive_unarchive_archive_cycle_produces_three_events() {
let store = memory_store();
let run_id = fixtures::RUN_1;
seed_succeeded(&store, &run_id).await;
let events_before = event_count(&store, &run_id).await;
archive(&store, &run_id, None).await.unwrap();
unarchive(&store, &run_id, None).await.unwrap();
archive(&store, &run_id, None).await.unwrap();
let events_after = event_count(&store, &run_id).await;
assert_eq!(events_after - events_before, 3);
assert_eq!(
current_status(&store, &run_id).await,
RunStatus::Succeeded {
reason: SuccessReason::Completed,
}
);
assert!(is_archived(&store, &run_id).await);
}
}

View file

@ -6,23 +6,24 @@
)
)]
use std::collections::{BTreeMap, HashMap};
use std::collections::HashMap;
use std::path::{Path, PathBuf};
use fabro_config::Storage;
use fabro_graphviz::graph::{AttrValue, Graph};
use fabro_store::platform_records::{
PlatformRecord, RunCreatedRecord, RunLifecycleKind, RunLifecycleRecord,
};
use fabro_store::{BlobStore, Database};
use fabro_template::TemplateContext;
use fabro_types::{
AutomationRef, BlobHash, ForkSourceRef, GitContext, ManifestPath, PetriAdmission, RunId,
RunProvenance, RunTarget, WorkflowSettings, WorkflowVersionId,
RunProvenance, RunStatus, RunTarget, WorkflowSettings, WorkflowVersionId,
};
use fabro_util::json::normalize_json_value;
use tokio::task::spawn_blocking;
use super::source::{ResolveWorkflowInput, WorkflowInput, resolve_workflow};
use crate::error::Error;
use crate::event::{self, Event, append_event};
use crate::pipeline::types::PersistOptions;
use crate::pipeline::{self, Persisted, TransformOptions, Validated};
use crate::records::RunSpec;
@ -378,6 +379,8 @@ pub async fn persist_create_run(
})
}
/// The run's first records: `run.created` with the spec Fabro built, and
/// the `submitted` lifecycle transition. Both wake the run's projector.
async fn persist_created_run(
store: &Database,
persisted: &Persisted,
@ -399,50 +402,32 @@ async fn persist_created_run(
write_optional_blob(&blob_store, definition_bytes.as_deref()),
async { blob_store.write(&spec_bytes).await.map_err(store_error) },
)?;
let _ = workflow_source;
let title = explicit_title.unwrap_or_else(|| fabro_types::infer_run_title(record.graph.goal()));
let first_event = Event::RunCreated {
run_id: record.run_id,
let mut spec = record.clone();
spec.definition_blob = definition_blob;
spec.spec_blob = Some(spec_blob);
let created = PlatformRecord::RunCreated(RunCreatedRecord {
spec,
title: Some(title),
settings: normalize_json_value(
serde_json::to_value(&record.settings).map_err(|err| Error::engine(err.to_string()))?,
),
graph: normalize_json_value(
serde_json::to_value(&record.graph).map_err(|err| Error::engine(err.to_string()))?,
),
workflow_source: (!workflow_source.is_empty()).then(|| workflow_source.to_string()),
labels: record
.labels
.clone()
.into_iter()
.collect::<BTreeMap<_, _>>(),
source_directory: record.source_directory.clone(),
workflow_slug: record.workflow_slug.clone(),
workflow_version_id: record.workflow_version_id,
target: record.target.clone(),
automation: record.automation.clone(),
provenance: record.provenance.clone(),
spec_blob: Some(spec_blob),
git: record.git.clone(),
fork_source_ref: record.fork_source_ref.clone(),
retried_from: None,
parent_id,
retried_from: None,
web_url,
admission: record.admission.clone(),
};
let run_store = event::create_run(
store,
&record.run_id,
&first_event,
record.run_id.created_at(),
)
.await
.map_err(|err| Error::engine_with_source("failed to create run store", err))?;
append_event(&run_store, &record.run_id, &Event::RunSubmitted {
definition_blob,
})
.await
.map_err(store_error)
});
let submitted = PlatformRecord::RunLifecycle(
RunLifecycleRecord::new(RunLifecycleKind::Submitted).with_status(RunStatus::Submitted),
);
let summaries = store.run_summary_store();
let platform_records = summaries.platform_records();
for platform_record in [created, submitted] {
platform_records
.append(&record.run_id, &platform_record, None)
.await
.map_err(store_error)?;
}
summaries.notify_platform_record(record.run_id);
Ok(())
}
async fn write_optional_blob(

View file

@ -1,19 +1,34 @@
mod archive;
mod create;
mod run_store;
mod source;
mod validate;
pub use archive::{
ArchiveOutcome, UnarchiveOutcome, archive, archived_rejection_message, ensure_not_archived,
unarchive,
};
pub use create::{
CompiledRun, CreateRunCompileInput, CreateRunPersistenceInput, CreateRunPersistenceMetadata,
CreatedRun, MaterializedRun, assemble_create_run_persistence_input, compile_admitted_run,
make_run_dir, materialize_admitted_run, persist_create_run,
};
use fabro_types::RunId;
pub use source::WorkflowInput;
pub use validate::{ValidateInput, validate};
pub use crate::error::Error;
pub use crate::transforms::RenderMode;
/// The canonical "run is archived — mutation rejected" error message. Shared
/// by the server's HTTP guards and the CLI so the user sees the same
/// actionable guidance everywhere.
#[must_use]
pub fn archived_rejection_message(run_id: &RunId) -> String {
format!("run {run_id} is archived; run `fabro unarchive {run_id}` to restore it and try again")
}
/// Returns `Err(Error::Precondition)` when the given status represents an
/// archived run. Use this at any mutation entry point that would otherwise
/// transition the run.
pub fn ensure_not_archived(archived: bool, run_id: &RunId) -> Result<(), Error> {
if archived {
Err(Error::Precondition(archived_rejection_message(run_id)))
} else {
Ok(())
}
}

View file

@ -1,11 +0,0 @@
use fabro_store::Error as StoreError;
use fabro_types::RunId;
use crate::error::Error;
pub(super) fn map_open_run_error(run_id: &RunId, err: StoreError) -> Error {
match err {
StoreError::RunNotFound(id) => Error::RunNotFound(id),
other => Error::engine(format!("failed to open run {run_id}: {other}")),
}
}

View file

@ -18,7 +18,6 @@ use tracing::{debug, info, warn};
use crate::outcome::format_cost as outcome_format_cost;
use crate::records::{Conclusion, RunSpec};
use crate::runtime_store::RunStoreHandle;
/// Maximum length of a PR title (Unicode scalar values).
const PR_TITLE_MAX_CHARS: usize = 72;
@ -333,7 +332,6 @@ pub async fn build_pr_content(
diff: &str,
goal: &str,
model: &str,
run_store: &RunStoreHandle,
llm_source: Arc<dyn CredentialProvider>,
catalog: Arc<Catalog>,
conclusion: Option<&Conclusion>,
@ -352,7 +350,6 @@ pub async fn build_pr_content(
diff,
goal,
model,
run_store,
catalog.as_ref(),
conclusion,
run_state,
@ -365,7 +362,6 @@ async fn build_pr_content_with_client(
diff: &str,
goal: &str,
model: &str,
run_store: &RunStoreHandle,
catalog: &Catalog,
conclusion: Option<&Conclusion>,
run_state: Option<&RunProjection>,
@ -373,18 +369,6 @@ async fn build_pr_content_with_client(
) -> Result<PrContent, String> {
info!("Building PR content");
let loaded_run_state = if run_state.is_none() {
run_store
.state()
.await
.inspect_err(|err| {
tracing::warn!(error = %err, "Failed to load run state from store for PR body");
})
.ok()
} else {
None
};
let run_state = run_state.or(loaded_run_state.as_ref());
let conclusion = conclusion.or_else(|| run_state.and_then(|state| state.conclusion.as_ref()));
let plan_text = run_state.and_then(read_plan_text);
let run_spec = run_state.map(|state| state.spec.clone());
@ -462,7 +446,6 @@ pub struct OpenPullRequestRequest<'a> {
pub model: &'a str,
pub draft: bool,
pub auto_merge: Option<AutoMergeOptions>,
pub run_store: &'a RunStoreHandle,
pub llm_source: Arc<dyn CredentialProvider>,
pub catalog: Arc<Catalog>,
pub conclusion: Option<&'a Conclusion>,
@ -616,7 +599,6 @@ pub async fn open_pull_request(
req.diff,
req.goal,
req.model,
req.run_store,
Arc::clone(&req.llm_source),
Arc::clone(&req.catalog),
req.conclusion,
@ -1073,12 +1055,11 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
#[tokio::test]
async fn build_pr_content_uses_in_memory_conclusion() {
let store = test_store();
let run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let _run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let PrContent { title, body } = build_pr_content_with_client(
"diff --git a/src/lib.rs b/src/lib.rs\n+fn new_feature() {}\n",
"Implement feature",
"mock-model",
&run_store.clone().into(),
&mock_catalog(),
Some(&make_test_conclusion()),
None,
@ -1152,7 +1133,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
"diff --git a/src/lib.rs b/src/lib.rs\n+fn new_feature() {}\n",
"Implement feature",
"mock-model",
&run_store.clone().into(),
&mock_catalog(),
Some(&make_test_conclusion()),
None,
@ -1251,7 +1231,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
"diff --git a/src/lib.rs b/src/lib.rs\n+fn new_feature() {}\n",
"Implement feature",
"mock-model",
&run_store.clone().into(),
&mock_catalog(),
Some(&make_test_conclusion()),
None,
@ -1271,12 +1250,11 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
#[tokio::test]
async fn build_pr_content_uses_explicit_llm_client() {
let store = test_store();
let run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let _run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let body = build_pr_content_with_client(
"diff --git a/src/lib.rs b/src/lib.rs\n+fn new_feature() {}\n",
"Implement feature",
"gpt-5.4",
&run_store.clone().into(),
&mock_catalog(),
Some(&make_test_conclusion()),
None,
@ -1327,14 +1305,12 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
let catalog = test_catalog_with_provider_base_url("openai", &server.url("/v1"));
let store = test_store();
let run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let run_store_handle: RunStoreHandle = run_store.into();
let _run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let PrContent { title, body } = build_pr_content(
"diff --git a/src/lib.rs b/src/lib.rs\n+fn new_feature() {}\n",
"Implement feature",
"gpt-5.4",
&run_store_handle,
llm_source,
catalog,
Some(&make_test_conclusion()),
@ -1522,7 +1498,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
model: "claude-sonnet-4-20250514",
draft: false,
auto_merge: None,
run_store: &harness.run_store,
llm_source: Arc::clone(&harness.llm_source),
catalog: harness.catalog.clone(),
conclusion: None,
@ -1554,14 +1529,13 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
#[tokio::test]
async fn build_pr_content_truncates_long_title() {
let store = test_store();
let run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let _run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let long_title = "x".repeat(200);
let payload = pr_content_json(&long_title, "Body content.");
let title = build_pr_content_with_client(
"diff --git a/src/lib.rs b/src/lib.rs\n+fn x() {}\n",
"Implement feature",
"mock-model",
&run_store.clone().into(),
&mock_catalog(),
Some(&make_test_conclusion()),
None,
@ -1578,13 +1552,12 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
#[tokio::test]
async fn build_pr_content_uses_default_title_when_generated_and_goal_titles_empty() {
let store = test_store();
let run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let _run_store = store.create_run(&fixtures::RUN_1).await.unwrap();
let payload = pr_content_json("", "Body content.");
let title = build_pr_content_with_client(
"diff --git a/src/lib.rs b/src/lib.rs\n+fn x() {}\n",
"## Plan:",
"mock-model",
&run_store.clone().into(),
&mock_catalog(),
Some(&make_test_conclusion()),
None,
@ -1675,7 +1648,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
"diff --git a/src/lib.rs b/src/lib.rs\n+fn x() {}\n",
"Implement feature",
"mock-model",
&run_store.clone().into(),
&mock_catalog(),
Some(&make_test_conclusion()),
None,
@ -1712,7 +1684,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
llm_source: Arc<dyn CredentialProvider>,
catalog: Arc<Catalog>,
creds: fabro_github::GitHubCredentials,
run_store: RunStoreHandle,
}
impl FallbackHarness {
@ -1912,7 +1883,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
llm_source,
catalog,
creds,
run_store: run_store.into(),
}
}
@ -1950,7 +1920,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
model: "gpt-5.4",
draft: false,
auto_merge: None,
run_store: &harness.run_store,
llm_source: Arc::clone(&harness.llm_source),
catalog: harness.catalog.clone(),
conclusion: None,
@ -1998,7 +1967,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
model: "gpt-5.4",
draft: false,
auto_merge: None,
run_store: &harness.run_store,
llm_source: Arc::clone(&harness.llm_source),
catalog: harness.catalog.clone(),
conclusion: None,
@ -2035,7 +2003,6 @@ capabilities = { text = true, tools = true, response_format = { json_object = tr
model: "gpt-5.4",
draft: false,
auto_merge: None,
run_store: &harness.run_store,
llm_source: Arc::clone(&harness.llm_source),
catalog: harness.catalog.clone(),
conclusion: None,