From 297b9e07a5ed063da65f10814810785a87b0adf0 Mon Sep 17 00:00:00 2001 From: Scott Werner Date: Thu, 24 Sep 2026 13:52:02 -0400 Subject: [PATCH] docs: correct storage upgrade behavior --- docs/public/reference/server-operations.mdx | 114 ++++++-------------- 1 file changed, 31 insertions(+), 83 deletions(-) diff --git a/docs/public/reference/server-operations.mdx b/docs/public/reference/server-operations.mdx index 8cedb9592..5b9c31fa1 100644 --- a/docs/public/reference/server-operations.mdx +++ b/docs/public/reference/server-operations.mdx @@ -53,97 +53,45 @@ Common flags: See [Server Configuration](/administration/server-configuration) for the full `settings.toml` reference. -### SQLite blob storage activation +### Storage and upgrades -On startup, Fabro activates SQLite as the only live content-addressed blob -store before it opens routes, schedulers, workers, webhooks, reapers, or the -ready callback. The activation inventories the exact legacy SlateDB blob -prefix and run history, then checks disk headroom for the rows not yet -imported, any required blob backup, and the projected post-import database -snapshot required by run-history activation. A warm restart with no pending -imports or backups only needs a small fixed headroom; on filesystems whose free -space cannot be determined the check is skipped with a warning. Fabro then -imports in bounded transactions, compares every legacy blob byte-for-byte with -SQLite, runs a live SQLite integrity check, and attempts a final WAL truncate -checkpoint. A busy final truncate logs a warning and startup continues so a -later checkpoint can finish after the blocking reader exits. -Boots that import new rows additionally re-verify every -legacy blob against SQLite and validate every SQLite blob row independently. -Any failure stops startup. Warm boots that import no rows skip that full target -scan: the import pass has already byte-compared every retained legacy row, and -SQLite-only blobs are hash-validated when read. Rows committed by an interrupted -import are retained so the next startup can resume, but the legacy source is -never modified and there is no fallback or dual read/write path. +Fabro stores current run history and content-addressed blobs in SQLite. +Artifact files use the separate local or S3 object store configured through +`[server.artifacts]` in [Server Configuration](/administration/server-configuration). -For a non-empty legacy inventory, the first activation also creates the -private sibling backup -`fabro.sqlite3.pre-blob-activation.bak`. Fabro writes the staging database -inside a private same-directory area, applies owner-only permissions, flushes -and validates it, then publishes the backup without overwriting an existing file. -A valid retained backup is revalidated on every warm restart and is preserved -as the original pre-activation safety artifact. If any legacy row is already -present in SQLite, a missing retained backup stops startup rather than silently -moving that rollback boundary forward. It is not a promise that an -older binary can safely resume after the activated server has accepted new -work; recovery after that boundary is forward-only. Empty legacy inventories -do not need this backup. +#### Legacy SlateDB settings -Keep both the unchanged legacy `blobs/sha256` prefix and the private activation -backup for at least 30 consecutive calendar days after the first successful -production activation. Cleanup is eligible only after a successful cold -activation, a later warm restart that revalidates the backup and byte-compares -every retained legacy blob against SQLite, and 30 days of production observation -with no unresolved inventory, import, verification, integrity, backup, or -checkpoint failure. Scott must review that evidence and explicitly authorize a -separate cleanup change. Day 30 is only the earliest eligibility date; nothing -is deleted automatically, and incomplete evidence extends the support window. +When Fabro loads active server settings that still contain `[server.slatedb]`, +it backs up the settings file and removes that section and its subtables. +For `settings.toml`, the backup is +`settings.toml.server-slatedb-migration.bak` beside the original file. Existing +backups are preserved; further backups use numbered suffixes such as +`settings.toml.server-slatedb-migration.1.bak`. A warning identifies the +rewritten file and its backup. Other settings are preserved by this cleanup, +and subsequent loads do nothing once the section is absent. -### SQLite run-history activation +This is configuration cleanup only. It neither imports nor deletes historical +SlateDB data, and the current server does not read that data. The settings +backup contains configuration, not run history or blobs. -Immediately after blob activation, and still before routes, schedulers, -workers, webhooks, reapers, or readiness are exposed, Fabro activates SQLite -as the sole authority for run existence, run events, and each run's current -projected row. The activation strictly validates and fingerprints the exact -legacy SlateDB run-event key/value stream, imports each complete run in its own -transaction, verifies every legacy history as an exact SQLite prefix, replays -and verifies every SQLite run independently, and runs a full SQLite integrity -check. It attempts a final WAL truncate checkpoint, but a blocking reader only -produces a warning because committed activation data remains durable in the -WAL. A source fingerprint or count change after activation stops startup. There -is no fallback or dual-read/write mode. +#### SQLite schema upgrades -For a non-empty legacy run history, the first activation creates and validates -the private sibling backup -`fabro.sqlite3.pre-run-history-activation.bak` before importing anything. The -backup is published without overwriting an existing file and is revalidated -on every restart. If import progress exists but that retained backup is -missing, startup stops. An empty legacy source is accepted without a backup -only when SQLite also has no unmarked run data. The activation marker stores -the source identity and first-success timestamp; retries preserve that -timestamp and repeat source, destination, backup, and integrity checks. +On startup, Fabro applies pending SQLite schema migrations before serving +normal traffic. Before changing a database that has previously applied +migrations, it writes a snapshot beside the database as +`fabro.sqlite3.pre-migration.bak`. A fresh database or a restart with no pending +migrations does not create a new snapshot. Each later schema upgrade replaces +this snapshot; failure to create it stops the upgrade. -After activation, creating a run commits `run.created`, the run's current row, -and its existence atomically. Later appends update the event log and current -row in one transaction, and live streams advance only after commit. Deleting -a migrated run commits a tombstone with the SQL deletion so the retained -legacy source cannot resurrect it during a restart. +The Petri transition drops the former SQL `run_events` table and its activation +bookkeeping. Those events are not converted into Petri run history. Current +runs use Petri records and Fabro platform records in SQLite. -Keep the unchanged legacy `runs/*/events/*` data and the private activation -backup for at least 30 consecutive calendar days after the persisted -first-success timestamp. Cleanup also requires successful cold and warm -activation evidence, production observation, backup and restore validation, -deletion/restart coverage, and explicit approval for a separate cleanup -change. Nothing is deleted automatically. The run-history activation backup -represents the database immediately before run-history import and can be used -to retry or recover the activation with a binary that knows the activated -schema. It is not a binary-downgrade artifact because it already contains the -new SQL migrations. - -To return to the older binary, stop the server and restore the database's -`.pre-migration.bak` snapshot instead, then remove any `-wal` and `-shm` -siblings before starting the older binary. That snapshot was taken before the -new migrations were applied. Either recovery path loses writes accepted after -its snapshot, so make the rollback boundary explicit before restoring it. +For a database rollback, stop the server and restore the `.pre-migration.bak` +snapshot, then remove any `-wal` and `-shm` siblings before starting the +corresponding older binary. The snapshot holds the database state immediately +before the most recent schema upgrade. Restoring it loses database writes +accepted after that snapshot. ## Submitting runs