docs: correct storage upgrade behavior

This commit is contained in:
Scott Werner 2026-09-24 13:52:02 -04:00
parent a2b39a2408
commit 297b9e07a5

View file

@ -53,97 +53,45 @@ Common flags:
See [Server Configuration](/administration/server-configuration) for the full `settings.toml` reference.
### SQLite blob storage activation
### Storage and upgrades
On startup, Fabro activates SQLite as the only live content-addressed blob
store before it opens routes, schedulers, workers, webhooks, reapers, or the
ready callback. The activation inventories the exact legacy SlateDB blob
prefix and run history, then checks disk headroom for the rows not yet
imported, any required blob backup, and the projected post-import database
snapshot required by run-history activation. A warm restart with no pending
imports or backups only needs a small fixed headroom; on filesystems whose free
space cannot be determined the check is skipped with a warning. Fabro then
imports in bounded transactions, compares every legacy blob byte-for-byte with
SQLite, runs a live SQLite integrity check, and attempts a final WAL truncate
checkpoint. A busy final truncate logs a warning and startup continues so a
later checkpoint can finish after the blocking reader exits.
Boots that import new rows additionally re-verify every
legacy blob against SQLite and validate every SQLite blob row independently.
Any failure stops startup. Warm boots that import no rows skip that full target
scan: the import pass has already byte-compared every retained legacy row, and
SQLite-only blobs are hash-validated when read. Rows committed by an interrupted
import are retained so the next startup can resume, but the legacy source is
never modified and there is no fallback or dual read/write path.
Fabro stores current run history and content-addressed blobs in SQLite.
Artifact files use the separate local or S3 object store configured through
`[server.artifacts]` in [Server Configuration](/administration/server-configuration).
For a non-empty legacy inventory, the first activation also creates the
private sibling backup
`fabro.sqlite3.pre-blob-activation.bak`. Fabro writes the staging database
inside a private same-directory area, applies owner-only permissions, flushes
and validates it, then publishes the backup without overwriting an existing file.
A valid retained backup is revalidated on every warm restart and is preserved
as the original pre-activation safety artifact. If any legacy row is already
present in SQLite, a missing retained backup stops startup rather than silently
moving that rollback boundary forward. It is not a promise that an
older binary can safely resume after the activated server has accepted new
work; recovery after that boundary is forward-only. Empty legacy inventories
do not need this backup.
#### Legacy SlateDB settings
Keep both the unchanged legacy `blobs/sha256` prefix and the private activation
backup for at least 30 consecutive calendar days after the first successful
production activation. Cleanup is eligible only after a successful cold
activation, a later warm restart that revalidates the backup and byte-compares
every retained legacy blob against SQLite, and 30 days of production observation
with no unresolved inventory, import, verification, integrity, backup, or
checkpoint failure. Scott must review that evidence and explicitly authorize a
separate cleanup change. Day 30 is only the earliest eligibility date; nothing
is deleted automatically, and incomplete evidence extends the support window.
When Fabro loads active server settings that still contain `[server.slatedb]`,
it backs up the settings file and removes that section and its subtables.
For `settings.toml`, the backup is
`settings.toml.server-slatedb-migration.bak` beside the original file. Existing
backups are preserved; further backups use numbered suffixes such as
`settings.toml.server-slatedb-migration.1.bak`. A warning identifies the
rewritten file and its backup. Other settings are preserved by this cleanup,
and subsequent loads do nothing once the section is absent.
### SQLite run-history activation
This is configuration cleanup only. It neither imports nor deletes historical
SlateDB data, and the current server does not read that data. The settings
backup contains configuration, not run history or blobs.
Immediately after blob activation, and still before routes, schedulers,
workers, webhooks, reapers, or readiness are exposed, Fabro activates SQLite
as the sole authority for run existence, run events, and each run's current
projected row. The activation strictly validates and fingerprints the exact
legacy SlateDB run-event key/value stream, imports each complete run in its own
transaction, verifies every legacy history as an exact SQLite prefix, replays
and verifies every SQLite run independently, and runs a full SQLite integrity
check. It attempts a final WAL truncate checkpoint, but a blocking reader only
produces a warning because committed activation data remains durable in the
WAL. A source fingerprint or count change after activation stops startup. There
is no fallback or dual-read/write mode.
#### SQLite schema upgrades
For a non-empty legacy run history, the first activation creates and validates
the private sibling backup
`fabro.sqlite3.pre-run-history-activation.bak` before importing anything. The
backup is published without overwriting an existing file and is revalidated
on every restart. If import progress exists but that retained backup is
missing, startup stops. An empty legacy source is accepted without a backup
only when SQLite also has no unmarked run data. The activation marker stores
the source identity and first-success timestamp; retries preserve that
timestamp and repeat source, destination, backup, and integrity checks.
On startup, Fabro applies pending SQLite schema migrations before serving
normal traffic. Before changing a database that has previously applied
migrations, it writes a snapshot beside the database as
`fabro.sqlite3.pre-migration.bak`. A fresh database or a restart with no pending
migrations does not create a new snapshot. Each later schema upgrade replaces
this snapshot; failure to create it stops the upgrade.
After activation, creating a run commits `run.created`, the run's current row,
and its existence atomically. Later appends update the event log and current
row in one transaction, and live streams advance only after commit. Deleting
a migrated run commits a tombstone with the SQL deletion so the retained
legacy source cannot resurrect it during a restart.
The Petri transition drops the former SQL `run_events` table and its activation
bookkeeping. Those events are not converted into Petri run history. Current
runs use Petri records and Fabro platform records in SQLite.
Keep the unchanged legacy `runs/*/events/*` data and the private activation
backup for at least 30 consecutive calendar days after the persisted
first-success timestamp. Cleanup also requires successful cold and warm
activation evidence, production observation, backup and restore validation,
deletion/restart coverage, and explicit approval for a separate cleanup
change. Nothing is deleted automatically. The run-history activation backup
represents the database immediately before run-history import and can be used
to retry or recover the activation with a binary that knows the activated
schema. It is not a binary-downgrade artifact because it already contains the
new SQL migrations.
To return to the older binary, stop the server and restore the database's
`.pre-migration.bak` snapshot instead, then remove any `-wal` and `-shm`
siblings before starting the older binary. That snapshot was taken before the
new migrations were applied. Either recovery path loses writes accepted after
its snapshot, so make the rollback boundary explicit before restoring it.
For a database rollback, stop the server and restore the `.pre-migration.bak`
snapshot, then remove any `-wal` and `-shm` siblings before starting the
corresponding older binary. The snapshot holds the database state immediately
before the most recent schema upgrade. Restoring it loses database writes
accepted after that snapshot.
## Submitting runs