Skip to content

Backup & restore

fdyno's backup APIs create table copies inside the same FoundationDB database. They provide a consistent table snapshot when the scan completes within FoundationDB's multiversion concurrency control (MVCC) history window. They do not provide an independent disaster-recovery copy or historical point-in-time recovery. Large restores can continue in the background, but there is no public job or progress API.

Choose the mechanism by failure domain:

Recovery need Use Reason
Copy or recover one table while the FoundationDB database is healthy fdyno CreateBackup and RestoreTableFromBackup Preserves a consistent item snapshot and selected table metadata.
Recover the current source into another table name CreateBackup, then RestoreTableFromBackup fdyno refuses to disguise a current clone as PITR.
Recover from FoundationDB database loss, corruption, or site loss FoundationDB-native backup to an independent destination fdyno backups share the source database's failure domain.

Historical PITR is unsupported

Enabling continuous backups and RestoreTableToPointInTime return explicit unavailable errors. fdyno does not retain 35 days of versions and never labels current source state with a requested historical timestamp.

Fresh-cluster native-backup drill (local only)

Commit 34cc0759 adds a disposable local mechanics drill. It starts two fresh, single-process ssd-2 FoundationDB clusters on the same host, backs up the source using fdbbackup to a file:// directory outside both cluster roots, then uses fdbrestore to populate the fresh target. It does not connect to or qualify a production cluster.

The retained run used FoundationDB 7.3.51 and a binary built from 18ea7e86; the retained drill script is byte-identical to the one committed in 34cc0759. Backup and restore reported completion without an error. A length-framed SHA-256 over all 87 FoundationDB keys and values matched between source and target. SDK checks verified table metadata, items, a 150 KB chunked item, a GSI, TTL configuration, tags, a CDC MODIFY old image, exact transaction-token replay after an intervening value change, and DeleteItem on the restored cluster. The off-repository evidence retains the binary, FDB configs, native backup, status/logs, script, and teardown; all 30/30 SHA256SUMS entries verified. The observed 35.600 s backup and 0.604 s restore are only same-host mechanics timings, not production RPO or RTO.

For an authorized disposable local rehearsal, with the local FoundationDB server/CLI, fdbbackup, fdbrestore, backup_agent, Go, Python, and uv available, run from the repository root:

FDYNO_DRILL_ARTIFACT_DIR=/path/to/new-evidence-dir \
  bash scripts/fdb-native-local-restore-drill.sh

After a successful run, verify SHA256SUMS, compare source-digest.txt and target-digest.txt, require completed backup/restore statuses, and inspect teardown.log for all three closed listeners. Keep the copied native backup and version/config records with that evidence; a command exit code alone is not a restore acceptance check. Treat backups and cluster files as sensitive artifacts rather than publishing them.

The evidence directory must not already exist. By default the script uses a new temporary evidence directory. It verifies process commands before signaling owned fdyno, backup agents, or FDB monitors, then checks that the source, target, and HTTP listeners have closed. If ownership checks fail, it refuses cleanup and preserves the disposable cluster root for manual inspection. Do not force kill a PID or delete its cluster files until its command and ownership are verified; retain the evidence and investigate the failed teardown.

This passes a local backup/restore mechanics check only. Both clusters, the file:// backup, and retained evidence are on one host. Off-host independent storage, site loss, a replicated FDB failover, network partitions, FDB TLS, recovery after actual data loss, and timed production RPO/RTO remain NOT RUN. A deployment owner must design and test those separately before relying on native backup for disaster recovery. See external components and exclusions.

CreateBackup contents

CreateBackup first opens the source table and pins one FoundationDB read version. It reads at most a 1 MB decoded page at that version. It serializes the page and writes generation-scoped 80 KB chunks with a size and SHA-256 manifest. The next page is read only after that commit. The process does not retain the full item set. A completed generation contains:

  • table metadata needed by the backup description and restore path;
  • TTL configuration;
  • independently verified item pages; and
  • item count and logical item-byte size.

The durable backup directory is dynodb_backups/<backup-arn>. A random generation is reserved as STAGING; each verified page and its progress update commit together, and one final transaction publishes COMPLETED. ListBackups and DescribeBackup read the small metadata record, and only a completed generation is usable.

The single pinned version makes the item set internally consistent: concurrent commits are either before or after the snapshot version. The backup does not copy live index-entry keys; restore rebuilds index entries from items. It also does not copy change-stream records, idempotency-token records, worker state, or other tables.

A backup name maps to a deterministic ARN. A second completed CreateBackup with the same name returns ResourceInUseException; it is not a successful idempotent replay. If the original HTTP response was lost, list or describe the expected ARN before choosing another name.

Snapshot size and MVCC limit

FoundationDB normally retains old read versions for only a short MVCC window (commonly about five seconds). fdyno pages reads at the pinned version with a five-second transaction timeout and a limited retry count. If the version ages out, backup creation returns a ValidationException rather than combining pages from different versions.

Paging removes the one-transaction scan limit, but it does not make the in-layer snapshot unbounded:

  • memory is bounded to one decoded page and one encoded page, not the table;
  • scan completion still depends on the pinned version remaining readable;
  • item decoding and FoundationDB load affect elapsed time;
  • backup chunks and source data consume capacity in the same cluster; and
  • there is no progress endpoint or cancellation API.

Use FoundationDB-native backup for tables that cannot reliably be scanned within the version window or for full-cluster recovery. The repository's local rehearsal above does not validate the commands, destination, security, or topology of an operator's production backup plan.

On-demand backup workflow

A safe table-backup procedure is:

  1. Check the failure domain. Confirm that an in-cluster table copy satisfies the recovery requirement. If cluster loss is in scope, use an independent native backup as well.
  2. Check headroom. Review FoundationDB health, storage, and process memory. A backup reads the full table and writes another full logical copy.
  3. Create a unique backup name. Record the returned ARN outside the source table. The call is synchronous and returns only after metadata is published.
  4. Verify the object. Use DescribeBackup; require BackupStatus=AVAILABLE and review item count and size. These checks prove publication, not recoverability.
  5. Exercise restore. Periodically restore to a new target name and run schema, count, sampled-data, index, and application-invariant checks.
  6. Apply retention. DeleteBackup removes a completed named backup. There is no scheduled retention worker for completed backups.

A process failure can leave a STAGING generation, but it is invisible to backup readers and cannot collide with another generation's keys. Failed owners remove their own reservation; a bounded stale-generation collector can reclaim abandoned work.

Restore from an on-demand backup

RestoreTableFromBackup requires a new target table name. It streams one verified backup page at a time into immutable, size-planned restore batches and rebuilds indexes from logical items rather than copying physical index keys.

For a small restore, target creation and all item writes fit in one FoundationDB transaction. A larger restore uses a durable job:

  1. stage all source batches before exposing a target;
  2. publish a target as CREATING, which ordinary data operations cannot open;
  3. apply each batch with an expiring lease, fencing generation, target identity check, and progress update in the same transaction; and
  4. atomically publish ACTIVE with the final batch.

Every server resumes READY or RUNNING jobs at startup and on its recovery interval. A process exit does not expose a partial table or require the original read version. DescribeTable can show CREATING; DeleteTable cancels that target and its job. The initiating HTTP request can still end before background recovery finishes, so wait for ACTIVE and validate before cutover.

Metadata restored and omitted

The restore path creates a new table identity. Current behavior is:

Property Restore behavior
Items Restored from the consistent backup snapshot.
Primary key and attribute definitions Preserved.
GSIs and LSIs Preserved and rebuilt, unless restore overrides replace their definitions.
Billing mode and provisioned throughput Preserved unless overridden.
TTL specification Preserved. Expired items in the snapshot may immediately be hidden or reclaimed after restore.
Tags Preserved.
Deletion protection and table class Preserved.
PITR-enabled flag Cleared; historical PITR is unsupported.
Contributor Insights and Kinesis destination state Preserved as metadata. Kinesis delivery is not implemented by fdyno.
Stream specification, stream ARN, and stream records Not restored. The backup captures stream metadata internally, but the target construction path does not apply it. Enable a new stream explicitly after validation if required.
Stored resource policy and revision Not restored. Resource policies are not enforced in any case.
Table ARN, table ID, name, creation time Recreated for the target.
Existing backup objects Not attached or copied to the target.

Restore writes do not generate change-stream records. If downstream consumers need a bootstrap, take it from the restored table separately and define the handoff to a new stream.

Historical PITR is rejected

fdyno has no historical version journal. UpdateContinuousBackups refuses enablement, DescribeContinuousBackups reports DISABLED without fictitious earliest/latest timestamps, and RestoreTableToPointInTime returns PointInTimeRecoveryUnavailableException after request/source/target validation.

This behavior prevents a requested historical time from silently producing current source data. Use immutable on-demand backups at tested recovery points. Other options are a validated change-replay system or FoundationDB-native backup and restore.

Validation after restore

Before cutover, check more than table status:

  1. Confirm ACTIVE and the expected key schema.
  2. Compare item counts cautiously; TTL can hide or remove expired items and counts can change if the source application remained active after the backup point.
  3. Query every restored secondary index and verify projection-sensitive fields.
  4. Sample large/chunked items and all DynamoDB attribute types used by the application.
  5. Verify tags, TTL, deletion protection, billing metadata, and intended overrides.
  6. Reapply any required resource policy and create a new stream explicitly.
  7. Run application-level invariants and a read-only canary before switching traffic.
  8. Keep the source and backup until rollback criteria expire.

There is no atomic traffic cutover, alias, or rename operation in fdyno. Endpoint or table-name switching belongs to the application/deployment layer.

Observability and incident checks

Backup and restore expose API status and table status only. They do not emit bytes copied, pages completed, estimated completion time, or dedicated metrics.

When creation fails, distinguish these cases:

  • ValidationException mentioning the read-version window: the consistent source scan aged out; reduce table size/load or use FoundationDB-native backup.
  • ResourceInUseException after a timeout: the first create may have completed; describe the expected backup before retrying.
  • Internal/FoundationDB error before publication: a STAGING generation may remain temporarily. It is not readable and the bounded staging collector can remove it.

When restore fails, inspect the target and server logs before retrying. A CREATING target is fenced from ordinary data operations. Recovery should resume its durable job; deleting the target cancels the job. Never mark the table active manually.

Finally, test full FoundationDB recovery independently. An fdyno backup that lists as AVAILABLE is still lost with the database that stores it.