Skip to content

Failure testing

fdyno includes a repeatable process-fault smoke test:

bash scripts/fault-smoke.sh

The test checks a narrow set of recovery invariants. It is not Jepsen, a general linearizability checker, a replicated FoundationDB failover test, or disaster-recovery validation.

Default topology

Without FDB_CLUSTER_FILE, the script creates and owns every process it uses:

Component Default
fdyno Two processes built from the checked-out source
FoundationDB One isolated process, single redundancy, ssd-2 storage engine
Network Unique loopback ports selected for the run
Data One unique table and deterministic seed
Deadline Three-second fdyno operation timeout and a 180-second harness deadman
Artifacts One run-specific directory under .tmp/fault-runs/

The script records the source commit and binary SHA-256. It writes process logs, history.jsonl, final state, checker output, and a run summary before cleanup.

Process safety

The harness stores the exact PIDs it starts. Before sending a signal, it verifies the PID is its child and that the executable is the run-specific fdyno binary. It never kills a process by name.

The default FoundationDB helper uses a unique root, cluster file, and port. Cleanup is registered for normal exit, failure, interruption, and deadman expiry. The compatibility entry point scripts/restart-durability-test.sh now delegates to this harness. scripts/extenddb-test.sh also uses owned PIDs and an isolated FoundationDB root.

scripts/fdb-local.sh also verifies the live PID against the expected fdbmonitor identity. If a PID marker names a live process that fails verification, both start and stop return nonzero without signaling the process. The refusal happens before file setup, so PID, lock, cluster, and monitor-configuration files remain byte-for-byte unchanged. Only markers for dead processes are removed as stale.

An existing cluster is rejected unless the operator sets both variables explicitly:

FDB_CLUSTER_FILE=/path/to/fdb.cluster \
FAULT_ALLOW_EXTERNAL_FDB=1 \
  bash scripts/fault-smoke.sh

External-cluster mode does not stop or restart FoundationDB. Use only a disposable cluster and review the generated table name and artifact location before running it.

Recorded campaign

The workload writes one invoke event and one completion event for every operation. A completion is classified as:

  • ok when an API response confirms success;
  • definitive_failed only for a bounded set of non-retryable request rejections; or
  • ambiguous for transport failures, server failures, timeouts, and other outcomes that do not prove whether a write committed.

The campaign then performs these checks:

  1. Start fdyno instances A and B against the same FoundationDB database.
  2. Create a run-specific table, accounts, TTL configuration, tags, and an expired item.
  3. Run conditional creates and tokened account transfers.
  4. Proxy one tokened transaction through a test server that waits for upstream HTTP 200, then drops the response before the client receives it.
  5. Replay the exact request digest through the second instance and reconcile the stored result.
  6. Send SIGKILL only to owned instance A, verify through B, restart A, and verify again.
  7. In default owned-cluster mode, stop the single FoundationDB process, record the ordered outage result, restart it, and verify recovery.
  8. Run the invariant checker and its sensitivity fixtures, then clean up.

The checker requires:

  • one winner and one definite loser for the conditional create;
  • conservation and non-negative balances across acknowledged transfers;
  • the ambiguous tokened effect applied exactly once;
  • byte-equivalent request digests for every replay;
  • preserved TTL configuration and tags;
  • no resurrection of the expired item; and
  • ordered stop, outage, recovery, and post-recovery evidence when the harness owns FDB.

The checker self-test proves it rejects a duplicated effect, missing completion, mismatched replay, duplicated invariant row, and missing FDB recovery event.

CI gate

For implementation and test changes, CI runs:

  1. deterministic FoundationDB integration tests;
  2. the two-instance fault smoke against the runner's existing local cluster; and
  3. artifact upload even when the job fails.

CI sets FAULT_ALLOW_EXTERNAL_FDB=1, so it does not stop the runner-managed FoundationDB process. The default local command additionally exercises owned single-process FoundationDB outage and restart.

Exact committed evidence

One local run of commit 80e64196 completed on 22 September 2026 UTC with seed 20260921.

The two-instance fault smoke job also passed for that commit in GitHub Actions run 35697229842. The overall workflow was not green: the Alternator job still failed on unsupported-PITR expectations and one compression-ratio assertion. The fault result is not presented as a complete release pass.

Evidence Result
Source 80e641966aacd4fe1a3945448c57887d1f93da6a
Binary SHA-256 b9792fc5a79e40369772227d5e737ffd1200c04cd0b411de7b1458bb21bb154a
History 71 events
Checker fdyno-narrow-invariants-v1, valid, zero errors
Checker scope Token reconciliation, transfer conservation, conditional-create uniqueness
fdyno fault Owned instance A killed and restarted
HTTP fault Response dropped only after upstream HTTP 200
FoundationDB fault Owned single process stopped and restarted
General linearizability claimed No

The retained evidence files have these identities:

File SHA-256
history.jsonl cc10a8d9916928cc2ee3abd634f1d5c6785c414b512090260676b599d8b41734
checker.json 05fc6d844403620d1e14844306bb0868a365f638921d16442549bf0cb52483f1
checker-self-test.txt bdedff28cce57ab6c24e8047d79d633e54a89da0296711e910a9ec34a05593b4
qualification.json bff530f0a652592b431973ba0c0a643665066a7507d9361d39b069e7255d5623
state.json a456d967f3d93e38c545a058223c5ea1d510f61c91d8a0c209363fdea01c5063

These files can contain run-specific request and process details. Review them before sharing outside the test environment.

Remaining release gates

The successful smoke test does not cover:

  • a real fdyno-to-FoundationDB network partition in an isolated network namespace;
  • multi-node, replicated FoundationDB failover;
  • disk, zone, or site failure;
  • a broad concurrent history checked by Jepsen or Elle;
  • sustained fault load or latency during failover; or
  • backup to independent storage and restore into a fresh cluster.

Stopping the only FoundationDB process proves outage and restart recovery. It does not prove availability or durability during replicated failover. Keep these items open in a production acceptance review.