Failure testing¶
fdyno includes a repeatable process-fault smoke test:
The test checks a narrow set of recovery invariants. It is not Jepsen, a general linearizability checker, a replicated FoundationDB failover test, or disaster-recovery validation.
Default topology¶
Without FDB_CLUSTER_FILE, the script creates and owns every process it uses:
| Component | Default |
|---|---|
| fdyno | Two processes built from the checked-out source |
| FoundationDB | One isolated process, single redundancy, ssd-2 storage engine |
| Network | Unique loopback ports selected for the run |
| Data | One unique table and deterministic seed |
| Deadline | Three-second fdyno operation timeout and a 180-second harness deadman |
| Artifacts | One run-specific directory under .tmp/fault-runs/ |
The script records the source commit and binary SHA-256. It writes process logs,
history.jsonl, final state, checker output, and a run summary before cleanup.
Process safety¶
The harness stores the exact PIDs it starts. Before sending a signal, it verifies the PID is its child and that the executable is the run-specific fdyno binary. It never kills a process by name.
The default FoundationDB helper uses a unique root, cluster file, and port. Cleanup is
registered for normal exit, failure, interruption, and deadman expiry. The compatibility
entry point scripts/restart-durability-test.sh now delegates to this harness.
scripts/extenddb-test.sh also uses owned PIDs and an isolated FoundationDB root.
scripts/fdb-local.sh also verifies the live PID against the expected fdbmonitor
identity. If a PID marker names a live process that fails verification, both start and
stop return nonzero without signaling the process. The refusal happens before file
setup, so PID, lock, cluster, and monitor-configuration files remain byte-for-byte
unchanged. Only markers for dead processes are removed as stale.
An existing cluster is rejected unless the operator sets both variables explicitly:
External-cluster mode does not stop or restart FoundationDB. Use only a disposable cluster and review the generated table name and artifact location before running it.
Recorded campaign¶
The workload writes one invoke event and one completion event for every operation. A completion is classified as:
okwhen an API response confirms success;definitive_failedonly for a bounded set of non-retryable request rejections; orambiguousfor transport failures, server failures, timeouts, and other outcomes that do not prove whether a write committed.
The campaign then performs these checks:
- Start fdyno instances A and B against the same FoundationDB database.
- Create a run-specific table, accounts, TTL configuration, tags, and an expired item.
- Run conditional creates and tokened account transfers.
- Proxy one tokened transaction through a test server that waits for upstream HTTP 200, then drops the response before the client receives it.
- Replay the exact request digest through the second instance and reconcile the stored result.
- Send
SIGKILLonly to owned instance A, verify through B, restart A, and verify again. - In default owned-cluster mode, stop the single FoundationDB process, record the ordered outage result, restart it, and verify recovery.
- Run the invariant checker and its sensitivity fixtures, then clean up.
The checker requires:
- one winner and one definite loser for the conditional create;
- conservation and non-negative balances across acknowledged transfers;
- the ambiguous tokened effect applied exactly once;
- byte-equivalent request digests for every replay;
- preserved TTL configuration and tags;
- no resurrection of the expired item; and
- ordered stop, outage, recovery, and post-recovery evidence when the harness owns FDB.
The checker self-test proves it rejects a duplicated effect, missing completion, mismatched replay, duplicated invariant row, and missing FDB recovery event.
CI gate¶
For implementation and test changes, CI runs:
- deterministic FoundationDB integration tests;
- the two-instance fault smoke against the runner's existing local cluster; and
- artifact upload even when the job fails.
CI sets FAULT_ALLOW_EXTERNAL_FDB=1, so it does not stop the runner-managed
FoundationDB process. The default local command additionally exercises owned
single-process FoundationDB outage and restart.
Exact committed evidence¶
One local run of commit
80e64196
completed on 22 September 2026 UTC with seed 20260921.
The two-instance fault smoke job also passed for that commit in GitHub Actions
run 35697229842.
The overall workflow was not green: the Alternator job still failed on unsupported-PITR
expectations and one compression-ratio assertion. The fault result is not presented as
a complete release pass.
| Evidence | Result |
|---|---|
| Source | 80e641966aacd4fe1a3945448c57887d1f93da6a |
| Binary SHA-256 | b9792fc5a79e40369772227d5e737ffd1200c04cd0b411de7b1458bb21bb154a |
| History | 71 events |
| Checker | fdyno-narrow-invariants-v1, valid, zero errors |
| Checker scope | Token reconciliation, transfer conservation, conditional-create uniqueness |
| fdyno fault | Owned instance A killed and restarted |
| HTTP fault | Response dropped only after upstream HTTP 200 |
| FoundationDB fault | Owned single process stopped and restarted |
| General linearizability claimed | No |
The retained evidence files have these identities:
| File | SHA-256 |
|---|---|
history.jsonl |
cc10a8d9916928cc2ee3abd634f1d5c6785c414b512090260676b599d8b41734 |
checker.json |
05fc6d844403620d1e14844306bb0868a365f638921d16442549bf0cb52483f1 |
checker-self-test.txt |
bdedff28cce57ab6c24e8047d79d633e54a89da0296711e910a9ec34a05593b4 |
qualification.json |
bff530f0a652592b431973ba0c0a643665066a7507d9361d39b069e7255d5623 |
state.json |
a456d967f3d93e38c545a058223c5ea1d510f61c91d8a0c209363fdea01c5063 |
These files can contain run-specific request and process details. Review them before sharing outside the test environment.
Remaining release gates¶
The successful smoke test does not cover:
- a real fdyno-to-FoundationDB network partition in an isolated network namespace;
- multi-node, replicated FoundationDB failover;
- disk, zone, or site failure;
- a broad concurrent history checked by Jepsen or Elle;
- sustained fault load or latency during failover; or
- backup to independent storage and restore into a fresh cluster.
Stopping the only FoundationDB process proves outage and restart recovery. It does not prove availability or durability during replicated failover. Keep these items open in a production acceptance review.