Consistency model¶
fdyno's normal data paths use FoundationDB transactions. Write preconditions and
query/scan reads use conflict-tracked, non-snapshot reads. A cached read-only
GetItem permits two narrow Snapshot Gets without making the transaction a
Snapshot. FoundationDB supplies strict serializability for each action
that completes in one transaction: one position in a serial history that
preserves real-time order for non-overlapping operations.
This guarantee belongs to a FoundationDB transaction boundary. It does not turn every multi-transaction workflow or sequence of HTTP requests into one atomic unit. That distinction is the center of this page.
Client-visible guarantee¶
For a one-transaction operation:
- all reads use one FoundationDB MVCC read version;
- writes remain private until commit;
- a successful commit makes all buffered keys visible together;
- a conflicting read/write history causes a retry rather than a successful stale read-modify-write;
- if operation A returns success before operation B starts, B cannot be serialized before A.
A read can exclude a concurrent write. “Current” means not stale with respect to writes that completed before the read began; it does not mean that a read waits for every concurrently in-flight writer.
sequenceDiagram
participant A as Writer A
participant F as FoundationDB
participant B as Reader B
A->>F: commit item + index + stream
F-->>A: success
Note over A,B: A completed before B began
B->>F: obtain read version and read
F-->>B: snapshot including A
Transaction scope¶
| Operation/workflow | Consistency boundary |
|---|---|
GetItem |
One read transaction. If TTL has expired, the result is hidden and cleanup may use a second transaction. |
One Query or Scan response page |
One read transaction/snapshot for the items materialized by that page. |
BatchGetItem |
All requested keys are read in one read transaction. |
PutItem, UpdateItem, DeleteItem |
Condition read, base write, index maintenance, and optional stream record share one transaction. |
TransactGetItems |
All target reads share one transaction/read version. |
TransactWriteItems |
All actions and optional idempotency token share one transaction. |
PartiQL ExecuteTransaction |
All accepted read-only statements or all accepted write statements share one transaction. |
BatchWriteItem |
Not an atomic API. Usually one transaction, but oversized valid batches can be split into independently committed groups. |
| Paginated query/scan | Snapshot per page, not one snapshot for all pages. |
| New GSI backfill | Transaction per bounded page while status is CREATING; atomic live maintenance once registered; complete only at ACTIVE. |
| Backup | Multiple read transactions pinned to one explicit read version, then multiple write transactions. The operation fails if the pinned version expires. |
| Restore/maintenance | Multiple bounded transactions with state/cursor rules; not one atomic request-wide transaction. |
See Transactions for the implementation of these boundaries.
FoundationDB conflict checks¶
Write callbacks and query/scan reads use normal FoundationDB key/range reads
that participate in conflict checking. Cached read-only GetItem may instead
Snapshot Get only its metadata revision and speculative item at the same read
version. It validates the revision before using those bytes and makes no write
based on them. These two Snapshot Gets do not establish write conflict
ranges; mutating operations retain non-snapshot revision and item reads.
No whole-transaction Snapshot is used.
For a cached, unfiltered base-table Query with explicit Limit≤100,
TTL disabled, and a non-CREATING table, fdyno overlaps an ordinary
metadata-revision Get with ordinary range reads in one read transaction.
It checks the authoritative revision before returning rows or a
schema-dependent error. A mismatch discards the speculative result and
re-resolves the table at the same read version. This changes scheduling,
not the transaction boundary or Snapshot policy. Other Query paths use
the authoritative fallback; some TTL/CREATING/cache paths may issue an
unused revision future before it. See Performance
for the current measurement boundary.
flowchart LR
R["transaction reads key/range at version V"]
W["concurrent transaction commits overlapping write"]
CHECK["FDB conflict check"]
RETRY["reject attempt and retry at a newer version"]
COMMIT["commit atomically"]
R --> CHECK
W --> CHECK
CHECK -->|"overlap"| RETRY
CHECK -->|"no invalidating overlap"| COMMIT
This mechanism covers both individual keys and range predicates used for partition
queries and scans. It is why a same-key ADD does not silently lose an accepted
increment, and why a condition such as attribute_not_exists(pk) has one winner
under contention.
Repository tests provide focused evidence:
- 50 concurrent conditional puts to one key produce exactly one winner.
- Concurrent same-item numeric increments reach the exact accepted-operation total.
- Concurrent list, set, and nested-path updates do not lose accepted changes.
- A failed transaction condition rolls back unrelated actions in that transaction.
These are targeted histories, not a complete formal verification suite.
States a transaction cannot expose¶
| State | Reason it is not visible |
|---|---|
| Dirty/intermediate read | Uncommitted buffered writes are not visible. |
| Fractured read | A reader sees either all or none of a committed multi-key write. |
| Lost read-modify-write update | The old item is read before the new image is written; an overlapping committed write conflicts and retries the attempt. |
| Write skew through read predicates | Non-snapshot key/range reads contribute conflict ranges; an invalidated attempt cannot commit unchanged. |
| Phantom within a transaction | Inserts into a range read by the transaction overlap that read range and force a retry. |
The guarantee applies to keys and ranges that the transaction reads through FoundationDB's normal transaction API. fdyno does not add a separate consistency algorithm; it depends on FoundationDB's conflict tracking and strict-serializable guarantees.
Read modes¶
DynamoDB exposes ConsistentRead, but fdyno does not switch to a weaker storage
path when it is false or omitted. GetItem, base-table Query/Scan, and batch
reads all use FoundationDB read transactions. The flag still affects compatibility
behavior such as consumed-capacity calculation.
For a GSI, fdyno rejects ConsistentRead=true with DynamoDB's validation behavior.
That API restriction does not imply an asynchronous fdyno index: active GSI entries
are still maintained in the base write transaction and read transactionally.
Avoid the shorthand “returns the latest value” without its concurrency qualifier. The rule is: a read cannot be ordered before a write that had already completed before the read began, but either order is valid for overlapping requests.
Secondary-index consistency¶
For an existing active GSI or LSI, a base-item change and all affected index entries commit together. Therefore:
- after a write returns success, a later index query cannot observe the old index entry while a base read observes the new item;
- deleting an item cannot leave a committed old index entry through a successful item mutation;
- a process restart cannot interrupt index maintenance between base and index commits.
This is stronger than an asynchronous propagation design, but it makes projection encoding and index writes part of write latency and transaction size.
Backfill exception¶
Adding a GSI to existing data is necessarily a multi-transaction workflow. The
index is CREATING while old items are backfilled. Live writes maintain it from the
metadata-registration commit onward, but some older items may still be absent until
the status becomes ACTIVE. The implementation only marks ACTIVE after all
backfill pages complete and resumes an interrupted CREATING index on startup.
Thus “indexes never lag” means normal mutations of a registered index are synchronous; it does not mean a newly declared index is instantly populated for pre-existing rows.
Streams and ordering¶
A change record is written with a FoundationDB versionstamped key in the same transaction as the state-changing item write. This gives two storage guarantees:
- A committed record cannot describe an aborted item write, and a committed state-changing write to a stream-enabled table does not depend on an asynchronous capture worker.
- Sequence numbers increase in FoundationDB commit order within a shard.
A generation with one shard has one table-wide stream order. With more than one, the table partition key selects a persisted shard. Changes to one partition stay ordered in that shard. There is no merged delivery order across shards.
This is a durable ordered log, not an exactly-once consumer protocol. Advancing the
returned cursor avoids re-reading a page in normal polling, and tests check that
behavior. A consumer that retries an old iterator or uses
AT_SEQUENCE_NUMBER can receive a record again and must checkpoint/idempotently
process records. See Change streams.
Pagination and filters¶
A Query or Scan page reads a bounded ordered prefix in one transaction. TTL and
key/range checks run while fdyno reads that prefix. Filters and projections then use
the copied page without changing its snapshot.
LastEvaluatedKey names where the next request should resume, not a retained read
version. Writes between pages can therefore cause the overall multi-page traversal
to reflect more than one database version. fdyno does not promise a repeatable
whole-table scan across HTTP page boundaries.
For an on-demand table backup, fdyno uses a different mechanism: it pins one FoundationDB read version and applies that version to each page transaction. If the version becomes too old, backup creation fails rather than silently continuing at a newer version.
Failure and availability behavior¶
Strict consistency removes the option to serve from a process-local stale copy. When FoundationDB cannot provide a read version or commit:
- the data request fails or reaches its deadline;
- the service does not acknowledge an uncommitted write;
- readiness reports the backend as unavailable, while liveness only reports process health;
- a commit-time cancellation may have an unknown outcome, which is a retry-safety question rather than an isolation violation.
fdyno does not accept operations while detached from FoundationDB. Fault tolerance and quorum behavior depend on the FoundationDB cluster configuration. fdyno adds no quorum or replication layer.
Not guaranteed¶
- No atomic multi-request session. Two calls are separate transactions. Use a transactional API when several items must change together.
- No cross-page snapshot by default. Pagination tokens are key positions, not read-version leases.
- No atomic large batch fallback.
BatchWriteItemmay partially succeed and reports the rest; useTransactWriteItemsfor all-or-nothing behavior. - No global order across configured stream shards. Ordering is per shard and, by hash routing, per partition key.
- No visibility guarantee for an index still
CREATING. Wait forACTIVE. - No availability without the backend. There is no local stale-read or queued write mode.
- No proof from conformance tests alone. API and concurrency tests exercise specific cases, but they are not a Jepsen/Elle proof under arbitrary process, network, and clock faults.
DynamoDB differences¶
| DynamoDB request | fdyno behavior |
|---|---|
| Eventually consistent/default read flag | Uses the same transactional storage read path as strong reads. |
| Strong base-table read | Transactional read; flag is accepted and reflected in compatibility accounting. |
| Strong GSI flag | Rejected for DynamoDB API compatibility, even though active index maintenance is synchronous. |
| Conditional single-item write | Condition and mutation are one FoundationDB transaction. |
| Transactional read/write | One FoundationDB transaction, subject to request and backend limits. |
| Stream consumption | Atomic record creation; cursor-based at-least-once-safe consumption rather than an exactly-once sink protocol. |
Code and tests¶
- Transaction selection and non-snapshot reads:
internal/dynodb/transaction_runner.go,internal/dynodb/item_ops.go,internal/dynodb/query_ops.go,internal/dynodb/txn_ops.go - Atomic index and stream maintenance:
internal/dynodb/fdb_store.go,internal/dynodb/item_ops.go,internal/dynodb/streams.go - GSI
CREATING/backfill/ACTIVEprotocol:internal/dynodb/table_ops.go - Page cursors and bounded reads:
internal/dynodb/fdb_store.go,internal/dynodb/query_ops.go - Pinned backup pages:
internal/dynodb/backup_ops.go - Conditional one-winner test:
test/extenddb/test_conditional_writes.py - Concurrent update tests:
test/extenddb/test_concurrency.py - Transaction rollback tests:
test/extenddb/test_transaction_operations.py,test/nubo-db/tests/tier2/transactions/transactWrite.test.ts - Stream iterator progression tests:
test/extenddb/test_streams.py
Strict serializability itself is inherited from FoundationDB rather than proved by fdyno code. The repository has no general-purpose history checker or injected network-partition test, so that remains the principal empirical evidence gap.
References¶
- FoundationDB documentation, Developer Guide: Transactions (transaction retries, conflict ranges, read versions, and versionstamps).
- Jingyu Zhou et al., FoundationDB: A Distributed Unbundled Transactional Key Value Store, SIGMOD 2021.
- Kyle Kingsbury, Consistency Models (terminology for strict serializability and weaker models).