Skip to content

Data model & keyspace

This page is the authoritative public inventory of fdyno-owned durable keys and values. fdyno maps DynamoDB items to FoundationDB's ordered byte keyspace. The layout supports direct item lookup, ordered range reads, and atomic updates to all stored forms of an item.

The paths below are conceptual tuple paths. FoundationDB's directory layer assigns opaque binary prefixes. A / separates tuple elements; it is not a byte stored in the key. FoundationDB's own directory metadata is outside this inventory.

Directory layout

Directory: dynodb/<table-name>
  ("m")                                             table metadata JSON
  ("i", tableHash [, tableSort])                    base item
  ("c", tableHash [, tableSort], chunkNumber)       base-item chunk
  ("x", indexName, indexHash [, indexSort],
        tableHash [, tableSort])                     secondary-index entry
  ("xc", indexName, rawIndexKey, chunkNumber)        index-value chunk
  ("v", versionstamp)                               stream record, one shard
  ("v", shardNumber, versionstamp)                  stream record, N > 1 shards
  ("vc", [shardNumber,] versionstamp, chunkNumber)   stream-value chunk

Service-wide directories
  dynodb_backups/<backup-arn>/...                     on-demand backups
  dynodb_restore_jobs/<job-id>/...                    durable restore jobs
  dynodb_txn_tokens/<client-request-token>            transaction deduplication
  dynodb_operation_receipts/...                       retry response receipts
  dynodb_ttl_cursors/<table-name>                     TTL sweep progress
  dynodb_worker_leases/<worker-name>                  maintenance ownership
flowchart TB
    ROOT["FoundationDB directory layer"]
    ROOT --> TABLES["dynodb"]
    TABLES --> TABLE["table directory"]
    TABLE --> META["m · metadata"]
    TABLE --> ITEM["i/c · items and chunks"]
    TABLE --> INDEX["x/xc · index entries and chunks"]
    TABLE --> CDC["v/vc · stream records and chunks"]
    ROOT --> BACKUPS["dynodb_backups"]
    ROOT --> RESTORES["dynodb_restore_jobs"]
    ROOT --> RETRY["tokens and operation receipts"]
    ROOT --> MAINT["cursors and worker leases"]

One directory per table separates table creation, deletion, and listing. Subspace tags prevent item, index, and stream range scans from reading unrelated records.

Primary keys and ordering

DynamoDB table and index key attributes may be string (S), number (N), or binary (B). attrToTupleElement maps them as follows:

DynamoDB key type Tuple element stored by fdyno Ordering mechanism
S FoundationDB tuple string UTF-8 byte order supplied by the tuple layer
B FoundationDB tuple byte string Byte order supplied by the tuple layer
N Custom order-preserving byte string Number is normalized, split into sign/exponent/digits, and encoded so byte order follows numeric order

The numeric transformation is necessary: tuple-encoding the number's text directly would put "100" before "2". sortableNumberKey gives negatives, zero, decimals, and positives an order that can be scanned in either direction. Query sort-order and reverse-pagination tests exercise numeric and binary keys.

A hash-only table item is:

("i", hashValue)

A table with a sort key is:

("i", hashValue, sortValue)

Every item in one partition shares the prefix ("i", hashValue), so Query uses a FoundationDB prefix range. fdyno converts ExclusiveStartKey back to the exact packed tuple and uses it as the next range boundary. It does not restart each page from the beginning of the partition.

Tuple order is not the entire number implementation

FoundationDB tuples delimit and order the byte elements. fdyno's own numeric encoding is what makes an N sort key numerically ordered inside that tuple.

Item values and the binary codec

The base-item value is encoded by codec.go, not stored as DynamoDB JSON. The format is type-tagged and supports all ten attribute forms:

S, N, B, BOOL, NULL, SS, NS, BS, L, and M.

Top-level attribute names and nested map keys are sorted before encoding. This makes map iteration order irrelevant and gives stable bytes for a given item representation. Lists and set slices retain their supplied order in the storage codec; the stronger canonicalization used for transaction request fingerprints is a separate path in idempotency.go.

The codec uses explicit lengths rather than delimiters, so arbitrary binary values and nested structures do not need escaping. Decode errors, such as an unknown type tag, truncated value, or missing chunk, fail the request rather than returning a partially decoded item. Unit tests cover empty items, every attribute type, BOOL=false, deterministic map encoding, and malformed payloads.

Large base items

FoundationDB values are bounded, so fdyno stores encoded base items in chunks of at most 10,000 bytes.

flowchart LR
    ENCODE["encoded item: 23,500 bytes"]
    MANIFEST["i/pk = chunked:3"]
    C0["c/pk/0 · 10,000 bytes"]
    C1["c/pk/1 · 10,000 bytes"]
    C2["c/pk/2 · 3,500 bytes"]

    ENCODE --> MANIFEST
    ENCODE --> C0
    ENCODE --> C1
    ENCODE --> C2

putItem first clears the old chunk range. A value at or below 10,000 bytes is stored directly under i/...; a larger value stores chunked:N under i/... and writes N values under c/.... All of those clears and writes occur in the caller's transaction, so readers cannot observe a new manifest with old or missing committed chunks. Reads issue the chunk Gets as FoundationDB futures, join them in chunk number order, then decode the reconstructed bytes.

Deleting an item clears both its i/... key and its complete c/... range in the same transaction.

Secondary indexes

GSIs and LSIs share the x subspace and maintenance code. An entry has this format:

("x", indexName, indexHash [, indexSort], tableHash [, tableSort])
    -> encoded projected item

Appending the complete base-table primary key is essential. If two items have the same index hash and sort values, their entries remain unique and pagination has a stable tiebreaker.

flowchart TB
    WRITE["state-changing item write"]
    subgraph TX["one FoundationDB transaction"]
      OLD["read old base item"]
      BASE["set or clear i/ and c/"]
      DROP["clear old x/ entries"]
      ADD["set new x/ projections"]
      STREAM["optionally set v/ record"]
      OLD --> BASE
      OLD --> DROP
      BASE --> ADD
      ADD --> STREAM
    end
    WRITE --> TX

Index behavior follows directly from the key and value design:

  • An item missing an index key is sparse and gets no entry.
  • Wrong-type, empty, or oversized index key values are rejected during the write.
  • ALL stores the full item; KEYS_ONLY stores table/index keys; INCLUDE adds the requested non-key attributes.
  • Querying one index partition is a prefix read on ("x", indexName, indexHash).
  • Base-item replacement first removes entries derived from the old image, then writes entries derived from the new image in the same transaction.

Large index projections

An index projection up to 10,000 encoded bytes is stored directly at its x key. A larger projection stores this tuple manifest at x:

("fdyno-index-chunks-v1", chunkCount, encodedSize, sha256)

The raw binary projection is split into values of at most 10,000 bytes under xc. Reads validate the marker, chunk count, size, SHA-256 digest, and decoded item. A missing or corrupt chunk fails the read. Overwrite, deletion, and index removal clear the corresponding chunk range.

Adding an index to existing data

A newly added GSI is first persisted as CREATING. From that commit onward, normal writes maintain it synchronously. Existing items are scanned in pages of 500; each page reads base items and writes corresponding index entries in one transaction. A concurrent update to an item read by the page conflicts and causes the page to be retried. Only after all pages finish does a final metadata transaction mark the index ACTIVE.

If the process stops during backfill, startup discovers CREATING indexes and re-runs the idempotent backfill from the beginning. Until ACTIVE, pre-existing items may not all have entries; strong index-read claims apply to an active, fully-backfilled index.

Change-stream records

For a stream-enabled table, a state-changing write adds a versioned binary change record under v using SetVersionstampedKey:

one shard:   ("v", incompleteVersionstamp) -> FDYNOC01 bytes or chunk manifest
N shards:    ("v", shardNumber, incompleteVersionstamp) -> FDYNOC01 bytes or chunk manifest

FoundationDB replaces the incomplete 10-byte transaction version at commit. fdyno adds the 2-byte user version and exposes the resulting 12 bytes as a fixed-width decimal sequence number.

The deterministic FDYNOC01 value stores an event byte, image-presence flags, creation time, and length-framed keys and images. Each item frame reuses the existing binary item encoding. An optional service identity uses length-prefixed strings.

A record up to 10,000 bytes is stored directly at v. A larger record stores this tuple manifest at v and raw binary chunks under vc:

("fdyno-cdc-chunks-v1", chunkCount, encodedSize, sha256, createdNanos)

Reads validate the chunk bounds, exact size, digest, and strict record decode. Unknown markers, events, flags, invalid item frames, and trailing bytes are rejected. Readers and the retention trimmer also accept stored JSON-formatted direct and chunked records.

The record and any chunks commit with the base item and index entries. Native item and transaction paths suppress an identical replacement. Deleting an absent item emits no record. For multiple shards, fdyno hashes canonical binary partition-key bytes so each partition remains in one shard. Ordering is commit order within a shard, not a merged order across shards. See Change streams.

Complete durable value inventory

Key Value format Bounds and integrity Lifetime
dynodb/<table>/m JSON snapshotTable One FoundationDB value; JSON and table state are validated Table lifetime
i/... and c/.../<n> Direct deterministic encodeItem, or ASCII chunked:<count> plus raw binary chunks 10,000-byte chunks; strict item decode; the base-item manifest has no size or digest Item lifetime
x/... and xc/.../<n> Direct encodeItem, or fdyno-index-chunks-v1 manifest plus raw chunks 10,000-byte chunks; count, size, SHA-256, and item decode Index-entry lifetime
v/... and vc/.../<n> Direct deterministic FDYNOC01, or fdyno-cdc-chunks-v1 manifest plus binary chunks; legacy direct and chunked JSON records remain readable 10,000-byte chunks; count, size, SHA-256, strict record and item decode Until stream retention removes it
dynodb_backups/<arn>/m JSON storage-format-3 metadata with generation and publication state Only COMPLETED metadata is readable as a backup Until DeleteBackup; stale staging work is collected
dynodb_backups/<arn>/generation/<id>/page/<page>/<n> FDYNOB01, item count, then length-prefixed encodeItem values 80,000-byte chunks; 8 MiB page limit; tuple manifest validates count, size, SHA-256, frames, and trailing bytes Immutable backup generation
dynodb_restore_jobs/<id>/m JSON restore job, version 1 Identity, state, counters, lease, and fence are validated Removed after completion; failed jobs retained for 24 hours
dynodb_restore_jobs/<id>/batch/<batch>/<n> Same FDYNOB01 framed item format as backup pages 80,000-byte chunks; 8 MiB batch limit; tuple manifest and strict item decode Restore-job lifetime
dynodb_txn_tokens/<token> Tuple ("v2-sha256", fingerprint, createdUnix) Exact tuple fields and 32-byte digest Ten-minute request-token window, then GC eligible
dynodb_operation_receipts/m/<id> and c/<id>/<n> v1 tuple metadata plus chunked JSON typed result; PutItem and DeleteItem results up to 4,096 bytes use v2 inline metadata with no chunks v1: 10,000-byte chunks, 8 MiB cap; v2: size and SHA-256 checked in one metadata value; both strictly decode typed JSON Request budget plus margin, at least one hour; pre-d489805b binaries cannot parse or collect v2
dynodb_worker_leases/<worker> Tuple (owner, expiryNanos) Tuple type checks Renewed while a worker owns the lease
dynodb_ttl_cursors/<table> Raw packed FoundationDB item key Used only as a scan position; item expiry is rechecked before deletion Cleared when a sweep wraps

Unknown markers, malformed manifests, missing chunks, digest mismatches, and trailing bytes fail closed.

The full writer and reader table is in docs/implementation/04-storage-and-fdb.md.

Example

Suppose Orders has table key (account S, orderNo N) and GSI ByStatus with (status S, createdAt N). This item:

{
  "account": {"S": "acct-7"},
  "orderNo": {"N": "100.50"},
  "status": {"S": "OPEN"},
  "createdAt": {"N": "1710000000"}
}

produces conceptual keys like:

("i", "acct-7", sortableNumber(100.5))
("x", "ByStatus", "OPEN", sortableNumber(1710000000),
      "acct-7", sortableNumber(100.5))

If the encoded item is larger than 10,000 bytes, the first key holds a chunk manifest and the bytes move to ("c", "acct-7", sortableNumber(100.5), n). If streams are enabled, the same transaction also writes a v/... record. A query for open orders reads the ByStatus/OPEN prefix in createdAt order, while the base-key suffix keeps equal timestamps distinct.

Storage invariants

  1. A committed base-item mutation and its synchronous index changes share one transaction.
  2. A stream record and its chunks, when emitted, share that transaction.
  3. A manifest and all of its chunks become visible atomically.
  4. Every index entry ends with the complete base key, so duplicate index values do not collide.
  5. Active index entries contain only their declared projection.
  6. Metadata under m is the durable source for schema, streams, tags, policies, TTL configuration, status, and restore identity.
  7. Backup and restore batches contain logical items. Restore derives current base and index keys instead of copying physical keys.

These invariants explain the consistency guarantees in Transactions and Consistency model.

Storage limits

  • The layout optimizes known primary-key and secondary-index access paths; it is not a general ad hoc indexing engine. A Scan remains a bounded walk over item or index ranges.
  • Synchronous index projections duplicate data and increase each write's transaction footprint in exchange for eliminating asynchronous index lag.
  • Chunking allows valid DynamoDB-sized base items to fit the value model, but adds multiple keys and reads per large item.
  • Numeric keys use fdyno's storage encoding. The raw FoundationDB keys are an internal format, not a public interchange or manual-editing API.
  • Directory prefixes and codec bytes are implementation details; clients should use the DynamoDB API, not read FoundationDB keys directly.

Code and tests

  • Table keys, item chunks, index keys, and index chunks: internal/dynodb/fdb_store.go and index_chunks_integration_test.go
  • Table metadata: internal/dynodb/fdb_persistence.go
  • Binary item codec and malformed values: codec.go, codec_test.go, and codec_fuzz_test.go
  • Framed backup/restore items and manifests: item_batch_codec.go and item_batch_codec_test.go
  • Backup generations: backup_generations.go and backup_generations_integration_test.go
  • Durable restore jobs: restore_jobs.go and restore_jobs_integration_test.go
  • Stream records and chunks: streams.go and cdc_chunks_integration_test.go
  • Request tokens and receipts: idempotency.go, operation_receipts.go, and their integration tests
  • Numeric key ordering: number.go, test/alternator/test_query.py, and test/extenddb/test_query_scan.py

These tests cover codec determinism, malformed input, missing chunks, digest mismatch, pagination, ordering, retry records, backup publication, and restore recovery. They do not make raw FoundationDB keys a public compatibility API.