Reliable writes & idempotency¶
A write can commit even when the client receives a timeout or loses the HTTP
response. Design retries around that ambiguity. For effects that must occur once,
use TransactWriteItems with a stable ClientRequestToken; do not rely on an error
to mean “nothing happened.”
A lost response makes the outcome unknown¶
The path has two independent commit boundaries from the caller's perspective:
client ── HTTP ──> fdyno ── FoundationDB transaction ──> committed state
<──────── response <──────────────────────────────
A failure before FoundationDB commit has no effect. A failure after commit leaves a
durable effect. If the fdyno-to-FoundationDB connection or the client response fails
at the boundary, neither the caller nor fdyno can always distinguish those cases.
FoundationDB names the database-side case commit_unknown_result.
fdyno retries retryable FoundationDB errors by running the transaction body again.
Foreground context-aware operations use a 30-second budget by default, configurable
with DYNODB_OPERATION_TIMEOUT. If the context expires after commit was attempted,
fdyno returns an InternalServerError whose message says the transaction outcome
may be unknown. A load balancer or client timeout can hide that message and create
the same ambiguity.
The FoundationDB transaction remains atomic in every case. “Unknown” means either all effects committed or none did. It never means that half a base/index/stream mutation is visible.
Classify the mutation before retrying¶
Idempotence is about applying the same request effect more than once. It does not mean a retry is safe in the presence of an unrelated concurrent writer.
| Pattern | Repeated state effect | Retry guidance |
|---|---|---|
PutItem with the same complete item |
Converges to the same value. | State-idempotent if overwriting a newer concurrent value is acceptable. |
DeleteItem for the same key |
The second delete sees no item. | State-idempotent, subject to concurrent recreation. |
SET path = :fixed or REMOVE path |
Converges to the same field state. | State-idempotent if concurrent updates cannot be overwritten. |
Numeric ADD, SET n = n + :delta |
Reapplies the delta. | Not idempotent; an unknown retry can double-apply. |
list_append or other old-value-derived update |
Derives another result from the already changed item. | Not idempotent. |
| Add fixed members to a set, or delete fixed set members | Union/removal converges for the same members. | State-idempotent, but still subject to concurrent writes to the same attribute. |
Conditional create with attribute_not_exists |
The effect occurs at most once. | Safe from duplicate creation, but a retry after a committed first attempt reports condition failure rather than the original success. |
BatchWriteItem |
Contains only puts/deletes, but may commit in chunks. | Retry only UnprocessedItems; after a lost whole response, reconcile concurrency-sensitive keys. |
TransactWriteItems without a token |
The transaction is atomic once, but the entire request can execute again. | Add a token before retrying non-idempotent effects. |
TransactWriteItems with a stable token |
The token and effects commit atomically. | Retry the identical request with the same token within ten minutes. |
Within one accepted mutation, fdyno allocates an internal operation ID and stores
a typed result receipt in the same FoundationDB transaction as the mutation.
On an ambiguous FoundationDB commit, its internal retry checks that receipt
before reapplying effects and can replay the original result. PutItem,
UpdateItem, DeleteItem, tokenless write transactions, PartiQL mutations, and
BatchWriteItem constituent transactions use this mechanism. It is not a client
token: a new HTTP request gets a new operation ID and may overwrite a concurrent
writer, reapply an increment, or return a different condition result. Use a
stable public token where available or reconcile application state.
Small PutItem and DeleteItem typed results now use inline v2 receipts;
larger results and other mutations use chunked v1. d489805b and later
binaries read both formats, but pre-d489805b binaries cannot parse v2.
Coordinated replacement, backup, and rollback restrictions are in
Deployment & operations.
Tokened transactions¶
For TransactWriteItems, fdyno computes a SHA-256 fingerprint of the request
parameters other than the token. It reads the token record, applies every mutation,
and stores the fingerprint and creation time in one FoundationDB transaction.
That atomic placement closes the unknown-outcome gap:
- If the transaction did not commit, neither effects nor token exist. The retry can execute.
- If it committed, both effects and token exist. A retry reads the token and returns a replay result without applying writes again.
- If the same live token is reused with different request parameters, fdyno returns
IdempotentParameterMismatchException.
The record is in a database-wide dynodb_txn_tokens directory, so deduplication
survives fdyno restart and works across replicas using the same FoundationDB
database. It also prevents a second stream mutation from the replayed transaction.
The window is ten minutes from the server-recorded creation time. After that,
readTxnToken treats the record as expired and the same token can execute as a new
request, whether or not token garbage collection has physically removed it. Keep
fdyno host clocks synchronized; the expiration decision uses process wall time.
Token GC is optional. DYNODB_TXN_TOKEN_GC_INTERVAL bounds storage growth, and
DYNODB_TXN_TOKEN_GC_MAX_SCAN bounds each pass. Semantic expiry does not depend on
the worker. Token records without valid timestamps fail closed and are not deleted
automatically.
Client retry procedure¶
- Generate a unique token before the first submission. The API accepts at most 36 characters.
- Persist the token and exact request payload in the same durable application state that records the intent. Do not generate a new token per network attempt.
- Submit all dependent changes in one
TransactWriteItemsrequest with that token. - On timeout, connection loss, or an explicit unknown-outcome error, resend the identical request with the same token.
- Complete the workflow only after success or after application-specific reconciliation proves the intended state.
- Stop automatic retries before the ten-minute server window expires. If recovery can exceed that window, add an application idempotency key/ledger in the data model; the fdyno token alone is insufficient.
The client must retain the original payload because changing return-capacity options, item values, expressions, or another fingerprinted parameter while keeping the token is a mismatch.
Conditional write retries¶
A conditional write is useful when the item itself carries a durable operation ID or state version. For example, “create only if absent” prevents duplicate creation, and “update only if version is 7” prevents an old retry from overwriting version 8.
Under an unknown result, however, the first attempt may have committed and the retry may fail its condition. The client must interpret that result by reading the item and checking the expected operation ID/version. Treating every condition failure as a previous success would be incorrect because another writer may have caused it.
This compare-and-set pattern is appropriate for one item. Use a tokened transaction when the invariant spans items or tables.
Concurrency and streams¶
A request can be idempotent in isolation yet unsafe against concurrency. Retrying a
fixed PutItem after another writer has changed the item restores the old request's
value. Use a condition on an application version when lost-update prevention matters.
Change streams record committed state changes, not client intents. A fixed-value retry that finds the same value emits no second record. A non-idempotent retry that commits another update emits another correctly ordered record. Stream consumers must still use durable checkpoints and idempotent sinks because reprocessing after a consumer crash is independent of write-side idempotency.
Batch writes¶
BatchWriteItem is non-atomic. fdyno attempts one FoundationDB
transaction, then falls back to validated, approximately 4 MiB chunks if the request
hits FoundationDB's transaction-size limit. Failed chunks are returned as
UnprocessedItems.
Retry the returned UnprocessedItems, not the entire original request. Internal
receipts guard each attempted FoundationDB transaction, including size-split
constituents, but there is no client-visible token or replay of an entire batch
after its HTTP response is lost. Puts and deletes converge only if no concurrent
writer makes replay harmful. For stronger semantics, split the workflow into
tokened transactions that fit the transaction limits.
Token scope and recovery¶
The fdyno token provides effect deduplication only:
- for
TransactWriteItems; - within ten minutes;
- while the same FoundationDB token state remains available; and
- for an identical request fingerprint.
An on-demand table backup excludes the database-wide token and operation-receipt directories. A FoundationDB-level backup can include them if the full fdyno keyspace is protected. Failover token and receipt continuity depends on restoring the corresponding global directories.
There is no public idempotency token for single-item writes, PartiQL mutations,
or BatchWriteItem. Internal receipts replay a typed result only inside one
accepted request's retry loop; they do not cache client responses across new HTTP
requests. Tokened transaction replay may report read rather than write capacity.
Metrics and tests¶
GET /metrics retains process-local JSON counters. Unauthenticated
GET /metrics/prometheus adds transaction attempts, FDB conflict (1020), unknown
result (1021), receipt-replay and token-replay/mismatch counters plus fixed-bucket
per-action latency histograms. A replay hit counts a transaction attempt, not
a unique HTTP operation; a 1021 result is distinct from terminal ambiguous timeout.
Both routes share the data listener and must be restricted at the proxy. Preserve
application token IDs and outcomes in separate telemetry without logging item
payloads or secrets. See Operations.
The codebase tests strong token fingerprints, expiry parsing, cross-instance replay, concurrent same-token submissions, and restart persistence. It does not publish a network-partition fault-injection or Jepsen result for the complete HTTP path. Validate client timeout, proxy failure, process termination, and FoundationDB recovery behavior in the environment you intend to operate.