Skip to content

Deployment & operations

fdyno is a stateless HTTP data plane over one FoundationDB database. Running more fdyno processes removes a process-level single point of failure, but it does not replace FoundationDB replication, backup, admission control, identity management, or operational monitoring.

Deployment topology

A multi-process topology places same-version fdyno processes behind a private load balancer. Every process uses a cluster file for the same FoundationDB database. FoundationDB replication and failover depend on the cluster configuration, not on the number of fdyno processes.

flowchart LR
    C["AWS SDK clients"] --> E["TLS termination, authentication perimeter,<br/>and load balancing"]
    E --> D1["fdyno process"]
    E --> D2["fdyno process"]
    E --> D3["fdyno process"]
    D1 --> F[("FoundationDB cluster")]
    D2 --> F
    D3 --> F

The process stores table data, metadata, secondary indexes, change records, on-demand backups, transaction idempotency tokens, operation receipts, and worker coordination records in FoundationDB. Connections, in-flight requests, metrics, and some compatibility configuration are process-local. A same-version process can read committed state from the same FoundationDB database without copying fdyno data. Losing the FoundationDB cluster loses every copy managed by fdyno.

Use a dedicated FoundationDB database when you need an isolation boundary. fdyno uses fixed top-level directory names and has no configurable keyspace prefix. Two fdyno deployments pointed at the same database share tables, backups, worker leases, and transaction-token state. Authentication also has no tenant or account isolation.

Startup and packaging

The server entry point is cmd/dynodb. It initializes the FoundationDB Go binding at API version 730 and opens the default FoundationDB database. A missing or invalid cluster configuration can fail startup; a cluster that becomes unreachable after open is detected by readiness and data-plane failures.

The repository's verified local workflow is documented in the Quickstart. The checked-in scripts/fdb-local.sh starts one FoundationDB process with the memory storage engine by default. This engine serves reads from memory but logs writes to disk; the helper stores data under its local root. It configures single redundancy: no replicated copy, failover, or off-host recovery. This Homebrew-oriented helper is a single-host development configuration.

No Dockerfile or deployment manifest is tracked. An operator must provide the binary packaging, FoundationDB client library, cluster file, process supervision, resource limits, TLS proxy, secret injection, and rollout strategy.

Opt-in strict startup profile

DYNODB_PRODUCTION_MODE=1 or true validates configuration before opening FoundationDB, starting workers, or binding HTTP. Unset, 0, or false keeps development defaults, including permissive binds and built-in test credentials. Other mode values fail startup. This profile validates listed local settings; it cannot inspect the external TLS proxy, FoundationDB redundancy, or identity controls.

Required setting in opt-in mode Startup check
DYNODB_CREDENTIALS Distinct custom keyId:secret pairs; secret length at least 32 bytes; no built-in dev IDs/secrets, malformed or duplicate pairs, or control characters. It does not assess entropy or implement IAM.
DYNODB_LISTEN_ADDR Explicit private/loopback literal IP and valid TCP port; no wildcard, public IP, hostname, or omitted bind.
DYNODB_PPROF_ADDR Optional, but loopback literal IP and port if enabled.
FDB_CLUSTER_FILE Explicit absolute path to a readable regular file; no topology, TLS, or content verification.
DYNODB_TRUSTED_SINGLE_TENANT=1 and DYNODB_TLS_TERMINATED=1 Required operator attestations; the flags do not verify a trusted perimeter or TLS proxy.
DYNODB_CDC_TRIM_INTERVAL and DYNODB_TXN_TOKEN_GC_INTERVAL Both explicitly configured as positive Go durations.
DYNODB_TTL_SWEEP_INTERVAL Default-on one-minute worker accepted; off and 0 refused.

When set, DYNODB_OPERATION_TIMEOUT, the four HTTP timeouts, TTL/CDC/token/receipt GC and restore/backup worker intervals must be positive Go durations. DYNODB_CDC_RETENTION must be at least 24 hours. Configured DYNODB_TTL_SWEEP_MAX_SCAN, DYNODB_CDC_TRIM_MAX_SCAN, DYNODB_TXN_TOKEN_GC_MAX_SCAN, DYNODB_OPERATION_RECEIPT_GC_MAX_SCAN, and DYNODB_BACKUP_STAGING_GC_MAX_SCAN must be positive integers. DYNODB_STREAM_SHARDS, when set, must be in 1–100. See the exact validator in cmd/dynodb/production_config.go.

For an evaluation run, provision external TLS termination and a trusted network boundary, inject a managed custom secret, and validate the FoundationDB cluster file and backup path separately. Set the opt-in variables before starting every fdyno process. Check rejection paths first in an isolated environment, then verify /readyz and one custom-key signed request. The startup tests exercise negative configurations and an early subprocess exit; a manual isolated-FDB smoke observed readiness HTTP 200, custom-key acceptance, and dev-key rejection. None of those observations verifies TLS, IAM, replication, or recovery.

Binary compatibility

Use the same fdyno version for all processes sharing one FoundationDB database. Mixed-version reads and receipt/stream-record garbage collection are not supported. Change-stream records and operation receipts are stored in FoundationDB and shared by every connected process. See Reliable writes and Change streams.

On SIGINT or SIGTERM, fdyno stops its maintenance-worker context and gives the HTTP server five seconds to shut down. A forced stop or a request still committing at the deadline can leave the caller with an unknown outcome; use the retry rules in Reliable writes.

Before accepting a deployment profile, run bash scripts/fault-smoke.sh in a disposable environment. It checks two fdyno instances, one dropped response, one owned process kill, exact token replay, and local outage recovery. It does not qualify network partitions or replicated FoundationDB failover. See Failure testing.

Configuration reference

Configuration is read from environment variables at process startup. Keep FDB_CLUSTER_FILE and credentials consistent across replicas. Use one DYNODB_STREAM_SHARDS default so stream enablement is independent of request routing.

Listener, backend, and request lifetime

Variable Default Operational effect
FDB_CLUSTER_FILE FoundationDB client default Cluster file used by the FoundationDB client; opt-in profile requires a readable absolute regular file path.
DYNODB_LISTEN_ADDR :8000 HTTP address; default binds all interfaces in development mode. Opt-in profile requires an explicit private/loopback literal IP and port.
DYNODB_OPERATION_TIMEOUT 30s FoundationDB database transaction timeout and request transaction budget used by context-aware operation paths. Invalid or non-positive values fall back to 30s.
DYNODB_READ_HEADER_TIMEOUT 10s Maximum time to read HTTP headers.
DYNODB_READ_TIMEOUT 60s Maximum time to read the request.
DYNODB_WRITE_TIMEOUT 120s Maximum time to write the response.
DYNODB_IDLE_TIMEOUT 120s HTTP keep-alive idle timeout.
DYNODB_ACCESS_LOG unset Any non-empty value enables one log line per request.
DYNODB_PPROF_ADDR unset Starts an unauthenticated pprof listener on this separate address.

The server additionally rejects a request line or header set over 16 KiB and a raw or decompressed body over 16 MiB. These limits are compiled in, not configurable. The Go HTTP server's MaxHeaderBytes is 1 MiB, but fdyno's earlier 16 KiB check is the effective application limit.

Authentication and stream layout

Variable Default Operational effect
DYNODB_CREDENTIALS built-in development keys Comma-separated accessKeyId:secret pairs. Development mode replaces defaults when at least one pair parses; opt-in profile requires strict custom credentials.
DYNODB_STREAM_SHARDS 1 Default count for a newly enabled stream generation. Valid values are 1–100; invalid values use 1. The selected topology is stored in table metadata.

Credentials load once during package initialization; rotation requires process replacement. In development mode invalid pairs are skipped, and a non-empty variable with no valid pair leaves built-in keys in memory. The opt-in profile rejects malformed pairs, duplicates, control characters, dev keys/secrets, and secrets shorter than 32 bytes before listen. Validate the effective key set through the trusted perimeter; the environment-variable presence alone proves nothing. Secrets containing commas cannot be represented by this format.

DYNODB_STREAM_SHARDS does not reroute an active stream. Enabling a generation stores its count and routing identity in table metadata. Changing the count requires disabling and re-enabling the stream, which starts an empty generation and invalidates old iterators. See Change streams.

TTL, stream, and token cleanup workers

The bounded TTL sweeper starts by default at one-minute intervals; development mode accepts DYNODB_TTL_SWEEP_INTERVAL=off or 0, but opt-in production mode rejects either. CDC trimming and transaction-token GC remain opt-in in development and required in the production profile. Each worker's first pass occurs after one interval, not immediately at startup.

Variable Development default Purpose
DYNODB_TTL_SWEEP_INTERVAL 1m Sweep expired items; explicit off/0 disables only outside production mode.
DYNODB_TTL_SWEEP_MAX_SCAN 1000 Maximum items scanned per table per pass.
DYNODB_CDC_TRIM_INTERVAL unset, disabled When set, default fallback 1h; trims change-stream records. Production mode requires a positive interval.
DYNODB_CDC_RETENTION 24h when trimmer enabled Wall-clock retention; at least 24h in production mode when configured.
DYNODB_CDC_TRIM_MAX_SCAN 1000 Maximum records examined per table and shard per pass.
DYNODB_TXN_TOKEN_GC_INTERVAL unset, disabled When set, default fallback 1h; deletes expired transaction idempotency records. Production mode requires a positive interval.
DYNODB_TXN_TOKEN_GC_MAX_SCAN 1000 Maximum token records examined per pass.

Always-on recovery and cleanup

Variable Default Purpose
DYNODB_OPERATION_RECEIPT_GC_INTERVAL 10m Remove retry receipts after each record's retention deadline.
DYNODB_OPERATION_RECEIPT_GC_MAX_SCAN 1000 Maximum receipt records examined per pass.
DYNODB_RESTORE_RECOVERY_INTERVAL 15s Find and advance durable restore jobs.
DYNODB_BACKUP_STAGING_GC_INTERVAL 10m Remove abandoned STAGING backup generations.
DYNODB_BACKUP_STAGING_GC_MAX_SCAN 100 Maximum backup directories examined per pass.

Every process runs these workers. Receipt and staging cleanup are bounded and safe to repeat. Restore workers use durable leases and fences. GSI backfill recovery also runs once when each process starts.

Durations use Go duration syntax. Opt-in production mode rejects invalid and nonpositive configured timeouts, intervals, and scan limits before startup. Without it, a parsed zero/negative interval may panic a ticker, and an invalid retention value may trim current stream records.

Workers use cooperative FoundationDB leases with a lifetime of three intervals, with a 30-second minimum. This normally elects one process, but the lease has no fencing token and uses process wall clocks. Clock skew or a pass that outlives its lease can produce overlapping work. The operations are designed to be idempotent: TTL deletion rechecks expiry transactionally, and repeated trim or token deletion is harmless. Keep clocks synchronized anyway.

The TTL cursor is persisted in FoundationDB. CDC trimming always resumes from the oldest record. Token-GC scan position is process-local and restarts from the beginning after replacement. Worker transactions use FoundationDB batch priority; this makes foreground work preferred, but workers still consume I/O, CPU, and transaction capacity.

Incomplete global secondary index (GSI) backfills are different from these workers. Every process starts an unleased recovery scan in the background. Concurrent startup can repeat backfill work; page transactions and index writes are idempotent, but a rolling restart can add load.

Health, readiness, and rollout behavior

All probe endpoints are unauthenticated and served on the data listener.

Endpoint Meaning FoundationDB access
GET /, /healthz, /livez The HTTP process is running. None
GET /readyz, /health/ready A FoundationDB read-version request completed. Yes, with a two-second context
GET /metrics Process-local JSON counters and averages. None
GET /metrics/prometheus Process-local Prometheus text with fixed-bucket request histograms and bounded counters. None

Use /livez only for process restart decisions and /readyz for load-balancer membership. Restarting a healthy process because FoundationDB is temporarily down adds churn without restoring the backend.

Readiness proves current cluster reachability; it does not validate redundancy, free storage, worker progress, stream-shard agreement, credentials, table invariants, or backup recoverability. fdyno also begins listening while the startup GSI recovery goroutine may still be running. Check GSI status before using a newly recovered index.

These read-only diagnostics match routes and tooling used by the repository:

curl -i http://127.0.0.1:8000/livez
curl -i http://127.0.0.1:8000/readyz
curl -sS http://127.0.0.1:8000/metrics
curl -sS http://127.0.0.1:8000/metrics/prometheus
FDB_CLUSTER_FILE=/path/to/fdb.cluster fdbcli --exec "status minimal"

A healthy local response is HTTP 200; readiness returns 503 when its FoundationDB round trip fails. fdbcli is the authoritative next check for cluster availability.

Security boundary

fdyno verifies AWS Signature Version 4 (SigV4) on POST requests. It does not implement IAM, account isolation, session-token validation, resource-policy enforcement, credential expiry, or per-table authorization. Stored resource policy documents are metadata only. Every accepted access key can perform every exposed action against every table.

Build the deployment boundary accordingly:

  • Terminate TLS before fdyno; the server itself serves plain HTTP.
  • Bind fdyno to loopback or a private interface and restrict both client and FoundationDB network paths.
  • Inject non-development credentials through a secret mechanism and restart to rotate them. Environment variables may be visible through process and deployment inspection tools.
  • Restrict unauthenticated health, JSON /metrics, and Prometheus /metrics/prometheus at the proxy. They share the data listener and cannot be bound separately.
  • Keep pprof disabled or bind it to a private loopback address. It has no authentication.
  • Apply browser origin policy at the proxy. Requests with an Origin header receive Access-Control-Allow-Origin: *; SigV4 is still required, but CORS is not an authorization control.
  • Configure at-rest encryption, replication, and transport security for FoundationDB itself. fdyno adds no encryption layer to stored values.

Without the strict startup profile, a non-loopback bind with built-in credentials logs a warning but does not stop startup. The strict profile rejects the listed inputs; it cannot verify external network or cryptographic controls.

Metrics and capacity

GET /metrics returns process-local JSON; unauthenticated GET /metrics/prometheus on the data listener exports per-action fixed-bucket request-duration histograms with count and sum. Buckets run from 1 ms through 10 s, plus +Inf. Counters include in-flight requests, transaction attempts, conflicts (FDB 1020), unknown results (FDB 1021), receipt/token replay and mismatch, and bounded worker lease/tick/item and reported-failure metrics. Unknown signed actions use one Unknown label; item, table, and token values are not metric labels. Six worker families report surfaced error attempts.

errors_total counts each final captured HTTP response with status at least 400 once. It is not a unique failed-operation count and does not change mutation retry semantics.

Interpret these process-local signals conservatively:

  • Counts reset on restart and do not aggregate across replicas. Request totals count dispatched actions, while in-flight includes health and metric scrapes.
  • errors_total counts each final captured HTTP response with status 400 or greater once, including shaped ServiceError and direct validation responses. It is a response count, not a count of unique failed client operations. Replay hits count transaction attempts, not unique operations. FDB 1021 unknown results and terminal timeout ambiguity are distinct counters, not two estimates of the same events.
  • Histograms use fixed client-facing handler buckets, not FoundationDB latency breakdowns, queue depth, or durable traces. They are not an SLO.
  • Worker items_processed counts completed work, not scanned candidates. reported_failures_total covers surfaced errors in six worker families, including corrupt receipt/token rows retained for investigation. It counts reported attempts, so retries, persistent bad rows, shutdown, or table-delete races may increment it. Zero does not prove worker health. A false lease gauge can mean another holder or backend trouble; the last lease-held tick age does not prove a successful pass. There is no durable-job backlog, oldest stream record/backup age, or DR-success metric.

DYNODB_ACCESS_LOG adds remote address, method, DynamoDB action, status, and total duration. It does not log item bodies or authorization headers. Centralize logs and metrics externally, and monitor FoundationDB separately.

Plan for these capacity risks:

  • The HTTP server has no application-level concurrency limit or rate limiter. Go creates work per connection/request, so overload reaches process memory and FoundationDB rather than a DynamoDB-style throttle response. Provisioned-capacity fields are compatibility metadata; fdyno does not enforce request units or emit DynamoDB provisioning throttles.
  • Authentication buffers each request body in memory; the maximum is 16 MiB per concurrent request.
  • Base items, secondary-index entries, and stream images amplify each write inside one FoundationDB transaction. FoundationDB's 10 MB transaction limit, five-second transaction lifetime, and single-value limit remain hard boundaries; fdyno's longer operation timeout does not remove them.
  • Backup creation keeps one decoded and one encoded page in memory. Large restores stage bounded durable batches and resume from FoundationDB progress after process loss. The pinned backup read version can still age out before scanning completes.
  • Disabled TTL, CDC, or token cleanup permits durable keyspaces to grow without bound. Enabled cleanup can still fall behind when its bounded work per interval is lower than the arrival rate.
  • More fdyno replicas increase front-end concurrency; they do not create additional FoundationDB storage or commit capacity.

The performance page describes current-revision measurement coverage. Load-test the exact software, requests, and FoundationDB deployment before setting admission limits.

External components and exclusions

Area fdyno provides Supply or verify separately
TLS and identity SigV4 against configured static credentials; opt-in strict startup checks. TLS termination, trusted network boundary, credential rotation, and per-identity authorization if required.
FoundationDB transport and redundancy Cluster-file connection and readiness round trip. FoundationDB TLS and peer identities, replication, coordinator quorum, and failure-domain placement.
Recovery Same-cluster table snapshots and durable restore jobs. Independent off-host backup and restore, integrity verification, and recovery timing.
Monitoring Process-local JSON/Prometheus metrics and optional access logs. Centralized dashboards, alerting, durable backlog signals, and workload-specific SLOs.
Packaging and rollout Server binary and same-version process behavior. Process supervision, deployment packaging, resource limits, and coordinated replacement.

Use FoundationDB's fault-tolerance, configuration, administration, TLS, and native backup/restore documentation to design and verify your own cluster. FoundationDB requires a reachable majority of configured coordinators to operate; its redundancy mode and failure-domain placement determine what failures can be survived. A startup acknowledgement or ready probe does not check these conditions.

Follow the instructions for the FoundationDB version you actually deploy; these source links do not constitute a tested fdyno node layout.

Incident checks

Use this sequence before restarting processes or retrying writes indiscriminately.

  1. Separate process failure from backend failure. Compare /livez and /readyz. If liveness is healthy and readiness is not, inspect FoundationDB status, cluster-file distribution, network reachability, and FoundationDB logs.
  2. Check scope. If one replica is unready, compare its cluster file, FoundationDB client library, environment, clock, and network policy with healthy replicas. If all are unready, treat FoundationDB as the primary incident.
  3. Preserve ambiguous writes. Search logs for operation timeouts and the text transaction outcome may be unknown. Do not blindly repeat numeric increments or other read-modify-write operations. Reconcile state or retry a tokened transaction with the same token.
  4. Check saturation. Compare action latency and error deltas across process restarts, then inspect FoundationDB queueing, disk space, fault tolerance, and transaction conflict signals. fdyno emits no synthetic throttling error to make overload visible.
  5. Check maintenance progress. Confirm exactly one expected lease holder under normal conditions and that run/item counters advance. No progress may mean no eligible records, a lease on another replica, or a swallowed worker error.
  6. Check stream configuration. Describe the active generation and confirm that consumers poll every persisted shard. Estimate whether trim capacity exceeds record arrival.
  7. Check recovery objects. A restore target remains fenced in CREATING until its durable job completes. Recovery runs at startup and on its interval. Inspect DescribeTable output and server logs if progress stops.
  8. Prove recovery periodically. Restore backups into a new table or isolated FoundationDB environment and validate application invariants. Backup existence is not evidence that restore has been exercised.

For write ambiguity, continue with Reliable writes. For storage ownership and disaster recovery, read Durability & recovery and Backup & restore.