Guides

Go to production

Everything here is a step you can verify. Do them in order — several of them fail loudly only once real traffic arrives.

Order of operations

  • Generate and set your secrets
  • Point storage at S3 if you will run more than one node
  • Build the image and confirm it boots locally
  • Mount volumes for the database and uploads
  • Wire the probes, with a grace period longer than the drain
  • Send real traffic, then watch the metrics

Secrets

generatebash
openssl rand -hex 32   # → STORAGE_SECRETopenssl rand -hex 32   # → AUTH_SECRET
environmentbash
NODE_ENV=productionPORT=4000 STORAGE_SECRET=…    # required; startup fails without itAUTH_SECRET=…       # do not leave the development default APP_URL=https://yourdomain.comRP_ID=yourdomain.com DATABASE_URL=/app/Database/app.db
Warning
STORAGE_SECRET is required in production — startup fails without it rather than starting with an ephemeral secret that invalidates every signed URL on each deploy.

The checklist

  • NODE_ENV=production — otherwise mail prints to the console and route hot-reloading stays on
  • STORAGE_SECRET and AUTH_SECRET set to real random values
  • Storage driver switched to s3 if more than one node
  • Database/ and storage/uploads on persistent volumes
  • Termination grace period longer than the 10-second drain
  • /healthz as liveness, /readyz as readiness
  • db.sql.exec("PRAGMA quick_check;") passes before traffic

Docker

terminalbash
docker build -t yatta . docker run -p 4000:4000 \  -e NODE_ENV=production \  -e STORAGE_SECRET="$(openssl rand -hex 32)" \  -e AUTH_SECRET="$(openssl rand -hex 32)" \  yatta

The image runs as the non-root bun user and ships a HEALTHCHECK pointed at /healthz.

Volumes

Without these, the database disappears when the container is replaced.

terminalbash
docker run -p 4000:4000 \  -v yatta-data:/app/Database \  -v yatta-uploads:/app/storage/uploads \  yatta

Probes and draining

Shutdown stops accepting connections, waits up to ten seconds for in-flight tasks, then terminates workers. The orchestrator has to allow that long or the drain never finishes.

kubernetesyaml
livenessProbe:  httpGet: { path: /healthz, port: 4000 }  initialDelaySeconds: 10  periodSeconds: 30 readinessProbe:  httpGet: { path: /readyz, port: 4000 }  initialDelaySeconds: 5  periodSeconds: 10 terminationGracePeriodSeconds: 20   # > the 10s drain
Note
Liveness uses /healthz, which never touches the runtime. If liveness checked readiness, a busy worker pool would make the orchestrator kill a healthy process and cause the very outage you were trying to avoid.

Before you scale

  • Multiple nodes. Local disk does not sync between machines. Switch the default storage disk to s3 first.
  • Job throughput. Each process runs its own workers. Raise concurrency only as far as the database tolerates.
  • Supervision is per process. If a whole container dies, the load balancer must replace it — no in-process supervisor can observe a process that is gone. That is what the readiness probe is for.

Verifying a deploy

after shippingbash
# readiness reflects the worker fleetcurl -s localhost:4000/readyz | jq '.status, .topology.totalWorkers' # job healthsqlite3 Database/jobs.db \  "SELECT state, COUNT(*) FROM _yatta_jobs GROUP BY state;" # a dead queue means handlers are failing after retriessqlite3 Database/jobs.db \  "SELECT name, COUNT(*) FROM _yatta_jobs WHERE state='dead' GROUP BY name;"

For the full picture see Deployment and Reliability.