Guides
Go to production
Everything here is a step you can verify. Do them in order — several of them fail loudly only once real traffic arrives.
Order of operations
- Generate and set your secrets
- Point storage at S3 if you will run more than one node
- Build the image and confirm it boots locally
- Mount volumes for the database and uploads
- Wire the probes, with a grace period longer than the drain
- Send real traffic, then watch the metrics
Secrets
openssl rand -hex 32 # → STORAGE_SECRETopenssl rand -hex 32 # → AUTH_SECRETNODE_ENV=productionPORT=4000 STORAGE_SECRET=… # required; startup fails without itAUTH_SECRET=… # do not leave the development default APP_URL=https://yourdomain.comRP_ID=yourdomain.com DATABASE_URL=/app/Database/app.dbWarning
STORAGE_SECRET is required in production — startup fails without it rather than starting with an ephemeral secret that invalidates every signed URL on each deploy.The checklist
NODE_ENV=production— otherwise mail prints to the console and route hot-reloading stays onSTORAGE_SECRETandAUTH_SECRETset to real random values- Storage driver switched to
s3if more than one node Database/andstorage/uploadson persistent volumes- Termination grace period longer than the 10-second drain
/healthzas liveness,/readyzas readinessdb.sql.exec("PRAGMA quick_check;")passes before traffic
Docker
docker build -t yatta . docker run -p 4000:4000 \ -e NODE_ENV=production \ -e STORAGE_SECRET="$(openssl rand -hex 32)" \ -e AUTH_SECRET="$(openssl rand -hex 32)" \ yattaThe image runs as the non-root bun user and ships a HEALTHCHECK pointed at /healthz.
Volumes
Without these, the database disappears when the container is replaced.
docker run -p 4000:4000 \ -v yatta-data:/app/Database \ -v yatta-uploads:/app/storage/uploads \ yattaProbes and draining
Shutdown stops accepting connections, waits up to ten seconds for in-flight tasks, then terminates workers. The orchestrator has to allow that long or the drain never finishes.
livenessProbe: httpGet: { path: /healthz, port: 4000 } initialDelaySeconds: 10 periodSeconds: 30 readinessProbe: httpGet: { path: /readyz, port: 4000 } initialDelaySeconds: 5 periodSeconds: 10 terminationGracePeriodSeconds: 20 # > the 10s drainNote
Liveness uses
/healthz, which never touches the runtime. If liveness checked readiness, a busy worker pool would make the orchestrator kill a healthy process and cause the very outage you were trying to avoid.Before you scale
- Multiple nodes. Local disk does not sync between machines. Switch the default storage disk to
s3first. - Job throughput. Each process runs its own workers. Raise
concurrencyonly as far as the database tolerates. - Supervision is per process. If a whole container dies, the load balancer must replace it — no in-process supervisor can observe a process that is gone. That is what the readiness probe is for.
Verifying a deploy
# readiness reflects the worker fleetcurl -s localhost:4000/readyz | jq '.status, .topology.totalWorkers' # job healthsqlite3 Database/jobs.db \ "SELECT state, COUNT(*) FROM _yatta_jobs GROUP BY state;" # a dead queue means handlers are failing after retriessqlite3 Database/jobs.db \ "SELECT name, COUNT(*) FROM _yatta_jobs WHERE state='dead' GROUP BY name;"For the full picture see Deployment and Reliability.