Clustering
Running several application nodes for capacity and to survive a failed host.
One application node serves a small team comfortably. Run several when you need more capacity, or when the loss of one host must not take the service down. The nodes form an Erlang cluster: each keeps serving its own users, and they share only what they must.
This page covers the application nodes. Clustering the services they depend on, PostgreSQL, the object store and the mail server, is a matter for those systems' own documentation; the application simply needs one reachable address for each.
What the cluster shares
- Live updates. When one node commits a change, it tells the others, so pages open on any node update.
- Who has a ticket open. Each node keeps that for its own pages and tells the others, so the ticket page's "Also here" list covers every node. It is held in memory only; a node that stops takes its entries with it.
- Background jobs. The job queue lives in PostgreSQL, and every node runs jobs from it. A job on a node that stops is picked up by another.
- Nothing else. Sessions are signed cookies, page state is rebuilt from the database, and there are no node-local caches that matter. A browser may reconnect to any node.
Because nodes are stateless, adding one is starting another copy with the same configuration, and losing one costs its users a reconnect of under a second.
Requirements
- Every node runs the same image version with the same
SECRET_KEY_BASE,CLOAK_KEYand other environment. A node with a differentSECRET_KEY_BASEcannot read the others' sessions. - A shared
RELEASE_COOKIE: a random string that Erlang uses to authenticate nodes to each other. Nodes with different cookies ignore each other. DNS_CLUSTER_QUERY: a DNS name that resolves to the address of every node. Each node resolves it regularly, connects to the addresses it finds, and so picks up new nodes and forgets stopped ones. Any orchestrator provides such a name: a Kubernetes headless service, an ECS service-discovery name, or Compose's service name.- A private network between the nodes. Erlang distribution uses port 4369 and a dynamic port range, is authenticated only by the cookie, and must not be reachable from the internet.
- A load balancer in front that spreads requests across nodes, passes WebSocket upgrades, and stops sending to a node whose
/healthfails. Sticky sessions are not needed.
The release names each node after its own IP address when DNS_CLUSTER_QUERY is set, so the addresses in DNS and the node names agree without further configuration.
With Docker Compose
compose.cluster.yaml is an override for the compose file from Installation that removes the application's port, adds the clustering variables, and puts a Traefik load balancer in front that discovers the replicas from Docker. Download it beside compose.yaml, then:
curl -O https://albaticket.com/files/compose.cluster.yaml
echo "RELEASE_COOKIE=$(openssl rand -base64 32)" >> .env
docker compose -f compose.yaml -f compose.cluster.yaml up -d --scale app=3
The service is then on the same port as before, served by the balancer. The migrate service runs the migrations once before any replica starts, so the replicas never run them themselves. The balancer only routes to a node once Docker reports it healthy, which takes a few seconds after the node starts and its health check passes; until then it answers 404.
Check that the nodes found each other:
docker compose exec --index 1 app /app/bin/alba rpc 'IO.inspect(Node.list())'
The list shows the other nodes. Scale up or down with --scale app=N; a node that is stopped drains its connections, and its browsers reconnect to the remaining ones.
On a single host this protects against a crashed process, not a failed machine. For that, run the nodes on different hosts with an orchestrator, keeping the requirements above, and give the database, storage and mail their own high-availability arrangements as their documentation describes. On Kubernetes, the headless Service and the two variables are in Installation.
Sizing
Start with two nodes of 2 CPU cores and 2 GB each for resilience, and add nodes when CPU stays busy. The database is usually the first limit: each node holds a pool of POOL_SIZE connections, so the total must stay under PostgreSQL's max_connections; put PgBouncer in session mode in front when it does not.
Rolling upgrades
Run the new version's migrations first (docker compose up migrate, or /app/bin/migrate from one container of the new image on other platforms); they are additive, so old nodes keep working. Then start nodes on the new version, wait for them to become healthy and join, and stop the old ones.