Deploying on AWS
Running an installation on ECS Fargate with RDS, S3 and SES.
This page deploys Alba Ticket on AWS from the published image, using the managed services it fits: a PostgreSQL database on RDS, an S3 bucket for attachments, SES for email, and ECS on Fargate for the application and for Typesense, behind an Application Load Balancer with a certificate from ACM. It assumes a VPC with public subnets for the load balancer and private subnets for everything else, and that you are comfortable in the console or with the aws command line; every step below can be done either way, and what matters is the values that end up in the task definition.
The result is one application task, one Typesense task with its data on EFS, and nothing else to run yourself. Clustering later is a matter of raising the task count.
Before you start
- A domain for the application,
tickets.example.combelow, in a Route 53 zone or anywhere you can add a record, and a certificate for it in ACM in the same region as the load balancer. - The two secrets, generated on your machine and kept somewhere safe; losing
CLOAK_KEYmakes stored credentials unreadable:
openssl rand -base64 64 | tr -d '\n' # SECRET_KEY_BASE
openssl rand -base64 32 # CLOAK_KEY
- Two security groups:
alba-lbfor the load balancer, open to the internet on 443 and 80, andalba-appfor the tasks, allowing 4000 fromalba-lband all TCP from itself (the application talks to Typesense, and later to its own replicas). The database's security group allows 5432 fromalba-app.
The published images at ghcr.io/garth/alba are public, so no registry login is needed. Pick the version to run from Published images; the examples below use 0.7.8, this release.
1. The database
Create an RDS for PostgreSQL instance, version 16, in the private subnets with the security group above and public access off. A db.t4g.small with 20 GB of storage suits a small team; RDS grows the storage when asked. Create a database called alba and note the master user's password, or create a user of its own for the application. The connection string is:
postgres://alba:<password>@<instance>.<id>.<region>.rds.amazonaws.com:5432/alba
Two things about TLS. RDS signs its certificates with Amazon's own RDS certificate authority, which is not in the public store the image trusts, so ?ssl=true on this URL, which verifies the certificate, fails against RDS. Recent RDS parameter groups also require TLS (rds.force_ssl is 1), which rejects the plain connection. Choose one:
- Plain inside the VPC. Create a parameter group with
rds.force_sslset to0and attach it to the instance. The connection is unencrypted, but it never leaves your private subnets. - Verified TLS. Add
?ssl=trueto the URL and start the container with a command that fetches the RDS certificate bundle and joins it to the system store, sinceSSL_CERT_FILEreplaces that store rather than adding to it:sh -c "curl -fsS -o /tmp/rds-ca.pem https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem && cat /etc/ssl/certs/ca-certificates.crt /tmp/rds-ca.pem > /tmp/ca.pem && SSL_CERT_FILE=/tmp/ca.pem exec /app/bin/server"(and the same with/app/bin/migratefor the migration task).
Turn on automated backups with a retention you are happy with; that is the database backup described in Backups.
2. The bucket
Create an S3 bucket in the same region, with Block Public Access left on: browsers reach objects through short-lived signed addresses, never public ones. Under the bucket's permissions, set this CORS configuration so browsers may upload to it and download from it:
[
{
"AllowedOrigins": ["https://tickets.example.com"],
"AllowedMethods": ["GET", "PUT", "HEAD"],
"AllowedHeaders": ["*"],
"ExposeHeaders": ["ETag"],
"MaxAgeSeconds": 3000
}
]
Create an IAM user for the application with programmatic access and only this policy, and note its access key:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:ListBucket", "s3:GetBucketLocation"],
"Resource": "arn:aws:s3:::alba-attachments"
},
{
"Effect": "Allow",
"Action": ["s3:PutObject", "s3:GetObject", "s3:DeleteObject"],
"Resource": "arn:aws:s3:::alba-attachments/*"
}
]
}
The storage variables are then STORAGE_ENDPOINT=https://s3.<region>.amazonaws.com, STORAGE_REGION=<region>, STORAGE_BUCKET=alba-attachments and the key pair; leave STORAGE_FORCE_PATH_STYLE and STORAGE_PUBLIC_BASE_URL unset, since S3 addresses the bucket by host name and browsers use the same address the application does. Turn on versioning if you want deleted attachments recoverable; see Backups.
3. Email
In SES, verify the domain you will send from and, while the account is in the SES sandbox, every address you will send to; request production access before inviting real people. Create SMTP credentials under SMTP settings, which makes an IAM user and gives you a user name and password for SMTP alone. The mail variables are SMTP_HOST=email-smtp.<region>.amazonaws.com, SMTP_PORT=587, SMTP_TLS=always, the two credentials, and MAIL_FROM on the verified domain.
4. Secrets
Put the values that must not sit in a task definition into Secrets Manager (or SSM Parameter Store, which is cheaper), one secret each: SECRET_KEY_BASE, CLOAK_KEY, DATABASE_URL, STORAGE_SECRET_ACCESS_KEY, SMTP_PASSWORD and a random TYPESENSE_API_KEY (openssl rand -base64 32). The task execution role needs secretsmanager:GetSecretValue (or ssm:GetParameters) on them; ECS reads them when a task starts and hands them to the container as environment variables, so the application needs no AWS permissions of its own.
5. The cluster, the file system and search
Create an ECS cluster for Fargate, and an EFS file system in the private subnets with a security group allowing NFS (2049) from alba-app, for Typesense's index alone: the application itself keeps nothing on disk. Give it one access point, typesense, root /typesense, owned by uid and gid 1000 with mode 755 so the container can write.
Typesense holds derived data only, rebuilt from the database on demand, so one task is enough. Register a task definition for it (0.5 vCPU, 1 GB, awsvpc) with the typesense/typesense:30.2 image, port 8108, the EFS volume at /data through the typesense access point, and this environment, which is how Typesense takes its settings:
| Name | Value |
|---|---|
TYPESENSE_DATA_DIR |
/data |
TYPESENSE_API_KEY |
from Secrets Manager |
Create a service from it with one task, in the private subnets with alba-app, and turn on Service Connect (or Cloud Map service discovery) in a namespace called alba, so the application reaches it as typesense.alba on port 8108. The application's TYPESENSE_URL is then http://typesense.alba:8108.
6. The application
Register a task definition. This one has everything the application needs; replace the account, region and secret ARNs, and the version if you pin one:
{
"family": "alba",
"networkMode": "awsvpc",
"requiresCompatibilities": ["FARGATE"],
"cpu": "1024",
"memory": "2048",
"executionRoleArn": "arn:aws:iam::123456789012:role/albaTaskExecutionRole",
"containerDefinitions": [
{
"name": "app",
"image": "ghcr.io/garth/alba:0.7.8",
"essential": true,
"portMappings": [{ "containerPort": 4000, "protocol": "tcp" }],
"environment": [
{ "name": "PHX_HOST", "value": "tickets.example.com" },
{ "name": "PHX_SCHEME", "value": "https" },
{ "name": "PORT", "value": "4000" },
{ "name": "TYPESENSE_URL", "value": "http://typesense.alba:8108" },
{ "name": "STORAGE_ENDPOINT", "value": "https://s3.eu-west-1.amazonaws.com" },
{ "name": "STORAGE_REGION", "value": "eu-west-1" },
{ "name": "STORAGE_BUCKET", "value": "alba-attachments" },
{ "name": "STORAGE_ACCESS_KEY_ID", "value": "AKIA..." },
{ "name": "SMTP_HOST", "value": "email-smtp.eu-west-1.amazonaws.com" },
{ "name": "SMTP_PORT", "value": "587" },
{ "name": "SMTP_USERNAME", "value": "AKIA..." },
{ "name": "SMTP_TLS", "value": "always" },
{ "name": "MAIL_FROM", "value": "tickets@example.com" }
],
"secrets": [
{ "name": "SECRET_KEY_BASE", "valueFrom": "arn:aws:secretsmanager:eu-west-1:123456789012:secret:alba/secret-key-base" },
{ "name": "CLOAK_KEY", "valueFrom": "arn:aws:secretsmanager:eu-west-1:123456789012:secret:alba/cloak-key" },
{ "name": "DATABASE_URL", "valueFrom": "arn:aws:secretsmanager:eu-west-1:123456789012:secret:alba/database-url" },
{ "name": "STORAGE_SECRET_ACCESS_KEY", "valueFrom": "arn:aws:secretsmanager:eu-west-1:123456789012:secret:alba/storage-secret" },
{ "name": "SMTP_PASSWORD", "valueFrom": "arn:aws:secretsmanager:eu-west-1:123456789012:secret:alba/smtp-password" },
{ "name": "TYPESENSE_API_KEY", "valueFrom": "arn:aws:secretsmanager:eu-west-1:123456789012:secret:alba/typesense-api-key" }
],
"healthCheck": {
"command": ["CMD-SHELL", "curl -fsS http://localhost:4000/health || exit 1"],
"interval": 10,
"timeout": 5,
"startPeriod": 30,
"retries": 3
},
"logConfiguration": {
"logDriver": "awslogs",
"options": {
"awslogs-group": "/ecs/alba",
"awslogs-region": "eu-west-1",
"awslogs-stream-prefix": "app",
"awslogs-create-group": "true"
}
}
}
]
}
ECS does not read the health check built into the image, which is why the task definition repeats it. The application's TYPESENSE_API_KEY is the same secret Typesense was given.
Migrations
The image does not migrate on start, so that several copies can start at once. Run the migrations as a one-off task from the same task definition, with the command replaced, and wait for it to stop:
aws ecs run-task --cluster alba --launch-type FARGATE --task-definition alba \
--network-configuration "awsvpcConfiguration={subnets=[subnet-a,subnet-b],securityGroups=[sg-app],assignPublicIp=DISABLED}" \
--overrides '{"containerOverrides":[{"name":"app","command":["/app/bin/migrate"]}]}'
aws ecs describe-tasks shows its exit code, and the /ecs/alba log group what it did. A non-zero code is almost always the database URL or its security group.
The service
Create an Application Load Balancer in the public subnets with alba-lb, an HTTPS listener with the ACM certificate, and an HTTP listener that redirects to HTTPS. Its target group is IP type, protocol HTTP on port 4000, health check path /health, and a deregistration delay of 30 seconds so a stopping task drains quickly. Then create the ECS service from the alba task definition: one task, Fargate, the private subnets with alba-app, public IP off, the target group attached to the app container on 4000, and a health check grace period of 60 seconds so the balancer does not kill a task that is still starting. Point tickets.example.com at the balancer with an alias record.
The balancer passes WebSocket connections as they are, which live updates depend on, and its idle timeout of 60 seconds is longer than the page's heartbeat, so nothing needs changing there. It can keep its default request size limits: attachments never pass through the application.
7. First visit
Open https://tickets.example.com. An installation with no users shows the setup page: it creates the first administrator, with a password so that login works before email delivery is confirmed, and names the installation. The bucket is registered as attachment storage on first start from the STORAGE_* variables; check under Administration, Storage that the location verifies, which writes, reads and deletes one small object. Send yourself an invitation from Invitations to confirm mail delivery, and run Rebuild index on the Search page under Administration if search returns nothing.
Upgrading
- Check that the automated RDS backups are in place, or take a snapshot.
- Register a new revision of the
albatask definition with the new image tag. - Run the migration task from the new revision, as above, and wait for it to stop with code 0.
- Update the service to the new revision. ECS starts a task on the new version, waits for the balancer to see it healthy, then stops the old one; people's browsers reconnect by themselves.
The notes in Upgrades about specific versions apply here too.
Several nodes
Raise the service's desired count and the tasks form a cluster, as Clustering describes, once they can find each other:
- Turn on Service Connect or service discovery for the
albaservice too, so the nameapp.albaresolves to every task, and setDNS_CLUSTER_QUERY=app.albain the task definition. - Add a
RELEASE_COOKIEsecret (openssl rand -base64 32), the same for every task. - Keep the security group rule that allows all TCP within
alba-app: Erlang distribution uses port 4369 and a dynamic range between the tasks, and must be reachable from nowhere else.
Run the migrations before rolling a new version, as above; ECS's rolling deployment then replaces tasks one at a time.
Importing from Jira
Upload the backup under Administration, Jira import. It goes straight from the browser to the bucket, under an imports/ key of its own, and the import reads it from there; nothing has to reach the task's filesystem. The bucket's CORS rules must allow the application's address, which the storage bootstrap sets when it makes the bucket and which step 3 covers for one made by hand. An upload is removed from the bucket after a day.
Operating notes
- Logs are in the
/ecs/albalog group in CloudWatch, one stream per task. Look there forPostgrexconnection errors, job failures and mail delivery errors. - Backups. RDS automated backups for the database, versioning or replication on the bucket for attachments, and the two secrets kept outside AWS as well; see Backups. Typesense's data on EFS is derived and needs no backup.
- Sizing. 1 vCPU and 2 GB for the application task and half that for Typesense suit a small team; see Requirements and, for more than one task, Clustering.
- Cost. The database and the load balancer are the fixed costs; Fargate is billed for the two tasks by the hour, EFS and S3 by what is stored.
Troubleshooting
- The migration task stops with a non-zero code. Its log stream says why.
no pg_hba.conf entry ... SSL offmeans the parameter group still requires TLS; a connection timeout means the database's security group does not allow 5432 fromalba-app, or the task is not in the private subnets that reach it. - The service never becomes healthy, or the balancer answers 502 or 503. The target group's health check must be
/healthon port 4000, and the grace period long enough for the task to start; the task's own log shows whether the application started. A task that starts and stops at once with no log is usually a secret the execution role may not read. - The storage location does not verify. The region in
STORAGE_ENDPOINTandSTORAGE_REGIONmust be the bucket's, and the IAM policy must name the right bucket. The error under Administration, Storage quotes S3's answer. - Uploads fail in the browser but the location verifies. The bucket's CORS configuration does not allow the application's origin, or lists
httpwhere the site ishttps. The browser console shows the blocked request. - No email arrives. SES is still in the sandbox,
MAIL_FROMis not on a verified domain, or the SMTP credentials are ordinary IAM keys rather than the SMTP ones. The application log shows the server's reply. - Pages load but nothing updates live. Something between the browser and the task is not passing WebSockets; a CloudFront distribution in front of the balancer, for example, needs its cache policy to forward the upgrade. The browser console shows the
/live/websocketconnection state.
See Troubleshooting for problems that are not specific to AWS.