Deploying the Web App to AWS (Infrastructure & Procedure)
Source:vignettes/dev-deployment.Rmd
dev-deployment.RmdThis article documents how the live EJAM Shiny web
app is hosted on AWS ECS Fargate, and the
routine procedure for deploying updates. It is a companion to Deploying the Web App, which covers the
higher-level choices (Docker/AWS vs. Posit Connect) and the
app-configuration options (isPublic, logos, titles); this
article is the concrete, current AWS procedure.
Not to be confused with the EJAM API. This is the Shiny app (the interactive website), hosted on AWS ECS Fargate. The separate EJAM REST API is hosted on Google Cloud Run and deployed from the EJAM-API repository.
Where the operational files live
The files that actually build and deploy the app are kept on the two
deploy branches (dev-deploy and
prod-deploy) of the EJAM repository, not on
main/development. This keeps the
branch-triggered deploy workflows and AWS/Terraform config out of the
package source (so they cannot misfire on a main push and
do not clutter the R package). This article is the human-readable guide
to those files; edit the files themselves on the deploy branches.
| File (on the deploy branches) | Purpose |
|---|---|
ejam-infra/main.tf |
Terraform — the full AWS infrastructure definition (VPC, ALB, ECR, ECS, IAM, logs) |
ejam-infra/prod.tfvars,
ejam-infra/dev.tfvars
|
Per-environment Terraform variables (task size, retention, etc.) |
Dockerfile |
Container image build for the Shiny app. Chooses the EJAM version
(ARG EJAM_VERSION) and the ejamdata release
(ARG EJAMDATA_VERSION), and its CMD starts the
app with EJAM::ejamapp(isPublic = ...)
|
.dockerignore |
Files excluded from the Docker build context |
.github/workflows/deploy.yaml |
GitHub Actions — build + deploy to prod (on push to
prod-deploy, or run manually) |
.github/workflows/deploy-dev.yaml |
GitHub Actions — build + deploy to dev (on push to
dev-deploy, or run manually with an optional
ejam_version input) |
.github/workflows/diag-ecs-dev.yaml |
GitHub Actions — read-only diagnostics for the dev ECS service (run manually) |
Apart from these, the deploy branches hold only a short
README.md, the licenses, and git config files. They do not
hold the EJAM package source: the Dockerfile has no
COPY step and installs EJAM from GitHub at the ref named by
ARG EJAM_VERSION. (Until September 2026 each branch also
carried an old, unused copy of the package source; it was removed in #633
and #634.)
Architecture at a glance
| Layer | Technology |
|---|---|
| App | R Shiny (rocker/rstudio base image) |
| PDF generation | Google Chrome stable + chromote/pagedown |
| Containerization | Docker |
| Container registry | AWS ECR (shared repo ejam; the prod stack owns it, dev
shares it) |
| Hosting | AWS ECS Fargate (region us-east-1) |
| Load balancing | AWS Application Load Balancer (ALB) |
| HTTPS / TLS | AWS ACM (DNS-validated) |
| Infrastructure-as-code | Terraform (S3 state backend, separate state per environment) |
| CI/CD | GitHub Actions |
| DNS | Squarespace (CNAME → ALB) |
Each deploy branch triggers its own GitHub Actions workflow, which builds the
image, pushes it to ECR, and updates that environment's ECS Fargate service:
dev-deploy ──push──► deploy-dev.yaml ──► ECR (ejam) ──► ECS Fargate: ejam-dev ──► dev ALB
prod-deploy ──push──► deploy.yaml ──► ECR (ejam) ──► ECS Fargate: ejam-prod ──► prod ALB
image tags: │
dev = dev-<sha>, prod = <sha> ▼
Squarespace CNAME ──► ejam.publicenvirodata.org
Each image installs EJAM from GitHub at the ref pinned by ARG EJAM_VERSION in
that branch's Dockerfile -- see "Branching and deploy model" below.
URLs
| Environment | URL |
|---|---|
| Production | https://ejam.publicenvirodata.org (Squarespace CNAME → prod ALB, ACM TLS) |
| Dev | the dev ALB DNS name printed by terraform output (an
…elb.amazonaws.com address) |
Task sizing (set in the .tfvars
files)
| vCPU / memory | Task count | Logs | |
|---|---|---|---|
| Prod | 2 vCPU / 7 GB | 2 | CloudWatch /ecs/ejam-prod, 30-day retention |
| Dev | 1 vCPU / 6 GB | 1 | CloudWatch /ecs/ejam-dev, 7-day retention |
The container serves the app on port 2000 and a
health-check endpoint on port 2001
(app_port / health_check_port in
ejam-infra/main.tf, which default to 2000/2001; the
Dockerfile EXPOSEs 2000 2001, and
its CMD starts a small health server on 2001 and then
EJAM::ejamapp(options = list(port = 2000))). The
CMD also sets the app mode: dev runs
isPublic = FALSE (the full app, so testing exercises every
feature) and prod runs isPublic = TRUE (the public
app).
Branching and deploy model
The deploy branches hold the build files, not the
code that gets deployed. Each branch’s Dockerfile names the
EJAM version to install, and the image installs that version from
GitHub:
dev-deploy/Dockerfile ARG EJAM_VERSION=<tag, branch, or SHA> ARG EJAMDATA_VERSION=vX.YYYY.Z
prod-deploy/Dockerfile ARG EJAM_VERSION=<release tag> ARG EJAMDATA_VERSION=vX.YYYY.Z
-
To change what is deployed, change those
ARGlines with a small pull request into the deploy branch. Don’t mergemainordevelopmentinto a deploy branch: that would bring the package source back onto the branch, and the image doesn’t use it. -
ARG EJAM_VERSIONcan be a release tag (e.g.v3.2022.3), a branch (e.g.development), or a commit SHA. Production should always use a release tag. Dev can use any ref, which lets you test a release candidate before it is tagged. -
ARG EJAMDATA_VERSIONnames theejamdatarelease; see Which EJAM version deploys below. -
Merging a pull request into
dev-deploydeploys dev; merging intoprod-deploydeploys prod, immediately (about 15–20 minutes). Each branch’s workflow targets its own ECS cluster and service. - The two
Dockerfiles are meant to differ only in the pins and inisPublic(devFALSE, prodTRUE). Make any other change (for example the font-cache line described below) on both branches. - Validating on dev and then setting the same tag on
prod deploys the same EJAM code. The prod image is still a fresh build,
though: it uses the current
rocker/rstudio:latestbase image and current CRAN packages, which may differ slightly from the dev image built earlier. - Both
dev-deployandprod-deployare protected: a PR is required (direct pushes are blocked for everyone, including admins); no approvals are required (a maintainer may merge their own PR).
Running the dev deploy for another EJAM ref. The dev
workflow can also be started by hand (GitHub → Actions → “Build &
Deploy to ECS Fargate (dev)” → Run workflow, or from the command line)
with an ejam_version input, which overrides
ARG EJAM_VERSION for that one build without editing the
branch:
gh workflow run "Build & Deploy to ECS Fargate (dev)" \
--repo Public-Environmental-Data-Partners/EJAM --ref dev-deploy \
-f ejam_version=developmentLeave ejam_version blank to build the version pinned in
the Dockerfile. The dispatch cannot override
EJAMDATA_VERSION: to test with a different data
release, change that pin in dev-deploy by pull request. The
prod workflow can also be run by hand, but it has no input; it rebuilds
the pinned version.
Deploys don’t queue. Neither deploy workflow limits concurrent runs, so a deploy started by a merge and one started by hand (or two started by hand) run at the same time, each updating the same ECS service. Either one can end up live, and it may not be the one you meant. Start one deploy, wait for it to finish, and then confirm the version shown in the app footer. In particular, don’t dispatch a dev deploy right after merging into
dev-deploy, since the merge already started one.
Routine deploy (GitHub Actions) — the normal path
- Release (or pick a candidate). Deploy a release tag. To test a release candidate on dev before tagging, use a branch or commit SHA instead.
-
Deploy to dev: open a PR into
dev-deploythat setsARG EJAM_VERSION(andARG EJAMDATA_VERSION, if the data release changed). On merge, GitHub Actions builds the image, pushes it to ECR, and updates the dev ECS service. For a quick test of an unreleased ref, you can dispatch the dev workflow withejam_versioninstead (see above). - Validate on dev at the dev URL: the footer version, one report (including the PDF), and readable chart labels. A human must green-light before promoting.
-
Deploy to prod: open a PR into
prod-deploythat sets the sameARG EJAM_VERSIONtag (andARG EJAMDATA_VERSION, if it changed). On merge, Actions deploys production. Then check the same things on the production app.
For the release steps around a deploy (tagging, the
ejamdata release, EJAM-API and EJSCREEN updates), see Releasing a New Version of EJAM.
PDF checks
PDF support is verified at three points because source-only testing does not exercise the image’s configuration:
-
During the image build: the package
configurescript performs a realpagedown::chrome_print()render whenever the build advertises Chrome throughCHROMOTE_CHROMEorPAGEDOWN_CHROME(or explicitly setsEJAM_VERIFY_PDF_BUILD=true). It verifies both the output size and%PDFmagic bytes. A Chrome executable lookup alone is not treated as success. -
In the Shiny browser suite: the full
latloncase requests both PDF and HTML report downloads. Missing Pandoc is an error for this test, not a silent pass. -
After deployment:
Verify deployed PDF generationruns after either successful ECS deploy workflow. It performs a one-point analysis against the deployed development or production app, downloads the report, and verifies the.pdffilename and%PDFbytes. The workflow can also be dispatched manually.
The post-deployment verification workflow lives on the default
branch, as required by GitHub’s workflow_run event, while
the operational build/deploy workflows stay on dev-deploy
and prod-deploy.
Required GitHub Actions secrets (repo → Settings → Secrets and variables → Actions):
| Secret | Purpose |
|---|---|
AWS_ACCESS_KEY_ID,
AWS_SECRET_ACCESS_KEY
|
deploy credentials for ECR/ECS |
No GitHub token secret is needed. Both workflows pass the workflow
run’s own short-lived token to the build
(--build-arg GITHUB_PAT=${{ github.token }}), which is
enough to read the public EJAM and ejamdata repositories.
(A stored personal access token, EJAMDATA_PAT, was used
until July 2026, when it expired and broke the builds; it is no longer
used.)
Which EJAM version deploys
-
EJAM: the
Dockerfileinstalls EJAM withremotes::install_github()atARG EJAM_VERSION(a tag, branch, or commit SHA; see Branching and deploy model). -
ejamdata: the 11 large.arrowfiles are downloaded separately, from theejamdataGitHub release named byARG EJAMDATA_VERSION. It must equal the installed EJAM’sejamdata_required_tag(in EJAM’sDESCRIPTION; for EJAM v3.2022.3 that isv3.2022.3). If they differ, EJAM finds that the data in the image is not the release it requires and downloads the whole set (about 1 GB) again every time a container starts. (An explicit emptyEJAMDATA_VERSIONfalls back to the “Latest”ejamdatarelease, which is not necessarily the one EJAM requires, so always set it.) - A new
ejamdatarelease, and so a newEJAMDATA_VERSION, is needed wheneverbgej, the FRS files, or the block geography files change. A code-only EJAM release keeps the existing data tag. See Releasing a New Version of EJAM for the full list of places that must name the same data tag. -
Font cache: the
DockerfileCMDmust start withsystem('fc-cache -r'), before the app starts (#621). On ECS Fargate the font cache built into the image did not match the fonts actually present, so R resolved “sans” to a decorative font fromtexlive-fonts-extraand chart text in reports came out as boxes (the same image looked fine under plain Docker). Rebuilding the cache at startup takes a few seconds. Keep this line on both deploy branches. (It was added todev-deployin #623; the matchingprod-deploychange is #624.)
First-time setup (manual / infrastructure)
Only needed when standing up the infrastructure or running Terraform locally; day-to-day app deploys use GitHub Actions (above).
Prerequisites (macOS shown; use your platform’s package manager):
# Homebrew, then:
brew install hashicorp/tap/terraform
brew install awscli
# plus Docker Desktop (docker.com), running before any Docker stepAWS credentials. Request credentials from the AWS account administrator, then:
Your IAM user needs the custom ejam-terraform-deploy
policy to provision infrastructure — request the policy JSON from the
deployment manager and attach it (AWS Console → IAM → Users → your user
→ Add permissions).
Provision / update infrastructure with Terraform.
Run from ejam-infra/. State lives in the S3 bucket
ejam-terraform-state-<ACCOUNT_ID>, with a
separate state key per environment:
cd ejam-infra
# Prod
terraform init -backend-config="key=prod/terraform.tfstate"
terraform plan -var-file=prod.tfvars -var="aws_account_id=<ACCOUNT_ID>"
terraform apply -var-file=prod.tfvars -var="aws_account_id=<ACCOUNT_ID>"
# Dev (separate state; -reconfigure switches the backend key)
terraform init -backend-config="key=dev/terraform.tfstate" -reconfigure
terraform apply -var-file=dev.tfvars -var="aws_account_id=<ACCOUNT_ID>"Get your account ID with
aws sts get-caller-identity --query Account --output text.
Manual Docker build & push (fallback)
Prefer GitHub Actions — the uncompressed image is large (~4 GB) and
local pushes are slow. If you must build/push by hand (from the deploy
branch, which has the Dockerfile):
# Authenticate Docker to ECR
aws ecr get-login-password --region us-east-1 \
| docker login --username AWS --password-stdin \
<ACCOUNT_ID>.dkr.ecr.us-east-1.amazonaws.com
# The ECR repo is created with image_tag_mutability = IMMUTABLE, so every push
# needs a UNIQUE tag -- re-pushing an existing tag (e.g. :latest) will fail.
# Match the workflows' commit-SHA convention:
TAG="manual-$(git rev-parse --short HEAD)"
docker build -t ejam:$TAG .
docker tag ejam:$TAG <ACCOUNT_ID>.dkr.ecr.us-east-1.amazonaws.com/ejam:$TAG
docker push <ACCOUNT_ID>.dkr.ecr.us-east-1.amazonaws.com/ejam:$TAGDon’t pass your own GitHub token as
--build-arg GITHUB_PAT=.... Docker saves build-arg values
in the image’s history, so anyone who can pull the image from ECR can
read them with docker history. No token is needed: EJAM and
ejamdata are public, and with EJAMDATA_VERSION
set the data files download without one. (The workflows pass the run’s
own github.token, which expires when the job ends, so the
copy left in the image history no longer works.)
The .dockerignore (on the deploy branch) excludes
.RData, .Rhistory, .Rproj.user,
.git, .github, ejam-infra/, and
several docs and test folders from the build context. (The build does
not copy anything from the context anyway; see Where the operational files
live.)
Infrastructure changes (Terraform)
App code deploys happen via GitHub Actions. AWS
infrastructure changes (resize a task, add HTTPS,
change retention) are made by editing the
.tf/.tfvars files and running
terraform apply locally from ejam-infra/ as
shown above.
Custom domain / HTTPS: set
domain_name = "ejam.yourdomain.com" in the relevant
.tfvars, run terraform apply; Terraform
outputs the CNAME records to add at the DNS provider (Squarespace) for
ACM certificate validation. Run terraform apply once more
after adding them — HTTP then redirects to HTTPS automatically.
Rollback
Point the ECS service back at an earlier task-definition revision:
# List recent revisions
aws ecs list-task-definitions --family-prefix ejam --sort DESC \
--query 'taskDefinitionArns[:5]' --output text
# Prod
aws ecs update-service --cluster ejam-prod-cluster \
--service ejam-prod-service --task-definition ejam:<REVISION>
# Dev
aws ecs update-service --cluster ejam-dev-cluster \
--service ejam-dev-service --task-definition ejam-dev:<REVISION>Logs and debugging
# Service health (running vs. desired)
aws ecs describe-services --cluster ejam-prod-cluster --services ejam-prod-service \
--query 'services[0].{Status:status,Running:runningCount,Desired:desiredCount}'
# Recent errors in the last 30 minutes (prod). `date +%s` (current epoch) is
# portable across macOS and Linux; 1800 s = 30 min. CloudWatch wants milliseconds.
aws logs filter-log-events --log-group-name /ecs/ejam-prod \
--filter-pattern "Error" \
--start-time $(( ($(date +%s) - 1800) * 1000 )) \
--query 'events[*].message' --output text| Symptom | Likely fix |
|---|---|
UnauthorizedOperation on an EC2/ECS/IAM action |
add the missing action to the ejam-terraform-deploy IAM
policy |
| Docker build fails | confirm Docker Desktop is running and .dockerignore is
present |
| ECS tasks failing health checks | confirm the container serves HTTP 200 on the health-check port (2001) and the app port (2000) matches the task definition |
| Slow local Docker push | use GitHub Actions instead |
| Chart text in reports shows as boxes | the Dockerfile CMD is missing
system('fc-cache -r') (see Which EJAM version deploys) |
Container logs show the ejamdata files downloading at
every start |
ARG EJAMDATA_VERSION does not equal the installed
EJAM’s ejamdata_required_tag
|
| The app footer shows a different version than the one just deployed | two deploys ran at once (a merge and a manual run); run one again and let it finish alone |
Testing against a local or draft API
The app computes analysis results in-process via
ejamit() – it does not call the EJAM REST API for analysis.
Where the API base URL matters is everywhere the app or package
builds URLs that point at the API: the per-site report
links in results tables and map popups (built by
url_ejamapi()), ejamapi() calls, and the
EJScreen-to-EJAM token handoff. All of those read the base URL from one
place, url_package("api"), which normally comes from the
url_api field of DESCRIPTION (the production
API, https://api.ejanalysis.com).
To test an app release candidate against a different
API – most usefully the local API served by
EJAM:::ejamapi_local(), which mirrors the latest EJAM-API
code before it is deployed anywhere (see the API
article) – override that one lookup. Precedence is:
options(ejam.api.baseurl=...) first, then the environment
variable EJAM_API_BASEURL, then
DESCRIPTION.
Local app + local API (no infrastructure needed):
apiproc <- EJAM:::ejamapi_local() # local API at http://127.0.0.1:3035
Sys.setenv(EJAM_API_BASEURL = "http://127.0.0.1:3035")
# or, equivalently: options(ejam.api.baseurl = "http://127.0.0.1:3035")
EJAM::ejamapp() # report links in tables/popups now
# hit the local API, end to end
# when done:
Sys.unsetenv("EJAM_API_BASEURL")
apiproc$kill()The same override works for one-off calls without the app, e.g.
ejamapi(fips = "10001", endpoint = "data") or
url_ejamapi(lat = 34, lon = -118) will target whatever base
is set.
Deployed app (dev server) + draft API: the
AWS-hosted dev app cannot reach a laptop’s localhost, so point its
EJAM_API_BASEURL (an environment variable in the ECS task
definition) at any reachable draft API deployment – e.g. a future
apidev.ejanalysis.com staging service (EJAM-API#47
tracks setting one up), or a temporary Cloudflare tunnel exposing a
locally-run API. Unsetting the variable (or omitting it) restores the
production API from DESCRIPTION.
Teardown
The prod ALB has deletion protection enabled;
disable it in the AWS Console first (EC2 → Load Balancers →
ejam-prod-alb → Edit attributes → Deletion protection: off)
before terraform destroy will succeed on prod.
Future direction
The long-lived dev-deploy / prod-deploy
branches are a workable pattern, but their two Dockerfiles
and deploy workflows must be kept in step by hand. The modern,
trunk-based alternative is to keep the infra + Dockerfile
in subdirectories on main (marked in
.Rbuildignore so they don’t affect the R package build) and
deploy via workflow_dispatch/tag triggers targeting
GitHub Environments
(dev/prod) with approval gates — which removes
the environment branches and their divergence entirely. This is a larger
change and is the deployment maintainer’s call; it is noted here as the
recommended target state.
About this document
This article consolidates two documents that previously lived only on
the deploy branches: DEPLOY-GUIDE.md (root; removed from
those branches in #633
and #634)
and ejam-infra/README.md. Where they disagreed, it follows
the actual deploy-branch files (main.tf,
Dockerfile, the deploy workflows): GitHub Actions
deploys are live (the older DEPLOY-GUIDE.md said
“coming soon”), and the container serves on ports
2000/2001 (as in DEPLOY-GUIDE.md and the current
main.tf/Dockerfile — the
ejam-infra/README.md wireframe’s “3838” was inaccurate).
DEPLOY-GUIDE.md also describes promoting code by merging
main → dev-deploy → prod-deploy;
that no longer applies, because each image installs EJAM from GitHub at
the ref in its Dockerfile (see Branching and deploy model).
Concrete AWS account IDs and individual contact names have been replaced
with <ACCOUNT_ID> and role descriptions; substitute
the real values from the deploy branch files. If you own the deployment,
treat this as the single source of truth and retire the two originals
(or replace them with a pointer here).