The decisions behind a production-ready ECS Fargate stack
Most ECS Terraform starters skip the choices that matter. Private tasks, per-environment sleep, alarms that page, and OIDC deploys, encoded so production cannot inherit the wrong default.
Most ECS Terraform starters have a feature list. They do not have a spine. Tasks land in public subnets because that is easier to demo. Staging runs all weekend because nobody parameterized sleep. Alarms exist as a screenshot. AWS keys sit in GitHub secrets. Then a team copies the repo into production and inherits every one of those defaults.
The problem is not Fargate. It is that the template never made the hard choices, so every apply becomes a negotiation. I published a production-ready ECS Fargate Terraform starter that encodes those choices. Here is what it refuses to leave as a comment in a README.
Put tasks where the internet cannot reach them
The load balancer belongs in public subnets. The tasks do not.
A Fargate task with assign_public_ip = true will run. It will also mean a security-group slip exposes the process, not just the ALB. The starter puts tasks in private subnets across two AZs, with no public IP. Egress goes through NAT. Image pulls and logs go through VPC endpoints for ECR, CloudWatch Logs, and S3, so the NAT gateway is not on the hot path for every deploy.
One NAT in non-prod. Two in production. That is a cost and availability trade-off you should make on purpose. A single NAT is cheaper and is also a single AZ of egress. The prod tfvars pay for the second gateway. Dev does not.
HTTPS is optional on the first apply so you can get a URL in an empty account. That is a bootstrap default, not a production posture. Set domain_name and hosted_zone_id before real users. The ALB then redirects HTTP to HTTPS with a current TLS policy.
Same module, different answers per environment
Copy-pasting a root module for staging and production is how drift starts. One module with flags is how you keep the blast radius explicit.
The three that matter:
allow_scale_to_zero— non-prod sleeps overnight and on weekends. Production isfalsein tfvars, not in a wiki.use_fargate_spot— fine for a sandbox. Not an availability strategy. Prod leaves it off.desired_count/min_capacity— one task is a demo. Production starts at two so a deploy or a zone blip does not take you to zero.
I scale non-production Fargate to zero on a schedule for the same reason. The schedule lives behind the flag so a staging overlay cannot leak into production. Application Auto Scaling scheduled actions set min and max to zero in the evening and restore them in the morning. AWS documents scheduled scaling for ECS.
State is split the same way. ecs-fargate/dev/terraform.tfstate is not ecs-fargate/prod/terraform.tfstate. Shared state across environments is how a destroy in dev becomes a story you tell later.
Alarms that page, and alarms you skip on purpose
Metrics on a dashboard are not observability. An alarm that reaches a human is.
The starter creates an SNS topic and, if you pass alarm_email, a subscription. It pages on ALB 5xx and unhealthy hosts, with M-out-of-N evaluation so a single spike does not wake anyone. Log groups have a retention period. Indefinite retention is a storage bill, not a strategy.
The important omission: a healthy-host count of zero is an outage in production. In an environment that is supposed to sleep, it is the schedule working. That alarm is skipped when allow_scale_to_zero is true. If you do not skip it, on-call learns to ignore the pager, which is worse than having no pager.
Treat missing data as notBreaching on error counts and breaching on "is anything healthy." That matches how the signals actually fail. I walk through the rest of this in CloudWatch defaults exist but nobody reads the alarms.
Deploy with OIDC, tag images with a SHA
Long-lived access keys in GitHub secrets are a standing grant. The starter's deploy workflow assumes a role with GitHub OIDC. The trust policy is pinned to repo:ORG/REPO:ref:refs/heads/main. Without the sub condition, any repository on GitHub can request credentials for your account.
The bootstrap stack creates that role once, with local credentials, so the pipeline is not trying to create the identity it needs to run. If the account already has a GitHub OIDC provider, you import it. You do not create a second one.
Images are immutable. The pipeline tags sha-<git sha> and rolls the task definition with a circuit breaker and rollback. latest is how two environments silently diverge. The ECR lifecycle policy keeps the last 20 images so the registry does not grow forever.
The workflow is manual dispatch into GitHub Environments named dev, staging, and prod. Put a required reviewer on production. That is the same gate I use when deploying to AWS from GitHub Actions with OIDC.
What the starter refuses to include
A template that provisions RDS, Redis, and a service mesh "just in case" is not production-ready. It is a bill with extra steps.
No database until you have a schema. No public tasks "for debugging." No Spot in prod because it is cheaper. No shared Terraform state. No latest tag. No alarm that fires every night in a sleeping environment.
The sample app is a small Go binary that serves / and /health. It exists so the first rollout is a container you built, not an unsourced nginx you forgot to replace. Swap the image. Keep the flags.
The takeaway
Production-ready means the dangerous defaults are not available. Private tasks, per-environment sleep and Spot, alarms that match how the environment is supposed to behave, and deploys that never store AWS keys. Clone the ECS Fargate Terraform starter, apply envs/dev.tfvars.example, and read DECISIONS.md next to the code. For the uptime layer that sits on top of this, read meeting a 98% SLA on AWS.
Frequently asked questions
Why not put tasks in public subnets and skip NAT?
It is cheaper, and it is how most demos work. It is also how a wide security group becomes a public container. Pay for NAT (and endpoints) or do not run this in an account that matters.
Can I turn on scale-to-zero in production?
No. Production should stay available. The flag exists so non-prod can sleep without a fork of the module. If you need lower prod cost, right-size CPU and use autoscaling. Do not sleep the service users hit.
Why is HTTPS optional?
So the first apply works without a hosted zone. HTTP on the ALB DNS name is for you. It is not for customers. Set a domain before you share the URL.
Do I need Container Insights?
Not to get paged. The starter alarms on ALB metrics, which exist without Insights. Enable Insights in prod when you want per-task CPU and memory. It costs extra.
How do I add a database later?
Put it in the private subnets, lock its security group to the task security group, and keep credentials in Secrets Manager. Do not add it to this repo until you have migrations to run.