Terraform workflows¶
Every image ships Terraform (latest at build time, with tfswitch to change it), Terragrunt, TFLint and Trivy. This page covers the day-to-day loop, remote state, Terragrunt, testing and drift detection. The examples use aws-devops, but they work the same in all-devops, and in gcp-devops with a GCS backend.
The loop¶
flowchart LR
I["init"] --> F["fmt + validate"]
F --> L["tflint + trivy config"]
L --> P["plan -out=tfplan"]
P --> R{"review"}
P -.-> AI["AI plan summary<br/>claude -p"]
AI -.-> R
R --> A["apply tfplan"]
A -.-> D["scheduled drift check<br/>plan -detailed-exitcode"]
classDef base fill:#0891b2,stroke:#0e7490,color:#fff
classDef ai fill:#d97706,stroke:#b45309,color:#fff
classDef aws fill:#ea7a0c,stroke:#c2410c,color:#fff
classDef all fill:#059669,stroke:#047857,color:#fff
classDef neutral fill:#334155,stroke:#1e293b,color:#fff
class I,F,L,P base
class AI ai
class R aws
class A all
class D neutral
Start a shell with your project and AWS config mounted:
docker run -it --rm \
-v "$PWD":/srv -w /srv \
-v ~/.aws:/root/.aws \
ghcr.io/jinalshah/devops/images/aws-devops:latest
Then, inside the container:
terraform init
terraform fmt -recursive
terraform validate
tflint --init && tflint --recursive
trivy config --severity HIGH,CRITICAL .
terraform plan -out=tfplan
terraform apply tfplan
The zsh config also has short tf* aliases for these; see Tool basics.
Remote state¶
terraform {
backend "s3" {
bucket = "my-terraform-state"
key = "production/terraform.tfstate"
region = "eu-west-2"
encrypt = true
use_lockfile = true
}
}
No DynamoDB table needed
use_lockfile = true uses S3's native locking. The old dynamodb_table argument is deprecated. When migrating, you can set both for a while, then remove dynamodb_table.
Useful state commands:
terraform state list
terraform state show aws_instance.web
terraform state mv aws_instance.old aws_instance.new
terraform force-unlock <LOCK_ID> # only if a crashed run left a lock behind
Multiple environments¶
A directory per environment, with shared modules, keeps state and blast radius separate:
terraform/
├── modules/
│ ├── vpc/
│ └── eks/
└── environments/
├── dev/
├── staging/
└── production/
Plan every environment in one container:
docker run --rm \
-v "$PWD":/srv -w /srv \
-v ~/.aws:/root/.aws \
ghcr.io/jinalshah/devops/images/aws-devops:latest \
bash -c 'for env in dev staging production; do
echo "==> $env"
terraform -chdir=terraform/environments/$env init -input=false
terraform -chdir=terraform/environments/$env plan -input=false -out=tfplan || exit 1
done'
Workspaces (terraform workspace new staging / select staging) also work, but they share one backend configuration and one set of credentials. That usually makes separate directories the safer choice for production.
Terragrunt¶
Terragrunt keeps backend and provider config DRY across many stacks.
infrastructure/
├── root.hcl
├── dev/
│ ├── vpc/terragrunt.hcl
│ └── eks/terragrunt.hcl
└── production/
├── vpc/terragrunt.hcl
└── eks/terragrunt.hcl
Run it:
cd infrastructure/dev
terragrunt run --all plan
terragrunt run --all --non-interactive apply
terragrunt dag graph # dependency graph in DOT format
cd eks && terragrunt plan # a single stack
Old Terragrunt commands
Terragrunt's CLI was redesigned, and the image ships a recent release that is bumped automatically. Update older scripts:
| Old | New |
|---|---|
terragrunt run-all plan |
terragrunt run --all plan |
--terragrunt-non-interactive |
--non-interactive |
terragrunt graph-dependencies |
terragrunt dag graph |
--terragrunt-include-external-dependencies |
--queue-include-external |
Terraform versions with tfswitch¶
The image has the latest Terraform at build time. If a project needs a different version, tfswitch reads required_version from your .tf files (or a .terraform-version file) and installs a match:
tfswitch # pick the version from required_version / .terraform-version
tfswitch 1.9.8 # or name it explicitly
terraform version
If terraform version doesn't change, write over the binary on PATH with tfswitch -b "$(command -v terraform)" 1.9.8. The switch only lasts for the life of the container, so put it at the start of each CI job.
Testing¶
Native tests (*.tftest.hcl) need nothing beyond Terraform:
Terratest is written in Go, and Go isn't in the image. Run Terratest from a Go toolchain image, or build your own image FROM a DevOps image and add Go.
Import existing resources¶
Use import blocks and let Terraform write the config for you:
terraform plan -generate-config-out=generated.tf
# review and tidy generated.tf, then:
terraform apply
Drift detection¶
-detailed-exitcode returns 0 for no changes, 2 for changes and 1 for an error:
#!/usr/bin/env bash
set -uo pipefail
terraform init -input=false >/dev/null
terraform plan -input=false -detailed-exitcode -out=drift.tfplan
case $? in
0) echo "No drift" ;;
2) echo "Drift detected"; terraform show -no-color drift.tfplan > drift.txt; exit 2 ;;
*) echo "terraform plan failed"; exit 1 ;;
esac
Run it on a schedule from your CI system (a schedule: trigger in GitHub Actions, a pipeline schedule in GitLab, a cron trigger in Jenkins or a scheduled pipeline in CircleCI), so it runs in the same image as your deploys.
AI help with plans and code¶
All four AI CLIs are in every image. In scripts, always use their non-interactive modes.
mkdir -p modules/vpc
claude -p "Write Terraform HCL for an AWS VPC with three public and three private subnets across three AZs. Output only HCL, with no commentary or code fences." \
> modules/vpc/main.tf
terraform -chdir=modules/vpc init -backend=false
terraform -chdir=modules/vpc validate
Always review generated code, then run validate, tflint and trivy config on it.
Plans can contain secrets
A plan can include sensitive values. Check your organisation's policy before sending one to an external AI service, and prefer sending the diff of your .tf files instead.
Authentication for each CLI is covered in AI-assisted DevOps.
Troubleshooting¶
Error acquiring the state lock
Another run holds the lock. Wait for it to finish, or, if a run crashed, use terraform force-unlock <LOCK_ID> with the ID from the error message. In CI, serialise runs per state (for example a concurrency group in GitHub Actions or a resource_group in GitLab).
Providers download on every run
The image doesn't set a provider cache. Set TF_PLUGIN_CACHE_DIR to a directory you mount or cache:
Unsupported Terraform Core version
Your required_version excludes the image's Terraform. Run tfswitch first (see Terraform versions with tfswitch).