Terraform
From the declarative model to a production three-tier build β providers, plans, state, modules, delivery pipelines and the failure modes that cause real outages.
Stage 1 β Foundations
3 lessonsWhat manual infrastructure management actually costs, why declarative beats procedural for provisioning, where Terraform sits relative to configuration management, and the two-part architecture that explains most of its behaviour.
The two blocks every configuration starts with, why pinning provider versions is not optional, how resource references build the dependency graph for free, and the lifecycle meta-arguments that control replacement.
The four plan symbols and which two should stop you, why 'forces replacement' is the most important string in the output, the graph and parallelism, and the difference between applying a saved plan file and using auto-approve.
Stage 2 β The Terraform Language
4 lessonsAll eight variable types including the structural ones, why declaring a type prevents a class of silent coercion bugs, validation blocks, and the assignment precedence order that decides which value actually wins.
The three constructs that stop a configuration repeating itself: named expressions, return values, and read-only lookups of infrastructure you do not manage β plus why you should never compute a value you could reference.
Conditional expressions, the built-in function library you actually reach for, heredoc strings, externalising policy files, and rendering templates with variables and loops.
How to create many resources from one block, why choosing count over for_each is the most common self-inflicted outage in Terraform, migrating between them safely, and when generating nested blocks stops being worth it.
Stage 3 β State
3 lessonsThe four reasons Terraform cannot work by reading the world on every run, what state actually contains, why it holds your secrets in plain text, and the inspection commands that tell you what Terraform believes.
Why local state stops working the moment a second person appears, the backend options and their locking mechanisms, partial configuration for keeping secrets out of the repo, and migrating an existing project without losing anything.
Bringing hand-built infrastructure under management, moving resources between projects without destroying them, detecting out-of-band changes, and why Terraform has no rollback command.
Stage 4 β Structure & Reuse
3 lessonsPackaging resources into reusable units, module inputs and outputs as a public interface, sourcing from registries and Git, why pinning with a ref tag is not optional, and how deep to nest before it stops paying.
How Terraform actually treats files and folders, the four ways to model dev/staging/prod and when each breaks down, and why splitting state along team boundaries is an organisational decision more than a technical one.
Multiple state instances from one configuration directory, referencing the active workspace in code, the delete command that removes state but not infrastructure, and the point at which workspaces stop scaling.
Stage 5 β Delivery
3 lessonsWhy every source warns that provisioners are a last resort, the structural reason behind it, creation-time versus destroy-time behaviour, and the three better alternatives ranked.
The three testing tiers and which tool serves each, static analysis that costs nothing, and the pipeline design whose single most important property is passing a reviewed plan file to apply.
Writing compliance rules that evaluate a plan before it applies, the two policy ecosystems, and turning drift detection and audit logging into continuous compliance rather than an annual scramble.
Stage 6 β Production & Operations
5 lessonsThe three places secrets leak in a Terraform workflow and the three different fixes, why sensitive = true is not a secrets strategy, and locking down the state backend as the production-admin boundary it actually is.
Managing several providers from one configuration, connecting on-premises networks to cloud, provisioning Kubernetes clusters and serverless functions, and where to draw the line between infrastructure and application delivery.
Encoding disaster recovery and high availability as code, automating backups and failover, the cost levers Terraform can pull, and wiring alerting alongside the resource it watches.
What slows down large estates and the four levers that fix it, managing state size, and writing a custom provider when nothing exists for the API you need.
A triage order that works, the eight root-cause classes behind virtually every Terraform error, a grouped catalogue of the hundred most common failures, and the resolution patterns that fix most of them.
Stage 7 β Capstone Project
1 lessonBuild a complete modular three-tier stack β network, compute, load balancer, database, artifacts β with remote state and locking, then migrate it from local state to a remote backend without recreating anything.
πΊοΈ Beginner β Expert Roadmap
7 stages with prerequisites and a concrete mastery check at each.
π― What You'll Learn
- β’ Describe infrastructure as a desired end state and let the dependency graph work out the order.
- β’ Read a plan the way it should be read: summary first, then destroys, then every forces-replacement marker.
- β’ Pin provider versions and module refs so nobody else's merge lands in your production apply.
- β’ Choose for_each over count deliberately β and explain the churn that makes it the most common self-inflicted outage.
- β’ Explain the four reasons state must exist, and treat the backend as the production-secrets store it actually is.
- β’ Import hand-built infrastructure and move resources between projects without destroying them.
- β’ Detect drift with refresh-only plans and choose between reconciling toward code, toward reality, or ignoring the field.
- β’ Build modules with a real interface, and know when nesting stops paying for itself.
- β’ Ship a pipeline that passes a reviewed plan file to apply instead of auto-approving whatever the world looks like.
- β’ Keep secrets out of configuration, out of logs, and understand why they are still in state.
- β’ Triage any Terraform failure in the right order, from validate through console to debug logging.
- β’ Assemble the whole path into a modular three-tier stack and migrate it to a remote backend without recreating anything.
π‘οΈ Best Practices in Production
The short version of this path. Every lesson also ends with the specific mistake it exists to prevent.
- β Read every plan summary first, then destroys, then every forces-replacement marker.
- β Pin provider and module versions; a floating version means someone else's release lands in your apply.
- β Use a remote backend with locking and versioning, and treat state as production-secret material.
- β Prefer
for_eachovercountso removing one item does not renumber the rest. - β Run
fmt,validateand a policy check in CI, and apply a reviewed plan file rather than re-planning. - β Keep modules small with a real interface; nest only while it still pays.
- β Import existing infrastructure rather than recreating it, and use
movedblocks when refactoring. - β Detect drift with scheduled refresh-only plans and decide deliberately which way to reconcile.
- β
terraform apply -auto-approvein a pipeline that has not shown anyone the plan. - β Editing state by hand, or running
taintwhen-replaceexpresses the same intent reviewably. - β One giant root module for the whole estate β every change plans everything and blast radius is total.
- β Secrets in variables or outputs. They are stored in plaintext in state regardless of provider support.
- β
counton a list that people insert into β the churn is the most common self-inflicted outage. - β Committing
.terraform/or a localterraform.tfstateto Git.
πΌ Interview Readiness
Once you reach the end of this path, test your engineering knowledge against real questions asked by top technical teams:
What are the four Terraform plan symbols, and which two should stop you?
Infrastructure as CodeWhat does applying a saved plan file guarantee that -auto-approve does not?
Infrastructure as CodeWhat is Terraform's variable definition precedence order?
Infrastructure as CodeDoes marking a variable or output sensitive keep the value out of state?
Infrastructure as CodeWhat is the difference between count and for_each, and when does the choice cause an outage?
Infrastructure as CodeWhy does Terraform need a state file β why can it not just read the live infrastructure?
Infrastructure as CodeWhat is the difference between terraform state rm and terraform destroy?
Infrastructure as CodeWalk through importing existing infrastructure into Terraform. What does import not do?
Infrastructure as CodeWhat is configuration drift, and what are your three valid responses?
Infrastructure as CodeHow do you pin a remote Terraform module, and where does the version argument silently not work?
Infrastructure as CodeWhat does terraform workspace delete remove β and what does it not?
Infrastructure as CodeWhy are provisioners considered a last resort in Terraform?
Infrastructure as Code