Terraform Roadmap
Seven stages, each gated by a capability rather than a lesson count. The path is ordered around the two things that actually cause incidents — misread plans and mishandled state — rather than around the order the language documentation happens to use.
Orderings that are not optional
Most lessons can be taken in any order within a stage. These four cannot.
- Reading a plan before applying anything Almost every self-inflicted Terraform incident is visible in a plan somebody scrolled past.
- State before modules and environments Both are decisions about how to split state. Without understanding what state is, you are splitting something you cannot describe.
- for_each over count before anything with a lifespan Index-based identity works perfectly until the first removal from the middle — usually months later, in production.
- A remote backend with locking before a second person Two people with two local state files is not a workflow. It is two records of the same infrastructure disagreeing silently.
Stage 1 — Foundations
~1hDescribe the end state and stop writing steps
A cloud account you can create and destroy cheap resources in. Comfort with a terminal.
The hardest part of this stage is unlearning procedural thinking. If your configuration reads as a sequence you mentally execute in order, you will fight the tool for months.
- ▸Explain why declarative provisioning derives the steps rather than taking them from you
- ▸Pin provider versions so a major release cannot land on a day you changed nothing
- ▸Build the dependency graph implicitly by referencing attributes instead of hard-coding values
- ▸Read a plan the way it should be read — summary, then destroys, then every forces-replacement marker
- ▸Use lifecycle meta-arguments to make replacement safe, deletion impossible, or drift ignorable
Handed a plan you did not write, you can say within a minute whether applying it causes downtime — and point at the exact attribute that decides it.
Stage 2 — The Terraform Language
~1.5hStop repeating yourself, without becoming unreadable
Stage 1.
- ▸Declare types deliberately and know the coercion rules that bite when you do not
- ▸Place values correctly — variable, local, output, or data source
- ▸Never compute a value you could reference, and explain the three things that costs you
- ▸Choose for_each over count on identity grounds, and migrate an existing count without churn
- ▸Render templates with loops, and build JSON with an encoder rather than string concatenation
- ▸Reach for the interactive console before reaching for a plan
You can take a configuration with twenty near-identical resource blocks and collapse it, and the reviewer finds the result easier to read rather than harder.
Stage 3 — State
~1hThe artifact whose loss ruins your day
Stage 2. A second person, or a second machine, to make the problem real.
State holds your secrets in plain text. Everything in this stage follows from that one fact.
- ▸Give four reasons state must exist, not one
- ▸Configure a remote backend with locking, and migrate to it without recreating anything
- ▸Treat backend read access as equivalent to production admin, because it is
- ▸Import hand-built infrastructure and move resources between projects without destroying them
- ▸Detect drift with a refresh-only plan and choose between three valid responses
- ▸Explain why there is no rollback command, and what you actually do instead
Given a project whose plan proposes destroying everything, you can diagnose whether it is a wrong backend key, a wrong workspace, or a genuine problem — before touching anything.
Stage 4 — Structure & Reuse
~1hMake change cheap in one place and consistent everywhere
Stage 3.
- ▸Build modules with an interface a stranger can use from the README alone
- ▸Pin remote modules by tag, and know why the version argument silently does nothing for Git sources
- ▸Judge when nesting stops paying for itself — usually one level down
- ▸Choose between workspaces, directory-per-environment and shared modules on isolation grounds
- ▸Split state along team and lifecycle boundaries before the plan takes fifteen minutes
- ▸Guarantee a dev run cannot write production state, with a control rather than a convention
You can defend your environment layout to someone who prefers the other one, in terms of blast radius rather than taste.
Stage 5 — Delivery
~1hMake the pipeline the thing that applies, and make it safe
Stage 4.
- ▸Rank the alternatives to provisioners and reach for one only when the first three fail
- ▸Run static checks that cost nothing on every commit
- ▸Split plan and apply into separate jobs so a pull request can never apply
- ▸Pass a reviewed plan file to apply, and explain what that guarantees over auto-approve
- ▸Write policy against the JSON plan rather than the source, and know why that distinction matters
Your pipeline fails a non-compliant change before it reaches an environment, and the failure message tells the author how to fix it.
Stage 6 — Production & Operations
~1.5hEverything that only shows up at scale
Stage 5.
- ▸Close the three separate secret leaks with three separate fixes
- ▸Span several providers, and separate state on team and security boundaries rather than convenience
- ▸Encode disaster recovery, cost controls and monitoring alongside the resource they protect
- ▸Diagnose slow plans and know when the answer is structural rather than a flag
- ▸Triage any failure in the right order, from validate through console to debug logging
Given an unfamiliar failure, you reach the right diagnostic in two steps rather than starting with debug logging and reading ten thousand lines.
Stage 7 — Capstone Project
~1hAssemble the whole path once
Stages 1–6.
- ▸Compose five modules through outputs rather than duplicated values
- ▸Chain security groups by identity instead of by address range
- ▸Migrate from local to remote state and prove the migration worked
- ▸Verify the account is genuinely empty after a destroy
The load balancer serves traffic, the app tier reaches the database, state lives remotely, plan reports no changes after migration — and destroy leaves nothing behind.
Where Terraform stops
It provisions infrastructure. It has no procedural logic, no rollback command, and no business being the deploy mechanism for a fast-moving application sitting behind an infrastructure state lock. Knowing the boundary is as valuable as knowing the tool.