Terraform stays fast until it does not. The transition is gradual, and by the time a plan takes fifteen minutes the fixes are structural rather than configurational.
Topic 1: Why Large Estates Get Slow
Every plan refreshes state against reality by default — one API call per resource. At fifty resources that is invisible; at two thousand it dominates, and providers start rate-limiting you.
time terraform plan
time terraform plan -refresh=false
The gap between those two numbers is your refresh cost, and it tells you which lever to pull.
Topic 2: The Four Levers
1. Parallelism.
terraform apply -parallelism=20 # default is 10
terraform apply -parallelism=5 # when the API is rate-limiting you
Raise it when the provider API can absorb the concurrency; lower it when you see throttling errors. This is a per-run flag, not a configuration setting, so it can differ between local runs and CI.
2. Skip the refresh.
terraform plan -refresh=false
terraform plan -target=module.network
Both trade accuracy for speed — Terraform will not detect out-of-band changes on that run. Acceptable for iteration, not before a production apply.
-target in particular is a debugging tool, not a workflow. Routinely targeting means your state is too large, which is lever three.
3. Split the state. The real fix. One enormous root module gives slow plans and unlimited blast radius. Split along team and lifecycle boundaries — the lines that decide who can break what and what changes at the same cadence.
data "terraform_remote_state" "network" {
backend = "s3"
config = {
bucket = "org-terraform-state"
key = "platform/network/terraform.tfstate"
region = "us-east-1"
}
}
resource "aws_instance" "app" {
subnet_id = data.terraform_remote_state.network.outputs.private_subnet_ids[0]
}
The cost of splitting is that cross-state references become explicit — an upstream state must output anything a downstream state needs. That constraint is healthy: it makes coupling visible instead of implicit.
4. Keep providers current. Provider releases carry both bug fixes and performance improvements, and older versions sometimes make N API calls where newer ones make one.
Topic 3: Managing State Size
State grows with every attribute of every resource, and large state is slow to download, parse, lock and write.
- Remove what you no longer manage.
terraform state rmfor resources handed to another team or decommissioned outside Terraform. - Use
ignore_changesfor churn attributes. A field that changes on every refresh — a timestamp, an autoscaler-managed count — bloats state and produces noisy plans. - Avoid outputting large objects unnecessarily. Outputting an entire resource is convenient for debugging and writes every attribute into state.
terraform state list | wc -l # resource count
terraform state pull | wc -c # state size in bytes
Those two numbers, tracked over time, are the earliest signal that a split is coming.
Topic 4: When to Write a Custom Provider
Three legitimate cases:
- Unsupported APIs. An internal platform, a niche vendor, a service with no community provider.
- Internal platforms. Exposing in-house tooling as Terraform resources so teams consume it the same way they consume everything else.
- Higher-level abstractions. Wrapping several API calls into one resource that represents a meaningful unit.
Topic 5: Provider Anatomy
Providers are Go binaries built against the plugin SDK. Four components:
- Resource schema — attributes, types, validation, and which fields force replacement.
- CRUD functions — Create, Read, Update, Delete.
- State handling — mapping between Terraform state and API objects.
- Import support — so existing objects can be adopted.
func Provider() *schema.Provider {
return &schema.Provider{
Schema: map[string]*schema.Schema{
"api_url": {
Type: schema.TypeString,
Required: true,
DefaultFunc: schema.EnvDefaultFunc("PLATFORM_API_URL", nil),
},
},
ResourcesMap: map[string]*schema.Resource{
"platform_widget": resourceWidget(),
},
}
}
func resourceWidget() *schema.Resource {
return &schema.Resource{
Schema: map[string]*schema.Schema{
"name": {Type: schema.TypeString, Required: true, ForceNew: true},
"size": {Type: schema.TypeInt, Optional: true, Default: 1},
},
CreateContext: resourceWidgetCreate,
ReadContext: resourceWidgetRead,
UpdateContext: resourceWidgetUpdate,
DeleteContext: resourceWidgetDelete,
Importer: &schema.ResourceImporter{
StateContext: schema.ImportStatePassthroughContext,
},
}
}
Two details worth noting:
ForceNew: trueis what produces the# forces replacementmarker in a plan. As a provider author you decide which attribute changes require a rebuild.- Update is optional. Omit it and Terraform falls back to destroy-and-create for every change. That is a legitimate choice for immutable resources and a poor one for anything holding state.
Read is the function that matters most. It is what detects drift. A Read that returns cached or incomplete data makes the provider silently unable to notice out-of-band changes — the single most common defect in home-grown providers.
Topic 6: Testing and Publishing a Provider
- Unit tests for CRUD logic with Go’s standard testing package.
- Acceptance tests using the SDK’s test helpers, which run real provider operations against a test environment.
- Linting for style and correctness.
- Version and publish — build per-platform binaries and push to a registry or an internal artifact repository, then consume it exactly like any other provider:
terraform {
required_providers {
platform = {
source = "registry.example.internal/platform/platform"
version = "~> 1.0"
}
}
}
Try it yourself: Run terraform state list | wc -l on your largest project. Above a few hundred resources, time a plan with and without -refresh=false. The difference tells you whether you are already paying the tax that splitting state removes.
Common mistake: Reaching for -parallelism and -refresh=false as a permanent fix for slow plans. They are painkillers. If every run needs them, the state file is too big, and no flag addresses the blast-radius half of that problem.