Troubleshooting & the Error Catalogue

A triage order that works, the eight root-cause classes behind virtually every Terraform error, a grouped catalogue of the hundred most common failures, and the resolution patterns that fix most of them.

advanced 22 min lesson hands-on task included

Terraform errors are unusually informative once you know which stage produced them. The stage tells you which tool to reach for, and that shortcut is worth more than memorising any individual message.


Topic 1: The Triage Order

                        ┌─────────────────────────────┐
   Error ──────────────▶│ terraform validate passes?  │
                        └───────────┬─────────────────┘
                            no │    │ yes
            ┌──────────────────┘    ▼
            ▼                 ┌─────────────────────────────┐
   syntax · schema · typo     │ terraform init succeeds?    │
                              └───────────┬─────────────────┘
                                  no │    │ yes
                ┌────────────────────┘    ▼
                ▼                   ┌─────────────────────────────┐
   provider · module · backend      │ terraform plan succeeds?    │
   resolution                       └───────────┬─────────────────┘
                                        no │    │ yes
                      ┌────────────────────┘    ▼
                      ▼                   ┌─────────────────────────────┐
   expressions · types · dependencies     │ terraform apply succeeds?   │
   variables                              └───────────┬─────────────────┘
                                              no │    │ yes
                            ┌────────────────────┘    ▼
                            ▼                   compare plan to expectation
   credentials · IAM · quota · timeout
   state lock

Reach for tools in this order:

terraform validate      # schema and syntax — instant, no network
terraform console       # inspect actual values and types
terraform plan          # what Terraform intends
terraform graph         # dependency cycles and ordering
TF_LOG=DEBUG terraform apply    # last resort, usually decisive
terraform show          # what Terraform believes exists
export TF_LOG=TRACE          # TRACE | DEBUG | INFO | WARN | ERROR
export TF_LOG_PATH=./tf.log  # persist it — TRACE output is enormous

TF_LOG=DEBUG shows every API request and response. It is the fastest way to find out that a provider is being told something different from what you think you wrote.


Topic 2: The Eight Root-Cause Classes

ClassTypical causes
SyntaxMissing brackets, commas, quotes; unsupported constructs
Provider configurationMissing block; outdated version; unsupported argument
DependenciesCircular references; missing depends_on; wrong ordering
BackendWrong or missing config; authentication or permission failure
VariablesNot passed; invalid value; type mismatch
StateConcurrent access and locking; manual edits causing corruption
Unsupported featuresArgument absent in the provider version in use
ExternalNetwork/API connectivity; quota and rate limits

Topic 3: The Catalogue

The hundred most common errors, grouped by cause with the resolution pattern for each group.

Credentials and permissions — No valid credential sources found; Insufficient IAM permissions; Provider requires authentication; Unauthorized to access remote backend; Operation not allowed for resource. Verify the credentials file exists and is well-formed; check the environment variables are actually exported; confirm the correct profile is selected; verify the caller identity the provider is really using; test the required actions against the policy; widen the policy to include the create/update/delete actions the plan needs.

Configuration syntax and schema — Invalid argument name; Missing required argument; Unsupported attribute; Unsupported argument; Unsupported block type; Unknown root level key; Invalid resource name; Invalid character in resource name; Duplicate resource definition; Duplicate output definition; Duplicate variable declaration; Argument or block required; Failed to parse configuration file; Invalid character in string. Check provider documentation for exact spelling and case; run terraform validate to localise; run terraform fmt to normalise; restrict names to alphanumerics and underscores; avoid reserved keywords; ensure names are unique within a module; upgrade the provider if the argument exists only in newer versions.

Provider and version resolution — The provider does not support resource type; Provider configuration not present; Provider requires explicit configuration; Required provider could not be found; Failed to query available provider packages; Unable to fetch provider; Provider binary not found; Invalid provider version constraint; Unsupported Terraform version; Incompatible versions between module and provider. Confirm the resource type exists at the pinned version; add or correct required_providers; run terraform init -upgrade; check registry reachability and proxy rules; run terraform providers to see what is actually installed; align the CLI to required_version with a version manager; consult changelogs for compatibility.

State and locking — State file is locked by another process; State lock already held; State file is corrupted; Failed to load state file; State file format not supported; Resource not found in state; Resource already exists. Check whether another run is mid-apply; inspect backend logs to identify the lock holder; terraform force-unlock <ID> only after confirming nothing is running; restore from backend versioning or terraform.tfstate.backup; terraform state pull for the authoritative copy; align the CLI version to the state format; terraform import to adopt something that exists but is not tracked; terraform state rm to drop a stale entry.

Backend configuration — Backend configuration not found; Invalid backend configuration; Unable to determine remote backend configuration; Failed to configure remote backend; Backend bucket does not exist; Remote backend configuration mismatch; Backend type not supported. Confirm every required parameter is present; verify the bucket and lock table exist and are reachable; check IAM permissions and the network path; re-run terraform init after any change; never change backend settings out of band without reconciling state.

Dependencies and ordering — Circular dependency; Cycle detected in dependencies; Resource dependencies are incomplete; Invalid resource dependency; Attribute cannot be computed; Cannot assign computed value to variable; Cannot assign computed value to output; Local value references itself. Run terraform graph to visualise the cycle; break it by refactoring rather than adding depends_on; split tightly coupled resources into modules; add explicit depends_on only where no implicit reference can exist; move computed values into outputs rather than variables; split self-referencing locals.

Expressions, types and collections — Inconsistent conditional expression; Invalid count argument; Invalid for_each argument; Invalid index; Invalid map key; Invalid type conversion; Invalid type for variable; Invalid function argument; Invalid interpolation syntax; Unsupported conditional operator; Cannot read property from null value; Cannot use a null value in this context; The object does not have an attribute; Attribute is read-only; Conflicts with other values; Invalid argument combination; Invalid lifecycle block. Both branches of a conditional must return the same type; count must resolve to a non-negative integer and must not be null; for_each needs a map or a set of strings — wrap a list in toset() and ensure keys are unique; check bounds with length() before indexing; use tolist()/toset()/tomap() for conversions; use coalesce() for null fallbacks; check documentation for mutually exclusive arguments. terraform console resolves this class faster than anything else.

Variables and validation — Invalid value for variable; Unknown variable; Variables not passed to module; Variable validation failed; Invalid output value; Module outputs not found. Confirm the variable is declared and spelled identically at both ends; check which assignment method is supplying it and whether precedence is overriding it; read the validation condition to understand the rejection; add defaults where genuinely optional; verify the module actually declares the referenced output.

Modules — Module not found; Failed to load module; Missing module source; Failed to resolve module source; Missing required module version. Verify the source path or URL; run terraform get; confirm the repository exists and authentication is configured for private sources; specify a version or ?ref= tag; re-run terraform init.

Resource runtime failures — Instance type not supported in region; Timeout while waiting for resource; Error creating resource; Insufficient memory for resource creation; Resource does not support import; Disk quota exceeded; Too many open files; Failed to create symbolic link. Check regional availability and service quotas; raise the timeouts block for slow resources; read the provider’s own error for the underlying API message; check disk space with df -h; raise the file-descriptor limit with ulimit -n; break very large configurations into smaller modules.

Data sources — Unsupported data source; Invalid data source configuration. Confirm the data source exists at the pinned provider version; check required arguments; inspect what it returns with terraform console.


Topic 4: Resolution Patterns

terraform validate                 # schema and syntax, before anything else
terraform console                  # inspect actual values and types
terraform init -upgrade            # provider and module resolution
terraform graph                    # dependency cycles
terraform plan -refresh-only       # drift detection
terraform import <addr> <id>       # adopt something that exists but is untracked
terraform state rm <addr>          # drop something Terraform should not own
terraform force-unlock <ID>        # stale lock, only after confirming
TF_LOG=DEBUG terraform apply       # the decisive last resort

Timeouts deserve a specific note:

resource "aws_db_instance" "app" {
  # ...
  timeouts {
    create = "40m"
    update = "80m"
    delete = "40m"
  }
}

Default timeouts assume a typical resource. Large databases, big clusters and cross-region replicas routinely exceed them, and the resulting error looks like a failure when the resource is actually still being created. Check the console before retrying — a retry on a half-created resource often makes things worse.


Topic 5: The Errors That Are Not Errors

Three situations that look like failures and are not:

  • “Resource already exists” on a first apply usually means the resource was created by hand or by a previous run whose state was lost. import it, do not rename around it.
  • A plan proposing to destroy everything almost always means Terraform is pointed at empty state — a wrong backend key, a wrong workspace, or a failed migration. Check terraform workspace show and the backend key before you touch anything.
  • Drift on every run for one attribute means another system owns that field. That is a configuration decision (ignore_changes), not a bug.

Try it yourself: Feed a list to for_each and read the error. Then paste the same expression into terraform console wrapped in toset(). The console answers in a second what the plan takes a full cycle to tell you.

Common mistake: Jumping straight to TF_LOG=DEBUG. It produces an enormous amount of output and buries a schema typo that terraform validate would have named in under a second. Work down the triage order — debug logging is the last step, not the first.