Provisioners & the Config-Management Boundary

Why every source warns that provisioners are a last resort, the structural reason behind it, creation-time versus destroy-time behaviour, and the three better alternatives ranked.

advanced 18 min lesson hands-on task included

A provisioner runs a script on the local machine or on a created resource. Every source teaching Terraform carries the same warning, and it is worth understanding the reason rather than just obeying the rule.


Topic 1: Why They Are a Last Resort

Provisioners are imperative. Terraform has no way to reason about a change to your script the way it does for a declarative resource.

Consider a provisioner that installs packages and runs a command. Change the script and Terraform sees… nothing. The resource’s attributes did not change, so there is no diff, so no plan, so no apply. The script is not re-run on existing machines.

Worse, when it does run on a rebuild, you must guarantee the new script works both on a fresh machine and on one that already ran the old version. That is exactly the headache Terraform exists to eliminate, reintroduced inside a Terraform resource.


Topic 2: The Three Types

ProvisionerRuns whereTypical use
fileCopies from the Terraform host to the resourceShip a config file
local-execOn the machine running TerraformTrigger an external tool
remote-execOn the created resource, via SSH or WinRMBootstrap a host
resource "aws_instance" "web" {
  ami                    = data.aws_ami.base.id
  instance_type          = "t3.micro"
  key_name               = aws_key_pair.deploy.key_name
  vpc_security_group_ids = [aws_security_group.web.id]

  provisioner "remote-exec" {
    inline = [
      "sudo amazon-linux-extras enable nginx1",
      "sudo yum -y install nginx",
      "sudo systemctl start nginx",
    ]
  }

  connection {
    type        = "ssh"
    host        = self.public_ip
    user        = "ec2-user"
    private_key = file(var.ssh_private_key_path)
  }
}

Connection block parameters: host, user, private_key or password, agent for SSH agent forwarding, and timeouts for long-running commands.

provisioner "local-exec" {
  command = "ansible-playbook -i '${self.public_ip},' site.yml"
}

Topic 3: Creation-Time vs Destroy-Time

By default provisioners run at creation only — not on update, not on any other lifecycle event.

If a creation-time provisioner fails, the resource is marked tainted and will be destroyed and recreated on the next apply. That is a reasonable default: a half-bootstrapped machine is worse than no machine.

provisioner "remote-exec" {
  when       = destroy          # run when the resource is destroyed
  on_failure = continue         # default is fail, which fails the whole apply
  inline     = ["/opt/app/deregister.sh"]
}

Destroy-time provisioners are the legitimate niche — deregistering from a service discovery system, draining connections, flushing a cache before teardown. Even here, a lifecycle hook on the platform itself is usually more reliable, because it runs whether or not Terraform is what destroyed the resource.


Topic 4: The Better Alternatives, Ranked

1. Immutable images. Bake the configuration into the machine image ahead of time. The instance boots ready. Nothing runs at provision time, so nothing can fail at provision time.

2. cloud-init / user data. The platform’s own first-boot mechanism:

resource "aws_instance" "web" {
  ami           = data.aws_ami.base.id
  instance_type = "t3.micro"

  user_data = <<-EOF
    #!/bin/bash
    set -ex
    yum update -y
    amazon-linux-extras enable nginx1
    yum -y install nginx
    systemctl start nginx
  EOF
}

Two differences matter operationally. User data runs as root, so the sudo prefixes needed under a limited SSH user disappear. And the shebang plus set -ex gives per-line tracing with abort-on-error, with output landing in the instance’s cloud-init log where you can actually read it.

Critically, user_data is an attribute. Change it and Terraform shows a diff — usually a replacement, which is honest about what actually needs to happen.

3. Configuration management after apply. Let Terraform provision, then hand off:

terraform apply
terraform output -json > inventory.json    # feed a dynamic inventory
ansible-playbook -i inventory.json site.yml

This is the clean boundary: Terraform owns infrastructure, the CM tool owns what runs inside it, and each does what it is good at.

4. Provisioners. Only when the first three genuinely cannot work.


Topic 5: null_resource

A null_resource is a no-op resource that exists solely to hang a provisioner on, usually with triggers controlling re-execution:

resource "null_resource" "deploy_notice" {
  triggers = {
    version = var.app_version     # changes here re-run the provisioner
  }

  provisioner "local-exec" {
    command = "curl -X POST $WEBHOOK -d 'deployed ${var.app_version}'"
  }
}

The triggers map is the only thing that makes it re-run, and it is the only reason to use one over a bare provisioner.

The honest assessment from the source material is worth repeating: the author could not construct a genuinely good example of when to use one. If you find yourself reaching for null_resource, you are almost certainly solving the problem in the wrong layer.


Topic 6: If You Must Use One

  • Make it idempotent. It may run again on a rebuild. yum install is fine; echo >> file is not.
  • Never embed secrets. Pull them from a secret manager inside the script.
  • Set on_failure deliberately. The default fails the apply; decide whether that is what you want for a non-critical step.
  • Log somewhere durable. Provisioner output appears in the apply log and then is gone.
  • Treat it as technical debt with a note explaining why the alternatives did not work.

Try it yourself: Add a remote-exec provisioner, apply, then change the script and run plan. Terraform reports no changes. That empty diff is the entire argument against provisioners, in one command.

Common mistake: Using remote-exec to install packages on every instance because it is the most obvious path from a shell-scripting background. It works on day one and silently stops matching your code on day two, which is the worst possible failure mode for infrastructure that is supposed to be described by that code.