Error Handling, Defensive Blocks & Dry-Run Testing

Master defensive playbook design: Task Blocks (block, rescue, always), error override flags (ignore_errors, failed_when), smoke testing with uri module, and --check mode.

intermediate 25 min lesson hands-on task included

In production automation, unhandled errors leave servers in inconsistent, half-configured states.

Ansible provides Task Blocks (block, rescue, always), error override flags (failed_when, changed_when), and dry-run execution modes (--check, --diff) to build fault-tolerant playbooks with automated rollbacks and pre-deployment validations.


Topic 1: Defensive Task Blocks (block, rescue, always)

DEFENSIVE TASK BLOCKS & ANSIBLE VAULT ENCRYPTION TASK BLOCK & RESCUE ERROR CONTROL block: Primary Critical Tasks • Apply database migrations • Restart production app service rescue: Rollback on Failure • Revert DB schema snapshot • Send Slack alert: "Deployment failed!" always: Cleanup (Runs Regardless) • Remove temporary download files • Close diagnostic SSH sessions ANSIBLE VAULT (AES-256 SECRETS) Vault Encrypted File (group_vars/vault.yml) $ANSIBLE_VAULT;1.1;AES256 6638303038623035343461623938... Decrypted In Memory During Playbook Run ansible-playbook site.yml \ --vault-password-file ~/.vault_pass • Secrets stay encrypted at rest in git! • In-memory decryption only during task execution SMOKE TESTING WITH URI MODULE Don't just start services — verify readiness using `uri: url=http://localhost/healthz status_code=200 until="res.status == 200" retries=10 delay=2`.
Task Blocks error control: block (primary tasks) → rescue (rollback on failure) → always (cleanup execution).

Task Blocks group related tasks together and function identically to try-catch-finally exception handling blocks in programming languages:

  1. block:: Defines the primary operational tasks to execute.
  2. rescue:: Executes ONLY if any task inside block fails. Ideal for triggering automated database rollbacks, restoring backup configs, or firing PagerDuty/Slack alerts.
  3. always:: Executes ALWAYS, regardless of whether block succeeded or rescue ran. Ideal for deleting temporary download archives or closing diagnostic sessions.
---
- name: Production Web Application Upgrade with Automated Rollback
  hosts: webservers
  become: true

  tasks:
    - name: Transactional Upgrade Block
      block:
        - name: Download application release package
          ansible.builtin.get_url:
            url: https://releases.acme.com/app-v2.4.tar.gz
            dest: /tmp/app-v2.4.tar.gz
            timeout: 10

        - name: Unarchive release payload to production directory
          ansible.builtin.unarchive:
            src: /tmp/app-v2.4.tar.gz
            dest: /opt/myapp/
            remote_src: yes

        - name: Verify application service started cleanly
          ansible.builtin.service:
            name: myapp
            state: started

      rescue:
        - name: LOG CRITICAL ERROR: Deployment Failed!
          ansible.builtin.debug:
            msg: "CRITICAL: Upgrade failed! Initiating automated rollback to previous release."

        - name: Restore previous release directory from backup
          ansible.builtin.copy:
            src: /opt/myapp_backup/
            dest: /opt/myapp/
            remote_src: yes

        - name: Restart application service on previous release
          ansible.builtin.service:
            name: myapp
            state: restarted

      always:
        - name: Clean up temporary download artifacts
          ansible.builtin.file:
            path: /tmp/app-v2.4.tar.gz
            state: absent

Topic 2: Error Control Flags (ignore_errors, failed_when, changed_when)

By default, if any task returns a non-zero exit code (rc != 0), Ansible halts execution for that host. You can customize this behavior:

1. ignore_errors: true

Continues playbook execution even if this specific task fails (useful for non-critical checks):

- name: Attempt to query optional legacy metric agent
  ansible.builtin.command: /opt/legacy_agent/status.sh
  ignore_errors: true              # Playbook continues even if script returns exit code 1

2. Custom Failure Criteria (failed_when:)

Overrides default exit code checks to fail based on specific output patterns:

- name: Check application log for fatal error strings
  ansible.builtin.command: cat /var/log/myapp/app.log
  register: log_output
  failed_when:
    - "'FATAL' in log_output.stdout"
    - "'Database connection lost' in log_output.stdout"

3. Overriding Changed State (changed_when:)

Prevents read-only commands or verification tasks from reporting a changed state (which would falsely trigger Handlers):

- name: Inspect disk space via shell
  ansible.builtin.command: df -h /var
  register: disk_info
  changed_when: false              # Read-only task; never reports changed state!

Topic 3: Smoke Testing & Service Health Verification

Don’t just start a service and assume it is working — write a Smoke Test using the ansible.builtin.uri module to verify real HTTP responses:

- name: Ensure Nginx web service is started
  ansible.builtin.service:
    name: nginx
    state: started

- name: Smoke Test — Verify Web Application returns HTTP 200 OK
  ansible.builtin.uri:
    url: http://localhost/healthz
    status_code: 200
    return_content: yes
  register: smoke_result
  until: "'HEALTHY' in smoke_result.content"
  retries: 10
  delay: 2                         # Retry up to 10 times with 2-second delay

Topic 4: Dry-Run Testing (--check & --diff) and CLI Troubleshooting Flags

Before applying playbooks to production hosts, validate the impact using dry-run flags:

CLI SwitchDescriptionUse Case
ansible-playbook --checkRuns playbook in Dry-Run Mode. Predicts changes without writing to target nodes.Pre-deployment sanity check
ansible-playbook --diffDisplays exact line-by-line configuration diffs (like git diff) for templates and files.Audit exact text modifications
ansible-playbook --syntax-checkParses YAML syntax and verifies module names without connecting over SSH.CI/CD pull request linter
ansible-playbook -vvvvMaximum verbose output displaying raw SSH commands, connection parameters, and JSON payloads.Deep SSH connection troubleshooting
ansible-playbook --start-at-task="Task Name"Resumes playbook execution at a specific task.Fast debugging on long playbooks
# Production Standard Execution: Check mode with line-by-line diff output
ansible-playbook -i inventories/production/hosts.ini site.yml --check --diff

Common mistake: Using ignore_errors: yes to get a play to green. It hides genuine failures and the play reports success while nothing was configured. failed_when with a real condition says what you actually mean, and block/rescue handles the cases where recovery is the point.