In production automation, unhandled errors leave servers in inconsistent, half-configured states.
Ansible provides Task Blocks (block, rescue, always), error override flags (failed_when, changed_when), and dry-run execution modes (--check, --diff) to build fault-tolerant playbooks with automated rollbacks and pre-deployment validations.
Topic 1: Defensive Task Blocks (block, rescue, always)
Task Blocks group related tasks together and function identically to try-catch-finally exception handling blocks in programming languages:
block:: Defines the primary operational tasks to execute.rescue:: Executes ONLY if any task insideblockfails. Ideal for triggering automated database rollbacks, restoring backup configs, or firing PagerDuty/Slack alerts.always:: Executes ALWAYS, regardless of whetherblocksucceeded orrescueran. Ideal for deleting temporary download archives or closing diagnostic sessions.
---
- name: Production Web Application Upgrade with Automated Rollback
hosts: webservers
become: true
tasks:
- name: Transactional Upgrade Block
block:
- name: Download application release package
ansible.builtin.get_url:
url: https://releases.acme.com/app-v2.4.tar.gz
dest: /tmp/app-v2.4.tar.gz
timeout: 10
- name: Unarchive release payload to production directory
ansible.builtin.unarchive:
src: /tmp/app-v2.4.tar.gz
dest: /opt/myapp/
remote_src: yes
- name: Verify application service started cleanly
ansible.builtin.service:
name: myapp
state: started
rescue:
- name: LOG CRITICAL ERROR: Deployment Failed!
ansible.builtin.debug:
msg: "CRITICAL: Upgrade failed! Initiating automated rollback to previous release."
- name: Restore previous release directory from backup
ansible.builtin.copy:
src: /opt/myapp_backup/
dest: /opt/myapp/
remote_src: yes
- name: Restart application service on previous release
ansible.builtin.service:
name: myapp
state: restarted
always:
- name: Clean up temporary download artifacts
ansible.builtin.file:
path: /tmp/app-v2.4.tar.gz
state: absent
Topic 2: Error Control Flags (ignore_errors, failed_when, changed_when)
By default, if any task returns a non-zero exit code (rc != 0), Ansible halts execution for that host. You can customize this behavior:
1. ignore_errors: true
Continues playbook execution even if this specific task fails (useful for non-critical checks):
- name: Attempt to query optional legacy metric agent
ansible.builtin.command: /opt/legacy_agent/status.sh
ignore_errors: true # Playbook continues even if script returns exit code 1
2. Custom Failure Criteria (failed_when:)
Overrides default exit code checks to fail based on specific output patterns:
- name: Check application log for fatal error strings
ansible.builtin.command: cat /var/log/myapp/app.log
register: log_output
failed_when:
- "'FATAL' in log_output.stdout"
- "'Database connection lost' in log_output.stdout"
3. Overriding Changed State (changed_when:)
Prevents read-only commands or verification tasks from reporting a changed state (which would falsely trigger Handlers):
- name: Inspect disk space via shell
ansible.builtin.command: df -h /var
register: disk_info
changed_when: false # Read-only task; never reports changed state!
Topic 3: Smoke Testing & Service Health Verification
Don’t just start a service and assume it is working — write a Smoke Test using the ansible.builtin.uri module to verify real HTTP responses:
- name: Ensure Nginx web service is started
ansible.builtin.service:
name: nginx
state: started
- name: Smoke Test — Verify Web Application returns HTTP 200 OK
ansible.builtin.uri:
url: http://localhost/healthz
status_code: 200
return_content: yes
register: smoke_result
until: "'HEALTHY' in smoke_result.content"
retries: 10
delay: 2 # Retry up to 10 times with 2-second delay
Topic 4: Dry-Run Testing (--check & --diff) and CLI Troubleshooting Flags
Before applying playbooks to production hosts, validate the impact using dry-run flags:
| CLI Switch | Description | Use Case |
|---|---|---|
ansible-playbook --check | Runs playbook in Dry-Run Mode. Predicts changes without writing to target nodes. | Pre-deployment sanity check |
ansible-playbook --diff | Displays exact line-by-line configuration diffs (like git diff) for templates and files. | Audit exact text modifications |
ansible-playbook --syntax-check | Parses YAML syntax and verifies module names without connecting over SSH. | CI/CD pull request linter |
ansible-playbook -vvvv | Maximum verbose output displaying raw SSH commands, connection parameters, and JSON payloads. | Deep SSH connection troubleshooting |
ansible-playbook --start-at-task="Task Name" | Resumes playbook execution at a specific task. | Fast debugging on long playbooks |
# Production Standard Execution: Check mode with line-by-line diff output
ansible-playbook -i inventories/production/hosts.ini site.yml --check --diff
Common mistake: Using ignore_errors: yes to get a play to green. It hides genuine failures and the play reports success while nothing was configured. failed_when with a real condition says what you actually mean, and block/rescue handles the cases where recovery is the point.