Ansible Patterns - Production Best Practices
Status: Active
Last Updated: 2026-01-30
Category: Infrastructure - Configuration Management
Prerequisites: ansible-roles, ansible-vault
Time: 3-4 hours
Tags: ansible, patterns, best-practices, production, deployment
Summary
Master production-ready Ansible patterns for reliable, maintainable infrastructure automation. Learn rolling deployments, blue-green patterns, immutable infrastructure, testing strategies, and organizational best practices used by successful engineering teams.
๐ฏ What You'll Learn
By the end of this article, you'll be able to:
- โ Implement rolling deployments
- โ Execute blue-green deployments
- โ Build immutable infrastructure patterns
- โ Organize large Ansible projects
- โ Test your automation
- โ Implement proper error handling
- โ Optimize playbook performance
๐ Rolling Deployments
Basic Rolling Update
---
- name: Rolling update web servers
hosts: webservers
serial: 2 # Update 2 servers at a time
pre_tasks:
- name: Remove from load balancer
haproxy:
state: disabled
host: "{{ inventory_hostname }}"
backend: web_backend
delegate_to: "{{ item }}"
loop: "{{ groups['loadbalancers'] }}"
tasks:
- name: Pull latest code
git:
repo: https://github.com/company/app.git
dest: /opt/app
version: "{{ app_version }}"
- name: Install dependencies
pip:
requirements: /opt/app/requirements.txt
- name: Restart application
systemd:
name: myapp
state: restarted
- name: Wait for app to be ready
uri:
url: "http://localhost:8080/health"
status_code: 200
register: result
until: result.status == 200
retries: 30
delay: 2
post_tasks:
- name: Add back to load balancer
haproxy:
state: enabled
host: "{{ inventory_hostname }}"
backend: web_backend
delegate_to: "{{ item }}"
loop: "{{ groups['loadbalancers'] }}"
Serial with Percentage
---
- name: Rolling update (25% at a time)
hosts: webservers
serial: "25%" # Update quarter of fleet at once
tasks:
- name: Deploy application
include_role:
name: app_deploy
Serial with Max Failure
---
- name: Rolling update with failure threshold
hosts: webservers
serial: 3
max_fail_percentage: 20 # Stop if >20% fail
tasks:
- name: Deploy
include_tasks: deploy.yaml
๐ต๐ข Blue-Green Deployments
Blue-Green Pattern
Inventory:
# inventories/production/hosts.yaml
all:
children:
webservers_blue:
hosts:
web1:
ansible_host: 10.0.1.10
app_version: v1.0.0
web2:
ansible_host: 10.0.1.11
app_version: v1.0.0
webservers_green:
hosts:
web3:
ansible_host: 10.0.1.20
app_version: v1.1.0
web4:
ansible_host: 10.0.1.21
app_version: v1.1.0
loadbalancers:
hosts:
lb1:
ansible_host: 10.0.1.100
Playbook:
---
- name: Deploy to green environment
hosts: webservers_green
tasks:
- name: Deploy new version
include_role:
name: app_deploy
vars:
app_version: "{{ new_version }}"
- name: Run smoke tests
uri:
url: "http://{{ ansible_host }}:8080/health"
status_code: 200
register: health_check
failed_when: health_check.status != 200
- name: Switch traffic to green
hosts: loadbalancers
tasks:
- name: Update load balancer config
template:
src: haproxy.cfg.j2
dest: /etc/haproxy/haproxy.cfg
vars:
active_backend: green
notify: Reload HAProxy
- name: Wait for traffic switch
pause:
seconds: 10
handlers:
- name: Reload HAProxy
systemd:
name: haproxy
state: reloaded
- name: Verify green is healthy
hosts: webservers_green
tasks:
- name: Check application logs
command: tail -n 100 /var/log/app/error.log
register: logs
failed_when: "'ERROR' in logs.stdout"
- name: Check error rate
uri:
url: "http://{{ ansible_host }}:8080/metrics"
register: metrics
failed_when: metrics.json.error_rate > 0.01
- name: Decommission blue (optional)
hosts: webservers_blue
tasks:
- name: Stop old version
systemd:
name: myapp
state: stopped
- name: Keep blue for rollback
debug:
msg: "Blue environment kept for 24h rollback window"
Rollback Playbook
---
- name: Rollback to blue
hosts: loadbalancers
tasks:
- name: Switch traffic back to blue
template:
src: haproxy.cfg.j2
dest: /etc/haproxy/haproxy.cfg
vars:
active_backend: blue
notify: Reload HAProxy
handlers:
- name: Reload HAProxy
systemd:
name: haproxy
state: reloaded
- name: Restart blue servers
hosts: webservers_blue
tasks:
- name: Ensure blue is running
systemd:
name: myapp
state: started
๐ฆ Immutable Infrastructure
Build AMI/Image Pattern
---
- name: Build application image
hosts: localhost
connection: local
tasks:
- name: Launch temporary instance
ec2_instance:
name: "image-builder-{{ ansible_date_time.epoch }}"
image_id: ami-ubuntu-22.04
instance_type: t3.medium
wait: yes
register: build_instance
- name: Add to inventory
add_host:
name: "{{ build_instance.instances[0].public_ip }}"
groups: image_builders
ansible_ssh_private_key_file: ~/.ssh/aws.pem
- name: Configure image
hosts: image_builders
become: yes
roles:
- common
- app_install
- monitoring
post_tasks:
- name: Clean up
command: cloud-init clean
- name: Remove SSH keys
file:
path: /home/ubuntu/.ssh/authorized_keys
state: absent
- name: Create AMI
hosts: localhost
connection: local
tasks:
- name: Create image
ec2_ami:
instance_id: "{{ build_instance.instances[0].instance_id }}"
name: "myapp-{{ app_version }}-{{ ansible_date_time.epoch }}"
wait: yes
register: new_ami
- name: Tag AMI
ec2_tag:
resource: "{{ new_ami.image_id }}"
tags:
Name: "myapp-{{ app_version }}"
Version: "{{ app_version }}"
BuildDate: "{{ ansible_date_time.iso8601 }}"
- name: Terminate build instance
ec2_instance:
instance_ids: "{{ build_instance.instances[0].instance_id }}"
state: absent
- name: Deploy new instances
hosts: localhost
connection: local
tasks:
- name: Update Auto Scaling Group
ec2_asg:
name: myapp-asg
launch_config_name: "myapp-lc-{{ app_version }}"
min_size: 3
max_size: 10
desired_capacity: 3
health_check_type: ELB
health_check_period: 300
replace_all_instances: yes
wait_for_instances: yes
๐ Project Organization
Large Project Structure
ansible-project/
โโโ ansible.cfg
โโโ requirements.yaml # Galaxy roles
โโโ .gitignore
โโโ .ansible-lint
โ
โโโ inventories/
โ โโโ production/
โ โ โโโ hosts.yaml
โ โ โโโ group_vars/
โ โ โ โโโ all/
โ โ โ โ โโโ vars.yaml
โ โ โ โ โโโ vault.yaml
โ โ โ โโโ webservers.yaml
โ โ โ โโโ databases.yaml
โ โ โโโ host_vars/
โ โ โโโ web1.yaml
โ โโโ staging/
โ โ โโโ ...
โ โโโ development/
โ โโโ ...
โ
โโโ playbooks/
โ โโโ site.yaml # Master playbook
โ โโโ webservers.yaml # Web server playbook
โ โโโ databases.yaml # Database playbook
โ โโโ deploy.yaml # Deployment playbook
โ โโโ rollback.yaml # Rollback playbook
โ
โโโ roles/
โ โโโ common/ # Base configuration
โ โโโ nginx/ # Web server
โ โโโ postgresql/ # Database
โ โโโ monitoring/ # Monitoring agents
โ โโโ security/ # Security hardening
โ
โโโ group_vars/ # Shared group vars
โ โโโ all.yaml
โ
โโโ library/ # Custom modules
โ โโโ my_custom_module.py
โ
โโโ filter_plugins/ # Custom filters
โ โโโ my_filters.py
โ
โโโ tasks/ # Reusable task files
โ โโโ ssl_setup.yaml
โ โโโ backup.yaml
โ
โโโ templates/ # Shared templates
โ โโโ maintenance.html.j2
โ
โโโ files/ # Shared files
โ โโโ company_ca.crt
โ
โโโ scripts/ # Helper scripts
โโโ deploy.sh
โโโ vault-pass.sh
Master Playbook Pattern
playbooks/site.yaml:
---
# Master playbook - orchestrates everything
- import_playbook: common.yaml
- import_playbook: security.yaml
- import_playbook: monitoring.yaml
- import_playbook: webservers.yaml
- import_playbook: databases.yaml
- import_playbook: cache.yaml
Run everything:
ansible-playbook playbooks/site.yaml -i inventories/production
Run specific parts:
ansible-playbook playbooks/webservers.yaml -i inventories/production
๐งช Testing Strategies
Pre-Flight Checks
---
- name: Pre-flight checks
hosts: all
gather_facts: yes
tasks:
- name: Verify connectivity
ping:
- name: Check disk space
assert:
that:
- ansible_mounts | selectattr('mount', 'equalto', '/') | map(attribute='size_available') | first > 1073741824
fail_msg: "Less than 1GB free space on /"
- name: Check memory
assert:
that:
- ansible_memfree_mb > 512
fail_msg: "Less than 512MB free memory"
- name: Verify required ports available
wait_for:
port: "{{ item }}"
state: stopped
timeout: 1
loop:
- 80
- 443
when: "'webservers' in group_names"
ignore_errors: yes
register: port_check
- name: Fail if ports in use
fail:
msg: "Port {{ item.item }} already in use"
when: item.failed is defined and not item.failed
loop: "{{ port_check.results }}"
Smoke Tests
---
- name: Smoke tests
hosts: webservers
tasks:
- name: Wait for application
wait_for:
port: 8080
timeout: 60
- name: Check health endpoint
uri:
url: "http://localhost:8080/health"
status_code: 200
register: health
retries: 10
delay: 5
until: health.status == 200
- name: Verify database connection
uri:
url: "http://localhost:8080/health/db"
status_code: 200
- name: Check application version
uri:
url: "http://localhost:8080/version"
register: version
failed_when: version.json.version != app_version
Integration Tests
---
- name: Integration tests
hosts: localhost
connection: local
tasks:
- name: Create test user
uri:
url: "https://api.example.com/users"
method: POST
body_format: json
body:
username: "test_{{ ansible_date_time.epoch }}"
email: "test@example.com"
status_code: 201
register: test_user
- name: Login as test user
uri:
url: "https://api.example.com/login"
method: POST
body_format: json
body:
username: "{{ test_user.json.username }}"
password: "testpass"
status_code: 200
register: login
- name: Create test data
uri:
url: "https://api.example.com/data"
method: POST
headers:
Authorization: "Bearer {{ login.json.token }}"
body_format: json
body:
name: "Test Item"
status_code: 201
register: test_data
- name: Retrieve test data
uri:
url: "https://api.example.com/data/{{ test_data.json.id }}"
headers:
Authorization: "Bearer {{ login.json.token }}"
status_code: 200
register: retrieved_data
failed_when: retrieved_data.json.name != "Test Item"
- name: Cleanup test data
uri:
url: "https://api.example.com/data/{{ test_data.json.id }}"
method: DELETE
headers:
Authorization: "Bearer {{ login.json.token }}"
status_code: 204
๐ฏ Error Handling Patterns
Graceful Degradation
---
- name: Deploy with fallback
hosts: webservers
tasks:
- name: Try to pull from primary registry
docker_image:
name: "registry1.example.com/myapp:{{ version }}"
source: pull
register: primary_pull
ignore_errors: yes
- name: Fall back to secondary registry
docker_image:
name: "registry2.example.com/myapp:{{ version }}"
source: pull
when: primary_pull is failed
register: secondary_pull
ignore_errors: yes
- name: Use cached image as last resort
docker_image:
name: "myapp:latest"
source: pull
when:
- primary_pull is failed
- secondary_pull is failed
Retry Logic
---
- name: Deploy with retries
hosts: webservers
tasks:
- name: Download artifact
get_url:
url: "https://releases.example.com/app-{{ version }}.tar.gz"
dest: "/tmp/app-{{ version }}.tar.gz"
register: download
retries: 5
delay: 10
until: download is succeeded
- name: Extract with verification
unarchive:
src: "/tmp/app-{{ version }}.tar.gz"
dest: /opt/app
remote_src: yes
register: extract
retries: 3
delay: 5
until: extract is succeeded
Atomic Operations
---
- name: Atomic deployment
hosts: webservers
tasks:
- name: Create temporary directory
tempfile:
state: directory
suffix: deploy
register: temp_dir
- block:
- name: Deploy to temporary location
synchronize:
src: /local/app/
dest: "{{ temp_dir.path }}/"
- name: Verify deployment
command: "{{ temp_dir.path }}/bin/verify.sh"
- name: Atomic switch
command: "mv {{ temp_dir.path }} /opt/app-new && mv /opt/app /opt/app-old && mv /opt/app-new /opt/app"
args:
removes: "{{ temp_dir.path }}"
rescue:
- name: Cleanup on failure
file:
path: "{{ temp_dir.path }}"
state: absent
- name: Rollback if needed
command: "mv /opt/app-old /opt/app"
when: app_old_exists
always:
- name: Remove old version
file:
path: /opt/app-old
state: absent
ignore_errors: yes
โก Performance Optimization
Parallel Execution
---
- name: Fast deployment
hosts: webservers
strategy: free # Don't wait for all hosts
tasks:
- name: Download (runs immediately when host is ready)
get_url:
url: "{{ artifact_url }}"
dest: /tmp/artifact.tar.gz
Fact Caching
ansible.cfg:
[defaults]
gathering = smart
fact_caching = jsonfile
fact_caching_connection = /tmp/ansible_facts
fact_caching_timeout = 3600
# Or use Redis
# fact_caching = redis
# fact_caching_connection = localhost:6379:0
Minimize Fact Gathering
---
- name: Quick tasks
hosts: all
gather_facts: no # Skip if not needed
tasks:
- name: Simple command
command: echo "Hello"
Pipeline Optimization
ansible.cfg:
[ssh_connection]
pipelining = True
ssh_args = -o ControlMaster=auto -o ControlPersist=60s
๐ Deployment Checklist
Pre-Deployment
---
- name: Pre-deployment checklist
hosts: localhost
connection: local
tasks:
- name: Verify version tag exists
uri:
url: "https://github.com/company/app/releases/tag/{{ version }}"
status_code: 200
- name: Check artifact is built
uri:
url: "https://releases.example.com/app-{{ version }}.tar.gz"
method: HEAD
status_code: 200
- name: Verify database migrations
command: "git diff {{ current_version }}..{{ version }} -- migrations/"
register: migration_check
changed_when: false
- name: Alert if migrations exist
debug:
msg: "WARNING: Database migrations detected!"
when: migration_check.stdout != ""
- name: Check production load
uri:
url: "https://monitoring.example.com/api/v1/query?query=rate(http_requests_total[5m])"
register: load_check
- name: Warn if high load
fail:
msg: "High load detected. Consider deploying off-peak."
when: load_check.json.data.result[0].value[1] | float > 1000
ignore_errors: yes
๐ What's Next?
Infrastructure as Code:
- terraform-basics - Provision infrastructure
- terraform-providers - Cloud providers
Testing:
- infrastructure-testing - Test your automation
GitOps:
- gitops-principles - Git as source of truth
๐ Resources
Books:
- "Ansible for DevOps" by Jeff Geerling
- "Ansible: Up and Running" by Lorin Hochstein
Best Practices:
Testing:
๐ Change Log
2026-01-30
- Created production patterns guide
- Covered rolling deployments
- Demonstrated blue-green deployments
- Included immutable infrastructure
- Provided project organization
- Added testing strategies
- Showed error handling patterns
- Included performance optimization
- Provided deployment checklist
Next Article: terraform-basics - Infrastructure as Code!