Automating VMware snapshots with Ansible playbooks

Managing snapshots across an estate of ESXi hosts used to be a manual chore carried out by a senior admin on the night shift. Today, with hybrid workloads stretching from Sydney CBD data centres to branch offices in Parramatta and Geelong, that approach no longer scales. Treating snapshots as a configuration artefact fits the way Australian IT teams already work with infrastructure-as-code. Bringing Ansible into the workflow turns ad-hoc right-click actions into version-controlled processes that align with the Essential Eight maturity model and APRA CPS 234 obligations.

A playbook-driven approach also dovetails with the way many local MSPs and enterprise teams run their quarterly change windows. Whether you are a sole operator in Hobart supporting a small retail chain or part of a larger crew in Brisbane managing hundreds of VMs, the same YAML files can enforce consistent snapshot behaviour, reduce storage sprawl, and keep audit trails tidy for the Notifiable Data Breaches scheme overseen by the Office of the Australian Information Commissioner.

Why snapshots still need governance

Snapshots are not backups, and the longer they live, the more risk they introduce to a vSphere environment. Delta disks grow, performance degrades, and consolidation failures become harder to clear as workloads age. In Australian financial services, where APRA CPS 234 demands demonstrable resilience controls, an unmanaged snapshot on a tier-one system can be the kind of weakness auditors flag in a s177 review.

The good news is that most painful snapshot scenarios are predictable. VMs that host databases, mail servers, or line-of-business applications tend to be the worst offenders because patching cycles attach and forget about them. A scheduled playbook that checks snapshot age, parent disk size, and backup job status can replace the human reminder that often slips through the cracks during the end-of-financial-year freeze in late June.

Preparing the Ansible control node and vCenter

Before writing a single task, the control node needs the right collections installed. Teams already familiar with the broader automation stack can lean on the Automating vSphere deployments with Ansible guide to lay down the foundation, including the community.vmware collection and the Python pyvmomi bindings.

Connection details belong in an encrypted vault file rather than plain group_vars. Many Australian organisations route their Ansible traffic through the AWS Sydney or Melbourne regions for low-latency API calls, while others keep it on-prem behind a privileged access workstation in the same datacentre as vCenter. Either way, use a dedicated service account with the Snapshot management privilege at the datastore or VM level, and rotate the password on the same cadence as any other privileged credential under your Cyber Security Strategy.

Crafting the snapshot playbook

The playbook itself reads cleanly when split into three plays: create, list, and remove. A minimal create play might pin a host pattern such as vcenter:children, reference a variable file with the VM name and snapshot name, and call the community.vmware.vmware_snapshot module with state: present. A descriptive snapshot_name like pre-patch-2024-08 pays off later when someone greps through reports trying to reconcile change tickets.

Retention belongs in the variable file too, not buried inside tasks. Pull the threshold from a group variable such as snapshot_max_age_days: 3, then compare it against the snapshot creation timestamp returned by a vmware_guest_snapshot_info lookup. This makes it easy to align with the internal change windows enforced across Australian government departments, where a typical patching window runs from Friday evening to Monday morning AEST.

When consolidation is needed, the same playbook can call state: remove with remove_children: yes and a snapshot_name filter so you do not accidentally blow away the wrong delta. Teams that already have PowerCLI scripts can apply the Creating custom PowerCLI scripts for playbook approach as a transitional step rather than throwing away years of working code.

Retention policies and reporting

Snapshot policies deserve their own vars file, especially when different business units have different recovery expectations. A retail tenant in Adelaide might be comfortable with a 24-hour snapshot, while a healthcare provider bound by the My Health Records Act needs stricter control over what data ever lands on a delta disk.

Variables that anchor the retention policy

Pair the playbook with a daily scheduled run and a Mailgun or SendGrid integration pointed at a local mailbox. Many shops in Melbourne and Sydney push snapshot warnings into the same Teams space where change tickets already live, creating a feedback loop the help desk can actually use rather than a dashboard no one opens.

Testing, troubleshooting, and PowerCLI fallback

Even with the cleanest playbook, ESXi connection issues can derail a run mid-way. The Troubleshooting common ESXi host connection guide covers the usual suspects that show up in Australian environments.

Pitfalls to anticipate before the first run

Treating these as known failure modes and writing idempotent tasks to handle them avoids the dreaded half-merged snapshot support ticket.

There are still scenarios where PowerCLI is the right hammer, especially when interacting with SRM or vSphere Replication. Calling a PowerCLI .ps1 from an Ansible task using win_powershell or a local command keeps the workflow unified. This hybrid pattern is common in Australian enterprises that have invested in PowerCLI and do not want to throw away working code while migrating toward pure-play automation.

The cleanest implementations treat PowerCLI as a fallback layer, invoked only when the Ansible collections do not expose a particular API. Over time, those PowerCLI scripts shrink, the YAML grows, and the next junior engineer joining the team in Perth or Canberra inherits a workflow that actually makes sense on a Monday morning.

Start small by automating one folder of non-critical VMs, watch the run logs for a fortnight, then widen the scope once the playbook proves itself. Once the daily digest lands in the team's inbox without anyone having to chase a stuck snapshot, the case for extending the same pattern to backups, patches, and capacity reports writes itself.