Automating ESXi patching with Ansible and VMware Update Manager
Anyone who has worn a pager for a Brisbane-based NOC knows the 02:00 alarm. A cluster needs patching, the vCenter is humming in a Sydney colocation, and there are eighteen hosts waiting for a reboot. Clicking through the vSphere client by hand at that hour is a fast track to mistakes, and a quiet way to burn out the one person who keeps answering the phone.
Tying Ansible to VMware Update Manager turns a tedious weekend ritual into something reproducible. You describe the desired state, the cluster converges, and you keep an audit trail without writing a Word doc. The same playbook that runs in a Melbourne office lab can be promoted to production with a single inventory swap.
Why manual ESXi patching breaks down
Australian IT teams often look after a spread-out estate. A retailer might run core in Sydney, a DR node in Adelaide, and small hyperconverged clusters in regional stores from Cairns to Geelong. The distances are not the only problem. Different time zones, limited onsite hands, and strict change windows at universities and government departments all make manual patching brittle.
Drift creeps in fast. One host misses a baseline because someone applied a hotfix manually, another has a stale driver, and the next maintenance window becomes a forensic exercise. Automation gives a single source of truth and a record auditors in healthcare, finance, and utilities will actually accept.
Mapping the environment with Ansible
Before any patching runs, Ansible needs to know what is there. A dynamic inventory script pointed at vCenter pulls the live list of ESXi hosts, their build numbers, and their clusters. Tagging hosts by site or business unit at this stage pays dividends later when you filter who gets what baseline.
For a typical mid-sized shop, grouping hosts by purpose - production, DR, dev, ROBO - is usually enough. The play can iterate cluster by cluster, and you can throttle concurrency so a single vCenter does not get hammered during business hours. If you are running a few clusters through the same controller out of a Perth office, that throttling is not optional.
Setting up VUM baselines the right way
VMware Update Manager still does the heavy lifting when it comes to staging the actual ESXi image and the vib packages. Baselines should be lean and named for purpose rather than vendor release notes. Think "Prod-Critical-Fixes", "GPU-Hosts-Image", or "ROBO-Cumulative" instead of "Baseline-July".
Attach baselines to clusters or folders, not individual hosts, unless you have a good reason. Folder-level attachment scales better when you onboard a new host and forget the manual step. Staging the patches ahead of time means the run against a host is just an apply-and-reboot call from the playbook.
It is worth pinning firmware and driver baselines for hosts running vSAN or NSX, since mismatches there cause more grief than a missed CPU microcode patch. The virtxpert disclaimer is a fair place to remind the team that baselines should match the HCL, not the latest vendor email.
Writing the playbook for the update run
The playbook is where the orchestration lives. A typical flow enters maintenance mode, evicts powered VMs, runs the remediate task, waits for the reboot, then exits maintenance mode. Both maintenance mode and remediate are first-class operations when you call into VUM from an Ansible task.
Notifications belong in the playbook too. A Slack or email step before and after the window stops people from "just rebooting it real quick" when something looks wedged. A shared chat channel works for your team and any managed service provider in the loop.
Variables keep the playbook flexible. Image profile name, target cluster, drain timeout, and reboot poll interval should sit in a group_vars file so the playbook copies between environments without rewriting tasks. A pre-flight block at the top of the play is also worth its weight.
Pre-flight checks before triggering a run:
- Confirm vCenter and VUM are reachable.
- Verify the baseline is attached and reports current.
- Snapshot host configuration or export a host profile.
- Check free capacity for VM evacuation.
A short checklist catches the dumb failures that wake you up on a Sunday, and it lives in version control next to the playbook itself.
Dealing with reboots, maintenance windows, and stubborn hosts
Reboots are where most homegrown automation dies. A host comes back up, but iDRAC or IPMI is slow on the first boot, and the playbook declares failure before the management agents are responsive. A retry loop with a ten to fifteen minute delay - longer for a remote site in Kalgoorlie - saves a lot of false alarms.
Maintenance windows in Australia often align with after-hours AEST, even when the host is physically in Perth. Make the playbook time-aware using the controller's clock and a TZ variable, then log start and end timestamps. That log becomes evidence the next morning when someone asks whether the change actually happened, and feeds straight into a change advisory board report.
Recovering when a patch goes sideways
VUM will roll a host back to its previous image if remediation fails, but only if you have ticked the rollback option on the baseline. It costs nothing and saves a 03:00 trip to a regional data centre when a NIC driver does not come back up.
Beyond the VUM rollback, your playbook should snapshot host configuration before the run, ideally with a host profile export. If a host emerges with a misbehaving PCI device, compare against the pre-change snapshot and decide whether to remediate, rebuild, or escalate.
Recovery steps to script into the playbook:
- Reattach the host to the cluster and re-check HA response.
- Validate VMFS datastores and vSAN health post-reboot.
- Rerun compliance and confirm the new baseline is reporting clean.
- Notify stakeholders with the actual outcome, not the planned one.
Beyond one-off runs: pipelines and reporting
A scheduled run via AWX or a GitHub Actions cron is the natural next step. Pull the latest group_vars, run the playbook against a staging cluster first, then promote to production. The same code path - tagged in Git, reviewed in a pull request - is what the rest of your infrastructure automation already does, so patching stops being a special case.
Reporting falls out of the same loop. AWX job templates log every host, baseline, reboot, and exit code into a database that a Grafana dashboard or CSV export can chew on. The first time a compliance officer in Sydney asks for proof of last quarter's patching, the answer is a saved query rather than a frantic inbox search.
If you want a deeper walkthrough of the surrounding automation patterns, the virtxpert VMware series covers the surrounding tooling in detail. Drop your playbook in a Git repo, run it on a single dev cluster this week, and iterate from there. Patching does not have to be the chore everyone dreads.