Troubleshooting VMware HA and DRS Failures

VMware vSphere High Availability (HA) and Distributed Resource Scheduler (DRS) are designed to make clusters resilient and efficient. HA restarts workloads after host failure, while DRS balances compute demand and supports intelligent VM placement. When either service reports an error, however, the cause may sit in networking, DNS, licensing, admission control, permissions, or an unhealthy host rather than in the feature itself.

A structured investigation reduces disruption and prevents repeated remediation. This guide covers common HA configuration errors, DRS migration failures, vCenter alarms, isolation responses, and practical validation steps for Australian environments where maintenance windows, data sovereignty, and limited after-hours support can influence the response.

Check The Cluster Foundations

Start with vCenter Server, ESXi host connectivity, and the cluster configuration. Confirm every host is connected, responding, and running a supported ESXi build. Review vCenter Tasks and Events before changing settings, then check whether the failure affects one host, several hosts, or the entire cluster.

DNS and time synchronisation are frequent causes of misleading HA errors. Forward and reverse DNS records should resolve consistently from vCenter and every ESXi host. Verify NTP configuration and compare host times, especially across geographically separated sites such as Sydney and Perth. A clock drift can disrupt certificates, host communication, and management operations.

Review the cluster’s licensing status and enabled features as well. DRS requires the appropriate vSphere edition, and some automated placement or migration capabilities depend on licensing and VM compatibility. For a broader configuration baseline, compare the environment with these vSphere best practices before treating an alarm as an isolated fault.

Investigate HA Agent And Network Health

VMware HA relies on the Fault Domain Manager, commonly referred to as FDM, running on ESXi hosts. If an agent fails to initialise, examine the host’s management network, firewall rules, certificates, and communication with the HA master. In vSphere Client, reconfigure HA on the affected host only after recording the current alarms and recent events.

Management network redundancy is important, but multiple paths can create confusion when VLANs or MTU settings are inconsistent. Test host-to-host connectivity on the management vmkernel interface and confirm that the same VLAN configuration reaches every host. Jumbo frames should be enabled end to end, or disabled consistently; a partial MTU configuration can make probes fail intermittently.

Also inspect physical switch ports, port-channel settings, and security policies. A change made by a network team in a Melbourne or Brisbane data centre may coincide with an HA failure without appearing in vCenter. If hosts can reach vCenter but cannot reliably reach one another, investigate the switching layer rather than repeatedly restarting agents.

Understand Isolation And Admission Control

An HA host isolation response determines what happens when a host loses contact with other cluster members. Depending on the configuration, VMs may be powered off, shut down, or left running. The correct choice depends on storage accessibility, application clustering, and the risk of split-brain behaviour. Document the decision instead of accepting the default without review.

Admission control can also prevent a VM from powering on or entering maintenance mode. HA reserves capacity for host failures, so a cluster that appears to have free resources may still reject operations. Check the configured host failure policy, slot calculations, dedicated failover capacity, and reservations on critical VMs.

A common mistake is disabling admission control to force a change through. That may hide capacity risk and leave the cluster unable to restart workloads during an outage. In Australian colocation environments, where procurement of extra capacity may take time, validate the failure scenario against current host utilisation and contractual recovery objectives.

Diagnose DRS Placement And Migration Errors

DRS recommendations depend on CPU and memory demand, VM-Host rules, affinity settings, reservations, shares, and available compatible hosts. When DRS refuses to move a VM, inspect the recommendation reason and the VM’s constraints. A hard affinity rule, pinned host, datastore access issue, or incompatible virtual hardware may be the real blocker.

Storage and networking are equally important. vMotion requires compatible CPU features, accessible VM files, working vmkernel adapters, and sufficient bandwidth. Check EVC settings when hosts have different processor generations, particularly after adding newer hardware to an established cluster. Validate vMotion network routing and confirm that the destination host can access every required datastore.

DRS automation should be increased gradually. Begin with manual or partially automated recommendations, confirm successful migrations, then enable more aggressive automation when the environment is stable. This approach is useful for busy Sydney production sites where an unexpected migration could affect latency-sensitive applications during local business hours.

Use Logs And Automation Carefully

The vSphere Client provides useful event history, but deeper diagnosis often requires ESXi and vCenter logs. Relevant files may include fdm.log for HA, hostd.log for host management, vpxa.log for vCenter communication, and vCenter service logs for task or inventory issues. Capture timestamps in Australian Eastern or local site time before correlating events across systems.

Avoid making several changes at once. Record the original state, apply one remediation, and test HA reconfiguration, vMotion, or VM power-on behaviour. If a host is repeatedly failing, evacuate workloads where safe and place it into maintenance mode rather than allowing automated recovery to mask a hardware, storage, or network fault.

Automation can provide repeatable checks for host state, cluster membership, datastores, and vmkernel configuration. For teams already using Git and Ansible, Ansible playbooks can standardise evidence collection and reduce manual drift, provided credentials, change control, and sensitive output are handled securely.

Apply A Repeatable Recovery Process

A reliable recovery process separates immediate service restoration from permanent correction. First protect workloads, identify the affected host or cluster service, and confirm whether HA restarts or DRS migrations are still operating. Then resolve the underlying fault, reconfigure only the necessary service, and monitor subsequent tasks and alarms.

Use the following operational recommendations during an incident:

The distinction between HA and DRS is useful when narrowing the fault:

Capability Primary purpose Typical failure clues First checks
HA Restart VMs after host failure FDM errors, isolation alarms, failed restarts Host communication, DNS, NTP, storage, admission control
DRS Balance workload and recommend or perform migrations No recommendation, vMotion failure, rule conflict VM-Host rules, EVC, datastore access, vMotion network
vMotion Move a running VM between hosts Timeout, compatibility, network or storage errors CPU compatibility, vmkernel adapters, MTU, shared storage

After the incident, retain a short post-incident record with the trigger, evidence, remediation, and validation results. Review host firmware, ESXi compatibility, switch changes, capacity trends, and monitoring coverage. In a regulated Australian environment, this record also supports change management, audit requirements, and data-centre operational reviews.

Use these checks as a runbook during the next HA or DRS alert, and adapt them to your cluster’s architecture, support agreements, and recovery objectives. Consistent evidence collection and controlled remediation will make VMware availability failures faster to diagnose and less likely to recur.