Designing a resilient vSAN cluster for small environments
Small organisations often need enterprise-grade availability without the budget, rack space or specialist staff of a large data centre. A compact VMware vSAN cluster can deliver that balance by combining local disks in several ESXi hosts into a shared, policy-driven datastore.
The design still needs careful planning. Two or three hosts may support production workloads, but a single failed disk, overloaded network link or poorly sized witness can reduce resilience quickly. High availability is a result of sound architecture, not simply enabling vSphere HA.
Australian conditions add practical considerations. A business in Sydney may have different connectivity and colocation options from one in regional Queensland, while Melbourne sites may need to plan for winter power and cooling costs. Local data residency, support availability and long hardware lead times can also influence the design.
The most effective approach is to define failure scenarios first, then select hardware, storage policies and operational controls that address them. This produces a cluster that remains usable during maintenance, component failure and short infrastructure outages.
Define the failure domain
Start by identifying what the cluster must survive. A resilient small environment should normally tolerate at least one host failure without losing access to critical virtual machines. It may also need to withstand a disk group failure, a network switch outage or planned maintenance on a single node.
A two-node vSAN cluster uses a dedicated witness appliance or host outside the data nodes. The witness supplies quorum and records metadata, but it does not store VM data. It should be placed in a separate failure domain, ideally in another room, building or site with independent power and network paths.
For a three-node cluster, vSAN can distribute components across the hosts without a traditional witness appliance. This is simpler operationally, although the third host increases capital cost. Avoid placing every component in the same rack if a rack-level power or cooling failure is a realistic risk.
Select hardware for predictable performance
Use vSAN-compatible hardware from the VMware Compatibility Guide, with supported controllers, firmware and drive models. Consumer SSDs can create inconsistent latency, premature wear and difficult support cases. Enterprise NVMe or SSD devices with power-loss protection are generally a safer choice for production workloads.
A typical small node includes a pair of cache or capacity devices, depending on the vSAN architecture and disk group design. All-flash configurations usually provide better performance and more consistent behaviour than hybrid arrangements, particularly for databases, VDI and automation platforms.
Memory and CPU capacity should include room for a host failure. If three hosts each run at 70–75 per cent utilisation, losing one node may cause excessive contention. Size the remaining hosts so essential workloads can restart without relying on aggressive memory overcommitment.
Build the network around failure tolerance
vSAN traffic is sensitive to latency, packet loss and congestion. Use a dedicated or clearly isolated network for vSAN, with redundant physical uplinks and correctly configured VLANs. A 10 GbE network is a practical baseline for many small clusters, although workload density and encryption requirements may justify faster links.
Separate vSAN, vMotion, management and virtual machine traffic logically. Physical separation is helpful where the budget allows, but correct QoS and switch configuration are more important than simply adding interfaces. Validate MTU settings end to end before enabling jumbo frames; a partial configuration can cause difficult intermittent faults.
Australian sites should also consider WAN dependence. A witness in a Melbourne office is not automatically suitable for a cluster in Brisbane if the inter-site connection has variable latency or limited service guarantees. Keep the witness path stable and use independent connectivity where a single NBN or carrier circuit represents a significant failure point.
Choose storage policies deliberately
vSAN storage policies determine the number of failures a virtual machine can survive. Failure to tolerate one failure (FTT=1) with mirroring is straightforward, but it consumes approximately twice the logical capacity, before overhead and free-space requirements are considered. RAID-5/6 erasure coding can improve usable capacity, though it has different performance and host requirements.
A small cluster should retain operational headroom rather than fill the datastore. Keep sufficient free capacity for rebuilds, resynchronisation and snapshots. Thin provisioning can delay the warning that capacity is being consumed, so monitor actual growth and set alerts before the cluster reaches a critical threshold.
Apply policies by workload. A domain controller, management appliance and file server may need different performance and availability settings from a test VM. Document the policy assigned to each critical service, and check that the policy is compliant after hardware changes or maintenance.
| Design choice | Practical small-environment approach | Main trade-off |
|---|---|---|
| Two-node cluster | Two data hosts plus an external witness | Lower cost, greater witness dependency |
| Three-node cluster | Three full vSAN hosts | Better fault tolerance, higher capital cost |
| FTT=1 mirroring | Two data copies | Strong protection, about 2x capacity use |
| RAID-5/6 | Erasure coding where supported | Better efficiency, more design constraints |
| 10 GbE networking | Redundant links and isolated traffic | Requires suitable switches and optics |
| Remote witness | Separate site or failure domain | WAN quality becomes important |
Plan operations before going live
Resilience includes safe maintenance. Confirm that a host can enter maintenance mode using the selected data migration option without exhausting capacity or causing excessive resynchronisation. Schedule firmware updates and disk replacements during periods of low workload activity.
Monitoring should cover vSAN health, capacity, latency, resync status, disk endurance, controller alerts and network errors. Integrate alerts with the team’s existing system rather than relying only on the vSphere Client. A practical runbook should explain how to identify a failed disk, evacuate a host and validate component compliance.
When an ESXi host disappears from vCenter, establish whether the issue is host, management network, switching or storage related. These troubleshoot ESXi connections steps can help isolate common connectivity faults before they become a wider availability event.
Test recovery and capacity assumptions
A design is only resilient if its recovery behaviour has been tested. Use a maintenance window to simulate a host failure, confirm that affected VMs restart, and measure the time required for vSAN components to resynchronise. Test witness loss separately so the team understands what continues operating and what becomes unavailable.
Review backup and disaster recovery independently from vSAN. vSAN protects against selected component and host failures; it does not replace application-aware backups or a second copy in another location. For organisations in Perth, Adelaide or regional areas, a cloud backup target or managed recovery service may be more practical than a second on-premises site.
Document the cluster’s IP addresses, VLANs, switch ports, disk layout, licensing, firmware versions and support contacts. A clear build record, such as this VMware build guide, reduces dependence on one administrator and makes future expansion safer.
A small vSAN cluster should be easy to understand under pressure. Validate the failure model, leave capacity for recovery, use supported hardware and test the procedures before a real outage occurs. Review the design annually as workloads, subscription costs and local service availability change, then invest in the next resilience improvement that removes the largest remaining risk.