Building a Kubernetes Cluster on VMware vSphere from the Ground Up
Kubernetes has quietly become the default orchestration layer for containerised workloads, and plenty of Australian IT teams are now running it alongside their existing virtualisation estate. If you already invest in VMware vSphere across your data centres in Sydney, Melbourne, or Brisbane, you have a head start: the hypervisor, the network fabric, and the storage tier can all be repurposed to host a production-grade cluster. This walkthrough is for the practitioner who wants a hands-on path rather than a managed service, drawing on practical experience shared across the Virtxpert home lab notes.
The appeal of running Kubernetes on vSphere is simple. You keep the operational muscle you have already built: vMotion for planned maintenance, HA for unplanned failures, and distributed switches for clean network segmentation. Many local organisations, from banks in the Sydney CBD to logistics firms in Perth, prefer this approach because it lets them meet internal governance and the Australian Cyber Security Centre's Essential Eight maturity targets without doubling their tooling footprint.
Before you start, it helps to map the work into a few discrete phases. We will cover environment preparation, network design, storage layout, the actual cluster deployment, and the validation steps that prove everything is healthy. The goal is a cluster that behaves well during a Tuesday arvo change window, not just one that boots cleanly in a lab.
A quick note on language: this guide uses Australian spelling throughout, and a few terms that resonate locally. When we talk about a "change window," think the typical AEST/AEDT evening cut-over when user impact is lowest. When we mention the platform team, we mean the same people who look after vCenter and the SAN, regardless of whether your operation runs out of a Macquarie Telecom facility or a private cage in a NextDC hall.
Prerequisites and Planning for Your Cluster
A Kubernetes cluster is only as good as the foundation underneath it, so start by confirming the building blocks. You will need a vCenter-managed environment running vSphere 7.0u3 or later, with sufficient CPU, memory, and storage headroom for at least three control plane nodes and a couple of worker nodes. The vSphere CSI driver and the vSphere Cloud Provider Interface (CPI) must be available, which usually means a valid vCenter licence with the Storage APIs enabled.
You will also want a load balancer in front of the API server. Many Australian shops reach for HAProxy on a small VM, while others use NSX Advanced Load Balancer or a F5 VE. Pick something your network team already trusts, because debugging load-balancer issues at 2am AEDT is no fun for anyone.
The bulleted list below summarises the core components to have in place before deployment:
- vCenter 7.0u3 or newer with appropriate licences
- At least three ESXi hosts with shared storage
- A load balancer capable of TCP health checks
- A DNS entry and routable IP range for cluster services
Spend an hour sketching the node sizing on a whiteboard before you deploy. A control plane node with 4 vCPU and 8 GB of RAM is a sensible starting point, and workers can be sized to match your workload rather than guesswork.
Preparing the vSphere Folders, Tags, and Permissions
A clean organisation of vSphere objects saves hours later. Create a dedicated VM folder for the cluster, apply tags for environment, owner, and backup policy, and map those tags to your existing automation. If you are using Ansible to manage your vSphere estate, you can fold the cluster provisioning into the same playbooks that build your application VMs, which keeps the operational model consistent.
Permissions matter too. A dedicated service account for Kubernetes, scoped to the cluster's resource pool, avoids the all-too-common situation where a runaway kubectl command accidentally rearranges production. Restrict the account to the data centre, cluster, and resource pool you intend to use, and document the scope in your change record.
Choosing Your Kubernetes Distribution
The distribution decision shapes everything that follows. Upstream kubeadm is free and familiar, and it gives you complete control over component versions, which is helpful when you need to align with internal vulnerability management. Vendor distributions such as Tanzu, OpenShift, or Rancher trade some of that freedom for integrated networking, ingress, and support contracts that map neatly to enterprise procurement.
For a regulated Australian business, the support question often tips the scales. If you operate under IRAP-aligned change processes or answer to a parent organisation that requires 24x7 vendor support, a paid distribution pays for itself the first time something breaks at an awkward hour. Match the choice to your risk appetite rather than to whatever happened to trend on social media that week.
Network Design and Address Planning
The network plan is where most home-grown deployments stumble. Decide early whether you will use a single flat network for all nodes, or separate networks for control plane and workers. Many local teams opt for the latter, often sitting on top of a distributed switch with VLAN-backed segments that map cleanly to their existing firewall zones.
If you operate across multiple sites, consider whether you need a stretched cluster. The latency between, say, a Macquarie Park data centre and a secondary site in Canberra matters: keep inter-node traffic under 10 ms if you can, and remember that synchronous storage replication will not forgive much more than that. Plan IP ranges, gateway addresses, and DNS resolvers before you touch the hypervisor.
Storage Configuration for Stateful Workloads
Kubernetes on vSphere can use VMFS, vSAN, or external SAN storage behind the scenes, but the cluster needs a default StorageClass to be useful. Install the vSphere CSI driver as part of your bootstrap, and define at least one StorageClass that maps to a datastore or datastore cluster suited to your performance profile.
For most Australian enterprises, the choice comes down to cost versus capability. A vSAN stretch cluster with erasure coding gives you resilience without the price of a dedicated all-flash array, and it plays nicely with the vSphere CSI driver's topology awareness. Test failure scenarios in a lab first: pull a host, lose a disk group, and watch the cluster recover before you trust it with production.
Deploying the Cluster
With the groundwork laid, the actual deployment is a matter of running the right tool against the right endpoints. Run the control plane first, join the additional control plane nodes, then add workers in batches that match your maintenance windows. Keep the join commands in your password vault rather than a shared Confluence page; the principle is no different from protecting any other privileged credential.
As nodes come online, watch the vSphere CPI and CSI pods carefully. They are the bridge between Kubernetes and your virtual infrastructure, and any permission or network hiccup will show up there first. If the pods stay in a crash loop, the fix is almost always a vSphere permission, a certificate, or a missing datastore mount rather than a Kubernetes manifest issue.
Post-Deployment Validation and Day-Two Operations
A cluster that boots is not yet a cluster you can trust. Validate node readiness, confirm that the vSphere CPI and CSI pods are running, and run a sample workload that exercises both compute and persistent storage. The list below is a good starting point:
- kubectl get nodes shows all nodes in Ready state
- A test pod binds a PVC and reads back written data
- Node drain and cordon behave as expected during simulated maintenance
- Metrics and logs flow to your existing observability stack
Once validation is complete, wire the cluster into your normal operational rhythm. That means patching aligned with your regular ESXi cycle, backups handled by the same tooling that protects the rest of your virtual estate, and runbooks that reference the same change advisory process you already use in Sydney or Melbourne. Keep an eye on certificate rotation, etcd backups, and Kubernetes version skew, and treat upgrades as a routine quarterly activity rather than a fire drill.
The most rewarding part of this work is watching a team that has spent years in vCenter start to feel just as comfortable with kubectl as they do with the vSphere client. Have a go in your home lab first, document the gotchas you hit, and bring that hard-won knowledge into your next production deployment. The Kubernetes on vSphere journey rewards practitioners who take the time to get the foundations right.