Troubleshooting Network Performance in VMware vDS
A VMware vSphere Distributed Switch (vDS) gives infrastructure teams consistent port groups, centralised policy, Network I/O Control and advanced visibility across ESXi hosts. When performance drops, however, the fault can sit anywhere between a virtual machine’s vNIC and the physical switch, making symptoms such as packet loss, jitter and slow application response difficult to isolate.
Australian environments add practical complications. A Sydney host connecting to storage or services in Melbourne may behave differently from a Perth workload crossing a longer WAN path, while NBN, carrier Ethernet and cloud interconnects each introduce their own limits. A methodical workflow prevents a “she’ll be right” diagnosis and identifies the failing layer with evidence.
Establish A Reliable Baseline
Start by recording normal throughput, round-trip latency, packet loss and retransmissions for the affected VM or service. Capture results during a quiet period and during the reported busy window. iperf3 is useful for controlled TCP and UDP testing, while application metrics reveal whether the problem is network capacity, server processing or storage wait time.
Check the vDS version, host build, port group configuration and physical adapter model before changing settings. Document the VM’s VLAN, MTU, teaming policy, security settings and active uplinks. A baseline from a Brisbane office or a Sydney data centre is only useful if the test path and workload are clearly identified.
Trace The Path From Guest To Uplink
Work outward from the guest operating system. Verify the VM’s IP configuration, default gateway, DNS response time and virtual NIC driver. Windows Performance Monitor, Linux ss and interface counters can expose retransmits or receive errors that are invisible from the vCenter summary screen.
In vCenter, inspect the VM’s connected port, distributed port group and host uplink. Confirm that the expected VLAN is available on every physical trunk carrying that vDS. A mismatched native VLAN or an omitted allowed VLAN can produce intermittent reachability that looks like poor performance rather than a straightforward configuration error.
Read VDS And ESXi Counters
The vDS networking view can show dropped packets, port utilisation and uplink health, but ESXi command-line tools add useful detail. Use esxtop in network mode to examine receive and transmit rates, dropped packets, queue pressure and NIC activity. A single saturated vmnic beside several idle adapters often points to teaming, hashing or switch-side configuration rather than a lack of total bandwidth.
For packet-level evidence, pktcap-uw can capture traffic at different points in the virtual path. Compare what enters the virtual switch with what leaves the uplink. This helps distinguish guest-generated retransmissions from losses introduced by a physical NIC, VLAN trunk or upstream security appliance. For latency-sensitive workloads, including casino game traffic, a few milliseconds of jitter can matter even when average bandwidth looks healthy.
Validate MTU And Physical Switching
An MTU mismatch is a common source of confusing performance. Standard Ethernet usually uses a 1500-byte MTU, while vMotion, storage and overlay networks may use jumbo frames. Test the complete path with a deliberately sized, non-fragmented ping rather than assuming that a configured 9000-byte value is working end to end.
Check the physical switch ports for CRC errors, input drops, output queue drops, speed or duplex negotiation problems and increasing interface resets. Review LACP, port-channel membership and load-balancing settings together with the vDS teaming policy. A port-channel configured on the switch but treated as independent links by ESXi, or the reverse, can create uneven traffic and unexpected failover behaviour.
Review NIOC And Traffic Contention
Network I/O Control can reserve bandwidth and apply shares to traffic types such as management, vMotion, vSAN, virtual machines and backup. Review active network resource pools and confirm that critical workloads have appropriate reservations. A backup window in a Melbourne cluster can consume an uplink and starve production traffic if priorities are left at their defaults.
Also examine physical switch QoS, firewall inspection, IDS appliances and WAN shaping. A VM may show low utilisation while its traffic is being queued downstream. If workloads run across a hybrid design, compare the vDS path with cloud gateway or SD-WAN metrics. Clear ownership between the VMware, network and security teams avoids repeated tests that examine only one segment.
Check Drivers, Firmware And Workload Design
Confirm that the ESXi host uses the supported driver and firmware combination for its physical NICs. Intel, Broadcom and Mellanox adapters can expose different offload, queue and RSS behaviour. Review VMware compatibility guidance before disabling TSO, LRO or checksum offloads; these features often improve performance, though a documented driver issue may justify a controlled test.
Virtual machine design matters as well. An undersized vCPU configuration, overloaded guest kernel or constrained application thread can be mistaken for network congestion. Kubernetes nodes hosted on vSphere add another layer of virtual interfaces, overlay encapsulation and service routing. When building such a platform, the Kubernetes creation guide provides useful infrastructure context, but packet capture and node-level metrics remain essential when diagnosing latency.
Use Controlled Changes And Evidence
Change one variable at a time and record the result. Move a test VM to another host, pin it temporarily to a different uplink, or compare a standard port group with a known-good one. Do not disable security controls, alter MTU values across production or rebuild a port channel without a rollback plan.
Keep screenshots, command output, switch counters and timestamps in the incident record. A short “arvo” test during business hours may reveal contention that an overnight benchmark misses. For repeatable checks, PowerCLI can collect vDS port, uplink and host details across a cluster; the free VMware tools collection can help accelerate that inventory work.
| Symptom | Likely area | Useful evidence | Practical check |
|---|---|---|---|
| High latency with low throughput | Queueing, QoS or WAN path | Switch drops, NIOC stats, ping variation | Test during and outside the busy period |
| Packet loss on one host | NIC, cable or switch port | esxtop, CRC counters, link events |
Move the VM or replace the physical path |
| Slow traffic after vMotion | Host-specific vDS or uplink issue | Port mapping, teaming state, MTU test | Compare source and destination host settings |
| Uneven uplink utilisation | Hashing or LACP mismatch | Port-channel counters, vDS policy | Align switch aggregation and teaming modes |
| Large transfers stall | MTU, offload or firewall inspection | Packet capture, fragmentation counters | Validate packet size across the entire path |
Use the workflow consistently across ESXi hosts, switches and applications, and treat every counter as part of a path rather than an isolated answer. Start with a baseline, prove where packets are lost or delayed, then make the smallest reversible correction. Build these checks into operational runbooks and use the evidence to keep VMware and network teams aligned.