VXLAN wraps an Ethernet frame inside a UDP packet so a layer 2 network can cross a routed layer 3 network. Each virtual network gets a 24 bit VNI, raising the ceiling from 4,094 VLANs to about 16 million. The device doing the wrapping is a VTEP, and the machines never know.
- Ethernet inside UDP, on port 4789
- The VNI is 24 bits, so about 16.7 million segments
- The wrapper costs 50 bytes, which the underlay MTU must allow
- EVPN replaced the original multicast flood and learn design
- It is encapsulation, not encryption. The inner frame is readable
On this page
Why it existsThe problem VXLAN was built for
Three limits of VLANs pushed cloud and data center networks toward an overlay, and none of them is the one people quote first.
The 4,094 ceiling. The VLAN tag carries a 12 bit identifier, so a physical network can hold 4,094 of them. That is generous for one company and nothing at all for a multi tenant cloud provider with thousands of customers, each wanting several networks of their own.
Layer 2 does not route. A VLAN is a broadcast domain, and stretching one across a building or between sites means stretching a broadcast domain, which means spanning tree, which means blocked links and a failure mode that takes everything down at once. The whole industry spent twenty years trying to avoid exactly this.
Virtual machines move. Server virtualization means a machine can migrate from one rack to another and keep its IP address, which means the network it lives on has to exist in both racks. Doing that with VLANs means trunking the VLAN everywhere it might ever be needed, which recreates the flat network the design was trying to avoid.
VXLAN answers all three by making the layer 2 network a passenger. The physical network infrastructure underneath is a plain routed IP fabric with no spanning tree and every link forwarding. The virtual networks ride on top of that infrastructure inside UDP, and a VXLAN segment can appear anywhere the fabric reaches.
The headersWhat the encapsulation actually looks like
A VXLAN packet is an ordinary Ethernet frame with a new set of headers wrapped around it. The original frame is untouched inside, which is what makes the network virtualization invisible to the machine that sent it.
original frame [ inner MAC ][ inner IP ][ payload ]
on the wire [ outer MAC ][ outer IP ][ UDP ][ VXLAN ][ inner MAC ][ inner IP ][ payload ]
14 20 8 8
Four headers are added and they total 50 bytes.
| Header | Bytes | What it carries |
|---|---|---|
| Outer Ethernet | 14 | The next hop on the physical network |
| Outer IP | 20 | Source and destination VTEP addresses |
| UDP | 8 | Destination port 4789, source port hashed per flow |
| VXLAN | 8 | The 24 bit VNI, plus flags |
| Total added | 50 | Which the underlay MTU has to allow for |
The outer Ethernet header, 14 bytes. Addressed to the next hop router on the physical network, exactly like any other packet.
The outer IP header, 20 bytes. Source is the sending VTEP, destination is the receiving VTEP. This is what makes the traffic routable: to every device in between, this is a normal IP packet between two hosts.
The UDP header, 8 bytes. Destination port 4789. The source port is deliberately randomized from a hash of the inner frame, which is the trick that makes equal cost paths work: routers hash on the UDP ports they can see, so different inner flows land on different physical links.
The VXLAN header, 8 bytes. Mostly the 24 bit VNI, which names the virtual network this frame belongs to. Everything else in the VXLAN header is flags and reserved space.
The receiving VTEP strips all four, reads the VNI to decide which VXLAN network the frame belongs to, and delivers the original frame into it. The machine at the far end sees an ordinary Ethernet frame from a machine it believes is on the same switch.
Underlay and overlayVTEPs, and the two networks that now exist
Once you run VXLAN there are two networks, and keeping them separate in your head is most of understanding the protocol.
The underlay is the physical network infrastructure: routers, switches, IP addresses, routing protocols. It carries VXLAN packets between VTEPs and knows nothing about the tenants riding on it. Its job is to be simple, fast and always available, which usually means a leaf and spine layout where every leaf reaches every spine and traffic is spread across all of them.
The overlay is the set of VXLAN networks, each identified by a VNI. This is where the tenants, the virtual machines and the broadcast domains live. Network virtualization means the overlay is created and torn down in software, without touching a cable.
A VTEP is the boundary between the two networks. It has an IP address on the underlay and it terminates VXLAN tunnels. Two kinds of VTEP exist and both are common.
A hardware VTEP is a switch, usually a leaf switch, doing the encapsulation in silicon. It is where physical servers and anything that is not virtualized connects.
A software VTEP is a hypervisor or a cloud host, doing the VXLAN encapsulation for the virtual machines it runs. The virtual machine sends an ordinary frame to a virtual switch, and the host wraps it before it ever reaches a cable.
The control planeFlood and learn, and why EVPN replaced it
A VTEP has to answer one question constantly: which other VTEP holds the machine this frame is addressed to. There are two ways to learn that, and the difference is the biggest thing that changed about VXLAN since the specification was written.
Flood and learn, the original VXLAN design. Each VNI is mapped to an IP multicast group in the underlay. When a VTEP receives a frame for an address it does not know, it floods the frame to that multicast group, every VTEP in the VNI receives it, and the one holding the machine responds.
Everybody learns from the reply. It works, and it requires multicast routing throughout the underlay, which many networks do not want to run.
EVPN, the control plane that replaced it. BGP carries the mapping instead. Each VTEP advertises the MAC addresses and IP addresses it holds, using BGP with an address family built for the purpose, and every other VTEP learns them without a single flooded frame.
No multicast, less flooding, faster convergence, and a control plane you can query when something is wrong.
| Flood and learn | EVPN | |
|---|---|---|
| How a VTEP learns addresses | By flooding and watching replies | From BGP advertisements |
| Needs multicast in the underlay | Yes | No |
| Broadcast traffic on the fabric | High | Much lower |
| Convergence after a move | Slow, on timers | Fast, on an update |
| Somewhere to look when it breaks | Nowhere | The BGP table |
| Used in current designs | Rarely | Almost always |
EVPN is what any current VXLAN deployment uses, to the point that the word in a vendor document usually means VXLAN with an EVPN control plane. The distinction still matters when reading older material, which describes a design almost nobody builds now.
Multi tenancyMulti tenancy, which is the commercial reason it exists
The technical case for VXLAN is the segment ceiling. The commercial case is multi tenancy, and it is why every cloud provider runs network virtualization of this kind.
A tenant in a shared infrastructure wants networks that behave as though nobody else is there: their own address ranges, their own broadcast domains, and no path to any other customer.
Two tenants both using 10.0.0.0/24 is not a conflict when each range lives in a different VNI, because the two VXLAN networks never meet. On a VLAN based fabric the same requirement means coordinating address space across every customer, which does not scale past a few dozen of them.
VXLAN separates the layer 2 networks and stops there. Keeping each tenant's routes separate as well is the job of a VRF, which holds its own routing table per tenant, and the two are normally configured together for that reason.
The separation is real at layer 2 and it is not security. Two VNIs cannot exchange a frame, but the underlay carries both, unencrypted, and anyone with access to the physical network infrastructure reads either one. Tenant isolation in a cloud is VXLAN plus the controls around it, never VXLAN alone.
The MTU trapThe MTU problem, which is the one that bites
The 50 bytes of VXLAN encapsulation are added to a frame that was already full size. This is the single most common way a VXLAN network fails, and the symptom never points at the cause.
A virtual machine sends a 1500 byte packet, which is normal and correct. The VTEP wraps it, producing a 1550 byte packet on the underlay. If the underlay is also configured for 1500, that packet is too large. It gets fragmented if the do not fragment bit permits, which costs performance, or it gets discarded, which costs the connection.
The symptom is precise and familiar. Small packets pass, so connections establish, pings succeed and the network looks healthy. Anything carrying real data hangs. It is the same pattern as any other MTU mismatch, and in a VXLAN fabric it appears the moment somebody moves a real workload onto it.
Two configurations avoid it, and one of them is the right answer.
Raise the underlay MTU. Set the physical network infrastructure to jumbo frames, commonly 9216, so 50 extra bytes are irrelevant. This is what every VXLAN design guide specifies and it is the correct answer.
Lower the MTU inside the overlay. Set the virtual machines to 1450 so the wrapped packet fits in 1500. It works, it is fragile, and it requires touching every guest, so it is what people do when they cannot change the physical network.
The failure to avoid is doing neither, which is also the default state of a fabric that nobody configured for this.
PitfallsWhere people go wrong
Leaving the underlay at 1500 bytes. The most common VXLAN networking failure by a wide margin, and it produces hangs rather than errors.
Treating network virtualization as a substitute for design. Sixteen million VXLAN segments does not make a flat network good. The segments still need addressing, routing between them, and rules about what may reach what.
Building it without EVPN because a tutorial used multicast. Flood and learn requires multicast routing in the underlay and gives you no visibility. Almost every current deployment uses a BGP control plane instead.
Forgetting that broadcast still exists. A VNI is a broadcast domain. Stretching one across sites means stretching broadcast traffic across sites, and a broadcast storm now travels over a wide area link.
Ignoring the underlay because the overlay is the interesting part. Every VXLAN problem that is not MTU is an underlay problem. If the physical network drops packets, the virtual network drops packets.
Assuming the VXLAN encapsulation encrypts anything. It does not. VXLAN is a wrapper, not a tunnel with security. Anyone able to read the underlay reads the inner frame in full, which matters in any multi tenant environment.
Stretching layer 2 between data centers because you can. VXLAN makes it possible and does not make it wise. A shared failure domain across two sites is still a shared failure domain.
ComparisonThree ways to separate networks, and the scale each one belongs at
| Criterion | VLAN | VXLAN | Plain routed |
|---|---|---|---|
| Segments available | 4,094 | 16.7 million | Not applicable |
| Identifier size | 12 bits | 24 bits | None |
| Crosses a router | No | Yes | Yes |
| Needs spanning tree | Yes | No | No |
| Overhead per frame | 4 bytes | 50 bytes | None |
| Needs a control plane | No | Yes, EVPN in practice | Yes, a routing protocol |
| Complexity | Low | High | Medium |
| Right for a small office | Yes | No | Yes |
The last row deserves saying plainly. VXLAN solves problems that appear at the scale of a data center or a cloud provider. In a company with one building and forty people, network virtualization adds a layer of encapsulation, a control plane and an MTU requirement in exchange for a segment ceiling nobody was going to reach.
FAQFrequently asked questions
What is VXLAN?
A network virtualization protocol that carries an Ethernet frame inside a UDP packet, so a layer 2 network can span a routed layer 3 network. Each VXLAN network is identified by a 24 bit VNI.
What is the difference between VLAN and VXLAN?
A VLAN is a tag inside an Ethernet frame that a switch reads, limited to 4,094 identifiers and unable to cross a router. VXLAN wraps the whole frame in UDP so it can be routed, and its identifier is 24 bits.
What is a VNI?
The VXLAN Network Identifier, a 24 bit number naming which virtual network a frame belongs to. It is the VXLAN equivalent of a VLAN ID, with about 16.7 million values instead of 4,094, which is what makes multi tenant clouds possible.
What is a VTEP?
A VXLAN Tunnel Endpoint: the device that wraps frames on the way out and unwraps them on the way in. It can be a switch doing it in hardware or a hypervisor doing it in software.
What port does VXLAN use?
UDP 4789 as the destination. The source port is randomized per flow so routers can spread traffic across equal cost paths.
How much overhead does VXLAN add?
50 bytes per frame: 14 for the outer Ethernet header, 20 for the outer IP header, 8 for UDP and 8 for the VXLAN header.
What MTU do I need for VXLAN?
At least 1550 on the underlay to carry a 1500 byte guest packet, and jumbo frames of 9216 in practice. Failing to raise it is the most common cause of a fabric that appears to work and hangs on real traffic.
What is the difference between the underlay and the overlay?
The underlay is the physical routed network carrying UDP between VTEPs. The overlay is the set of virtual layer 2 networks riding inside it. The underlay knows nothing about the overlay.
Does VXLAN encrypt traffic?
No. It is encapsulation, not protection. Anyone who can read packets on the underlay can read the original frame inside them.
What is EVPN and do I need it?
A BGP based control plane that distributes which addresses live behind which VTEP. It replaced the original multicast flood and learn design, and any current deployment uses it.
Does VXLAN replace spanning tree?
In the underlay, yes, because the underlay is routed rather than switched. Inside a single VNI you still have a broadcast domain, with the behavior that implies.
Do I need VXLAN in a small office?
Almost certainly not. VXLAN solves a segment ceiling and a virtualization mobility problem that appear at cloud and data center scale, and it costs a control plane and an MTU requirement.
Can VXLAN stretch a network between two sites?
Technically yes, and it is worth asking whether it should. Stretching layer 2 also stretches the failure domain and the broadcast traffic across the link between them.
Keep readingRelated concepts
Read next · Routing What Is BGP? EVPN is a BGP address family, so the control plane behind a modern fabric is the protocol that runs the internet. Open this next12 min- Switching · 11 min Spanning Tree Protocol A routed underlay has no loops to block, which is most of why the fabric is built this way.
- Switching · 14 min What Is a VLAN? The 12 bit tag and its 4,094 ceiling are the limit VXLAN was built to move.
- Routing · 11 min What Is VRF? Where VRFs and overlays are deployed together.
- Diagnostics · 12 min MTU and Fragmentation, and the Numbers Worth Memorizing Another fifty bytes, and the same requirement to raise the underlay.
- Switching · 10 min Top of Rack Switching, and the Number Nobody Puts on the Datasheet The physical layout the fabric is built on.