Datacenter & Leaf-Spine • 2253 words • 10 min read

Modern Data Center Network Architecture First Principles: Clos Topologies, Leaf-Spine Fabrics, BGP Underlay, and EVPN-VXLAN Large Layer-2 Virtualization

Why can't traditional three-tier architectures and Spanning Tree Protocol (STP) handle modern cloud computing and large-scale AI clusters? Deconstruct the evolution of data center networking: from classical 3-tier bottlenecks (Access, Aggregation, Core) with STP link blocking and oversubscription, to Charles Clos's 1953 telephone switching theory and modern non-blocking 2-Tier / 3-Tier Leaf-Spine mathematical modeling; analyze RFC 7938 eBGP Underlay routing designs, ECMP 5-tuple hash polarization hazards and ASIC-level solutions; and dissect large Layer-2 Overlay mechanics: VXLAN MAC-in-UDP encapsulation (24-bit VNI, 9000 MTU Jumbo Frames) and MP-BGP EVPN Type-2/Type-3/Type-5 control plane route signaling.

Networking Series Part 8 / 12

Computer Networking Masterclass: From Ethernet Principles to Hyperscale InfiniBand Architecture

A ground-zero masterclass to 10,000-GPU AI networking: from physical voltages, twisted pairs, and optical fibers to hubs, collision domains, switches, and MAC/ARP; step-by-step binary derivations of IP addressing, subnet masks, default gateways, VLANs, DHCP, and NAT; in-depth DNS, Socket 5-tuples, TCP 11-state machines, and BBR congestion control; Linux kernel NAPI, sk_buff, and eBPF XDP; datacenter traditional 3-tier vs 2-tier Spine-Leaf fabrics and EVPN-VXLAN; leading up to AI supercomputing: lossless RoCEv2, native InfiniBand NDR/XDR link speeds, credit flow control, Rail-Optimized Fat-Trees, Adaptive Routing, SHARP, and bare-metal OFED/NCCL performance tuning.

Browse all 12 chapters in this series ▾
  1. 01 Computer Networking from Scratch: Bits, Physical Media, Hub Collision Domains, and Switch MAC Addressing First Principles
  2. 02 IP Addressing and Subnetting First Principles: Binary Arithmetic, Subnet Masks, CIDR, and Default Gateway Routing
  3. 03 Enterprise LAN Infrastructure: VLAN Segmentation (802.1Q), Dynamic DHCP, and NAT Port Forwarding
  4. 04 Application & Transport Layer Bridges: DNS Resolution, Sockets, Ports, and UDP vs TCP Foundations
  5. 05 Network & Transport Layers in Depth: IP Routing, CIDR, TCP 11-State Machine, and Sliding Window First Principles
  6. 06 TCP Congestion Control Evolution & High-Performance Transport: From Reno and Cubic to BBR Mathematical Models, and HTTP/2 to HTTP/3 (QUIC)
  7. 07 Linux Kernel Networking Subsystem in Depth: From NIC Drivers, NAPI, and Ring Buffers to eBPF XDP Wire-Speed Forwarding
  8. 08 Modern Data Center Network Architecture First Principles: Clos Topologies, Leaf-Spine Fabrics, BGP Underlay, and EVPN-VXLAN Large Layer-2 Virtualization Reading
  9. 09 RDMA High-Performance Networking Foundations: Kernel Bypass, Zero-Copy, Queue Pairs, and Lossless RoCEv2 (PFC/ECN) Architecture
  10. 10 InfiniBand Architecture First Principles: Physical Link Rates, Credit-Based Link Flow Control, and Subnet Manager Fabric Orchestration
  11. 11 InfiniBand AI Cluster Networking in Practice: Fat-Tree Topologies, Rail-Optimized Architecture, Adaptive Routing (AR), and In-Network Reduction (SHARP)
  12. 12 Production InfiniBand Deployment, Cluster Operations, and NCCL Tuning: OFED Drivers, OpenSM HA, ibdiagnet Fabric Auditing, and GPUDirect RDMA

Introduction: From North-South Enterprise Traffic to East-West AI Computation Torrents

In Chapter 7 of our masterclass, Linux Kernel Networking Subsystem in Depth: From NIC Drivers, NAPI, and Ring Buffers to eBPF XDP Wire-Speed Forwarding, we explored how a host operating system optimizes packet paths and leverages eBPF for microsecond forwarding.

However, when thousands of servers are racked into modern cloud facilities and 10,000-GPU AI compute clusters, the physical network challenges scale exponentially.

In the traditional enterprise IT era, data center traffic was predominantly North-South (client-to-server). Physical topologies were structured around the classic three-tier hierarchy: Access, Aggregation, and Core.

To prevent Layer-2 broadcast loops, switches ran STP (Spanning Tree Protocol), deliberately blocking half of all physical links:

  • 50% stranded bandwidth: Half of all costly optical links remained idle as passive standbys;
  • Severe oversubscription bottlenecks: As rack densities increased, upstream bandwidth between aggregation and core tiers suffered from severe oversubscription ratios (10:1 or 20:1);
  • Unpredictable East-West latency: When Server A communicated with Server B in the same facility, packets had to climb all the way to the core switches and descend back, enduring multiple hops of queueing delays.

With the rise of cloud computing, microservices, distributed object stores, and the explosion of 10,000-GPU distributed AI training clusters in 2026, East-West traffic (server-to-server) now accounts for over 80% of all data center packets!

Traditional three-tier STP networks collapsed under this load. To deliver predictable, high-bandwidth, ultra-low-latency connectivity between any pair of hosts, modern data center fabrics underwent a ground-up revolution.


1. Architectural Paradigm Shift: Traditional 3-Tier Networks vs Modern 2-Tier Leaf-Spine

Understanding modern hyperscale fabrics requires understanding what they replaced: the legacy three-tier network.

flowchart TD
    subgraph Traditional3Tier["Legacy Three-Tier Hierarchy (Access - Aggregation - Core)"]
        Core1["Core Switch 1"] --- Core2["Core Switch 2"]
        Agg1["Agg Switch 1"] --- Agg2["Agg Switch 2"]
        Acc1["Access Switch 1"]
        Acc2["Access Switch 2"]

        Core1 --- Agg1 & Agg2
        Core2 --- Agg1 & Agg2
        Agg1 === Acc1 & Acc2
        Agg2 -.-|"STP blocks 50% of links to prevent loops!"| Acc1 & Acc2
    end

1.1 Three Fatal Flaws of Traditional Three-Tier Networks

  1. STP Link Waste: To prevent Layer-2 broadcast loops, Spanning Tree Protocol (STP) disables half of all physical uplink connections. Half of the expensive optical cabling sits idle in blocking state, cutting bandwidth utilization to 50%;
  2. Severe Oversubscription: 48 servers connect to an Access switch (48 Gbps aggregate), but the uplinks to Aggregation switches often total only 2 to 4 Gbps—an oversubscription ratio of 12:1 to 24:1! When multiple servers burst concurrently, aggregation buffers quickly overflow;
  3. East-West Latency Jitter: Inter-rack server communication requires 4 switch hops (Access →\to Agg →\to Core →\to Agg →\to Access), amplifying queueing jitter across multiple tiers.

2. Clos Switching Theory and Modern Leaf-Spine Topologies

The mathematical foundation of modern data center networking was established in 1953 by Bell Labs mathematician Charles Clos in his seminal paper on multi-stage telephone switches—the Clos Network. Modern fabrics implement this as a 2-Tier Leaf-Spine architecture:

1.1 2-Tier Leaf-Spine (Folded Clos Architecture)

flowchart TD
    subgraph SpineTier["Spine Tier Switches (Pure High-Speed Backbones, No End-Hosts)"]
        Spine1["Spine 1"]
        Spine2["Spine 2"]
        Spine3["Spine 3"]
        Spine4["Spine 4"]
    end

    subgraph LeafTier["Leaf Tier Switches / Top-of-Rack (ToR)"]
        Leaf1["Leaf 1 (Rack 1)"]
        Leaf2["Leaf 2 (Rack 2)"]
        Leaf3["Leaf 3 (Rack 3)"]
    end

    subgraph Servers["Physical Compute Nodes (GPU / CPU Servers)"]
        S1["Server A"] --> Leaf1
        S2["Server B"] --> Leaf1
        S3["Server C"] --> Leaf2
        S4["Server D"] --> Leaf3
    end

    Leaf1 === Spine1 & Spine2 & Spine3 & Spine4
    Leaf2 === Spine1 & Spine2 & Spine3 & Spine4
    Leaf3 === Spine1 & Spine2 & Spine3 & Spine4

Topological Rules

  1. Full Mesh Interconnection: Every Leaf switch connects directly to every Spine switch;
  2. Horizontal Tier Isolation: No Spine-to-Spine links, and no Leaf-to-Leaf links;
  3. Deterministic Latency: For Server A in Rack 1 to communicate with Server D in Rack 3, the packet traversal path is strictly fixed at 3 switch hops (Leaf 1 →\to any Spine →\to Leaf 3), completely eliminating hop-count jitter.

1.2 Non-blocking Fabrics and Oversubscription Ratio Mathematics

The standard metric for potential network congestion is the Oversubscription Ratio (RR):

R=Total Downlink Bandwidth Facing ServersTotal Uplink Bandwidth Facing SpinesR = \frac{\text{Total Downlink Bandwidth Facing Servers}}{\text{Total Uplink Bandwidth Facing Spines}}

Consider a 64-port 400 Gbps Leaf switch:

  • Downlinks: 32 ports connected to servers at 400 Gbps each = 32×400 Gbps=12.8 Tbps32 \times 400\text{ Gbps} = 12.8\text{ Tbps};
  • Uplinks: 32 ports connected to 32 Spine switches at 400 Gbps each = 32×400 Gbps=12.8 Tbps32 \times 400\text{ Gbps} = 12.8\text{ Tbps}.

The resulting oversubscription ratio is:

R=12.8 Tbps12.8 Tbps=1:1R = \frac{12.8\text{ Tbps}}{12.8\text{ Tbps}} = 1:1
  • 1:1 Non-blocking Fabric: Even if all servers in the rack transmit full wire-speed traffic simultaneously to external nodes, zero uplink contention occurs! For mission-critical AI training fabrics, a 1:1 ratio is mandatory;
  • Hyperscale 3-Tier Clos Fabrics: When server counts exceed the port capacity of a 2-tier design (scaling from thousands to tens of thousands of hosts), data centers deploy Pod-based architectures with Super-Spine switches, forming a 5-stage (3-Tier Folded) Clos network.

2. Underlay Routing: eBGP and ECMP Hash Polarization

With physical cabling in place, how do we distribute traffic evenly across all Spine switches without loops? Modern hyperscalers employ Layer 3 routing down to the Top-of-Rack (ToR) switch.

RFC 7938 (authored by engineers from Facebook/Meta and Arista) defined the industry standard for hyperscale Underlays: eBGP (external BGP).

flowchart TD
    subgraph SpinePlane["Spine Shared Autonomous System (AS 65000)"]
        Spine1["Spine 1 (ASN 65000)"]
        Spine2["Spine 2 (ASN 65000)"]
    end

    subgraph LeafPlane["Each Leaf Owns a Dedicated Private ASN"]
        Leaf1["Leaf 1 (ASN 65001)"]
        Leaf2["Leaf 2 (ASN 65002)"]
    end

    Leaf1 -- eBGP Session --- Spine1 & Spine2
    Leaf2 -- eBGP Session --- Spine1 & Spine2

2.1 Why eBGP Over OSPF/IS-IS?

  1. Minimal Blast Radius: OSPF and IS-IS are Link-State protocols using Dijkstra's algorithm. A single flapping optical link triggers fabric-wide LSA floods and shortest-path (SPF) recalculations, risking CPU exhaustion. eBGP is a Path-Vector protocol with built-in route damping and step-by-step filtering;
  2. Loop Prevention via AS_PATH: BGP nodes discard any route advertisement containing their own ASN, eliminating routing loops natively.

2.2 ECMP and Hash Polarization

When forwarding packets toward a remote subnet, a Leaf switch sees multiple equal-cost paths across the Spines. Switch ASICs compute a hash over the packet's 5-tuple (Source IP, Destination IP, Protocol, Source Port, Destination Port):

Egress Port=Hash(SrcIP, DstIP, Protocol, SrcPort, DstPort)(modN)\text{Egress Port} = \text{Hash}(\text{SrcIP, DstIP, Protocol, SrcPort, DstPort}) \pmod N

The Hazard: Hash Polarization

If Leaf switches and Spine switches execute the identical, unperturbed polynomial hash function:

  • The Leaf switch splits odd-hashed flows to Spine 1 and even-hashed flows to Spine 2;
  • When packets reach the Spine, the Spine evaluates the same hash function on the same 5-tuple, causing flows to concentrate onto the same downstream egress port, while parallel links sit starved!
  • Engineering Solution: Modern switch ASICs introduce unique, tier-specific Hash Seeds (salts) and support Dynamic Load Balancing (DLB), which monitors queue depths at microsecond granularity to steer sub-flows (Flowlets) away from congested links.

3. Overlay Virtualization: EVPN-VXLAN in Depth

While the physical Underlay provides high-speed Layer 3 transport, multi-tenant cloud environments and Kubernetes container fabrics require seamless Layer 2 network abstractions across physical nodes.

3.1 VXLAN: MAC-in-UDP Encapsulation

VXLAN (RFC 7348) encapsulates guest Layer-2 Ethernet frames inside standard physical UDP datagrams:

VXLAN Packet Encapsulation Hierarchy:
+-------------------------------------------------------------------------+
| Outer Ethernet Header (Outer MAC): Egress NIC MAC -> Next-Hop MAC      | 14B
+-------------------------------------------------------------------------+
| Outer IP Header       (Outer IP) : VTEP_A IP -> VTEP_B IP (L3 Underlay) | 20B
+-------------------------------------------------------------------------+
| Outer UDP Header      (Outer UDP): SrcPort(5-Tuple Hash) -> DstPort 4789| 8B
+-------------------------------------------------------------------------+
| VXLAN Header          (Flags + 24-bit VNI, 16 Million Tenants)          | 8B
+-------------------------------------------------------------------------+
| Inner Ethernet Header (Inner MAC): VM_A MAC -> VM_B MAC                | 14B
+-------------------------------------------------------------------------+
| Inner Payload         (Inner IP/TCP/Payload): Original Guest Packet     | ~1460B
+-------------------------------------------------------------------------+
  1. Shattering the 4096 VLAN Limit: VXLAN provides a 24-bit VNI (VXLAN Network Identifier), scaling to over 16.7 million16.7\text{ million} isolated virtual tenant networks;
  2. ECMP Load Balancing via Outer UDP Source Port: The VTEP hashes the inner 5-tuple into the outer UDP source port. Intermediate Underlay Spine/Leaf switches do not need to understand VXLAN; standard ECMP distributes encapsulated packets across all spine links;
  3. Mandatory Jumbo Frames: The outer headers introduce 14+20+8+8=5014 + 20 + 8 + 8 = 50 bytes of encapsulation overhead. To prevent 1500-byte guest frames from fragmenting, Underlay physical switch ports must be configured with an MTU of 9000 or 9216 bytes!

3.2 The Control Plane Solution: MP-BGP EVPN

Early VXLAN relied on flood-and-learn multicast to discover MAC addresses, frequently flooding the physical fabric with broadcast ARP storms.

EVPN (RFC 7432 / RFC 8365) employs MP-BGP as an out-of-band control plane. VTEPs exchange host MAC and IP reachability via structured BGP Network Layer Reachability Information (NLRI):

sequenceDiagram
    autonumber
    actor VTEP1 as Leaf 1 (VTEP A)
    actor Spine as BGP Route Reflector (Spine)
    actor VTEP2 as Leaf 2 (VTEP B)

    Note over VTEP1: Virtual Machine VM-A (IP: 10.1.1.2, MAC: AA:AA) boots up
    VTEP1->>Spine: BGP Type-2 Route: VNI=100, MAC=AA:AA, IP=10.1.1.2, NextHop=VTEP_A
    Spine->>VTEP2: Reflects Type-2 Route
    Note over VTEP2: Injects VM-A into local forwarding table with zero broadcast flooding!
  • Type-2 Route (MAC/IP Advertisement Route): Advertises host MAC and IP bindings, eliminating multicast ARP flooding;
  • Type-3 Route (Inclusive Multicast Ethernet Tag Route): Dynamically establishes head-end replication tunnels between VTEPs for unavoidable BUM (Broadcast, Unknown Unicast, Multicast) traffic;
  • Type-5 Route (IP Prefix Route): Carries subnet prefixes for inter-subnet and cross-VPC Layer-3 routing.

4. Hands-On: Configuring Native VXLAN on Linux

To establish a point-to-point VXLAN tunnel between two Linux hosts (Host A: 192.168.10.1, Host B: 192.168.10.2):

# 1. On Host A: Create VXLAN device (VNI 100, physical interface eth0, port 4789)
ip link add vxlan100 type vxlan \
    id 100 \
    dev eth0 \
    local 192.168.10.1 \
    remote 192.168.10.2 \
    dstport 4789

# 2. Assign virtual private IP and bring the interface up
ip addr add 10.100.0.1/24 dev vxlan100
ip link set vxlan100 up

# 3. On Host B: Configure reciprocal endpoint
ip link add vxlan100 type vxlan \
    id 100 \
    dev eth0 \
    local 192.168.10.2 \
    remote 192.168.10.1 \
    dstport 4789
ip addr add 10.100.0.2/24 dev vxlan100
ip link set vxlan100 up

# 4. Verify end-to-end Layer-2 reachability across physical Layer-3
ping 10.100.0.2

5. Summary and Next Steps

By combining Clos topologies, eBGP Underlay routing, ECMP load balancing, and EVPN-VXLAN virtualization, modern hyperscale data centers decouple physical infrastructure from logical multi-tenant services.

However, even with 400 Gbps / 800 Gbps Leaf-Spine bandwidth, traditional Ethernet networking encounters severe bottlenecks when supporting large-scale AI distributed training:

  • CPU Overhead and Context Switching: At 400 Gbps, TCP kernel interrupts and memory copies consume over 80% of CPU cores;
  • Strict Microsecond Latency Demands: In distributed training with 10,000 GPUs, All-Reduce operations synchronize across every node. Millisecond TCP jitter leaves billions of dollars of GPU hardware sitting idle;
  • What is RDMA (Remote Direct Memory Access)? How do Kernel Bypass and Zero-Copy bring end-to-end latency below 1 microsecond?
  • How does RoCEv2 create "Lossless Ethernet" using Priority Flow Control (PFC) and ECN/DCQCN without requiring specialized InfiniBand switches?

In Chapter 9 of our masterclass, we explore high-performance AI networking: RDMA High-Performance Networking Foundations: Kernel Bypass, Zero-Copy, Queue Pairs, and Lossless RoCEv2 (PFC/ECN) Architecture!


Frequently Asked Questions (FAQ)

Q1: Why do hyperscale data center Underlays favor eBGP over OSPF or IS-IS?

Key factors: Blast radius isolation, fine-grained traffic policy control, and SPF calculation overhead.

  1. Flap dampening: OSPF and IS-IS are Link-State protocols. In massive Clos fabrics with tens of thousands of links, a single flapping transceiver triggers fabric-wide Link State Advertisement (LSA) flooding and shortest-path tree (SPF) recalculations, risking control-plane CPU collapse. eBGP operates on Path-Vector principles with hop-by-hop dampening and scoped route propagation;
  2. Operational control: Allocating a unique private Autonomous System Number (ASN) to each Top-of-Rack Leaf switch allows engineers to enforce deterministic traffic engineering, route tagging via BGP Communities, and graceful maintenance drains (e.g., prepending ASNs to redirect traffic with zero packet loss).

Q2: What causes ECMP "Hash Polarization" in multi-tier Clos networks, and how do modern switch ASICs resolve it?

Definition: Hash polarization occurs when switches across multiple tiers (Leaf, Spine, Super-Spine) run the identical mathematical hash algorithm over the same packet 5-tuple. Flows assigned to a specific spine link in Tier 1 produce identical hash outputs in Tier 2, concentrating traffic onto a single downstream egress port while parallel links sit idle. ASIC-Level Solutions:

  1. Unique Hash Seeds: Each switch tier incorporates a unique polynomial seed/salt, ensuring hash distributions differ between tiers;
  2. Dynamic Load Balancing (DLB): Modern ASICs (e.g., Broadcom Tomahawk 4/5) replace static per-packet hashing with flowlet-level load balancing. The hardware monitors queue egress depths in real time, automatically routing sub-flows (Flowlets) to less congested paths.

Q3: In EVPN-VXLAN design, what is the difference between Symmetric IRB and Asymmetric IRB for inter-subnet routing?

The distinction lies in whether both source and destination VTEPs must configure the destination Layer-2 VNI.

  • Asymmetric IRB: The source VTEP routes the packet into the destination's Layer-2 VNI and encapsulates it. The destination VTEP simply bridges the packet locally. Drawback: Every VTEP must configure every Layer-2 VNI and learn all tenant MAC/ARP entries across the entire data center, quickly exhausting switch hardware forwarding tables (TCAM);
  • Symmetric IRB: Employs a dedicated Layer-3 VNI (Tenant VRF) for inter-subnet transit. The ingress VTEP routes the packet into the L3 VNI tunnel. The egress VTEP unpacks the packet and routes it into the local subnet. Advantage: VTEPs only need to maintain local subnets and the shared transit L3 VNI, reducing forwarding table footprint by an order of magnitude. This is the industry-standard architecture for modern cloud networks.

Related Articles

Start with the same topic, then continue with the latest deep dives.

Production InfiniBand Deployment, Cluster Operations, and NCCL Tuning: OFED Drivers, OpenSM HA, ibdiagnet Fabric Auditing, and GPUDirect RDMA

What bare-metal operational challenges arise when translating network architecture into physical 10,000-GPU AI data centers? A comprehensive guide to production InfiniBand operations: installing and managing Mellanox OFED / DOCA driver stacks and firmware tools (flint/mlxlink); configuring high-availability Master/Standby Subnet Manager topologies with OpenSM; auditing fabric health and diagnosing dirty optical links (Symbol Errors) and speed renegotiation drops using ibdiagnet; and diving into the GPU communication layer to configure GPUDirect RDMA (nvidia-peermem) and tune mission-critical NCCL parameters (NCCL_IB_HCA, NCCL_NET_GDR_LEVEL=5) for wire-speed All-Reduce performance.

InfiniBand AI Cluster Networking in Practice: Fat-Tree Topologies, Rail-Optimized Architecture, Adaptive Routing (AR), and In-Network Reduction (SHARP)

How can thousands of 8-GPU servers be interconnected with tens of thousands of optical links into a non-blocking, deadlock-free high-performance fabric? Deconstruct modern AI supercluster topologies: from Charles Leiserson's 1985 Fat-Tree mathematical model and port count k derivations for 2-Tier and 3-Tier non-blocking ceilings to the 8-plane Rail-Optimized architecture tailored for DGX H100/H200/B200 clusters; analyze how FTree and Up/Down routing engines forbid 'Down-then-Up' turns to eliminate credit loop deadlocks; and discover how hardware Adaptive Routing (AR) and SHARP in-network aggregation achieve a 2x throughput boost during GPU All-Reduce operations.

InfiniBand Architecture First Principles: Physical Link Rates, Credit-Based Link Flow Control, and Subnet Manager Fabric Orchestration

Why does native InfiniBand remain the dominant fabric for 10,000-GPU AI compute clusters and top-tier supercomputers in 2026? Deconstruct the layered InfiniBand protocol stack and physical link evolution: from EDR 100G, HDR 200G, and NDR 400G (Quantum-2) to the 2026 mass production of XDR 800G (Quantum-X800/ConnectX-8); analyze link-layer first principles: Flit-level Credit-Based hardware flow control and sub-100ns Cut-Through switching mechanics; and explore the control engine: Subnet Manager (OpenSM) fabric discovery, dynamic GUID/LID/LMC allocation, and Linear Forwarding Table (LFT) hardware orchestration.

← Prev Application & Transport Layer Bridges: DNS Resolution, Sockets, Ports, and UDP vs TCP Foundations Next → RDMA High-Performance Networking Foundations: Kernel Bypass, Zero-Copy, Queue Pairs, and Lossless RoCEv2 (PFC/ECN) Architecture
← Back to Articles