IP Routing & TCP State Machine • 2475 words • 10 min read

Network & Transport Layers in Depth: IP Routing, CIDR, TCP 11-State Machine, and Sliding Window First Principles

How can an absolutely reliable, ordered stream of data be built on top of an unreliable physical network? Deconstruct core Layer 3 and Layer 4 mechanics: from IPv4/IPv6 packet structures and CIDR subnets to Linux kernel Longest Prefix Match (LPM) and FIB/RIB table lookups; dissect the classic 11 TCP state transitions, 3-way handshake mechanics, SYN Flood & SYN Cookies defense, 4-way teardown with TIME_WAIT rationale; and derive sliding window flow control, zero-window probing, and TCP_NODELAY production tuning from first principles.

Networking Series Part 5 / 12

Computer Networking Masterclass: From Ethernet Principles to Hyperscale InfiniBand Architecture

A ground-zero masterclass to 10,000-GPU AI networking: from physical voltages, twisted pairs, and optical fibers to hubs, collision domains, switches, and MAC/ARP; step-by-step binary derivations of IP addressing, subnet masks, default gateways, VLANs, DHCP, and NAT; in-depth DNS, Socket 5-tuples, TCP 11-state machines, and BBR congestion control; Linux kernel NAPI, sk_buff, and eBPF XDP; datacenter traditional 3-tier vs 2-tier Spine-Leaf fabrics and EVPN-VXLAN; leading up to AI supercomputing: lossless RoCEv2, native InfiniBand NDR/XDR link speeds, credit flow control, Rail-Optimized Fat-Trees, Adaptive Routing, SHARP, and bare-metal OFED/NCCL performance tuning.

Browse all 12 chapters in this series ▾
  1. 01 Computer Networking from Scratch: Bits, Physical Media, Hub Collision Domains, and Switch MAC Addressing First Principles
  2. 02 IP Addressing and Subnetting First Principles: Binary Arithmetic, Subnet Masks, CIDR, and Default Gateway Routing
  3. 03 Enterprise LAN Infrastructure: VLAN Segmentation (802.1Q), Dynamic DHCP, and NAT Port Forwarding
  4. 04 Application & Transport Layer Bridges: DNS Resolution, Sockets, Ports, and UDP vs TCP Foundations
  5. 05 Network & Transport Layers in Depth: IP Routing, CIDR, TCP 11-State Machine, and Sliding Window First Principles Reading
  6. 06 TCP Congestion Control Evolution & High-Performance Transport: From Reno and Cubic to BBR Mathematical Models, and HTTP/2 to HTTP/3 (QUIC)
  7. 07 Linux Kernel Networking Subsystem in Depth: From NIC Drivers, NAPI, and Ring Buffers to eBPF XDP Wire-Speed Forwarding
  8. 08 Modern Data Center Network Architecture First Principles: Clos Topologies, Leaf-Spine Fabrics, BGP Underlay, and EVPN-VXLAN Large Layer-2 Virtualization
  9. 09 RDMA High-Performance Networking Foundations: Kernel Bypass, Zero-Copy, Queue Pairs, and Lossless RoCEv2 (PFC/ECN) Architecture
  10. 10 InfiniBand Architecture First Principles: Physical Link Rates, Credit-Based Link Flow Control, and Subnet Manager Fabric Orchestration
  11. 11 InfiniBand AI Cluster Networking in Practice: Fat-Tree Topologies, Rail-Optimized Architecture, Adaptive Routing (AR), and In-Network Reduction (SHARP)
  12. 12 Production InfiniBand Deployment, Cluster Operations, and NCCL Tuning: OFED Drivers, OpenSM HA, ibdiagnet Fabric Auditing, and GPUDirect RDMA

Introduction: Forging Determinism over an Untrusted Medium

In Chapter 4 of our masterclass, Application & Transport Layer Bridges: DNS Resolution, Sockets, Ports, and UDP vs TCP Foundations, we examined the global DNS hierarchy, port numbering, and OS socket 5-tuples.

However, the global Internet is a vast, heterogeneous web of millions of local area networks bridged together by intermediate routers:

  • The physical universe is inherently unreliable: Transoceanic submarine cables can be severed, router hardware queues frequently experience buffer overflow during traffic spikes, and wireless links suffer severe electromagnetic interference;
  • Stateless Best-Effort Delivery: The IP protocol is strictly best-effort. It does not guarantee in-order delivery, does not guarantee against loss, and does not prevent packets from being duplicated or reordered by multi-path routing fabrics.

Confronted with this chaotic substrate fraught with packet loss, reordering, duplication, and latency jitter, computer scientists conceived one of the greatest software marvels in engineering history: TCP (Transmission Control Protocol).

Through ingenious Sequence Number acknowledgments, bidirectional Sliding Windows, and a rigorous 11-state transition machine, TCP conjures an error-free, non-lossy, deduplicated, and strictly ordered full-duplex byte stream out of a fundamentally unreliable IP substrate.

As Part 5 of our masterclass Computer Networking: From Ethernet Principles to Hyperscale InfiniBand Architecture, this article guides you through Layer 3 and Layer 4: starting from IP routing lookups and CIDR, dissecting the TCP 3-way handshake and 4-way teardown state machines, and unraveling the first principles of sliding windows and flow control.


1. Network Layer (Layer 3): IP Packet Structure and Longest Prefix Match (LPM)

1.1 Physical Anatomy of the IPv4 Header

IPv4 Header Structure (Fixed 20 bytes without Options):
 0                   1                   2                   3
 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|Version|  IHL  |Type of Service|          Total Length         |
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|         Identification        |Flags|      Fragment Offset    |
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|  Time to Live |    Protocol   |         Header Checksum       |
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|                       Source IP Address                       |
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
|                    Destination IP Address                     |
+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+
  1. TTL (Time to Live, 8 bits): Every router hop decrements TTL by 1. When it hits 0, the router drops the packet and transmits an ICMP Time Exceeded message back to the source. This is the fundamental safeguard preventing rogue packets from looping infinitely in cyclic routing topologies;
  2. Identification, Flags, and Fragment Offset: When an IP datagram exceeds the egress link MTU (e.g., 1500 bytes) and the DF (Don't Fragment) bit is 0, the router fragments the payload. All fragments share the same Identification. The receiver reconstructs the original datagram using the Offset. In modern high-throughput networks, fragmentation introduces severe tail-latency amplification and head-of-line blocking; modern stacks enforce PMTUD (Path MTU Discovery) with DF=1 to eliminate network-layer fragmentation entirely;
  3. Protocol (8 bits): Identifies the transport protocol: 0x06 for TCP, 0x11 for UDP, and 0x01 for ICMP.

1.2 Classless Inter-Domain Routing (CIDR) and Longest Prefix Match (LPM)

In the early days of ARPANET and early Internet, addresses were categorized into rigid Classes (Class A, B, C), resulting in astronomical address wastage. In 1993, CIDR (Classless Inter-Domain Routing, e.g., 192.168.1.0/24) decoupled addressing from rigid octet boundaries.

When the Linux kernel or a hardware switch inspects its Forwarding Information Base (FIB), multiple rules may simultaneously match an incoming packet:

  • Route A: 10.0.0.0/8 via eth0
  • Route B: 10.1.0.0/16 via eth1
  • Route C: 10.1.2.0/24 via eth2

Longest Prefix Match (LPM) Rule: When the destination IP is 10.1.2.55, all three routes match, but the kernel always selects Route C (/24) because it has the longest subnet mask and the most specific address scope! In the Linux kernel, this lookup is driven by highly optimized Level-Compressed Trie (LC-Trie) data structures.


2. Transport Layer (Layer 4): The 11 TCP States, Handshake, and Teardown

TCP models network communication as a synchronized state machine between two endpoints. Its complete lifecycle comprises 11 distinct states:

stateDiagram-v2
    [*] --> CLOSED
    CLOSED --> LISTEN: Passive Open (bind and listen)
    CLOSED --> SYN_SENT: Active Open (Send SYN)
    
    LISTEN --> SYN_RCVD: Receive SYN, Send SYN+ACK
    SYN_SENT --> ESTABLISHED: Receive SYN+ACK, Send ACK
    SYN_RCVD --> ESTABLISHED: Receive ACK (Handshake Complete!)
    
    ESTABLISHED --> FIN_WAIT_1: Active Close (Send FIN)
    ESTABLISHED --> CLOSE_WAIT: Passive Close (Receive FIN, Send ACK)
    
    FIN_WAIT_1 --> FIN_WAIT_2: Receive peer ACK
    FIN_WAIT_1 --> CLOSING: Simultaneous Close (Receive peer FIN)
    CLOSE_WAIT --> LAST_ACK: Passive app closed (Send FIN)
    
    FIN_WAIT_2 --> TIME_WAIT: Receive peer FIN, Send ACK
    CLOSING --> TIME_WAIT: Receive ACK
    LAST_ACK --> CLOSED: Receive final ACK
    
    TIME_WAIT --> CLOSED: Wait 2MSL timeout (Socket destroyed)

2.1 Dissecting the 3-Way Handshake: Why Not Two? Why Not Four?

sequenceDiagram
    autonumber
    actor Client as Client
    actor Server as Server

    Note over Server: listen() in LISTEN state
    Client->>Server: 1. SYN: seq=X (Client picks Initial Sequence Number ISN_c)
    Note over Client: Enters SYN_SENT state
    Note over Server: Receives SYN, enters SYN_RCVD state

    Server->>Client: 2. SYN+ACK: seq=Y (Server picks ISN_s), ack=X+1
    Note over Client: Receives SYN+ACK, enters ESTABLISHED state

    Client->>Server: 3. ACK: ack=Y+1, seq=X+1 (May carry application payload)
    Note over Server: Receives ACK, enters ESTABLISHED state
    Note over Client,Server: Bidirectional Full-Duplex Connection Established!

Why Must It Be Exactly Three Packets?

  1. Preventing Stale Duplicate Connection Initialization: Imagine an old, delayed SYN packet traversing a congested router that arrives at the server long after the client has given up. With only a 2-way handshake, the server would immediately allocate memory and consider the connection live. In a 3-way handshake, when the client receives the server's SYN+ACK with an unexpected acknowledgment number, it immediately dispatches an RST packet to abort the invalid connection;
  2. Synchronizing Both Initial Sequence Numbers (ISN): Because TCP is full-duplex, both parties must independently declare their own sequence numbers (XX and YY) and verify that the peer has confirmed them. The client announces XX and receives confirmation (Steps 1 & 2); the server announces YY and receives confirmation (Steps 2 & 3). Bundling the server's ACK and SYN into Step 2 collapses 4 interactions cleanly into 3!

2.2 SYN Queue, Accept Queue, and SYN Flood Mitigation

During the three-way handshake, the Linux kernel manages two queues for each listening socket:

flowchart LR
    SYN_Packet["Incoming Client SYN"] --> SynQueue["SYN Queue (Incomplete Connection Queue)
State: SYN_RCVD
Size governed by tcp_max_syn_backlog"] SynQueue -->|"Final Client ACK received"| AcceptQueue["Accept Queue (Completed Connection Queue)
State: ESTABLISHED
Size = min(somaxconn, listen backlog)"] AcceptQueue -->|"Application invokes accept()"| AppThread["Worker thread processes socket"]

The Lethal Threat: SYN Flood Attack

Attackers spoof millions of random source IPs and flood the server with SYN packets, deliberately withholding the final ACK. The server's SYN Queue exhausts its memory budget within milliseconds, starving legitimate user connections.

The Production Remedy: SYN Cookies

Enabled via /etc/sysctl.conf:

net.ipv4.tcp_syncookies = 1
  • How it works: When the SYN Queue overflows, the kernel refuses to allocate socket buffers (struct sock) in memory;
  • Instead, it hashes the 4-tuple (Client IP, Client Port, Server IP, Server Port) combined with a local secret timestamp to compute a cryptographic token, used directly as the server's Initial Sequence Number YY (the SYN Cookie);
  • Only when the client responds with a valid ACK does the server reconstruct and verify the Cookie, allocating connection resources just-in-time—completely neutralizing memory exhaustion attacks!

2.3 4-Way Teardown and the First Principles of TIME_WAIT

Terminating a TCP connection requires four steps because TCP supports "half-close": when an endpoint sends a FIN, it indicates that it has finished transmitting data, but it remains fully capable of receiving data sent by the peer.

Why Must the Active Closer Linger in TIME_WAIT for 2MSL?

MSL (Maximum Segment Lifetime) represents the longest duration an IP datagram can survive in the network before being dropped (defaulting to 30 seconds on Linux; 2MSL is approximately 60 seconds).

  1. Ensuring Reliable Delivery of the Final ACK: If the client's final ACK packet is dropped by an intermediate router, the server will time out and retransmit its FIN. Remaining in TIME_WAIT allows the client to catch the retransmitted FIN and reissue the ACK. If the client destroyed its socket immediately, the server's retransmitted FIN would trigger an unexpected RST;
  2. Purging Stale Packets from the Network: Dwelling in TIME_WAIT for 2MSL guarantees that any duplicate packets wandering through multi-path routing fabrics will expire and die naturally. This prevents a newly spawned connection reusing the same 4-tuple from inadvertently consuming corrupted, delayed packets from the previous session.

3. Sliding Window and Flow Control

TCP achieves high network throughput through its pipelined sliding window mechanism, abandoning the sluggish stop-and-wait paradigm.

TCP Sender Sliding Window Model:
          Sent & Acked            Sent & Unacked        Usable Window (Unsent)        Cannot Send Yet
        [ ... 1 2 3 4 ]       [ 5 6 7 8 ]        [ 9 10 11 12 ]       [ 13 14 15 ... ]
                               |<---------- Send Window (SND.WND) ---------->|
                               ^                  ^                    ^
                            SND.UNA            SND.NXT              SND.UNA + SND.WND
  1. SND.UNA (Send Unacknowledged): Sequence number of the earliest unacknowledged byte;
  2. SND.NXT (Send Next): Sequence number of the next byte scheduled to be sent;
  3. Advertised Window (win): Announced dynamically by the receiver in every TCP header, indicating the remaining byte capacity in its kernel receive buffer.

3.1 Breaking the 64KB Ceiling: Window Scaling

The legacy TCP header allocated only 16 bits for the window size, capping the maximum advertised window at 216−1=65,535 Bytes2^{16} - 1 = 65,535\,\text{Bytes} (64KB). Over modern high-bandwidth long-delay fiber links (large BDP), a 64KB window cannot even saturate a 100Mbps link!

  • Modern networks negotiate TCP Window Scale (RFC 7323) during the 3-way handshake, bit-shifting the window value up to 14 bits leftward, expanding the effective sliding window ceiling to 1 GB;
  • Production requirement: ensure net.ipv4.tcp_window_scaling = 1 is turned on.

3.2 Silly Window Syndrome and TCP_NODELAY in Practice

In low-latency RPC gateways, financial order routing, and gaming backends, engineers frequently suffer from sudden latency spikes. This is almost always caused by the deadly interaction between Nagle's Algorithm and Delayed ACK:

  • Nagle's Algorithm: Buffers outgoing small packets until a full MSS (Maximum Segment Size) is accumulated or until an outstanding ACK arrives;
  • Delayed ACK: The receiver postpones sending ACKs (by 40ms to 200ms), hoping to piggyback the acknowledgment onto an outgoing response payload;
  • The 40ms Standoff: The sender waits for an ACK, while the receiver waits for more data. Both nodes stall in silence for 40ms!
  • Production Standard: All latency-sensitive TCP sockets (such as gRPC, Redis clients, reverse proxies) must explicitly disable Nagle's algorithm at the socket layer:
    int flag = 1;
    setsockopt(sock, IPPROTO_TCP, TCP_NODELAY, (char *)&flag, sizeof(int));

4. Production-Grade TCP Kernel Parameter Tuning

# 1. Expand Accept Queue and SYN Queue limits
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535

# 2. Allow outbound port reuse for TIME_WAIT sockets (Client-side only)
net.ipv4.tcp_tw_reuse = 1

# 3. Optimize dynamic buffer scaling (min / default / max memory in bytes)
net.ipv4.tcp_rmem = 4096 87380 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216

# 4. Tune keepalive timers to reap dead zombie sockets
net.ipv4.tcp_keepalive_time = 300     # Send probe after 5 minutes of idle
net.ipv4.tcp_keepalive_intvl = 15     # Interval between probes: 15s
net.ipv4.tcp_keepalive_probes = 3     # Disconnect after 3 consecutive failed probes

5. Summary and Next Steps

The design of the Network and Transport Layers represents a masterclass in distributed systems engineering:

  • IP & CIDR establish the global addressing coordinate system, while Longest Prefix Matching (LPM) enables line-rate packet forwarding;
  • TCP 3-Way Handshake & 4-Way Teardown conquer asynchronous transport hazards and rogue duplicate packets using a rigorous 11-state machine;
  • Sliding Windows and Flow Control ensure the sender does not overwhelm the receiver's buffer space.

However, end-to-end flow control only protects the receiver's memory—it possesses zero visibility into the congestion state of intermediate routers and optical switches:

  • When thousands of hosts concurrently pump traffic into the network, switch buffers overflow, causing massive packet loss and throughput collapse;
  • How did congestion control evolve from early loss-driven models (Reno, Cubic) to Google's revolutionary delay-bandwidth model BBR?
  • Why is the industry moving past TCP toward UDP-based HTTP/3 (QUIC) in 2026?

In Chapter 6 of our masterclass, we dive straight into network control theory: TCP Congestion Control Evolution & High-Performance Transport: From Reno and Cubic to BBR Mathematical Models, and HTTP/2 to HTTP/3 (QUIC)!


Frequently Asked Questions (FAQ)

Q1: Why does TCP require three packets to establish a connection, but four packets to close it?

The fundamental reason is the asymmetry between data transmission completion and peer application termination in full-duplex communication.

  • Handshakes can be combined (3 packets): During connection initialization, when the server receives the client's SYN, it can simultaneously bundle its acknowledgment (ACK) and its own connection request (SYN) into a single packet (SYN+ACK), completing setup in 3 flights;
  • Teardowns cannot be combined immediately (4 packets): When the client sends a FIN, it only indicates that the client has no further data to transmit. The server may still have queued data or pending application requests that must be delivered to the client. Thus, the server must immediately respond with an ACK to confirm receipt of the client's FIN, finish processing its workload, and only later issue its own FIN when invoking close(). Because the server's ACK and FIN are separated in time, the teardown requires 4 distinct steps.

Q2: With thousands of sockets stuck in `TIME_WAIT` on a production server, should we enable `tcp_tw_recycle` to reclaim ports rapidly?

Never! Enabling tcp_tw_recycle in modern production environments causes catastrophic, random connection drops for real clients! tcp_tw_recycle relies on monotonically increasing PAWS (Protect Against Wrapped Sequences) timestamps per remote IP. In the real Internet, client traffic almost invariably passes through corporate or mobile NAT gateways. Different client devices behind the same public NAT address exhibit slight clock skews. When tcp_tw_recycle is active, the server interprets packets from a device with a slightly slower clock as stale duplicate segments and drops them silently, causing widespread connection timeouts and 502 errors. Note: Linux kernel 4.12+ has completely removed tcp_tw_recycle. The correct remedy: Enable net.ipv4.tcp_tw_reuse = 1 (which safely reuses outbound client sockets) and deploy reverse proxies (such as Nginx or Envoy) with persistent connection pooling toward backend servers.

Q3: What causes the "40ms latency standoff" between Nagle's algorithm and Delayed ACK, and how can microservices eliminate it?

Nagle's algorithm mandates that a sender buffer subsequent small packets until an ACK for previously transmitted data has arrived. Conversely, the receiver's Delayed ACK timer waits (typically 40ms) before responding, hoping to attach the ACK to an upcoming application reply. When a microservice issues an RPC request consisting of a small header followed by a small payload body, the first chunk triggers Nagle's waiting condition, while the receiver triggers Delayed ACK. Both endpoints sit idle waiting on each other, causing inter-service RPC latency to jump by 40ms even within the same rack! The solution: In all latency-critical TCP services (e.g., gRPC, Redis drivers, Kafka clients, Netty), unconditionally invoke setsockopt(..., TCP_NODELAY) to disable Nagle's algorithm, allowing the network interface card to dispatch small packets without artificial serialization delay.

Related Articles

Start with the same topic, then continue with the latest deep dives.

Production InfiniBand Deployment, Cluster Operations, and NCCL Tuning: OFED Drivers, OpenSM HA, ibdiagnet Fabric Auditing, and GPUDirect RDMA

What bare-metal operational challenges arise when translating network architecture into physical 10,000-GPU AI data centers? A comprehensive guide to production InfiniBand operations: installing and managing Mellanox OFED / DOCA driver stacks and firmware tools (flint/mlxlink); configuring high-availability Master/Standby Subnet Manager topologies with OpenSM; auditing fabric health and diagnosing dirty optical links (Symbol Errors) and speed renegotiation drops using ibdiagnet; and diving into the GPU communication layer to configure GPUDirect RDMA (nvidia-peermem) and tune mission-critical NCCL parameters (NCCL_IB_HCA, NCCL_NET_GDR_LEVEL=5) for wire-speed All-Reduce performance.

InfiniBand AI Cluster Networking in Practice: Fat-Tree Topologies, Rail-Optimized Architecture, Adaptive Routing (AR), and In-Network Reduction (SHARP)

How can thousands of 8-GPU servers be interconnected with tens of thousands of optical links into a non-blocking, deadlock-free high-performance fabric? Deconstruct modern AI supercluster topologies: from Charles Leiserson's 1985 Fat-Tree mathematical model and port count k derivations for 2-Tier and 3-Tier non-blocking ceilings to the 8-plane Rail-Optimized architecture tailored for DGX H100/H200/B200 clusters; analyze how FTree and Up/Down routing engines forbid 'Down-then-Up' turns to eliminate credit loop deadlocks; and discover how hardware Adaptive Routing (AR) and SHARP in-network aggregation achieve a 2x throughput boost during GPU All-Reduce operations.

InfiniBand Architecture First Principles: Physical Link Rates, Credit-Based Link Flow Control, and Subnet Manager Fabric Orchestration

Why does native InfiniBand remain the dominant fabric for 10,000-GPU AI compute clusters and top-tier supercomputers in 2026? Deconstruct the layered InfiniBand protocol stack and physical link evolution: from EDR 100G, HDR 200G, and NDR 400G (Quantum-2) to the 2026 mass production of XDR 800G (Quantum-X800/ConnectX-8); analyze link-layer first principles: Flit-level Credit-Based hardware flow control and sub-100ns Cut-Through switching mechanics; and explore the control engine: Subnet Manager (OpenSM) fabric discovery, dynamic GUID/LID/LMC allocation, and Linear Forwarding Table (LFT) hardware orchestration.

← Prev Computer Networking from Scratch: Bits, Physical Media, Hub Collision Domains, and Switch MAC Addressing First Principles Next → IP Addressing and Subnetting First Principles: Binary Arithmetic, Subnet Masks, CIDR, and Default Gateway Routing
← Back to Articles