Database Storage • 1783 words • 8 min read

Production Databases and Storage on Kubernetes: Local NVMe Passthrough, RocksDB/WAL Tuning, and Fencing Split-Brain Protection

Why do enterprise engineering teams migrating databases to Kubernetes abandon cloud network block storage in favor of bare-metal Local NVMe SSDs? A comprehensive exploration of running stateful distributed storage on Kubernetes: overcoming cloud disk multi-hop latency bottlenecks with Local Persistent Volumes; XFS formatting tuning and resolving cgroups v2 OOM cascades caused by Linux page cache dirty page backlogs; RocksDB/MySQL WAL Direct I/O and fdatasync optimizations; and mitigating split-brain dual-primary corruption using hardware IPMI STONITH and lease-based Node Fencing.

Container & K8s Series Part 8 / 13

From Docker to Kubernetes: Cloud-Native Container & Cluster Orchestration Handbook

Master containerization and cluster orchestration from first principles: Linux namespaces, cgroups v2, SwarmKit, K8s control plane; real-world software architecture on K8s (gRPC/zero-downtime, Local NVMe/RocksDB/fencing); hacking K8s internals (Scheduling Framework plugins, NRI runtimes, Aggregated APIServer); and 2026 AI GPU scheduling.

Browse all 13 chapters in this series ▾
  1. 01 Docker Internals from First Principles: Linux Namespaces, cgroups v2, and OverlayFS Union Mounts Deep Dive
  2. 02 Container Networking Deep Dive: veth-pair, Linux Bridge, iptables NAT, and Cross-Host Topology
  3. 03 Lightweight Cluster Orchestration: Docker SwarmKit Architecture, Raft Consensus, and Ingress Routing Mesh
  4. 04 Kubernetes Control Plane Deep Dive: Declarative APIs, etcd Consensus, Scheduler, and Controller Reconciliation Loops
  5. 05 Kubernetes Networking Panorama: CNI Specification, Calico BGP, Cilium eBPF, and Gateway API Architecture
  6. 06 Kubernetes Storage Architecture: CSI Specification, Dynamic PV/PVC Provisioning, and StatefulSet Guarantees
  7. 07 Application Networking on Kubernetes: The gRPC Load Balancing Trap, Service Mesh, and Zero-Downtime Draining Sequences
  8. 08 Production Databases and Storage on Kubernetes: Local NVMe Passthrough, RocksDB/WAL Tuning, and Fencing Split-Brain Protection Reading
  9. 09 Hacking the Kubernetes Scheduler: Custom Plugins via Scheduling Framework, Volcano DRF Math, and Descheduler Dynamic Rebalancing
  10. 10 Hacking Kubernetes Nodes & Runtimes: Breaking PLEG Bottlenecks, NRI Plugins, Kata/gVisor Sandboxes, and cgroups v2 Tuning
  11. 11 Hacking the Kubernetes Control Plane: WatchCache Internals, Aggregated APIServers, APF Shuffle Sharding, and 10k-Node etcd Sharding
  12. 12 Kubernetes Extensibility: The Operator Pattern, Custom Resource Definitions (CRD), and KubeBuilder in Production
  13. 13 2026 AI Computing Infrastructure: Kubernetes GPU Operator, MIG Partitioning, Topology-Aware Scheduling, and Auto-scaling Inference Engines

Introduction: The Performance Deadlines and Durability Floors of Kubernetes Databases

In Chapter 06, Kubernetes Storage Architecture, we broke down the CSI plugin specification and StatefulSet invariants.

However, in demanding production database deployments, teams that configure a generic cloud StorageClass (such as AWS gp3 or standard managed cloud disks) to deploy MySQL, PostgreSQL, TiDB, Kafka, or Elasticsearch regularly encounter two operational traps:

  1. The Performance Cliff: Write latency P99 degrades from bare-metal 1ms levels up to 15ms. Under peak ingestion loads, disk I/O queues saturate, exhausting connection pools and cascading into microservice outages;
  2. Silent Data Corruption (Split-Brain): A physical host suffers a transient network blip. The Kubernetes scheduler instantiates a replacement primary Pod on another node; however, because the isolated original node continues writing to its local filesystem without clean unmounting, both instances accept writes, causing split-brain data corruption!

Datastores represent the foundational persistence tier of modern software. They cannot merely be containerized passively; their underlying kernel interactions, disk channels, and failover topologies must be re-engineered for cloud-native infrastructure.

As the eighth chapter of From Docker to Kubernetes: Cloud-Native Container & Cluster Orchestration Handbook, this guide deconstructs distributed database engineering on Kubernetes: Local NVMe SSD passthrough with optimized XFS parameters, Linux Page Cache and WAL flush tuning, and Node Fencing / IPMI STONITH split-brain mitigation.


1. The Storage Hierarchy: Why High-Performance Databases Mandate Local PVs

Storage architectures in Kubernetes fall into three distinct tiers:

flowchart TD
    subgraph NetStorage["1. Network Block Storage (EBS / Ceph / Managed Disks)"]
        direction LR
        App1["Database Pod"] --> VFS1["Linux VFS Subsystem"]
        VFS1 --> TCP1["TCP/IP Kernel Encapsulation"]
        TCP1 --> NIC1["Host NIC"]
        NIC1 -->|"2~3 Network Switch Hops"| SAN["Remote Storage Cluster / SAN Controller"]
        SAN --> Disk1["Remote Physical Disk (P99: 3~15 ms)"]
    end

    subgraph LocalStorage["2. Local Hardware Passthrough (Local Persistent Volumes)"]
        direction LR
        App2["Database Pod"] --> VFS2["Linux VFS (XFS / O_DIRECT)"]
        VFS2 --> PCIe["Host PCIe 5.0 / NVMe Controller"]
        PCIe --> Disk2["Local Enterprise NVMe SSD (P99: 50~100 μs)"]
    end
Storage ParadigmReference ImplementationsHardware Transit PathRandom Write P99 LatencyFailure Recovery ModelRecommended Workload
Network Block StorageAWS EBS (gp3/io2), Cloud DisksHost NIC →\to VPC Fabric →\to Storage Controller1 ~ 5 msUpon host failure, cloud disk detaches and re-attaches to new nodeMedium-tier web datastores without sub-millisecond SLA requirements
Distributed Network FilesystemCephFS, NFS, PortworxHost Kernel →\to FUSE / Network Driver →\to Multi-replica network3 ~ 15 msSoftware storage layer handles replicationStatic assets, shared media, multi-reader Pod scenarios
Local Persistent Volume (Local PV)Host PCIe NVMe SSDs (Direct Passthrough)Host PCIe bus direct, zero network encapsulation50 ~ 100 μs (20x~50x lower latency!)Rebuilt via application-layer consensus (Raft/Paxos)Mission-Critical Datastores (TiDB, MySQL, Kafka, ClickHouse)

Architectural Trade-Offs:

  • Network block storage provides 3x storage-layer redundancy, but introduces significant network transit hop latency and bandwidth throttling;
  • Modern distributed datastores (TiDB TiKV, Apache Kafka, Elasticsearch) already implement application-level replication protocols (such as Raft or ISR);
  • Running an application-level Raft database atop an underlying 3x replicated storage SAN results in double-replication amplification, compounding write amplification and network saturation;
  • Consequently, high-scale database architectures adopt Local Persistent Volumes (Local PV), delegating durability and failover entirely to application-level consensus algorithms.

2. Production Local PV Architecture & XFS Filesystem Tuning

Using Local PVs in Kubernetes requires avoiding raw hostPath volumes. Dedicated, node-pinned Local PersistentVolumes must be provisioned.

2.1 Disk Initialization: Tuning XFS Formatting

When formatting physical NVMe drives, optimize XFS parameters to prevent metadata write lock contention:

# Format enterprise NVMe SSD
mkfs.xfs -f \
  -n ftype=1 \          # Mandatory: ftype=1 (Prerequisite for container runtimes and OverlayFS)
  -l size=128m \        # Enlarge log buffer to 128MB to eliminate journal lock contention during WAL flushes
  -d agcount=32 \       # Increase allocation groups to 32 to maximize parallel write allocation
  /dev/nvme0n1

# Production mount parameters
mkdir -p /mnt/disks/nvme0n1
mount -o noatime,nodiratime,allocsize=64M,logbufs=8,logbsize=256k /dev/nvme0n1 /mnt/disks/nvme0n1
  • noatime,nodiratime: Disables file access timestamp updates, eliminating up to 30% of unnecessary disk writes;
  • allocsize=64M: Extends buffer pre-allocation, substantially reducing large WAL file fragmentation.

2.2 Local PV Manifest Declaration

apiVersion: v1
kind: PersistentVolume
metadata:
  name: local-nvme-pv-node1
spec:
  capacity:
    storage: 1.6Ti
  volumeMode: Filesystem
  accessModes:
  - ReadWriteOnce
  persistentVolumeReclaimPolicy: Retain
  storageClassName: local-nvme-sc
  local:
    path: /mnt/disks/nvme0n1     # Mount point of dedicated NVMe block device
  nodeAffinity:
    required:
      nodeSelectorTerms:
      - matchExpressions:
        - key: kubernetes.io/hostname
          operator: In
          values:
          - k8s-worker-node-1     # Physical node pinning
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: local-nvme-sc
provisioner: kubernetes.io/no-provisioner
# Mandatory requirement: defer volume binding!
volumeBindingMode: WaitForFirstConsumer

3. Kernel and Storage Engine Tuning

Running databases within containers introduces operational risks regarding Linux Page Cache dirty page accumulation and cgroups v2 OOM kills.

3.1 The Memory Trap: Page Cache and Container Eviction

Under default Linux kernel settings, file write operations cache within the operating system Page Cache as "dirty pages," asynchronously flushed by background pdflush/flusher threads.

  • Under cgroups v2, memory allocated to the Page Cache counts directly toward the container's memory.current consumption;
  • If the database ingests data faster than the physical medium can flush, Page Cache buffers surge, reaching the container's resources.limits.memory;
  • The Linux kernel triggers synchronous writeback (Direct Reclaim), freezing worker threads for several seconds or terminating the database process via the OOMKiller!
Host and Container Kernel Tuning Guidelines (/etc/sysctl.conf):
# 1. Lower threshold for background asynchronous dirty page flushing (from 10% to 5%)
vm.dirty_background_ratio = 5

# 2. Bound maximum dirty page volume to prevent IO stalls (cap at 10%)
vm.dirty_ratio = 10

# 3. Shorten dirty page expiration window (centiseconds: drop from 3000 to 500 = 5 seconds)
vm.dirty_expire_centisecs = 500

3.2 Bypassing the Page Cache: Direct I/O (`O_DIRECT`) and `fdatasync`

For consistent throughput across write-heavy engines (MySQL InnoDB, RocksDB WAL engines), applications should enable Direct I/O (O_DIRECT):

  • MySQL Configuration: innodb_flush_method = O_DIRECT;
  • RocksDB Configuration: use_direct_io_for_flush_and_compaction = true, use_direct_reads = true;
  • WAL Flush Optimization: Replace standard fsync() with fdatasync(). While fsync() forces synchronous metadata updates (triggering disk controller stalls), fdatasync() flushes only raw data blocks, delivering up to 40% higher throughput;
  • Data transfers bypass the OS Page Cache entirely, eliminating double-buffering memory overhead and guarding against container OOM kills.

4. Split-Brain Mitigation: Node Fencing Architecture

In distributed architectures, few failures cause more data corruption than unmitigated split-brain conditions resulting from network partitions:

sequenceDiagram
    autonumber
    participant NodeA as Node A (Original Primary MySQL)
    participant K8sCtrl as K8s Control Plane / Scheduler
    participant NodeB as Node B (Newly Promoted Primary)
    participant Client as Application Traffic

    Note over NodeA: Node A experiences asymmetric network split (Cannot reach API Server, but local clients can reach it)
    K8sCtrl->>K8sCtrl: Heartbeat timeout: Mark Node A as NotReady
    K8sCtrl->>NodeB: Launch and promote replacement primary instance on Node B
    
    Critical The Split-Brain Disaster!
        Client->>NodeB: Client issues write to Node B (Generates Transaction ID 101)
        Client->>NodeA: Network-partitioned client issues write to Node A (Generates Transaction ID 101)
        Note over NodeA,NodeB: Both nodes write conflicting transactions to storage, consistency destroyed!
    end

4.1 Principles of Node Fencing

Mitigating split-brain scenarios requires enforcing Fencing: "Before a replacement primary node is permitted to accept write transactions, the system must guarantee the prior primary is incapacitated!"

Three tiers of Fencing mechanisms are deployed in production:

  1. Storage Fencing (SCSI-3 PR / CSI Locks): Leverages SCSI-3 Persistent Reservation locks or cloud VolumeAttachment constraints. Before assuming primary duties, the replacement node revokes the previous node's access token at the storage controller level. Any subsequent write attempts by the isolated node encounter kernel I/O errors.
  2. Node Fencing (STONITH - Shoot The Other Node In The Head): When the control plane flags Node A as failed, automated agents issue power-cutoff commands via server out-of-band management interfaces (IPMI, iLO, or Redfish APIs), cutting power to the chassis.
  3. Software Lease Fencing: The primary process maintains a short-lived distributed lease in memory (e.g., a Kubernetes Lease object or Raft consensus lease refreshed every 2 seconds). If the primary loses contact with the cluster and fails to renew the lease within 5 seconds, it must execute an immediate process abort/panic, halting disk write routines.

5. Summary and Transition

Operating distributed databases on Kubernetes requires systems-level discipline:

  • Storage Selection: High-throughput workloads should prioritize Local PV NVMe passthrough, using application-level consensus to handle hardware redundancy;
  • Filesystem & Kernel Tuning: Formatting with mkfs.xfs -l size=128m eliminates journal lock contention, while tuning dirty page parameters and enabling O_DIRECT protects against cgroups v2 OOM cascades;
  • Split-Brain Defense: Enforcing Node Fencing, IPMI STONITH chassis power cutoffs, and lease-based heartbeat locks shields databases against split-brain corruption during network partitions.

However, as workloads scale beyond basic microservices and databases, platform architects confront internal control-plane bottlenecks:

  • Why does the standard kube-scheduler struggle with batch AI jobs, causing deadlocks in distributed training workloads?
  • How can platform engineers extend the Kubernetes scheduling core using the official Scheduling Framework?
  • What are the design principles of Gang Scheduling, and how does the Descheduler maintain runtime cluster balance?

In Chapter 09, Hacking the Kubernetes Scheduler: Custom Plugins via Scheduling Framework, Gang Scheduling, and Descheduler Dynamic Rebalancing, we customize the Kubernetes scheduler!


Frequently Asked Questions (FAQ)

Q1: Why do distributed systems like TiDB, Kafka, and Elasticsearch prefer Local PVs over network block storage (Ceph or AWS EBS)?

Core Reason: Eliminates Double-Replication Overhead and Delivers Low Latency. These distributed platforms implement data redundancy at the application layer via Raft, ISR, or primary-replica topologies. Running them atop 3x replicated network block storage multiplies every write operation into two network transit phases and nine separate disk writes. Local PVs leverage direct host PCIe buses, reducing latencies to microseconds while eliminating network transfer charges.

Q2: Why is configuring `-l size=128m` critical when formatting XFS volumes for database workloads?

XFS maintains an internal journaling area for filesystem transactions. By default, this log buffer is restricted to a few megabytes. Under high-frequency concurrent writes (such as RocksDB WAL flushes or MySQL Binlog writes), multiple threads concurrently submitting transaction commits saturate the internal XFS log ticket lock, freezing threads in kernel space. Expanding the log buffer to 128MB provides sufficient queuing capacity to eliminate up to 80% of metadata lock contention.

Q3: What is Node Fencing, and why are Kubernetes Liveness Probes insufficient to prevent database split-brain conditions?

Liveness probes evaluate process health locally within a node; they cannot distinguish between container failure and an asymmetric network partition. If a primary node becomes isolated from the control plane while remaining reachable by local clients, promoting a secondary node without fencing causes both instances to accept writes simultaneously. Node Fencing guarantees that the original node is cut off from storage or power before the new primary is promoted.

Related Articles

Start with the same topic, then continue with the latest deep dives.

Production InfiniBand Deployment, Cluster Operations, and NCCL Tuning: OFED Drivers, OpenSM HA, ibdiagnet Fabric Auditing, and GPUDirect RDMA

What bare-metal operational challenges arise when translating network architecture into physical 10,000-GPU AI data centers? A comprehensive guide to production InfiniBand operations: installing and managing Mellanox OFED / DOCA driver stacks and firmware tools (flint/mlxlink); configuring high-availability Master/Standby Subnet Manager topologies with OpenSM; auditing fabric health and diagnosing dirty optical links (Symbol Errors) and speed renegotiation drops using ibdiagnet; and diving into the GPU communication layer to configure GPUDirect RDMA (nvidia-peermem) and tune mission-critical NCCL parameters (NCCL_IB_HCA, NCCL_NET_GDR_LEVEL=5) for wire-speed All-Reduce performance.

InfiniBand AI Cluster Networking in Practice: Fat-Tree Topologies, Rail-Optimized Architecture, Adaptive Routing (AR), and In-Network Reduction (SHARP)

How can thousands of 8-GPU servers be interconnected with tens of thousands of optical links into a non-blocking, deadlock-free high-performance fabric? Deconstruct modern AI supercluster topologies: from Charles Leiserson's 1985 Fat-Tree mathematical model and port count k derivations for 2-Tier and 3-Tier non-blocking ceilings to the 8-plane Rail-Optimized architecture tailored for DGX H100/H200/B200 clusters; analyze how FTree and Up/Down routing engines forbid 'Down-then-Up' turns to eliminate credit loop deadlocks; and discover how hardware Adaptive Routing (AR) and SHARP in-network aggregation achieve a 2x throughput boost during GPU All-Reduce operations.

InfiniBand Architecture First Principles: Physical Link Rates, Credit-Based Link Flow Control, and Subnet Manager Fabric Orchestration

Why does native InfiniBand remain the dominant fabric for 10,000-GPU AI compute clusters and top-tier supercomputers in 2026? Deconstruct the layered InfiniBand protocol stack and physical link evolution: from EDR 100G, HDR 200G, and NDR 400G (Quantum-2) to the 2026 mass production of XDR 800G (Quantum-X800/ConnectX-8); analyze link-layer first principles: Flit-level Credit-Based hardware flow control and sub-100ns Cut-Through switching mechanics; and explore the control engine: Subnet Manager (OpenSM) fabric discovery, dynamic GUID/LID/LMC allocation, and Linear Forwarding Table (LFT) hardware orchestration.

← Prev 2026 AI Computing Infrastructure: Kubernetes GPU Operator, MIG Partitioning, Topology-Aware Scheduling, and Auto-scaling Inference Engines Next → Hacking the Kubernetes Control Plane: WatchCache Internals, Aggregated APIServers, APF Shuffle Sharding, and 10k-Node etcd Sharding
← Back to Articles