Inference Engines • 2282 words • 10 min read

Production LLM Inference Engines Teardown: vLLM v1 vs. SGLang vs. TensorRT-LLM vs. llama.cpp Internals

From Python scheduling to asynchronous C++/Rust execution engines: an in-depth codebase teardown of the open-source world's top four inference engines. Comparing vLLM v1 zero-overhead multi-step scheduling, SGLang RadixAttention tree-based execution loops, TensorRT-LLM fused kernel pipelines with in-flight batching, and llama.cpp zero-dependency edge quantization, with an enterprise engineering decision matrix.

Inference Engine Series Part 9 / 10

Production LLM Inference Engines: Evolution, Architecture, and Future

Master LLM inference systems from first principles: Roofline modeling of Prefill/Decode, KV Cache & MLA memory optimization, PagedAttention, Continuous Batching, P/D disaggregation, and in-depth engine architectures across vLLM, SGLang, TensorRT-LLM, and llama.cpp.

Browse all 10 chapters in this series ▾

Introduction: From Toy Scripts to Enterprise-Grade Serving Runtimes

In the early days of the generative AI boom in 2023, most developers initiated LLM inference on a single GPU using a simple Python loop:

# The Naive PyTorch Autoregressive Loop
for _ in range(max_new_tokens):
    outputs = model(input_ids)
    next_token = torch.argmax(outputs.logits[:, -1, :], dim=-1)
    input_ids = torch.cat([input_ids, next_token.unsqueeze(-1)], dim=-1)

While clean and educational, running raw PyTorch loops in high-concurrency production deployments proves disastrous:

  1. Static Memory Allocations Waste 60%∼80%60\% \sim 80\% of GPU Memory: Memory fragmentation prevents handling concurrent user requests;
  2. Padding Inefficiencies: Uneven request lengths waste substantial compute cycles processing meaningless padding tokens;
  3. Severe Python GIL and Scheduling Latency: In autoregressive generation where single-step GPU kernel execution completes in milliseconds, Python interpretation overhead starves the GPU, leaving Tensor Cores idle;
  4. Memory Bandwidth Choke: Independent kernel invocations bounce tensors between High Bandwidth Memory (HBM) and SRAM, plunging Model FLOPs Utilization (MFU) below 15%15\%.

To extract maximum performance from modern accelerators, open-source communities and hardware vendors engineered four industry-standard inference engines:

  • vLLM (UC Berkeley / LMSYS): The pioneer that standardized modern inference via PagedAttention memory virtualization, now fundamentally overhauled with its vLLM v1 C++ engine core;
  • SGLang (LMSYS / Stanford): Tailored for multi-turn agentic interactions, pioneering RadixAttention tree-based caching with low-overhead execution loops;
  • TensorRT-LLM (NVIDIA): NVIDIA's high-performance powerhouse, deeply fusing CUTLASS and FlashAttention kernels with static engine graphs for peak hardware throughput;
  • llama.cpp (Georgi Gerganov): The minimalist Swiss Army knife built in pure zero-dependency C/C++, introducing the GGUF container standard and dominating edge and on-device computing.

In this ninth chapter of the Production LLM Inference Engines series, we bypass public benchmarks to dissect the architectural internals, scheduler execution loops, and underlying design philosophies of these four runtimes, concluding with an enterprise decision matrix.


1. Architectural Philosophy and Top-Level Topology

To evaluate these runtimes fairly, one must first recognize their contrasting first-principle objectives:

flowchart TD
    subgraph ArchComparison["Core Topological Comparison of Leading Inference Engines"]
        subgraph VLLM["vLLM v1 Architecture"]
            V_IN["HTTP / gRPC Ingress"] --> V_ENG["C++ Async Engine Core (Zero-Overhead)"]
            V_ENG --> V_SCHED["Multi-Step Batch Scheduler"]
            V_SCHED --> V_RUN["Worker Runner (CUDA Graph + PagedAttn)"]
        end

        subgraph SGL["SGLang Architecture"]
            S_IN["Complex Multi-Turn Prompts / Agents"] --> S_ROUTER["RadixTree Global Router"]
            S_ROUTER --> S_TREE["RadixAttention LRU Tree State Machine"]
            S_TREE --> S_RUN["SGL-Kernel + FlashInfer Backend"]
        end

        subgraph TRT["TensorRT-LLM Architecture"]
            T_IN["High-Concurrency Enterprise Workloads"] --> T_MGR["C++ GptManager (In-Flight Batching)"]
            T_MGR --> T_FUSED["CUTLASS / FMHA Fully Fused Graphs"]
            T_FUSED --> T_GPU["NVIDIA Tensor Core Bare-Metal Execution"]
        end

        subgraph LLAMA["llama.cpp Architecture"]
            L_IN["Edge / CPU / Personal Devices"] --> L_CORE["ggml-backend Abstraction Layer"]
            L_CORE --> L_QUANT["GGUF On-the-Fly Dequantization (Q4_K, Q2_K)"]
            L_CORE --> L_METAL["Metal / Vulkan / AVX-512 SIMD Execution"]
        end
    end
DimensionvLLM v1SGLangTensorRT-LLMllama.cpp
Primary GoalHigh-throughput, universally extensible standard ecosystemMulti-turn agent loops, structured decoding, and prefix reusePeak hardware throughput and energy efficiency on NVIDIA siliconZero-dependency, ultra-quantized portability across edge devices
Language StackC++ core scheduler + Python ecosystem extensionsPython scheduler + C++/CUDA kernels (SGL-Kernel)Pure C++ engine runtime + Python build-time declarationsPure C/C++ (Zero external library dependencies)
Memory ArchitecturePagedAttention virtual block memory poolRadixAttention tree-structured prefix caching poolStatic pre-allocated flat memory + In-Flight Paged KVMemory-mapped (mmap) unified memory with split tensors
Target HardwareNVIDIA GPU, AMD ROCm, Intel Gaudi, Ascend NPUNVIDIA GPU, AMD ROCmNVIDIA Silicon Exclusively (Hopper, Blackwell, Ada)Apple Silicon, x86 CPU, ARM Mobile, embedded SoCs

2. vLLM v1: Escaping the Python Bottleneck with C++ Core Engines

2.1 The Latency Penalty of the v0 Architecture

In early vLLM v0 iterations, despite revolutionary memory improvements via PagedAttention, the scheduler operated within Python's Global Interpreter Lock (GIL).

On modern accelerators (e.g., H100/H800), a single autoregressive decoding step executes in 1.5∼3 ms1.5 \sim 3 \text{ ms}. However, the Python scheduling loop consumed 2∼4 ms2 \sim 4 \text{ ms} of CPU time per step updating sequence metadata, checking block allocations, and dispatching tensors. The GPU continually stalled waiting for the CPU scheduler, bottlenecking Model FLOPs Utilization (MFU).

2.2 Deep Dive into v1: Zero-Overhead C++ Schedulers and Multi-Step Execution

The vLLM v1 redesign migrated the core scheduler and request lifecycle managers into a high-performance C++ async engine, introducing Multi-Step Scheduling.

In traditional step-by-step engines, CPU and GPU synchronize after every single generated token:

Traditional Step-by-Step Execution:
[CPU: Schedule Step 1] -> [GPU: Kernel Launch] -> [CPU Wait & Sync] -> [CPU: Schedule Step 2] ...

In vLLM v1, the scheduler pre-checks memory capacity and dispatches execution plans for the next NN steps (e.g., N=8N=8 or 1616) in a single batch:

// vLLM v1 Core Scheduling Decision Logic (Scheduler::schedule)
SchedulerOutputs Scheduler::schedule() {
    // 1. Prioritize active decoding sequences in the Running Queue
    std::vector<SequenceGroup*> running_batch;
    for (auto& seq_group : running_queue_) {
        // Pre-allocate blocks for the next num_steps iterations
        if (block_allocator_->can_allocate_multi_steps(seq_group, num_steps_)) {
            block_allocator_->allocate_for_multi_steps(seq_group, num_steps_);
            running_batch.push_back(seq_group);
        } else {
            // Memory pressure threshold reached: trigger preemptive swap-out
            preempt(seq_group);
        }
    }

    // 2. Consume remaining token budget with pending prefill requests
    std::vector<SequenceGroup*> prefill_batch;
    size_t remaining_token_budget = max_num_batched_tokens_ - get_num_tokens(running_batch);
    while (!waiting_queue_.empty() && remaining_token_budget > 0) {
        auto next_req = waiting_queue_.front();
        if (can_prefill(next_req, remaining_token_budget)) {
            prefill_batch.push_back(next_req);
            remaining_token_budget -= next_req->get_prompt_len();
            waiting_queue_.pop_front();
        } else {
            break;
        }
    }

    return SchedulerOutputs(running_batch, prefill_batch, num_steps_);
}

By executing multiple iterations directly on the GPU within persistent CUDA Graph streams, CPU scheduling overhead is amortized across NN steps. This delivers a 30%∼50%30\% \sim 50\% throughput increase under low-batch conditions.


3. SGLang: Tree-Structured Prefix State Machines for Fast Agent Loops

Where vLLM excels as a general-purpose serving platform, SGLang specializes in multi-turn agent workflows, extensive few-shot prompts, and structured output generation.

3.1 RadixAttention Codebase Internals

As explored conceptually in Chapter 04 on Prefix Caching and RadixAttention, SGLang models KV Cache memory as a dynamic Radix Tree:

# SGLang RadixTree Core Matching and Eviction Logic
class RadixTree:
    def __init__(self):
        self.root = TreeNode(key_tokens=[])
        self.eviction_lru = DoublyLinkedList()

    def match_prefix(self, token_ids: List[int]) -> Tuple[List[int], TreeNode]:
        """
        Locates the longest shared prefix in O(K) time complexity, where K is prefix depth
        """
        curr = self.root
        matched_blocks = []
        idx = 0
        while idx < len(token_ids):
            child = curr.find_child(token_ids[idx])
            if child is None:
                break
            matched_len = child.match_prefix(token_ids[idx:])
            matched_blocks.extend(child.get_block_ids(matched_len))
            idx += matched_len
            if matched_len < len(child.key_tokens):
                # Partial match identified: split tree node
                break
            curr = child
            
        # Update LRU cache recency
        self.eviction_lru.touch(curr)
        return matched_blocks, curr

3.2 Tight Integration with SGL-Kernel and FlashInfer

SGLang pairs this tree data structure directly with optimized custom kernels:

  • In-Kernel Constrained Sampling: For structured JSON generation, SGLang tracks grammar state machines directly within C++ logit processors, avoiding costly host-device transfers of full vocabulary tensors;
  • Zero-Copy Branching: In tree search applications (MCTS or multi-agent voting), child branches share memory block references with ancestors using reference counting, performing Copy-on-Write only when tokens diverge.

4. TensorRT-LLM: Extreme Kernel Fusion on NVIDIA Silicon

For deployments exclusively targeting modern NVIDIA data center hardware (e.g., 8x H100/H800/B300 servers with NVSwitch fabrics), TensorRT-LLM provides unparalleled compute efficiency.

4.1 In-Flight Batching (IFB) and Pre-Compiled Computation Graphs

Rather than relying on dynamic graph generation, TensorRT-LLM compiles model architectures into static .engine binaries tailored to specific GPU models and tensor-parallel widths:

  • A pure C++ GptManager controls In-Flight Batching (IFB), decoupling layer execution into concurrent CUDA streams to hide inter-GPU communication latency;
  • Memory allocation maps onto pre-computed flat arenas, virtually eliminating runtime allocation jitter.

4.2 Extreme Kernel Fusion Pipelines

TensorRT-LLM achieves peak throughput by eliminating intermediate memory reads and writes:

Standard Transformer Layer Execution:
[RMSNorm] --(Write HBM)--> [QKV GEMM] --(Write HBM)--> [RoPE] --(Write HBM)--> [FlashAttention] ...

TensorRT-LLM Deeply Fused Pipeline:
┌────────────────────────────────────────────────────────────────────────┐
│ Fused RMSNorm + QKV Projection + RoPE (Executed in L1/SRAM registers)  │
└───────────────────────────────────┬────────────────────────────────────┘
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│ Fused Paged-KV Attention + Output Projection Linear Kernel             │
└───────────────────────────────────┬────────────────────────────────────┘
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│ Fused SwiGLU MLP (Gate + Up Projection + Activation fully integrated)  │
└────────────────────────────────────────────────────────────────────────┘

By leveraging NVIDIA CUTLASS, RMSNorm, QKV GEMM, and RoPE positional encoding execute within a single fused kernel. Intermediate activations remain entirely within streaming multiprocessor (SM) register files and shared memory, bypassing HBM entirely.

Trade-off: Modifying non-standard architectures requires authoring custom C++ TensorRT plugins and re-running extended compilation pipelines.


5. llama.cpp: The Minimalist Swiss Army Knife for Edge Runtimes

While enterprise engines focus on multi-GPU server clusters, llama.cpp takes a completely different path: enabling large language models to run efficiently on commodity CPUs, mobile devices, and consumer laptops.

5.1 Pure C/C++ Architecture with Zero External Dependencies

Authored by Georgi Gerganov, llama.cpp requires no PyTorch, no CUDA runtime, and no external package managers:

  • The core tensor library, ggml, maps computations directly to hardware SIMD vector intrinsics;
  • It utilizes AVX-512 and AVX2 on x86 CPUs;
  • On Apple Silicon, it invokes Metal Performance Shaders (MPS), achieving direct zero-copy memory reads via unified memory architectures (UMA);
  • On mobile and embedded platforms, it supports Vulkan, OpenCL, and ARM NEON.

5.2 The GGUF Format and Non-Uniform k-Quants

llama.cpp standardized the universal GGUF container format, packaging model weights, tokenizers, hyperparameter configs, and quantization metadata into a single self-contained binary.

Furthermore, its custom k-quants block-quantization algorithms enable aggressive compression:

llama.cpp k-quants Precision Hierarchy:
┌───────────────────────────────────────────────────────────────┐
│ Q4_K_M: 4.5-bit attention layers with 4-bit feed-forward grids│
│ Q3_K_S: 3-bit compression reducing memory footprint by ~75%   │
│ Q2_K  : 2-bit quantization running 70B models on 24GB GPUs    │
└───────────────────────────────────────────────────────────────┘

Weights are dequantized on-the-fly during matrix-vector multiplications using fast lookup tables, enabling multi-billion parameter models to execute smoothly on personal workstations.


6. Enterprise Production Decision Matrix: Which Engine to Choose?

When deploying production infrastructure, system architects must select runtimes aligned with operational realities:

flowchart TD
    START{"What is your deployment target and workload profile?"}
    
    START --> C1["Personal devices / Edge hardware / Mac Studio / On-device apps"]
    C1 ==> R_LLAMA["Select: llama.cpp (GGUF quantization, zero dependencies)"]
    
    START --> C2["Homogeneous NVIDIA clusters (8x H100/H800/B300), maximizing pure throughput"]
    C2 ==> R_TRT["Select: TensorRT-LLM (Fused kernels, peak hardware MFU)"]
    
    START --> C3["Complex multi-turn agents / Structured JSON outputs / High prefix sharing"]
    C3 ==> R_SGL["Select: SGLang (RadixTree caching, lightweight execution)"]
    
    START --> C4["General-purpose Model-as-a-Service (MaaS) / Multi-hardware (AMD, Ascend)"]
    C4 ==> R_VLLM["Select: vLLM v1 (C++ core, broadest model ecosystem)"]

Comprehensive Technical Evaluation Matrix (1 to 5 Stars)

Evaluation CriterionvLLM v1SGLangTensorRT-LLMllama.cpp
Peak Throughput Capacity⭐️⭐️⭐️⭐️☆⭐️⭐️⭐️⭐️☆⭐️⭐️⭐️⭐️⭐️ (Industry Peak)⭐️⭐️☆☆☆
Time to First Token (TTFT)⭐️⭐️⭐️⭐️☆⭐️⭐️⭐️⭐️⭐️ (Superb Prefix Reuse)⭐️⭐️⭐️⭐️☆⭐️⭐️⭐️☆☆
Agent / Multi-Turn Suitability⭐️⭐️⭐️☆☆⭐️⭐️⭐️⭐️⭐️ (Optimal Choice)⭐️⭐️☆☆☆⭐️⭐️⭐️☆☆
Model Ecosystem Coverage⭐️⭐️⭐️⭐️⭐️ (Fastest Community Adoption)⭐️⭐️⭐️⭐️☆⭐️⭐️☆☆☆ (Official Ports Only)⭐️⭐️⭐️⭐️☆
Heterogeneous Silicon (AMD/Intel/ARM)⭐️⭐️⭐️⭐️⭐️ (Widest Accelerator Support)⭐️⭐️⭐️☆☆⭐️☆☆☆☆ (Strictly NVIDIA)⭐️⭐️⭐️⭐️⭐️ (Runs Anywhere)
Custom Extensibility & DevOps Overhead⭐️⭐️⭐️⭐️☆ (Well Documented)⭐️⭐️⭐️⭐️☆ (Lightweight Codebase)⭐️☆☆☆☆ (Steep Compilation Curve)⭐️⭐️⭐️⭐️⭐️ (Clone & Compile)

Summary and Series Finale Outlook

From vLLM v1's C++ zero-overhead scheduler to SGLang's dynamic Radix Trees, TensorRT-LLM's fused kernel pipelines, and llama.cpp's edge minimalism, modern inference runtimes represent specialized trade-offs across the compute landscape.

With our examination of kernel architectures, batching schedulers, and runtime engines complete, we arrive at the frontier of serving systems. In the tenth and final chapter of this series, we investigate: The Future of LLM Inference: Test-Time Compute, Reasoning Scheduling, and Tiered Cluster Storage, exploring how long reasoning chains (16k~128k generated thinking tokens) will redefine memory hierarchies and speculative decoding schedulers.


Frequently Asked Questions (FAQ)

Q1: Why do enterprise Model-as-a-Service (MaaS) platforms often deploy both vLLM and SGLang concurrently?

Deploying both represents a practical division of labor based on workload characteristics:

  • vLLM v1 serves as the universal gateway: It powers general single-turn Q&A, document summarization, code generation, and long-tail open-source models, leveraging its extensive community model support and cross-hardware compatibility (e.g., AMD ROCm, Ascend NPUs);
  • SGLang powers multi-turn agent platforms: For agentic workflows sharing massive system prompts (frequently thousands of tokens containing tool definitions and rules), RadixAttention achieves 90%+90\%+ prefix cache hits, significantly reducing latency and compute costs.

Q2: If TensorRT-LLM delivers such high throughput, why has it not completely eclipsed alternative runtimes?

TensorRT-LLM achieves peak performance through significant operational trade-offs:

  1. Prolonged Compilation Overhead: Building optimized engines requires compiling ahead of time for specific GPU architectures, tensor-parallel widths, and maximum batch sizes. Recompiling across varied hardware topologies can take hours;
  2. Debugging Complexity: The core runtime is embedded within low-level C++ drivers, making runtime debugging challenging when non-standard inputs cause failures;
  3. Lagging Support for Novel Architectures: When academic or open-source labs introduce novel attention mechanisms (e.g., DeepSeek MLA or KDA), community engines like vLLM and SGLang frequently ship Triton implementations within days, whereas TensorRT-LLM requires engineering custom C++ plugins.

Q3: How does llama.cpp's Unified Memory zero-copy differ from discrete data center GPU HBM?

On Apple Silicon devices (e.g., Mac Studio with M-series chips), the CPU and GPU share a unified pool of high-bandwidth LPDDR memory (achieving 400~800 GB/s bandwidth). In standard PC or data center servers, model weights must be copied from host system RAM across the PCIe bus into discrete GPU HBM before computation can begin. Under unified memory architectures, llama.cpp uses operating system mmap() calls to allow the GPU to read weights directly from memory in place. This eliminates weight transfer overhead and allows running models whose parameter sizes exceed traditional single-GPU VRAM limits (e.g., running 70B models in 128GB unified RAM), though raw double-precision and half-precision compute density remains lower than dedicated data center Tensor Cores.

Related Articles

Start with the same topic, then continue with the latest deep dives.

Production InfiniBand Deployment, Cluster Operations, and NCCL Tuning: OFED Drivers, OpenSM HA, ibdiagnet Fabric Auditing, and GPUDirect RDMA

What bare-metal operational challenges arise when translating network architecture into physical 10,000-GPU AI data centers? A comprehensive guide to production InfiniBand operations: installing and managing Mellanox OFED / DOCA driver stacks and firmware tools (flint/mlxlink); configuring high-availability Master/Standby Subnet Manager topologies with OpenSM; auditing fabric health and diagnosing dirty optical links (Symbol Errors) and speed renegotiation drops using ibdiagnet; and diving into the GPU communication layer to configure GPUDirect RDMA (nvidia-peermem) and tune mission-critical NCCL parameters (NCCL_IB_HCA, NCCL_NET_GDR_LEVEL=5) for wire-speed All-Reduce performance.

InfiniBand AI Cluster Networking in Practice: Fat-Tree Topologies, Rail-Optimized Architecture, Adaptive Routing (AR), and In-Network Reduction (SHARP)

How can thousands of 8-GPU servers be interconnected with tens of thousands of optical links into a non-blocking, deadlock-free high-performance fabric? Deconstruct modern AI supercluster topologies: from Charles Leiserson's 1985 Fat-Tree mathematical model and port count k derivations for 2-Tier and 3-Tier non-blocking ceilings to the 8-plane Rail-Optimized architecture tailored for DGX H100/H200/B200 clusters; analyze how FTree and Up/Down routing engines forbid 'Down-then-Up' turns to eliminate credit loop deadlocks; and discover how hardware Adaptive Routing (AR) and SHARP in-network aggregation achieve a 2x throughput boost during GPU All-Reduce operations.

InfiniBand Architecture First Principles: Physical Link Rates, Credit-Based Link Flow Control, and Subnet Manager Fabric Orchestration

Why does native InfiniBand remain the dominant fabric for 10,000-GPU AI compute clusters and top-tier supercomputers in 2026? Deconstruct the layered InfiniBand protocol stack and physical link evolution: from EDR 100G, HDR 200G, and NDR 400G (Quantum-2) to the 2026 mass production of XDR 800G (Quantum-X800/ConnectX-8); analyze link-layer first principles: Flit-level Credit-Based hardware flow control and sub-100ns Cut-Through switching mechanics; and explore the control engine: Subnet Manager (OpenSM) fabric discovery, dynamic GUID/LID/LMC allocation, and Linear Forwarding Table (LFT) hardware orchestration.

← Prev From model.generate() to the Roofline Model: The Physics of Prefill and Decode Divergence Next → Evolution of Attention Kernels: From FlashAttention-1/2/3 SRAM Tiling to FlashInfer Unified Heterogeneous Serving
← Back to Articles