
vLLM
High-throughput inference and serving engine for open-weights LLMs
From the vLLM blog
Published by vLLM, not by us. Every card opens the original post.
Taking vLLM Apart: A Practical Guide to Disaggregated Serving
What disaggregated serving actually buys you, how to run it end to end in vLLM today with prefill/decode plus the new GPU-less frontend and the things we're still working on.
Watermarking in vLLM
How vLLM implements distribution-preserving Gumbel-max text watermarking with efficient GPU kernels, statistical detection, speculative decoding, and repeated-context safeguards.
Announcing vllm-metal: Concurrent Serving on Apple Silicon
vllm-metal brings vLLM's paged, continuously batched serving stack to Apple Silicon, with flatter TTFT under concurrent agent load, batched MTP, and automatic M5 prefill acceleration.
PD Serving of Qwen3.8-2.4T
How vLLM reaches 5K throughput and 180 interactivity on Qwen3.8-2.4T with GB300 NVL72 PD serving and how to reproduce results yourself.
Scaling Multi-GPU Video Captioning with PyNvVideoCodec and vLLM
How to leverage NVIDIA Hardware Video Decoders to Achieve Multi-GPU Scaling in Video Captioning and Description tasks.
How we trained the fastest DSpark for Kimi-K3 using GB300 NVL72
How Speculators and Mooncake enabled multi-node DSpark training for Kimi K3.
vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300
Novita AI has open-sourced Chord, a high-performance W4A16 MoE CUDA kernel for Kimi K2.x serving shapes, with a Humming-compatible indexed path and grouped SM90 operators.
vime × RL-Kernel × AMD: Bitwise Train–Rollout Consistency on ROCm
vime and RL-Kernel align selected-token logprobs bit for bit across Megatron training and vLLM rollout on AMD Instinct MI300X, with zero mismatches across 200 GRPO steps.
Kimi K3 Performance Optimizations in vLLM: The Road to 2.8× Throughput
Kimi K3 serving optimizations across scheduling, KDA prefix caching, ReplaySSM state recovery, PD disaggregation and state offload, parallelism, MoE, and GPU kernels.
Following the Bottleneck: Optimizing MiniMax M3 on AMD Instinct MI355X
A performance model for LLM serving: inspect local shapes, remove repeated work, verify data movement and dispatch, then follow the queue.
Tiered KV Cache Offloading in vLLM
A host-centric framework for scaling KV cache across host memory, filesystems, object stores, and remote peers — reducing recomputation and increasing serving capacity.
GLM 5.3 Optimizations, Part 1: Hybrid HiSparse Offloading in vLLM
vLLM integrates HiSparse as a pressure-driven memory tier that composes with the Hybrid Memory Allocator and KV offloading, letting GLM 5.3 requests keep decoding when their KV no longer fits in GPU memory, so concurrency stays high.