Code GenerationOpen_source
vLLM logo

vLLM

High-throughput inference and serving engine for open-weights LLMs

OverviewArticles

From the vLLM blog

Published by vLLM, not by us. Every card opens the original post.

Taking vLLM Apart: A Practical Guide to Disaggregated Serving

What disaggregated serving actually buys you, how to run it end to end in vLLM today with prefill/decode plus the new GPU-less frontend and the things we're still working on.

vllm.ai

Watermarking in vLLM

How vLLM implements distribution-preserving Gumbel-max text watermarking with efficient GPU kernels, statistical detection, speculative decoding, and repeated-context safeguards.

vllm.ai

Announcing vllm-metal: Concurrent Serving on Apple Silicon

vllm-metal brings vLLM's paged, continuously batched serving stack to Apple Silicon, with flatter TTFT under concurrent agent load, batched MTP, and automatic M5 prefill acceleration.

vllm.ai

PD Serving of Qwen3.8-2.4T

How vLLM reaches 5K throughput and 180 interactivity on Qwen3.8-2.4T with GB300 NVL72 PD serving and how to reproduce results yourself.

vllm.ai

Scaling Multi-GPU Video Captioning with PyNvVideoCodec and vLLM

How to leverage NVIDIA Hardware Video Decoders to Achieve Multi-GPU Scaling in Video Captioning and Description tasks.

vllm.ai

How we trained the fastest DSpark for Kimi-K3 using GB300 NVL72

How Speculators and Mooncake enabled multi-node DSpark training for Kimi K3.

vllm.ai

vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300

Novita AI has open-sourced Chord, a high-performance W4A16 MoE CUDA kernel for Kimi K2.x serving shapes, with a Humming-compatible indexed path and grouped SM90 operators.

vllm.ai

vime × RL-Kernel × AMD: Bitwise Train–Rollout Consistency on ROCm

vime and RL-Kernel align selected-token logprobs bit for bit across Megatron training and vLLM rollout on AMD Instinct MI300X, with zero mismatches across 200 GRPO steps.

vllm.ai

Kimi K3 Performance Optimizations in vLLM: The Road to 2.8× Throughput

Kimi K3 serving optimizations across scheduling, KDA prefix caching, ReplaySSM state recovery, PD disaggregation and state offload, parallelism, MoE, and GPU kernels.

vllm.ai

Following the Bottleneck: Optimizing MiniMax M3 on AMD Instinct MI355X

A performance model for LLM serving: inspect local shapes, remove repeated work, verify data movement and dispatch, then follow the queue.

vllm.ai

Tiered KV Cache Offloading in vLLM

A host-centric framework for scaling KV cache across host memory, filesystems, object stores, and remote peers — reducing recomputation and increasing serving capacity.

vllm.ai

GLM 5.3 Optimizations, Part 1: Hybrid HiSparse Offloading in vLLM

vLLM integrates HiSparse as a pressure-driven memory tier that composes with the Hybrid Memory Allocator and KV offloading, letting GLM 5.3 requests keep decoding when their KV no longer fits in GPU memory, so concurrency stays high.

vllm.ai