
vLLM
High-throughput inference and serving engine for open-weights LLMs
About vLLM
vLLM is the layer between a downloaded model and something you can actually send requests to. It loads open-weights models and serves them behind a drop-in OpenAI-compatible API, so code written against the OpenAI SDK points at your own GPU with a changed base URL and nothing else. It started in the Sky Computing Lab at UC Berkeley, is now a PyTorch Foundation project, and has grown into one of the most active open source AI projects with over 3,000 contributors from academic institutions and companies.
What makes it worth the operational effort is memory management. PagedAttention treats the attention key and value cache the way an operating system treats virtual memory, in pages rather than one contiguous block per request, which removes most of the fragmentation that otherwise wastes GPU memory and caps how many requests fit at once. Combined with continuous batching, chunked prefill and prefix caching, this is the difference between a GPU serving a handful of concurrent users and the same GPU serving many. That is the whole value proposition: not a nicer interface, but more throughput per dollar of hardware you already bought. It is free and Apache-2.0 with no hosted tier and no paid edition, funded by donations through GitHub and OpenCollective, with compute for development contributed by AMD, AWS, Google Cloud, NVIDIA, Red Hat and others.
Install with uv pip install vllm or pull the Docker image, then start a server pointing at a Hugging Face model identifier or a local path. vLLM loads the weights, allocates its paged KV cache, and exposes an OpenAI-compatible HTTP server, with Anthropic Messages API and gRPC endpoints available as well. Requests join a continuously batched queue rather than waiting for a fixed batch to fill, so a short request does not sit behind a long one. For models too large for one card, you configure tensor, pipeline, data, expert or context parallelism to spread the work across GPUs or nodes. Quantization is set at load time across a long list of formats including FP8, NVFP4, INT8, INT4, GPTQ, AWQ and GGUF. Speculative decoding, structured output through xgrammar or guidance, tool calling and reasoning parsers, and multi-LoRA serving are all configuration rather than custom code.
- •PagedAttention - Paged management of the attention KV cache that removes fragmentation and raises how many concurrent requests a GPU can hold
- •Continuous Batching - Requests enter and leave the batch independently, with chunked prefill and prefix caching, so latency does not depend on the slowest neighbour
- •OpenAI-Compatible Server - A drop-in API plus Anthropic Messages and gRPC support, so existing client code works against your own hardware
- •200+ Model Architectures - Dense and mixture-of-expert LLMs, hybrid state-space models, multimodal models, and embedding and reranking models from Hugging Face
- •Broad Hardware Support - NVIDIA and AMD GPUs, x86, ARM and PowerPC CPUs, plus plugins for Google TPU, Intel Gaudi, IBM Spyre, Huawei Ascend and Apple Silicon
- •Distributed Serving - Tensor, pipeline, data, expert and context parallelism, with disaggregated prefill, decode and encode for large deployments
Teams running open-weights models on their own hardware or rented GPUs, and anyone whose inference bill has grown large enough that throughput per GPU is a budget line rather than a detail. It suits platform and ML infrastructure engineers, and organisations serving models on-premises for privacy or compliance reasons. It is not a consumer product and should not be mistaken for one: there is no signup, no hosted tier and no interface, and running it well means understanding GPU memory, quantization formats and parallelism strategies. A developer who simply wants a model running locally on a laptop is better served by Ollama or LM Studio, both of which are simpler and are listed in this directory. The 7,000-plus open issues are what a project of this surface area looks like, but it is worth knowing about before you build a platform on top of it.
Pricing
No pricing page found
Free and Apache-2.0. There is no pricing page, no hosted tier and no paid edition; the project is funded by donations through GitHub and OpenCollective, and compute for development and testing is contributed by sponsors. Your cost is the hardware you run it on.
Checked 2026-09-18













