Code GenerationOpen_source
vLLM logo

vLLM

High-throughput inference and serving engine for open-weights LLMs

OverviewArticles

About vLLM

vLLM is the layer between a downloaded model and something you can actually send requests to. It loads open-weights models and serves them behind a drop-in OpenAI-compatible API, so code written against the OpenAI SDK points at your own GPU with a changed base URL and nothing else. It started in the Sky Computing Lab at UC Berkeley, is now a PyTorch Foundation project, and has grown into one of the most active open source AI projects with over 3,000 contributors from academic institutions and companies.

What makes it worth the operational effort is memory management. PagedAttention treats the attention key and value cache the way an operating system treats virtual memory, in pages rather than one contiguous block per request, which removes most of the fragmentation that otherwise wastes GPU memory and caps how many requests fit at once. Combined with continuous batching, chunked prefill and prefix caching, this is the difference between a GPU serving a handful of concurrent users and the same GPU serving many. That is the whole value proposition: not a nicer interface, but more throughput per dollar of hardware you already bought. It is free and Apache-2.0 with no hosted tier and no paid edition, funded by donations through GitHub and OpenCollective, with compute for development contributed by AMD, AWS, Google Cloud, NVIDIA, Red Hat and others.

Pricing

No pricing page found

No free tier

Free and Apache-2.0. There is no pricing page, no hosted tier and no paid edition; the project is funded by donations through GitHub and OpenCollective, and compute for development and testing is contributed by sponsors. Your cost is the hardware you run it on.

Checked 2026-09-18

Details

GitHub Stars 90,229
Forks 21,294
Data from: GitHub • Website•Updated: Aug 27, 2026
devopsinfrastructuredeploymentmodel-servingopen-source