vLLM screenshot
Code GenerationOpen_source
vLLM logo

vLLM

High-throughput inference and serving engine for open-weights LLMs

The Apache-2.0 serving engine that most self-hosted LLM stacks run on, exposing any of 200+ open-weights architectures behind an OpenAI-compatible API.

GitHub Stars 90,229
Forks 21,294
Data from: GitHubWebsiteUpdated: Aug 27, 2026

Pricing

No pricing page found

No free tier

Free and Apache-2.0. There is no pricing page, no hosted tier and no paid edition; the project is funded by donations through GitHub and OpenCollective, and compute for development and testing is contributed by sponsors. Your cost is the hardware you run it on.

Checked 2026-08-28

About vLLM

Tags

devopsinfrastructuredeploymentmodel-servingopen-source