vLLM screenshot
Code GenerationOpen_source
vLLM logo

vLLM

High-throughput inference and serving engine for open-weights LLMs

The Apache-2.0 serving engine that most self-hosted LLM stacks run on, exposing any of 200+ open-weights architectures behind an OpenAI-compatible API.

GitHub Stars 90,229
Forks 21,294
Data from: GitHubWebsiteUpdated: Aug 27, 2026

About vLLM

Tags

devopsinfrastructuredeploymentmodel-servingopen-source