vLLM climbs GitHub Trending among Python projects
vLLM is a high-throughput, memory-efficient inference and serving engine for large language models, originally developed in the Sky Computing Lab at UC Berkeley and now maintained by a community of more than 2,000 contributors. It manages attention key-value memory with PagedAttention and supports continuous batching, chunked prefill, prefix caching, speculative decoding and quantization formats including FP8, INT8/INT4, GPTQ/AWQ and GGUF, alongside an OpenAI-compatible API server. The engine runs on NVIDIA, AMD and Intel GPUs plus x86/ARM/PowerPC CPUs, and supports 200+ model architectures on Hugging Face.