vLLM
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- vLLM v0.31.0 expands serving, speculation, and model support
vLLM v0.31.0 adds broad serving and model-support updates, including DeepSeek-V4.1-Flash optimizations, Model Runner V2 speculative decoding, larger-scale expert parallelism, and improved multimodal support. It also introduces the `vllm preload` weight-cache daemon for faster engine restarts and experimental CRIU-based initialized-engine snapshots, alongside several security and breaking changes.
vLLM Releases · 🔥 4
Experience and discussion from the community
Share my vLLM experienceAsk about vLLM
Nobody has shared their experience with vLLM yet.