Qwen3.8-Flash-Next
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 2
- vLLM v0.31.0 expands serving, speculation, and model support
vLLM v0.31.0 adds broad serving and model-support updates, including DeepSeek-V4.1-Flash optimizations, Model Runner V2 speculative decoding, larger-scale expert parallelism, and improved multimodal support. It also introduces the `vllm preload` weight-cache daemon for faster engine restarts and experimental CRIU-based initialized-engine snapshots, alongside several security and breaking changes.
vLLM Releases · 🔥 3 - Strata reportedly runs 125B Qwen model on 12GB GPUs
Developer Niko1221 has open-sourced the Strata engine, which reportedly runs a quantized Qwen3.8-Flash-Next model with 125 billion parameters on consumer GPUs with at least 12GB of VRAM. The engine keeps the MoE model in system RAM, loads frequently used experts into VRAM, and uses a lightweight model for speculative decoding; reported tests reached 94 tokens per second on an RTX 5070 with a 2-bit quantization.
IT之家 AI · 🔥 3
Experience and discussion from the community
Share my Qwen3.8-Flash-Next experienceAsk about Qwen3.8-Flash-Next
Nobody has shared their experience with Qwen3.8-Flash-Next yet.