Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

Qwen3.8-Flash-Next

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 2

  1. vLLM v0.31.0 expands serving, speculation, and model support

    vLLM v0.31.0 adds broad serving and model-support updates, including DeepSeek-V4.1-Flash optimizations, Model Runner V2 speculative decoding, larger-scale expert parallelism, and improved multimodal support. It also introduces the `vllm preload` weight-cache daemon for faster engine restarts and experimental CRIU-based initialized-engine snapshots, alongside several security and breaking changes.

    vLLM Releases · 🔥 3
  2. Strata reportedly runs 125B Qwen model on 12GB GPUs

    Developer Niko1221 has open-sourced the Strata engine, which reportedly runs a quantized Qwen3.8-Flash-Next model with 125 billion parameters on consumer GPUs with at least 12GB of VRAM. The engine keeps the MoE model in system RAM, loads frequently used experts into VRAM, and uses a lightweight model for speculative decoding; reported tests reached 94 tokens per second on an RTX 5070 with a 2-bit quantization.

    IT之家 AI · 🔥 3

Experience and discussion from the community

Share my Qwen3.8-Flash-Next experienceAsk about Qwen3.8-Flash-Next

Nobody has shared their experience with Qwen3.8-Flash-Next yet.