Strata
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- Strata reportedly runs 125B Qwen model on 12GB GPUs
Developer Niko1221 has open-sourced the Strata engine, which reportedly runs a quantized Qwen3.8-Flash-Next model with 125 billion parameters on consumer GPUs with at least 12GB of VRAM. The engine keeps the MoE model in system RAM, loads frequently used experts into VRAM, and uses a lightweight model for speculative decoding; reported tests reached 94 tokens per second on an RTX 5070 with a 2-bit quantization.
IT之家 AI · 🔥 3
Experience and discussion from the community
Share my Strata experienceAsk about Strata
Nobody has shared their experience with Strata yet.