Strata reportedly runs 125B Qwen model on 12GB GPUs
Developer Niko1221 has open-sourced the Strata engine, which reportedly runs a quantized Qwen3.8-Flash-Next model with 125 billion parameters on consumer GPUs with at least 12GB of VRAM. The engine keeps the MoE model in system RAM, loads frequently used experts into VRAM, and uses a lightweight model for speculative decoding; reported tests reached 94 tokens per second on an RTX 5070 with a 2-bit quantization.
Why it matters: It could lower the hardware barrier for running very large open models locally, although the reported performance comes from third-party coverage and specific configurations.
IT之家 10 月 6 日消息,科技媒体 gigazine 今天(10 月 6 日)报道,报道称开发者 Niko1221 开源推出 Strata 引擎,可以在 12GB 及以上显存的消费级显卡上,运行量化的 Qwen3.8-Flash-Next 模型(1250 亿参数)。

IT之家注:Qwen3.8-Flash-Next 模型是阿里巴巴 Qwen 团队于 2026 年 8 月推出的多模态混合专家(MoE)模型的压缩版本,配有 125B 参数,另含 51B 参数的 n-gram 嵌入表,原生支持 262K token 上下文长度。

Strata 为了降低显存占用,主要采用两项关键技术:其一是将 MoE(混合专家)模型整体载入 RAM,仅将高频使用的专家模型载入 VRAM。其二使用轻量模型进行投机解码,预测下一 Token 以加速推理过程。
硬件要求方面,运行 Qwen3.8-Flash-Next 需满足最低配置:12GB 以上 VRAM 的 NVIDIA 或 AMD 显卡、32GB 以上 RAM、80GB 以上存储空间(推荐 SSD)、Windows 10/11 或 Linux 系统。
性能方面,在 NVIDIA GeForce RTX 5070(12GB VRAM)、Ryzen 5 7600、64GB RAM 配置下,2 比特量化版 Q2_0 达到 94 词元 / 秒推理速度,3 比特量化版 IQ3_S 为 53 词元 / 秒。AMD Radeon RX 9070 XT(16GB VRAM)环境下,Q2_0 版本达到 60 词元 / 秒。

英伟达主机性能表现:
硬件配置: NVIDIA RTX 5070(12 GB 显存),AMD Ryzen 5 7600,64 GB 系统内存。
| 量化版本(Size) | 写答案速度(Writes answers) | 阅读提示速度(Reads your prompt) |
|---|---|---|
| Q2_0 | 94 个词元 / 秒 | 2,650 个词元 / 秒 |
| IQ2_XS | 79 个词元 / 秒 | 2,090 个词元 / 秒 |
| IQ3_XXS | 62 个词元 / 秒 | 1,750 个词元 / 秒 |
| IQ3_S | 53 个词元 / 秒 | 1,620 个词元 / 秒 |
| Coder | 55 个词元 / 秒 | 2,180 个词元 / 秒 |
AMD 主机性能表现:
硬件配置:AMD RX 9070 XT(16 GB 显存),AMD Ryzen 9 3900X,47 GB 内存。
| 量化版本(Size) | 写答案速度(Writes answers) | 阅读提示速度(Reads your prompt) |
|---|---|---|
| Q2_0 | 60 个词元 / 秒 | 1,160 个词元 / 秒 |
| IQ2_XS | 52 个词元 / 秒 | 1,110 个词元 / 秒 |
| Coder | 44 个词元 / 秒 | 1,420 个词元 / 秒 |
这里的“写答案”对应生成回复(输出)的速度,“阅读提示”对应处理你输入的 prompt(上下文)的速度。
参考
广告声明:文内含有的对外跳转链接(包括不限于超链接、二维码、口令等形式),用于传递更多信息,节省甄选时间,结果仅供参考,IT之家包含外链的文章均包含本声明。
How we got here
- Ghost AI raises $11M to build Core, a $3,499 personal AI agent computerSiliconANGLE AI · Qwen
- TypeSafe AI’s Jev turns frequent AI tasks into fast decisions创业邦 科技 · Qwen