Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

Niko1221

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 1

  1. Strata reportedly runs 125B Qwen model on 12GB GPUs

    Developer Niko1221 has open-sourced the Strata engine, which reportedly runs a quantized Qwen3.8-Flash-Next model with 125 billion parameters on consumer GPUs with at least 12GB of VRAM. The engine keeps the MoE model in system RAM, loads frequently used experts into VRAM, and uses a lightweight model for speculative decoding; reported tests reached 94 tokens per second on an RTX 5070 with a 2-bit quantization.

    IT之家 AI · 🔥 3

Experience and discussion from the community

Share my Niko1221 experienceAsk about Niko1221

Nobody has shared their experience with Niko1221 yet.