DeepGEMM adds Ascend support and new kernels, trends on GitHub
DeepGEMM is DeepSeek's open-source, high-performance tensor-core kernel library that unifies many core LLM computation primitives — FP8/FP4/BF16 GEMMs, fused MoE with overlapped communication (Mega MoE), MQA scoring for the lightning indexer (including a sparse version) and HyperConnection (HC) — in a single CUDA codebase, with all kernels compiled at runtime via DeepJIT and no CUDA compilation at install time. According to the repository's changelog, DeepGEMM-Ascend became available on 2026.09.30 alongside new optimizations such as locality domain features, after earlier additions of the Sparse Indexer, Mega Gate, Mega mHC and MoE/Indexer optimizations. The repo also says its performance matches or exceeds expert-tuned libraries across various matrix shapes, and that it reached up to 1550 TFLOPS on H800 in April 2025.