DeepGEMM
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 2
- DeepGEMM adds Ascend support and new kernels, trends on GitHub
DeepGEMM is DeepSeek's open-source, high-performance tensor-core kernel library that unifies many core LLM computation primitives — FP8/FP4/BF16 GEMMs, fused MoE with overlapped communication (Mega MoE), MQA scoring for the lightning indexer (including a sparse version) and HyperConnection (HC) — in a single CUDA codebase, with all kernels compiled at runtime via DeepJIT and no CUDA compilation at install time. According to the repository's changelog, DeepGEMM-Ascend became available on 2026.09.30 alongside new optimizations such as locality domain features, after earlier additions of the Sparse Indexer, Mega Gate, Mega mHC and MoE/Indexer optimizations. The repo also says its performance matches or exceeds expert-tuned libraries across various matrix shapes, and that it reached up to 1550 TFLOPS on H800 in April 2025.
GitHub Trending(每日) · 🔥 26 - vLLM v0.31.0 expands serving, speculation, and model support
vLLM v0.31.0 adds broad serving and model-support updates, including DeepSeek-V4.1-Flash optimizations, Model Runner V2 speculative decoding, larger-scale expert parallelism, and improved multimodal support. It also introduces the `vllm preload` weight-cache daemon for faster engine restarts and experimental CRIU-based initialized-engine snapshots, alongside several security and breaking changes.
vLLM Releases · 🔥 4
Experience and discussion from the community
Share my DeepGEMM experienceAsk about DeepGEMM
Nobody has shared their experience with DeepGEMM yet.