Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

DeepGEMM

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 2

  1. DeepGEMM adds Ascend support and new kernels, trends on GitHub

    DeepGEMM is DeepSeek's open-source, high-performance tensor-core kernel library that unifies many core LLM computation primitives — FP8/FP4/BF16 GEMMs, fused MoE with overlapped communication (Mega MoE), MQA scoring for the lightning indexer (including a sparse version) and HyperConnection (HC) — in a single CUDA codebase, with all kernels compiled at runtime via DeepJIT and no CUDA compilation at install time. According to the repository's changelog, DeepGEMM-Ascend became available on 2026.09.30 alongside new optimizations such as locality domain features, after earlier additions of the Sparse Indexer, Mega Gate, Mega mHC and MoE/Indexer optimizations. The repo also says its performance matches or exceeds expert-tuned libraries across various matrix shapes, and that it reached up to 1550 TFLOPS on H800 in April 2025.

    GitHub Trending(每日) · 🔥 26
  2. vLLM v0.31.0 expands serving, speculation, and model support

    vLLM v0.31.0 adds broad serving and model-support updates, including DeepSeek-V4.1-Flash optimizations, Model Runner V2 speculative decoding, larger-scale expert parallelism, and improved multimodal support. It also introduces the `vllm preload` weight-cache daemon for faster engine restarts and experimental CRIU-based initialized-engine snapshots, alongside several security and breaking changes.

    vLLM Releases · 🔥 4

Experience and discussion from the community

Share my DeepGEMM experienceAsk about DeepGEMM

Nobody has shared their experience with DeepGEMM yet.