DeepSeek’s kernel library shipped an Ascend backend on September 30. DeepGEMM-Ascend is the headline item in the repo’s news feed, and it lands alongside a second September 30 entry adding “locality domain features” and a September 10 batch covering a sparse indexer, Mega Gate, Mega mHC, DeepJIT, and more MoE and indexer optimizations. A CUDA library from a Chinese lab now compiles for Huawei’s accelerator stack. That is the news, and it is bigger than any single kernel.

The rest of the changelog is a paper trail. DeepGEMM began as a clean FP8 GEMM library, hit 1550 TFLOPS on an H800 in April 2025, added SM90/SM100 support in July 2025, added lightning-indexer scoring kernels for DeepSeek v3.2 in September 2025, then shipped Mega MoE, FP8xFP4 GEMM, a FP4 indexer, PDL, and faster JIT compilation in April 2026. The September 2026 batch is the latest layer. This is not a research artifact that gets one commit and dies. It is a maintained production dependency, and the Ascend port is the first time DeepSeek has publicly aimed it at non-NVIDIA silicon.

What DeepGEMM actually is

Read the repo’s own framing: “a unified, high-performance tensor core kernel library that brings together the key computation primitives of modern large language models.” GEMMs in FP8, FP4, and BF16. Fused MoE with overlapped communication (Mega MoE). MQA scoring for the lightning indexer. HyperConnection. All in one CUDA codebase, all compiled at runtime through DeepJIT, no CUDA compilation at install.

The design choices are the interesting part. DeepGEMM “leverages some concepts from CUTLASS and CuTe, but avoids heavy reliance on their templates or algebras.” It keeps a small number of core kernel functions on purpose, and the README says the goal is to be “a clean and accessible resource for learning NVIDIA GPU kernel optimization techniques.” A frontier lab is publishing its kernel internals as teaching material. That is unusual. Most labs treat the fused MoE kernel as a moat.

The performance claim is blunt: “Despite its lightweight design, DeepGEMM’s performance matches or exceeds expert-tuned libraries across various matrix shapes.” The 1550 TFLOPS H800 figure from April 2025 is the concrete number behind that. The library also offers set_tc_util and set_num_sms for tuning tensor core utilization and SM count, which tells you the authors expect users to squeeze the last few percent rather than accept defaults.

Mega MoE is the load-bearing piece

The Mega MoE kernel fuses and overlaps expert-parallel dispatch, linear 1, linear 2 (FP8xFP4 or FP8xFP8), SwiGLU, and EP combine into a single mega-kernel, overlapping NVLink communication with tensor core computation. It requires multi-process launch with symmetric memory, and the repo notes PyTorch 2.9 or newer for the symmetric buffer path.

That is a communication-bound workload collapsed into one kernel. For a sparse MoE model, the expensive part is not the matmul, it is moving tokens between experts across NVLink and waiting. Fusing dispatch, both linears, the activation, and the combine means the tensor cores keep working while the fabric moves data. DeepGEMM’s grouped GEMM API is built for the same shape: it groups only the M-axis, with N and K fixed, and the README says the design is “tailored for scenarios where experts in an MoE model share the same shape.” The masked variant exists because during decoding with CUDA graphs enabled, “the CPU is unaware of the number of tokens each expert receives,” so the kernel computes only the valid portions from a mask tensor.

None of this is generic BLAS work. Every API is shaped around one architecture: DeepSeek’s own MoE inference path.

The Ascend port changes the calculus

Here is the take. DeepGEMM-Ascend is not a portability exercise. It is an insurance policy.

DeepSeek has spent two years building a kernel library that extracts maximum performance from NVIDIA’s SM90 and SM100 parts, then extended it to Huawei’s stack. The company that made its name by training a frontier model at a reported fraction of Western cost is now hedging on the hardware side too. Export controls have made high-end NVIDIA accelerators hard to buy in China for years. A lab that can run its fused MoE kernel on Ascend has an option that most Western labs do not.

The caveat is that we cannot verify from the repo alone how complete the Ascend backend is, what performance it reaches, or which Ascend part it targets. The September 30 entry says only “Check DeepGEMM-Ascend for more details.” Treat the port as announced, not benchmarked. {/* TODO: verify DeepGEMM-Ascend performance figures and target chip (910B/910C) against Huawei or DeepSeek disclosures */}

There is a second, quieter signal. The repo’s citation lists nine authors: Chenggang Zhao, Zhean Xu, Liang Zhao, Jiashi Li, Chenhao Xu, Anyi Xu, Shengyu Liu, Kexing Zhou, and Kuai Yu. That is a kernel team, not a research group. DeepSeek is staffing the unglamorous layer, the part that decides whether a model runs at 40% or 80% of peak.

The licensing matters too. DeepGEMM is MIT. A competitor can lift the Mega MoE fusion strategy, the contiguous grouped GEMM layout, and the masked decoding path wholesale, and run them on their own hardware. DeepSeek is giving away the thing that makes its inference economics work.

What this means for builders

If you run MoE inference, the contiguous grouped GEMM API and the masked variant are worth reading even if you never call them. The pattern, group only M, keep N and K fixed, align each expert segment to the M block size, is a clean answer to the ragged-token problem that most MoE serving stacks handle with worse code.

If you write kernels, DeepJIT is the part to study. Runtime compilation with no CUDA build at install, a cache directory that defaults to $HOME/.dj and accepts a colon-separated search path, and environment flags like DG_JIT_CHECK_NO_SPILLS and DG_JIT_CHECK_NO_LOCAL_MEMORY that turn register-spill and local-memory usage into hard assertions. That is a testing harness for kernel correctness disguised as a build system.

The number to watch is whether DeepGEMM-Ascend shows up in a DeepSeek training or inference run, and what it costs against the H800 baseline. Until then, the port is a hedge, and the hedge is the story.