News · High impactBack to News

DeepSeek Open-Sources Full Ascend AI Infrastructure Stack: Parity with CUDA Tools and 128-Card SuperPoD Support

DeepSeek has open-sourced its production-grade AI infrastructure stack for Huawei Ascend platforms, including the TileLang DSL compiler, DeepGEMM, DeepEP communication library, and FlashMLA sparse attention kernels, achieving 1:1 functional parity with CUDA tools and near-peak hardware utilization on 128-card SuperPoD Flex clusters.

Core Open-Source Components and Parity

On September 30, 2026, DeepSeek officially open-sourced its complete AI infrastructure stack for the Huawei Ascend computing platform. Extracted directly from DeepSeek's live production environments, this software foundation achieves 1 functional parity with the core toolchain previously released for NVIDIA CUDA platforms, providing an industrial-grade solution for efficient training and low-latency inference on domestic compute clusters.

The open-sourced Ascend software stack covers the entire critical path from low-level kernel authoring and memory optimization to large-scale distributed communication, comprising five core tools and operator libraries:

  • TileLang DSL Compiler: A lightweight Python-based domain-specific language and compiler that encapsulates low-level Ascend C primitives, allowing engineers to define tiling and memory layout strategies in high-level code while automating thread binding and instruction pipelining.
  • DeepGEMM-Ascend Matrix Library: A high-throughput kernel library built for Mixture-of-Experts and dense workloads supporting FP8, FP4, and BF16 precisions, abstracting Ascend matrix multiply-add primitives to approach theoretical hardware ceilings.
  • DeepEP-Ascend Communication Library: An inter-device communication library engineered for Expert Parallelism and MoE workloads, leveraging Ascend C, HCCL, and UBMEM to maximize all-to-all token dispatch and combine efficiency.
  • FlashMLA Attention Kernel: A specialized operator optimized for Multi-head Latent Attention in DeepSeek-V3 and V4 series models, utilizing paged KV cache management to accelerate decoding during long-context inference.
  • TileKernels and DeepSelect: Utility operator suites offering high-throughput vector arithmetic, memory-bound access kernels, and lightweight data filtering for complex dynamic data streams.

SuperPoD Interconnect and Communication Co-Design

In trillion-parameter Mixture-of-Experts models, cross-device and cross-node network latency frequently forms the primary bottleneck that caps linear scaling efficiency. Huawei and DeepSeek engineered deep co-optimizations around the Ascend 950 (including 950DT) platform, introducing SuperPoD Flex architectures and UBL128 all-to-all networking.

This super-node interconnect topology delivers highly competitive scaling performance across multiple structural tiers:

  • Single-Tier Scale-up: Employs fully connected UBL128 fabric to achieve single-tier direct switching across 128 NPUs with 3.2Tbps per-card interconnect bandwidth, eliminating multi-hop routing penalties.
  • Two-Tier Scale-out: Supports scaling up to 256K accelerator nodes via a two-tier network hierarchy, satisfying the training and real-time scheduling throughput required by frontier foundation models.
  • Hardware-Optimized Communication Primitives: Integrated with Ascend's ASC-COMM communication primitives, DeepEP fully saturates physical fabric bandwidth during inter-card token routing, minimizing compute stall times.

Benchmark Results and Production Validation

Rather than a synthetic prototype, this infrastructure stack represents a battle-tested production system proven under enterprise workloads. In offline inference benchmarks using an EP32 parallel strategy with DeepSeek-V4.1-Flash, the software stack maintained high concurrent throughput and minimal latency while driving compute and communication utilization close to physical limits.

In terms of production deployment, DeepSeek's multi-thousand-card inference clusters in Inner Mongolia have planned extensive rollouts powered by Ascend 950DT processors. By bypassing generic framework abstractions and generating shape-specific kernels via DeepJIT runtime compilation, the system achieves minimal tail latency and optimal memory footprints.

Repository Index and Engineering Significance

For years, domestic computing platforms encountered engineering friction characterized by functional compatibility without full hardware saturation, alongside steep communication overheads across multi-node topologies. By releasing its battle-tested infrastructure, DeepSeek directly resolves ecosystem deficiencies in high-performance DSL compilation and distributed communication, freeing practitioners from rewriting low-level assembly kernels.

All components from this release are now available on GitHub with complete build guides and Ascend deployment instructions:

ComponentCore Role and CapabilitiesGitHub Repository Path
TileLangCross-hardware kernel compiler and high-level DSLtile-ai/tilelang
DeepGEMM-AscendAscend matrix multiplication library optimized for MoEdeepseek-ai/DeepGEMM-Ascend
DeepEP-AscendHigh-throughput low-latency communication library for EPdeepseek-ai/DeepEP-Ascend
FlashMLAMulti-head Latent Attention kernel for long-context decodingdeepseek-ai/FlashMLA
TileKernelsVector arithmetic, memory optimization, and data selectiondeepseek-ai/TileKernels

Next step

Turn the update into a next step

Continue along the same topic.

Browse practical guides