Core Open-Source Components and Parity
On September 30, 2026, DeepSeek officially open-sourced its complete AI infrastructure stack for the Huawei Ascend computing platform. Extracted directly from DeepSeek's live production environments, this software foundation achieves 1 functional parity with the core toolchain previously released for NVIDIA CUDA platforms, providing an industrial-grade solution for efficient training and low-latency inference on domestic compute clusters.
The open-sourced Ascend software stack covers the entire critical path from low-level kernel authoring and memory optimization to large-scale distributed communication, comprising five core tools and operator libraries:
- TileLang DSL Compiler: A lightweight Python-based domain-specific language and compiler that encapsulates low-level Ascend C primitives, allowing engineers to define tiling and memory layout strategies in high-level code while automating thread binding and instruction pipelining.
- DeepGEMM-Ascend Matrix Library: A high-throughput kernel library built for Mixture-of-Experts and dense workloads supporting FP8, FP4, and BF16 precisions, abstracting Ascend matrix multiply-add primitives to approach theoretical hardware ceilings.
- DeepEP-Ascend Communication Library: An inter-device communication library engineered for Expert Parallelism and MoE workloads, leveraging Ascend C, HCCL, and UBMEM to maximize all-to-all token dispatch and combine efficiency.
- FlashMLA Attention Kernel: A specialized operator optimized for Multi-head Latent Attention in DeepSeek-V3 and V4 series models, utilizing paged KV cache management to accelerate decoding during long-context inference.
- TileKernels and DeepSelect: Utility operator suites offering high-throughput vector arithmetic, memory-bound access kernels, and lightweight data filtering for complex dynamic data streams.
SuperPoD Interconnect and Communication Co-Design
In trillion-parameter Mixture-of-Experts models, cross-device and cross-node network latency frequently forms the primary bottleneck that caps linear scaling efficiency. Huawei and DeepSeek engineered deep co-optimizations around the Ascend 950 (including 950DT) platform, introducing SuperPoD Flex architectures and UBL128 all-to-all networking.
This super-node interconnect topology delivers highly competitive scaling performance across multiple structural tiers:
- Single-Tier Scale-up: Employs fully connected UBL128 fabric to achieve single-tier direct switching across 128 NPUs with 3.2Tbps per-card interconnect bandwidth, eliminating multi-hop routing penalties.
- Two-Tier Scale-out: Supports scaling up to 256K accelerator nodes via a two-tier network hierarchy, satisfying the training and real-time scheduling throughput required by frontier foundation models.
- Hardware-Optimized Communication Primitives: Integrated with Ascend's ASC-COMM communication primitives, DeepEP fully saturates physical fabric bandwidth during inter-card token routing, minimizing compute stall times.
Benchmark Results and Production Validation
Rather than a synthetic prototype, this infrastructure stack represents a battle-tested production system proven under enterprise workloads. In offline inference benchmarks using an EP32 parallel strategy with DeepSeek-V4.1-Flash, the software stack maintained high concurrent throughput and minimal latency while driving compute and communication utilization close to physical limits.
In terms of production deployment, DeepSeek's multi-thousand-card inference clusters in Inner Mongolia have planned extensive rollouts powered by Ascend 950DT processors. By bypassing generic framework abstractions and generating shape-specific kernels via DeepJIT runtime compilation, the system achieves minimal tail latency and optimal memory footprints.
Repository Index and Engineering Significance
For years, domestic computing platforms encountered engineering friction characterized by functional compatibility without full hardware saturation, alongside steep communication overheads across multi-node topologies. By releasing its battle-tested infrastructure, DeepSeek directly resolves ecosystem deficiencies in high-performance DSL compilation and distributed communication, freeing practitioners from rewriting low-level assembly kernels.
All components from this release are now available on GitHub with complete build guides and Ascend deployment instructions:
| Component | Core Role and Capabilities | GitHub Repository Path |
|---|---|---|
| TileLang | Cross-hardware kernel compiler and high-level DSL | tile-ai/tilelang |
| DeepGEMM-Ascend | Ascend matrix multiplication library optimized for MoE | deepseek-ai/DeepGEMM-Ascend |
| DeepEP-Ascend | High-throughput low-latency communication library for EP | deepseek-ai/DeepEP-Ascend |
| FlashMLA | Multi-head Latent Attention kernel for long-context decoding | deepseek-ai/FlashMLA |
| TileKernels | Vector arithmetic, memory optimization, and data selection | deepseek-ai/TileKernels |