DeepSeek released open-source software designed for Huawei's Ascend processors on September 30, expanding their collaboration into the software layer of AI infrastructure. The updates bring Ascend support to six open-source projects that handle low-level operations required for building and optimizing AI workloads, according to a Channel Insider report. The move addresses a core challenge for channel partners: determining how much existing development work transfers when customers consider alternative accelerator platforms.
Six components now work with Huawei's chip platform. Two releases hold particular relevance for developers switching between hardware systems: DeepGEMM-Ascend preserves the API and development workflow from DeepGEMM on other hardware, while TileKernels exposes identical Python APIs across Nvidia GPUs and Huawei NPUs. Four additional components broaden support across the stack—TileLang adds native support for the Ascend 950 with tools for writing and optimizing low-level AI operations, DeepEP-Ascend manages communication between accelerators including data transfers for large mixture-of-experts models, FlashMLA provides sparse-attention kernels for processing prompts and generating tokens on Ascend 950 NPUs, and DeepSelect adds TopK operations for selecting relevant data during model processing. DeepSeek worked with Huawei to optimize computation and communication for a supernode configuration built on 128 Ascend 950 chips, Reuters reported. DeepEP-Ascend achieved roughly 90% to 95% of the physical payload bandwidth limit in expert-parallel dispatch tests with up to 32 participating ranks, though larger configurations and combine operations remain under optimization.
DeepSeek's repositories credit Huawei with technical and engineering support. The report notes that software maturity has represented a significant obstacle for the chip platform as it competes with an Nvidia ecosystem built around years of CUDA development. Huawei also states it has upgraded Ascend C for the Ascend 950 generation and made progress opening its PTO instruction set through broader developer tooling efforts. The company has been linked to a planned 160,000-chip Ascend deployment in Inner Mongolia, stretching the relationship from software and cluster optimization into plans for substantially larger AI infrastructure.
For systems integrators, shared APIs can reduce some redevelopment without making Nvidia and Ascend environments interchangeable, the report finds. Teams evaluating the platform should start with a representative customer workload, checking model and framework compatibility, required versions of Huawei's CANN software toolkit, networking requirements, then validating monitoring and performance under production-like loads. VARs and infrastructure advisers should include software maturity in hardware comparisons, since migration effort and long-term support can alter deployment economics even when the underlying accelerator meets compute requirements. MSPs supporting mixed accelerator environments should keep platform-specific dependencies visible, and providers of managed AI infrastructure may need separate runbooks and escalation paths when CUDA and CANN coexist in the same customer estate. For China-based partners, broader software support could make Ascend easier to package with integration and managed services rather than as hardware-only sales, with more value coming from services attached to deployment and ongoing support. The platform's long-term viability depends on whether shared development workflows can offset the ecosystem advantages that years of proprietary tooling have created, and whether partners will treat interoperability as a feature worth engineering for or a migration tax best avoided.

