On September 30, 2026, DeepSeek took one of its boldest steps yet by publicly releasing a complete AI chip programming toolkit developed in close collaboration with Huawei Technologies. The package includes TileLang —a high-level language that acts as a direct alternative to CUDA, Nvidia's de facto standard that has dominated AI development for over a decade— alongside specialized libraries: DeepGEMM for matrix multiplication, FlashMLA for LLM inference, TileKernel, DeepSelect, and DeepEP. Everything is open-source, free to download, and optimized for the Huawei Ascend 950 chips. This move marks an inflection point in the AI hardware ecosystem, with direct implications for the costs that businesses —including SMBs— pay when using AI cloud services.
What Did DeepSeek and Huawei Actually Release?
Today's toolkit is not an academic experiment: it is production software that DeepSeek already used internally to run its own models on Huawei hardware. TileLang provides a high-level abstraction layer similar to what Nvidia's CUDA offers, meaning a developer familiar with the Nvidia ecosystem can migrate their models to Huawei Ascend chips without learning an entirely alien system. DeepGEMM and FlashMLA are low-level optimizations for the most compute-intensive operations in LLMs (transformers), while DeepEP handles inter-chip communication in multi-GPU configurations. Additionally, DeepSeek and Huawei co-developed a 128-chip Ascend 950 'supernode' solution that lets them compete in performance with Nvidia H100/H200 GPU clusters for training and inference of large models. The toolkit generalizes that earlier effort: any company or cloud provider can now attempt the same migration with their own models.
"When the software layer connecting models to hardware stops being a monopoly, AI compute prices drop. Not tomorrow, but within 18 months. SMBs planning their AI infrastructure today have the opportunity to choose from more providers at better prices."
Davarion Group & LabsReal Impact for SMBs
- 01Lower inference costs in the medium term: Cloud providers that adopt Huawei Ascend 950 chips will be able to offer cheaper GPU-hours, potentially reducing the cost of running proprietary AI models or using APIs by 30–50% within 12–18 months.
- 02More cloud AI provider options: Current reliance on AWS, Azure, and Google (which predominantly run on Nvidia hardware) will decrease as new data centers —especially in Latin America and Asia— adopt the DeepSeek + Ascend 950 stack.
- 03Open-source does not mean unstable: TileLang, DeepGEMM, and the rest already ran production models at DeepSeek. The battle-tested code significantly reduces early adoption risk.
- 04Immediate recommended action: If your business hosts AI models or plans to run proprietary models on-premise, ask your cloud provider whether they plan to support Ascend 950; early movers will be able to negotiate better compute contracts.
To understand the magnitude of this move, it helps to appreciate what CUDA represents: since 2007, Nvidia's software ecosystem has been the only practical path for training and inferring AI models at scale. Frameworks like PyTorch and JAX were built on top of CUDA. This gave Nvidia a near-unbreachable competitive moat —not because of the hardware itself, but because the cost of rewriting software for other chips was prohibitive. TileLang attempts to solve exactly that problem: by offering an abstraction compatible with CUDA patterns, it dramatically reduces migration friction. If the open-source community adopts it (which depends on whether Ascend chips become accessible outside China), downward pressure on AI compute pricing will intensify significantly.
At Davarion Group & Labs, we help businesses in Houston TX and across Latin America design their AI infrastructure independently of any single vendor. If your business currently spends more than 15% of its tech budget on AI APIs or cloud compute, it's worth reviewing your architecture before the hardware market shifts again. Contact us at davarion.com for an AI cost optimization consultation.