Beyond CUDA: What DeepSeek's Move to Huawei Hardware Actually Changes
DeepSeek and Huawei have jointly opened a practical path for Chinese AI developers to train and run models on domestic hardware rather than relying on Nvidia GPUs. DeepSeek, one of China's most prominent AI laboratories, released open-source software tools built specifically for Huawei's Ascend 950 chips. Huawei backed the initiative, and the two companies collaborated to fine-tune a 128-chip cluster. The tools cover the layers developers need to write, optimise and scale heavy AI workloads across non-Nvidia hardware, and they are free for anyone to use or adapt.
That is the news. The interesting part is what it says about where the real battle in AI infrastructure sits: not in silicon, but in the software stacked around it.
The real moat was never silicon
Nvidia's commercial dominance has long rested less on the chips themselves and more on CUDA, the proprietary software environment surrounding them. Once a team has built its training pipelines, kernels and tooling around CUDA, migrating to alternative hardware becomes prohibitively expensive. The lock-in is the product.
US export restrictions already limit access to Nvidia's top-tier data centre GPUs in China. What this release does is lower the switching cost for organisations that are looking, or are now required, to keep training and inference on domestic hardware. When the cost of leaving a platform falls, the platform's strongest defence weakens with it.
A moat is only a moat until someone builds a bridge. DeepSeek and Huawei just published the blueprints.
What changes, and for whom
For Huawei
A much larger market
Open, capable software dramatically expands the commercial case for its AI accelerator systems. Hardware is only sellable when the software around it works.
The 128-chip reference run is the proof developers ask for before committing.
For Chinese developers
A second path
Less critical reliance on a single foreign supplier, and a credible stack for teams that cannot, or will not, depend on Nvidia.
Choice is leverage, even when the alternative is not yet equal.
For Nvidia
A regional alternative
A formidable alternative is forming in a market where software lock-in has traditionally been the strongest defence of all.
The moat still holds. It is simply no longer unbridged.
What this does not mean
It does not mean Huawei's chips have closed the raw performance gap with Nvidia's flagship GPUs. It does not mean DeepSeek is designing its own silicon. What it demonstrates is narrower, and more important: a leading Chinese AI lab now treats Huawei hardware as a primary target for cutting-edge workloads, and shares the supporting software openly so the wider ecosystem can follow.
Readout: what the release actually contains
- Open-source tools for writing and optimising AI workloads on Ascend 950 chips.
- Support for scaling across multi-chip clusters, proven on a 128-chip fine-tune.
- Free to use or adapt, with Huawei's backing behind the project.
Why it matters beyond China
The lesson travels. Every serious platform eventually meets the same economics: the software layer, not the hardware, decides how expensive it is to leave. That is as true of cloud APIs as it is of GPUs, and it is the reason I keep making the same argument to clients about owning the serving layer rather than renting it by the token.
Portability is becoming a strategy, not a nicety. Teams that keep their tooling abstracted, their evaluation harnesses hardware-agnostic and their options open will be the ones able to move when the economics shift. Organisations that hard-code themselves to one vendor's stack will pay for the privilege, whichever vendor wins this round.
If you're weighing up where your own inference should run, and what it would take to keep your options open, email brandon@kreostudio.co.uk.
Reader signal
Was this useful?
Work with KREO Studio
AI engineering, data science and design architecture, from Plymouth to the wider UK.
