DeepSeek Bets Big on Huawei Chips to Bypass U.S. Export Controls

DeepSeek CEO Liang Wenfeng told investors that a major priority for his company is to use more domestic chips to train its models, adding that he expects Huawei Technologies to start delivering training chips to DeepSeek as early as the fourth quarter this year, according to two people with direct knowledge of the matter.

DeepSeek’s growing adoption of Huawei chips during AI model training—a process that requires more powerful chips compared to running the models—demonstrates how both companies are joining hands to break away from the stranglehold imposed by U.S. export controls on advanced chips. Liang’s most recent remarks, made during a closed-door, in-person meeting on Sunday, came as DeepSeek finalizes its second funding round, which aims to raise 50 billion yuan ($7.5 billion) at a valuation of 500 billion yuan.

To attend the meeting, investors were required to come to DeepSeek’s Beijing or Hangzhou offices, while Liang dialled in online. Investors were also asked to hand over their phones, other electronic devices and their bags during the meeting, and were provided with a pen and pieces of paper for taking notes, as the company sought to prevent similar leaks that happened during its first fundraising. The transcript of a hourslong video call between Liang and investors in May —-the first time the elusive founder talked to a large group of investors—was later circulated by Chinese media outlets and on social media, prompting the company to put on hold its second fundraising for a week.

During the Sunday meeting, Liang described training models on Huawei chips or other Chinese processors as one of DeepSeek’s biggest bets, saying the effort “has to work,” according to the people. He also said DeepSeek expected to receive a new batch of Huawei chips in the fourth quarter this year or the first quarter next year that can be used in training its latest models, without elaborating how much supply Huawei could offer to the company.

DeepSeek and Huawei didn’t immediately respond to requests for comments.

The Information reported previously that DeepSeek started experimenting with using more Huawei chips last year, in light of persistent constraints with advanced Nvidia chips. Liang believes that it’s only a matter of a few years before Huawei chips will become as good as Nvidia’s and that DeepSeek should move first by adapting to semiconductor hardware outside Nvidia’s dominance.

Yet the move toward Huawei doesn’t mean DeepSeek can completely wean itself off Nvidia’s hardware anytime soon. The company is still using Nvidia chips, likely obtained from the black market, to train its new models. Liang told investors that the company is investing about 30 billion yuan to expand its computing capacity. DeepSeek raised 50 billion yuan in June.

DeepSeek is joining the race to develop increasingly larger models—the size of which is typically measured by the number of parameters, the numerical values a model learns during training that determine how it transforms input into output.

DeepSeek is currently training a 2-trillion-parameter model, which is larger than its existing V4 flagship model’s 1.4 trillion parameters, and the company also plans to develop a much larger model with 8 trillion parameters, Liang told investors at the meeting. By comparison, Moonshot AI’s Kimi K3 model has 2.8 trillion parameters, and is also developing an even larger model. Anthropic and OpenAI don’t disclose the number of parameters in their models, but various industry estimates suggest their latest models are at least several trillion parameters.

The larger the models, the more chips and computing power are required during both training and inference—the process when AI models perform tasks and answer questions.

In May, during DeepSeek’s online meeting with investors, Liang estimated that training a model comparable in scale with the largest systems such as those being developed by OpenAI would require about 50,000 Nvidia GB300 processors or 200,000 Huawei Ascend 950 chips, according to the leaked meeting transcript, the veracity of which has been confirmed by The Information. Liang said at the time that DeepSeek wanted to use more Huawei processors for training but couldn’t get enough supply from Huawei.

Huawei has encountered production constraints because of limited access to advanced memory—an essential component that allows AI processors to move large amounts of data quickly—as well as shortages of other components, according to one of the two people with knowledge of Liang’s Sunday remarks and another person with direct knowledge of the situation.

While Liang didn’t specify what Huawei chips DeepSeek is expecting to receive, the timing broadly matches Huawei’s product roadmap announced last week: its Ascend 950DT, designed for training and part of the inference process, is due to become available in the fourth quarter, while the more powerful 960DT is scheduled for the first quarter of 2027.

Huawei said in April that its AI chips had handled part of the training of DeepSeek’s smaller V4-Flash model and that the companies had worked closely to adapt DeepSeek’s V4 model family to Huawei’s Ascend computing systems.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论