The Chinese AI Infrastructure Boom: Introducing the SemiAnalysis China Datacenter Model

China sits at the frontier of the global model race. GLM 5.3 and Kimi K3 are the latest in a run of striking open-weights releases. ByteDance's Doubao serves 345M monthly users as China's ChatGPT, and Seedance is the State-Of-The-Art video generation model.

Every one of those models runs on a datacenter, and China has been building them at a pace that has gone largely unmeasured outside the country. The biggest tenant files no 10-K. Several of the largest landlords have never listed. Most of the primary sources are in Chinese. So the market settled on two lazy assumptions: China is big, and China is empty. Published estimates of China’s datacenter capacity differ by 15x, and reports keep citing high vacancy rates.

As with our flagship SemiAnalysis Datacenter Model, we use building-level data to show which of the most widely cited narratives hold, starting with how large the market is.

The global model tracks over 5,000 facilities across four regions. The US leads the world with 56GW of capacity as of 2026YE, followed by ~15GW for APAC ex-China, ~14GW for EMEA, and ~2GW for Latin America. Until now, it stopped at the border of China. Today we cross it.

Our tracking of 1,000+ datacenter facilities across over 60 players shows that China alone boasts a fleet of over 24GW. Bigger than EMEA. Bigger than the rest of Asia. This excludes ~20GW of dated pipeline and another ~30GW of announced projects.

Growth is at an inflection point. In 2Q26, the combined capex of Alibaba, Tencent, and Baidu (“BAT”) reached $20B, more than doubling YoY, and for the first time on record, all three posted negative free cash flow. This is the largest capex step-up in the sector’s history.

That total leaves out the largest spender of all, ByteDance, which remains private. According to our China Datacenter Model, ByteDance alone occupies roughly a fifth of delivered datacenter capacity in China, and it rents nearly all of it, making it the single most important customer for all wholesale colocation players in the country.

The listed players get the attention, but they are only the visible tip of the iceberg. This is also true for datacenter landlords. GDS and VNET, the only two Chinese datacenter landlords listed in the US, signed 1.3GW of wholesale orders in 1H26. But according to our China Datacenter Model, they captured barely a third of ByteDance and Alibaba orders in 2024–2026YTD.

The buildout is a nationwide effort, and the state-owned carriers are also part of it. China’s datacenter market was historically telecom-dominated. The state carriers held the majority of the market share in the 2010s, and still own a third of national capacity today. The national power grid companies invest aggressively too. Their combined capex accelerated from 2024, ending the 14th Five-Year Plan 24% over the original blueprint. The 15th plan (2026–2030) layers another 40% on top, to over $746B (¥5T).

Largely free of power constraints, labor shortages, and public protests, China routinely delivers 100MW datacenter facilities in under 12 months. Modular DCs are the standard playbook, which Tencent deployed its third-gen modular design in 2014. What is old news in China is only now being adopted at scale in America, as we detailed in .

The growth does not stop at China’s border.Overseas leasing by Chinese hyperscalers is set to double from 2026 to 2029 and approach ~4GW of leased capacity in our SemiAnalysis Datacenter Model. That number still understates their offshore compute, because it excludes the hundreds of thousands of GPUs they rent from Western clouds, which we covered at length in How Oracle Is Winning the AI Compute Market.

However, the market also faces some dark clouds. Vacancy rates are high, and datacenter developers compete heavily on price. Chip supply is also constrained, due to export restrictions. But as explained below, neither has stopped AI datacenters from being built and filled at a remarkable speed.

In this report we cover, in order:

  • Introduction to China’s datacenter industry: the world’s second-largest market, built retail-first, and why that history produced overbuild and high vacancy alongside an AI capacity shortage.
  • The westward buildout: how the “Eastern Data, Western Compute” policy and hyperscale AI demand are moving the industry inland, at construction speeds and unit costs the West cannot match.
  • Behind the paywall: we go tenant by tenant through ByteDance, Alibaba, Tencent, Baidu, and Huawei, from their AI and datacenter strategy to leasing and self-build footprints.

Introducing the SemiAnalysis China Datacenter Model

This report is a preview of the SemiAnalysis China Datacenter Model. It is the China counterpart to our global SemiAnalysis Datacenter Model, built to the same standard, bottom-up and building by building.

Contact our sales team at sales@semianalysis.com for more information.

What you’ll find inside:

  • Tracking over 1,000 facilities across 60+ operators
  • Quarterly delivery curves through 2032 for every facility, with construction status and timelines
  • Tenant identification per building, supplemented by tenant allocations per customer-mix for leading operators
  • Hyperscaler self-build vs leased splits and per-tenant capacity paths
  • A lease signing tracker and a lease delivery tracker, on one reconciled basis
  • Eastern Data, Western Compute (“EDWC”) hub and cluster analytics, load-growth and power-draw forecasts
  • An interactive full map view showing where the capacity is concentrated

Consider it the opening chapter. As with our global Datacenter Model, we will keep adding layers (AI capacity, power, suppliers, supply/demand analysis) with every update.

Let’s dive in.

China’s datacenter industry: built retail-first, flipped by AI

One of America’s key advantages in the AI infrastructure race is that its datacenter industry grew up hyperscale. Multi-hundred-MW campuses have existed for years, purpose-built for a handful of enormous tenants. China’s industry was shaped by telecom carriers and thousands of small retail deployments. Understanding that history explains what’s strange about the market today, from overall high vacancy coexisting with an AI capacity shortage, to why the carriers still own a third of everything.

The following chart previews the whole story. Between 2010 and 2022, a leading colocation player in China watched its utilization rate shrink from the mid-70% range to the mid-50% range. Then, from 2023 onwards, its segment disclosure reveals a two-speed market. Wholesale buildings are filled fast on AI demand, back above 70% utilization, while the legacy retail racks sit near 60%.

Era 1 — Carrier hosting (pre-2015): datacenters as telecom real estate

Unlike the rest of the world, where hyperscalers and colos dominate, China’s datacenter market was historically telecom-dominated. The three state carriers (China Mobile, China Telecom, China Unicom) held a combined 60–70% market share, with GDS and VNET, the largest third parties, below 5% each. The carriers enjoyed structural advantages in licensing and networking and benefited from captive demand. A large share of Chinese datacenter demand comes from government, SOEs and regulated industries, which default to state-owned suppliers.

For its first two decades, a Chinese datacenter was essentially telecom real estate, featuring low-density retail racks (~90% of facilities ran below 2kW) rented to websites, online gaming and CDN customers. This was often offered as a bundled product with bandwidth, served by small urban switching centers around Beijing, Shanghai and the Pearl River Delta. Third-party operators existed largely as resellers of carrier space and bandwidth. The hyperscalers were barely a factor, as Alibaba Cloud was launched in 2009 and Tencent Cloud only opened to the public in 2013.

This era matters today because its stock is precisely the capacity now sitting vacant. The legacy retail racks no longer match what an H20 or Ascend server demands, and it’s impractical to retrofit given partial occupancy and high cost. The headline utilization number of ~50% is largely a monument to Era 1 (and Era 3, as we’ll see).

Era 2 — The cloud land grab (2015–2021): growth at any price

The cloud era began around 2015. The two leading players, Alibaba and Tencent, concluded that it was a winner-take-all market and opened a capex race. Alibaba Cloud revenue compounded at a ~110% CAGR across 2015–2019, at the cost of a brutal price war and EBITA losses. Its core products were marked down more than 50% after 17 consecutive price cuts during a 12-month window. Tencent responded in kind, driving combined hyperscaler capex to a 2021 peak.

Third-party colocation platforms (GDS, VNET, Chindata) scaled alongside them. Retail developers, meanwhile, kept building small urban sites on the assumption that cloud demand would spill over indefinitely.

Era 3 — Digestion (2022–2023): the music stops

The wild expansion stopped abruptly in 2022. Policy headwinds, macro slowdown, and share gains by state-owned clouds pushed Alibaba Cloud growth from +50% in 2021 to +23% in 2022, then to effectively zero. Tencent management framed it as a repositioning “from pursuing revenue growth to healthy growth” and saw a similar slowdown. Tencent cut capex budgets by nearly half in the middle of 2022. Alibaba did so in 2023.

Almost simultaneously, in February 2022, the National Development and Reform Commission (“NDRC”) launched Eastern Data, Western Compute guidelines and named ten computing clusters within eight national hubs. These regions, especially the western ones, enjoy fast track permitting and cheap green power.

The result was an oversupply despite a slowdown in delivery. Pricing was hit most. Power-exclusive market rates were cut in half from ~$80/kW/month. On the August 2026 earnings call, GDS management estimated another 18 months to fully digest the price reset, even though it has been going on for a few years.

Era 4 — The AI supercycle (2024–): full steam ahead

DeepSeek and the hyperscalers’ own GenAI ambitions woke the market up over the course of 2024-2025. Combined BATB (ByteDance/Alibaba/Tencent/Baidu) capex went from roughly $35B in 2024 to over $50B in 2025 and is on a trajectory to $100B for 2026.

Wholesale orders returned at a scale never seen before, and the western hubs finally found their tenant. GDS and VNET together signed roughly 1.3GW of wholesale capacity in 1H26, with less than 10MW of retail orders. The AI era is wholesale. The legacy retail stock is not participating.

But watching only these two listed names badly understates total demand, as they only capture ~1/3 of ByteDance and Alibaba orders in 2024–2026YTD per our China Datacenter Model.

To recap, the four-stage history explains China’s datacenter market today: (1) a retail-focused legacy stock around Tier-1 cities, running 30-50% utilization, often physically unsuitable for AI densities; (2) a wholesale AI-driven buildout happening somewhere else entirely; and (3) a carrier share in structural decline, as hyperscaler self-build and third-party wholesale take over the growth.

SemiAnalysis is a reader-supported publication. To receive new posts and support our work, consider becoming a subscriber.

Eastern Data, Western Compute: the pipeline is heading west

The EDWC policy designates eight national computing hubs, implemented as ten clusters, spanning both the demand centers (Beijing-Tianjin-Hebei, Yangtze River Delta, Greater Bay Area, Chengdu-Chongqing) and the resource-rich west (Inner Mongolia, Ningxia, Gansu, Guizhou).

This is the industry’s top-level design blueprint, and unlike most “guidance” documents, it actually changed how the industry develops.

The policy itself has been covered extensively. What is missing, and what matters, is the transmission mechanism.

In China, the first permit a datacenter developer applies for after securing land is an energy-consumption review filed with the local Development and Reform Commission. That quota is the binding cap on how large a facility can be. Any project consuming over 10,000 tons of coal equivalent (>10MW) cannot legally begin construction without passing this review.

The review regime dates to 2010 and became a rationing tool over time. When EDWC arrived in 2022, the quota system became the most effective mechanism because the same offices that issue quotas are the ones implementing the EDWC policy, so quotas effectively stopped construction flowing to non-hub locations.

Guangdong is the cleanest example. The province stopped issuing datacenter energy quotas in 2021-2022 and reopened applications only for the EDWC hub region in Shaoguan. Similar quota control governs the other eastern metros. Meanwhile, EDWC regions hand out energy quotas, land approvals and tax incentives generously.

Despite the mechanism, for the first two years, migration was slow because the demand of the day (cloud, gaming, e-commerce) required low latency to population, so hyperscalers kept fighting for scarce quota near Tier-1 cities.

A little-noticed episode illustrates how the quota system played out in practice. In mid-2023, Guangdong’s Energy Bureau inspected Tencent’s two self-built campuses in Qingyuan and found each running at roughly 13–14x its approved scale. The projects were labeled “approved small, built big”, and ordered rectification within six months. The local grid company was instructed to hold power supply to the approved quota. Even for Tencent, incremental builds now belong to the province’s designated hub node.

Today, the campuses keep operating, but as our China Datacenter Model shows, Tencent’s subsequent capacity went to Shaoguan. The reason is not simply regulatory pressure.

Two other things changed the picture.

First, telecom-led fiber investments have significantly reduced latency. Shaoguan again is the model case. Since 2022, the local government has invested ~$600M (~¥4B) in network infrastructure, building 13 ultra-high-speed 400G all-optical routes covering the Greater Bay Area, achieving 1.3ms one-way latency to Guangzhou, 1.66ms to Shenzhen, and under 3ms to Hong Kong. A direct Shenzhen-Guangzhou-Shaoguan hollow-core fiber corridor has been planned to cut latency a further ~30% from 2027 onwards. At those numbers, a Shaoguan rack is functionally a Shenzhen rack for all but the most latency-critical workloads.

Second, AI demand did the rest. Training workloads don’t care about latency, but they care about power prices.

Inner Mongolia is the best example here. The region has China’s largest renewable fleet and, critically, western Inner Mongolia is served by the Mengxi grid, China’s only provincial grid outside State Grid and Southern Grid, with its own multilateral trading mechanism and independent pricing.

As a result, Inner Mongolia has emerged as China’s Johor, the default destination for hyperscale AI capacity, with generous quota allocations, land measured in square kilometers, and, most importantly, power prices at roughly half the Tier-1 city level.

The Johor parallel runs deeper than land and power. On economics, Chinese rents are far lower, but so is the capex, while Johor rents are higher, but so is the cost to build. Our data shows that the two markets' yield-on-cost ranges overlap, with the best Inner Mongolia campuses underwriting to the return of Johor's international developers. However, the trends diverge. Inner Mongolia construction costs keep falling as the prefab approach and domestic substitution compound, while Johor developers face capex inflation and overrun risk in a supply chain running hot. Chinese operators building in Johor enjoy geographical arbitrage leveraging the Chinese supply chain and earn higher yields.

Zoom out from Inner Mongolia, though, and the ten clusters tell very different stories. It is easiest to see when the map is weighted by what is actually live and what is committed, rather than by what the policy designated, as the map below shows.

A recurring media take holds that EDWC is a boondoggle of empty western shells. The simplest rebuttal is to look at where the hyperscalers put their own money. As our tracking shows, ByteDance currently operates three self-build clusters: Datong in Shanxi (its earliest), Wuhu in Anhui (an EDWC cluster, serving east-China inference) and a twin-site Inner Mongolia program (Horinger and Ulanqab), with two Ningxia entities registered in Zhongwei (another EDWC region) in July 2026 pointing to a fourth.

Outside of Inner Mongolia, development across the clusters is sharply divergent, but much of that divergence is by design. Gui’an was positioned from the start as a storage hub, as cold data doesn’t mind the mountains. The most famous anchor tenant is Apple’s iCloud China datacenter, operated with Guizhou-Cloud. Less reported is that Huawei and Tencent both have self-build clusters in this region.

Chengdu and Chongqing are different stories. They are pitched as national compute hubs but lost the AI-demand contest to Inner Mongolia on power price and grid flexibility. EDWC succeeded in concentrating the buildout inside the map, and the market then picked the winners within it.

Shanxi is the opposite. It is the largest datacenter region outside the EDWC hub system entirely (the red diamond on the map above), but attracted close to 2GW commitment from ByteDance, Baidu, and JD across self-build sites and colocation leases. The reasons are abundant cheap power (Shanxi is China’s coal province), proximity to Beijing, and first mover advantage. Hyperscalers went there before EDWC existed and kept building out, proof that the quota regime channels the buildout, but electrons and geography still make the site selection.

SemiAnalysis is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Speed and cost: a different sport

The US buildout is power-gated with interconnection queues, transformer lead times, behind-the-meter workarounds, which we have covered at length . China’s is chip-gated: generation is abundant, state-owned utilities treat datacenter load as a development win, permitting is a policy funnel rather than a bottleneck, and construction labor is plentiful. With no physical friction on the facility side, the build cycle is driven purely by internet platform demand.

On speed, the standard delivery time for a 100MW facility has compressed from ~18 months in the cloud era to ~12 months in the AI era, with state-of-the-art projects finishing faster still.

For reference, our US modular deep dive found the best-case full-modular buildout in the US still takes 12–18 months of construction, sitting behind 12–13 months of permitting that cannot overlap. For China, the permitting period before construction is typically 3–6 months, and faster in EDWC regions.

Several structural factors contribute to the overall faster speed.

First, speculative shells ahead of orders. Because hyperscalers now run T+6-month delivery requirements in their tenders, colocation developers pre-build civil works at their own risk, making the tenant competition effectively a race among half-finished shells. When GDS signed a 152MW AI mega-order in early 2025, the customer required delivery within six months and got it.

Second, building greenfield sites in western EDWC hubs changed the construction method. Away from Tier-1 land constraints, developers are moving from conventional multi-story concrete to low-rise, long-span steel halls. The steel is made in a factory while foundations are dug on site, then trucked in and lifted into place by crane. Concrete is used only in the ground, so there is no waiting for it to harden before the next floor goes up. VNET’s Ulanqab footage shows a hall go from bare foundations to a finished building within one summer-to-winter season.

Tencent’s T-Block takes the same logic to its extreme. The entire datacenter arrives as factory-built modules, with the “building” shrinking to a leveled site and a simple steel shell, sometimes no building at all.

Third, prefabrication turns a serial schedule into a parallel one. The bulk of MEP work moves off-site into factories, so electrical and mechanical fit-out no longer waits for the building. The headline example is Alibaba’s CUBE 5.0 “100-day datacenter”, comprising 30 days of factory production and site prep, 50 days of on-site installation, and 20 days of commissioning, measured from a finished shell to the day servers arrive. This roughly halves M&E and commissioning stage compared with market standard 6-month, and brings ground-break to commissioning down to roughly 7–9 months.

Every factor above points in the same direction, which is more factory, less site. Follow that logic to its endpoint, and you would expect Tencent-style full-modular container campuses to be the national standard by now. They aren’t, for two very Chinese reasons.

First, the labor math is different. Modularization’s biggest payoff in the West is dodging scarce, expensive skilled trades. Chinese site labor is affordable and abundant, so a traditional steel-and-prefab build already hits aggressive schedules. The marginal speed gain from going full-container is small.

Second, the financing dynamics play a role. Chinese colocation developers pledge their datacenters to banks as loan collateral. A proper building with a 20-year design life appraises well, but a yard of containers appraises poorly and carries lower residual value. When your business model is “borrow against the asset, lease it to ByteDance,” you build a real building.

Today, full-modular builds remain limited to hyperscaler self-builds and Southeast Asia projects, but overall construction speed is as fast as needed.

On cost, capex keeps falling even as speed rises. Continued supply-chain compression and rising domestic content have pushed shell-and-fit-out costs to a fraction of comparable US wholesale costs, as we show in our model. Alibaba credits its modular approach with a further >10% cost reduction compared to its earlier version on top of the schedule gains.

SemiAnalysis is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

The demand: five hyperscalers, multiple playbooks

The demand side of China’s AI buildout is concentrated in a handful of names: ByteDance, Alibaba, Tencent, Baidu, and Huawei.

Like their American counterparts, they span social media, e-commerce, search, gaming and short video. However, the consumer internet that built these companies has run out of new users. Today, China’s internet penetration already sits at ~80% and grows only ~1.5% a year.

The growth engines of the past are all in single digits.

  • Social media: Tencent’s core asset, WeChat, grew MAU just 2% YoY in 2025.
  • Search engine: Baidu’s 2025 revenue fell 3% with online marketing down double digits.
  • E-commerce: Alibaba’s customer management revenue fell 7% YoY in the quarter ended June 2026.
  • Short-video: Kuaishou’s 2025 DAU grew just 2.7% YoY.

Outside of consumer apps, the same titans lead the cloud industry, with a structural handicap. The market is dominated by Alibaba, Huawei and Tencent alongside the three telecom carriers. But unlike the US hyperscalers, which generate roughly a third or more of their cloud revenue outside the US, Chinese cloud revenue is largely domestic. Alibaba Cloud, the most global of the group, still earns only ~20% of cloud revenue overseas, and the carrier clouds are effectively 100% domestic. In terms of “addressable GDP,” the core market of the US tech titans is simply much larger.

Against that backdrop, AI has become the growth engine, which drives the capex inflection point in 2Q26 we mentioned earlier.

Beyond funding their own model R&D, the big platforms have built an investment web across China’s GenAI labs. Alibaba invested in five LLM unicorns. Tencent backed six and wrote the largest external ticket in DeepSeek’s Series A.

The labs, in turn, train and serve on the hyperscalers’ clouds (in the case of the Alibaba/Moonshot relationship) or have their models served inside the hyperscalers’ products (in the case of the Tencent/DeepSeek relationship). Either way, the hyperscaler datacenter expansion is the compute foundation for the entire domestic AI ecosystem. Hence the deep dives that follow.

We provide a snapshot below showing where the hyperscalers stand:

The leasing: Today, ByteDance and Alibaba each issue annual leasing demand at gigawatt scale and both are actively expanding overseas. ByteDance's is the largest and leases majority of its footprint, so its tender calendar effectively sets the order book for every wholesale operator in the country. Alibaba, as we show below, splits the demand across self-build, BOT and wholesale. Tencent is the most disciplined. It spent the 2022–23 downcycle closing small leased metro sites and consolidating into its own campuses, before leasing picked up again in the AI era. Baidu tenders at a fraction of the others’ scale.

The self-builds: In the 2010s, Chinese hyperscalers began self-building datacenters as their businesses reached sufficient scale, and the strategies have diverged since. Alibaba, Tencent, and Baidu self-build their strategic campuses and lease the rest. Huawei self-builds almost everything because of its business model. ByteDance, the primary tenant, is now ramping its own campuses fast. Our China Datacenter Model tracks all self-builds, and the map below shows the current landscape.

Between pure leasing and pure self-build, a hybrid structure called BOT also emerged. The real anatomy is a two-layer ownership split. The hyperscaler owns the land and the civil shell, while the colocation partner invests in, owns and operates the M&E layer and recovers its capital through 8–10-year service contracts. In this way, the hyperscaler keeps control and speed, the partner’s balance sheet carries the capex, and the IRR is contracted rather than merchant. All three of BAT run some version of this playbook. Alibaba partners with AtHub, Tencent with Kehua Data, and Baidu with Aofei. Alibaba takes the mixing furthest, running self-build, BOT, and wholesale leases side by side within the same cluster, as we discuss below.

SemiAnalysis is a reader-supported publication. To receive new posts and support my work, consider becoming asubscriber.

ByteDance: TikTok parent, China’s largest tenant, among the world’s largest GPU users

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论