Why The AI Compute Crunch Is Hitting Neolabs Especially Hard

The AI chip shortage is causing headaches for tech giants such as OpenAI. Among those hardest hit, though, are the startups trying to build their own AI models—so-called neolabs that need to use their limited funding to secure access to enormous clusters of graphics processing units for the compute-intensive task of model training.

“It’s like VC currency right now to know the current price of GPUs,” said Evan Morikawa, who recently finished talking with about 17 different AI cloud providers in search of compute for his startup Generalist, which trains AI models for robotics.

Morikawa compared searching for AI chips to exploring distant lands: “At a San Francisco dinner party, this is the equivalent of returning from the Far East,” he joked.

Startups training their own models require hundreds or thousands of GPUs or similar AI chips. When Morikawa was in the market for hundreds of chips last winter, he said cloud computing providers were offering contracts at reasonable prices with durations as short as 1 year. Six months into a 1-year contract, Generalist needed more compute, so Morikawa set out again to hunt for around a thousand chips.

He found during his discussions in May and June that prices had increased sharply, contract terms had typically reached 3 to 5 years, and there were 12 to 18 months of delay to access some clusters of thousands of the latest chips. “Now 1-year contracts are almost impossible to get,” he said.

That makes contract duration “the most painful but least talked about” part of compute deals for neolabs, he said. Cloud providers are pushing years-long deals to lock in revenue, which “has made it very very difficult if not impossible for the startups,” said Morikawa. “If your company is 1 year old, signing a 5 year contract makes no sense. The value of that 5 year contract is probably 3 times what you have raised, so you're already underwater on paper.”

Such an expensive and longterm commitment can be especially daunting because neolabs are betting on AI research agendas that may not succeed, and training new models from scratch remains more of a secret art than an established recipe in the AI industry. Compute is the main use of new funds for neolabs, said Morikawa.

Market data further illustrates the problem. The hourly rental price for one Nvidia H100 GPU as part of a 1-year contract rose 50% to $2.60 in mid-June from $1.73 in mid-December, according to Semianalysis. The price increase is especially remarkable given that Nvidia released the H100 in 2022 and has released two new generations of chips since then. For H100 chips available on-demand, Semianalysis doesn’t report prices at all for many recent months; instead it just says “sold out.”

While some neolabs are still able to secure compute for training their models, it can cost an enormous amount of money. Reflection AI, an Nvidia-backed company developing open-source foundation models, recently signed a deal to rent chips from SpaceX for $150 million per month and a multiyear deal with Nebius totaling over $1 billion. That represents a hefty chunk of the $2.5 billion the startup raised in its last funding round.

Not too long ago, startups could get compute on an as-needed basis, but as of February, much of that capacity has dried up, according to the infrastructure lead for a mid-size AI safety research organization who talked to over 20 compute providers between March and July in search of about 60 chips. Previously, the organization relied on multiple providers, including neocloud Runpod.

In response to dwindling capacity, the organization sought new compute deals for a larger number of chips that will allow it to customize larger models and focus more on pre-training, the compute-intensive earliest stage of training a large language model. Reserving chips 1 to 3 years in advance is especially tricky for a company doing such training runs because they require bursts of computing capacity, and future compute needs depend on the results of earlier experiments. For those reasons, renting GPUs on-demand for a few weeks was better suited to their organization’s work, this person said.

Hyperscalers are often focused on larger customers—my colleague Catherine Perloff wrote recently about how a growing number of startups said Amazon Web Services no longer meets their needs. And the smaller neoclouds, such as CoreWeave, Runpod and Nebius, haven’t had much on-demand capacity available since March, this person said.

A Runpod spokesperson said users who wait in line for a request for compute on average get access within a few hours. At the same time, as my colleague Martin Peers wrote yesterday, CoreWeave and Nebius are enjoying the price increases. CoreWeave CEO Michael Intrator in an earnings call Tuesday said “our near-term capacity remains effectively sold out,” despite its investment in increasing capacity. That “is translating into signed commitments on increasingly favorable terms from a broadening set of customers,” he said.

On at least three occasions, the research organization scheduled meetings with providers that were advertising availability, and within 2 days of the meeting, the capacity was sold to someone else, the infrastructure official said. The organization ultimately signed a deal worth several million dollars with Prime Intellect, which is reselling compute from data centers located in India. One employee at the organization bought a cake to celebrate when the deal closed.

“It was many months of so much stress. It is not pleasant to not have an answer as to whether your organization can keep doing research for months,” this person said. “It‘s a new landscape. We’re all adjusting.”

Here’s what else is going on…

Earnings Report

Neocloud Nebius reported a 454% expansion in second quarter revenues to $582 million thanks to surging demand for AI computing.

Tencent Holdings said its capital expenditures in the second quarter nearly tripled from a year earlier to 52.8 billion yuan ($7.8 billion), as the Chinese tech giant beefed up its computing infrastructure to train better AI models and meet growing demand for its coding and productivity tools.

Cerebras, which designs AI chips meant to run AI models extra fast, reported rapid sales growth for the latest quarter but investors sent its shares down sharply on competition and other concerns.

Cisco Systems shares dropped 5% after its fourth quarter earnings, even after the company reported strong revenue growth and said it is being fueled by cloud provider customers increasing their spending on its AI networking chips and switches.

Deals and Debuts

See The Information’s Generative AI Database for an exclusive list of private companies and their investors.

Thrive Holdings, a New York City-based company that buys professional-services businesses and infuses them with AI, raised $2 billion from D1 Capital Partners, Altimeter Capital, SoftBank Group and others.

Kalshi, the biggest prediction market, is in advanced talks to raise at least $750 million in a new financing round at a $40 billion valuation, according to people familiar with the matter.

Lovable, a startup whose AI software lets anyone build, launch and maintain working web apps, raised $400 million in a Series C funding round at a $13.3 billion valuation led by Menlo Ventures and the Scaleup Europe Fund.

Legora, an AI legal startup, is in talks to raise new funding at a valuation of at least $10 billion after growing its annual recurring revenue to $150 million, the Financial Times reported.

Anthropic is in talks to buy Decart AI, a world model developer that also develops tech to make AI chips run more efficiently, for about $6 billion, Bloomberg reported.

CodeRabbit, a startup whose AI software automatically reviews software code, raised $143 million in a Series C funding round at a $1.5 billion valuation led by Atomico and Smash Capital.

Skan AI, a startup whose AI software watches how employees actually use their business applications, raised $63 million in a Series C funding round led by Cathay Innovation and Dell Technologies Capital.

Gravity, a startup developing an advertising system for a web run by AI, raised $30.5 million in a Series A funding round led by Lightspeed Venture Partners and Committed Capital.

Silicon Data, a startup that provides pricing and market data for the AI computing economy, raised $30.5 million in the first close of a Series A funding round led by the Valor Atreides AI Fund.

ClearJet, a startup that develops packages faster and cheaper by using AI to route them onto empty cargo space on flights already scheduled between US cities, raised $25 million in a funding round led by Edison Partners.

Preview, a San Francisco-based AI-video production software company, raised $12 million in funding led by Sequoia Capital.

Diald, a startup whose AI software helps real estate investors do their homework on a property, raised $1 million in a funding round led by Feedback Ventures.

Mindgard, a startup whose software helps companies find and fix security weaknesses in their AI systems, raised $30 million in a Series A funding round led by Album VC.

Chinese AI developer DeepSeek has launched its flagship model, V4-Pro, to mixed reviews from users, with some expressing disappointment. The company officially released the model on Wednesday. A leaderboard compiled by U.S. AI benchmarking startup Vals AI ranked V4-Pro second among all open-source models, just behind Moonshot AI’s popular Kimi K3.

Google raised the prices of its newest Pixel smartphones by $100, a sign of how rising memory chip prices are driving up costs of consumer hardware.

Twitch, the streaming platform, announced that it will now use creators’ content to train AI models for its parent company, Amazon, unless they opt out.

Thank you for reading the AI Agenda Newsletter! I’d love your feedback, ideas and tips: [email protected].

If you think someone else might enjoy this newsletter, please pass it forward or they can sign up here.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论