Fireworks, Fal Consider New Rounds as Inference Demand Soars

Startups such as Fal and Fireworks AI that sell access to AI models and servers have been ringing up sales as developers use them to run models quickly, sparking a new investor rush to pour money into them.
Fal, which specializes in providing such inference services for video and image generation models such as Google’s Nano Banana models, has spoken to investors about raising new funding at a $15 billion valuation, according to two people familiar with the discussions. Talks are early and the company could aim for a higher valuation, of $17 billion to $20 billion, according to a third person.
Either way, the round would roughly double the five-year-old’s $8 billion valuation notched in a funding round this spring. The talks follow a jump in annualized revenue to $800 million, double its March pace, according to one of the people.
Fireworks AI, one of the biggest inference providers by revenue, is also considering a new funding round and could target a $30 billion valuation if it does, according to a third person. That would nearly double its valuation from a round announced in July.
Fireworks said on Friday that it had completed a $132 million sale of its employees’ stock led by Atreides Management, at a $17.5 billion valuation—the same price as its last fundraising round, announced in July. Its annualized revenue jumped to $1 billion by July, up five times from what it was generating a year earlier. Its recent revenue couldn’t be learned.
Baseten, another inference provider, is also in talks with investors to raise new funding at a $26 billion valuation, reported Axios. And Modal, another such provider, is in talks to raise funding at a roughly $15 billion valuation, about triple what it was worth in a round four months ago, Bloomberg reported.
The flurry of fundraising discussions reflects the boom times for startups that provide access to frontier models and help them customize open-source models.
These startups have seen usage grow as developers take advantage of advances in open-source AI, which can often be cheaper than those from OpenAI and Anthropic. Rather than running open-source models on their own chips, which can be difficult to set up and expensive due to a shortage of such chips, many developers prefer to access the models through an easy-to-use application programming interface, such as the ones offered by Fireworks and Fal.
However, such companies could face pressure as developers of closed-source models like OpenAI and Anthropic become more aggressive with cutting prices. Amjad Masad, CEO of coding startup Replit, said at The Information’s AI Agenda Live conference Wednesday that his company is likely using less open-source AI these days than at the beginning of the year, due to how cheap some OpenAI models have gotten.
What’s more, the inference providers are at risk of a margin squeeze as the cost of the compute they rent from cloud providers spikes. These inference companies have historically had gross margins of around 50%, below the level of best-in-class software providers, which usually see 70%-plus gross margins.
Inference providers argue that they will be able to improve their margins through optimizing the ways the models run on chips. Together AI, which provides inference as well as rents out AI servers, has also been buying chip servers and renting them out from its own data centers.