OpenAI, Anthropic Costs Push More Startups to Build Off Cheaper Open Models

The $15.6 billion legal startup Harvey built its business around training AI models like OpenAI’s GPT-4 to do specialized work for lawyers. Recently, however, soaring artificial intelligence costs have nudged it to rethink its dependence on the AI giants.

After a March update to its AI agents, Harvey’s customer usage spiked, but its gross margins dropped sharply — from about 50% at the beginning of the year to -50% by June, according to a person familiar with the matter.Now, the startup has joined a growing group of software companies that are embracing open-weight AI, including increasingly capable alternatives from China. These offerings are typically cheaper than proprietary technology from US AI developers and let firms like Harvey create custom models with their own data.

Investors like Sequoia Capital and General Catalyst are backing the trend, which is giving firms a way to save on one of their biggest costs, while granting them more control over their technology — instead of outsourcing it to OpenAI and Anthropic. This push risks cutting into the AI giants’ revenue as both gear up for highly anticipated initial public offerings in the near future. And as debate swirls over the need to pace cutting-edge AI, for some startups, it’s also becoming a hedge against a future in which the most advanced models could be slowed or restricted.

Harvey released a model of its own in August, powered by China-based Moonshot AI’s Kimi K3, which can perform close to Anthropic’s best offerings at a fraction of the cost. That launch, plus other tweaks to its AI usage, have made Harvey’s gross margins positive again, according to people familiar with the efforts. The company declined to comment on specific financials for this story.

Similarly, healthtech startup Abridge recently announced it was building a custom foundation model for clinical settings, trained on Nvidia’s open models. AI customer support startup Decagon said it now flows 80% of queries through its own models. In fintech, startups including Ramp and Rogo are exploring training their own models for the first time. Coding companies including Cursor, now part of SpaceX, and $48 billion Cognition were some of the first AI applications to release bespoke models.

While model-building was previously too expensive for Ramp, it’s considering those efforts now after raising $750 million in June because open-weight systems have improved so dramatically, the startup’s co-CEO Karim Atiyeh said. “It made absolutely no sense a year ago. It’s starting to make a lot more sense now.”

As compute costs climb, “it’s completely reckless not to be thinking about that and optimizing for it," said Atiyeh, whose firm is currently in talks to bring in more capital.

Not all investors are bought in, however. Matt Kraning, a partner at Anthropic backer Menlo Ventures, argued that custom model-building doesn’t make sense for every company, since it comes with specialized talent needs and higher upfront costs. For some, he said, it’s merely a marketing tactic.

“Having your own model or not is such the wrong question,” he said. “In most cases, it tends to be a lot of cosplay.”

The open-weight debate

Until late last year, AI applications’ performance mostly depended on the models underneath them, said Harvey president and co-founder Gabe Pereyra. Training those base models was an extremely expensive effort usually left to the biggest AI labs. Any efforts to train custom models weren’t so meaningful in comparison, he said, making it important for companies like Harvey to pay up for the best base options.But the rising costs of the underlying models were putting pressure on some of those companies. So far this year, Harvey has seen a twenty-fold increase in AI token usage, according to the company, a spike its founders knew wouldn’t be sustainable if they kept relying on the higher-cost models from OpenAI and Anthropic.

While Anthropic and OpenAI have more recently released lower-cost models in response to customer concerns about AI costs, both now charge enterprises for their model usage on top of base subscription fees, a shift that’s punished “tokenmaxxing” pushes. Uber, for example, burned through its full-year AI budget by April after encouraging its engineers to maximize their use of Anthropic’s Claude Code.

“If you don’t optimize your costs by fine-tuning your own model, you are by definition not efficient,” said Dr. Lan Xuezhao, the founder and managing partner of San Francisco-based venture firm Basis Set.

She added: “I don’t think the company will be fundable” if it doesn’t look to build models for its own use.

Anthropic and OpenAI have also moved more assertively into startup territory. This year, both have been hiring aggressively, releasing plug-ins and launching pilot programs across industries including legal, finance and healthcare.

At a recent investor forum at Anthropic's offices, attendees discussed the rise of startups building their own models. Anthropic mentioned companies including Rogo and healthcare startup OpenEvidence on a slide about how to "ride the curve" of value, according to a source familiar with the matter. It also referenced Harvey's internal work with AI, with a reminder that the startup still needs Claude's Opus, one of Anthropic’s most advanced models, for its hardest tasks.Still, some startups worry about what could happen if those model makers cut off their access. Only a week after SpaceX finalized its acquisition of Cursor, OpenAI announced it was suspending the coding startup’s access to its models, citing past instances of Elon Musk’s companies violating OpenAI’s terms of service.

Open-weight offers another option. Closed models from OpenAI and Anthropic keep their internal parameters under lock and key, but open-weight ones publicize those parameters for external developers to download and modify in a process known as post-training.

Some, like Kimi K3 developer Moonshot, even offer their own engineers to help customers fine-tune their off-the-shelf models into customized internal ones, according to two people familiar with the discussions.

The challenges ahead

A shift to open-weight models has its sticking points.

Concerns of data privacy and cybersecurity risks have prompted US lawmakers to weigh various restrictions to open-weight models. Both US security agencies and US-based labs have accused Moonshot and fellow Chinese AI developer DeepSeek of distilling the labs’ AI, which is the process of training a new, lower-cost model by continuously querying an older “teacher” model.

Rogo’s training efforts, which haven’t previously been reported, have necessitated discussions with some of its large customers with hesitations about Chinese models, according to the startup’s head of product Strib Walker.

AI data provider Mercor’s Chief Executive Officer Brendan Foody said his customers’ questions about Chinese models usually come from worries over data usage, rather than potential regulatory crackdowns. Some developers argue, however, that the training process for custom models can be so extensive that the model’s parameters ultimately don’t look anything like when they started. “They really go from being an open-source model to a Rogo model,” Rogo’s Walker said.

Also, post-training isn’t exactly an easy lift. Menlo’s Kraning pointed to the competitive market for specialized AI talent, with engineers that can command salaries into the millions of dollars and could be easily poached by top companies like OpenAI and Anthropic.

Plus, he said, companies fine-tuning their own models need large swaths of proprietary data, or else they must be able and willing to pay for that data. Harvey, which can’t access its customers’ sensitive legal data to train its models, buys that expertise from Mercor instead.

Not all companies that attempt building their own models have found it worth the squeeze. Salespeak, a company that has mostly avoided venture capital, previously announced plans to build its own large-language models. It ultimately ditched the effort after a few months of trial and error. It did not see “major advantages” over using market-ready large language models, such as those offered by Anthropic and OpenAI, according to co-founder and CEO Omer Gotlieb.

For example, Elorian CEO Andrew Dai, who is building visual AI models, said the economics of running open-source models don’t always make sense. Downloading open-weight models and managing the computer infrastructure to host the models can be expensive. For some nascent companies with lower traffic, such as his, it may be cheaper to just pay for a closed model and spend money based on usage, he said.

Startups building their own models generally aren’t under the impression that they can stop using Anthropic or OpenAI’s tech entirely, even when these labs aspire to compete directly with those startups.

Logan Bartlett, a managing director at Redpoint Ventures which is invested in Anthropic, Abridge and Ramp, only expects the push to reduce reliance on Anthropic’s models to grow. Still, he says, startups will continue to use the best intelligence available, even when it costs them. “They’re not going to cut off their nose to spite their face,” he said.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论