OpenAI Math Result Stokes Data-Sharing Concerns

In case you missed it: we got an eye-catching X post on Tuesday night from Evan Hubinger, a leader of Anthropic's alignment efforts. Responding to the resignation announcement of an Anthropic employee who warned that AI companies aren’t doing enough to ensure AI won't kill humanity, Hubinger wrote that the employee is “correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” Hubinger added that he was worried about future AI models, not those currently available, but still, yikes!

Meanwhile, our colleague Jyoti sizes up the significance of Meta’s new personal AI agent, Muse, unveiled on Tuesday. That’s below. But before we get to that:

This may go down as the year the AI industry got a crash course in mathematics. Breakthroughs in AI models have yielded announcements on such arcane subjects as the “Erdős unit distance problem” and the “Jacobian Conjecture,” which observers have watched as an indicator of AI progress in other domains.

The latest announcement, OpenAI's claim on Tuesday that it had solved the vexing “Navier-Stokes existence and smoothness problem,” has drawn attention for another reason: questions over whether its math-genius models may have gained an edge by incorporating data from two human mathematicians that had used OpenAI's Codex in their work on the same problem. Those questions fed into concerns about the possibility of AI companies learning from customers’ use of their AI models to develop competing products.

At issue is the work of New York University mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge, who had been laboring on the Navier-Stokes problem and related problems over the past year. They had been using several AI products as part of their efforts, including OpenAI’s Codex. As the pair were preparing to publish their breakthroughs, OpenAI heard about their work and decided to see if it could solve Navier-Stokes, the company said. OpenAI used “an internal model that is significantly more capable than GPT‑6 Astra” powering on the order of 10,000 agents working at the same time.

Alpöge and Buckmaster published their work before OpenAI released its own, more expansive results on Tuesday. In an accompanying statement, Buckmaster said that in his conversations with OpenAI over the past week, he asked whether its models “had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.”

OpenAI responded in a post on X Tuesday that the company couldn’t rule out the possibility that its models had trained on Buckmaster and Alpöge’s work—though the company deemed that possibility “unlikely.”

In response, Alpöge offered a bit of snark, saying “props to them for straight coming clean” in a post on X.

The dispute was clouded by competing accusations of bad faith over claims of credit for the math work and the broader acrimony of the bitter OpenAI-Anthropic rivalry.

But it’s not clear whether OpenAI violated any business promises (or academic norms).

Both OpenAI and Anthropic say they don’t use enterprise customers’ prompts or responses—essentially the chat logs between those firms and Claude or ChatGPT—to improve their models. They also both say they do improve their models using chat logs from some individual users, but only those who haven’t opted against sharing their data in the apps’ settings.

Many businesses worry the AI companies are getting their proprietary data anyway, possibly because workers are incidentally giving it away when they use their personal chatbot accounts for work.

OpenAI said its employees and agents “did not see any of their work through any means until they released it publicly—in particular, no specific user data was accessed in order to solve this problem.” The company and Alpöge and Buckmaster dispute the similarity of their proofs and results.

OpenAI’s chief research officer, Mark Chen, weighed in again later Tuesday in response to Alpöge’s post, saying: “Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.”

That still leaves unanswered a key question that hangs over the controversy: whether Alpöge and Buckmaster were using enterprise or individual AI accounts—and if individual, whether they had opted out of sharing their data. Buckmaster and Alpöge didn’t respond to requests for comment Tuesday.

If they had failed to opt out, it is possible that OpenAI’s models ingested some of the work they placed in Codex as part of their training data. That might mean OpenAI deserves somewhat less credit for its Navier-Stokes breakthrough, but OpenAI is within its rights to train on such information.

On the other hand, if they were using enterprise accounts or they had opted out of sharing their data, OpenAI’s inability to rule out training on their work starts to look more concerning.

The answer will help determine whether this dustup has implications for OpenAI’s business beyond what it says about its models’ math skills.

(As for the math itself: this is the part where we’ll endeavor to explain for the intrepid readers among you, Lord help us. The Navier-Stokes equations describe how fluids move. Until Tuesday, it was unknown whether solutions to these equations are always “smooth,” and answering that question comes with a $1 million prize. Alpöge and Buckmaster focused on a version of the problem that assumes fluids have zero viscosity. For the more general case, OpenAI showed that solutions are not always smooth by constructing “a vortex, a spinning swirl of fluid, that spirals inward and gets increasingly elongated, like spaghetti.” These are the sorts of things academic mathematicians spend their time on, apparently.)

Here’s what else is going on…

Meta’s New Muse is a Big Test for Consumer AI

Meta Platforms’ new personal AI agent is a big step in its efforts to build new products out of its mountains of AI Investment. It’s also a broader test of whether there's a big business in consumer AI.

Muse, released Tuesday, is an attempt to make one of the hottest facets of today’s AI development—bots that don’t just answer questions but take actions on our behalf—accessible to the masses. It’s designed to do everything from sending emails and booking travel to helping with shopping and carrying out tasks across the services people use every day.

With Muse, Meta CEO Mark Zuckerberg is returning to the roots of the AI boom, which began as a consumer phenomenon with ChatGPT’s launch nearly four years ago. The industry has since shifted hard toward business users, especially this year with the explosive growth of Anthropic’s Claude Code. That has left the market for agentic AI apps for consumers unconquered.

Other companies are trying. The agentic AI apps Instinct, from startup Spear Street Technology, and Townies, from Town, are targeting consumers by automating tasks like booking restaurants, ordering products and scheduling appointments.

Meta has an obvious advantage. It can put Muse in front of billions of people that use its social media apps, and Zuckerberg made clear that it is prioritizing popularity over immediate revenue. There are paid versions at $20 and $100 a month for users who want higher usage limits, but the free version will allow up to 100 million tokens per week. ”Muse is built to help deliver personal superintelligence to everyone over time,” Zuckerberg wrote in a post on X in touting the free version.

The allure of individualized agents was illustrated by OpenClaw, the open-source agent service that became red-hot with Silicon Valley types earlier this year but was too complicated for most non-techies to use.

It remains to be seen whether the average person is anywhere near as interested in automating their workflows and optimizing productivity as those who flocked to OpenClaw. Especially since users must give agents access to a lot of other personal information to be effective: email, calendars, travel information, shopping accounts and potentially financial services accounts.

That’s one reason why Meta has emphasized its efforts to make Muse safe. As I laid out last week, Muse—developed internally under the codename Hatch—was tested for months by Meta employees, who flagged some significant safety and alignment concerns. Meta said it has built in guardrails to address those, and that Muse has privacy and security features baked in such as needing user permission before carrying out a purchase or sending an email.

Zuckerberg’s emphasis on Muse’s free version means Meta is essentially planning to give away a lot of expensive AI compute to find out whether it can turn a sizable portion of its social media users into AI agent users. If it fails to catch on or security and privacy issues draw new negative attention to Meta, it could prove a costly stumble. But if enough people start using Muse, it’ll be worth Zuckerberg’s bet.-Jyoti Mann

Overheard

OpenAI’s price cut on its GPT-5.6 Luna model shortly after its release in July resulted in a tenfold increase in model usage, said OpenAI chief financial officer Sarah Friar in an interview on Monday at Goldman Sachs’ Communacopia + Technology conference in San Francisco. Usage of the model is now “higher than even the next Chinese model,” giving OpenAI the highest market share on OpenRouter, a platform that allows users to access AI models from several different companies in one place, Friar said.

The U.S. government has accused six Chinese AI firms of using large-scale distillation to copy the capabilities of American models, in a move that highlights the intensifying AI rivalry between the two nations. The six Chinese firms under accusation are DeepSeek, Moonshot AI, Alibaba Group, MiniMax, StepFun and Z.ai. The U.S. National Security Agency, the Federal Bureau of Investigation and the Cybersecurity and Infrastructure Security Agency issued a joint statement on Tuesday to alert U.S. companies and organizations about the Chinese companies’ activities.

Policy Watch

Attempts by Senate Democrats to get more clarity about the White House’s AI policies have gone nowhere. A White House official responded last Friday to an August letter from a group of Senate Democrats asking for more details about the White House’s voluntary AI testing framework without a direct response to any of the questions.

Deals and Debuts

See The Information’s Generative AI Database for an exclusive list of private companies and their investors.

Qualcomm announced a data center infrastructure partnership with Amazon Web Services, and in a securities filing said it had issued warrants to Amazon to buy 25 million Qualcomm shares at $161.26 each—a $4 billion investment.

Mistral, the AI lab that develops open-weight large language models, raised €3 billion (about $3.5 billion) in a Series D funding round at a $24 billion valuation.

Cognition, the artificial intelligence startup behind the Devin coding assistant, has raised more than $2 billion at a $48 billion valuation including the new funding, nearly double its valuation in May, the company said Tuesday. Andreessen Horowitz and Accel led the Series E. The Information first reported on its financials here.

Celero Communications, a company that designs the chips that push data down the fiber-optic cables linking AI data centers together, raised $275 million in a Series C funding round. Atreides managing partner and chief investment officer Gavin Baker is joining the board, according to Bloomberg.

Forus, which uses AI to shorten the gap between a doctor writing a prescription and a patient actually getting the drug, raised $150 million in a Series C funding round, according to Bloomberg.

Antioch, a company whose software develops a simulation of a customer's actual robot or drone hardware, raised $32 million in a Series A funding round led by Greylock.

Cymphony, an AI cybersecurity startup, raised $30 million in funding, including a $25 million Series A round co-led by Sequoia Capital and SMBC Fin Atlas Beyond Fund, which valued the startup at more than $100 million including the investment.

Blee, a company whose software reviews the marketing material that banks, brokerages and insurers put out, raised $27 million across two rounds—a $20 million Series A funding round led by Fin Capital and SMBC, and a $7 million seed, according to Axios.

Ollie, a company developing an AI assistant that plugs into a company's email, calendar and documents and then manages schedules, raised $7.5 million in a seed funding round led by Khosla Ventures.

Flamingo, a company that runs an open-source system to automate IT and security services raised $4.5 million in a seed funding round led entirely by Vertex Ventures.

Outline, a company whose AI agents read a company's accounting, billing, payroll and sales systems, raised $3 million in a seed funding round led by Founders Future.

Keep Converting, a company whose software rewrites an online shop's product pages for each individual visitor in real time and keeps whichever version sells best, raised $2 million in a pre-seed funding round from Nuwa Capital and COTU Ventures.

Gaia, a company selling Saudi businesses an AI system that searches and answers questions across their own files, email and internal systems without the data leaving the company, raised $1.5 million in a pre-seed funding round led by Seedra Ventures.

OpenAI said it is working with Samsung Electronics to develop next-generation chips, deepening a collaboration with the South Korean technology giant that already spans data centers and enterprise sales, Reuters reported.

Data labeling and curation firm Turing is launching CEO Bench, a new benchmark designed to test whether frontier AI agents can handle the complexity of real-world knowledge work, rather than isolated benchmark questions. Turing is developing an entire simulated company including over 1,100 company files such as financial records, contracts, HR files, board materials, system exports and months of internal Slack.

Meta Platforms unveiled Muse, its first consumer-facing AI agent, giving users an AI assistant designed to carry out tasks on their behalf across the web.

Suno, an AI music model maker and app developer, released a new model family on Wednesday called Suno v6, which it says was developed using licensed data from music labels and distributors such as Warner Music Group, BMG and Believe.

Accenture and Google Cloud launched the Accenture Gemini Enterprise Business Group to help companies deploy agentic artificial intelligence. Google Cloud will help train up to 1,000 Accenture forward-deployed engineers to work with clients on-site to plan and build AI applications

Thank you for reading the AI Agenda Newsletter! I’d love your feedback, ideas and tips: [email protected].

If you think someone else might enjoy this newsletter, please pass it forward or they can sign up here.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论