How OpenAI Beat Math

Mathematicians are likely feeling the kind of anxiety that software engineers have experienced for months in the wake of OpenAI release on Tuesday of a whopping 722 research papers with AI-created solutions to a range of problems across math subjects.

AI’s success solving ultra-difficult math problems—even though it still makes mistakes doing simple math—is due to several factors. Solutions to math problems can often be verified, much like software code, and there’s a lot of overlap between mathematics and the work that AI researchers do (and which these researchers are trying to automate in pursuit of RSI).

And with enough computing resources, the AI can essentially do the heavy-duty number crunching that would take humans years to do, and it doesn’t give up. One mathematician who is part of a group advising OpenAI on how to publicly disclose its math breakthroughs told me that they refer to these abilities as brute force and stamina.

Also, while human mathematicians struggle to track dozens of variables at once as they pursue tough problems, AI models handle lots of parameters simultaneously with relative ease, this person said.

Moreover, AI models don’t fall prey to the psychological biases that afflict humans. There have been multiple instances of human mathematicians overlooking simple solutions to a tough math problem simply because they thought it would be impossible for such a difficult problem to have such an obvious solution, according to the mathematician in the OpenAI advisory group. In several of these cases, OpenAI’s models investigated those simple solutions and found out that they actually worked, they said.

Similarly, mathematicians often give up too quickly when looking for proof that a theory is false—especially when that theory supports the argument they want to make, the mathematician from the advisory board said. That’s not the case with AI models, which are willing to find such evidence as long as humans give them the server power to do so.

There could be a limit to how far AI can go in the field of mathematics. Models tend to miss the forest for the trees, the mathematician said. The point of mathematics isn’t necessarily to solve a single hard problem, but to use that solution and problem to come up with new fields of mathematics or to find general principles that can be applied to other types of problems.

If models are hyperfocused on solving one-off problems, they might never have the capability to connect those ideas and solutions to come up with more broad, meaningful discoveries, this person said. This is where human mathematicians might still have an edge, the mathematician said, as long as they don’t become too dependent on AI to give them the answers for everything.

But the mathematician’s assertion remains to be seen. AI has already shown some ability to connect the dots between concepts from different fields to suggest new types of solutions or experiments that human experts in those individual fields might overlook.

AI Deep Dive: Trial and Error Won’t Fix AI’s Bad Behavior, Says Berkeley Prof. Stuart Russell

David Robinson, the OpenAI employee who led the writing of the company’s safety reports until his recent resignation, wrote in an essay that OpenAI’s “culture is broken.” He argued that “OpenAI has thrived by trial and error (which it calls ‘iterative deployment’), looking for problems and improving its guardrails in response. But this approach, by its very nature, guarantees periodic failures—and the scale of those failures is growing as systems get more capable.”

On this point, he and UC Berkeley professor Stuart Russell agree. Russell, who co-authored the leading textbook on AI, joined the latest episode of AI Deep Dive to break down the challenges involved in steering the goals of AI models. I asked Russell whether this trial and error approach to safety will work in the longterm. His answer was short: “no.”

But, he went on, trial and error may be the best that AI companies can hope for on the current path of AI development. There is “a total lack of rigorous science or engineering” that goes into steering the goals of AI models, he said. The current state of the art methods are “not even alchemy.”—Rocket Drew

Here’s what else is going on…

Overheard

OpenAI announced Intelligent UI, a new interface for ChatGPT that will integrate more visuals into users’ experiences, including graphics, tappable buttons, forms, charts and more.

Microsoft on Wednesday unveiled its latest effort to run AI in PCs powered by its Windows software, rather than running the AI in the cloud, which the company said would bring down costs for customers.

Elon Musk announced on Tuesday night that SpaceX’s AI unit will now use some AI models from competitors to power Grok Bot, signaling a shift away from relying exclusively on models developed in-house.

Sriram Krishnan, who advised President Donald Trump on AI policy before leaving the White House in June, is raising a venture fund that is targeting about $500 million to invest in growth- and later-stage AI companies, The Information reported.

Mike Kubzansky, the former CEO of philanthropic investment firm Omidyar Network, is launching a group for large investors that will facilitate information sharing about the AI industry to foster more trustworthy practices, he told The Information.

Policy Watch

The U.S. Treasury Department said on Wednesday it had fined the parent company of startup accelerator Plug and Play Tech Center, the first penalty issued under an outbound investment restriction program targeting China’s tech sector.

Deals and Debuts

See The Information’s Generative AI Database for an exclusive list of private companies and their investors.

Broadcom has been working in recent weeks to arrange more than $50 billion in financing for a custom AI chip it is developing with OpenAI, with Apollo Global Management and Blackstone among the lenders it has approached, the Wall Street Journal reported.

Isomorphic Labs, the AI drug-discovery startup spun out of Google DeepMind, is in early talks to raise new funds at a valuation of at least $40 billion, which could reach as high as $50 billion, Bloomberg reported.

Biren, a Shanghai-based AI chipmaker, conditionally agreed to place 130 million new shares on Hong Kong's exchange for gross proceeds of about HK$4.04 billion (roughly $515 million), through CICC Hong Kong Securities and UBS's Hong Kong branch.

The parent company of Manus, the AI agent startup that recently separated from Meta Platforms, said Thursday that it has raised more than $500 million in a new funding round. Butterfly Effect said in a social media post on WeChat that the funding round was led by new investors Boyu Capital and IDG Capital. Its existing investors Tencent, HSG and ZhenFund also participated in the round.

Ledgebrook, which uses AI to underwrite specialty insurance, raised $200 million in primary equity financing co-led by Allianz X and Rockefeller Capital Management.

General Medicine, which runs an AI-assisted healthcare marketplace, raised $120 million in a Series B funding round led by Andreessen Horowitz.

Parallel Systems, which builds autonomous, battery-electric freight-rail vehicles, raised $100 million in a Series C funding round led by AVP.

Nous Research, which makes the open-source Hermes AI agent, raised $90 million in a Series B funding round at a $1.5 billion valuation led by Robot Ventures.

Mecka, which pays people to record themselves doing everyday tasks to train humanoid robots, raised $60 million in a Series B funding round led by Sequoia.

Stuut, which uses AI agents to automate enterprise order-to-cash collections, raised $52.5 million in a Series B funding round led by Insight Partners.

Hadrian, an Amsterdam-based startup that uses AI to test companies’ systems for cybersecurity vulnerabilities, raised $40 million in funding led by Forgepoint Capital International and SmartFin.

Healthleap, which develops AI to screen patients for often-missed conditions, raised $38 million in funding across seed and Series A rounds from Sequoia Capital, First Round Capital and Hummingbird Ventures.

Vesta, an AI startup that helps lenders originate mortgages, raised a $30 million round led by Conversion Capital.

An unnamed stealth AI lab founded by Keyu Tian, a former ByteDance intern, raised about $30 million at a $200 million valuation from 5Y Capital and IDG Capital. The roughly 10-person team is building “world models” aimed at robotics, interactive video and autonomous driving.

Wally Health, which runs an AI-powered subscription oral-care platform, raised $25 million in a Series A funding round led by Maveron.

Rivercell, which is building an AI “virtual cell” world model to predict how human cells respond to drugs, raised $25 million in seed funding led by HV.

Ampersand, which builds integration infrastructure for enterprise AI agents, raised $15 million in a Series A funding round led by Bessemer Venture Partners.

Anjney Midha, a former general partner at Andreessen Horowitz, and former executives at Google, Apple and Nvidia have launched National Compute, a company that aims to make it easier and more affordable for smaller companies or startups to access compute.

Google launched SynthID, a website allowing users to identify AI-generated media, as well as Playground, a platform that allows users to create games using AI powered by Gemini, Nano Banana and Lyria.

Thank you for reading the AI Agenda Newsletter! I’d love your feedback, ideas and tips: [email protected].

If you think someone else might enjoy this newsletter, please pass it forward or they can sign up here.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论