General Catalyst Backs Coding Data Startup Proximal As It Reaches Revenue Milestone

It’s no secret that leading AI developers are on the hunt for data to improve their models, especially in lucrative areas like software engineering. The insatiable appetite for data has lifted a whole host of so-called data labeling or data curation startups, including older ones like Scale and Surge and newer ones such as Fleet and Micro1.
The latest beneficiary is San Francisco-based Proximal, which is coming out of stealth with $15 million in seed funding led by General Catalyst at a $300 million valuation, the company told The Information exclusively.
Proximal raised its seed round at the end of last year, when it was generating a negligible amount of revenue, co-founder and CEO Calvin Chen said. Chen had cofounded the 40-person startup last fall with his roommate Justus Mattern, who previously worked at compute startup Prime Intellect, which itself is only two years old.
Now, just 10 months later, the company has surpassed $200 million in annualized revenue, he said. (Proximal calculates annualized revenue by taking the revenue generated in the quarter so far and multiplying it by four, Chen said.)
That makes it a small but fast-growing firm in the data labeling industry. Mercor surpassed $2 billion in annualized revenue in June and Handshake was generating nearly $1 billion in annualized revenue in April, for instance.
Many of the other data labeling firms hire hundreds of thousands of experts in fields like math, biology and software engineering to come up with uber-difficult questions for AI models to answer. But Proximal has tried to take a more automated approach, Chen said. That involves—naturally—using AI models to generate coding tasks to train AI models on.
For instance, Proximal might notice that, when an engineer asks a model to recreate a website, the model can build a functioning site but struggles to make it look like the original, Chen said. To help improve the model’s capabilities in that situation, Proximal can set up a task asking the model to recreate a specific site, then compare its version with the original to see whether it succeeded. Rather than having a person write the code to apply that test to a variety of different websites, Proximal can ask an AI model to do that instead, which helps to automate the process, he said.
Proximal has said it uses AI to help it with other data generation tasks, like using AI coding agents to simulate the back-and-forth communication between human software engineers as they develop a large, complex codebase over multiple days.
The more-automated approach helped Proximal minimize costs. While typical data labeling startups have gross margins of between 30% to 50%, Chen said, Proximal has “software-like” margins, he said. (He wouldn’t expand on what exactly Proximal’s margins are, but traditional software companies typically have gross margins of 70% or higher.) The company is also profitable on a net-income basis, he said.
Another factor that apparently boosts the bottom line: Chen says Proximal doesn’t typically respond to requests that major AI developers send to multiple data labelers when they are seeking specific training data, Chen said. Such requests are often highly competitive and leave the data providers with little leverage to negotiate good contracts, he said.
Proximal tries to flip the script. It evaluates frontier models, sees gaps in their capabilities, and then comes up with its own interesting examples of training data that it believes will help close those gaps, he said. Those include long-duration coding tasks that might take hours or days rather than minutes, as well as tasks that frontier models like GPT-6 Astra or Fable-5.1 might still struggle with, Chen said.
Such tasks are obviously difficult to come up with. As models have gotten smarter, for instance, Chen said that it’s gotten more difficult to prevent issues like “reward hacking,” where models cheat their way to the answer without learning the “right” way to do something. To try to prevent this, Proximal and other data companies have to carefully specify to models what they’re allowed and not allowed to do and monitor their chains of thought to verify that the models solved the question correctly when they're generating training data, he said.
Here’s what else is going on…
Cognition Taps Ex-Meta Security Chief Amid Cyber Push
OpenAI isn’t the only AI startup beefing up its cybersecurity credentials. Cognition, the $48 billion startup behind the Devin coding agent, has hired cybersecurity veteran Alex Stamos to serve as its new chief information security officer, the company told The Information.
The hire signals that Cognition is serious about expanding its sales of AI security products, which has recently become a booming market among enterprises. It also shows how the startup is trying to burnish the security of its broader suite of products amid a backdrop of growing concerns about the industry’s preparedness for autonomous hacks.
In an interview, Stamos—who previously served as Meta’s security chief, and more recently as an executive at the security firm SentinelOne and at the AI startup Corridor—said he will oversee both Cognition’s internal security and the growing suite of security AI products it sells to customers.
Cognition already sells AI tools that automatically scan customers’ code for vulnerabilities and suggest ways to patch it, and Stamos said the company is plotting more AI products that appeal to both software developers and cybersecurity professionals.
“You have an interesting merge happening between product security teams and developer teams already—the personas are blending,” Stamos said.
That strategy mirrors the approach of firms like OpenAI and Anthropic that are increasingly selling cyber-capable models like Astra and Mythos to cybersecurity teams. But Stamos said Cognition thinks Devin can give customers similar performance by harnessing a mix of models—including open source ones—that are less expensive than models like Mythos, which are proving budget-breakers for heavy users.
“Defenders are spending a ton of money on Mythos-class models and that's not sustainable,” Stamos said. “So one of the things we have to do as an industry is provide more options for customers to defend themselves much more cheaply.” –Aaron Holmes
AI Deep Dive: Without Faster Chip Design, AI Models Are Stuck in a Rut, say Ricursive Founders
New kinds of AI models could be much better than the best models today, but these new designs are being held back by the slow pace of developing new AI chips, said Anna Goldie and Azalia Mirhoseini on the latest episode of The Information’s AI Deep Dive. They are the co-founders of Ricursive Intelligence, a year-old startup developing AI systems for designing AI chips.
“There’s this deadlock right now,” said Goldie, in which “current models are designed for current chips, and we kind of can’t break out of that because each of these edges is too slow.”
For example, if a new idea for an AI model is “2X better than some existing idea, but the chip is like 4X worse for that, then you’re at a disadvantage,” said Mirhoseini.
Right now, it can take years to create a new chip, but AI could speed up that process by automating the stages known as physical design and design verification, they argued. That faster pace could allow new chips to be “co-designed” alongside new AI models.—Rocket Drew
Big Numbers
Anthropic said nearly a quarter of its $4.6 billion in revenue last year came from just two customers, according to a confidentially filed IPO prospectus reported by Reuters on Monday. That’s one metric that could give investors pause ahead of its IPO, expected later this year. The report didn’t name the customers.
The company also posted an eye-popping $42 billion net loss last year, although about $34 billion of that total stemmed from an accounting charge related to the increase in per-share value of outside funding it raised, which investors will likely overlook. The Information previously reported its net loss was about $15 billion in the first half of this year, with about $12 billion related to the same accounting charge.
Overheard
OpenAI has decided not to release the AI model that it had intended to release as GPT-6.1 Astra due to its results on safety tests, a spokesperson said Monday. The model performed worse than GPT-6 Astra—OpenAI’s newest model, released earlier this month—on tests measuring whether it pursues users’ goals and is transparent with users about its actions.
OpenAI also apologized on Monday to the Australian government for not immediately notifying its administration that its agents had breached some public services websites. “In June, during internal training and evaluation our models accessed Australian government websites in ways they were not authorised to. We also should have handled our response better. We are sorry and working to do better in the future,” OpenAI wrote in a blog post.
People on the Move
Instinct, Silicon Valley’s hot AI assistant of the moment, is hiring Ben Schwerin, an investor at Coatue Management, as its first chief business officer. In Schwerin, the year-old company—which Monday said it raised $1 billion at a $10-billion valuation—is getting a hot consumer tech connection and operator.
Meta Platforms is launching a new division to sell its AI tools to businesses, tapping MongoDB Chief Executive Chirantan Desai to lead the unit as chief enterprise platform officer.
Court Watch
Florida Attorney General James Uthmeier on Monday asked a Florida state court for an injunction that would bar OpenAI from developing new AI models “without third-party approved safety guardrails” and from allowing minors to use ChatGPT. The injunction would be effective throughout Florida.
Deals and Debuts
See The Information’s Generative AI Database for an exclusive list of private companies and their investors.
EliseAI, a voice agent startup, raised $350 million in funding at a $4 billion valuation led by Andreessen Horowitz and Bessemer Venture Partners. The company said it has surpassed $200 million in annualized revenue.
SiMa.ai, which makes chips for physical AI devices like robots and drones, raised $150 million in an oversubscribed Series C funding round led by Fidelity Management & Research Company and Amplify, at a $1.45 billion valuation. The Information first reported on the round here.
Samsung Electronics and five of its affiliates are investing $1 billion in Helix Digital Infrastructure, an AI infrastructure venture established in June by private equity firm KKR, the companies announced on Tuesday.
Maven Robotics, which builds robots that handle goods in warehouses, raised $100 million in a Series A funding round raised from RoboStrategy, LocalGlobe, Vine Ventures, XTX Ventures and others.
Reco, an AI security startup, raised $55 million in funding from AT&T, Forestay and Quadrille Capital.
Knowin, a Chinese startup building AI-powered home robots, raised 300 million yuan (about $45 million) in an angel funding round led by JD.com.
Outmarket AI, an AI platform that automates insurance paperwork, raised $34.5 million in a Series B funding round led by SignalFire.
Modulate, which builds AI models that analyze human voice, raised $25 million in new funding led by Future Ventures.
Advanced Micro Devices said Monday it has agreed to buy Dr. Fei-Fei Li’s World Labs, the two-year-old maker of models that can simulate three-dimensional spaces, for about $8.2 billion in an all-stock transaction. The deal is expected to close at the end of this year, AMD said.
Mitratech Legal acquired BotDojo, an AI agent-orchestration startup, to accelerate development of its ARIES AI platform for legal teams.
Meta signed AI infrastructure agreements with Australia's Firmus for contracted GPU compute capacity, built on Nvidia‘s DSX platform, at Firmus’s Southeast Asia data centers to support Meta's AI research and model training.
Anthropic released Claude Sonnet 5.5, saying it runs 30% faster and costs 30% less for most work than Claude Sonnet 5, the previous model in the series.
Nvidia on Monday released new software and hardware products to prevent hacking incidents such as the breach in July of Hugging Face and other companies by OpenAI’s AI. Other AI developers such as Anthropic and Google have disclosed similar hacking incidents involving their AI agents, which took unintended actions such as improperly accessing certain data.
NinjaTech AI, an early-stage startup backed by Amazon and led by a former Google product management executive, unveiled a new product that promises to help large companies take the guesswork out of forecasting the costs of running agents around the clock to complete long-running tasks.
Manus, the AI agent startup that recently separated from Meta Platforms, launched on Monday a new standalone app for personal agents that could compete with Meta’s Muse. The new app, Cue, provides each user with multiple personal AI agents.
LiveRamp, which helps advertisers use customer data to target and measure ads, said Monday it is expanding its partnership with OpenAI so advertisers can target their ads in ChatGPT based on customer data they collect, a feature already available with Google and Meta.
Thank you for reading the AI Agenda Newsletter! I’d love your feedback, ideas and tips: [email protected].
If you think someone else might enjoy this newsletter, please pass it forward or they can sign up here.