Startups Find Old Slack Threads, IT Tickets Are Suddenly in High Demand

Eight days after AI agent startup Warmly in late June agreed to be acquired by HubSpot, CEO Maximus Greenwald found an unusual email in his inbox. Mercor, a San Francisco startup that pays contractors to train AI models for labs such as OpenAI, Anthropic and Google, was offering to buy or license Warmly’s code base and records of its tasks and staff communication, for as much as $300,000.
Warmly fielded four such approaches from companies like Mercor in the days after the HubSpot deal was announced, Greenwald said. He turned them down. They asked for every dataset the acquisition wasn’t expected to include: Slack messages, GitHub and Asana content, Google Drive documents, and years of transcripts of staff meetings that had long outlived their usefulness.
For some time, firms that gather data for AI labs have been buying the code bases of failed or sold startups and selling them to AI labs to improve their models’ coding capabilities. But in recent months, demand for data beyond the code bases has soared, according to more than two dozen interviews with founders who have been approached with offers, as well as startups that gather and clean up the data for the labs and financial advisers who handle transactions for large tech companies.
The buyers scrub that data of all personally identifying information and sell it to companies training AI models, say these people. Most of these sales and licensing deals aren’t made public, since they contribute to the secret sauce of model training. Spokespeople for the AI labs didn’t have a comment or respond to requests for comment.
Driving this demand is increased use of AI agents that carry out tasks such as answering customers’ technical questions, managing clinical drug trials and verifying invoices on behalf of users. To train models for these agents, AI companies are eager for communication that shows how humans navigate an IT problem or recordings of people using software. It’s the same kind of hunger for data that led Meta to try monitoring its employees’ computer use to feed real-world information into its models.
“The labs have scraped the whole internet. Now they need real interactions between real people, and that is the kind of data they’re looking for,” said Bobby Samuels, CEO at Protege, a startup that vets whether a dataset can legally be sold—say, by checking that the seller holds the intellectual property rights.
The total value of transactions handled by Protege jumped to at least $100 million so far this year, up from $30 million for all of last year, he said. The startup takes 25% to 30% of the transaction value.
The most highly prized materials are corporate records such as IT requests and Slack threads as well as communications with customers, Warmly’s Greenwald said.
“If you want to train the ultimate AI employee, you have to be able to train on day-to-day employee conversations,” Greenwald said.
Hubspot, Warmly’s new owner, said in a statement: “This was a talent-focused acquisition for us. No data of any kind was acquired as part of the transaction.”
Companies on the hunt for data are looking at startups that are shutting down or being acquired, or startups that haven’t raised venture funding in a couple of years. The latter may welcome a way to make money off everyday customer and staff exchanges—as long as such information isn’t protected by rules, such as those on healthcare patient privacy.
The AI data companies either buy or license the data on behalf of clients like OpenAI or resell it to the labs. In some cases, they act as intermediaries, cleaning up the data or making sure it’s privacy compliant.
‘Hunting for Records’
Beyond Slack threads, buyers are interested in corporate content such as slide deck presentations made on Asana or stored in Google Drive, emails showing engineers talking through a problem, a chief financial officer on Zoom talking to colleagues about a company’s performance, and engineers’ changes to a code base on GitHub.
The labs can pair those exchanges with other events, like an IT ticket for a problem—say, a website login failing on Safari—to train a model on how to solve a problem.
The labs have scraped the whole internet. Now they need real interactions between real people.
The labs are also hunting for proprietary datasets that differentiate a company from its competitors. In logistics, for instance, such strategic data connects thousands of carriers and shipments, describing how the business fits together. Warmly’s proprietary dataset consisted of 330 million customer contacts developed over seven years, said CEO Greenwald.
“The demand for proprietary datasets in every vertical is exploding,” said Shubh Sinha, co-founder and CEO of Integral, a startup that prepares corporate data for licensing to AI labs.
He’s seen increased demand for datasets from healthcare, energy and financial services companies in the last year amid efforts by Anthropic, OpenAI and others to develop apps for specific industries, such as life sciences. But sellers of such data need to make sure they don’t violate customers’ terms of service or fall afoul of privacy regulations.
Integral strips identifying information out of the records—swapping an Indian name for a different Indian name, for example—keeping the record usable while making the person unidentifiable. Sinha said labs have also been staffing up teams to source specific datasets and negotiate their anonymization. “If you scrub too much, the data loses its value,” he said.
Mercor uses an anonymization process to mask personally identifiable data when it obtains data on behalf of clients, according to a person close to the company.
Data companies and their clients are also increasingly scouring for video recordings that can train agents how to act like humans.
Every week, Michael Litt, founder and CEO of Vidyard, fields a version of the same question: Would you sell us your data? The 15-year-old company sells software for video marketing and tutorials and hosts the videos. Vidyard has amassed an archive of roughly 3.7 million hours of videos, growing 50,000 to 100,000 hours a month.
In particular, potential buyers are seeking examples of workers demonstrating software. For instance, labs are specifically looking for footage of humans using Salesforce, Netsuite and HubSpot, Litt said. The videos also capture what happens when things go wrong.
But licensing or sales agreements may amount to only a fraction of what a startup believes it’s worth—prompting it to balk at such deals.
Jason Pressman, managing director at Shasta Ventures, said the big companies are typically offering to license data for a few million to $10 million a year, for companies he would value at $100 million to $300 million in an acquisition.
“They’re trying to cut sweetheart deals to license the data, stripping the company of its most valuable asset, leaving nothing behind,” he said. Pressman believes they should acquire the company outright in such a scenario. Such acquisitions haven’t happened yet, but big tech companies looking for data have approached three companies his firm has backed, he said.
Alix Coutures is a reporter covering telecom and startups based in New York City. She graduated from the M.A Business concentration at Columbia Journalism School in May 2026. She previously covered tech and telecom in Paris, at the French business publication Challenges. She can be reached at [email protected] or on Signal at Alix_coutures.34