'Sketchy AF': What to Know About How OpenAI Staff Discussed Book-Pirating

Authors including John Grisham, David Baldacci, Jodi Picoult and Jonathan Franzen sued OpenAI and Microsoft three years ago for copyright infringement.

A court filing unsealed Thursday details internal messages and testimony laying out how OpenAI employees—including executives—talked about the company’s use of pirated books to train an early ChatGPT model, as well as their technology’s potential impact on authors. The filing was made in support of the authors’ request that a federal judge rule in their favor ahead of a trial.

The tech companies have argued that their actions constituted “fair use,” which allows for some copyright material to be used without explicit permission, and that their use of the material was ultimately transformative.

Here’s what to know:

Starting in 2019, OpenAI began using books from file-sharing site Library Genesis—LibGen for short—to help train its large language model. The site, which provides free access to books and scholarly articles, had faced global accusations of pirating copyright material.

OpenAI’s then-general counsel David Lansky was among those who recommended free book downloads from LibGen as a possible data source, according to the filing.

Tom Brown, then a top GPT-3 engineer, and Ben Mann, another member of the technical staff, described LibGen as “sketchy AF,” according to messages and testimony cited in the filing. A federal court ordered LibGen to shut down in 2015 and academic publisher Elsevier received a $15 million judgment against the site in 2017.

OpenAI staffers discussed removing references to LibGen from papers that could later be public, according to the filing.

In one instance, Dario Amodei, then a senior researcher at OpenAI, who is now chief executive of rival Anthropic, asked in Slack whether it was “sketchy to call our corpuses ‘Books1’ and ‘Books2’ and not say what they are, particularly when in fact they are Libgen (which is a slightly sketchy source).”

Separately, Mann wrote that a description of the data sets was “deliberately vague since it’s libgen.”

A number of those involved in OpenAI’s early model training now work at Anthropic.

OpenAI said employees who used LibGen to create the earlier ChatGPT models are no longer at the company, and the LibGen data set wasn’t used to power the company’s current ChatGPT models. An Anthropic spokeswoman declined to comment. Convergent Research, where Lansky currently works, didn’t respond to a request for comment.

OpenAI priced out the potential cost of buying large quantities of books on multiple occasions.

Greg Brockman, OpenAI president, wrote to OpenAI co-founder Ilya Sutskever in late 2022 that they could purchase lots of books, “but they are expensive per token so we haven’t prioritized.” (Tokens are the basic measurement unit for AI use.)

Varun Shetty, OpenAI’s vice president of media partnerships, testified in the lawsuit that he wasn’t aware of OpenAI’s purchasing books at scale and scanning them for training purposes, according to the filing.

Safe Superintelligence, Sutskever’s current AI lab, didn’t respond to a request for comment.

Bob McGrew, then vice president of research, wrote in a June 2022 Slack channel, “now is the right time to excise Libgen from our systems and storage,” given how much OpenAI was in the news.

Removing LibGen would prevent researchers from being able to reproduce earlier GPT-3 or GPT-3.5 results, he said, adding, “but would be very valuable for legal reasons.”

OpenAI removed the data sets from its library.

“The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon,” Jack Clark, OpenAI’s former policy director, said in an internal warning in 2020. “There will be a point where a bunch of artists express worry about what we’re doing here and we’ll likely ignore their concerns and release anyway.”

Clark, a co-founder at Anthropic, also acknowledged that the company’s “work in this area will make people unemployed.”

Technical staffer Tarun Gogineni wrote on X in December 2022 that he knew artists were complaining about increased AI-generated competition, but viewed potential job losses as “acceptable economic disruption.”

OpenAI CEO Sam Altman testified to Congress in 2023 that the company “does not want to replace creators.”

A number of other author lawsuits against AI companies are winding their way through courts against AI companies, including from textbook writers.

Meanwhile, Anthropic agreed last year to pay at least $1.5 billion in a major settlement with authors who alleged the company had pirated their work to train its models. A judge ruled the AI company could use books for training in some circumstances, but the “fair use” argument didn’t cover millions of books that Anthropic obtained from known ebook piracy sites.

Wall Street Journal parent News Corp has a content deal with OpenAI.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论