AirTag Investigation Tracks Rare Book Shipment to Amazon Facility Scanning Books for AI Training

A $29 Apple AirTag hidden inside a rare book helped track a shipment to an Amazon warehouse in Las Vegas, where employees take apart books and scan their pages for AI training data.

This finding contradicts an earlier Amazon statement denying that it destroyed books during scanning. Reporter Emanuel Maiberg placed the AirTag in one book from an order of about 1,000 volumes purchased through Biblio, a marketplace that keeps buyers anonymous.

Amazon's public response avoided any reference to artificial intelligence.

AirTag Tracking Led to Amazon's Las Vegas Book-Scanning Facility

A bookseller agreed to participate after noticing unusually large orders with no clear connection between the titles. The seller suspected these books were being bought for AI training, not for resale.

In the weeks that followed, the AirTag's location data showed the package moving through four states. It first flew to Milwaukee, then stayed at a distribution center near Kenosha, Wisconsin, for two weeks.

After that, a truck took it west, stopping overnight in Grand Junction, Colorado. The last recorded location was the northern part of an Amazon facility called LAS8 in northeast Las Vegas.

LAS8 is mainly a print-on-demand site where books are made. The northern part of the building houses a separate unit called VGT3, according to employees in an online forum. Its entrance features an image of a Tyrannosaurus rex holding an open book.

One employee said, "All we do is scan books." Other workers described a constant flow of large book shipments. Staff cut off the spines to make scanning faster, then throw away the printed books after they are digitized.

At VGT3, employees scan each book's ISBN barcode before scanning the pages. This creates a running list of titles the unit has already processed. The record helps them identify which books have not yet been scanned and focus on those.

Amazon's Response and Why AI Companies Want Pre-2022 Books

In its only public response, Amazon did not mention artificial intelligence. A spokesperson said Amazon "purchases books through commercial channels" to develop and improve its products and services, but did not say what happens to the physical books after scanning.

Two competing AI developers have drawn a firmer line on the practice. Anthropic has said its data-acquisition efforts avoid purchasing and destroying rare or antiquarian books, a position xAI has echoed publicly.

No equivalent assurance has come from Amazon. The seller who placed the AirTag characterized the shipment as niche or slow-moving inventory rather than collector-grade material. "There are different types of value," the bookseller said, referencing historical and sentimental worth left out of any scanning operation's calculations.

AI developers prefer pre-2022 books because of a problem called model collapse. This happens when AI systems are trained too often on their own generated text, causing them to lose accuracy and variety over time.

A July 2024 study from Oxford, Cambridge, and Imperial College London tracked this decline. Another analysis by Ahrefs found that by April 2025, 74.2% of nearly 900,000 new web pages already included AI-generated content.

ISBNdb, a company that provides book metadata, briefly marketed pre-2022 print books to AI companies as material free from AI-generated text or data poisoning. The marketing page was taken down in July 2026. The company later said it was only a test to see if there was market interest.

Anthropic's Similar Book-Scanning Program and the US Legal Framework

This is not the first known operation of its kind. Court filings unsealed during Bartz v. Anthropic revealed an internal program called Project Panama, overseen by Tom Turvey, previously head of partnerships for Google Books, which followed the same process of cutting spines, scanning pages, and discarding the originals.

Anthropic finalized a copyright settlement in July 2026, paying around $3,000 per work across 482,460 titles, the largest known payout of its kind in a US class-action case.

The same litigation also addressed more than seven million books Anthropic had obtained from piracy sources such as Library Genesis and Pirate Library Mirror.

US District Judge William Alsup ruled in June 2025 that turning a legally acquired physical book into a digital training file qualifies as transformative fair use under copyright law.

The decision originated from a single district court, was never taken up on appeal, and holds no binding authority beyond that jurisdiction. A May 2025 report from the US Copyright Office separately cautioned that commercial AI training does not automatically qualify as fair use in every case.

Current law does not require book buyers to reveal their identity or purpose. The first-sale doctrine allows owners to do what they want with a physical copy they have bought. Together with the Alsup ruling, this means operations like VGT3 can continue without breaking US copyright law. Authors whose secondhand books are bought and scanned are not notified and do not receive compensation.

AI Book-Scanning Faces Stricter Rules Outside the US

Legal standards are different outside the US. An intellectual property expert at the University of Oxford said UK law requires consent to compile a training dataset and to run the training process.

There is no fair use equivalent in the UK. In Germany, the publishers' and booksellers' association said scanning books for AI would break German copyright law, no matter how the books were obtained.

Similar legal challenges have targeted other AI developers. In July 2026, Hachette Book Group, Cengage Learning, Elsevier, and author Scott Turow brought a lawsuit against Google in a New York federal court concerning Gemini's training data, claiming the company stripped or modified copyright information from the works it incorporated.

Amazon's subsidiary Twitch drew separate attention after introducing an opt-out setting for generative AI training on August 12, 2026, once users discovered every account had already been enrolled automatically before that choice existed.

Amazon has not confirmed that the scanned books are used for AI training, referring only to purchases "through commercial channels" without addressing what happens to the physical books.

The full scope of the VGT3 operation, including how many titles have been processed and which specific AI systems the data supports, has not been detailed. The AirTag investigation traced one shipment to the facility, but Amazon has not publicly accounted for the destruction of the books or reconciled it with its earlier denial.

Thank you for being a Ghacks reader. The post appeared first on gHacks.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论