22 Sources
[1]
Hidden Airtag reveals Amazon is trashing rare books to train AI
For the past year or so, booksellers have suspected that AI firms are buying up huge lots of rare books, then destroying them after scanning them to train AI. But this was hard to prove until now, as 404 Media reports that an Airtag hidden in a rare book shows that at least one tech giant, in the race to advance its frontier models, is behind some of the bulk orders: Amazon. On Monday, 404 Media revealed that it had connected with a bookseller who agreed to plant an Airtag in a rare book that was part of a bulk order. That Airtag was then tracked to an Amazon AI training facility in Las Vegas that housed a team focused on tearing books from their spines and scanning pages, 404 Media reported. Apparently tone-deaf to the escalating backlash over destructive book scanning, a logo on the door of that team's warehouse, VGT3, showed a Tyrannosaurus rex preparing to devour a book, 404 Media documented. Amazon deflects Amazon declined to comment on 404 Media's findings, only providing Ars with the same statement it gave to 404 Media, which does not mention AI training specifically. "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use," Amazon's statement said. However, Amazon is developing what it considers frontier AI models, which require a massive amount of unique training data to stay competitive with leading firms like Google, OpenAI, or Anthropic. Right now, firms carefully guard their training data to avoid losing an edge. And training models on text from rare books that are difficult to find would seemingly offer an advantage for Amazon, especially since rivals like Anthropic and xAI have publicly stated that they are not training on rare or antique books. Further, it seems that Amazon needed a new source of original text. 404 Media flagged discussions in online forums where VGT3 workers suggested that earlier this year Amazon had run low on books to scan. The shortage was so alarming that they worried the warehouse might shut down if Amazon couldn't find more books to scan. At one point, the supply completely ran out, workers said. But the facility is still operational, 404 Media reported. And it's now confirmed that bulk orders of rare books delivered there are systematically destroyed by this crew. Many book lovers are horrified by destructive book-scanning (since there is an alternative), with one staunch critic, Michael Burry, even reportedly labeling the practice to be "evil incarnate." However, Amazon workers reported in forums that they consider the gig to be a "nice" opportunity for those drawn to a humdrum job with flexible hours. Debate rages over rare books On top of revealing that rare works that booksellers value are getting chewed up and swallowed by Amazon's AI machine, 404 Media suggested that its investigation helped firm up another bookseller theory about why AI firms might be ordering certain rare books and not others. After a reportedly historic year of sales, booksellers had suspected that AI firms were targeting books with ISBN numbers in order to ensure that the highest volume of unique works were present in training data sets. And 404 Media's review of Amazon workers' online discussions indicated that they were trained to scan barcodes or ISBNs before scanning books. That practice, 404 Media reported, "gives further credence" to booksellers' theory that "AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs." For booksellers, the money may be good, but the risk that their carefully sourced collections will be destined for destructive book scanning like Amazon's raises an ethical dilemma. They know how to assess a wide range of rare books to determine their value, and AI firms seem to be skipping that step in hunting low-cost, unique ISBNs to complete their checklists. Right now, the books that AI firms are apparently buying up aren't necessarily the kind of prized first editions of celebrated works that are typically valued quite highly. Instead, AI firms often target older books with lower monetary value, such as books that were never translated from a foreign language that's not widely used today or books that were never popular enough to be widely distributed. However, these works may still have "historical value, intellectual value, sentimental value" that AI firms overlooked, the bookseller who planted the Airtag told 404 Media. A rare book's value can be derived from "all sorts of things" that "the AI companies don't care about. They just want the content as a bunch of words strung together." Redditors weigh in On Reddit, some book fans debated whether it was that problematic that companies are destroying rare books to train AI, especially since, as the BBC reported, some of these books have been sitting on booksellers' shelves for decades gathering dust. "You were not going to buy that old paper book," one Redditor commented in response to a post lamenting that "an obscure book from 1700 is now a museum piece and may reveal day-to-day stuff that we didn't know." In that thread, the original poster said that the real problem was that tiny details and even major historical insights that can be gleaned from reviewing rare books will be lost to AI greed. AI firms will "never share the contents" of books they scan, the poster said, "as they don't want anyone else to be able to train their AI" on the same works. "This sounds like propaganda from the AI haters," another Redditor pushed back, but a subsequent commenter shared similar fears. Although people might clash with someone who argues that training AI on these works will make all the knowledge that a work contains more accessible online, the commenter suggested instead that AI models would only make available "warped, censored, and paywalled fragments of ideas" from forgotten works. "Even if you love the technology, you can admit that the concept of an AI literally eating books to become more powerful is pretty dystopian," someone else on the thread said. Some booksellers agree that not every text needs to be saved, the BBC reported. However, Scottish bookseller Derek Walker told the BBC that AI firms should be striving to distinguish between works that won't be missed much -- such as little-known academic texts that may technically be rare -- and lesser-known antique works that may be "the only known surviving example of an edition." "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long," Walker said. For AI firms guarding their training data and using services that mask their identities as buyers, there's likely little desire to discuss their bulk buying directly with booksellers. Allowing booksellers to weigh in on the works they plan to feed into their models would require a level of transparency that may seem riskier than a possible reputation hit if it's ever proven that a treasured first edition was destroyed in the name of advancing AI.
[2]
Amazon, once an online bookseller, is destroying rare books to train AI models
Amazon is buying tons of rare books, cutting off their spines, and scanning them for AI training, according to 404 Media, which placed a tracking device in a rare book that ultimately arrived at an Amazon facility in Las Vegas. The facility, known as VGT3, identifies itself with a symbol of a dinosaur holding a book in its claws. Amazon told 404 Media in a statement that it "purchases books through commercial channels to improve the products and services customers use." Companies like Amazon need unfathomably large amounts of text to train their LLMs, which have already ingested what they can from the internet (and, in Anthropic's case, illegally pirated books). Rare books, especially ones that are out of print or impossible to find on the internet, offer a new source of coveted training data. These texts are especially valuable since there's no chance that anything published before 2022 was written by an LLM. When LLMs train on AI-generated text, they risk "model collapse," which can occur when the quality of an LLM's outputs degrade after ingesting too much AI-generated text.
[3]
Here's a balm if the idea of destroying books to train AI breaks your heart
If you can truly appreciate an old book -- and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them -- then headlines about tech companies destroying books to train AI likely torture a tender part of your soul. It's indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that's the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality, long-form texts that can only be found in books. So book lovers fear it's likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever. What makes this destruction extra painful, though, is that it doesn't have to be this way. Google patented a non-destructive book-scanning technology in 2009 that AI firms could use to efficiently scan books -- if they were willing to slow down and invest in the process. It's not perfect, however; studies have found that the curve of the page can distort text, and pages can be missed. As Wired reported, glitches can happen when workers move too quickly, including disembodied hands obscuring pages. Overall, the trade-offs in cost and speed may not appeal to AI firms looking for the cheapest way to scan millions of titles, and Google's method may not be the best way to handle rare books anyway. The Internet Archive, which helps libraries preserve aging collections, has long understood that scanning old texts takes time and attention to limit handling, and that's why it's considered such a human job. For AI companies bent on finding shortcuts, the Internet Archive's method may feel almost alien. But if you're a book lover looking for a balm while parsing rumors of AI-driven book destruction, an old post from 2021 describes how the Internet Archive scans books to avoid damaging even the most precious rare books. In it, a book scanner who has been with the Archive since 2010, Eliza Zhang, offered a moment of zen by explaining what she likes so much about scanning books "the hard way." The Internet Archive did try going the automated route, the post said, even testing out "commercial book scanners that feature a vacuum-powered page-turning arm." But "it turns out those automated scanners didn't really work well for brittle books, rare volumes, and other special collections -- the kinds of material our library partners ask us to digitize," the Internet Archive said. "The job requires keen concentration," Zhang said, since the pages of "very old, fragile books" are "paper thin." In the post, Andrea Mills, who helps lead the Archive's book-scanning operations, explained that "clean, dry human hands are the best way to turn pages." To ensure each rare book only has to go through the scanning process once, Zhang takes her time. She carefully raises the scanner glass with a foot pedal each time she turns a page, then adjusts the cameras and ensures the page is readable. She also takes note of any fold-outs, setting a reminder to go back and scan the inserts so that bonus materials aren't lost while documenting the main pages of the work. The whole time, she understands that if a page is skipped or an image is too blurry, Internet Archive's proprietary software will stop the process and prompt her to scan it again. Practice makes perfect, though, and she reports a low error rate after more than a decade of finding a rhythm in the job. At the time, Zhang had scanned "more than 3 million pages, 14,000 foldouts, and 18,000 items (mostly books)," the post said, with the goal of guaranteeing "zero errors." Ars asked the Internet Archive for comment, but Chris Freeland, the director of library services, said the post detailing Zhang's work is "still the best description of our scanning process today." The post came after a video of Zhang's book scanning got 1.5 million views on what was then Twitter, accompanied by a caption that would resonate with book lovers appalled by AI-driven book destruction today. "At the Internet Archive, this is how we digitize a book," the tweet said. "We never destroy a book by cutting off its binding. Instead, we digitize it the hard way -- one page at a time." Rare booksellers flag suspicious bulk orders Ever since a lawsuit in summer 2025 outed Anthropic for destroying millions of print books to train its AI models, book lovers have moved to defend some of the most precious collections from what feels like AI firms' endless quest to feed all the books in the world into their large language models. The biggest fear for people who want to see books preserved through the training process is that AI firms will callously pulp rare books that can never be replaced. As The Atlantic reported last week, social media "raged" after two recent reports indicated that AI was already endangering rare books. First, a Telegraph report accused Silicon Valley of destroying millions of rare books and "shredding the originals," then 404 Media reported that a book-database company called ISBNdb was advertising that it could help AI firms source books in bulk. This backlash was expected, with ISBNdb reportedly warning its clients that the optics were bad. Anthropic started using a codename -- "Project Panama" -- for its destructive book-scanning in an effort to keep it hidden from the public. There is no indication that Anthropic ever destroyed rare books, and the company has denied doing so in statements. It's also true that some AI firms are helping preserve rare books. OpenAI and Microsoft, for example, are working with Harvard librarians on an initiative to train AI models on about 1 million public-domain books dating back to the 15th century. Others, including Elon Musk's xAI, have publicly said they won't destroy rare books to train AI. In a post on X, Musk said that he "asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning." But not everyone is convinced that the biggest AI firms are sincere in their promises to preserve rare books, even if copyright law allows them to destroy copies of books they've purchased. Critics on social media and Reddit have pointed out that Musk didn't say the xAI team would not destroy any books, and it's unclear how his company defines a "rare" book. Meanwhile, rare booksellers continue to flag strange orders, recognizing that most sincere buyers tend to seek out single books, not order large batches of widely varying selections. Just yesterday, an Irish bookstore called Kennys flagged a "bananas" order for 5,000 obscure titles, The Irish Times reported. Unlike bulk orders placed by universities or big libraries, this order seemed suspicious because the buyer made no attempt to haggle on the price -- which has now become an obvious "red flag" that AI is potentially involved, the Times reported. Although some booksellers said they could benefit in the short term from large orders, they fear that selling to AI firms indiscriminately destroying their collections "may be signing their own death warrant," the Times reported. Tomás Kenny of Kennys Bookshop told the Times that AI is "terrifying our industry," and he plans to refuse orders likely destined for AI training in an effort to protect authors' rights to their work and readers' access to hard-to-find titles. "Booksellers tend to be at the forefront of a lot of moral dilemmas and political and ethical issues. In most scenarios, it is not for me to decide," he said. "However, if they are taking a book and scanning it with a view to taking intellectual property, then Kennys do not want to be a part of that." It's possible that public backlash will pressure the biggest firms to adopt methods like the Internet Archive's, which focus as much on preserving rare books physically as they do digitally. But it's heartbreaking to imagine the alternative, where rare books become even rarer, especially for those who love the feeling of holding an old book in their hands. Zhang knows the allure of rare books, occasionally pausing her work to read parts of the titles she scans. When asked what she liked most about her job, Zhang responded with the same gusto that many book lovers feel while browsing the shelves in stores like Kennys. "Everything! I find everything interesting," Zhang said. "I don't feel it is boring. Every collection is important to me."
[4]
Secret tracking device placed in rare book ends up in Amazon processing facility -- destroying books to train AI models is 'all' the Vegas warehouse does
AI companies are reportedly buying up millions of books, some of which are rare, and using them to train AI models, often destroying the book in the process. We now have more specifics about how this process works, however, thanks to an investigation from 404 Media. The outlet placed a tracking device inside a rare book that was part of a large, 1,000-book order, and it ended up in a Las Vegas Amazon warehouse, named VGT3. Employees who spoke to 404 Media said they spend their time at the facility cutting the spines off of books and feeding them into a scanner. VGT3 is reportedly part of another, larger facility in Las Vegas called LAS8. Its purpose, according to the employees who spoke to 404 media, is solely to cut the spines off books and scan them. "All we do is scan books," one employee told the outlet. If there was any doubt about the purpose of the facility, a logo with VGT3 shows a dinosaur holding a book, looking like it's about to rip it up and take a bite. At the very least, the dinosaur isn't reading the book. You can see that image captured on 404's website above. The investigation started in July, when 404 asked a bookseller if the outlet could place an Apple AirTag in a book inside a 1,000-book shipment. A shipment this large is unusual, though they've reportedly become increasingly common in the AI era. Book sales are up, and many suspect that isn't due to a rise in readership, but for training frontier AI models. The order was submitted through Biblio, an online marketplace that claims to be the "largest independent book marketplace in the world." Often, these large orders contain a scattershot of different books. An independent bookseller in Ireland, for example, received an order for 5,000 books that they suspected was for training AI. "Some was high-quality non-fiction, like A History of Connemara, and the next thing might be The Eddie Hobbs Guide to your SSIA," Tomás Kenny of Kennys Bookshop said at the time. 404 Media says that employees at VGT3 are required to scan the ISBN (barcode) of each book, lending credibility to the theory that AI companies are working through a list of every book that has ever been published. Or, at least, every book with an ISBN. A bookseller told 404 Media that these large orders "never" include rare books that don't have an ISBN. Earlier this year, court filings revealed Anthropic's 'Project Panama,' which kicked off in 2024. Documents as part of those filings make the project's purpose clear: "Project Panama is our effort to destructively scan all the books in the world," said an internal planning document. A judge ruled that Anthropic was allowed to use books to train its AI models. However, it was fined $1.5 billion for keeping 7 million pirated books in a central library. In a 2025 lawsuit against Meta, it was revealed that it had pirated nearly 82TB of books. Scanning books en masse is nothing new. In 2005, Google spearheaded Google Books by scanning out-of-copyright titles from an on-campus library using specialized scanners. The books were then returned. Presumably, Google made use of some type of V-shaped scanner, which lays a book out naturally so as not to disturb the spine and binding. In the case of VGT3, 404 reports that employees cut off the spine of the book before scanning, presumably to feed the pages flat into an industrial scanner. The insatiable hunger for data in frontier AI models seems to be moving at a faster pace than Google's early book digitization efforts. It's hard to say why a facility like VGT3 operates in this way, though it likely comes down to cost. As major companies like Meta and Anthropic have been caught with a library of pirated books, they now need to buy them. And when purchasing thousands of books at a time, it's probably much cheaper to get secondhand copies from marketplaces like Biblio than it is to spend full price on digital versions of those books (if digital versions exist in the first place). Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[5]
Independent bookstores in Europe receive suspicious orders for thousands of books, prompting fears they'll be destroyed to train AI -- sellers believe acquisitions are part of AI tech companies' push to get more data
Some fear that these books will be scanned and destroyed by machines, with no human ever setting their eyes on their content. One independent book retailer in Galway, Ireland, received an online order for 5,000 books, which would be elating for many shop owners, especially at a time when people prefer shopping online from major platforms like Amazon or ditch physical books altogether and buy eBooks instead. However, according to The Irish Times, it's not the number of titles that went into the order that raised red flags, but the obscure books that went with the orders, prompting fears the books are being acquired to train AI, possibly resulting in their destruction. "Some was high-quality non-fiction, like A History of Connemara, and the next thing might be The Eddie Hobbs Guide to your SSIA," Tomás Kenny of Kennys Bookshop told the publication. In this example, the former is a history book that covers a region in western Ireland while the latter is about the Special Savings Incentive Account unique to the country. Berlin bookshops are also reportedly seeing similar orders that contain titles that wouldn't make sense for a person to purchase today, like "Pass Your Driving Test, 2018 Edition." It follows a report last month that AI companies are reportedly shredding millions of books after using them to train AI models, using destructive scanners to quickly digitize the books. Massive orders like these aren't that unique, with many private institutions looking to build libraries often ordering this huge number of books from independent bookstores. However, aside from weird titles included in the order, many large buyers connected with educational and other public institutions often negotiate a bulk price. When this is combined with the weird titles being included in the orders, the sellers cannot help but suspect that these orders are, in fact, made by AI companies looking to ingest more books to add to their training data. The biggest AI companies have been hit with multiple lawsuits regarding book piracy -- for example, court records revealed that Meta torrented 82TB of pirated books for AI training, while Nvidia is in hot water as a judge said that its NeMo Framework "have no other purpose" than to speed up infringement. More recently, Anthropic was hit with a $1.5-billion settlement for infringing the rights of authors and their publishers. Unfortunately, the penalty here is based on the startup's use of pirated books -- the court has ruled that the AI firms' use of these written works to train AI is considered fair use. This isn't applicable in Germany, though. "Under German law scanning books - regardless of the purpose - would not be permissible and constitute a clear violation of copyright law," said Thomas Koch, spokesperson of Germany's Publishers and Booksellers' Association. "That this is now happening with second-hand books is another highly troubling example of this practice." They've also surmised that the purchases are driven by AI bots that troll the internet for ISBNs (the unique numerical code assigned to each book title) that they don't have in their library yet. The orders are shipped to local addresses, but because some European nations have laws that prevent book scanning, the association thinks that these are just collection points for eventual bulk shipping to the U.S. It's not clear if the bookshops are fulfilling these orders or not. On the one hand, these bulk orders are a lifesaver for these stores and could help make surviving society's transition towards eBooks and readers much easier. On the other hand, AI is scaring them, and some fear it wouldn't just be the end of bookstores but could also lead to the degradation of human critical thinking abilities. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[6]
Secondhand book sales are booming. Is it because of AI?
Something mysterious has been happening in the world of secondhand books. Over the last few months independent booksellers have been observing a strange pattern of purchases which have seen scores of their novels being shipped to far-off warehouses. In a typical week, Stuart Manley, from Barter Books, in Northumberland, would sell two or three thousand books. But one recent single bulk order from a Canadian company equalled what he would expect to move in seven days. He says he has "never seen the like of this after 30 years in the second-hand book trade". Mr Manley is not alone. Booksellers from around the world have been reporting similarly unusual mega orders. They aren't sure what the final destination for their books is. But the suspicion is that they are not being bought for avid overseas readers but instead something else with a voracious appetite for new information: AI. The idea that the secondhand sales boom is being driven by the explosive growth of AI can be traced back to a court ruling in the US. In 2025, a judge ruled using books purchased in this way to train AI software was not a violation of US copyright law. The decision was the result of a lawsuit brought against AI firm Anthropic by three authors. In his ruling, Judge William Alsup said Anthropic's use of the authors' books was "exceedingly transformative" and therefore allowed under US law. When court documents were unsealed last month it also emerged books were being destroyed in the process of training Anthropic's AI chatbot, Claude. "Claude is trained on a mix of publicly available web data, commercially acquired datasets, and data we generate ourselves," a spokesperson said. They insisted sourcing books for training was a widely used approach across the AI industry. "None of our data acquisition programs buy and destroy rare or antiquarian books," they added. Nonetheless, the idea that books are being pulped is causing unease. David Tobin, runs Walden Books in north London, has also had unusual sales. "In some ways it's very nice to sell some of these titles which haven't been sold for many years, but it would be sad if they are ultimately destroyed," he says. The court documents relating to the Anthropic case revealed the project of ingesting old books was referred to in internal company communications as "Project Panama." The documents indicated the company's aim was to "destructively scan all the books in the world". Destructive scanning is the process of shipping books to locations where they can be digitised at an industrial scale. It includes removing a book's spine so all the pages can be scanned rapidly - and the remains recycled. "A lot of mystery surrounds Project Panama," Manley, from Barter Books, says. "The name is new to me, but the reality of the project is not, and has been the subject of much discussion on the bookseller forums." He does not know that his books are being bought for it or similar projects by other AI firms. But he says it's also difficult to account for the sales, which appear random with "no rhyme or reason." They have varied from obscure Latin texts to cowboy novels. Experts say the diverse subject matter also points to AI, as unusual and rare texts could provide fresh material to improve the training of large language models, the tech which underpins generative AI tools like chatbots. Professor Emily Hudson, intellectual property specialist at Oxford University, says copyright laws in the UK are different to those in the US. "The starting point in the UK is that all these acts of copying - creating the training library and doing the training - require the permission of the copyright owner," she says. For the booksellers, it poses a dilemma. They are uncomfortable with the idea of books being destroyed - even if they admit not every title needs to be saved. "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's. "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century. "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long." And Manley says recycling books is a good solution for many titles which the public no longer want on their shelves - especially when it comes with a bump in trade. "Some may have ethical concerns about where the books end up and if they're destroyed," he says. "But the world no longer needs five million copies of The Da Vinci Code. "I've had books advertised for 20 years on the web which haven't sold until now". Sign up for our Tech Decoded newsletter to follow the world's top tech stories and trends. Outside the UK? Sign up here.
[7]
How an AirTag planted by a reporter led to a secret Amazon site where old books are cut apart and scanned
Amazon is reportedly cutting the spines off old books and scanning the pages at a Las Vegas facility, presumably to train AI models on text that exists almost nowhere else. The company won't confirm that's the reason. It gave us the same statement it provided to 404 Media, which broke the story: it "purchases books through commercial channels to help develop and improve the products and services our customers use." But what really got my attention (and professional admiration) was 404 Media's means of discovering this was happening at all: reporter Emanuel Maiberg put an Apple AirTag in a rare book and watched where it went, like a biologist tracking an endangered salmon. Maiberg, a co-founder of 404 Media, has been digging into this topic for a while. He reported in July that booksellers were seeing a massive surge in bulk orders from buyers who didn't haggle. According to Maiberg's latest story, a seller informed him that they'd received an order for about 1,000 books through the marketplace Biblio, and agreed to slip an AirTag supplied by 404 Media into one of them. 404 Media granted the seller anonymity because the seller was worried the disclosure would hurt their business. The book flew out of a California airport to Milwaukee, sat for two weeks in a distribution warehouse outside Kenosha, Wis., then went west by truck, making an overnight stop in Grand Junction, Colo., before arriving at an Amazon warehouse in Las Vegas known as LAS8. As Maiberg recounts in the story, he was initially confused. LAS8 is largely a print-on-demand operation. It prints and ships books as customers order them, the opposite of destroying them. But the AirTag put the book at the north end of the building, which Amazon employees who posted on a workers' forum described as a separate operation with its own code: VGT3. Its logo, painted at the entrance, is a T. rex with an open book in its hands. Employees described a split operation: some workers cut books, others received them and scanned bar codes. Booksellers told Maiberg the bulk orders never included the very rarest books, the ones old enough to predate ISBNs, suggesting that buyers were working methodically through the serial numbers assigned to every published book. A history of reportorial tracking This technique of journalistic investigation has actually been around for a while. The Basel Action Network, a Seattle nonprofit, started planting GPS trackers inside old printers and monitors in 2014, dropping them at Goodwill locations and recyclers around the country to find out where America's electronic waste actually ends up. Nearly a third of the tracked devices were exported. Two old TVs dropped at Oregon recyclers traveled to a warehouse in south Seattle, then to the Port of Seattle, and to junkyards in Hong Kong. BAN's trackers led to federal conspiracy charges against Total Reclaim, the Seattle recycler that had been handling that Oregon e-waste. Over the years, others have adopted the same tactics. ABC News put trackers in plastic bags dropped at Walmart and Target recycling bins in 10 states, and Finland's public broadcaster hid them in used clothing to trace where donated fast fashion actually ends up. What's different now is the hardware. BAN worked with MIT and used cellular trackers that needed a data plan. Maiberg used a $29 AirTag that reports its position by pinging any nearby iPhone. What's going on at VGT3? Sure, it's possible that the slicing and scanning at Amazon's VGT3 could be for something other than training AI models. Amazon has digitized books for two decades, for example, going back to Search Inside the Book. But that program runs on files publishers submit themselves. It doesn't require buying used copies on the open market and cutting the spines off. Still, the circumstantial evidence is strong. The books are rare titles with almost no resale market, but that's exactly what makes them valuable as training data. The text was never digitized, and books printed before the AI boom are free of the machine-generated writing that degrades AI models trained on it. The bookseller who sold the tracked shipment put it plainly to Maiberg: the books have historical and sentimental value, and the AI companies destroying them don't care about that. Cutting the spine isn't just faster for scanning. It's the specific act that made Anthropic's version of this legal: in June 2025, a federal judge ruled that buying print books, stripping the bindings and scanning them was fair use, because the digital copy replaced an original that no longer existed. The same ruling went against Anthropic on books it had downloaded from pirate sites, which is the claim the company has since settled for $1.5 billion. As someone who has covered Amazon for a while, I should note that this could be some "peculiar" project that actually looks nothing like anything people are speculating about, which will only become clear "in the fullness of time," to use some of the favorite phrases inside a company known for being "willing to be misunderstood for long periods of time." But in the meantime, it's pretty fascinating to see everyday technology being used in a creative way to uncover something that otherwise might have never come to light.
[8]
Used Booksellers Are Racking Up Huge Sales. They're Not Thrilled About the Likely Reason
Another bookseller, David Tobin, who runs north London's Walden Books -- not to be mistaken for Waldenbooks -- said he was also in the midst of a business boom. "In some ways it's very nice to sell some of these titles which haven't been sold for many years," he told the BBC. But he added "it would be sad if they are ultimately destroyed." As 404 Media wrote last month, old books are becoming a premium product because anything published after the adoption of LLMs might have LLM outputs in it, which is bad. Up until about 2022, only humans wrote books. And it's human writing that these companies need in order to train future LLMs. Anthropic got in big trouble for using illegal copies of books as training data, eventually paying out $1.5 billion to settle a class action lawsuit. Recently unearthed court documents reveal that the company's new plan for ingesting the data in books, codenamed "Project Panama." is nice and legal, but a PR nightmare: Destructive scanning, which means taking a physical copy of a book and scanning it, but destroying the physical copy in the process. Scanning books to train AI falls under the fair use doctrine, a court has found. But Anthropic is obligated to destroy books as well, because only one copy of the book exists before the scan, and only one copy exists after, meaning no piracy is occurring. Perhaps you've seen a report recently about this practice featuring this grisly video of books having their spines sliced off. That's part of destructive scanning. But in case you think all of this is a big overreaction, or some kind of naive romanticization of physical books, the booksellers in the BBC's report do note the difference between shredding one of five million copies of the Da Vinci Code, and destroying something of actual value. As the owner of an Edinburgh bookshop called McNaughtan's, Derek Walker, points out to the BBC, there are even some relatively rare books the destruction of which wouldn't be a huge catastrophe. For instance, Walker notes, "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market -- but it is perhaps not such a great loss if one copy is destroyed." However, Walker says he has been known to sell "the only known surviving example of an edition from the 18th century." Rare technical books that have never even been scanned are a sort of holy grail for competing LLMs, since they remain exactly what they are in movies: dusty old artifacts that contain unique, secret knowledge that's been lost to history. Frontier AI companies seeking to create so-called "superintelligence" will no doubt want this knowledge. And if AI companies really do find, absorb, and destroy such knowledge in pursuit of training data, then I guess by definition we'll never know it happened.
[9]
Hidden AirTag appears to confirm suspicions about illicit AI book scanning
Suspicious bulk orders of used books have led many booksellers to believe they are being purchased on behalf of AI companies so that they can be scanned and used as training material without author consent or payment. An AirTag hidden inside a rare book sold as part of a bulk order appears to confirm the theory after it ended up at an Amazon AI training facility ... Booksellers specializing in the sale of used books normally receive orders for individual titles or perhaps a small number of related books. For quite some time, however, a number of them have reported receiving massive bulk orders for used books, with no apparent connection between the titles. The suspicion is that they're being purchased by AI companies who strip the bindings and use high-speed scanning equipment to ingest the text and feed it into AI systems for use as training material. Doing so without the consent of the authors and without offering any payment for the use of their work is at the very least unethical, but there hasn't yet been any concrete proof that this is happening. 404 Media's investigation appears to change this. We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books. The report is paywalled, but TNW reports on the findings, noting that the tracking device used was an AirTag, with staff at the facility confirming what goes on there. Employees told 404 Media what the work involves. Large shipments of printed books arrive, staff slice the bindings off so they can scan the pages faster, and the book does not survive it. The logo on the door is a dinosaur holding a book and baring its teeth. Used books are especially valuable because anything published after 2022 may potentially include text written or edited by AI systems; effectively, they could be training AI systems on AI output. While publishing contracts invariably explicitly prohibit authors from using AI-generated content, there are likely to be some who ignore this. There are also no effective controls when it comes to self-published books. Part of the problem is that intellectual property laws have not yet caught up with the practice. While this would likely be illegal in the UK, a 2025 ruling in the US somehow declared it lawful as a "transformative" use of the copyrighted material.
[10]
We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility
We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books. Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process. A 404 Media investigation was able to reveal Amazon's book buying operation, which hasn't been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination. That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands. "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use," an Amazon spokesperson told me in a statement. The world's AI companies are constantly looking for, and spending extreme resources to locate, more material to train their AI models. With books, that sometimes means destroying them in the process, something that large parts of the public have spoken up against, and which we can now confirm Amazon is doing. In July, I published a story about booksellers who reported a historical spike in sales starting in the past year. They suspected this spike in sales was due to AI companies acquiring any books they can in search of new training data. Printed books are valuable as training data because a lot of the text they contain is not readily available on the internet, which AI companies have already scraped. The data is also conveniently organized and, if the book was printed before 2022, is guaranteed to be free of AI-generated text, which can make any AI model that is trained on it worse via a recursive process called "model collapse." These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all. But booksellers couldn't say for certain who was behind the large purchases because the marketplaces where they sell their books keep the buyers anonymous. When an order comes in, a bookseller ships the sold books to a warehouse operated by the marketplaces, where books are sorted and then sent to the buyer. In July, one bookseller told me they received a very large order of around 1,000 books on Biblio, one of these marketplaces. The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order. 404 Media granted the bookseller anonymity because they worried sharing this information would harm their business. Biblio did not respond to a request for comment. The bookseller sent the shipment to Biblio, and it arrived at a California airport. It then traveled by plane to Milwaukee International Airport in Wisconsin. Later that day, the book traveled to a warehouse belonging to a specialized shipping and distribution company called Trifinity, right outside of Kenosha Regional Airport, about 30 miles south. The book remained there for about two weeks, at which point it began traveling west by what appeared to be a truck. I could see the book travel via the highway and spend a night over at a trucking travel center around Grand Junction, Colorado. The next day, the book arrived at its final destination, an Amazon warehouse called LAS8 in Las Vegas. LAS8 is one of several large Amazon warehouses in the area, each with a different specialty. LAS7, right across the street, for example, is a fulfilment center, while LAS8 appears to mostly operate as one of Amazon's "print on demand" operations, which will print and ship books to Amazon shoppers as they are buying them. Initially, I was confused about why the book would arrive there, but Amazon employees who work at this location and who discuss working conditions there with other Amazon employees online explain that the the north end of the LAS8 warehouse, where I saw the book arrived, housed a different Amazon operation with a different code: VGT3. Many Amazon warehouses have unique symbols to represent that specific site. VGT3's symbol, painted on the entrance to the site and inside, is of a Tyrannosaurus rex, the massive carnivorous dinosaur, with its mouth open, holding an open book. "I work at VGT3 here in Vegas, and all we do is scan books," one Amazon employee wrote on a forum for Amazon workers. "Some are assigned to cut books, and others go to receive where they get books and scan the bar codes. We didn't have rates, but now we do, but it's not stressful." "Working at VGT3 is nice all we do is scan books," another Amazon employee wrote. "It's so cool." I saw VGT3 employees online talk about how working in this operation is a good but boring job because workers do the same, easy, repetitive tasks all day. I also saw some discussion indicating that VGT3 jobs are desirable for this reason, and because the warehouse sometimes offered night shifts. Earlier this year, some employees expressed concerns that Amazon would shut down the site because they had worked through the books they had and weren't getting enough shipments of new books, though the site is still operational today. That employees are scanning the barcodes or ISBNs on books -- a unique serial number given to every published book -- before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs. Elsewhere online in 2024, booksellers who sell their books on Amazon said they received a spike in orders to be shipped directly to the VGT3 location for a customer named "Amazon FC." This was so unusual given its size they suspected the orders were some kind of scam. An Amazon representative chimed in to say they checked and that the orders were legitimate. Like other major tech companies, Amazon is developing its own large language models (LLMs). Amazon considers its family of models branded Nova to be "frontier" models, meaning it considers them to be competitive with other cutting edge LLMs from Google, OpenAI, and Anthropic. These models require massive amounts of training data in the form of human-written text. Internally, Amazon also used an AI coding agent called Kiro to develop its own software. We first learned that AI companies wanted to scan millions of books for training data because of a lawsuit from book authors against Anthropic, which revealed Anthropic's "Project Panama." The goal of the project was to acquire books from commercial bookselling marketplaces, cut the spines off the books, and scan them. It's possible to scan books without destroying them, but cutting the spine makes it cheaper and faster. Additionally, the judge in the lawsuit ruled it was fair use and not a copyright violation for Anthropic to scan a book for training data in part because it destroyed the original, printed copy. Essentially, it's the customer's right to take physical media and store it digitally, and destroying the original copy means that copy isn't duplicated and resold, and isn't cutting into the publisher's business. We're not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that's because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak. As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn't mean they're not valuable. "There are different types of value," the bookseller said. "There's monetary value, obviously, but there are a lot of other types of value. There's historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don't care about. They just want the content as a bunch of words strung together." Amazon provided its statement above but did not respond to questions about why it cut the books it scanned, how many facilities like this it has across the world, and their response to the backlash from people upset that AI companies are destroying books.
[11]
Secondhand booksellers in UK and Ireland suspect AI firms behind 'strange' bulk orders
Development comes after Anthropic was found to have spent millions on books to scan for 'data acquisition' Secondhand booksellers in the UK and Ireland are reporting a flurry of bulk orders from mystery buyers, amid speculation AI companies are acquiring the tomes for their data. Bookshops contacted by the Guardian say they have been receiving orders from buyers in the US, Canada, continental Europe and the UK. The increase has been replicated around the world, with booksellers in the US, Australia and Europe also reporting unusual orders. Stuart Manley, the co-owner of Barter Books in Alnwick, Northumberland, said the orders, which began arriving three months ago, were unusual because they did not follow the typical pattern of being grouped into themes such as sport or motoring. "It's definitely a strange combination of books," said Manley. "Normally a book order is on a theme but this is all over the place." A recent order included requests for an Estonian translation of John le Carré's The Mission Song, a copy of Anne Brontë's Agnes Grey from a specific imprint and the October 1983 edition of Warship, a monthly magazine. Manley said he had sold "hundreds" of random books to three buyers for a total of about £4,000. Jim Shaughnessy, the owner of MW Books in Claregalway, Ireland, said his business had been receiving large, varied orders since May. "They are sporadic, presumably machine-driven and have little in terms of thematic currency; books on agricultural implements in 18th century Africa through to biographies of 1950s racing car drivers," he said. "Personally, I don't see we have the right to dictate what any customer does with their purchases but it is more than curious to see what's being selected." David Gower-Spence, the owner of BookLovers of Bath, said he had received a rush of orders between May and July for about 200 books from buyers who have also been cited by other sellers contacted by the Guardian for this article. One, the Canada-based Zoom Books, which has featured in other media reports, describes itself as "North America's leading book recycling company". Zoom has been contacted for comment by the Guardian. "It was overwhelming for a while and I'm relatively well organised," said Gower-Spence. One UK-based secondhand bookseller, speaking on condition of anonymity, said some random orders could come from different buyers but be delivered to the same address. "The aliases behind these orders are very opaque even to experienced booksellers," the source said. One delivery postcode seen by the Guardian and used by different buyers points to a group of freight warehouses near Heathrow airport. The bookseller, who has received orders for 6,000 books from similar buyers since January worth thousands of pounds, said the buyers were paying "top whack" and were not seeking a discount despite requesting titles in bulk. "It's disrupting the secondhand bookselling business because this isn't typical for the trade," the bookseller said. "It appears to be either a human or a machine dipping into the sandbox of secondhand books and what they are buying makes no sense in terms of the wider trends in the market." The Guardian reported this month that Australian booksellers had also received a burst of orders, with one saying the sudden scale of demand was "driving me nuts". The spate of orders has fuelled speculation among booksellers that AI companies are acquiring the books to scan their contents before pulping them. Manley said: "My theory is that it's masses and masses of AI money being squandered." The Washington Post revealed this year that Anthropic, a leading AI company and developer of the Claude chatbot, had spent tens of millions of dollars acquiring books and slicing off their spines so their contents could be scanned before the books were sent for recycling. Anthropic said: "Claude is trained on a mix of publicly available web data, commercially acquired datasets, and data we generate ourselves. Sourcing books is a widely used approach for training large language models across the AI industry. None of our data acquisition programs buy and destroy rare or antiquarian books." Last month, the tech news site 404 Media reported that a books database was telling booksellers "the world's best AI training data is sitting on a shelf", adding that books published before 2022 were ideal material for AI companies because their text was unpolluted by material that could have been generated by chatbots. The database in question, ISBNdb, said the report was based on a webpage that was a "test of market interest" and had since been taken down. AI tools such as chatbots are "trained" on vast amounts of data to help them generate responses, a requirement that has set tech companies searching for fresh sources of material to build new models. Secondhand books, some of them published decades before ChatGPT exploded on to the scene and not available online, could be a source of fresh data. One post on an online forum for booksellers, seen by the Guardian, said: "Shame these are all destined for AI training, and will therefore get the covers stripped and ultimately end up being pulped." Other posts expressed puzzlement at the erratic orders, while some questioned whether they could be intended for AI training when requests included books available for nothing in the public online domain such as Gulliver's Travels, published in 1726. Orders are typically placed through book marketplaces including Biblio and the Amazon-owned AbeBooks. Biblio and AbeBooks have been contacted by the Guardian for comment.
[12]
AppleInsider.com
An AirTag hidden in a rare book followed its journey to a Las Vegas warehouse, where it was likely destroyed in an effort to supply fresh data to hungry LLMs. I don't know if you've noticed, but everything is becoming increasingly AI-generated. Whether it's news, music, or social media, AI has seemingly stuck all of its nasty little fingers into any pie it can find. Experts have long warned that eventually we'd get to the point where AI was training itself on data that AI had created, and we're pretty sure that point is behind us. And that's not actually particularly useful, it turns out. It's harder to pass the Turing test when everything sounds exactly like it fell out of the back of a ChatGPT prompt. And that's why AI companies don't really want the AI version of the human centipede. They want to train AI off of human-produced content. Which, as a human who produces human-produced content, you'd assume I'd be relieved to hear that I at least have a spot in the pipeline. Unfortunately, those same companies are telling other companies that they can get rid of the human in favor of their hyper-efficient, human-like product. So, in the human centipede analogy, where does that leave me, specifically? Am I at the front, or did they try to find a space for me somewhere in the middle? Just kidding, they really don't want humans involved in the process at all. That's why they've gone straight to destroying books. In a warehouse in northeast Las Vegas... Rare booksellers have seen historical spikes in sales in the past year. So 404 Media reached out to one to see if they could track who was buying these books. The seller agreed to hide an AirTag in a rare book that was part of a larger 1,000-book order. The book traveled from California to Milwaukee, then from Milwaukee to Grand Junction, Colorado. After a few weeks, it arrived at its final destination: Las Vegas, Nevada. Specifically, it wound up at an Amazon warehouse known as "LAS8." Typically, LAS8 operates as a print-on-demand book-selling operation. But 404 learned that the north end of the warehouse doesn't create books. Instead, it destroys them. The location in question is known as VGT3. It is there that books are received, sorted, and then processed. Processing requires the spines to be cut off and the pages to be fed through scanners that convert the scans into data. The data is not available anywhere, they are not being scanned into PDFs to read. It is used exclusively to train AI. The remnants of the books are discarded. The contents are never meant to be viewed by humans, only by AI. If this were a movie, I would have called the writer out for being a hack. But unfortunately, this is just the reality we're living in now. Welcome to the future, dear reader. Is it as exciting as you'd hoped it would be? 21st Century book burning Data is finite, unfortunately. Especially as far as large-language models are concerned. There are lots of reasons for this, but two major ones that seemingly have AI companies and researchers on edge. The first is model collapse. When AI trains on its own data repeatedly, eventually the data condenses down so much that it becomes useless. When this happens, AI will produce generic, unusable information. Effectively, it's the same thing that happens when you photocopy a photocopy of a photocopy. The second is that AIs hallucinate. It hallucinates a lot, actually. So if one model trains off of another model's hallucination, the hallucination is treated as the baseline. Now imagine that happening over and over and over again. Eventually you arrive at the same point as model collapse, but this time with the added danger of AI citing hallucinations as fact before they eventually become unusable. All roads lead to the same location. And it's not a location anyone is really interested in going to, or alternatively, staying there, like my editor thinks we are now. Nearly all of the human-produced training data available on the internet has already been slurped up by LLMs. And, as mentioned before, new data is overwhelmingly being produced by AI. So AI companies need to figure out how to circumvent the digital ouroboros effect. But if humans can't create fresh data fast enough, where do you find new information to feed into the machine? Well, as we learned above, it's apparently by destroying books. Books are a goldmine where AI training data is concerned. Physical books are even better. Much of the information contained within hasn't been digitized, which means it's raw, virgin data. And the data doesn't really need to be "factual"; it just needs to be unique, unsullied by AI. It just needs to train the math behind the algorithm. This is all is why rare booksellers are seeing an increase in sales. And judges have ruled that it's fine for a company to do so, provided they're not selling the digitized copy afterward. Amazon isn't the first company to do this. Anthropic has also done this. In fact, I'd be willing to bet that most big commercial AI outfits have been doing this for a while now. Eventually, the companies will run out of books. And likely, some exceedingly rare books will be entirely lost in the process. At some point, the stone will cease to yield blood, but I severely doubt that the AI powers-that-be will ever stop squeezing.
[13]
Amazon Caught Destroying Rare Books to Train AI
Can't-miss innovations from the bleeding edge of science and tech Amazon has been caught scanning and destroying a large shipment of rare books so it can train its AI models. In an investigation by 404 Media, a bookseller suspected that an order they received for 1,000 books on Biblio, a marketplace where buyers can remain anonymous, was for an AI company. So to get to the bottom of the mystery, they agreed to place an Apple AirTag between the pages of one of the volumes. They tracked the shipment to a site at a massive Amazon warehouse in Nevada, where workers say their job is to strip the spines off of books they receive so they can quickly scan their pages, destroying them in the process. "I work at VGT3 here in Vegas, and all we do is scan books," an Amazon employee wrote on a forum for Amazon workers, as quoted by 404. "Some are assigned to cut books, and others go to receive where they get books and scan the barcodes. We didn't have rates, but now we do, but it's not stressful." Despite Amazon clearly trying to keep this practice under wraps, the logo for the part of the warehouse the book scanners work in, called VGT3, is so on the nose that it borders on farce: a T-Rex holding a book in its hand that it's about to eat -- which, we have to say, kind of looks AI-generated. The company only had this to say in a statement to 404: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use." The public only became aware of this AI industry practice in a lawsuit against Anthropic, which revealed that the Claude chatbot maker was using industrial equipment to cut the pages out of millions of books and scan them, before they were disposed of. More alarming was the upshot of that lawsuit: a judge ruled that because Anthropic was turning the physical texts into digital ones and destroying the original copies, it was "transformative" and therefore didn't violate copyright law. Since then, book sellers have often suspected that their wares are being used, and then destroyed, to train AI models. In previous 404 reporting, one said that some of the giveaways were the size of the orders, the seemingly random selection of books, and that the books all had ISBNs. But they couldn't prove this hunch because the orders are anonymous. As 404 notes, that the workers at the Amazon site are apparently scanning the ISBNs of all the books they process gives credence to a spooky theory by booksellers: that AI companies are trying to scan every printed book in the world by going through a list of all their serial numbers. The bookseller 404 worked with sold books that are rare, meaning that there are few copies left in circulation. That can be for a myriad of reasons -- they're not necessarily first editions of old classics -- but they're still valuable. "There are different types of value," the bookseller told 404. "There's monetary value, obviously, but there are a lot of other types of value. There's historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don't care about. They just want the content as a bunch of words strung together."
[14]
AI companies may be gobbling up old books, and I really hope they aren't destroying them
Booksellers are seeing mysterious bulk orders for old books, and some suspect AI companies Secondhand booksellers have noticed a curious pattern over the past few months. Large orders are coming in for old books that often share no obvious subject, author, or genre, leaving sellers wondering who wants them and why. According to The Guardian, booksellers in the U.K. and Ireland have received bulk orders covering everything from agricultural texts to racing biographies. Some buyers reportedly use opaque aliases, send orders to the same freight warehouses, and pay full price without negotiating bulk discounts. Similar activity has also surfaced in the U.S., Australia, and Europe. Recommended Videos There is no proof that AI companies are behind most of these purchases. But take a look at what has been happening elsewhere, and the theory doesn't seem quite so far-fetched. AI companies really do want old books 404 Media previously found an ISBNdb landing page advertising "Printed Books Sourcing for Your AI LLMs Dataset Needs." ISBNdb removed the page on July 28, saying it had been part of an effort to explore demand and that the company had decided to move away from the idea. Two days later, ISBNdb issued a more detailed clarification. It said it has never purchased, scanned, or sold a book for AI training, has never trained AI models, and that the proposed service was never actually launched. There are other clues. The Atlantic reports that Singapore-based AI data company 2077AI has been linked to requests for thousands of specialized books, although that still doesn't connect it to the wider buying spree. One theory is that books published before the generative AI boom offer something increasingly valuable to model developers: large amounts of human-written material with little risk of containing AI-generated text. The removed ISBNdb page had also pitched older printed books as valuable training material for this reason. Hopefully these books aren't being destroyed Unfortunately, we already know destructive scanning happens. Court records previously revealed that Anthropic spent millions acquiring physical books, cutting off their spines, scanning the contents, and then sending the books for recycling while building its training library. Anthropic says its acquisition programs do not buy and destroy rare or antiquarian books. As someone who spends a fair amount of time hunting for old and out-of-print books, though, the possibility bothers me. Some books only get limited print runs, and once those copies disappear, finding another one can be difficult. The idea that one of those remaining copies could be destroyed simply to extract its text for AI training doesn't sit well with me. If AI companies really are behind these mystery orders, I just hope someone is checking what they're cutting apart first.
[15]
Rare Books Traced to Amazon AI Training Facility to Be Scanned and Destroyed
Amazon confirmed it "purchases books through commercial channels" and said the operation hadn't been reported before. AI companies are scanning truckloads of books, sometimes rare ones, and often destroying them in the process to train their models -- and now we have some idea of which companies and how they're doing it. Independent tech media publication 404 Media placed a tracking device inside a shipment of rare books and watched it travel to an Amazon facility in Las Vegas where the company scans and destroys printed books to train AI. The investigation, published Monday, is the first to publicly pinpoint where bulk book purchases tied to AI training end up. The final stop was Amazon's VGT3 team. Employees told the outlet all they do is receive massive shipments of printed books, then cut the bindings off so the pages feed a scanner more quickly. The book is destroyed in the process. The team's extremely fitting logo is a dinosaur clutching a book in its claws. Why printed books are suddenly in demand Books printed before 2022 carry text that isn't mostly readily available online and, crucially, isn't machine-written. Training a model on its own kind of output risks "model collapse," where quality degrades with each recursive loop. So, pre-2022 paper is clean fuel for AI training. Amazon isn't the only one. A cottage industry has formed to supply AI labs with physical books that get stripped, scanned, and discarded -- the "Fahrenheit 451" scene of texts destroyed to feed machines that Anthropic's internal "Project Panama" already ran at scale, digitizing millions of books through destructive scanning. That scramble for clean text has already landed in court. A federal judge ruled that training AI on legally purchased books can qualify as fair use, a partial win for Anthropic that Meta and OpenAI also claimed. Pirated copies are a different matter: Anthropic agreed to a $1.5 billion settlement over scanned pirated titles, and Salesforce now faces a class action over alleged book piracy. The Amazon operation stayed quiet for a structural reason. Booksellers noticed a spike in bulk orders over the past year and suspected AI companies were behind them -- the buyers weren't price-sensitive and picked titles seemingly at random. Their concerns were finally confirmed by an AirTag revealing the destination. Amazon said it "purchases books through commercial channels to help develop and improve the products and services our customers use."
[16]
AI companies are buying used books by the thousands. Some may be destroyed for training
Bulk orders are giving used bookstores a surprising sales boost as AI companies seek legally acquired books to scan for model training. Not too many years ago, before the dawn of Kindles, used bookstores were a refuge for budget-minded, voracious readers. As the industry evolved, many suffered. Now they're seeing a resurgence in sales, but the customers increasingly aren't bookworms. They're AI companies. Barter Books, for instance, typically sells 2,000 to 3,000 books per week, which isn't bad for a used bookstore in Northumberland, the least densely populated county in England. Recently, however, the owner got a bulk order from a company in Canada for that many books in a single day. Several other bookshops in the area are seeing similar demand. And while they can't definitively say what the books are being used for, a recent court ruling offers a clue. The fate of those books may also be a dire one. AI companies have been in hot water with the literary world in recent months. Anthropic recently reached one of the largest copyright infringement settlements in history, agreeing to pay $1.5 billion to more than 300,000 writers who filed a complaint against the AI giant two years ago. OpenAI, meanwhile, is still facing several copyright infringement lawsuits from a variety of sources, including The New York Times, Encyclopedia Britannica, comedian Sarah Silverman, and several nonfiction authors, who accuse the company of copying their books to train its large language model and saying OpenAI has "enjoyed enormous financial gain from their exploitation of copyrighted material." The Anthropic suit, however, may have provided a path forward for AI companies. The court ruled that it was not illegal for Anthropic to train its AI on copyrighted works, as long as it paid for the books it used. A separate judge ruled similarly last year for Meta. Used books offer AI companies a cheaper way to acquire source material. Some publishers might also refuse to supply copies directly to an AI company if they know what it plans to do with them. The Anthropic suit highlighted a program at the company called "Project Panama," which Anthropic described as its "effort to destructively scan all the books in the world." A memo highlighted as an exhibit in the trial explained that the company used a codename for the training method "because we don't want it to be known that we are working on this."
[17]
Amazon Caught Destroying 1,000 Rare Books to Feed Its AI Models
If you've ever sold an old or rare book online, there's a chance it didn't end up on someone's shelf. It may have ended up in a Las Vegas warehouse. There, its spine gets sliced off and its pages get fed through a scanner so Amazon can train its AI. That's what a new investigation from 404 Media found. 404 Media worked with a bookseller who had gotten a strange order for 1,000 books through Biblio, an online marketplace where buyers can stay anonymous. Together they slipped an Apple AirTag between the pages of one book and watched where it went. Weeks later, it landed at an Amazon facility known as VGT3. An AirTag Blew the Whistle Workers at VGT3 say their entire job is scanning books. One employee wrote on an Amazon workers' forum that some staff cut the spines off books. Others scan the barcodes so the pages can run through machines and turn into digital text. The physical book doesn't survive the process. The part of the warehouse where this happens marks itself with a logo of a dinosaur holding a book in its claws. It's on the nose for what's actually going on inside. Amazon gave 404 Media a short statement, saying it buys books through normal commercial channels to help improve its products. The company didn't explain why the originals get destroyed instead of kept, resold, or donated. This isn't the first time an AI company has done this. A lawsuit last year revealed that Anthropic, the company behind the Claude chatbot, ran a similar operation, cutting apart millions of printed books to scan them. A judge later ruled that practice didn't break copyright law. Turning a physical book into digital text, and destroying the original in the process, counted as legally "transformative." Why Rare Books Are the Target Rare and out-of-print books are especially valuable to AI companies because their text usually isn't available anywhere online. That makes it fresh training material. It was also written before AI chatbots existed, so it hasn't been diluted by AI-generated text. One bookseller who spoke with 404 Media put it plainly: books carry historical and sentimental value that AI companies simply don't factor in. They just want the words. There's no indication Amazon is doing anything illegal here. Courts have so far sided with AI companies on this kind of book scanning. But if you're planning to sell or donate an old book through a third-party marketplace, it's worth knowing it might never reach another reader.
[18]
Amazon Is Buying Up Older Books and Destroying Them After Scanning. The Orders Go Back to 2024
Amazon is buying physical books, cutting off their bindings, scanning the pages for AI training, and destroying the originals, according to a new investigation that traced a roughly 1,000-book order to one of the company's Las Vegas warehouses. 404 Media worked with a bookseller who received the order through Biblio, an online marketplace for used and rare books. The publication hid an AirTag inside one book and followed it to Amazon's LAS8 warehouse in North Las Vegas, where an operation called VGT3 scans books for AI training. "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use," an Amazon spokesperson told Inc. The company did not answer questions about how many books it has purchased, when the program began, which products use the scans, or whether it screens out rare or scarce copies before destroying them. Unusual book orders go back to 2024 Inc. found signs that unusual book shipments were reaching the same North Las Vegas warehouse nearly two years ago. In a September 2024 discussion on Amazon's Seller Forums, booksellers reported sudden bursts of orders destined for LAS8. One seller said he received 68 orders from a single buyer. Another said 48 orders arrived around 4 a.m. and consisted entirely of University Press titles. A third said the same account bought books from him "every day." An Amazon forum representative said the company's team was investigating and later told the sellers he had "confirmed the validity of these orders."
[19]
Report: Amazon Is Buying Rare Books in Bulk and Destroying Them to Train AI -- A Hidden AirTag Proved It
Opinions expressed by Entrepreneur contributors are their own. A rare bookseller suspected AI companies were quietly buying up rare books to train their models, then destroying them. To find out for sure, they planted an AirTag inside a shipment and tracked where it went. It led to a warehouse in Las Vegas called VGT3, part of Amazon, according to a 404 Media investigation. Workers there cut the spines off incoming books to scan pages faster, destroying the original in the process. The team's logo shows a Tyrannosaurus rex devouring a book. Amazon wouldn't confirm the books are being used for AI training. Its statement only said the company "purchases books through commercial channels to help develop and improve the products and services our customers use." But Amazon is building competitive frontier AI models that need massive, unique training data, and workers reportedly said the facility nearly shut down earlier this year after running out of books to scan. 404 Media also found evidence supporting a theory that AI firms are systematically working through lists of ISBNs to make sure every unique book gets scanned, Ars Technica reported. Workers said they're trained to check barcodes before scanning. Not every rare book is worth a fortune, but booksellers say many still carry real historical or sentimental value, the kind of value AI companies "don't care about," one told 404 Media. "They just want the content as a bunch of words strung together."
[20]
Amazon's massacre of rare books is a betrayal of the brand's origins
Fahrenheit 451 isn't a great look for the "Earth's biggest bookstore". Today, Amazon is one of the world's biggest tech companies, with vast cloud computing and AI operations as well as its mammoth global online retail platforms. But its brand still has a close connection with where it all began as an online bookseller. As well as still selling books directly, Amazon's marketplace is used by many third-party booksellers, and the company remains the leader in e-readers with its Kindle line launched back in 2007. That's why it's so harmful for the self-proclaimed "Earth's biggest bookstore" to have, apparently, just been caught destroying old and rare books to train AI - and even celebrating the fact with a fun cartoon logo design. Amazon built its business on selling books, first coming online as a bookstore back in 1994. Now, according to an investigation by 404 Media, it's tearing old and rare books apart so it can feed them its AI models. Journalists hid an Apple Airtag inside a book in a bulk shipment of 1,000 that were ordered via Biblio, a marketplace where buyers can remain anonymous, after the seller said they suspected that the order was made by an AI company. The tag was then traced to Amazon's LAS8 warehouse in Las Vegas, reportedly the location of a scanning operation called VGT3. 404 reports cites workers at the facility, who told it that their jobs involved scanning books to provide training data for Amazon's AI. To speed up the process, workers would cut the bindings off books. In online forums, VGT3 employees have said that "all we do is scan books," and that "some are assigned to cut books, and others go to receive where they get books and scan the bar codes." That practice would be damaging enough for Amazon's brand image, but to make it ever worse, the the warehouse team reportedly has a logo that depicts a Tyrannosaurus rex devouring a book - a tone-deaf image that reinforces the irony of a bookseller now consuming books for profit. Amazon built its brand on making books more universally accessible, yet it now appears to be destroying rare and irreplaceable volumes to feed AI models. It reminds me of the backlash when Apple released (and later withdrew) an iPad Pro advert that showed a machine crushing physical objects like musical instruments, books and camera lenses, only in Amazon's case it seems it's on an industrial scale as a core part of its AI developments. The books bought via Biblio were not necessarily highly valuable in monetary terms, but they were defined as rare. "There are a lot of other types of value," the bookseller is quoted as saying. "There's historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don't care about. They just want the content as a bunch of words strung together." Rare books not only offer "clean" pre-2022 text not polluted by machine-written content, but could also contain data that rivals don't have, making them potentially valuable training material. Courts have ruled that training AI on legally purchased books can qualify as fair use, but that doesn't change the fact that the optics of shredding rare works are damaging even if technically legal. Alternative paths ignored: Non-destructive scanning or licensing digital archives could achieve the same goal without destroying cultural artifacts. Rivals also destroy books to train AI, but Anthropic has claimed that it avoids using rare or antique books, which makes Amazon's approach look particularly reckless. The revelation of such cultural vandalism raises more doubts about AI developers' claims that the technology is good for human creativity. It's also likely to alienate authors, bibliophiles, and customers, clashing with Amazon's identity as a platform for readers. "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use," an Amazon spokesperson said in a statement. Various booksellers have suspected that AI companies are behind some bulk purchases due to the size of orders and seeming random selection of books. 404 suggests that Amazon workers' scanning the ISBNs of the books they process backs up a theory that it aims to scan every published book on record by logging every serial number.
[21]
List of AI companies turning to old and rare books for AI training - MEDIANAMA
Amazon is buying old, obscure and rare books and sending them for destructive scanning at a facility used for AI-related work, according to an investigation by 404 Media. Booksellers and preservationists have raised concerns that destroying these physical copies could make certain books irreplaceable. 404 Media said it placed a tracking device inside a rare book. It then followed the shipment to an Amazon facility called VGT3 in Las Vegas. Amazon confirmed that it purchases books. However, it did not explain how many books it buys or whether it preserves rare titles. "Amazon purchases books through commercial channels to improve the products and services customers use," the company told 404 Media. The concern goes beyond digitisation. Workers reportedly remove book spines and separate the pages for faster scanning. They then discard the physical copies. This matters more for old and rare books than for widely available titles. Some out-of-print works exist in only a limited number of copies. They may also contain handwritten notes, inserts, bindings or other physical details. A digital text file cannot preserve these features. European booksellers have reported unusual bulk orders involving thousands of unrelated and obscure titles. Some sellers suspect AI companies or intermediaries are behind the purchases. The orders do not resemble normal collecting or library buying. An Irish bookseller, Kennys, recently flagged an order for 5,000 obscure titles. Other sellers have also reported requests involving specialist, foreign-language and out-of-print books. Amazon is not the first major AI company to turn to books. Anthropic, Meta, OpenAI and Google have all used book collections for AI development. Some companies have also supported projects that make digitised book collections available for AI research. However, their methods of obtaining and processing the material differ. How AI companies are using books Anthropic: The Claude developer provides the clearest earlier example of destructive book scanning. Under an internal programme known as Project Panama, Anthropic bought millions of physical books, removed their bindings, scanned their pages and discarded the originals. Court documents described the project as an effort to "destructively scan all the books in the world." The digitised material was added to Anthropic's library, and sets and subsets of books were used to train its LLMs. Meta: Meta has used books to train its Llama models. It has also faced lawsuits over allegations that it obtained millions of copyrighted books through LibGen. The repository contains pirated books and research papers. Before that controversy, Meta executives also discussed other ways to secure book data. According to recordings of internal meetings reported by The New York Times, Meta staff discussed buying publisher Simon & Schuster in 2023. Such an acquisition could have given Meta access to a large catalogue of books for AI training. Some employees also considered paying $10 per book for licensing rights to new titles. Meta has maintained that its use of information to train AI models is consistent with existing law. OpenAI: OpenAI has faced separate copyright lawsuits over material allegedly used for AI training. It has also supported Harvard's Institutional Data Initiative alongside Microsoft. The project released nearly one million digitised public-domain books in 254 languages for AI research, including works dating to the 15th century. Unlike destructive scanning, the initiative is based on library collections that have already been preserved and digitised. Google: Google's large-scale use of digitised books predates the generative AI boom. Through Google Books, the company began working with libraries in the 2000s. It helped digitise millions of library books and make their contents searchable. The project triggered a long copyright battle that ultimately ended in Google's favour. More recently, Google worked with Harvard to recover public-domain volumes. Google Books had previously digitised these books. The companies cleared them for release as part of a dataset for AI researchers. Unlike the Amazon and Anthropic cases described here, there is no evidence that Google destroyed rare physical books for AI training. The distinction is important. These companies are all part of the wider push to obtain book-based data. However, there is stronger evidence that Amazon and Anthropic acquired and dismantled physical books. Meta's controversy centres on allegedly pirated digital books. The OpenAI- and Google-linked projects described here involve digitised public-domain collections. The Anthropic ruling addressed destructive scanning Anthropic's book programme became particularly important after a US federal court ruled in June 2025 on whether the company's use of copyrighted books amounted to infringement. The court separated three issues: training AI on copyrighted books, digitising lawfully purchased physical books, and obtaining books through piracy. It ruled that using copyrighted works to train Anthropic's LLMs was fair use because the training was transformative. "Like any reader aspiring to be a writer, Anthropic's LLMs trained upon works not to race ahead and replicate or supplant them -- but to turn a hard corner and create something different," the court said. The court also found fair use when Anthropic bought physical books, converted them into digital copies and destroyed the originals. It noted that the process replaced one physical copy with one digital copy rather than creating additional copies for distribution. "The print original was destroyed. One replaced the other," the ruling said. But the court rejected Anthropic's fair-use defence for more than seven million books. Anthropic had downloaded them from pirate sources and kept them in its wider library. This included books it did not ultimately use for AI training. The court said obtaining an illegal copy could not be justified by a later transformative use. Anthropic later agreed to a $1.5 billion settlement with authors over pirated books. The settlement received final court approval in July 2026. Why AI companies want physical books AI developers need huge amounts of high-quality text to train large language models. Older books are increasingly attractive because they contain long-form, edited human writing and were produced before the rapid spread of generative AI. Many rare, specialist and out-of-print books are also missing from the public internet or do not exist as clean digital files. That has become more important as researchers warn about "model collapse". It can occur when newer AI models repeatedly train on material generated by earlier models. This process can make AI systems less reliable. A 2024 Nature study found that indiscriminately training models on synthetic, machine-generated data can gradually distort the underlying distribution of information learned from human-created material. In simple terms, old books offer something AI companies increasingly want: text that is clearly human-written and less likely to be contaminated by AI-generated content. Preservation versus speed: Destructive scanning is attractive because it is fast. Removing a book's spine allows loose pages to pass through high-speed scanners and generally produces cleaner images for optical character recognition. But it is not the only way to digitise books. The Internet Archive uses non-destructive scanning for old and fragile collections, with workers turning pages manually so the physical books remain intact. Google has also developed non-destructive scanning systems, although these methods can be slower and more expensive. The debate therefore goes beyond copyright. A company may have the legal right to destroy a book it buys. But that right does not address the preservation consequences if multiple AI firms buy and dismantle uncommon books at scale. That concern is now spreading through the rare-books trade. Some booksellers have said they may refuse suspicious bulk orders because a short-term sale could contribute to the disappearance of difficult-to-find physical works.
[22]
Amazon joins other tech giants in gobbling up books to train AI, report finds
Amazon is reportedly buying massive quantities of rare books, scanning their pages to train its AI models and destroying the tomes in the process -- joining Anthropic, Meta and other tech giants in operating massive book scan-and-destroy operations. The goal, according to a report by news site 404 Media, is to train AI on well-written texts as opposed to low-quality internet slop, which AI models have already for the most part devoured. Slicing machines sever the spines of the books to speed up the process and cut costs, according to reports that have detailed operations by Amazon and others playing out in non-descript warehouses. 404 Media said it uncovered Amazon's book-buying operation by hiding a tracking device in a rare volume that it suspected an AI firm would purchase. Its movements were traced across states to its final destination - an Amazon warehouse in Las Vegas. "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use," an Amazon spokesperson said. The bookseller - which 404 Media kept anonymous - agreed to slip an Apple AirTag provided by the publication into a huge order the vendor received in July for around 1,000 works through a book marketplace. Warehouse workers at the Las Vegas location said their days are spent taking in heaping shipments of printed books and preparing them for AI scanning. The process - also detailed in a January Washington Post article on Anthropic and other AI labs' own book-scanning efforts - involves automated machines that chop the bindings off the orphaned volumes. The mascot of the Amazon team that works at the warehouse, named VGT3, is a Tyrannosaurus rex showing its chompers while holding a book, according to 404 Media. "I work at VGT3 here in Vegas, and all we do is scan books," one Amazon employee wrote on an online forum for the company's workers. "Some are assigned to cut books, and others go to receive where they get books and scan the bar codes." Another laborer wrote: "Working at VGT3 is nice all we do is scan books," adding, "It's so cool." Some booksellers logged huge spikes in sales over the past year - which they suspected were driven by AI companies seeking to get ahold of new training data, according to 404 Media. While they couldn't be sure, booksellers believed AI giants were behind the bulk orders because of the exceptional number of books they were buying and the seemingly random selection of volumes, 404 Media reported. Another giveaway was that the buyers didn't appear to care about prices, unlike libraries and universities. Books have become particularly valuable to AI labs, especially ones authored prior to 2022 - the year OpenAI first released ChatGPT - as they are certain to not contain AI-generated text. Companies have said training models on AI text can actually wreck havoc on models, a concept known as "model collapse." One Anthropic co-founder believed that training AI on books could teach them "how to write well" rather than imitating "low-quality internet speak," according to a January Washington Post report. A 2024 Meta internal email called securing a digital cache of books "essential" to competing in the AI race. Anthropic executives in early 2024 ramped up their book-scanning project - code named "Project Panama" - and sought to keep it under wraps, according to the Washington Post. "Project Panama is our effort to destructively scan all the books in the world," Anthropic executives said in an internal planning document that was unsealed in legal filings in January. "We don't want it to be known that we are working on this." The report noted that Anthropic initially considered purchasing books from used bookstores and libraries including the New York Public Library -- or "a new library that is chronically underfunded," according to documents cited by WaPo.
Share
Copy Link
An investigation using a hidden AirTag has exposed Amazon's systematic destruction of rare books at its Las Vegas VGT3 facility. The tech giant is buying bulk orders of rare books, cutting their spines, and scanning them to train AI models, raising ethical concerns among book lovers and independent bookstores worldwide.
An investigative report by
1
has confirmed what booksellers suspected for over a year: AI companies are destroying rare books to train AI models. The investigation placed an Apple AirTag inside a rare book that was part of a 1,000-book bulk order, tracking it directly to Amazon's VGT3 facility in Las Vegas. The warehouse, which operates within the larger LAS8 facility, exists solely to process books for AI training data. Employees told2
that their entire job involves cutting spines off books and feeding pages into industrial scanners. The facility's logo—a Tyrannosaurus rex clutching a book in its claws—leaves little doubt about the fate of these literary works. "All we do is scan books," one employee4
.
Source: 9to5Mac
Amazon declined to specifically address the AI training allegations, providing only a vague statement: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use." However, the company is developing frontier AI models that require massive amounts of unique training data to compete with Google, OpenAI, and Anthropic. The need is so urgent that VGT3 workers reported in online forums that Amazon nearly ran out of books to scan earlier this year, raising fears the warehouse might shut down. The supply shortage was temporarily resolved, but the facility's continued operation depends on a steady stream of rare books. Scanning books for AI training offers Amazon a crucial advantage: these texts predate 2022, ensuring they weren't AI-generated. When large language models train on AI-generated text, they risk model collapse, where output quality degrades significantly.
The investigation revealed that AI companies are systematically working through lists of ISBNs to ensure comprehensive coverage of published works. VGT3 employees are trained to scan barcodes or ISBNs before processing each book, giving credence to booksellers' theory that "AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs,"
1
reported. Independent bookstores across Europe have received suspicious bulk orders for thousands of books containing seemingly random titles. Tomás Kenny of Kennys Bookshop in Galway, Ireland, received an order for 5,000 books that included eclectic selections like "A History of Connemara" alongside "The Eddie Hobbs Guide to your SSIA." Berlin bookshops reported similar orders containing outdated titles like "Pass Your Driving Test, 2018 Edition." Booksellers noted these orders "never" include rare books lacking ISBNs, suggesting automated systems are driving the5
.
Source: Gizmodo
The AI companies destroying physical books employ destructive scanning methods that stand in stark contrast to established preservation techniques. Amazon's VGT3 facility cuts book spines before feeding pages flat into industrial scanners—the cheapest and fastest method available. Yet alternatives exist. Google Books patented non-destructive book-scanning technology in 2009 using V-shaped scanners that lay books open naturally without disturbing bindings. The Internet Archive has refined careful, manual scanning processes since 2010. Scanner Eliza Zhang has digitized more than 3 million pages, 14,000 foldouts, and 18,000 items by turning pages with "clean, dry human hands"—considered the best method for brittle, fragile books. The Archive's approach prioritizes "zero errors" and ensures each rare book only undergoes scanning once. Chris Freeland, director of library services, confirmed this remains their standard
3
. However, AI firms racing to advance frontier models appear unwilling to invest in slower, more expensive non-destructive techniques.
Source: Fast Company
Related Stories
The practice of Amazon is trashing rare books has intensified ethical concerns and legal battles. Anthropic faced a $1.5 billion settlement in 2025 for maintaining 7 million pirated books in its library. Internal documents from "Project Panama" stated plainly: "Project Panama is our effort to destructively scan all the books in the world." Meta was revealed to have torrented 82TB of pirated books for AI training. While U.S. courts have ruled that using books to train AI constitutes fair use, the penalties focused on copyright infringement from pirated copies. Germany takes a stricter stance. "Under German law scanning books - regardless of the purpose - would not be permissible and constitute a clear violation of copyright law," said Thomas Koch, spokesperson of Germany's Publishers and Booksellers' Association. European bookstores suspect bulk orders are shipped to local addresses as collection points before eventual transport to the U.S. for processing.
The bookseller who planted the AirTag told
1
that rare books possess "historical value, intellectual value, sentimental value" derived from "all sorts of things" that "the AI companies don't care about. They just want the content as a bunch of words strung together." AI firms often target older books with lower monetary value—foreign language works never translated or titles never widely distributed. While these may not be prized first editions, they represent irreplaceable cultural artifacts. Michael Burry reportedly labeled the practice "evil incarnate." Yet some Reddit users questioned whether destruction matters for books gathering dust on shelves for decades. For independent bookstores, bulk orders offer financial lifelines during the transition to digital reading, creating a painful ethical dilemma. The orders help stores survive, but contributing to the systematic destruction of literary heritage troubles many sellers. As AI training data demands continue escalating, the tension between preserving physical books and feeding AI's insatiable appetite for text shows no signs of resolution.Summarized by
Navi
[1]
[4]
23 Jul 2026•Policy and Regulation
30 Jul 2026•Policy and Regulation

26 Jun 2025•Technology

1
Technology

2
Policy and Regulation

3
Technology
