At this point, you’re probably used to hearing all about AI and how it's constantly hungry for more information, right? Industry leaders are desperate to get their hands on any and every piece of high-quality data to better train their increasingly sophisticated systems. Well, now it looks like they’ve taken that desperation to the next level.
Booksellers all over Australia, Europe, and the United Kingdom have started reporting weird bulk orders for thousands of obscure titles. The buyers aren’t identifying themselves, and they aren’t reading the books or selling them on. Instead, they’re scanning the books so they can be digitized, fed into AI systems, and then pulped into mush.
Court documents have already proven this is a common practice, and companies like Anthropic have apparently already destroyed millions of books.
If you’re an avid reader, this sounds like a crime against humanity. But even if you never pick up a book, there’s a bigger story at play here. After all, if AI companies are willing to destroy physical copies of knowledge to obtain data, what does that say about the economics behind the current AI boom?
Why Do AI Companies Want Your Books?
The math is simple for AI advocates and developers. Agentic AI needs data. Large language models (LLMs) learn and become more advanced by processing collections of text. This enables them to absorb the information and generate more complex and organic-sounding responses.
That’s why the quality of training materials matters. If you just feed low-quality internet spam back into an AI model, it’s never going to get better. You’ve got to train it on innovative academic research, authoritative documents, and well-edited works of literature. As a result, demand across the second-hand book market has skyrocketed over the past couple of years.
Not only do physical books contain structured information that’s already withstood quality control measures. They also haven’t been tainted by all the AI-generated slop that now plagues our newsfeeds. Older, human-written material is now some of the most valuable training data available for AI companies.
Unfortunately, it looks like they’ll stop at nothing to get that data.
We all know the AI industry moves fast, and nobody wants to get outpaced by the competition. That’s why Anthropic came up with a new practice as part of its “Project Panama” initiative.
Project Panama was the codename of a drive to purchase copyrighted books and then cut off the bindings so they could be scanned and fed into Claude easier. After Anthropic’s team digitized the pages, most of the physical books were destroyed.
Why use a codename? According to an internal memo, executives didn’t want anybody to know about the company’s “effort to destructively scan all the books in the world.”
Sounds shady, right? That’s why loads of authors decided to take Anthropic to court. But judges have ruled in the company’s favor. They decided the way these books were being used was legal and didn’t have much to say about how they’re getting destroyed out of convenience.
From a cultural perspective, that’s difficult to swallow. It’s also given other AI companies carte blanche to keep on bulk-buying books to scan and then pulp forever.
Not every book they purchase is valuable, sure. But booksellers have said a lot of the suspicious orders they’re getting are for nonfiction titles and academic works that have been out of print for decades. This creates an issue of preservation if one private company is scanning a rare piece of information and then discarding copies so that nobody else can get their hands on it.
But let’s set the cultural ramifications of this new type of gatekeeping to one side for a minute.
There’s another element at play here, and that’s the enormous economic value being attached to digitizing all the world’s literary works. Anthropic has already agreed to a $1.5 billion settlement with authors over allegations the company has pirated their books to train AI systems, and a similar lawsuit was filed last month against Google (GOOG) (GOOGL).
Courts are essentially being forced to decide where the line sits between copying information to create transformative technology and commercially exploiting the work these companies are scanning. The decisions those courts make are going to have enormous financial consequences, and that’s something markets need to prepare for.
Why Should Markets Care About Any of This?
AI might be a revolutionary technology, but it’s not cheap to run. Servers already consume a huge volume of resources, and that’s something investors have to price in when they're working out the value of AI companies. One consideration that most of them haven’t been pricing in is training costs.
Anthropic has essentially just shelled out $1.5 billion as an unexpected training cost because they hadn’t licensed the copyrighted material they’re using. Now, other publishing firms are out for blood, and it’s inevitably going to hit some of these other AI companies just as hard. Training programs will need to start absorbing far greater costs, and those costs can ultimately eat away at a company’s overall value.
Then again, let's suppose more judges do decide these mass consumption projects are a “fair use” of copyrighted materials. There’s still an underlying risk of ownership. As freely available data continues to wither, AI companies are going to have to compete harder and pay more to get hold of new training data. That creates a perfect storm of issues concerning supply, economic sustainability, and quality.
It'll take more money than ever if companies like Anthropic want to continue to innovate. At what point do those costs become too great?
At the end of the day, that’s why you should be paying attention to this whole drama over “Project Panama” and how humanity's rarest books are being treated. Even if you don’t care about the ethical dilemma here, you can’t deny there’s a cost dilemma. That’s bad news for both readers and shareholders.
On the date of publication, Nash Riggins did not have (either directly or indirectly) positions in any of the securities mentioned in this article. All information and data in this article is solely for informational purposes. For more information please view the Barchart Disclosure Policy here.