AI companies now confront mounting public criticism after evidence emerged that some firms purchase rare physical books, slice them open with industrial cutters, scan the pages for training data, and then destroy the originals to expand their models. A circulating video demonstrates that fully intact digitization remains possible through specialized equipment that leaves volumes undamaged, underscoring an available alternative many labs have chosen not to prioritize. This approach raises urgent questions about knowledge ownership in an era when closed systems from major developers continue to dominate while open weight alternatives struggle for equivalent high quality sources.
The practice entered public view through Anthropic’s internal “Project Panama,” a bulk book acquisition and destructive scanning program detailed in court filings unsealed during the Bartz v. Anthropic copyright litigation. According to the Washington Post’s investigation into Anthropic’s Project Panama, the company hired Tom Turvey, the former head of partnerships for Google Books, to source millions of pre-2022 titles from used-book sellers including Better World Books and World of Books. Workers used hydraulic machines to sever spines, flatten the pages for rapid industrial imaging, and discard what remained so only the digital version would exist. The legal backdrop is the June 23, 2025, federal court decision in the Northern District of California, recorded in the Bartz v. Anthropic summary judgment order on fair use, in which Judge William Alsup held that digitizing lawfully purchased print books for AI training qualified as fair use, with the destruction-to-digital-replacement pattern cited as a factor in the analysis. Anthropic later settled related claims involving pirated digital books for $1.5 billion.
The video circulating on social media appears drawn from promotional material produced by the Austrian firm Treventus to showcase its ScanRobot system. According to the Treventus product documentation for the ScanRobot non-destructive scanner, the machines perform automatic page turns with controlled air flow inside a narrow V-shaped cradle that never forces the binding beyond safe limits. Prism optics capture both pages at once without glass pressure or distortion, achieving high throughput while the original volume stays completely intact and can return to a shelf or archive. Such equipment has existed for years yet remains underused by the largest training operations that prefer speed and legal simplicity over preservation.
Elon Musk addressed the issue directly by instructing the SpaceXAI team to place rare volumes in a library and apply the slower careful method rather than spine removal. In his X post directing the SpaceXAI team to preserve rare books, Musk wrote:
“I’ve asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning.”
This stance stands in contrast to the closed model strategies pursued by developers of ChatGPT, Claude, and Google systems, which retain exclusive control over the resulting data advantages. Advocates for open weight models have long argued that concentrated private corpora create unfair barriers, yet the resource intensity of bulk physical acquisition makes it difficult for smaller or community driven projects to match the depth now flowing into proprietary systems. The result is an accelerating split where a few companies absorb unique historical text while open efforts rely more heavily on noisier public web sources.
Book knowledge enters the models as dense, structured sequences that improve factual grounding, long form reasoning, and domain expertise far beyond typical internet scrapes. Once ingested, the content helps reduce hallucinations in specialized subjects such as history, science, and literature, enabling agents to draw on centuries of accumulated human insight during complex tasks. At the same time, autonomous AI agents have faced repeated security incidents involving prompt manipulation and unauthorized tool use, raising the possibility that highly refined book derived knowledge could be extracted or misdirected if agent frameworks remain vulnerable. This concern parallels broader AI safety research, including a recent study suggesting AI models may choose to resist shutdown to preserve their own operation. The combination of richer internal libraries and imperfect agent safeguards therefore heightens both capability and risk.
For humans the effects reach beyond technical performance. Physical destruction of scarce editions erodes the shared material record that libraries and collectors have protected for generations, leaving future readers dependent on whatever filtered digital version a private model chooses to surface. Concentration of that knowledge inside closed systems can deepen information asymmetry, limit independent verification, and shape cultural memory according to corporate priorities rather than open scholarly consensus. The wider cultural conversation around AI has surfaced similar tensions, from a London theater canceling the premiere of an AI-written film to Timbaland’s debut of the first fully AI-generated music artist. While better trained models may accelerate discovery and education, the irreversible loss of original objects removes a tangible link to the past that no synthetic reconstruction can fully restore, placing long term human heritage in tension with short term competitive gains in artificial intelligence.
Broader AI development continues to wrestle with these same trade-offs across energy demands, copyright enforcement, and the push for safer agent architectures. The copyright question itself has spawned parallel claims, including Addison Rae’s copyright claim against the Department of Homeland Security over the use of her song in a recruitment video, illustrating how artists and rights holders are increasingly pushing back against unauthorized appropriation in both commercial and governmental contexts. The current episode around rare books simply makes visible a deeper contest over who controls the highest quality training material and whether preservation of the physical world will remain a design constraint or an optional afterthought. Choices made now will determine whether the next generation of systems expands human understanding while still safeguarding the artifacts that made that understanding possible in the first place.


