Artificial intelligence developers have begun mass-purchasing printed books published before 2022 to ensure the purity of training data and protect models from degradation.
What Happened
AI development companies are shifting toward large-scale procurement of physical printed editions released before 2022. To accelerate the digitization process, a destructive scanning method is being used, where book spines are cut to allow pages to pass through scanners as quickly as possible.
Context
The primary reason for this approach is the need to avoid "contaminating" training sets with synthetic content, known as AI slop. Using data generated by modern language models can lead to "model collapse"—a progressive degradation of the quality and diversity of AI responses.
Why It Matters for the Industry
This process is creating a new market for high-quality physical data and is changing approaches to copyright protection, shifting the focus toward the fair use doctrine. In the long term, this could create a technological divide between models trained on "clean" historical data and those that rely primarily on synthetic web content.
Why It Matters for Users
For readers, this means a risk of the irreversible loss of rare physical editions due to their mass destruction during the digitization process. On the other hand, it guarantees that future generations of AI will possess higher quality knowledge and fewer hallucinations caused by errors in internet texts.
What Is Not Yet Known / Limitations
There are disagreements regarding the consequences: technical specialists see this as the development of new data pipelines, while legal and cultural experts point to the threat of losing physical cultural heritage.
Sources
Author
Look at AI, Editorial Team
