As artificial intelligence (AI) keeps evolving every day, from viral meme videos of AI models eating weirdly to now producing hyper-realistic videos that make you look twice, all of it came from existing data that was fed to AI models.
But everything has a limit. Every time an AI trains data that was created by AI, the quality can begin to degrade.
This is why, in a recent report, AI companies are reportedly buying tons of physical books, especially books published before 2022, as these are seen as valuable sources of human-created data with less exposure to AI-generated content.
AI Inbreeding
Just like in genetics, where inbreeding can reduce genetic diversity and increase the chance of harmful traits being passed down, AI models can also lose diversity when they repeatedly train data generated by AI.
With each cycle, AI models can also lose information when they repeatedly train using AI-generated data.
Their products, whether images, videos, or writing, can become more predictable and average. Models can also begin to forget rare or less common information as each cycle passes.
This problem was discussed in a published paper titled “The Curse of Recursion: Training on Generated Data Makes Models Forget.”
For the next generation of AI models, this could mean having a much smaller pool of reliable sources to learn from.
This is why older books can carry special value for AI companies. Books published before the widespread explosion of generative AI are seen as useful sources of human-created information, with less risk of synthetic data being mixed into the training material.
ISBNdb: From Book Dealer to AI Middleman?
According to a report, ISBNdb, which operates one of the largest international databases for printing book metadata with more than 113 million titles, has reportedly been helping AI companies bulk-buy physical books.
The company, which previously focused on helping libraries, distributors, and bookstores find and sell books, was reportedly offering purchases ranging from one thousand to one million books, with the promise that the transactions would remain private.
The unusual purchases have left mixed feelings among booksellers. While sellers confirmed the unusual demand and financially benefiting from clearing out of old inventory, some also worried that uncommon books, including limited editions, could end up being destroyed rather than preserved.
The planned AI training through the bought physical books would reportedly be processed by removing their bindings with hydraulic cutters before being scanned. After the pages are digitized, the physical copies can simply be disposed of.
Under U.S. law, digitizing lawfully acquired printed books and using the resulting copies to train language models can potentially fall under fair use, although the legal questions surrounding AI training remain complicated.
However, the story took an unexpected turn. On July 28, ISBNdb removed the promotion from its website. Two days later, its management released a statement saying that the company was never actually purchasing or scanning old books for AI training.
According to ISBNdb, the promotion was only a market-demand test, and the program was never launched.




