Amazon's AI Training Raises a New Question: What Happens to Rare Books?
Artificial intelligence companies need enormous amounts of information to train increasingly capable AI models.
But a new report involving Amazon has raised an unusual question: how much of humanity's physical knowledge should be sacrificed to build artificial intelligence?
Rare books can be valuable sources of information because many historical texts are not available online.
Why Books Matter to AI
Large language models learn from huge collections of text.
Much of the material used for training already exists online, but there are still enormous amounts of information locked inside physical books, archives and other documents.
Historical books can contain information that may not exist in easily accessible digital databases.
The Controversy
The situation raises several important questions.
Should companies be allowed to digitize or destroy physical copies of books for AI training?
Who owns the rights to the information?
And what happens when the physical source is extremely rare?
These questions are becoming more important as AI companies search for increasingly diverse training data.
The Bigger AI Problem
AI development depends heavily on data.
The industry has already faced debates about copyrighted websites, books, images, music and other creative material being used to train AI systems.
As AI models become more powerful, the competition for high-quality training data is becoming increasingly intense.
Final Thought
AI may eventually become one of humanity's most powerful technologies, but the way it obtains its knowledge matters.
The Amazon situation shows that the AI revolution isn't only about processors and software.
It is also about who controls information and how that information is preserved.
