TechCrunch refers to 404 Media reports that the Amazon is purchasing a large number of rare books and removing post-ridge scanning to obtain text data that can be used to train the AI model. A rare book that was placed in a tracker was eventually sent to an Amazon facility in Las Vegas.
Book tracked to VGT3
Media states that the book finally reached the Amazon facility called VGT3. The site is marked with a “book by dinosaur claws”. The Amazon indicated to 404 Media that the company would buy books through commercial channels to improve the products and services used by its clients.
Incisive text as new language source
Reports indicate that training large-language models require very large amounts of text data, and that public content directly available on the Internet has been extensively used. For companies like Amazonia, hard-to-reach, rare or online books are becoming new sources of training.
Model training shift to belowline information
The value of such texts is also that content that was published earlier than 2022 does not normally contain AI-generated texts. It was mentioned that if models continued to ingestion too much AI-generated content, there could be “model collapse”, i.e. a gradual decline in output quality. It also makes it possible for AI to focus again on how to get quality original text.
