Google is moving forward with an acquisition transaction related to AI training data. According to foreign media reports, the company proposes to pay $10 million to purchase part of the business data of the American Airlines Fairy to improve the product and the AI model. The United States bankruptcy court is still required to approve the transaction.
Involving internal communications and business records
Google indicated to the media that the company purchased a part of the Spirit enterprise data set that could be used to upgrade product performance and modelling capabilities. Google also stated that the transaction would not allow the company to obtain personal information.
Reuter quoted relevant documents as stating that the data to be sold covered not only general business records but also Microsoft Teams news, calendar content, and marketing, productivity and business-related information. This means that AI ' s access to training data is further extending from open Internet content to inside the enterprise.
Disposal of surplus assets after the insolvency of Spirit
Spirit ceased operations in May this year and sold the remaining assets in insolvency proceedings. When the airline enters insolvency, the data are usually packed and disposed of together with other assets and ultimately taken over by the buyer.
This data transaction is not without competitors. According to reports, the AI Data Company Mercor had made a $7.5 million offer, but Google's bid was higher. The United States bankruptcy judge Sean Lane planned to consider the deal at a hearing on Wednesday.
AI accelerated the purchase of training data
This transaction also reflects that AI is continuously expanding the range of paid data. In February 2024, Reddit entered into a collaboration with Google to open up user content to the latter for almost 20 years, with a reported value of approximately $6 million per year. In May 2024, OpenAI entered into a similar agreement with Reddit to introduce it into ChatGPT.
Earlier this January, Wikimedia announced that agreements had been signed with a number of AI companies, including Microsoft, Google, Amazon and Meta, to provide them with access to Wikipedia content on a large scale.
The procurement of training resources is expanding, from public content to enterprise data. As model development costs rise, who has continued to have access to high-quality, available and compliant data is becoming a realistic condition in AI competition.
