Head AI Laboratory and Large Enterprises Training in Continuous Adding Models, leading to the rapid expansion of the training data and data labelling industry. TechCrunch cites sources who claim that the four-year-old Micro1 has increased the annualized total flow of water from $100 million to $500 million over the past eight months, indicating that this track is still being rapidly released.
Raised to $500 million in eight months.
According to informed sources, Micro1 currently has an annualized total of about $500 million. Companies typically employ specialists in the fields of doctors, lawyers, scientists, etc., to participate in data production and labelling on a contractual basis. According to this model, Micro1 actually retained between 60 per cent and 70 per cent of its income, with corresponding annualized income between $150 million and $200 million.
While the scale still lags behind some of their peers, it is clear that market demand is sufficient to support the growth of multiple suppliers at the same time. According to the report, Mercor's annualized income this summer has reached $2 billion, and Handshanke reached $1 billion earlier.
- Total annual flow of water: about $500 million
- Annual clean-up income: approximately $150 million to $200 million
- Start of the last eight months: approximately $100 million
Synthetic data pulls up profitability
Micro1 is now making more use of synthetic data production methods that do not require manual participation, such as automatic generation of video content descriptions. According to the source, such operations continue to expand the size of the company ' s contracts and contribute to subsequent increases in profitability.
Another change is that some of the data products can be re-sold to multiple customers rather than serving only a single customer. For this type of “ready-made dataset”, the Māori rate is said to be between 80 and 90 per cent. This means that training data companies are moving from traditional manpower-intensive labelling to more standardized and replicable product models.
- Māori rates in existing data sets: about 80 to 90 per cent
- Business orientation: synthetic data, reusable sales data Set
Data sales leading to discussion
However, data sets that can be resold are also controversial. Critics believe that if such standardized data were sold to Chinese AI developers, it might help to bring the model closer to the level of capability of the head model in the United States.
Micro1 founder Ali Ansari stated on platform X last month that the company did not sell data to Chinese model developers as some of its peers did. He also stated that some manpower data companies were working with “foreign opponents” and that Micro1 would not do so.
Moving from recruitment to data operations
Like Mercor, Micro1 was originally an AI recruitment firm. Ansari found that some of the data indicated that clients would use their AI platform to screen and recruit engineers to participate in the labelling, and subsequently decided to shift the focus of operations to training data.
In addition to the model output assessment, Micro1 is also building a robotic pre-training data set, which will allow hundreds of ordinary participants to record the day-to-day interaction of objects at home. Such data are usually used to train robots in understanding and implementing basic actions in the real environment.
It was also mentioned that Micro1 completed the A round of finance in September last year with a valuation of $500 million. TechCrunch learned that the company may have completed another round of financing in the near future, with a significant increase in valuations over the previous round. For the relevant information, Micro1 did not respond to the request for comment.
