AI reasoning speed up competition is warming. The French original company Kog did not opt for self-study of the chip, but rather took note of the optimization of the bottom software to further unleash the reasoning of the existing data centre of the enterprise GPU.

First aim at the high-time scenario.

This company received attention in May this year for a technology preview. Kog then demonstrated that using standard data centres like AMD MI300X and Nvidia H200 GPU could also achieve very fast single request decoder speed. The company subsequently stated that the exhibition had brought about 200 clear business leads.

Kog currently gives priority attention to professional workflows that require high response speed, particularly software engineering scenarios. According to the company, some of the heavy-intensity codes can be generated by users for hours, and the delay in reasoning has become an important issue affecting the experience and cost of use.

The company is also working with some design partners to facilitate the landing of applications. Such customers allow users to generate games or applications through hints, which, if the pace of reasoning increases, usually implies higher conversion efficiency and income space.

The small model is out 30000 TPS

Kog had previously proposed the goal of “30 times faster for large model reasoning”, but public demonstrations are still based mainly on small models. It was displayed using Laneformer 2B of about 2 billion parameters, requesting a single deduction speed of 30000 tokens per second. The model is now open.

However, market demand did not stop at small model fine-tuning. Kog indicated that potential customers were not prepared to focus on small models and that, since the presentation of the demonstration, the company had shifted its main effort to accelerating the development of larger models to match actual needs.

Validation of large model capabilities around September

Kog believes that the GPU's potential on the code of understanding has not yet been fully released. The CEO of the company, Gaël Delalleau, stated that there was still much room for optimization of the software layer as a new generation of GPU memory bandwidth continued to rise.

The cost of this approach is a longer R & D cycle. Kog states that for each appropriate new GPU, teams often need to invest weeks or even months in hardware-level research. The number of chips that the company can support in the short term remains limited in the current 11-person team size.

Kog plans to gradually access the methodology to intelligent-based development processes in the future to support more chips and models. The more critical node at the moment is to prove that the route is built on a large model. Delalleau anticipates that the company will complete the first major model at a rate of about 10 times faster around September, thereby demonstrating the progress of its clients and facilitating the follow-up round A.