Flash News

SemiAnalysis: Kimi K3 will reverse the pull of Yvette and HBM

On 19 July, Kimi K3, a large model under the cover of the month, introduced linear attention mechanisms, raising concerns about the possible weakening of the demand for Yin Weida, HBM and network equipment. However, SemiAnalysis, a semiconductor research institute, has recently made quite different judgements: the large size of K3 parameters and the need for a reasoning structure not only does not weaken high-end AI hardware requirements, but may further enhance demand for high-end GPU, HBM and high-speed interconnection equipment. SemiAnalysis states that the K3 parameter size exceeds 2.8 trillion and the model weight capacity exceeds 1.5 TB HBM. Even in a relatively limited number of user scenarios, KV caches still require large amounts of offloading to CPU DDR5 memory and NVME storage, and HBM space does not generate significant surplus. More importantly, the dark side of the month had previously revealed that the efficient reasoning deployment of K3 required a large-scale extended-area structure of at least 64 chips. This hardware requirement corresponds to the design direction of a stage AI system such as GB200/GB300 NVL72 in Weida. SemiAnalysis argued that the logic of the market's prior understanding of linear attention as “dividing GPU demand” was biased. The real impact could be the opposite: a more efficient model structure reduces AI reasoning costs and will drive more applications to land, thus stimulating the long-term demand for GPU, HBM, DRAM and network infrastructure. (Grunting)

OKX - Unlock Rewards