Flash News

SemiAnalysis: Kimi K3 lower KV bandwidth, and IA network demand is not diminished

According to SimiAnalysis, according to the early release of In 2021, Kimi K3 uses KDA technology to reduce the KV cache bandwidth by up to 10 times, but the AI network demand will not contract. Kimi K3 has 2.8 trillion parameters and requires about 1.5 TB HBM bandwidth for each forward calculation and still needs to connect chips through high bandwidth networks such as GB300 NVL72. Wideep spreads 896 expert models to multiple GPUs, requiring more than 120 one-way calculations, and large expert models generate more extended network needs than KDA saves. More efficient attention mechanisms may drive the context length from 1 million tokens to more than 5 million tokens, which will increase the size of AI and increase network demand。

OKX - Unlock Rewards