AMD release full open source MoE large model Instella-MoE, 16B parameter size challenge mainstream open source model
On 25 July, AMD announced the launch of the Full Open Source Mixed Expert Model (MoE) Instella-MoE, which has a total parameter of 16 billion, each Token activates a parameter of 2.8 billion and is described as leading performance in the same size open source language model. AMD states that the Instella-MoE is based entirely on its own AMD Institute MI300X and MI325X GPU, and that the ROCm repository was completed from zero training and introduced structural innovations such as Gated Multi-head Late Action and FarSkip-Collective to improve the efficiency of training and reasoning. Performance tests show that the Instella-MoE-16B-A3B base model scored an average of 76.7 points, ranked ahead of the SmolLM3-3B, OLMo-3-7B models and competed with larger models when only 2.8 billion parameters were activated. In addition, the model supports 64K Token context processing and completes a complete training process such as pre-training, medium-term training, long-term context extension, oversight fine-tuning (SFT), direct preference optimization (DPO) and enhanced learning (RL). This time, AMD synchronises the release of all of the Instella-MoE models weights, training configurations, data matching, intermediate check points and reasoning codes to promote open AI model research and replication. AMD indicates that Instella-MoE demonstrates the ability to conduct large-scale MoE model training based on AMD hardware and open software ecology, and will continue to advance the development of larger, more reasoned and more efficient open-source language models in the future。
