industryMarkTechPostAug 1, 2026
AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs
Read original ↗Sentiment: neutral
TL;DR
AMD has released Instella-MoE-16B-A3B, an open-source Mixture-of-Experts language model trained on Instinct GPUs, which could advance research by providing transparency in the training process.
Detailed Summary
AMD has released Instella-MoE-16B-A3B, an open-source Mixture-of-Experts language model trained on Instinct MI300X and MI325X GPUs, featuring 16 billion total parameters with only 2.8 billion active per token. The model utilizes Gated MLA and FarSkip-Collective techniques and is fully accessible for research and development purposes. This release aims to advance the field of large language models by providing transparency in training stages and hardware optimization methods.
Key Points
- • AMD released Instella-MoE-16B-A3B as an open-source Mixture-of-Experts LLM.
- • It features 2.8B active parameters per token during inference.
- • Trained on Instinct MI300X and MI325X GPUs from scratch.
- • Uses Gated MLA and FarSkip-Collective techniques for efficiency.