← Back to News
industryMarkTechPostAug 1, 2026

AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs

Read original ↗

Sentiment: neutral

TL;DR

AMD has released Instella-MoE-16B-A3B, an open-source Mixture-of-Experts language model trained on Instinct GPUs, which could advance research by providing transparency in the training process.

Detailed Summary

AMD has released Instella-MoE-16B-A3B, an open-source Mixture-of-Experts language model trained on Instinct MI300X and MI325X GPUs, featuring 16 billion total parameters with only 2.8 billion active per token. The model utilizes Gated MLA and FarSkip-Collective techniques and is fully accessible for research and development purposes. This release aims to advance the field of large language models by providing transparency in training stages and hardware optimization methods.

Key Points

  • • AMD released Instella-MoE-16B-A3B as an open-source Mixture-of-Experts LLM.
  • • It features 2.8B active parameters per token during inference.
  • • Trained on Instinct MI300X and MI325X GPUs from scratch.
  • • Uses Gated MLA and FarSkip-Collective techniques for efficiency.

Source: MarkTechPost

Score: 24