Training Schemes and Features
MindSpeed LLM includes distributed pretraining and distributed fine-tuning schemes.
Distributed Pretraining
The measured pretraining performance of MindSpeed LLM is as follows.
| Model Family | Experimental Model | Hardware | Cluster Size | Throughput (Tokens/s) |
|---|---|---|---|---|
| Qwen3 | 8B | Atlas 900 A3 SuperPoD | 1x16 | 7617.002 |
| 30B | Atlas 900 A2 PODc | 2x8 | 2318.373 | |
| DeepSeek-V3 | 671B | Atlas 900 A3 SuperPoD | 32x16 | 914.97 |
Pretraining Schemes
| Scheme Type | MCore | Released | Contributor |
|---|---|---|---|
| Multi-Sample Dataset Pretraining | ✅ | ✅ | 【Ascend】 |
| Multi-Sample Pack-Mode Pretraining | ✅ | ❌ |
Acceleration Features
| Scenario | Feature Name | MCore | Released | Contributor |
|---|---|---|---|---|
| SPTD Parallelism | Tensor Parallelism | ✅ | ✅ | 【Ascend】 |
| Pipeline Parallelism | ✅ | ✅ | ||
| Virtual Pipeline Parallelism | ✅ | ✅ | ||
| Sequence Parallelism | ✅ | ✅ | ||
| noop layers | ✅ | ✅ | ||
| Long-Sequence Parallelism | Ascend Ring Attention Long-Sequence Parallelism | ✅ | ✅ | |
| Ulysses Long-Sequence Parallelism | ✅ | ✅ | ||
| Hybrid Long-Sequence Parallelism | ✅ | ✅ | ||
| MoE | MoE Expert Parallelism | ✅ | ✅ | |
| MoE Reordering Communication Optimization | ✅ | ✅ | ||
| Memory Optimization | Parameter Replica Reuse | ✅ | ✅ | |
| Distributed Optimizer | ✅ | ✅ | ||
| Swap Attention | ✅ | ✅ | ||
| Recompute | ✅ | ✅ | ||
| Norm Recompute | ✅ | ✅ | ||
| O2 BF16 Optimizer | ✅ | ❌ | ||
| Fused Operators | Flash Attention | ✅ | ✅ | |
| Flash Attention Variable Length | ✅ | ✅ | ||
| Fused RMSNorm | ✅ | ✅ | ||
| Fused SwiGLU | ✅ | ✅ | ||
| Fused Rotary Position Embedding | ✅ | ✅ | ||
| GMM | ✅ | ✅ | ||
| Matmul Add | ✅ | ✅ | ||
| Communication Optimization | Gradient Reduce Communication Overlap | ✅ | ✅ | |
| Recompute in Advance | ✅ | ✅ | ||
| Weight All-Gather Communication Overlap | ✅ | ✅ | ||
| MC2 | ✅ | ❌ | ||
| CoC | ✅ | ❌ | ||
| Ascend Gloo Archive-to-Drive Optimization | ✅ | ✅ |
Distributed Fine-Tuning
The measured instruction fine-tuning performance of MindSpeed LLM is as follows.
| Model | Hardware | Cluster | Scheme | Sequence | Throughput (Tokens/s) |
|---|---|---|---|---|---|
| Qwen3-30B | Atlas 900 A3 SuperPoD | 8x16 | Full Parameters | 256K | 3774.914 |
| Qwen3-32B | Atlas 900 A3 SuperPoD | 8x16 | Full Parameters | 256K | 1435.603 |
| DeepSeek-V3-671B | Atlas 900 A2 PODc | 8x8 | LoRA | 4K | 978.914 |
Fine-Tuning Schemes
| Scheme Name | MCore | LoRA | Released | Contributor |
|---|---|---|---|---|
| Single-Sample Fine-Tuning | ✅ | ✅ | ✅ | 【Ascend】 |
| Multi-Sample Pack Fine-Tuning | ✅ | ✅ | ❌ | 【Ascend】 |
| Multi-Turn Conversation Fine-Tuning | ✅ | ✅ | ❌ | 【Ascend】 |
Acceleration Features
| Scenario | Feature | MCore | Released | Contributor |
|---|---|---|---|---|
| LoRA Fine-Tuning | CCLoRA | ✅ | ✅ | 【Ascend】 |
| Long-Sequence Fine-Tuning | Long-Sequence CP | ✅ | ❌ | 【Ascend】 |