任务地址:https://gitcode.com/mindspore/community/issues/2099 GLM-5 是智谱 AI 于 2026 年 2 月发布的最新 decoder-only 稀疏大语言模型(arXiv:2602.15763),在中文 NLP 生态中占据核心地位。当前 HyperParallel 仅支持 Qwen 系列模型,缺少 GLM 系模型覆盖。
tests/torch/integration/glm5/
examples/glm5/
本次交付聚焦训练基础设施接入(Phase 1),使用简化 dense 架构验证注册机制、Trainer 对接、checkpoint 闭环、精度对齐。完整 GLM-5 架构(MoE + MLA + DSA + MTP)在 Phase 2–5 逐步叠加。
GLM-5 完整架构(供参考):
核心价值:
core/context_parallel/
modules/moe.py
GLM-5 采用 MLA + DSA + MoE + MTP 四合一架构:
weight * normed
GLM-5 完整架构复杂度较高,采用分层交付策略——每个 Phase 交付一组可独立训练/验证的组件:
Phase 1 为最小可用交付:简化架构(dense + GQA + SwiGLU + RMSNorm)即可跑通训练闭环,同时验证注册机制、Trainer 对接、checkpoint 保存恢复。后续 Phase 在此基础逐步叠加真实 GLM-5 组件。
遵循现有 ModelSpec + register_spec 注册模式。model.name: glm5 触发 auto-discovery,Universal fields 由 _resolve_overrides() 映射到 GLM5Config。
model.name: glm5
_resolve_overrides()
GLM5Config
支持两套命名方案(GLM-5 标准布局 + GLM-4 旧版布局),tie 场景自动合成 lm_head.weight。后续 Phase(MoE/MLA/DSA)需扩展键映射逻辑处理 expert 权重、MLA 潜变量参数。
checkpoint_wrapper
_tp_plan
core/expert_parallel/
DSAIndexerContextParallel
DSASparseAttentionContextParallel
hyper_parallel/models/glm5/ ├── __init__.py # register_spec("glm5") + _build() + _resolve_overrides() ├── model.py # GLM5Config, GLM5ForCausalLM(Phase 1 dense 版本) ├── decoder.py # GLM5Decoder(Phase 2 加入 MoE 分支) ├── attention.py # MLA attention(Phase 3) ├── moe.py # MoE router + expert 层(Phase 2,复用 modules/moe.py) ├── dsa.py # DSA 索引/边界(Phase 4,对接 core/context_parallel/) ├── mtp.py # MTP 参数共享(Phase 5) ├── checkpoint.py # HF safetensors 权重加载 + 键映射 ├── parallelize.py # AC + FSDP + EP/CP 策略 └── state_dict.py # StateDictAdapter
@dataclass class GLM5Config: # ── 基础参数(Phase 1 使用) ── vocab_size: int = 151936 hidden_size: int = 1024 intermediate_size: int = 3072 num_hidden_layers: int = 24 # 训练用小模型;GLM-5 实际为 80 num_attention_heads: int = 16 num_key_value_heads: int = 4 # Phase 1 GQA;Phase 3 MLA 后废弃 head_dim: int = 64 # Phase 1;MLA 后为 kv_lora_rank max_position_embeddings: int = 131072 rms_norm_eps: float = 1e-6 rope_theta: float = 500000.0 tie_word_embeddings: bool = True # ── MoE 参数(Phase 2 启用) ── num_experts: int = 256 num_experts_per_tok: int = 8 num_dense_layers: int = 3 # 前 N 层为 dense,其余为 MoE moe_intermediate_size: int = 1024 # ── MLA 参数(Phase 3 启用) ── kv_lora_rank: int = 576 # KV 压缩秩 qk_rope_head_dim: int = 64 # RoPE 维度(MLA 中 Q/K 分离) v_head_dim: int = 128 # ── DSA 参数(Phase 4 启用) ── dsa_topk: int = 2048 # 每 token 选中的 Top-K 历史 token dsa_indexer_dim: int = 64 # ── MTP 参数(Phase 5 启用) ── num_mtp_layers: int = 3 # MTP 共享层数
Phase 1 仅使用基础参数。MoE/MLA/DSA/MTP 参数在对应 Phase 启用。
GLM5Config (@dataclass) GLM5RMSNorm (nn.Module) — 标准 weight * normed GLM5Decoder (nn.Module): ├─ input_layernorm: GLM5RMSNorm ├─ self_attn: GroupQueryAttention(Phase 1)/ MLA(Phase 3) ├─ post_attention_layernorm: GLM5RMSNorm └─ mlp: SwiGLUMLP(Phase 1 dense)/ MoEExperts(Phase 2 MoE) GLM5TextModel (nn.Module): embed + layers + norm + rotary_emb GLM5ForCausalLM (nn.Module): model + lm_head + _tp_plan + _cp_modules
def forward(self, input_ids, labels=None, position_ids=None, attention_mask=None, **kwargs): # Phase 1: embed → dense layers (GQA + SwiGLU) → norm → lm_head → loss # Phase 2+: MoE routing 在部分层中替代 SwiGLU # Phase 3+: MLA 替代 GQA # Phase 4+: DSA 稀疏 mask 叠加到 attention_mask return {"loss": loss, "logits": logits}
目标:满足全部验收标准。使用简化 dense 架构(全层 GQA + SwiGLU + RMSNorm),验证注册机制、Trainer 对接、checkpoint 闭环、精度对齐。
model.py
GLM5RMSNorm
GLM5Decoder
GLM5TextModel
GLM5ForCausalLM
__init__.py
parallelize.py
checkpoint.py
state_dict.py
examples/glm5_dense/train.yaml
train.yaml
验证(直接对应验收标准):
__post_init__
examples/glm5/train.yaml
目标:将 dense MLP 替换为 MoE(前 num_dense_layers 层 dense + 后续 MoE 层),支持 EP。
num_dense_layers
layer_type
layer_types
num_experts=256
topk=8
MoEExperts
_ep_modules = ["*.experts"]
验证(Ascend A2,EP=2):
目标:用 MLA 替代 Phase 1 的 GQA。MLA 使用 576 维 KV 压缩潜变量替代标准 KV Cache,显存降低 ~75%。
MLA
attention.py
attn_type
generate/kv_cache.py
验证:
目标:集成 DSA,支持 200K 长上下文高效推理。复用 core/context_parallel/ 中已有的 DSAIndexerContextParallel / DSASparseAttentionContextParallel。
dsa.py
验证(Ascend A2,CP=2):
目标:实现 3 层 MTP 参数共享,支持推测解码。
mtp.py
from hyper_parallel.models.glm5 import GLM5Config, GLM5ForCausalLM # GLM5Config 完整参数集(Phase 1 仅使用基础参数) # 默认值:vocab_size=151936, hidden_size=1024, num_hidden_layers=24, # num_attention_heads=16, num_key_value_heads=4, head_dim=64 config = GLM5Config(num_hidden_layers=4) # 小模型快速验证 model = GLM5ForCausalLM(config) output = model(input_ids, labels=labels) # {"loss": ..., "logits": ...}
# examples/glm5_dense/train.yaml model: name: glm5 weights_path: null tokenizer_path: null config_overrides: num_hidden_layers: 4 data: type: preset_pt train_path: /path/to/preset_batches.pt max_seq_len: 64 train: max_steps: 100 global_batch_size: 4 micro_batch_size: 1 seed: 1234 backend: torch init_device: meta accelerator: dp_shard: 2 comm_fusion: true optimizer: type: adamw lr: 1.0e-4 loss_aggregation: rank_average mixed_precision: enabled: true param_dtype: bfloat16 reduce_dtype: float32 gradient_checkpointing: activation_checkpoint: full checkpoint: output_dir: outputs/glm5 save_steps: 50 debug: deterministic: true
python scripts/train_lm.py --config examples/glm5_dense/train.yaml
MoE/MLA/DSA/MTP 通过 config_overrides 和新增 YAML 字段逐 Phase 启用:
config_overrides
model: name: glm5 config_overrides: num_experts: 256 num_experts_per_tok: 8 num_dense_layers: 3 kv_lora_rank: 576 # Phase 3 MLA dsa_topk: 2048 # Phase 4 DSA num_mtp_layers: 3 # Phase 5 MTP
dsa_context_parallel.py
kv_lora_rank=576
num_attention_heads
num_key_value_heads
tp_size
tests/torch/integration/llamafactory/
hyper_parallel/models/qwen3_5/
hyper_parallel/models/qwen3_5_moe/
hyper_parallel/core/context_parallel/dsa_context_parallel.py
hyper_parallel/models/modules/moe.py
hyper_parallel/models/spec/
hyper_parallel/trainer/base.py
RFC: HyperParallel Trainer 新增 GLM-5 系列模型 #2099
需求背景 & 价值
任务地址:https://gitcode.com/mindspore/community/issues/2099
GLM-5 是智谱 AI 于 2026 年 2 月发布的最新 decoder-only 稀疏大语言模型(arXiv:2602.15763),在中文 NLP 生态中占据核心地位。当前 HyperParallel 仅支持 Qwen 系列模型,缺少 GLM 系模型覆盖。
验收标准
tests/torch/integration/glm5/可运行examples/glm5/提供 YAML 配置模板交付范围
本次交付聚焦训练基础设施接入(Phase 1),使用简化 dense 架构验证注册机制、Trainer 对接、checkpoint 闭环、精度对齐。完整 GLM-5 架构(MoE + MLA + DSA + MTP)在 Phase 2–5 逐步叠加。
GLM-5 完整架构(供参考):
核心价值:
core/context_parallel/)和 MoE 模块(modules/moe.py),为后续完整架构接入铺垫功能描述
1. 模型架构
GLM-5 采用 MLA + DSA + MoE + MTP 四合一架构:
weight * normed2. 分阶段交付策略
GLM-5 完整架构复杂度较高,采用分层交付策略——每个 Phase 交付一组可独立训练/验证的组件:
Phase 1 为最小可用交付:简化架构(dense + GQA + SwiGLU + RMSNorm)即可跑通训练闭环,同时验证注册机制、Trainer 对接、checkpoint 保存恢复。后续 Phase 在此基础逐步叠加真实 GLM-5 组件。
3. 模型注册与发现
遵循现有 ModelSpec + register_spec 注册模式。
model.name: glm5触发 auto-discovery,Universal fields 由_resolve_overrides()映射到GLM5Config。4. Checkpoint 权重转换
支持两套命名方案(GLM-5 标准布局 + GLM-4 旧版布局),tie 场景自动合成 lm_head.weight。后续 Phase(MoE/MLA/DSA)需扩展键映射逻辑处理 expert 权重、MLA 潜变量参数。
5. 并行策略
checkpoint_wrapper_tp_plan声明(Phase 1 dense 层完全兼容,Phase 2+ 需适配 MoE gate/experts)modules/moe.py+core/expert_parallel/core/context_parallel/已有的DSAIndexerContextParallel/DSASparseAttentionContextParallel设计方案
1. 文件组织
2. GLM5Config(完整参数集)
@dataclass class GLM5Config: # ── 基础参数(Phase 1 使用) ── vocab_size: int = 151936 hidden_size: int = 1024 intermediate_size: int = 3072 num_hidden_layers: int = 24 # 训练用小模型;GLM-5 实际为 80 num_attention_heads: int = 16 num_key_value_heads: int = 4 # Phase 1 GQA;Phase 3 MLA 后废弃 head_dim: int = 64 # Phase 1;MLA 后为 kv_lora_rank max_position_embeddings: int = 131072 rms_norm_eps: float = 1e-6 rope_theta: float = 500000.0 tie_word_embeddings: bool = True # ── MoE 参数(Phase 2 启用) ── num_experts: int = 256 num_experts_per_tok: int = 8 num_dense_layers: int = 3 # 前 N 层为 dense,其余为 MoE moe_intermediate_size: int = 1024 # ── MLA 参数(Phase 3 启用) ── kv_lora_rank: int = 576 # KV 压缩秩 qk_rope_head_dim: int = 64 # RoPE 维度(MLA 中 Q/K 分离) v_head_dim: int = 128 # ── DSA 参数(Phase 4 启用) ── dsa_topk: int = 2048 # 每 token 选中的 Top-K 历史 token dsa_indexer_dim: int = 64 # ── MTP 参数(Phase 5 启用) ── num_mtp_layers: int = 3 # MTP 共享层数Phase 1 仅使用基础参数。MoE/MLA/DSA/MTP 参数在对应 Phase 启用。
3. Phase 1 模型类层次(简化 dense 版本)
4. 前向接口
def forward(self, input_ids, labels=None, position_ids=None, attention_mask=None, **kwargs): # Phase 1: embed → dense layers (GQA + SwiGLU) → norm → lm_head → loss # Phase 2+: MoE routing 在部分层中替代 SwiGLU # Phase 3+: MLA 替代 GQA # Phase 4+: DSA 稀疏 mask 叠加到 attention_mask return {"loss": loss, "logits": logits}实施计划
Phase 1 — Dense GQA 最小训练闭环(对应验收标准)
目标:满足全部验收标准。使用简化 dense 架构(全层 GQA + SwiGLU + RMSNorm),验证注册机制、Trainer 对接、checkpoint 闭环、精度对齐。
GLM5Configdataclass(完整参数集,含 MoE/MLA/DSA/MTP 预留)model.pyGLM5RMSNorm+GLM5Decoder(dense GQA + SwiGLU)model.pyGLM5TextModel+GLM5ForCausalLMmodel.py__init__.pyregister_spec +parallelize.pyAC/FSDP__init__.py+parallelize.pycheckpoint.py+state_dict.py(兼容 GLM-4/GLM-5 布局)checkpoint.py+state_dict.pyexamples/glm5_dense/train.yaml(参照 qwen3.5 dense 模板)train.yaml验证(直接对应验收标准):
__post_init__校验examples/glm5/train.yaml可被 HyperTrainerConfig 解析tests/torch/integration/glm5/可运行Phase 2 — MoE 架构
目标:将 dense MLP 替换为 MoE(前
num_dense_layers层 dense + 后续 MoE 层),支持 EP。GLM5Decoder支持layer_type调度(dense / moe)layer_typesnum_experts=256,topk=8modules/moe.py的MoEExperts_ep_modules = ["*.experts"]core/expert_parallel/checkpoint.py验证(Ascend A2,EP=2):
Phase 3 — MLA 注意力
目标:用 MLA 替代 Phase 1 的 GQA。MLA 使用 576 维 KV 压缩潜变量替代标准 KV Cache,显存降低 ~75%。
MLA类:Q/KV 分离投影 + RoPE 分离 + KV 压缩/解压attention.pyGLM5Decoder支持attn_type调度(gqa / mla)generate/kv_cache.pycheckpoint.py验证:
Phase 4 — DSA 稀疏注意力
目标:集成 DSA,支持 200K 长上下文高效推理。复用
core/context_parallel/中已有的DSAIndexerContextParallel/DSASparseAttentionContextParallel。dsa.pyattention.pyDSAIndexerContextParallel+ mask 构造验证(Ascend A2,CP=2):
Phase 5 — MTP 推测解码
目标:实现 3 层 MTP 参数共享,支持推测解码。
mtp.pymodel.py验证:
对外 API
模型构建(Phase 1)
from hyper_parallel.models.glm5 import GLM5Config, GLM5ForCausalLM # GLM5Config 完整参数集(Phase 1 仅使用基础参数) # 默认值:vocab_size=151936, hidden_size=1024, num_hidden_layers=24, # num_attention_heads=16, num_key_value_heads=4, head_dim=64 config = GLM5Config(num_hidden_layers=4) # 小模型快速验证 model = GLM5ForCausalLM(config) output = model(input_ids, labels=labels) # {"loss": ..., "logits": ...}YAML 训练(Phase 1)
# examples/glm5_dense/train.yaml model: name: glm5 weights_path: null tokenizer_path: null config_overrides: num_hidden_layers: 4 data: type: preset_pt train_path: /path/to/preset_batches.pt max_seq_len: 64 train: max_steps: 100 global_batch_size: 4 micro_batch_size: 1 seed: 1234 backend: torch init_device: meta accelerator: dp_shard: 2 comm_fusion: true optimizer: type: adamw lr: 1.0e-4 loss_aggregation: rank_average mixed_precision: enabled: true param_dtype: bfloat16 reduce_dtype: float32 gradient_checkpointing: activation_checkpoint: full checkpoint: output_dir: outputs/glm5 save_steps: 50 debug: deterministic: truePhase 2+ 扩展
MoE/MLA/DSA/MTP 通过
config_overrides和新增 YAML 字段逐 Phase 启用:model: name: glm5 config_overrides: num_experts: 256 num_experts_per_tok: 8 num_dense_layers: 3 kv_lora_rank: 576 # Phase 3 MLA dsa_topk: 2048 # Phase 4 DSA num_mtp_layers: 3 # Phase 5 MTP使用约束
dsa_context_parallel.py)和 MoE(modules/moe.py),Phase 2/4 直接复用kv_lora_rank=576与标准 GQA 的 KV Cache 格式不兼容,generate 模块需适配(随 Phase 3)num_attention_heads/num_key_value_heads被tp_size整除测试设计
Phase 1 单元测试(CPU)
Phase 1 Checkpoint 测试(CPU)
Phase 1 精度验证(Ascend A2)
Phase 2 分布式测试(Ascend A2)
回归测试
tests/torch/integration/llamafactory/全部通过规格 & 约束
参考
hyper_parallel/models/qwen3_5/、hyper_parallel/models/qwen3_5_moe/hyper_parallel/core/context_parallel/dsa_context_parallel.pyhyper_parallel/models/modules/moe.pyhyper_parallel/models/spec/hyper_parallel/trainer/base.py