已关闭
[Bug] nightly SP replay 与 MXFP4 GMM 回归(!754 / !699) #405
lutean创建于 26 天前关闭于 17 天前
kai1949
26 天前 评论:
26 天前 评论:
/label add triaged
👋 您好,欢迎向 MindStudio-Modeling 提交 Issue!
我们已收到您的反馈,感谢你对开源社区的支持。🎉
📅处理时效: 维护团队将在8小时内 查看并回复您的问题(工作日)。
🔍自助查询: 在等待期间,建议您先查阅以下资料,可能已有解决方案:
📖 MindStudio-Modeling官方文档
📝 贡献者指南
请确保 Issue 描述清晰,包含复现步骤和日志,这将帮助我们更快定位问题。谢谢!


26 天前 添加了label:triaged
jiangruitao
26 天前 评论:
26 天前 评论:
/label add resolved


ascend-robot
26 天前 评论:
26 天前 评论:
26 天前 添加了label:resolved
ascend-robot
19 天前 评论:
19 天前 评论:
您好,当前Issue标记为resolved且有一段时间未进一步更新,因此我们将其标记为'stale'(闲置)状态。若您认为这是误操作,可通过添加任意评论来去除'stale'标签。标记为stale的Issue在4天内无更新活动将自动关闭。


19 天前 添加了label:stale
问题概述
当前
master在 nightly 中稳定复现两条与 !754 / !699 相关的 TensorCast 回归:all_reduce。grouped_matmul_mxfp4_swiglu。本 Issue 仅覆盖上述两条 SP/MXFP4 问题;DeepSeek PD ratio compile 图拓扑乱序(
pow_2在消费者之后定义)是独立问题,不在本 Issue 范围内。环境
.venvHF_ENDPOINT=https://hf-mirror.comAscend/msmodeling/master,commitf34b7ce复现
Case A:SP residual leftover all_reduce
参数为
_0 = (tp_size=2, expected_local_seq=64, disable_repetition=False),模型为Qwen/Qwen3-32B,query_len=128,开启enable_sequence_parallel=True、row word-embedding TP 和 Region Replay。实际性能表同时出现:
disable_repetition=True的_1参数通过。Case B:MXFP4 GMM expectation stale
参数为
LinearQuantType.MXFP4,模型为Qwen/Qwen3-235B-A22B,num_tokens=100、单层覆盖。测试当前断言:
实际图/运行时事件为:
W8A8、W4A8 和 FP8 对照参数通过。
根因分析
Case A
!754 引入
_internal_copy_region_v2后,Region Replay 会保留 representative region output 作为 formal graph output。SP P3 matcher 在_find_norm_after_add/_find_moe_norm_after_add中把该 bookkeeping 用户误当成数据路径用户,无法穿过region_end/copy boundary 找到后续 norm,导致 residual down-proj 的all_reduce没有被 P3 替换。Case B
!699 将
grouped_matmul + SwiGLU + dynamic MXFP4 quantization融合为grouped_matmul_mxfp4_swiglu_quant,旧测试仍断言已被替换的grouped_matmul_mxfp4_swiglu。修复建议
outputformal bookkeeping 用户,兼容_internal_copy_region和_internal_copy_region_v2,并补充带 formal output alias 的 replay 单测。grouped_matmul_mxfp4_swiglu_quant,保留基础grouped_matmul_mxfp4断言。验收标准
关联信息
tensor_cast/compilation/passes/sequence_parallel_pass.py、tests/regression/tensor_cast/test_sequence_parallel_pass.py、tests/regression/tensor_cast/test_gmm_pass.pyAI 协作