HyperParallel 门禁流水线中,MindSpore fully_shard replicate_params backward prefetch 回归用例在主体执行通过后,msrun worker 在进程退出/GC 阶段触发 native segmentation fault,导致流水线失败。
失败流水线: https://build.mindspore.cn/blue/organizations/jenkins/Hyper-parallel_Atomgit_Gate/detail/Hyper-parallel_Atomgit_Gate/4372/pipeline
失败用例: tests/mindspore/st/fully_shard/test_fully_shard_replicate_prefetch_regression.py::test_ms_fully_shard_replicate_prefetch_regression_suite
tests/mindspore/st/fully_shard/test_fully_shard_replicate_prefetch_regression.py::test_ms_fully_shard_replicate_prefetch_regression_suite
实际失败的子用例: test_ms_fully_shard_replicate_params_backward_prefetch_regression
test_ms_fully_shard_replicate_params_backward_prefetch_regression
_test_fully_shard_replicate_prefetch_regression.py::test_ms_fully_shard_replicate_params_backward_prefetch_regression ======================= 1 passed, 30 warnings in 38.07s ======================== Fatal Python error: Segmentation fault Current thread ...: Garbage-collecting <no Python frame> Worker process 24341 exit with exception. Error code: -11. RuntimeError: Distributed job exited with exception. AssertionError: List cases failed: ['test_ms_fully_shard_replicate_params_backward_prefetch_regression']
同一 suite 中 test_ms_hsdp_replicate_dtensor_state_visible_after_backward 已通过,并打印:
test_ms_hsdp_replicate_dtensor_state_visible_after_backward
rank: 0, max_logits_sum_after_backward: 112.0
test_ms_fully_shard_replicate_params_backward_prefetch_regression 也已打印 loss:
rank: 0, regression step loss: 2.7763614654541016
说明失败不是 Python 断言不满足,而是用例结束后的 MindSpore 分布式运行时清理阶段 native crash。
PR #851 临时跳过该 ST,避免当前门禁被 teardown 阶段 segfault 阻塞。待 msrun worker 退出阶段 core dump 修复后,需要恢复该用例。
parallel_run
-11
test_ms_fully_shard_replicate_prefetch_regression_suite
问题描述
HyperParallel 门禁流水线中,MindSpore fully_shard replicate_params backward prefetch 回归用例在主体执行通过后,msrun worker 在进程退出/GC 阶段触发 native segmentation fault,导致流水线失败。
失败流水线:
https://build.mindspore.cn/blue/organizations/jenkins/Hyper-parallel_Atomgit_Gate/detail/Hyper-parallel_Atomgit_Gate/4372/pipeline
失败用例:
tests/mindspore/st/fully_shard/test_fully_shard_replicate_prefetch_regression.py::test_ms_fully_shard_replicate_prefetch_regression_suite实际失败的子用例:
test_ms_fully_shard_replicate_params_backward_prefetch_regression关键日志
同一 suite 中
test_ms_hsdp_replicate_dtensor_state_visible_after_backward已通过,并打印:test_ms_fully_shard_replicate_params_backward_prefetch_regression也已打印 loss:说明失败不是 Python 断言不满足,而是用例结束后的 MindSpore 分布式运行时清理阶段 native crash。
临时规避
PR #851 临时跳过该 ST,避免当前门禁被 teardown 阶段 segfault 阻塞。待 msrun worker 退出阶段 core dump 修复后,需要恢复该用例。
期望结果
test_ms_fully_shard_replicate_params_backward_prefetch_regression执行通过后,msrun worker 正常退出。parallel_run不应因 worker-11退出而失败。test_ms_fully_shard_replicate_prefetch_regression_suite的门禁覆盖。