已关闭
[Bug-Report|缺陷反馈]: MindSpeed-MM训练Qwen2.5VL-7B视频问答时出现报错 #173
weixin_57734188创建于  2025年11月21日关闭于  2025年12月15日
weixin_57734188
weixin_57734188
2025年11月21日 创建

Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.

Describe the current behavior / 问题描述 (Mandatory / 必填)

MindSpeed训练过程中长时间卡住不出现第一个epoch,数据集为长视频问答数据集,大多数视频都在3分钟-5分钟,我设置"video_fps": 1.0,"video_maxlen": 128,这里设置最大帧数是128帧。在多图场景下,已验证130张以下的多图问答数据可以正常训练,在多图场景下,已经验证过模型可以训练,AI Core一直在变化,并且最终也训练出来了。但是在视频问答的情况下,AI core长时间为0,偶尔某两张卡可能功率上升,其他时间都是沉寂状态,并且控制台也没有输出。在长时间的静默下,最终会出现报错

Environment / 环境信息 (Mandatory / 必填)

训练设备:8*910B 64G
软件版本:

torch 2.7.1
torch_npu 2.7.1
torchaudio 2.7.1
torchvision 0.22.1
transformers 4.49.0
MindSpeed v2.2.0_core_r0.12.1
MindSpeed-MM v2.2.0
CANN版本 8.3.RC1

系统版本

NAME="openEuler"
VERSION="24.03 (LTS-SP2)"
ID="openEuler"
VERSION_ID="24.03"
PRETTY_NAME="openEuler 24.03 (LTS-SP2)"
ANSI_COLOR="0;31"
Package                   Version     Editable project location
------------------------- ----------- ---------------------------------------------
absl-py                   2.3.1
accelerate                0.32.1
aiohappyeyeballs          2.6.1
aiohttp                   3.13.2
aiosignal                 1.4.0
annotated-types           0.7.0
antlr4-python3-runtime    4.9.3
anyio                     4.11.0
apex                      0.1+ascend
attrs                     25.3.0
auto_tune                 0.1.0
av                        16.0.1
beartype                  0.22.5
beautifulsoup4            4.14.2
bs4                       0.0.2
certifi                   2025.8.3
cffi                      2.0.0
charset-normalizer        3.4.3
Cython                    3.1.3
dataflow                  0.0.1
datasets                  4.4.1
decorator                 5.2.1
decord                    0.6.0       /data/gWX1447552/T_cot/decord/python
diffusers                 0.30.3
dill                      0.4.0
docstring_parser          0.17.0
einops                    0.8.1
ffmpeg-python             0.2.0
filelock                  3.20.0
frozenlist                1.8.0
fsspec                    2025.10.0
ftfy                      6.3.1
future                    1.0.0
gpytorch                  1.14.2
greenlet                  3.2.4
h11                       0.16.0
hccl                      0.1.0
hccl_parser               0.1
hf-xet                    1.2.0
httpcore                  1.0.9
httpx                     0.28.1
huggingface-hub           0.36.0
idna                      3.10
ImageIO                   2.37.2
imageio-ffmpeg            0.6.0
importlib_metadata        8.7.0
importlib_resources       6.5.2
iniconfig                 2.3.0
jaxtyping                 0.3.3
Jinja2                    3.1.6
joblib                    1.5.2
jsonargparse              4.42.0
jsonnet                   0.21.0
jsonschema                4.25.1
jsonschema-specifications 2025.9.1
linear-operator           0.6
llm_datadist              0.0.1
llm_datadist_v1           0.0.1
MarkupSafe                3.0.3
mindspeed                 0.12.1      /data/gWX1447552/T_cot/MindSpeed-MM/MindSpeed
mindspeed-mm              0.1         /data/gWX1447552/T_cot/MindSpeed-MM
mpmath                    1.3.0
msobjdump                 0.1.0
multidict                 6.7.0
multiprocess              0.70.18
networkx                  3.5
ninja                     1.13.0
numpy                     1.26.0
omegaconf                 2.3.0
op_compile_tool           0.1.0
op_gen                    0.1
op_test_frame             0.1
opc_tool                  0.1.0
opencv-python             4.12.0.88
packaging                 25.0
pandarallel               1.6.5
pandas                    2.0.3
pathlib2                  2.3.7.post1
peft                      0.7.1
pillow                    12.0.0
pip                       25.3
pluggy                    1.6.0
propcache                 0.4.1
protobuf                  3.20.0
psutil                    7.0.0
pyarrow                   22.0.0
pybind11                  3.0.1
pycparser                 2.23
pydantic                  2.12.4
pydantic_core             2.41.5
Pygments                  2.19.2
pytest                    8.4.2
pytest-mock               3.15.1
python-dateutil           2.9.0.post0
python-json-logger        4.0.0
pytz                      2025.2
PyYAML                    6.0.2
qwen-vl-utils             0.0.14
reconplogger              4.18.0
referencing               0.37.0
regex                     2025.11.3
requests                  2.32.5
rpds-py                   0.28.0
ruamel.yaml               0.18.16
ruamel.yaml.clib          0.2.14
safetensors               0.6.2
schedule_search           0.0.1
scikit-learn              1.7.2
scipy                     1.15.3
sentencepiece             0.2.1
setuptools                80.9.0
show_kernel_debug_data    0.1.0
six                       1.17.0
sniffio                   1.3.1
soupsieve                 2.8
SQLAlchemy                2.0.44
sympy                     1.14.0
te                        0.4.0
threadpoolctl             3.6.0
timm                      1.0.8
tokenizers                0.21.4
toml                      0.10.2
torch                     2.7.1
torch_npu                 2.7.1
torchaudio                2.7.1
torchvision               0.22.1
tqdm                      4.67.1
transformers              4.49.0
typeshed_client           2.8.2
typing_extensions         4.15.0
typing-inspection         0.4.2
tzdata                    2025.2
urllib3                   2.5.0
wadler_lindig             0.1.7
wcwidth                   0.2.14
wheel                     0.46.1
xxhash                    3.6.0
yarl                      1.22.0
zipp                      3.23.0

Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)

.

Describe the expected behavior / 预期结果 (Mandatory / 必填)

模型可以正常训练,AI core不长时间为0

以下为日志,长时间卡在Loading extension module npu_matmul_add_fp32...的控制台输出下,没有任何反应,不会输出第一个epoch信息,当长时间运行的情况下,会出现报错.

Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module npu_matmul_add_fp32...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Emitting ninja build file /root/.cache/torch_extensions/py311_cpu/npu_matmul_add_fp32/build.ninja...
Building extension module npu_matmul_add_fp32...
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module npu_matmul_add_fp32...
Loading extension module npu_matmul_add_fp32...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Emitting ninja build file /root/.cache/torch_extensions/py311_cpu/npu_matmul_add_fp32/build.ninja...
Building extension module npu_matmul_add_fp32...
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module npu_matmul_add_fp32...
Loading extension module npu_matmul_add_fp32...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Emitting ninja build file /root/.cache/torch_extensions/py311_cpu/npu_matmul_add_fp32/build.ninja...
Building extension module npu_matmul_add_fp32...
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module npu_matmul_add_fp32...
Loading extension module npu_matmul_add_fp32...

长时间没有反应的情况下出现的报错

[rank6]:[E1121 05:42:35.405473440 compiler_depend.ts:722] [Rank 3] Watchdog caught collective operation timeout: WorkHCCL(SeqNum=0, OpType=RECV, NumelIn=3, NumelOut=3, Timeout(ms)=1200000) ran for 1200778 milliseconds before timing out.
[rank6]:[E1121 05:42:35.405543310 compiler_depend.ts:798] Some HCCL operations have failed or timed out. Due to the asynchronous nature of ASCEND kernels, subsequent NPU operations might run on corrupted/incomplete data.
[rank6]:[E1121 05:42:35.405552320 compiler_depend.ts:804] To avoid data inconsistency, we are taking the entire process down.
[rank6]:[E1121 05:42:35.405683200 compiler_depend.ts:1675] [Rank 3] HCCL watchdog thread terminated with exception: [Rank 3] Watchdog caught collective operation timeout: WorkHCCL(SeqNum=0, OpType=RECV, NumelIn=3, NumelOut=3, Timeout(ms)=1200000) ran for 1200778 milliseconds before timing out.
terminate called after throwing an instance of 'std::runtime_error'
  what():  [Rank 3] HCCL watchdog thread terminated with exception: [Rank 3] Watchdog caught collective operation timeout: WorkHCCL(SeqNum=0, OpType=RECV, NumelIn=3, NumelOut=3, Timeout(ms)=1200000) ran for 1200778 milliseconds before timing out.
[rank7]:[E1121 05:42:35.459500480 compiler_depend.ts:722] [Rank 3] Watchdog caught collective operation timeout: WorkHCCL(SeqNum=0, OpType=RECV, NumelIn=3, NumelOut=3, Timeout(ms)=1200000) ran for 1200832 milliseconds before timing out.
[rank7]:[E1121 05:42:35.459592260 compiler_depend.ts:798] Some HCCL operations have failed or timed out. Due to the asynchronous nature of ASCEND kernels, subsequent NPU operations might run on corrupted/incomplete data.
[rank7]:[E1121 05:42:35.459601600 compiler_depend.ts:804] To avoid data inconsistency, we are taking the entire process down.
[rank7]:[E1121 05:42:35.459771300 compiler_depend.ts:1675] [Rank 3] HCCL watchdog thread terminated with exception: [Rank 3] Watchdog caught collective operation timeout: WorkHCCL(SeqNum=0, OpType=RECV, NumelIn=3, NumelOut=3, Timeout(ms)=1200000) ran for 1200832 milliseconds before timing out.
terminate called after throwing an instance of 'std::runtime_error'
  what():  [Rank 3] HCCL watchdog thread terminated with exception: [Rank 3] Watchdog caught collective operation timeout: WorkHCCL(SeqNum=0, OpType=RECV, NumelIn=3, NumelOut=3, Timeout(ms)=1200000) ran for 1200832 milliseconds before timing out.
`Qwen2VLImageProcessor` works only with image inputs and doesn't process videos anymore. This is a deprecated behavior and will be removed in v5.0. Your videos should be forwarded to `Qwen2VLVideoProcessor`.
W1121 05:42:43.938000 193866 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 193935 closing signal SIGTERM
W1121 05:42:43.939000 193866 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 193936 closing signal SIGTERM
W1121 05:42:43.939000 193866 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 193937 closing signal SIGTERM
W1121 05:42:43.940000 193866 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 193938 closing signal SIGTERM
W1121 05:42:43.940000 193866 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 193939 closing signal SIGTERM
W1121 05:42:43.940000 193866 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 193940 closing signal SIGTERM
W1121 05:42:43.941000 193866 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 193942 closing signal SIGTERM
E1121 05:42:46.009000 193866 site-packages/torch/distributed/elastic/multiprocessing/api.py:874] failed (exitcode: -6) local_rank: 6 (pid: 193941) of binary: /usr/local/python3.11.13/bin/python3.11
Traceback (most recent call last):
  File "/usr/local/python3.11.13/bin/torchrun", line 7, in <module>
    sys.exit(main())
             ^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 357, in wrapper
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/run.py", line 901, in main
    run(args)
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/run.py", line 892, in run
    elastic_launch(
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/launcher/api.py", line 143, in __call__
    return launch_agent(self._config, self._entrypoint, list(args))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/launcher/api.py", line 277, in launch_agent
    raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
=======================================================
pretrain_vlm.py FAILED
-------------------------------------------------------
Failures:
  <NO_OTHER_FAILURES>
-------------------------------------------------------
Root Cause (first observed failure):
[0]:
  time      : 2025-11-21_05:42:43
  host      : worker-66
  rank      : 6 (local_rank: 6)
  exitcode  : -6 (pid: 193941)
  error_file: <N/A>
  traceback : Signal 6 (SIGABRT) received by PID 193941
=======================================================
[ERROR] 2025-11-21-05:42:46 (PID:193866, Device:-1, RankID:-1) ERR99999 UNKNOWN applicaiton exception
awk: cmd. line:1: BEGIN{printf "%.3f\n", 20*1000/}
awk: cmd. line:1:                                ^ syntax error
Elapsed Time Per iteration:
Average Samples per Second:

top显示的cpu信息

(base) [root@worker-66 /]# top
top - 17:11:24 up 35 days,  1:43, 11 users,  load average: 31.27, 33.14, 34.72
Tasks: 2854 total,   9 running, 2843 sleeping,   2 stopped,   0 zombie
%Cpu(s):  3.8 us,  7.5 sy,  0.0 ni, 88.4 id,  0.0 wa,  0.2 hi,  0.2 si,  0.0 st
MiB Mem : 515147.7 total,  64167.4 free,  82388.7 used, 368591.6 buff/cache
MiB Swap:   8192.0 total,   1668.3 free,   6523.7 used. 425133.1 avail Mem

    PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ COMMAND
1133276 root      20   0 8201.1g   3.1g 633644 R 104.5   0.6   7:37.74 /usr/local/python3.11.13/bin/python3.11 -u pretrain_vlm.py --use-mcore-models --tensor-model-parallel-size 2 --pipeline-model-parallel-size 4 --context-p+
1133277 root      20   0 8201.1g   3.1g 639768 R 104.5   0.6   7:42.22 /usr/local/python3.11.13/bin/python3.11 -u pretrain_vlm.py --use-mcore-models --tensor-model-parallel-size 2 --pipeline-model-parallel-size 4 --context-p+
1133279 root      20   0 8200.9g   2.9g 633584 R 104.5   0.6   6:08.85 /usr/local/python3.11.13/bin/python3.11 -u pretrain_vlm.py --use-mcore-models --tensor-model-parallel-size 2 --pipeline-model-parallel-size 4 --context-p+
1133274 root      20   0 8201.1g   3.1g 642292 R 104.1   0.6   9:24.60 /usr/local/python3.11.13/bin/python3.11 -u pretrain_vlm.py --use-mcore-models --tensor-model-parallel-size 2 --pipeline-model-parallel-size 4 --context-p+
1133275 root      20   0 8201.1g   3.0g 628804 R 104.1   0.6   8:31.67 /usr/local/python3.11.13/bin/python3.11 -u pretrain_vlm.py --use-mcore-models --tensor-model-parallel-size 2 --pipeline-model-parallel-size 4 --context-p+
1133278 root      20   0 8200.9g   2.9g 631488 R 104.1   0.6   5:55.97 /usr/local/python3.11.13/bin/python3.11 -u pretrain_vlm.py --use-mcore-models --tensor-model-parallel-size 2 --pipeline-model-parallel-size 4 --context-p+

npu-smi info

+------------------------------------------------------------------------------------------------+
| npu-smi 25.3.rc1.b099            Version: 25.3.rc1.b099                                        |
+---------------------------+---------------+----------------------------------------------------+
| NPU   Name                | Health        | Power(W)    Temp(C)           Hugepages-Usage(page)|
| Chip                      | Bus-Id        | AICore(%)   Memory-Usage(MB)  HBM-Usage(MB)        |
+===========================+===============+====================================================+
| 0     910B3               | Warning       | 103.7       38                0    / 0             |
| 0                         | 0000:C1:00.0  | 0           0    / 0          15663/ 65536         |
+===========================+===============+====================================================+
| 1     910B3               | OK            | 98.6        39                0    / 0             |
| 0                         | 0000:C2:00.0  | 0           0    / 0          15642/ 65536         |
+===========================+===============+====================================================+
| 2     910B3               | OK            | 98.4        38                0    / 0             |
| 0                         | 0000:81:00.0  | 0           0    / 0          20529/ 65536         |
+===========================+===============+====================================================+
| 3     910B3               | Warning       | 106.1       38                0    / 0             |
| 0                         | 0000:82:00.0  | 0           0    / 0          20529/ 65536         |
+===========================+===============+====================================================+
| 4     910B3               | Warning       | 93.0        43                0    / 0             |
| 0                         | 0000:01:00.0  | 0           0    / 0          17869/ 65536         |
+===========================+===============+====================================================+
| 5     910B3               | Warning       | 93.1        43                0    / 0             |
| 0                         | 0000:02:00.0  | 0           0    / 0          17869/ 65536         |
+===========================+===============+====================================================+
| 6     910B3               | Warning       | 103.5       41                0    / 0             |
| 0                         | 0000:41:00.0  | 0           0    / 0          32099/ 65536         |
+===========================+===============+====================================================+
| 7     910B3               | OK            | 97.0        41                0    / 0             |
| 0                         | 0000:42:00.0  | 0           0    / 0          32092/ 65536         |
+===========================+===============+====================================================+
+---------------------------+---------------+----------------------------------------------------+
| NPU     Chip              | Process id    | Process name             | Process memory(MB)      |
+===========================+===============+====================================================+
| 0       0                 | 1133274       | python3.11               | 12288                   |
+===========================+===============+====================================================+
| 1       0                 | 1133275       | python3.11               | 12287                   |
+===========================+===============+====================================================+
| 2       0                 | 1133276       | python3.11               | 17176                   |
+===========================+===============+====================================================+
| 3       0                 | 1133277       | python3.11               | 17176                   |
+===========================+===============+====================================================+
| 4       0                 | 1133278       | python3.11               | 14516                   |
+===========================+===============+====================================================+
| 5       0                 | 1133279       | python3.11               | 14515                   |
+===========================+===============+====================================================+
| 6       0                 | 1133280       | python3.11               | 28739                   |
+===========================+===============+====================================================+
| 7       0                 | 1133283       | python3.11               | 28739                   |
+===========================+===============+====================================================+


Special notes for this issue/备注 (Optional / 选填)

data_7b.json

{
    "dataset_param": {
        "dataset_type": "huggingface",
        "preprocess_parameters": {
            "model_name_or_path": "/data/gWX1447552/models/huggingface/Qwen2.5-VL-7B-Instruct",
            "use_fast_tokenizer": true,
            "split_special_tokens": false,
            "image_max_pixels": 307200,
            "image_min_pixels": 4096,
            "video_max_pixels": 307200,
            "video_min_pixels": 4096,
            "video_fps": 1.0,
            "video_maxlen": 128
        },
        "basic_parameters": {
            "template": "qwen2vl",
            "dataset_dir": "/data/gWX1447552/data/StreamingBench",
            "dataset": "/data/gWX1447552/data/StreamingBench/train_data/mini_train.json",
            "cache_dir": "/data/gWX1447552/data/StreamingBench/data/cache_dir",
            "train_on_prompt": false,
            "mask_history": false,
            "preprocessing_batch_size": 500,
            "preprocessing_num_workers": 16,
            "max_samples": null,
            "tool_format": null
        },
        "attr": {
            "system": "system",
            "images": null,
            "videos": "videos",
            "messages": "messages",
            "role_tag": "role",
            "content_tag": "content",
            "user_tag": "user",
            "assistant_tag": "assistant",
            "observation_tag": null,
            "function_tag": null,
            "system_tag": null
        }
    },
    "dataloader_param": {
        "dataloader_mode": "sampler",
        "drop_last": true,
        "sampler_type": "BaseRandomBatchSampler",
        "collate_param": {
            "model_name": "qwen2vl",
            "ignore_pad_token_for_loss": true
        },
        "pin_memory": true,
        "shuffle": true
    }
}

finetune_qwen2_5_vl_7b.sh

#!/bin/bash
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
# 该变量只用于规避megatron对其校验,对npu无效
export CUDA_DEVICE_MAX_CONNECTIONS=1
export ASCEND_SLOG_PRINT_TO_STDOUT=0
export ASCEND_GLOBAL_LOG_LEVEL=3
export TASK_QUEUE_ENABLE=2
export COMBINED_ENABLE=1
export CPU_AFFINITY_CONF=1
# 这两个数字修改过,记不清报错时这两个数字设置的是多少了
# export HCCL_CONNECT_TIMEOUT=1200
export HCCL_EXEC_TIMEOUT=1200
export NPU_ASD_ENABLE=0
export ASCEND_LAUNCH_BLOCKING=0
export ACLNN_CACHE_LIMIT=100000
export PYTORCH_NPU_ALLOC_CONF="expandable_segments:True"
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export ASCEND_PROCESS_LOG_PATH=1

NPUS_PER_NODE=8
MASTER_ADDR=localhost
MASTER_PORT=6000
NNODES=1
NODE_RANK=0
WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES))


MM_DATA="./examples/qwen2.5vl/data_7b_sow.json"
MM_MODEL="./examples/qwen2.5vl/model_7b_sow.json"
MM_TOOL="./mindspeed_mm/tools/tools.json"
LOAD_PATH="/data/gWX1447552/data/StreamingBench/models/MindSpeed/Qwen2.5-VL-7B-Instruct-MM"
SAVE_PATH="/data/gWX1447552/data/StreamingBench/lora_adapter/Qwen2.5VL-7B-SRC"
TP=2
PP=4
CP=1
MBS=1 
GRAD_ACC_STEP=4
DP=$(($WORLD_SIZE/$TP/$PP/$CP))
GBS=$(($MBS*$GRAD_ACC_STEP*$DP))

DISTRIBUTED_ARGS="
    --nproc_per_node $NPUS_PER_NODE \
    --nnodes $NNODES \
    --node_rank $NODE_RANK \
    --master_addr $MASTER_ADDR \
    --master_port $MASTER_PORT
"
# --fp16 \--bf16 \
# GPT_ARGS中模型相关参数具体配置在example/qwen2.5vl/model_7b.json中,训练相关参数配置在这里
GPT_ARGS="
    --use-mcore-models \
    --tensor-model-parallel-size ${TP} \
    --pipeline-model-parallel-size ${PP} \
    --context-parallel-size ${CP} \
    --context-parallel-algo ulysses_cp_algo \
    --micro-batch-size ${MBS} \
    --global-batch-size ${GBS} \
    --tokenizer-type NullTokenizer \
    --vocab-size 152064 \
    --seq-length 128000 \
    --make-vocab-size-divisible-by 1 \
    --normalization RMSNorm \
    --use-fused-rmsnorm \
    --swiglu \
    --use-fused-swiglu \
    --no-masked-softmax-fusion \
    --lr 1.0e-5 \
    --lr-decay-style cosine \
    --weight-decay 0 \
    --train-iters 50 \
    --lr-warmup-fraction 0.1 \
    --clip-grad 0.0 \
    --adam-beta1 0.9 \
    --adam-beta2 0.999 \
    --seed 42 \
    --bf16 \
    --load $LOAD_PATH \
    --use-flash-attn \
    --variable-seq-lengths \
    --use-distributed-optimizer \
    --no-load-optim \
    --no-load-rng \
    --no-save-optim \
    --no-save-rng \
    --num-workers 0 \
    --distributed-timeout-minutes 20 \
    --calculate-per-token-loss \
"

LORA_ARGS="
    --lora-r 8 \
    --lora-alpha 16 \
    --lora-dropout 0 \
    --lora-target-modules linear_qkv linear_proj linear_fc1 linear_fc2 \
"

MM_ARGS="
    --mm-data $MM_DATA \
    --mm-model $MM_MODEL \
    --mm-tool $MM_TOOL
"

OUTPUT_ARGS="
    --log-interval 1 \
    --save-interval 5 \
    --eval-interval 5 \
    --eval-iters 5 \
    --save $SAVE_PATH \
    --ckpt-format torch \
"
logfile=$(date +%Y%m%d)_$(date +%H%M%S)
mkdir -p logs
torchrun $DISTRIBUTED_ARGS pretrain_vlm.py \
    $GPT_ARGS \
    $MM_ARGS \
    $LORA_ARGS \
    $OUTPUT_ARGS \
    --distributed-backend nccl \
    2>&1 | tee logs/train_${logfile}.log
chmod 440 logs/train_${logfile}.log
find $SAVE_PATH -type d -exec chmod 750 {} \;
find $SAVE_PATH -type f -exec chmod 640 {} \;
STEP_TIME=`grep "elapsed time per iteration" logs/train_${logfile}.log | awk -F ':' '{print$5}' | awk -F '|' '{print$1}' | head -n 150 | tail -n 100 | awk '{sum+=$1} END {if (NR != 0) printf("%.1f",sum/NR)}'`
SAMPLES_PER_SECOND=`awk 'BEGIN{printf "%.3f\n", '${GBS}'*1000/'${STEP_TIME}'}'`
echo "Elapsed Time Per iteration: $STEP_TIME"
echo "Average Samples per Second: $SAMPLES_PER_SECOND"
LOG_TOKENS_PER_SECOND=`grep "tokens per sample" logs/train_${logfile}.log`
if [ "$LOG_TOKENS_PER_SECOND" ]; then
    AVERAGE_TOKENS=`grep "tokens per sample" logs/train_${logfile}.log | awk -F 'tokens per sample:' '{print$2}' | awk -F '|' '{print$1}' | head -n 150 | tail -n 100 | awk '{sum+=$1} END {if (NR != 0) printf("%.1f",sum/NR)}'`
    TOKENS_PER_SECOND=`awk 'BEGIN{printf "%.3f\n", '${SAMPLES_PER_SECOND}'*'${AVERAGE_TOKENS}'}'`
    echo "Consumed Tokens per Second: $TOKENS_PER_SECOND"
fi

model_7b.json

{
    "model_id": "qwen2_5vl",
    "img_context_token_id": 151656,
    "vision_start_token_id": 151652,
    "image_encoder": {
        "vision_encoder": {
            "model_id": "qwen2vit",
            "num_layers": 32,
            "hidden_size": 1280,
            "ffn_hidden_size": 3420,
            "llm_hidden_size": 3584,
            "gated_linear_unit": true,
            "bias_activation_fusion": true,
            "num_attention_heads": 16,
            "hidden_dropout": 0.0,
            "attention_dropout": 0.0,
            "in_channels": 3,
            "patch_size": 14,
            "spatial_merge_size": 2,
            "temporal_patch_size": 2,
            "layernorm_epsilon": 1e-06,
            "normalization": "RMSNorm",
            "fp16": false,
            "bf16": true,
            "params_dtype": "bf16",
            "activation_func": "silu",
            "freeze": true,
            "use_fused_rotary_pos_emb": true,
            "post_layer_norm": false,
            "pipeline_num_layers": [32,0,0,0],
            "intermediate_size": 3420,
            "tokens_per_second": 2,
            "window_attn_size": 112,
            "fullatt_block_indexes": [
                7,
                15,
                23,
                31
            ],
            "attention_mask_type": "general"
        },
        "vision_projector": {
            "model_id": "lnmlp",
            "num_layers": 1,
            "gated_linear_unit": false,
            "bias_activation_fusion": false,
            "add_bias_linear": true,
            "input_size": 1280,
            "hidden_size": 3584,
            "ffn_hidden_size": 5120,
            "activation_func": "gelu",
            "bf16": true,
            "params_dtype": "bf16",
            "freeze": true,
            "layernorm_epsilon": 1e-06,
            "normalization": "RMSNorm"
        }
    },
    "text_decoder": {
        "model_id": "qwen2_5_lm",
        "num_layers": 28,
        "pipeline_num_layers": [1,10,10,7],
        "recompute_granularity": "full",
        "recompute_method": "uniform",
        "recompute_num_layers": 1,
        "hidden_size": 3584,
        "ffn_hidden_size": 18944,
        "num_attention_heads": 28,
        "seq_length": 128000,
        "max_position_embeddings": 128000,
        "vocab_size": 152064,
        "rope_theta": 1000000.0,
        "untie_embeddings_and_output_weights": true,
        "disable_bias_linear": true,
        "attention_dropout": 0.0,
        "init_method_std": 0.01,
        "hidden_dropout": 0.0,
        "position_embedding_type": "mrope",
        "normalization": "RMSNorm",
        "activation_func": "silu",
        "use_fused_rotary_pos_emb": true,
        "attention_softmax_in_fp32": true,
        "params_dtype": "bf16",
        "bf16": true,
        "parallel_output": true,
        "group_query_attention": true,
        "num_query_groups": 4,
        "mrope_section": [16, 24, 24],
        "rope_scaling": null,
        "gated_linear_unit": true,
        "bias_activation_fusion": true,
        "layernorm_epsilon": 1e-06,
        "add_bias_linear":false,
        "add_qkv_bias": true
    },
    "text_encoder": null,
    "video_encoder": null
}

likedislike
js1234567成员
2025年11月22日 评论:

感谢您使用Mindspeed-MM,您可以先试试缩小数据集,看看是否因为数据集过大,前期加载过久

likedislike
xubin成员
2025年11月22日 评论:

你好,从报错看已经出现Loading extension module npu_matmul_add_fp32... ,根据经验,出现这个打印的时候,已经加载完数据集,下一步就是出loss。如果还没加载数据集,请尝试
js1234567所提到的排查方法。如果已经加载完数据集,可以尝试如下方法

1 . 模型减层,都减到一层,降低计算量,看是否正常出loss。减到一层,还是不正常,请通过减少视频数量缩减序列长度。
2 . 设置export ASCEND_LAUNCH_BLOCKING=1并增加环境变量export ASCEND_PROCESS_LOG_PATH= 指定plog日志落盘路径,将plog发出。现在怀疑可能是序列长度太长,算子执行时间太长,超时了。

最后,这么长的序列,你可以开启CP试试。

likedislike
js1234567成员
2025年11月22日 评论:
likedislike
weixin_57734188
weixin_57734188
2025年11月24日 评论:

感谢您使用Mindspeed-MM,您可以先试试缩小数据集,看看是否因为数据集过大,前期加载过久

@js1234567

你好,我把数据集换成了一个平均40s的视频问答数据集,然后每秒抽一帧,进行全量微调,然后就可以正常训练了,128帧的无法训练可能是序列过长?有没有什么办法解决

likedislike
yangx_sy成员
2025年11月24日 评论:

感谢您使用Mindspeed-MM,您可以先试试缩小数据集,看看是否因为数据集过大,前期加载过久

@js1234567

你好,我把数据集换成了一个平均40s的视频问答数据集,然后每秒抽一帧,进行全量微调,然后就可以正常训练了,128帧的无法训练可能是序列过长?有没有什么办法解决

@weixin_57734188

序列过长可能需要采集一下profiling结果看看优化方向了,推测可能出现了长序列fa降频问题

likedislike
weixin_57734188
weixin_57734188
2025年11月26日 评论:

你好,从报错看已经出现Loading extension module npu_matmul_add_fp32... ,根据经验,出现这个打印的时候,已经加载完数据集,下一步就是出loss。如果还没加载数据集,请尝试
js1234567所提到的排查方法。如果已经加载完数据集,可以尝试如下方法

1 . 模型减层,都减到一层,降低计算量,看是否正常出loss。减到一层,还是不正常,请通过减少视频数量缩减序列长度。
2 . 设置export ASCEND_LAUNCH_BLOCKING=1并增加环境变量export ASCEND_PROCESS_LOG_PATH= 指定plog日志落盘路径,将plog发出。现在怀疑可能是序列长度太长,算子执行时间太长,超时了。

最后,这么长的序列,你可以开启CP试试。

@MoCuishle-M

这个MindSpeed-MM更换成了master版本,当我设置export ASCEND_LAUNCH_BLOCKING=1和export ASCEND_PROCESS_LOG_PATH=1后下面为报错的日志信息,设置数据集最大帧数为120帧,TP=2,PP=4。之前尝试开启CP会报错

training ...
(min, max) time across ranks (ms):
    model-and-optimizer-setup ......................: (3027.20, 3038.69)
    train/valid/test-data-iterators-setup ..........: (52762.04, 52763.42)
[before the start of training step] datetime: 2025-11-26 07:56:12 
[2025-11-26 07:56:12] [WARNING] [745] profiler.py: Invalid parameter export_type: None, reset it to text.
[2025-11-26 07:56:12] [WARNING] [745] profiler.py: Profiler won't be using warmup, this can skew profiler results
...
/usr/local/python3.11.13/lib/python3.11/site-packages/torch/autograd/function.py:575: UserWarning: Cannot create tensor with interal format while allow_internel_format=False, tensor will be created with base format. (Triggered internally at build/CMakeFiles/torch_npu.dir/compiler_depend.ts:335.)
  return super().apply(*args, **kwargs)  # type: ignore[misc]
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Emitting ninja build file /root/.cache/torch_extensions/py311_cpu/npu_matmul_add_fp32/build.ninja...
Building extension module npu_matmul_add_fp32...
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module npu_matmul_add_fp32...
Loading extension module npu_matmul_add_fp32...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Emitting ninja build file /root/.cache/torch_extensions/py311_cpu/npu_matmul_add_fp32/build.ninja...
Building extension module npu_matmul_add_fp32...
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module npu_matmul_add_fp32...
Loading extension module npu_matmul_add_fp32...
[rank4]: Traceback (most recent call last):
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/pretrain_vlm.py", line 252, in <module>
[rank4]:     pretrain(
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 292, in pretrain
[rank4]:     iteration, num_floating_point_operations_so_far = train(
[rank4]:                                                       ^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/utils/auto_setting.py", line 113, in wrapper
[rank4]:     return step_fn(*args, **kwargs)
[rank4]:            ^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 546, in train
[rank4]:     loss_dict, skipped_iter, grad_norm, num_zeros_in_grad = train_step(
[rank4]:                                                             ^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/utils/auto_setting.py", line 126, in wrapper
[rank4]:     ret = step_fn(*args, **kwargs)
[rank4]:           ^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 753, in train_step
[rank4]:     losses_reduced = forward_backward_func(
[rank4]:                      ^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/schedules.py", line 1940, in forward_backward_pipelining_without_interleaving
[rank4]:     input_tensor = send_backward_recv_forward(
[rank4]:                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/schedules.py", line 1712, in send_backward_recv_forward
[rank4]:     input_tensor = p2p_communication.send_backward_recv_forward(
[rank4]:                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/p2p_communication.py", line 554, in send_backward_recv_forward
[rank4]:     input_tensor, _, _ = _communicate(
[rank4]:                          ^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/adaptor.py", line 77, in mindspeed_communicate
[rank4]:     return communicate_impl(
[rank4]:            ^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/communicate.py", line 205, in communicate_impl
[rank4]:     recv_prev_shape, recv_next_shape = communicate_shapes_impl(
[rank4]:                                        ^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/communicate.py", line 102, in communicate_shapes_impl
[rank4]:     reqs = p2p_func(
[rank4]:            ^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/p2p_communication.py", line 198, in _p2p_ops
[rank4]:     send_prev_req = torch.distributed.isend(
[rank4]:                     ^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/distributed_c10d.py", line 2375, in isend
[rank4]:     return group.send([tensor], group_dst, tag)
[rank4]:            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]: RuntimeError: InnerRunOpApi:build/CMakeFiles/torch_npu.dir/compiler_depend.ts:281 OPS function error: HcclSend, error code is 6
[rank4]: [ERROR] 2025-11-26-08:00:42 (PID:749, Device:4, RankID:4) ERR01100 OPS call acl api failed.
[rank4]: EI0006: [PID: 749] 2025-11-26-08:00:12.177.688 Getting socket times out. Reason: Remote Rank did not send the data in time. Please check the reason for the rank being stuck
[rank4]:         Solution: 1. Check the rank service processes with other errors or no errors in the cluster.2. If this error is reported for all NPUs, check whether the time difference between the earliest and latest errors is greater than the connect timeout interval (120s by default). If so, adjust the timeout interval by using the HCCL_CONNECT_TIMEOUT environment variable.3. Check the connectivity of the communication link between nodes. (For example, run the 'hccn_tool -i $devid -tls -g' command to check the TLS status of each NPU).
[rank4]:         TraceBack (most recent call last):
[rank4]:         Transport init error. Reason: [Create][DestLink]Create Dest error! createLink para:rank[1]-localUserrank[1]-localIpAddr[10.44.154.66], dst_rank[0]-remoteUserrank[0]-remote_ip_addr[10.44.154.66]
likedislike
weixin_57734188
weixin_57734188
2025年11月26日 评论:
[rank3]: Traceback (most recent call last):
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/pretrain_vlm.py", line 252, in <module>
[rank3]:     pretrain(
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 292, in pretrain
[rank3]:     iteration, num_floating_point_operations_so_far = train(
[rank3]:                                                       ^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/utils/auto_setting.py", line 113, in wrapper
[rank3]:     return step_fn(*args, **kwargs)
[rank3]:            ^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 546, in train
[rank3]:     loss_dict, skipped_iter, grad_norm, num_zeros_in_grad = train_step(
[rank3]:                                                             ^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/utils/auto_setting.py", line 126, in wrapper
[rank3]:     ret = step_fn(*args, **kwargs)
[rank3]:           ^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 753, in train_step
[rank3]:     losses_reduced = forward_backward_func(
[rank3]:                      ^^^^^^^^^^^^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/schedules.py", line 1912, in forward_backward_pipelining_without_interleaving
[rank3]:     output_tensor_grad = send_forward_recv_backward(
[rank3]:                          ^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/schedules.py", line 1695, in send_forward_recv_backward
[rank3]:     output_tensor_grad = p2p_communication.send_forward_recv_backward(
[rank3]:                          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/p2p_communication.py", line 529, in send_forward_recv_backward
[rank3]:     _, output_tensor_grad, _ = _communicate(
[rank3]:                                ^^^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/adaptor.py", line 77, in mindspeed_communicate
[rank3]:     return communicate_impl(
[rank3]:            ^^^^^^^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/communicate.py", line 205, in communicate_impl
[rank3]:     recv_prev_shape, recv_next_shape = communicate_shapes_impl(
[rank3]:                                        ^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/communicate.py", line 102, in communicate_shapes_impl
[rank3]:     reqs = p2p_func(
[rank3]:            ^^^^^^^^^
[rank3]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/p2p_communication.py", line 223, in _p2p_ops
[rank3]:     recv_next_req = torch.distributed.irecv(
[rank3]:                     ^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]:   File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/distributed_c10d.py", line 2420, in irecv
[rank3]:     return group.recv([tensor], group_src, tag)
[rank3]:            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank3]: RuntimeError: InnerRunOpApi:build/CMakeFiles/torch_npu.dir/compiler_depend.ts:281 OPS function error: HcclRecv, error code is 6
[rank3]: [ERROR] 2025-11-26-08:04:56 (PID:748, Device:3, RankID:3) ERR01100 OPS call acl api failed.
[rank3]: EI0009: [PID: 748] 2025-11-26-08:04:56.668.126 Transport init error. Reason: [Create][DestLink]Create Dest error! createLink para:rank[0]-localUserrank[0]-localIpAddr[10.44.154.66], dst_rank[1]-remoteUserrank[1]-remote_ip_addr[10.44.154.66]
[rank3]:         Solution: Check other NPUs that are not reporting errors, or check if there are any abnormalities on the other end's NPU.
sys:1: DeprecationWarning: builtin type swigvarlink has no __module__ attribute
sys:1: DeprecationWarning: builtin type swigvarlink has no __module__ attribute
W1126 08:05:52.697000 676 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 745 closing signal SIGTERM
W1126 08:05:52.698000 676 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 746 closing signal SIGTERM
W1126 08:05:52.699000 676 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 747 closing signal SIGTERM
W1126 08:05:52.699000 676 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 748 closing signal SIGTERM
W1126 08:05:52.700000 676 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 750 closing signal SIGTERM
W1126 08:05:52.700000 676 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 751 closing signal SIGTERM
W1126 08:05:52.701000 676 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 752 closing signal SIGTERM
Exception in thread Thread-1:
Traceback (most recent call last):
  File "/usr/local/python3.11.13/lib/python3.11/threading.py", line 1045, in _bootstrap_inner
    self.run()
  File "/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/tbe/common/repository_manager/utils/multiprocess_util.py", line 88, in run
    item = self.task_q.get()
           ^^^^^^^^^^^^^^^^^
  File "<string>", line 2, in get
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/managers.py", line 822, in _callmethod
    kind, result = conn.recv()
                   ^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 250, in recv
    buf = self._recv_bytes()
          ^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 430, in _recv_bytes
    buf = self._recv(4)
          ^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 399, in _recv
    raise EOFError
EOFError
Exception in thread Thread-1:
Traceback (most recent call last):
  File "/usr/local/python3.11.13/lib/python3.11/threading.py", line 1045, in _bootstrap_inner
    self.run()
  File "/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/tbe/common/repository_manager/utils/multiprocess_util.py", line 88, in run
    item = self.task_q.get()
           ^^^^^^^^^^^^^^^^^
  File "<string>", line 2, in get
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/managers.py", line 822, in _callmethod
    kind, result = conn.recv()
                   ^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 250, in recv
    buf = self._recv_bytes()
          ^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 430, in _recv_bytes
    buf = self._recv(4)
          ^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 399, in _recv
    raise EOFError
EOFError
likedislike
weixin_57734188
weixin_57734188
2025年11月26日 评论:
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
E1126 08:05:55.239000 676 site-packages/torch/distributed/elastic/multiprocessing/api.py:869] failed (exitcode: 1) local_rank: 4 (pid: 749) of binary: /usr/local/python3.11.13/bin/python3.11
Traceback (most recent call last):
  File "/usr/local/python3.11.13/bin/torchrun", line 7, in <module>
    sys.exit(main())
             ^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 355, in wrapper
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/run.py", line 918, in main
    run(args)
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/run.py", line 909, in run
    elastic_launch(
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/launcher/api.py", line 138, in __call__
    return launch_agent(self._config, self._entrypoint, list(args))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/launcher/api.py", line 269, in launch_agent
    raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError: 
============================================================
pretrain_vlm.py FAILED
------------------------------------------------------------
Failures:
  <NO_OTHER_FAILURES>
------------------------------------------------------------
Root Cause (first observed failure):
[0]:
  time      : 2025-11-26_08:05:52
  host      : worker-66
  rank      : 4 (local_rank: 4)
  exitcode  : 1 (pid: 749)
  error_file: <N/A>
  traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
============================================================
[ERROR] 2025-11-26-08:05:55 (PID:676, Device:-1, RankID:-1) ERR99999 UNKNOWN applicaiton exception

likedislike
weixin_57734188
weixin_57734188
2025年11月26日 评论:

这里我设置最大帧数是30的时候,可以训练了,但是训练速度很慢。是不是视频的长度和训练速度有关?我之前是40s的视频抽40帧,可以正常训练;这里是三四分钟的视频抽40帧,发现训练的非常慢,并且AI core几乎不动

Loading extension module npu_matmul_add_fp32...
/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/fp8_utils.py:11: UserWarning: Currently, it is not supported to Cast shard fp32 main params to fp8 model params
  warnings.warn("Currently, it is not supported to Cast shard fp32 main params to fp8 model params")
/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/fp8_utils.py:11: UserWarning: Currently, it is not supported to Cast shard fp32 main params to fp8 model params
  warnings.warn("Currently, it is not supported to Cast shard fp32 main params to fp8 model params")
/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/fp8_utils.py:11: UserWarning: Currently, it is not supported to Cast shard fp32 main params to fp8 model params
  warnings.warn("Currently, it is not supported to Cast shard fp32 main params to fp8 model params")
/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/fp8_utils.py:11: UserWarning: Currently, it is not supported to Cast shard fp32 main params to fp8 model params
  warnings.warn("Currently, it is not supported to Cast shard fp32 main params to fp8 model params")
/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/fp8_utils.py:11: UserWarning: Currently, it is not supported to Cast shard fp32 main params to fp8 model params
  warnings.warn("Currently, it is not supported to Cast shard fp32 main params to fp8 model params")
/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/fp8_utils.py:11: UserWarning: Currently, it is not supported to Cast shard fp32 main params to fp8 model params
  warnings.warn("Currently, it is not supported to Cast shard fp32 main params to fp8 model params")
/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/fp8_utils.py:11: UserWarning: Currently, it is not supported to Cast shard fp32 main params to fp8 model params
  warnings.warn("Currently, it is not supported to Cast shard fp32 main params to fp8 model params")
/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/fp8_utils.py:11: UserWarning: Currently, it is not supported to Cast shard fp32 main params to fp8 model params
  warnings.warn("Currently, it is not supported to Cast shard fp32 main params to fp8 model params")
 [2025-11-26 11:28:23] iteration        1/     150 | consumed samples:           10 | elapsed time per iteration (ms): 2477289.2 | learning rate: 0.000000E+00 | global batch size:    10 | loss: 1.401060E-01 | loss scale: 1.0 | grad norm: 18.441 | num zeros: 0.0 | number of skipped iterations:   0 | number of nan iterations:   0 |
[Rank 4] (after 2 iterations) memory (MB) | allocated: 20306.8779296875 | max allocated: 24916.529296875 | reserved: 25866.0 | max reserved: 25866.0
 [2025-11-26 12:03:45] iteration        2/     150 | consumed samples:           20 | elapsed time per iteration (ms): 2121580.4 | learning rate: 1.333333E-06 | global batch size:    10 | loss: 1.388947E-01 | loss scale: 1.0 | grad norm: 15.587 | num zeros: 0.0 | number of skipped iterations:   0 | number of nan iterations:   0 |
Number of parameters in transformer block in billions:  7.14
Number of parameters in embedding layers in billions: 0.55
Total number of parameters in billions: 7.69
Number of parameters in most loaded shard in billions: 1.1653
Number of parameters in other shards in billions: 0.8928
Theoretical memory footprints: weight and optimizer=20003.44 MB
[Rank 5] (after 2 iterations) memory (MB) | allocated: 20306.8779296875 | max allocated: 24916.529296875 | reserved: 25866.0 | max reserved: 25866.0
[Rank 2] (after 2 iterations) memory (MB) | allocated: 20306.8779296875 | max allocated: 25901.46044921875 | reserved: 26826.0 | max reserved: 26826.0
[Rank 7] (after 2 iterations) memory (MB) | allocated: 18983.4248046875 | max allocated: 31476.9443359375 | reserved: 35006.0 | max reserved: 35006.0
[Rank 6] (after 2 iterations) memory (MB) | allocated: 18983.4248046875 | max allocated: 31476.9443359375 | reserved: 35006.0 | max reserved: 35006.0
[Rank 3] (after 2 iterations) memory (MB) | allocated: 20306.8779296875 | max allocated: 25901.46044921875 | reserved: 26826.0 | max reserved: 26826.0
[Rank 0] (after 2 iterations) memory (MB) | allocated: 7555.14892578125 | max allocated: 12718.77294921875 | reserved: 12668.0 | max reserved: 12784.0
[Rank 1] (after 2 iterations) memory (MB) | allocated: 7555.14892578125 | max allocated: 12718.77294921875 | reserved: 12668.0 | max reserved: 12784.0
 [2025-11-26 12:25:42] iteration        3/     150 | consumed samples:           30 | elapsed time per iteration (ms): 1317859.7 | learning rate: 2.666667E-06 | global batch size:    10 | loss: 1.073673E-01 | loss scale: 1.0 | grad norm: 11.223 | num zeros: 0.0 | number of skipped iterations:   0 | number of nan iterations:   0 |
 [2025-11-26 12:37:36] iteration        4/     150 | consumed samples:           40 | elapsed time per iteration (ms): 713827.5 | learning rate: 4.000000E-06 | global batch size:    10 | loss: 8.372531E-02 | loss scale: 1.0 | grad norm: 7.456 | num zeros: 0.0 | number of skipped iterations:   0 | number of nan iterations:   0 |
 [2025-11-26 13:01:39] iteration        5/     150 | consumed samples:           50 | elapsed time per iteration (ms): 1443274.1 | learning rate: 5.333333E-06 | global batch size:    10 | loss: 5.440528E-02 | loss scale: 1.0 | grad norm: 9.984 | num zeros: 0.0 | number of skipped iterations:   0 | number of nan iterations:   0 |
 [2025-11-26 13:32:51] iteration        6/     150 | consumed samples:           60 | elapsed time per iteration (ms): 1871255.3 | learning rate: 6.666667E-06 | global batch size:    10 | loss: 1.172730E-01 | loss scale: 1.0 | grad norm: 9.300 | num zeros: 0.0 | number of skipped iterations:   0 | number of nan iterations:   0 |
W1126 13:58:49.707000 4472 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 4541 closing signal SIGTERM
W1126 13:58:49.708000 4472 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 4542 closing signal SIGTERM
W1126 13:58:49.708000 4472 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 4543 closing signal SIGTERM
W1126 13:58:49.709000 4472 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 4544 closing signal SIGTERM
W1126 13:58:49.710000 4472 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 4545 closing signal SIGTERM
W1126 13:58:49.710000 4472 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 4547 closing signal SIGTERM
W1126 13:58:49.711000 4472 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 4548 closing signal SIGTERM
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
Exception in thread Thread-1:
Traceback (most recent call last):
  File "/usr/local/python3.11.13/lib/python3.11/threading.py", line 1045, in _bootstrap_inner
    self.run()
likedislike
weixin_57734188
weixin_57734188
2025年11月27日 评论:
File "/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/tbe/common/repository_manager/utils/multiprocess_util.py", line 88, in run
    item = self.task_q.get()
           ^^^^^^^^^^^^^^^^^
  File "<string>", line 2, in get
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/managers.py", line 822, in _callmethod
    kind, result = conn.recv()
                   ^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 250, in recv
    buf = self._recv_bytes()
          ^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 430, in _recv_bytes
    buf = self._recv(4)
          ^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 399, in _recv
    raise EOFError
EOFError
Exception in thread Thread-1:
Traceback (most recent call last):
  File "/usr/local/python3.11.13/lib/python3.11/threading.py", line 1045, in _bootstrap_inner
    self.run()
  File "/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/tbe/common/repository_manager/utils/multiprocess_util.py", line 88, in run
    item = self.task_q.get()
           ^^^^^^^^^^^^^^^^^
  File "<string>", line 2, in get
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/managers.py", line 822, in _callmethod
    kind, result = conn.recv()
                   ^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 250, in recv
    buf = self._recv_bytes()
          ^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 430, in _recv_bytes
    buf = self._recv(4)
          ^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 399, in _recv
    raise EOFError
EOFError
Traceback (most recent call last):
  File "/usr/local/python3.11.13/lib/python3.11/threading.py", line 1045, in _bootstrap_inner
Exception in thread Thread-1:
Traceback (most recent call last):
  File "/usr/local/python3.11.13/lib/python3.11/threading.py", line 1045, in _bootstrap_inner
    self.run()
    self.run()
  File "/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/tbe/common/repository_manager/utils/multiprocess_util.py", line 88, in run
  File "/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/tbe/common/repository_manager/utils/multiprocess_util.py", line 88, in run
    item = self.task_q.get()
           ^^^^^^^^^^^^^^^^^
  File "<string>", line 2, in get
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/managers.py", line 822, in _callmethod
    item = self.task_q.get()
           ^^^^^^^^^^^^^^^^^
  File "<string>", line 2, in get
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/managers.py", line 822, in _callmethod
    kind, result = conn.recv()
                   ^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 250, in recv
    kind, result = conn.recv()
    buf = self._recv_bytes()
                   ^^^^^^^ ^ ^ ^ ^
     ^^  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 250, in recv
^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 430, in _recv_bytes
    buf = self._recv_bytes()
       buf = self._recv(4)
      ^^^^^^^^^^^^^^^^ ^ ^
       File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 430, in _recv_bytes
  ^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 399, in _recv
        buf = self._recv(4)raise EOFError

          ^^^^^^^EOFError^^
^^^^
  File "/usr/local/python3.11.13/lib/python3.11/multiprocessing/connection.py", line 399, in _recv
    raise EOFError
EOFError

/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
/usr/local/python3.11.13/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
E1126 13:58:52.744000 4472 site-packages/torch/distributed/elastic/multiprocessing/api.py:869] failed (exitcode: -9) local_rank: 5 (pid: 4546) of binary: /usr/local/python3.11.13/bin/python3.11
Traceback (most recent call last):
  File "/usr/local/python3.11.13/bin/torchrun", line 7, in <module>
    sys.exit(main())
             ^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 355, in wrapper
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/run.py", line 918, in main
    run(args)
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/run.py", line 909, in run
    elastic_launch(
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/launcher/api.py", line 138, in __call__
    return launch_agent(self._config, self._entrypoint, list(args))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/launcher/api.py", line 269, in launch_agent
    raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
=====================================================
pretrain_vlm.py FAILED
-----------------------------------------------------
Failures:
  <NO_OTHER_FAILURES>
-----------------------------------------------------
Root Cause (first observed failure):
[0]:
  time      : 2025-11-26_13:58:49
  host      : worker-66
  rank      : 5 (local_rank: 5)
  exitcode  : -9 (pid: 4546)
  error_file: <N/A>
  traceback : Signal 9 (SIGKILL) received by PID 4546
=====================================================
[ERROR] 2025-11-26-13:58:52 (PID:4472, Device:-1, RankID:-1) ERR99999 UNKNOWN applicaiton exception
find: '/data/gWX1447552/data/StreamingBench/models/trained/MindSpeed/Qwen2.5VL-7B-SRC': No such file or directory
find: '/data/gWX1447552/data/StreamingBench/models/trained/MindSpeed/Qwen2.5VL-7B-SRC': No such file or directory
Elapsed Time Per iteration: 1657514.4
Average Samples per Second: 0.006



likedislike
xubin成员
2025年11月27日 评论:

你好,从报错看已经出现Loading extension module npu_matmul_add_fp32... ,根据经验,出现这个打印的时候,已经加载完数据集,下一步就是出loss。如果还没加载数据集,请尝试
js1234567所提到的排查方法。如果已经加载完数据集,可以尝试如下方法

1 . 模型减层,都减到一层,降低计算量,看是否正常出loss。减到一层,还是不正常,请通过减少视频数量缩减序列长度。
2 . 设置export ASCEND_LAUNCH_BLOCKING=1并增加环境变量export ASCEND_PROCESS_LOG_PATH= 指定plog日志落盘路径,将plog发出。现在怀疑可能是序列长度太长,算子执行时间太长,超时了。

最后,这么长的序列,你可以开启CP试试。

@MoCuishle-M

这个MindSpeed-MM更换成了master版本,当我设置export ASCEND_LAUNCH_BLOCKING=1和export ASCEND_PROCESS_LOG_PATH=1后下面为报错的日志信息,设置数据集最大帧数为120帧,TP=2,PP=4。之前尝试开启CP会报错

training ...
(min, max) time across ranks (ms):
   model-and-optimizer-setup ......................: (3027.20, 3038.69)
   train/valid/test-data-iterators-setup ..........: (52762.04, 52763.42)
[before the start of training step] datetime: 2025-11-26 07:56:12 
[2025-11-26 07:56:12] [WARNING] [745] profiler.py: Invalid parameter export_type: None, reset it to text.
[2025-11-26 07:56:12] [WARNING] [745] profiler.py: Profiler won't be using warmup, this can skew profiler results
...
/usr/local/python3.11.13/lib/python3.11/site-packages/torch/autograd/function.py:575: UserWarning: Cannot create tensor with interal format while allow_internel_format=False, tensor will be created with base format. (Triggered internally at build/CMakeFiles/torch_npu.dir/compiler_depend.ts:335.)
 return super().apply(*args, **kwargs)  # type: ignore[misc]
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Emitting ninja build file /root/.cache/torch_extensions/py311_cpu/npu_matmul_add_fp32/build.ninja...
Building extension module npu_matmul_add_fp32...
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module npu_matmul_add_fp32...
Loading extension module npu_matmul_add_fp32...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Using /root/.cache/torch_extensions/py311_cpu as PyTorch extensions root...
Emitting ninja build file /root/.cache/torch_extensions/py311_cpu/npu_matmul_add_fp32/build.ninja...
Building extension module npu_matmul_add_fp32...
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module npu_matmul_add_fp32...
Loading extension module npu_matmul_add_fp32...
[rank4]: Traceback (most recent call last):
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/pretrain_vlm.py", line 252, in <module>
[rank4]:     pretrain(
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 292, in pretrain
[rank4]:     iteration, num_floating_point_operations_so_far = train(
[rank4]:                                                       ^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/utils/auto_setting.py", line 113, in wrapper
[rank4]:     return step_fn(*args, **kwargs)
[rank4]:            ^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 546, in train
[rank4]:     loss_dict, skipped_iter, grad_norm, num_zeros_in_grad = train_step(
[rank4]:                                                             ^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/utils/auto_setting.py", line 126, in wrapper
[rank4]:     ret = step_fn(*args, **kwargs)
[rank4]:           ^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/mindspeed_mm/training.py", line 753, in train_step
[rank4]:     losses_reduced = forward_backward_func(
[rank4]:                      ^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/schedules.py", line 1940, in forward_backward_pipelining_without_interleaving
[rank4]:     input_tensor = send_backward_recv_forward(
[rank4]:                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/schedules.py", line 1712, in send_backward_recv_forward
[rank4]:     input_tensor = p2p_communication.send_backward_recv_forward(
[rank4]:                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/p2p_communication.py", line 554, in send_backward_recv_forward
[rank4]:     input_tensor, _, _ = _communicate(
[rank4]:                          ^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/adaptor.py", line 77, in mindspeed_communicate
[rank4]:     return communicate_impl(
[rank4]:            ^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/communicate.py", line 205, in communicate_impl
[rank4]:     recv_prev_shape, recv_next_shape = communicate_shapes_impl(
[rank4]:                                        ^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/MindSpeed/mindspeed/core/pipeline_parallel/variable_seq_length/communicate.py", line 102, in communicate_shapes_impl
[rank4]:     reqs = p2p_func(
[rank4]:            ^^^^^^^^^
[rank4]:   File "/data/gWX1447552/ACL/SFT/MindSpeed-MM/megatron/core/pipeline_parallel/p2p_communication.py", line 198, in _p2p_ops
[rank4]:     send_prev_req = torch.distributed.isend(
[rank4]:                     ^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]:   File "/usr/local/python3.11.13/lib/python3.11/site-packages/torch/distributed/distributed_c10d.py", line 2375, in isend
[rank4]:     return group.send([tensor], group_dst, tag)
[rank4]:            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank4]: RuntimeError: InnerRunOpApi:build/CMakeFiles/torch_npu.dir/compiler_depend.ts:281 OPS function error: HcclSend, error code is 6
[rank4]: [ERROR] 2025-11-26-08:00:42 (PID:749, Device:4, RankID:4) ERR01100 OPS call acl api failed.
[rank4]: EI0006: [PID: 749] 2025-11-26-08:00:12.177.688 Getting socket times out. Reason: Remote Rank did not send the data in time. Please check the reason for the rank being stuck
[rank4]:         Solution: 1. Check the rank service processes with other errors or no errors in the cluster.2. If this error is reported for all NPUs, check whether the time difference between the earliest and latest errors is greater than the connect timeout interval (120s by default). If so, adjust the timeout interval by using the HCCL_CONNECT_TIMEOUT environment variable.3. Check the connectivity of the communication link between nodes. (For example, run the 'hccn_tool -i $devid -tls -g' command to check the TLS status of each NPU).
[rank4]:         TraceBack (most recent call last):
[rank4]:         Transport init error. Reason: [Create][DestLink]Create Dest error! createLink para:rank[1]-localUserrank[1]-localIpAddr[10.44.154.66], dst_rank[0]-remoteUserrank[0]-remote_ip_addr[10.44.154.66]

@weixin_57734188
看了上下文,是一台8卡 910B3在跑那么长的序列吗?序列长度太长,计算量很多,等待超时了。调大HCCL_CONNECT_TIMEOUT 试试(不过,一步时间太长了,真能训练吗.......),这是这个环境变量的说明文档:https://www.hiascend.com/document/detail/zh/canncommercial/83RC1/maintenref/envvar/envref_07_0077.html 。 如果能有更多的硬件资源,请参考该模型的readme https://gitcode.com/Ascend/MindSpeed-MM/blob/master/examples/qwen2.5vl/README.md#长序列支持

likedislike
hhhzhuyizhi成员
2025年12月15日 评论:

您好,此issue已经超过一周没有更新,现先将issue关闭,有需要可以重新打开,谢谢!

likedislike
Hhhhzhuyizhi成员
2025年12月15日 issue状态由 TODO 改变为 WIP
Hhhhzhuyizhi成员
2025年12月15日 issue状态由 WIP 改变为 DONE
Hhhhzhuyizhi成员
2025年12月15日 关闭了 issue