已开启
【实践文档】MindSpeed-MM体验Qwen3.5-4B模型微调 #483
JeffDing创建于  7月6日
JeffDing
JeffDing
7月6日 创建

体验介绍

本次体验主要体验MindSpeed-MM微调Qwen3.5-4B模型,主要目的体验MindSpeed-MM工具体验Qwen模型,一般来说4B大小的模型能够拉起训练,其他的模型其实主要就是模型参数需要修改不一致以外,其他的过程其实是一样的,会一个大小的其他大小的同类模型一样可以拉起。

环境信息

环境:HiDevLabhttps://hidevlab.huawei.com/online-develop-intro
NPU:1 * Ascend910C(单卡2NPU)
CANN版本:CANN9.1.0

设置国内镜像源

pip config --user set global.index https://mirrors.huaweicloud.com/repository/pypi
pip config --user set global.index-url https://mirrors.huaweicloud.com/repository/pypi/simple
pip config --user set global.trusted-host mirrors.huaweicloud.com

安装torch

pip install torch==2.7.1 torch-npu==2.7.1post6 torchvision

安装MindSpeed-MM

git clone https://gitcode.com/Ascend/MindSpeed-MM.git
cd MindSpeed-MM
bash scripts/install.sh --msbranch master && bash examples/qwen3_5/install_extensions.sh
pip install transformers==5.2.0 triton-ascend==3.2.0 accelerate==1.2.0

设置环境变量

source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh --cxx_abi=0

降级setuptools(非必须)

Tips: 如果setuptools版本高于82.0的话可能会报错找不到pkg_package,这时候需要降级setuptools

pip install setuptools==81.0

安装Triton-Ascend

# 注意:triton-ascend 3.2.0 及以下 Triton-Ascend 和 Triton 不能同时存在。需要先卸载社区 Triton,再安装 Triton-Ascend。
pip install triton-ascend==3.2.1 --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple

安装fla-npu以适配AscendC

git clone https://github.com/flashserve/flash-linear-attention-npu
cd flash-linear-attention-npu
git checkout c2e3d83f

安装步骤参考:https://github.com/flashserve/flash-linear-attention-npu/blob/v26.1.0/README.md

# source 实际的cann路径
source /usr/local/Ascend/cann/set_env.sh

# 编译算子 run 包,--soc 需指定为当前机器芯片类型 {ascend910b/ascend910_93/ascend950}
bash build.sh --soc=ascend910b --pkg --vendor_name=fla_npu
bash build_out/fla-npu-*.run
cd torch_custom/fla_npu/
bash build.sh

注意:请确保操作系统已安装 gawk,否则后续安装会失败。可参考以下命令安装

# Ubuntu / Debian
apt-get update
apt-get install gawk

# openEuler / CentOS / RHEL
yum update
yum install gawk

检测

pip list | grep fla_npu

数据集准备及处理

数据集介绍

COCO2017 是微软推出的通用场景计算机视觉标准数据集,包含 118287 张训练、5000 张验证、40670 张测试图片,拥有 80 类日常物体,配套目标检测框、实例分割掩码、人体关键点、图像描述等多维度标注,支持目标检测、实例分割、姿态估计、图文生成等多种任务,采用多 IoU 平均 AP 作为评测指标,场景复杂、遮挡丰富,是当下深度学习视觉算法训练与对比最常用的基准数据集。
官网:https://cocodataset.org/

安装modelscope

pip install modelscope

下载数据集

modelscope download --dataset 'PAI/COCO2017' --local_dir '/workspace/dataset/COCO2017'
modelscope download --dataset 'AI-ModelScope/LLaVA-Instruct-150K' --local_dir '/workspace/dataset/LLaVA-Instruct-150K'

cd COCO2017
unzip train2017.zip
unzip val2017.zip
unzip annotations_trainval2017.zip 
rm -rf *.zip

数据格式转换

cd /workspace/MindSpeed-MM/
python mindspeed_mm/fsdp/tools/data_tool/llava_instruct_2_mllm_demo_format.py \
    --coco_path /workspace/dataset/COCO2017 \
    --llava_json_path /workspace/dataset/LLaVA-Instruct-150K/llava_instruct_150k.json \
    --output_json_path /workspace/dataset/COCO2017/mllm_format_llava_instruct_data.json

Qwen3.5-4B体验

权重处理

权重介绍

Qwen3.5-4B 是阿里云通义千问团队推出的40 亿参数开源原生多模态轻量大模型,采用高效混合架构,原生支持图文、视频理解,内置 262K 超长上下文窗口并可扩展至百万 token,覆盖 201 种语言;它兼顾推理、代码、文档解析能力,量化后仅需 3GB 左右显存即可运行,在同等参数量模型中综合性能突出,采用 Apache2.0 开源协议,适合本地端侧部署、轻量智能体开发与低成本私有化场景使用。

权重下载

modelscope download --model Qwen/Qwen3.5-4B --local_dir /workspace/models/Qwen3.5-4B

权重转换(hf to dcp)

mm-convert Qwen35Converter hf_to_dcp \
--hf_dir /workspace/models/Qwen3.5-4B \
--dcp_dir /workspace/models/Qwen3.5-4B-dcp \
--tie_weight_mapping '{"lm_head.weight":"model.language_model.embed_tokens.weight"}' \
--num_workers 0

配置文件修改

修改配置文件

cd /workspace/MindSpeed-MM/examples/qwen3_5

修改数据集配置

训练开始前修改qwen3_5_4B_config.yaml

# 并行策略
parallel:
  fully_shard_parallel_size: auto
  fsdp_plan:
    apply_modules: # 如果要开prefetch的话,请不要随意修改apply_modules的顺序
      - model.visual
      - model.visual.blocks.{*}
      - model.language_model
      - model.language_model.embed_tokens
      - model.language_model.layers.{*}
      - lm_head
      - mtp
    param_dtype: bf16
    reduce_dtype: fp32
    num_to_forward_prefetch: 1
    num_to_backward_prefetch: 1
  ulysses_parallel_size: 1 # 开启 ulysses-cp 时, 请将 model 的 attn_implementation 设置为 flash_attention_2

### 数据相关配置
data:
  dataset_param:
    dataset_type: huggingface
    #数据集属性
    attr:
      images: images
      messages: messages
      role_tag: role
      content_tag: content
      user_tag: user
      assistant_tag: assistant

    # 数据预处理
    preprocess_parameters:
      model_name_or_path: &HF_MODEL_LOAD_PATH /workspace/models/Qwen3.5-4B # 替换为原始hf权重
      use_fast_tokenizer: true
      split_special_tokens: false
      image_max_pixels: 262144
      image_min_pixels: 1024
      video_max_pixels: 16384
      video_min_pixels: 0
      video_fps: 2.0
      video_maxlen: 64

    basic_parameters:
      cutoff_len: 1024
      template: qwen3_vl_nothink
      enable_thinking: false
      train_on_prompt: false
      mask_history: false
      dataset_dir: /workspace/dataset/COCO2017
      dataset: &DATASET_PATH /workspace/dataset/COCO2017/mllm_format_llava_instruct_data.json
      cache_dir: ./cache_dir/
      overwrite_cache: false
      preprocessing_batch_size: 1000
      preprocessing_num_workers: 16
      max_samples: null

  # 数据加载
  dataloader_param:
    pin_memory: true
    shuffle: true
    dataloader_mode: sampler
    drop_last: true
    sampler_type: BaseRandomBatchSampler
    num_workers: 8
    collate_param:
      model_name: qwen3vl
      ignore_pad_token_for_loss: true
    enable_preload: false

# 模型配置
model:
  model_id: qwen3_5
  model_name_or_path: *HF_MODEL_LOAD_PATH
  trust_remote_code: true
  attn_implementation: flash_attention_2
  # freeze:
  #   - model.visual
  # 融合算子配置
  gdn_implementation: triton
  causal_conv1d_implementation: eager
  # skip_recompute
  skip_gdn_recompute: false
  skip_flash_attn_recompute: false

# 优化特性配置
features:
  # loss 配置
  loss_cfg:
    loss_type: default   # If you want raw loss in model, loss_type can be set to "raw".
    router_aux_loss_coef: 0.0
  # 重计算配置
  recompute: true
  recompute_plan:
      apply_modules:
        - model.visual.blocks.{*}
        - model.language_model.layers.{*}
  # chunkloss配置
  enable_chunk_loss: true
  chunkloss_plan:
    apply_module: lm_head
    chunk_size: 1024
  # activation offload 配置
  enable_activation_offload: false
  activation_offload_plan:
    apply_modules:
     - model.visual.blocks.{*}
     - model.language_model.layers.{*}
  # chunkmbs配置
  enable_chunk_mbs: false
  chunkmbs_plan:
    apply_modules:
     - model.language_model.layers.{*}
    chunk_mbs: 2 # 这个表示的是chunk之后的micro batchsize
    batch_dim: 0
    chunk_arg_indexs: [0]
    chunk_kwarg_names: ["position_embeddings", "position_ids", "rope_deltas", "attention_mask"]

# 训练配置
training:
  micro_batch_size: 1
  gradient_accumulation_steps: 1
  seed: 42
  lr: 1.0e-5
  lr_decay_style: cosine
  lr_warmup_ratio: 0.1
  weight_decay: 0
  train_iters: 1000
  clip_grad: 0.0
  init_model_with_meta_device: true
  optimizer: adamw
  adam_fused: true
  save_interval: 1000
  no_load_optim: true  # Do not load optimizer state; remove if loading is needed.
  no_load_rng: true  # Do not load RNG state; remove if loading is needed.
  no_save_optim: true  # Do not save optimizer state; remove if saving is needed.
  no_save_rng: true  # Do not save RNG state; remove if saving is needed.
  load: /workspace/models/Qwen3.5-4B-dcp  # 替换为转换后的dcp权重
  save: /workspace/models/Qwen3.5-4B-save
  use_deter_comp: false
  plugin:
    - mindspeed_mm/fsdp/models/qwen3_5
    - mindspeed_mm/fsdp/data/datasets/huggingface

# 工具配置
tools:
  profile:
    enable: false
    profile_type: static
    ranks: [0]
    static_param:
      level: level1
      with_stack: false
      with_memory: false
      record_shapes: false
      with_cpu: true
      save_path: ./profiling
      start_step: 10
      end_step: 11
      data_simplification: false
      aic_metrics_type: PipeUtilization
  memory_profile:
      enable: false
      start_step: 1
      end_step: 2
      save_path: ./memory_snapshot
      dump_ranks: [0]
      stacks: all
      max_entries: null
      mem_info: false

启动微调

修改finetune_qwen3_5_4B.sh

# 根据实际情况修改 ascend-toolkit 路径
source /usr/local/Ascend/cann/set_env.sh
export NON_MEGATRON=true
export MULTI_STREAM_MEMORY_REUSE=2
export TASK_QUEUE_ENABLE=2
export ASCEND_LAUNCH_BLOCKING=0
export ACLNN_CACHE_LIMIT=100000
export CPU_AFFINITY_CONF=1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True

# 删除triton的cache
# export TRITON_CACHE_DIR=./triton_cache
# rm -rf $TRITON_CACHE_DIR/*

NPUS_PER_NODE=2
MASTER_ADDR=localhost
MASTER_PORT=6000
NNODES=1
NODE_RANK=0
WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES))

DISTRIBUTED_ARGS="
    --nproc_per_node $NPUS_PER_NODE \
    --nnodes $NNODES \
    --node_rank $NODE_RANK \
    --master_addr $MASTER_ADDR \
    --master_port $MASTER_PORT
"

logfile=$(date +%Y%m%d)_$(date +%H%M%S)
config_path=examples/qwen3_5/qwen3_5_4B_config.yaml
mkdir -p logs
torchrun $DISTRIBUTED_ARGS mindspeed_mm/fsdp/train/trainer.py \
    ${config_path} \
    2>&1 | tee logs/train_${logfile}.log

STEP_TIME=`grep "elapsed time per iteration" logs/train_${logfile}.log | awk -F 'elapsed time per iteration [(]ms[)]:' '{print$2}' | awk -F '|' '{print$1}' | head -n 200 | tail -n 100 | awk '{sum+=$1} END {if (NR != 0) printf("%.1f",sum/NR)}'`
GBS=`grep "global batch size" logs/train_${logfile}.log | awk -F 'global batch size:' '{print$2}' | awk -F '|' '{print$1}' | head -n 1 | awk '{print $1}'`
SAMPLES_PER_SECOND=`awk 'BEGIN{printf "%.3f\n", '${GBS}'*1000/'${STEP_TIME}'}'`
echo "Elapsed Time Per iteration (ms): $STEP_TIME" | tee -a logs/train_${logfile}.log
echo "Average Samples per Second: $SAMPLES_PER_SECOND" | tee -a logs/train_${logfile}.log

启动微调

bash examples/qwen3_5/finetune_qwen3_5_4B.sh

运行截图

image.png

权重转换(dcp to hf)

mm-convert Qwen35Converter dcp_to_hf \
--save_hf_dir /workspace/models/Qwen3.5-4B-hf \
--dcp_dir /workspace/models/Qwen3.5-4B-save/iter_0001000 \
--origin_hf_dir /workspace/models/Qwen3.5-4B \
--to_bf16 false

Qwen3.6-27B体验

权重处理

权重下载

modelscope download --model Qwen/Qwen3.6-27B --local\_dir /workspace/models/Qwen3.6-27B

如果使用Hidevlab在路径:/workspace/shared\_assets/models/Qwen/Qwen3.6-27B已经预置了模型,可以直接使用

权重转换(hf_to_dcp)

mm-convert Qwen35Converter hf_to_dcp \
--hf_dir /workspace/shared_assets/models/Qwen/Qwen3.6-27B \
--dcp_dir /workspace/models/Qwen3.6-27B-dcp \
--num_workers 0

配置文件修改

修改配置文件

cd /workspace/MindSpeed-MM/examples/qwen3_6

修改数据集配置

训练开始前修改qwen3_6_27B_config.yaml

# 并行策略
parallel:
  fully_shard_parallel_size: auto
  fsdp_plan:
    apply_modules: # 如果要开prefetch的话,请不要随意修改apply_modules的顺序
      - model.visual
      - model.visual.blocks.{*}
      - model.language_model
      - model.language_model.embed_tokens
      - model.language_model.layers.{*}
      - lm_head
    param_dtype: bf16
    reduce_dtype: fp32
  ulysses_parallel_size: 1 # 开启 ulysses-cp 时, 请将 model 的 attn_implementation 设置为 flash_attention_2

### 数据相关配置
data:
  dataset_param:
    dataset_type: huggingface
    #数据集属性
    attr:
      images: images
      messages: messages
      role_tag: role
      content_tag: content
      user_tag: user
      assistant_tag: assistant

    # 数据预处理
    preprocess_parameters:
      model_name_or_path: &HF_MODEL_LOAD_PATH /workspace/shared_assets/models/Qwen/Qwen3.6-27B # 替换为原始hf权重
      use_fast_tokenizer: true
      split_special_tokens: false
      image_max_pixels: 262144
      image_min_pixels: 1024
      video_max_pixels: 16384
      video_min_pixels: 0
      video_fps: 2.0
      video_maxlen: 64

    basic_parameters:
      cutoff_len: 1024
      template: qwen3_vl_nothink
      enable_thinking: false
      train_on_prompt: false
      mask_history: false
      dataset_dir: /workspace/dataset/COCO2017/
      dataset: &DATASET_PATH /workspace/dataset/COCO2017/mllm_format_llava_instruct_data.json
      cache_dir: ./cache_dir/
      overwrite_cache: false
      preprocessing_batch_size: 1000
      preprocessing_num_workers: 16
      max_samples: null

  # 数据加载
  dataloader_param:
    pin_memory: true
    shuffle: false
    dataloader_mode: sampler
    drop_last: true
    sampler_type: BaseRandomBatchSampler
    num_workers: 8
    collate_param:
      model_name: qwen3vl
      ignore_pad_token_for_loss: true
    enable_preload: false

# 模型配置
model:
  model_id: qwen3_5
  model_name_or_path: *HF_MODEL_LOAD_PATH
  trust_remote_code: true
  attn_implementation: flash_attention_2
  freeze:
    - model.visual
  # 融合算子配置
  gdn_implementation: triton
  causal_conv1d_implementation: eager

# 优化特性配置
features:
  # loss 配置
  loss_cfg:
    loss_type: default   # If you want raw loss in model, loss_type can be set to "raw".
    router_aux_loss_coef: 0.0
  # 重计算配置
  recompute: true
  recompute_plan:
      apply_modules:
        - model.visual.blocks.{*}
        - model.language_model.layers.{*}
  # chunkloss配置
  enable_chunk_loss: true
  chunkloss_plan:
    apply_module: lm_head
    chunk_size: 1024
  # activation offload 配置
  enable_activation_offload: false
  activation_offload_plan:
    apply_modules:
     - model.visual.blocks.{*}
     - model.language_model.layers.{*}

# 训练配置
training:
  micro_batch_size: 1
  gradient_accumulation_steps: 1
  seed: 42
  lr: 1.0e-5
  lr_decay_style: cosine
  lr_warmup_ratio: 0.1
  weight_decay: 0
  train_iters: 1000
  clip_grad: 0.0
  init_model_with_meta_device: true
  optimizer: adamw
  adam_fused: true
  save_interval: 1000
  no_load_optim: true  # Do not load optimizer state; remove if loading is needed.
  no_load_rng: true  # Do not load RNG state; remove if loading is needed.
  no_save_optim: true  # Do not save optimizer state; remove if saving is needed.
  no_save_rng: true  # Do not save RNG state; remove if saving is needed.
  load: /workspace/models/Qwen3.6-27B-dcp  # 替换为转换后的dcp权重
  save: /workspace/models/Qwen3.6-27B-save
  use_deter_comp: false
  plugin:
    - mindspeed_mm/fsdp/models/qwen3_5
    - mindspeed_mm/fsdp/data/datasets/huggingface

# 工具配置
tools:
  profile:
    enable: false
    profile_type: static
    ranks: [0]
    static_param:
      level: level1
      with_stack: false
      with_memory: false
      record_shapes: false
      with_cpu: true
      save_path: ./profiling
      start_step: 10
      end_step: 11
      data_simplification: false
      aic_metrics_type: PipeUtilization
  memory_profile:
      enable: false
      start_step: 1
      end_step: 2
      save_path: ./memory_snapshot
      dump_ranks: [0]
      stacks: all
      max_entries: null
      mem_info: false

修改finetune_qwen3_6_27B.sh:

# 根据实际情况修改 ascend-toolkit 路径
source /usr/local/Ascend/cann/set_env.sh
export NON_MEGATRON=true
export MULTI_STREAM_MEMORY_REUSE=2
export TASK_QUEUE_ENABLE=2
export ASCEND_LAUNCH_BLOCKING=0
export ACLNN_CACHE_LIMIT=100000
export CPU_AFFINITY_CONF=1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True

# 删除triton的cache
# export TRITON_CACHE_DIR=./triton_cache
# rm -rf $TRITON_CACHE_DIR/*

NPUS_PER_NODE=2
MASTER_ADDR=localhost
MASTER_PORT=6000
NNODES=1
NODE_RANK=0
WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES))

DISTRIBUTED_ARGS="
    --nproc_per_node $NPUS_PER_NODE \
    --nnodes $NNODES \
    --node_rank $NODE_RANK \
    --master_addr $MASTER_ADDR \
    --master_port $MASTER_PORT
"

logfile=$(date +%Y%m%d)_$(date +%H%M%S)
config_path=examples/qwen3_6/qwen3_6_27B_config.yaml
mkdir -p logs
torchrun $DISTRIBUTED_ARGS mindspeed_mm/fsdp/train/trainer.py \
    ${config_path} \
    2>&1 | tee logs/train_${logfile}.log

STEP_TIME=`grep "elapsed time per iteration" logs/train_${logfile}.log | awk -F 'elapsed time per iteration [(]ms[)]:' '{print$2}' | awk -F '|' '{print$1}' | head -n 200 | tail -n 100 | awk '{sum+=$1} END {if (NR != 0) printf("%.1f",sum/NR)}'`
GBS=`grep "global batch size" logs/train_${logfile}.log | awk -F 'global batch size:' '{print$2}' | awk -F '|' '{print$1}' | head -n 1 | awk '{print $1}'`
SAMPLES_PER_SECOND=`awk 'BEGIN{printf "%.3f\n", '${GBS}'*1000/'${STEP_TIME}'}'`
echo "Elapsed Time Per iteration (ms): $STEP_TIME" | tee -a logs/train_${logfile}.log
echo "Average Samples per Second: $SAMPLES_PER_SECOND" | tee -a logs/train_${logfile}.log

启动微调

bash examples/qwen3_6/finetune_qwen3_6_27B.sh

HiDevLab报错

/dev/shm 空间不足问题

使用hidevlab会出现/dev/shm报错

RuntimeError: unable to write to file </torch_73197_2762417848_6>: No space left on device (28)

[rank0]: RuntimeError: DataLoader worker (pid 73193) is killed by signal: Bus error. It is possible that dataloader's workers are out of shared memory. Please try to raise your shared memory limit.

NPU显存不足报错

[rank0]: torch.OutOfMemoryError: NPU out of memory. Tried to allocate 2.37 GiB (NPU 0; 61.27 GiB total capacity; 59.25 GiB already allocated; 59.25 GiB current active; 945.48 MiB free; 59.33 GiB reserved in total by PyTorch).If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.

[rank1]: torch.OutOfMemoryError: NPU out of memory. Tried to allocate 2.37 GiB (NPU 1; 61.28 GiB total capacity; 59.46 GiB already allocated; 59.46 GiB current active; 992.16 MiB free; 59.54 GiB reserved in total by PyTorch).If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.
[ERROR] 2026-08-09-00:19:41 (PID:98893, Device:0, RankID:-1) ERR99999 UNKNOWN applicaiton exception
[ERROR] 2026-08-09-00:19:42 (PID:98894, Device:1, RankID:-1) ERR99999 UNKNOWN applicaiton exception
W0809 00:19:45.258000 98724 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 98894 closing signal SIGTERM
E0809 00:19:46.024000 98724 site-packages/torch/distributed/elastic/multiprocessing/api.py:874] failed (exitcode: 1) local_rank: 0 (pid: 98893) of binary: /usr/local/python3.12.13/bin/python3.12

原因分析

1. shm 报错根因:
容器内 /dev/shm 仅 64MB(df -h /dev/shm 确认),且无 CAP_SYS_ADMIN 无法 remount。DataLoader num_workers=8 时,worker 子进程通过 /dev/shm 传递 tensor,64MB 瞬间写满 → bus error。PyTorch 的 SharedMemory 在 Linux 上硬编码用 /dev/shm,file_system sharing strategy 也无法绕开(验证过 _share_filename_cpu_() 仍写 /dev/shm)。
2. 解决 shm:
将 yaml 中 num_workers: 8 改为 num_workers: 0,数据加载在主进程内完成,完全不使用共享内存。已验证 num_workers=0 时 /dev/shm 文件数不增加。
3. 次生 OOM:
shm 解决后,暴露出 27B 模型在 2×64GB NPU 上 FSDP 反向传播时 HBM 显存不足(每卡 ~59.5GB 已分配,再申请 2.37GB 失败)。开启 enable_activation_offload: true 仅省 ~0.2GB 不足。
4. 解决 OOM:
在 fsdp_plan 下开启 cpu_offload: true,把参数/梯度/优化器状态卸载到 CPU(本机 2TB 内存充足)。HBM 占用从 ~59GB 降到 ~26GB,OOM 消除。

修改qwen3_6_27B_config.yaml文件修复报错**

num_workers: 8 → num_workers: 0(消除 shm 报错)
enable_activation_offload: false → true(省激活显存)
fsdp_plan 新增 cpu_offload: true(参数/优化器状态 offload 到 CPU,消除 HBM OOM)

测试微调过程

image.png

NPU使用情况
image.png

训练速度分析

  1. NPU 0: 26252/65536 MB,NPU 1: 25983/65536 MB,两卡均在使用,但是HBM均没有用足,可能还存在优化空间
  2. CPU 内存 557GB/2013GB,健康
  3. 无 shm 报错,无 OOM
  4. 每 iteration ~30 秒,100 步约 50 分钟完成

微调速度优化

修改文件qwen3_6_27B_config.yaml

  1. fsdp_plan 新增 prefetch(第 17-18 行)
num_to_forward_prefetch: 1
num_to_backward_prefetch: 1

开启前向/反向预取,隐藏 cpu_offload 时 H2D 搬运延迟。

  1. enable_activation_offload: true → false(第 108 行)

关闭激活卸载。recompute 已经省了激活显存,无需再 offload 到 CPU,省掉反向时 D2H/H2D 搬运。

  1. micro_batch_size: 1 → 8(第 116 行)【核心加速项】

这是提速的关键。cpu_offload 每步的 H2D 搬运是固定开销(搬整个 FSDP shard,与 batch 无关)。batch=1 时搬运开销被 2 个样本分摊,batch=8 时被 16 个样本分摊,计算密度大幅提升。

优化并且将训练iter恢复成1000后的最终qwen3_6_27B_config.yaml

# 并行策略
parallel:
  fully_shard_parallel_size: auto
  fsdp_plan:
    apply_modules: # 如果要开prefetch的话,请不要随意修改apply_modules的顺序
      - model.visual
      - model.visual.blocks.{*}
      - model.language_model
      - model.language_model.embed_tokens
      - model.language_model.layers.{*}
      - lm_head
    param_dtype: bf16
    reduce_dtype: fp32
    # 27B 模型在 2×64GB NPU 上 FSDP 优化器状态单卡 108GB,必须 offload 到 CPU
    cpu_offload: true
    # 开启 prefetch 隐藏 H2D 搬运延迟(cpu_offload 时的主要优化手段)
    num_to_forward_prefetch: 1
    num_to_backward_prefetch: 1
  ulysses_parallel_size: 1 # 开启 ulysses-cp 时, 请将 model 的 attn_implementation 设置为 flash_attention_2

### 数据相关配置
data:
  dataset_param:
    dataset_type: huggingface
    #数据集属性
    attr:
      images: images
      messages: messages
      role_tag: role
      content_tag: content
      user_tag: user
      assistant_tag: assistant

    # 数据预处理
    preprocess_parameters:
      model_name_or_path: &HF_MODEL_LOAD_PATH /workspace/shared_assets/models/Qwen/Qwen3.6-27B # 替换为原始hf权重
      use_fast_tokenizer: true
      split_special_tokens: false
      image_max_pixels: 262144
      image_min_pixels: 1024
      video_max_pixels: 16384
      video_min_pixels: 0
      video_fps: 2.0
      video_maxlen: 64

    basic_parameters:
      cutoff_len: 1024
      template: qwen3_vl_nothink
      enable_thinking: false
      train_on_prompt: false
      mask_history: false
      dataset_dir: /workspace/dataset/COCO2017/
      dataset: &DATASET_PATH /workspace/dataset/COCO2017/mllm_format_llava_instruct_data.json
      cache_dir: ./cache_dir/
      overwrite_cache: false
      preprocessing_batch_size: 1000
      preprocessing_num_workers: 16
      max_samples: null

  # 数据加载
  dataloader_param:
    pin_memory: false
    shuffle: false
    dataloader_mode: sampler
    drop_last: true
    sampler_type: BaseRandomBatchSampler
    # num_workers 设为 0:容器内 /dev/shm 仅 64MB 且无法 remount,
    # DataLoader worker(num_workers>0) 会通过 /dev/shm 传递 tensor,
    # 64MB 不够会触发 "bus error / No space left on device"。
    # num_workers=0 时数据加载在主进程内完成,不使用共享内存,彻底规避该问题。
    num_workers: 0
    collate_param:
      model_name: qwen3vl
      ignore_pad_token_for_loss: true
    enable_preload: false

# 模型配置
model:
  model_id: qwen3_5
  model_name_or_path: *HF_MODEL_LOAD_PATH
  trust_remote_code: true
  attn_implementation: flash_attention_2
  freeze:
    - model.visual
  # 融合算子配置
  gdn_implementation: triton
  causal_conv1d_implementation: eager

# 优化特性配置
features:
  # loss 配置
  loss_cfg:
    loss_type: default   # If you want raw loss in model, loss_type can be set to "raw".
    router_aux_loss_coef: 0.0
  # 重计算配置:开启 recompute 省激活显存(关 cpu_offload 后显存紧张),
  # 反向重算比 H2D 搬运快得多,是关 offload 后的合理折中
  recompute: true
  recompute_plan:
      apply_modules:
        - model.visual.blocks.{*}
        - model.language_model.layers.{*}
  # chunkloss配置
  enable_chunk_loss: true
  chunkloss_plan:
    apply_module: lm_head
    chunk_size: 1024
  # activation offload 配置:关闭,recompute 已省激活,无需再 offload 到 CPU
  enable_activation_offload: false
  activation_offload_plan:
    apply_modules:
     - model.visual.blocks.{*}
     - model.language_model.layers.{*}

# 训练配置
training:
  micro_batch_size: 8
  gradient_accumulation_steps: 1
  seed: 42
  lr: 1.0e-5
  lr_decay_style: cosine
  lr_warmup_ratio: 0.1
  weight_decay: 0
  train_iters: 1000
  clip_grad: 0.0
  init_model_with_meta_device: true
  optimizer: adamw
  adam_fused: true
  save_interval: 1000
  no_load_optim: true  # Do not load optimizer state; remove if loading is needed.
  no_load_rng: true  # Do not load RNG state; remove if loading is needed.
  no_save_optim: true  # Do not save optimizer state; remove if saving is needed.
  no_save_rng: true  # Do not save RNG state; remove if saving is needed.
  load: /workspace/models/Qwen3.6-27B-dcp  # 替换为转换后的dcp权重
  save: /workspace/models/Qwen3.6-27B-save
  use_deter_comp: false
  plugin:
    - mindspeed_mm/fsdp/models/qwen3_5
    - mindspeed_mm/fsdp/data/datasets/huggingface

# 工具配置
tools:
  profile:
    enable: false
    profile_type: static
    ranks: [0]
    static_param:
      level: level1
      with_stack: false
      with_memory: false
      record_shapes: false
      with_cpu: true
      save_path: ./profiling
      start_step: 10
      end_step: 11
      data_simplification: false
      aic_metrics_type: PipeUtilization
  memory_profile:
      enable: false
      start_step: 1
      end_step: 2
      save_path: ./memory_snapshot
      dump_ranks: [0]
      stacks: all
      max_entries: null
      mem_info: false

优化后的NPU情况
image.png

训练完成截图
image.png

权重转换dcp_to_hf

mm-convert Qwen35Converter dcp_to_hf \
--save_hf_dir /workspace/models/Qwen3.6-27B-hf \
--dcp_dir /workspace/models/Qwen3.6-27B-save/iter_0001000 \
--origin_hf_dir /workspace/shared_assets/models/Qwen/Qwen3.6-27B \
--to_bf16 false

建议

  1. 数据集相关
    最初在尝试体验Qwen3.6-27B模型过程中,在阅读qwen3.6的文档的时候,他有一部分openai数据集的处理说明,这一部分建议标注一下如果默认COCO2017数据集可以跳过这部分设置,有些初学者或者新手会以为COCO2017的数据集也是openai格式的,然后去做了这部分处理,然后就会出现微调的时候标签数量和图片数量不一致的问题。所以建议还是建议做一个说明如果是自定义数据集并且是openai格式的,可以参考这一部分,如果使用默认COCO2017数据集的话这部分可以跳过。

  2. 在安装mindspore-mm的时候他会自动安装torch-npu,但是他安装的版本是torch-npu==2.7.1,这个包在有些环境安装的时候会遇到找不到torch-npu==2.7.1 这个包,是不是可以考虑使用torch-npu==2..7.1post6这个替代,目前我体验下来torch-npu==2..7.1post6 这个版本的包在大部分环境下都能安装成功。

  3. qwen3.5文档中的pip list | grep fla_npu 这个命令是不是错了,我这边体验的时候安装好是fla-npu这个包。

  4. 根据这个文档https://atomgit.com/Ascend/MindSpeed-MM/blob/master/examples/qwen3_5/README.md 安装MindSpeedMM的时候,有些环境会遇到下面报错:

 ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts.
mindspeed-mm 0.1 requires transformers==4.57.0, but you have transformers 5.2.0 which is incompatible.
Successfully installed huggingface-hub-1.27.0 transformers-5.2.0 typer-slim-0.24.0
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning.
Looking in indexes: https://mirrors.huaweicloud.com/repository/pypi/simple
ERROR: Could not find a version that satisfies the requirement triton-ascend==3.2.0 (from versions: none)
ERROR: No matching distribution found for triton-ascend==3.2.0

这个报错我pip list 看了下MindSpeed-MM这个插件是安装成功的,继续安装文档的下一步Triton-Ascend 是不会有问题的,但是担心开发者会搞不清以为出现了报错MindSpeed-MM并没有安装成功。所以感觉是不是可以考虑把 triton-ascend==3.2.0 安装的部分提到这个之前安装。这样感觉可能可以规避一下这个报错。

likedislike
JeffDingJeffDing
7月6日 修改标题为 “【实践文档】MindSpeed-MM体验Qwen3.5-4B模型微调”,原标题为“【实践文档】MindSpeed-MM体验Qwen”
JeffDingJeffDing
7月6日 修改了issue 的描述
JeffDingJeffDing
7月6日 修改了issue 的描述
JeffDingJeffDing
7月6日 修改了issue 的描述
ascend-robotascend-robot成员
7月7日 关联了看板:MindStudio ISSUE管理
young256young256成员
7月8日 关联了里程碑:MindSpeed 26.2.0
yeqm成员
7月14日 评论:

感谢您对 MindSpeed-MM 的体验与实践分享!您反馈的镜像使用及数据集文档说明问题我们已收到,后续将由社区 Maintainer 评估完善。

likedislike
且奏长歌
且奏长歌成员
8月4日 评论:

可以发散思考下:如何在有限资源的环境下跑通较大尺寸的模型如27B/35B等,qwen3.6和qwen3.5本质上结构相同,跑测起来没有区别;
另外:建议模块写的比较散乱,可以分点梳理清楚;以及当前qwen3.5系列适配了AscendC算子,见qwen3.5下的README说明,可以针对这部分再深入体验下,提出相关的易用性建议和使用体验,感谢您对MindSpeed-MM的关注

likedislike
JeffDing
JeffDing
8月4日 评论:

好的,那后续计划增加27B的单卡/2卡推理实践,计划在原有基础上增加一个环节。然后再增加一个环节体验qwen3.5系列适配了AscendC算子。最后全部完成后,建议这块我会再重新梳理一下,预计2周内完成。

可以发散思考下:如何在有限资源的环境下跑通较大尺寸的模型如27B/35B等,qwen3.6和qwen3.5本质上结构相同,跑测起来没有区别;
另外:建议模块写的比较散乱,可以分点梳理清楚;以及当前qwen3.5系列适配了AscendC算子,见qwen3.5下的README说明,可以针对这部分再深入体验下,提出相关的易用性建议和使用体验,感谢您对MindSpeed-MM的关注

@WendongPang

likedislike
JeffDingJeffDing
8月4日 修改了issue 的描述
此处折叠了19条事件消息 查看更多
Yyeqm成员
10 天前 issue状态由 Feedback 改变为 TODO
ascend-robot
ascend-robot成员
10 天前 评论:

您好,当前Issue标记为resolved且有一段时间未进一步更新,因此我们将其标记为'stale'(闲置)状态。若您认为这是误操作,可通过添加任意评论来去除'stale'标签。标记为stale的Issue在4天内无更新活动将自动关闭。

likedislike
ascend-robotascend-robot成员
10 天前 添加了label:stale
ascend-robotascend-robot成员
6 天前 关闭了 issue
ascend-robotascend-robot成员
6 天前 issue状态由 TODO 改变为 CLOSED
Yyeqm成员
23 小时前 issue状态由 CLOSED 改变为 TODO
Yyeqm成员
23 小时前 重新打开了 issue
Yyeqm成员
23 小时前 删除了label:stale
ascend-robot
ascend-robot成员
21 小时前 评论:

您好,当前Issue标记为resolved且有一段时间未进一步更新,因此我们将其标记为'stale'(闲置)状态。若您认为这是误操作,可通过添加任意评论来去除'stale'标签。标记为stale的Issue在4天内无更新活动将自动关闭。

likedislike
ascend-robotascend-robot成员
21 小时前 添加了label:stale