已开启
【实践文档】MindSpeed-MM体验Qwen3.5-4B模型微调 #483
JeffDing创建于 7月6日
7月6日 修改标题为 “【实践文档】MindSpeed-MM体验Qwen3.5-4B模型微调”,原标题为“【实践文档】MindSpeed-MM体验Qwen”
7月6日 修改了issue 的描述
7月6日 修改了issue 的描述
7月6日 修改了issue 的描述
7月7日 关联了看板:MindStudio ISSUE管理
7月8日 关联了里程碑:MindSpeed 26.2.0
且奏长歌
8月4日 评论:
8月4日 评论:
可以发散思考下:如何在有限资源的环境下跑通较大尺寸的模型如27B/35B等,qwen3.6和qwen3.5本质上结构相同,跑测起来没有区别;
另外:建议模块写的比较散乱,可以分点梳理清楚;以及当前qwen3.5系列适配了AscendC算子,见qwen3.5下的README说明,可以针对这部分再深入体验下,提出相关的易用性建议和使用体验,感谢您对MindSpeed-MM的关注


JeffDing
8月4日 评论:
8月4日 评论:
好的,那后续计划增加27B的单卡/2卡推理实践,计划在原有基础上增加一个环节。然后再增加一个环节体验qwen3.5系列适配了AscendC算子。最后全部完成后,建议这块我会再重新梳理一下,预计2周内完成。
可以发散思考下:如何在有限资源的环境下跑通较大尺寸的模型如27B/35B等,qwen3.6和qwen3.5本质上结构相同,跑测起来没有区别;
另外:建议模块写的比较散乱,可以分点梳理清楚;以及当前qwen3.5系列适配了AscendC算子,见qwen3.5下的README说明,可以针对这部分再深入体验下,提出相关的易用性建议和使用体验,感谢您对MindSpeed-MM的关注


8月4日 修改了issue 的描述
此处折叠了19条事件消息 查看更多
ascend-robot
10 天前 评论:
10 天前 评论:
您好,当前Issue标记为resolved且有一段时间未进一步更新,因此我们将其标记为'stale'(闲置)状态。若您认为这是误操作,可通过添加任意评论来去除'stale'标签。标记为stale的Issue在4天内无更新活动将自动关闭。


10 天前 添加了label:stale
6 天前 关闭了 issue
6 天前 issue状态由 TODO 改变为 CLOSED
ascend-robot
21 小时前 评论:
21 小时前 评论:
您好,当前Issue标记为resolved且有一段时间未进一步更新,因此我们将其标记为'stale'(闲置)状态。若您认为这是误操作,可通过添加任意评论来去除'stale'标签。标记为stale的Issue在4天内无更新活动将自动关闭。


21 小时前 添加了label:stale
体验介绍
本次体验主要体验MindSpeed-MM微调Qwen3.5-4B模型,主要目的体验MindSpeed-MM工具体验Qwen模型,一般来说4B大小的模型能够拉起训练,其他的模型其实主要就是模型参数需要修改不一致以外,其他的过程其实是一样的,会一个大小的其他大小的同类模型一样可以拉起。
环境信息
环境:HiDevLabhttps://hidevlab.huawei.com/online-develop-intro
NPU:1 * Ascend910C(单卡2NPU)
CANN版本:CANN9.1.0
设置国内镜像源
pip config --user set global.index https://mirrors.huaweicloud.com/repository/pypi pip config --user set global.index-url https://mirrors.huaweicloud.com/repository/pypi/simple pip config --user set global.trusted-host mirrors.huaweicloud.com安装torch
安装MindSpeed-MM
git clone https://gitcode.com/Ascend/MindSpeed-MM.git cd MindSpeed-MM bash scripts/install.sh --msbranch master && bash examples/qwen3_5/install_extensions.sh pip install transformers==5.2.0 triton-ascend==3.2.0 accelerate==1.2.0设置环境变量
source /usr/local/Ascend/ascend-toolkit/set_env.sh source /usr/local/Ascend/nnal/atb/set_env.sh --cxx_abi=0降级setuptools(非必须)
Tips: 如果setuptools版本高于82.0的话可能会报错找不到pkg_package,这时候需要降级setuptools
安装Triton-Ascend
# 注意:triton-ascend 3.2.0 及以下 Triton-Ascend 和 Triton 不能同时存在。需要先卸载社区 Triton,再安装 Triton-Ascend。 pip install triton-ascend==3.2.1 --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple安装fla-npu以适配AscendC
git clone https://github.com/flashserve/flash-linear-attention-npu cd flash-linear-attention-npu git checkout c2e3d83f安装步骤参考:https://github.com/flashserve/flash-linear-attention-npu/blob/v26.1.0/README.md
# source 实际的cann路径 source /usr/local/Ascend/cann/set_env.sh # 编译算子 run 包,--soc 需指定为当前机器芯片类型 {ascend910b/ascend910_93/ascend950} bash build.sh --soc=ascend910b --pkg --vendor_name=fla_npu bash build_out/fla-npu-*.run cd torch_custom/fla_npu/ bash build.sh注意:请确保操作系统已安装 gawk,否则后续安装会失败。可参考以下命令安装
# Ubuntu / Debian apt-get update apt-get install gawk # openEuler / CentOS / RHEL yum update yum install gawk检测
数据集准备及处理
数据集介绍
COCO2017 是微软推出的通用场景计算机视觉标准数据集,包含 118287 张训练、5000 张验证、40670 张测试图片,拥有 80 类日常物体,配套目标检测框、实例分割掩码、人体关键点、图像描述等多维度标注,支持目标检测、实例分割、姿态估计、图文生成等多种任务,采用多 IoU 平均 AP 作为评测指标,场景复杂、遮挡丰富,是当下深度学习视觉算法训练与对比最常用的基准数据集。
官网:https://cocodataset.org/
安装modelscope
下载数据集
modelscope download --dataset 'PAI/COCO2017' --local_dir '/workspace/dataset/COCO2017' modelscope download --dataset 'AI-ModelScope/LLaVA-Instruct-150K' --local_dir '/workspace/dataset/LLaVA-Instruct-150K' cd COCO2017 unzip train2017.zip unzip val2017.zip unzip annotations_trainval2017.zip rm -rf *.zip数据格式转换
cd /workspace/MindSpeed-MM/ python mindspeed_mm/fsdp/tools/data_tool/llava_instruct_2_mllm_demo_format.py \ --coco_path /workspace/dataset/COCO2017 \ --llava_json_path /workspace/dataset/LLaVA-Instruct-150K/llava_instruct_150k.json \ --output_json_path /workspace/dataset/COCO2017/mllm_format_llava_instruct_data.jsonQwen3.5-4B体验
权重处理
权重介绍
Qwen3.5-4B 是阿里云通义千问团队推出的40 亿参数开源原生多模态轻量大模型,采用高效混合架构,原生支持图文、视频理解,内置 262K 超长上下文窗口并可扩展至百万 token,覆盖 201 种语言;它兼顾推理、代码、文档解析能力,量化后仅需 3GB 左右显存即可运行,在同等参数量模型中综合性能突出,采用 Apache2.0 开源协议,适合本地端侧部署、轻量智能体开发与低成本私有化场景使用。
权重下载
权重转换(hf to dcp)
mm-convert Qwen35Converter hf_to_dcp \ --hf_dir /workspace/models/Qwen3.5-4B \ --dcp_dir /workspace/models/Qwen3.5-4B-dcp \ --tie_weight_mapping '{"lm_head.weight":"model.language_model.embed_tokens.weight"}' \ --num_workers 0配置文件修改
修改配置文件
cd /workspace/MindSpeed-MM/examples/qwen3_5修改数据集配置
训练开始前修改
qwen3_5_4B_config.yaml# 并行策略 parallel: fully_shard_parallel_size: auto fsdp_plan: apply_modules: # 如果要开prefetch的话,请不要随意修改apply_modules的顺序 - model.visual - model.visual.blocks.{*} - model.language_model - model.language_model.embed_tokens - model.language_model.layers.{*} - lm_head - mtp param_dtype: bf16 reduce_dtype: fp32 num_to_forward_prefetch: 1 num_to_backward_prefetch: 1 ulysses_parallel_size: 1 # 开启 ulysses-cp 时, 请将 model 的 attn_implementation 设置为 flash_attention_2 ### 数据相关配置 data: dataset_param: dataset_type: huggingface #数据集属性 attr: images: images messages: messages role_tag: role content_tag: content user_tag: user assistant_tag: assistant # 数据预处理 preprocess_parameters: model_name_or_path: &HF_MODEL_LOAD_PATH /workspace/models/Qwen3.5-4B # 替换为原始hf权重 use_fast_tokenizer: true split_special_tokens: false image_max_pixels: 262144 image_min_pixels: 1024 video_max_pixels: 16384 video_min_pixels: 0 video_fps: 2.0 video_maxlen: 64 basic_parameters: cutoff_len: 1024 template: qwen3_vl_nothink enable_thinking: false train_on_prompt: false mask_history: false dataset_dir: /workspace/dataset/COCO2017 dataset: &DATASET_PATH /workspace/dataset/COCO2017/mllm_format_llava_instruct_data.json cache_dir: ./cache_dir/ overwrite_cache: false preprocessing_batch_size: 1000 preprocessing_num_workers: 16 max_samples: null # 数据加载 dataloader_param: pin_memory: true shuffle: true dataloader_mode: sampler drop_last: true sampler_type: BaseRandomBatchSampler num_workers: 8 collate_param: model_name: qwen3vl ignore_pad_token_for_loss: true enable_preload: false # 模型配置 model: model_id: qwen3_5 model_name_or_path: *HF_MODEL_LOAD_PATH trust_remote_code: true attn_implementation: flash_attention_2 # freeze: # - model.visual # 融合算子配置 gdn_implementation: triton causal_conv1d_implementation: eager # skip_recompute skip_gdn_recompute: false skip_flash_attn_recompute: false # 优化特性配置 features: # loss 配置 loss_cfg: loss_type: default # If you want raw loss in model, loss_type can be set to "raw". router_aux_loss_coef: 0.0 # 重计算配置 recompute: true recompute_plan: apply_modules: - model.visual.blocks.{*} - model.language_model.layers.{*} # chunkloss配置 enable_chunk_loss: true chunkloss_plan: apply_module: lm_head chunk_size: 1024 # activation offload 配置 enable_activation_offload: false activation_offload_plan: apply_modules: - model.visual.blocks.{*} - model.language_model.layers.{*} # chunkmbs配置 enable_chunk_mbs: false chunkmbs_plan: apply_modules: - model.language_model.layers.{*} chunk_mbs: 2 # 这个表示的是chunk之后的micro batchsize batch_dim: 0 chunk_arg_indexs: [0] chunk_kwarg_names: ["position_embeddings", "position_ids", "rope_deltas", "attention_mask"] # 训练配置 training: micro_batch_size: 1 gradient_accumulation_steps: 1 seed: 42 lr: 1.0e-5 lr_decay_style: cosine lr_warmup_ratio: 0.1 weight_decay: 0 train_iters: 1000 clip_grad: 0.0 init_model_with_meta_device: true optimizer: adamw adam_fused: true save_interval: 1000 no_load_optim: true # Do not load optimizer state; remove if loading is needed. no_load_rng: true # Do not load RNG state; remove if loading is needed. no_save_optim: true # Do not save optimizer state; remove if saving is needed. no_save_rng: true # Do not save RNG state; remove if saving is needed. load: /workspace/models/Qwen3.5-4B-dcp # 替换为转换后的dcp权重 save: /workspace/models/Qwen3.5-4B-save use_deter_comp: false plugin: - mindspeed_mm/fsdp/models/qwen3_5 - mindspeed_mm/fsdp/data/datasets/huggingface # 工具配置 tools: profile: enable: false profile_type: static ranks: [0] static_param: level: level1 with_stack: false with_memory: false record_shapes: false with_cpu: true save_path: ./profiling start_step: 10 end_step: 11 data_simplification: false aic_metrics_type: PipeUtilization memory_profile: enable: false start_step: 1 end_step: 2 save_path: ./memory_snapshot dump_ranks: [0] stacks: all max_entries: null mem_info: false启动微调
修改
finetune_qwen3_5_4B.sh# 根据实际情况修改 ascend-toolkit 路径 source /usr/local/Ascend/cann/set_env.sh export NON_MEGATRON=true export MULTI_STREAM_MEMORY_REUSE=2 export TASK_QUEUE_ENABLE=2 export ASCEND_LAUNCH_BLOCKING=0 export ACLNN_CACHE_LIMIT=100000 export CPU_AFFINITY_CONF=1 export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True # 删除triton的cache # export TRITON_CACHE_DIR=./triton_cache # rm -rf $TRITON_CACHE_DIR/* NPUS_PER_NODE=2 MASTER_ADDR=localhost MASTER_PORT=6000 NNODES=1 NODE_RANK=0 WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES)) DISTRIBUTED_ARGS=" --nproc_per_node $NPUS_PER_NODE \ --nnodes $NNODES \ --node_rank $NODE_RANK \ --master_addr $MASTER_ADDR \ --master_port $MASTER_PORT " logfile=$(date +%Y%m%d)_$(date +%H%M%S) config_path=examples/qwen3_5/qwen3_5_4B_config.yaml mkdir -p logs torchrun $DISTRIBUTED_ARGS mindspeed_mm/fsdp/train/trainer.py \ ${config_path} \ 2>&1 | tee logs/train_${logfile}.log STEP_TIME=`grep "elapsed time per iteration" logs/train_${logfile}.log | awk -F 'elapsed time per iteration [(]ms[)]:' '{print$2}' | awk -F '|' '{print$1}' | head -n 200 | tail -n 100 | awk '{sum+=$1} END {if (NR != 0) printf("%.1f",sum/NR)}'` GBS=`grep "global batch size" logs/train_${logfile}.log | awk -F 'global batch size:' '{print$2}' | awk -F '|' '{print$1}' | head -n 1 | awk '{print $1}'` SAMPLES_PER_SECOND=`awk 'BEGIN{printf "%.3f\n", '${GBS}'*1000/'${STEP_TIME}'}'` echo "Elapsed Time Per iteration (ms): $STEP_TIME" | tee -a logs/train_${logfile}.log echo "Average Samples per Second: $SAMPLES_PER_SECOND" | tee -a logs/train_${logfile}.log启动微调
运行截图
权重转换(dcp to hf)
mm-convert Qwen35Converter dcp_to_hf \ --save_hf_dir /workspace/models/Qwen3.5-4B-hf \ --dcp_dir /workspace/models/Qwen3.5-4B-save/iter_0001000 \ --origin_hf_dir /workspace/models/Qwen3.5-4B \ --to_bf16 falseQwen3.6-27B体验
权重处理
权重下载
modelscope download --model Qwen/Qwen3.6-27B --local\_dir /workspace/models/Qwen3.6-27B如果使用
Hidevlab在路径:/workspace/shared\_assets/models/Qwen/Qwen3.6-27B已经预置了模型,可以直接使用权重转换(hf_to_dcp)
配置文件修改
修改配置文件
cd /workspace/MindSpeed-MM/examples/qwen3_6修改数据集配置
训练开始前修改
qwen3_6_27B_config.yaml# 并行策略 parallel: fully_shard_parallel_size: auto fsdp_plan: apply_modules: # 如果要开prefetch的话,请不要随意修改apply_modules的顺序 - model.visual - model.visual.blocks.{*} - model.language_model - model.language_model.embed_tokens - model.language_model.layers.{*} - lm_head param_dtype: bf16 reduce_dtype: fp32 ulysses_parallel_size: 1 # 开启 ulysses-cp 时, 请将 model 的 attn_implementation 设置为 flash_attention_2 ### 数据相关配置 data: dataset_param: dataset_type: huggingface #数据集属性 attr: images: images messages: messages role_tag: role content_tag: content user_tag: user assistant_tag: assistant # 数据预处理 preprocess_parameters: model_name_or_path: &HF_MODEL_LOAD_PATH /workspace/shared_assets/models/Qwen/Qwen3.6-27B # 替换为原始hf权重 use_fast_tokenizer: true split_special_tokens: false image_max_pixels: 262144 image_min_pixels: 1024 video_max_pixels: 16384 video_min_pixels: 0 video_fps: 2.0 video_maxlen: 64 basic_parameters: cutoff_len: 1024 template: qwen3_vl_nothink enable_thinking: false train_on_prompt: false mask_history: false dataset_dir: /workspace/dataset/COCO2017/ dataset: &DATASET_PATH /workspace/dataset/COCO2017/mllm_format_llava_instruct_data.json cache_dir: ./cache_dir/ overwrite_cache: false preprocessing_batch_size: 1000 preprocessing_num_workers: 16 max_samples: null # 数据加载 dataloader_param: pin_memory: true shuffle: false dataloader_mode: sampler drop_last: true sampler_type: BaseRandomBatchSampler num_workers: 8 collate_param: model_name: qwen3vl ignore_pad_token_for_loss: true enable_preload: false # 模型配置 model: model_id: qwen3_5 model_name_or_path: *HF_MODEL_LOAD_PATH trust_remote_code: true attn_implementation: flash_attention_2 freeze: - model.visual # 融合算子配置 gdn_implementation: triton causal_conv1d_implementation: eager # 优化特性配置 features: # loss 配置 loss_cfg: loss_type: default # If you want raw loss in model, loss_type can be set to "raw". router_aux_loss_coef: 0.0 # 重计算配置 recompute: true recompute_plan: apply_modules: - model.visual.blocks.{*} - model.language_model.layers.{*} # chunkloss配置 enable_chunk_loss: true chunkloss_plan: apply_module: lm_head chunk_size: 1024 # activation offload 配置 enable_activation_offload: false activation_offload_plan: apply_modules: - model.visual.blocks.{*} - model.language_model.layers.{*} # 训练配置 training: micro_batch_size: 1 gradient_accumulation_steps: 1 seed: 42 lr: 1.0e-5 lr_decay_style: cosine lr_warmup_ratio: 0.1 weight_decay: 0 train_iters: 1000 clip_grad: 0.0 init_model_with_meta_device: true optimizer: adamw adam_fused: true save_interval: 1000 no_load_optim: true # Do not load optimizer state; remove if loading is needed. no_load_rng: true # Do not load RNG state; remove if loading is needed. no_save_optim: true # Do not save optimizer state; remove if saving is needed. no_save_rng: true # Do not save RNG state; remove if saving is needed. load: /workspace/models/Qwen3.6-27B-dcp # 替换为转换后的dcp权重 save: /workspace/models/Qwen3.6-27B-save use_deter_comp: false plugin: - mindspeed_mm/fsdp/models/qwen3_5 - mindspeed_mm/fsdp/data/datasets/huggingface # 工具配置 tools: profile: enable: false profile_type: static ranks: [0] static_param: level: level1 with_stack: false with_memory: false record_shapes: false with_cpu: true save_path: ./profiling start_step: 10 end_step: 11 data_simplification: false aic_metrics_type: PipeUtilization memory_profile: enable: false start_step: 1 end_step: 2 save_path: ./memory_snapshot dump_ranks: [0] stacks: all max_entries: null mem_info: false修改
finetune_qwen3_6_27B.sh:# 根据实际情况修改 ascend-toolkit 路径 source /usr/local/Ascend/cann/set_env.sh export NON_MEGATRON=true export MULTI_STREAM_MEMORY_REUSE=2 export TASK_QUEUE_ENABLE=2 export ASCEND_LAUNCH_BLOCKING=0 export ACLNN_CACHE_LIMIT=100000 export CPU_AFFINITY_CONF=1 export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True # 删除triton的cache # export TRITON_CACHE_DIR=./triton_cache # rm -rf $TRITON_CACHE_DIR/* NPUS_PER_NODE=2 MASTER_ADDR=localhost MASTER_PORT=6000 NNODES=1 NODE_RANK=0 WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES)) DISTRIBUTED_ARGS=" --nproc_per_node $NPUS_PER_NODE \ --nnodes $NNODES \ --node_rank $NODE_RANK \ --master_addr $MASTER_ADDR \ --master_port $MASTER_PORT " logfile=$(date +%Y%m%d)_$(date +%H%M%S) config_path=examples/qwen3_6/qwen3_6_27B_config.yaml mkdir -p logs torchrun $DISTRIBUTED_ARGS mindspeed_mm/fsdp/train/trainer.py \ ${config_path} \ 2>&1 | tee logs/train_${logfile}.log STEP_TIME=`grep "elapsed time per iteration" logs/train_${logfile}.log | awk -F 'elapsed time per iteration [(]ms[)]:' '{print$2}' | awk -F '|' '{print$1}' | head -n 200 | tail -n 100 | awk '{sum+=$1} END {if (NR != 0) printf("%.1f",sum/NR)}'` GBS=`grep "global batch size" logs/train_${logfile}.log | awk -F 'global batch size:' '{print$2}' | awk -F '|' '{print$1}' | head -n 1 | awk '{print $1}'` SAMPLES_PER_SECOND=`awk 'BEGIN{printf "%.3f\n", '${GBS}'*1000/'${STEP_TIME}'}'` echo "Elapsed Time Per iteration (ms): $STEP_TIME" | tee -a logs/train_${logfile}.log echo "Average Samples per Second: $SAMPLES_PER_SECOND" | tee -a logs/train_${logfile}.log启动微调
HiDevLab报错
/dev/shm 空间不足问题
使用
hidevlab会出现/dev/shm报错RuntimeError: unable to write to file </torch_73197_2762417848_6>: No space left on device (28) [rank0]: RuntimeError: DataLoader worker (pid 73193) is killed by signal: Bus error. It is possible that dataloader's workers are out of shared memory. Please try to raise your shared memory limit.NPU显存不足报错
[rank0]: torch.OutOfMemoryError: NPU out of memory. Tried to allocate 2.37 GiB (NPU 0; 61.27 GiB total capacity; 59.25 GiB already allocated; 59.25 GiB current active; 945.48 MiB free; 59.33 GiB reserved in total by PyTorch).If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. [rank1]: torch.OutOfMemoryError: NPU out of memory. Tried to allocate 2.37 GiB (NPU 1; 61.28 GiB total capacity; 59.46 GiB already allocated; 59.46 GiB current active; 992.16 MiB free; 59.54 GiB reserved in total by PyTorch).If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. [ERROR] 2026-08-09-00:19:41 (PID:98893, Device:0, RankID:-1) ERR99999 UNKNOWN applicaiton exception [ERROR] 2026-08-09-00:19:42 (PID:98894, Device:1, RankID:-1) ERR99999 UNKNOWN applicaiton exception W0809 00:19:45.258000 98724 site-packages/torch/distributed/elastic/multiprocessing/api.py:900] Sending process 98894 closing signal SIGTERM E0809 00:19:46.024000 98724 site-packages/torch/distributed/elastic/multiprocessing/api.py:874] failed (exitcode: 1) local_rank: 0 (pid: 98893) of binary: /usr/local/python3.12.13/bin/python3.12原因分析
1. shm 报错根因: 容器内 /dev/shm 仅 64MB(df -h /dev/shm 确认),且无 CAP_SYS_ADMIN 无法 remount。DataLoader num_workers=8 时,worker 子进程通过 /dev/shm 传递 tensor,64MB 瞬间写满 → bus error。PyTorch 的 SharedMemory 在 Linux 上硬编码用 /dev/shm,file_system sharing strategy 也无法绕开(验证过 _share_filename_cpu_() 仍写 /dev/shm)。 2. 解决 shm: 将 yaml 中 num_workers: 8 改为 num_workers: 0,数据加载在主进程内完成,完全不使用共享内存。已验证 num_workers=0 时 /dev/shm 文件数不增加。 3. 次生 OOM: shm 解决后,暴露出 27B 模型在 2×64GB NPU 上 FSDP 反向传播时 HBM 显存不足(每卡 ~59.5GB 已分配,再申请 2.37GB 失败)。开启 enable_activation_offload: true 仅省 ~0.2GB 不足。 4. 解决 OOM: 在 fsdp_plan 下开启 cpu_offload: true,把参数/梯度/优化器状态卸载到 CPU(本机 2TB 内存充足)。HBM 占用从 ~59GB 降到 ~26GB,OOM 消除。修改qwen3_6_27B_config.yaml文件修复报错**
num_workers: 8 → num_workers: 0(消除 shm 报错) enable_activation_offload: false → true(省激活显存) fsdp_plan 新增 cpu_offload: true(参数/优化器状态 offload 到 CPU,消除 HBM OOM)测试微调过程
NPU使用情况

训练速度分析
微调速度优化
修改文件
qwen3_6_27B_config.yaml开启前向/反向预取,隐藏 cpu_offload 时 H2D 搬运延迟。
关闭激活卸载。recompute 已经省了激活显存,无需再 offload 到 CPU,省掉反向时 D2H/H2D 搬运。
这是提速的关键。cpu_offload 每步的 H2D 搬运是固定开销(搬整个 FSDP shard,与 batch 无关)。batch=1 时搬运开销被 2 个样本分摊,batch=8 时被 16 个样本分摊,计算密度大幅提升。
优化并且将训练iter恢复成1000后的最终qwen3_6_27B_config.yaml
# 并行策略 parallel: fully_shard_parallel_size: auto fsdp_plan: apply_modules: # 如果要开prefetch的话,请不要随意修改apply_modules的顺序 - model.visual - model.visual.blocks.{*} - model.language_model - model.language_model.embed_tokens - model.language_model.layers.{*} - lm_head param_dtype: bf16 reduce_dtype: fp32 # 27B 模型在 2×64GB NPU 上 FSDP 优化器状态单卡 108GB,必须 offload 到 CPU cpu_offload: true # 开启 prefetch 隐藏 H2D 搬运延迟(cpu_offload 时的主要优化手段) num_to_forward_prefetch: 1 num_to_backward_prefetch: 1 ulysses_parallel_size: 1 # 开启 ulysses-cp 时, 请将 model 的 attn_implementation 设置为 flash_attention_2 ### 数据相关配置 data: dataset_param: dataset_type: huggingface #数据集属性 attr: images: images messages: messages role_tag: role content_tag: content user_tag: user assistant_tag: assistant # 数据预处理 preprocess_parameters: model_name_or_path: &HF_MODEL_LOAD_PATH /workspace/shared_assets/models/Qwen/Qwen3.6-27B # 替换为原始hf权重 use_fast_tokenizer: true split_special_tokens: false image_max_pixels: 262144 image_min_pixels: 1024 video_max_pixels: 16384 video_min_pixels: 0 video_fps: 2.0 video_maxlen: 64 basic_parameters: cutoff_len: 1024 template: qwen3_vl_nothink enable_thinking: false train_on_prompt: false mask_history: false dataset_dir: /workspace/dataset/COCO2017/ dataset: &DATASET_PATH /workspace/dataset/COCO2017/mllm_format_llava_instruct_data.json cache_dir: ./cache_dir/ overwrite_cache: false preprocessing_batch_size: 1000 preprocessing_num_workers: 16 max_samples: null # 数据加载 dataloader_param: pin_memory: false shuffle: false dataloader_mode: sampler drop_last: true sampler_type: BaseRandomBatchSampler # num_workers 设为 0:容器内 /dev/shm 仅 64MB 且无法 remount, # DataLoader worker(num_workers>0) 会通过 /dev/shm 传递 tensor, # 64MB 不够会触发 "bus error / No space left on device"。 # num_workers=0 时数据加载在主进程内完成,不使用共享内存,彻底规避该问题。 num_workers: 0 collate_param: model_name: qwen3vl ignore_pad_token_for_loss: true enable_preload: false # 模型配置 model: model_id: qwen3_5 model_name_or_path: *HF_MODEL_LOAD_PATH trust_remote_code: true attn_implementation: flash_attention_2 freeze: - model.visual # 融合算子配置 gdn_implementation: triton causal_conv1d_implementation: eager # 优化特性配置 features: # loss 配置 loss_cfg: loss_type: default # If you want raw loss in model, loss_type can be set to "raw". router_aux_loss_coef: 0.0 # 重计算配置:开启 recompute 省激活显存(关 cpu_offload 后显存紧张), # 反向重算比 H2D 搬运快得多,是关 offload 后的合理折中 recompute: true recompute_plan: apply_modules: - model.visual.blocks.{*} - model.language_model.layers.{*} # chunkloss配置 enable_chunk_loss: true chunkloss_plan: apply_module: lm_head chunk_size: 1024 # activation offload 配置:关闭,recompute 已省激活,无需再 offload 到 CPU enable_activation_offload: false activation_offload_plan: apply_modules: - model.visual.blocks.{*} - model.language_model.layers.{*} # 训练配置 training: micro_batch_size: 8 gradient_accumulation_steps: 1 seed: 42 lr: 1.0e-5 lr_decay_style: cosine lr_warmup_ratio: 0.1 weight_decay: 0 train_iters: 1000 clip_grad: 0.0 init_model_with_meta_device: true optimizer: adamw adam_fused: true save_interval: 1000 no_load_optim: true # Do not load optimizer state; remove if loading is needed. no_load_rng: true # Do not load RNG state; remove if loading is needed. no_save_optim: true # Do not save optimizer state; remove if saving is needed. no_save_rng: true # Do not save RNG state; remove if saving is needed. load: /workspace/models/Qwen3.6-27B-dcp # 替换为转换后的dcp权重 save: /workspace/models/Qwen3.6-27B-save use_deter_comp: false plugin: - mindspeed_mm/fsdp/models/qwen3_5 - mindspeed_mm/fsdp/data/datasets/huggingface # 工具配置 tools: profile: enable: false profile_type: static ranks: [0] static_param: level: level1 with_stack: false with_memory: false record_shapes: false with_cpu: true save_path: ./profiling start_step: 10 end_step: 11 data_simplification: false aic_metrics_type: PipeUtilization memory_profile: enable: false start_step: 1 end_step: 2 save_path: ./memory_snapshot dump_ranks: [0] stacks: all max_entries: null mem_info: false优化后的NPU情况

训练完成截图

权重转换dcp_to_hf
mm-convert Qwen35Converter dcp_to_hf \ --save_hf_dir /workspace/models/Qwen3.6-27B-hf \ --dcp_dir /workspace/models/Qwen3.6-27B-save/iter_0001000 \ --origin_hf_dir /workspace/shared_assets/models/Qwen/Qwen3.6-27B \ --to_bf16 false建议
数据集相关
标签数量和图片数量不一致的问题。所以建议还是建议做一个说明如果是自定义数据集并且是openai格式的,可以参考这一部分,如果使用默认COCO2017数据集的话这部分可以跳过。
最初在尝试体验
Qwen3.6-27B模型过程中,在阅读qwen3.6的文档的时候,他有一部分openai数据集的处理说明,这一部分建议标注一下如果默认COCO2017数据集可以跳过这部分设置,有些初学者或者新手会以为COCO2017的数据集也是openai格式的,然后去做了这部分处理,然后就会出现微调的时候在安装mindspore-mm的时候他会自动安装torch-npu,但是他安装的版本是
torch-npu==2.7.1,这个包在有些环境安装的时候会遇到找不到torch-npu==2.7.1这个包,是不是可以考虑使用torch-npu==2..7.1post6这个替代,目前我体验下来torch-npu==2..7.1post6这个版本的包在大部分环境下都能安装成功。qwen3.5文档中的
pip list | grep fla_npu这个命令是不是错了,我这边体验的时候安装好是fla-npu这个包。根据这个文档
https://atomgit.com/Ascend/MindSpeed-MM/blob/master/examples/qwen3_5/README.md安装MindSpeedMM的时候,有些环境会遇到下面报错:这个报错我
pip list看了下MindSpeed-MM这个插件是安装成功的,继续安装文档的下一步Triton-Ascend是不会有问题的,但是担心开发者会搞不清以为出现了报错MindSpeed-MM并没有安装成功。所以感觉是不是可以考虑把triton-ascend==3.2.0安装的部分提到这个之前安装。这样感觉可能可以规避一下这个报错。