Pull Request已成功合入, 合并人@ascend-robot
(感谢 LKONE 的贡献)变更摘要
本 PR 主要调整了 LoRA 的保存模式:在默认的全量 checkpoint 保存流程中支持将 LoRA 权重(lora_A / lora_B)合并回基础权重后再导出,并新增 lora_save_only 开关,仅在显式开启时走"只导出 LoRA adapter"的路径;同时将 LoRA-only 保存改为写入按迭代号组织的 checkpoint 子目录,便于 PEFT/vLLM/SGLang 直接加载。
主要改动
-
LoRA 权重合并导出:在
hf_utils.py中新增merge_lora_weights函数,遍历.lora_A.default.weight/.lora_B.default.weight键,通过delta = matmul(lora_b, lora_a)计算增量并加到对应基础权重(base + scaling * delta),随后从状态字典中删除 LoRA 键;遇到DTensor时会先redistribute为Replicate再合并。 -
Checkpointer 参数与预处理扩展:
HuggingFaceCheckpointer新增lora_alpha、lora_rank参数,enable_lora时以lora_alpha / lora_rank作为缩放系数调用merge_lora_weights,并在保存前将浮点权重统一转换为save_ckpt_dtype。 -
新增
lora_save_only配置项:LoraArguments新增lora_save_only(默认False)字段,用于控制是否仅导出 adapter 的 safetensors 及其配置。 -
保存逻辑分流:
train_engine.py的 checkpoint 保存逻辑改为仅当args.training.lora.enable且args.training.lora.lora_save_only时调用save_lora_only,随后执行torch.distributed.barrier()并提前返回;否则走默认全量模型保存流程,并将lora_alpha、lora_rank传入 checkpointer。 -
LoRA-only 保存目录调整:
LoraWeightManager.save_lora_only改用get_checkpoint_name(save_path, iteration, release=False)生成带迭代号的输出目录,lora_adapter.safetensors与adapter_config.json均写入该目录而非原save_path下。


Pull Request 已合并或已关闭。
If you want to solve this problem, you can click here to do it in the FAQs.


Pull Request 已合并或已关闭。
If you want to solve this problem, you can click here to do it in the FAQs.


Pull Request 已合并或已关闭。
If you want to solve this problem, you can click here to do it in the FAQs.


What this PR does / why we need it?
Does this PR introduce any user-facing change?
无
How was this patch tested?
测试lora训练场景下,单独保存lora权重,导出为HF权重和DCP权重的正确性