| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
chore(fsdp2): develop longcat-flash-lite model in fsdp2 Co-authored-by: guihaowen666<guihaowen@huawei.com> # message auto-generated for no-merge-commit merge: !4344 merge br_master_longcat_flash_lite_fsdp2 into master chore(fsdp2): develop longcat-flash-lite model in fsdp2 Created-by: guihaowen666 Commit-by: guihaowen666 Merged-by: ascend-robot Description: ## What this PR does / why we need it? develop longcat-flash-lite model in fsdp2 ## Does this PR introduce any user-facing change? new model development, no user-facing change ## How was this patch tested? Run the inference task and check whether the model can perform normal dialogs. See merge request: Ascend/MindSpeed-LLM!4344 | 3 个月前 | |
chore(fsdp2): develop longcat-flash-lite model in fsdp2 Co-authored-by: guihaowen666<guihaowen@huawei.com> # message auto-generated for no-merge-commit merge: !4344 merge br_master_longcat_flash_lite_fsdp2 into master chore(fsdp2): develop longcat-flash-lite model in fsdp2 Created-by: guihaowen666 Commit-by: guihaowen666 Merged-by: ascend-robot Description: ## What this PR does / why we need it? develop longcat-flash-lite model in fsdp2 ## Does this PR introduce any user-facing change? new model development, no user-facing change ## How was this patch tested? Run the inference task and check whether the model can perform normal dialogs. See merge request: Ascend/MindSpeed-LLM!4344 | 3 个月前 | |
feat(pytorch): refactor fsdp2 ckpt converter Co-authored-by: HanhuiChen<chenhanhui1@h-partners.com> # message auto-generated for no-merge-commit merge: !4568 merge bugfix into master feat(pytorch): refactor fsdp2 ckpt converter Created-by: HANHU1CHEN Commit-by: HanhuiChen Merged-by: ascend-robot Description: ## What this PR does / why we need it? This PR adds a weight conversion utility for the MiniMax-M2.7 model that converts its FP8-quantized HuggingFace checkpoint into BF16 format, so the weights can be loaded and used on Ascend NPU. ## Does this PR introduce any user-facing change? No change to existing functionality. ## How was this patch tested? Verified by running the script on a MiniMax-M2.7 FP8 checkpoint and confirming the output directory loads correctly and produces the expected BF16 weights with the fused MoE layout. See merge request: Ascend/MindSpeed-LLM!4568 | 3 个月前 | |
feat(pytorch): add glm52 in fsdp2 Co-authored-by: guozhihua2<guozhihua2@huawei.com> # message auto-generated for no-merge-commit merge: !4614 merge add_glm52_in_fsdp2 into master feat(pytorch): add glm52 in fsdp2 Created-by: guozhihua2 Commit-by: guozhihua2 Merged-by: ascend-robot Description: ## What this PR does / why we need it? This PR adds GLM52 model support in the FSDP2 training framework. Main changes include: 1. Add the GLM52 model implementation for FSDP2, including model definition, configuration adaptation, and registration logic. 2. Support GLM52 pretraining under the FSDP2 framework, enabling users to launch GLM52 training with the FSDP2 training entry and related scripts. 3. Support GLM52 chat/inference flow, so the adapted GLM52 model can be used for basic generation and chat validation after loading. 4. Adapt GLM52-specific model logic in FSDP2, including attention/indexer-related behavior and model forward compatibility required by GLM52. 5. Provide related ST coverage to ensure the GLM52 FSDP2 model path can be built, loaded, and executed correctly. ## Does this PR introduce any user-facing change? Yes. Users can now use the GLM52 model in the FSDP2 framework, including GLM52 pretraining and chat/inference workflows. Existing model usage is not expected to be affected. ## How was this patch tested? Verified by running the GLM52 ST test. pipeline/st/glm52/pretrain_glm52_38b_4k_fsdp2_A3.sh The test covers the GLM52 FSDP2 model build and execution path, including pretraining-related runtime validation and basic chat/inference functionality. See merge request: Ascend/MindSpeed-LLM!4614 | 3 个月前 |